The intelligent voice ASR system for the iberspeech 2018 speech to text transcription challenge

Nazim Dugan; Cornelius Glackin; Gérard Chollet; Nigel Cannings

doi:10.21437/IberSPEECH.2018-57

Communication Dans Un Congrès Année : 2018

The intelligent voice ASR system for the iberspeech 2018 speech to text transcription challenge

, , (1, 2, 3) ,

1
2
3

Nazim Dugan

Fonction : Auteur

Cornelius Glackin

Fonction : Auteur

Gérard Chollet

Fonction : Auteur
PersonId : 176991
IdHAL : gerard-chollet
ORCID : 0000-0003-4245-146X
IdRef : 078020824

Institut Polytechnique de Paris

Département Electronique et Physique

ARMEDIA

Nigel Cannings

Fonction : Auteur

Résumé

This paper describes the system developed by the Empathic team for the open set condition of the Iberspeech 2018 Speech to Text Transcription Challenge. A DNN-HMM hybrid acoustic model is developed, with MFCC's and iVectors as input features, using the Kaldi framework. The provided ground truth transcriptions for training and development are cleaned up using customized clean-up scripts and then realigned using a two-step alignment procedure which uses word lattice results coming from a previous ASR system. 261 hours of data is selected from train and dev1 subsections of the provided data, by applying a selection criterion on the utterance level scoring results. The selected data is merged with the 91 hours of training data used to train the previous ASR system with a factor 3 times data augmentation by reverberation using a noise corpus on the total training data, resulting a total of 1057 hours of final …

Mots clés

ASR system Transcription Script

Domaines

Informatique [cs] Réseau de neurones [cs.NE] Traitement du signal et de l'image [eess.SP]

TelecomParis HAL : Connectez-vous pour contacter le contributeur

https://telecom-paris.hal.science/hal-02288554

Soumis le : samedi 14 septembre 2019-18:56:33

Dernière modification le : jeudi 21 décembre 2023-11:34:41

Dates et versions

hal-02288554 , version 1 (14-09-2019)

Identifiants

HAL Id : hal-02288554 , version 1
DOI : 10.21437/IberSPEECH.2018-57

Citer

Nazim Dugan, Cornelius Glackin, Gérard Chollet, Nigel Cannings. The intelligent voice ASR system for the iberspeech 2018 speech to text transcription challenge. IberSPEECH 2018, Nov 2018, Barcelone, Spain. pp.272-276, ⟨10.21437/IberSPEECH.2018-57⟩. ⟨hal-02288554⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

INSTITUT-TELECOM TELECOM-SUDPARIS PARISTECH UNIV-PARIS-SACLAY

55 Consultations

0 Téléchargements

The intelligent voice ASR system for the iberspeech 2018 speech to text transcription challenge

Résumé

Mots clés

Domaines

Dates et versions

Identifiants

Citer

Exporter

Collections

Altmetric

Partager