Topic Identification from Spoken TED-Talks

Táto práca sa zaoberá problémom spracovania prirodzeného jazyka a následnej klasifikácie. Použité systémy boli modelované na TED-LIUM korpuse. Systém automatického spracovania jazyka bol modelovaný s použitím sady nástrojov Kaldi. Vo výsledku bol dosiahnutý WER s hodnotou 16.6%. Problém klasifikácie textu bol adresovaný s pomocou metód na lineárnu klasifikáciu, konkrétne Multinomial Naive Bayes a Linear Support Vector Machines, kde druhá technika dosiahla vyššiu presnosť klasifikácie.
This thesis deals with the problems of language recognition and topic classification, using TED-LIUM corpus to train both the ASR and classification models. The ASR system is built using the Kaldi toolkit, achieving the WER of 16.6%. The classification problem is addressed using linear classification methods, specifically Multinomial Naive Bayes and Linear Support Vector Machines, the latter method achieving higher topic classification accuracy.

Keywords

TED , talks , identifikácia tém , strojové učenie , klasifikácia , transkripcia , lineárna klasifikácia , Kaldi , support vector machines , akustický model , lingvistický model , TED-LIUM , ASR , TED , talks , topic identification , machine learning , classification , transcription , linear classification , Kaldi , support vector machines , acoustic modeling , language modeling , TED-LIUM , ASR

Citation

VAŠŠ, A. Topic Identification from Spoken TED-Talks [online]. Brno: Vysoké učení technické v Brně. Fakulta informačních technologií. 2019.

Language of document

en

Study field

Informační technologie

Comittee

doc. Ing. Richard Růžička, Ph.D., MBA (předseda) doc. Ing. Ondřej Ryšavý, Ph.D. (místopředseda) Ing. Jaroslav Dytrych, Ph.D. (člen) Ing. Bohuslav Křena, Ph.D. (člen) doc. Ing. Michal Španěl, Ph.D. (člen)

Date of acceptance

2019-08-29

Defence

Student nejprve prezentoval výsledky, kterých dosáhl v rámci své práce. Komise se poté seznámila s hodnocením vedoucího a posudkem oponenta práce. Student následně odpověděl na otázky oponenta a na další otázky přítomných. Komise se na základě posudku oponenta, hodnocení vedoucího, přednesené prezentace a odpovědí studenta na položené otázky rozhodla práci hodnotit stupněm " C ". Otázky u obhajoby: * How to describe in a few sentences the main components of the ASR system? * How to analyze the results of the topic identification system?Is there any comparable results already published on similar corpus? * Why the results from the ASR-TID system are sometimes better than the text based TID system.

Result of defence

práce byla úspěšně obhájena

URI

http://hdl.handle.net/11012/187233

Collections

2019

Citace PRO

Full item page

Topic Identification from Spoken TED-Talks

Files

Date

Authors

Advisor

Referee

Mark

Journal Title

Journal ISSN

Volume Title

Publisher

ORCID

Abstract

Description

Keywords

Citation

Document type

Document version

Date of access to the full text

Language of document

Study field

Comittee

Date of acceptance

Defence

Result of defence

DOI

URI

Collections

Endorsement

Review

Supplemented By

Referenced By

Citace PRO