Konferensartikel

Part-of-speech tagging of Swedish texts in the neural era

Yvonne Adesam

Aleksandrs Berdicevskis

Ladda ner artikel

Ingår i: Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa), May 31-June 2, 2021.

Linköping Electronic Conference Proceedings 178:20, s. 200-209

NEALT Proceedings Series 45:20, p. 200-209

Visa mer +

Publicerad: 2021-05-21

ISBN: 978-91-7929-614-8

ISSN: 1650-3686 (tryckt), 1650-3740 (online)

Abstract

We train and test five open-source taggers, which use different methods, on three Swedish corpora, which are of comparable size but use different tagsets. The KB-Bert tagger achieves the highest accuracy for part-of-speech and morphological tagging, while being fast enough for practical use. We also compare the performance across tagsets and across different genres in one of the corpora. We perform manual error analysis and perform a statistical analysis of factors which affect how difficult specific tags are. Finally, we test ensemble methods, showing that a small (but not significant) improvement over the best-performing tagger can be achieved.

Nyckelord

tagging, machine learning, evaluation, morphology, part of speech error analysis, annotation, corpus, ensemble

Referenser

Inga referenser tillgängliga

Citeringar i Crossref