Abstract
Annotatsiya. Maqolada kod almashinuvli (bir gapda ikki tildan foydalanilgan) matnlarda nomlangan obyektlarni aniqlash (NER) masalasi uchun belgilangan ma'lumotlar tanqisligini bartaraf etishga qaratilgan uchta ma'lumotlarni ko'paytirish usuli ko'rib chiqiladi: so'z embeddinglari asosida almashtirish (klassik FastText va kontekstli BERT, KERMIT modellari), ikki tilni qo'llab-quvvatlashga moslashtirilgan EDA usuli hamda oraliq tillar orqali teskari tarjima. Usullar arab-ingliz kod almashinuvli korpusda BiLSTM-CRF arxitekturali NER modeli yordamida ekstrinsik baholangan. Eng yaxshi natija teskari tarjima va embedding-analogiyalar asosidagi obyekt almashtirishning ketma-ket kombinatsiyasida erishilib, F-ball 77,69 foizdan 79,20 foizga (+1,51 foiz punkt) oshgan; o'qitish to'plami 5306 gapdan 10 612 gapga kengaygan. Natijalar sintetik namunalar sifati ularning miqdoridan muhimroq ekanini ko'rsatadi: kamroq obyekt qo'shgan, biroq semantik jihatdan to'g'ri usul eng yuqori samarani bergan.
References
1. Sabty C., Omar I., Wasfalla F., Islam M., Abdennadher S. Data augmentation techniques on Arabic data for named entity recognition // Procedia Computer Science. – 2021. – Vol. 189. – P. 292–299. DOI: 10.1016/j.procs.2021.05.092.
2. Sabty C., Sherif A., Elmahdy M., Abdennadher S. Techniques for named entity recognition on Arabic-English code-mixed data // International Journal of Computational Linguistics and Applications. – 2019. – Vol. 1, No. 1. – P. 44–63.
3. Ratner A.J., Ehrenberg H., Hussain Z., Dunnmon J., Ré C. Learning to compose domain-specific transformations for data augmentation // Advances in Neural Information Processing Systems 30. – Long Beach, 2017. – P. 3236–3246.
4. Wei J., Zou K. EDA: Easy data augmentation techniques for boosting performance on text classification tasks // arXiv:1901.11196. – 2019.
5. Kobayashi S. Contextual augmentation: Data augmentation by words with paradigmatic relations // arXiv:1805.06201. – 2018.
6. Xie Q., Dai Z., Hovy E., Luong M.-T., Le Q.V. Unsupervised data augmentation for consistency training // arXiv:1904.12848. – 2019.
7. Anaby-Tavor A., Carmeli B., Goldbraich E. et al. Do not have enough data? Deep learning to the rescue // Proc. of the 34th AAAI Conference on Artificial Intelligence. – New York, 2020. – P. 7383–7390.
8. Kumar V., Choudhary A., Cho E. Data augmentation using pre-trained transformer models // arXiv:2003.02245. – 2020.
9. Mathew J., Fakhraei S., Ambite J.L. Biomedical named entity recognition via reference-set augmented bootstrapping // arXiv:1906.00282. – 2019.
10. Sabty C., Elmahdy M., Abdennadher S. Named entity recognition on Arabic-English code-mixed data // Proc. of the 13th IEEE International Conference on Semantic Computing (ICSC). – 2019. – P. 93–97.
11. Sabty C., Islam M., Abdennadher S. Contextual embeddings for Arabic-English code-switched data // Proc. of the Fifth Arabic Natural Language Processing Workshop. – 2020. – P. 215–225.
12. Bojanowski P., Grave E., Joulin A., Mikolov T. Enriching word vectors with subword information // Transactions of the Association for Computational Linguistics. – 2017. – Vol. 5. – P. 135–146.
13. Miller G.A. WordNet: A lexical database for English // Communications of the ACM. – 1995. – Vol. 38, No. 11. – P. 39–41.
14. ElKateb S., Black W., Rodríguez H. et al. Building a WordNet for Arabic // Proc. of the International Conference on Language Resources and Evaluation. – 2006. – P. 29–34.
15. Finkel J.R., Grenager T., Manning C.D. Incorporating non-local information into information extraction systems by Gibbs sampling // Proc. of the 43rd Annual Meeting of the ACL. – 2005. – P. 363–370.
16. Benajiba Y., Rosso P., Benedíruiz J. ANERsys: An Arabic named entity recognition system based on maximum entropy // Computational Linguistics and Intelligent Text Processing. – 2007. – P. 143–153.
17. Vaswani A., Bengio S., Brevdo E. et al. Tensor2Tensor for neural machine translation // arXiv:1803.07416. – 2018.
18. Ziemski M., Junczys-Dowmunt M., Pouliquen B. The United Nations parallel corpus v1.0 // Proc. of the Tenth International Conference on Language Resources and Evaluation. – 2016. – P. 3530–3534.
19. Hamed I., Elmahdy M., Abdennadher S. Collection and analysis of code-switch Egyptian Arabic-English speech corpus // Proc. of the Eleventh International Conference on Language Resources and Evaluation. – 2018.