Article published In: International Journal of Corpus Linguistics
Vol. 5:2 (2000) ► pp.179–197
The Process of Designing a Multidisciplinary Monolingual Sample Corpus
Published online: 30 May 2001
https://doi.org/10.1075/ijcl.5.2.05das
https://doi.org/10.1075/ijcl.5.2.05das
This paper discusses the approach of developing a sample of printed corpus in Bangla, one of the national languages of India and the only national language of Bangladesh. It is designed from the data collected from various published documents. The paper highlights different issues related to corpus generation, data-file preparation, language analysis, and processing as well as application potentials to different areas of pure and applied linguistics. It also includes statistical studies on the corpus along with some interpretation of the results. The difficulties that one may face during corpus generation are also pointed out.
Keywords: dictionary, word forms, concordance, NLP, machine translation, graphic symbol, corpus, diacritic, data-file
Cited by (5)
Cited by five other publications
Wynne, Hilary S.Z., Beinan Zhou, Sandra Kotzor & Aditi Lahiri
Dash, Niladri Sekhar
Pal, Alok Ranjan, Diganta Saha, Sudip Kumar Naskar & Niladri Sekhar Dash
Parameswarappa, S., V. N. Narayana & G. N. Bharathi
This list is based on CrossRef data as of 12 december 2025. Please note that it may not be complete. Sources presented here have been supplied by the respective publishers. Any errors therein should be reported to them.
