Texts, corpora, and language models: natural language processing in practice 1500-SDN-TKMJCPJNWP
The primary objective of this course is to provide participants with a practical introduction to computational linguistics and natural language processing (NLP), enabling them to acquire the skills necessary to create and process text corpora with custom annotation layers. Corpus annotation involves enriching raw text data with additional linguistic information, such as identifying the part-of-speech (POS) for individual text segments. While existing corpus tools allow for the automatic annotation of morphological and syntactic data – and occasionally semantic and pragmatic levels – the ability to create custom annotation layers allows participants to overcome the inherent limitations of general-purpose tools and develop bespoke solutions tailored to their specific research projects.
The first module will introduce the basics of Python programming, including libraries essential for NLP and text processing (e.g., spaCy, Stanza, pandas, ollama), as well as the datasets used to build corpora and perform automated analysis and annotation (such as Korpusomat or SketchEngine).
The second module will focus on the methodology and practice of text annotation – the creation of proprietary data resources – and machine learning techniques (including supervised and unsupervised learning, deep learning, and various text classification tasks). Resources prepared by participants using the Label Studio application will be used to train custom large language models (LLMs) based on modern neural network architectures, which allow for relatively high efficiency even with small training datasets.
The third module will cover the fundamentals of corpus-based research, focusing on the creation of a personal corpus using the BlackLab search engine, incorporating the annotation layers developed in previous stages. Text corpora serve as a natural foundation for quantitative and qualitative research in linguistics, literary studies, and sociology. However, their utility extends beyond these fields; they are equally valuable for researchers in other disciplines, for instance, in constructing disciplinary corpora or conducting literature reviews. Using a properly prepared corpus, participants will be able to formulate queries, generate statistical reports, and create data visualizations relevant to their own research.
While the methods and examples presented will focus primarily on the Polish language, they are generally transferable to other languages (e.g., English, German, French, or Italian).
No prior programming experience is required; however, a strong commitment to acquiring these technical competencies is expected.
The course acknowledges and permits the extensive use of generative AI tools. An honor code will be established to precisely define the areas where the use of AI is encouraged and those where it is discouraged or prohibited, as it may conflict with the learning objectives. Specifically, a distinction will be made between “learning mode”, where AI serves as a personal assistant, and “verification mode”, where participants must demonstrate independent proficiency and mastery of the subject matter.
Course coordinators
Type of course
Learning outcomes
Knowledge (The student knows and understands)
WG_01 The theoretical and technological foundations of tools used for the processing and analysis of textual data.
WG_01 Fundamentals of Python programming and the libraries utilized for the automated processing and analysis of textual data.
WG_03 Research methodologies in computational linguistics and natural language processing (NLP) within the context of humanities and interdisciplinary research.
WK_01 Fundamental dilemmas of contemporary civilization regarding the development and application of language models, including their use in research practices and academic contexts.
Skills (The student is able to)
K_U01 Communicate on specialized topics within the field of natural language processing to a level that enables active participation in the international scientific community.
UW_01 Acquire and process textual data for qualitative and quantitative analyses using programming techniques.
UW_01 Perform annotation of text corpora, taking into account methodological requirements and their implications for the performance and effectiveness of trained machine learning models.
Social competences (The student is ready to)
KK_03 Recognize the significance of knowledge in natural language processing for solving cognitive and practical problems, and apply NLP methods to achieve their own research objectives.
Assessment criteria
Active participation in classes (maximum of two absences permitted).
A minimum score of 60% on tests and assignments completed on the Kampus platform.
Completion of an individual or group project involving data collection and annotation (creation of a micro-corpus), training an AI model, and developing a corpus with a custom annotation layer.
Bibliography
Alammar, J., Grootendorst, M. (2024) Hands-on Large Language Models. Language Understanding and Generation, Sebastopol, CA: O’Reilly Media.
Altinuk, D. (2021) Mastering spaCy: An end-to-end practical guide to implementing NLP applications using the Python ecosystem. Birmingham: Packt Publishing.
Does, J. de, Niestadt, J., Depuydt, K. (2017), Creating research environments with BlackLab, w: J. Odijk, A. van Hessen (eds.), CLARIN in the Low Countries, s. 151–165, London: Ubiquity Press.
Hobson, L., Cole, H., Hannes, H. (2021) Przetwarzanie języka naturalnego w akcji. Rozumienie, analiza i generowanie tekstu w Pythonie na przykładzie języka angielskiego. Warszawa: PWN.
Kinsley, H., Kukieła, D. (2020) Neural Networks from Scratch in Python. Sentdex, Kinsley Enterprises.
Pustejovsky, J., Stubbs, A. (2013) Natural Language Annotation for Machine Learning. Sebastopol, CA: O’Reilly Media.
Sweigart, A. (2020) Automatyzacja nudnych zadań z Pythonem. Nauka programowania. Gliwice: Helion.
Tkachenko, M., Malyuk, M., Holmanyuk, A., Liubimov, N. (2020–2026), Label Studio: Data labeling software. https://github.com/HumanSignal/label-studio