LinguistCore

About

Muhammad Shoaib Tahir

Applied linguist working in corpus and computational linguistics, teaching at university level in Pakistan and building language resources for Punjabi.

01 — What this site is

Everything in one place

LinguistCore gathers information and resources across linguistics, with particular depth in corpus linguistics and computational linguistics: books, free MCQ practice papers with answer keys, browser-based analysis tools, and teaching material for students and researchers.

It is built for people who work with language data — students preparing for examinations, teachers putting a course together, and researchers who want to run an analysis without installing anything.

Corpus linguistics

Corpus design, concordancing, frequency, collocation, keyness and the statistics behind them.

Computational linguistics

Text processing, tagging, morphology, WordNet and practical NLP for linguists.

Study resources

MCQ practice papers, analysis tools and teaching material, free to use.

02 — Background

A short profile

I hold an M.Phil. in Applied Linguistics from Government College University, Faisalabad, where my thesis developed a Punjabi synset using an expansion approach from English WordNet.

I currently teach Computational Linguistics as a visiting lecturer at COMSATS University Islamabad, Lahore Campus, and have taught Corpus Linguistics, Morphology, Syntax and Computer Assisted Language Learning at Government College University Faisalabad, the University of Education, and the University of Agriculture, Faisalabad.

Before that I worked as a research assistant on an HEC-NRPU project developing linguistic tools for Shahmukhi Punjabi — work that continues in the projects below.

At a glance

Field
Corpus & computational linguistics
Teaching
Computational and corpus linguistics, morphology, syntax
Languages of research
Punjabi (Shahmukhi), Urdu, English
Tools
Python, corpus software, web scraping
03 — Selected work

Projects and writing

The books →

Language resource projects

  • Punjabi POS tagger
  • Punjabi WordNet
  • Morphological analyser for Punjabi
  • Rule-based stemmer for Punjabi

Selected publications

  • Morphological Description of Nouns in Shahmukhi Punjabi: A Corpus Based Study (2023)
  • Developing Digital Resources of Shahmukhi Punjabi (2024)
  • Mapping of English and Punjabi Verb Classes: A Corpus Based Study (2024)

Teaching and talks

Workshops and invited talks on corpus tools, statistics in corpus linguistics and Python for linguists, delivered at universities across Pakistan including NUST, NED, FAST, UMT, NUML and the Linguistic Association of Pakistan.