Qaamuuska-NLP
Share this post:

Ogaal Labs is pleased to introduce the research preview of Qaamuuska-NLP, a structured lexical resource developed from Qaamuuska Af-Soomaaliga, the 2012 monolingual Somali dictionary edited by Annarita Puglielli and Cabdalla Cumar Mansuur.
Qaamuuska-NLP converts the dictionary into 46,314 machine-readable records designed to support Somali natural language processing, computational lexicography and language-technology research.
Unlike OCR-based digitisation projects, Qaamuuska-NLP was extracted directly from the PDF’s native text layer. The project uses a transparent, deterministic and rule-based pipeline without machine-learning or generative-AI extraction.
What the resource contains
The structured resource preserves information explicitly recorded in the source dictionary, including:
Headwords and homonym indices
Parts of speech
Noun gender and plural information
Verb conjugation classes and transitivity
Numbered definitions
Synonym references
Dictionary cross-references
Subject-domain labels
Original extracted entry text
Resource highlights
46,314 dictionary records
34,726 nouns
11,445 verbs
32,801 extracted definitions
21,032 redirect records
20,075 successfully resolved redirects
11,854 synonym links
9,212 homonym-marked records
16 specialised subject domains
The project retains the dictionary’s internal structure instead of reducing it to a simple word list. This makes the resource useful for lexical lookup, morphological research and future Somali language technologies.
Research findings
The accompanying study also examines whether grammatical labels can be predicted from Somali word forms.
Using only character suffixes, a simple classifier achieved:
93.5% accuracy for verb conjugation class
86.4% accuracy for noun gender
86.9% accuracy for coarse part of speech
These results show that Somali word-final patterns contain strong signals for several grammatical categories represented in the dictionary.
Potential applications
Qaamuuska-NLP may support future work in:
Somali morphological analysis
Lemmatization and part-of-speech processing
Search and information retrieval
Spell-checking and writing assistance
Digital dictionary applications
Educational language tools
Lexical-semantic research
Training and evaluation of Somali NLP systems
Lexical-standard exports such as TEI Lex-0 and OntoLex-Lemon
Responsible release
Qaamuuska-NLP is currently being released as a research preview.
The original extraction code, schema, aggregate statistics and validation methodology are being prepared for public release under appropriate software terms. Public redistribution of the complete structured lexical dataset remains subject to written rightsholder permission.
The software license does not apply to dictionary definitions or other content derived from the original publication.
Research team
Abdullahi Mohamud Ahmed¹, Abdulkadir Ugas¹˒², Abdirashid Omar¹ and Abdulhakim Ismail¹
¹ Ogaal Labs
² Department of Bigdata, Chungbuk National University
Acknowledgement:
Qaamuuska-NLP is derived from:
Annarita Puglielli and Cabdalla Cumar Mansuur, editors. Qaamuuska Af-Soomaaliga. Roma Tre Press, 2012. ISBN 978-88-97524-02-1. DOI: 10.13134/978-88-97524-02-1.
Interactive research dashboard: qaamuuska-nlp.ogaaldata.com