A dataset of kazakh handwritten characters
Аннотация
Handwriting still appears in classrooms, paperwork, and personal records, even where most communication has moved to a screen. For Kazakh written in Cyrillic, the main blocker for character-level recognition research is not a lack of ideas but a lack of carefully described data that other groups can reuse with confidence. This paper treats data quality as the central engineering problem and walks through how a usable dataset was put together. We describe form design and writer recruitment, scanning at 300 dpi, semi-automatic segmentation of grid forms into character crops, normalization and resizing to 64×64, and a two-pass annotation workflow with class-wise spot checks. The resulting dataset holds approximately 140,000 grayscale images across 42 classes from around 110 writers in two age groups. We also report a baseline convolutional neural network that reaches 84.3% test accuracy and a macro-averaged F1 of 0.836, mainly to confirm the dataset is usable and to give later work a comparable reference point. The contribution is the methodology more than the size: per-class auditing, structured handling of segmentation failures, and a strict ambiguity-removal rule during labeling did more for reliability than additional samples would have.
Автор
Abdurakhim Bakytzhan
Dinara Kaibassova
DOI
https://doi.org/10.48081/BGQF2031
Ключевые слова
handwritten character dataset
Kazakh Cyrillic
data quality
segmentation
annotation workflow
label noise
convolutional neural network
low-resource scripts
Год
2026
Номер
Выпуск 2
Для цитирования:
Abdurakhim Bakytzhan, Dinara Kaibassova A dataset of kazakh handwritten characters // Вестник Торайгыров университета Серия: физика, математика и компьютерные науки - 2026 - №2 - Р.30-41. https://doi.org/10.48081/BGQF2031
Скопировано!