Publication details
This post summarizes the peer-reviewed work cited below.
- Paper
- Türkçe Metinlerin Sınıflandırılmasında Öznitelik Olarak Eklerin de Kullanılması
- Authors
- Ahmet Erkan Çelik
- Presented at
- International Symposium on Advanced Engineering Technologies (ISADET 2), Kahramanmaraş, 16–18 June 2022 — ISADET 2 Abstracts and Proceedings Book
- Pages
- 422–427
- Presentation date
- June 17, 2022
- Full text (PDF)
Today, documents collected at scale from many different sources are analyzed with machine-learning approaches, and these methods are widely used for text classification. In natural language processing pipelines, stemming is commonly applied during pre-processing: words are stripped of their suffixes in order to increase the value of the resulting features. In agglutinative languages such as Turkish, however, this step is not always efficient.
This study examines the effect on classification performance of evaluating words together with their suffixes. The experiments used customer complaint and request data from the Next4biz CSM product. When the data were analyzed with deep-learning algorithms, keeping stems and suffixes together was observed to contribute to classification accuracy, while processing with the Snowball stemmer produced a number of misclassifications. The proposed model was then integrated into the Next4biz CSM product as a cloud-based service and is now in use by several Next4biz customers.
The results indicate that evaluating stems and suffixes together can yield better text-classification outcomes for agglutinative languages.
For more detail, you can read the full paper (in Turkish): Türkçe Metinlerin Sınıflandırılmasında Öznitelik Olarak Eklerin de Kullanılması (PDF)