The Role of Suffixes in Turkish Text Classification and NLP Applications

As the Next4biz R&D team, we presented our study Türkçe Metinlerin Sınıflandırılmasında Öznitelik Olarak Eklerin de Kullanılması (Using Suffixes as Features in Turkish Text Classification) at the International Symposium on Advanced Engineering Technologies (ISADET), held in Kahramanmaraş on 16–18 June 2022. The study examines how the distinctive structure of Turkish can be used more effectively in natural language processing (NLP) models.

Open the PDF in a new tab

Publication details

This post summarizes the peer-reviewed work cited below.

Paper
Türkçe Metinlerin Sınıflandırılmasında Öznitelik Olarak Eklerin de Kullanılması
Authors
Ahmet Erkan Çelik
Presented at
International Symposium on Advanced Engineering Technologies (ISADET 2), Kahramanmaraş, 16–18 June 2022 — ISADET 2 Abstracts and Proceedings Book
Pages
422–427
Presentation date
June 17, 2022
PDF
Full text (PDF)

All our academic publications →

Today, documents collected at scale from many different sources are analyzed with machine-learning approaches, and these methods are widely used for text classification. In natural language processing pipelines, stemming is commonly applied during pre-processing: words are stripped of their suffixes in order to increase the value of the resulting features. In agglutinative languages such as Turkish, however, this step is not always efficient.

This study examines the effect on classification performance of evaluating words together with their suffixes. The experiments used customer complaint and request data from the Next4biz CSM product. When the data were analyzed with deep-learning algorithms, keeping stems and suffixes together was observed to contribute to classification accuracy, while processing with the Snowball stemmer produced a number of misclassifications. The proposed model was then integrated into the Next4biz CSM product as a cloud-based service and is now in use by several Next4biz customers.

The results indicate that evaluating stems and suffixes together can yield better text-classification outcomes for agglutinative languages.

For more detail, you can read the full paper (in Turkish): Türkçe Metinlerin Sınıflandırılmasında Öznitelik Olarak Eklerin de Kullanılması (PDF)

Prof. Dr. Akhan Akbulut
Prof. Dr. Akhan Akbulut
Professor Doctor Akhan Akbulut worked in the Computer Engineering departments of Istanbul Kültür University and NC State University. He serves institutions such as TÜBİTAK, Ministry of Industry and Technology, TÜSEB, and KOSGEB. He researches Distributed Systems and Artificial Intelligence and has over 100 international journal articles and conference proceedings from his work.