Empirical evaluation of compounds indexing for Turkish texts

Chedi Bechikh Ali; Hatem Haddad; Yahya Slimani

doi:10.1016/j.csl.2019.01.004

Back

Empirical evaluation of compounds indexing for Turkish texts

Journal article

Peer reviewed

Empirical evaluation of compounds indexing for Turkish texts

Chedi Bechikh Ali, Hatem Haddad and Yahya Slimani

Computer speech & language, Vol.56, pp.95-106

07/2019

DOI: https://doi.org/10.1016/j.csl.2019.01.004

Abstract

Compound

Indexing

Information retrieval

Syntactic patterns

Turkish language

•Explore whether compounds can help improve Turkish IR systems.•Define the syntactic patterns used to index compounds for the Turkish Language.•Compare the impact of different compounds types on the retrieval performances.•Study the impact of state-of-the-art model on the results of the retrieval with compounds•Show that using compounds as index can improve retrieval performances. In this article, we describe an empirical evaluation of compounds indexing for Turkish texts. We dive beyond the keyword indexing to propose a framework for Turkish compounds extraction and indexing. We identify twelve Turkish compounds pattern types that we classify in six categories. To extract Turkish compounds, we rely on a light natural language processing approach based on syntactic pattern recognition. We compare different compounds indexing strategies. We also investigate the effectiveness of using one compounds type and the effectiveness of combining different compound types. We conduct experiments over the Milliyet test dataset. The results of our experiments show that using compounds as index terms can improve retrieval performances. However, not all the compound types have a positive impact on the retrieval process.

Metrics

1 Record Views

Details

Title: Empirical evaluation of compounds indexing for Turkish texts
Creators - without role: Chedi Bechikh Ali - LISI, INSAT, Universié de Carthage, Centre Urbain Nord BP 676, Tunis 1080, Tunisia
Hatem Haddad - Université Libre de Bruxelles
Yahya Slimani - Manouba University
Publication Details: Computer speech & language, Vol.56, pp.95-106
Publisher: Elsevier Ltd
Identifiers: 9914846408331
Academic Unit: Imam Abdulrahman Bin Faisal University
Language: English
Resource Type: Journal article