A Fully Semantic Approach to Large Scale Text Categorization.

DESSI, NICOLETTA;DESSI', STEFANIA;PES, BARBARA
2013-01-01

Abstract

Text categorization is usually performed by supervised algorithms on the large amount of hand-labelled documents which are labor-intensive and often not available. To avoid this drawback, this paper proposes a text categorization approach that is designed to fully exploiting semantic resources. It employs the ontological knowledge not only as lexical support for disambiguating terms and deriving their sense inventory, but also to classify documents in topic categories. Specifically, our work relates to apply two corpus-based thesauri (i.e. WordNet and WordNet Domains) for selecting the correct sense of words in a document while utilizing domain names for classification purposes. Experiments presented show how our approach performs well in classifying a large corpus of documents. A key part of the paper is the discussion of important aspects related to the use of surrounding words and different methods for word sense disambiguation
2013
Information Sciences and Systems 2013
Gelenbe, E; Lent, R
Erol Gelenbe
264
149
157
9
Springer International Publishing
BERLIN
978-3-319-01603-0
http://link.springer.com/chapter/10.1007%2F978-3-319-01604-7_15
Esperti anonimi
info:eu-repo/semantics/bookPart
2.1 Contributo in volume (Capitolo o Saggio)
Dessi, Nicoletta; Dessi', Stefania; Pes, Barbara
2 Contributo in Volume::2.1 Contributo in volume (Capitolo o Saggio)
3
268
none
Files in This Item:
There are no files associated with this item.

Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.

Questionnaire and social

Share on:
Impostazioni cookie