Sentence similarity based on semantic nets and corpus statistics

Li, Yuhua, McLean, David A., Bandar, Zuhair A., O'Shea, James and Crockett, Keeley (2006) Sentence similarity based on semantic nets and corpus statistics. IEEE transactions on knowledge and data engineering, 18 (8). pp. 1138-1150. ISSN 1558-2191

Preview

Available under License In Copyright.
Download (384kB) | Preview

Abstract

Sentence similarity measures play an increasingly important role in text-related research and applications in areas such as text mining, Web page retrieval, and dialogue systems. Existing methods for computing sentence similarity have been adopted from approaches used for long text documents. These methods process sentences in a very high-dimensional space and are consequently inefficient, require human input, and are not adaptable to some application domains. This paper focuses directly on computing the similarity between very short texts of sentence length. It presents an algorithm that takes account of semantic information and word order information implied in the sentences. The semantic similarity of two sentences is calculated using information from a structured lexical database and from corpus statistics. The use of a lexical database enables our method to model human common sense knowledge and the incorporation of corpus statistics allows our method to be adaptable to different domains. The proposed method can be used in a variety of applications that involve text knowledge representation and discovery. Experiments on two sets of selected sentence pairs demonstrate that the proposed method provides a similarity measure that shows a significant correlation to human intuition.

Item Type:	Article
Peer-reviewed:	No
Date Deposited:	24 Mar 2010 14:49
Publisher:	IEEE Press
Additional Information:	This article was originally published following peer-review in IEEE Transactions on Knowledge and Data Engineering, published by and copyright IEEE Press.
Divisions:	Faculties > Science and Engineering
URI:	https://e-space.mmu.ac.uk/id/eprint/94900
DOI:	https://doi.org/10.1109/TKDE.2006.130
ISSN	1558-2191

Impact and Reach

Statistics

DownloadsShow export options

Activity Overview

6 month trend

3,029Downloads

6 month trend

655Hits

Additional statistics for this dataset are available via IRStats2.

Altmetric

Repository staff only

Edit record