Accurate Document Retrieval with Sub-Topic Attention Using Term Frequency – Inverse Document Frequency
Main Article Content
Abstract
In a real-world search environment, the queries that users submit are typically vague or involve
many aspects. In addition to being pertinent to the query, it is anticipated that the documents
that are obtained will be as varied as is humanly feasible in order to cater to a variety of
information requirements. In this paper, a document retrieval system is proposed based on
term-frequency – inverse document frequency (TF-IDF). The TF-IDF is a measurement for
statistical analysis to provide the quantified results in terms of a word significance for a given
document from the collection of database. In user modeling, text mining, and information
retrieval searches, this statistical parameter is used as a weight factor frequently. TF-IDF can
help rank the significance of a word in a context. Thus, this feature will be very useful while
retrieving the information which is relevant to the query based on the subtopic. The proposed
model compares with the earlier techniques. The accuracy of documents retrieval obtained by
the proposed approach is 98.92%.