<?xml version="1.0"?>
<records>
  <record>
    <language>eng</language>
    <publisher>Ansari Education and Research Society</publisher>
    <journalTitle>Journal of Ultra Scientist of Physical Sciences</journalTitle>
    <issn/>
    <eissn/>
    <publicationDate>December 2009</publicationDate>
    <volume>21</volume>
    <issue>3</issue>
    <startPage>981</startPage>
    <endPage>988</endPage>
    <doi>jusps-A</doi>
    <publisherRecordId>1288</publisherRecordId>
    <documentType>article</documentType>
    <title language="eng">Comparison of similarity measure for web document clustering&#xA0;</title>
    <authors>
      <author>
        <name>Gunjan Ansariu00a0(gunjan_ansari@yahoo.co.in)</name>
        <affiliationId>1</affiliationId>
      </author>
    </authors>
    <affiliationsList>
      <affiliationName affiliationId="1">Department of Information Technology, JSS Academy of Technical Education, C-20/1, Sector-62 Noida</affiliationName>
    </affiliationsList>
    <abstract language="eng">&lt;p style="text-align: justify;"&gt;With the rapid growth of the World Wide Web (www), it becomes a critical issue to design and organize the vast amounts of on-line documents on the web according to their topic. Even for the search engines it is very important to group similar documents in order to improve their performance when a query is submitted to the system. Clustering is useful for taxonomy design and similarity search of documents on such a domain.&lt;/p&gt;&#xD;
&#xD;
&lt;p style="text-align: justify;"&gt;Similarity or distance measures play important role in the performance of clustering algorithms. This paper compares three term based similarity measures for web document clustering. The similarity measures used are Euclidean distance, cosine measure and jaccard measure. The clustering algorithm used is the so-called k-means clustering to cluster web documents. These three different similarity measures are used to find the similarity between documents. The clusters are formed based on their similarity measure calculation using k-means clustering algorithm. Overall similarity measure is used to evaluate the clusters formed using different similarity measure. Tested with web data, we observe that the Euclidean measure outperforms the other similarity measures in clustering accuracy.&lt;/p&gt;&#xD;
</abstract>
    <fullTextUrl format="html">https://www.ultrascientist.org/paper/1288/</fullTextUrl>
    <keywords>
      <keyword language="eng">Comparison </keyword>
    </keywords>
    <keywords>
      <keyword language="eng">similarity measure</keyword>
    </keywords>
    <keywords>
      <keyword language="eng">document clustering</keyword>
    </keywords>
  </record>
</records>
