On evaluation and training-set construction for duplicate detection

DOI:

关键词:

摘要: A variety of experimental methodologies have been used to evaluate the accuracy of duplicate-detection systems. We advocate presenting precision-recall curves as the most informative evaluation methodology. We also discuss a number of issues that arise when evaluating and assembling training data for adaptive systems that use machine learning to tune themselves to specific applications. We consider several different application scenarios and experimentally examine the effectiveness of alternative methods of collecting training data under each scenario. We propose two new approaches to collecting training data called static-active learning and weaklylabeled non-duplicates, and present experimental results on their effectiveness.

uni-leipzig.de PDF 下载加速

参考文章(0)

On evaluation and training-set construction for duplicate detection

来源期刊

我的账户

On evaluation and training-set construction for duplicate detection

来源期刊

相似文章 0

我的账户