作者: Xiao Wei , Chenglei Qin , Zheng Xu
DOI: 10.4018/IJCINI.2015040101
关键词:
摘要: Faceted search is an efficient search method to use the big data and one of its key issues is to extract facets from unstructured webpages automatically. It is still a problem to extract facets from massive unstructured webpages exactly and automatically. To solve the problem, this paper first proposed a novel index structure of webpages, the Multidimensional Semantic Index (MDSI), which holds rich semantics and are helpful to extract facets. In MDSI, the differently dimensional semantic indexes are bridged by mining the semantic mapping between them. Then, an automatic facet extraction method is proposed by analysing semantic mapping relations in MDSI. At last, to validate the effect of the proposed method, two datasets are constructed and the experimental results show that the proposed method is feasible and comparatively precise.