文章摘要
庞宁.基于网页特征的特征词提取技术[J].西南民族大学自然科学版,2014,40(1):137-141
基于网页特征的特征词提取技术
Signature word extracting retrieval based on web feature
  
中文关键词: 特征词提取  网页  元数据  加权函数
英文关键词: signature word extracting  web  metadata  weighting function
基金项目:山西省自然科学基金(2012011011-4).
作者单位
庞宁 太原科技大学应用科学学院 
摘要点击次数: 3435
全文下载次数: 2340
中文摘要:
      特征词提取是一项提炼整个web页面内容的实用技术, 同时也为文本分类, 信息抽取应用提供了技术支持. 在web页面内容上, 利用段落间语义关系划分出网页内容的篇章结构, 并以此为基础使用网页的元数据和特殊标签, 设计了一个特征词的加权函数, 综合考虑了词频、词长和位置因子, 最后, 实验对比了各类位置因子对系统的贡献度. 实验结果表明, 改进方法的F1值比传统的TFIDF提取技术提高了15.5%, 其中, 位置因子中的标题, 关键词和摘要因素对系统的贡献最大.
英文摘要:
      Signature word extracting of the text is a useful technique which can abstract web page text, and it provides technical support for text classification, information extraction tasks. A web hierarchical structure is extracted through parsing the semantic relation between each adjacent paragraph in the web page contents. On the basis of the hierarchical structure, this paper uses the HTML metadata and special tags to design a weighting function, which is a combination of the factor of the frequency, length and location for a word. Meanwhile, an initial contrast analysis is carried out of various position factor about contributing degree to the system. Experimental results show that F1value of improved method has increased by 15.5% than that of the traditional TFIDF extraction method. The contributing degree to the system of the title, abstract and keywords in the location factor are the largest.
查看全文   查看/发表评论  下载PDF阅读器
关闭