Automatic text collection method for specific field

An automatic text, field-specific technology, applied in special data processing applications, network data indexing, instruments, etc., can solve the problem of not using the semantic relationship between words

Pending Publication Date: 2022-07-29
WUHAN UNIV OF TECH
View PDF0 Cites 0 Cited by
  • Summary
  • Abstract
  • Description
  • Claims
  • Application Information

AI Technical Summary

Problems solved by technology

Therefore, it is necessary to advance data cleaning and analysis to the acquisition runtime. In the classic information matching system, the calculation of similarity is based on strict matching, and the request text is strictly matched with the document library after word segmentation, without using the inter-lexical Semantic relationship

Method used

the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
View more

Image

Smart Image Click on the blue labels to locate them in the text.
Viewing Examples
Smart Image
  • Automatic text collection method for specific field
  • Automatic text collection method for specific field
  • Automatic text collection method for specific field

Examples

Experimental program
Comparison scheme
Effect test

Embodiment Construction

[0027] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, so as to facilitate a clear understanding of the present invention, but they do not limit the present invention.

[0028] like Figure 1-3 As shown, the technical solution adopted in the present invention is: an automatic text collection method oriented to a specific field, and the method includes the following steps:

[0029] S1, establish a url scheduling link pool and a task manager according to a preset target site; the url scheduling link pool is used to temporarily store the url of the webpage to be crawled; the task manager is used to obtain a url from the url scheduling link pool each time, Generate task thread access to the page content of the url;

[0030] S2, map the input keywords and the words in the Chinese thesaurus to a high-dimensional vector space through word2vec, and calculate and generate the subject word group;

...

the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
Login to View More

PUM

No PUM Login to View More

Abstract

The invention provides an automatic text collection method oriented to a specific field. The method comprises the following steps: establishing a url scheduling link pool and a task manager according to a preset target site; the method comprises the following steps: mapping input keywords and words in a Chinese synonym library to a high-dimensional vector space through word2vec, and calculating to generate a subject word group; the task manager cleans the html unstructured data in the accessed url page to obtain a long text; extracting Chinese feature sentences of the long text, extracting feature words in the Chinese feature sentences of the long text, and generating a long text feature word group; mapping the subject word group and the long text feature word group to a high-dimensional vector space, and calculating a semantic similarity value of the subject word group and the long text feature word group; and if the semantic similarity value reaches a set threshold value, writing the structured data in the url page corresponding to the long text feature word group into a database. According to the method, domain focusing is realized by converting rule matching of Chinese words and long texts into semantic distance calculation.

Description

technical field [0001] The invention belongs to the technical field of computer information and natural language processing, and in particular relates to an automatic text collection method oriented to a specific field. Background technique [0002] With the development of the Internet, rich and diverse text data are spread on the Internet. As a mainstream information, Internet text information has greater research value than other information sources. It is very necessary to focus and collect Internet news accurately and efficiently. It is of great significance in the fields of information retrieval and data mining. [0003] Text capture is a program that automatically crawls web pages and extracts web content, the purpose of which is to obtain information resources from the Internet. The general collection method does not distinguish which data the user wants in the collection process, so it will cause a large amount of invalid data to be stored, which not only brings cha...

Claims

the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
Login to View More

Application Information

Patent Timeline
no application Login to View More
IPC IPC(8): G06F16/951G06F16/955G06F16/215
CPCG06F16/951G06F16/955G06F16/215
Inventor旷海兰宋永超马小林刘新华
OwnerWUHAN UNIV OF TECH