Crawler content extraction method based on machine learning large model

By adopting a machine learning large model-based method in crawler content extraction, the problems of low accuracy, high cost, low scalability, poor generalization ability and low computing resource utilization in the prior art are solved, and higher accuracy, lower cost, higher scalability and stronger generalization ability are achieved, and computing resources are effectively utilized.

CN120045765APending Publication Date: 2025-05-27SSE INFORMATION NETWORK LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510131002.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing crawler content collection solutions have problems such as low accuracy, high cost, low scalability, poor generalization capability and low computing resource utilization.

Method used

The crawler content extraction method based on the machine learning big model is adopted, and the content positioned by URL is classified through the Transformer model processed in natural language. The SimHash text deduplication algorithm is used for deduplication and freshness detection. A preliminary prompt word template is generated according to the content type, the context of the content is input into the machine learning big model, and the prompt word is dynamically generated, and the prompt word is dynamically adjusted according to feedback through the policy gradient model, and the content is finally parsed and JSON structured data is output.

Benefits of technology

Improves the accuracy of crawler extraction, reduces costs, improves scalability and generalization capabilities, and effectively utilizes computing resources to enable them to operate effectively in resource-constrained environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045765A_ABST
    Figure CN120045765A_ABST
Patent Text Reader

Abstract

The invention relates to the field of machine learning, and provides a crawler content extraction method based on a machine learning large model, and the method comprises the steps: obtaining a uniform resource locator URL from a content queue, and carrying out the classification of the content located by the uniform resource locator URL through a natural language processing Transform model, and obtaining a content type; selecting a corresponding cue word template according to a classification result, generating a preliminary cue word template according to a content type, inputting context of the content into a machine learning large model, dynamically generating cue words by the machine learning large model, and dynamically adjusting the cue words according to feedback through a strategy gradient method model; and according to the cue word and the classification result, analyzing the content and outputting JSON structured data, carrying out confidence scoring on the structured data, giving out scoring reason analysis, and carrying out visual display on the structured data. The method is high in accuracy, low in cost, high in expandability, high in generalization ability and high in computing resource utilization rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of machine learning, and particularly to a method for extracting crawler content based on a large machine learning model. Background Art

[0002] Existing crawler content collection solutions mainly analyze the crawled content through manual parsing and extract the content through the following methods.

[0003] 1. XPath / CSS Selectors: Use the XML Path Language (XPath) or Cascading Style Sheets (CSS) selectors to locate and extract specific elements in an HTML document; 2. Regular Expressions: Write specific patterns to match the text content in an HTML document; 3. HTML Parsing Libraries: Such as BeautifulSoup, lxml, etc. These libraries can parse HTML documents and allow developers to access and extract document content programmatically; 4. The technical languages used are mainly development languages such as JAVA and Python.

[0004] The above methods have the disadvantages of low accuracy, high cost, low scalability, poor generalization ability, and low utilization rate of computing resources. Summary of the Invention

[0005] To help solve the above technical problems, this application provides a method for extracting crawler content based on a large machine learning model, adopting the following technical solutions: A method for extracting crawler content based on a large machine learning model, wherein the method includes: Step S11: Obtain a Uniform Resource Locator (URL) from a content queue, and use the Transformer model of natural language processing to classify the content located by the Uniform Resource Locator (URL) to obtain a content type; Step S12: Perform duplicate removal and freshness detection on the crawled content through the SimHash text duplicate removal algorithm; Step S13: Classify the crawled content according to the content type; Step S2: Select a corresponding prompt word template according to the classification result of Step S13, generate a preliminary prompt word template according to the content type, input the context of the content into the large machine learning model, the large machine learning model dynamically generates prompt words, and dynamically adjusts the prompt words according to the feedback through the policy gradient method model; Step S3: Parse the content according to the prompt words and the classification result, output JSON structured data, perform a confidence score on the structured data and give an analysis of the scoring reason, and visually display the structured data.

[0006] Preferably, in the step S3, a confidence score is performed on the structured data through identifiers such as confidence thresholds, distributions, labels, etc.

[0007] Preferably, in the step S2, the prompt words include background, preference, introduction, objective, constraint, skill, and output format.

[0008] In summary, the present application has the following beneficial effects: 1. Improve accuracy: Use an advanced large model to enhance the semantic understanding of web content, thereby improving the accuracy of crawler extraction.

[0009] 2. Reduce costs: Use the trained large model to develop a technology that can adapt to changes in web page structures, thereby reducing the need for crawl rule updates and program maintenance work due to changes in the web page structure of the website.

[0010] 3. Improve scalability: Develop a scalable content extraction solution that can adapt to large-scale and diverse web data.

[0011] 4. Enhance generalization ability: Design a content extraction method that can generalize to different website structures and content types, reduce dependence on specific websites, and reduce the risk of crawlers being blocked.

[0012] 5. Effectively utilize computing resources: Design a content extraction method with low resource consumption so that it can operate effectively in resource-constrained environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 It is a schematic flowchart of an embodiment of a crawler content extraction method based on a machine learning large model of the present application; Figure 2 It is a schematic diagram of the working principle of the machine learning large model of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0014] The present application will be further described below with reference to the accompanying drawings. The structure and principle of the present application are very clear to those skilled in the art. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0015] Figure 1 It is a schematic flowchart of an embodiment of a crawler content extraction method based on a machine learning large model of the present application, Figure 2Schematic diagram of the working principle of the large machine learning model of this application.

[0016] Combined with Figure 1 and Figure 2 It can be understood that the crawler content extraction method of this application may include: Step S1: Data preprocessing, including: Step S11: Obtain a Uniform Resource Locator (URL) from the content queue, and use the Transformer model in natural language processing to classify the content located by the URL to obtain the content type. In step S11, the Transformer model (such as BERT, etc.) in natural language processing is used to analyze this content to determine its type (such as news, blog, product description, etc.). This process uses a deep learning model that can understand a large number of text features.

[0017] Step S12: Use the SimHash text deduplication algorithm to deduplicate and detect the freshness of the crawled content, avoiding repeated processing of the crawled content. In step S12, the SimHash algorithm is used to deduplicate the crawled content, reducing the storage and processing of duplicate data. SimHash can efficiently identify similar texts, thus ensuring that the obtained content is new and unique.

[0018] Step S13: Classify the crawled content according to the content type, ensuring that complex and detailed content enters the detailed parsing module first, while simple content is processed quickly, optimizing the scheduling.

[0019] Step S2 is used for model training, including selecting corresponding prompt templates according to the classification results of step S13, generating preliminary prompt templates according to the content type, inputting the context of the content into the large machine learning model, the large machine learning model dynamically generates prompt words, and the policy gradient method model dynamically adjusts the prompt words according to the feedback to improve the quality of the prompt words. In step S2, select a suitable prompt template. The prompt words include elements that may be related to the content, such as background, preferences, introductions, goals, constraints, skills, and output formats, etc. Generate a preliminary prompt template, then input the context of the content into the large machine learning model to dynamically generate appropriate prompt words. Through the policy gradient method model, optimize and adjust the prompt words according to the real-time feedback to ensure that the prompt words generated by the model are more in line with the actual situation and requirements of the content.

[0020] Step S3 is used for structured output, including parsing the content and outputting JSON structured data according to the prompt words and classification results, performing a confidence score on the structured data and giving an analysis of the scoring reasons, and visually displaying the structured data for users to view and provide feedback. In step S3, the structured data can be confidence scored through identifiers such as confidence thresholds, distributions, and labels.

[0021] The above prompt words can include background, preferences, introduction, goals, constraints, skills, and output format.

[0022] Specifically, the prompt words can parse web page content and output structured content as required, such as JSON and XML. Taking a web page parser as an example, the prompt words of this application are introduced as follows: ## Background - Background: A web page parser is a tool that can convert web page content into a structured document. It can extract elements such as text, images, and links from a web page and organize these elements into an ordered document format. Web page parsers are usually used in fields such as data collection, content management, and information retrieval.

[0023] ## Preferences - Preferences: The web page parser prefers simple and clear structured data formats, such as JSON, XML, etc. It tends to efficiently and accurately extract key information from web pages and maintain the integrity and consistency of the data.

[0024] ## Introduction - Profile: Extract web page content and convert it into a structured document.

[0025] ## Goals - Goals: 1. Extract elements such as text, images, and links from a web page.

[0026] 2. Organize the extracted elements into an ordered document format.

[0027] ## Constrains - Constrains: 1. The integrity and consistency of the data must be maintained.

[0028] 2. The copyright and usage terms of the web page must be complied with.

[0029] ## Skills - Skills: 1. Be able to parse HTML and CSS.

[0030] 2. Be able to handle dynamic content generated by JavaScript.

[0031] 3. Be able to convert the extracted data into JSON or XML format.

[0032] ## OutputFormat - OutputFormat: 1. Parse web page content and extract elements such as text, images, and links.

[0033] 2. Organize the extracted elements into a structured document in JSON or XML format.

[0034] Replace the traditional xpath for web content extraction with a large machine learning model, and output formatted documents such as Json, xml, and markdown through the large machine learning model.

Claims

1. A crawler content extraction method based on a machine learning large model, characterized in that: The method comprises: Step S11: obtaining a uniform resource locator URL from a content queue, and using a natural language processing Transformer model to classify the content located by the uniform resource locator URL to obtain a content type; Step S12: De-duplication and freshness detection of crawled content using SimHash text de-duplication algorithm; Step S13: Categorize the crawled content according to the content type; Step S2: selecting a corresponding prompt word template according to the classification result of step S13, and generating a preliminary prompt word template according to the content type, inputting the context of the content into the machine learning model, the machine learning model dynamically generates prompt words, and dynamically adjusts the prompt words according to feedback through the policy gradient method model; Step S3: according to the prompt words and the classification results, the content is parsed and JSON structured data is output, a confidence score is performed on the structured data and an analysis of the reasons for the score is given, and the structured data is visualized.

2. The crawler content extraction method based on a machine learning large model according to claim 1 is characterized in that: In step S3, the structured data is scored for confidence using confidence thresholds, distributions, labels, and other identifiers.

3. The crawler content extraction method based on machine learning big model according to claim 1 is characterized in that: In step S2, the prompt words include background, preference, introduction, goal, constraint, skill and output format.