Method, device, storage medium and program product for implementing intelligent AI to obtain web content based on LLM

Through the intelligent AI method based on LLM, converting user search instructions and analyzing HTML source code, the problem of failure of traditional web page data acquisition methods in extracting complex structure web pages is solved, and efficient and intelligent web page data acquisition is achieved.

CN119988713BActive Publication Date: 2025-06-27SHENZHEN SMARTCITY TECH DEV GRP CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510466015.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-06-27
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

Traditional web page data acquisition methods rely on preset rules, making it difficult to deeply understand and semantically analyze web page content, resulting in failure or inaccurate data extraction in complex or non-standardized web pages, increasing time cost.

Method used

Using an intelligent AI method based on LLM, the search instructions entered by the user are converted into query parameters through the semantic analysis model, sent to the search engine to obtain the web page, and the target web page is determined using preset filtering rules. Then, based on the LLM model, the HTML source code is analyzed, the target code tags and their contents are automatically identified, and the web page structure changes are dynamically adapted to changes.

Benefits of technology

It realizes intelligent data extraction of complex or non-standardized web pages, reduces the time cost of manually modifying crawler codes, and improves the intelligence and efficiency of web page content acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988713B_ABST
    Figure CN119988713B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device, storage medium and program product for an intelligent AI to obtain web page content based on an LLM, which relates to the technical field of data processing. The above method receives a search instruction input by a user, inputs the search instruction into a trained semantic parsing model, after the semantic parsing model converts the search instruction into query parameters, sends the query parameters to a search engine, then receives the web page obtained by the search engine according to the query parameters, determines a target web page from the web page according to a preset screening rule, crawls the HTML source code of the target web page based on a pre-obtained authorization result, inputs the HTML source code into a trained LLM model, and after the LLM model determines target code tags according to the HTML source code, obtains the text content within the target code tags. Among them, the LLM model has powerful language understanding ability, can dynamically adapt to changes in the web page structure, and reduces the time cost of crawling web page content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular, to a method, device, storage medium, and program product for an intelligent AI to obtain web page content based on LLM. Background Art

[0002] In traditional methods for obtaining web page data, first, a traditional crawler program is used to automatically access web pages and download their HyperText Markup Language (HTML) content based on a predefined list of Uniform Resource Locators (URLs) or seed URLs. Subsequently, regular expressions or XML Path Language (XPath) are used to parse the HTML page to locate and extract the required data fields contained in specific elements on the page.

[0003] However, the above methods mainly rely on preset rules or patterns to extract web page data and lack the ability to deeply understand and semantically analyze web page content. This results in the crawler program being difficult to adapt to changes in the web page structure when faced with complex or non-standardized web pages, and thus data extraction fails or the extracted data is inaccurate. At this time, complex conversion logic and manual intervention are often required to continuously update and maintain the crawler rules to adapt to web page changes, resulting in a high time cost for obtaining web page content.

[0004] The above content is only used to assist in understanding the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0005] The main purpose of this application is to provide a method, device, storage medium, and program product for an intelligent AI to obtain web page content based on LLM, aiming to solve the technical problem of high time cost for obtaining web page content.

[0006] To achieve the above objective, this application proposes a method for an intelligent AI to obtain web page content based on LLM, and the method includes:

[0007] Receiving a search instruction input by a user, inputting the search instruction into a trained semantic parsing model, and after the semantic parsing model converts the search instruction into query parameters, sending the query parameters to a search engine;

[0008] Receiving the web pages obtained by the search engine according to the query parameters, and determining target web pages from the web pages according to preset screening rules;

[0009] Based on the pre-obtained authorization result, crawl the HTML source code of the target web page, input the HTML source code into the trained LLM model, and after the LLM model determines the target code label according to the HTML source code, obtain the text content within the target code label.

[0010] In one embodiment, the semantic parsing model receives the search instruction, tokenizes the search instruction, and obtains at least one word.

[0011] The semantic parsing model identifies the entity words in the search instruction according to a preset named entity recognition algorithm, and uses the entity words as the query parameters.

[0012] In one embodiment, the step of determining the target web page from the web pages according to a preset screening rule includes:

[0013] Obtain preset screening elements, where the screening elements include the number of times the query parameter appears in the web page, the update time of the web page, and the source of the web page.

[0014] Based on the preset element weights, determine the screening element scores corresponding to each web page, perform weighted summation on the screening element scores, and obtain the score of each web page.

[0015] Use the web pages with scores greater than a preset threshold as the target web pages.

[0016] In one embodiment, the step in which the LLM model determines the target code label according to the HTML source code includes:

[0017] Receive the HTML source code, parse the HTML source code into a DOM tree, and construct label relationship graph data according to the DOM tree.

[0018] According to the preset internal cluster tightness, perform hierarchical clustering analysis on the label nodes in the DOM tree to obtain label hierarchical clustering data.

[0019] Match according to the label relationship graph data and the label hierarchical clustering data in a preset knowledge base to determine the web page type of the target web page.

[0020] Based on the web page type of the target web page, determine the target code label in the HTML source code.

[0021] In one embodiment, the step of determining the target code label in the HTML source code based on the web page type of the target web page includes:

[0022] According to the preset knowledge base, determine the valid code labels corresponding to the web page type.

[0023] Based on the valid code label, determine the target code label in the DOM tree.

[0024] In one embodiment, the step of determining the target code label in the DOM tree based on the valid code label includes:

[0025] Read the first class name corresponding to the code label in the DOM tree and the second class name corresponding to the valid code label;

[0026] Convert the first class name and the second class name into embedding vectors, and calculate the similarity between the first class name and the second class name;

[0027] If the similarity between the first class name and the second class name is greater than a preset threshold, determine the code label corresponding to the first class name as the target code label.

[0028] In one embodiment, the step of extracting the text content within the target code label includes:

[0029] Traverse the child nodes of the target code label node in the DOM tree;

[0030] Access the node type attribute of the child node;

[0031] If the node type attribute of the child node is a text node, extract the text content corresponding to the child node.

[0032] In addition, to achieve the above object, the present application also proposes a device for an intelligent AI to obtain web page content based on an LLM. The device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. The computer program is configured to implement the steps of the method for an intelligent AI to obtain web page content based on an LLM as described above.

[0033] In addition, to achieve the above object, the present application also proposes a storage medium. The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the method for an intelligent AI to obtain web page content based on an LLM as described above.

[0034] In addition, to achieve the above object, the present application also provides a computer program product. The computer program product includes a computer program. When the computer program is executed by a processor, it implements the steps of the method for an intelligent AI to obtain web page content based on an LLM as described above.

[0035] The present application provides a method, device, storage medium, and program product for an intelligent AI to obtain web page content based on an LLM. By receiving a search instruction input by a user, the search instruction is input into a trained semantic parsing model. After the semantic parsing model converts the search instruction into query parameters, the query parameters are sent to a search engine. Then, the web pages obtained by the search engine according to the query parameters are received, and target web pages are determined from the web pages according to a preset filtering rule. Based on a pre-obtained authorization result, the HTML source code of the target web page is obtained, and the HTML source code is input into a trained LLM model. After the LLM model determines target code tags according to the HTML source code, the text content within the target code tags is obtained. This method analyzes the HTML source code of the target web page through a trained LLM model, can automatically identify the target code tags and their content. Among them, the LLM model has strong language understanding ability, can dynamically adapt to changes in the web page structure, and does not need to rely on fixed rules or patterns, reducing the time cost of manually modifying crawler code to obtain web page content. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0037] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0038] Figure 1 It is a schematic flowchart provided for the first embodiment of the method for an intelligent AI to obtain web page content based on an LLM in the present application;

[0039] Figure 2 It is a schematic flowchart provided for the second embodiment of the method for an intelligent AI to obtain web page content based on an LLM in the present application;

[0040] Figure 3 It is a schematic flowchart provided for the third embodiment of the method for an intelligent AI to obtain web page content based on an LLM in the present application;

[0041] Figure 4 It is a schematic diagram of the device structure of the hardware operating environment involved in the method for an intelligent AI to obtain web page content based on an LLM in the embodiments of the present application.

[0042] The implementation, functional features, and advantages of the purpose of the present application will be further described with reference to the embodiments and the drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not used to limit this application.

[0044] To better understand the technical solutions of this application, the following will be described in detail in conjunction with the accompanying drawings of the specification and specific implementation manners.

[0045] In the traditional method of obtaining web page data, first, a traditional crawler program is used to automatically access web pages and download their HyperText Markup Language (hereinafter uniformly described as HTML) content according to a predefined list of Uniform Resource Locators (hereinafter uniformly described as URLs) or seed URLs. Subsequently, regular expressions or Extensible Markup Language Path Language (hereinafter uniformly described as XPath) are used to parse the HTML page to locate and extract the required data fields contained in specific elements on the page.

[0046] However, the above method mainly relies on preset rules or patterns to extract web page data and lacks the ability to deeply understand and semantically analyze web page content. This results in the crawler program being difficult to adapt to changes in the web page structure when facing complex or non-standardized web pages, and further causes data extraction failures or inaccurate extracted data. At this time, complex conversion logic and manual intervention are often required to continuously update and maintain the crawler rules to adapt to the changes in the web page, resulting in a high time cost for obtaining web page content.

[0047] In view of the above problems, this application proposes a method for an intelligent AI to obtain web page content based on an LLM (Large Language Model). By receiving a search instruction input by the user, the search instruction is input into a trained semantic parsing model. After the semantic parsing model converts the search instruction into query parameters, the query parameters are sent to a search engine. Then, the web pages obtained by the search engine according to the query parameters are received, and the target web page is determined from the web pages according to preset screening rules. Based on the pre-obtained authorization result, the HTML source code of the target web page is crawled, and the HTML source code is input into a trained LLM model. After the LLM model determines the target code tags according to the HTML source code, the text content within the target code tags is obtained. This method analyzes the HTML source code of the target web page through a trained LLM model, can automatically identify the target code tags and their content. Among them, the LLM model has a powerful language understanding ability, can dynamically adapt to changes in the web page structure, does not need to rely on fixed rules or patterns, and reduces the time cost spent on manually modifying crawler code to obtain web page content.

[0048] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, etc., or an electronic device capable of implementing the above functions. Hereinafter, taking an artificial intelligence (AI) system for obtaining web page content based on the LLM model as an example, this embodiment and the following embodiments will be described.

[0049] Based on this, the first embodiment proposed in this application provides a method for an intelligent AI to obtain web page content based on the LLM. Referring to Figure 1 , in this embodiment, the method for an intelligent AI to obtain web page content based on the LLM includes steps S10 to S30:

[0050] Step S10, receive a search instruction input by a user, input the search instruction into a trained semantic parsing model, and after the semantic parsing model converts the search instruction into query parameters, send the query parameters to a search engine.

[0051] In the case of obtaining website authorization and performing desensitization processing on privacy data, traditional crawler programs usually rely on a fixed URL list or rules to crawl web pages, lacking the ability to deeply understand the user's search intent and perform semantic analysis. By introducing a semantic parsing model in this embodiment, it is possible to achieve intelligent parsing and conversion of the user's search instruction, enabling the crawler program to more flexibly adapt to different search engines and complex web page structures, and improving the intelligent level of crawling web page content.

[0052] Exemplarily, the user inputs a search instruction through a user interface or interface. The semantic parsing model performs in-depth semantic analysis on the instruction input by the user, understands the user's true intent, and converts this intent into query parameters that can be understood by a computer, and then sends the query parameters to the search engine. The search engine executes a search task based on these parameters and returns relevant search results.

[0053] Optionally, a user interface is set, and a text input box is provided on the user interface. The user can input a search instruction in the text input box, or the user is allowed to input a search instruction through a microphone, and the speech recognition algorithm is used to convert the user's speech into text, and then this text is used as the search instruction.

[0054] Optionally, step S10 includes steps S11 to S12:

[0055] Step S11, the semantic parsing model receives the search instruction and performs word segmentation on the search instruction to obtain at least one word.

[0056] Step S12: The semantic parsing model identifies the entity words in the search instruction according to a preset named entity recognition algorithm, and uses the entity words as the query parameters.

[0057] Before inputting the search instruction into the semantic parsing model, the search instruction can be preprocessed, including removing irrelevant characters, word segmentation, part-of-speech tagging, etc., so that the semantic parsing model can better understand and process the search instruction. The preprocessed search instruction is input into the trained semantic parsing model, and the semantic parsing model uses deep learning techniques such as Long Short-Term Memory (LSTM) and Bidirectional Encoder Representations from Transformers (BERT) to deeply understand and analyze the search instruction. Among them, the deep learning-based model can automatically learn the deep features and context information of words, so as to perform word segmentation more accurately and obtain a list of words. Then, the semantic parsing model extracts the text representation of words, the context information features such as the words before and after them, etc., according to a preset named entity recognition algorithm such as Hidden Markov Model (HMM) and Conditional Random Field (CRF), classifies and judges each word in the word list to determine whether it belongs to an entity word. By identifying the entity words in the sentence, the query parameters are extracted. For example, from "I want to read the latest technology news", the semantic parsing model can identify "technology" and "news" as the query parameters.

[0058] Step S20: Receive the web pages obtained by the search engine according to the query parameters, and determine the target web pages from the web pages according to preset screening rules.

[0059] After receiving the query parameters, the search engine will search for web pages related to the query parameters in the web page database according to its internal indexing and ranking algorithms, and return the web pages that match the query parameters to the artificial intelligence system for obtaining web page content based on the LLM model after the search is completed. After receiving the web pages returned by the search engine, the artificial intelligence system for obtaining web page content based on the LLM model further screens these web pages according to specific requirements or goals. To achieve this screening process, screening rules can be preset in advance. The screening rules can be based on multiple dimensions such as the content, source, publication time, relevance score, and user evaluation of the web pages. According to these preset screening rules, the artificial intelligence system for obtaining web page content based on the LLM model will evaluate the received list of web pages one by one, and finally determine the web pages that meet the requirements as the target web pages.

[0060] Optionally, step S20 includes steps S21 to S23:

[0061] Step S21, obtaining preset screening elements, where the screening elements include the number of times the query parameter appears in the web page, the update time of the web page, and the source of the web page.

[0062] The screening elements are indicators used to evaluate whether a web page meets specific conditions. In this embodiment, the screening elements include, but are not limited to, the number of times the query parameter appears in the web page, the update time of the web page, and the source of the web page. For each screening element, the artificial intelligence system for obtaining web page content based on the LLM model will obtain its corresponding value. For example, a text matching algorithm is used to calculate the frequency of the query parameter appearing in the web page content; the update time of the web page is extracted from the metadata or HTML tags of the web page; the source of the website is judged according to the domain name of the web page, the URL structure, or a list of known authoritative websites.

[0063] Exemplarily, a screenshot tool is used to take a screenshot of the web page, and the text content in the screenshot is recognized. The text content is segmented into independent words or phrases through a word segmentation tool such as a Chinese word segmenter, and then the number of times the query parameter appears is counted. The screening elements may also include the relevance between the query parameter and the web page title. For example, a crawler program is used to first crawl the title of the web page, and a semantic matching algorithm such as word2vec (word to vector), bidirectional Transformer encoder representation, etc. is used to calculate the relevance between the query parameter and the web page title. For the update time of the web page, the HTML code of the web page can be parsed through an HTML parser to find tags containing update time information, such as the "last-modified" attribute in the " <meta> " tag, the "Last-Modified" field in the "" section, etc. The metadata related to the update time is extracted from the parsed HTML. For the source of the web page, the URL in the web page can be parsed through a regular expression or a URL parsing library, and the website source of the web page is judged according to the URL.

[0064] Step S22, determining the screening element score corresponding to each web page based on the preset element weights, and performing weighted summation on the screening element scores to obtain the score of each web page.

[0065] Step S23, using the web pages with scores greater than a preset threshold as the target web pages.

[0066] Exemplarily, an element weight is determined for the screening element, and the screening element score corresponding to the web page is calculated according to the element weight. Among them, the more times the query parameter appears in the web page, and the closer the update time of the web page is to the query time, the more valuable the web page content is, and the larger the corresponding screening element score is. A preset website list is established inside the web page content acquisition artificial intelligence system based on the LLM model. The website list is sorted according to the authority of the website, and the sorted website list is used as the basis for the subsequent screening element score. Different intervals are divided into different screening element scores for the website list. If the web page source is located at the front of the list, a higher screening element score is assigned. If the web page source is not in the website list, a lower screening element score is assigned. Finally, the score of each screening element is multiplied by its corresponding element weight, and then weighted summed to obtain the total score of each web page, and the web page with a score greater than the preset threshold is used as the target web page.

[0067] Step S30, based on the authorization result obtained in advance, crawl the HTML source code of the target web page, input the HTML source code into the trained LLM model, and after the LLM model determines the target code tag according to the HTML source code, extract the text content in the target code tag.

[0068] After obtaining website authorization and desensitizing private data, the crawler program is used to capture the HTML source code of the target web page, and the HTML source code is passed as input data to a trained LLM model. The LLM model will analyze the input HTML source code, identify the various HTML tags in it, and determine the target code tags from which text content needs to be extracted based on the trained algorithm logic. Once the target code tags are determined, the LLM model will extract text content from these tags, including web page titles, paragraph text, link text, etc.

[0069] The LLM model determines the web page type of the target web page according to the HTML source code, and determines the target code tag in the HTML source code according to the web page type.

[0070] It should be noted that different types of web pages have different structures. For example, news web pages usually have a title, text, author, release date, etc., product web pages usually have a product name, price, specifications, evaluation, etc., and forum web pages usually have a topic, author, release date, comments, etc. The LLM model can infer the type of web page based on its structure. Different types of web pages also use different HTML tags. For example, news web pages usually use " <h1>"The label indicates the title and is commonly used on product web pages" ”Labels represent product information, which are usually used on forum web pages" ”The label represents a user comment. Therefore, compared with the traditional method of extracting web page data through preset regular expressions or XPath, the method of using an LLM model to analyze HTML source code and crawl text content according to the type of web page in this embodiment can automatically adapt to web pages with complex or non-standard structures without the need for manual updating of the crawler program code, which can reduce the time cost of web page content crawling.

[0071] Specifically, first collect a large number of HTML source codes of different types of web pages, annotate these HTML source codes to clarify the type of each web page and the HTML tags in each type of web page. Use the collected annotated data to train the LLM model to learn how to identify the type of web page from the HTML source code and understand the uses of HTML tags in different types of web pages. When a new HTML source code of a web page is input into the LLM model, the LLM model first analyzes the structure and content of the HTML web page, uses the knowledge it has learned during the training process to infer the type of this web page, and then identifies the corresponding target code tags according to the type of web page, and extracts the text content from the target code tags.

[0072] Based on the first embodiment of this application, in the second embodiment of this application, the same or similar content as in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 2 , in step S30, the steps for the LLM model to determine the target code tags according to the HTML source code include steps S31 to S34:

[0073] Step S31, receive the HTML source code, parse the HTML source code into a DOM tree, and construct label relationship graph data according to the DOM tree.

[0074] Exemplarily, after receiving the HTML source code, the LLM model parses the HTML source code into a DOM (Document Object Model) tree through an HTML parser. The DOM tree is a tree structure, where each node represents an element, attribute, or text node in the HTML document. Based on the DOM tree, construct label relationship graph data. The nodes in the label relationship graph data represent the tags in the DOM tree such as " ”、" ”、" ", etc. Edges represent the parent-child relationship, sibling relationship, or other custom relationships between tags. Parsing the HTML source code into a DOM tree helps to understand the structure and content of a web page, and the tag relationship graph data can more intuitively represent the relationships between tags.

[0075] Specifically, the LLM model first traverses the DOM tree, using algorithms such as depth-first search or breadth-first search to traverse from the root node. For each tag node in the DOM tree, a corresponding graph node is created in the graph structure, and other relevant content within the tag is associated with the graph node as child nodes. After traversing the DOM tree and adding all necessary edges, the constructed tag relationship graph data is output.

[0076] Step S32: Perform hierarchical clustering analysis on the tag nodes in the DOM tree according to a preset internal cluster tightness to obtain tag hierarchical clustering data.

[0077] It should be noted that hierarchical clustering analysis is an unsupervised learning algorithm that can divide a dataset into multiple levels or clusters, and these clusters have similar characteristics under a specific metric, while the internal cluster tightness is used to determine the tightness of the tag nodes within each cluster in the tag hierarchical clustering data.

[0078] Exemplarily, assign a unique cluster identifier to each tag node in the DOM tree, calculate the similarity or distance between each pair of clusters, and select the pair of clusters with the highest similarity or the smallest distance for merging. Calculate the internal cluster tightness of the tag hierarchical clustering data formed by the merger. When the internal tightness of the merged cluster reaches or exceeds the preset threshold, stop the merging. For example, in a news web page, " <header> ”Labels are usually used to contain the introductory content of a page, which includes" <h1>”、"< / h1> <h2>”、"< / h2> <h3>” tags, representing first-level headings, second-level headings, and third-level headings respectively. Through tag hierarchical clustering, these headings can be merged into a cluster to more clearly display the hierarchical structure between the tags.

[0079] Among them, the internal cluster tightness can be calculated based on the distance from the tag nodes within the cluster to the cluster center, such as the mean squared error or the mean Manhattan distance, and the cluster center can be user-defined. The tag hierarchical clustering data can more clearly display the hierarchical relationship between each tag in the HTML source code, thus reflecting the structure of the web page.

[0080] Step S33, according to the tag relationship graph data and the tag hierarchical clustering data, perform matching in a preset knowledge base to determine the web page type of the target web page.

[0081] Exemplarily, build a knowledge base within the LLM model, covering various web page types and the corresponding tag relationship graphs and hierarchical clustering features for each web page type. Adopt a specific similarity calculation method such as cosine similarity, etc., and compare the tag relationship graph data and tag hierarchical clustering data of the target web page with the tag relationship graphs and hierarchical clustering features corresponding to each web page type in the knowledge base one by one to calculate the similarity scores. According to the similarity scores, select the web page type with the highest score as the matching result for output.

[0082] Step S34, based on the web page type of the target web page, determine the target code tags in the HTML source code.

[0083] It should be noted that different types of web pages have different structural and content characteristics, so the corresponding HTML tags will also be different. Determining the target code tags based on the web page type can more accurately locate the required HTML elements, thereby improving the accuracy of data extraction or web page parsing. For example, for news web pages, key information such as titles, texts, and publication times can be mainly extracted; while for e-commerce web pages, product names, prices, pictures, and purchase links need to be concerned. Different types of web pages have different information display requirements and layout styles, so different HTML tags are used to implement them. For example, news web pages usually use "< / h3> <h1>”、"< / h1> <h2>” and other tags to define headings, while e-commerce web pages will use " ”、" Use tags such as "etc." to layout product information and purchase buttons.

[0084] Exemplarily, through model training, the LLM model is made to recognize the HTML tags corresponding to the web page content in different types of web pages. First, collect a large amount of web page data of different types, including news, e-commerce, blogs, etc. The web page data includes HTML source code and corresponding tag information. Manually annotate the key features from the HTML source code, including tag names, tag attributes, tag hierarchy relationships, text content, etc. Divide these web page data into a training set and a validation set, select appropriate machine learning or deep learning algorithms, such as support vector machines, random forests, neural networks, etc., and use the training set to train the LLM model. During the training process, the LLM model will learn how to predict the output tags based on the input HTML source code. Then use the validation set to evaluate the performance of the model through cross-validation or other evaluation methods, and optimize the model according to the evaluation results, including adjusting model parameters, increasing the number of features, or improving the feature extraction method, etc.

[0085] Optionally, step S34 includes steps S341 to S342:

[0086] Step S341, determine the valid code tags corresponding to the web page type according to the preset knowledge base.

[0087] Exemplarily, in the knowledge base within the LLM model, for each web page type, list the valid code tags used to identify the text content in the web page. For example, news web pages may focus on " <title>”、"< / title> <h1>” and other tags. After determining the web page type of the target web page, the LLM model retrieves a list of valid code tags related to this type from the preset knowledge base.

[0088] Step S342, based on the valid code tags, determine the target code tag in the DOM tree.

[0089] Exemplarily, traverse the DOM tree. According to the list of valid code tags obtained from the knowledge base, recursively traverse the nodes of the DOM tree and check whether the tag name of each node matches the tags in the list.

[0090] Optionally, step S342 includes steps S3421 to S3423:

[0091] Step S3421, read the first class name corresponding to the code tag in the DOM tree and the second class name corresponding to the valid code tag.

[0092] Step S3422, convert the first class name and the second class name into embedding vectors and calculate the similarity between the first class name and the second class name.

[0093] Step S3423, if the similarity between the first class name and the second class name is greater than the preset threshold, determine the code tag corresponding to the first class name as the target code tag.

[0094] It should be noted that in HTML, the class name is an attribute that can be attached to HTML elements such as HTML code tags to specify the style or behavior of the element. The class name uses "class" as the attribute name, and its value is one or more class identifiers separated by spaces. For example: " This is a header ”, where " "An element (i.e., a code tag) has two class names: "header" and "main-header". In web design and development, developers usually follow certain naming conventions to imply or indicate the content, function, or style of an element through class names. For example, a class name "product-description" is very likely to indicate that the element contains product description information. By analyzing the class names, the semantics of the text content within the code tag can be identified and deduced.

[0095] In traditional methods for obtaining web page data, when parsing an HTML page through regular expressions or XPath to locate and extract code tags in the page, it is usually in a hard-coded manner to query tags with specific sub-elements or attributes. However, when the structure of the HTML page changes, or when the naming of class names by developers changes, it will cause the originally valid regular expression or XPath expression to become invalid and unable to obtain the content in the web page. Therefore, the semantic information of the text content within the code tag can be analyzed through the class name of the code tag, so as to determine whether a code tag is the target code tag.

[0096] Exemplarily, traverse the DOM tree to find the value of the class attribute of each code tag, which is the first class name. At the same time, obtain the valid code tags and their corresponding expected class attribute values, which are the second class names, from a preset knowledge base. By comparing the actual class attribute value of the code tag with the expected value, it can be determined whether the tag meets the conditions of the target code tag.

[0097] Specifically, use a pre-trained word embedding model to convert the first class name and the second class name into embedding vectors. Then, calculate the similarity between these embedding vectors. The embedding vectors can capture the semantic relationships between class names. By calculating the similarity, the semantic proximity between the two names can be evaluated. Compare the calculated similarity with a preset threshold. If the similarity is greater than the preset threshold, it is considered that the code tag corresponding to the first class name is the target code tag.

[0098] In this embodiment, the HTML source code is parsed by the LLM model to construct a tag relationship graph and hierarchical clustering data, and the web page type is matched in combination with a preset knowledge base, and then the target code tag is determined. Specifically, the model first parses the HTML source code into a DOM tree and constructs a tag relationship graph to clearly show the parent-child, sibling, and other relationships between tags; then, hierarchical clustering analysis is performed on the tag nodes according to the preset internal cluster tightness to divide the tag clusters. Subsequently, in combination with the tag relationship graph and hierarchical clustering data, the target web page type is matched in the knowledge base. Through hierarchical clustering and the tag relationship graph, the web page structure can be more clearly reflected, and the target code tag can be accurately located. Finally, according to the web page type, the effective code tags are obtained from the knowledge base, and the target code tags are further determined in the DOM tree through the embedding vector similarity of the class names. Based on the semantic analysis of the class names, the parsing failure caused by the change of the HTML structure or class names can be avoided, thereby improving the flexibility of web page content extraction and reducing the time cost of web page content acquisition.

[0099] Based on the above embodiments of the present application, in the third embodiment of the present application, the same or similar content as the above embodiments can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 3 , step S30 further includes steps S35 to S36:

[0100] Step S35, traverse the child nodes of the target code tag node in the DOM tree.

[0101] Step S36, access the node type attribute of the child node.

[0102] Step S37, if the node type attribute of the child node is a text node, extract the text content corresponding to the child node.

[0103] Exemplarily, in the DOM tree, the nodes mounted under the target code tag node can be element nodes, text nodes, comment nodes, etc. Write a recursive function that receives a node as a parameter and accesses all the child nodes of the node, starting from the target code tag node and sequentially accessing its direct child nodes.

[0104] Among them, in the DOM, each node has a type attribute, which is used to indicate the type of the node, such as an element node, a text node, a comment node, etc. Exemplarily, in JavaScript and other programming languages and libraries that support DOM operations, the node type attribute such as "nodeType" is used to indicate the type of the node, and each node type corresponds to a specific integer value. For example, the value of the element node (ELEMENT_NODE) is 1, representing an element in an HTML or XML document; the value of the text node (TEXT_NODE) is 3, representing the text content in an element or attribute, which is used to store actual text data. The value of the comment node (COMMENT_NODE) is 8, representing a comment in an HTML or XML document. In JavaScript, the node type attribute can be obtained by accessing the "nodeType" attribute of the node. If the node type attribute of the accessed node is a text node, the "nodeValue" attribute or the "textContent" attribute can be used to obtain the text content within the target code tag.

[0105] Optionally, obtain the text content within the target tag and the class name corresponding to the target code tag, use the class name as an attribute of the text content, and save it as structured data in the form of key-value pairs with the text content, store it in the database of the web content acquisition artificial intelligence system, and provide an application programming interface for other systems to call.

[0106] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the method for the intelligent AI to obtain web content based on the LLM in this application. Any simple transformation in more forms based on this technical concept is within the protection scope of this application.

[0107] This application provides a device for the intelligent AI to obtain web content based on the LLM. The device for the intelligent AI to obtain web content based on the LLM includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for the intelligent AI to obtain web content in the first embodiment above.

[0108] Next, refer to Figure 4 , which shows a schematic structural diagram of a device suitable for implementing the web content acquisition based on the LLM in the embodiments of this application. The device for the intelligent AI to obtain web content in the embodiments of this application may include, but is not limited to, mobile terminals such as laptop computers and digital broadcast receivers, and fixed terminals such as desktop computers. Figure 4 The device shown for implementing an intelligent AI to obtain web content based on an LLM is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present application.

[0109] As Figure 4 shown, the device for implementing an intelligent AI to obtain web content based on an LLM may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM, Read Only Memory) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM, Random Access Memory) 1004. In the random access memory 1004, various programs and data required for the operation of the device for implementing an intelligent AI to obtain web content based on an LLM are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, etc.; an output device 1008 including, for example, a liquid crystal display (LCD, Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the device for implementing an intelligent AI to obtain web content based on an LLM to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a device for implementing an intelligent AI to obtain web content based on an LLM having various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be implemented or had alternatively.

[0110] Specifically, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are executed.

[0111] The device for an intelligent AI to obtain web content based on an LLM provided by this application adopts the method for an intelligent AI to obtain web content based on an LLM in the above-mentioned embodiment, and can solve the technical problem of relatively high time cost for web content crawling. Compared with the prior art, the beneficial effects of the device for an intelligent AI to obtain web content based on an LLM provided by this application are the same as those of the method for an intelligent AI to obtain web content based on an LLM provided by the above-mentioned embodiment, and other technical features in the device for an intelligent AI to obtain web content based on an LLM are the same as the features disclosed in the method of the previous embodiment, and will not be elaborated here.

[0112] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0113] As mentioned above, only the specific implementation manners of this application are described, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all of them should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claimed rights.

[0114] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the method for an intelligent AI to obtain web content based on an LLM in the above-mentioned embodiment.

[0115] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM), or a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination of the above.

[0116] The above computer-readable storage medium may be included in a device for an intelligent AI based on an LLM to obtain web content; or it may exist independently and not be assembled into a device for an intelligent AI based on an LLM to obtain web content.

[0117] The above computer-readable storage medium carries one or more programs. When the above one or more programs are executed by a device for an intelligent AI based on an LLM to obtain web content, the device for an intelligent AI based on an LLM to obtain web content can write computer program code for performing the operations of this application in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages - such as Java, Smalltalk, C++; and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, or executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0118] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutively represented blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0119] The modules described in the embodiments of the present application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the unit itself in some cases.

[0120] The readable storage medium provided by this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the method for implementing intelligent AI to obtain web page content based on LLM as described above, and can solve the technical problem of relatively high time cost for web page content crawling. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the method for implementing intelligent AI to obtain web page content based on LLM provided by the above embodiments, and will not be elaborated here.

[0121] This application also provides a computer program product, including a computer program, and the steps of the method for implementing intelligent AI to obtain web page content based on LLM as described above are implemented when the computer program is executed by a processor.

[0122] The computer program product provided by this application can solve the technical problem of relatively high time cost for web page content crawling. Compared with the prior art, the beneficial effects of the computer program product provided by this application are the same as those of the method for implementing intelligent AI to obtain web page content based on LLM provided by the above embodiments, and will not be elaborated here.

[0123] The above are only partial embodiments of this application, and do not limit the patent scope of this application accordingly. Any equivalent structural transformation made by using the content of the specification and drawings of this application under the technical concept of this application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of this application. < / h1> < / h2> < / header> < / h1>

Claims

1. A method for realizing intelligent AI acquisition of webpage content based on LLM, characterized in that: The method comprises: Receive a search instruction input by a user, input the search instruction into a trained semantic parsing model, and after the semantic parsing model converts the search instruction into a query parameter, send the query parameter to a search engine; Receiving web pages acquired by the search engine according to the query parameters, and determining a target web page from the web pages according to a preset screening rule; Based on the authorization result obtained in advance, crawl the HTML source code of the target webpage, and input the HTML source code into the trained LLM model, wherein the LLM model receives the HTML source code, parses the HTML source code into a DOM tree, and constructs label relationship graph data according to the DOM tree, wherein the nodes in the label relationship graph data represent labels in the DOM tree, and the edges represent the parent-child relationship, sibling relationship or other custom relationship between the labels, assign a cluster identifier to each label node in the DOM tree, calculate the similarity or distance between each pair of clusters, select the cluster pair with the highest similarity or the smallest distance for merging, calculate the internal cluster density of the merged label hierarchical clustering data, and stop merging when the internal density of the merged cluster reaches or exceeds a preset threshold, match the label relationship graph data and the label hierarchical clustering data in a preset knowledge base, determine the webpage type of the target webpage, determine the valid code tag corresponding to the webpage type according to the preset knowledge base, and determine the target code tag in the DOM tree based on the valid code tag; Traversing the child nodes of the target code tag node in the DOM tree, and accessing the node type attributes of the child nodes; If the node type attribute of the child node is a text node, the text content corresponding to the child node is extracted.

2. The method for realizing intelligent AI acquisition of webpage content based on LLM according to claim 1, characterized in that: The step of converting the search instruction into a query parameter by the semantic parsing model comprises: The semantic parsing model receives the search instruction and performs word segmentation on the search instruction to obtain at least one word; The semantic parsing model identifies entity words in the search instruction according to a preset named entity recognition algorithm, and uses the entity words as the query parameters.

3. The method for realizing intelligent AI acquisition of webpage content based on LLM according to claim 2, characterized in that: The step of determining the target web page from the web pages according to the preset screening rules comprises: Obtaining a preset screening element, wherein the screening element includes the number of times the query parameter appears in the webpage, the update time of the webpage, and the source of the webpage; Determine the screening element score corresponding to each of the web pages based on the preset element weights, and perform weighted summation on the screening element scores to obtain the score of each of the web pages; The web page with a score greater than a preset threshold is used as the target web page.

4. The method for realizing intelligent AI acquisition of webpage content based on LLM according to claim 1, characterized in that: The step of determining the target code tag in the DOM tree based on the valid code tag comprises: Read the first category name corresponding to the code tag in the DOM tree, and the second category name corresponding to the valid code tag; Convert the first category name and the second category name into an embedding vector, and calculate the similarity between the first category name and the second category name; If the similarity between the first category name and the second category name is greater than a preset threshold, the code label corresponding to the first category name is determined to be the target code label.

5. A device for realizing intelligent AI acquisition of webpage content based on LLM, characterized in that: The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the method for acquiring web page content by intelligent AI based on LLM as described in any one of claims 1 to 4.

6. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the method for realizing intelligent AI acquisition of web page content based on LLM as described in any one of claims 1 to 4 are implemented.

7. A computer program product, characterized in that The computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the method for acquiring web page content by intelligent AI based on LLM as described in any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Network searching method and device

    CN104679783A

  • Accurate searching method and device for mobile terminal

    CN111460307A

  • Webpage data capturing method and device, storage medium and electronic equipment

    CN118535785A