LLM question and answer method and device for HTML document

By parsing the metadata and structured tags of HTML documents, combined with the large language model LLM, the problems of information loss and low generation efficiency in HTML documents are solved, and more accurate and efficient answer generation is achieved.

CN120067397APending Publication Date: 2025-05-30BEIJING BANGCLE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311596924.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When processing HTML documents, it is difficult to efficiently generate comprehensive and accurate answer results, and there are problems such as loss of information and low generation efficiency.

Method used

By obtaining the target HTML document and user question text, parsing the HTML document to extract metadata and structured tags, using a large language model LLM to determine the target metadata that matches the user question text, and generating answer text based on that metadata.

Benefits of technology

By analyzing the structured features and metadata of HTML documents, we can more accurately locate and extract key information in the document, improve the accuracy and efficiency of information retrieval, and enhance the effectiveness of LLM's answers generated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067397A_ABST
    Figure CN120067397A_ABST
Patent Text Reader

Abstract

The invention discloses an LLM question and answer method and device for an HTML document. The LLM question and answer method and device are used for improving the effectiveness of generating answers to the HTML document by LLM. According to the scheme, the method comprises the steps of obtaining a target HTML document and a user question text of the target HTML document; analyzing the target HTML document to obtain metadata in the target HTML document and a structured tag of the metadata in the target HTML document; target metadata matched with the user question text is determined through a large language model LLM, the target metadata comprises first metadata matched with user question text semantics and second metadata associated with the first metadata, and a structured tag of the first metadata and a structured tag of the second metadata are associated based on an HTML structure; and generating an answer text corresponding to the user question text according to the content corresponding to the target metadata in the target HTML document.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language processing, and in particular, to an LLM question-answering method and apparatus for HTML documents. Background Art

[0002] In the field of natural language processing, large language models can be widely applied to various scenarios such as chatbots, text generation, and automatic summarization. In the question-answering scenario, large language models can generate answers with rich content to the questions raised by users.

[0003] In practical applications, an HTML (HyperText Markup Language) document is a structured document, and HTML documents are characterized by a large amount of text and complex structures. Affected by the complexity of HTML documents, the answers generated by large language models to questions are often relatively one-sided, with defects such as information loss and low generation efficiency, and it is difficult to efficiently generate comprehensive and accurate answer results based on long HTML documents.

[0004] How to improve the effectiveness of LLM in generating answers for HTML documents is the technical problem to be solved by this application. Summary of the Invention

[0005] The purpose of the embodiments of this application is to provide an LLM question-answering method and apparatus for HTML documents to improve the effectiveness of LLM in generating answers for HTML documents.

[0006] In the first aspect, an LLM question-answering method for HTML documents is provided, including:

[0007] Obtain a target HTML document and the user question text of the target HTML document;

[0008] Parse the target HTML document to obtain the metadata in the target HTML document and the structured tags of the metadata in the target HTML document;

[0009] Determine, through a large language model LLM, the target metadata that matches the user question text, where the target metadata includes the first metadata that semantically matches the user question text and the second metadata associated with the first metadata, and among them, the structured tags of the first metadata and the structured tags of the second metadata are associated based on the HTML structure;

[0010] Generate an answer text corresponding to the user question text according to the content corresponding to the target metadata in the target HTML document.

[0011] In the second aspect, an LLM question-answering apparatus for HTML documents is provided, including:

[0012] An acquisition module that acquires a target HTML document and the user question text of the target HTML document;

[0013] An analysis module that analyzes the target HTML document to obtain the metadata in the target HTML document and the structured tags of the metadata in the target HTML document;

[0014] A determination module that determines, through a large language model (LLM), target metadata that matches the user question text. The target metadata includes first metadata that semantically matches the user question text and second metadata associated with the first metadata. Among them, the structured tags of the first metadata and the structured tags of the second metadata are associated based on the HTML structure;

[0015] A generation module that generates an answer text corresponding to the user question text based on the content corresponding to the target metadata in the target HTML document.

[0016] In a third aspect, an electronic device is provided. The electronic device includes a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps of the method in the first aspect are implemented.

[0017] In a fourth aspect, a computer-readable storage medium is provided. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the steps of the method in the first aspect are implemented.

[0018] In the embodiments of the present application, by acquiring a target HTML document and the user question text of the target HTML document; analyzing the target HTML document to obtain the metadata in the target HTML document and the structured tags of the metadata in the target HTML document; determining, through a large language model (LLM), target metadata that matches the user question text. The target metadata includes first metadata that semantically matches the user question text and second metadata associated with the first metadata. Among them, the structured tags of the first metadata and the structured tags of the second metadata are associated based on the HTML structure; generating an answer text corresponding to the user question text based on the content corresponding to the target metadata in the target HTML document. This solution parses metadata and corresponding structured tags from HTML. When performing question matching, not only the first metadata that is closely related semantically is matched, but also the second metadata is associated based on the structured tags, which is beneficial to enabling the special structure information in HTML to be more comprehensively included in the answer content through the large language model (LLM), thereby achieving the purpose of improving the effectiveness of the answers generated by the LLM. Description of the Drawings

[0019] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0020] Figure 1a This is one of the flowcharts of an LLM question-answering method for HTML documents in one embodiment of the present application;

[0021] Figure 1b This is a second flow chart of an LLM question-answering method for HTML documents in an embodiment of the present application;

[0022] Figure 2 This is a flowchart of an LLM question-answering method for HTML documents in one embodiment of the present application;

[0023] Figure 3 This is a fourth flow chart of an LLM question-answering method for HTML documents in an embodiment of the present application;

[0024] Figure 4 This is a flowchart of an LLM question-answering method for HTML documents in one embodiment of the present application;

[0025] Figure 5 This is a sixth flow chart of an LLM question-answering method for HTML documents in an embodiment of the present application;

[0026] Figure 6 This is the seventh flow chart of an LLM question-answering method for HTML documents in one embodiment of the present application;

[0027] Figure 7 This is a flowchart of an LLM question-answering method for HTML documents in one embodiment of the present application;

[0028] Figure 8 It is a schematic diagram of the structure of an LLM question-and-answer device for HTML documents in one embodiment of the present application. DETAILED DESCRIPTION

[0029] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application. The numbering of the drawings in this application is only used to distinguish the various steps in the scheme, and is not used to limit the execution order of the various steps. The specific execution order is subject to the description in the specification.

[0030] In the field of computer science and natural language processing (NLP), large language models (LLMs) have gradually become the focus of research and application in recent years. Multiple GPT models have demonstrated excellent performance in various benchmark tests. They have made remarkable progress in understanding and generating natural language text, and have shown strong practical value and broad application prospects in a variety of application scenarios, such as chatbots, text generation, and automatic summarization.

[0031] As the scale and complexity of the models continue to grow to provide more accurate and diverse outputs, multi-modal learning that combines various data types such as text, images, and sounds is becoming a research hotspot, aiming to enable the models to understand and generate information more comprehensively.

[0032] At the application level, LLMs will play a greater role in fields such as personalized recommendations, content creation, and online customer service, further promoting the digital transformation of all industries.

[0033] Especially in the application of QA (Question Answering), LLMs have demonstrated remarkable capabilities. By understanding and analyzing the questions raised by users, LLMs can generate accurate, relevant, and information-rich answers, greatly improving the efficiency of information retrieval and knowledge sharing.

[0034] However, when dealing with long documents for question answering, LLMs face certain challenges because they usually have a maximum token processing limit. When the document length exceeds this limit, information may need to be truncated, reduced, or preprocessed in other ways to adapt to the capabilities of the model, which often reduces the quality and accuracy of the answers.

[0035] Although existing retrieval-enhanced LLMs perform well in question answering, they still face significant challenges when dealing with specific types of documents, such as structured documents like web pages and presentations. These documents usually contain rich structural information such as pages, tables, and various tags. If these documents are directly converted into plain text for processing, ignoring the structural information of the documents, it usually leads to information loss or inaccurate understanding.

[0036] In practical applications, businesses usually ask questions based on the structure of the document when querying these structured documents, such as questions about specific parts or tables. When the content of the document cannot fully fit the limited context length of the LLM, the problem of document question answering emerges.

[0037] There are some technical challenges and limitations in retrieval-augmented LLMs. First, the efficiency and accuracy of the information retrieval stage largely depend on the retrieval technology. If the retrieved information is inaccurate or incomplete, or if the IR (Information Retrieval) technology fails to precisely retrieve information relevant to the query, the answers generated by the LLM may not be accurate enough. Second, the information completeness of the text retrieved by information retrieval may be lost. In addition, complex retrieval systems increase the complexity of the system. Thus, in the Q&A application scenario, there are the following drawbacks:

[0038] a. Dependence on retrieval technology: The efficiency and accuracy of retrieval-augmented LLMs highly depend on the capabilities of the information retrieval technology used.

[0039] b. Increased system complexity: Introducing an information retrieval system increases the complexity of the system.

[0040] c. Failure to fully utilize the structured nature of HTML documents: When HTML documents cannot adapt to the context length of the LLM, the method of retrieving relevant context from the document and representing it as plain text ignores the structured natural structure of HTML documents, which does not conform to the user's mental model of these structurally rich documents and may lead to incoordination when the system queries the document context.

[0041] Thus, in the process of processing document Q&A, if directly retrieving relevant context from the document, although the content can also be converted into plain text, this method ignores the natural structure of web documents. Representing such structured documents as plain text does not conform to the user's mental model of these structurally rich documents. When the system queries the document context, this incoordination becomes particularly obvious, seriously hindering the performance of the Q&A system.

[0042] To solve the problems existing in the prior art, the embodiments of the present application provide an LLM Q&A method for HTML documents. This method can be applied to Q&A systems, text analysis systems, or other systems involving language processing, and can be specifically executed by computing devices such as servers and computers, as Figure 1a shown, including:

[0043] S11: Obtain a target HTML document and the user question text of the target HTML document.

[0044] Both the above-mentioned target HTML document and the user question text can be provided by the user. For example, the user can generate associated information from the link address of the target HTML document and the corresponding user question text, and by providing this associated information, both the target HTML document and the user question text can be provided simultaneously. In this step, the information provided by the user is parsed to obtain the target HTML document and the corresponding user question text.

[0045] In practical applications, there are many ways to associate the target HTML document with the user question text. In addition to the above way of the user actively providing associated information. The user can also type the user question text within the web page during the process of browsing an HTML document. In this case, the document of the web page can be obtained as the target HTML document, and the content typed by the user can be obtained as the user question text.

[0046] In addition, the above user question text can also be in the form of text obtained by converting through methods such as image character recognition and speech character recognition.

[0047] See Figure 1b In this step, it includes Figure 1b the first step in

[0048] S12: Parse the target HTML document to obtain the metadata in the target HTML document and the structured tags of the metadata in the target HTML document.

[0049] This step can be specifically implemented through HTML Triage. This HTML Triage extracts key structured tags, such as titles, subtitles, lists, tables, etc., by parsing the tags and attributes in the HTML document. These structured tags are used to identify the structure of the corresponding metadata in the target HTML document.

[0050] The metadata in this solution refers to HTML metadata, which is a kind of data used to describe an HTML document, and this metadata can be called through services.

[0051] See Figure 1b In this step, it includes Figure 1b the second step and the third step in

[0052] S13: Determine the target metadata that matches the user question text through the large language model LLM. The target metadata includes the first metadata that semantically matches the user question text and the second metadata associated with the first metadata, where the structured tags of the first metadata and the structured tags of the second metadata are associated based on the HTML structure.

[0053] In this step, through HTML Triage based on natural language processing technology, the user question text is parsed to obtain the actual question raised by the user, and it is matched with the content corresponding to each metadata in the above steps to determine the target metadata. The content corresponding to the target metadata is the content with the highest relevance to the user question text.

[0054] Specifically, metadata matching can be achieved in various ways. For example, first perform matching on the content corresponding to each metadata and the user question text through text semantics, and combine other analysis methods such as keywords to improve the matching accuracy, and determine the first metadata that matches the user question text. Then, according to the structured tags of the first metadata, determine the second metadata associated with the first metadata. Among them, the association relationships between metadata are diverse. For example, the first metadata corresponds to the content in a table, and the second metadata corresponds to the associated annotation of the table. Another example is that the first metadata corresponds to a subtitle, and the second metadata corresponds to the content linked through a link in the subtitle.

[0055] In this step, through the large language model LLM, the target metadata matching the user question text is determined. The target metadata includes the content corresponding to the first metadata matched based on text semantics and the associated second metadata, and can perform matching according to structured tags during the process of matching data, obtaining more complete matching content and avoiding the problem of missing matching content due to the lack of web page structure.

[0056] Optionally, the user question text can also be optimized by performing matching according to a preset question library. Specifically, through methods such as keyword matching and semantic parsing, the user question text is matched and retrieved with multiple preset question texts in the preset question library. When the matching degree is higher than the preset matching degree, the matched preset question text is used to correct or replace the user question text. The method of matching based on the preset question library can reduce the overall processing volume of the system. When a preset question can be matched, the preset question is directly used to avoid problems such as inaccurate understanding of the user question text and inaccurate generation of answers due to expression defects such as typos and ungrammatical sentences entered by the user.

[0057] See Figure 1b , this step includes Figure 1b the fourth step in

[0058] S14: Generate the answer text corresponding to the user question text according to the content corresponding to the target metadata in the target HTML document.

[0059] In this step, an answer text is generated based on the content corresponding to the target metadata. Among them, the main answer text can be generated first according to the first metadata, and then the main answer text can be supplemented and corrected according to the second metadata to generate a complete answer text.

[0060] See Figure 1b , this step includes Figure 1b the fifth step in

[0061] Through the solution provided by the embodiments of the present application, the metadata of the HTML document and the corresponding structured tags are deeply analyzed. When performing question matching, not only the first metadata that is closely related is matched according to semantics, but also the second metadata is associated according to the structured tags, which is beneficial to more comprehensively include the special structure information in the HTML in the answer content, so as to achieve the purpose of improving the effectiveness of the answer. This solution can be applied to the application scenario of LLM question and answer for HTML page content. This solution makes full use of the structured characteristics of the HTML page to construct the basic information of HTML metadata. When performing LLM question and answer for HTML page content, relevant text content is obtained through HTML metadata and input into the LLM, effectively improving the effectiveness of the answer generated by the LLM and being beneficial to improving the business usability.

[0062] Based on the solution provided by the above embodiments, optionally, as Figure 2 shown, before the above step S13, that is, before determining the target metadata matching the user question text, it further includes:

[0063] S21: Generate a data dictionary according to the structured tags corresponding to the metadata in the target HTML document, where the data dictionary includes the correspondence between the metadata and the structured tags.

[0064] In practical applications, the element content of each metadata in the target HTML document can be traversed to construct a data dictionary. Among them, it is judged whether the element of the metadata has text content. If it exists, the text content is extracted and associated with the element of the metadata to which it belongs to construct a data dictionary.

[0065] Furthermore, it can also be recognized whether the element of the metadata contains various other forms of content such as pictures, sounds, videos, etc. If there is content, the content expressed in other forms is converted into text content through methods such as image recognition and speech recognition, and then associated with the element of the metadata to which it belongs to construct a data dictionary.

[0066] In addition, if the element of the metadata does not contain text content and no text content is obtained through extraction and conversion, it can be marked that there is no text content in the element of the metadata.

[0067] In a data dictionary, associated data can be stored in the form of multiple data items. Such data items may include, for example, elements of metadata and corresponding text content.

[0068] Among them, in the above step S13, determining the target metadata that matches the user question text through the large language model LLM includes:

[0069] S22: Perform a retrieval on the data dictionary to determine the target metadata that matches the user question text through the large language model LLM.

[0070] In this step, matching retrieval can be performed by means of keywords, semantics, etc., and the metadata in the data item with the highest matching degree in the data dictionary is determined as the target metadata.

[0071] In the example of this application, by constructing a data dictionary, the content in the target HTML document can be structurally transformed, thereby improving the speed and accuracy of subsequent associated queries of the large language model LLM. In practical applications, after constructing a data dictionary for a certain target HTML document, this data dictionary can be used to perform matching retrievals on multiple different questions of multiple different users, especially suitable for application scenarios with multiple users and multiple questions and answers.

[0072] Based on the solution provided in the above embodiment, optionally, as Figure 3 shown, in the above step S22, performing a retrieval on the data dictionary to determine the target metadata that matches the user question text includes:

[0073] S31: Retrieve the first metadata that matches the user question text in the data dictionary based on semantic analysis.

[0074] In this step, matching retrieval is performed based on the data dictionary of the target HTML document, which is implemented by means of semantic analysis.

[0075] Optionally, first perform semantic analysis on the user question text to express the user question text through semantic features, and the semantic features can be in vector form. Subsequently, match the semantic features of the user question with the semantic features of the text content in each data item in the data dictionary to determine the first metadata with the highest matching degree.

[0076] Optionally, extract the keywords corresponding to the user question text through semantic analysis, and then match the text content of each data item in the data dictionary by means of keyword matching to determine the first metadata with the highest matching degree.

[0077] Alternatively, matching can also be performed in the form of combining multiple semantic analyses to comprehensively determine the first metadata that matches the user question text.

[0078] S32: Retrieve, in the data dictionary, a second structured tag associated with the first structured tag based on the HTML structure according to the first structured tag corresponding to the first metadata.

[0079] After determining the first metadata, query for the second structured tag in the data dictionary based on the first structured tag. There can be multiple second structured tags. For example, if the first structured tag is a major title, the second structured tag can be a subtitle under that major title. If the first structured tag is the content under a subtitle, the second structured tag can be other content belonging to the same subtitle or a note for the content in that subtitle, etc.

[0080] S33: Determine the second metadata corresponding to the second structured tag.

[0081] After querying for the second structured tag associated with the first structured tag in the above step, the content of the second metadata corresponding to the second structured tag can be queried through the data dictionary. If there are multiple second structured tags, the second metadata corresponding to each second structured tag is obtained separately. To generate the content corresponding to the user question text more accurately based on the HTML structure, the association relationship between the content of the first metadata and the first structured tag and the association relationship between the content of the second metadata and the second structured tag can be stored, which is conducive to generating more accurate answer content according to the HTML structure corresponding to the structured tag.

[0082] Based on the solution provided in the above embodiments, optionally, as Figure 4 shown, in the above step S12, parsing the target HTML document to obtain the metadata in the target HTML document and the structured tag of the metadata in the target HTML document includes:

[0083] S41: Obtain the tree structure of the target HTML document.

[0084] In this step, the tree structure can be obtained through the link of the target HTML document. In practical applications, libraries such as Python's lxml or BeautifulSoup can be used to parse the HTML document to obtain the tree structure of the target HTML document.

[0085] S42: Recursively extract HTML elements based on the tree structure to obtain the metadata in the target HTML document and the structured tag of the metadata in the target HTML document.

[0086] Specifically, the extract_structure function can be used to recursively extract the structured information of an HTML document. This function can handle the nested structure of HTML elements and convert it into a nested dictionary for subsequent analysis and processing.

[0087] Based on the solution provided in the above embodiments, optionally, as Figure 5 shown, in the above step S42, based on the tree structure, recursively extract HTML elements to obtain the metadata in the target HTML document and the structured tags of the metadata in the target HTML document, including:

[0088] S51: Extract the text content in the target HTML element based on the tree structure.

[0089] The tree structure can express the association relationship between the structures in the target HTML text. For example, "head" and "body" are included in HTML. In this step, based on the tree structure, the content extraction of HTML elements is performed, and each structure can be comprehensively traversed according to the structure of HTML, and the content in the target HTML text can be fully extracted.

[0090] S52: Construct a corresponding relationship between the text content of the target HTML element and the structured tags of the target HTML element.

[0091] In this step, a corresponding relationship is constructed between the text content corresponding to the same HTML element and the structured tags. For example, it can be expressed in the form of {structured tags of the element: text content}. The text content and structured tags with a corresponding relationship constructed in this step can be used to implement the construction of a data dictionary.

[0092] Based on the solution provided in the above embodiments, optionally, as Figure 6 shown, after the above step S52, that is, after constructing the corresponding relationship between the text content of the target HTML element and the structured tags of the target HTML element, it further includes:

[0093] S61: If the target HTML element has HTML sub-elements, determine the sub-structured tags corresponding to the HTML sub-elements;

[0094] S62: If there is text content in the HTML sub-element, construct a corresponding relationship between the text content of the HTML sub-element and the sub-structured tags of the HTML sub-element.

[0095] Next, a specific example of pseudo-code is used to illustrate this solution for implementing dictionary construction.

[0096] The pseudo-code is as follows:

[0097] Define the function extract_structure(element):

[0098] # If the tag type of the element is not a string, return directly

[0099] If the tag type of the element is not a string:

[0100] Return

[0101] # Extract the text content of the current element (if it exists)

[0102] Current text = the text content of the element (after removing leading and trailing whitespace) if the element has text content, otherwise None

[0103] If the tag of the element is "body" or "strong":

[0104] Current text = ""

[0105] # If the element has no child tags

[0106] If the element has no child tags:

[0107] # If the current text exists, return a dictionary containing the tag and text content of the element

[0108] # Otherwise, only return the tag of the element

[0109] Return {the tag of the element: current text} if the current text exists, otherwise the tag of the element

[0110] # If the element has child tags

[0111] # Recursively call the extract_structure function to extract the structure of each child element

[0112] Sub - structure = [extract_structure(sub - element) for each sub - element of the element]

[0113] # If the current element also has text content

[0114] If the current text exists:

[0115] # Add the current text to the front of the sub - structure

[0116] Sub - structure = [current text]+sub - structure

[0117] # Return a dictionary containing the tag and sub - structure of the element

[0118] Return {the tag of the element: sub - structure}

[0119] In the example of this application, first, the tag type of the element is judged to identify the element of string type. Then, for the element containing text content, its text content is extracted. If the element has sub-tags, the content of the sub-elements is further recursively extracted until all the content of the sub-tags is extracted, generating the dictionary content containing the element tag and text. In addition, for the element with a sub-structure, the dictionary content containing the element tag and sub-structure is generated to express the relationship between the element structures.

[0120] Based on the solution provided in the above embodiment, optionally, as Figure 7 shown, in the above step S11, obtaining the target HTML document and the user question text of the target HTML document includes:

[0121] S71: Obtain the target URL corresponding to the user question text;

[0122] S72: Start a headless browser and instruct the headless browser to open a new page;

[0123] S73: Navigate the new page to the target URL;

[0124] S74: After the page rendering of the target URL is completed, obtain the page content of the target URL as the target HTML document.

[0125] Next, a solution will be described in combination with an example of pseudo-code for obtaining the target HTML document and the user question text. This solution can use asynchronous programming based on Python and browser automation tools (such as pyppeteer), and define an asynchronous function get_page_content_after_js. The pseudo-code is as follows:

[0126] Function get_page_content_after_js(input: url):

[0127] Try:

[0128] # Start a headless browser

[0129] Browser = Asynchronously start a headless browser()

[0130] # Open a new page

[0131] Page = Asynchronously open a new page in the browser()

[0132] # Asynchronously navigate to the specified URL and wait for JavaScript rendering to complete

[0133] Asynchronously page navigate to(url, wait for network idle)

[0134] # Get and return the rendered page content

[0135] content = AsynchronouslyGetPageContent()

[0136] Capture any exceptions as exception:

[0137] # Print the exception information

[0138] Output("An error occurred: " + str(exception))

[0139] # Set the content to None

[0140] content = None

[0141] Finally:

[0142] # Close the browser

[0143] AsynchronouslyCloseBrowser()

[0144] # Return the obtained content

[0145] Return content

[0146] Based on the pseudocode shown in the embodiments of the present application, the function in this solution implements the following functions:

[0147] Asynchronously start a headless browser instance; asynchronously access a specified URL and wait for the execution of JavaScript and the network to be idle; get and return the rendered page content.

[0148] The solution provided by the embodiments of the present application realizes efficient question answering of HTML long documents using LLM by parsing the structural features of HTML documents and constructing a metadata structure of the document based on HTML tags, which can effectively improve the accuracy and efficiency of information retrieval.

[0149] In terms of accuracy, this solution can more accurately locate and extract key information and structures in the document by precisely parsing the HTML structure and using metadata, thereby improving the accuracy of information retrieval.

[0150] In terms of efficiency, by using the structured metadata, it is possible to quickly locate the key parts of the document, avoiding full-text search of the entire document, thereby significantly improving the efficiency of information retrieval.

[0151] In addition, this solution effectively improves the efficiency of business analysis of HTML documents by introducing an LLM to analyze HTML structured documents. Through HTML document parsing and information extraction technologies, key information and structures in the document can be accurately identified and extracted. Moreover, this solution can achieve in-depth question understanding and matching, using natural language processing technologies to deeply understand user questions and precisely match them with the data extracted from HTML documents.

[0152] To solve the problems existing in the prior art, such as Figure 8 As shown, an LLM question-and-answer device 80 for HTML documents is further provided in an embodiment of the present application, including:

[0153] An acquisition module 81 that acquires a target HTML document and the user question text of the target HTML document;

[0154] An analysis module 82 that analyzes the target HTML document to obtain metadata in the target HTML document and the structured tags of the metadata in the target HTML document;

[0155] A determination module 83 that determines target metadata matching the user question text through a large language model LLM, where the target metadata includes first metadata semantically matching the user question text and second metadata associated with the first metadata, and the structured tags of the first metadata and the structured tags of the second metadata are associated based on the HTML structure;

[0156] A generation module 84 that generates an answer text corresponding to the user question text according to the content corresponding to the target metadata in the target HTML document.

[0157] The device provided in the embodiment of the present application acquires a target HTML document and the user question text of the target HTML document; analyzes the target HTML document to obtain metadata in the target HTML document and the structured tags of the metadata in the target HTML document; determines target metadata matching the user question text through a large language model LLM, where the target metadata includes first metadata semantically matching the user question text and second metadata associated with the first metadata, and the structured tags of the first metadata and the structured tags of the second metadata are associated based on the HTML structure; generates an answer text corresponding to the user question text according to the content corresponding to the target metadata in the target HTML document. This solution parses metadata and corresponding structured tags from HTML. When performing question matching, not only the closely related first metadata is matched according to semantics, but also the second metadata is associated according to the structured tags, which is beneficial to more comprehensively including the special structure information in HTML in the answer content through the large language model LLM, thereby achieving the purpose of improving the effectiveness of the answers generated by the LLM.

[0158] Among them, the above modules in the device provided by the embodiments of the present application can also implement the method steps provided by the above method embodiments. Alternatively, the device provided by the embodiments of the present application may further include other modules in addition to the above modules to implement the method steps provided by the above method embodiments. And the device provided by the embodiments of the present application can achieve the technical effects that the above method embodiments can achieve.

[0159] Preferably, the embodiments of the present application further provide an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements each process of the above embodiments of the method for LLM question and answer for HTML documents and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0160] The embodiments of the present application further provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the above embodiments of the method for LLM question and answer for HTML documents and can achieve the same technical effects. To avoid repetition, it will not be elaborated here. Among them, the computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0161] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0162] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one or more processes in the flowchart and / or one or more blocks in the block diagram.

[0163] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the function specified in one or more blocks of a flowchart and / or one or more blocks of a block diagram.

[0164] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing steps for implementing the function specified in one or more blocks of a flowchart and / or one or more blocks of a block diagram by the instructions executed on the computer or other programmable apparatus.

[0165] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0166] Memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.

[0167] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storage of information such as computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0168] It should also be noted that the term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising said element.

[0169] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, system or computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0170] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. An LLM Q&A method for HTML documents, characterized in that, it includes: Obtain a target HTML document and the user question text of the target HTML document; Parse the target HTML document to obtain the metadata in the target HTML document and the structured tags of the metadata in the target HTML document; Determine, through a large language model (LLM), target metadata that matches the user question text, where the target metadata includes first metadata that semantically matches the user question text and second metadata associated with the first metadata, and where the structured tags of the first metadata and the structured tags of the second metadata are associated based on the HTML structure; Generate an answer text corresponding to the user question text based on the content corresponding to the target metadata in the target HTML document.

2. The method according to claim 1, characterized in that, before determining, through a large language model (LLM), target metadata that matches the user question text, it further includes: Generate a data dictionary based on the structured tags corresponding to the metadata in the target HTML document, where the data dictionary includes the correspondence between the metadata and the structured tags; where determining, through a large language model (LLM), target metadata that matches the user question text includes: Perform a retrieval on the data dictionary and determine, through a large language model (LLM), target metadata that matches the user question text.

3. The method according to claim 2, characterized in that, performing a retrieval on the data dictionary and determining, through a large language model (LLM), target metadata that matches the user question text includes: Retrieve, based on semantic analysis, first metadata in the data dictionary that matches the user question text; According to the first structured tag corresponding to the first metadata, retrieve, in the data dictionary, a second structured tag that is associated with the first structured tag based on the HTML structure; Determine the second metadata corresponding to the second structured tag.

4. The method according to any one of claims 1 to 3, characterized in that, parsing the target HTML document to obtain the metadata in the target HTML document and the structured tags of the metadata in the target HTML document includes: Obtain the tree structure of the target HTML document; Recursively extract HTML elements based on the tree structure to obtain the metadata in the target HTML document and the structured tags of the metadata in the target HTML document.

5. The method according to claim 4, characterized in that, recursively extracting HTML elements based on the tree structure to obtain the metadata in the target HTML document and the structured tags of the metadata in the target HTML document includes: Extract the text content in the target HTML element based on the tree structure; Construct a correspondence between the text content of the target HTML element and the structured tag of the target HTML element.

6. The method according to claim 5, characterized in that, After establishing the correspondence between the text content of the target HTML element and the structured tags of the target HTML element, the following steps are further included: If the target HTML element has HTML child elements, determine the corresponding sub-structured tags of the HTML child elements; If there is text content in the HTML child elements, establish the correspondence between the text content of the HTML child elements and the sub-structured tags of the HTML child elements.

7. The method according to claim 1, wherein, obtaining a target HTML document and the user question text of the target HTML document includes: obtaining the target URL corresponding to the user question text; starting a headless browser and instructing the headless browser to open a new page; navigating the new page to the target URL; after the page rendering of the target URL is completed, obtaining the page content of the target URL as the target HTML document.

8. An LLM question-answering device for an HTML document, wherein, it includes: an acquisition module that acquires a target HTML document and the user question text of the target HTML document; a parsing module that parses the target HTML document to obtain the metadata in the target HTML document and the structured tags of the metadata in the target HTML document; a determination module that determines, through a large language model (LLM), the target metadata that matches the user question text, where the target metadata includes first metadata that semantically matches the user question text and second metadata associated with the first metadata, and wherein the structured tags of the first metadata and the structured tags of the second metadata are associated based on the HTML structure; a generation module that generates an answer text corresponding to the user question text according to the content corresponding to the target metadata in the target HTML document.

9. An electronic device, wherein, it includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, wherein, a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.