Enhanced question answering method and system based on document slicing and marking
By adopting enhanced question-and-answer methods based on document slicing and marking in the document analysis and question-and-answer system, the existing system's inefficiency and insufficient accuracy when processing large and complex documents is solved, efficient document processing and question-and-answer feedback are achieved, and user experience is improved.
Patent Information
- Application Number
- CN202510109953.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-23
AI Technical Summary
The existing document analysis and question-and-answer systems are inefficient in processing large and complex documents, lack accuracy, and it is difficult to understand the complex structure and content within the document, and the lack of an effective caching mechanism leads to a decrease in response speed.
Using enhanced Q&A methods based on document slicing and marking, through multiple functions such as document upload, intelligent segmentation, content summary, QA generalization and user Q&A processing, documents in HTML, DOC, DOCX and PDF formats can be parsed, forming an end-to-end process, and improving the efficiency and accuracy of document processing and Q&A feedback.
It realizes automated in-depth analysis and slice management of documents, can identify and distinguish complex elements such as tables and pictures in the document, provide efficient content summary extraction and QA generalization process, and improves the response speed and user experience of the user's Q&A system.
Smart Images

Figure CN120030123A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information processing and artificial intelligence technology, and in particular to an enhanced question-answering method and system based on document slicing and marking. Background Art
[0002] In current document parsing and question-answering systems, most methods focus on keyword searches or simple text matching of the entire document, and lack the ability to deeply understand the complex structure and content within the document. These methods are often prone to inefficiency and lack of accuracy when dealing with large and complex documents. Specifically, existing systems are usually unable to effectively distinguish between different types of information in documents, such as text, tables, images, etc., and it is also difficult to automatically identify and classify the value of this information for question answering. In addition, when dealing with a large number of concurrent requests, the question-answering system often suffers from a slow response speed and poor user experience due to the lack of an effective caching mechanism. Summary of the invention
[0003] In order to solve the above technical problems, the present invention provides an enhanced question-and-answer method and system based on document slicing and labeling, which integrates multiple functions such as document uploading, intelligent segmentation, content summary, QA generalization, user question-and-answer processing, and can parse HTML format, DOC format, DOCX format or PDF format, forming a complete end-to-end process, greatly improving the efficiency and accuracy of document processing and feedback on user questions and answers.
[0004] In a first aspect, the present invention provides an enhanced question-answering method based on document slicing and tagging, which adopts the following technical solution:
[0005] An enhanced question answering method based on document slicing and tagging includes the following steps:
[0006] S1. Receive the initial document uploaded by the user and detect the format of the initial document. If it is in HTML format, perform HTML parsing. If it is in DOC format, DOCX format or PDF format, perform PDF parsing.
[0007] HTML parsing: parse the initial document in HTML format into a DOM tree;
[0008] PDF parsing: convert the initial document into a standard PDF document in a unified PDF format, and then perform PDF parsing on the standard PDF document. PDF parsing is to identify the standard PDF document and organize it into PDF parsed text;
[0009] Secondarily segment the DOM tree or PDF parsed text and archive the slices, perform hierarchical summary extraction and QA generalization on the content of the PDF parsed text, and store them in the Milvus vector library;
[0010] Among them, hierarchical summary extraction includes extracting multi-level sub-content summaries until the entire summary of the entire document is generated;
[0011] S2. Receive the user's query content, and retrieve the top five data with the highest relevance in the Milvus vector library according to the user's query content. The top five data include text information or tables or pictures. If the top five data include tables or pictures, mark the tables or pictures and temporarily store them in the cache.
[0012] After receiving the TOP5 data, the big model combines its powerful semantic understanding and reasoning to generate corresponding text answers. If a table or picture is referenced in the process of generating the corresponding answer, the big model will generate an identifier corresponding to the table or picture, and then take out the corresponding table and picture from the cache based on the identifier corresponding to the table or picture and fill it into the corresponding text answer, and generate a graphic answer.
[0013] Preferably, the large model generates an identification corresponding to the table or picture including a link or identifier.
[0014] Preferably, HTML parsing includes the following steps:
[0015] Get child nodes: Starting from the current element, get a list of all child nodes;
[0016] Traverse child nodes: traverse each child node, detect the child node type, and then perform corresponding processing according to the child node type. The child node type includes text node or element node;
[0017] If the child node type is a text node, create a storage object to store the text node. Since the text node has no HTML tag, set its tag to an empty string and set the content of the text node to the storage object.
[0018] If the child node type is an element node, it is processed according to the tag name of the element node, including the following categories:
[0019] If the tag name is title, a function is recursively called to extract the text content of the element and all its child elements, then an HtmlElement object is created to store the title text, and the title text is added to the node list;
[0020] If the tag name is table, the HTML content of the table is converted to Markdown format. The corresponding start and end tags for distinction are added before and after the conversion result. The formatted string is encapsulated in an HtmlElement object and added to the node list.
[0021] If the tag name is picture: If the document comes from a ZIP file and the picture resource is available, extract the picture and construct the picture text including the picture path, then create an HtmlElement object to store the picture URL and add it to the node list;
[0022] If the tag name is other elements: other elements are non-title, non-table and non-picture elements, check whether other elements need to be filtered out. If not, recursively call the function of processing element nodes to process the other elements. For p and br tags, create HtmlElement objects including line breaks to simulate the conversion of other elements in HTML and add them to the node list;
[0023] Recursively build and generate a DOM tree based on HTML parsing.
[0024] Preferably, HTML parsing further includes a filtering mechanism to store filtered elements in a filter set.
[0025] Preferably, the DOM tree includes a starting heading level, an HtmlElement list including all HTML elements, a root node of a text object tree, and a parent text object currently being processed.
[0026] In a second aspect, the present invention provides an enhanced question-answering system based on document slicing and marking, which adopts the following technical solution:
[0027] An enhanced question-answering system based on document slicing and tagging, comprising:
[0028] A receiving and detecting unit, used for receiving an initial document uploaded by a user and detecting the format of the initial document. If it is in HTML format, the HTML parsing unit is used; if it is in DOC format, DOCX format or PDF format, the PDF parsing unit is used;
[0029] An HTML parsing unit, used to parse an initial document in HTML format into a DOM tree;
[0030] PDF parsing unit: used for converting the initial document into a standard PDF document in a unified PDF format, and then performing PDF parsing on the standard PDF document, wherein the PDF parsing is to identify the standard PDF document and organize it into PDF parsed text;
[0031] The secondary processing module is used to perform secondary segmentation on the DOM tree or PDF parsed text and archive the slices, as well as perform hierarchical summary extraction and QA generalization on the content of the PDF parsed text and store it in the Milvus vector library;
[0032] A user query receiving module, used to receive user query content;
[0033] The retrieval module is used to retrieve the top five data with the highest relevance in the Milvus vector library according to the user's query content. The top five data include text information, tables or pictures.
[0034] Whether to include tables or pictures: a module for judging whether the TOP5 data includes tables or pictures. If so, the tables or pictures are marked and temporarily stored in the cache.
[0035] The large model module is used to receive the TOP5 data and generate corresponding text answers based on its powerful semantic understanding and reasoning;
[0036] The table or picture reference module is used to detect whether the text answer generated by the large model module references a table or picture. If so, it enters the table or picture supplement module;
[0037] The table or picture supplement module is used to call the large model module to generate labels corresponding to the table or picture, and then take out the corresponding table and picture from the cache based on the labels corresponding to the table or picture to fill in the corresponding text answer, and generate a graphic answer.
[0038] Preferably, the HTML parsing unit includes:
[0039] The child node acquisition module is used to obtain child nodes and obtain a list of all child nodes starting from the current element;
[0040] The child node traversal detection module is used to traverse each child node and detect the child node type, and then perform corresponding processing according to the child node type. The child node type includes a text node or an element node. If it is a text node, it enters the text node module, and if it is an element node, it enters the element node module;
[0041] The text node module is used to create a storage object to store the text node. Since the text node has no HTML tag, its tag is set to an empty string, and the content of the text node is set to the storage object;
[0042] Element nodes are used to process according to the tag names of element nodes. Title names include titles, tables, pictures and other elements. Other elements are non-title, non-table and non-picture elements. If it is a title, it enters the title module. If it is a table, it enters the table module. If it is a picture, it enters the picture module. If it is other elements, it enters the other element module.
[0043] The title module is used to recursively call a function, which is used to extract the text content of an element and all its child elements, then create an HtmlElement object to store the title text, and add the title text to the node list;
[0044] Table module converts the HTML content of the table into Markdown format. It adds corresponding start and end tags before and after the conversion result. The formatted string is encapsulated in an HtmlElement object and added to the node list.
[0045] The image module is used to determine whether the document comes from a ZIP file. If so and the image resource is available, it extracts the image and constructs an image text including the image path, then creates an HtmlElement object to store the image URL and adds it to the node list;
[0046] Other element module, used to check whether other elements need to be filtered out. If not, the function of processing element nodes is recursively called to process the other elements. For p and br tags, HtmlElement objects including line breaks are created to simulate the conversion of other elements in HTML and added to the node list.
[0047] DOM tree generation module, used to recursively build and generate DOM tree based on HTML parsing.
[0048] Preferably, the HTML parsing unit further includes a filtering mechanism module for storing filtered elements in a filtering set.
[0049] In a third aspect, the present application provides an electronic device, which adopts the following technical solution:
[0050] An electronic device, comprising:
[0051] one or more processors;
[0052] Memory;
[0053] One or more applications, wherein one or more applications are stored in a memory and configured to be executed by one or more processors, and the one or more programs are configured to: execute an enhanced question-answering method based on document slicing and tagging as shown in any possible implementation of the first aspect.
[0054] In a fourth aspect, the present application provides a computer-readable storage medium, which adopts the following technical solution:
[0055] A computer-readable storage medium, comprising: a computer program that can be loaded and executed by a processor to implement an enhanced question-answering method based on document slicing and tagging as shown in any possible implementation of the first aspect.
[0056] In summary, the present invention includes the following beneficial technical effects:
[0057] 1. The present invention integrates multiple functions such as document uploading, intelligent segmentation, content summarization, QA generalization, user question and answer processing, and can parse HTML format, DOC format, DOCX format or PDF format, forming a complete end-to-end process, which greatly improves the efficiency and accuracy of document processing and feedback on user questions and answers.
[0058] 2. The present invention realizes the automatic deep analysis and slice management of documents. It can not only identify the text content, but also accurately identify and distinguish the complex elements such as tables, pictures, lists, etc. in the document, and divide the document into multiple slices that are closely related logically and relatively independent in content through intelligent algorithms. Each slice contains a complete and coherent information unit, which is convenient for subsequent processing.
[0059] 3. The present invention provides an efficient content summary extraction method, which recursively constructs a DOM tree based on the hierarchical relationship of titles, ensures that each title can be correctly placed at its proper hierarchical position, and forms a clear and hierarchical document structure representation. The content summary extraction method can deeply understand the document content, automatically extract key information, and generate an accurate, concise DOM tree that maintains the original intention. The content summary extraction method takes into account the overall structure of the document, the logical relationship between paragraphs, and the importance of keywords, so that the extracted summary is both comprehensive and concise.
[0060] 4. The present invention constructs an intelligent QA generalization process, which can automatically extract problem points from document slices and generate a series of inspiring questions that are closely related to the document content. These inspiring questions can cover the key knowledge points, difficult points and possible doubts of the document, providing a high-quality corpus basis for subsequent manual question and answer or automatic question and answer systems.
[0061] 5. The present invention targets user queries by efficiently utilizing advanced search engines such as Milvus to quickly lock in the most relevant diversified data, such as text, tables, images, etc., and intelligently preprocesses and caches non-text elements to ensure the continuity and efficiency of data processing. It also utilizes the powerful semantic understanding and reasoning capabilities of high-performance large models to generate accurate and detailed text answers. And when non-text elements are referenced in the text answer, a graphic answer with rich text and information is generated by generating a unique identifier and automatically backfilling the filling process, which greatly improves the user experience and the practicality of the user question and answer system.
[0062] 6. The present invention adopts a recursive method to traverse HTML document nodes and performs different processing according to the node type. In particular, there is a special processing logic for key elements such as titles, tables, and pictures to ensure accurate parsing and extraction of the document structure. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 It is a flow chart of an enhanced question-answering method based on document slicing and tagging in an embodiment of the present invention.
[0064] Figure 2 It is a specific example of "secondary segmentation of the DOM tree or PDF parsed text and archiving of the slices, and hierarchical summary extraction and QA generalization of the content of the PDF parsed text" in the enhanced question-answering method based on document slicing and tagging in an embodiment of the present invention. DETAILED DESCRIPTION
[0065] The present invention is further described in detail below in conjunction with the accompanying drawings.
[0066] The embodiment of the present invention discloses an enhanced question-answering method based on document slicing and marking.
[0067] Reference Figure 1 , the enhanced question answering method based on document slicing and tagging includes the following steps:
[0068] S1. Receive the initial document uploaded by the user and detect the format of the initial document. If it is in HTML format, perform HTML parsing. If it is in DOC format, DOCX format or PDF format, perform PDF parsing.
[0069] Specifically, the user uploads an initial document, receives the initial document uploaded by the user and saves it, for example, to nfs, the initial document includes the following formats: HTML format, DOC format, DOCX format or PDF format, reads the saved initial document and detects the format of the initial document, if it is in HTML format, performs HTML parsing, if it is in DOC format, DOCX format or PDF format, performs PDF parsing;
[0070] HTML parsing: Parse the initial document in HTML format into a DOM tree, and filter it in advance for user-specified tags;
[0071] Specifically, HTML parsing includes the following steps:
[0072] Get child nodes: Starting from the current element, get a list of all child nodes;
[0073] Traverse child nodes: traverse each child node, detect the child node type, and then perform corresponding processing based on the child node type. The child node type includes text node or element node;
[0074] If the child node type is a text node, create a storage object to store the text node. Since the text node has no HTML tag, set its tag to an empty string and set the content of the text node to the storage object.
[0075] If the child node type is an element node, it is processed according to the tag name of the element node, including the following categories:
[0076] If the tag name is title, a function is recursively called to extract the text content of the element and all its child elements, then an HtmlElement object is created to store the title text, and the title text is added to the node list;
[0077] If the tag name is table, the HTML content of the table is converted to Markdown format. The corresponding start and end tags for distinction are added before and after the conversion result. The formatted string is encapsulated in an HtmlElement object and added to the node list.
[0078] If the tag name is picture: If the document comes from a ZIP file and the picture resource is available, extract the picture and construct the picture text including the picture path, then create an HtmlElement object to store the picture URL and add it to the node list;
[0079] If the tag name is other elements: other elements are non-title, non-table and non-picture elements, check whether other elements need to be filtered out. If not, recursively call the function of processing element nodes to process the other elements. For p and br tags, create HtmlElement objects including line breaks to simulate the conversion of other elements in HTML and add them to the node list;
[0080] HTML parsing also includes a filtering mechanism that stores filtered elements in a filter collection.
[0081] Recursively build and generate a DOM tree based on HTML parsing.
[0082] In the above collated HTML node list, an automated process is used to detect and process all heading elements from h1 to h7. This process recursively builds a DOM tree based on the hierarchical relationship of the headings to ensure that each heading is correctly placed in its proper hierarchical position. This process is implemented through a function that receives four key parameters: the starting heading level, the HtmlElement list containing all HTML elements, the root node of the text object tree, and the parent text object currently being processed.
[0083] Inside the function, a loop is started, going through all heading levels from the specified start level to h7. For each level in the loop, the function checks if a heading of that level exists in the list of HTML elements. If a heading of the specified level is detected, an array of sublists is generated using a function that can split the list of HTML elements according to these headings. Each sublist in the array of sublists contains a heading element and all the content after the heading, thus clearly dividing the structure of the document.
[0084] After the split is completed, in order to avoid repeated operations on elements that have already been processed, the function will clear the original HTML element list.
[0085] Then, the function traverses these sublists and checks each sublist to confirm whether it contains the title of the current level. If the sublist does contain the title of the current level, the function will extract the title and the content behind the title, create a new TextObj object, and configure the corresponding properties for the TextObj object, such as title text, associated content, etc. Then the newly created TextObj object will be added to the sublist of its parent object, thus establishing a hierarchical relationship.
[0086] If the newly created TextObj object has a valid title, the function will further recursively call itself, using the content after the current title as a new HTML element list, and setting this TextObj object as the new parent object. The above process is repeated to deeply explore and process lower-level titles until a complete nested title structure is constructed.
[0087] If the current level of heading is not found in a sublist, the function will intelligently increase the heading level and recursively try to find the next level of heading to ensure that no potential nested heading structure is missed. In this way, the heading elements in the entire HTML element list can be effectively organized into the DOM tree to form a clear and hierarchical document structure representation.
[0088] PDF parsing: convert the initial document into a standard PDF document in a unified PDF format, and then perform PDF parsing on the standard PDF document. PDF parsing is to identify the standard PDF document and organize it into PDF parsed text;
[0089] Secondary segmentation of the DOM tree or PDF parsed text and archiving of the slices, as well as hierarchical summary extraction and QA generalization of the content of the PDF parsed text, and storage in the Milvus vector library; hierarchical summary extraction includes extracting multi-level sub-content summaries until the entire summary of the entire document is generated;
[0090] The parsed document in the figure is a DOM tree or PDF parsed text.
[0091] Specifically, refer to Figure 2 , the DOM tree or PDF parsed text can be processed in the following ways:
[0092] 1. Perform secondary segmentation, correct multiple table headers, correct slice classification, archive slices, and then store them in the Milvus vector library;
[0093] 2. Extract the abstracts, build the article directory tree, the bottom chapter abstract, and the parent chapter abstract, and then store them in the Milvus vector library;
[0094] 3. Perform QA generalization in sequence, obtain all slices, generalize the slice content into a large model, manually confirm, and then store it in the Milvus vector library;
[0095] S2. Receive the user's query content, and use the ANN algorithm to retrieve the top five data with the highest relevance in the Milvus vector library according to the user's query content. The TOP5 data includes text information or tables or pictures. If the TOP5 data includes tables or pictures, the tables or pictures are marked and temporarily stored in the cache.
[0096] After receiving the TOP5 data, the big model combines its powerful semantic understanding and reasoning to generate corresponding text answers. If a table or picture is referenced in the process of generating the corresponding answer, the big model will generate an identifier corresponding to the table or picture, and then take out the corresponding table and picture from the cache based on the identifier corresponding to the table or picture and accurately fill it into the corresponding text answer, and generate a picture and text answer with rich information.
[0097] Among them, the large model generates identification corresponding to the table or picture, including links or identifiers.
[0098] An enhanced question-answering system based on document slicing and labeling is also disclosed in an embodiment of the present invention, which is used to implement an enhanced question-answering method based on document slicing and labeling disclosed in the above embodiment. The enhanced question-answering system based on document slicing and labeling includes:
[0099] A receiving and detecting unit, used for receiving an initial document uploaded by a user and detecting the format of the initial document. If it is in HTML format, the HTML parsing unit is used; if it is in DOC format, DOCX format or PDF format, the PDF parsing unit is used;
[0100] The HTML parsing unit is used to parse the initial document in HTML format into a DOM tree and filter the tags specified by the user;
[0101] The HTML parsing unit includes:
[0102] The child node acquisition module is used to obtain child nodes and obtain a list of all child nodes starting from the current element;
[0103] The child node traversal detection module is used to traverse each child node and detect the child node type, and then perform corresponding processing according to the child node type. The child node type includes a text node or an element node. If it is a text node, it enters the text node module, and if it is an element node, it enters the element node module;
[0104] The text node module is used to create a storage object to store the text node. Since the text node has no HTML tag, its tag is set to an empty string, and the content of the text node is set to the storage object;
[0105] Element nodes are used to process according to the tag names of element nodes. Title names include titles, tables, pictures and other elements. Other elements are non-title, non-table and non-picture elements. If it is a title, it enters the title module. If it is a table, it enters the table module. If it is a picture, it enters the picture module. If it is other elements, it enters the other element module.
[0106] The title module is used to recursively call a function, which is used to extract the text content of an element and all its child elements, then create an HtmlElement object to store the title text, and add the title text to the node list;
[0107] Table module converts the HTML content of the table into Markdown format. It adds corresponding start and end tags before and after the conversion result. The formatted string is encapsulated in an HtmlElement object and added to the node list.
[0108] The image module is used to determine whether the document comes from a ZIP file. If so and the image resource is available, it extracts the image and constructs an image text including the image path, then creates an HtmlElement object to store the image URL and adds it to the node list;
[0109] Other element module, used to check whether other elements need to be filtered out. If not, the function of processing element nodes is recursively called to process the other elements. For p and br tags, HtmlElement objects including line breaks are created to simulate the conversion of other elements in HTML and added to the node list.
[0110] The filtering mechanism module is used to store the filtered elements in the filter collection.
[0111] DOM tree generation module, used to recursively build and generate DOM tree based on HTML parsing;
[0112] PDF parsing unit: used for converting the initial document into a standard PDF document in a unified PDF format, and then performing PDF parsing on the standard PDF document, wherein the PDF parsing is to identify the standard PDF document and organize it into PDF parsed text;
[0113] The secondary processing module is used to perform secondary segmentation on the DOM tree or PDF parsed text and archive the slices, as well as perform hierarchical summary extraction and QA generalization on the content of the PDF parsed text and store it in the Milvus vector library;
[0114] A user query receiving module, used to receive user query content;
[0115] The retrieval module is used to retrieve the top five data with the highest relevance in the Milvus vector library according to the user's query content. The top five data include text information, tables or pictures.
[0116] Whether to include tables or pictures: a module for judging whether the TOP5 data includes tables or pictures. If so, the tables or pictures are marked and temporarily stored in the cache.
[0117] The large model module is used to receive the TOP5 data and generate corresponding text answers based on its powerful semantic understanding and reasoning;
[0118] The table or picture reference module is used to detect whether the text answer generated by the large model module references a table or picture. If so, it enters the table or picture supplement module;
[0119] The table or picture supplement module is used to call the large model module to generate labels corresponding to the table or picture, and then take out the corresponding table and picture from the cache based on the labels corresponding to the table or picture to fill in the corresponding text answer, and generate a graphic answer.
[0120] An embodiment of the present invention further discloses an electronic device, which includes: a processor and a memory. The processor and the memory are connected, such as through a bus. Optionally, the electronic device may also include a transceiver. It should be noted that in actual applications, the transceiver is not limited to one, and the structure of the electronic device does not constitute a limitation on the embodiment of the present invention.
[0121] The processor may be a CPU central processing unit, a general processor, a DSP data signal processor, an ASIC application specific integrated circuit, an FPGA field programmable gate array or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It may implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of the present invention. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0122] The bus may include a path to transmit information between the above components. The bus may be a PCI peripheral component interconnect standard bus or an EISA extended industry standard architecture bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc.
[0123] The memory can be a ROM read-only memory or other types of static storage devices that can store static information and instructions, a RAM random access memory or other types of dynamic storage devices that can store information and instructions, or an EEPROM electrically erasable programmable read-only memory, a CD-ROM read-only disc or other optical disc storage, an optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to this.
[0124] The memory is used to store the application code for executing the solution of the present invention, and the execution is controlled by the processor. The processor is used to execute the application code stored in the memory to implement the contents shown in the enhanced question-answering method based on document slicing and marking disclosed in the above embodiment.
[0125] The electronic devices include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and fixed terminals such as digital TVs, desktop computers, etc. It can also be a server, etc.
[0126] An embodiment of the present invention discloses a computer-readable storage medium having a computer program stored thereon. When the computer-readable storage medium is run on a computer, the computer can execute the corresponding content of an enhanced question-answering method based on document slicing and tagging disclosed in the above embodiment.
[0127] The above are all preferred embodiments of the present invention, and are not intended to limit the protection scope of the present invention. Therefore, any equivalent changes made based on the structure, shape, and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. An enhanced question-answering method based on document slicing and tagging, characterized in that: The following steps are involved: S1. Receive the initial document uploaded by the user and detect the format of the initial document. If it is in HTML format, perform HTML parsing. If it is in DOC format, DOCX format or PDF format, perform PDF parsing. HTML parsing: parse the initial document in HTML format into a DOM tree; PDF parsing: convert the initial document into a standard PDF document in a unified PDF format, and then perform PDF parsing on the standard PDF document. PDF parsing is to identify the standard PDF document and organize it into PDF parsed text; Secondarily segment the DOM tree or PDF parsed text and archive the slices, perform hierarchical summary extraction and QA generalization on the content of the PDF parsed text, and store them in the Milvus vector library; Among them, hierarchical summary extraction includes extracting multi-level sub-content summaries until the entire summary of the entire document is generated; S2. Receive the user's query content, and retrieve the top five data with the highest relevance in the Milvus vector library according to the user's query content. The top five data include text information or tables or pictures. If the top five data include tables or pictures, mark the tables or pictures and temporarily store them in the cache. After receiving the TOP5 data, the big model combines its powerful semantic understanding and reasoning to generate corresponding text answers. If a table or picture is referenced in the process of generating the corresponding answer, the big model will generate an identifier corresponding to the table or picture, and then take out the corresponding table and picture from the cache based on the identifier corresponding to the table or picture and fill it into the corresponding text answer, and generate a graphic answer.
2. The enhanced question-answering method based on document slicing and marking according to claim 1, characterized in that: The large model generates tags corresponding to the table or picture, including links or identifiers.
3. The enhanced question-answering method based on document slicing and marking according to claim 1, characterized in that: HTML parsing consists of the following steps: Get child nodes: Starting from the current element, get a list of all child nodes; Traverse child nodes: traverse each child node, detect the child node type, and then perform corresponding processing according to the child node type. The child node type includes text node or element node; If the child node type is a text node, create a storage object to store the text node. Since the text node has no HTML tag, set its tag to an empty string and set the content of the text node to the storage object. If the child node type is an element node, it is processed according to the tag name of the element node, including the following categories: If the tag name is title, a function is recursively called to extract the text content of the element and all its child elements, then an HtmlElement object is created to store the title text, and the title text is added to the node list; If the tag name is table, the HTML content of the table is converted to Markdown format. The corresponding start and end tags for distinction are added before and after the conversion result. The formatted string is encapsulated in an HtmlElement object and added to the node list. If the tag name is picture: If the document comes from a ZIP file and the picture resource is available, extract the picture and construct the picture text including the picture path, then create an HtmlElement object to store the picture URL and add it to the node list; If the tag name is other elements: other elements are non-title, non-table and non-picture elements, check whether other elements need to be filtered out. If not, recursively call the function of processing element nodes to process the other elements. For p and br tags, create HtmlElement objects including line breaks to simulate the conversion of other elements in HTML and add them to the node list; Recursively build and generate a DOM tree based on HTML parsing.
4. The enhanced question-answering method based on document slicing and marking according to claim 3 is characterized in that: HTML parsing also includes a filtering mechanism that stores filtered elements in a filter collection.
5. The enhanced question-answering method based on document slicing and marking according to claim 1, characterized in that: The DOM tree includes the starting heading level, the HtmlElement list that includes all HTML elements, the root node of the text object tree, and the parent text object currently being processed.
6. An enhanced question-answering system based on document slicing and tagging, characterized in that: include: A receiving and detecting unit, used for receiving an initial document uploaded by a user and detecting the format of the initial document. If it is in HTML format, the HTML parsing unit is used; if it is in DOC format, DOCX format or PDF format, the PDF parsing unit is used; An HTML parsing unit, used to parse an initial document in HTML format into a DOM tree; PDF parsing unit: used for converting the initial document into a standard PDF document in a unified PDF format, and then performing PDF parsing on the standard PDF document, wherein the PDF parsing is to identify the standard PDF document and organize it into PDF parsed text; The secondary processing module is used to perform secondary segmentation on the DOM tree or PDF parsed text and archive the slices, as well as perform hierarchical summary extraction and QA generalization on the content of the PDF parsed text and store it in the Milvus vector library; A user query receiving module, used to receive user query content; The retrieval module is used to retrieve the top five data with the highest relevance in the Milvus vector library according to the user's query content. The top five data include text information, tables or pictures. Whether to include tables or pictures: a module for judging whether the TOP5 data includes tables or pictures. If so, the tables or pictures are marked and temporarily stored in the cache. The large model module is used to receive the TOP5 data and generate corresponding text answers based on its powerful semantic understanding and reasoning; The table or picture reference module is used to detect whether the text answer generated by the large model module references a table or picture. If so, it enters the table or picture supplement module; The table or picture supplement module is used to call the large model module to generate labels corresponding to the table or picture, and then take out the corresponding table and picture from the cache based on the labels corresponding to the table or picture to fill in the corresponding text answer, and generate a graphic answer.
7. The enhanced question-answering system based on document slicing and marking according to claim 6, characterized in that: The HTML parsing unit includes: The child node acquisition module is used to obtain child nodes and obtain a list of all child nodes starting from the current element; The child node traversal detection module is used to traverse each child node and detect the child node type, and then perform corresponding processing according to the child node type. The child node type includes a text node or an element node. If it is a text node, it enters the text node module, and if it is an element node, it enters the element node module; The text node module is used to create a storage object to store the text node. Since the text node has no HTML tag, its tag is set to an empty string, and the content of the text node is set to the storage object; Element nodes are used to process according to the tag names of element nodes. Title names include titles, tables, pictures and other elements. Other elements are non-title, non-table and non-picture elements. If it is a title, it enters the title module. If it is a table, it enters the table module. If it is a picture, it enters the picture module. If it is other elements, it enters the other element module. The title module is used to recursively call a function, which is used to extract the text content of an element and all its child elements, then create an HtmlElement object to store the title text, and add the title text to the node list; Table module converts the HTML content of the table into Markdown format. It adds corresponding start and end tags before and after the conversion result. The formatted string is encapsulated in an HtmlElement object and added to the node list. The image module is used to determine whether the document comes from a ZIP file. If so and the image resource is available, it extracts the image and constructs an image text including the image path, then creates an HtmlElement object to store the image URL and adds it to the node list; Other element module, used to check whether other elements need to be filtered out. If not, the function of processing element nodes is recursively called to process the other elements. For p and br tags, HtmlElement objects including line breaks are created to simulate the conversion of other elements in HTML and added to the node list. DOM tree generation module, used to recursively build and generate DOM tree based on HTML parsing.
8. The enhanced question-answering system based on document slicing and marking according to claim 7, characterized in that: The HTML parsing unit also includes a filtering mechanism module for storing filtered elements in a filtering set.
9. An electronic device, characterized in that: It includes: One or more processors; Memory; one or more applications; One or more of the applications are stored in the memory and configured to be executed by one or more of the processors, and one or more of the applications are configured to: execute an enhanced question-answering method based on document slicing and tagging according to any one of claims 1 to 5.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the enhanced question-answering method based on document slicing and tagging as described in any one of claims 1 to 5.