A large model document question-answering method and system based on graph structure
By introducing a combination of graph structure and large models in the document Q&A system, the problem of low information extraction efficiency in document Q&A is solved, and accurate Q&A and user experience improvement in the home field is achieved.
Patent Information
- Application Number
- CN202411323514.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-23
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-09-23
AI Technical Summary
In the field of document Q&A/understanding, it is difficult for the prior art to quickly and accurately extract all information fragments related to the problem, resulting in a decrease in output accuracy and excessive resource usage, especially in the field of home furnishings.
A large-scale document question-and-answer method based on graph structure is proposed. By obtaining target product documents, identifying product information and generating preliminary knowledge graphs, searching for graph structure similarity in historical product knowledge bases, integrating or supplementing knowledge graphs, extracting features related to user questions, and performing inference answers through large models.
It realizes accurate response to related questions in the home field, improves the accuracy and user experience of answers, reduces resource usage, and improves the efficiency of document understanding.
Smart Images

Figure CN119377361B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a large-model document question-answering method and system based on a graph structure. Background Art
[0002] In the field of document question answering / understanding, how to quickly and accurately obtain all the information fragments related to a certain question or a certain content of interest has always been the focus of business attention. In the era of big models and big data, the understanding of documents often needs to rely on the ability of big models to embedd vectorization of the entire document, and then divide it into blocks according to manual or other rules and store it in the vector database. When encountering a question about a certain information in the document, it is necessary to extract the parts related to the question from all the blocks in the vector database, merge these document blocks with the question and send them to the big model for reasoning and answering. The big model document understanding method using graph structure can extract all related parts from the vector database more accurately and quickly, reduce the resources consumed by document understanding, improve the accuracy of document understanding, and thus improve work efficiency.
[0003] Document question answering generally requires obtaining all information fragments related to the question and information. Information fragments are obtained by using some algorithms such as cosine similarity algorithm and Euclidean distance algorithm to traverse each document block in the vector database and obtain all document blocks that meet the similarity requirements of the question. However, the obtained document blocks contain a lot of irrelevant information, and the key information that is helpful for answering questions can easily be hidden and submerged by useless text, resulting in a decrease in output accuracy. At the same time, more useless information also occupies a large number of large model input characters, and GPU resources are too high and cannot be used efficiently.
[0004] In particular, for the home furnishing field, there is no large-model document question-and-answer method to be applied to this application scenario. Summary of the invention
[0005] The main purpose of the embodiments of the present invention is to propose a large-model document question-answering method and system based on a graph structure, which can accurately answer relevant questions in the home field, has high accuracy and helps to improve user experience.
[0006] To achieve the above purpose, an embodiment of the present invention provides a large model document question answering method based on a graph structure, comprising the following steps:
[0007] Obtain a target product document, identify product information described in the target product document, and generate a preliminary knowledge graph of the graph structure of the target product;
[0008] Based on the product information and the preliminary knowledge graph, a product graph structure similarity search is performed in an existing historical product knowledge base to determine whether the target product document already exists in the historical product knowledge base;
[0009] If the target product exists in the historical product knowledge base, the preliminary knowledge graph is merged with the knowledge graph in the historical product knowledge base to obtain a target knowledge graph;
[0010] If the target product does not exist in the historical product knowledge base, the most similar product in the historical product knowledge base is matched to guide the target product document to generate a more comprehensive target knowledge graph;
[0011] Encode the user's question information and the information of the target knowledge graph, and calculate the similarity to extract all features related to the question information;
[0012] The summary and reasoning capabilities of the large model are used to provide feedback on the extracted features.
[0013] In some embodiments, the step of obtaining a target product document, identifying product information described in the target product document, and generating a preliminary knowledge graph of a graph structure of the target product includes the following steps:
[0014] For the PDF product document input by the user, the PDF product document is converted into a target image in PNG format, and the image is preprocessed, wherein the preprocessing operation includes image grayscale processing and binarization processing;
[0015] Through the feature detection algorithm of deep learning, the contour line segments of the target image are detected to screen out the target contour, and the target contour is used to represent the table structure and the CAD curve structure;
[0016] Based on the analysis of line segment angles and lengths, the identified structural contours are connected and analyzed, multiple line segments are combined into a complete structural contour, and contour analysis results for each content type are generated;
[0017] According to the structural outline and the outline analysis result, the product name contained in the document and the attribute information corresponding to the product are obtained by adjusting the prompt word prompt in the home furnishing field, wherein the attribute information includes color, style, material, and space;
[0018] Generate a preliminary graph structure knowledge graph based on the attribute information; wherein the main node of the graph structure knowledge graph is the product name, and the child node is the attribute information of the product;
[0019] According to the information of the graph structure knowledge graph of the obtained nodes, the Neo4j graph database is used for storage.
[0020] In some embodiments, the analysis based on line segment angles and lengths, connecting and analyzing the identified structural contours, combining multiple line segments into a complete structural contour, and generating contour parsing results for each content type includes the following steps:
[0021] Based on the obtained contour information, different analysis methods are performed:
[0022] For tabular type profiles, row and column analysis is performed;
[0023] For outlines of the text cell type, the content of the text cell is parsed according to the original layout format of the table;
[0024] For outlines of picture type, the pictures are individually identified and saved to the corresponding positions on the document layout.
[0025] In some embodiments, the method of performing a product graph structure similarity search in an existing historical product knowledge base based on the product information and the preliminary knowledge graph to determine whether the target product document already exists in the historical product knowledge base includes the following steps:
[0026] Traverse and search the existing historical product knowledge base and perform graph similarity matching, specifically: determine whether the target product document already exists in the library, and if so, extract the graph structure of the corresponding product in the historical product knowledge base; if the target product document does not exist in the existing historical product knowledge base, use the path matching algorithm to calculate the similarity between the preliminary knowledge graph and the knowledge graph of the existing historical product knowledge base, and extract the most similar product graph structure in the historical product knowledge base;
[0027] The graph of the target product document is fused, specifically: if the target product document exists in the historical product knowledge base, then the attributes that do not exist in the target product document but also belong to the target product are extracted from the matched historical knowledge base product graph, and integrated into the knowledge graph of the target product document for information supplement and graph alignment; if the target product document does not exist in the historical product knowledge base, then the attributes contained in the matched similar knowledge base product graph are summarized, and these summaries are passed as prompt words to the large model to regenerate the product knowledge graph of the target product document, thereby realizing the guidance of the product knowledge base on the generation of new document graphs.
[0028] In some embodiments, encoding the user's question information with the information of the target knowledge graph and calculating the similarity to extract all features related to the question information includes the following steps:
[0029] Encode the user's question through the embedding model to obtain the user's question embedding vector;
[0030] Use the graph embedding model to vectorize the final product graph and obtain the product graph embedding vector;
[0031] The similarity between the user question embedding vector and the product graph embedding vector is calculated to extract all nodes related to the user question and meeting the threshold and all attribute information of the nodes in the graph.
[0032] In some embodiments, the step of providing answer information based on the extracted feature feedback by using the summarization and reasoning capabilities of the large model includes the following steps:
[0033] The user's question is spliced with all the node information related to the question extracted from the graph, and all are input into the natural language dialogue model for reasoning and analysis, and the big model returns the answer to the question.
[0034] Another aspect of the embodiment of the present invention further provides a large model document question-answering system based on a graph structure, including:
[0035] The first module is used to obtain a target product document, identify the product information described in the target product document, and generate a preliminary knowledge graph of the graph structure of the target product;
[0036] The second module is used to perform a product graph structure similarity search in an existing historical product knowledge base based on the product information and the preliminary knowledge graph, and determine whether the target product document already exists in the historical product knowledge base;
[0037] The third module is used for fusing the preliminary knowledge graph with the knowledge graph in the historical product knowledge base to obtain a target knowledge graph if the target product exists in the historical product knowledge base;
[0038] The fourth module is used for matching the most similar product in the historical product knowledge base to guide the target product document if the target product does not exist in the historical product knowledge base, so as to generate a more comprehensive target knowledge graph;
[0039] The fifth module is used to encode the user's question information and the information of the target knowledge graph, and calculate the similarity to extract all features related to the question information;
[0040] The sixth module is used to provide feedback and answer information based on the extracted features through the summarization and reasoning capabilities of the large model.
[0041] Another aspect of an embodiment of the present invention further provides an electronic device, including a processor and a memory;
[0042] The memory is used to store programs;
[0043] The processor executes the program to implement the method described above.
[0044] Another aspect of the embodiments of the present invention further provides a computer-readable storage medium, wherein the storage medium stores a program, and the program is executed by a processor to implement the method described above.
[0045] Another aspect of an embodiment of the present invention further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method described above is implemented.
[0046] The embodiments of the present invention at least include the following beneficial effects: the present invention provides a large model document question-answering method and system based on graph structure, the scheme obtains the target product document, identifies the product information described in the target product document and generates the preliminary knowledge graph of the graph structure of the target product; according to the product information and the preliminary knowledge graph, the product graph structure similarity search is performed in the existing historical product knowledge base to determine whether the target product document already exists in the historical product knowledge base; if the target product exists in the historical product knowledge base, the preliminary knowledge graph is merged with the knowledge graph in the historical product knowledge base to obtain the target knowledge graph; if the target product does not exist in the historical product knowledge base, the most similar product in the historical product knowledge base is matched to guide the target product document to generate a more comprehensive target knowledge graph; the user's question information and the information of the target knowledge graph are encoded, and the similarity is calculated to extract all the features related to the question information; the answer information is fed back for the extracted features through the summary and reasoning ability of the large model. The embodiments of the present invention can accurately answer relevant questions in the field of home furnishing, with high accuracy and help to improve user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 It is a schematic diagram of an implementation environment provided by an embodiment of the present invention;
[0048] Figure 2 is a flow chart of the overall steps provided by an embodiment of the present invention;
[0049] Figure 3 It is a question-answering flow chart based on a graph structure knowledge graph provided by an embodiment of the present invention;
[0050] Figure 4 It is a data model structure diagram of the domain knowledge graph provided by an embodiment of the present invention;
[0051] Figure 5 A method for generating a product library map guide provided by an embodiment of the present invention;
[0052] Figure 6 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0053] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present invention. They are only examples of devices and methods consistent with some aspects of the embodiments of the present invention as detailed in the attached claims.
[0054] It is understood that the terms "first", "second", etc. used in the present invention may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0055] The terms "at least one", "multiple", "each", "any", etc. used in the present invention, at least one includes one, two or more, multiple includes two or more, each refers to each of the corresponding multiple, and any refers to any one of the multiple.
[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present invention belongs. The terms used herein are only for the purpose of describing the embodiments of the present invention and are not intended to limit the present invention.
[0057] The large-model document question-and-answer method and system based on a graph structure provided in an embodiment of the present invention relate to the field of computer technology. The large-model document question-and-answer method based on a graph structure provided in an embodiment of the present invention can be applied to a terminal, can be applied to a server, or can be software running in a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, and a car terminal, etc., but is not limited to this; the server side can be configured as an independent physical server, or a server cluster or distributed system composed of multiple physical servers, and can also be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application that implements a large-model document question-and-answer method based on a graph structure, etc., but is not limited to the above forms.
[0058] The present invention can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0059] like Figure 1 FIG. 1 is a schematic diagram of an implementation environment provided by an embodiment of the present invention. Figure 1 The implementation environment includes at least one terminal 102 and a server 101. The terminal 102 and the server 101 can be connected to a network wirelessly or wired to complete data transmission and exchange.
[0060] Server 101 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), as well as big data and artificial intelligence platforms.
[0061] In addition, the server 101 can also be a node server in the blockchain network. Blockchain is a new application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm.
[0062] The terminal 102 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc. The terminal 102 may also be a vehicle-mounted terminal of various device types described above, but is not limited thereto. The terminal 102 and the server 101 may be directly or indirectly connected via wired or wireless communication, which is not limited in this embodiment of the present invention.
[0063] Based on the example Figure 1 In the implementation environment shown, an embodiment of the present invention provides a large model document question and answer method based on a graph structure. The following is explained using the large model document question and answer method based on a graph structure applied in the server 101 as an example. It can be understood that the method can also be applied to the terminal 102.
[0064] Reference Figure 2 , Figure 2 A flowchart of a graph-structured large-model document question-and-answer method applied to a server provided in an embodiment of the present invention, wherein the execution subject of the method may be any of the aforementioned computer devices (including a server or a terminal).
[0065] Reference Figure 2 , the method may include the following steps:
[0066] Obtain a target product document, identify product information described in the target product document, and generate a preliminary knowledge graph of the graph structure of the target product;
[0067] Based on the product information and the preliminary knowledge graph, a product graph structure similarity search is performed in an existing historical product knowledge base to determine whether the target product document already exists in the historical product knowledge base;
[0068] If the target product exists in the historical product knowledge base, the preliminary knowledge graph is merged with the knowledge graph in the historical product knowledge base to obtain a target knowledge graph;
[0069] If the target product does not exist in the historical product knowledge base, the most similar product in the historical product knowledge base is matched to guide the target product document to generate a more comprehensive target knowledge graph;
[0070] Encode the user's question information and the information of the target knowledge graph, and calculate the similarity to extract all features related to the question information;
[0071] The summary and reasoning capabilities of the large model are used to provide feedback on the extracted features.
[0072] In some embodiments, the step of obtaining a target product document, identifying product information described in the target product document, and generating a preliminary knowledge graph of a graph structure of the target product includes the following steps:
[0073] For the PDF product document input by the user, the PDF product document is converted into a target image in PNG format, and the image is preprocessed, wherein the preprocessing operation includes image grayscale processing and binarization processing;
[0074] Through the feature detection algorithm of deep learning, the contour line segments of the target image are detected to screen out the target contour, and the target contour is used to represent the table structure and the CAD curve structure;
[0075] Based on the analysis of line segment angles and lengths, the identified structural contours are connected and analyzed, multiple line segments are combined into a complete structural contour, and contour analysis results for each content type are generated;
[0076] According to the structural outline and the outline analysis result, the product name contained in the document and the attribute information corresponding to the product are obtained by adjusting the prompt word prompt in the home furnishing field, wherein the attribute information includes color, style, material, and space;
[0077] Generate a preliminary graph structure knowledge graph based on the attribute information; wherein the main node of the graph structure knowledge graph is the product name, and the child node is the attribute information of the product;
[0078] According to the information of the graph structure knowledge graph of the obtained nodes, the Neo4j graph database is used for storage.
[0079] In some embodiments, the analysis based on line segment angles and lengths, connecting and analyzing the identified structural contours, combining multiple line segments into a complete structural contour, and generating contour parsing results for each content type includes the following steps:
[0080] Based on the obtained contour information, different analysis methods are performed:
[0081] For tabular type profiles, row and column analysis is performed;
[0082] For outlines of the text cell type, the content of the text cell is parsed according to the original layout format of the table;
[0083] For outlines of picture type, the pictures are individually identified and saved to the corresponding positions on the document layout.
[0084] In some embodiments, the method of performing a product graph structure similarity search in an existing historical product knowledge base based on the product information and the preliminary knowledge graph to determine whether the target product document already exists in the historical product knowledge base includes the following steps:
[0085] Traverse and search the existing historical product knowledge base and perform graph similarity matching, specifically: determine whether the target product document already exists in the library, and if so, extract the graph structure of the corresponding product in the historical product knowledge base; if the target product document does not exist in the existing historical product knowledge base, use the path matching algorithm to calculate the similarity between the preliminary knowledge graph and the knowledge graph of the existing historical product knowledge base, and extract the most similar product graph structure in the historical product knowledge base;
[0086] The graph of the target product document is fused, specifically: if the target product document exists in the historical product knowledge base, then the attributes that do not exist in the target product document but also belong to the target product are extracted from the matched historical knowledge base product graph, and integrated into the knowledge graph of the target product document for information supplement and graph alignment; if the target product document does not exist in the historical product knowledge base, then the attributes contained in the matched similar knowledge base product graph are summarized, and these summaries are passed as prompt words to the large model to regenerate the product knowledge graph of the target product document, thereby realizing the guidance of the product knowledge base on the generation of new document graphs.
[0087] In some embodiments, encoding the user's question information with the information of the target knowledge graph and calculating the similarity to extract all features related to the question information includes the following steps:
[0088] Encode the user's question through the embedding model to obtain the user's question embedding vector;
[0089] Use the graph embedding model to vectorize the final product graph and obtain the product graph embedding vector;
[0090] The similarity between the user question embedding vector and the product graph embedding vector is calculated to extract all nodes related to the user question and meeting the threshold and all attribute information of the nodes in the graph.
[0091] In some embodiments, the step of providing answer information based on the extracted feature feedback by using the summarization and reasoning capabilities of the large model includes the following steps:
[0092] The user's question is spliced with all the node information related to the question extracted from the graph, and all are input into the natural language dialogue model for reasoning and analysis, and the big model returns the answer to the question.
[0093] The implementation process of the embodiment of the present invention in a specific application scenario is described in detail below with reference to the accompanying drawings of the specification:
[0094] In order to quickly and accurately answer users' questions about document content information, the text understanding and summarization capabilities of the big model are used to generate graph structure information for the document. Then, the user's questions are matched with the graph structure information through a similar algorithm, and all graph node information related to the question is accurately summarized. The big model's reasoning ability is used to analyze and answer the questions, thereby improving the accuracy of the answers.
[0095] The embodiment of the present invention proposes a method for document question answering using a graph structure. The core is to combine the existing graph structure information in the product library and the reasoning and analysis capabilities of the large model to quickly and accurately aggregate and search the information related to the user's question to achieve the deep understanding requirements of complex documents.
[0096] The embodiment of the present invention proposes a large-scale model household product document question-answering method based on a graph structure, and the method flow chart is as follows: Figure 3 First, AI is used to analyze the documents sent by users, identify the product names described in the documents and the corresponding attributes such as style, color, space, etc., and generate a preliminary Figure 4 The product graph structure knowledge graph shown in the figure. Secondly, perform a product graph structure similarity search in the existing product knowledge base to determine whether the product already exists in the product knowledge base. If the product exists in the product knowledge base, merge the knowledge graph of the document with the knowledge graph in the product base to achieve feature complementation and alignment. If the product does not exist in the product knowledge base, match the closest product in the knowledge base to guide the document and achieve the generation of a more comprehensive knowledge graph. Encode the user's question and the information in the knowledge graph and calculate the similarity to extract all the features related to the question. Finally, give an answer through the summary and reasoning ability of the large model to complete the deep understanding and question-answering of the document.
[0097] Specifically, the implementation process of the embodiment of the present invention includes the following steps:
[0098] 1. Document analysis and knowledge graph construction, reference Figure 5 The graph generation process of the embodiment of the present invention includes the following steps:
[0099] I. Document Analysis
[0100] a. For the PDF product document input by the user, first convert the PDF into a PNG format image, and then perform operations such as image grayscale and binarization to facilitate subsequent content extraction and table parsing.
[0101] b. Through deep learning feature detection algorithms, such as CNN, all possible contour segments in the document image are detected. Among these contour segments, the contours that may represent structures such as tables and CAD curves are screened out.
[0102] c. Connect and analyze the identified structural contours. Combine multiple line segments into a complete structural contour through methods such as line segment angle and length analysis. Perform different analyses based on the obtained contour information, perform row and column analysis on the table, and parse the content of the text grid according to the original layout format of the table; identify the images separately and save them to the corresponding position of the document layout.
[0103] II. Knowledge Graph Construction
[0104] d. Combine the parsed document file with the analysis capability of the big model, and adjust the prompt word in the home furnishing field to obtain the product name contained in the document and the corresponding color, style, material, space and other attribute information of the product.
[0105] e. Based on the attribute information obtained above, a graph structure knowledge graph is initially generated. The main node of the knowledge graph is the product name, and the sub-nodes are the attribute information of the product such as color, style, material, space, etc., and these attribute nodes each contain specific information nodes.
[0106] f. The graph structure knowledge graph information of the nodes is obtained and stored in the Neo4j graph database, which facilitates the subsequent retrieval, addition and similarity calculation of the graph information.
[0107] 2. Graph matching and fusion
[0108] I. Graph Similarity Matching
[0109] g. First, traverse and search the existing product knowledge base to determine whether the document product already exists in the library. If it does, extract the graph structure of the product in the knowledge base.
[0110] h. If the document product does not exist in the existing product knowledge base, use the path matching algorithm to calculate the similarity between the document product knowledge graph obtained above and the existing product library knowledge graph, and extract the most similar product graph structure in the library.
[0111] II. Graph Fusion
[0112] i. If the document product exists in the product knowledge base, then extract the attributes that do not exist in the document but also belong to the product from the matched knowledge base product graph, and integrate them into the knowledge graph of the document for information supplement and graph alignment.
[0113] j. If the document product does not exist in the product knowledge base, the attributes contained in the matched similar knowledge base product graph will be summarized, and these summaries will be passed to the large model as prompt words to regenerate the product knowledge graph of the product document, so as to realize the guidance of the product knowledge base on the generation of new document graph.
[0114] 3. Question encoding and node extraction
[0115] I. Question Coding
[0116] k. In order to answer users’ questions better and faster, the same embedding model as the large model is used to encode and represent users’ questions, which is convenient for storage and subsequent similarity calculation.
[0117] II. Node extraction
[0118] l. Use the graph embedding model to vectorize the final product graph, calculate the similarity with the user question embedding vector obtained above, and extract all nodes related to the question and meeting the threshold and all attribute information of the nodes in the graph.
[0119] 4. Large model reasoning answer
[0120] m. The user's question is combined with all the node information related to the question extracted from the graph, and all are input into the natural language dialogue model for reasoning and analysis, and the big model returns the answer to the question. Summarizing information in the form of a graph solves the problem that reference information must be at a similar reading position in traditional big model knowledge base question answering / understanding, while also reducing the redundancy of reference information and improving retrieval efficiency and answer speed.
[0121] Another aspect of the embodiment of the present invention further provides a large model document question-answering system based on a graph structure, including:
[0122] The first module is used to obtain a target product document, identify the product information described in the target product document, and generate a preliminary knowledge graph of the graph structure of the target product;
[0123] The second module is used to perform a product graph structure similarity search in an existing historical product knowledge base based on the product information and the preliminary knowledge graph, and determine whether the target product document already exists in the historical product knowledge base;
[0124] The third module is used for fusing the preliminary knowledge graph with the knowledge graph in the historical product knowledge base to obtain a target knowledge graph if the target product exists in the historical product knowledge base;
[0125] The fourth module is used for matching the most similar product in the historical product knowledge base to guide the target product document if the target product does not exist in the historical product knowledge base, so as to generate a more comprehensive target knowledge graph;
[0126] The fifth module is used to encode the user's question information and the information of the target knowledge graph, and calculate the similarity to extract all features related to the question information;
[0127] The sixth module is used to provide feedback and answer information based on the extracted features through the summarization and reasoning capabilities of the large model.
[0128] It can be understood that the contents of the above method embodiments are all applicable to the present system embodiments, the functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0129] The embodiment of the present invention further provides an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the above-mentioned large model document question-and-answer method based on a graph structure when executing the computer program. The electronic device can be any intelligent terminal including a tablet computer, a car computer, etc.
[0130] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0131] See also Figure 6 , Figure 6 The hardware structure of an electronic device of another embodiment is illustrated, and the electronic device includes:
[0132] The processor 601 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present invention;
[0133] The memory 602 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 602 can store an operating system and other application programs. When the technical solution provided in the embodiment of this specification is implemented by software or firmware, the relevant program code is stored in the memory 602, and the processor 601 calls and executes the large model document question-and-answer method based on the graph structure of the embodiment of the present invention;
[0134] Input / output interface 603, used to implement information input and output;
[0135] Communication interface 604, used to realize communication interaction between the device and other devices, which can be realized by wired mode (such as USB, network cable, etc.) or wireless mode (such as mobile network, WIFI, Bluetooth, etc.);
[0136] Bus 605 , which transmits information between various components of the device (e.g., processor 601 , memory 602 , input / output interface 603 , and communication interface 604 );
[0137] The processor 601 , the memory 602 , the input / output interface 603 and the communication interface 604 are connected to each other in communication within the device via a bus 605 .
[0138] An embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned large model document question and answer method based on the graph structure.
[0139] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiments, the functions specifically implemented by the present storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0140] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0141] It should be noted that in various specific embodiments of the present invention, when it comes to the need to perform relevant processing based on data related to user identity or characteristics such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present invention needs to obtain the user's sensitive personal information, it will obtain the user's separate permission or consent through a pop-up window or jump to a confirmation page, and after clearly obtaining the user's separate permission or consent, it will obtain the necessary user-related data for the normal operation of the embodiment of the present invention.
[0142] The embodiments described in the embodiments of the present invention are intended to more clearly illustrate the technical solutions of the embodiments of the present invention, and do not constitute a limitation on the technical solutions provided by the embodiments of the present invention. Those skilled in the art can appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present invention are also applicable to similar technical problems.
[0143] Those skilled in the art will appreciate that the technical solutions shown in the figures do not limit the embodiments of the present invention and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0144] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0145] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.
[0146] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0147] It should be understood that in the present invention, "at least one (item)" refers to one or more, and "plurality" refers to two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can represent: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0148] In the several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0149] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0150] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0151] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including multiple instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store programs.
[0152] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the embodiments of the present invention is not limited thereby. Any modification, equivalent substitution and improvement made by those skilled in the art without departing from the scope and essence of the embodiments of the present invention shall be within the scope of the rights of the embodiments of the present invention.
Claims
1. A large-model document question-answering method based on a graph structure, characterized in that: The following steps are involved: Obtain a target product document, identify product information described in the target product document, and generate a preliminary knowledge graph of the graph structure of the target product; Based on the product information and the preliminary knowledge graph, a product graph structure similarity search is performed in an existing historical product knowledge base to determine whether the target product document already exists in the historical product knowledge base; If the target product exists in the historical product knowledge base, the preliminary knowledge graph is merged with the knowledge graph in the historical product knowledge base to obtain a target knowledge graph; If the target product does not exist in the historical product knowledge base, the most similar product in the historical product knowledge base is matched to guide the target product document to generate a more comprehensive target knowledge graph; Encode the user's question information and the information of the target knowledge graph, and calculate the similarity to extract all features related to the question information; Feedback information based on extracted features through the summarization and reasoning capabilities of the large model; The step of performing a product graph structure similarity search in an existing historical product knowledge base based on the product information and the preliminary knowledge graph to determine whether the target product document already exists in the historical product knowledge base comprises the following steps: Traverse and search the existing historical product knowledge base and perform graph similarity matching, specifically: determine whether the target product document already exists in the library, and if so, extract the graph structure of the corresponding product in the historical product knowledge base; if the target product document does not exist in the existing historical product knowledge base, use the path matching algorithm to calculate the similarity between the preliminary knowledge graph and the knowledge graph of the existing historical product knowledge base, and extract the most similar product graph structure in the historical product knowledge base; The graph of the target product document is fused, specifically: if the target product document exists in the historical product knowledge base, then the attributes that do not exist in the target product document but also belong to the target product are extracted from the matched historical knowledge base product graph, and integrated into the knowledge graph of the target product document for information supplement and graph alignment; if the target product document does not exist in the historical product knowledge base, then the attributes contained in the matched similar knowledge base product graph are summarized, and these summaries are passed as prompt words to the large model to regenerate the product knowledge graph of the target product document, thereby realizing the guidance of the product knowledge base on the generation of new document graphs.
2. According to the graph-structured large-model document question-answering method of claim 1, it is characterized in that: The step of obtaining the target product document, identifying the product information described in the target product document, and generating a preliminary knowledge graph of the graph structure of the target product includes the following steps: For the PDF product document input by the user, the PDF product document is converted into a target image in PNG format, and the image is preprocessed, wherein the preprocessing operation includes image grayscale processing and binarization processing; Through the feature detection algorithm of deep learning, the contour line segments of the target image are detected to screen out the target contour, and the target contour is used to represent the table structure and the CAD curve structure; Based on the analysis of line segment angles and lengths, the identified structural contours are connected and analyzed, multiple line segments are combined into a complete structural contour, and contour analysis results for each content type are generated; According to the structural outline and the outline analysis result, the product name contained in the document and the attribute information corresponding to the product are obtained by adjusting the prompt word prompt in the home furnishing field, wherein the attribute information includes color, style, material, and space; Generate a preliminary graph structure knowledge graph based on the attribute information; wherein the main node of the graph structure knowledge graph is the product name, and the child node is the attribute information of the product; According to the information of the graph structure knowledge graph of the obtained nodes, the Neo4j graph database is used for storage.
3. According to the graph-structured large-model document question-answering method of claim 2, it is characterized in that: The analysis based on the line segment angle and length connects and analyzes the identified structural contours, combines multiple line segments into a complete structural contour, and generates contour analysis results for each content type, including the following steps: Based on the obtained contour information, different analysis methods are performed: For tabular type profiles, row and column analysis is performed; For outlines of the text cell type, the content of the text cell is parsed according to the original layout format of the table; For outlines of picture type, the pictures are individually identified and saved to the corresponding positions on the document layout.
4. According to the graph-structured large-model document question-answering method of claim 1, it is characterized in that: The step of encoding the user's question information and the target knowledge graph information and calculating the similarity to extract all features related to the question information includes the following steps: Encode the user's question through the embedding model to obtain the user's question embedding vector; Use the graph embedding model to vectorize the final product graph and obtain the product graph embedding vector; The similarity between the user question embedding vector and the product graph embedding vector is calculated to extract all nodes related to the user question and meeting the threshold and all attribute information of the nodes in the graph.
5. According to the graph-structured large-model document question-answering method of claim 1, it is characterized in that: The method of providing answer information based on the extracted features through summarization and reasoning capabilities of the large model includes the following steps: The user's question is spliced with all the node information related to the question extracted from the graph, and all are input into the natural language dialogue model for reasoning and analysis, and the big model returns the answer to the question.
6. A large-scale document question-answering system based on a graph structure, characterized in that: include: The first module is used to obtain a target product document, identify the product information described in the target product document, and generate a preliminary knowledge graph of the graph structure of the target product; The second module is used to perform a product graph structure similarity search in an existing historical product knowledge base based on the product information and the preliminary knowledge graph, and determine whether the target product document already exists in the historical product knowledge base; The third module is used for fusing the preliminary knowledge graph with the knowledge graph in the historical product knowledge base to obtain a target knowledge graph if the target product exists in the historical product knowledge base; The fourth module is used for matching the most similar product in the historical product knowledge base to guide the target product document if the target product does not exist in the historical product knowledge base, so as to generate a more comprehensive target knowledge graph; The fifth module is used to encode the user's question information and the information of the target knowledge graph, and calculate the similarity to extract all features related to the question information; The sixth module is used to provide feedback and answer information based on the extracted features through the summarization and reasoning capabilities of the large model; Wherein, the second module is specifically used for: Traverse and search the existing historical product knowledge base and perform graph similarity matching, specifically: determine whether the target product document already exists in the library, and if so, extract the graph structure of the corresponding product in the historical product knowledge base; if the target product document does not exist in the existing historical product knowledge base, use the path matching algorithm to calculate the similarity between the preliminary knowledge graph and the knowledge graph of the existing historical product knowledge base, and extract the most similar product graph structure in the historical product knowledge base; The graph of the target product document is fused, specifically: if the target product document exists in the historical product knowledge base, then the attributes that do not exist in the target product document but also belong to the target product are extracted from the matched historical knowledge base product graph, and integrated into the knowledge graph of the target product document for information supplement and graph alignment; if the target product document does not exist in the historical product knowledge base, then the attributes contained in the matched similar knowledge base product graph are summarized, and these summaries are passed as prompt words to the large model to regenerate the product knowledge graph of the target product document, thereby realizing the guidance of the product knowledge base on the generation of new document graphs.
7. An electronic device, characterized in that: including a processor and a memory; The memory is used to store programs; The processor executes the program to implement the method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that: The storage medium stores a program, and the program is executed by a processor to implement the method according to any one of claims 1 to 5.
9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Book knowledge question-answering method based on knowledge graph enhanced large language model
CN117891930A
Large language model knowledge question-answering method and system fused with multi-modal knowledge graph
CN118627628A