Data generation method and device and data query method and device

By sharding the document and generating Q&A pairs, graph structure elements and community vectors, the problem of error information generated by the big model in knowledge Q&A and lagging knowledge updates is solved, achieving a more comprehensive and accurate retrieval effect.

CN120011492APending Publication Date: 2025-05-16BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411864855.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-05-16

Smart Images

  • Figure CN120011492A_ABST
    Figure CN120011492A_ABST
Patent Text Reader

Abstract

The invention provides a data generation method and device, relates to the technical field of information, in particular to the technical fields of natural language processing, large models, RAG (Retrieval Enhanced Generation) and the like, and can be applied to the fields of intelligent questions and answers, intelligent medical inquiry, educational training, legal consultation, news interpretation and the like. According to the specific implementation scheme, an obtained document is subjected to fragmentation processing, and a text unit set is obtained; based on the text unit set, obtaining a question-answer pair set and a graph structure element set, and storing a graph structure mapping relation in a graph database; based on the question and answer pair set and the graph structure element set, obtaining a question vector and a graph structure vector, taking the question vector and the graph structure vector as text vectors, and storing the text vectors and a text mapping relation in a vector database; on the basis of the graph database, community vectors are obtained, community mapping relations and the community vectors are stored in the vector database, and the community mapping relations are used for representing the relation between the community vectors and the graph structure element set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of information technology, specifically to technical fields such as natural language processing, large models, and retrieval enhancement generation (RAG). It can be applied to fields such as intelligent question and answer, intelligent medical consultation, education and training, legal consultation, and news interpretation, and in particular to a data generation method and device, a data query method and device, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0002] Retrieval-augmented generation means that when a large model answers a question or generates text, it first retrieves relevant information from a large document library, and then uses the retrieved information to generate an answer or text.

[0003] With the rise of big model technology, the introduction of the RAG (Retrieval-Augmented Generation) technology paradigm into the contextual learning process of big models can effectively alleviate the problems of big models generating wrong information and lagging knowledge updates in knowledge question answering, and increase the reliability of big model knowledge question answering. In the RAG system, "R" stands for retrieval. Building an efficient retriever is a complex task, and its goal is to retrieve the text containing the answer and rank it highly. Summary of the invention

[0004] The present disclosure provides a data generation method and device, a data query method and device, an electronic device, a computer-readable storage medium, and a computer program product.

[0005] According to a first aspect, a data generation method is provided, the method comprising: performing segmentation processing on an acquired document to obtain a text unit set; obtaining a question-answer pair set and a graph structure element set based on the text unit set, and storing the graph structure element set and the graph structure mapping relationship in a graph database, wherein the graph structure mapping relationship is used to characterize the relationship between the text unit set and the graph structure element set; obtaining a question vector and a graph structure vector based on the question-answer pair set and the graph structure element set, using the question vector and the graph structure vector as text vectors, and storing the text mapping relationship and the text vector in a vector database, wherein the text mapping relationship is used to characterize the relationship between the text vector and the question-answer pair set; obtaining a community vector based on the graph database, and storing the community mapping relationship and the community vector in the vector database, wherein the community mapping relationship is used to characterize the relationship between the community vector and the graph structure element set.

[0006] According to the second aspect, another data query method is provided, which includes: obtaining a question to be processed; obtaining a graph structure element to be queried and a question to be queried based on the question to be processed, a graph database and a vector database; obtaining a query result for the question to be processed based on the graph structure element to be queried, the question to be queried, a graph structure mapping relationship, a text mapping relationship and a community mapping relationship; the graph structure mapping relationship is used to characterize the relationship between a text unit set and a graph structure element set, the text mapping relationship is used to characterize the relationship between a text vector and a question-answer pair set, and the community mapping relationship is used to characterize the relationship between a community vector and a graph structure element set.

[0007] According to the third aspect, a data generating device is provided, which includes: a segmentation unit configured to segment the acquired document to obtain a text unit set; an analysis unit configured to obtain a question-answer pair set and a graph structure element set based on the text unit set, and store the graph structure element set and the graph structure mapping relationship in a graph database, wherein the graph structure mapping relationship is used to characterize the relationship between the text unit set and the graph structure element set; a quantization unit configured to obtain a question vector and a graph structure vector based on the question-answer pair set and the graph structure element set, use the question vector and the graph structure vector as text vectors, and store the text mapping relationship and the text vector in a vector database, wherein the text mapping relationship is used to characterize the relationship between the text vector and the question-answer pair set; a community unit configured to obtain a community vector based on the graph database, and store the community mapping relationship and the community vector in the vector database, wherein the community mapping relationship is used to characterize the relationship between the community vector and the graph structure element set.

[0008] According to a fourth aspect, a data query device is provided, which includes: an acquisition unit, configured to acquire a question to be processed; a question acquisition unit, configured to obtain a graph structure element to be checked and a question to be checked based on the question to be processed, a graph database and a vector database; a result acquisition unit, configured to obtain a query result of the question to be processed based on the graph structure element to be checked, the question to be checked, a graph structure mapping relationship, a text mapping relationship and a community mapping relationship; the graph structure mapping relationship is used to represent the relationship between a text unit set and a graph structure element set, the text mapping relationship is used to represent the relationship between a text vector and a question-answer pair set, and the community mapping relationship is used to represent the relationship between a community vector and a graph structure element set.

[0009] According to the fifth aspect, an electronic device is provided, which includes: at least one processor; and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in any implementation of the first aspect or the second aspect.

[0010] According to a sixth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the method described in any implementation of the first aspect or the second aspect.

[0011] According to a seventh aspect, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the method described in any implementation of the first aspect or the second aspect.

[0012] The data generation method and device provided by the embodiments of the present disclosure firstly perform segmentation processing on the acquired documents to obtain a text unit set; secondly, based on the text unit set, obtain a question-answer pair set and a graph structure element set, and store the graph structure element set and the graph structure mapping relationship in a graph database; thirdly, based on the question-answer pair set and the graph structure element set, obtain a question vector and a graph structure vector, use the question vector and the graph structure vector as text vectors, and store the text mapping relationship and the text vector in a vector database; finally, based on the graph database, obtain a community vector, and store the community mapping relationship and the community vector in the vector database. Thus, on the basis of the acquired documents, various forms of information to be retrieved, such as graph structure elements, text vectors, and community vectors, can be generated, providing a reliable query basis for data query.

[0013] The data query method and device provided by the embodiments of the present disclosure first obtain the problem to be processed; secondly, based on the problem to be processed, the graph database and the vector database, the graph structure elements to be checked and the problem to be checked are obtained; finally, based on the graph structure elements to be checked, the problem to be checked, the graph structure mapping relationship, the text mapping relationship and the community mapping relationship, the query result of the problem to be processed is obtained. In this way, more comprehensive information can be obtained from the graph structure elements and the problem than by querying the problem alone; based on the different mapping relationships of the graph structure elements and the problem query, more accurate query results can be obtained than text vector retrieval, which improves the comprehensiveness of the retrieval; when the query results are given to the large model, the large model can generate accurate results.

[0014] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0016] Figure 1 is a flow chart of an embodiment of a data generation method according to the present disclosure;

[0017] Figure 2is a schematic diagram of the structure of the data generation process in the present disclosure;

[0018] Figure 3 is a flow chart of an embodiment of a data query method according to the present disclosure;

[0019] Figure 4 It is a structural diagram of the data query process in the present disclosure;

[0020] Figure 5 is a structural schematic diagram of an embodiment of a data generating device according to the present disclosure;

[0021] Figure 6 is a structural schematic diagram of an embodiment of a data query device according to the present disclosure;

[0022] Figure 7 It is a block diagram of an electronic device used to implement the data generation method or data query method of the embodiment of the present disclosure. DETAILED DESCRIPTION

[0023] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0024] Currently, the main retrieval technologies in the industry include keyword retrieval technology based on keyword matching and text vector retrieval technology based on text vectors.

[0025] Keyword search technology is a search method based on word matching. The query question entered by the user is converted into a keyword list after passing through the word segmenter. Algorithms such as TFIDF and BM25 are used to match the content containing these keywords in the document and return the results.

[0026] Text vector retrieval technology is based on the vector space model, which represents query questions and text units as vectors in the vector space, and matches and sorts them by calculating the similarity between vectors.

[0027] Disadvantages of keyword search technology:

[0028] The accuracy is limited, only keywords are matched, and word meanings and contextual information are ignored, which may lead to missed detections or false detections; semantic understanding is not supported, and semantic relevance such as synonyms and antonyms cannot be understood.

[0029] Disadvantages of text vector retrieval technology:

[0030] The vectorization model has limitations on query questions and document length, which will cause loss of semantic information. In addition, when the text involves multiple semantic centers, the features captured by the vectorization model are lacking in semantic completeness.

[0031] In addition, both of the above technologies cannot solve the problem of multi-topic summarization. Multi-topic summary retrieval is a summary retrieval focused on the query, rather than a direct and explicit text vector or keyword retrieval.

[0032] In view of the defects in traditional technologies, this disclosure proposes a data generation method. Figure 1 A process 100 according to an embodiment of a data generation method of the present disclosure is shown. The data generation method comprises the following steps:

[0033] Step 101, segment the acquired document to obtain a text unit set.

[0034] In this embodiment, the acquired document is a subject that records information content, for example, a document is mainly used to store text, format settings and other information. The execution subject on which the data generation method runs can obtain the document in a variety of ways. For example, the execution subject can obtain the document stored in the database server through a wired connection or a wireless connection. For another example, a user can obtain the document collected by the terminal by communicating with the terminal.

[0035] In this embodiment, the text unit set is a set of text units, and the text unit set includes at least one text unit, which is the smallest unit of text, such as a character, a word, or a number. The acquired document may include multiple text units, and the text unit set including at least one text unit is obtained by performing information recognition on the acquired document and dividing the recognition result into text units.

[0036] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of the documents involved are carried out after authorization and comply with relevant laws and regulations.

[0037] Step 102, based on the text unit set, obtain a question-answer pair set and a graph structure element set, and store the graph structure element set and the graph structure mapping relationship in a graph database.

[0038] In this embodiment, the question-answer pair set includes at least one question-answer pair, and the question-answer pair includes a question and an answer to the question. Question-answer pair mining is performed on each text unit in the text unit set to obtain a question-answer pair of each text unit, and the question-answer pairs of all text units are collected to obtain a question-answer pair set including at least one question-answer pair. Among them, question-answer pair mining on text units means filling text units according to a certain question-answer format to obtain question-answer pairs.

[0039] In this embodiment, the graph structure element set includes at least one graph structure element, and the graph structure elements include: entities, relationships, and claims, where entities can be physical objects, concepts, or events; relationships define how two entities are associated; and claims represent the semantic relevance of entities and relationships.

[0040] In this embodiment, graph structure element mining can be performed on the text units in the text unit set to obtain a graph structure element set including at least one graph structure element, wherein the graph structure element mining on the text units in the text unit set can use a classification model to classify the text units into entities, relationships, and claims, and determine the text units in the text units that belong to the graph structure elements.

[0041] In this embodiment, the graph structure mapping relationship is used to characterize the relationship between the text unit set and the graph structure element set. When the graph structure element is obtained based on the text unit, the text unit and the graph structure element have a mapping relationship. Specifically, the graph structure mapping relationship can characterize the correspondence between each text unit in the text unit set and each graph structure element in the graph structure element set. The graph structure mapping relationship can be recorded in a relationship table. When querying data, the graph structure element can be directly queried based on the obtained text unit using the graph structure mapping relationship, or the text unit can be queried based on the obtained graph structure element using the graph structure mapping relationship.

[0042] Step 103, based on the question-answer pair set and the graph structure element set, obtain a question vector and a graph structure vector, use the question vector and the graph structure vector as text vectors, and store the text mapping relationship and the text vector in a vector database.

[0043] In this embodiment, the questions in the question-answer pair set can be directly vectorized to obtain a question vector; and the graph structure element set can be vectorized to obtain a graph structure vector.

[0044] Optionally, detect the non-overlapping parts of the questions and graph structure elements in the question-answer pair set, vectorize the non-overlapping questions in the question-answer pair set to obtain a question vector; and vectorize the non-overlapping graph structure elements in the graph structure elements to obtain a graph structure vector.

[0045] In this embodiment, the text mapping relationship is used to characterize the relationship between the text vector and the question-answer pair set. When a question vector in a text vector is obtained based on the question of a question-answer pair in the question-answer pair set, the question-answer pair has a mapping relationship with the question vector. Specifically, the text mapping relationship can characterize the correspondence between each question vector in the text vector and each question-answer pair in the question-answer pair set. The text mapping relationship can be recorded in a relationship table. When querying data, the question-answer pair can be directly queried based on the obtained question vector using the text mapping relationship, or the text vector can be queried based on the obtained question using the text mapping relationship.

[0046] Step 104: obtain a community vector based on the graph database, and store the community mapping relationship and the community vector in the vector database.

[0047] In this embodiment, the community is the division result obtained by clustering the graph structure elements of the same type with certain interactive relationships and common cultural ties using a clustering algorithm. The community can be a cluster name generated by any type of multiple entities, multiple relationships, and multiple claims, or the community is a cluster name generated by any type of entities, relationships, and claims.

[0048] In this embodiment, the above step 104 includes: extracting all graph structure elements from the graph database, clustering all graph structure elements to obtain communities; vectorizing the communities to obtain community vectors. When a question vector in a text vector is obtained based on the question of a question-answer pair in a question-answer pair set, the question-answer pair has a mapping relationship with the question vector. Specifically, the text mapping relationship can represent the correspondence between each question vector in the text vector and each question-answer pair in the question-answer pair set. The text mapping relationship can be recorded in a relationship table. When querying data, the question-answer pair can be directly queried based on the obtained question vector using the text mapping relationship, or the text vector can be queried based on the obtained question using the text mapping relationship.

[0049] In this embodiment, the community mapping relationship is used to characterize the relationship between the community vector and the graph structure element set. Specifically, the community mapping relationship can characterize the corresponding relationship between the community vector and each graph structure element in the graph structure element set. The community mapping relationship can be recorded in a relationship table. When querying data, the community mapping relationship can be directly used to query the graph structure element in the graph structure element set based on the obtained community vector, or the community mapping relationship can be used to query the community vector based on the obtained graph structure element.

[0050] The data generation method provided by the embodiment of the present disclosure first performs segmentation processing on the acquired document to obtain a text unit set; secondly, based on the text unit set, obtains a question-answer pair set and a graph structure element set, and stores the graph structure element set and the graph structure mapping relationship in a graph database; thirdly, based on the question-answer pair set and the graph structure element set, obtains a question vector and a graph structure vector, uses the question vector and the graph structure vector as text vectors, and stores the text mapping relationship and the text vector in a vector database; finally, based on the graph database, obtains a community vector, and stores the community mapping relationship and the community vector in the vector database. Thus, on the basis of the acquired document, various forms of information to be retrieved, such as graph structure elements, text vectors, and community vectors, can be generated, providing a reliable query basis for data query.

[0051] In some embodiments of the present disclosure, the above-mentioned slicing processing of the acquired document to obtain a text unit set includes: parsing the acquired document to obtain a parsed text; slicing the parsed text to obtain a text unit set including at least one text unit.

[0052] In this optional implementation, data parsing refers to the process of converting one data format into another readable format. Specifically, data parsing is to convert the given data into a form that is more suitable for use or analysis by analyzing the relationship between the various components in the given data. For example, data in HTML format can be converted into a more understandable form through a parser.

[0053] In this optional implementation, slicing the parsed text refers to dividing the parsed text into multiple text units, where a text unit is a basic unit of text, such as a text unit is a word or a number.

[0054] The method for obtaining a text unit set provided by this optional implementation parses the acquired document to obtain a parsed text; and slices the parsed text to obtain a text unit set including at least one text unit. By slicing the parsed text, information can be refined, thereby improving the subsequent model's understanding of the text units in the text unit set.

[0055] Optionally, the above-mentioned segmenting the acquired document to obtain the text unit set further includes: preprocessing the text unit set.

[0056] In some embodiments of the present disclosure, the above-mentioned obtaining a question-answer pair set and a graph structure element set based on a text unit set, and storing the graph structure element set and the graph structure mapping relationship in a graph database includes: generating a question-answer pair set including at least one question-answer pair based on a large language model and a text unit set; generating a graph structure element set including at least one graph structure element based on a large language model and a text unit set; taking the relationship between the text unit set and the graph structure element set as a graph structure mapping relationship, and storing the graph structure mapping relationship and the graph structure element set in the graph database.

[0057] In this optional implementation, the question-answer pair set includes at least one question-answer pair, which is a question and an answer.

[0058] The method for obtaining a question-answer pair set and a graph structure element provided by this optional implementation method generates a question-answer pair set including at least one question-answer pair based on a large language model and a text unit set; generates a graph structure element set including at least one graph structure element based on a large language model and a text unit set; uses the relationship between the text unit set and the graph structure element set as a graph structure mapping relationship, and stores the graph structure mapping relationship and the graph structure element set in a graph database, and simultaneously generates a question-answer pair set and a graph structure element set through a large model, thereby improving the reliability and accuracy of the question-answer pair set, the graph structure element set and the graph structure mapping relationship.

[0059] In some optional implementations of the present disclosure, the above-mentioned obtaining question vectors and graph structure vectors based on the question-answer pair set and the graph structure element set, taking the question vectors and the graph structure vectors as text vectors, and storing the text mapping relationship and the text vectors in the vector database includes: extracting questions from the question-answer pair set; using a text vectorization model to vectorize the questions to obtain question vectors; using a text vectorization model to vectorize the graph structure element set to obtain a graph structure vector, taking the question vector and the graph structure vector as text vectors; taking the relationship between the text vector and the question-answer pair set as a text mapping relationship, and storing the text mapping relationship and the text vector in the vector database.

[0060] In this optional implementation, the text vectorization model is a pre-trained model for converting text into vectors. The training process of the text vectorization model is a conventional process and will not be repeated here.

[0061] In this optional implementation, the above-mentioned use of a text vectorization model to vectorize the problem to obtain a problem vector refers to: inputting the problem into the vectorization model to obtain the problem vector output by the vectorization model.

[0062] In this optional implementation, the graph structure elements in the graph structure element set belong to text, that is, the graph structure elements can be text units or text unit groups, and the graph structure elements can be converted into graph structure vectors through a text vectorization model.

[0063] In this optional implementation, the above-mentioned use of a text vectorization model to vectorize the graph structure element set to obtain a graph structure vector refers to: inputting the graph structure element set into the text vectorization model to obtain the graph structure vector output by the text vectorization model.

[0064] This optional implementation provides a method for obtaining text vectors and text mapping relationships, extracting questions from a question-answer pair set; using a text vectorization model to vectorize the questions to obtain question vectors; using a text vectorization model to vectorize a graph structure element set to obtain a graph structure vector, and using the question vector and the graph structure vector as text vectors; using the relationship between the text vector and the question-answer pair set as a text mapping relationship, and storing the text mapping relationship and the text vector in a vector database, so that the text vector includes the question vector and the graph structure element vector, thereby improving the diversity of the text vector content.

[0065] Optionally, the above-mentioned obtaining a question vector and a graph structure vector based on the question-answer pair set and the graph structure element set, taking the question vector and the graph structure vector as text vectors, and storing the text mapping relationship and the text vector in the vector database includes: extracting questions from the question-answer pair set; using a text-to-vector algorithm to vectorize the questions to obtain a question vector; using a text-to-vector algorithm to vectorize the graph structure element set to obtain a graph structure vector, taking the question vector and the graph structure vector as text vectors; taking the relationship between the text vector and the question-answer pair set as a text mapping relationship, and storing the text mapping relationship and the text vector in the vector database.

[0066] In some optional implementations of the present disclosure, the above-mentioned obtaining community vectors based on a graph database and storing community mapping relationships and community vectors in a vector database include: clustering the nodes in the graph database to obtain communities; obtaining community reports based on the communities; obtaining community summaries based on the community reports; obtaining community vectors based on the community summaries; taking the relationship between the community vector and a set of graph structure elements as a community mapping relationship, and storing the community mapping relationship and community vectors in the vector database.

[0067] In this optional implementation, the nodes are graph structure elements in the graph database. By clustering the nodes, clustering information of graph structure elements of the same type can be obtained, and the clustering information is used as the community.

[0068] In this optional implementation, the community report is a description of the relevant content of the community, and the community report can be a structured description. For example, a community report in an insurance contract is as follows:

[0069] {"title":"Termination Application and Supporting Documents","summary":"The community revolves around termination application and supporting documents. Termination application refers to the written application document for termination of insurance contract, while supporting documents are documents used to prove the identity of the applicant. There is a close relationship between these two entities, and supporting documents are required when applying for termination."}

[0070] In this optional implementation, after obtaining the community, a community report may be obtained by describing specific information of the community.

[0071] In this optional implementation, the community summary is a short text of the community report that briefly and clearly records the important contents of the document. The community summary is obtained by summarizing the community report.

[0072] In this optional implementation, the community summary is vectorized to obtain a community vector.

[0073] The method for obtaining community vectors and community mapping relationships provided by this optional implementation method clusters the nodes in a graph database to obtain communities; obtains community reports based on the communities; obtains community summaries based on the community reports; obtains community vectors based on the community summaries; uses the relationship between the community vector and a set of graph structure elements as a community mapping relationship, and stores the community mapping relationship and the community vector in a vector database; combines the social vector after clustering the graph data elements with the problem vector, and stores them in the vector database to improve the diversity of data in the vector database.

[0074] In some optional implementations of the present disclosure, the above-mentioned clustering the nodes in the graph database to obtain the community includes: in the graph database, clustering the nodes using the Leiden community detection algorithm to obtain the community.

[0075] In this optional implementation, the Leiden community detection algorithm (Leiden algorithm) generates the community division results in the network. The algorithm finds communities by optimizing the modularity of the network, where modularity is defined as "the number of edges in the community minus the expected number of edges in an equivalent network with randomly placed edges".

[0076] The method for obtaining a community provided by this optional implementation adopts the Leiden community detection algorithm to cluster the nodes in the graph database to obtain the community, thereby improving the reliability of obtaining the community.

[0077] In some optional implementations of the present disclosure, the above-mentioned obtaining a community report based on the community includes: inputting the community into a report macro model, and obtaining the community report output by the report macro model.

[0078] In this optional implementation, the report big model is a big model that describes the community in detail and generates a community report. When the community is input into the report big model, report prompt words can be simultaneously input into the report big model so that the report big model generates a community report based on the currently input community through the report prompt words.

[0079] The method for obtaining a community report provided by this optional implementation adopts a large report model to obtain a community report, thereby improving the reliability of obtaining the community report.

[0080] In some optional implementations of the present disclosure, obtaining the community summary based on the community report includes: inputting the community report into a summary macro model to obtain the community summary output by the summary macro model.

[0081] In this optional implementation, the summary big model is a big model that analyzes the main content of the community report and generates a community summary. When the community report is input into the summary big model, summary prompt words can be simultaneously input into the summary big model to prompt the summary big model through the summary prompt words to generate a community summary based on the currently input community report.

[0082] The method for obtaining a community summary provided by this optional implementation adopts a large summary model to obtain the community summary, thereby improving the reliability of obtaining the community summary.

[0083] In an example of the present disclosure, a data generation method corresponding to a data generation process structure diagram is shown as follows: Figure 2 As shown in the figure, the entire indexing process requires two storage components, vector database and graph database, and two sets of model services, large model LLM (Large Language Model) and text vectorization model Emb. The specific process is as follows:

[0084] The acquired documents are parsed, sliced, and other operations to obtain a text unit set docs. For each text unit in the text unit set, the large model is used to perform question-answer pair mining and entity, relationship, and claim mining, respectively, to obtain a question-answer pair set FAQs and a graph structure element set including entities, relationships, and claims, and a graph structure mapping relationship (not shown in the figure). For each question in the question-answer pair set FAQs, the text vectorization model is used to obtain the question vector and store it in the vector database. For each graph structure element in the graph structure element set, namely, entities, relationships, and claims, on the one hand, the text vectorization model is used to obtain the corresponding graph structure vector, and the above question vector and graph structure vector are stored in the vector database as text vectors, and the text mapping relationship is stored in the vector database (not shown in the figure). On the other hand, the graph structure element set is stored in the graph database.

[0085] In the graph database, the Leiden community detection algorithm is used to cluster nodes to obtain communities, and a community report is generated through a large model. The large model then generates a community summary, which is then vectorized and stored in a vector database. The community mapping relationship (not shown in the figure) is also stored in the vector database.

[0086] Figure 3 A process 300 according to an embodiment of a data query method of the present disclosure is shown. The data query method comprises the following steps:

[0087] Step 301, obtaining issues to be processed.

[0088] In this embodiment, the pending issues are issues raised by users or other entities, and the execution entity on which the data generation method runs can obtain the pending issues in a variety of ways. For example, the execution entity can obtain the pending issues stored in the database server through a wired connection or a wireless connection. For another example, the execution entity obtains the pending issues raised by the user through the terminal by communicating with the terminal.

[0089] Step 302, based on the problem to be processed, the graph database and the vector database, obtain the graph structure elements to be checked and the problem to be checked.

[0090] In this embodiment, the graph database and the vector database may be databases generated by the embodiment of the above-mentioned data generation method.

[0091] In this embodiment, the above step 302 includes: performing vectorization processing on the problem to be processed to obtain the vector to be processed, matching the vector to be processed with the information in the graph database and the vector database respectively to obtain the graph structure element to be checked and the problem vector to be checked; performing vector inverse operation on the problem vector to be checked to obtain the problem to be checked. The vector inverse operation is also called reverse quantization or devectorization. In practice, reverse quantization often involves data restoration and visualization.

[0092] In this embodiment, the graph structure element to be checked is a graph structure element matched by the problem to be processed and belongs to the graph database; the problem to be checked is a problem matched by the problem to be processed, belongs to the vector database and is obtained through vectorized inverse operation.

[0093] Step 303, based on the graph structure elements to be checked, the questions to be checked, the graph structure mapping relationships, the text mapping relationships and the community mapping relationships, obtain the query results of the questions to be processed.

[0094] In this embodiment, the graph structure mapping relationship is used to characterize the relationship between a set of text units and a set of graph structure elements, the text mapping relationship is used to characterize the relationship between a text vector and a set of question-answer pairs, and the community mapping relationship is used to characterize the relationship between a community vector and a set of graph structure elements; the graph structure mapping relationship, the text mapping relationship, and the community mapping relationship may be the graph structure mapping relationship, the text mapping relationship, and the community mapping relationship generated by the embodiment of the above-mentioned data generation method.

[0095] In this embodiment, the graph structure mapping relationship is used to characterize the relationship between a text unit set and a graph structure element set; the text mapping relationship is used to characterize the relationship between a text vector and a question-answer pair set; and the community mapping relationship is used to characterize the relationship between a community vector and a graph structure element set.

[0096] In this embodiment, the above step 303 includes: querying the graph structure mapping relationship through the graph structure element to be checked, and obtaining a set of text units corresponding to the graph structure element to be checked; querying the text mapping relationship through the text unit set corresponding to the graph structure element to be checked, and obtaining a set of question-answer pairs corresponding to the text unit set; querying the community mapping relationship through the graph structure element to be checked, and obtaining a community vector corresponding to the graph structure element to be checked, performing a vector inverse operation on the community vector to obtain a community; and using the text unit set corresponding to the graph structure element to be checked, the question-answer pair set corresponding to the text unit set, and the community as the query result of the problem to be processed.

[0097] The data query method provided by the embodiment of the present disclosure first obtains the problem to be processed; secondly, based on the problem to be processed, the graph database and the vector database, the graph structure elements to be checked and the problem to be checked are obtained; finally, based on the graph structure elements to be checked, the problem to be checked, the graph structure mapping relationship, the text mapping relationship and the community mapping relationship, the query result of the problem to be processed is obtained. In this way, more comprehensive information can be obtained from the graph structure elements and the problem than by querying the problem alone; based on the different mapping relationships of the graph structure elements and the problem query, more accurate query results can be obtained than text vector retrieval, which improves the comprehensiveness of the retrieval; when the query results are given to the large model, the large model can generate accurate results.

[0098] In some embodiments of the present disclosure, the above-mentioned obtaining the graph structure elements to be checked and the questions to be checked based on the questions to be processed, the graph database and the vector database includes: extracting keywords in the questions to be processed to obtain a keyword set; based on the keyword set, querying the graph database to obtain a first entity list; vectorizing the questions to be processed to obtain the vectors to be processed; based on the vectors to be processed and the vector database, obtaining the questions to be checked and the second entity list; based on the first entity list and the second entity list, obtaining the graph structure elements to be checked.

[0099] In this optional implementation, keywords are words that reflect the core of the problem to be processed, and a keyword set of the problem to be processed can be extracted through a commonly used keyword extraction algorithm; the keyword set includes at least one keyword.

[0100] In this optional implementation, the above-mentioned querying the graph database based on the keyword set to obtain the first entity list includes: matching each keyword in the keyword set with the graph structure elements in the graph database to obtain a set of matching graph structure elements; selecting entities in the matching graph structure element set to obtain the first entity list.

[0101] In this optional implementation, the above-mentioned obtaining of questions to be checked and the second entity list based on the vector to be processed and the vector database includes: matching the vector to be processed with the question vector in the vector database to obtain a matched question vector; performing a vector inverse operation on the matched question vector to obtain questions to be checked; matching the vector to be processed with the graph structure element vector in the vector database to obtain a matched graph structure element vector, and filtering the entities in the matched graph structure element vector to obtain a second entity list.

[0102] In this optional implementation, obtaining the graph structure element to be checked based on the first entity list and the second entity list includes: combining the first entity list and the second entity list to obtain the graph structure element to be checked.

[0103] This optional implementation provides a method for obtaining graph structure elements to be checked and questions to be checked, extracting keywords from the questions to be processed to obtain a keyword set; based on the keyword set, querying a graph database to obtain a first entity list; vectorizing the questions to be processed to obtain vectors to be processed; based on the vectors to be processed and the vector database, obtaining questions to be checked and a second entity list; based on the first entity list and the second entity list, obtaining graph structure elements to be checked, thereby improving the reliability of obtaining graph structure elements to be checked and questions to be checked.

[0104] In some optional implementations of the present disclosure, obtaining the graph structure element to be queried based on the first entity list and the second entity list includes: combining the first entity list and the second entity list to obtain a mixed entity list; and removing duplicate entities in the mixed entity list to obtain the graph structure element to be queried.

[0105] In this optional implementation, the mixed entity list is a list obtained by combining the first entity list and the second entity list.

[0106] The method for obtaining the graph structure element to be checked provided by this optional implementation combines the first entity list and the second entity list to obtain a mixed entity list; removes duplicate entities in the mixed entity list to obtain the graph structure element to be checked, which can make the content in the graph structure element to be checked non-repetitive and improve the quality of obtaining the graph structure element to be checked.

[0107] Optionally, the above-mentioned obtaining of the graph structure element to be queried based on the first entity list and the second entity list includes: combining the first entity list and the second entity list to obtain a mixed entity list; removing duplicate entities in the mixed entity list to obtain a mixed entity set, and selecting other graph structure elements related to each mixed entity in the mixed entity set from the graph database to obtain the graph structure element to be queried.

[0108] In some optional implementations of the present disclosure, the query results of the questions to be processed based on the graph structure elements to be checked, the questions to be checked, the graph structure mapping relationships, the text mapping relationships and the community mapping relationships include: obtaining a set of candidate text units based on the questions to be checked and the text mapping relationships; obtaining a set of candidate graph structures based on the graph structure elements to be checked and the graph structure mapping relationships; obtaining a set of candidate community reports based on the graph structure elements to be checked and the community mapping relationships; filtering and reordering the candidate text unit set, the candidate graph structure set and the candidate community report set to obtain the query results of the questions to be processed.

[0109] In this optional implementation, the above-mentioned obtaining of the candidate text unit set based on the question to be checked and the text mapping relationship includes: matching the question to be checked with the questions in the question-answer pair set in the text mapping relationship to obtain matching questions, querying the text mapping relationship through the matching questions to obtain a text unit set corresponding to the matching questions, and using the text unit set corresponding to the matching questions as the candidate text unit set.

[0110] In this optional implementation, the above-mentioned obtaining of the candidate graph structure set based on the graph structure elements to be checked and the graph structure mapping relationship includes: matching the graph structure elements to be checked with the graph structure elements in the graph structure mapping relationship to obtain a set of matched graph structure elements, and the set of matched graph structure elements is used as the candidate graph structure set.

[0111] In this optional implementation, the above-mentioned obtaining of a candidate community report set based on the graph structure elements to be queried and the community mapping relationship includes: matching the graph structure elements to be queried with the set of graph structure elements in the community mapping relationship to obtain a matching community vector; based on the matching community vector, generating a matching community report, and using the matching community report as a candidate community report set.

[0112] In this optional implementation, the above-mentioned filtering and reordering of the candidate text unit set, the candidate graph structure set, and the candidate community report set to obtain the query result of the problem to be processed includes: filtering the candidate text unit set, the candidate graph structure set, and the candidate community report set that are not related to the problem to be processed, and removing duplicate content after filtering to obtain the query result of the problem to be processed.

[0113] The method for obtaining query results provided by this optional implementation method obtains a set of candidate text units based on the question to be queried and the text mapping relationship; obtains a set of candidate graph structures based on the graph structure elements to be queried and the graph structure mapping relationship; obtains a set of candidate community reports based on the graph structure elements to be queried and the community mapping relationship; filters and reorders the candidate text unit set, the candidate graph structure set, and the candidate community report set to obtain the query result of the question to be processed, and obtains the query result of the question to be processed from multiple aspects, thereby improving the comprehensiveness of the information in the query result.

[0114] In an example of the present disclosure, a data query process structure diagram corresponding to the data query method is as follows: Figure 4As shown, the problem to be processed is mined through a large model to obtain a set of keywords involved, and the keyword set is used as a candidate entity list to query the first entity list that is actually hit from the graph database; the problem to be processed obtains a text vector through a text vectorization model, and this text vector is used to query the vector database for related questions to be queried in the question-answer pair set, and obtains a related second entity list, and the second entity list is combined with the first entity list to remove duplications to obtain the hit graph structure elements to be queried.

[0115] According to the mapping relationship in the records in the database (such as Figure 4 The problem to be processed and the graph structure elements to be queried are respectively converted into a set of candidate text units, a set of candidate community reports, and a set of candidate graph structures. After filtering and re-ranking, the final query result is obtained.

[0116] The data generation method and data query method provided in the present disclosure are key links in the RAG technical paradigm. The above two methods can be applied in the following fields:

[0117] In the enterprise knowledge management system, realize intelligent knowledge retrieval and sharing, intelligent question and answer and problem solving, knowledge graph construction and intelligent recommendation and other functions

[0118] In the online question-and-answer system, it provides functions such as automatic question-and-answer and customer service, internal knowledge sharing and collaboration, education and learning assistance, etc.

[0119] In the customer service system, it can help answer customer questions and ensure the accuracy of answers by real-time retrieval of document libraries. It is suitable for the field of customer service systems.

[0120] In addition, it can also play a practical value and potential in areas such as smart medical consultation, education and training, legal consultation and news interpretation.

[0121] Further references Figure 5 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a data generating device, which is similar to Figure 1 Corresponding to the method embodiment shown, the device can be specifically used in various electronic devices.

[0122] like Figure 5As shown, the data generation device 500 provided in this embodiment includes: a slicing unit 501, an analysis unit 502, a quantization unit 503, and a community unit 504. The slicing unit 501 can be configured to perform slicing processing on the acquired document to obtain a text unit set. The analysis unit 502 can be configured to obtain a question-answer pair set and a graph structure element set based on the text unit set, and store the graph structure element set and the graph structure mapping relationship in the graph database, and the graph structure mapping relationship is used to characterize the relationship between the text unit set and the graph structure element set. The quantization unit 503 can be configured to obtain a question vector and a graph structure vector based on the question-answer pair set and the graph structure element set, and use the question vector and the graph structure vector as text vectors, and store the text mapping relationship and the text vector in the vector database, and the text mapping relationship is used to characterize the relationship between the text vector and the question-answer pair set. The community unit 504 can be configured to obtain a community vector based on the graph database, and store the community mapping relationship and the community vector in the vector database, and the community mapping relationship is used to characterize the relationship between the community vector and the graph structure element set.

[0123] In this embodiment, the specific processing of the sharding unit 501, the analysis unit 502, the quantification unit 503, and the community unit 504 and the technical effects thereof can be referred to in detail. Figure 1 The relevant descriptions of step 101, step 102, step 103, and step 104 in the corresponding embodiment are not repeated here.

[0124] In some optional implementations of the present disclosure, the above-mentioned slicing unit 501 is configured to: parse the acquired document to obtain a parsed text; and slice the parsed text to obtain a text unit set including at least one text unit.

[0125] In some optional implementations of this embodiment, the above-mentioned analysis unit 502 is configured to: generate a question-answer pair set including at least one question-answer pair based on a large language model and a text unit set; generate a graph structure element set including at least one graph structure element based on a large language model and a text unit set; use the relationship between the text unit set and the graph structure element set as a graph structure mapping relationship, and store the graph structure mapping relationship and the graph structure element set in a graph database.

[0126] In some optional implementations of the present embodiment, the quantization unit 503 is configured to: extract questions from the question-answer pair set; use a text vectorization model to vectorize the questions to obtain a question vector; use a text vectorization model to vectorize a set of graph structure elements to obtain a graph structure vector, and use the question vector and the graph structure vector as text vectors; use the relationship between the text vector and the question-answer pair set as a text mapping relationship, and store the text mapping relationship and the text vector in a vector database.

[0127] In some optional implementations of this embodiment, the community unit 504 is configured to: cluster the nodes in the graph database to obtain communities; obtain community reports based on the communities; obtain community summaries based on the community reports; obtain community vectors based on the community summaries; use the relationship between the community vector and the set of graph structure elements as a community mapping relationship, and store the community mapping relationship and the community vector in a vector database.

[0128] In some optional implementations of this embodiment, the community unit 504 is configured to: in a graph database, cluster nodes using a Leiden community detection algorithm to obtain communities.

[0129] In some optional implementations of this embodiment, the community unit 504 is configured to: input the community into the reporting model to obtain the community report output by the reporting model, and the reporting model generates the community report based on the input community.

[0130] In some optional implementations of this embodiment, the community unit 504 is configured to: input the community report into the summary model to obtain the community summary output by the summary model, and the summary model is used to generate the community summary based on the input community report.

[0131] The data generation device provided by the embodiment of the present disclosure is as follows: first, the slicing unit 501 performs slicing processing on the acquired document to obtain a text unit set; secondly, the analyzing unit 502 obtains a question-answer pair set and a graph structure element set based on the text unit set, and stores the graph structure element set and the graph structure mapping relationship in the graph database; thirdly, the quantization unit 503 obtains a question vector and a graph structure vector based on the question-answer pair set and the graph structure element set, uses the question vector and the graph structure vector as text vectors, and stores the text mapping relationship and the text vector in the vector database; finally, the community unit 504 obtains a community vector based on the graph database, and stores the community mapping relationship and the community vector in the vector database. Thus, on the basis of the acquired document, various forms of information to be retrieved, such as graph structure elements, text vectors, and community vectors, can be generated, providing a reliable query basis for data query.

[0132] Further references Figure 6As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a data query device, which is similar to Figure 3 Corresponding to the method embodiment shown, the device can be specifically used in various electronic devices.

[0133] like Figure 6 As shown, the data query device 600 provided in this embodiment includes: an acquisition unit 601, a question acquisition unit 602, and a result acquisition unit 603. Among them, the above-mentioned acquisition unit 601 can be configured to acquire the question to be processed. The above-mentioned question acquisition unit 602 can be configured to obtain the graph structure elements to be checked and the question to be checked based on the question to be processed, the graph database and the vector database. The above-mentioned result acquisition unit 603 can be configured to obtain the query result of the question to be processed based on the graph structure elements to be checked, the question to be checked, the graph structure mapping relationship, the text mapping relationship and the community mapping relationship; the graph structure mapping relationship is used to characterize the relationship between the text unit set and the graph structure element set, the text mapping relationship is used to characterize the relationship between the text vector and the question-answer pair set, and the community mapping relationship is used to characterize the relationship between the community vector and the graph structure element set.

[0134] In this embodiment, in the data generating device 600, the specific processing of the acquisition unit 601, the problem obtaining unit 602, and the result obtaining unit 603 and the technical effects thereof can be referred to respectively. Figure 3 The relevant descriptions of step 301, step 302, and step 303 in the corresponding embodiment are not repeated here.

[0135] In some embodiments of the present disclosure, the above-mentioned question obtaining unit 602 is configured to: extract keywords in the question to be processed to obtain a keyword set; based on the keyword set, query the graph database to obtain a first entity list; vectorize the question to be processed to obtain a vector to be processed; based on the vector to be processed and the vector database, obtain the question to be checked and the second entity list; based on the first entity list and the second entity list, obtain the graph structure element to be checked.

[0136] In some optional implementations of this embodiment, the problem obtaining unit 602 is configured to: combine the first entity list and the second entity list to obtain a mixed entity list; remove duplicate entities in the mixed entity list to obtain the graph structure element to be checked.

[0137] In some optional implementations of the present embodiment, the above-mentioned result obtaining unit 603 is configured to: obtain a set of candidate text units based on the question to be checked and the text mapping relationship; obtain a set of candidate graph structures based on the graph structure elements to be checked and the graph structure mapping relationship; obtain a set of candidate community reports based on the graph structure elements to be checked and the community mapping relationship; filter and reorder the candidate text unit set, the candidate graph structure set and the candidate community report set to obtain the query result of the question to be processed.

[0138] The data query device provided by the embodiment of the present disclosure, first, the acquisition unit 601 acquires the problem to be processed; secondly, the problem acquisition unit 602 obtains the graph structure elements to be checked and the problem to be checked based on the problem to be processed, the graph database and the vector database; finally, the result acquisition unit 603 obtains the query result of the problem to be processed based on the graph structure elements to be checked, the problem to be checked, the graph structure mapping relationship, the text mapping relationship and the community mapping relationship. Thus, more comprehensive information can be obtained from the graph structure elements and the problem than by using the problem query alone; based on the different mapping relationships of the graph structure elements and the problem query, more accurate query results can be obtained than the text vector retrieval, which improves the comprehensiveness of the retrieval; when the query results are given to the large model, the large model can generate accurate results.

[0139] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0140] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0141] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0142] like Figure 7As shown, the device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0143] A number of components in the device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0144] The computing unit 701 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 701 performs the various methods and processes described above, such as a data generation method or a data query method. For example, in some embodiments, the data generation method or the data query method may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the data generation method or the data query method described above may be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to execute the data generating method or the data query method in any other appropriate manner (eg, by means of firmware).

[0145] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0146] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer or other programmable data generation device or a data query device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, partially on the machine as a stand-alone software package and partially on a remote machine, or entirely on a remote machine or server.

[0147] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0148] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0149] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an information server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communications network). Examples of communications networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0150] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0151] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0152] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A data generation method, the method comprising: Slice the acquired document to obtain a set of text units; Based on the text unit set, a question-answer pair set and a graph structure element set are obtained, and the graph structure element set and the graph structure mapping relationship are stored in a graph database, where the graph structure mapping relationship is used to characterize the relationship between the text unit set and the graph structure element set; Based on the question-answer pair set and the graph structure element set, a question vector and a graph structure vector are obtained, the question vector and the graph structure vector are used as text vectors, and a text mapping relationship and the text vector are stored in a vector database, wherein the text mapping relationship is used to characterize the relationship between the text vector and the question-answer pair set; Based on the graph database, a community vector is obtained, and a community mapping relationship and the community vector are stored in the vector database, wherein the community mapping relationship is used to characterize the relationship between the community vector and the graph structure element set.

2. The method according to claim 1, wherein: The obtained document is segmented to obtain a text unit set including: Parse the acquired document to obtain the parsed text; The parsed text is sliced ​​to obtain a text unit set including at least one text unit.

3. The method according to claim 1, wherein: The step of obtaining a question-answer pair set and a graph structure element set based on the text unit set, and storing the graph structure element set and the graph structure mapping relationship in a graph database includes: Based on the large language model and the text unit set, generating a question-answer pair set including at least one question-answer pair; Based on the large language model and the text unit set, generating a graph structure element set including at least one graph structure element; The relationship between the text unit set and the graph structure element set is used as a graph structure mapping relationship, and the graph structure mapping relationship and the graph structure element set are stored in a graph database.

4. The method according to claim 1, wherein: The obtaining of a question vector and a graph structure vector based on the question-answer pair set and the graph structure element set, taking the question vector and the graph structure vector as text vectors, and storing the text mapping relationship and the text vector in a vector database comprises: Extracting questions from the question-answer pair set; Using a text vectorization model to vectorize the question to obtain a question vector; The text vectorization model is used to vectorize the graph structure element set to obtain a graph structure vector, and the question vector and the graph structure vector are used as text vectors; The relationship between the text vector and the question-answer pair set is used as a text mapping relationship, and the text mapping relationship and the text vector are stored in the vector database.

5. The method according to claim 1, wherein: The obtaining of the community vector based on the graph database and storing the community mapping relationship and the community vector in the vector database comprises: For the nodes in the graph database, clustering the nodes to obtain communities; Based on the community, a community report is obtained; Based on the community report, a community summary is obtained; Based on the community summary, a community vector is obtained; The relationship between the community vector and the graph structure element set is used as a community mapping relationship, and the community mapping relationship and the community vector are stored in the vector database.

6. The method according to claim 5, wherein: The clustering of the nodes in the graph database to obtain communities includes: In the graph database, the Leiden community detection algorithm is used to cluster the nodes to obtain communities.

7. The method according to claim 5, wherein: The obtaining of a community report based on the community includes: The community is input into the reporting model to obtain a community report output by the reporting model. The reporting model generates a community report based on the input community.

8. The method according to claim 5, wherein: The obtaining of a community summary based on the community report comprises: The community report is input into a summary macromodel to obtain a community summary output by the summary macromodel, wherein the summary macromodel is used to generate a community summary based on the input community report.

9. A data query method, the method comprising: Get pending issues; Based on the problem to be processed, the graph database and the vector database, a graph structure element to be checked and a problem to be checked are obtained; Based on the graph structure elements to be checked, the questions to be checked, the graph structure mapping relationship, the text mapping relationship and the community mapping relationship, a query result of the question to be processed is obtained; The graph structure mapping relationship is used to characterize the relationship between a text unit set and a graph structure element set, the text mapping relationship is used to characterize the relationship between a text vector and a question-answer pair set, and the community mapping relationship is used to characterize the relationship between a community vector and the graph structure element set.

10. The method according to claim 9, wherein: The obtaining of the graph structure elements to be checked and the questions to be checked based on the questions to be processed, the graph database and the vector database comprises: Extracting keywords from the problem to be processed to obtain a keyword set; Based on the keyword set, query the graph database to obtain a first entity list; Vectorizing the problem to be processed to obtain a vector to be processed; Based on the vector to be processed and the vector database, obtaining the questions to be checked and the second entity list; Based on the first entity list and the second entity list, a graph structure element to be searched is obtained.

11. The method according to claim 10, wherein: The obtaining of the to-be-queried graph structure element based on the first entity list and the second entity list comprises: Combining the first entity list and the second entity list to obtain a mixed entity list; The duplicate entities in the mixed entity list are removed to obtain the graph structure elements to be searched.

12. The method according to claim 9, wherein: The query result of the problem to be processed obtained based on the graph structure element to be checked, the problem to be checked, the graph structure mapping relationship, the text mapping relationship and the community mapping relationship includes: Based on the question to be checked and the text mapping relationship, a candidate text unit set is obtained; Based on the graph structure element to be checked and the graph structure mapping relationship, a candidate graph structure set is obtained; Based on the to-be-queried graph structure element and the community mapping relationship, a candidate community report set is obtained; The candidate text unit set, the candidate graph structure set, and the candidate community report set are filtered and reordered to obtain a query result of the problem to be processed.

13. A data generating device, comprising: A fragmentation unit is configured to fragment the acquired document to obtain a text unit set; an analyzing unit configured to obtain a question-answer pair set and a graph structure element set based on the text unit set, and store the graph structure element set and the graph structure mapping relationship in a graph database, wherein the graph structure mapping relationship is used to characterize the relationship between the text unit set and the graph structure element set; a quantization unit configured to obtain a question vector and a graph structure vector based on the question-answer pair set and the graph structure element set, use the question vector and the graph structure vector as text vectors, and store a text mapping relationship and the text vector in a vector database, wherein the text mapping relationship is used to characterize a relationship between the text vector and the question-answer pair set; The community unit is configured to obtain a community vector based on the graph database, and store a community mapping relationship and the community vector in the vector database, wherein the community mapping relationship is used to characterize the relationship between the community vector and the graph structure element set.

14. A data query device, comprising: An acquisition unit, configured to acquire a problem to be processed; A question obtaining unit is configured to obtain a graph structure element to be checked and a question to be checked based on the question to be processed, the graph database and the vector database; A result obtaining unit is configured to obtain a query result of the problem to be processed based on the graph structure element to be checked, the problem to be checked, the graph structure mapping relationship, the text mapping relationship and the community mapping relationship; The graph structure mapping relationship is used to characterize the relationship between a text unit set and a graph structure element set, the text mapping relationship is used to characterize the relationship between a text vector and a question-answer pair set, and the community mapping relationship is used to characterize the relationship between a community vector and the graph structure element set.

15. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 12.

16. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable the computer to execute the method according to any one of claims 1 to 12.

17. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method according to any one of claims 1 to 12.