Text generation graph query statement-based model training method and device, computer equipment, readable storage medium and program product
By generating sample graph query statements from knowledge graphs and using fine-tuning models, the accuracy and efficiency issues of enterprise data queries are solved, achieving efficient and accurate graph query statement generation and meeting the business needs of enterprise data queries.
Patent Information
- Application Number
- CN202511010200.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-11-07
AI Technical Summary
Existing technologies lack sufficient semantic understanding and execution precision for multi-table join queries in enterprise data retrieval, resulting in low query accuracy. Furthermore, existing text-to-image query statements have low accuracy and are difficult to meet business needs.
By acquiring enterprise data from a pre-built knowledge graph, sample graph query statements are generated. Then, by using a fine-tuned graph query statement to text model and text to graph query statement model, graph query statements corresponding to user questions are generated. The model parameters are optimized by combining preset prompts and preference data to improve query accuracy.
It improves the efficiency and accuracy of enterprise data queries, reduces model prediction time, covers a variety of query scenarios, and generates more accurate graph query statements to meet business needs.
Smart Images

Figure CN120910309A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a training method, apparatus, computer device, computer-readable storage medium, and computer program product for a model based on text-generated graph query statements. Background Technology
[0002] In enterprise data management scenarios, databases are typically used to store massive amounts of enterprise data. This data includes information on hundreds of dimensions such as shareholder structure, patent technology, legal cases, and product services, forming hundreds or even thousands of interconnected SQL (Structured Query Language) tables.
[0003] To provide users with the ability to query enterprise data within chat conversations, it is necessary to understand the semantics of user questions. The traditional process involves first converting the user question into an SQL query statement, then obtaining the results and generating an answer through multi-table join queries. However, existing technologies have significant shortcomings in semantic understanding and execution accuracy of multi-table join queries, making it difficult to meet actual business needs.
[0004] In contrast, graph queries have significant advantages. The characteristics of knowledge graphs, which represent entities with nodes and relationships with edges, can intuitively present the complex relationships between enterprise data. At the same time, key information can be located through path search. However, the accuracy of converting text into graph query statements by existing technologies is low, resulting in low query accuracy for enterprise data. Summary of the Invention
[0005] Therefore, it is necessary to provide a training method, apparatus, computer equipment, computer-readable storage medium, and computer program product for a text-based graph query statement model that can improve the accuracy of enterprise data querying, addressing the aforementioned technical problems.
[0006] Firstly, this application provides a training method for a model based on text-generated graph query statements, the method comprising:
[0007] Enterprise data is obtained from a pre-built knowledge graph, and corresponding sample graph query statements are generated based on the enterprise data;
[0008] The graph query statement to text conversion model is called to process the sample graph query statement to obtain multiple corresponding sample question texts. The graph query statement to text conversion model is obtained by fine-tuning the first pre-trained model.
[0009] Fine-tune a pre-trained second large model based on the sample graph query statement and the corresponding plurality of sample question texts, to obtain a text-to-graph query statement model; wherein the text-to-graph query statement model is used to process question texts input by a user, and outputs corresponding graph query statements, the question texts are used to request target enterprise data, and the graph query statements are used to query the target enterprise data in the knowledge graph.
[0010] In one of the embodiments, after the calling of the graph query statement-to-text model processing the sample graph query statement to obtain the corresponding plurality of sample question texts, the following steps are included:
[0011] Based on the preset prompt words and the pre-trained third large model, the plurality of sample question texts are respectively rewritten and expanded to obtain updated plurality of sample question texts.
[0012] In one of the embodiments, the training method of the graph query statement-to-text model includes:
[0013] Obtain preference data for each sample question text in the plurality of sample question texts;
[0014] Based on the sample graph query statement, the plurality of sample question texts corresponding to the sample graph query statement, and the preference data of each sample question text, update the graph query statement-to-text model.
[0015] In one of the embodiments, based on the sample graph query statement, the plurality of sample question texts corresponding to the sample graph query statement, and the preference data of each sample question text, updating the graph query statement-to-text model includes:
[0016] Based on the sample graph query statement, the plurality of sample question texts corresponding to the sample graph query statement, and the preference data of each sample question text, the graph query statement-to-text model is updated by using a direct preference optimization algorithm.
[0017] In one of the embodiments, before the step of obtaining enterprise data from the pre-constructed knowledge graph and generating corresponding sample graph query statements based on the enterprise data, the following step is included:
[0018] Based on a plurality of data tables to be queried, a knowledge graph is constructed, and the plurality of data tables are used to store enterprise data.
[0019] In one of the embodiments, after the step of fine-tuning a pre-trained second large model based on the sample graph query statement and the corresponding plurality of sample question texts to obtain a text-to-graph query statement model, the following step is included:
[0020] In the case of user inputting question text, the text-to-graph query statement model is invoked to process the question text, to obtain a corresponding graph query statement, the question text being used to request target enterprise data;
[0021] Based on the graph query statement, the target enterprise data is queried in the knowledge graph;
[0022] A reply text including the target enterprise data is generated and output.
[0023] In a second aspect, the application further provides a training device of a model for generating a graph query statement based on text, comprising:
[0024] The device comprises:
[0025] A training data generation module is configured to obtain enterprise data from a pre-constructed knowledge graph, and generate a corresponding sample graph query statement based on the enterprise data;
[0026] The training data generation module is further configured to invoke a graph query statement-to-text model to process the sample graph query statement, to obtain a plurality of sample question texts corresponding thereto, the graph query statement-to-text model being obtained by fine-tuning a pre-trained first large model;
[0027] A training module is configured to fine-tune a pre-trained second large model based on the sample graph query statement and the plurality of sample question texts, to obtain a text-to-graph query statement model; wherein the text-to-graph query statement model is used to process user input question text, to output a corresponding graph query statement, the question text being used to request target enterprise data, and the graph query statement being used to query the target enterprise data in the knowledge graph.
[0028] In a third aspect, the application further provides a computer device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the steps of the method of any one of the above aspects when executing the computer program.
[0029] In a fourth aspect, the application further provides a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the steps of the method of any one of the above aspects.
[0030] In a fifth aspect, the application further provides a computer program product comprising a computer program, the computer program being executed by a processor to implement the steps of the method of any one of the above aspects.
[0031] The training method and device of the model for generating a graph query statement based on text, the computer device, the computer readable storage medium, and the computer program product can generate sample graph query statements through knowledge graph sampling, which can provide materials for paired generation of training data and comprehensively cover various supported query scenarios. The text-to-graph query statement model trained based on these data can accurately convert user problems into graph query statements, and has small model parameter size and shorter prediction time compared with a model relying on a prompt word and a large model, thereby improving the efficiency and accuracy of business data query. BRIEF DESCRIPTION OF DRAWINGS
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the drawings needed to be used in the description of the embodiments of the present application or the related art will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other related drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0033] Figure 1 A flowchart of a training method of a model for generating a graph query statement based on text in an embodiment;
[0034] Figure 2 A flowchart of a training method of a model for generating a graph query statement based on text in another embodiment;
[0035] Figure 3 A flowchart of a training method of a model for generating a graph query statement based on text in another embodiment;
[0036] Figure 4 A flowchart of a training method of a model for generating a graph query statement based on text in another embodiment;
[0037] Figure 5 A flowchart of a training method of a model for generating a graph query statement based on text in another embodiment;
[0038] Figure 6 A structural block diagram of a training device of a model for generating a graph query statement based on text in an embodiment;
[0039] Figure 7 An internal structure diagram of a computer device in an embodiment. DETAILED DESCRIPTION
[0040] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.
[0041] It should be noted that the terms "first", "second", etc. used in the present application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "include" and "have" and any variations thereof used in the present application are intended to cover non-exclusive inclusion. The term "a plurality of" used in the present application means two and more than two. The term "and / or" used in the present application means one of the options or any combination of multiple options.
[0042] If hundreds of SQL tables are constructed into a knowledge graph, a large parameter model such as deepseek-v3, qwen2.5-72b-instruct is used, and a prompt is used to let it directly generate a cypher statement according to a user question (graph query), the accuracy is poor. Moreover, when the number of nodes and edges of the knowledge graph is large, the length of the prompt will increase dramatically if the information is converted into context and included in the prompt, and the model needs to process too much content, which leads to low accuracy of the generated cypher statement in actual business, and cannot meet the business use requirements. Even if the prompt is optimized and a small amount of samples are added to improve the effect, the actual improvement is very limited. In addition, this way does not support training the model through data preparation, and the prediction time is long.
[0043] If a small parameter model such as qwen2.5-7b-instruct is selected for fine-tuning to obtain a special model with the ability to generate graph query statements based on text (text2cypher), a training data set containing pairs of user questions and cypher statements needs to be constructed. But this way faces the problems of difficulty in constructing the training data set, high cost of manual annotation, and lack of user question samples, and the amount of data that can be prepared is limited, and the coverage of user questions is narrow.
[0044] Based on this, the embodiment of the present application provides a training method of a model for generating graph query statements based on text. The method is applied to a server for example in the embodiment, and it can be understood that the method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is realized through the interaction of the terminal and the server. As shown in the embodiment, Figure 1 The method includes the following steps:
[0045] Step S102, obtaining enterprise data from the pre-built knowledge graph, and generating a corresponding sample graph query statement based on the enterprise data.
[0046] In the pre-built knowledge graph, the enterprise data can be represented in the form of nodes and edges.
[0047] Exemplarily, the corresponding node data and / or edge data can be sampled from the pre-built knowledge graph multiple times according to preset strategy rules to construct corresponding sample graph query statements. The preset strategy rules can include node and / or edge filtering conditions. The sample graph query statement can be a cypher statement or other GQL (Graph Query Language) statement.
[0048] Step S104, calling a graph query statement to text model to process the sample graph query statement to obtain a corresponding plurality of sample question texts, the graph query statement to text model being obtained based on fine-tuning of a pre-trained first large model.
[0049] The pre-trained first large model includes but is not limited to qwen2.5-7b-instruct or qwen2.5-14b-instruct. The graph query statement to text model is a proprietary model with the ability to generate text based on graph query statements (cypher2text).
[0050] Exemplarily, the graph query statement to text model can be trained in a cold start manner. Specifically, a small batch of question texts can be prepared, and paired data of text and cypher statements can be obtained by manual annotation. The pre-trained first large model is fine-tuned based on these data to obtain the graph query statement to text model. The graph query statement to text model can be used to generate multiple texts based on one cypher statement.
[0051] Further, for each text in the plurality of texts, rewriting and expansion can be performed to obtain a corresponding plurality of updated texts.
[0052] Step S106, fine-tuning a pre-trained second large model based on the sample graph query statement and the corresponding plurality of sample question texts to obtain a text to graph query statement model; wherein the text to graph query statement model is used to process user input question texts to output corresponding graph query statements, the question texts are used to request target enterprise data, and the graph query statements are used to query target enterprise data in the knowledge graph.
[0053] Exemplarily, quality evaluation of the multiple sample question texts corresponding to each sample graph query sentence can be firstly performed in terms of semantic matching degree, grammatical correctness and the like, and labels can be added to the multiple sample question texts according to the quality evaluation results. Then, the pre-trained second large model is iteratively updated based on the multiple sample graph query sentences and the multiple sample question texts corresponding thereto with labels, to obtain a text-to-graph query sentence model.
[0054] In the training method of the text-to-graph query sentence model, sample graph query sentences are generated by means of knowledge graph sampling, which can not only provide materials for paired generation of training data, but also comprehensively cover various supported query scenarios; a large number of texts corresponding to sample graph query sentences are generated by means of the fine-tuned graph query sentence-to-text model, which can effectively expand the training data pairs. The text-to-graph query sentence model trained based on these data can accurately convert user questions into graph query sentences, and has small model parameter size and shorter prediction time compared with a model relying on a prompt word and a large model, thereby improving the efficiency and accuracy of business data query.
[0055] In an exemplary embodiment, to further increase the number and diversity of training data, as shown in Figure 2 The training method of the text-to-graph query sentence model further includes:
[0056] In step S105, the multiple sample question texts are respectively rewritten and expanded based on the preset prompt word and the pre-trained third large model, to obtain updated multiple sample question texts.
[0057] In a possible implementation, for each sample question text in the multiple sample question texts, the pre-trained third large model can be used to rewrite the sample question text in a synonymous manner, convert a sentence of the sample question text, or parse the semantics of the sample question text, and then reorganize the sample question text based on a grammar rule, to obtain the corresponding updated multiple sample question texts.
[0058] It should be noted that the pre-set first large model, the second large model and the third first large model can be the same or different large models.
[0059] In an exemplary embodiment, the text-to-graph query sentence model can be iteratively optimized based on the obtained training data, as shown in Figure 3 The training method of the text-to-graph query sentence model includes:
[0060] In step A1, preference data for each sample question text in the multiple sample question texts is obtained.
[0061] Step A2, updating the graph query sentence to text model based on the sample graph query sentence, the plurality of sample question texts corresponding to the sample graph query sentence, and the preference data of each sample question text.
[0062] The preference data can be determined based on manual annotation. Alternatively, the quality of the sample question texts can be verified by using the trained text to graph query sentence model in reverse, and the preference data can be indirectly determined.
[0063] For example, the graph query sentence to text model can be updated by using reinforcement learning based on the sample graph query sentence, the plurality of sample question texts corresponding to the sample graph query sentence, and the preference data of each sample question text. For example, the graph query sentence to text model can be updated by using a proximal policy optimization (PPO) algorithm.
[0064] Optionally, as shown in Figure 4 Step A2 can include:
[0065] Step A21, updating the graph query sentence to text model by using a direct preference optimization algorithm based on the sample graph query sentence, the plurality of sample question texts corresponding to the sample graph query sentence, and the preference data of each sample question text.
[0066] For example, the graph query sentence to text model can be trained by using multiple rounds of self-iteration direct preference optimization (DPO) based on the sample graph query sentence, the plurality of sample question texts corresponding to the sample graph query sentence, and the preference data of each sample question text.
[0067] In a possible implementation, after the first large model is fine-tuned based on the small-batch text and cypher sentence pair data, and the graph query sentence to text model is obtained, the cypher sentences generated by sampling the knowledge graph can be input into the graph query sentence to text model to obtain a plurality of corresponding texts; the preferred data pairs (each cypher sentence can correspond to a high-quality text and a low-quality text) can be determined from the data, the DPO training is performed on the graph query sentence to text model, and an updated graph query sentence to text model is obtained; and the step of expanding the training data and iteratively optimizing the updated graph query sentence to text model is repeatedly performed. Specifically, the step of expanding the training data and iteratively optimizing the updated graph query sentence to text model can include: sampling the knowledge graph, inputting the generated cypher sentences into the updated graph query sentence to text model to obtain a plurality of corresponding texts; obtaining the preferred data pairs by manual annotation, rewriting and expanding based on the preset prompt and the large model, and using the data pairs to perform DPO training on the updated graph query sentence to text model.
[0068] In the embodiment, the model is fine-tuned by preparing small-batch training data to obtain an initial graph query sentence to text capability, and then the graph query sentences generated by sampling the knowledge graph are expanded to generate paired data, and the model is reinforced to generate higher-quality user questions, so that a graph query sentence to text proprietary model with excellent generation quality is obtained, and reliable support is provided for the generation of paired data.
[0069] In an example embodiment, as shown in Figure 5 The training method of the text-based graph query sentence generation model further includes:
[0070] In step S101, a knowledge graph is constructed based on a plurality of data tables to be queried, and the plurality of data tables are used to store enterprise data.
[0071] In a possible implementation, enterprise data can be read from a plurality of data sources and preliminarily cleaned; then the fields in different data tables are mapped to entities and relationship types of the knowledge graph; and the Neo4j graph database is used to store the knowledge graph, and the processed data is imported into the knowledge graph.
[0072] In an example embodiment, please continue to refer to Figure 5 The training method of the text-based graph query sentence generation model further includes:
[0073] Step S1071, in the case of user inputting question text, calling a text-to-graph query statement model to process the question text to obtain a corresponding graph query statement, the question text being used to request target enterprise data.
[0074] Step S1072, querying the target enterprise data in the knowledge graph based on the graph query statement.
[0075] Step S1073, generating and outputting a reply text including the target enterprise data.
[0076] Exemplarily, the above-mentioned model for generating a graph query statement based on text can be used in a chat dialogue robot in the business search field to provide the user with the ability of retrieving a company and querying specific information of the company.
[0077] In summary, in the training method of the above-mentioned model for generating a graph query statement based on text, sample graph query statements are generated by knowledge graph sampling, which can not only provide materials for the paired generation of training data, but also comprehensively cover various supported query scenarios; a large number of sample graph query statements corresponding to texts are generated by means of the graph query statement-to-text specific model obtained through fine-tuning, which can effectively expand the training data pairs. The text-to-graph query statement specific model trained based on these data can accurately convert user questions into graph query statements, and the model parameter size is small, the time-consuming of prediction is shorter compared with relying on prompt words and large models, thereby improving the efficiency and accuracy of business data query.
[0078] It should be understood that, although each step in the flowchart involved in each embodiment as described above is displayed in sequence according to the arrow, these steps are not necessarily executed in sequence according to the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, at least part of the steps in the flowchart involved in each embodiment as described above can include multiple steps or stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these steps or stages is not necessarily sequential, but can be alternately or alternately executed with at least part of other steps or steps or stages in other steps. It can be understood that the steps in different embodiments can be freely combined as needed, and various non-contradictory schemes formed by combination are within the scope of protection of the present application.
[0079] Based on the same inventive concept, the embodiment of the present application also provides a device for training a model for generating a graph query statement based on text. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more device embodiments for training a model for generating a graph query statement based on text provided below can be referred to the limitations of the training method for the model for generating a graph query statement based on text in the above, which will not be repeated here.
[0080] In one exemplary embodiment, as shown in Figure 6 A device 200 for training a model for generating a graph query statement based on text is provided, comprising: a training data generation module 201 and a training module 202, wherein:
[0081] The training data generation module 201 is configured to obtain enterprise data from a pre-built knowledge graph, and generate corresponding sample graph query statements based on the enterprise data.
[0082] The training data generation module 201 described above is further configured to call a graph query statement to text model to process the sample graph query statements, to obtain corresponding multiple sample question texts, and the graph query statement to text model is obtained based on fine-tuning of a pre-trained first large model.
[0083] The training module 202 is configured to fine-tune a pre-trained second large model based on the sample graph query statements and the corresponding multiple sample question texts, to obtain a text to graph query statement model; wherein the text to graph query statement model is configured to process question texts input by a user, and output corresponding graph query statements, the question texts are configured to request target enterprise data, and the graph query statements are configured to query the target enterprise data in the knowledge graph.
[0084] In one embodiment, the training data generation module 201 described above is further configured to:
[0085] Based on a pre-set prompt word and a pre-trained third large model, the multiple sample question texts are respectively rewritten and expanded, to obtain updated multiple sample question texts.
[0086] In one embodiment, the training module 202 described above is further configured to:
[0087] Obtain preference data for each sample question text in the multiple sample question texts;
[0088] Based on the sample graph query statements, the multiple sample question texts corresponding to the sample graph query statements, and the preference data of each sample question text, update the graph query statement to text model.
[0089] In one embodiment, the training module 202 described above is further configured to:
[0090] Based on the sample graph query statement, the plurality of sample question texts corresponding to the sample graph query statement, and the preference data of each sample question text, a direct preference optimization algorithm is used to update the graph query statement to text model.
[0091] In one embodiment, the training device 200 of the model for generating a graph query statement based on text described above further comprises:
[0092] A knowledge graph construction module is configured to construct a knowledge graph based on a plurality of data tables to be queried, and the plurality of data tables are used to store enterprise data.
[0093] In one embodiment, the training device 200 of the model for generating a graph query statement based on text described above further comprises:
[0094] A graph query statement generation module is configured to, in the case that a user inputs a question text, call the text to graph query statement model to process the question text, and obtain a corresponding graph query statement, and the question text is used to request target enterprise data.
[0095] A query module is configured to query target enterprise data in the knowledge graph based on the graph query statement.
[0096] A response module is configured to generate and output a response text including the target enterprise data.
[0097] Each module in the training device of the model for generating a graph query statement based on text described above can be realized by software, hardware, and a combination thereof, in whole or in part. Each module described above can be embedded in or independent of a processor in a computer device in a hardware form, or can be stored in a memory in a computer device in a software form, so as to be called and executed by a processor to perform the operations corresponding to each module.
[0098] In one exemplary embodiment, a computer device is provided, which can be a server, and an internal structure diagram of the computer device can be as shown in Figure 7As shown in the figure. The computer device includes a processor, a memory, an input / output interface (I / O for short) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through the system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capability. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the text-to-map query statement model. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with the terminal outside through the network connection. The computer program is executed by the processor to implement a method for training a model for generating a map query statement based on text.
[0099] Those skilled in the art can understand that, Figure 7 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0100] In one exemplary embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps in each of the method embodiments described above.
[0101] In one embodiment, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement the steps in each of the method embodiments described above.
[0102] In one embodiment, a computer program product is provided, including a computer program, and the computer program is executed by a processor to implement the steps in each of the method embodiments described above.
[0103] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant regulations.
[0104] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. The non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. The volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, the RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., without being limited thereto.
[0105] The technical features of the above embodiments can be combined in any manner. To make the description concise, not all possible combinations of the technical features in the above embodiments are described, but as long as the combinations of the technical features do not exist contradictions, they should be considered as the scope of the present application.
[0106] The above-described embodiments are merely illustrative of several embodiments of the present application, and the description is relatively specific and detailed, but should not be understood as a limitation on the scope of the patent. It should be noted that for those skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the scope of the present application. Therefore, the scope of protection of the present application should be subject to the appended claims.
Claims
1. A method for training a model for generating a graph query statement based on text, characterized in that, The method comprises: obtaining enterprise data from a pre-built knowledge graph, and generating a corresponding sample graph query statement based on the enterprise data; calling a graph query statement to text model to process the sample graph query statement to obtain a plurality of corresponding sample question texts, the graph query statement to text model being obtained based on fine-tuning of a first large model; fine-tuning a second large pre-trained model based on the sample graph query statement and the plurality of corresponding sample question texts to obtain a text to graph query statement model, wherein the text to graph query statement model is used to process user input question texts to output corresponding graph query statements, the question texts are used to request target enterprise data, and the graph query statements are used to query the target enterprise data in the knowledge graph.
2. The method of claim 1, wherein, After the calling of the graph query statement to text model to process the sample graph query statement to obtain the plurality of corresponding sample question texts, the method comprises: rewriting and expanding the plurality of sample question texts based on preset prompt words and a third large pre-trained model to obtain updated plurality of sample question texts.
3. The method according to claim 1 or 2, characterized in that, The training method of the graph query statement to text model comprises: obtaining preference data for each sample question text in the plurality of sample question texts; updating the graph query statement to text model based on the sample graph query statement, the plurality of corresponding sample question texts of the sample graph query statement, and the preference data of each sample question text.
4. The method of claim 3, wherein, The updating of the graph query statement to text model based on the sample graph query statement, the plurality of corresponding sample question texts of the sample graph query statement, and the preference data of each sample question text comprises: updating the graph query statement to text model using a direct preference optimization algorithm based on the sample graph query statement, the plurality of corresponding sample question texts of the sample graph query statement, and the preference data of each sample question text.
5. The method of claim 1, wherein, Before the obtaining of the enterprise data from the pre-built knowledge graph and the generating of the corresponding sample graph query statement based on the enterprise data, the method comprises: constructing a knowledge graph based on a plurality of data tables to be queried, wherein the plurality of data tables are used to store enterprise data.
6. The method of claim 1, wherein, After the fine-tuning of the second large pre-trained model based on the sample graph query statement and the plurality of corresponding sample question texts to obtain the text to graph query statement model, the method comprises: in the case that a user inputs a question text, calling the text to graph query statement model to process the question text to obtain a corresponding graph query statement, wherein the question text is used to request target enterprise data; querying the target enterprise data in the knowledge graph based on the graph query statement; generating and outputting a reply text comprising the target enterprise data. 7.A device for training a model for generating a graph query statement based on text, characterized by, The device comprises: a training data generation module configured to obtain enterprise data from a pre-built knowledge graph, and generate a corresponding sample graph query statement based on the enterprise data; The training data generation module is further configured to invoke a graph query statement to text model to process the sample graph query statement, to obtain a plurality of sample question texts corresponding to the sample graph query statement, and the graph query statement to text model is obtained based on fine-tuning of a first large pre-trained model; The training module is configured to fine-tune a second large pre-trained model based on the sample graph query statement and the plurality of sample question texts, to obtain a text to graph query statement model, wherein the text to graph query statement model is configured to process a question text input by a user, and output a corresponding graph query statement, the question text is configured to request target enterprise data, and the graph query statement is configured to query the target enterprise data in the knowledge graph. 8.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-7. The processor executes the computer program to implement the steps of the method of any one of claims 1 to 6.
9. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 6.
10. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 6.
Citation Information
Patent Citations
Information query method
CN116595026A
Knowledge graph query method and device, equipment and storage medium
CN118760775A
Mechanical intelligent question and answer maintenance method and system based on large language model and knowledge graph double-wheel driving
CN119003709A
Knowledge graph analysis and question answering method and device based on large language model
CN119691243A
Retrieval enhancement generation method and device based on knowledge graph, equipment and medium
CN120277206A