Text-to-Cypher semantic analysis model generation method based on graph enhancement and LLM fine tuning

By building a knowledge graph summary vector database and a LoRA fine-tuning model, the problems of entity linking and syntax parsing in graph databases are solved, and efficient and accurate Text-to-Cypher semantic parsing is achieved, which is suitable for graph database queries.

CN120633667APending Publication Date: 2025-09-12HUAQIAO UNIVERSITY
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510745310.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing technologies make it difficult to accurately perform entity linking and grammatical parsing in graph databases, resulting in low accuracy and efficiency of Text-to-Cypher semantic parsing.

Method used

By building a knowledge graph summary vector database, combining the question type classifier and ICIO framework prompt words, and using the LoRA fine-tuning model for semantic parsing, we ensure entity linking and grammatical correctness.

Benefits of technology

It achieves efficient and accurate semantic parsing in graph databases, lowers the query threshold, enables non-technical personnel to easily use graph databases, and improves the accuracy and reliability of queries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633667A_ABST
    Figure CN120633667A_ABST
Patent Text Reader

Abstract

A Text-to-Cypher semantic analysis model generation method based on graph enhancement and LLM fine tuning relates to the technical field of semantic analysis, and comprises the following steps: extracting triples in a knowledge graph to construct a knowledge graph abstract vector database; inputting the retrieval question sentences into a trained question sentence type classifier for classification to obtain question sentence types; vectorizing the retrieval question sentences and sending the vectorized retrieval question sentences into a knowledge graph abstract vector database to obtain a retrieval question sentence related triple set; integrating the retrieval question, the retrieval question related triple set and the question type into an ICIO framework cue word; supervised data pairs formed by the ICIO framework cue words and the Cypher statements corresponding to the retrieval question sentences are sent into an analysis generation module, LoRA fine tuning is carried out, and a Text-to-Cypher semantic analysis model is obtained. According to the method and the device, the retrieval question can be automatically converted into the Cypher query statement, so that efficient query of the graph database is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of semantic parsing technology in natural language processing, and in particular to a method for generating a Text-to-Cypher semantic parsing model based on graph enhancement and LLM fine-tuning. Background Art

[0002] Text-to-Cypher semantic parsing involves mapping natural language questions into graph queries, which are then used to retrieve the desired results in the graph database Neo4j. This approach plays a crucial role in applications such as knowledge graph question answering. While extensive research has been conducted on text-to-SQL query generation, this research primarily focuses on query generation for relational databases. However, query patterns and syntax differ significantly between relational databases and graph data, making them difficult to translate well to text-to-Cypher semantic parsing. Furthermore, improper entity linking and grammatical errors are long-standing issues. Large language models (LLMs) struggle to accurately link entities without detailed domain knowledge. However, if knowledge about natural language questions is fed into the LLM as partial context, the LLM can effectively perform entity linking. Furthermore, incorporating the type of natural language questions into the model input can improve the grammatical accuracy of semantic parsing. Therefore, to address these issues and shortcomings, a text-to-Cypher semantic parsing model that can properly link entities and address grammatical issues is urgently needed. Summary of the Invention

[0003] The purpose of the present invention is to provide a method for generating a Text-to-Cypher semantic parsing model based on graph enhancement and LLM fine-tuning, so as to solve the problems of improper entity linking and grammar in Text-to-Cypher semantic parsing.

[0004] The present invention adopts the following technical solutions:

[0005] A method for generating a Text-to-Cypher semantic parsing model based on graph enhancement and LLM fine-tuning, including:

[0006] S101, extracting triples from the knowledge graph to construct a knowledge graph summary vector database;

[0007] S102, inputting the search question into the trained question type classifier for classification to obtain the question type;

[0008] S103, vectorizing the search question and sending it to the knowledge graph summary vector database to obtain a set of triples related to the search question;

[0009] S104, integrating the search question, the set of triples related to the search question, and the question type into an ICIO framework prompt word;

[0010] S105 , the supervised data pairs consisting of the ICIO framework prompt words and the Cypher statements corresponding to the search questions are sent to the parsing generation module for LoRA fine-tuning to obtain a Text-to-Cypher semantic parsing model.

[0011] Preferably, the S101 specifically includes:

[0012] S1011, obtain the given knowledge graph G,

[0013] Among them, h i represents the i-th head entity in the knowledge graph G, h i _label represents the entity type of the i-th head entity in the knowledge graph G, r i represents the i-th relationship, r i _type indicates the relationship type of the i-th relationship, t i represents the i-th tail entity in the knowledge graph G, t i _label represents the entity type of the i-th tail entity in the knowledge graph G, and n represents the number of all triples;

[0014] S1012, extract all triples from the knowledge graph G. Each triple can be structurally represented as: triple i =(h i ,h i _label,r i ,r i _type,t i ,t i _label);

[0015] S1013, triple i The 3 element head entities h i , relationship r i and tail entity t i Spliced ​​into a complete sentence of knowledge, represented as text i , and organize this triple information into: tripledata i ={"text":text i ,"entity":[[h i _label,h i ],[t i _label,t i ]],"relation":[ri _type,r i ]}, and finally get the processed aggregated data D vec , D vec ={tripledata i |i=1,2,...,n};

[0016] S1014, to D vec In each data, an "id" field is added as a unique identifier. Then each data contains 4 fields {"id","text","entity","relation"}. i Indicates: data i ={"id":i,"text":text i ,"entity":[[h i _label,h i ],[t i _label,t i ]],"relation":[r i _type,r i ]}, then update D vec ={data i |i=1,2,...,n};

[0017] S1015, for D vec , tripledata i text in i Enter the text embedding model to get text i The vector representation W i ; data i Further update to data i ′=(id i ,W i ,text i ,metadata i ), id i Represents the i-th unique identifier; metadata i Refers to the set of entity and relation in the i-th data; then each data i ′ is stored in the vector database to obtain the knowledge graph summary vector database.

[0018] Preferably, the S102 specifically includes:

[0019] S1021, for category type train The training retrieval question q train, first call process_classify(·) to process, and segment, serialize and fill / crop to form the input sequence k. The calculation formula is shown in (1):

[0020] k=process_classify(q train )#(1)

[0021] S1022, the input sequence k is forward propagated and normalized by calling map_label(·) to obtain the classification output vector O, whose dimension is equal to T, where T is the total number of categories in the category table of the known search question; thereby mapping k to its index label type in the category table. train , the formula is shown in (2):

[0022] type train =map_label(k)#(2)

[0023] S1023, use cross entropy loss to calculate the loss between the true label and the model prediction output. The cross entropy loss calculation formula for a single sample is shown in (3):

[0024]

[0025] Among them, Y is the category label type train One-hot encoding, whose dimension is equal to the number of categories T, only when t is the true category y t =1, otherwise y t =0,y t Represents the binary value after one-hot encoding; o t is the predicted probability value of the model for each category;

[0026] For the question type classification training set, iterative optimization of multiple batches of data was carried out to finally obtain a trained question type classifier;

[0027] S1024, for any search question q in the graph search dataset new , after classification by the trained question type classifier, the question classification result is obtained, namely the question type type new .

[0028] Preferably, the S103 specifically includes:

[0029] S1031, the search question q new Feed it into the text embedding model to get the converted query vector W q ;

[0030] S1032, according to W qSimilarity search is performed in the construction of knowledge graph summary vector database. The similarity search process involves W q and construct vector W in the knowledge graph summary vector database i Cosine similarity between cosine_si·milarity(W q ,W i ) is calculated using formula (4), and returns the K data results data-K that are most similar to the query vector;

[0031]

[0032] S1033, extract all metadata in data-K k |k=1,2,...,K}, which includes the related entity set {entity k |k=1,2,...,K} and the relation set {relation k |k=1,2,...,K}; remove duplicates from the entity set and the relationship set to obtain the query-related triple sets entities and relations.

[0033] Preferably, the S104 specifically includes:

[0034] Call gen_prompt(·) to retrieve the question q new , retrieve the question-related triple sets entities and relations, as well as the question type type new The three parts are integrated into the ICIO framework prompt word prompt q , the calculation formula is shown in (5):

[0035] prompt q =gen_prompt(q new ,entities,relations,type new )#(5).

[0036] Preferably, the S105 specifically includes:

[0037] S1051, design LoRA fine-tuning model: select specific layers of the large model for LoRA fine-tuning; set the parameter weight of a single selected original linear layer to Both d and k represent dimensions. During fine-tuning, W0 remains unchanged and the weights are updated using the following formula (6). The new weight matrix W is used for inference:

[0038] W=W0+ΔW#(6)

[0039] in, is the weight update matrix in the fine-tuning training phase; ΔW is transformed into a matrix, that is, ΔW is represented by low-rank decomposition, which is expressed by formula (7):

[0040] ΔW=B·A#(7)

[0041] in, r represents rank;

[0042] The weight matrix after LoRA adjustment is expressed by formula (8):

[0043] W=W0+B·A#(8)

[0044] S1052, call the parsing generation module process_parse(·) to process prompt q And Cypher query statement C, prompt q After word segmentation, serialization, and padding / cropping with C, the input and output token IDs are obtained, which are the input sequence x=[x1,x2,...,x p ] and the output sequence y=[y1,y2,...,y c ], and take the output sequence y as the true label sequence, where p and c represent the input sequence length and output sequence length respectively; as shown in formulas (9) and (10):

[0045] x=process_parse(prompt q )#(9)

[0046] y=process_parse(C)#(10)

[0047] S1053, input x into the designed LoRA fine-tuning model, and obtain the output of the LoRA fine-tuning model through forward propagation

[0048] S1054, calculate the output using the cross entropy loss function The difference between the true label sequence y; assuming that the LoRA fine-tuning model predicts a probability distribution when generating the i-th token Indicates generating each possible token The probability of; let the size of the vocabulary be V, then Is a probability vector of length V; for the true label y i , maximize P(y i Specifically, first, the cross entropy loss of a single token in y is expressed as formula (11):

[0049] L i =-logP(yi |x)#(11)

[0050] Then calculate the total cross entropy loss of y using the formula shown in (12):

[0051]

[0052] described Used to normalize the loss to ensure fair comparison between sequences of different lengths;

[0053] S1055, calculate the total cross entropy loss L of the batch ′ , then the gradient of the low-rank matrices A and B and After multiple iterations of model optimization, the parameters of the low-rank matrices A and B are updated; finally, a trained Text-to-Cypher semantic parsing model based on graph enhancement using LoRA technology and LLM fine-tuning is obtained.

[0054] Preferably, after S1055, the step further includes:

[0055] S1056, for new arbitrary search question q test The processed prompt word prompt test , perform word segmentation and serialization to obtain the input sequence P test ;P test Send it to the trained Text-to-Cypher semantic parsing model to get the model output Will Convert to query statement C test .

[0056] Compared with the prior art, the present invention has the following beneficial effects:

[0057] (1) The knowledge graph summary vector database of the present invention stores triple information in a structured manner to ensure the accuracy of entity linking; the question type classifier can identify the search question type and guide Cypher to generate templates based on the question type to avoid grammatical structure errors. The ICIO framework prompt word combines the two to ensure that the LLM large language model receives the corresponding information, ultimately solving the problems of improper entity linking and grammar in Text-to-Cypher semantic parsing;

[0058] (2) The semantic parsing model fine-tuned by LoRA in the present invention can automatically convert the natural language questions input by the user into relevant and accurate Cypher statements, and use the statements to query the graph database without the need for professional learning of the Cypher language, thereby achieving efficient query of the graph database. It can not only improve the query efficiency and lower the query threshold, so that non-technical personnel can also easily use the graph database, but also reduce the error rate and improve the accuracy and reliability of the query. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 A flowchart of a method for generating a Text-to-Cypher semantic parsing model based on graph enhancement and LLM fine-tuning provided by an embodiment of the present invention;

[0060] Figure 2 A structural diagram of the Text-to-Cypher semantic parsing model based on graph enhancement and LLM fine-tuning provided in an embodiment of the present invention;

[0061] Figure 3 A schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0062] Below in conjunction with specific embodiment, further set forth the present invention.Should be understood that these embodiments are only used to illustrate the present invention and are not used in limiting the scope of the present invention.In addition, should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms fall equally within the scope limited by the appended claims of the application.

[0063] See also Figure 1 and Figure 2 As shown, this embodiment provides a method for generating a Text-to-Cypher semantic parsing model based on graph enhancement and LLM fine-tuning, including the following steps.

[0064] S101, extracting triples from the knowledge graph to construct a knowledge graph summary vector database;

[0065] S102, inputting the search question into the trained question type classifier for classification to obtain the question type;

[0066] S103, vectorizing the search question and sending it to the knowledge graph summary vector database to obtain a set of triples related to the search question;

[0067] S104, integrating the search question, the set of triples related to the search question, and the question type into an ICIO framework prompt word;

[0068] S105 , the supervised data pairs consisting of the ICIO framework prompt words and the Cypher statements corresponding to the search questions are sent to the parsing generation module for LoRA fine-tuning to obtain a Text-to-Cypher semantic parsing model.

[0069] In this example, the experimental data comes from the open-source SpCQL dataset for Text-to-Cypher research. This dataset consists of two main components: a knowledge graph G0 stored in Neo4j and 10,008 supervised datasets consisting of query sentences and their corresponding Cypher statements D0.

[0070] D0={(query1,cypher1,answer1),(query2,cypher2,answer2),…,(query i ,cypher i ,answer i ),…,(query 10008 ,cypher 10008 ,answer 10008 )}; Reduce the size of the knowledge graph G0 to G, which consists of 80,000 relation triples, 15,000 entities, 1 entity label, 3 attributes, and 4 relation types. Then, filter out the portion of D0 that is unrelated to the entities in the knowledge graph G to form a subset D1. Use the large-model distillation method to construct a question type classification dataset for the query queries in D1. The question category table has 8 categories: [["single-hop fact", 0], ["multiple facts", 1], ["constrained fact", 2], ["multiple intents", 3], ["string operation", 4], ["count", 5], ["boolean", 6], ["sort", 7]]. The remaining portion of D0, D2, is used as the graph query dataset, and D2 is divided into training and test sets in a ratio of 9:1.

[0071] The S101 is specifically implemented as follows.

[0072] S1011, given a knowledge graph G, The knowledge graph G includes 80,000 relation triples, 15,000 entities, 1 entity label, 3 attributes, and 4 relation types.

[0073] S1012, extract all triples from the knowledge graph G. The content of each triple can be structured as follows: Set triple i=("2012 Guangzhou International Yacht Show","ENTITY","Exhibition Venue","Relationship","China Import and Export Fair Pazhou Complex","ENTITY").

[0074] S1013, triple i The three elements in the head entity "2012 Guangzhou International Yacht Show", the relationship "exhibition location" and the tail entity "China Import and Export Fair Pazhou Complex" are spliced ​​into a complete sentence knowledge, which is represented as text i = "The venue of the 2012 Guangzhou International Yacht Show is the Pazhou Complex of the China Import and Export Fair." This triplet information is organized and represented as: tripledata i ={"text":"The venue of the 2012 Guangzhou International Yacht Show is the Pazhou Complex of the China Import and Export Fair","entity":[["ENTITY":"2012 Guangzhou International Yacht Show"],["ENTITY","Pazhou Complex of the China Import and Export Fair"]],"relation":["Relationship","Exhibition Venue"]}, and finally obtain the processed aggregated data D vec , D vec ={tripledata i |i=1,2,...,80000}.

[0075] S1014, to D vec In each data, an "id" field is added as a unique identifier. Then each data contains 4 fields {"id","text","entity","relation"}. i Indicates: data i ={"id":i,"text":"The venue of the 2012 Guangzhou International Yacht Show is the Pazhou Complex of the China Import and Export Fair","entity":[["ENTITY","2012 Guangzhou International Yacht Show"],["ENTITY","Pazhou Complex of the China Import and Export Fair"]],"relation":["Relationship","Exhibition Venue"]}, then update D vec ={data i |i=1,2,...,80000}.

[0076] S1015, for D vec , tripledata i text in i("The venue of the 2012 Guangzhou International Yacht Show is the Pazhou Complex of the China Import and Export Fair") is fed into the text embedding model TE-DOLLM-Qwen2.5-7B, and the text is obtained. i The vector representation W i ; data i Further update to data i ′={"id":i,"embeddings":W i ,"text":"The venue of the 2012 Guangzhou International Yacht Show is the Pazhou Complex of the China Import and Export Fair","entity":[["ENTITY","2012 Guangzhou International Yacht Show"],["ENTITY","Pazhou Complex of the China Import and Export Fair"]],"relation":["Relationship","Exhibition Venue"]}; Then each data i ′ is stored in the vector database to obtain the knowledge graph summary vector database NKGVecDB.

[0077] The S102 is specifically implemented as follows.

[0078] S1021, for the question type classification dataset, for a retrieval question q = "I gained a lot of knowledge at the 2011 Guangzhou International Energy Saving and Environmental Protection Technology Exhibition. Do you know the exhibition location?" with category = "single-hop fact", it is segmented, serialized, and padded (or cropped) to form an input sequence k.

[0079] S1022, the input sequence k is forward propagated and normalized by calling the map_label(·) module to obtain a classification output vector O, whose dimension is equal to 8, thereby mapping k to its index label y=0 in the category table.

[0080] In step S1023, the cross-entropy loss is used to calculate the loss between the true label and the model's predicted output. The cross-entropy loss for a single sample is calculated as H. For the question type classification training set D1, iterative optimization of multiple batches of data is performed to ultimately obtain a specialized question type classifier.

[0081] S1024, for any search question q in the known graph search dataset new , after classification by the trained question type classifier, the question classification result is obtained, namely the question type type new .

[0082] The S103 is specifically implemented as follows.

[0083] S1031: Send the query sentence q into the text embedding model TE-DOLLM-Qwen2.5-7B to obtain the converted query vector W.q .

[0084] S1032, according to W q Similarity search is performed in the constructed NKGVecDB. The similarity search process involves W q Calculate the cosine similarity between the vectors in NKGVecDB and return the four data results data-4 that are most similar to the query vector. data-4 = {"ids":[[100031,103022,42751,42753]],"distances":[[0.4510642886161804,0.7689331769943237,0.7702440023422241,0.7814088463783264]],"metadatas":[[ {"entity":[["ENTITY","2011 Guangzhou International Energy Saving and Environmental Protection Technology Exhibition"]],"relation":[["Relationship","Exhibition Venue"]]},{"entity":[["ENTITY","2010 China Guangzhou International Home Decorations and Home Textiles Exhibition"]],"relation":[["Relationship","Exhibition Venue"]]},{"entity":[["ENTITY","2012 China [Guangzhou] International Wire, Cable, Materials and Equipment Exhibition"],["ENTITY","China Import and Export Fair Pazhou Complex"]],"relation":["Relationship","Exhibition Venue"]},{"entity":[["ENTITY","2012 China [Guangzhou] International Wire, Cable, Materials and Equipment Exhibition"],["ENTITY","China Import and Export Fair Pazhou Complex"]],"relation":["Relationship","Exhibition Venue"]}] ],"documents":[["The 2011 Guangzhou International Energy Saving and Environmental Protection Technology Exhibition will be held at the China Import and Export Fair Complex.","The 2010 China Guangzhou International Home Decorations and Home Textile Fabrics Exhibition will be held at the China Import and Export Fair Complex.","The 2012 China [Guangzhou] International Wire, Cable, Materials and Equipment Exhibition will be held at the China Import and Export Fair Pazhou Complex.","The 2012 China [Guangzhou] International Wire, Cable, Materials and Equipment Exhibition will be held at Pazhou, Guangzhou City, Guangdong Province, China."]]}.

[0085] S1033, extract all metadata in data-4 "metadatas": [[{"entity":[["ENTITY","2011 Guangzhou International Energy Saving and Environmental Protection Technology Exhibition"]],"relation":[["Relationship","Exhibition Location"]]},{"entity":[["ENTITY","2010 China Guangzhou International Home Decorations and Home Textiles Exhibition"]],"relation":[["Relationship","Exhibition Location"]]},{"entity":[["ENTITY","2012 China [Guangzhou] International Wire & Cable and Materials Equipment Exhibition"],["ENTITY","China Import and Export Fair Pazhou Complex"]],"relation":["Relationship","Exhibition Venue"]},{"entity":[["ENTITY","2012 China [Guangzhou] International Wire & Cable and Materials Equipment Exhibition"],["ENTITY","Pazhou, Guangzhou, Guangdong, China"]],"relation":['Relationship","Exhibition Venue"]}]], which includes the related entity collection [["ENTITY","2011 Guangzhou International Energy Conservation and Environmental Protection Technology Exhibition ["ENTITY","2010 China Guangzhou International Home Decorations and Home Textiles Exhibition"],["ENTITY","2012 China [Guangzhou] International Wire and Cable and Materials Equipment Exhibition"],["ENTITY","China Import and Export Fair Pazhou Complex"],["ENTITY","2012 China [Guangzhou] International Wire and Cable and Materials Equipment Exhibition"],["ENTITY","Pazhou, Guangzhou, Guangdong, China"]] and the relationship set [["Relationship","Exhibition Venue"],["Relationship","Exhibition Venue"],["Relationsh ip","Exhibition Venue"],["Relationship","Exhibition Venue"]]; remove duplicates from the entity set and relationship set to obtain the related triple sets entities = ["2011 Guangzhou International Energy Saving and Environmental Protection Technology Exhibition","2010 China Guangzhou International Home Decorations and Home Textile Fabrics Exhibition","2012 China [Guangzhou] International Wire and Cable and Materials Equipment Exhibition","China Import and Export Fair Pazhou Complex","2012 China [Guangzhou] International Wire and Cable and Materials Equipment Exhibition","Pazhou, Guangzhou, Guangdong, China"] and relations = ["Exhibition Venue","Exhibition Venue","Exhibition Venue"].

[0086] The S104 is specifically implemented as follows.

[0087] S1041, according to the ICIO framework design concept, explains the semantic parsing instructions respectively, provides the background information of the graph database, informs the semantic parsing model of the input data to be processed, and guides the output query statement Cypher, Cypher = "match(:ENTITY{name:'2011 Guangzhou International Energy Saving and Environmental Protection Technology Exhibition'})-[:Relationship{name:'Exhibition Location'}]->(n)return n.name"; then calls the gen_prompt prompt word generation module to retrieve the question q, the query question related triples entities and relations, and the question type type new The three parts are integrated into the ICIO framework prompt word prompt, prompt = {

[0088] "instruction":

[0089] "You are an expert in neo4j knowledge graph question answering. You will receive a text containing a natural language question and the entity nodes and relationships that may be involved in the question. Please output the Cypher query statement corresponding to this natural language question."

[0090] "context":

[0091] "Possibly required related entities: \"2011 Guangzhou International Energy Saving and Environmental Protection Technology Exhibition\", \"2010 China Guangzhou International Home Decorations and Home Textile Fabrics Exhibition\", \"China Import and Export Fair Pazhou Complex\", \"2012 China [Guangzhou] International Wire and Cable and Materials Equipment Exhibition\", \"Pazhou, Guangzhou, Guangdong, China\"; Possibly required related relationships: \"Exhibition Venue\", \"Exhibition Venue\", \"Exhibition Venue Hall\"; Question Category: \"Single-hop Fact\"",

[0092] "input":

[0093] Question: I learned a lot at the 2011 Guangzhou International Energy Conservation and Environmental Protection Technology Exhibition. Do you know where the exhibition was held?

[0094] "output":

[0095] }.

[0096] The S105 is specifically implemented as follows.

[0097] S1051, select specific layers of the large model (such as word embedding layer, attention layer, etc.) for LoRA fine-tuning, including query projection layer (q_proj), key projection layer (k_proj), value projection layer (v_proj), output projection layer (o_proj), gated projection layer (gate_proj), upsampling projection layer (up_proj) and downsampling projection layer (down_proj), a total of 7 layers. The parameter weights of the original linear layer are LoRA transforms the weight update matrix ΔW of the original linear layer, that is, it is decomposed into and After LoRA adjustment, the parameter weight W0 of the original linear layer is updated to be

[0098] S1052, call the parsing generation module process_parse(·) to process prompt q And Cypher query statement C, prompt q After word segmentation, serialization and padding (or clipping) with C, the input and output token IDs are obtained, which are the input sequence x=[x1,x2,...,x p ] and the output sequence y=[y1,y2,...,y c ], where p and c represent the input sequence length and output sequence length respectively.

[0099] S1053, input x into the LoRA fine-tuning model designed in step 5.1, and obtain the output of the model through forward propagation

[0100] S1054, calculate model output using cross entropy loss function The difference between the true label y. Assume that when the model generates the i-th token, it predicts a probability distribution Indicates generating each possible token The probability of . Assuming the vocabulary size is 152064, then Is a probability vector of length 152064. For the true label y i , we hope that the model maximizes P(y i |x), which is the probability of generating the correct token. First, the cross entropy loss L for a single token in y is i Then calculate the normalized total cross entropy loss L of y.

[0101] S1055, for the graph query dataset D2, calculate the total cross entropy loss L for each batch ′ , then the gradient of the low-rank matrices A and B and After multiple iterations of model optimization, the parameters of the low-rank matrices A and B are updated; ultimately, a Text-to-Cypher semantic parsing model based on graph enhancement using LoRA technology and LLM fine-tuning is obtained.

[0102] S1056, for new arbitrary search question q test The processed prompt word prompt test , perform word segmentation and serialization to obtain the input sequence P test ;P test Send it to the trained Text-to-Cypher semantic parsing model to get the model output Will Convert to Cypher query statement test .

[0103] Figure 3 FIG. 1 is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention. Figure 3 As shown, the electronic device of this embodiment includes: a processor 301 and a memory 302; the memory 302 is configured to store computer-executable instructions; and the processor 301 is configured to execute the computer-executable instructions stored in the memory to implement the various steps performed by the electronic device in the above embodiment. For details, please refer to the description of the embodiment of the method for generating a text-to-cypher semantic parsing model based on graph enhancement and LLM fine-tuning.

[0104] Optionally, the memory 302 may be independent or integrated with the processor 301 .

[0105] When the memory 302 is independently provided, the electronic device further includes a bus 303 for connecting the memory 302 and the processor 301 .

[0106] An embodiment of the present invention also provides a computer storage medium, which stores computer execution instructions. When the processor 301 executes the computer execution instructions, the above-mentioned Text-to-Cypher semantic parsing model generation method based on graph enhancement and LLM fine-tuning is implemented.

[0107] An embodiment of the present invention further provides a computer program product, including a computer program. When the computer program is executed by the processor 301, the above method is implemented.

[0108] In the embodiments provided herein, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the module division is merely a logical functional division. In actual implementation, other division methods may be used. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not implemented. In addition, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interface, device or module, which may be electrical, mechanical or other forms.

[0109] Modules described as separate components may or may not be physically separate, and components shown as modules may or may not be physical units, that is, they may be located in one place or distributed across multiple network elements. Some or all of these modules may be selected to implement the solution of this embodiment based on actual needs.

[0110] In addition, the functional modules in various embodiments of the present invention may be integrated into a single processing unit, each module may exist physically separately, or two or more modules may be integrated into a single unit. The units formed by the above modules may be implemented in the form of hardware or hardware plus software functional units.

[0111] The above-mentioned integrated module implemented in the form of a software function module can be stored in a computer-readable storage medium. The above-mentioned software function module is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or processor 301 to perform some steps of the methods of various embodiments of the present application.

[0112] It should be understood that the processor 301 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), or application-specific integrated circuits (ASIC). A general-purpose processor may be a microprocessor, or the processor 301 may be any conventional processor 301. The steps of the method disclosed in the present invention may be directly implemented by the hardware processor 301 or implemented by a combination of hardware and software modules in the processor 301.

[0113] The memory 302 may include a high-speed RAM memory, and may also include a non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disk.

[0114] Bus 303 can be an Industry Standard Architecture (ISA), a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Bus 303 can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, the bus 303 in the drawings of this application is not limited to a single bus 303 or a single type of bus 303.

[0115] The storage medium may be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage medium may be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0116] An exemplary storage medium is coupled to the processor 301, so that the processor 301 can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor 301. The processor 301 and the storage medium can be located in an application-specific integrated circuit (ASIC). Of course, the processor 301 and the storage medium can also exist as discrete components in an electronic device or a host control device.

[0117] Those skilled in the art will appreciate that all or part of the steps in the above-described method embodiments can be implemented using hardware associated with program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0118] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for generating a Text-to-Cypher semantic parsing model based on graph enhancement and LLM fine-tuning, characterized by: include: S101, extracting triples from the knowledge graph to construct a knowledge graph summary vector database; S102, inputting the search question into the trained question type classifier for classification to obtain the question type; S103, vectorizing the search question and sending it to the knowledge graph summary vector database to obtain a set of triples related to the search question; S104, integrating the search question, the set of triples related to the search question, and the question type into an ICIO framework prompt word; S105 , the supervised data pairs consisting of the ICIO framework prompt words and the Cypher statements corresponding to the search questions are sent to the parsing generation module for LoRA fine-tuning to obtain a Text-to-Cypher semantic parsing model.

2. The method for generating a Text-to-Cypher semantic parsing model based on graph enhancement and LLM fine-tuning according to claim 1, characterized in that: The S101 specifically includes: S1011, obtain the given knowledge graph G, Among them, h i represents the i-th head entity in the knowledge graph G, h i _label represents the entity type of the i-th head entity in the knowledge graph G, r i represents the i-th relationship, r i _type indicates the relationship type of the i-th relationship, t i represents the i-th tail entity in the knowledge graph G, t i _label represents the entity type of the i-th tail entity in the knowledge graph G, and n represents the number of all triples; S1012, extract all triples from the knowledge graph G. Each triple can be structurally represented as: triple i =(h i ,h i _label,r i ,r i _type,t i ,t i _label); S1013, triple i The 3 element head entities h i , relationship r i and tail entity t i Spliced ​​into a complete sentence of knowledge, represented as text i , and organize this triple information into: tripledata i ={"text":text i ,"entity":[[h i _label,h i ],[t i _label,t i ]],"relation":[r i _type,r i ]}, and finally get the processed aggregated data D vec , D vec ={tripledata i |i=1,2,...,n}; S1014, to D vec In each data, an "id" field is added as a unique identifier. Then each data contains 4 fields {"id","text","entity","relation"}. i Indicates: data i ={"id":i,"text":text i ,"entity":[[h i _label,h i ],[t i _label,t i ]],"relation":[r i _type,r i ]}, then update D vec ={data i |i=1,2,...,n}; S1015, for D vec , tripledata i text in i Enter the text embedding model to get text i The vector representation W i ; data i Further update to data i ′=(id i ,W i ,text i ,metadata i ), id i Represents the i-th unique identifier; metadata i Refers to the set of entity and relation in the i-th data; then each data i ′ is stored in the vector database to obtain the knowledge graph summary vector database.

3. The method for generating a Text-to-Cypher semantic parsing model based on graph enhancement and LLM fine-tuning according to claim 1, characterized in that: The S102 specifically includes: S1021, for category type train The training retrieval question q train First, the question classification module process_classify(·) is called to process, and the input sequence k is formed by word segmentation, serialization, and padding / cropping. The calculation formula is shown in (1): k=process_classify(q train )#(1) S1022, the input sequence k is forward propagated and normalized by calling map_label(·) to obtain the classification output vector O, whose dimension is equal to T, where T is the total number of categories in the category table of the known search question; thereby mapping k to its index label type in the category table. train , the formula is shown in (2): type train =map_label(k)#(2) S1023, use cross entropy loss to calculate the loss between the true label and the model prediction output. The cross entropy loss calculation formula for a single sample is shown in (3): Among them, Y is the category label type train One-hot encoding, whose dimension is equal to the number of categories T, only when t is the true category y t =1, otherwise y t =0,y t Represents the binary value after one-hot encoding; o t is the predicted probability value of the model for each category; For the question type classification training set, iterative optimization of multiple batches of data was carried out to finally obtain a trained question type classifier; S1024, for any search question q in the graph search dataset new , after classification by the trained question type classifier, the question classification result is obtained, namely the question type type new .

4. The method for generating a Text-to-Cypher semantic parsing model based on graph enhancement and LLM fine-tuning according to claim 1, characterized in that: The S103 specifically includes: S1031, the search question q new Feed it into the text embedding model to get the converted query vector W q ; S1032, according to W q Similarity search is performed in the construction of knowledge graph summary vector database. The similarity search process involves W q and construct vector W in the knowledge graph summary vector database i Cosine similarity between cosine_si·milarity(W q ,W i ) is calculated using formula (4), and returns the K data results data-K that are most similar to the query vector; S1033, extract all metadata in data-K k |k=1,2,...,K}, which includes the related entity set {entity k |k=1,2,...,K} and the relation set {relation k |k=1,2,...,K}; remove duplicates from the entity set and the relationship set to obtain the query-related triple sets entities and relations.

5. The method for generating a Text-to-Cypher semantic parsing model based on graph enhancement and LLM fine-tuning according to claim 4, characterized in that: The S104 specifically includes: Call gen_prompt(·) to retrieve the question q new , retrieve the question-related triple sets entities and relations, as well as the question type type new The three parts are integrated into the ICIO framework prompt word prompt q , the calculation formula is shown in (5): prompt q =gen_prompt(q new ,entities,relations,type new )#(5)。 6. The method for generating a Text-to-Cypher semantic parsing model based on graph enhancement and LLM fine-tuning according to claim 1, characterized in that: The S105 specifically includes: S1051, design LoRA fine-tuning model: select specific layers of the large model for LoRA fine-tuning; set the parameter weight of a single selected original linear layer to Both d and k represent dimensions. During fine-tuning, W0 remains unchanged and the weights are updated using the following formula (6). The new weight matrix W is used for inference: W=W0+ΔW#(6) in, is the weight update matrix in the fine-tuning training phase; ΔW is transformed into a matrix, that is, ΔW is represented by low-rank decomposition, which is expressed by formula (7): ΔW=B·A#(7) in, r represents rank; The weight matrix after LoRA adjustment is expressed by formula (8): W=W0+B·A#(8) S1052, call the parsing generation module process_parse(·) to process prompt q And Cypher query statement C, prompt q After word segmentation, serialization, and padding / cropping with C, the input and output token IDs are obtained, which are the input sequence x=[x1,x2,...,x p ] and the output sequence y=[y1,y2,...,y c ], and take the output sequence y as the true label sequence, where p and c represent the input sequence length and output sequence length respectively; as shown in formulas (9) and (10): x=process_parse(prompt q )#(9) y=process_parse(C)#(10) S1053, input x into the designed LoRA fine-tuning model, and obtain the output of the LoRA fine-tuning model through forward propagation S1054, calculate the output using the cross entropy loss function The difference between the true label sequence y; assuming that the LoRA fine-tuning model predicts a probability distribution when generating the i-th token Indicates generating each possible token The probability of; let the size of the vocabulary be V, then Is a probability vector of length V; for the true label y i , maximize P(y i Specifically, first, the cross entropy loss of a single token in y is expressed as formula (11): L i =-logP(y i |x)#(11) Then calculate the total cross entropy loss of y using the formula shown in (12): described Used to normalize the loss to ensure fair comparison between sequences of different lengths; S1055, calculate the total cross entropy loss L of the batch ′ , then the gradient of the low-rank matrices A and B and After multiple iterations of model optimization, the parameters of the low-rank matrices A and B are updated; finally, a trained Text-to-Cypher semantic parsing model based on graph enhancement using LoRA technology and LLM fine-tuning is obtained.

7. The method for generating a Text-to-Cypher semantic parsing model based on graph enhancement and LLM fine-tuning according to claim 6, characterized in that: After S1055, the following steps are also included: S1056, for the new arbitrary search question q test The processed prompt word prompt test , perform word segmentation and serialization to obtain the input sequence P test ;P test Send it to the trained Text-to-Cypher semantic parsing model to get the model output Will Convert to query statement C test .

Citation Information

Cited By

  • Method for converting natural text into graph database query language based on attention mechanism

    CN118113731A

  • Text2SQL training and reasoning method based on structured database atlas enhancement, electronic equipment, computer readable storage medium and computer program product

    CN121501818A

  • A Text2SQL training and reasoning method based on structured database graph enhancement, an electronic device, a computer readable storage medium and a computer program product

    CN121501818B