Knowledge graph question and answer-oriented SPARQL generation optimization method and system
By using domain mask vectors and inference chain prediction in knowledge graph questions and answers, SPARQL queries are directly generated and the fallback mechanism is adopted, the invalid query and empty result problems caused by domain errors are solved, and high-quality and robust answer generation is achieved.
Patent Information
- Application Number
- CN202510230025.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art problem of invalid SPARQL query and query results being empty due to domain errors in knowledge graph questions and answers.
The domain mask vector is used to predict the domain of input problems, and the domain prediction inference chain is based on the prediction, and SPARQL query and direct answers are generated through the dual-path generation module, and a fallback mechanism is used to ensure effective output when the query result is empty.
By directly generating SPARQL queries, the errors in the intermediate logical form are eliminated, the quality of logical form generation is improved, the search space is effectively pruned, irrelevant information interference is reduced, and robust answer generation is ensured.
Smart Images

Figure CN120144712A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of knowledge graphs, and particularly relates to an optimization method and system for SPARQL generation for knowledge graph question answering. Background Art
[0002] Knowledge graph question answering (KBQA) aims to use the knowledge in a knowledge graph (KB) to answer natural language questions (NLQs). Existing KBQA methods are generally divided into information retrieval (IR-based) methods and semantic parsing (SP-based) methods. Information retrieval-based methods focus on retrieving relevant facts from the knowledge graph to improve reasoning performance. For example, UniK-QA retrieves triples from the knowledge graph and generates answers by combining the input question with the retrieved triples; ROG introduces a planning-retrieval-reasoning framework, generates relation paths as plans, and then retrieves reasoning paths from the knowledge graph to facilitate reliable reasoning. Semantic parsing-based methods aim to convert NLQs into logical forms (LFs), which are ultimately executed by a query engine to retrieve answers from the knowledge graph. For example, RnG-KBQA enumerates potential S-expressions based on the main entity of the question and applies a ranking and generation framework to generate the final S-expression; KG-Agent uses a collaborative enhancement strategy to gradually generate a query graph; CHATKBQA adopts a generation-retrieval framework, first fine-tuning LLMs to generate S-expressions, and then using an unsupervised retrieval method to refine entities and relations.
[0003] Since information retrieval-based methods lack the reasoning depth supported by structured queries, while semantic parsing-based methods can handle complex reasoning through interpretable paths, they are considered more powerful. However, most semantic parsing-based methods require additional conversion steps, and common methods for generating logical forms include λ-DCS, query graphs, and S-expressions, etc., and there is usually a problem that the generated logical form is inconsistent with the final executable query. Even if the logical form is consistent with the actual situation, errors in the process of converting to SPARQL (the standard interface for accessing the knowledge graph) may lead to execution failures or incorrect answers, thus reducing reliability. In addition, domain errors can interfere with the construction of subsequent reasoning chains, resulting in invalid SPARQL queries. And due to syntax errors or incomplete data in the knowledge graph, empty results may also occur. Summary of the Invention
[0004] Aiming at the above defects or improvement requirements of the prior art, the present invention aims to optimize the granularity of the SPARQL generation system to solve the problems of invalid SPARQL queries caused by domain errors and empty query results in knowledge graph question answering.
[0005] To achieve the above object, according to one aspect of the present invention, there is provided an optimized SPARQL generation system for knowledge graph question answering, the system includes a domain prediction module, an inference chain prediction module, a dual-path generation module, and a fallback module.
[0006] The domain prediction module is used to predict the domain of the input question by using a domain mask vector.
[0007] The inference chain prediction module is used to predict the inference chain according to the predicted question domain.
[0008] The dual-path generation module is used to fine-tune the LLMs model according to the predicted inference chain to generate SPARQL queries and direct answers.
[0009] The fallback module is used to, when the result of executing the SPARQL query on the knowledge graph is empty, use the fallback mechanism to supplement the direct answer as the final answer.
[0010] According to another aspect of the present invention, there is provided an optimized SPARQL generation method for knowledge graph question answering, the method includes the following steps:
[0011] Step 1: Predict the domain of the input question by using a domain mask vector.
[0012] Step 2: Predict the inference chain according to the predicted question domain.
[0013] Step 3: Fine-tune the LLMs model according to the predicted inference chain to generate SPARQL queries and direct answers.
[0014] Step 4: Execute the SPARQL query on the knowledge graph. If the query result is empty, go to the next step; otherwise, use the query result as the final answer.
[0015] Step 5: Use the fallback mechanism to supplement the direct answer as the final answer.
[0016] Further, the step of predicting the domain of the input question by using a domain mask vector in Step 1 specifically includes:
[0017] Collect domain labels from the benchmark dataset.
[0018] Build a domain prediction model based on the domain mask vector.
[0019] Train the domain prediction model by using the domain labels.
[0020] Use the trained domain prediction model to predict the domain of the input question.
[0021] Further, the domain labels include a complete domain set, a domain candidate set for each question, and a golden domain for each question.
[0022] Furthermore, the collection of domain labels from the benchmark dataset specifically includes:
[0023] Extract all the relationships connected to the main entity in the benchmark dataset to form a complete relationship set R all , and derive the complete domain set as D all = Domain(R all );
[0024] Collect the relationships connecting each question to its main entity to form a relationship candidate set R c ∈ R all , and calculate the corresponding domain candidate set as D c = Domain(R c ), where D c ∈ D all ;
[0025] Extract the gold relationship of each question from the true SPARQL provided by the benchmark dataset. The gold relationship is the relationship directly connected to the main entity in the SPARQL query, and then determine the gold domain as d g = Domain(r g ), where d g ∈ D c .
[0026] Furthermore, the construction of the domain prediction model based on the domain mask vector specifically includes:
[0027] Use the domain mask vector M = {m 1 , m 2 , …, m n} as the input vector D all = {d 1 , d 2 , … d n} to assign a boolean value m i to each domain d i ;
[0028] Calculate the probability p i of each domain d i :
[0029]
[0030] Take the domain with the highest probability as the predicted domain of the given question.
[0031] Furthermore, the prediction of the inference chain according to the predicted question domain in step two specifically includes:
[0032] Collect the inference chains from the benchmark dataset;
[0033] Construct an inference chain prediction model;
[0034] Construct training samples with each question, its domain, and the inference chain, and train the inference chain prediction model;
[0035] For the input question and its predicted domain, use the inference chain prediction model to predict the correct inference chain.
[0036] Further, the collecting of the inference chain from the benchmark dataset specifically includes:
[0037] Use regular expressions to extract all triples from the SPARQL query;
[0038] Convert the subgraph formed by the triples into an undirected graph;
[0039] Adopt depth-first search to find a path in the undirected graph based on the main entity and the answer entity, and the combination of relationships on the path is the inference chain.
[0040] Further, the constructing of the inference chain prediction model specifically includes:
[0041] Encode the question Q using a pre-trained language model and incorporate its domain d as context information:
[0042] Q v =PLM(Q|d)
[0043] Pass through the classification layer to generate a probability distribution P on the candidate inference chains ic 1 ,ic 2 ,…ic n : ic :
[0044] P ic =[P(ic 1 |Q v ),P(ic 2 |Q v ),…,P(ic n |Q v )]
[0045] Select the inference chain ic with the highest probability h as the correct inference chain ic:
[0046]
[0047] Further, the fine-tuning of the LLMs model in step three to generate the SPARQL query and the direct answer according to the predicted inference chain specifically includes:
[0048] Input two prompts generated by combining the prefix with the given question Q and the predicted inference chain ic into the LLMs model:
[0049] Prompt SP = concat(prefix SP , Q, ic)
[0050] Prompt ANS = concat(prefix ANS , Q, ic)
[0051] Among them, the prefix prefix SP is defined as "generating an executable SPARQL query for retrieving information corresponding to a given question and reasoning chain", and the prefix prefix ANS is defined as "generating one or more answers corresponding to a given question and reasoning chain";
[0052] Train the LLMs model using a multi-task learning framework;
[0053] Use beam search to generate SPARQL queries and direct answers based on the LLMs model.
[0054] Compared with the prior art, the present invention can achieve the following beneficial effects:
[0055] (1) By directly generating SPARQL queries, the present invention eliminates the intermediate logical form, avoids errors that may occur in the conversion step, and improves the quality of logical form generation.
[0056] (2) The present invention uses a masked vector for domain prediction, effectively pruning the search space and reducing the interference of irrelevant information when constructing the reasoning chain.
[0057] (3) The present invention decomposes SPARQL generation into three gradually optimized granularities. As the granularity is gradually optimized, the performance in the generation stage is correspondingly improved.
[0058] (4) The present invention introduces a dual-path fallback mechanism for robust answer generation, ensuring effective output even when the SPARQL execution returns an empty result. Description of the Drawings
[0059] Figure 1 It is the overall architecture diagram of the SPARQL generation optimization system GOSP for knowledge graph question answering in a specific embodiment of the present invention. Specific Embodiments
[0060] To make the objectives, technical solutions and advantages of the present invention more comprehensible, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0061] Knowledge graphs (KBs) are sets composed of triples (s, r, o), where each triple consists of a head entity s, a relation r, and a tail entity o. Formally, it is represented as where E and R represent the sets of entities and relations respectively. Among them, the representation of the relation r spans multiple levels, and each level is separated by dots, such as location.location.adjoin_s.
[0062] The domain candidate set is determined according to the main entity involved in the given question. In the knowledge graph, each main entity is associated with multiple relations, and these relations are mapped to the corresponding domains through a predefined function Domain(r) = d. This function selects the first two levels of the relation r as the domain d. For a main entity m, it is associated with a set of relations R c ={r 1 , r 2 , …, r n}, and the domain candidate set D c can be calculated by the following formula: D c = Domain(R c ) = {Domain(r 1 ), Domain(r 2 ), …, Domain(r n )}.
[0063] An inference chain represents a sequence of relations connecting the head entity and the answer entity, denoted as ic = r 1 → r 2 → … → r n , where r i ∈R.
[0064] The present invention provides an SPARQL generation optimization system for knowledge graph question answering - GOSP, including:
[0065] A domain prediction module for predicting the domain of the input question using a domain mask vector;
[0066] An inference chain prediction module for predicting the inference chain according to the predicted question domain;
[0067] A dual - path generation module for fine - tuning LLMs to generate SPARQL queries and direct answers according to the predicted inference chain;
[0068] A fallback module, which is used to supplement the direct answer as the final answer by using a fallback mechanism when the result of executing a SPARQL query on the knowledge graph is empty.
[0069] As Figure 1 shown, GOSP decomposes SPARQL generation into three gradually optimized granularities: (1) Domain prediction using a mask vector: For the input question, GOSP uses a domain mask vector to predict the domain of the question in its corresponding domain candidate set. This step effectively reduces the reasoning chain errors caused by irrelevant domains. (2) Domain-aware reasoning chain prediction: Based on the predicted domain, identify the reasoning chain to closely align it with the structure of the underlying knowledge graph. This step lays a solid foundation for SPARQL query generation. (3) Dual-path generation with a fallback mechanism: Use the predicted reasoning chain as auxiliary information to fine-tune large language models (LLMs) to generate SPARQL queries and direct answers. For questions where the SPARQL execution returns an empty result, the fallback mechanism uses the direct answer to ensure an effective result is obtained.
[0070] The SPARQL generation optimization method for knowledge graph question answering according to the present invention specifically includes the following steps:
[0071] First, predict the domain of the input question using a mask vector.
[0072] In GOSP, domain prediction is crucial because it can prune the search space and reduce the interference of irrelevant information when constructing the reasoning chain.
[0073] 1. Collect domain labels from the benchmark dataset.
[0074] The collected domain labels include the complete domain set D all , the domain candidate set D c for each question, and the golden domain d g for each question. In this embodiment, two widely used knowledge graph question answering benchmark datasets, WebQSP and CWQ, are selected. WebQSP contains 4,737 NLQs and their corresponding SPARQL queries, while CWQ includes 34,689 pairs of such questions and queries. Freebase is the knowledge graph on which these two benchmark datasets are based. The specific steps for collecting domain labels are as follows: First, extract all the relationships connected to the main entity in the benchmark dataset to form a complete relationship set R all . The complete domain set is derived as D all = Domain(R all ), which consists of 7,507 unique labels. Second, for each question, collect the relationships connected to its main entity to form a relationship candidate set R c ∈ Rall The corresponding domain candidate set is calculated as D c =Domain(R c ),D c ∈D all Finally, the golden relations of each question are extracted from the real SPARQL provided by the benchmark dataset. The golden relations are defined as the relations directly connected to the main entity in the SPARQL query. The golden fields are then determined as d g =Domain(r g ), where d g ∈D c .
[0075] 2. Build a domain prediction model based on the domain mask vector.
[0076] The candidate domain set for each question is only a small subset of the complete domain set. In order to avoid assigning probabilities to irrelevant domains, the domain mask vector is used in the domain prediction model to improve the formula of the Softmax function. The domain mask vector M = {m 1 ,m 2 ,…,m n} is the input vector D all ={d 1 ,d 2 ,…d n Each field d in i Assign a Boolean value m i , that is, in the original Softmax calculation formula An additional mask element m is added i For the specified element d in the input vector i Masking is performed to prevent the model from predicting irrelevant domains outside the candidate set of domains for the current problem.
[0077] Each field i The probability p i Calculated as:
[0078]
[0079] The domain with the highest probability is taken as the predicted domain for the given problem.
[0080] The principle of Boolean value assignment is to assign Boolean value 1 to the fields in the field candidate set for the current problem, and assign Boolean value 0 to other fields. Through Boolean value assignment, the probability of non-candidate fields is made zero, ensuring that the model is not disturbed by irrelevant fields when predicting the field of the current problem, thereby improving the accuracy of field prediction.
[0081] 3. Use domain labels to train domain prediction models.
[0082] Taking each question in the domain label and its domain candidate set as input, and its golden domain as output, train the domain prediction model.
[0083] 4. Use the trained domain prediction model to predict the domain of the input question.
[0084] II. Problem domain prediction inference chain based on prediction.
[0085] The inference chain is the core of SPARQL. Taking it as auxiliary information can improve the generation quality, enhance the overall performance and efficiency.
[0086] 1. Collect inference chains from the benchmark dataset.
[0087] Extract the inference chain of each question from the true SPARQL queries in the benchmark dataset. The steps are as follows: First, use regular expressions to extract all triples from the SPARQL query. Second, convert the subgraph formed by the above triples into an undirected graph because some inference chains involve bidirectional relationships. Third, use depth-first search (DFS) to find a path in the undirected graph based on the main entity and the answer entity. The combination of relationships on the path is defined as the inference chain. In this embodiment, a total of 2,786 inference chain labels are identified in two benchmark datasets, WebQSP and CWQ.
[0088] 2. Build an inference chain prediction model.
[0089] Encode the question Q using a pre-trained language model (PLM) and incorporate its domain d as context information:
[0090] Q v = PLM(Q|d)
[0091] Q v represents the vector of the question Q after being encoded by the PLM, which is passed through the classification layer to generate a probability distribution P 1 , ic 2 ,… ic n over the candidate inference chains ic ic :
[0092] P ic = [P(ic 1 |Q v ), P(ic 2 |Q v ), …, P(ic n |Q v )]
[0093] Select the inference chain ic h with the highest probability as the correct inference chain:
[0094]
[0095] 3. Construct training samples with each question, its domain, and the reasoning chain, and train the reasoning chain prediction model.
[0096] To train the reasoning chain prediction model, we construct training samples with each question and its domain as the input and the reasoning chain of each question as the output.
[0097] 4. For the input question and its predicted domain, use the reasoning chain prediction model to predict the correct reasoning chain.
[0098] III. Generate a dual-path with a fallback mechanism based on the predicted reasoning chain and output the final answer.
[0099] The present invention uses two different prompts to guide the dual-path generation:
[0100] Prompt SP = concat(prefix SP , Q, ic)
[0101] Prompt ANS = concat(prefix ANS , Q, ic)
[0102] where prefix SP is defined as "generate an executable SPARQL query to retrieve information corresponding to the given question and reasoning chain", and prefix ANS is defined as "generate one or more answers corresponding to the given question and reasoning chain". These prefixes are combined with the given question Q and the predicted reasoning chain ic to generate two prompts.
[0103] During the training process, we adopt a multi-task learning framework, where each question contributes two training pairs - SPARQL generation and direct answer generation. To simplify SPARQL generation, the main entity is replaced with a placeholder token [ENT] (e.g., m.09c7w0 → [ENT]), enabling the generation module to focus on identifying the position of the main entity rather than the meaningless ID. Similarly, the generation module generates corresponding text labels for the constraint entity IDs (e.g., m.01mp → Country).
[0104] During the reasoning process, the post-processing step converts the [ENT] token and the text labels back to their respective entity IDs. In addition, we use beam search (beam size B) to generate multiple candidates for the SPARQL query SP and the direct answer ANS.
[0105] When in the knowledge graph When the result of executing a SPARQL query is empty, the fallback mechanism will supplement the direct answer to ensure the generation of the final answer ANS final robustness of:
[0106]
[0107] In this embodiment, LLaMA-2-7B is used as the backbone model for dual-path generation, and LoRA is used to fine-tune it to achieve parameter-efficient training. The model is fine-tuned for 5 epochs on the combined WebQSP and CWQ benchmark datasets, with a batch size of 16 and a learning rate of 5e-5. The beam size B of beam search is 10.
[0108] To analyze and verify the performance of GOSP, we conducted a series of experimental analyses on the WebQSP and CWQ benchmark datasets, and the results are as follows.
[0109] (1) In this embodiment, two metrics, Hits@1 and F1, are used to evaluate the GOSP model based on the WebQSP and CWQ benchmark datasets. Among them, Hits@1 evaluates the accuracy of the highest-ranked answer, and the F1 score measures the coverage of all correct answers. The results show that the Hits@1 of GOSP on WebQSP is 89.3%, and the F1 is 85.4%; the Hits@1 on CWQ is 88.0%, and the F1 is 84.6%. By querying the relevant literature of other existing models, the Hits@1 and F1 evaluation results of the existing models are obtained. It shows that the most advanced evaluation result among the existing models is the ChatKBQA model. The Hits@1 of ChatKBQA on WebQSP is 86.4%, and the F1 is 83.5%; the Hits@1 on CWQ is 86.0%, and the F1 is 81.3% (from "Chatkbqa: A generate-then-retrieve framework for knowledge base question answering with fine-tuned large language models," in Findings of the Association for Computational Linguistics: ACL 2024, 2024, pp. 2039–2056). Thus, it can be seen that the results of GOSP on the two benchmark datasets are better than those of the ChatKBQA model. Specifically, GOSP increases Hits@1 by 2.9% and F1 by 1.9% on WebQSP; on CWQ, it increases Hits@1 by 2.0% and F1 by 3.3%. This shows the robustness of GOSP in dealing with complex reasoning.
[0110] (2) To explore how granularity optimization affects model performance, this embodiment specifically analyzes the impact of the domain and the reasoning chain. In this analysis, we only use real information instead of predicted information. Table 1 shows the results under three settings: None means no auxiliary information is used; Gold Dom. (Domain) means using the gold domain as auxiliary information for fine-tuning the large language model; Gold InCh. (Inference Chain) means providing the gold reasoning chain for fine-tuning the large language model.
[0111] As shown in Table 1, first, for SPARQL generation performance, without auxiliary information, the F1 scores of the model on WebQSP and CWQ are 83.4% and 83.9% respectively. Introducing the gold domain increases the F1 scores to 88.1% and 85.3%, indicating that domain-level guidance effectively prunes the search space for SPARQL generation. When the gold reasoning chain is provided, the F1 score on WebQSP reaches 93.6% and on CWQ reaches 90.2%, highlighting the importance of step-by-step guidance for generating accurate SPARQL queries. Second, for direct answer generation performance, the F1 scores on WebQSP and CWQ are relatively low, at 52.1% and 43.1% respectively, reflecting the difficulty of directly generating answers without context guidance. The introduction of the gold domain has a minimal impact, indicating that it is not sufficient for answer generation. Providing the gold reasoning chain increases the F1 score on WebQSP to 57.9% and on CWQ to 46.7%, indicating that the reasoning chain plays a greater role. Finally, as the granularity is refined, the combined performance gradually improves, highlighting the key role of granularity optimization reasoning in enhancing GOSP performance.
[0112] Table 1 Impact of Granularity Optimization
[0113]
[0114]
[0115] (3) This embodiment analyzes the performance of the upstream prediction module. Table 2 shows the accuracy of the domain and reasoning chain prediction models. The accuracy of the domain prediction model on WebQSP is 89.3% and on CWQ is 96.6%, while the accuracy of the reasoning chain prediction model reaches 88.7% on both benchmark datasets.
[0116] Table 2 Accuracy of Domain and Reasoning Chain Prediction Models
[0117] Task WebQSP CWQ Domain Classification 89.3% 96.6% Inference Chain Classification 88.7% 88.7%
[0118] (4) To evaluate the impact of prediction in the modular process, we analyzed the performance of GOSP under predicted and true auxiliary information. When switching from true input to predicted input, the F1 score on WebQSP decreased from 93.9% in Table 1 to 85.4% in (1), a decrease of 8.5%; the F1 score on CWQ decreased from 92.4% in Table 1 to 84.6% in (1), a decrease of 7.8%. However, despite the small inaccuracies leading to a performance drop, the final performance of GOSP is still superior.
[0119] The embodiments described above only represent the implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent for the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent for the present invention should be subject to the appended claims.
Claims
1. A SPARQL generation optimization system for knowledge graph question answering, characterized in that: The system includes a domain prediction module, a reasoning chain prediction module, a dual path generation module, and a fallback module. The domain prediction module is used to predict the domain of the input question using the domain mask vector; The inference chain prediction module is used to predict the inference chain according to the predicted problem domain; The dual-path generation module is used to fine-tune the LLMs model to generate SPARQL queries and direct answers according to the predicted reasoning chain; The fallback module is used to use a fallback mechanism to supplement the direct answer as the final answer when the result of executing the SPARQL query on the knowledge graph is empty.
2. A SPARQL generation optimization method for knowledge graph question answering, characterized in that: The method comprises the following steps: Step 1: Use the domain mask vector to predict the domain of the input problem; Step 2: Predict the reasoning chain based on the predicted problem domain; Step 3: Fine-tune the LLMs model to generate SPARQL queries and direct answers based on the predicted reasoning chain; Step 4: Execute SPARQL query on the knowledge graph. If the query result is empty, proceed to the next step. Otherwise, take the query result as the final answer. Step 5: Use the fallback mechanism to supplement the direct answer into the final answer.
3. The SPARQL generation optimization method for knowledge graph question answering according to claim 2 is characterized in that: The domains of the input problem predicted by the domain mask vector in step 1 specifically include: Collect domain labels from benchmark datasets; Build a domain prediction model based on the domain mask vector; Use domain labels to train domain prediction models; Use the trained domain prediction model to predict the domain of the input question.
4. The SPARQL generation optimization method for knowledge graph question answering according to claim 3 is characterized in that: The domain labels include the complete domain set, the domain candidate set for each question, and the golden domain for each question.
5. The SPARQL generation optimization method for knowledge graph question answering according to claim 4 is characterized in that: The collecting of domain labels from the benchmark dataset specifically includes: Extract all the relations connected to the main entity in the benchmark dataset to form a complete relation set R all , the complete domain set is derived as D all =Domain(R all ); Collect the relationship between each question and its main entity to form a relationship candidate set R c ∈R all , calculate the corresponding domain candidate set as D c =Domain(R c ), where D c ∈D all ; The golden relation of each question is extracted from the real SPARQL provided by the benchmark dataset. The golden relation is the relation directly connected to the main entity in the SPARQL query, and then the golden field is determined as d g =Domain(r h ), where d g ∈D c .
6. The SPARQL generation optimization method for knowledge graph question answering according to claim 5 is characterized in that: The constructing of a domain prediction model based on a domain mask vector specifically includes: Use the domain mask vector M = {m1,m2,…,m n } is the input vector D all ={d1,d2,…d n Each field d in i Assign a Boolean value m i ; Calculate each field d i The probability p i : The domain with the highest probability is taken as the predicted domain for the given problem.
7. The SPARQL generation optimization method for knowledge graph question answering according to claim 2 is characterized in that: The prediction reasoning chain according to the predicted problem domain described in step 2 specifically includes: Collect reasoning chains from benchmark datasets; Constructing inference chain prediction models; Construct training samples based on each problem and its domain and reasoning chain, and train the reasoning chain prediction model; For the input question and its prediction domain, the reasoning chain prediction model is used to predict the correct reasoning chain.
8. The SPARQL generation optimization method for knowledge graph question answering according to claim 7 is characterized in that: The collecting of reasoning chains from the benchmark dataset specifically includes: Extract all triples from a SPARQL query using regular expressions; Converting the subgraph formed by the triples into an undirected graph; A depth-first search is used to find a path based on the main entity and the answer entity in an undirected graph, and the relationships on the path are combined into a reasoning chain.
9. The SPARQL generation optimization system for knowledge graph question answering according to claim 7, characterized in that: The construction of the inference chain prediction model specifically includes: The question Q is encoded using a pre-trained language model and its domain d is incorporated as context information: Q v =PLM(Q|d) Passed through the classification layer, in the candidate inference chain ic1, ic2, ...ic n Generate probability distribution P ic : P ic =[P(ic1|Q v ),P(ic2|Q v ),…,P(ic n |Q v )] Select the inference chain with the highest probability h As a correct chain of reasoning:
10. The SPARQL generation optimization method for knowledge graph question answering according to claim 2, characterized in that: Step 3, fine-tuning the LLMs model to generate SPARQL queries and direct answers based on the predicted reasoning chain, specifically includes: The LLMs model is fed with two prompts generated by combining the prefix with the given question Q and the predicted reasoning chain ic: Prompt SP =concat(prefix SP ,Q,ic) Prompt ANS =concat(prefix ANS ,Q,ic) Among them, the prefix prefix SP Defined as "generates an executable SPARQL query to retrieve information corresponding to a given question and reasoning chain", prefix prefix ANS Defined as "generating an answer or answers corresponding to a given question and chain of reasoning"; The LLMs model is trained using a multi-task learning framework; Generate SPARQL queries and direct answers based on LLMs models using beam search.