Generative multi-hop law rule chain mining and evaluating method

Through the generative multi-hop legal rule chain mining and evaluation method, combined with statistical mining and logical reasoning of large language model, the problems of missing relationships and sparse connections in the knowledge graph are solved, and efficient and robust legal rule chain mining and evaluation are achieved.

CN119990135AActive Publication Date: 2025-05-13CHINALAWINFO CO LTD

Patent Information

Application Number
CN202411817052.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-11
Publication Date
2025-05-13
Estimated Expiration
2044-12-11

AI Technical Summary

Technical Problem

The existing technology faces the problems of lack of relationships and sparse connections in the construction of knowledge graphs. Traditional algorithms are difficult to effectively mine frequent item sets, high-dimensional problems have high computational complexity, lack of semantic understanding, and the black box characteristics of the neural network part make the inference rules difficult to explain, the training cost is high, and multi-step models are easily overfitted, and are sensitive to noise.

Method used

A generative multi-hop legal rule chain mining and evaluation method is proposed. A potential legal rule chain is generated through two paths: chain rule mining based on statistics and legal ontology logic reasoner based on large language model. Comprehensive evaluation is carried out by filtering and filtering out highly credible legal rule chains.

Benefits of technology

It enhances semantic understanding ability, improves interpretability and reduces training costs, enhances the generalization ability and robustness of the model, and realizes a high robustness and rational multi-hop legal rule chain mining technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990135A_ABST
    Figure CN119990135A_ABST
Patent Text Reader

Abstract

The invention discloses a generative multi-hop legal rule chain mining and evaluating method, which comprises the following steps of: generating a candidate potential legal rule chain set Z through dual-path mining; the method specifically comprises the steps of generating a potential legal rule chain set I Z1 based on statistical chain rule mining and generating a potential legal rule chain set II Z2 based on a legal ontology logical reasoning model of a large language model; evaluating all potential legal rule chains from two dimensions of voting by utilizing legal logic rationality of the large language model and voting by utilizing truth-oriented truth reasoning chain credibility of the large language model; and carrying out legal rule chain rationality arbitration, and filtering and screening out a highly credible legal rule chain from the potential legal rule chain set Z. According to the method, the semantic understanding of entities and relationships in the legal knowledge graph can be enhanced, the interpretability of the pre-trained generative language model is improved, the training cost is reduced, and a multi-hop legal rule chain mining technology with high robustness and rationality is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer-related artificial intelligence and natural language processing technology, and in particular, relates to a method for automatically generating a comprehensive report on legislative opinions for a long period and across fields. Background Art

[0002] The construction of a legal knowledge graph is a complex task that aims to create an extensive network of relationships between entities and concepts to support a variety of intelligent applications. However, during the construction process, problems such as missing relationships and sparse connections are often encountered. This means that the associations between certain entities may not be fully identified or recorded, resulting in gaps or incompleteness in the knowledge graph when representing complex interactions in the real world. This sparsity not only affects the coverage and accuracy of the knowledge graph, but also poses challenges to graph-based queries and data mining. To address these issues, researchers and developers have adopted a variety of strategies, including using text mining techniques to automatically extract relationships from large amounts of unstructured data, using statistical and machine learning methods to infer implicit relationships, and filling missing connections through manual intervention and expert knowledge. Despite this, missing relationships and sparse connections remain one of the important challenges in the field of knowledge graph construction, which requires continuous efforts and innovation to overcome.

[0003] At present, traditional association rule mining algorithms such as Apriori and FP-Growth can be used to mine chain rules in knowledge graphs to count the frequency of occurrence between entities and relationships. Recently, there are also methods such as RNNLogic, which combines hybrid methods of neural networks and symbolic reasoning. There are also methods that transform this problem into a multi-step knowledge graph link prediction problem, by mapping entities and relationships to vector representations in low-dimensional space. Representative models such as GNN, GCN, TransE, DistMult and ComplEx predict possible relationships between entities by learning vector representations of entities and relationships.

[0004] The disadvantages of the prior art derived by causal reasoning are:

[0005] 1) Problem description of traditional association rule mining algorithms (such as Apriori and FP-Growth):

[0006] Data sparsity: In a large-scale knowledge graph, there may be a large number of entities and relationships, but each combination of entities and relationships may appear relatively few times, making it difficult to effectively mine frequent itemsets.

[0007] High-dimensional problem: When the number of entities and relationship types is huge, the potential association rule space that traditional algorithms need to explore is huge and the computational complexity is high.

[0008] Lack of semantic understanding: These algorithms are based only on statistical frequencies and ignore the underlying semantic connections between entities and relations.

[0009] 2) Problem description of hybrid methods of neural networks and symbolic reasoning (such as RNNLogic):

[0010] Explanation issues: Although it combines the powerful fitting ability of neural networks and the rule expression ability of symbolic logic, the black box characteristics of the neural network part may make the final inference rules difficult to explain.

[0011] High training cost: Neural networks require a large amount of data and computing resources for training, which is not friendly to environments with limited resources.

[0012] 3) Multi-step knowledge graph link prediction methods (such as GNN, GCN, TransE, etc.) Problem description:

[0013] Overfitting risk: These models are prone to overfitting on small-scale or highly homogeneous datasets, and may not be able to generalize well to complex and changing real-world environments.

[0014] Robustness issue: Sensitive to noisy data and outliers, which may lead to inaccurate learned vector representations.

[0015] The relevant terms and their explanations of the present invention are as follows:

[0016]

Chain rule of knowledge graph

[0017]

[0018] [Statistical chain rule mining] Statistical chain rule mining statistically analyzes the frequency of occurrence of entities and relationships, and summarizes reasonable chain rules based on the probability of rule occurrence to reason and fill in missing links in the knowledge graph. For example, in the knowledge graph in the field of medical health, by analyzing the common patterns between symptoms, treatments, and results, rules can be generated to help doctors quickly diagnose and treat diseases.

[0019]

Legal Ontology

[0020] [Thinking Chain Tips] Thinking Chain is a technology that allows large models to explicitly output intermediate steps during the reasoning process to improve their reasoning ability. This concept refers to the way humans solve complex problems, breaking down the problem into a series of reasoning steps in the form of natural language until the final conclusion is reached. Specifically, the goal of thinking chain is to allow large models to explicitly output intermediate step-by-step reasoning steps before outputting the final answer, thereby enhancing their conceptual understanding and reasoning ability. This is of great significance for solving tasks that require precise reasoning, such as mathematical problems, common sense reasoning, etc. Through thinking chaining, large models can get closer to the way humans think when solving problems, thereby improving their intelligence level.

[0021] [Generative Language Model] A generative language model is a deep learning model that uses a large amount of unlabeled text or labeled text for training to learn natural language processing tasks. These models learn a general representation of language in a pre-training phase, which allows them to be used for a wide range of natural language processing tasks without requiring too much labeled data. GPT3 and ChatGPT are representative works of generative language models, which aim to simulate human natural language processing capabilities by using neural network models for text generation and processing. By using these models, users can have more natural and realistic conversations with computers and achieve natural language processing tasks such as speech recognition, text summarization, machine translation, text classification, computer code generation, movie script creation, etc. Summary of the invention

[0022] The purpose of the present invention is to propose a generative multi-hop legal rule chain mining and evaluation method, which obtains a credible mining result of filtering out highly credible legal rule chains from the potential legal rule chain set Z by conducting a comprehensive evaluation of the potential legal rule chain set generated by dual-path mining, combining the rationality of legal logic and the credibility of legal fact reasoning chain.

[0023] The present invention is achieved by the following technical solutions:

[0024] The present invention provides a generative multi-hop legal rule chain mining and evaluation method, comprising:

[0025] Step 1, using two paths, chain rule mining based on statistics and legal ontology logic reasoning based on a large language model, to generate possible legal rules, so as to generate a candidate potential legal rule chain set Z, wherein the potential legal rule chain set Z further includes a potential legal rule chain set Z1 generated by chain rule mining based on statistics and a potential legal rule chain set Z2 generated by a legal ontology logic reasoning model based on a large language model;

[0026] Step 2: Evaluate the rationality of the potential chain legal rule set Z generated by the dual-path mining in step 1. Specifically, evaluate all potential legal rule chains from two dimensions, including:

[0027] i. Use the legal logic rationality voting of the large language model to evaluate the legal logic rationality of the generated potential legal rule chain set, including using the legal ontology prompter to integrate the legal ontology semantic knowledge of each legal rule chain into the prompt of the large language model, and through reasoning with the large language model, complete the identification of the rationality of the legal rule chain at the level of legal common sense, and generate a logical rationality confidence score as the result of the logical rationality voting;

[0028] ii. Using the fact-oriented factual reasoning chain credibility voting of the large language model to evaluate the credibility of the legal fact reasoning chain of the generated potential legal rule chain set, including using the fact chain prompter to calculate the credibility confidence score of the legal fact reasoning chain, constructing the knowledge graph retrieval statement through the aforementioned mining of incomplete legal rule chains, searching for qualified triple fact data in the knowledge graph and performing random sampling, using the large language model to infer and calculate the credibility confidence score of each legal fact reasoning chain, and generating the corresponding legal fact reasoning chain credibility confidence score as the legal logic rationality voting result;

[0029] iii. Combine the confidence scores of legal logic rationality and legal fact reasoning chain credibility to generate a weighted confidence assessment score for each potential legal rule chain;

[0030] Step 3: Conduct arbitration on the rationality of the legal rule chain. Specifically, perform weighted synthesis on the evaluation results of the legal logic rationality of the legal rule chain obtained in step 2 and the credibility evaluation results of the factual reasoning chain of the legal rule chain obtained in step 3, and filter out highly credible legal rule chains from the potential legal rule chain set Z.

[0031] In some embodiments, the process of generating a potential legal rule chain set Z1 based on statistical chain rule mining further includes:

[0032] Construct a legal rule mining model based on statistics, take the potential legal rule chain set Z as the input of the model, select any chain rule {r1,r2,...,r l} as the legal rule mining chain, the legal rule mining query operation is represented by q1, and the result of the query operation q1 obtained through mining is equivalent to rule 1 (target relationship) {r1,r2,...,r l}→r a1 As the output of the mining model, the probability distribution p of the potential legal rule chain set Z1 generated by the statistical chain rule mining in this step is calculated according to the output of the model. g,e (A|G,q), the expression is as follows:

[0033] p g,e (A|G,q)=∑p e (A|G,q,Z)·p g (Z|G,q)

[0034] Among them, p g To generate the probability distribution of potential legal rule chains through statistical chain-based legal rule mining, p e is the evaluation result of the potential legal rule chain, A is the equivalent legal rule r a The evaluation results are as follows: G is the legal knowledge graph representation, and Z is the set representation of potential legal rule chains.

[0035] In some embodiments, the related process of generating a potential legal rule chain set Z2 based on the legal ontology logic reasoning model of the large language model further includes:

[0036] Construct a legal ontology logic reasoning model to obtain a legal ontology logic reasoning model based on a large language model. The legal ontology logic reasoning model based on a large language model takes the legal knowledge ontology sorted out by experts as the input of the model, selects any chain rule {r1,r2,...,r l} as the legal rule mining chain, the legal rule reasoning query operation is represented by q2, and the result of the query operation q2 obtained by reasoning is the equivalent rule 2 (target relationship) is {r1,r2,...,r l}→r a2 As the output 2 of the model, the probability distribution p of the potential legal rule chain set 2 Z2 generated by the legal ontology logic reasoning model based on the large language model is calculated according to the output 2 of the reasoning model. gLM (Z|G,O,q), the expression is as follows:

[0037]

[0038] Among them, p LM () is the probability of the large language model, T1(O) is the function of converting the legal knowledge ontology O sorted by experts into LLM prompts, z <i is all the legal rules before the i-th order legal rule in Z, p gLM To generate probability distributions of possible legal rule chains by mining using statistics-based chain rules.

[0039] In some embodiments, the legal ontology logic reasoning model based on the large language model specifically includes obtaining prompts from the legal knowledge ontology previously sorted out by experts, and then using the prompts to enrich the general large language model, and then defining the reasoning ability of the enriched large language model in dealing with legal issues.

[0040] In some embodiments, the union of the legal rule chain set Z1 and the legal rule chain set Z2 is taken to obtain the legal rule chain set Z, which is expressed as follows:

[0041] Z=Z1∪Z2.

[0043] In some embodiments, the thought chain process in the legal ontology prompter is represented as follows:

[0044] p(C|R,E)=p(C|R,E,A)p(C|R,E)

[0045]

[0046] p(C|R,E)=p(C|R,E,A)p(A|R,E)

[0047] Where p(C|R,E) is the value of the given chain rule relation R, the chain rule interpretation E and the equivalent rule r. a1 、r a2 The total probability of evaluating the confidence score C under the evaluation result A, p(A|R,E) is the evaluation probability of the chain rule and its explanation.

[0048] In some embodiments, in the fact chain prompter, a Cypher query statement of the knowledge graph is generated by the mined multi-hop legal rule chain, the query results of the graph database are randomly sampled, and the fact chain prompter is enabled to traverse these sampled records and construct appropriate fact evaluation prompts.

[0049] In some embodiments, the weighted confidence assessment score AR of each potential legal rule chain is expressed as follows:

[0050] AR=ω1·N AR +ω2·F AR

[0051] In the formula, ω1 and ω2 are weight parameters, and ω1+ω2=1, F AR is the credibility confidence score of the legal fact reasoning chain, N AR It is the confidence score of legal logic rationality.

[0052] The method for automatically generating a comprehensive report on cross-domain legislative opinions as described in claim 1 is characterized in that the agent natural language memory stream module is used to save the observation value text data of the current rule of law situation perceived quantitatively.

[0053] Compared with the prior art, the advantages of the present invention are mainly reflected in the following aspects:

[0054] 1) Enhanced semantic understanding: Traditional algorithms usually ignore the potential semantic connections between entities and relationships and only rely on frequency statistics. The "legal ontology prompter" and "fact chain prompter" of the present invention can evaluate the legal logic and the rationality of the fact chain, thereby deepening the semantic understanding of entities and relationships in the legal knowledge graph;

[0055] 2) Improve interpretability and reduce training costs: Compared with the hybrid method of neural network and symbolic reasoning, the generative multi-hop chain rule mining method of the present invention provides better interpretability through clear logic and fact chain evaluation, and uses a pre-trained generative language model, which also relatively reduces the training cost;

[0056] 3) Enhance the generalization and robustness of the model: In view of the problem that the multi-step knowledge graph link prediction method is prone to overfitting and is sensitive to noise, the method of the present invention can effectively improve the generalization and robustness of the model through a dual evaluation strategy and a legal rule rationality arbitrator, ensuring that the generated legal rules are logically rigorous and consistent with actual legal facts;

[0057] 4) A multi-hop legal rule chain mining technology with high robustness and rationality is implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 It is an overall flow chart of a generative multi-hop legal rule chain mining and evaluation method of the present invention;

[0059] Figure 2 This is the detailed flow chart of step 1;

[0060] Figure 3 This is the detailed flow chart of step 2;

[0061] Figure 4 It is a framework diagram of a specific embodiment of a generative multi-hop legal rule chain mining and evaluation algorithm of the present invention. DETAILED DESCRIPTION

[0062] In order to make the purpose, technical solution and advantages of the embodiments of the present invention more clear, the technical solution in the embodiments of the present invention will be clearly described below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0063] like Figure 1 As shown, the generative multi-hop legal rule chain mining and evaluation method proposed in the present invention is divided into four steps:

[0064] Step 1: Generate a set of candidate potential legal rule chains. The specific operation of this step includes using two paths, chain rule mining based on statistics and legal ontology logic reasoner based on large language model, to generate possible legal rules, and then using large language model LLM to generate a set of candidate potential legal rule chains from these possible legal rules.

[0065] The formal definition includes: the legal knowledge graph is represented as G, the legal knowledge ontology sorted out by experts is represented as O, the set of potential legal rule chains is represented as Z, and any chain legal rule is represented as {r1,r2,...,r l}, the legal rule mining query operation is represented as q1, and the legal rule reasoning query operation is represented as q2; Figure 2 As shown, it further includes the following detailed processes:

[0066] Step 1.1, construct a statistical-based legal rule mining model, generate a potential legal rule chain set Z1 through the statistical-based legal rule mining model, and specifically use probabilistic rule mining algorithms such as RNNLogic, RLogic, and NCRL to take the potential legal rule chain set Z as the input of the model, and select any chain rule {r1, r2, ..., r ... l} as the legal rule mining chain, the legal rule mining query operation is represented by q1, and the result of the query operation q1 obtained through mining is the equivalent rule 1 {r1,r2,...,r l}→r a1 As the output of the mining model, the probability distribution p of the potential legal rule chain set Z1 generated by the statistical chain rule mining in this step is calculated according to the output of the model. g,e (A|G,q), the expression is as follows:

[0067] p g,e (A|G,q)=∑p e (A|G,q,Z)·p g (Z|G,q)

[0068] Among them, p g To generate the probability distribution of potential legal rule chains through statistical chain-based legal rule mining, p e is the evaluation result of the potential legal rule chain, A is the equivalent legal rule r a the assessment results;

[0069] Step 1.2, constructing a legal ontology logic reasoning model, generating a potential legal rule chain set Z2 through a legal ontology logic reasoning model based on a large language model, including obtaining a legal ontology logic reasoning model based on a large language model, wherein the legal ontology logic reasoning model based on a large language model uses the legal knowledge ontology sorted out by experts as the input of the model,

[0070] Specifically, we obtain hints from the legal knowledge ontology previously sorted out by experts, and then use the hints to enrich the general large language model. Then, we define the reasoning ability of the enriched large language model in dealing with legal issues as a legal ontology logic reasoner based on the large language model. We select any chain rule {r1, r2, ..., r3} from the potential legal rule chain set Z. l} as the legal rule mining chain, the legal rule reasoning query operation is represented by q2, and the result of the query operation q2 obtained by reasoning is the equivalent rule 2 {r1,r2,...,r l}→r a2 As the output 2 of the model, the probability distribution p of the potential legal rule chain set 2 Z2 generated by the legal ontology logic reasoning model based on the large language model is calculated according to the output 2 of the reasoning model. gLM (Z|G,O,q), the expression is as follows:

[0071]

[0072] Among them, p LM () is the probability of the large language model, T1(O) is the function of converting the legal knowledge ontology O sorted by experts into LLM prompts, z <i is all the legal rules before the i-th order legal rule in Z, p gLMTo generate a probability distribution of possible legal rule chains by mining using statistics-based chain rules;

[0073] Among them, the equivalent rule is the definition of the rule that is equivalent in logical reasoning. For example, after a multi-step reasoning, it is equivalent to a single-step reasoning. There is a logical reasoning relationship between title 1 and title 2. For example: individual 1 is title 1 of individual 2, individual 2 is title 1 of individual 3, then individual 1 is title 2 of individual 3, and the equivalent rule is title 2 of title 2 is title 3; formally expressed: r1 is title 2, r2 is title 3, and the equivalent rule is: {r1, r1}->r2. The equivalent rule is also the target relationship of logical reasoning;

[0074] Step 1.3, take the union of legal rule chain set 1 Z1 and legal rule chain set 2 Z2 to obtain legal rule chain set Z:

[0075] Z=Z1∪Z2

[0076] Step 2, evaluate the rationality of the potential chain legal rule set Z generated by the dual-path mining in step 1. Specifically, evaluate all output legal rule chains in two dimensions, such as Figure 3 As shown, it further includes the following detailed processes:

[0077] Step 2.1, using the legal logic rationality voting of the large language model to evaluate the legal logic rationality of the generated potential legal rule chain set, specifically including using the legal ontology prompter to integrate the legal ontology semantic knowledge of each legal rule chain into the prompt of the large language model, and through reasoning with the large language model, complete the identification of the rationality of the potential legal rule chain at the level of legal common sense, and generate a logical rationality confidence score as the result of the logical rationality voting; wherein, the legal ontology prompter is used to generate prompts according to the semantic information of the legal ontology (CLCF);

[0078] Specifically, before using a large language model to process the potential legal rule chain data, each first-order reasoning logical relationship (such as individual 1 is the title 1 of individual 2) is decomposed and the semantic information of the domain and scope of the relationship is extracted from the legal ontology. This semantic information is written in the form of triples (legal entity subject, legal entity predicate, legal entity object), such as "[entity:subject value][link:predicate value][entity:objectvalue]", and then the ontology triples corresponding to these relationships are connected.

[0079] To illustrate the specific process, take the chain rule "LawChanges, LawChanges, PromulgateDecree→PromulgateDecree" as an example: the legal meaning of the LawChanges relationship is "Law A is an amendment to Law B", written as [entity:LawA][link:LawChanges][entity:LawB]; the legal meaning of the PromulgateDecree relationship is "the legislative body of the law is a certain body", written as [entity:Law][link:PromulgateDecree][entity:State Organs]. This chain rule can be rewritten into a format that includes ontology information:

[0080] [entity:LawA][link:lawChanges][entity:LawB]^

[0081] [entity:LawB][link:lawChanges][entity:LawC]^

[0082] [entity:LawC][link:PromulgateDecree][entity:State0rgamsD]→

[0083] [entity:LawA][link:PromulgateDecree][entity:State0rgamsD].

[0084] In the legal ontology prompter, a large language model is used to analyze and explain whether this logical reasoning chain is reasonable by using the thinking chain method. The prompt is divided into two steps: a) Based on legal common sense, understand the mutual connection between triples, analyze the relationship and logical reasoning between triples, follow the laws of logical reasoning such as transitivity, as well as similar case judgments and causal relationships in legal reasoning, to build a logical chain showing how to reason from one triple to the next, including the legal rules or principles used in reasoning.

[0085] The following are the instructions for the large language model in step a), Instruction 1: "As a legal expert, explain the input triple rule chain and its reasoning from a legal perspective, and provide a step-by-step explanation. The explanation should be as concise and detailed as possible; the rule chain example is as follows: [Entity: LawA][Link: LawChanges][Entity: LawB]: Indicates that Law A has undergone changes, resulting in the emergence of Law B; [Entity: LawB][Link: LawChanges][Entity: LawC]: Indicates that Law B is transformed into Law C after some changes; [Entity: LawC][Link: PromulgateDecree][Entity: StateOrgansD]: Indicates that Law C is promulgated by the decree issuing agency D; [Entity: LawA][Link: PromulgateDecree][Entity: StateOrgansD]: Indicates that Law A is ultimately promulgated by the decree issuing agency D."

[0086] T1 = LLM (instruction 1, triple chain rule to be evaluated)

[0087] b) Evaluate the legitimacy of the legal rule chain, evaluate its rationality based on the established logical chain under legal common sense and the existing legal system, and whether all inferences are based on legal grounds. Calculate the legal logic rationality confidence score, and provide a legal logic rationality confidence score that reflects the logical strength of the legal rule chain and takes into account the legal uncertainty in the real world, given the rigor of the legal rule chain and the rationality of the legal logic.

[0088] The result of concatenating step a) is used as the input instruction of the large language model in step b), instruction 2: "As a legal expert, [triple chain rule to be evaluated, interpreted as {T1}], evaluate the rationality of the reasoning chain based on the rationality of legal logic; the rule chain provides a confidence level between 0 and 1."

[0089] T2=LLM(instruction 2)

[0090] Repeat this process multiple times, and take the arithmetic mean of the legal logic rationality confidence score results as the final legal logic rationality voting result.

[0091] The thought chain process representation in the legal ontology prompter is as follows:

[0092] p(C|R,E)=p(C|R,E,A)p(A|R,E)

[0093]

[0094] p(C|R,E)=p(C|R,E,A)p(A|R,E)

[0095] Where p(C|R,E) is the value of the given chain rule relation R, the chain rule interpretation E and the equivalent rule r. a1 、r a2 The total probability of the evaluation of the confidence score C under the evaluation result A, p(A|R,E) is the evaluation probability of the chain rule and its interpretation;

[0096] Repeat the execution of the legal ontology prompter to use the large language model to analyze and explain whether this logical reasoning chain is reasonable by using the thinking chain method, and record the legal logic rationality confidence score results, and take the arithmetic mean as the final legal logic rationality confidence score N AR , as shown below:

[0097]

[0098] Step 2.2, using the fact-oriented factual reasoning chain credibility voting of the large language model to evaluate the credibility of the legal factual reasoning chain of the generated potential legal rule chain set, specifically including using the factual chain prompter to calculate the credibility confidence score of the legal factual reasoning chain, through the aforementioned mining in a sparsely linked legal knowledge graph, obtaining an incomplete potential legal rule chain to construct a knowledge graph retrieval statement, searching for qualified triple fact data in the knowledge graph and performing random sampling, using the large language model to infer and calculate the credibility confidence score of each legal factual reasoning chain, and generating the corresponding legal factual reasoning chain credibility confidence score as the legal logic rationality voting result;

[0099] Specifically, based on the mined multi-hop legal rule chain, a Cypher query statement of the knowledge graph is generated to query the graph database, and the results are then randomly sampled to produce K records; if there are fewer than K results, the complete data set will be retained; then the fact chain prompter is enabled to traverse these sampled records and construct appropriate fact evaluation prompts.

[0100] For example, consider the query of the chain rule “LawChanges,LawChanges,PromulgateDecree→PromulgateDecree”, which produces the following sample facts: “[Law Change One][link:LawChanges][Law Change Two]^[Law Change Two)][link:LawChanges][Law Change Three]^[Law Change Three][link:PromulgateDecree][Law Promulgation Behavior]→[Law Change Version Four][link:PromulgateDecree][Law Promulgation Behavior]”.

[0101] The fact chain prompter follows a three-step thinking chain to determine the validity of facts: (1) The interpretation of the fact chain requires a clear definition and explanation of each element, their respective legal provisions and their interrelationships; (2) The evaluation of the validity of factual inferences involves analyzing the likelihood of factual conclusions based on the previously explained chain; (3) The confidence level calculation is completed by assigning a value from 0 to 1 after the inference validity evaluation, thereby quantifying the reliability of the inference. The formal expression of the thinking chain process of fact evaluation is consistent with the formula of the legal ontology prompter, except that the input data is facts. The credibility confidence score F of the legal fact reasoning chain is obtained by taking the arithmetic average of the confidence scores of K fact chains AR , as shown below:

[0102]

[0103] Step 2.3, integrate the confidence score of legal logic rationality and the confidence score of legal fact reasoning chain credibility, specifically including weighted fusion of the confidence scores of the above two steps, and generate the weighted confidence assessment score AR of each potential legal rule chain by weighting and integrating the confidence score of legal logic rationality and the confidence score of legal fact reasoning chain credibility, as shown in the following formula:

[0104] AR=ω1·N AR +ω2·F AR

[0105] In the formula, ω1 and ω2 are weight parameters, and ω1+ω2=1.

[0106] This weighted confidence score AR is compared with the predefined reliability identification threshold L1, and the legal rule chains above the threshold are filtered out, thereby constructing a reliable legal rule chain set; thus, the generative multi-hop legal rule chain mining and evaluation method finally achieves the extraction of reliable multi-hop legal reasoning chain rules from the legal knowledge graph;

[0107] Specifically, the fact-oriented credibility assessment of the legal fact reasoning chain further includes: using a fact chain prompter to convert a candidate potential legal rule chain into a graph query statement, randomly extracting a matching triple chain from the legal knowledge graph, and using the fact chain prompter to generate fact assessment prompt words to assess and score the credibility of the fact reasoning chain;

[0108] Step 3 is to conduct rationality arbitration of the legal rule chain. Specifically, for example, a rationality arbitrator is constructed to perform weighted synthesis on the legal logic rationality evaluation of the legal rule chain obtained in step 2 and the factual reasoning chain credibility evaluation results of the legal rule chain obtained in step 3, and highly credible legal rule chains are filtered out from the candidate rule set.

[0109] like Figure 4As shown, the overall framework of a specific implementation scheme of a generative multi-hop legal rule chain mining and evaluation algorithm of the present invention is shown, and the numbers ①-④ in the figure correspond to the above steps.

[0110] In summary, the key points of the present invention are summarized as follows:

[0111] (1) Generative multi-hop chain rule mining: Using advanced generative language models and statistical methods to extract multi-hop chain legal rules from incomplete legal knowledge graphs, this method can discover implicit and complex association rules and fill in the missing connections between entities in the existing legal knowledge graph.

[0112] (2) A novel technical framework for automatic mining and evaluation of legal rule chains was constructed, including four parts: "candidate chain rule generator", "legal ontology prompter", "fact chain prompter" and "legal rule rationality arbitrator". The unique design of the functional settings and interaction methods of these four components is the key to achieving high-quality legal rule mining.

[0113] (3) Dual-path mining to generate a set of potential legal rule chains as candidates: Through the dual paths of "chain rule generator based on generative language model" and "rule mining algorithm based on statistical method", it helps to improve the quality and coverage of rule generation.

[0114] (4) Dual evaluation strategy: Combining the evaluation of legal logic by the “legal ontology prompter” and the evaluation of the rationality of the fact chain by the “fact chain prompter”, a comprehensive rule verification mechanism is provided, which significantly improves the accuracy and reliability of rule mining.

[0115] (5) Legal rule rationality arbitrator: This component comprehensively considers the evaluation results of legal logic and factual rationality, and conducts the final rule rationality arbitration to ensure that the generated legal rules are logically rigorous and consistent with actual legal facts.

[0116] The present invention is not limited to the description of the above technical solutions. Any modifications and variations made by those skilled in the art without departing from the spirit and scope of the present invention shall fall within the scope defined by the attached claims.

Claims

1. A generative multi-hop legal rule chain mining and evaluation method, characterized in that: include: Step 1, using two paths, chain rule mining based on statistics and legal ontology logic reasoning based on a large language model, to generate possible legal rules, so as to generate a candidate potential legal rule chain set Z, wherein the potential legal rule chain set Z further includes a potential legal rule chain set Z1 generated by chain rule mining based on statistics and a potential legal rule chain set Z2 generated by a legal ontology logic reasoning model based on a large language model; Step 2: Evaluate the rationality of the potential chain legal rule set Z generated by the dual-path mining in step 1. Specifically, evaluate all potential legal rule chains from two dimensions, including: i. Use the legal logic rationality voting of the large language model to evaluate the legal logic rationality of the generated potential legal rule chain set, including using the legal ontology prompter to integrate the legal ontology semantic knowledge of each legal rule chain into the prompt of the large language model, and through reasoning with the large language model, complete the identification of the rationality of the legal rule chain at the level of legal common sense, and generate a logical rationality confidence score as the result of the logical rationality voting; ii. Using the fact-oriented factual reasoning chain credibility voting of the large language model to evaluate the credibility of the legal fact reasoning chain of the generated potential legal rule chain set, including using the fact chain prompter to calculate the credibility confidence score of the legal fact reasoning chain, constructing the knowledge graph retrieval statement through the aforementioned mining of incomplete legal rule chains, searching for qualified triple fact data in the knowledge graph and performing random sampling, using the large language model to infer and calculate the credibility confidence score of each legal fact reasoning chain, and generating the corresponding legal fact reasoning chain credibility confidence score as the legal logic rationality voting result; iii. Combine the confidence scores of legal logic rationality and legal fact reasoning chain credibility to generate a weighted confidence assessment score for each potential legal rule chain; Step 3: Conduct arbitration on the rationality of the legal rule chain. Specifically, perform weighted synthesis on the evaluation results of the legal logic rationality of the legal rule chain obtained in step 2 and the credibility evaluation results of the factual reasoning chain of the legal rule chain obtained in step 3, and filter out highly credible legal rule chains from the potential legal rule chain set Z.

2. A generative multi-hop legal rule chain mining and evaluation method as claimed in claim 1, characterized in that: The relevant process of generating a potential legal rule chain set Z1 based on statistical chain rule mining further includes: Construct a legal rule mining model based on statistics, take the potential legal rule chain set Z as the input of the model, select any chain rule {r1,r2,...,r l } as the legal rule mining chain, the legal rule mining query operation is represented by q1, and the result of the query operation q1 obtained through mining is equivalent to rule 1 (target relationship) {r1,r2,...,r l }→r a1 As the output of the mining model, the probability distribution p of the potential legal rule chain set Z1 generated by the statistical chain rule mining in this step is calculated according to the output of the model. g,e (A|G,q), the expression is as follows: p g,e (A|G,q)=∑p e (A|G,q,Z)·p g (Z|G,q) Among them, p g To generate the probability distribution of potential legal rule chains through statistical chain legal rule mining, p e is the evaluation result of the potential legal rule chain, A is the equivalent legal rule r a The evaluation results are as follows: G is the legal knowledge graph representation, and Z is the set representation of potential legal rule chains.

3. A generative multi-hop legal rule chain mining and evaluation method as claimed in claim 1, characterized in that: The relevant process of generating the potential legal rule chain set Z2 based on the legal ontology logic reasoning model of the large language model further includes: Construct a legal ontology logic reasoning model to obtain a legal ontology logic reasoning model based on a large language model. The legal ontology logic reasoning model based on a large language model takes the legal knowledge ontology sorted out by experts as the input of the model, selects any chain rule {r1,r2,...,r l } as the legal rule mining chain, the legal rule reasoning query operation is represented by q2, and the result of the query operation q2 obtained by reasoning is the equivalent rule 2 (target relationship) is {r1,r2,...,r l }→r a2 As the output 2 of the model, the probability distribution p of the potential legal rule chain set 2 Z2 generated by the legal ontology logic reasoning model based on the large language model is calculated according to the output 2 of the reasoning model. gLM (Z|G,O,q), the expression is as follows: Among them, p LM () is the probability of the large language model, T1(O) is the function of converting the legal knowledge ontology O sorted by experts into LLM prompts, z <i is all the legal rules before the i-th order legal rule in Z, p gLM To generate probability distributions of possible legal rule chains by mining using statistics-based chain rules.

4. A generative multi-hop legal rule chain mining and evaluation method as claimed in claim 1, characterized in that: The legal ontology logic reasoning model based on the large language model specifically includes obtaining prompts from the legal knowledge ontology previously sorted out by experts, and then using the prompts to enrich the general large language model, and then defining the reasoning ability of the enriched large language model in dealing with legal issues.

5. A generative multi-hop legal rule chain mining and evaluation method as claimed in claim 1, characterized in that: Take the union of legal rule chain set 1 Z1 and legal rule chain set 2 Z2 to obtain legal rule chain set Z, the expression is as follows: Z=Z1∪Z2.

6. A generative multi-hop legal rule chain mining and evaluation method as claimed in claim 1, characterized in that: The thought chain process in the legal ontology prompter is represented as follows: p(C|R,E)=p(C|R,E,A)p(A|R,E) p(C|R,E)=p(C|R,E,A)p(C|R,E) Where p(C|R,E) is the value of the given chain rule relation R, the chain rule interpretation E and the equivalent rule r. a1 、r a2 The total probability of evaluating the confidence score C under the evaluation result A, p(A|R,E) is the evaluation probability of the chain rule and its explanation.

7. A generative multi-hop legal rule chain mining and evaluation method as claimed in claim 1, characterized in that: In the fact chain prompter, the Cypher query statement of the knowledge graph is generated by the mined multi-hop legal rule chain, the query results of the graph database are randomly sampled, and the fact chain prompter is enabled to traverse these sampled records and construct appropriate fact evaluation prompts.

8. A generative multi-hop legal rule chain mining and evaluation method as claimed in claim 1, characterized in that: The expression of the weighted confidence assessment score AR for each potential legal rule chain is as follows: AR=ω1·N AR +ω2·F AR In the formula, ω1 and ω2 are weight parameters, and ω1+ω2=1, F AR is the credibility confidence score of the legal fact reasoning chain, N AR It is the confidence score of legal logic rationality.

9. A generative multi-hop legal rule chain mining and evaluation method as claimed in claim 1, characterized in that: The agent natural language memory stream module is used to store the observation value text data of the current rule of law situation perceived quantitatively.

Citation Information

Patent Citations

  • Large legal data management system based on fuzzy inference

    CN106204366A

  • Automatic analysis method for foreign-related legal information text based on retrieval enhanced thinking chain

    CN118761865A

  • Legal element analysis method and system based on large language model and knowledge graph

    CN119046476A

  • Knowledge graph construction method and device

    US20190019088A1

  • Voting-based consensus method

    WO2019232789A1

Cited By

  • Ternary fusion knowledge processing system and method based on ontology, graph database and large language model

    CN122242682A