A Legal Concept Interpretation Method Based on Multi-Perspective Clustering and Hybrid Retrieval

By employing a legal concept interpretation method that combines multi-perspective clustering and hybrid retrieval, this approach addresses the issues of unreliable LLM-generated content and illusion phenomena in existing technologies. It achieves efficient, accurate, and diverse interpretations of legal concepts, meeting the rapid comprehension needs of legal professionals and general users.

CN119760113BActive Publication Date: 2025-11-14SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411843859.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-14
Publication Date
2025-11-14
Estimated Expiration
2044-12-14

AI Technical Summary

Technical Problem

Existing methods for interpreting legal concepts rely on Large Language Models (LLMs), which suffer from weak professionalism in generated content, outdated knowledge, unreliable generated content, and illusion phenomena. Furthermore, existing Retrieval-Augmented Generation (RAG) models suffer from issues such as user ambiguity affecting output, and LLMs being limited by information retrieval databases.

Method used

This approach employs multi-perspective clustering and hybrid retrieval methods. Through steps such as semantic hierarchical deconstruction, dynamic unit parsing, recursive information retrieval, depth-first aggregation, cross-source information fusion, multi-level knowledge mapping, and reflective reasoning, combined with a large language model, it acquires a diverse set of cases, reduces the probability of hallucinations, and improves the credibility of the generated data.

Benefits of technology

It enhances the credibility of legal concept interpretations, improves the quality and efficiency of generated answers, meets the needs of legal professionals and ordinary users for quick understanding, reduces users' time and effort for repeated searches, and provides a unified presentation of diverse information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119760113B_ABST
    Figure CN119760113B_ABST
Patent Text Reader

Abstract

This application discloses a legal concept interpretation method based on multi-perspective clustering and hybrid retrieval. The method first parses user input into legal concepts. Second, it extracts diverse relevant legal cases using multi-perspective clustering, then extracts the relevant legal provisions from these cases and retrieves them from a database. A large language model is then used to generate the legal provisions that the user query might involve (i.e., the hybrid retrieval step). Next, the specific definition of the parsed concept is retrieved from the legal concept database, and keyword information given in relevant legal cases and provisions is searched. These keywords are used as relevant concepts and retrieved, and the set of parsed concepts and keywords is taken as the total set of relevant concepts. Finally, the relevant concepts and relevant legal provisions are organized into prompts and input into the large language model to generate initial results. The generated results, input prompts, and initial questions are then fed into the large language model for further processing, resulting in the final result. The legal cases are then appended to the generated results and output to the user for viewing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a legal concept interpretation method based on multi-perspective clustering and hybrid retrieval, belonging to the field of natural language processing technology. Background Technology

[0002] A significant source of legal concept interpretation is the application of the term in past court cases. Therefore, traditional methods of legal concept interpretation typically revolve around human legal professionals who consult relevant legal cases to understand past interpretations of legal concepts and then summarize them in light of current circumstances to arrive at a practical legal interpretation. However, the sheer number of legal concepts makes it impossible for human legal professionals to cover everything comprehensively, potentially leading to biased interpretations that compromise objectivity.

[0003] Large Language Models (LLMs) refer to Transformer language models containing hundreds of billions (or more) of parameters. They exhibit powerful natural language understanding capabilities and the ability to solve complex tasks through text generation. The main challenges faced by human legal professionals revolve around the sheer volume of cases and limited memory and reasoning abilities. Because LLMs are not limited by the latter, they are used to solve the problem of legal concept interpretation. Humans act as users, providing prompts to the LLM. Current LLM legal concept interpretation models break down user input into one or more legal concepts, then organize the legal concepts, user input, and specially designed formatted information into prompts input into the LLM, and finally, the LLM provides a formatted legal interpretation.

[0004] LLM, as a legal concept interpretation model, also faces numerous challenges. These include weak professionalism in generated content, outdated knowledge, unreliable or fabricated content. A key cause of these problems is the lack of quality assurance in the relevant cases automatically invoked by LLM to interpret legal concepts. Furthermore, LLM itself has a tendency to hallucinate, i.e., "unconsciously" using internal and external invocations lacking precise control. These problems are partly due to an insufficient number of quality-assured cases. Therefore, previous researchers proposed the Retrieval-Augmented Generation (RAG) method. A common approach is to acquire large amounts of data through an Information Retrieval (IR) system. This data comes from a specially maintained database, ensuring a certain level of quality and thus guaranteeing the quality of the generated legal interpretations.

[0005] However, the RAG model still has a series of problems, such as the user's ambiguous expression of true intent seriously affecting the output, the LLM being severely limited by the database of IR calls, and the lack of a good mechanism to organize the generation of illusions, etc.

[0006] To mitigate these issues, it is insufficient to rely solely on the individual databases invoked by IR (Investigational Retrieval) or to use only LLM (Limited Law Management) itself. Therefore, a hybrid retrieval strategy is formed by combining IR retrieval based on multiple databases with LLM retrieval. Furthermore, highly diverse case information helps reduce the likelihood of illusions; thus, multi-perspective clustering methods for obtaining highly diverse case sets from search results have also been applied to the problem of legal concept interpretation. Summary of the Invention

[0007] This invention addresses the problem of legal concepts being prone to illusion, thus reducing the credibility of concept interpretations. It proposes a legal concept interpretation model based on multi-perspective clustering and hybrid retrieval. A hybrid retrieval approach combining database (IR system) retrieval and LLM retrieval is employed to broadly obtain high-quality legal content, while multi-perspective clustering ensures case diversity, thereby reducing the probability of illusion and improving the credibility of the generated interpretations.

[0008] The technical solution of the present invention is as follows:

[0009] Step 1) Semantic separation of queries based on semantic hierarchical deconstruction, used to distinguish between legal and non-legal concepts;

[0010] Step 2) Semantic fragment extraction based on dynamic unit parsing, performed according to the semantic separation method in Step 1), is used to capture legal concepts describing the relationship between legal terms;

[0011] Step 3) A case acquisition mechanism based on recursive information retrieval and depth-first aggregation is used to systematically collect relevant legal cases;

[0012] Step 4) A case enhancement mechanism based on multi-perspective clustering, implemented according to the case acquisition mechanism in Step 3), is used to integrate and optimize case information;

[0013] Step 5) Preliminary screening of legal provisions based on cross-source information fusion, used to filter relevant legal provisions from multiple data sources;

[0014] Step 6) Extracting legal provisions based on semantic enhancement retrieval, which is performed according to the preliminary screening of legal provisions in Step 5), is used to filter legal content that is more consistent with the query.

[0015] Step 7) Construct a knowledge mapping system based on multi-level partitioning to systematically classify legal concepts hierarchically;

[0016] Step 8) Based on the knowledge mapping constructed in Step 7), dynamic concept matching and supplementation based on similarity analysis is used to match and supplement relevant legal concepts;

[0017] Step 9) A feedback mechanism based on dynamic prompts is used to optimize the presentation of search results in steps 4), 6), and 8.

[0018] Step 10) An output adjustment strategy based on reflective reasoning is used to dynamically adjust the output of step 9) to improve performance;

[0019] Specifically, the query semantic separation based on semantic hierarchical deconstruction is as follows:

[0020] Let Q be the user query (i.e., the text entered by the user). To effectively parse this query, we first deconstruct its semantic structure hierarchically, constructing two query categories: the first category, S1, represents the set of indivisible legal concept queries, and the second category, S2, represents the set of indivisible non-legal concept queries. Where:

[0021] S1={c1,c2,…,c m} refers to a set containing m legal terms, where c i This represents the i-th legal term.

[0022] S2={n1,n2,…,n k} refers to a set containing k non-legal terms, where n j This represents the j-th non-legal term.

[0023] The semantic fragment extraction based on dynamic unit parsing is as follows:

[0024] For user query Q, the legal terms involved are extracted and organized into first-type query instances using the dynamic unit parsing method. For the query text Q = (x1, x2, ..., x... n Introduce a labeled sequence y = (y1, y2, ..., y...) n ), where each y i It is for x i After categorizing, a label is assigned, indicating whether the corresponding term is a legal term. Then, the probability is maximized.

[0025]

[0026] Subsequently, the text words corresponding to the labels identified as legal terms in the obtained y are extracted to form the first type of query instance q1. All the first type of query instances are extracted to form the first type of query instance set Q1. Then, based on the relationships between different legal terms in Q, the second type of query instances are organized using a relation extraction method. Given a query Q and two entities e within it. i and e j Find r ij The value of is chosen to maximize the probability.

[0027] P(r ij |Q,e i ,e j )

[0028] Where r ij Represents e i and e j If there is a correlation between them, set the result to 1; otherwise, set it to 0. Then extract the result that makes r... ij For ordered pairs (i,j) where 1 = 1, extract the two corresponding text words and connect them with "AND" to form a second type of query instance q2. Extract all second type of query instances to form a second type of query instance set Q2. Subsequently, use the first type of query instance set Q1 and the second type of query instance set Q2 to query the legal knowledge base to obtain the first type of query concept C1 and the second type of query concept C2.

[0029] The union of all legal concept sets is the analytic concept set C. p C p =C1∪C2.

[0030] The case acquisition mechanism based on recursive information retrieval and depth-first aggregation is as follows:

[0031] 3-1 Concept Search and Relevance Assessment

[0032] Establish a legal case database When retrieving legal cases related to user query Q, first extract a subset of concepts from the parsed concept set. For each concept A collection of legal cases was obtained by searching relevant legal cases.

[0033]

[0034] Where rel(d,c) represents the relevance function between case d and concept c, and at the same time

[0035]

[0036] in, and θ1 represents the embeddings of case d and concept c obtained through Word2Vec, respectively. θ1 represents the relevance threshold, which is set to 0.5. This parameter varies depending on the actual situation and can be obtained using the grid search method.

[0037] 3-2 Depth-First Aggregation

[0038] The acquired legal case set is aggregated to obtain a related case set under the output depth limit.

[0039]

[0040] Among them, dep(d i (For case d) i The depth evaluation function, at the same time

[0041]

[0042] Where, level(d) i ,d j ) represents d i and d j In legal case database The shortest distance on the [parameter range]. This represents the hierarchical relationship with the user query, where k is the depth-first traversal constraint.

[0043] The case enhancement mechanism based on multi-perspective clustering is as follows:

[0044] 4-1 Classification and Sampling

[0045] To obtain more representative legal cases, a categorized set of entries was established based on this. These entries indicate the classification criteria for each legal case. For each classification entry t i From the case Extracting from t i Content of the corresponding entry The feature v is obtained by embedding using Word2Vec. i (d):

[0046]

[0047] From all cases A feature vector obtained from the embedding is randomly selected as the initial cluster center ce1. Then, if there are j cluster centers, the calculation is performed for each case. The Euclidean distance D to the nearest cluster center i (d):

[0048]

[0049] Then, based on the selection probability P c To select the next cluster center ce j+1 The probability of selection P c Defined as

[0050]

[0051] Suppose there are K cluster centers.

[0052] Ce = {ce1,ce2,…,ce} K}

[0053] Classification result function r i for:

[0054]

[0055] Where, r i For the classification result, indicate the case d in the classification entry t. i The classification results were then integrated into a single classification index, cr.

[0056]

[0057] The parentheses represent the floor function, which rounds the relevant case sets based on the clustering criterion cr. Divided into classification result set Right now

[0058]

[0059] 4-2 Output a collection of relevant legal cases

[0060] from A case is randomly selected from the data. As representative cases, the relevant legal case collection for:

[0061]

[0062] Among them, the preliminary screening of legal provisions based on cross-source information fusion is as follows:

[0063] 5-1 Construction of a Legal Provisions Database

[0064] First, data from multiple legal provisions. Extract legal text data and organize it into a database.

[0065]

[0066] Each data source The set of clauses in the text is represented as Among them l ij This represents the j-th clause.

[0067] 5-2 Text Extraction and Information Fusion

[0068] The proposed extraction method ε(l) ijc), this method indicates from clause l ij Extracting the concept c∈C p Related content:

[0069]

[0070] Among them, rel(l ij c) represents a measure of the relevance between the clause and the concept.

[0071]

[0072] Here, `Tokenizer(·)` represents using an LLM-based tokenizer to segment the text, resulting in a set of words. `θ2` represents the relevance threshold, set to 0.3. This parameter varies depending on the specific situation and can be obtained using a grid search method. Based on this, a cross-source information fusion method is used to merge relevant entries from various data sources into a unified set of entries.

[0073]

[0074] The precise extraction of legal provisions based on semantic enhancement retrieval is detailed below:

[0075] 6-1 Semantic Enhancement Retrieval and Entry Extraction

[0076] In the integrated set of articles Then, semantic enhancement retrieval methods were further used. To extract subsets of related concepts Relevant legal provisions. The goal of this method is to extract the most relevant provisions from the fused set of provisions based on their degree of matching with the concepts:

[0077]

[0078] Where sim(l,c) represents the similarity measure between clause l and legal concept c, and θ3 is the similarity threshold, set to 0.5. This parameter varies depending on the actual situation and can be obtained using the grid search method.

[0079] 6-2 Output of the final set of clauses

[0080] The legal provisions retrieved through semantic enhancement are represented as follows:

[0081]

[0082] This collection of articles will serve as the final output of the legal provisions extraction, ensuring that it covers the legal provisions relevant to the user's query.

[0083] The concept classification based on multi-level knowledge mapping is as follows:

[0084] 7-1 Construction of Knowledge Maps

[0085] First, utilize legal case collections Extracted set of relevant legal concepts and from the collection of legal provisions Extracted related concept set Constructing a preliminary knowledge mapping

[0086]

[0087] in, It represents the set of all related concepts.

[0088] 7-2 Hierarchical Division Concept Extraction

[0089] For legal concept set The concept c is defined using a hierarchical partitioning method H(c), which is based on an existing legal knowledge graph G. k The level of concept c h c Defined as the longest path length from c to the root node

[0090]

[0091] Where PathToRoot(c) is the set of all possible paths from c to the root node in the graph. len(p) represents the length of path p. The hierarchical partitioning tree T is constructed as follows:

[0092] First, categorize all concept nodes according to level h. c The categories are as follows:

[0093]

[0094] Represents all concepts at level i, and Where V g It is G k The set of points for G. k edge set E in k The edge set in the hierarchical partitioning tree is E. T :

[0095]

[0096] Finally, a hierarchical partitioning tree is constructed.

[0097] T = (V g E T )

[0098] For a specific concept c, its direct subconcepts are:

[0099] Child(c) = {c j ∈V g |(c,c j )∈E T}

[0100] Then the final knowledge mapping is formed.

[0101]

[0102] The dynamic concept matching and supplementation based on similarity analysis is as follows:

[0103] 8-1 Similarity Analysis and Conceptual Association

[0104] For the constructed hierarchical knowledge mapping, the similarity between two legal concepts is evaluated, with a similarity threshold of 0.6. This parameter varies depending on the specific circumstances and can be obtained using a grid search method. This yields a set of associated concepts.

[0105]

[0106] Detailed explanation and output of the 8-2 concept

[0107] Finally, a detailed explanation of each concept was retrieved from a legal concept database. Its detailed explanation is d c The detailed explanations of all concepts are compiled to form the final output set of related concepts.

[0108]

[0109] The feedback mechanism based on dynamically generated prompts is as follows:

[0110] Assume the user queries Q and the dynamic content generated during the execution of the steps: Step 6) Collection of relevant legal provisions Step 8) Collection of relevant legal concepts and step 4) a collection of relevant legal cases

[0111] The initial output O0 is

[0112] O0 = LLM(Q)

[0113] The feedback A from the model combined with dynamically generated content is:

[0114] concat(O0, "Relevant legal provisions: ", "Collection of relevant legal concepts:"

[0115] , "Collection of relevant legal cases:"

[0116] Here, concat represents the string concatenation operation.

[0117] The output adjustment strategy based on reflective reasoning is as follows:

[0118] The user query is Q, and the output of step 9) is A. Organizational reflection prompt P r :

[0119] concat(“Please check query”, Q, “and answer”, A′, “Does the subject match? Please answer yes or no”)

[0120] Here, `concat` represents the string concatenation function. Reflecting on the conclusion, O satisfies...

[0121] O = LLM(P r )

[0122] Here, LLM(·) represents content generation using a large language model. If the reflection result is "yes", then the reflection conclusion O at this time is taken as the final output; if the reflection conclusion is "no", then a reflection reasoning hint P is generated. i :

[0123] concat(“query”,Q,“and answer”,A′,“subject does not match, please modify the inconsistent parts of the answer”)

[0124] Then update A′:

[0125] A′=LLM(P i )

[0126] And repeat the reflection prompts and their follow-up content.

[0127] Compared with existing technologies, the advantages of this invention are:

[0128] 1) This application is the first to propose applying multi-perspective clustering to ensure diversity in the mixed retrieval of relevant cases in the interpretation of legal concepts.

[0129] 2) This application adopts a hybrid search strategy, which combines a large language model and an information retrieval system for interpreting legal concepts. This approach utilizes the reliable information sources of the information retrieval system to reduce the illusion problem caused by the quality of citations in the large language model, and can also avoid the serious decline in the quality of answers caused by relying solely on an information retrieval system with quality issues.

[0130] 3) This application adopts a reflective strategy, that is, in multiple stages of the execution process involving the large language model, the large language model is re-engineered.

[0131] Improving the quality and accuracy of self-generated data can enhance the quality of generated answers to some extent. In practical applications, legal professionals and ordinary users often need to quickly find relevant cases, legal provisions, and conceptual definitions within a limited time. This method integrates the retrieval of legal cases, provisions, and conceptual definitions, as well as the expansion of related concepts, into an organic process. Through multi-perspective clustering and hybrid retrieval methods, it provides diverse information at once, reducing the time and effort users spend repeatedly switching search directions, thus significantly improving retrieval efficiency. Furthermore, for ordinary users without a legal background, simple legal provisions or professional dictionary explanations are often obscure and difficult to understand. This method presents authoritative definitions of concepts, related cases, and relevant legal provisions together, and finally uses a large language model to generate easy-to-understand explanations and summaries. This helps users quickly grasp the background and context of legal concepts, allowing them to understand legal concepts more intuitively and in a more contextualized way. Attached Figure Description

[0132] Figure 1 This is a diagram illustrating the implementation architecture of a legal concept interpretation method based on multi-perspective clustering and hybrid retrieval.

[0133] Figure 2 This is a flowchart showing the different steps involved in the implementation of this plan. Detailed Implementation

[0134] The implementation process of the present invention will be described in detail below with reference to the embodiments and the accompanying drawings.

[0135] Example 1: A legal concept interpretation method based on multi-perspective clustering and hybrid retrieval, with the following architecture diagram: Figure 1 As shown, it includes the following steps:

[0136] Step 1) Separation of query semantics based on semantic hierarchical deconstruction, as follows:

[0137] Let Q be the user query (i.e., the text entered by the user). To effectively parse this query, we first deconstruct its semantic structure hierarchically, constructing two query categories: the first category, S1, represents the set of indivisible legal concept queries, and the second category, S2, represents the set of indivisible non-legal concept queries. Where:

[0138] S1={c1,c2,…,c m} refers to a set containing m legal terms, where c i This represents the i-th legal term.

[0139] S2={n1,n2,…,nk} refers to a set containing k non-legal terms, where n j This represents the j-th non-legal term.

[0140] Step 2) Semantic fragment extraction based on dynamic unit parsing, such as Figure 1 As shown on the left, the details are as follows:

[0141] For user query Q, the legal terms involved are extracted and organized into first-type query instances using the dynamic unit parsing method. For the query text Q = (x1, x2, ..., x... n Introduce a labeled sequence y = (y1, y2, ..., y...) n ), where each y i It is for x i After categorizing, a label is assigned, indicating whether the corresponding term is a legal term. Then, the probability is maximized.

[0142]

[0143] Subsequently, the text words corresponding to the labels identified as legal terms in the obtained y are extracted to form the first type of query instance q1. All the first type of query instances are extracted to form the first type of query instance set Q1. Then, based on the relationships between different legal terms in Q, the second type of query instances are organized using a relation extraction method. Given a query Q and two entities e within it. i and e j Find r ij The value of is chosen to maximize the probability.

[0144] P(r ij |Q,e i ,e j )

[0145] Where r ij Represents e i and e j If there is a correlation between them, set the result to 1; otherwise, set it to 0. Then extract the result that makes r... ij For ordered pairs (i,j) where 1 = 1, extract the two corresponding text words and connect them with "AND" to form a second type of query instance q2. Extract all second type of query instances to form a second type of query instance set Q2. Subsequently, use the first type of query instance set Q1 and the second type of query instance set Q2 to query the legal knowledge base to obtain the first type of query concept C1 and the second type of query concept C2.

[0146] The union of all legal concept sets is the analytic concept set C. p C p =C1∪C2.

[0147] Step 3) The case acquisition mechanism based on recursive information retrieval and depth-first aggregation is as follows:

[0148] 3-1) Concept retrieval and relevance assessment

[0149] Establish a legal case database When retrieving legal cases related to user query Q, first extract a subset of concepts from the parsed concept set. For each concept A collection of legal cases was obtained by searching relevant legal cases.

[0150]

[0151] Where rel(d,c) represents the relevance function between case d and concept c, and at the same time

[0152]

[0153] in, and θ1 represents the embeddings of case d and concept c obtained through Word2Vec, respectively. θ1 represents the relevance threshold, which is set to 0.5. This parameter varies depending on the actual situation and can be obtained using the grid search method.

[0154] 3-2) Depth-first aggregation

[0155] The acquired legal case set is aggregated to obtain a related case set under the output depth limit.

[0156]

[0157] Among them, dep(d i ) is case d i The depth evaluation function, at the same time

[0158]

[0159] Wherein, level(d) i ,d j ) represents d i and d j In legal case database The shortest distance on the [parameter range]. This represents the hierarchical relationship with the user query, where k is the depth-first traversal constraint.

[0160] Step 4) Case enhancement mechanism based on multi-view clustering, such as Figure 1 As shown in the middle, the details are as follows:

[0161] 4-1) Classification and Sampling

[0162] To obtain more representative legal cases, a categorized set of entries was established based on this. These entries indicate the classification criteria for each legal case. For each classification entry t i From the case Extracting from t i Content of the corresponding entry The feature v is obtained by embedding using Word2Vec. i (d):

[0163]

[0164] From all cases A feature vector obtained from the embedding is randomly selected as the initial cluster center ce1. Then, if there are j cluster centers, the calculation is performed for each case. The Euclidean distance D to the nearest cluster center i (d):

[0165]

[0166] Then, based on the selection probability P c To select the next cluster center ce j+1 The probability of selection P c Defined as

[0167]

[0168] Suppose there are K cluster centers.

[0169] Ce = {ce1,ce2,…,ce} K}

[0170] Classification result function r i for:

[0171]

[0172] Where, r i For the classification result, indicate the case d in the classification entry t. i The classification results were then integrated into a single classification index, cr.

[0173]

[0174] The parentheses represent the floor function, which rounds the relevant case sets based on the clustering criterion cr. Divided into classification result set Right now

[0175]

[0176] 4-2) Output a set of relevant legal cases

[0177] from A case is randomly selected from the data. As representative cases, the relevant legal case collection for:

[0178]

[0179] Step 5) Preliminary screening of legal provisions based on cross-source information fusion, as follows:

[0180] 5-1) Construction of a legal provisions database

[0181] First, data from multiple legal provisions. Extract legal text data and organize it into a database.

[0182]

[0183] Each data source The set of clauses in the text is represented as Among them l ij This represents the j-th clause.

[0184] 5-2) Text Extraction and Information Fusion

[0185] The proposed extraction method ε(l) ij c), this method indicates from clause l ij Extracting the concept c∈C p Related content:

[0186]

[0187] Among them, rel(l ij c) represents a measure of the relevance between the clause and the concept.

[0188]

[0189] Here, `Tokenizer(·)` represents using an LLM-based tokenizer to segment the text, resulting in a set of words. `θ2` represents the relevance threshold, set to 0.3. This parameter varies depending on the specific situation and can be obtained using a grid search method. Based on this, a cross-source information fusion method is used to merge relevant entries from various data sources into a unified set of entries.

[0190]

[0191] Step 6) Precise extraction of legal provisions based on semantically enhanced retrieval, such as... Figure 1 As shown in the middle, the details are as follows:

[0192] 6-1) Semantic Enhancement Retrieval and Item Acquisition

[0193] In the integrated set of articles Then, semantic enhancement retrieval methods were further used. To extract subsets of related concepts Relevant legal provisions. The goal of this method is to extract the most relevant provisions from the fused set of provisions based on their degree of matching with the concepts:

[0194]

[0195] Where sim(l,c) represents the similarity measure between clause l and legal concept c, and θ3 is the similarity threshold, set to 0.5. This parameter varies depending on the actual situation and can be obtained using the grid search method.

[0196] 6-2) Output of the final set of clauses

[0197] The legal provisions retrieved through semantic enhancement are represented as follows:

[0198]

[0199] This set of provisions will serve as the final output of the legal provision extraction, ensuring that it covers legal provisions relevant to the user's query. Step 7) Concept classification based on multi-level knowledge mapping, as follows:

[0200] 7-1) Construction of Knowledge Mapping

[0201] First, utilize legal case collections Extracted set of relevant legal concepts and from the collection of legal provisions Extracted related concept set Constructing a preliminary knowledge mapping

[0202]

[0203] in, It represents the set of all related concepts.

[0204] 7-2) Hierarchical Division of Concepts

[0205] For legal concept set The concept c is defined using a hierarchical partitioning method H(c), which is based on an existing legal knowledge graph G. k The level of concept c hc Defined as the longest path length from c to the root node

[0206]

[0207] Where PathToRoot(c) is the set of all possible paths from c to the root node in the graph. len(p) represents the length of path p. The hierarchical partitioning tree T is constructed as follows:

[0208] First, categorize all concept nodes according to level h. c The categories are as follows:

[0209]

[0210] Represents all concepts at level i, and Where V g It is G k The set of points for G. k edge set E in k The edge set in the hierarchical partitioning tree is E. T :

[0211]

[0212] Finally, a hierarchical partitioning tree is constructed.

[0213] T = (V g E T )

[0214] For a specific concept c, its direct subconcepts are:

[0215] Child(c) = {c j ∈V g |(c,c j )∈E T}

[0216] Then the final knowledge mapping is formed.

[0217]

[0218] Step 8) Dynamic concept matching and supplementation based on similarity analysis, such as... Figure 1 As shown in the middle, the details are as follows:

[0219] 8-1) Similarity Analysis and Concept Association

[0220] For the constructed hierarchical knowledge mapping, the similarity between two legal concepts is evaluated, with a similarity threshold of 0.6. This parameter varies depending on the specific circumstances and can be obtained using a grid search method. This yields a set of associated concepts.

[0221]

[0222] 8-2) Detailed explanation and output of the concept

[0223] Finally, a detailed explanation of each concept was retrieved from a legal concept database. Its detailed explanation is d c The detailed explanations of all concepts are compiled to form the final output set of related concepts.

[0224]

[0225] Step 9) The feedback mechanism based on the dynamically generated prompts is as follows:

[0226] Assume the user queries Q and the dynamic content generated during the execution of the steps: Step 6) Collection of relevant legal provisions Step 8) Collection of relevant legal concepts and step 4) a collection of relevant legal cases

[0227] The initial output O0 is

[0228] O0 = LLM(Q)

[0229] The feedback A from the model combined with dynamically generated content is:

[0230] concat(O0, "Relevant legal provisions: ", "Collection of relevant legal concepts:"

[0231] , "Collection of relevant legal cases:"

[0232] Here, concat represents the string concatenation operation.

[0233] Step 10) Output adjustment strategies based on reflective reasoning, such as... Figure 1 As shown on the right, the details are as follows:

[0234] The user query is Q, and the output of step 9) is A. Organizational reflection prompt P r :

[0235] concat(“Please check query”, Q, “and answer”, A′, “Does the subject match? Please answer yes or no”)

[0236] Here, `concat` represents the string concatenation function. Reflecting on the conclusion, O satisfies...

[0237] O = LLM(P r )

[0238] Here, LLM(·) represents content generation using a large language model. If the reflection result is "yes", then the reflection conclusion O at this time is taken as the final output; if the reflection conclusion is "no", then a reflection reasoning hint P is generated. i :

[0239] concat(“query”,Q,“and answer”,A′,“subject does not match, please modify the inconsistent parts of the answer”)

[0240] Then update A′:

[0241] A′=LLM(P i )

[0242] And repeat the reflection prompts and their follow-up content.

[0243] Evaluation method:

[0244] Legal concepts and their corresponding interpretations were reviewed by human experts, focusing on different indicators x. i They were scored. The scoring range was 1-5 points: 1 point represents that the indicator was not met, and 5 points represents that the indicator was met. The average score of each expert was calculated. To represent this explanation in indicator x i The degree of satisfaction on a particular indicator. The overall satisfaction level on a particular indicator is the score for all indicators. Then calculate the average.

[0245] The evaluation model is primarily based on five metrics: Accuracy of concept resolution: Can the system accurately resolve user input to obtain appropriate legal concepts? Completeness of related concept discovery: Can the system resolve all legal concepts related to user input? Relevance of legal cases: Are the legal cases invoked by the system relevant to user input? Diversity of legal cases: Are the sources and content of the legal cases invoked by the system diverse? Depth of legal knowledge: Does the legal knowledge invoked by the system possess sufficient professionalism? Timeliness of concept explanation: Is the legal knowledge invoked by the system up-to-date?

[0246] Explanation method accuracy Integrity Correlation diversity depth Timeliness GPT-4 3.8 3.7 3.6 3.4 3.5 4.2 Humans 4.7 4.6 4.6 4.5 4.7 4.9 This method 4.4 4.5 4.5 4.3 4.2 4.6

[0247] The experimental results above show that this method performs in line with human legal experts in terms of explanatory methods, accuracy, completeness, and relevance, and surpasses them in terms of diversity and timeliness. This demonstrates that this method can effectively explore massive external knowledge and case databases, achieving more accurate conceptual explanations that meet the requirements of legal practice, and exhibiting significant advantages in improving response speed and multi-perspective analytical capabilities. This indicates that while maintaining high efficiency, this method can enhance the comprehensive handling of complex legal issues.

[0248] Next, an example of a legal concept interpretation generated by this method will be given.

[0249]

[0250]

[0251]

[0252] It should be noted that the above embodiments are not intended to limit the scope of protection of the present invention. Equivalent transformations or substitutions made based on the above technical solutions all fall within the scope of protection of the claims of the present invention.

Claims

1. A method for interpreting legal concepts based on multi-perspective clustering and hybrid retrieval, characterized in that, The method includes the following steps: Step 1) Semantic separation of queries based on semantic hierarchical deconstruction, used to distinguish between legal and non-legal concepts; Step 2) Semantic fragment extraction based on dynamic unit parsing, performed according to the semantic separation method in Step 1), is used to capture legal concepts describing the relationship between legal terms; Step 3) A case acquisition mechanism based on recursive information retrieval and depth-first aggregation is used to systematically collect relevant legal cases; Step 4) A case enhancement mechanism based on multi-perspective clustering, implemented according to the case acquisition mechanism in Step 3), is used to integrate and optimize case information; Step 5) Preliminary screening of legal provisions based on cross-source information fusion, used to filter relevant legal provisions from multiple data sources; Step 6) Extracting legal provisions based on semantic enhancement retrieval, which is performed according to the preliminary screening of legal provisions in Step 5), is used to filter legal content that is more consistent with the query. Step 7) Construct a knowledge mapping system based on multi-level partitioning to systematically classify legal concepts hierarchically; Step 8) Based on the knowledge mapping constructed in Step 7), dynamic concept matching and supplementation based on similarity analysis is used to match and supplement relevant legal concepts; Step 9) A feedback mechanism based on dynamic prompts is used to optimize the presentation of search results in steps 4), 6), and 8. Step 10) An output adjustment strategy based on reflective reasoning is used to dynamically adjust the output of step 9) to improve performance; Step 1) is query semantic separation based on semantic hierarchical deconstruction, as follows: Let the user query be Q. To effectively parse this query, we first deconstruct its semantic structure hierarchically, constructing two query categories: the first category, S1, represents the set of indivisible legal concept queries, and the second category, S2, represents the set of indivisible non-legal concept queries, where: S1={c1,c2,…,c m } refers to a set containing m legal terms, where c i Represents the i-th legal term. S2={n1,n2,…,n k } refers to a set containing k non-legal terms, where n j Represents the j-th non-legal term; Step 2) Semantic fragment extraction based on dynamic unit parsing, as follows: For user query Q, the legal terms involved are extracted and organized into first-type query instances using the dynamic unit parsing method. All first-type query instances constitute the first-type query instance set Q1. Then, based on the relationships between different legal terms in Q, they are organized into second-type query instances using a relation extraction method. All instances of the second type of query constitute the second type of query instance set Q2. Subsequently, the first type of query instance set Q1 and the second type of query instance set Q2 are used to query the legal knowledge base to obtain the first type of query concept C1 and the second type of query concept C2. The union of all legal concept sets is the analytic concept set C. p C p =C1∪C2; Step 3) The case acquisition mechanism based on recursive information retrieval and depth-first aggregation is as follows: 3-1 Concept retrieval and relevance assessment Establish a legal case database When retrieving legal cases related to user query Q, first extract a subset of concepts from the parsed concept set. For each concept A collection of legal cases was obtained by searching relevant legal cases. Where rel(d,c) represents the correlation function between case d and concept c, and θ1 represents the correlation threshold set to 0.

5. This parameter varies depending on the actual situation and can be obtained using the grid search method. 3-2 Depth-first aggregation, The acquired legal case set is aggregated to obtain a related case set under the output depth limit. Among them, dep(d i (For case d) i The depth evaluation function represents the hierarchical relationship with the user query, where k is the depth traversal limit; Step 4) Case enhancement mechanism based on multi-view clustering, as follows: 4-1 Classification and Sampling To obtain more representative legal cases, a categorized set of entries was established based on this. These entries indicate the classification criteria for each legal case; for each classification entry t i Propose a classification result function. for: Where, r i For the classification result, indicate the case d in the classification entry t. i The classification results were then integrated into a single classification index, cr. The parentheses represent the floor function, which rounds the relevant case sets based on the clustering criterion cr. Divided into classification result set Right now 4-2 Output a set of relevant legal cases. from A case is randomly selected from the data. As representative cases, the relevant legal case collection for: Step 5) Preliminary screening of legal provisions based on cross-source information fusion, as follows: 5-1 Construction of a Legal Provisions Database First, data from multiple legal provisions. Extract legal text data and organize it into a database. Each data source The set of clauses in the text is represented as Among them l ij This represents the j-th clause. 5-2 Text Extraction and Information Fusion The proposed extraction method ε(l) ij c), this method indicates from clause l ij Extracting the concept c∈C p Related content: Among them, rel(l ij c) represents the relevance measure between the clause and the concept, and θ2 represents the relevance threshold, set to 0.

3. This parameter varies depending on the actual situation and can be obtained using the grid search method. Based on this, the cross-source information fusion method is used to merge the relevant clauses from various data sources into a unified clause set. Step 6) Accurate extraction of legal provisions based on semantically enhanced retrieval, as detailed below: 6-1 Semantic Enhancement Retrieval and Entry Extraction In the integrated set of articles Then, semantic enhancement retrieval methods were further used. To extract subsets of related concepts The goal is to extract the most relevant legal provisions from the merged set of provisions based on their degree of matching with the concepts. Where sim(l,c) represents the similarity measure between clause l and legal concept c, and θ3 is the similarity threshold, set to 0.

5. This parameter varies depending on the actual situation and can be obtained using a grid search method. 6-2 Output of the final set of clauses The legal provisions retrieved through semantic enhancement are represented as follows: This collection of articles will serve as the final output of the legal article extraction, ensuring that it covers the legal articles relevant to the user's query. Step 7) Construct a knowledge mapping based on multi-level partitioning, as follows: 7-1 Construction of Knowledge Maps First, utilize legal case collections Extracted relevant legal concept set and from the collection of legal provisions Extracted related concept set Constructing a preliminary knowledge mapping in, Represents the set of all related concepts. 7-2 Hierarchical Division Concept Extraction For legal concept set The concept c is mapped to the hierarchy h using the hierarchical partitioning method H(c). c Above, a hierarchical partitioning tree T is constructed. For a specific concept c, it and its direct child concepts Child(c) are added to the knowledge mapping to form the final knowledge mapping. Step 8) Dynamic concept matching and supplementation based on similarity analysis, as follows: 8-1 Similarity Analysis and Conceptual Association For the constructed hierarchical knowledge mapping, the similarity between two legal concepts is evaluated, with a similarity threshold of 0.

6. This parameter varies depending on the specific circumstances and can be obtained using a grid search method, thereby yielding a set of associated concepts. Detailed explanation and output of the 8-2 concept. Finally, a detailed explanation of each concept was retrieved from a legal concept database. Its detailed explanation is d c This involves compiling detailed explanations of all concepts to form the final output set of related concepts. Step 9) The feedback mechanism based on the dynamically generated prompts is as follows: Assume the user queries Q and the dynamic content generated during the execution of the steps: Step 6) Collection of relevant legal provisions Step 8) Collection of relevant legal concepts and step 4) a collection of relevant legal cases The initial output O0 is O0 = LLM(Q). The feedback A from the model combined with dynamically generated content is: Here, concat represents the string concatenation operation; Step 10) The output adjustment strategy based on reflective reasoning is as follows: The user query is Q, the output of step 10) is A, and the organizational reflection prompt is P. r : The function `concat("Please check if the query ", Q," and answer ", A'," whether the subject matches, please answer yes or no")` performs string concatenation. The conclusion is that O satisfies the condition. O=LLM(P r ) Here, LLM(·) represents content generation using a large language model. If the reflection result is "yes", then the reflection conclusion O at this time is taken as the final output; if the reflection conclusion is "no", then the reflection reasoning prompt P is organized. i : The function `concat("Query", Q,"and Answer", A'), Subject mismatch, please modify the inconsistent part of the answer")` is then updated to `A'`. A′=LLM(P i ) And repeat the reflection prompts and their follow-up content.

2. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, it implements the legal concept interpretation method based on multi-perspective clustering and hybrid retrieval as described in claim 1.

Citation Information

Patent Citations

  • Case query method and device, computer equipment and storage medium

    CN110209828A

  • An incomplete multi-view data clustering method and electronic equipment

    CN113705603A