Knowledge base enhancement-based large model class case recommendation method and system

By introducing large-scale model technology based on knowledge base enhancement into the case recommendation system, combined with the BERT and T5 models, the problem that the existing system cannot effectively utilize legal logic is solved, and more accurate and efficient case recommendations are achieved, which improves the efficiency and effectiveness of dispute mediation.

CN120123581APending Publication Date: 2025-06-10GUANGDONG UNIV OF EDUCATION +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510175023.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing case recommendation system cannot effectively utilize the complex semantics and legal logic behind the case, resulting in the inability to provide mediators with truly useful case recommendation services.

Method used

The big model case recommendation method based on knowledge base enhancement is adopted, and a comprehensive professional legal knowledge base and vector database is constructed, combined with the technology of BERT and T5 models, in-depth analysis and optimization sorting are carried out to provide more accurate case recommendations.

Benefits of technology

It has achieved more accurate and efficient case recommendations, which can have a deeper understanding of the content of the dispute, improve the relevance and practicality of the recommendations, and improve the efficiency and effectiveness of mediation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123581A_ABST
    Figure CN120123581A_ABST
Patent Text Reader

Abstract

The invention relates to a knowledge base enhancement-based large model class case recommendation method and system. The method comprises the following steps of: S1, collecting dispute mediation case data and preprocessing the dispute mediation case data; s2, constructing a knowledge base, and updating the knowledge base regularly; s3, converting the sorted dispute case data into high-dimensional vector representation by using an advanced vector embedding technology, and importing the high-dimensional vector representation into a vector database; information input by a user is converted into vectors in the same format as case vectors in the vector database, and similarity matching is carried out; s4, providing a mediation suggestion for the user according to the mediation result and the mediation process of the historical case; s5, performing deep analysis and optimization sorting by using a class case retrieval and large model optimization sorting module; and S6, the mediator performs dispute mediation according to the finally recommended class case information and the mediation suggestions. According to the invention, a more accurate and efficient class case recommendation service is provided for mediation personnel, and the mediation personnel are assisted to better deal with disputes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of recommendation systems and judicial mediation, and particularly to a method and system for recommending similar cases based on a knowledge base-enhanced large model. Background Art

[0002] In existing similar case recommendation systems, methods based on rules, clustering, or classification are mainly used to perform similarity matching on cases. However, these methods often only recommend based on semantic features, behavioral features, interaction features, etc. related to the cases, ignoring the complex semantics and legal logic behind the cases. Therefore, these methods cannot provide truly useful similar case recommendation services for mediators.

[0003] Therefore, it is urgent to introduce more advanced technical means to solve the current deficiencies in dispute mediation, provide more accurate and efficient similar case recommendation services for mediators, and help mediators better handle disputes. Summary of the Invention

[0004] To solve the technical problems existing in the prior art, the present invention provides a method and system for recommending similar cases based on a knowledge base-enhanced large model. By combining a legal domain knowledge base and large model technology, it can more accurately understand the semantics and logic of legal texts, thereby providing more precise similar case recommendations.

[0005] The method of the present invention is implemented by the following technical solutions: A method for recommending similar cases based on a knowledge base-enhanced large model, including the following steps:

[0006] S1. Collect and organize dispute mediation case data, removing noise and invalid information, including party information, case reasons, mediation processes, and mediation results;

[0007] S2. Through a knowledge base construction module, collect legal terms, classic cases, legal regulations, and judicial interpretation content to build a comprehensive professional legal knowledge base, and regularly update the knowledge base to ensure that its content is synchronized with the latest laws, regulations, and judicial practices;

[0008] S3. Use advanced vector embedding technology to convert the organized dispute case data into high-dimensional vector representations and import them into a vector database; when receiving user input of dispute information, including party information, dispute type, and dispute facts, perform vectorization processing on it, convert the user input information into a vector in the same format as the case vectors in the vector database, and perform similarity matching;

[0009] S4. Relying on the efficient similarity retrieval ability of the vector database, find a set of historical cases similar to this dispute, and provide mediation suggestions for the user based on the mediation results and processes of the historical cases.

[0010] S5. Use the similar case retrieval and large model optimization ranking module for in-depth analysis and optimization ranking, including the semantic retrieval module, generative candidate generation, and fine ranking module;

[0011] S6. The mediator conducts dispute mediation based on the finally recommended similar case information and mediation suggestions.

[0012] The system of the present invention adopts the following technical solutions to implement: a large model similar case recommendation system based on knowledge base enhancement, including:

[0013] Data preprocessing module: By collecting and organizing dispute mediation case data, removing noise and invalid information, including party information, case reasons, mediation process, and mediation results;

[0014] Knowledge base construction module: Used to construct a knowledge base, collect legal terms, classic cases, regulations, and judicial interpretation content, construct a comprehensive professional legal knowledge base, and update the knowledge base regularly to ensure that its content is synchronized with the latest laws, regulations, and judicial practices;

[0015] Text vectorization module: By using advanced vector embedding technology, convert the sorted dispute case data into high-dimensional vector representations and import them into the vector database; when receiving user input dispute information, including party information, dispute type, and dispute facts, perform vectorization processing on it, convert the user input information into vectors in the same format as the case vectors in the vector database, and perform similarity matching;

[0016] Mediation suggestion module: By leveraging the efficient similarity retrieval ability of the vector database, find a set of historical cases similar to the dispute, and provide mediation suggestions for users based on the mediation results and processes of historical cases;

[0017] Similar case retrieval and large model optimization ranking module: Used for in-depth analysis and optimization ranking, including the semantic retrieval module, generative candidate generation, and fine ranking module;

[0018] Dispute mediation module: The mediator conducts dispute mediation based on the finally recommended similar case information and mediation suggestions.

[0019] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0020] 1. The present invention combines the multi-hop retrieval of BERT, the generative candidates of the T5 model, and the instruction optimization of the large model, and can understand the dispute content more deeply, so as to provide more accurate similar case recommendations.

[0021] 2. The present invention utilizes advanced vector embedding technology to enable the system to efficiently convert dispute case data into high-dimensional vector representations, facilitating rapid retrieval and processing. Through the efficient similarity retrieval ability of the vector database, the system can quickly identify historical cases similar to the user's dispute, improving the relevance and practicality of the recommendations.

[0022] 3. Finally, in the present invention, the mediator conducts dispute mediation based on the similar case information and mediation suggestions provided by the system. Due to the high degree of automation of the system, the mediator can handle disputes more efficiently, enhancing the efficiency and effectiveness of mediation. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 is a flowchart of the method of the present invention;

[0024] Figure 2 is a schematic diagram of the knowledge base construction of this embodiment;

[0025] Figure 3 is a schematic diagram of the structure of the system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] The present invention will be further described in detail below in conjunction with the embodiments and the drawings, but the embodiments of the present invention are not limited thereto.

[0027] Embodiment

[0028] As Figure 1 shown, the method for recommending similar cases based on a large model enhanced by a knowledge base in this embodiment includes the following steps:

[0029] S1. Collect and organize dispute mediation case data, removing noise and invalid information, including party information, case reasons, mediation processes, mediation results, etc.;

[0030] S2. As Figure 2 shown, through the knowledge base construction module, widely collect contents such as legal terms, classic cases, regulations, and judicial interpretations, construct a comprehensive, authoritative, and continuously updated professional legal knowledge base, and regularly update the knowledge base to ensure that its content is synchronized with the latest laws, regulations, and judicial practices;

[0031] S3. Apply advanced vector embedding technology to convert the sorted dispute case data into high-dimensional vector representations and import them into the vector database; when receiving user input dispute information, including party information, dispute type, and dispute facts, perform vectorization processing on it, convert the user input information into a vector in the same format as the case vectors in the vector database, and perform similarity matching;

[0032] S4. Leverage the efficient similarity retrieval ability of the vector database to find a set of historical cases similar to the dispute, and provide mediation suggestions for the user based on the mediation results and processes of the historical cases;

[0033] S5. Use case retrieval and large model optimization ranking module for in-depth analysis and optimization ranking, including semantic retrieval module (multi-hop retrieval based on BERT), generative candidate generation (based on T5 model), and fine ranking module (large model optimized based on instructions);

[0034] S6. The mediator conducts dispute mediation based on the finally recommended similar case information and mediation suggestions.

[0035] Specifically, in this embodiment, the specific process of step S1 includes: widely collecting various dispute mediation case data from channels such as court judgment databases, legal information platforms, and law firm case libraries, covering different types of disputes and detailed content, and performing deduplication, denoising, etc. to organize them to meet the system requirements.

[0036] Specifically, in this embodiment, the specific process of step S3 includes:

[0037] S31. Use advanced vector embedding technology to convert the sorted dispute case data into high-dimensional vector representations;

[0038] S32. Select a suitable vector embedding model to ensure that the vectors accurately reflect the semantic and legal features of the text;

[0039] S33. Import the vectorized case data into the vector database, establish an efficient indexing mechanism, and quickly retrieve similar cases;

[0040] S34. Continuously optimize the storage and retrieval algorithms of the vector database to improve the response speed and accuracy of the system.

[0041] Specifically, in this embodiment, the specific process of step S4 includes:

[0042] S41. Leverage the vector database to retrieve a set of historical cases similar to the dispute input by the user, and set a similarity threshold to ensure relevance;

[0043] S42. Use the cosine distance between vectors as a distance metric to judge similarity. Let the dispute vector input by the user be The historical case vector be If the cosine distance:

[0044]

[0045] where ε is the similarity threshold, then it is considered that the historical case is similar to the user's dispute and is included in the candidate set.

[0046] Specifically, in this embodiment, the specific process of step S5 includes:

[0047] The large model optimization and ranking module uses the Prompt instruction technology to guide the large model to deeply analyze and optimize the ranking of the candidate case set according to factors such as case similarity and legal complexity; the specific processes of each module it contains are as follows:

[0048] Semantic retrieval module:

[0049] The input query information Q entered by the user is embedded through the BERT model to obtain the semantic representation of the query Embedding(Q) = BERT(Q);

[0050] The data set {D 1 , D 2 , D 3 ,..., D N} in the knowledge base is embedded through the BERT model to obtain the semantic representation of each data Embedding(D I ) = BERT(D I );

[0051] Retrieval is performed according to the cosine similarity calculation formula of semantic similarity. Let the query vector be and the knowledge base data vector be then calculate the cosine similarity:

[0052]

[0053] Select the top most relevant documents D s1 , D s2 ,..., D sk , and quickly screen out a preliminary candidate set that is relatively relevant to the user's query from the knowledge base, providing a basis for subsequent generative candidate generation and fine ranking;

[0054] Generative candidate generation:

[0055] After obtaining the candidate set {D s1 , D s2 ,..., D sk} based on BERT multi-hop retrieval, use the T5 generation model for supplementary expansion; use the candidate set as the context input, combined with the user's query, to generate more reference candidate sets

[0056] C = T5(Q, {D s1 , D s2 ,..., D sk});

[0057] Fine ranking module:

[0058] Use a large model to finely sort the generated candidate set; define the candidate set sorting score S(D_sj), calculated based on the instruction optimization model: S(D_sj) = InstructionModel(Q, D_sj); select the top n candidates with the highest scores as the output result R = Top-n(S(D_sj)), j = 1, 2,..., k;

[0059] Relevance evaluation model:

[0060] Define the relevance scoring function: R = f(S, K, L), where R represents the relevance score, S represents the case similarity, K represents the keyword matching degree, and L represents the legal logic matching degree;

[0061] Calculation of case similarity: Use methods such as cosine similarity in the vector space model to calculate the similarity between the user input dispute vector and the historical case vector; assume the user input dispute vector is the historical case vector is then the similarity

[0062]

[0063] Calculation of keyword matching degree: By extracting the keyword set K in the user input dispute u and the keyword set K in the historical case c , calculate the matching degree of keywords:

[0064]

[0065] Calculation of legal logic matching degree: Based on legal rules and logical reasoning models, judge the consistency of the user's dispute and the historical case in legal logic; by constructing a legal logic graph, compare the node and path matching situations of the user's dispute and the historical case in the graph, and give the corresponding score L;

[0066] The relevance scoring function R = f(S, K, L) obtains the relevance score of each candidate case by synthesizing the above calculation results.

[0067] Specifically, in this embodiment, the process of generating and outputting the mediation suggestions in step S6 is as follows: According to the mediation results and mediation process of historical cases, combined with the analysis results of the large model, the mediation suggestion module provides targeted and practical mediation suggestions for users; the mediation suggestions include content such as legal basis, solution suggestions, negotiation strategies, etc., and output the similar case information and mediation suggestions to the user in a clear and easy-to-understand manner for the user to understand and reference.

[0068] Specifically, in this embodiment, by designing an AB experiment mechanism, the mediation success rate of different random experiments is automatically evaluated, and a better combination of vectorization parameters and large model tuning instructions is selected. This process can be continuously iterated to obtain the optimal similar case recommendation method based on the vector database and large model.

[0069] Specifically, when designing the AB experimental mechanism, this embodiment will conduct refined grouping based on multiple factors such as dispute type, complexity, and region; for example, contract disputes will be divided into groups of different complexity levels according to the amount, and labor disputes in different regions will be grouped by region for experiments, so as to more accurately evaluate the mediation success rate.

[0070] Specifically, this embodiment dynamically adjusts vectorization parameters, large model tuning instruction combinations, etc. according to experimental feedback and data analysis. For example, if it is found that a certain type of dispute case does not perform well under specific parameters, it will automatically adjust the parameters and adjust the large model instructions based on user feedback.

[0071] Specifically, this embodiment uses multi-dimensional indicators such as mediation success rate, recommendation accuracy (quantified by precision, recall, etc.), timeliness (recording processing time), and user satisfaction (collected by questionnaires, etc.) to comprehensively evaluate the performance of the recommendation method.

[0072] Based on the same inventive concept, Figure 3 As shown, the present invention also provides a large model similar case recommendation system based on knowledge base enhancement in this embodiment, including:

[0073] Data preprocessing module: collects and organizes dispute mediation case data to remove noise and invalid information, including party information, case causes, mediation process, mediation results, etc.;

[0074] Knowledge base construction module: used to build a knowledge base, widely collect legal terms, classic cases, regulations, judicial interpretations and other content, build a comprehensive, authoritative and constantly updated professional legal knowledge base, and regularly update the knowledge base to ensure that its content keeps pace with the latest laws, regulations and judicial practices;

[0075] Text vectorization module: By using advanced vector embedding technology, the sorted dispute case data is converted into high-dimensional vector representation and imported into the vector database; when receiving the dispute information input by the user, including the party information, dispute type, and dispute facts, it is vectorized and the user-entered information is converted into a vector in the same format as the case vector in the vector database for similarity matching;

[0076] Mediation suggestion module: by using the efficient similarity search capability of the vector database, it finds a collection of historical cases similar to the dispute, and provides users with mediation suggestions based on the mediation results and mediation process of the historical cases;

[0077] Case Retrieval and Large Model Optimization Sorting Module: Used for in-depth analysis and optimization sorting, including a semantic retrieval module (multi-hop retrieval based on BERT), a generative candidate generation (based on the T5 model), and a fine sorting module (large model optimized based on instructions);

[0078] Dispute Mediation Module: The mediator conducts dispute mediation based on the finally recommended case information and mediation suggestions.

[0079] The method of the present invention is deployed on a server and provides services externally through an API interface or a web page. The system is regularly maintained and upgraded to repair potential problems and vulnerabilities to ensure the normal operation of the system.

[0080] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A large model similar case recommendation method based on knowledge base enhancement, characterized in that: The following steps are involved: S1. Collect and organize dispute mediation case data, remove noise and invalid information, including party information, case causes, mediation process, and mediation results; S2. Through the knowledge base construction module, collect legal terms, classic cases, regulations, and judicial interpretations to build a comprehensive professional legal knowledge base, and regularly update the knowledge base to ensure that its content keeps pace with the latest laws, regulations, and judicial practices; S3. Use advanced vector embedding technology to convert the sorted dispute case data into high-dimensional vector representation and import it into the vector database; when receiving the dispute information input by the user, including the party information, dispute type, and dispute facts, vectorize it and convert the user input information into a vector with the same format as the case vector in the vector database for similarity matching; S4. With the help of the efficient similarity search capability of the vector database, a collection of historical cases similar to the dispute is found, and mediation suggestions are provided to users based on the mediation results and mediation process of the historical cases; S5. Use similar case retrieval and large model optimization sorting modules for in-depth analysis and optimization sorting, including semantic retrieval module, generative candidate generation, and refined sorting module; S6. The mediator conducts dispute mediation based on the final recommended similar case information and mediation suggestions.

2. The large model similar case recommendation method based on knowledge base enhancement according to claim 1 is characterized in that: The specific process of step S1 includes: collecting various dispute mediation case data from court judgment databases, legal information platforms, and law firm case libraries, covering different types of disputes and detailed content, and performing deduplication and denoising to make them meet system requirements.

3. The large model similar case recommendation method based on knowledge base enhancement according to claim 1 is characterized in that: The specific process of step S3 includes: S31. Use advanced vector embedding technology to transform the sorted dispute case data into high-dimensional vector representation; S32. Select an appropriate vector embedding model to make the vector accurately reflect the semantic and legal characteristics of the text; S33, importing the vectorized case data into the vector database, establishing an efficient indexing mechanism, and quickly retrieving similar cases; S34. Continuously optimize the storage and retrieval algorithms of the vector database to improve the response speed and accuracy of the system.

4. The large model similar case recommendation method based on knowledge base enhancement according to claim 1 is characterized in that: The specific process of step S4 includes: S41, using the vector database to retrieve a collection of historical cases similar to the dispute input by the user, and setting a similarity threshold to ensure relevance; S42, using the distance metric cosine distance between vectors to determine similarity, assuming that the user input dispute vector is The historical case vector is If the cosine distance: Among them, ε is the similarity threshold. If this historical case is considered similar to the user dispute, it will be included in the candidate set.

5. The large model similar case recommendation method based on knowledge base enhancement according to claim 1 is characterized in that: The specific process of step S5 includes: The big model optimization sorting module uses prompt instruction technology to guide the big model to deeply analyze and optimize the candidate case set according to case similarity and legal complexity factors; the specific processes of each module are as follows: Semantic retrieval module: Input the dispute information Q entered by the query user and embed it through the BERT model to obtain the semantic representation of the query: Embedding(Q) = BERT(Q); For the datasets {D1, D2, D3, ..., D N }Embedding through the BERT model to obtain the semantic representation of each data I )=BERT(DI); According to the semantic similarity calculation formula cosine similarity, the query vector is The knowledge base data vector is Then calculate the cosine similarity: Select the most relevant previous document D s1 , D s2 , ..., D sk , quickly filter out a preliminary candidate set that is more relevant to the user query from the knowledge base, providing a basis for subsequent generative candidate generation and refined ranking; Generative candidate generation: Get the candidate set based on BERT multi-hop retrieval {D s1 , D s2 , ..., D sk }, use the T5 generative model for supplementary expansion; use the candidate set as context input, combined with user queries, to generate more reference candidate sets C=T5(Q,{D s1 ,D s2 ,...,D sk }); Fine sorting module: Use the large model to finely sort the generated candidate set; define the candidate set sorting score S(D_sj), calculated based on the instruction optimization model: S(D_sj) = InstructionModel(Q, D_sj); select the n candidates with the highest scores as the output result R = Top-n(S(D_sj)), j = 1, 2, ..., k; Relevance evaluation model: Define the relevance scoring function: R = f(S, K, L), where R represents the relevance score, S represents the case similarity, K represents the keyword matching degree, and L represents the legal logic matching degree; Case similarity calculation: The cosine similarity method in the vector space model is used to calculate the similarity between the user input dispute vector and the historical case vector; assuming that the user input dispute vector is The historical case vector is The similarity Keyword matching calculation: By extracting the keyword set K from the user input dispute u and the keyword set K in the historical cases c , calculate the matching degree of keywords: Legal logic matching degree calculation: Based on legal rules and logical reasoning models, judge the consistency of user disputes and historical cases in legal logic; by constructing a legal logic graph, compare the node and path matching between user disputes and historical cases in the graph, and give a corresponding score L; The relevance scoring function R=f(S, K, L) is used to obtain the relevance score of each candidate case by combining the above calculation results.

6. The large model similar case recommendation method based on knowledge base enhancement according to claim 1 is characterized in that: The process of generating and outputting mediation suggestions in step S6 is as follows: based on the mediation results and mediation process of historical cases and combined with the analysis results of the large model, the mediation suggestion module provides users with targeted and practical mediation suggestions; the mediation suggestions include legal basis, solution suggestions, and negotiation strategy content, and similar case information and mediation suggestions are output to users for their understanding and reference.

7. A large-model similar case recommendation system based on knowledge base enhancement, characterized by: include: Data preprocessing module: collects and organizes dispute mediation case data to remove noise and invalid information, including party information, case causes, mediation process, and mediation results; Knowledge base construction module: used to build a knowledge base, collect legal terms, classic cases, regulations, and judicial interpretations, build a comprehensive professional legal knowledge base, and regularly update the knowledge base to ensure that its content keeps pace with the latest laws, regulations, and judicial practices; Text vectorization module: By using advanced vector embedding technology, the sorted dispute case data is converted into high-dimensional vector representation and imported into the vector database; when receiving the dispute information input by the user, including the party information, dispute type, and dispute facts, it is vectorized and the user-entered information is converted into a vector in the same format as the case vector in the vector database for similarity matching; Mediation suggestion module: by using the efficient similarity search capability of the vector database, it finds a collection of historical cases similar to the dispute, and provides users with mediation suggestions based on the mediation results and mediation process of the historical cases; Similar case retrieval and large model optimization sorting module: used for in-depth analysis and optimization sorting, including semantic retrieval module, generative candidate generation, and fine sorting module; Dispute mediation module: The mediator conducts dispute mediation based on the final recommended similar case information and mediation suggestions.

Citation Information

Cited By

  • Network appeal text keyword extraction and embedding method and application

    CN121597803A