Similar case retrieval method based on large model fusion multi-level legal element generation
By building a multi-level legal element knowledge base and adopting large-scale model technology, combining multiple recall strategies and fine-scheduling models, the problem that traditional search methods are difficult to accurately find similar cases is solved, efficient and accurate search of similar cases is achieved, and judicial efficiency and fairness are improved.
Patent Information
- Application Number
- CN202411891827.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-13
AI Technical Summary
Traditional manual search methods are difficult to accurately find cases similar to current cases in massive cases, resulting in the impact of judicial efficiency and fairness.
A similar case search method is adopted based on the large model fusion of multi-level legal elements, and efficient and accurate similar case search is achieved by building a multi-level legal element knowledge base, training a multi-level case classification model and problem classification model, adopting a multi-channel recall strategy, and sorting based on fine-scheduling model.
It improves the accuracy and efficiency of similar case searches, can cover potential similar cases more comprehensively, and enhances the supportiveness of judicial decision-making.
Smart Images

Figure CN119989074A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of legal artificial intelligence technology, and in particular to a similar case retrieval method based on a large model integrating multi-level legal elements. Background Art
[0002] Against the backdrop of rapid economic and social development, the number of cases has shown a trend of rapid growth, which has undoubtedly brought unprecedented pressure to the judicial system. Faced with a large number of cases and complex document content, traditional manual search methods seem to be unable to cope with the situation. Even experienced legal practitioners find it difficult to accurately find cases similar to the current case from a large number of cases in a short period of time. This not only affects judicial efficiency, but also poses certain challenges to judicial fairness.
[0003] In order to meet this challenge, similar case retrieval technology came into being and quickly became a research hotspot in the judicial field. As an important part of smart justice, similar case retrieval is based on quickly finding the most similar case in the database by comparing the similarity between two cases, thereby providing strong support for judicial decision-making. The emergence of this technology can help legal practitioners find relevant legal basis and similar cases more quickly and accurately, and improve the quality and efficiency of trials.
[0004] At present, there are mainly similar case retrieval strategies based on keywords, similar case retrieval strategies based on judgment results, similar case retrieval strategies based on cited legal provisions, etc., but they all have certain limitations. Now, due to the emergence of big models, new ideas are brought to similar case retrieval. Therefore, the present invention designs a similar case retrieval method based on big models and multi-level legal elements. Summary of the invention
[0005] In view of the technical problems existing in the prior art, the purpose of the present invention is to provide a similar case retrieval method based on a large model integrating multi-level legal elements. The method achieves efficient and accurate similar case retrieval by building a multi-level legal element knowledge base, understanding the intention of legal elements at different levels based on a large model, training a multi-level case classification model and a problem classification model, adopting a multi-way recall strategy, and sorting based on a refined ranking model. The specific process steps are as follows: Figure 1 shown.
[0006] A similar case retrieval method based on a large model integrating multi-level legal elements includes the following steps:
[0007] Step (1) Build a multi-level legal element knowledge base based on the characteristics of the document structure and combine large-model fine-tuning training and iterative optimization of the large-model prompt instruction method.
[0008] (1.1) Based on the characteristics of the document structure and the search habits of users, we have carefully constructed a multi-level and multi-type legal element knowledge system. This system covers legal elements at the original text level, legal elements at the case feature label level, legal elements at the dispute focus level generated through the combination of litigation and defense, and legal elements at the judgment essence level extracted from the full text. On this basis, we further constructed corresponding element knowledge bases based on different types of legal elements.
[0009] (1.2) Constructing the knowledge base of original legal elements: Obtain the original data of the judicial documents database, divide the documents into sections and structure them, and set up key sections such as "claim section", "defense section", "this court's findings section", "this court's opinion section", "judgment result section", "full text section", etc., to form a basic knowledge base of original legal elements, and enter them into the corresponding database, such as ElasticSearc, etc.
[0010] (1.3) Constructing a knowledge base of legal elements of case characteristics: Since the "claim paragraph" and "this court has found" are often used in case descriptions, which contain descriptions of the actor's behavior and purpose, etc., and cover more "case characteristic elements" information, based on the above two paragraphs and combined with the case characteristic legal element system sorted out by legal personnel in each case, the knowledge base of case characteristic legal elements is constructed by using the large model fine-tuning training method.
[0011] (1.4) Constructing a knowledge base of legal elements of dispute focus: Based on the structural characteristics of the document, the "claim paragraph" and "defense paragraph" are the self-statement stages of the parties, which can reflect the implicit points of the contradiction between the two parties. Then the focus of the dispute can be extracted. In this regard, the above-mentioned "claim paragraph" and "defense paragraph" are used as input texts, based on the legal element system of dispute focus sorted out by legal personnel, using the powerful summary generation ability of the big model, continuously iteratively testing and optimizing the Prompt command, and finally completing the database running work according to the optimized instructions to form a knowledge base of legal elements of dispute focus.
[0012] (1.5) Constructing a knowledge base of legal elements of adjudication essence: Based on the structure of the People's Court case database, we know that a document can be refined into a "judgment essence" that reflects the most important core elements of the document. Based on this, we combined the five paragraphs of "claim paragraph", "defense paragraph", "this court's findings", "this court's opinion", and "judgment result", and completed the construction of the knowledge base of legal elements of adjudication essence by using large model fine-tuning training.
[0013] (1.6) Construct a vector version of the legal element knowledge base for all the above elements: Use the BGE M3-Emebdding (BAAI General Embedding) vector model to vectorize the contents of all the above legal element knowledge bases to form a vector version of the legal element knowledge base. Since the combined search method of sparse vector search and dense vector search is more accurate, two types of vectors, sparse vectors and dense vectors, are set here to store. Since dense vectors occupy a large amount of memory when vectorizing long texts, we use PCA (Principal Component Analysis) vector dimensionality reduction technology here to successfully reduce its vector dimension and reduce memory usage.
[0014] Step (2): Based on the UTC (Universal Text Classification) classification model, the case classification model is trained. The case classification model is used to predict the case category of the input question and associate the name of the upper-level case to provide a basis for the subsequent search. Based on the UTC (Universal Text Classification) classification model, the question classification model is trained to classify the input questions. Different types of questions have different focuses on understanding the intentions of different types of questions. Different legal element knowledge bases can also be searched based on different focuses.
[0015] Step (3): Based on different types of questions, the big model is used to understand the legal elements of the input questions at different levels: the types of questions mentioned above are mainly divided into three categories: "questions of ordinary users", "questions of professional legal professionals", and "questions of original documents or long text cases". Based on the characteristics of the above three types of questions, the big model Prompt instruction is used to fine-tune the extraction and summary of legal elements at different levels. In this way, the core elements can be found closer to the search intent, and then the core features can be found.
[0016] (3.1) If the input question is a "normal user type question", a purely literal keyword legal element information will be generated based on the input. It can exclude some unimportant information such as stop words, and then a "case characteristic" legal element will be summarized based on the case characteristic element system.
[0017] (3.2) If the input question is a "professional legal professional type of question", the big model will first be used to rewrite the question into multiple forms of expression based on the prompt instruction, and then the pure literal keyword elements and "case characteristics" legal elements will be extracted based on the logic of (3.1).
[0018] (3.3) If the input question is a "question of the original document or long text case type", a "judgment essence" legal element will be generated first, then the "dispute focus" legal element will be summarized, and finally a "case characteristics" legal element will be generated.
[0019] Step (4): Construct a multi-way recall strategy route: The multi-way recall strategy mainly includes "slop retrieval route based on phrase matching", "retrieval route based on BM25 algorithm", and "hybrid retrieval route based on vector library". Slop retrieval based on phrase matching mainly focuses on literal meaning search, which allows skip word search. The retrieval route based on BM25 algorithm mainly focuses on the relevance and importance of the question and document content. The hybrid retrieval route based on vector library mainly focuses on converting high-dimensional vector representation to measure the semantic similarity between the question and the document.
[0020] Step (5): Recall search based on the input question and different levels of element knowledge: Based on the input question type and the different levels of legal elements extracted by the large model in step (3), according to the recall route constructed in step (4), enter the corresponding "original text legal element knowledge base", "case feature legal element knowledge base", "dispute focus legal element knowledge base", "judgment summary legal element knowledge base", and vector version of the legal element knowledge base to search and recall.
[0021] Step (6): According to the weight ratio of different recall routes, combined with the number of return values of "slop retrieval route based on phrase matching", "retrieval route based on BM25 algorithm", and "hybrid retrieval route based on vector library", the recall subset samples of each route are dynamically selected. If the number of return values of slop retrieval route based on phrase matching is greater than topk, the final recall set is obtained. If the return value is less than topk, the "retrieval route based on BM25 algorithm" and the "hybrid retrieval route based on vector library" are combined to obtain the final recall set.
[0022] Step (7): If the number of recall sets returned by the "slop search route based on phrase matching" is enough for the topk, directly return the similar case set; if the number returned is not enough for the topk, then the recall set formed by combining the "search route based on the BM25 algorithm" and the "hybrid search route based on the vector library" is accurately sorted to obtain the final similar cases.
[0023] The present invention also provides a server, characterized in that it includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing the above method.
[0024] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the above method when executed by a processor.
[0025] Compared with the prior art, the present invention has the following beneficial effects:
[0026] The present invention provides a similar case retrieval method based on the generation of large-scale model-integrated multi-level legal elements. By constructing a multi-level legal element knowledge base and adopting a large-scale model technology method, the expansion of multi-level case elements is completed. The search scope is limited by distinguishing multiple types of cases and multiple types of questions, and a multi-way recall strategy is adopted to retrieve and match legal elements from different angles and dimensions, thereby more comprehensively covering potential similar cases. Finally, the completeness and reliability of the retrieval results are improved by combining the precision sorting technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 A flowchart of the specific implementation steps.
[0028] Figure 2 This is the UTC model structure diagram.
[0029] Figure 3 This is the BGE M3 model structure diagram. DETAILED DESCRIPTION
[0030] To further illustrate the technical solution of the present invention, the above steps are described in detail below through drawings and specific examples, but the embodiments are not intended to limit the present invention.
[0031] The process of the present invention is as follows Figure 1 As shown, the steps include:
[0032] Step (1): Based on the characteristics of the document structure and combined with the large model fine-tuning training and iterative optimization of the large model prompt instruction method, a multi-level legal element knowledge base is constructed.
[0033] 1) Construct the knowledge base of original legal elements: Specifically, obtain a case database of first-instance judgments with a volume of tens of millions {X1, X2, X3...X n}, where case types include but are not limited to five types of cases: civil, criminal, administrative, execution, and compensation. By writing regular expressions, the paragraph division is completed, and key paragraphs such as "claim paragraph", "defense paragraph", "this court's findings paragraph", "this court's opinion paragraph", "judgment result paragraph", and "full text paragraph" are obtained. A standard knowledge base of the original text of the judgment document is formed.
[0034] 2) Construct a knowledge base of case characteristics and legal elements: Combine the claim paragraph and the court's findings paragraph in each judgment case, annotate them using the case characteristics and legal elements system, and then input them into the big model for fine-tuning, extract the case characteristics of each judgment case, and construct a knowledge base of case characteristics and legal elements; specifically, combine the "claim paragraph" {SC1, SC2, SC3...SC n} and "This Court has found that" {BYCM1,BYCM2,BYCM3...BYCM n} The paragraph texts are combined, and based on the manually sorted case feature element system of each case cause, the large model is fine-tuned and trained using the training set generated by regular self-annotation. Then the elements of the above combined paragraphs are refined to form a case feature legal element knowledge base.
[0035] The following strategies are mainly used to generate training sets using regular self-annotation: select some simple element labels based on the manually sorted case feature element system to generate a part of regular expressions, run a large number of document labels based on regular expressions, and then train and fine-tune the large model based on this batch of data to obtain the fine-tuned case feature element extraction large model. The following is an example of generating regular expressions based on simple element labels;
[0036]
[0037] 3) Constructing a knowledge base of dispute focus legal elements: Combine the claim paragraph and the defense paragraph in each judgment case and use the dispute focus legal element system to continuously optimize and iterate to generate a large model prompt instruction. Based on this, extract the dispute focus of each judgment case and construct a knowledge base of dispute focus legal elements; specifically, combine the "claim paragraph" {SC1, SC2, SC3...SC n} and "defense paragraph" {BC1,BC2,BC3...BC n} paragraph texts, based on the manually sorted dispute focus element system of each case cause, the large model is optimized by iterative optimization of the prompt instruction, and then the elements of the above combined paragraphs are refined to form a knowledge base of legal elements of dispute focus. The examples of manually sorted case characteristics and legal element system of dispute focus are as follows:
[0038]
[0039] The format of the Prompt instruction after iterative optimization is as follows:
[0040] Role
[0041] -You are an expert in extracting dispute focus, well-versed in legal knowledge, and proficient in the theoretical basis of law.
[0042] background
[0043] -xxxxxxxxxx
[0044] Task
[0045] -xxxxxxxxxx
[0046] Target
[0047] -xxxxxxxxxx
[0048] Example
[0049] -xxxxxxxxxx
[0050] Output
[0051] -xxxxxxxxxx
[0052] 4) Constructing a knowledge base of legal elements of adjudication essence: Based on the structural example of the People's Court case database, it can be known that any document will have a "adjudication essence" that can reflect the most important core elements of this document. The description style of adjudication essence is as follows:
[0053]
[0054] Specifically, the “claim segment” {SC1, SC2, SC3...SC n}, "Defense Section" {BC1, BC2, BC3...BC n}, "This Court believes that paragraph" {BYRW1,BYRW2,BYRW3...BYRW n}, "This court found out" {BYCM1,BYCM2,BYCM3...BYCM n}, "Judgment Result Section" {CPJG1, CPJG2, CPJG3...CPJG n} and other texts, and use the structural characteristics of the People's Court case library data to generate a training set to fine-tune the training model to form a fine-tuning model for the judgment essence. Then, the above combined paragraphs are used as input to generate the judgment essence, forming a knowledge base of legal elements of the judgment essence.
[0055] 5) Construct a vector version of the legal element knowledge base: Specifically, all the above contents related to the legal element knowledge base are vectorized using the BGE M3-Emebdding (BAAI General Embedding) vector model. For the subsequent hybrid retrieval method, two vector storage types are selected, namely sparse vector and dense vector. Since the default dense vector dimension of BGEM3-Emebdding is 1024 dimensions, it will occupy a lot of memory when saving tens of millions of documents and hundreds of millions of vector-level data. Therefore, the dimensionality reduction technology we use here reduces it to 128 dimensions, greatly reducing its memory footprint.
[0056] Here we choose PCA dimensionality reduction method, and its calculation formula is as follows: Assume there is a data set X containing m samples and n features, where X = {X1, X2, X3...X i ...X m}, each sample X i is an n-dimensional vector: the mean vector formula is
[0057]
[0058] The covariance matrix formula is:
[0059]
[0060] Based on the above covariance, the first k eigenvalues are selected as principal components to form a matrix W, and the data is projected into the space of the first K eigenvectors:
[0061]
[0062] Step (2): Based on the UTC model, the case cause hierarchical classification model is trained. First, the case cause hierarchical classification is manually set based on legal standards and specifications. The manually sorted case cause hierarchy is as follows:
[0063]
[0064] Then, based on the Prompt instruction for generating questions and the case name, the query question is generated using the large model. The format of the Prompt instruction for generating query questions is roughly similar to the instruction format in step (a) step 3), except that the roles, tasks, background, and examples are closer to the content of the question generation. The case and query questions are then combined to generate a training set, and finally the UTC model is trained to form a case hierarchical classification model. When querying questions, the model can determine multi-level case names, which can solve both the limitations of searching in a single case and the generalization of unrestricted cases.
[0065] Step (3): Train the classification model of query questions based on the UTC model. Because different types of query questions have different focuses on the core legal elements, we should focus on understanding their intentions based on the focus of the questions, and we can also search different legal element knowledge bases based on different focuses. The training method is the same as step (2). The structure diagram of UTC is as follows: Figure 2 shown.
[0066] The specific query question types are as follows:
[0067] 1) The query question categories are mainly divided into three categories: "general user type questions", "professional legal person type questions", and "original document or long text case type questions". Based on the characteristics of the above three types of questions, the large model Prompt instruction fine-tuning method is used to extract and summarize different types of legal elements.
[0068] 2) If the input question is a "common user type question", such as "I accidentally hit an elderly person walking on the road while driving a two-wheeled motorcycle. What should I do?" First, the big model will be used to extract a set of legal element information of pure literal keywords based on the Prompt instruction of the pure literal keywords. For example, keywords such as "driving a two-wheeled motorcycle", "hitting", "elderly", etc. Then, the big model fine-tuned based on the legal element system of the case characteristics will extract and summarize a set of "case characteristics" legal elements, such as "driving a non-motor vehicle", "the victim is an elderly person" and other legal elements.
[0069] 3) If the input question is a "problem for professional legal professionals", such as "custody of children under two years old", the big model will be used to rewrite the question into multiple forms of expression based on the prompt instruction for rewriting the question, such as "custody of infants and young children", "determination of custody of children under two years old", "disputes over custody of children under two years old", and then the pure literal keyword elements and "case characteristics" legal elements will be extracted based on the logic of step 2).
[0070] 4) If the input question is "original document or long text case type question", firstly, a "judgment essence" legal element will be generated based on the macro model fine-tuned by the judgment essence combined with the input question, and then the "dispute focus" legal element will be summarized by the macro model based on the prompt instruction of the dispute focus. Finally, a "case characteristics" legal element will be extracted and summarized based on the macro model fine-tuned by the legal element system of the case characteristics.
[0071] Step (4): Construct a multi-way recall strategy type: The multi-way recall routes mainly include "slop retrieval route based on phrase matching", "retrieval route based on BM25 algorithm", and "hybrid retrieval route based on vector library".
[0072] 1) Recall route 1 is a slop search route based on phrase matching, where slop search refers to the maximum number of positions that a word in a phrase can move when it matches in a document. Specifically, when using a phrase matching query, if the slop parameter is set, Elasticsearch will allow the words in the phrase to move a certain number of positions in the document during retrieval in an attempt to find a matching document. The number of positions moved is the slop value. This method allows a certain degree of spacing and order changes between query terms, which makes the query results more flexible and inclusive, and the accuracy is relatively high.
[0073] 2) Recall route 2 is a retrieval route based on the BM25 algorithm. This retrieval method comprehensively considers factors such as term frequency (TF), inverse document frequency (IDF), and document length normalization to score documents. It can more accurately measure the relevance between documents and query statements, thereby returning more accurate search results. The BM25 retrieval calculation formula is as follows:
[0074]
[0075] Where: Q represents a given query statement containing keywords q{1},...,q{n}, D is a document, f() is a formula for calculating word frequency, IDF() is an inverse document frequency calculation formula, |D| represents the length of the document, avgdl represents the average field length, and k1 and b are adjustable hyperparameters.
[0076] 3) Recall route 3 is a hybrid retrieval route based on a vector library. This retrieval method combines the advantages of dense vectors and sparse vectors. Dense vectors are usually generated by neural network models and can capture the semantic similarity of texts. Sparse vectors can be generated by traditional literal text retrieval algorithms and are more suitable for word frequency statistics and keyword matching. By mixing these two vector representations, the retriever can strike a balance between semantic relevance and keyword matching, thereby improving the accuracy and comprehensiveness of retrieval. When using this route, the input question should first be vectorized based on the BGE M3-Emebdding vector model, with restrictions on parameters such as the cause of action, and then searched in the vector version of the legal element knowledge base.
[0077] Step (5): Recall search based on input questions and knowledge of elements at different levels: First, use step (2) to determine the name of the case level based on the input question, then extract key legal elements in combination with the large model in step (3), and finally, combine the three recall routes in step (4) to enter the corresponding legal element knowledge base to search for recall. The specific search process is as follows:
[0078] 1) If the question is a "common user type question", based on the restrictions of the case, combined with the key element information such as case characteristics and pure literal, based on the three recall routes in step (4), search for cases in the original legal element knowledge base, the case feature legal element knowledge base, and the vector version of the legal element knowledge base to obtain the recall set.
[0079] 2) If the question is a "professional legal professional type of question", based on the restrictions of the case, combined with the content of the rewritten question, case characteristics, pure literal key information and other factor information, based on the three recall routes in step (4), search for cases in the original legal factor knowledge base, the case feature legal factor knowledge base, and the vector version of the legal factor knowledge base to obtain the recall set.
[0080] 3) If the question is "a question of the type of original document or long text case", based on the restrictions of the case cause, combined with the case characteristics, dispute focus, judgment essence and other element information, based on the three recall routes in step (4), search for cases in the case characteristics legal element knowledge base, dispute focus legal element knowledge base, judgment essence legal element knowledge base, and vector version legal element knowledge base to obtain the recall set.
[0081] Step (6): According to the weight ratio of different recall routes and the number of return values of the three recall routes in step (4), the recall subset samples of each route are dynamically selected. The specific method of selecting the recall subset is as follows:
[0082] 1) Since the matching rule required by recall route 1 is that all phrases of all elements in the input question must be matched before the case can be returned, the accuracy is relatively high. If the number of recalls of route 1 reaches the required topk, the recall set is directly returned.
[0083] 2) If the return value of route 1 is 0, since both routes 2 and 3 are considered from the overall semantic level, half of the topk will be taken from routes 2 and 3 respectively. If the number of returned items in route 2 is less than half of the topk or is 0, it will be supplemented from route 3. Since route 3 is a hybrid retrieval route based on the vector library, the return value will not be less than topk.
[0084] 3) If the return value of route 1 is greater than 0 or less than topk, the recall set will be selected according to the method in step 2, and then the return cases of route 1 will be added to the front of the recall set in step 2, and the cases with the same number of cases at the end of the recall set in step 2 will be deleted. Then the final recall set will be obtained.
[0085] Step (7): If the number of recall sets returned by "Route 1" is enough for topk, directly return the similar case set. If the number of recall sets returned is not enough for topk, then the recall set of "Route 2 and Route 3" combined is brought into the BGE M3-Emebdding fine sorting model to complete the vector-based precise sorting and obtain the final similar cases. The structure of the fine sorting model is as follows: Figure 3 shown.
[0086] The present invention is tested by professional testers of the company in accordance with the above process in a manually verified data set, and the accuracy rate has increased from more than 50% (similar case retrieval based on rule keywords) to 80% now, which greatly improves the accuracy rate of similar case retrieval.
[0087] The embodiments described above are only specific ways of presenting the present invention. Any technical solutions that can be easily obtained by simply changing or equivalently replacing the above embodiments belong to the protection scope of the present invention.
Claims
1. A similar case retrieval method based on a large model integrating multi-level legal elements, the steps of which include: 1) Based on the structural characteristics of legal documents and combined with the big model, a multi-level legal element knowledge base is designed, including the original legal element knowledge base, the case feature legal element knowledge base, the dispute focus legal element knowledge base, the judgment essence legal element knowledge base, and the vector version legal element knowledge base; 2) receiving a question input by a user when searching for similar cases of a target case, first predicting the case category of the question using the case category hierarchical classification model, and associating the predicted case category with its parent case category name; Then the question recognition classification model is used to predict the question type of the question; 3) Based on the type of the question and the multi-level types of the legal element knowledge base, the large model is used to generate and extract legal element information at different levels to form a corresponding set of legal elements to be searched; 4) Based on the case category, question type and the legal element sets at all levels obtained by searching in step 3), according to the different recall strategies set, search and recall the corresponding original legal element knowledge base, case feature legal element knowledge base, dispute focus legal element knowledge base, judgment essence legal element knowledge base, and vector version legal element knowledge base; 5) Generate the final recall set according to the weight ratio of each recall strategy and the number of recalls corresponding to each recall strategy; 6) Based on the final recall set, similar cases of the target case are obtained and returned to the user.
2. The method according to claim 1, characterized in that The method for constructing the multi-level legal element knowledge base is: The construction of the original text element knowledge base uses regular expressions combined with the document structure to structure the original text of each judgment case, obtain the claim section, the defense section, the court's findings section, the court's opinion section, the judgment result section, and the full text information, and construct the original text legal element knowledge base; To construct the knowledge base of case characteristics and legal elements, we first combine the alleged paragraphs and the ascertained paragraphs of some judgment cases to obtain training samples, and then form regular expressions based on the case characteristics and legal elements system to complete the annotation of some documents to form a training set. After that, we bring the training set into the big model for fine-tuning training to obtain the fine-tuned big model of case characteristics. Then, we complete the database running of the stock documents based on the fine-tuned big model of case characteristics to form a knowledge base of case characteristics and legal elements. The construction of the dispute focus legal element knowledge base first combines the claim and defense sections in any judgment case to form a sample to be input, and uses the labels in the dispute focus legal element system as the output result, then continuously iterates and updates the combination examples to obtain the optimized Prompt instruction, and then completes the extraction of dispute focus elements in each existing judgment case based on the optimized Prompt instruction to build the dispute focus legal element knowledge base; The construction of the knowledge base of legal elements of judgment essence first selects the claim section, defense section, this court's investigation section, this court's opinion section, and judgment result section of some judgment cases based on the sample data structure characteristics of the People's Court case database, and combines them, inputs them into the big model to generate judgment essence information, and obtains the training set after manual screening. Each training data uses the judgment essence section as output and other sections as input; then the screened training set is brought into the big model for fine-tuning training to obtain the fine-tuned judgment essence legal elements big model; finally, the claim section, defense section, this court's investigation section, this court's opinion section, and judgment result section of each judgment case in the stock documents are combined and brought into the judgment essence legal elements big model, extract the judgment essence of each judgment case, and construct the judgment essence legal elements knowledge base; A vector version of the legal element knowledge base is constructed by vectorizing the original paragraph information, case characteristic element information, dispute focus element information, and judgment subject element information of each judgment case to obtain a vector version of the legal element knowledge base.
3. The method according to claim 1 or 2, characterized in that: The big model is used to generate and extract legal element information at different levels based on the problem and its case category, the name of the related superior case, and the type of the problem, to form a set of sentences to be searched with case and legal elements. The specific method is as follows: 31) If the input question is a "normal user type question", then according to the optimized Prompt instruction of pure literal keywords, the big model is used in combination with the input question to extract a set of legal element information of pure literal keywords, and then a set of legal element information containing several case features is generated based on the fine-tuned case feature element big model; 32) If the input question is a "professional legal professional type question", the Prompt instruction rewritten according to the question is used to generate multiple forms of expression using the big model; then for each form of expression, the Prompt instruction of pure literal keywords is used to extract a piece of legal element information of pure literal keywords in combination with the expression using the big model, and then a set of legal element information containing several case features is generated based on the fine-tuned case feature element big model; 33) If the input question is "a question of the original document or long text case type", the big model fine-tuned based on the judgment essence will generate a "judgment essence" legal element based on the question, and then use the big model to summarize the "dispute focus" legal elements based on the Prompt instruction of the dispute focus, and then use the big model fine-tuned based on the case feature legal element system to summarize and generate a legal element information set containing several case characteristics.
4. The method according to claim 3, characterized in that The recall strategy includes a slop search route based on phrase matching, a search route based on the BM25 algorithm, and a hybrid search route based on a vector library; if the question is a "common user type question", the pure literal keyword information and the legal element information set at the case feature level are combined to search for cases in the original legal element knowledge base, the case feature legal element knowledge base, and the vector version of the legal element knowledge base to obtain a recall set; if the question is a "professional legal person type question", the question rewritten content, the pure literal keyword information, and the legal element information set at the case feature level are combined to search for cases in the original legal element knowledge base, the case feature legal element knowledge base, and the vector version of the legal element knowledge base to obtain a recall set; if the question is an "original document or long text case type question", the legal element information at the dispute focus level, the legal element information at the judgment essence level, and the legal element information set at the case feature level are combined to search for cases in the dispute focus legal element knowledge base, the judgment essence legal element knowledge base, the case feature legal element knowledge base, and the vector version of the legal element knowledge base to obtain a recall set.
5. The method according to claim 4, characterized in that The method for obtaining similar cases of the target case based on the final recall set is as follows: if the number of recalls returned by the "slop retrieval route based on phrase matching" reaches the set value topk, then the cases corresponding to the objects recalled by the slop retrieval route based on phrase matching are directly returned as the final recall set; otherwise, the recall results of the "retrieval route based on BM25 algorithm" and the "hybrid retrieval route based on vector library" are combined to obtain a recall set that is input into the BGE M3-Emebdding precise ranking model to complete the precise sorting based on vectors and obtain similar cases of the target case.
6. The method according to claim 1 or 2, characterized in that: The case cause hierarchical classification model is obtained based on UTC classification model training, and the method is: first, the hierarchical category of the case cause is set based on the legal standard specifications, and then the query question is generated by combining the case cause name with the big model based on the Prompt instruction of generating the question, and then the case cause and the query question are combined together to generate a training set, and the UTC classification model is trained to obtain the case cause hierarchical classification model.
7. The method according to claim 1 or 2, characterized in that: The query question classification model is obtained based on UTC classification model training, and the method is: first set the category of the query question, then use the large model to construct a training set in combination with the optimized Prompt instruction, and finally train the UTC classification model to obtain the question classification model.
8. A server, characterized in that: The invention comprises a memory and a processor, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program comprises instructions for executing the method according to any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Question and answer method and device, equipment, medium and program product
CN121456083A
Class case retrieval method, device and equipment based on large language model and vector retrieval
CN122262205A