Engineering case retrieval method and equipment based on large language model and machine learning
By using thinking chains and a prompt word template designed with a small sample technology in case search and using a large language model for multi-dimensional analysis, the problem of being unable to comprehensively evaluate case similarity in the existing technology is solved, and higher retrieval accuracy and interpretability are achieved.
Patent Information
- Application Number
- CN202510285118.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-13
AI Technical Summary
The existing case search methods are mainly based on text matching or single-dimensional similarity comparison, and cannot comprehensively evaluate case similarity from multiple dimensions, and cannot effectively utilize human legal expertise.
Chain-of-Thought (CoT) and small sample technology are used to design prompt word templates, and the case is structured from four dimensions: case fact, dispute focus, legal application, and judgment reasoning through a large language model to generate multi-dimensional similarity scores.
It improves the accuracy and interpretability of the model's problem solving, can more comprehensively simulate human thinking and judgment processes, and enhances the accuracy and reliability of similar case searches.
Smart Images

Figure CN120144740A_ABST
Abstract
Description
1. Technical Field
[0001] The present invention relates to the technical field of classification and retrieval, in particular to a case retrieval method and device for engineering cases based on large language models and machine learning. 2. Background Art
[0002] Case retrieval is of great significance in judicial practice. The "Guiding Opinions of the Supreme People's Court on Unifying the Application of Law and Strengthening Case Retrieval (Trial)" (hereinafter referred to as the "Guiding Opinions"), which came into effect on July 31, 2020, positions case retrieval as a specific system under the statutory law system, aiming to achieve the unified application of law. The cases referred to in the "Guiding Opinions" refer to cases that are similar to the case to be adjudicated in terms of basic facts, disputed issues, legal application issues, etc., and have been adjudicated and become effective by the people's court.
[0003] Case retrieval plays an important role in case adjudication. However, traditional case retrieval systems mostly conduct case search by matching tags, and the matching accuracy and the quality of the pushed cases still cannot meet the actual requirements. To improve case retrieval technology, it is necessary to think hard about how to combine advanced algorithms in the field of artificial intelligence with the characteristics of the legal industry.
[0004] With the development of artificial intelligence technology, the application of natural language processing technology to case retrieval has also achieved a large number of results. Dan et al. proposed a legal event context model, aiming to integrate legal events and background information based on the attention mechanism so that legal events can be associated with corresponding contexts. Sansone et al. analyzed the latest artificial intelligence methods in the legal field, focusing on legal information retrieval systems based on natural language processing, machine learning, and knowledge extraction technologies. Chalkidis et al. studied the application of deep learning in legal analysis, including text classification, information extraction, and information retrieval, where semantic feature representation is a key tool for the successful application of deep learning in natural language processing.
[0005] Generally speaking, the current case retrieval based on artificial intelligence mainly adopts the following technical means: shallow text representation methods that extract small-grained text information such as words, phrases, etc. as features, mainly considering the information contained in the words themselves, and deep text representation methods based on neural networks, pre-trained models, and knowledge graphs. However, the existing case retrieval methods mainly rely on text matching (such as keyword, sentence vector) or only conduct similarity comparison in a single dimension (such as legal provisions), ignoring the complexity of cases. They cannot comprehensively evaluate the similarity of cases from multiple dimensions, cannot well simulate the human thinking and judgment process, and cannot effectively utilize human legal expertise.
[0006] In recent years, generative large language models represented by GPT have achieved success in natural language processing. By predicting missing words or phrases in sentences, large models can perform self-supervised learning training on a large amount of unlabeled text data, thereby extracting the statistical laws and semantic relationships of language. Under the guidance of optimized prompt words, large models can demonstrate excellent reasoning and information extraction capabilities, bringing a new mode for text processing. However, large models may output incorrect results, and simply using large models for case retrieval cannot effectively avoid this "hallucination" phenomenon of large models. Therefore, it is imperative to invent a new case retrieval method. III. Summary of the Invention
[0007] In view of the above situation, to solve the defects of the prior art, the purpose of the present invention is to provide an engineering case retrieval method and device based on large language models and machine learning. The present invention designs a prompt word template using the Chain-of-Thought (CoT) and few-shot techniques. The Chain-of-Thought can simulate the thinking process of humans when solving problems, guiding the model to generate answers through a series of logical reasoning steps. The model can construct a logical reasoning chain and finally obtain the answer. This method not only improves the accuracy of the model in solving problems but also enhances its interpretability.
[0008] One of the technical solutions provided by the present invention is an engineering case retrieval method based on large language models and machine learning, including the following steps:
[0009] Step1: Construct a data set of judgment documents
[0010] Construct a data set D total , d i ={query:query_file,candidate:[candidate_file 1 ,candidate_file 2 ,…,candidate_file ni ,gt_idx:[ID 1 ,ID 2 ,…,ID ni}, where d iIt is a dictionary that includes three keys, namely "query", "candidate", and "gt_idx". "query" represents the query case, "candidate" represents the candidate cases, and "gt_idx" is the true sorting index of the similarity of the candidate cases to the query case. The value of "query" is the case file "query_file" of the query case. The value of "candidate" is a list, and the elements in the list are the judgment files of multiple candidate cases. The judgment files include information such as the case situation, case classification, legal provisions, and judgment reasoning process. Having the same case cause or the same type is the primary step in case retrieval. Type screening measures are adopted to make the candidate cases have the same case cause or type as the query case. The value of "gt_idx" is a list, and the elements in the list are the actual sorting order of the candidate case indices;
[0011] Step2: Design a prompt template using Chain-of-Thought (CoT) and few-shot techniques, and use a large language model to construct the following scoring chains:
[0012] Case situation scoring chain: Chain detail = Prompt detail |LLM|OutputParsers
[0013] Dispute focus scoring chain: Chain focus = Prompt foucs |LLM|OutputParsers
[0014] Legal application scoring chain: Chain law = Prompt law |LLM|OutputParsers
[0015] Judgment reasoning scoring chain: Chain reason = Prompt reason |LLM|OutputParsers
[0016] Among them, Prompt detail is the prompt template for case situation similarity scoring, Prompt foucs is the prompt template for dispute focus similarity scoring, Prompt law is the prompt template for the applicability scoring of the legal provisions cited by the candidate cases to the query case, Prompt reason is the prompt template for the applicability scoring of the reasoning part of the candidate cases to the query case. LLM is the large model adopted, and OutputParsers is the output parser used to convert the output of the large model into a specific format;
[0017] Step3: Score the query case and the candidate cases
[0018] For each data d in D total input the values of query and candidate into Chain i ,Chain detail ,Chain focus ,Chain law ,Chain reason for scoring;
[0019] Step4: Train and cross-validate the ranking machine learning model
[0020] Use the LightGBM model to complete the ranking task. LightGBM is an efficient gradient boosting tree algorithm that can effectively reduce computational complexity and memory usage through histogram-based decision tree learning methods. Use LightGBM as the ranking model and the LambdaRank algorithm to rank according to the scores of candidate cases;
[0021] Use the TPE (Tree-structured Parzen Estimator) algorithm based on Bayesian optimization to optimize the parameters of the LightGBM model. This algorithm constructs a probability model to predict the performance of hyperparameters, and thus selects the most promising hyperparameter combination for evaluation in each iteration. This method can effectively balance exploring new hyperparameter regions and exploiting known excellent hyperparameter regions; To use all data for model training and evaluation, improve the accuracy and reliability of model performance evaluation through multiple rounds of cross-validation;
[0022] Step5: Construct a similar case judgment chain
[0023] Rank candidate cases according to their similarity to the query case. However, whether the candidate case with the highest similarity constitutes a similar case still needs further judgment. To enable the use of large models for similar case judgment, construct the following similar case judgment chain:
[0024] Chain judgment = Prompt judgment |LLM|OutputParsers
[0025] According to the definition of similar cases in the "Guiding Opinions of the Supreme People's Court on Unifying the Application of Law and Strengthening the Retrieval of Similar Cases (Trial)", that is, the similar cases referred to in this opinion refer to cases that are similar to the case to be decided in terms of basic facts, dispute focuses, legal application issues, etc., and have been adjudicated and become effective by the people's court. Combining the importance of the discussion on reasoning and basis in the judgment gist of adjudicated cases in practice for similar case judgment, design a similar case judgment prompt template;
[0026] Step6: Conduct retrieval and judgment of similar cases
[0027] Input the query case and candidate cases into Chain detail 、Chain focus 、Chain law 、Chain reason to obtain the scoring result. Input the scoring result into the trained LightGBM ranking model to obtain the ranking result of candidate cases. Input the query case and candidate cases with relatively high similarity into Chain judgment to obtain the judgment result of similar cases.
[0028] Furthermore, in the above-mentioned Step2, the prompt template is specifically as follows:
[0029] Prompt detail :
[0030] You are a legal expert. {input} is the
basic situation of the case
[0031] Step 1: Extract the basic situation of the cases in {query} and {input} respectively.
[0032] Step 2: Judge whether the basic situation of the cases in {query} and {input} is the same or similar. When making the similarity judgment, mainly consider whether there is a strong similarity in the actions of the parties in the cases of {query} and {input}. Analyze step by step and give the similarity score of the basic situation of the cases in {query} and {input}. The score range is from 0 to 100, with 100 being the most similar and 0 being completely different.
[0033] Prompt foucs :
[0034] You are a legal expert. {input} is the
basic situation of the case
[0035] Step 1: You need to summarize the core disputes in the cases of {query} and {input} respectively. Analyze step by step. The steps for summarizing the focus of the dispute are as follows:
[0036] (1) Clarify the background and basic facts of the case.
[0037] Determine what legal relationships are involved in the case (such as contract relationships, debt relationships, etc.). Sort out the factual disputes between the plaintiff and the defendant, including the conclusion and performance of the contract, the amount in dispute, etc.
[0038] (2) Analyze the litigation requests of both parties.
[0039] What does the plaintiff request the court to decide? What does the defendant oppose?
[0040] Mainly focus on the litigation requests of the plaintiff and the content of the defendant's defenses.
[0041] (3) Examine the core issues in dispute.
[0042] Among the claims of both parties, which issue is the core issue that the court needs to judge?
[0043] Determine whether there is actual evidence to support the claim of a certain party, especially the key points in dispute.
[0044] (4) Analyze the application of law and contract terms.
[0045] Check the relevant legal provisions and the specific content of the contract to determine whether the focus of the dispute involves the performance, interpretation, or payment conditions of the contract terms, etc.
[0046] (5) Judge the focus of the dispute.
[0047] Based on the analysis of the above steps, determine the most critical point in dispute in the case.
[0048] Step 2: You need to judge whether the focus of the dispute in {query} is the same or similar to that in the case of {input}, analyze step by step, and give the similarity score of the focus of the dispute in {query} and the case in {input}, with the score range from 0 to 100, a score of 100 means exactly the same, and a score of 0 means completely different. Scoring criteria: 90 - 100 points: The focus of the dispute is essentially the same; 60 - 89 points: The main aspects of the focus of the dispute are similar; 30 - 59 points: Some parts of the focus of the dispute are relevant; 1 - 29 points: The relevance of the focus of the dispute is very low; 0 points: Completely irrelevant; The scoring should be based on the similarity of the core content of the focus of the dispute in {query} and {input}, not just based on the similarity of words.
[0049] Step 3: Generate a similarity score represented by a number.
[0050] Step 4: Output a similarity score represented by only one number.
[0051] Prompt law :
[0052] You are a legal expert. {input} is [the legal provisions on which the judge's judgment reason and basis of the case], and you are required to evaluate the similarity between the case in {input} and {query} based on the application of law. The similarity needs to be based on the application of law, and the specific steps are as follows:
[0053] Step 1: You need to extract the legal provisions and judicial interpretation provisions cited by the judge in the judgment from {input}.
[0054] Step 2: You find the original legal texts of the legal provisions extracted in Step 1.
[0055] Step 3: You need to analyze whether each legal provision is applicable to {query} according to the original legal provisions and judicial interpretation provisions cited in the case in {input}. Analyze step by step and give the similarity score between {query} and the legal provisions and judicial interpretation provisions cited in the case in {input} (the similarity score range is 0 to 100, and the scoring method is: the number of legal provisions and judicial interpretation provisions cited in the case in {input} that are applicable to {query} / the total number of legal provisions and judicial interpretation provisions cited in the case in {input}).
[0056] Step 4: Generate a similarity score represented by a number.
[0057] Step 5: Output a similarity score represented by only one number.
[0058] Prompt reason :
[0059] You are a legal expert. {input} is [the judge's judgment reasoning and the legal provisions on which it is based in the case], and it is necessary to evaluate the applicability between the judgment reasoning in {input} and {query}. The specific steps are as follows:
[0060] Step 1: Extract the part of the judge's judgment reasoning in the case of {input}.
[0061] Step 2: Judge whether the reasoning in the judge's judgment in the case of {input} is applicable to {query}. Analyze step by step and give the applicability score. The similarity score range is 0 to 100, with a score of 100 indicating full applicability and a score of 0 indicating full inapplicability.
[0062] Step 3: Generate a similarity score represented by a number.
[0063] Step 4: Output a similarity score represented by only one number.
[0064] Preferably, in the said Step 3, the scoring method adopts an adaptive control mechanism, which specifically includes the following steps:
[0065] (1) Initialize the number of calculation rounds \(n\leftarrow0\),
[0066] (2) Establish the large model parameter calculation formula:
[0067] Temperatur = max(IV t ×f(n) t , LL t )
[0068] where IV t is the initial value of Temperatur, IV t = 0.75; f(n) t is the adjustment function of Temperatur; the adjustment function adopted in the present invention is LL t = 0;
[0069] Top P = max(IV p ×f(n) p , LL p )
[0070] where IV p is the initial value of Top P, IV p = 0.8, f(n) p is the adjustment function of Top P, the adjustment function adopted in the present invention is LL p = 0.05;
[0071] Substitute Temperatur and Top P into the large model;
[0072] (3) Input the query case and a single candidate case into the scoring chain, repeat the scoring 3 times to obtain the scoring matrix
[0073]
[0074] According to S, calculate the coefficient of variation by column to obtain CV detal 、CV focus 、CV law 、CV reason , measure the stability of the large model scoring according to the coefficient of variation of each dimension scoring. The closer the coefficient of variation is to 0, the better the stability;
[0075] Calculate the index ICC3 based on S as an index to measure the consistency of the large model's scores for each dimension. The Intraclass Correlation Coefficient (ICC) is a statistical index used to evaluate the consistency of measurement results of the same group of objects by different evaluators or at different time points. ICC3 is used to evaluate the scoring consistency of fixed evaluators for the same group of objects. The value range of ICC3 is from 0 to 1. The closer the value is to 1, the higher the consistency. In the present invention, ICC3 is used to evaluate the consistency of multiple scoring results of the large model, which can effectively control the quality of the model output. If the ICC3 value is lower than the set threshold, it indicates that the scoring result consistency is poor, and the model parameters or scoring strategy need to be adjusted to improve the reliability of the output;
[0076] (4) Determine the threshold threshold of the coefficient of variation CV = 0.15 and the ICC3 threshold threshold ICC3 = 0.8, and set the following rules:
[0077] if(max(CV detal ,CV focus ,CV law ,CV reason )>threshold CV )∨(ICC3<threshold ICC3 ):
[0078] Let n = n + 1,
[0079] Repeat steps (2) to (4);
[0080] else:
[0081] Output the mean of the scores for each dimension as the final score;
[0082] Obtain the corresponding score score under the guidance of the prompt word detail 、score focus 、score law 、score reason , and then obtain the data set D' total , Among them,
[0083]
[0084] Preferably, in the said Step4, the cross-validation method adopts 10-fold cross-validation, including the following steps:
[0085] (1) Data set division:
[0086] After scoring the data set D' total is divided into 10 subsets D 1 , D 2 , …, D 10 , that is, the size of each subset is N / 10, where N is the total number of samples in the data set.
[0087] (2) Training and validation of the model:
[0088] In each round i (i ∈ [1, 10]), let the test set D' test = D i , and the remaining data is D -i = D' total \D i . To optimize the model, D -i is further divided into a training set D' train and a validation set D' validation in the ratio of 80% / 20%.
[0089] Use the training set D' train to train the LightGBM model and perform hyperparameter tuning on the validation set D' validation .
[0090] Predict with the trained model on the test set D' test , calculate the ranking evaluation metrics, and the evaluation metrics include: P@5, P@10, MAP@10, NDCG@10, NDCG@20, NDCG@30, MRR. The larger each metric is, the better.
[0091] (3) Repeating process:
[0092] Repeat step (2), each time selecting a different subset as the test set to ensure that each subset has the opportunity to be used as the test set to evaluate the model performance;
[0093] (4) Comprehensive evaluation:
[0094] After completing all 10 rounds of calculations, calculate the mean and coefficient of variation of each evaluation metric to measure the stability and generalization ability of the model on different test sets. The final evaluation metric is the average of the 10-fold evaluation metrics.
[0095] Furthermore, in Step 5 mentioned above, the specific template of the similar case judgment prompt is as follows:
[0096] Prompt judgment :
[0097] You are a legal expert. Similar cases refer to cases that are similar to the case to be decided in terms of basic facts, disputed issues, legal application, etc., and have been adjudicated and become effective by the people's court. Please determine whether the following {query} and {input} constitute similar cases.
[0098] Thought process:
[0099] Step 1: Comparison of basic facts:
[0100] Compare the main facts of {query} and {input}, including aspects such as the background of the case, the parties, the time and place of the incident, the nature of the act, the damage result, etc.
[0101] Do {query} and {input} have similar factual situations? If so, please briefly explain.
[0102] Step 2: Comparison of disputed issues:
[0103] Analyze whether the disputed issues of {query} and {input} are the same.
[0104] Are the disputed points of {query} and the disputed points of {input} consistent in terms of legal issues and actual operations? If the same, please briefly explain the specific similar parts, such as contract interpretation, liability for breach of contract, liability for tort, etc.
[0105] Step 3: Comparison of legal application:
[0106] Judge whether {query} and {input} apply the same or similar legal provisions, especially in the application of legal articles, judicial interpretations, guiding cases and relevant regulations.
[0107] Are {query} and {input} consistent in terms of legal application? If consistent, please briefly explain the applicable legal articles and judicial interpretations.
[0108] Step 4: Applicability of the gist of the judgment:
[0109] Analyze whether the gist of the judgment in {input} is applicable to {query}, focusing on whether the legal reasoning, judgment basis and conclusion in the gist of the judgment can provide a reference basis for resolving the disputed issues of {query}.
[0110] If applicable, please briefly explain the core content of the gist of the judgment and the reasons for its applicability to the disputed issues of {query}.
[0111] Step 5: Judgment:
[0112] Based on the above analysis, especially by focusing on whether the gist of the judgment in {input} is applicable to the focus of the dispute in {query}, it is thus determined whether {query} and {input} constitute similar cases.
[0113] If they constitute similar cases, please state the reasons; if they do not constitute similar cases, please state the main differences. Even if there are differences in some aspects, if the core issues (such as the nature of the legal relationship, the essence of the focus of the dispute, etc.) are similar, they may still constitute similar cases. However, if the core content of the gist of the judgment is not fully applicable to resolving the focus of the dispute in {query}, it cannot be determined as a similar case.
[0114] Based on the same inventive concept, the second technical solution provided by the present invention is a computer electronic device corresponding to the first technical solution, and this device includes the following components:
[0115] Memory: used to store computer programs and related data, the computer program includes the code for implementing the method for retrieving similar cases of the present invention, as well as the related interface code for calling the large model API.
[0116] Processor: used to execute the computer program to implement the method for retrieving similar cases of the present invention.
[0117] Further, the processor executes the following steps:
[0118] 1) Call the large model API through the Internet;
[0119] 2) Score the similarity between the query case and the candidate cases from four dimensions;
[0120] 3) Input the scores of the four dimensions into the machine learning ranking model to generate the final result of retrieving similar cases.
[0121] Internet communication device: used to communicate with other devices or services through the Internet, and the Internet communication device includes but is not limited to the following types:
[0122] Network Interface Card (NIC): used to connect to a local area network or a wide area network, supporting high-speed wired or wireless network communication.
[0123] Load Balancer: used to distribute network traffic to ensure the high availability and stability of the system.
[0124] Security Gateway: used to protect the security of data transmission and prevent unauthorized access.
[0125] The beneficial technical effects of the present invention:
[0126] (1) Multi-dimensional Case Similarity Analysis Based on Prompt Words
[0127] Compared with traditional methods that usually rely on single-dimensional text matching or semantic embedding, the present invention decomposes the task of similar case retrieval into similarity scores of query cases and candidate cases from different dimensions. It introduces the chain of thought prompt words into the field of similar case retrieval, and through the prompt words, guides the large model to conduct a structured analysis of cases from four dimensions: case facts, dispute focus, legal application, and judgment reasoning, generating multi-dimensional similarity scores. The professional knowledge of humans is integrated into the prompt word template, which can better achieve human-computer collaboration.
[0128] (2) Adaptive Control Mechanism for the Output of Large Model
[0129] The determination of the similarity of legal cases involves multiple dimensions, and the scoring of each dimension may be affected by background and model bias. By scoring each dimension multiple times and calculating the coefficient of variation and ICC3, the stability and consistency of the scoring can be effectively measured. The coefficient of variation reflects the stability of the scoring. A low coefficient of variation indicates that the model judgment is stable, while a higher one indicates large fluctuations in the scoring and the need to adjust the model. ICC3 quantifies the consistency among scorers, ensuring high reliability of the scoring. When the coefficient of variation or ICC3 exceeds the threshold, the system can adjust the model parameters to re-score, thereby improving the stability and consistency of the output. By combining the coefficient of variation and ICC3 and adaptively controlling the output quality of the large model, the reliability and accuracy of similar case retrieval can be improved.
[0130] (3) Combining Large Model with Machine Learning Ranking Model
[0131] The present invention combines the semantic understanding ability of the large model with the machine learning ranking model. The large model is responsible for generating multi-dimensional similarity scores, and the machine learning ranking model ranks the candidate cases based on these scores, making use of both the powerful reasoning ability of the large model and achieving efficient and accurate ranking through the machine learning model.
[0132] (4) Interpretability and Scalability
[0133] The present invention can provide users with clear bases for case similarity through multi-dimensional scores and the machine learning ranking model, enhancing the interpretability of the retrieval results.
[0134] The framework of the present invention has high scalability and can expand the analysis dimensions according to specific needs (such as time, region, court level, etc.). For cases with fewer precedents, the present invention can guide the large model to give similarity scores through prompt words, and then use the trained machine learning ranking model to rank the candidate cases, which can avoid the decline in retrieval accuracy caused by insufficient training data. IV. Description of the Drawings
[0135] Figure 1This is the flow chart of case retrieval for the present invention. V. Specific Embodiments
[0136] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts belong to the scope of protection of the present invention.
[0137] Refer to Figure 1 , in a preferred embodiment, an engineering case retrieval method based on large language models and machine learning specifically includes the following steps:
[0138] Step1: Construct a collection of judgment document data
[0139] Construct a data set D total , d i ={query:query_file,candidate:[candidate_file 1 ,candidate_file 2 ,…,candidate_file ni ,gt_idx:[ID 1 ,ID 2 ,…,ID ni}, where d i is a dictionary, including three keys, namely query, candidate, and gt_idx. Query represents the query case (pending case), candidate represents the candidate case (decided case), and gt_idx is the true sorting index of the similarity of the candidate case to the query case. The value of query is the case file query_file of the query case, the value of candidate is a list, and the elements in the list are judgment documents of multiple candidate cases. The judgment documents include information such as case facts, case classification, legal provisions, and judgment reasoning process. Having the same case name or type is the primary step in case retrieval. Here, a type screening measure is adopted, that is, the candidate case should have the same case name or type as the query case. The value of gt_idx is a list, and the elements in the list are the actual sorting order of the candidate case indexes;
[0140] Step2: Design a prompt template using the Chain-of-Thought (CoT) and few-shot techniques, and use a large language model to construct the following scoring chain:
[0141] Case fact scoring chain: Chaindetail = Prompt detail |LLM|OutputParsers
[0142] Dispute focus scoring chain: Chain focus = Prompt foucs |LLM|OutputParsers
[0143] Legal application scoring chain: Chain law = Prompt law |LLM|OutputParsers
[0144] Adjudication reasoning scoring chain: Chain reason = Prompt reason |LLM|OutputParsers
[0145] Among them, Prompt detail is the similarity scoring prompt template for case facts, Prompt foucs is the similarity scoring prompt template for dispute focus, Prompt law is the applicability scoring prompt template for the legal provisions cited in candidate cases to the query case, Prompt reason is the applicability scoring prompt template for the reasoning part of candidate cases to the query case. LLM is the large model adopted. OutputParse is the output parser used to convert the output of the large model into a specific format.
[0146] The specific prompt templates are as follows:
[0147] Prompt detail :
[0148] You are a legal expert. {input} is the
basic situation of the case
[0149] Step 1: Extract the basic situation of the cases in {query} and {input} respectively.
[0150] Step 2: Judge whether the basic situation of the cases in {query} and {input} is the same or similar. When making the similarity judgment, mainly consider whether there is a strong similarity in the actions of the parties in the cases of {query} and {input}. Analyze step by step and give the similarity score of the basic situation of the cases in {query} and {input} (the score range is 0 to 100, with a score of 100 being the most similar and a score of 0 being completely different).
[0151] Refer to the following examples:
[0152] Example 1:
[0153] Basic situation of {query}: Plaintiff Peng Zhengkun signed an equity transfer agreement with defendants Ran Qibin and Ma Weichi, transferring 75% of the equity of Chongqing Kekexi Ecological Fruit Industry Co., Ltd. and paying 3.75 million yuan. The plaintiff found that the company could not operate, demanded to terminate the contract, refund the equity payment and deposit, compensate for the loss of interest, and demanded that the three defendants bear joint liability and litigation costs.
[0154] Basic situation of {input}: Plaintiff Zhao Yong signed an equity transfer agreement with defendants Chen Feng and Li Jun. Chen Feng and Li Jun transferred 70% of the equity of East China Environmental Protection Technology Co., Ltd. they held to Zhao Yong and received a transfer payment of 3.5 million yuan. Zhao Yong found that the company could not continue to operate, demanded to terminate the contract, refund the equity payment and deposit, and demanded that the defendants bear joint liability and litigation costs.
[0155] Score: 100
[0156] Judgment process: After the equity transfer in both {query} and {input}, it was found that the transferred company had major defects and could not operate. The assignee demanded to terminate the contract and compensate for the losses. Therefore, the two cases are exactly similar and the score is 100.
[0157] Other examples are omitted.
[0158] Step 3: Generate a similarity score represented by a number.
[0159] Step 4: Output a similarity score represented by only one number.
[0160] Prompt foucs :
[0161] You are a legal expert. {input} is [the basic situation of the case]. You are required to conduct a similarity assessment of the case in {input} with {query}. The similarity assessment is based on the focus of the dispute, and the specific steps are as follows:
[0162] Step 1: You need to respectively summarize the core disputes in the cases of {query} and {input}. For step-by-step analysis, the steps for summarizing the focus of the dispute are as follows:
[0163] (1) Clarify the background and basic facts of the case.
[0164] Determine what legal relationship the case involves (such as contract relationship, debt relationship, etc.). Sort out the factual disputes between the plaintiff and the defendant, including the signing and performance of the contract, the amount in dispute, etc.
[0165] (2) Analyze the litigation requests of both parties.
[0166] What does the plaintiff request the court to decide? What does the defendant oppose?
[0167] Mainly focus on the litigation requests of the plaintiff and the content of the defendant's defenses.
[0168] (3) Examine the core issues in dispute.
[0169] Among the claims of both parties, which issue is the core issue that the court needs to judge?
[0170] Determine whether there is actual evidence to support the claim of a certain party, especially the key points in dispute.
[0171] (4) Analyze the application of law and contract terms.
[0172] Check the relevant legal provisions and the specific content of the contract to determine whether the focus of the dispute involves the performance, interpretation, or payment conditions of the contract terms, etc.
[0173] (5) Determine the focus of the dispute.
[0174] Based on the analysis of the above steps, determine the most critical point in dispute in the case.
[0175] Examples of summarizing the focus of the dispute are as follows:
[0176] Basic case: The plaintiff, Anhui Mingyuan Electric Power Equipment Manufacturing Co., Ltd., alleged that: The plaintiff and the defendant signed a "Custom Manufacturing Contract" on September 1, 2014. After the contract was signed, the plaintiff supplied goods to the defendant as agreed. As of July 2017, the defendant owed 536,511 yuan in payment for goods. Now the plaintiff sues to request the court to order the defendant to pay 536,511 yuan in payment for goods and interest. The defendant admitted the facts claimed by the plaintiff, but believed that the payment for goods claimed by the plaintiff included a quality guarantee deposit of 5% of the total contract price, and the payment conditions for the quality guarantee deposit have not been fulfilled.
[0177] Thought process of summarizing the focus of the dispute:
[0178] (1) Background and basic facts of the case
[0179] Contract signing: The plaintiff, Anhui Mingyuan Electric Power Equipment Manufacturing Co., Ltd., and the defendant signed a "Custom Manufacturing Contract" on September 1, 2014.
[0180] Contract performance: After the contract was signed, the plaintiff fulfilled the obligation of supplying goods according to the contract.
[0181] Dispute over payment for goods: The plaintiff claimed that as of July 2017, the defendant owed 536,511 yuan in payment for goods and requested payment.
[0182] (2) Litigation requests of both parties
[0183] The plaintiff's lawsuit request: Requiring the defendant to pay 536,511 yuan for the goods and interest.
[0184] The defendant's defense: The defendant admits the facts claimed by the plaintiff, but believes that 5% of the total contract price is included as the quality warranty deposit in the payment for the goods, and the payment conditions for the quality warranty deposit have not been fulfilled. Therefore, this part of the amount should not be paid.
[0185] (3) Examine the core issues in dispute
[0186] Whether the payment for goods includes the quality warranty deposit: The core of the dispute is that the defendant believes that 5% of the quality warranty deposit is included in the payment for goods, and the payment conditions for the quality warranty deposit are not met.
[0187] Payment conditions for the quality warranty deposit: The defendant believes that the payment conditions for the quality warranty deposit have not been met, while the 536,511 yuan payment for goods required by the plaintiff includes the quality warranty deposit.
[0188] (4) Analyze the applicable law and contract terms
[0189] According to the relevant provisions of the Contract Law, a party to a contract can request the other party to pay the payment for goods after fulfilling the contract obligations. However, if special items such as the quality warranty deposit are involved in the payment for goods, whether to pay depends on the payment conditions stipulated in the contract.
[0190] It is necessary to check the terms in the "Customized Contract" to confirm whether there are clear payment conditions for the quality warranty deposit, such as conditions like the goods being accepted as qualified or the service period expiring.
[0191] (5) Determine the focus of the dispute
[0192] Focus of the dispute: Whether the quality warranty deposit included in the contract should be paid, especially whether the payment conditions for the quality warranty deposit have been fulfilled. The defendant claims that the payment conditions for the quality warranty deposit have not been met, while the plaintiff requests the payment of the full payment for goods. Therefore, the court needs to determine whether the payment conditions for the quality warranty deposit have been fulfilled to decide whether the quality warranty deposit part should be paid.
[0193] Step 2: You need to determine whether the focus of the dispute in {query} is the same or similar to that in the case in {input}, analyze step by step, and give the similarity score of the focus of the dispute in {query} and the case in {input} (the score range is 0 to 100, with a score of 100 being exactly the same, a score of 0 being completely different, and the scoring criteria: 90 - 100 points: The focus of the dispute is essentially the same; 60 - 89 points: The main aspects of the focus of the dispute are similar; 30 - 59 points: There are some correlations in the focus of the dispute; 1 - 29 points: The relevance of the focus of the dispute is very low; 0 points: Completely irrelevant). The scoring should be based on the similarity of the core content of the focus of the dispute in {query} and {input}, not just on the similarity of words.
[0194] Examples of determining the similarity of the focus of the dispute are as follows:
[0195] Example 1:
[0196] The focus of controversy in {query}: Whether the quality warranty deposit in the contract should be paid, especially whether the payment conditions for the quality warranty deposit have been met.
[0197] The focus of controversy in {input}: Whether the quality deposit can be refunded.
[0198] Score: 100
[0199] Thought process: The focus of controversy in both {query} and {input} involves whether the quality warranty deposit can be refunded. Although the two focuses of controversy are expressed differently, their basic meanings are the same, so the score is 100.
[0200] Other examples are omitted.
[0201] Step 3: Generate a similarity score represented by a number.
[0202] Step 4: Output a similarity score represented by only one number.
[0203] Prompt law :
[0204] You are a legal expert. {input} is [the judge's reasoning for the judgment in the case and the legal provisions relied on]. You are required to evaluate the similarity between the case in {input} and {query} based on the application of the law. The similarity evaluation is based on the following specific steps:
[0205] Step 1: You need to extract from {input} the legal provisions and judicial interpretation provisions cited by the judge in the judgment. Examples are as follows:
[0206] Example: {input} is [the judge's reasoning process for the judgment in the case and the legal provisions relied on]: The court holds that in October 2014, the plaintiff had delivered the customized goods, and the defendant had actually installed and used the customized goods. According to the contract, if no quality problems are raised within 24 months after the installation and acceptance of the customized goods, the payment should be made in a timely manner. The defendant's agent's claim that the payment conditions for the quality warranty deposit have not been met does not conform to the facts. Accordingly, in accordance with Article 263 of the Contract Law of the People's Republic of China, the judgment is as follows.
[0207] The legal provisions and judicial provisions you extracted are: Article 263 of the Contract Law of the People's Republic of China.
[0208] Step 2: You find the original legal text of the legal provisions extracted in Step 1. Examples are as follows:
[0209] Example: The legal provisions extracted in the first step are Article 263 and Article 268 of the "Contract Law of the People's Republic of China". The original text of Article 263 of the "Contract Law of the People's Republic of China" is: omitted. The original text of Article 268 of the "Contract Law of the People's Republic of China" is: omitted.
[0210] Step 3: You need to analyze whether each legal provision and judicial interpretation provision cited in the case in {input} applies to {query} according to their original texts. Analyze step by step and give the similarity score between {query} and the legal provisions and judicial interpretation provisions cited in the case in {input} (the similarity score ranges from 0 to 100, and the scoring method is: the number of legal provisions and judicial interpretation provisions cited in the case in {input} that apply to {query} / the total number of legal provisions and judicial interpretation provisions cited in the case in {input}). The following is an example:
[0211] Example 1:
[0212] The original texts of the legal provisions and judicial interpretation provisions cited in the case in {input} are: The original text of Article 263 of the "Contract Law of the People's Republic of China" is: omitted. The original text of Article 268 of the "Contract Law of the People's Republic of China" is: omitted.
[0213] {query} is: The plaintiff, Anhui Mingyuan Electric Power Equipment Manufacturing Co., Ltd., alleged that: The plaintiff and the defendant signed a "Custom Manufacturing Contract" on September 1, 2014. After the contract was signed, the plaintiff supplied goods to the defendant as agreed. As of July 2017, the defendant owed 536,511 yuan in payment for goods. Now the plaintiff sues to request that the defendant pay 536,511 yuan in payment for goods and interest. The defendant admitted the facts claimed by the plaintiff, but believed that 5% of the total contract price in the payment for goods claimed by the plaintiff was the quality warranty deposit, and the payment condition for the quality warranty deposit has not been fulfilled.
[0214] Similarity score: 50
[0215] Analysis process:
[0216] (1) The original text of Article 263 of the "Contract Law of the People's Republic of China" is: omitted. {query} belongs to a dispute over the payment of custom-made goods. In October 2014, the plaintiff had delivered the custom-made goods, and the defendant had actually installed and used the custom-made goods. According to the contract, if no quality problems are raised within 24 months after the installation and acceptance of the custom-made goods, the payment should be made in a timely manner. Therefore, Article 263 of the "Contract Law of the People's Republic of China" applies to {query}.
[0217] (2) The original text of Article 268 of the "Contract Law of the People's Republic of China" is: omitted. {query} does not involve the dissolution of the contract for work, so Article 268 of the "Contract Law of the People's Republic of China" does not apply to {query}.
[0218] (3) There are two legal provisions and judicial interpretation provisions cited in the case of {input}, and one of them is applicable to {query}, so the score is 1 / 2 = 50.
[0219] Step 4: Generate a similarity score represented by a number.
[0220] Step 5: Output a similarity score represented by only one number.
[0221] Prompt reason :
[0222] You are a legal expert. {input} is [the judge's reasoning in the case judgment and the legal provisions relied on], and it is necessary to conduct an applicability assessment based on the judgment reasoning in {input} and {query}. The specific steps are as follows:
[0223] Step 1: Extract the judge's reasoning part in the case of {input}.
[0224] Step 2: Determine whether the judge's reasoning in the {input} case judgment is applicable to {query}. Analyze step by step and give an applicability score (the similarity score range is 0 to 100, with a score of 100 indicating full applicability and a score of 0 indicating no applicability at all).
[0225] For example:
[0226] Description: The judgment reasoning of {input} is that in October 2014, the plaintiff had delivered the customized products, and the defendant had actually installed and used the customized products. According to the contract agreement, if no quality problems were raised within 24 months after the installation and acceptance of the customized products, the payment should be made in a timely manner. The defendant's agent's claim that the payment condition for the quality warranty deposit was not fulfilled did not conform to the facts. Conduct an applicability analysis on {query} in the following examples respectively.
[0227] Example 1:
[0228] {query} is: The plaintiff, Anhui Mingyuan Electric Power Equipment Manufacturing Co., Ltd., alleges that: The plaintiff and the defendant signed a "Customized Contract" on September 1, 2014. After the contract was signed, the plaintiff supplied goods to the defendant as agreed. As of July 2017, the defendant owed 536,511 yuan in payment for goods. Now the plaintiff sues to request that the defendant pay 536,511 yuan in payment for goods and interest. The defendant admits the facts claimed by the plaintiff, but believes that 5% of the total contract price in the payment for goods claimed by the plaintiff is the quality warranty deposit, and the payment condition for the quality warranty deposit has not been fulfilled.
[0229] Score: 100
[0230] Analysis process:
[0231] (1) Contract Performance and Payment Terms: The statement in the judgment reasoning that "if no quality issues are raised within 24 months after installation and acceptance, payment should be made in a timely manner" indicates that there may be clear payment terms in the contract, that is, the defendant should pay the full amount provided that no quality issues occur within 24 months after acceptance. If such a clause indeed exists in the contract and there is no record of quality issues within 24 months, then the defendant should pay the full purchase price, including the quality warranty deposit.
[0232] (2) Dispute over the Quality Warranty Deposit: The defendant believes that the payment conditions for the quality warranty deposit have not been met, but the judgment reasoning points out that this claim is inconsistent with the facts. If the quality warranty period has expired and no quality issues have occurred, the payment conditions for the quality warranty deposit should be considered as having been met.
[0233] (3) Timeline Analysis: The contract in question was signed on September 1, 2014, and the customized product was delivered in October 2014. As of July 2017, the 24 - month quality warranty period has passed. If there is no record of quality issues during this period, the payment conditions for the quality warranty deposit should be considered as having been met.
[0234] (4) Application of Law: According to the relevant provisions of the "Contract Law", the client shall pay the remuneration as agreed. As part of the contract, the quality warranty deposit should be paid if there are no quality issues.
[0235] Conclusion:
[0236] Applicability Judgment: The judgment reasoning is consistent with the core dispute points of the case in question, that is, whether the payment conditions for the quality warranty deposit have been met.
[0237] Rationality: If the contract terms and facts of the case in question support the reasoning in the judgment reasoning, then the judgment reasoning is applicable to the case.
[0238] The Defendant's Claim: The defendant's claim that the quality warranty deposit has not been met in the case in question needs to be supported by specific factual basis (such as records of quality issues), otherwise the judgment reasoning is reasonable.
[0239] In summary, the judgment reasoning is applicable to the case, provided that the contract terms are clear and the facts support that the payment conditions for the quality warranty deposit have been met. So the score is 100.
[0240] Other examples are omitted.
[0241] Step 3: Generate a similarity score represented by a number.
[0242] Step 4: Output a similarity score represented by only one number.
[0243] Step3: Score the query case and candidate cases
[0244] For D total for each data d i , input the values of query and candidate into Chain detail , Chain focus , Chain law , Chain reason , and conduct scoring.
[0245] During scoring, to control the scoring quality and ensure the stability and consistency of scoring, an adaptive control mechanism is adopted. The specific approach is as follows:
[0246] (1) Initialize the number of calculation rounds n←0,
[0247] (2) Establish the large model parameter calculation formula:
[0248] Temperatur = max(IV t ×f(n) t , LL t )
[0249] where IV t is the initial value of Temperatur, IV t = 0.75; f(n) t is the adjustment function of Temperatur; the adjustment function adopted in the present invention is LL t = 0.
[0250] Top P = max(IV p ×f(n) p , LL p )
[0251] where IV p is the initial value of Top P, IV p = 0.8, f(n) p is the adjustment function of Top P, and the adjustment function adopted in the present invention is LL p = 0.05.
[0252] Substitute Temperatur and Top P into the large model.
[0253] (3) Input the query case and a single candidate case into the scoring chain, repeat scoring 3 times, and obtain the scoring matrix
[0254]
[0255] Calculate the coefficient of variation (CV) column by column according to S to obtain CV detal and CV focus and CV law and CV reason Measure the stability of the large model's scores based on the coefficient of variation of each dimension's scores. The closer the coefficient of variation is to 0, the better the stability.
[0256] Calculate the index ICC3 according to S as an index to measure the consistency of the large model's scores for each dimension. The Intraclass Correlation Coefficient (ICC) is a statistical index used to evaluate the consistency of measurement results of the same group of objects by different evaluators or at different time points. ICC3 is used to evaluate the score consistency of fixed evaluators for the same group of objects. The value range of ICC3 is from 0 to 1, and the closer the value is to 1, the higher the consistency. The present invention uses ICC3 to evaluate the consistency of the multiple scoring results of the large model, which can effectively control the quality of the model output. If the ICC3 value is lower than the set threshold, it indicates that the consistency of the scoring results is poor, and the model parameters or scoring strategies need to be adjusted to improve the reliability of the output.
[0257] (4) Determine the threshold threshold of the coefficient of variation CV = 0.15 and the threshold threshold of ICC3 ICC3 = 0.8, and set the following rules:
[0258] if (max(CV detal , CV focus , CV law , CV reason ) > threshold CV ) ∨ (ICC3 < threshold ICC3 ):
[0259] Let n = n + 1,
[0260] Repeat steps (2) to (4).
[0261] else:
[0262] Output the mean of the scores for each dimension as the final score.
[0263] Obtain the corresponding score score under the guidance of the prompt detail and score focus and score law and score reason Furthermore, obtain the data set D' total , where
[0264]
[0265] Step 4: Train and cross-validate the ranking machine learning model
[0266] The present invention uses the LightGBM model to complete the ranking task. LightGBM is an efficient gradient boosting tree algorithm that can effectively reduce computational complexity and memory usage through a histogram-based decision tree learning method. The present invention uses LightGBM as the ranking model and the LambdaRank algorithm to rank according to the scores of candidate cases.
[0267] The TPE (Tree-structured Parzen Estimator) algorithm based on Bayesian optimization is used to optimize the parameters of the LightGBM model. This algorithm constructs a probability model to predict the performance of hyperparameters, and thus selects the most promising hyperparameter combination for evaluation in each iteration. This method can effectively balance exploring new hyperparameter regions and exploiting known excellent hyperparameter regions.
[0268] In order to use all the data for model training and evaluation and improve the accuracy and reliability of model performance evaluation through multiple rounds of cross-validation, the present invention adopts 10-fold cross-validation, and the specific method is as follows:
[0269] (1) Dataset division:
[0270] Divide the scored dataset D' total into 10 subsets D 1 , D 2 , …, D 10 , that is, the size of each subset is N / 10, where N is the total number of samples in the dataset.
[0271] (2) Model training and validation:
[0272] In each round i (i ∈ [1, 10]), let the test set D' test = D i , and the remaining data is D -i = D' total \ D i . To optimize the model, D -i is further divided into a training set D' train and a validation set D' validation in a ratio of 80% / 20%.
[0273] Use the training set D' train to train the LightGBM model and use it on the validation set D' validationPerform hyperparameter tuning on it.
[0274] Perform predictions on the trained model using the test set D′ test and calculate the ranking evaluation metrics. The evaluation metrics include: P@5, P@10, MAP@10, NDCG@10, NDCG@20, NDCG@30, MRR. The larger each metric is, the better.
[0275] (3) Repeat process:
[0276] Repeat step (2), each time selecting a different subset as the test set to ensure that each subset has the opportunity to be used as the test set to evaluate the model performance.
[0277] (4) Comprehensive evaluation:
[0278] After completing all 10 rounds of calculations, calculate the mean and coefficient of variation of each evaluation metric to measure the stability and generalization ability of the model on different test sets. The final evaluation metric is the average of the 10-fold evaluation metrics.
[0279] Step5: Construct a similar case judgment chain
[0280] The provided similar case retrieval method of the present invention can rank candidate cases according to their similarity to the query case. However, it is still necessary to further determine whether the candidate case with the highest similarity constitutes a similar case. In order to be able to use the large model for similar case judgment, construct the following similar case judgment chain:
[0281] Chain judgment = Prompt judgment |LLM|OutputParsers
[0282] According to the definition of similar cases in the "Guiding Opinions of the Supreme People's Court on Unifying the Application of Law and Strengthening the Retrieval of Similar Cases (Trial)", that is, similar cases referred to in this opinion refer to cases that are similar to the case to be decided in terms of basic facts, dispute focus, legal application issues, etc., and have been adjudicated and become effective by the people's court. Combining the importance of the discussion on reasoning and basis in the judgment gist of the decided cases in practice for similar case judgment, design the following similar case judgment prompt template:
[0283] Prompt judgment :
[0284] You are a legal expert. Similar cases refer to cases that are similar to the case to be decided in terms of basic facts, dispute focus, legal application, etc., and have been adjudicated and become effective by the people's court. Please determine whether the following {query} and {input} constitute similar cases.
[0285] Thought process:
[0286] Step 1: Comparison of Basic Facts:
[0287] Compare the main facts of {query} and {input}, including aspects such as the background of the case, the parties, the time and place of the event, the nature of the act, the damage result, etc.
[0288] Do {query} and {input} have similar factual situations? If so, please briefly explain.
[0289] Step 2: Comparison of Dispute Focuses:
[0290] Analyze whether the dispute focuses of {query} and {input} are the same.
[0291] Are the dispute points of {query} and the dispute points of {input} consistent in legal issues and practical operations? If the same, briefly explain the specific similar parts, such as contract interpretation, liability for breach of contract, liability for tort, etc.
[0292] Step 3: Comparison of Applicable Laws:
[0293] Judge whether {query} and {input} apply the same or similar legal provisions, especially in the application of legal articles, judicial interpretations, guiding cases and relevant regulations.
[0294] Are {query} and {input} consistent in the issue of applicable laws? If consistent, briefly explain the applicable legal articles and judicial interpretations.
[0295] Step 4: Applicability of the Key Points of Judgment:
[0296] Analyze whether the key points of judgment in {input} are applicable to {query}, with a focus on whether the legal reasoning, judgment basis and conclusion in the key points of judgment can provide a reference basis for resolving the dispute focus of {query}.
[0297] If applicable, briefly explain the core content of the key points of judgment and the reasons for applying to the dispute focus of {query}.
[0298] Step 5: Judgment:
[0299] Based on the above analysis, especially paying attention to whether the key points of judgment in {input} are applicable to the dispute focus of {query}, thus judge whether {query} and {input} constitute similar cases.
[0300] If it constitutes a similar case, please state the reasons; if it does not constitute a similar case, please state the main differences. Even if there are differences in some aspects, if the core issues (such as the nature of the legal relationship, the essence of the dispute focus, etc.) are similar, it may still constitute a similar case. However, if the core content of the ruling gist is not fully applicable to resolve the dispute focus of {query}, it cannot be determined as a similar case.
[0301] Example:
[0302] {query}: A certain company (the plaintiff) filed a lawsuit against another company (the defendant) due to a construction project contract dispute. The plaintiff claimed that the defendant failed to pay the project payment and performance bond as per the contract and requested the court to rule for payment.
[0303] {input}: A certain construction company (the plaintiff) filed a lawsuit against a certain construction company (the defendant) regarding a construction project contract dispute. The court ruled that the defendant failed to pay the project payment as per the contract and supported the plaintiff's lawsuit request. Ruling gist: In construction project contract dispute cases, the court should, based on the contract agreement and the performance of both parties, determine whether there is a breach of contract. If one party fails to pay the project payment as per the contract without justifiable reasons, it should be determined that it constitutes a breach of contract. The court should, in accordance with the relevant provisions of the "Contract Law", support the plaintiff's lawsuit request, order the defendant to pay the project payment, and bear the corresponding liability for breach of contract. In this case, the defendant, a certain construction company, failed to pay the project payment as per the contract and failed to fulfill its contractual obligations, constituting a breach of contract. The plaintiff, a certain construction company, legally filed a lawsuit with the court requesting payment of the project payment. After trial, the court held that the plaintiff's lawsuit request was legal, in line with the contract agreement and legal provisions, and should be supported. Therefore, the court ruled that the defendant should pay the project payment and perform it after the judgment takes effect.
[0304] Thought process:
[0305] (1) Comparison of basic facts:
[0306] Both cases involve construction project contract disputes and the issue of unpaid project payments, and the basic facts are similar.
[0307] (2) Comparison of dispute focuses:
[0308] The dispute focuses of the pending case and the decided case are both whether to fulfill the contract and pay the project payment, and the dispute focuses are the same.
[0309] (3) Comparison of legal applications:
[0310] Both cases apply the relevant provisions of the "Civil Code" and the "Construction Law", and the legal applications are the same.
[0311] (4) Applicability of ruling gist:
[0312] The gist of the judgment in {input} involves the judgment of breach of contract in construction project contract disputes and the general situation of legal consequences, which explains the judge's judgment reasoning process and judgment result, and is fully applicable to resolving disputes in {query}.
[0313] (5) Judgment: Based on the similarity of basic facts, disputes, and legal applications, especially the applicability of the judgment gist to resolving disputes in {query}, it is determined that {query} and {input} constitute similar cases.
[0314] Step6: Conduct similar case retrieval and judgment
[0315] Input the query case and candidate cases into Chain detail 、Chain focus 、Chain law 、Chain reason to obtain the scoring results. Input the scoring results into the trained LightGBM ranking model to obtain the ranking results of candidate cases. Input the query case and candidate cases with higher similarity into Chain judgment to obtain the similar case judgment result.
[0316] The present invention adopts a similar case retrieval method based on a large model, including a scoring module, a ranking model training and optimization module, and a testing module. With the help of prompt words, the large model scores the similarity between the query case and candidate cases from four dimensions: case situation, dispute focus, legal application, and judgment reasoning. Adjust the parameters of the large model according to the coefficient of variation and ICC3 index output by the large model to ensure the stability and consistency of the large model output. Train the machine learning model LightGBM with the labeled similar case retrieval data, use the LambdaRank algorithm to achieve candidate case ranking, test the trained LightGBM ranking model, and the tested LightGBM can perform similar case ranking and retrieval according to the similarity scores of the four dimensions. The test results are as follows:
[0317] In this embodiment, the present invention implements the foregoing technical solution in the following manner:
[0318] (1) Large model:
[0319] Adopt the Tongyi Qianwen - Turbo large model developed by Alibaba Cloud, and complete the natural language generation task by calling the large model through the API.
[0320] (2) Programming framework:
[0321] Use LangChain as the large model programming framework to construct prompt word templates and chains.
[0322] (3) Similar case retrieval data set:
[0323] The data set is from the Chinese Civil Case Retrieval Dataset (C3RD) (https: / / aistudio.baidu.com / datasetdetail / 205651).
[0324] All case texts in C3RD are from publicly available Chinese civil case judgments. The determination and annotation of similar cases are completed by Baidu's legal expert team. Each query case and the corresponding candidate case pool constitute a retrieval data item.
[0325] (4) Programming libraries used
[0326] The machine learning model for case ranking uses the LightGBM library, parameter optimization uses the Optuna library, and calculating the ranking metrics of test data uses the Pytrec_eval library.
[0327] The test results are as follows:
[0328] (1) Model comparison results
[0329] According to the foregoing technical solution, the ranking performance of the model is tested using 10-fold cross-validation. The mean and coefficient of variation of the cross-validation metric values obtained are shown in Table 1.
[0330] Table 1 Test metric values
[0331]
[0332]
[0333] In Table 1, except for the method proposed in the present invention, other methods are baseline models given by C3RD. It can be seen that the mean values of the various indicators of the method proposed in the present invention are better than those of the baseline models, and the coefficients of variation are all less than 0.05, indicating that the cross-validation technical results are relatively stable and the model has good generalization ability. The MRR indicator of the method proposed in the present invention is equal to 1, indicating that this method can accurately retrieve the most similar cases.
[0334] (2) Ablation experiment
[0335] To prove the effectiveness of the solution proposed in the present invention, an ablation experiment was conducted, which includes two comparison schemes.
[0336] Scheme 1: The few-shot and chain-of-thought are removed from the used prompt template, and the rest is the same as the solution proposed in the present invention. 10-fold cross-validation is used.
[0337] Solution 2: The adaptive control mechanism of the large model output is not adopted, the Temperatur of the large model is fixed at 0.75, the Top P is fixed at 0.8, and the other parts are the same as the solution proposed in the present invention. 10-fold cross validation is adopted.
[0338] Solution 3: Instead of implementing item-by-item scoring, prompt words are used to guide the large model to comprehensively evaluate the similarity between candidate cases and query cases from four aspects: case facts, focus of dispute, application of law, and judicial reasoning. The candidate cases are arranged in descending order according to the similarity, and the ranking results are given directly without the need to train a machine learning ranking model.
[0339] After 10-fold cross validation, the calculation indicators of the ablation experiment are shown in Table 2.
[0340] Table 2 Ablation experiment index values
[0341]
[0342] According to the ablation experiment results in Table 2, it can be seen that the method adopted by the present invention can improve the performance of similar case retrieval. The average value of the indicator in the 10-fold cross validation is better than that of other schemes, indicating that the present scheme has good generalization ability.
[0343] The following is a specific application example:
[0344] The following query cases are available:
[0345] The plaintiff Company A and the defendant Company B signed a "Construction Project Subcontract Contract" in 2018. According to the contract, the defendant Company B, as the general contractor of a commercial building project in Nanjing, subcontracted part of the civil engineering project to the plaintiff. The contract stipulates that the defendant should pay the project fee, performance bond and related overdue interest on time. However, the plaintiff claimed that the defendant failed to pay the above-mentioned amount in accordance with the contract, resulting in the plaintiff's failure to receive the relevant amount after repeated reminders, and believed that the defendant had breached the contract.
[0346] After many unsuccessful attempts at communication, the plaintiff filed a lawsuit with the Nanjing Intermediate People's Court of Jiangsu Province, requesting the court to order the defendant to pay the remaining project fee, overdue interest, and return the performance guarantee deposit and other amounts.
[0347] The defendant, Company B, raised an objection to the jurisdiction, arguing that the dispute in this case should involve the construction quality, project acceptance and completion of the construction project, and that the plaintiff did not perform the acceptance procedures. The defendant believed that the contract clearly stipulated that any disputes should be submitted to the jurisdiction of the court where the contract was signed, and the contract in question was signed in Hangzhou, Zhejiang Province, so the case should be heard by the Hangzhou Intermediate People's Court.
[0348] It is determined that this case is a construction project contract dispute. After searching the "Case Database of the People's Courts", 13 candidate cases of construction project contract disputes are found, and the case numbers are respectively:
[0349] 2024-08-2-115-001, 2023-07-2-115-002, 2023-07-2-115-00, 2023-07-2-115-008, 2024-07-2-115-001, 2024-08-2-115-002, 2021-18-2-115-001, 2023-07-2-115-005, 2024-07-2-115-003, 2023-16-2-115-007, 2024-01-2-115-003, 2023-16-2-115-002, 2023-07-2-115-007.
[0350] Using the similar case retrieval method proposed by the present invention to sort these 13 candidate cases, the result is:
[0351] 2024-01-2-115-003, 2023-07-2-115-001, 2024-08-2-115-002, 2023-16-2-115-002, 2024-08-2-115-001, 2023-07-2-115-002, 2023-07-2-115-008, 2024-07-2-115-001, 2021-18-2-115-001, 2023-07-2-115-005, 2024-07-2-115-003, 2023-16-2-115-007, 2023-07-2-115-007.
[0352] Using the similar case judgment chain Chain judgment Judge whether the first 3 candidate cases constitute similar cases. Under the guidance of the prompt words, the judgment results of the large model are as follows:
[0353] Table 3 Similar case judgment results
[0354] Candidate case number Whether it constitutes a similar case and the reasons 2024-01-2-115-003 Yes, both involve objections to jurisdiction 2023-07-2-115-001 Yes, both involve defaults in payment of project funds 2024-08-2-115-002 No, there are significant differences in specific facts and gist of the judgment
[0355] The large model can judge whether the candidate cases constitute similar cases according to the prompt words and give the judgment reasons.
[0356] Based on the same inventive concept, the present invention also provides a computer electronic device corresponding to the similar case retrieval method provided in the above embodiment. This device includes but is not limited to the following components:
[0357] (1) Memory:
[0358] Function: Used to store computer programs and related data. The computer program includes the code for implementing the case retrieval method of the present invention, as well as the related interface code for calling the large model API.
[0359] Specific implementation:
[0360] Cache: Used to temporarily store frequently accessed data to speed up data processing.
[0361] Solid State Drive (SSD): Used to store case data and computer program code for a long time to ensure fast reading and writing.
[0362] Database Management System (DBMS): Used to efficiently manage and query data and support complex retrieval requirements.
[0363] (2) Processor:
[0364] Function: Used to execute the computer program to implement the case retrieval enhancement method of the present invention. Specifically, the processor performs the following steps:
[0365] 1) Call the large model API through the Internet;
[0366] 2) Score the similarity between the query case and the candidate cases from four dimensions;
[0367] 3) Input the scores of the four dimensions into the machine learning ranking model to generate the final case retrieval result.
[0368] Specific implementation: Multi-core CPU architecture: Utilize the multi-core parallel processing ability to significantly improve data processing efficiency.
[0369] (3) Internet communication device:
[0370] Function: Used to communicate with other devices or services through the Internet. The Internet communication device includes but is not limited to the following types:
[0371] Network Interface Card (NIC): Used to connect to a local area network or a wide area network and support high-speed wired or wireless network communication.
[0372] Load Balancer: Used to distribute network traffic to ensure the high availability and stability of the system.
[0373] Security Gateway: Used to protect the security of data transmission and prevent unauthorized access.
[0374] The following gives an implementation example:
[0375] In actual application scenarios, the computer electronic device can be deployed locally or on a cloud server. The following are the specific implementation steps:
[0376] (1) Receive a query request: The user sends a query request to the computer electronic device through a client device (such as a personal computer).
[0377] (2) Invoke the large model API: The processor establishes a connection with an external large model API service through an Internet communication device (such as a network interface card NIC, router, etc.). Then, it sends the query case and candidate case information by invoking the API of the large model and receives the returned similarity score.
[0378] (3) Similarity scoring: The large model scores the similarity between the query case and the candidate case from four dimensions based on the received data. The scoring mechanism for each dimension is as follows:
[0379] Case similarity: Evaluate the similarity of the case situation between the query case and the candidate case.
[0380] Similarity of dispute focus: Evaluate the similarity of the dispute focus between the query case and the candidate case.
[0381] Similarity of legal application: Evaluate the applicability of the legal provisions cited in the candidate case to the query case.
[0382] Similarity of judgment reasoning: Evaluate the applicability of the judgment reasoning part in the candidate case to the query case.
[0383] During the scoring process, the processor issues instructions based on the coefficient of variation and ICC3 index output by the large model to adjust the large model parameters to ensure the stability and consistency of the large model output.
[0384] (4) Machine learning ranking model:
[0385] The processor inputs the scores of the four dimensions into a pre-trained machine learning ranking model to generate a case ranking result. This model has been trained with a large amount of labeled case data and can accurately rank cases.
[0386] (5) Invoke the large model API to perform a similar case judgment on the candidate cases ranked at the top.
[0387] (6) Return the result: The processor returns the final similar case retrieval result to the user through a display device.
[0388] In summary, the present invention introduces the chain of thought and few-shot prompting into case retrieval, constructs a prompting template in combination with legal professional knowledge, refines the case retrieval task, guides the large model to construct multi-dimensional vectors, adopts an adaptive mechanism to control the output of the large model, and uses the large model scoring vector as the input of the machine learning ranking model, so as to combine the semantic understanding ability of the large model with the machine learning ranking model. The present invention can provide users with a clear basis for case similarity, enhance the interpretability of retrieval results, and has high scalability.
Claims
1. A method for retrieving similar engineering cases based on a large language model and machine learning, characterized in that: The following steps are involved: Step 1: Construct a referee file data set Construct the data set D total , d i ={query:query_file,candidate:[candidate_file1,candidate_file2,…,candidate_file ni ],gt_idx:[ID1,ID2,…,ID ni ]}, where d i It is a dictionary, including three keys, namely qurey, candidate and gt_idx. Query represents the query case, candidate represents the candidate case, gt_idx is the real ranking index of the similarity between the candidate case and the query case, the value of query is the query case file query_file, the value of candidate is a list, the elements in the list are the adjudication files of multiple candidate cases, the adjudication files include case facts, case classification, legal provisions, adjudication reasoning process information, the same cause of action or the same type is the first step in similar case retrieval, and the type screening measures are adopted to make the candidate case have the same cause of action or type as the query case, the value of gt_idx is a list, the elements in the list are the actual sorting order of the candidate case index; Step 2: Use thought chain and few-shot technology to design prompt word templates, and use a large language model to build the following scoring chain: Case scoring chain: Chain detail =Prompt detail |LLM|OutputParsers Controversial focus: Rating chain: Chain focus =Prompt foucs |LLM|OutputParsers Legal Applicability Rating Chain: Chain law =Prompt law |LLM|OutputParsers Judge reasoning scoring chain: Chain reason =Prompt reason |LLM|OutputParsers Among them, Prompt detail Prompt word template for case similarity scoring, foucs Prompt word template for scoring the similarity of the focus of dispute, law Prompt word template for scoring the applicability of the legal provisions cited in the candidate case to the query case, reason It is the template for scoring the applicability of the candidate case reasoning part to the query case. LLM is the large model used. OutputParsers is the output parser used to convert the output of the large model into a specific format. Step 3: Score query cases and candidate cases For D total Each data d in i , enter the query and candidate values into Chain detail 、Chain focus 、Chain law 、Chain reason , and make a score; Step 4: Train and cross-validate the sorting machine learning model The LightGBM model is used to complete the sorting task. LightGBM is an efficient gradient boosting tree algorithm that can effectively reduce computational complexity and memory usage through a histogram-based decision tree learning method. LightGBM is used as a sorting model and the LambdaRank algorithm is used to sort candidate cases according to their scores. The TPE algorithm based on Bayesian optimization is used to optimize the parameters of the LightGBM model. This algorithm predicts the performance of hyperparameters by building a probability model, so as to select the most promising hyperparameter combination for evaluation in each iteration. This method can effectively balance the exploration of new hyperparameter areas and the use of known good hyperparameter areas. In order to use all the data for model training and evaluation, multiple rounds of cross-validation are used to improve the accuracy and reliability of model performance evaluation. Step 5: Build a similar case judgment chain Candidate cases are sorted according to their similarity to the query case. However, whether the candidate case with the highest similarity is a similar case requires further judgment. In order to use the large model to judge similar cases, the following similar case judgment chain is constructed: Chain judgment =Prompt judgment |LLM|OutputParsers Based on the definition that similar cases refer to cases that are similar to pending cases in terms of basic facts, dispute focus, and applicable law issues, and have been adjudicated by the People's Court, and combined with the importance of reasoning and basis in the adjudication essentials of decided cases in practice to similar case judgment, a similar case judgment prompt word template is designed; Step 6: Search and judge similar cases Input query cases and candidate cases into Chain detail 、Chain focus 、Chain law 、Chain reason , get the scoring results, input the scoring results into the trained LightGBM ranking model, get the ranking results of the candidate cases, and input the query cases and the candidate cases with high similarity into Chain judgment , and obtain similar case judgment results.
2. The engineering case retrieval method based on large language model and machine learning according to claim 1 is characterized in that: In the aforementioned Step 2, the prompt word template is as follows: Prompt detail : You are a legal expert. {input} is [basic information of the case]. You are required to evaluate the similarity between the case in {input} and {query}. The similarity evaluation must be based on the facts of the case. The specific steps are: Step 1: Extract the basic information of the cases in {query} and {input} respectively; Step 2: Determine whether the basic circumstances of the cases in {query} and {input} are consistent or similar. When making similarity judgments, we mainly consider whether the cases in {query} and {input} have strong similarities in the behavior of the parties involved. We analyze step by step and give a similarity score for the basic circumstances of the cases in {query} and {input}. The score range is 0 to 100, with a score of 100 being the most similar and a score of 0 being completely different. Prompt foucs : You are a legal expert. {input} is [basic information of the case]. You are required to conduct a similarity assessment based on the case in {input} and {query}. The similarity assessment is based on the focus of the dispute. The specific steps are as follows: Step 1: You need to summarize the core disputes in the cases in {query} and {input} respectively, analyze them step by step, and summarize the focus of the disputes in the following steps: (1) Clarify the background and basic facts of the case Determine the legal relationship involved in the case and clarify the factual disputes between the plaintiff and the defendant, including the signing and performance of the contract and the amount in dispute; (2) Analyze the litigation requests of both parties What does the plaintiff ask the court to rule on? What does the defendant object to? Focus on the plaintiff's claims and the defendant's defense; (3) Review the core issues of the dispute Among the claims of both parties, which issue is the core issue that the court needs to judge? Determine whether there is actual evidence to support a party's claims, especially key disputed points; (4) Analyze applicable law and contract terms Check the relevant legal clauses and the specific contents of the contract to determine whether the focus of the dispute involves the performance, interpretation or payment conditions of the contract terms; (5) Determine the focus of the dispute Based on the analysis of the above steps, determine the most critical dispute points in the case; Step 2: You need to determine whether the dispute focus of the case in {query} and {input} is the same or similar, and analyze step by step to give a similarity score between the dispute focus of the case in {query} and {input}. The score range is 0 to 100, with 100 being exactly the same and 0 being completely different. The scoring criteria are: 90-100 points: the dispute focus is substantially the same; 60-89 points: the dispute focus is similar in major aspects; 30-59 points: the dispute focus is partially related; 1-29 points: the dispute focus has a very low correlation; 0 points: completely unrelated; the score should be based on the similarity of the core content of the dispute focus of {query} and {input}, not just based on word similarity; Step 3: Generate a numerical similarity score; Step 4: Output a similarity score represented by only one number; Prompt law : You are a legal expert, {input} is [the judge's reasoning for the case and the legal provisions based on it], and you are required to conduct a similarity assessment based on the case in {input} and {query}. Similarity must be based on legal application, and the specific steps are as follows: Step 1: You need to extract the legal provisions and judicial interpretation provisions cited in the judge's judgment from {input}; Step 2: You find the original legal text from which the legal provisions were extracted in step 1; Step 3: You need to analyze whether the legal provisions and judicial interpretation provisions cited in the cases in {input} are applicable to {query}. Step by step analysis will give the similarity score between {query} and the legal provisions and judicial interpretation provisions cited in the cases in {input}. The similarity score range is 0 to 100. The score is calculated as: the number of legal provisions and judicial interpretation provisions cited in the cases in {input} that are applicable to {query} / the total number of legal provisions and judicial interpretation provisions cited in the cases in {input}; Step 4: Generate a numerical similarity score; Step 5: Output a similarity score represented by only one number; Prompt reason : You are a legal expert. {input} is [the judge's reasoning and legal text of the case]. You need to conduct an applicability assessment based on the reasoning of the judgment in {input} and {query}. The specific steps are as follows: Step 1: Extract the judge’s reasoning in the case of {input}; Step 2: Determine whether the judge's reasoning in the {input} case is applicable to {query}, analyze step by step, and give an applicability score. The similarity score range is 0 to 100, with a score of 100 being completely applicable and a score of 0 being completely inapplicable; Step 3: Generate a numerical similarity score; Step 4: Output a similarity score represented by just one number.
3. The engineering case retrieval method based on large language model and machine learning according to claim 1 is characterized in that: In the Step 3, the scoring method adopts an adaptive control mechanism, which specifically includes the following steps: (1) Initialize the number of calculation rounds n←0, (2) Establish the large model parameter calculation formula: Temperature=max(IV t ×f(n) t ,LL t ) Among them, IV t is the initial value of Temperatur, IV t =0.75; f(n) t is the adjustment function of Temperatur; the adjustment function used is LL t =0; Top P=max(IV p ×f(n) p ,LL p ) Among them, IV p is the initial value of Top P, IV p =0.8, f(n) p is the adjustment function of Top P. The adjustment function used in the present invention is LL p =0.05; Substitute Temperatur and Top P into the large model; (3) Input the query case and a single candidate case into the scoring chain and repeat the scoring three times to obtain the scoring matrix According to S, calculate the coefficient of variation by column and get CV detal 、CV focus 、CV law 、CV reason , the stability of the large model score is measured according to the coefficient of variation of each dimension score. The closer the coefficient of variation is to 0, the better the stability; The indicator ICC3 is calculated according to S as a consistency indicator for measuring the large model's scoring of each dimension. The intra-group correlation coefficient is a statistical indicator used to evaluate the consistency of the measurement results of the same group of objects by different evaluators or at different time points. ICC3 is used to evaluate the consistency of the scores of the same group of objects by a fixed evaluator. The value range of ICC3 is from 0 to 1. The closer the value is to 1, the higher the consistency. The present invention uses ICC3 to evaluate the consistency of multiple scoring results of the large model, which can effectively control the quality of the model output. If the ICC3 value is lower than the set threshold, it means that the consistency of the scoring results is poor, and the model parameters or scoring strategy need to be adjusted to improve the reliability of the output; (4) Determine the threshold of the coefficient of variation CV =0.15 and ICC3 threshold ICC3 =0.8, set the following rules: if(max(CV detal ,CV focus ,CV law ,CV reason )>threshold CV )∨(ICC3<threshold ICC3 ): Let n = n + 1, Repeat steps (2) to (4); else: Output the mean of the scores of each dimension as the final score; Get the corresponding score under the guidance of the prompt word detail 、score focus 、score law 、score reason , and then get the data set D' total , d i ={input:score,label:[ID1,ID2,…,ID ni ]},in, 4. The engineering case retrieval method based on large language model and machine learning according to claim 1 is characterized in that: In the Step 4, the cross-validation method uses 10-fold cross-validation, including the following steps: (1) Dataset division: The scored data set D' total Divide into 10 subsets of equal size D1, D2, ..., D 10 , that is, the size of each subset is N / 10, where N is the total number of samples in the data set; (2) Model training and verification: In each round i (i∈[1,10]), let the test set D' test =D i , the rest of the data is D -i =D' total \D i In order to tune the model, D -i Further divided into training set D' in the ratio of 80% / 20% train and verification set D' validation ; Using the training set D' train Train the LightGBM model and use it on the validation set D' validation Perform hyperparameter tuning on Put the trained model in the test set D' test Predictions are made on the , and ranking evaluation indicators are calculated. The evaluation indicators include: P@5, P@10, MAP@10, NDCG@10, NDCG@20, NDCG@30, and MRR. The larger the indicator, the better. (3) Repeat the process: Repeat step (2), selecting a different subset as the test set each time to ensure that each subset has the opportunity to be used as a test set to evaluate model performance; (4) Comprehensive evaluation: After completing all 10 rounds of calculations, the mean and coefficient of variation of each evaluation indicator are calculated to measure the stability and generalization ability of the model on different test sets. The final evaluation indicator is the average value of the 10-fold evaluation indicator.
5. The engineering case retrieval method based on large language model and machine learning according to claim 1, characterized in that: In the aforementioned Step 5, the template for similar case judgment prompt words is as follows: Prompt judgment : You are a legal expert. Similar cases refer to cases that are similar to pending cases in terms of basic facts, dispute focus, and applicable law, and have been adjudicated by the People's Court. Please determine whether the following {query} and {input} constitute similar cases; Thought Process: Step 1: Basic facts comparison: Compare the main facts of {query} and {input}, including the background of the case, the parties involved, the time and place of the incident, the nature of the behavior, and the damage results; Do {query} and {input} have similar factual contexts? If so, please briefly describe; Step 2: Contrast of controversial issues: Analyze whether the dispute focus of {query} and {input} is the same; Are the dispute points of {query} and {input} consistent in terms of legal issues and practical operations? If they are the same, please briefly describe the specific similarities, such as contract interpretation, liability for breach of contract, and liability for tort; Step 3: Comparison of applicable laws: Determine whether {query} and {input} are subject to the same or similar legal provisions, especially in terms of the application of legal provisions, judicial interpretations, guiding cases and related regulations; Are {query} and {input} consistent in terms of applicable law? If so, briefly describe the applicable legal provisions and judicial interpretations; Step 4: Applicability of the Judgment Points: Analyze whether the judgment summary in {input} is applicable to {query}, focusing on whether the legal reasoning, judgment basis and conclusion in the judgment summary can provide reference for resolving the dispute focus of {query}; If applicable, briefly explain the core content of the ruling and the reasons for the dispute applicable to {query}; Step 5: Judgement: Based on the above analysis, we should pay special attention to whether the judgment in {input} is applicable to the dispute focus of {query}, so as to judge whether {query} and {input} constitute similar cases; If it is a similar case, please explain the reason; if it is not a similar case, please explain the main differences. Even if there are differences in some aspects, if the core issues are similar, it may still constitute a similar case. However, if the core content of the adjudication summary is not completely applicable to resolving the focus of the dispute of {query}, it cannot be judged as a similar case.
6. A computer electronic device corresponding to the engineering case retrieval method based on a large language model and machine learning as described in claim 1, characterized in that: The device includes the following components: Memory: used to store computer programs and related data, wherein the computer programs include codes for implementing the similar case retrieval method of the present invention and related interface codes for calling the large model API; Processor: used to execute the computer program to implement the similar case retrieval method of the present invention; Internet communications devices: used to communicate with other devices or services over the Internet.
7. A computer electronic device corresponding to the engineering case retrieval method based on a large language model and machine learning according to claim 6, characterized in that: The processor performs the following steps: 1) Call the big model API through the Internet; 2) Score the similarity between the query case and the candidate case from four dimensions; 3) Input the scores of the four dimensions into the machine learning ranking model to generate the final similar case retrieval results.
8. A computer electronic device corresponding to the engineering case retrieval method based on a large language model and machine learning according to claim 6, characterized in that: The Internet communication equipment includes the following types: Network interface card: used to connect to a local area network or wide area network, supporting high-speed wired or wireless network communications; Load balancer: used to distribute network traffic and ensure high availability and stability of the system; Security Gateway: Used to protect the security of data transmission and prevent unauthorized access.
Citation Information
Cited By
Large model tool calling method, model training data processing method and intelligent vehicle
CN121166237A
Large model tool calling method, model training data processing method and intelligent vehicle
CN121166237B
Class case retrieval method, device and equipment based on large language model and vector retrieval
CN122262205A