Case Retrieval Enhanced Claim Generation and Evaluation Method, System, Medium and Device
By constructing a legal data set and corpus covering diversity, and using a case-like search enhanced method, the problem of insufficient data set scope and diversity in the existing technology is solved, and high-quality appeal generation and evaluation are achieved.
Patent Information
- Application Number
- CN202410836011.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-26
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-06-26
AI Technical Summary
The insufficient scope and diversity of existing legal data sets lead to poor appeal generation and lack of fine-grained evaluation criteria.
A legal data set and corpus covering 100 cases is constructed, fine-grained standards of authenticity, clarity and consistency are introduced, and a similar case search enhancement method is adopted to generate and evaluate litigation by obtaining civil judgment documents, screening relevant paragraphs, conducting similar case searches and learning in small samples.
It effectively improves the quality of generating petitions, provides more detailed evaluation standards, and improves the ability to generate and evaluate petitions.
Smart Images

Figure CN118861264B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of legal artificial intelligence, and particularly relates to a method, system, medium and device for claim generation and evaluation enhanced by similar case retrieval. Background Art
[0002] In the field of legal artificial intelligence, traditional tasks aim to assist individuals in completing various legal tasks, including legal judgment prediction, court trial view generation, similar case matching, legal language understanding, and legal question answering. However, existing research mainly focuses on court trials and judge assistance, with less attention paid to pre-trial situations or non-professional needs (such as claim generation). A claim refers to the claims of the plaintiff in a case and plays an important role in the judicial process as they have a significant impact on the outcome of the dispute. However, the complex legal language and the financial burden of hiring a lawyer often put non-professionals at a considerable disadvantage. Despite the government's efforts to expand the coverage of legal services by training local legal practitioners, there are still many challenges in some areas where legal professionals are scarce. Proposing and researching the problem of claim generation based on case facts can promote the rule of law and increase the accessibility of legal support.
[0003] There are some legal data sets in the research of legal artificial intelligence in the field of civil litigation. However, existing work faces challenges such as insufficient data coverage, lack of diverse and detailed evaluation metrics, and poor claim generation effects. Constructing a legal data set and corpus covering 100 case types can address the limitations in the scope and diversity of existing legal data sets. Introducing fine-grained criteria including authenticity, clarity, and consistency can better evaluate the performance of the model in claim generation. A claim generation method enhanced by similar cases can effectively improve the quality of generated claims by leveraging previous similar cases. Summary of the Invention
[0004] The purpose of the present invention is to overcome the deficiencies of existing data and technologies, and provide a method, system, medium and device for claim generation and evaluation enhanced by similar case retrieval.
[0005] To achieve the above-mentioned invention purpose, the present invention specifically adopts the following technical solutions:
[0006] In the first aspect, the present invention provides a method for claim generation and evaluation enhanced by similar case retrieval, including the following steps:
[0007] S1. Obtain civil judgment documents. Each civil judgment document contains the unstructured content of the entire document and additional information. Screen the obtained civil judgment documents and only retain the first-instance civil judgment documents. Use regular expressions to find the text parts containing keywords in the unstructured content of the first-instance civil judgment documents, segment the text parts before and after the keywords, divide the unstructured content of the first-instance civil judgment documents into specific segments, and only retain the paragraphs related to the claim generation task in the specific segments to obtain civil case corpus data containing only the plaintiff's facts and the plaintiff's claims. Screen out the civil case corpus data in which the court supports the plaintiff's claims from the civil case corpus data, and form a civil case corpus database with the screened civil case corpus data; among them, the additional information includes the case type, title, start time, and end time; the specific segment includes the introduction, plaintiff's facts, plaintiff's claims, defendant's defense, court findings, and judgment.
[0008] S2. Input the plaintiff's facts of the case for which the claim is to be generated and the civil case corpus database into the retrieval model together. The retrieval model conducts similar case retrieval and outputs similar cases related to the case for which the claim is to be generated.
[0009] S3. Use the few-shot learning method to input the plaintiff's facts of the case for which the claim is to be generated, the plaintiff's facts of the similar cases, and the plaintiff's claims of the similar cases into the large language model together, and at the same time input the first prompt word into the large language model. The large language model generates the plaintiff's claims.
[0010] S4. Obtain the true plaintiff's claims of the case for which the claim is to be generated from the civil case corpus database. Input the true plaintiff's claims and the plaintiff's claims generated by the large language model into the large language model, and input the second prompt word into the large language model. The large language model conducts claim understanding and evaluates the generated plaintiff's claims on different evaluation indicators, and outputs the scoring results of the plaintiff's claims on each evaluation indicator.
[0011] Based on the above solutions, each step can be implemented in the following preferred specific ways.
[0012] As a preference of the first aspect above, the keywords used in the segmentation process include: filing a lawsuit with this court, facts and reasons, defense, the court holds that, and judgment as follows.
[0013] As a preference of the first aspect above, retain the top 100 common civil case types in the first-instance civil judgment documents as all the case types of the civil case corpus database.
[0014] As a preference of the first aspect above, the output scoring result is a positive integer between 1 and 5, where 1 represents the lowest quality and 5 represents the highest quality.
[0015] Second aspect, the present invention provides a case - similar retrieval - enhanced claim generation and evaluation system, including:
[0016] A data acquisition module, configured to acquire civil judgment documents. Each civil judgment document contains the unstructured content of the entire document and additional information. Screen the acquired civil judgment documents, and only retain the first - instance civil judgment documents; Use regular expressions to find the text parts containing keywords in the unstructured content of the first - instance civil judgment documents, segment the text parts before and after the keywords, divide the unstructured content of the first - instance civil judgment documents into specific segments, and only retain the paragraphs related to the claim generation task in the specific segments to obtain civil case corpus data containing only the plaintiff's facts and the plaintiff's claims; Screen out the civil case corpus data in which the court supports the plaintiff's claims from the civil case corpus data, and the screened civil case corpus data constitutes a civil case corpus database; Among them, the additional information includes the case cause, title, start time, and end time; The specific segment includes an introduction, plaintiff's facts, plaintiff's claims, defendant's defense, court findings, and judgment.
[0017] A retrieval module, configured to jointly input the plaintiff's facts of the case for which a claim is to be generated and the civil case corpus database into a retrieval model, and the retrieval model conducts similar case retrieval and outputs similar cases related to the case for which a claim is to be generated.
[0018] A claim generation module, configured to use the few - shot learning method to jointly input the plaintiff's facts of the case for which a claim is to be generated, the plaintiff's facts of the similar cases, and the plaintiff's claims of the similar cases into a large - language model, and at the same time input a first prompt word into the large - language model, and the large - language model generates the plaintiff's claims.
[0019] An evaluation module, configured to obtain the true plaintiff's claims of the case for which a claim is to be generated from the civil case corpus database, input the true plaintiff's claims and the plaintiff's claims generated by the large - language model into the large - language model, and input a second prompt word into the large - language model. The large - language model conducts claim understanding and evaluates the generated plaintiff's claims on different evaluation metrics, and outputs the scoring results of the plaintiff's claims on each evaluation metric.
[0020] Third aspect, the present invention provides a computer - readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the case - similar retrieval - enhanced claim generation and evaluation method as described in any one of the above - mentioned first - aspect solutions.
[0021] Fourth aspect, the present invention provides a computer electronic device, characterized by including a memory and a processor;
[0022] The memory is used to store a computer program;
[0023] The processor is used to implement the case retrieval enhanced claim generation and evaluation method as described in any solution of the first aspect above when executing the computer program.
[0024] In a fifth aspect, the present invention provides a computer program product, including a computer program / instructions, characterized in that when the computer program / instructions are executed by a processor, they can implement the case retrieval enhanced claim generation and evaluation method as described in any solution of the first aspect above.
[0025] Compared with the prior art, the present invention has the following beneficial effects:
[0026] The present invention proposes a case retrieval enhanced claim generation and evaluation method, system, medium and device for the claim generation requirements in the legal scenario. From the perspective of practical application, this method effectively addresses the limitations in the scope and diversity of existing legal data sets, and uses previous similar cases to effectively improve the quality of generated claims. Introducing fine-grained criteria including authenticity, clarity and consistency better evaluates the performance of the model in claim generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 is a flowchart of the method of the present invention;
[0028] Figure 2 is a system block diagram of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] In order to make the above objects, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below. The technical features in the various embodiments of the present invention can be combined correspondingly without conflict.
[0030] As Figure 1 shown, in a preferred implementation manner of the present invention, the above-mentioned case retrieval enhanced claim generation and evaluation method includes the following steps S1 to S4. The specific implementation processes thereof will be described in detail below.
[0031] S1. Obtain civil judgment documents. Each civil judgment document contains the unstructured content of the entire document and additional information. Screen the obtained civil judgment documents and only retain the first-instance civil judgment documents. Use regular expressions to find the text parts containing keywords in the unstructured content of the first-instance civil judgment documents, segment the text parts before and after the keywords, divide the unstructured content of the first-instance civil judgment documents into specific segments, and only retain the paragraphs related to the claim generation task in the specific segments to obtain civil case corpus data containing only the plaintiff's facts and the plaintiff's claims. Screen out the civil case corpus data in which the court supports the plaintiff's claims from the civil case corpus data, and the screened civil case corpus data constitutes the civil case corpus database. Among them, the additional information includes the cause of action, title, start time, and end time. The specific segment includes introduction, plaintiff's facts, plaintiff's claims, defendant's defense, court findings, and judgment.
[0032] It should be noted that step S1 of the present invention constructs a civil case corpus database D (i.e., Figure 1 the case corpus in
[0033] S11: Obtain publicly available civil judgment documents from the China Judgments Online ( http: / / wenshu.court.gov.cn / ). Each civil judgment document contains the unstructured content of the entire document and additional information, and the additional information includes the cause of action, title, start time, and end time.
[0034] S12: Original data processing:
[0035] 1) To maintain material consistency, only retain the first-instance civil judgment documents among all the obtained civil judgment documents and filter out the documents that cannot be made public for certain reasons.
[0036] 2) To ensure the integrity of the content, segment the unstructured content of the first-instance civil judgment documents and only retain the paragraphs related to the claim generation task. To evaluate the integrity, use the continuous inclusion criterion, and the keywords include: "filed a lawsuit with this court", "facts and reasons", "defended", "the court holds that", and "judgment is as follows". Subsequently, the unstructured content of the first-instance civil judgment documents is divided into specific segments, including introduction, plaintiff's facts, plaintiff's claims, defendant's defense, court findings, and judgment, and finally obtain civil case corpus data containing only the plaintiff's facts ( Figure 1 the facts and reasons in
[0037] 3) To obtain reasonable claims, screen out the civil case corpus data in which the court supports the plaintiff's claims from the civil case corpus data, and these civil case corpus data form the civil case corpus database D.
[0038] In addition, it should be noted that in step S1 of the present invention, there are 134 civil case types in the first-instance civil judgment documents. Due to insufficient data in some categories, the first 100 most common civil case types are retained as all case types in the civil case corpus database.
[0039] S2. Input the plaintiff's facts of the case to be generated and the civil case corpus database into the retrieval model together. The retrieval model conducts a similar case search and outputs similar cases related to the case to be generated.
[0040] It should be noted that in step S2 of the present invention, the BM25 retrieval model method is used. Based on the plaintiff's facts f of the case to be generated g to retrieve the set of similar cases S = BM25(D, f g ). The set of similar cases S = {s1, s2,..., s m} contains m similar cases. Among them, s1, s2,..., s m represent the 1st, 2nd,..., mth similar cases respectively, and m is the number of retrieved similar cases.
[0041] In the present invention, the number of similar cases m is a positive integer and can be determined by those skilled in the art according to actual needs. In this embodiment, m can be 1 or 2 or 3.
[0042] S3. Use the few-shot learning method to input the plaintiff's facts of the case to be generated, the plaintiff's facts of the similar cases, and the plaintiff's claims of the similar cases into the large language model together, and at the same time input the first prompt into the large language model. The large language model generates the plaintiff's claim.
[0043] It should be noted that in step S3 of the present invention, the few-shot learning method is used. Input the plaintiff's facts f of the case to be generated g and the plaintiff's facts and claims of the similar cases in set S into the large language model, and at the same time input the first prompt to generate the plaintiff's claim C = LLM(f g , S) of the case to be generated. Among them, LLM represents the large language model.
[0044] S4. Obtain the true plaintiff's claim of the case to be generated from the civil case corpus database. Input the true plaintiff's claim and the plaintiff's claim generated by the large language model into the large language model, and input the second prompt into the large language model. The large language model conducts claim understanding and evaluates the generated plaintiff's claim on different evaluation indicators, and outputs the scoring results of the plaintiff's claim on each evaluation indicator.
[0045] It should be noted that in step S4 of the present invention, the plaintiff's claims generated through evaluation on different evaluation metrics of the large language model are scored. In this embodiment, a total of three evaluation metrics are selected, namely factuality, clarity, and consistency. The final output score is a positive integer between 1 and 5, where 1 represents the lowest quality and 5 represents the highest quality. In this embodiment, as Figure 1 shown, the full score for factuality, clarity, and consistency is 5 points each, and the scores obtained by the plaintiff's claims generated by the large language model on each evaluation metric are all 2 points.
[0046] This shows that the present invention can generate and evaluate the claims for civil litigation, obtain the claims corresponding to the facts and their scores, thereby enhancing the generation and evaluation capabilities for application scenarios.
[0047] Next, the present invention will generate and score claims through a specific civil litigation case to demonstrate the application effect of the case retrieval enhanced claim generation and evaluation method described in S1 - S4 in the above embodiment on a specific dataset, so as to facilitate the understanding of the essence of the present invention.
[0048] Embodiment
[0049] The specific steps of the case retrieval enhanced claim generation and evaluation method are as described in S1 - S4, which will not be elaborated here, and mainly the specific parameters and technical effects will be shown.
[0050] This embodiment verifies the effect of the method of the present invention in claim generation and scoring for civil litigation. Tables 1 and 2 respectively show the claim generation and scoring results for two different groups of samples.
[0051] Table 1 Claim Generation and Scoring Results for the First Group of Samples
[0052]
[0053]
[0054] Table 2 Claim Generation and Scoring Results for the Second Group of Samples
[0055]
[0056] In Tables 1 and 2, some comparison methods are selected in this embodiment to compare with the method of the present invention. These comparison methods include: the Bidirectional Auto-Regressive Transformer model BART, which is from the prior art literature: Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, Luke Zettlemoyer: BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. ACL 2020: 7871-7880; the Text-to-Text Transfer Transformer model T5, which is from the prior art literature: Raffel C, Shazeer N, Roberts A, et al. Exploring the limits of transfer learning with a unified text-to-text transformer[J]. Journal of machine learning research, 2020, 21(140): 1-67.
[0057] In addition, in Tables 1 and 2 of this embodiment, three large language models are also selected, and the plaintiff's claims are directly generated by these large language models. Specifically, GPT3.5-0shot, GPT3.5-1shot, and GPT3.5-2shot all represent methods of applying the generative pre-trained transformer model GPT3.5. Among them, GPT3.5-0shot is a comparative method that directly uses the generative pre-trained transformer model to generate the plaintiff's claims; GPT3.5-1shot and GPT3.5-2shot are the methods of the present invention. GPT3.5-1shot means setting 1 similar case when generating the plaintiff's claims, and GPT3.5-2shot means setting 2 similar cases when generating the plaintiff's claims. Gemini-0shot, Gemini-1shot, and Gemini-2shot all represent methods of applying the multimodal large language model Gemini. Among them, Gemini-0shot is a comparative method that directly uses the multimodal large language model to generate the plaintiff's claims; Gemini-1shot and Gemini-2shot are the methods of the present invention. Gemini-1shot means setting 1 similar case when generating the plaintiff's claims, and Gemini-2shot means setting 2 similar cases when generating the plaintiff's claims. LLaMA2-0shot, LLaMA2-1shot, and LLaMA2-2shot represent methods of applying the large language model LLaMA2. Among them, LLaMA2-0shot is a comparative method that directly uses the multimodal large language model to generate the plaintiff's claims; LLaMA2-1shot and LLaMA2-2shot are the methods of the present invention. LLaMA2-1shot means setting 1 similar case when generating the plaintiff's claims, and LLaMA2-2shot means setting 2 similar cases when generating the plaintiff's claims.
[0058] It should be noted that the implementation methods of the above three large language models all belong to the prior art. The generative pre-trained transformer model GPT3.5 is from the prior art literature: Brown T, Mann B, Ryder N, et al. Language models are few-shot learners[J]. Advances in neural information processing systems, 2020, 33: 1877-1901; the multimodal large language model Gemini is from the prior art literature: Team G, Anil R, Borgeaud S, et al. Gemini: a family of highly capable multimodal models[J]. arXiv preprint arXiv:2312.11805, 2023; the large language model Llama 2 is from the prior art literature: Touvron H, Martin L, Stone K, et al. Llama 2: Open foundation and fine-tuned chat models[J]. arXiv preprint arXiv:2307.09288, 2023.
[0059] As can be seen from the above results, on each baseline model, the method of the present invention can improve the factuality, clarity and consistency of the claims generated by prediction.
[0060] It should be noted that the method for claim generation and evaluation with enhanced case retrieval in the above embodiments can essentially be executed by a computer program or module. Therefore, similarly, based on the same inventive concept, another preferred embodiment of the present invention also provides a system for claim generation and evaluation with enhanced case retrieval corresponding to the method for claim generation and evaluation with enhanced case retrieval provided in the above embodiments, as Figure 2 shown, which includes:
[0061] A data acquisition module, which is used to acquire civil judgment documents. Each civil judgment document contains the unstructured content of the entire document and additional information. The acquired civil judgment documents are screened to retain only the first-instance civil judgment documents. Regular expressions are used to find the text parts containing keywords in the unstructured content of the first-instance civil judgment documents, and the text parts before and after the keywords are segmented. The unstructured content of the first-instance civil judgment documents is divided into specific segments, and only the paragraphs related to the claim generation task are retained in the specific segments to obtain civil case corpus data containing only the plaintiff's facts and the plaintiff's claims. From the civil case corpus data, the civil case corpus data in which the court supports the plaintiff's claims is screened, and the civil case corpus database is composed of the screened civil case corpus data. Among them, the additional information includes the cause of action, title, start time, and end time. The specific segment includes an introduction, plaintiff's facts, plaintiff's claims, defendant's defense, court findings, and judgment.
[0062] A retrieval module, which is used to jointly input the plaintiff's facts of the case for which a claim is to be generated and the civil case corpus database into a retrieval model. The retrieval model conducts a similar case retrieval and outputs similar cases related to the case for which a claim is to be generated.
[0063] A claim generation module, which is used to jointly input the plaintiff's facts of the case for which a claim is to be generated, the plaintiff's facts of the similar cases, and the plaintiff's claims of the similar cases into a large language model using the few-shot learning method. At the same time, a first prompt word is input into the large language model, and the large language model generates the plaintiff's claims.
[0064] An evaluation module, which is used to obtain the true plaintiff's claims of the case for which a claim is to be generated from the civil case corpus database, input the true plaintiff's claims and the plaintiff's claims generated by the large language model into the large language model, and input a second prompt word into the large language model. The large language model conducts claim understanding and evaluates the generated plaintiff's claims on different evaluation indicators, and outputs the scoring results of the plaintiff's claims on each evaluation indicator.
[0065] It can be understood that the method for claim generation and evaluation enhanced by similar case retrieval described in S1 to S4 above can essentially be implemented through a computer program. Therefore, based on the same inventive concept, in another preferred embodiment of the present invention, a computer program product corresponding to the method for claim generation and evaluation enhanced by similar case retrieval provided in the above embodiment is also provided. It includes computer programs / instructions, and when the computer programs / instructions are executed by a processor, they can implement the method for claim generation and evaluation enhanced by similar case retrieval as described in the above embodiment.
[0066] Similarly, based on the same inventive concept, in another preferred embodiment of the present invention, a computer electronic device corresponding to the case retrieval enhanced claim generation and evaluation method provided in the above embodiment is further provided, which includes a memory and a processor;
[0067] The memory is used to store computer programs;
[0068] The processor is used to implement the case retrieval enhanced claim generation and evaluation method in the above embodiment when executing the computer program.
[0069] In addition, when the logical instructions in the above memory are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0070] Therefore, based on the same inventive concept, in another preferred embodiment of the present invention, a computer-readable storage medium corresponding to the case retrieval enhanced claim generation and evaluation method provided in the above embodiment is further provided. The storage medium stores a computer program, and when the computer program is executed by the processor, it can implement the case retrieval enhanced claim generation and evaluation method in the above embodiment.
[0071] It can be understood that the above storage medium may include a random access memory (RAM), and may also include a non-volatile memory (NVM), such as at least one disk memory. At the same time, the storage medium may also be various media such as a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc that can store program codes.
[0072] It can be understood that the above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0073] In addition, it should be noted that those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the above-described system can refer to the corresponding process in the foregoing method embodiments, and will not be elaborated herein. In the embodiments provided in the present application, the division of steps or modules in the system and method is only a logical function division, and there may be other division methods in actual implementation. For example, multiple modules or steps can be combined or integrated together, and a module or step can also be split.
[0074] The above-described embodiments are only a preferred solution of the present invention, but they are not intended to limit the present invention. Those of ordinary skill in the relevant art can still make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all technical solutions obtained by adopting equivalent replacement or equivalent transformation fall within the protection scope of the present invention.
Claims
1. A method for generating and evaluating claims with enhanced similar case retrieval, characterized in that: The following steps are involved: S1. Obtain civil judgment documents. Each civil judgment document contains the unstructured content of the entire document and additional information. The obtained civil judgment documents are screened and only the first-instance civil judgment documents are retained. Regular expressions are used to find the text part containing keywords in the unstructured content of the first-instance civil judgment documents, and the text part before and after the keywords is segmented. The unstructured content of the first-instance civil judgment documents is divided into specific segments. Only the paragraphs related to the claim generation task are retained in the specific segments to obtain civil case corpus data containing only the plaintiff's facts and the plaintiff's claims. The civil case corpus data in which the court supports the plaintiff's claims are screened out from the civil case corpus data, and the screened civil case corpus data constitute a civil case corpus database; wherein the additional information includes the cause of action, title, start time and end time; the specific segments include introduction, plaintiff's facts, plaintiff's claims, defendant's defense, court findings and judgments; S2. The plaintiff facts of the case to be generated and the civil case corpus database are input into the retrieval model, and the retrieval model searches for similar cases and outputs similar cases related to the case to be generated; S3, using a few-shot learning method to input the plaintiff's facts of the case to be generated, the plaintiff's facts of similar cases, and the plaintiff's claims of similar cases into a large language model, and at the same time input the first prompt word into the large language model, and the large language model generates the plaintiff's claim; S4. Obtain the real plaintiff's claim of the case to be generated from the civil case corpus database, input the real plaintiff's claim and the plaintiff's claim generated by the large language model into the large language model, and input the second prompt word into the large language model. The large language model understands the claim and evaluates the generated plaintiff's claim based on different evaluation indicators, and outputs the scoring results of the plaintiff's claim on each evaluation indicator.
2. The method for generating and evaluating claims with enhanced similar case retrieval as claimed in claim 1, characterized in that: The key words used in the paragraphing process include: bringing a lawsuit to this court, facts and reasons, defense, this court holds and the judgment is as follows.
3. The method for generating and evaluating claims with enhanced similar case retrieval as claimed in claim 1, characterized in that: The top 100 common civil causes of action in the first instance civil judgment documents are retained as all the causes of action in the civil case corpus database.
4. The method for generating and evaluating claims with enhanced similar case retrieval as claimed in claim 1, characterized in that: The output score is a positive integer between 1 and 5, where 1 represents the lowest quality and 5 represents the highest quality.
5. A system for generating and evaluating claims with enhanced similar case retrieval, characterized in that: include: The data acquisition module is used to obtain civil judgment documents. Each civil judgment document contains the unstructured content of the entire document and additional information. The obtained civil judgment documents are screened and only the first-instance civil judgment documents are retained; regular expressions are used to find the text part containing keywords in the unstructured content of the first-instance civil judgment document, and the text part before and after the keywords is segmented, and the unstructured content of the first-instance civil judgment document is divided into specific segments. In the specific segments, only the paragraphs related to the claim generation task are retained to obtain civil case corpus data containing only the plaintiff's facts and the plaintiff's claims; civil case corpus data in which the court supports the plaintiff's claims are screened from the civil case corpus data, and the screened civil case corpus data constitute a civil case corpus database; wherein the additional information includes the cause of action, title, start time and end time; the specific segments include introduction, plaintiff's facts, plaintiff's claims, defendant's defense, court findings and judgments; A retrieval module is used to input the plaintiff facts of the case to be generated and the civil case corpus database into the retrieval model, and the retrieval model searches for similar cases and outputs similar cases related to the case to be generated; A claim generation module, used to input the plaintiff facts of the case to be generated, the plaintiff facts of similar cases and the plaintiff claims of similar cases into a large language model using a few-shot learning method, and input the first prompt word into the large language model, so that the large language model generates the plaintiff claim; The evaluation module is used to obtain the real plaintiff's claim of the case to be generated from the civil case corpus database, input the real plaintiff's claim and the plaintiff's claim generated by the large language model into the large language model, and input the second prompt word into the large language model. The large language model understands the claim and evaluates the generated plaintiff's claim based on different evaluation indicators, and outputs the scoring results of the plaintiff's claim on each evaluation indicator.
6. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method for generating and evaluating claims with enhanced similar case retrieval as described in any one of claims 1 to 4 is implemented.
7. A computer electronic device, characterized in that: including memory and processor; The memory is used to store computer programs; The processor is used to implement the claim generation and evaluation method enhanced by similar case retrieval as described in any one of claims 1 to 4 when executing the computer program.
8. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, it can implement the claim generation and evaluation method enhanced by similar case retrieval as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Course trial query generation method and device for generating Seq2Seq model based on pointer, and medium
CN112417155A
Text retrieval method, device and equipment and computer readable storage medium
CN117725160A