A legal fact completion method for non-expert users

CN122817431APending Publication Date: 2026-09-25ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610914205.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-24
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]本发明要解决的技术问题是:现有LegalAI系统在面对非专家用户模糊查询时性能不佳;现有的主动提问方法多采用开放式问题,增加了非专家用户的认知负担且难以获取精确法律事实;简单问卷缺乏法律实质性;以及在缺乏人工标注的“查询-问卷”数据的情况下,如何训练出能够生成兼具法律专业性和用户友好性的高质量事实补全问卷的模型

Benefits of technology

[0035]第一,本发明提出了一种无需人工标注数据的无标签迭代训练范式,通过辅助法律专家模型间接评估和跨案例生成策略,有效解决了法律领域“查询-问卷”数据稀缺的问题,实现了提问模型的自动优化。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817431A_ABST
    Figure CN122817431A_ABST
Patent Text Reader

Abstract

The application discloses a legal fact completion method for non-expert users. In the training stage, the application takes legal documents as data sources, simulates the situation of user query missing facts, evaluates the completion effect with the help of auxiliary expert models, automatically generates training corpus through the unlabeled iteration paradigm and fine-tunes the questioning model; during reasoning, the conditional logic type law is converted into a multiple-choice question, similar cases are retrieved based on user queries, the frequency of law citation is counted, and the multiple-choice question corresponding to the high-frequency law is used as the context; the fine-tuned model dynamically generates a structured questionnaire in combination with the context, guides the user to supplement key facts, and provides the downstream expert model to generate responses. The method does not require manual annotation, bridges the semantic gap through the "case-law-multiple-choice question" cascade retrieval, reduces the cognitive burden of users through structured questioning, and improves the expression support while maintaining ease of understanding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence technology, and in particular relates to a method for completing legal facts for non-expert users. Background Technology

[0002] With the widespread application of artificial intelligence in the legal field, Legal AI systems, such as those for legal consultation and Q&A, legal judgment prediction, and case retrieval, have made significant progress. However, in practical applications, non-expert users often struggle to accurately express their legal needs, and their queries are typically vague or incomplete, lacking crucial legal factual details (such as the absence of a labor contract or missing pay records). This information gap prevents Legal AI systems from providing reliable or accurate assistance, thereby undermining user trust in the system.

[0003] While existing technologies, such as large language models, can clarify user intent through proactive questioning, they exhibit significant limitations in legal scenarios. On one hand, open-ended clarifying questions place an excessive cognitive load on non-expert users, who may struggle to understand the question's meaning or provide precise legal details. On the other hand, simple structured questionnaires lacking legal basis result in "yes / no" questions that contribute little to downstream reasoning. Furthermore, the current lack of well-labeled query-questionnaire datasets and the extremely high cost of manual annotation make it difficult to train questioning models that generate both legally meaningful questions and easily understandable to non-experts. Addressing the challenges of non-expert user expression, the poor effectiveness of open-ended questions, and the lack of labeled training data in existing technologies are pressing technical issues that require immediate resolution. Summary of the Invention

[0004] The technical problems this invention aims to solve are: the poor performance of existing LegalAI systems when faced with fuzzy queries from non-expert users; the prevalence of open-ended questions in existing proactive questioning methods, which increases the cognitive burden on non-expert users and makes it difficult to obtain accurate legal facts; the lack of substantive legal information in simple questionnaires; and how to train a model capable of generating high-quality fact-complete questionnaires that are both legally professional and user-friendly, given the absence of manually labeled query-questionnaire data. To address these issues, this invention proposes a legal fact-complete method for non-expert users.

[0005] To achieve the above-mentioned objectives, the present invention specifically adopts the following technical solution:

[0006] In a first aspect, the present invention provides a method for supplementing legal facts for non-expert users, comprising the following steps:

[0007] S1. In the unlabeled iterative training paradigm, legal documents containing factual descriptions and court opinions are used as data sources. By simulating the process of non-expert users querying missing factual descriptions, and combining the auxiliary legal expert model to evaluate the effectiveness of the questionnaire in completing the factual descriptions, high-quality training corpus can be generated through multiple rounds of iteration without manual annotation. Then, the pre-trained question model is fine-tuned on the training corpus.

[0008] S2. In the reasoning stage, conditional logic legal provisions are transformed into multiple-choice questions, so that different options in the multiple-choice questions correspond to different legal consequences. Based on user queries, similar cases are retrieved from the case database, and the frequency of citation of each conditional logic legal provision in similar cases is statistically analyzed. The multiple-choice question corresponding to the conditional logic legal provision with the highest citation frequency is used as the reference context when the question model generates the questionnaire.

[0009] S3. By combining the finely tuned question model with the reference context, a structured questionnaire is dynamically generated for user queries, guiding non-expert users to supplement key facts and generating a completed user query, which is then used by the downstream LegalAI expert model to generate legal service responses.

[0010] Based on the above scheme, each step can be implemented in the following preferred manner.

[0011] As a preferred embodiment of the first aspect mentioned above, in S1, the first... The specific process of generating training corpus through rounds of iteration is as follows:

[0012] S11. Use a large language model to rewrite the factual descriptions in legal documents, discard some factual descriptions and restate them as incomplete user queries to simulate the fuzzy questioning scenarios of non-expert users.

[0013] S12, the first The training corpus generated by rounds of iterations is divided into Each sub-corpus is used to train a separate question-asking model. Number of training corpora for children

[0014] S13. Iterate through all the question models trained in S12, for the first... The first question model is updated using a cross-case generation strategy. Individual training corpus; among them All are query model indexes and satisfy ;

[0015] S14. After the traversal is complete, combine all the updated sub-training corpora from S13 as the first... The training corpus from the first iteration is used for the next iteration.

[0016] As a preferred embodiment of the first aspect mentioned above, in S13, the first case generation strategy is updated. The specific process of training the sub-corpus is as follows: the question model generates the sub-corpus for user queries. The system uses factual descriptions to simulate the process of non-expert users selecting answers for each candidate questionnaire, generating user responses for each candidate questionnaire. Then, each candidate questionnaire generated by the user query and question model, along with its corresponding user responses, is input into an auxiliary legal expert model to generate court opinions, outputting the probability that each candidate questionnaire generates a correct court opinion. The candidate questionnaire with the highest probability is selected as the new optimal questionnaire, replacing the original optimal questionnaire for that user query. The user query and the new optimal questionnaire are then combined into a training corpus and added to the first... In each sub-training corpus, the update of the sub-training corpus is completed; among them, This represents the number of candidate questionnaires.

[0017] As a preferred option in the first aspect mentioned above, the specific implementation process of converting conditional logic legal provisions into multiple-choice questions in S2 is as follows: non-normative legal provisions that do not support legal reasoning are filtered out from the original legal provision database. Then, the enumerated items in the remaining conditional logic legal provisions are mapped one by one to the options of the multiple-choice questions. The conditional logic legal provisions and the set of options of the multiple-choice questions are input into a large language model. The large language model generates a natural language question corresponding to the set of options. The natural language question and the set of options constitute a multiple-choice question, thus completing the conversion of conditional logic legal provisions into multiple-choice questions.

[0018] As a preferred embodiment of the first aspect mentioned above, the specific implementation process for generating the reference context in S2 is as follows:

[0019] S21, Based on user query Perform vector retrieval to identify the most similar cases from the case database. Several similar cases; among them, The number of similar cases;

[0020] S22. Calculate the conditional logic type legal provisions in the legal provisions database using the following formula. Citation frequency in similar cases :

[0021] ;

[0022] in, This is a similar case; A collection of similar cases; For indicator functions, when When the value is 1, the indicator function is 1; otherwise, it is 0.

[0023] S23. Sort conditional logic legal provisions according to their calculated citation frequency, starting with the most frequently cited. Each conditional logic-type legal provision is selected as a high-frequency legal provision, and the multiple-choice questions corresponding to each high-frequency legal provision are used as reference contexts when the question model generates the questionnaire.

[0024] As a preferred option of the first aspect mentioned above, the specific implementation process of S3 is as follows: The user query and reference context are input into the finely tuned question model to generate a structured questionnaire containing multiple multiple-choice questions. Each multiple-choice question in the structured questionnaire has a preset number of options. Non-expert users answer the multiple-choice questions in the structured questionnaire by selecting options. Finally, the answers to each multiple-choice question form the questionnaire answer, which is then combined with the user query to form the completed user query.

[0025] Secondly, the present invention provides a legal fact completion system for non-expert users, comprising:

[0026] The model optimization module is used in the unlabeled iterative training paradigm to use legal documents containing factual descriptions and court opinions as data sources. By simulating the process of non-expert users querying missing factual descriptions, and combining it with an auxiliary legal expert model to evaluate the effectiveness of the questionnaire in completing the factual descriptions, high-quality training corpus can be generated through multiple rounds of iteration without manual annotation. Then, the pre-trained question model is fine-tuned on the training corpus.

[0027] The information acquisition module is used in the reasoning stage to transform conditional logic legal provisions into multiple-choice questions, so that different options in the multiple-choice questions correspond to different legal consequences; it retrieves similar cases from the case database based on user queries, counts the citation frequency of each conditional logic legal provision in similar cases, and uses the multiple-choice question corresponding to the conditional logic legal provision with the highest citation frequency as the reference context when the question model generates the questionnaire.

[0028] The query completion module combines a finely tuned question model with the reference context to dynamically generate a structured questionnaire for user queries. This guides non-expert users to supplement key facts, generating a completed user query that is then used by the downstream LegalAI expert model to generate legal service responses.

[0029] Thirdly, the present invention provides a computer program product, including a computer program / instruction, which, when executed by a processor, can implement the legal fact completion method for non-expert users as described in any of the solutions in the first aspect above.

[0030] Fourthly, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the legal fact completion method for non-expert users as described in any of the solutions of the first aspect above.

[0031] Fifthly, the present invention provides a computer electronic device, which includes a memory and a processor;

[0032] The memory is used to store computer programs;

[0033] The processor is configured to, when executing the computer program, implement the legal fact completion method for non-expert users as described in any of the solutions of the first aspect above.

[0034] Compared with the prior art, the present invention has the following advantages:

[0035] First, this invention proposes a label-free iterative training paradigm that does not require manually labeled data. By assisting legal expert models in indirect evaluation and cross-case generation strategies, it effectively solves the problem of scarce "query-questionnaire" data in the legal field and realizes automatic optimization of the question model.

[0036] Second, this invention designs a cascading retrieval mechanism of "case-legal provisions-multiple choice questions", which overcomes the semantic gap between users' colloquial queries and the formal text of legal provisions, ensuring that the generated questionnaire questions have substantial legal significance and that different options correspond to clear legal consequences.

[0037] Third, this invention uses a structured questionnaire to replace open-ended questions, which significantly reduces the cognitive burden on non-expert users while ensuring legal professionalism. Experiments show that this method improves users' expressive support while maintaining comprehensibility. Attached Figure Description

[0038] Figure 1 This is a flowchart of the steps of the method of the present invention;

[0039] Figure 2 This is a schematic diagram illustrating the training and reasoning methods of the question-and-answer model.

[0040] Figure 3 This is a system block diagram of the present invention;

[0041] Figure 4 This is a schematic diagram of a computer electronic device provided by the present invention. Detailed Implementation

[0042] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below. Technical features in the various embodiments of the present invention can be combined accordingly without mutual conflict.

[0043] In the description of this invention, it should be understood that the terms "first" and "second" are used only for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first" and "second" may explicitly or implicitly include at least one of those features.

[0044] like Figure 1 As shown, in a preferred embodiment of the present invention, the above-mentioned legal fact completion method for non-expert users includes the following steps S1 to S3. The specific implementation process of each step will be described in detail below.

[0045] S1. In the unlabeled iterative training paradigm, legal documents containing factual descriptions and court opinions are used as data sources. By simulating the process of non-expert users querying missing factual descriptions, and combining this with an auxiliary legal expert model to evaluate the effectiveness of the questionnaire in completing the factual descriptions, high-quality training corpus can be generated through multiple rounds of iterations without manual annotation. Then, the pre-trained question model is fine-tuned on this training corpus.

[0046] It should be noted that in S1 of this invention, the question-asking model is essentially a large language model capable of performing question-and-answer tasks, i.e., taking a user query as input and outputting a questionnaire. It has been pre-trained and fine-tuned using the training corpus of this invention. Specifically, selectable question-asking models include Qwen3.6-35B-A3B, DeepSeek-V4-Flash, gemma-3-27b-it, etc., and are not limited in this invention.

[0047] It should be noted that in S1 of the present invention, for the auxiliary legal expert model, the objective function of this embodiment is defined as maximizing the logarithmic probability of generating a correct court opinion under the given conditions of "incomplete user query + generated candidate questionnaire + simulated user response", which is used to indirectly evaluate the quality of the questionnaire. Its architecture is not limited to a court opinion generation model and can be replaced by any legal artificial intelligence model that can reflect "fact-legal consequences", such as LegalOne, wisdomInterrogatory, LawLLM, etc., which are not limited in the present invention.

[0048] like Figure 2 As shown, it should be noted that in S1 of the present invention, the first... The specific process of generating training corpus through rounds of iteration is as follows:

[0049] S11. Utilizing large language models to describe facts in legal documents Rewrite the query, discarding some factual descriptions and restating it as an incomplete user query. This is to simulate vague questioning scenarios from non-expert users.

[0050] S12, the first Training corpus generated by rounds of iteration Divided into Each sub-corpus is used to train a separate question-asking model. The number of training corpora for children.

[0051] S13. Iterate through all the question models trained in S12, for the first... The first question model is updated using a cross-case generation strategy. Individual training corpus; among them All are query model indexes and satisfy .

[0052] Furthermore, in S13, the first case is updated through a cross-case generation strategy. The specific process of training the sub-corpus is as follows: the question model generates the sub-corpus for user queries. A candidate questionnaire is generated, using factual descriptions to simulate the process of non-expert users selecting answers for each candidate questionnaire, thus generating user responses for each candidate questionnaire; then, the user queries are... Each candidate questionnaire generated by the questioning model and the corresponding user response input-assisted legal expert model Generate court opinions Output the probability that each candidate questionnaire generates the correct court opinion. The candidate questionnaire with the highest probability is selected as the new optimal questionnaire, replacing the original optimal questionnaire for this user's query. The user's query and the new optimal questionnaire are then combined into a training corpus, which is added to the first... In each sub-training corpus, the update of the sub-training corpus is completed; among them, This represents the number of candidate questionnaires.

[0053] S14. After the traversal is complete, combine all the updated sub-training corpora from S13 as the first... round-by-round training corpus This is used for the next iteration.

[0054] It should be noted that in S1 of this invention, when the preset maximum number of iterations is reached, the training corpus generated in the last round is used as the final training corpus, and the pre-trained question model is fine-tuned based on it. It is worth noting that some question models are also involved in the update process of the training corpus. Although these question models have undergone parameter updates, they do not participate in the final fine-tuning process. That is to say, these question models are essentially used to assist in building a high-quality training corpus, while this invention uses a new question model that has only undergone pre-training for fine-tuning.

[0055] S2. In the reasoning stage, conditional logic legal provisions are transformed into multiple-choice questions, so that different options in the multiple-choice questions correspond to different legal consequences; similar cases are retrieved from the case database based on user queries, the citation frequency of each conditional logic legal provision in similar cases is counted, and the multiple-choice question corresponding to the conditional logic legal provision with the highest citation frequency is used as the reference context when the question model generates the questionnaire.

[0056] It should be noted that in the above S2 of this invention, conditional logic type legal provisions refer to legal provisions that specifically explain the scope of application, constituent elements, legal consequences or rights and obligations of a certain legal norm by listing multiple matters, situations, conditions or behaviors. They are usually organized into multiple enumerated items in the form of numbering (I), (II), (III), etc.

[0057] It should be noted that, in S2 of the present invention, the specific implementation process of converting conditional logic legal provisions into multiple-choice questions is as follows: non-normative legal provisions that do not support legal reasoning are filtered out from the original legal provision database; the remaining enumerated items in the conditional logic legal provisions are mapped one by one to the options of the multiple-choice questions; the conditional logic legal provisions and the set of options of the multiple-choice questions are input into a large language model; the large language model generates a natural language question corresponding to the set of options; and the natural language question and the set of options constitute a multiple-choice question, thus completing the conversion of conditional logic legal provisions into multiple-choice questions.

[0058] like Figure 2 As shown, it should be noted that the specific implementation process for generating the reference context in S2 of the present invention is as follows:

[0059] S21, Based on user query Perform vector retrieval to identify the most similar cases from the case database. Several similar cases; among them, This represents the number of similar cases.

[0060] It should be noted that when performing a search in S21, it is not only possible to search by vector; it can be replaced by search methods such as BM25 and TF-IDF. This invention does not impose any restrictions on these methods.

[0061] S22. Calculate the conditional logic type legal provisions in the legal provisions database using the following formula. Citation frequency in similar cases :

[0062]

[0063] in, This is a similar case; A collection of similar cases; For indicator functions, when When the value is 1, the indicator function is 1; otherwise, it is 0.

[0064] S23. Sort conditional logic legal provisions according to their calculated citation frequency, starting with the most frequently cited. Each conditional logic-based legal provision is selected as a high-frequency legal provision, and the multiple-choice questions corresponding to each high-frequency legal provision are used as reference context when the question model generates the questionnaire, ensuring that the generated questionnaire has clear legal basis and differentiation.

[0065] S3. By combining the finely tuned question model with the reference context, a structured questionnaire is dynamically generated for user queries, guiding non-expert users to supplement key facts and generating a completed user query, which is then used by the downstream LegalAI expert model to generate legal service responses.

[0066] It should be noted that in S3 of this invention, the LegalAI expert model is essentially a large language model capable of performing question-answering tasks, i.e., inputting a user query with completion and outputting a legal service response. Specifically, selectable LegalAI expert models include Qwen3.7-max, Lawformer, LegalOne, etc., and are not limited in this invention.

[0067] It should be noted that the specific implementation process of S3 in the present invention is as follows: the user query and reference context are input into the finely tuned question model to generate a structured questionnaire containing multiple multiple-choice questions. Each multiple-choice question in the structured questionnaire has a preset number of options. Non-expert users answer the multiple-choice questions in the structured questionnaire by selecting options. Finally, the answers to each multiple-choice question form the questionnaire answer and are combined with the user query to form the completed user query.

[0068] It should be noted that in this invention, the options for each multiple-choice question include "other" to support flexibility. This invention, by having users select options to answer multiple-choice questions, reduces the cognitive load required for users to independently organize their thoughts, recall factual details, and understand legal terminology. This allows non-legal professionals to supplement and describe case facts with a lower barrier to entry, significantly reducing cognitive load compared to open-ended questions.

[0069] The present invention will now demonstrate the application effect of the legal fact completion method for non-expert users described in S1~S3 of the above embodiments on a specific dataset through a specific example, so as to facilitate understanding of the essence of the present invention.

[0070] Example

[0071] The specific implementation process of the legal fact completion method for non-expert users adopted in this embodiment is as described above and will not be repeated here.

[0072] This embodiment is validated based on 17,865 judgments collected from China Judgments Online (covering 90 common causes of action) and a retrieval database containing 82,257 cases. Both the query model and the auxiliary model are fine-tuned based on Qwen2.5-7B-Instruct, simulating user interaction with GPT-4o-mini.

[0073] In the manual evaluation experiment, as shown in Table 1, this embodiment invited 10 legal experts and 50 non-expert participants to rate the quality of the questionnaire. The results showed that the FactFiller method of this invention was significantly better than the baseline method Multiple-Choice in the dimensions of relevance (4.19 vs 3.89), comprehensiveness (3.85 vs 3.51), and usefulness (4.04 vs 3.69) in expert evaluation (p<0.0001); in non-expert evaluation, FactFiller maintained comparable comprehensibility to the baseline (3.95 vs 3.97, no significant difference), while significantly improving expression support (4.03 vs 3.93, p=0.0492).

[0074] Table 1. Manual Evaluation Experiment

[0075]

[0076] In downstream task performance verification, as shown in Tables 2, 3, and 4, the FactFiller method of this invention improved the ROUGE series indicators by an average of 2.15 and the BLEU series indicators by an average of 4.61 in the court opinion generation task; in the legal case retrieval task, Recall@50, Recall@500, and Recall@1000 improved by 3.50%, 7.62%, and 8.30%, respectively; and in the five subsets of the confusion charge prediction task, the average accuracy improved by 1.84%.

[0077] Table 2. Generation of Court Opinions

[0078]

[0079] Table 3. Results of Case Study Search

[0080]

[0081] Table 4. Predictive Effect of Easily Confused Crimes

[0082]

[0083] It should also be noted that the legal fact completion method for non-expert users in the above embodiments can essentially be executed by a computer program or module. Therefore, similarly, based on the same inventive concept, another preferred embodiment of the present invention also provides a legal fact completion system for non-expert users, corresponding to the legal fact completion method for non-expert users provided in the above embodiments, such as... Figure 3 As shown, it includes:

[0084] The model optimization module is used in the unlabeled iterative training paradigm to use legal documents containing factual descriptions and court opinions as data sources. By simulating the process of non-expert users querying missing factual descriptions, and combining it with an auxiliary legal expert model to evaluate the effectiveness of the questionnaire in completing the factual descriptions, high-quality training corpus can be generated through multiple rounds of iteration without manual annotation. Then, the pre-trained question model is fine-tuned on the training corpus.

[0085] The information acquisition module is used in the reasoning stage to transform conditional logic legal provisions into multiple-choice questions, so that different options in the multiple-choice questions correspond to different legal consequences; it retrieves similar cases from the case database based on user queries, counts the citation frequency of each conditional logic legal provision in similar cases, and uses the multiple-choice question corresponding to the conditional logic legal provision with the highest citation frequency as the reference context when the question model generates the questionnaire.

[0086] The query completion module combines a finely tuned question model with the reference context to dynamically generate a structured questionnaire for user queries. This guides non-expert users to supplement key facts, generating a completed user query that is then used by the downstream LegalAI expert model to generate legal service responses.

[0087] It is understood that the legal fact completion method for non-expert users described in S1-S3 above can essentially be implemented by a computer program. Therefore, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer program product corresponding to the legal fact completion method for non-expert users provided in the above embodiments, which includes a computer program / instructions. When the computer program / instructions are executed by a processor, they can implement the legal fact completion method for non-expert users as described in the above embodiments.

[0088] Similarly, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer electronic device corresponding to the legal fact completion method for non-expert users provided in the above embodiments, such as... Figure 4 As shown, it includes a memory and a processor;

[0089] The memory is used to store computer programs;

[0090] The processor is configured to implement the legal fact completion method for non-expert users in the above embodiments when executing the computer program.

[0091] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0092] Therefore, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer-readable storage medium corresponding to the legal fact completion method for non-expert users provided in the above embodiments. The storage medium stores a computer program, which, when executed by a processor, can realize the legal fact completion method for non-expert users in the above embodiments.

[0093] Specifically, in the computer-readable storage medium of the above three embodiments, the stored computer program is executed by a processor, which can perform the aforementioned steps S1 to S3.

[0094] It is understood that the aforementioned storage media may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Furthermore, the storage media may also be various media capable of storing program code, such as USB flash drives, external hard drives, magnetic disks, or optical discs.

[0095] It is understood that the processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0096] It should also be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the system described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here. In the embodiments provided in this application, the division of steps or modules in the system and method is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple modules or steps may be combined or integrated together, and a module or step may also be split.

[0097] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the invention. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the invention. Therefore, all technical solutions obtained through equivalent substitution or transformation fall within the protection scope of the present invention.

Claims

1. A method for completing legal facts for non-expert users, characterized in that, Includes the following steps: S1. In the unlabeled iterative training paradigm, legal documents containing factual descriptions and court opinions are used as data sources. By simulating the process of non-expert users querying missing factual descriptions, and combining the auxiliary legal expert model to evaluate the effectiveness of the questionnaire in completing the factual descriptions, high-quality training corpus can be generated through multiple rounds of iteration without manual annotation. Then, the pre-trained question model is fine-tuned on the training corpus. S2. In the reasoning stage, conditional logic legal provisions are transformed into multiple-choice questions, so that different options in the multiple-choice questions correspond to different legal consequences. Based on user queries, similar cases are retrieved from the case database, and the frequency of citation of each conditional logic legal provision in similar cases is statistically analyzed. The multiple-choice question corresponding to the conditional logic legal provision with the highest citation frequency is used as the reference context when the question model generates the questionnaire. S3. By combining the finely tuned question model with the reference context, a structured questionnaire is dynamically generated for user queries, guiding non-expert users to supplement key facts and generating a completed user query, which is then used by the downstream LegalAI expert model to generate legal service responses.

2. The legal fact completion method for non-expert users as described in claim 1, characterized in that, In S1, the first The specific process of generating training corpus through rounds of iteration is as follows: S11. Use a large language model to rewrite the factual descriptions in legal documents, discard some factual descriptions and restate them as incomplete user queries to simulate the fuzzy questioning scenarios of non-expert users. S12, the first The training corpus generated by rounds of iterations is divided into Each sub-corpus is used to train a separate question-asking model. Number of training corpora for children S13. Iterate through all the question models trained in S12, for the first... The first question model is updated using a cross-case generation strategy. Individual training corpus; among them All are query model indexes and satisfy ; S14. After the traversal is complete, combine all the updated sub-training corpora from S13 as the first... The training corpus from the first iteration is used for the next iteration.

3. The legal fact completion method for non-expert users as described in claim 2, characterized in that, In S13, the first case is updated through a cross-case generation strategy. The specific process of training the sub-corpus is as follows: the question model generates the sub-corpus for user queries. A candidate questionnaire is generated by using factual descriptions to simulate the process of non-expert users selecting answers for each candidate questionnaire, and user responses are generated for each candidate questionnaire. Then, each candidate questionnaire generated by the user query and question model, along with its corresponding user responses, is input into the auxiliary legal expert model to generate court opinions. The model outputs the probability that each candidate questionnaire generates a correct court opinion. The candidate questionnaire with the highest probability is selected as the new optimal questionnaire, replacing the original optimal questionnaire for that user query. The user query and the new optimal questionnaire are then combined into a training corpus, which is added to the first... In each sub-training corpus, the update of the sub-training corpus is completed; among them, This represents the number of candidate questionnaires.

4. The legal fact completion method for non-expert users as described in claim 1, characterized in that, In S2, the specific implementation process for converting conditional logic legal provisions into multiple-choice questions is as follows: Non-normative legal provisions that do not support legal reasoning are filtered out from the original legal provision database. Then, the enumerated items in the remaining conditional logic legal provisions are mapped one by one to the options of the multiple-choice questions. The conditional logic legal provisions and the set of options of the multiple-choice questions are input into the large language model. The large language model generates a natural language question corresponding to the set of options. The natural language question and the set of options constitute a multiple-choice question, thus completing the conversion of conditional logic legal provisions into multiple-choice questions.

5. A method for completing legal facts for non-expert users as described in claim 4, characterized in that, In S2, the specific implementation process for generating the reference context is as follows: S21, Based on user query Perform vector retrieval to identify the most similar cases from the case database. Several similar cases; among them, The number of similar cases; S22. Calculate the conditional logic type legal provisions in the legal provisions database using the following formula. Citation frequency in similar cases : ; in, This is a similar case; A collection of similar cases; For indicator functions, when When the value is 1, the indicator function is 1; otherwise, it is 0. S23. Sort conditional logic legal provisions according to their calculated citation frequency, starting with the most frequently cited. Each conditional logic-type legal provision is selected as a high-frequency legal provision, and the multiple-choice questions corresponding to each high-frequency legal provision are used as reference contexts when the question model generates the questionnaire.

6. The legal fact completion method for non-expert users as described in claim 1, characterized in that, The specific implementation process of S3 is as follows: The user query and reference context are input into the finely tuned question model to generate a structured questionnaire containing multiple multiple-choice questions. Each multiple-choice question in the structured questionnaire has a preset number of options. Non-expert users answer the multiple-choice questions in the structured questionnaire by selecting options. Finally, the answers to each multiple-choice question form the questionnaire answer, which is then combined with the user query to form the completed user query.

7. A legal fact completion system for non-expert users, characterized in that, include: The model optimization module is used in the unlabeled iterative training paradigm to use legal documents containing factual descriptions and court opinions as data sources. By simulating the process of non-expert users querying missing factual descriptions, and combining it with an auxiliary legal expert model to evaluate the effectiveness of the questionnaire in completing the factual descriptions, high-quality training corpus can be generated through multiple rounds of iteration without manual annotation. Then, the pre-trained question model is fine-tuned on the training corpus. The information acquisition module is used in the reasoning stage to transform conditional logic legal provisions into multiple-choice questions, so that different options in the multiple-choice questions correspond to different legal consequences; it retrieves similar cases from the case database based on user queries, counts the citation frequency of each conditional logic legal provision in similar cases, and uses the multiple-choice question corresponding to the conditional logic legal provision with the highest citation frequency as the reference context when the question model generates the questionnaire. The query completion module combines a finely tuned question model with the reference context to dynamically generate a structured questionnaire for user queries. This guides non-expert users to supplement key facts, generating a completed user query that is then used by the downstream LegalAI expert model to generate legal service responses.

8. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it can implement the legal fact completion method for non-expert users as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the legal fact completion method for non-expert users as described in any one of claims 1 to 6.

10. A computer electronic device, characterized in that, Including memory and processor; The memory is used to store computer programs; The processor is configured to, when executing the computer program, implement the legal fact completion method for non-expert users as described in any one of claims 1 to 6.