Training sample generation method and device, electronic equipment and storage medium
By extracting multi-perspective question-and-answer data, hierarchical structured data, and contextualized terminology knowledge items from legal documents, a comprehensive training sample set is constructed, which solves the problems of scarce and homogeneous training data and improves the professionalism and robustness of the legal big data model.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WU HAN XIN ZHI SHU ZI KE JI YOU XIAN GONG SI
- Filing Date
- 2026-02-09
- Publication Date
- 2026-05-12
AI Technical Summary
Currently, when training large-scale legal models, there are problems of data scarcity and homogeneity, resulting in insufficient training samples that cannot meet the requirements of professionalism, accuracy, and robustness for high-performance models.
Through a systematic and automated data processing workflow, multi-dimensional and high-quality training samples are extracted from legal documents. Multiple preset models are used to generate multi-perspective question-and-answer data, hierarchical structured data structures, and contextualized knowledge items for terminology, thereby constructing a comprehensive training sample set.
It significantly enhances the training accuracy of the legal big data model, improves the model's ability to understand the expressions of different roles, its logical reasoning ability, and the accuracy of its terminology explanation, and is better able to adapt to questions from diverse users and provide professional and accurate legal services.
Smart Images

Figure CN122020177A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and more specifically, to a training sample generation method, apparatus, electronic device, and storage medium. Background Technology
[0002] As large language models are increasingly applied in legal services, judicial assistance, and intelligent consultation, the requirements for their professionalism, accuracy, and robustness are also rising dramatically. The core of training a high-performance legal domain model lies in obtaining high-quality, multi-dimensional training data rich in professional logic. However, current model training suffers from data scarcity and homogeneity, resulting in insufficient training samples.
[0003] Therefore, a solution is needed. Summary of the Invention
[0004] In view of this, embodiments of this application provide a training sample generation method, apparatus, electronic device, and storage medium, which aim to efficiently mine and construct multi-dimensional, high-quality training samples from legal document sources through a systematic and automated data processing process, so as to significantly enhance the training accuracy of the legal large model.
[0005] In a first aspect, embodiments of this application provide a training sample generation method, the method comprising: Obtain initial question-and-answer data pairs, which are generated based on pre-given legal documents; The initial question-and-answer data pairs are input into a first preset model to obtain target question-and-answer data pairs with different perspectives generated by the first preset model based on the initial question-and-answer data pairs; wherein, the perspectives include at least two of the prosecution's perspective, the defense's perspective, and the judge's perspective; The legal document is input into the second preset model to obtain a hierarchical data structure output by the second preset model after performing hierarchical structured extraction on the legal document; the hierarchical data structure includes at least: a basic information layer, a logical chain layer of facts and evidence, and a judgment logic layer; The legal document is input into a third preset model, and multiple knowledge entries are output by the third preset model after performing terminological contextual analysis on the legal document; each knowledge entry includes at least: the applicable context of each term in the legal document, and the incorrect and correct meanings generated based on the applicable context; The target question-and-answer data pairs, the hierarchical data structure, and the knowledge entries are integrated to form a comprehensive training sample set for training a large legal model.
[0006] In one feasible implementation, the initial question-and-answer data pair is input into a first preset model to obtain target question-and-answer data pairs with different perspectives generated by the first preset model based on the initial question-and-answer data pair, including: The first pre-defined script is invoked to perform the following steps: Construct a prompt template that includes perspective control instructions and legal constraint instructions; the legal constraint instructions are used to ensure that the content generated by the first preset model is logically consistent with the initial question-and-answer data pair. Call the interface of the first preset model and use the initial question-and-answer data pair and the prompt template as input to the first preset model; Receive the first output content returned by the first preset model; the first output content includes question-answer pairs restated from multiple different perspectives on the initial question-answer data pairs; The first output content is parsed and its format is validated to obtain the target question-and-answer data pair that meets the preset requirements.
[0007] In one feasible implementation, the legal document is input into a second preset model to obtain a hierarchical data structure output by the second preset model after performing hierarchical structured extraction of the legal document, including: The pre-defined second script is invoked to perform the following steps: Construct a hierarchical structured extraction prompt template, which includes: prompt information of the basic information layer, prompt information of the fact and evidence layer, and prompt information of the adjudication logic layer; Call the interface of the second preset model, and use the legal document and the prompt template as input to the second preset model; Receive and parse the second output content returned by the second preset model; The second output content is filled with missing values and standardized to obtain the hierarchical data structure.
[0008] In one feasible implementation, the legal document is input into a third preset model, resulting in multiple knowledge items output by the third preset model after performing contextualized terminology analysis on the legal document, including: Invoke the preset third script and execute the following steps: A prompt template for contextualized terminology analysis is constructed, which includes instructions for associating terminology recognition with context, instructions for extracting original text explanations and applicability analysis, and instructions for generating incorrect and correct meanings. Call the interface of the third preset model, and use the legal document and the prompt template as input to the third preset model; Receive and parse the third output content returned by the third preset model based on the prompt template; The third output content is structured and organized to obtain the knowledge entries.
[0009] In a feasible implementation, the target question-answer data pairs, the hierarchical data structure, and the knowledge entries are integrated, including: The target question-answer data pairs, the hierarchical data structure, and the knowledge entries are converted into a unified training data format, and then mixed and sampled to form the comprehensive training sample set.
[0010] In one feasible implementation, the prompt template further includes: The program stage parameters are used to instruct the first preset model to generate question-and-answer content by combining multiple different processing stages; and the legal constraint instructions are used to instruct the first preset model to return a self-check result of consistency between the crime and the sentencing in the first output content, so that the first script can filter the generated content based on the self-check result.
[0011] In one feasible implementation, the prompt template includes: For each initial question-and-answer data pair, the first script automatically generates multiple prompt templates containing different expression perspectives and program stage parameter combinations to drive the first preset model to generate the target question-and-answer data pair in multiple dimensions.
[0012] Secondly, embodiments of this application also provide a training sample generation apparatus, the apparatus comprising: The acquisition module is used to acquire initial question-and-answer data pairs, which are generated based on pre-given legal documents; The first module is used to input the initial question-and-answer data pair into a first preset model to obtain target question-and-answer data pairs with different expression perspectives generated by the first preset model based on the initial question-and-answer data pair; wherein, the expression perspective includes at least two of the prosecution's perspective, the defense's perspective, and the judge's perspective; The second module is used to input the legal documents into the second preset model to obtain a hierarchical data structure output by the second preset model after performing hierarchical structured extraction on the legal documents; the hierarchical data structure includes at least: a basic information layer, a logical chain layer of facts and evidence, and a judgment logic layer; The third module is used to input the legal document into a third preset model to obtain multiple knowledge entries output by the third preset model after performing terminology contextualization analysis on the legal document; each knowledge entry includes at least: the applicable context of each term in the legal document, and the incorrect and correct meanings generated based on the applicable context; The integration module is used to integrate the target question-and-answer data pairs, the hierarchical data structure, and the knowledge entries to form a comprehensive training sample set for training the legal big data model.
[0013] In one feasible implementation, the first module is used to input the initial question-and-answer data pair into a first preset model to obtain target question-and-answer data pairs with different perspectives generated by the first preset model based on the initial question-and-answer data pair, for the purpose of: The first pre-defined script is invoked to perform the following steps: Construct a prompt template that includes perspective control instructions and legal constraint instructions; the legal constraint instructions are used to ensure that the content generated by the first preset model is logically consistent with the initial question-and-answer data pair. Call the interface of the first preset model and use the initial question-and-answer data pair and the prompt template as input to the first preset model; Receive the first output content returned by the first preset model; the first output content includes question-answer pairs restated from multiple different perspectives on the initial question-answer data pairs; The first output content is parsed and its format is validated to obtain the target question-and-answer data pair that meets the preset requirements.
[0014] In one feasible implementation, the second module is used to input the legal document into a second preset model to obtain a hierarchical data structure output by the second preset model after performing hierarchical structured extraction on the legal document, for the following purposes: The pre-defined second script is invoked to perform the following steps: Construct a hierarchical structured extraction prompt template, which includes: prompt information of the basic information layer, prompt information of the fact and evidence layer, and prompt information of the adjudication logic layer; Call the interface of the second preset model, and use the legal document and the prompt template as input to the second preset model; Receive and parse the second output content returned by the second preset model; The second output content is filled with missing values and standardized to obtain the hierarchical data structure.
[0015] In one feasible implementation, the third module is used to input the legal document into a third preset model, and obtain multiple knowledge items output by the third preset model after performing terminological contextualization analysis on the legal document, for the following purposes: Invoke the preset third script and execute the following steps: A prompt template for contextualized terminology analysis is constructed, which includes instructions for associating terminology recognition with context, instructions for extracting original text explanations and applicability analysis, and instructions for generating incorrect and correct meanings. Call the interface of the third preset model, and use the legal document and the prompt template as input to the third preset model; Receive and parse the third output content returned by the third preset model based on the prompt template; The third output content is structured and organized to obtain the knowledge entries.
[0016] In one feasible implementation, the integration module is used to integrate the target question-answer data pairs, the hierarchical data structure, and the knowledge entries, for the following purposes: The target question-answer data pairs, the hierarchical data structure, and the knowledge entries are converted into a unified training data format, and then mixed and sampled to form the comprehensive training sample set.
[0017] In one feasible implementation, the prompt template further includes: The program stage parameters are used to instruct the first preset model to generate question-and-answer content by combining multiple different processing stages; and the legal constraint instructions are used to instruct the first preset model to return a self-check result of consistency between the crime and the sentencing in the first output content, so that the first script can filter the generated content based on the self-check result.
[0018] In one feasible implementation, the prompt template includes: For each initial question-and-answer data pair, the first script automatically generates multiple prompt templates containing different expression perspectives and program stage parameter combinations to drive the first preset model to generate the target question-and-answer data pair in multiple dimensions.
[0019] Thirdly, embodiments of this application also provide an electronic device, including: a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of the method as described in any one of the first aspects.
[0020] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the method as described in any one of the first aspects.
[0021] The present application provides a training sample generation method, apparatus, electronic device, and storage medium. By processing the initial question-and-answer pairs through a first preset model, diverse question-and-answer data from different perspectives such as the prosecution, defense, and judge are generated. This enables the model obtained from subsequent training to understand and simulate the expression methods and focus of different roles, thereby significantly enhancing the model's generalization ability when facing questions from diverse users.
[0022] By using a second pre-defined model to extract layers from the full text of legal documents, a multi-layered data structure was obtained, encompassing basic information, a logical chain of facts and evidence, and a judgment logic layer. This structure not only extracts surface-level fields but also reveals the inherent reasoning logic of the case, making the training data itself contain judicial logic. This effectively enhances the model's understanding of complex legal documents and its logical reasoning ability.
[0023] By employing a third pre-defined model to deeply analyze terminology in legal documents, a knowledge set is constructed that includes specific application scenarios for each term, along with explanations of common misunderstandings and corrections. This moves terminology learning away from abstract definitions and anchors it in real-world case contexts, significantly improving the accuracy and practicality of the model's explanations of legal terms and effectively correcting users' common-sense misunderstandings.
[0024] Finally, by systematically integrating three types of content—diverse question-and-answer pairs, logically rich structured data, and contextualized knowledge items—a comprehensive training sample set was created, which enhanced the three dimensions of dialogue ability, logical depth, and knowledge accuracy.
[0025] Compared with existing technologies, the embodiments of this application do not stop at the simple transformation of legal texts or the extraction of information from a single dimension. Instead, through the aforementioned multi-model collaborative generation method, it systematically and automatically mines multi-dimensional complementary training data from the same source of legal documents. This directly solves several problems existing in current training samples, such as the single perspective of training data, lack of logical structure, and detachment of terminology from specific contexts, providing a high-quality data foundation for training a more comprehensive, professional, and robust legal model.
[0026] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0027] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 A flowchart of a training sample generation method provided in an embodiment of this application is shown.
[0029] Figure 2 A flowchart of another training sample generation method provided in an embodiment of this application is shown.
[0030] Figure 3 A flowchart of another training sample generation method provided in an embodiment of this application is shown.
[0031] Figure 4 A flowchart of another training sample generation method provided in an embodiment of this application is shown.
[0032] Figure 5 A schematic diagram of a training sample generation device provided in an embodiment of this application is shown.
[0033] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation
[0034] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0035] As large language models are increasingly applied in legal services, judicial assistance, and intelligent consultation, the requirements for their professionalism, accuracy, and robustness are also rising dramatically. The core of training a high-performance legal domain model lies in obtaining high-quality, multi-dimensional training data rich in professional logic. However, current model training suffers from data scarcity and homogeneity, resulting in insufficient training samples.
[0036] Based on this, embodiments of this application provide a training sample generation method, apparatus, electronic device, and storage medium, which are described below through embodiments.
[0037] To facilitate understanding of this embodiment, a training sample generation method disclosed in this application will first be described in detail. For example... Figure 1 As shown, it includes the following steps: Step 101: Obtain initial question-and-answer data pairs, which are generated based on pre-given legal documents.
[0038] The pre-given legal document forms the basis of all training samples. It typically refers to a legally binding judicial document that contains complete facts of the case, evidence, arguments from the prosecution and defense, and the court's reasoning. Alternatively, it can be a legal provision, a contract template, a complaint, etc., without limitation.
[0039] After obtaining the legal document, extract or construct at least one pair containing the legal question and its corresponding answer from the legal document.
[0040] A common approach is for legal professionals to manually design several question-and-answer pairs based on the core content of the judgment (e.g., case focus, applicable law, judgment result). Alternatively, automated methods can be used to assist in generation; for example, a basic natural language processing model can be used to automatically generate potential questions based on chapter titles or key sentences in the legal document, which are then manually verified and corrected. Regardless of the specific method used, the final "initial question-and-answer data pairs" should ensure that the questions are clear, the answers are accurate, and they strictly adhere to the facts and conclusions stated in the original document.
[0041] The key to this step is that the initial question-and-answer pair is not an arbitrarily constructed generic legal question, but rather closely anchored to a specific, authentic judicial document. This establishes an immutable "factual anchor" and "legal conclusion anchor" for all subsequent processing steps, ensuring that regardless of subsequent shifts in perspective or expression, the core legal principles remain stable and consistent. It provides a credible source and a consistent benchmark for the entire methodology.
[0042] Step 102: Input the initial question-and-answer data pair into the first preset model to obtain target question-and-answer data pairs with different expression perspectives generated by the first preset model based on the initial question-and-answer data pair; wherein, the expression perspective includes at least two of the prosecution's perspective, the defense's perspective, and the judge's perspective.
[0043] This step expands the perspective and enriches the expression of a single "seed" question and answer. The initial question and answer pair usually presents the question and answer in a relatively objective or universal way. This step introduces a processing unit called the "first presupposition model," whose core function is to simulate the cognitive positions and expression habits of different legal roles in litigation.
[0044] Here, the first pre-defined model typically refers to a large language model with strong text understanding and generation capabilities, whose output style can be guided by specific instructions. By inputting the initial question-and-answer pairs into this first pre-defined model and providing explicit role-playing instructions, the model can generate new question-and-answer content that is semantically equivalent but has a different tone, based on the same set of facts and legal conclusions.
[0045] For example, regarding the question of "whether the defendant's actions constitute surrender," the model can generate: Questions and answers presented from the prosecution's perspective may be more accusatory, focusing on listing evidence to prove that the defendant's actions do not fully meet the legal conditions for surrender. Questions and answers presented from the defense's perspective tend to be more defensive and explanatory, potentially emphasizing the defendant's subjective intent and the specific circumstances of their voluntary surrender, arguing that it constitutes surrender. Questions and answers presented from the judge's perspective should reflect neutrality and discretion, synthesizing the opinions of both the prosecution and the defense, citing relevant legal provisions for analysis, and ultimately providing the court's reasoning for its determination.
[0046] This step expands what was originally a single question-and-answer document into a set of dialogue samples examining the same legal issue from different perspectives. The resulting "target question-and-answer data pairs" significantly increase the diversity and realism of subsequent training data. When used to train legal models, this enables the models not only to learn to answer "standard questions" but also to understand how the same legal issue might be raised, debated, and adjudicated under different roles and perspectives, thereby significantly improving the model's adaptability and the relevance of its responses in real-world applications. Using this diverse data as training samples is key to enhancing the dialogue capabilities and generalization performance of subsequent large-scale legal models.
[0047] Step 103: Input the legal document into the second preset model to obtain a hierarchical data structure output by the second preset model after performing hierarchical structured extraction on the legal document; the hierarchical data structure includes at least: a basic information layer, a logical chain layer of facts and evidence, and a judgment logic layer.
[0048] This step aims to perform in-depth analysis and logical reconstruction of legal documents, transforming them from coherent natural language text into a cognitive representation with clear hierarchies that is easier for machines to understand and process. The second pre-defined model introduced here is typically a processing unit designed or guided to perform complex text understanding and structured information extraction tasks.
[0049] The purpose of inputting the complete original legal document into this second pre-set model is not to obtain a new answer, but to make it "understand" and "deconstruct" the legal document like a professional legal assistant. The task that the second pre-set model needs to accomplish is "hierarchical structured extraction," which means that it cannot simply find scattered information points, but must reorganize the content of the legal document according to a pre-set logical hierarchy.
[0050] In other words, it first extracts the outermost layer of basic information, which is similar to the "identity card" of a document, including objective fields such as case number, basic information of the parties involved, court of trial, date of judgment, charge, and sentence. This step is equivalent to creating a standardized index for the case.
[0051] Next, the second pre-defined model delves into the narrative portion of the document, constructing a logical chain of facts and evidence. This is no longer a simple listing, but rather identifying key behavioral nodes in the development of the case (e.g., preparation, execution, and subsequent actions), and associating various types of evidence mentioned in the document (such as documentary evidence, physical evidence, and witness testimony) with these specific behavioral nodes. In this way, a dynamic chain of facts is formed, with "actions" as the skeleton and "evidence" as the support, clearly demonstrating "what happened" and "how to prove it."
[0052] Finally, the second presupposition model needs to extract the adjudicative logic layer of the legal document. This is the core and soul of the legal document, involving the judge's reasoning process. The second presupposition model needs to identify the points of contention in the case (e.g., whether it constitutes legitimate self-defense, whether sentencing mitigating circumstances are valid), and extract the judge's analytical opinions on each point of contention, the reasons for accepting specific evidence, and the various aggravating and mitigating factors considered in the final sentencing. This layer reveals the complete reasoning path from "fact-finding" to "application of law" and then to "final judgment."
[0053] Through this step, a complex legal document with intricate reasoning is transformed into a hierarchical data structure with clear layers and relationships. This structured data provides subsequent training with an information density and logical explicitness far exceeding that of the original text. Training a large-scale legal model with such data helps it transcend simply matching surface-level text, truly learning to understand the inherent development logic of cases and the deep reasoning patterns of judicial decisions, thereby acquiring stronger legal logic analysis and reasoning capabilities.
[0054] Step 104: Input the legal document into the third preset model to obtain multiple knowledge entries output by the third preset model after performing terminology contextualization analysis on the legal document; each knowledge entry includes at least: the applicable context of each term in the legal document, and the incorrect and correct meanings generated based on the applicable context.
[0055] This step aims to contextualize and pedagogically analyze legal terminology in legal documents, transforming abstract legal concepts into three-dimensional knowledge units that relate to specific cases and include common cognitive corrections. The "Third Presupposition Model" introduced here is a processing unit specifically guided to perform tasks of in-depth understanding of terminology and knowledge construction.
[0056] Its object of study is also the full text of legal documents, but its focus shifts from the macro-level logic of the case to the micro-level practical application of specialized terminology in legal practice. The third presupposition model does not simply list terms and their dictionary definitions, but rather performs a "contextualized analysis." This means it locates the position of each key term within the vast sea of legal documents and carefully analyzes its context.
[0057] First, the third presupposition model needs to summarize the specific applicable context of each term in this case. For example, regarding the case of "surrender," the third presupposition model needs to analyze: in this case, at what time, in what way, to which agency did the defendant surrender, and to what extent did he / she make a confession? And based on what specific circumstances did the court ultimately determine (or not determine) that he / she had surrendered? This is equivalent to taking a snapshot of the term "at the scene of the crime," recording how it is activated and used in actual judgment.
[0058] More importantly, the third pre-defined model also needs to proactively construct comparative teaching materials based on the specific contexts uncovered. It generates one or more common "misinterpretations" of the term—misunderstandings that typically simulate cognitive biases arising from everyday experiences or one-sided interpretations of legal provisions. For example, in the aforementioned "surrender" scenario, a possible misinterpretation is: "As long as you voluntarily tell others after committing a crime, it counts as surrendering."
[0059] Next, the third pre-defined model will combine the specific context in which the term is actually applied in this case with relevant legal provisions to generate a corresponding "correct meaning," thus constituting a direct and targeted correction. For example, the correct interpretation (correct meaning) would state: "Surrender requires both voluntary surrender and truthful confession, and must be made to judicial authorities, relevant organizations, or individuals. In this case, although the defendant voluntarily confessed to his family, he did not surrender to judicial authorities within a reasonable time, therefore he does not meet the requirement of voluntary surrender." Through this step, each legal term is no longer an isolated label, but is encapsulated into a complete knowledge entry containing "real-world usage cases," "common misconceptions," and "authoritative contextual corrections." This type of data allows subsequent training of the legal big data model to not only learn the "static definitions" of terms, but also understand their "flexible boundaries" and "discriminatory points" in dynamic judicial practice, and to pre-emptively identify and correct common user misunderstandings. This significantly improves the accuracy, practicality, and user-friendliness of the legal big data model when providing legal consultation, answering questions, or providing explanations.
[0060] Step 105: Integrate the target question-answer data pairs, the hierarchical data structure, and the knowledge entries to form a comprehensive training sample set for training the legal big model.
[0061] The preceding steps (102, 103, 104) are like three parallel production lines, producing intermediate data products with different forms and functions: multi-perspective question-and-answer pairs for dialogue training, structured cognitive frameworks for logical understanding training, and contextualized terminology units for knowledge deepening training.
[0062] The task of this step is not simply to mechanically pile these three types of data together, but to organically integrate them according to the final training objective. This integration is a purposeful data engineering process aimed at building a comprehensive training sample set that is internally rich and dimensionally complementary, providing a full and balanced "nourishment" for the large legal model.
[0063] One approach to integration involves format standardization and hybridization. For example, multi-perspective question-and-answer pairs can be organized into standard instruction-following or dialogue format samples; hierarchical data structures can be transformed into corresponding information extraction, text summarization, or reading comprehension task samples based on their different levels (such as basic information, logical chains, and reasoning); and terminology knowledge items can be constructed into samples in the form of definition explanations, correct / false judgments, or case-based question-and-answer sessions. Subsequently, these samples in different task formats are arranged together according to certain strategies (such as proportional mixing or progressive mixing based on curriculum learning concepts) to form a large and diverse training sample set.
[0064] Another approach to integration is to establish internal connections. For example, when constructing training samples, question-and-answer pairs, structured data fragments, and terminological knowledge generated from the same original legal document can be consciously linked. In this way, when the model learns, it may simultaneously encounter factual questions and answers about the same case, the logical chain behind them, and the analysis of key legal concepts involved, thereby helping it to build a more complete and self-consistent legal cognitive map.
[0065] Through this integration, the resulting training sample set possesses advantages unmatched by single-type data. It is no longer "specialized" data excelling in only one task, but rather capable of simultaneously providing the model with experience in dialogue interaction, paradigms for logical reasoning, and precise, professional legal knowledge. Training with such a comprehensive sample set is expected to guide the development of a more balanced and powerful comprehensive capability in large-scale legal models, enabling them not only to engage in fluent dialogue but also to perform in-depth analysis and provide answers that are both professional and easy to understand, thus better meeting the needs of real-world, complex legal service scenarios.
[0066] The present application provides a training sample generation method, apparatus, electronic device, and storage medium. By processing the initial question-and-answer pairs through a first preset model, diverse question-and-answer data from different perspectives such as the prosecution, defense, and judge are generated. This enables the model obtained from subsequent training to understand and simulate the expression methods and focus of different roles, thereby significantly enhancing the model's generalization ability when facing questions from diverse users.
[0067] By using a second pre-defined model to extract layers from the full text of legal documents, a multi-layered data structure was obtained, encompassing basic information, a logical chain of facts and evidence, and a judgment logic layer. This structure not only extracts surface-level fields but also reveals the inherent reasoning logic of the case, making the training data itself contain judicial logic. This effectively enhances the model's understanding of complex legal documents and its logical reasoning ability.
[0068] By employing a third pre-defined model to deeply analyze terminology in legal documents, a knowledge set is constructed that includes specific application scenarios for each term, along with explanations of common misunderstandings and corrections. This moves terminology learning away from abstract definitions and anchors it in real-world case contexts, significantly improving the accuracy and practicality of the model's explanations of legal terms and effectively correcting users' common-sense misunderstandings.
[0069] Finally, by systematically integrating three types of content—diverse question-and-answer pairs, logically rich structured data, and contextualized knowledge items—a comprehensive training sample set was created, which enhanced the three dimensions of dialogue ability, logical depth, and knowledge accuracy.
[0070] Compared with existing technologies, the embodiments of this application do not stop at the simple transformation of legal texts or the extraction of information from a single dimension. Instead, through the aforementioned multi-model collaborative generation method, it systematically and automatically mines multi-dimensional complementary training data from the same source of legal documents. This directly solves several problems existing in current training samples, such as the single perspective of training data, lack of logical structure, and detachment of terminology from specific contexts, providing a high-quality data foundation for training a more comprehensive, professional, and robust legal model.
[0071] Furthermore, by using the first, second, and third preset models for processing, the efficiency and quality of training sample set processing can be improved, and the amount of manual work can be reduced.
[0072] In one feasible implementation, the initial question-and-answer data pair is input into a first preset model to obtain target question-and-answer data pairs with different perspectives generated by the first preset model based on the initial question-and-answer data pair, including: Call the preset first script, such as Figure 2 As shown, the following steps are performed: Step 201: Construct a prompt template that includes perspective control instructions and legal constraint instructions; the legal constraint instructions are used to ensure that the content generated by the first preset model is logically consistent with the initial question-and-answer data pair.
[0073] This step is the core design regarding how to specifically guide the "first preset model" in its operation. To generate diverse and compliant question-and-answer responses, a clear "task specification," or prompt template, needs to be provided to the first preset model. This template is not simply a question input, but rather contains sophisticated control logic.
[0074] Perspective control instructions are the creative part of the template. They explicitly require models to move away from generic or neutral response patterns and instead mimic the thought processes and language styles of specific legal roles (such as prosecutors, defense attorneys, and judges). The instructions specifically define the typical tone, focus, and expressive tendencies of these roles; for example, instructing the model to "speak as a defense attorney, emphasizing mitigation of the defendant's liability."
[0075] Legal constraints serve as the cornerstone of stability within the template, combating the inherent "illusion" or risk of arbitrary interpretation inherent in large language models. It mandates that the first pre-defined model, when creatively shifting perspectives, must strictly adhere to the factual basis and legal conclusions established by the "initial question-and-answer data pair." The instructions explicitly require that generated content not alter the core facts of the case, key evidence, the determined charges, or the final sentencing outcome, ensuring that all perspectives are developed within the same legal factual framework. This ensures that the diversity of generated content does not come at the expense of legal accuracy, maintaining "logical consistency" between the generated content and the source.
[0076] Step 202: Call the interface of the first preset model and use the initial question-and-answer data pair and the prompt template as input to the first preset model.
[0077] This step is an automated operation to put the designed "blueprint" into practice. The script calls the application programming interface (API) of the selected large language model (i.e., the first preset model). The script combines two key pieces of material—the "initial question-and-answer data pair" as the content benchmark and the "prompt template" as the behavioral guideline—into a single complete request and sends it to the first preset model. This is equivalent to providing the first preset model with specific case material (question-and-answer pair) along with detailed work requirements (prompt template), instructing it to begin performing a creative but constrained text generation task.
[0078] Step 203: Receive the first output content returned by the first preset model; the first output content includes question-answer pairs restated from multiple different perspectives on the initial question-answer data pairs.
[0079] This step involves receiving the model's initial output. After processing the input, the first preset model returns its generated text, i.e., the "first output content." According to the prompt template in step 201, this output content is expected to contain multiple restatements of the original question-and-answer pairs. Each version should reflect a specific perspective; for example, one version might be a question-and-answer exchange from the prosecutor's perspective, arguing their indictment, while another version might be a question-and-answer exchange from the defense attorney's perspective, offering a defense argument. The content received at this stage is a direct product of the model under instruction and has not yet undergone subsequent quality control.
[0080] Step 204: Parse and format the first output content to obtain the target question-and-answer data pair that meets the preset requirements.
[0081] This step is an automated part of quality control. Because the output of a large language model may have inconsistent formats, partially follow instructions, or contain irrelevant information, the directly received "first output content" needs to be post-processed to become usable high-quality data.
[0082] Parsing typically refers to accurately extracting structured fields such as viewpoint labels, restated questions, and answers from a text block returned by the model (which may be JSON, XML, or text with specific tags).
[0083] Format validation ensures that the extracted content conforms to preset data specifications, such as whether each question-and-answer pair is complete and whether the viewpoint tags belong to the allowed set.
[0084] More importantly, this step can (and is often designed to) include consistency verification logic. For example, it can quickly check whether the generated content contains core elements that are prohibited from being tampered with by legally binding directives (such as charges or sentences) through simple rules (like keyword matching) or by calling a lightweight verification model. Generated results that fail the verification will be automatically filtered out.
[0085] Finally, the parsed and verified content is standardized into "target question-and-answer data pairs" with uniform format and controllable quality. Thus, the automated generation process, from initial questions and answers from a single perspective to multi-perspective, high-quality extended questions and answers, is completed.
[0086] It should be noted that, in one feasible embodiment, the prompt template further includes: The program stage parameters are used to instruct the first preset model to generate question-and-answer content by combining multiple different processing stages; and the legal constraint instructions are used to instruct the first preset model to return a self-check result of consistency between the crime and the sentencing in the first output content, so that the first script can filter the generated content based on the self-check result.
[0087] Furthermore, the prompt template includes: For each initial question-and-answer data pair, the first script automatically generates multiple prompt templates containing different expression perspectives and program stage parameter combinations to drive the first preset model to generate the target question-and-answer data pair in multiple dimensions.
[0088] In other words, the prompt template constructed in step 201 above can be further refined to introduce richer scenario dimensions and embed automated quality assurance mechanisms. This mainly involves enhancements on two levels: I. Dimensional Deepening: Introducing “Program Stage Parameters” and “First Preset Model Self-Check”.
[0089] Procedural Stage Parameters: In addition to controlling the perspective of expression ("who is speaking"), the prompt template can also incorporate procedural stage parameters to control the timing of the litigation process simulated in the question-and-answer content ("when it is spoken"). For example, the parameter can instruct the first preset model to generate simulated content such as "question and answer during the lawyer's first meeting with the suspect during the investigation stage," "question and answer during the prosecutor's interrogation during the review and prosecution stage," "question and answer during the judge's questioning during the trial," or "question and answer when the party consults with the lawyer after the verdict." The focus, level of information mastery, and language style are completely different at different stages. Adding this parameter allows the generated data to cover the entire life cycle of legal events, greatly enhancing the scenario diversity and temporal realism of the training data, and helping the model understand the different manifestations and answer focuses of the same legal issue at different procedural nodes.
[0090] Strengthening Legal Constraints: Self-Check and Result Return of the First Preset Model: To ensure that the generated content does not deviate from legal facts, the aforementioned legal constraint instructions can be designed to be more proactive and verifiable. Specifically, the instructions can explicitly require the first preset model to perform a key check on its output while generating restatement content: that is, to self-determine whether the charge and sentencing conclusions in the generated content are consistent with the initial question-and-answer pair. The first preset model needs to include the result of this self-check (e.g., "consistent," "inconsistent," or a confidence score) as a specific field in the returned output content. This design moves the responsibility for consistency verification to the generation stage of the first preset model, and its returned self-check results provide a direct and structured basis for subsequent automated filtering.
[0091] II. Automated scaling: Combinatorial generation and batch-driven.
[0092] Based on the aforementioned prompt template design that incorporates multiple dimensions such as "perspective" and "stage," the entire data expansion process can achieve a high degree of automation and scalability. Specifically, an automation script (i.e., the "first script") can be written, the logic of which is: For each initial question-and-answer data pair, the script does not just build a single prompt template. Instead, it automatically performs a Cartesian product combination based on a preset set of perspectives, such as {prosecution, defense, judge}, and a set of procedural stages, such as {investigation, prosecution review, trial}, to generate a series of (e.g., 3 perspectives × 3 stages = 9) different specific prompt templates.
[0093] Each template combination precisely specifies an expression task for a "specific role at a specific stage of litigation." Subsequently, the script automatically and cyclically submits the original question-and-answer pairs and this series of template combinations to the first preset model for invocation.
[0094] In this way, a single initial question-and-answer data pair can automatically and in batches generate dozens or even hundreds of target question-and-answer data pairs with different combinations of dimensions. This allows the dataset to expand exponentially, and this expansion is systematic, covering all predefined important scene dimensions, rather than random or aimless. Ultimately, the resulting dataset is a high-dimensional target question-and-answer dataset that is fully expanded in both the perspective and procedural dimensions, laying a solid data foundation for training models capable of adapting to complex scenarios.
[0095] In one feasible implementation, the legal document is input into a second preset model to obtain a hierarchical data structure output by the second preset model after performing hierarchical structured extraction of the legal document, including: Call the preset second script, such as Figure 3As shown, perform the following steps: Step 301: Construct a hierarchical structured extraction prompt template, which includes: prompt information of the basic information layer, prompt information of the fact and evidence layer, and prompt information of the adjudication logic layer.
[0096] This step aims to create a clear, structured "analysis guide" for the model. The script dynamically generates a hierarchical, structured extraction prompt template. This template is not a single instruction, but a well-defined task description that explicitly tells the model to interpret the document from three levels: The basic information layer prompts the second preset model to extract the most superficial objective information from the document, such as the case number, court, parties involved, crime, sentence, and the specific legal provisions on which it is based, just like filling out a form.
[0097] The information provided by the fact and evidence layer indicates that the second pre-set model should delve deeper into the case narrative, like sorting out clues, to find the key behavioral nodes in the development of the event (when, where, and what was done), and to identify the various types of evidence appearing in the text (such as physical evidence, documentary evidence, and witness testimony). It is also necessary to associate each piece of evidence with the specific behavior it aims to prove.
[0098] The prompts in the judgment logic layer instruct the second preset model to focus on the court's reasoning, identify the points of contention in the case (such as whether it constitutes surrender), and extract the judge's analytical views on each point of contention, the key evidence ultimately accepted and the reasons for it, as well as the aggravating or mitigating circumstances considered in sentencing.
[0099] Step 302: Call the interface of the second preset model and use the legal document and the prompt template as input to the second preset model.
[0100] The script calls the "second preset model" via an application programming interface (API). It sends the complete legal document along with the prompt template built in step 301 as input to the second preset model. This is equivalent to handing a complex case file and a detailed analysis outline to a professional analyst, instructing them to complete the analysis report according to the outline's structure.
[0101] Step 303: Receive and parse the second output content returned by the second preset model.
[0102] After analyzing the document according to the instructions, the second preset model will return the generated text, i.e., the second output content. Ideally, this content should be a structured text block (e.g., JSON format) containing information organized according to three levels: basic information, chain of factual evidence, and adjudication logic. The script needs to parse this content, that is, identify and extract the values of each field, and convert them into a structured data object that can be processed internally by the program.
[0103] Step 304: Perform missing value filling and standardization on the second output content to obtain the hierarchical data structure.
[0104] Since the output of the second preset model may contain omissions, inconsistent formats, or non-standard expressions, the received data needs to be post-processed.
[0105] Missing value imputation: The script checks the parsed data structure. For fields that the template requires to be extracted but the model does not provide or explicitly indicates as "unknown", the script will uniformly fill them with a specific null value marker (such as null) to maintain the integrity of the data structure.
[0106] Standardization processing: For key fields, such as legal citations, the script applies preset normalization rules. This ensures that similar information extracted from different documents has a consistent expression, facilitating subsequent retrieval, comparison, and learning by the legal model.
[0107] After processing in step 304, the potentially rough and inconsistent second output content is transformed into a hierarchical data structure with a unified format, complete fields, and standardized expression. This structured data clearly reveals the whole picture of the case and the internal logic of the judgment, and can be directly used for subsequent data integration and training of legal big data models.
[0108] In one feasible implementation, the legal document is input into a third preset model, resulting in multiple knowledge items output by the third preset model after performing contextualized terminology analysis on the legal document, including: Call a pre-defined third script, such as Figure 4 As shown, perform the following steps: Step 401: Construct a prompt template for contextualized terminology parsing. The prompt template includes instructions for associating terminology recognition with context, instructions for extracting original text explanations and applicability analysis, and instructions for generating incorrect and correct meanings.
[0109] This step aims to design a multi-stage, task-oriented "operation manual" for the "Third Preset Model," namely, a prompt template for contextualized terminology analysis. The core goal of this template is not to allow the model free rein, but to guide it like an experienced legal knowledge engineer, executing a systematic process of mining, analyzing, and constructing knowledge about the technical terms in legal documents. Therefore, the template typically contains three progressive instruction modules: The instruction on terminology identification and contextual association: This is the starting point for knowledge discovery. This instruction requires the third presupposition model to act as a "terminology explorer." It instructs the model to thoroughly read the given legal document text and automatically identify all professional terms with specific legal meanings. More importantly, the instruction requires the third presupposition model not only to list the terms but also to briefly describe the specific factual context in which each identified term is associated within the relevant paragraph. For example, for the identified term "legitimate self-defense," the third presupposition model needs to point out that in this case, it is associated with the specific contextual description that "the defendant claimed that the other party attacked first with a weapon, therefore his counterattack constituted self-defense." This initially establishes an anchoring relationship between terms and specific cases.
[0110] The instruction for extracting original text interpretation and applicability analysis: After identifying the terms and their general contexts, this instruction requires the third presupposition model to act as a "legal analyst." It instructs the third presupposition model, for each identified term, to search and extract original text fragments from the entire legal document (especially reasoning sections such as "The Court holds") that directly explain the court's arguments or reasons for applying the term. For example, finding and extracting the judge's original judgment explaining why the circumstances in this case "constitute" or "do not constitute" surrender. Simultaneously, the instruction also requires the third presupposition model to extract one or two concise sentences of applicability analysis based on these original texts, summarizing "the key points of how the term was specifically applied or excluded in this case." This step deepens the term from an abstract concept into a practical case supported by specific judicial rules.
[0111] Instructions for Generating Incorrect and Correct Meanings: Finally, this instruction requires the third pre-defined model to act as a "teaching material compiler." It instructs the third pre-defined model to actively construct comparative teaching materials based on the specific contexts and authoritative explanations obtained in the first two steps. Specifically, it requires the third pre-defined model to generate one or more typical "incorrect meanings" of the term—these erroneous understandings should simulate common misunderstandings that non-legal professionals might have based on common sense, literal meaning, or limited information (e.g., believing that "robbery" is "forcibly taking things"). Following this, the instruction requires the third pre-defined model to generate targeted "correct meanings" by referencing or based on the extracted original text explanations and applicability analyses, precisely identifying and correcting the aforementioned erroneous understandings. This step completes a closed loop of knowledge encapsulation from "identifying the term" to "explaining the term" and then to "preventing misunderstandings."
[0112] Through such a structured prompt template, the third pre-set model is systematically guided to transform the professional terms scattered in the judgment into independent knowledge units rich in pedagogical value, which include "scenario anchors", "authoritative evidence" and "correctness analysis".
[0113] Step 402: Call the interface of the third preset model and use the legal document and the prompt template as input to the third preset model.
[0114] The system (usually controlled by a "third script") calls the "third preset model" through an application programming interface (API). It encapsulates the complete legal document text and the prompt template for contextualized terminology parsing constructed in step 401 into a structured request and sends it to the third preset model.
[0115] This means that the input received by the third pre-defined model contains two core parts: The object of analysis is the full text of the original legal document that records the complete facts, evidence, and reasoning of the judgment in the case. This provides a unique and authoritative textual source for all subsequent analyses.
[0116] A detailed workflow guide: This is the prompt template containing instructions for the three stages of "terminology recognition - text extraction - error correction generation". This template precisely specifies the sequence of tasks that the third preset model needs to perform, the objectives of each stage, and the expected knowledge form of the final output.
[0117] This process can be likened to this: We submit a complex case file (legal document), along with a work instruction (prompt template) that requires systematically extracting and analyzing specific types of knowledge points from the file and organizing them into teaching cards according to a specific format, to an intelligent assistant (third pre-set model) with strong text comprehension capabilities. Through this invocation, the third pre-set model is officially activated and begins to perform in-depth scanning, understanding, and creative reorganization of the legal document according to the given instructions, in order to complete the task of constructing contextualized terminological knowledge.
[0118] Step 403: Receive and parse the third output content returned by the third preset model based on the prompt template.
[0119] Once the third preset model has completed processing based on the input (legal documents and complex prompt template) from step 402, it will return all the text content it generated via API, i.e., the third output content. This content is a response to the multi-instruction prompt template from step 401.
[0120] Since the prompt template itself is phased, the ideal output of the third preset model should also be a structured response. For example, it might be returned as a list or nested objects: The response to the "Term Recognition and Context Association" instruction is: a list in which each element contains the identified term and a brief description of the corresponding factual context.
[0121] Response to the "Interpretation and Applicability Analysis" instruction: For each term in the previous list, provide further relevant excerpts from the original text and key points regarding its applicability in this case.
[0122] Response to the "Generate Incorrect and Correct Meanings" instruction: Similarly, for each term, provide the generated common misinterpretations and the correct meaning based on the original text.
[0123] The receiving operation is when the system obtains the original text block containing the expected structure described above.
[0124] Then, the parsing operation begins. The purpose of parsing is to transform this raw text, which may be organized in the form of specific tags (such as JSON, XML) or natural language paragraphs, into discrete data fields that the program can accurately recognize and process. Specifically, the parser (usually part of a "third-party script") will: Identify the overall structure of the output content (e.g., identify that it is a JSON array).
[0125] Traverse this structure and extract the components under each term entry: term name, factual context, original text citation, applicable analysis, incorrect meaning, correct meaning, etc.
[0126] These extracted values are then assigned to the corresponding properties of structured objects (such as dictionaries or class instances) defined internally by the program.
[0127] This step is crucial, as it transforms the human-readable "textual response" generated by the third pre-defined model into a machine-manageable structured data object suitable for subsequent automated processing. At this point, the knowledge content output by the model has been initially "unpacked" and organized, preparing it for the final formatting process in the next step.
[0128] Step 404: The third output content is structured and organized to obtain the knowledge entries.
[0129] While the raw parsing results returned by the third preset model already possess a preliminary structure, their content may contain redundant information, inconsistent expressions, or loose logical connections. This step refines, correlates, and formats the aforementioned results through structured processing, ultimately producing standardized "knowledge entries."
[0130] Specifically, the process includes: denoising and refining the extracted original text fragments and generated content to ensure the professionalism and accuracy of their expression; logically binding and semantically integrating all information elements associated with the same term—including its specific applicable context, the original judgment supporting its usage, the applicability analysis based on the case, and the corresponding correct and incorrect semantic comparisons—and encapsulating them into a cohesive, self-consistent, independent data object; and finally, standardizing the format of all data objects to ensure that their fields are complete and their types are consistent.
[0131] After this step, the third output is a standardized set consisting of a series of knowledge items. Each item systematically presents a specific application example of a legal term in judicial practice and its cognitive boundaries, making it high-quality training data that can be directly used to enhance the model's ability to understand and interpret legal concepts.
[0132] Based on the same technical concept, embodiments of this application also provide a training sample generation device, such as... Figure 5 As shown, the device includes: The acquisition module 501 is used to acquire initial question-and-answer data pairs, which are generated based on pre-given legal documents.
[0133] The first module 502 is used to input the initial question-and-answer data pair into the first preset model to obtain target question-and-answer data pairs with different expression perspectives generated by the first preset model based on the initial question-and-answer data pair; wherein, the expression perspective includes at least two of the prosecution's perspective, the defense's perspective, and the judge's perspective.
[0134] The second module 503 is used to input the legal document into the second preset model to obtain a hierarchical data structure output by the second preset model after performing hierarchical structured extraction on the legal document; the hierarchical data structure includes at least: a basic information layer, a logical chain layer of facts and evidence, and a judgment logic layer.
[0135] The third module 504 is used to input the legal document into the third preset model to obtain multiple knowledge entries output by the third preset model after performing terminology contextualization analysis on the legal document; each knowledge entry includes at least: the applicable context of each term in the legal document, and the incorrect and correct meanings generated based on the applicable context.
[0136] The integration module 505 is used to integrate the target question-answer data pairs, the hierarchical data structure, and the knowledge entries to form a comprehensive training sample set for training the legal big model.
[0137] In one feasible implementation, the first module is used to input the initial question-and-answer data pair into a first preset model to obtain target question-and-answer data pairs with different perspectives generated by the first preset model based on the initial question-and-answer data pair, for the purpose of: The first pre-defined script is invoked to perform the following steps: Construct a prompt template that includes perspective control instructions and legal constraint instructions; the legal constraint instructions are used to ensure that the content generated by the first preset model is logically consistent with the initial question-and-answer data pair.
[0138] The interface of the first preset model is called, and the initial question-and-answer data pair and the prompt template are used as input to the first preset model.
[0139] Receive the first output content returned by the first preset model; the first output content includes question-answer pairs restated from multiple different perspectives on the initial question-answer data pairs.
[0140] The first output content is parsed and its format is validated to obtain the target question-and-answer data pair that meets the preset requirements.
[0141] In one feasible implementation, the second module is used to input the legal document into a second preset model to obtain a hierarchical data structure output by the second preset model after performing hierarchical structured extraction on the legal document, for the following purposes: The pre-defined second script is invoked to perform the following steps: A hierarchical structured extraction prompt template is constructed, which includes: prompt information of the basic information layer, prompt information of the fact and evidence layer, and prompt information of the adjudication logic layer.
[0142] Call the interface of the second preset model and use the legal document and the prompt template as input to the second preset model.
[0143] Receive and parse the second output content returned by the second preset model.
[0144] The second output content is filled with missing values and standardized to obtain the hierarchical data structure.
[0145] In one feasible implementation, the third module is used to input the legal document into a third preset model, and obtain multiple knowledge items output by the third preset model after performing terminological contextualization analysis on the legal document, for the following purposes: Invoke the preset third script and execute the following steps: A prompt template for contextualized terminology parsing is constructed. The prompt template includes instructions for associating terminology recognition with context, instructions for extracting original text explanations and applicability analysis, and instructions for generating incorrect and correct meanings.
[0146] The interface of the third preset model is invoked, and the legal document and the prompt template are used as inputs to the third preset model.
[0147] Receive and parse the third output content returned by the third preset model based on the prompt template.
[0148] The third output content is structured and organized to obtain the knowledge entries.
[0149] In one feasible implementation, the integration module is used to integrate the target question-answer data pairs, the hierarchical data structure, and the knowledge entries, for the following purposes: The target question-answer data pairs, the hierarchical data structure, and the knowledge entries are converted into a unified training data format, and then mixed and sampled to form the comprehensive training sample set.
[0150] In one feasible implementation, the prompt template further includes: The program stage parameters are used to instruct the first preset model to generate question-and-answer content by combining multiple different processing stages; and the legal constraint instructions are used to instruct the first preset model to return a self-check result of consistency between the crime and the sentencing in the first output content, so that the first script can filter the generated content based on the self-check result.
[0151] In one feasible implementation, the prompt template includes: For each initial question-and-answer data pair, the first script automatically generates multiple prompt templates containing different expression perspectives and program stage parameter combinations to drive the first preset model to generate the target question-and-answer data pair in multiple dimensions.
[0152] Figure 6 A schematic diagram of an electronic device provided in this application embodiment includes: a processor 601, a storage medium 602, and a bus 603. The storage medium 602 stores machine-readable instructions executable by the processor 601. When the electronic device runs the training sample generation method as described in the embodiment, the processor 601 communicates with the storage medium 602 via the bus 603, and the processor 601 executes the machine-readable instructions to perform the steps as described in the embodiment.
[0153] In this embodiment, the storage medium 602 may also execute other machine-readable instructions to perform other methods as described in the embodiment. For details on the specific execution steps and principles, please refer to the description of the embodiment, which will not be repeated here.
[0154] This application also provides a computer-readable storage medium storing a computer program that is executed by a processor to perform the steps as described in the embodiments.
[0155] In this embodiment, the computer program, when run by the processor, can also execute other machine-readable instructions to perform other methods as described in the embodiments. For details on the specific execution steps and principles, please refer to the description of the embodiments, which will not be repeated here.
[0156] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the coupling or direct coupling or communication connection shown or discussed may be through some communication interface; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.
[0157] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0158] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0159] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0160] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for generating training samples, characterized in that, The method includes: Obtain initial question-and-answer data pairs, which are generated based on pre-given legal documents; The initial question-and-answer data pairs are input into a first preset model to obtain target question-and-answer data pairs with different perspectives generated by the first preset model based on the initial question-and-answer data pairs; wherein, the perspectives include at least two of the prosecution's perspective, the defense's perspective, and the judge's perspective; The legal document is input into the second preset model to obtain a hierarchical data structure output by the second preset model after performing hierarchical structured extraction on the legal document; the hierarchical data structure includes at least: a basic information layer, a logical chain layer of facts and evidence, and a judgment logic layer; The legal document is input into a third preset model, and multiple knowledge entries are output by the third preset model after performing terminological contextual analysis on the legal document; each knowledge entry includes at least: the applicable context of each term in the legal document, and the incorrect and correct meanings generated based on the applicable context; The target question-and-answer data pairs, the hierarchical data structure, and the knowledge entries are integrated to form a comprehensive training sample set for training a large legal model.
2. The method according to claim 1, characterized in that, The initial question-and-answer data pairs are input into a first preset model to obtain target question-and-answer data pairs with different perspectives generated by the first preset model based on the initial question-and-answer data pairs, including: The first pre-defined script is invoked to perform the following steps: Construct a prompt template that includes perspective control instructions and legal constraint instructions; the legal constraint instructions are used to ensure that the content generated by the first preset model is logically consistent with the initial question-and-answer data pair. Call the interface of the first preset model and use the initial question-and-answer data pair and the prompt template as input to the first preset model; Receive the first output content returned by the first preset model; the first output content includes question-answer pairs restated from multiple different perspectives on the initial question-answer data pairs; The first output content is parsed and its format is validated to obtain the target question-and-answer data pair that meets the preset requirements.
3. The method according to claim 1, characterized in that, The legal document is input into a second preset model, resulting in a hierarchical data structure output by the second preset model after performing hierarchical structured extraction of the legal document, including: The pre-defined second script is invoked to perform the following steps: Construct a hierarchical, structured, extracted prompt template; the prompt template includes: prompt information at the basic information layer, prompt information at the fact and evidence layer, and prompt information at the adjudication logic layer; Call the interface of the second preset model, and use the legal document and the prompt template as input to the second preset model; Receive and parse the second output content returned by the second preset model; The second output content is filled with missing values and standardized to obtain the hierarchical data structure.
4. The method according to claim 1, characterized in that, The legal document is input into a third preset model, resulting in multiple knowledge items output by the third preset model after performing contextualized terminology analysis on the legal document, including: Invoke the preset third script and execute the following steps: Construct a prompt template for contextualized terminology analysis; the prompt template includes instructions for associating terminology recognition with context, instructions for extracting original text explanations and applicability analysis, and instructions for generating incorrect and correct meanings; Call the interface of the third preset model, and use the legal document and the prompt template as input to the third preset model; Receive and parse the third output content returned by the third preset model based on the prompt template; The third output content is structured and organized to obtain the knowledge entries.
5. The method according to claim 1, characterized in that, Integrating the target question-and-answer data pairs, the hierarchical data structure, and the knowledge entries includes: The target question-answer data pairs, the hierarchical data structure, and the knowledge entries are converted into a unified training data format, and then mixed and sampled to form the comprehensive training sample set.
6. The method according to claim 2, characterized in that, The prompt template also includes: The program stage parameters are used to instruct the first preset model to generate question-and-answer content by combining multiple different processing stages; and the legal constraint instructions are used to instruct the first preset model to return a self-check result of consistency between the crime and the sentencing in the first output content, so that the first script can filter the generated content based on the self-check result.
7. The method according to claim 6, characterized in that, The prompt template includes: For each initial question-and-answer data pair, the first script automatically generates multiple prompt templates containing different expression perspectives and program stage parameter combinations to drive the first preset model to generate the target question-and-answer data pair in multiple dimensions.
8. A training sample generation device, characterized in that, The device includes: The acquisition module is used to acquire initial question-and-answer data pairs, which are generated based on pre-given legal documents; The first module is used to input the initial question-and-answer data pair into a first preset model to obtain target question-and-answer data pairs with different expression perspectives generated by the first preset model based on the initial question-and-answer data pair; wherein, the expression perspective includes at least two of the prosecution's perspective, the defense's perspective, and the judge's perspective; The second module is used to input the legal documents into the second preset model to obtain a hierarchical data structure output by the second preset model after performing hierarchical structured extraction on the legal documents; the hierarchical data structure includes at least: a basic information layer, a logical chain layer of facts and evidence, and a judgment logic layer; The third module is used to input the legal document into a third preset model to obtain multiple knowledge entries output by the third preset model after performing terminology contextualization analysis on the legal document; each knowledge entry includes at least: the applicable context of each term in the legal document, and the incorrect and correct meanings generated based on the applicable context; The integration module is used to integrate the target question-and-answer data pairs, the hierarchical data structure, and the knowledge entries to form a comprehensive training sample set for training the legal big data model.
9. An electronic device, characterized in that, include: The device includes a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is in operation, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of the training sample generation method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the training sample generation method as described in any one of claims 1 to 7.