Sample data generation method and electronic equipment

Through the combination of the large language model and ELO-ranking technology, high-quality legal sample data is generated, which solves the problems of scarce and poor quality training sample data in the legal field, and improves the effectiveness of intelligent legal services.

CN120086601AInactive Publication Date: 2025-06-03HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510582871.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-06-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The scarcity and poor quality of training sample data in the legal field has led to the inability to meet the increasing demand for legal Q&A.

Method used

Multiple legal questions are generated by a multiple language model based on the target legal provisions and prompt words. The second largest language model determines candidate questions. The first largest language model generates multiple answers. The answer scores are compared with the second largest language model and ELO-ranking technology to generate the best question-and-answer pair as sample data.

Benefits of technology

Automatically generate high-quality sample data, avoid deviations caused by manual intervention, improve the efficiency of sample data generation, and ensure the accuracy of answers through target legal provisions and ELO-ranking mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086601A_ABST
    Figure CN120086601A_ABST
Patent Text Reader

Abstract

The invention discloses a sample data generation method. The method comprises the steps of generating a plurality of legal questions through a first large language model based on target legal provisions and cues; inputting the plurality of legal questions into a second large language model to determine candidate questions; the first large language model generates a plurality of answers based on the candidate questions, the target legal provisions and the cue words; and inputting the plurality of answers and the candidate questions into the second large language model to obtain answer scores of the plurality of answers, comparing the answer scores based on an ELO-Ranking technology to obtain an optimal question and answer pair, taking the optimal question and answer pair as sample data, and solving the problems of scarcity of training samples and poor quality in the legal field.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data processing, and in particular, to a method for generating sample data and an electronic device. Background Art

[0002] With the rapid development of big data and artificial intelligence technologies, people have begun to use LLM (Large Language Model) to provide intelligent legal services. The demand for intelligent legal services is increasing day by day. However, the training sample data in the legal field is not only scarce but also of poor quality for the most part. This has led to the fact that the effect of intelligent legal services cannot meet the increasing demand for legal Q&A. Summary of the Invention

[0003] The purpose of the present application is to provide a method for generating sample data, which can solve the problems of scarce and poor-quality training samples in the legal field.

[0004] In a first aspect, the present application provides a method for generating sample data, including: Generating a plurality of legal questions through a first large language model based on a target legal provision and a prompt; Inputting the plurality of legal questions into a second large language model to determine candidate questions; The first large language model generates a plurality of answers based on the candidate questions, the target legal provision, and the prompt; Inputting the plurality of answers and the candidate questions into the second large language model to obtain answer scores for the plurality of answers, and comparing the answer scores based on the ELO-ranking technology to generate the best Q&A pairs as sample data.

[0005] Optionally, inputting the plurality of legal questions into the second large language model to determine candidate questions includes: Inputting the plurality of legal questions into the second large language model, and obtaining question scores for the plurality of legal questions based on the legal provision compliance and expression clarity of the plurality of legal questions; Comparing the question scores based on the ELO-ranking technology, and taking the legal question with the highest question score as the candidate question.

[0006] Optionally, the first large language model adopts the FT-Llama3 framework, and the second large language model adopts the RGA framework.

[0007] Optionally, the first large language model generates a plurality of answers based on the candidate questions, the target legal provision, and the prompt, including: By adjusting the generation parameters of the first large language model, the first large language model generates multiple answers according to the candidate question, the target legal provision, and the prompt words, where the generation parameters include at least one of a temperature parameter, a random seed, a Top-K sampling parameter, and a Top-P sampling parameter.

[0008] Optionally, when the generation parameters include a temperature parameter, the temperature parameter is set between 0.7 and 1.5.

[0009] Optionally, the first large language model generates multiple answers based on the candidate question, the target legal provision, and the prompt words, including: The value of the temperature parameter in the softmax function of the first large language model is 0.9, In each generation cycle, the random seed in the first large language model is randomly adjusted, and the first large language model combines the candidate question, the target legal provision, and the prompt words to generate diverse answers based on the random seed.

[0010] Optionally, comparing the scores of the answers based on the ELO-ranking technique to obtain the best Q&A pair includes: Obtain the question score of the candidate question and the answer scores of the multiple answers, and integrate the scores of the legal question and the corresponding answer through a preset weight; Combining the integrated scores, use the ELO-ranking technique to rank the Q&A pairs composed of legal questions and corresponding answers, where the question scores and answer scores include at least one of legal provision compliance, Q&A accuracy, Q&A logic, and expression clarity; Take the Q&A pair with the highest score as the best Q&A pair.

[0011] Optionally, inputting the multiple answers and the candidate question into the second large language model to obtain the answer scores of the multiple answers includes: The second large language model analyzes the legal provision consistency between the multiple answers and the candidate question according to the target legal provision; The second large language model analyzes the precedent relevance of the multiple answers according to the target legal provision and its own knowledge base; The second large language model analyzes the logical coherence of the multiple answers according to the context association relationship between the multiple answers and the candidate question; For each answer and the candidate question, the second large language model integrates the legal provision consistency, precedent relevance, and logical coherence of the Q&A pair to obtain the score of the Q&A pair.

[0012] In a second aspect, an embodiment of the present application further provides a question-answering method based on a large language model, including: Input a legal question and a prompt word into the large language model through an input device, where the large language model is trained with the sample data generated by the above sample data generation method; Display the answer output by the large language model through a display device.

[0013] In a third aspect, an embodiment of the present application further provides an electronic device, including: An input device for collecting legal questions; At least one processor and at least one memory, the at least one memory can be used to store a computer program, and the at least one processor can execute the computer program to implement the above sample data generation method.

[0014] The embodiment of the present application uses a first large language model, combines target legal provisions and prompt words to generate candidate questions and multiple answers, uses a second large language model to compare the scores of candidate questions and multiple answers in combination with the ELO-ranking technology, determines the best question-answer pair as sample data. Since the screening process is carried out automatically without manual intervention, it avoids the deviation caused by manual intervention and greatly improves the generation efficiency of sample data. In addition, the accuracy of the answer is ensured by combining the target legal provisions and the ELO-ranking mechanism. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a schematic flowchart of a model generation method provided by an embodiment of the present application; Figure 2 It is a schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0016] The following will describe the present application in detail with reference to the specific embodiments shown in the drawings, but these embodiments do not limit the present application. Any structural, method, or functional transformation made by those of ordinary skill in the art based on these embodiments is included in the protection scope of the present application.

[0017] Please refer to Figure 1 , an embodiment of the present application provides a model generation method, including the following steps: Step 101: Generate multiple legal questions through a first large language model based on target legal provisions and prompt words; Step 102: Input the multiple legal questions into a second large language model to determine candidate questions; Step 103: The first large language model generates multiple answers based on the candidate questions, target legal provisions, and prompt words; Step 104: Input multiple answers and candidate questions into the second large language model to obtain answer scores for the multiple answers. Based on the ELO-ranking technology, compare the answer scores to obtain the best question-answer pair, which is used as sample data. Exemplarily, a question-answer pair represents a pair of a question and an answer, including a question and an answer.

[0018] The above-mentioned first large language model can adopt pre-trained models with various transformer structures such as GPT and Llama. In an alternative embodiment, the first large language model adopts the FT-Llama3 framework, and the second large language model adopts the RAG framework. The second large language model can include multiple large language models that adopt the RAG framework.

[0019] In this way, by using two large language models with different roles, namely the generative large language model and the discriminative large language model, to replace a single generative model, the quality of synthetic data can be improved. During the generation stage, the diversity can be controlled by adjusting the temperature parameter to ensure that the multiple generated texts have sufficient diversity in terms of language expression, information organization, style, etc. In the screening stage, a prompt design that balances quality and diversity is introduced to prevent the model from repeatedly selecting candidate items with similar styles.

[0020] Exemplarily, the first large language model can collect and organize legal provisions to form a structured database. For example, collect and organize Chinese legal provision information, classify the legal provisions according to categories such as civil law and criminal law, so as to form a structured database. This structured database can serve as the knowledge base of the first large language model, providing basic data support for the subsequent generation of prompt words and answers.

[0021] For example, the first large language model can collect Chinese legal provisions through channels such as public legal databases and legal literature, and use natural language processing technology to clean and structure the collected Chinese legal provisions. Each legal provision is structured and classified according to information such as provision ID, provision content, promulgation date, and legal category, and stored in the knowledge base of the first large language model. This knowledge base supports full-text retrieval and quick positioning functions, so as to quickly provide basic data support for the subsequent generation of prompt words and answers.

[0022] The prompt words can be designed according to the target generation task and can be generated in various ways such as templated prompts, dynamic prompt generation, and context examples. The prompt words generated in these ways are in a format that is convenient for the large language model to understand the prompt and generate content related to the question as an answer.

[0023] After receiving the prompt words, the first large language model combines the above-mentioned knowledge base and the target legal provisions to generate legal questions and answers, and initially filters out content that is not relevant to the target legal provisions or legal questions.

[0024] In an embodiment of the present application, the second large language model can be used as an AI judge to evaluate the quality of legal questions and answers generated by the first large language model, and score them in terms of prompt relevance, legal provision relevance, logic, language expression quality, etc., so as to screen out candidate questions and the best Q&A pairs.

[0025] Exemplarily, inputting multiple legal questions into the second large language model to determine candidate questions may include: Inputting multiple legal questions into the second large language model, and obtaining question scores of the multiple legal questions based on the compliance of legal provisions and the clarity of expression of the multiple legal questions; Comparing the question scores based on the ELO-ranking technology, and taking the legal question with the highest question score as the candidate question.

[0026] For example, adopting a dynamic scoring strategy based on the ELO-ranking technology, comparing multiple candidate questions pairwise and scoring them, and screening out the question with the best score as the candidate question.

[0027] Embodiments of the present application can design different prompt templates according to different generation tasks. In the prompt, context examples can be introduced to enhance the guiding ability of the prompt. Exemplarily, the template of the prompt can but is not limited to adopting the following format: For example: When generating legal questions, the prompt is: "Based on the following legal regulations, please generate relevant legal questions:\n{Content of legal regulations}".

[0028] When answering legal questions, the prompt is: "According to the following legal regulations, please answer the question:\n{Content of legal regulations}\nQuestion: {Generated question}". After the prompt template is designed, the prompt template is input into the first large language model, so that the first large language model can dynamically adjust parameters according to the prompt template to ensure the relevance and accuracy of the generated content.

[0029] In embodiments of the present application, the relevance and accuracy of the generated content can be ensured by but not limited to the following methods.

[0030] Judging the consistency of legal provisions of the Q&A pair according to the target legal provisions; Judging the relevance of precedents of the Q&A pair based on sample data and the corresponding knowledge base; Judging the logical coherence of the Q&A pair based on the context correlation between the candidate question and the candidate answer; Combining the consistency of legal provisions, the relevance of precedents and the logical coherence of the Q&A pair to obtain the score of the Q&A pair.

[0031] In the embodiments of the present application, multiple answers can be generated by adjusting the generation parameters of the first large language model, where the generation parameters include at least one of a temperature parameter, a random seed, a Top-K sampling parameter, and a Top-P sampling parameter. Exemplarily, the temperature parameter and the random seed in the first large language model can be adjusted to generate multiple different answers.

[0032] For example, the value of the temperature parameter in the softmax function of the first large language model is 0.9; in each answer generation cycle, the first large language model randomly adjusts the random seed in its generation parameters, and combines the candidate question, the target legal provision, and the prompt words to generate different answers. After N such generation cycles, N different answers can be obtained.

[0033] Exemplarily, during the process of generating a question by the first large language model, after each word is calculated, the probability distribution of the next word output by the first large language model is predicted through the softmax function in the first large language model.

[0034] Exemplarily, assuming that the input sequence of the question is x = [x 1 , x 2 , …, x n , and the first large language model needs to predict the probability of the next word x n+1 through the softmax function, then the probability of x n+1 can be obtained through the following formula:

[0035] where z i represents the score of the i-th word (which can be the logits output by the fully connected layer in the large language model), that is, Z i represents the relative probability of the i-th word, T represents the temperature parameter (temperature), and e is the base of the natural logarithm. In a large language model (LLM), logits refer to the unnormalized scores generated by the model in the output layer (usually a fully connected layer), that is, the value of Zi, which can be used to represent the "score" of each possible next word.

[0036] Since the temperature parameter (temperature) is applied in the large language model and acts on the probability distribution in the softmax function, it can adjust the "smoothness" of the probability distribution, thereby affecting the diversity, randomness, and accuracy of the question-and-answer data generated by the large language model. And in large language models (such as GPT series, LLaMA, etc.), the process of generating question-and-answer data is usually based on conditional probability. Therefore, adjusting the temperature parameter can affect the output of the large language model.

[0037] Exemplarily, assume that the Q&A data generated by the large language model is text. After a previous word or a piece of context is given, the large language model calculates the probability distribution of the next word based on the conditional calculation reflected by the softmax function. The influence of the temperature parameter in the softmax function on the probability distribution of the softmax function can be but is not limited to the following: Low temperature (T < 1): Reducing the temperature makes the probability distribution of the softmax function become more "sharp". In this case, the large language model is more inclined to select the word with the highest probability, thereby generating more accurate and stable text. When the temperature is low, the output of the large language model is more conservative. Therefore, the low temperature of the temperature data reduces the chance of generating unexpected content.

[0038] High temperature (T > 1): Increasing the temperature makes the probability distribution of the softmax function become more flat. In this case, the large language model's choice of vocabulary is more random, and the generated text is more diverse and innovative.

[0039] T = 1: In this case, the weights of the probability distribution of the softmax function do not change, and the output of the large language model usually maintains a certain degree of diversity and a certain degree of quality.

[0040] The first large language model can generate multiple answers based on the prompt words in the following way: Assume that the prompt word prompts the output target of the large language model to generate corresponding legal answers according to legal provisions. The output of the large language model can include multiple candidate words. Assume that there are 3 candidate words, namely: "correct" (corresponding probability 0.7), "wrong" (corresponding probability 0.2), "possible" (corresponding probability 0.1).

[0041] The inventor found that when using a low temperature (T = 0.5), the probability distribution of the softmax function becomes sharper, and the model tends to select the candidate word "correct" with the highest probability. This means that the generated text is more stable and consistent, suitable for scenarios that require accuracy. At this time, the probability distribution of the softmax function can adopt the following formula

[0042] where z i represents the score of the i-th word (which can be the logits output in the large language model), represents the relative probability of the i-th word, and e is the base of the natural logarithm.

[0043] When using a higher temperature (T = 1.5), the probability distribution of the softmax function will be smoother. Due to the reduced difference in the probability distribution, the probability of the large language model choosing the two candidate words "wrong" or "possible" relatively increases. Therefore, the answers generated in this case may be more creative or random, suitable for tasks that require more diverse answers. At this time, the probability distribution of the softmax function can be expressed by the following formula

[0044] where z i represents the score of the i-th word (which can be the logits output in the large language model), represents the relative probability of the i-th word, and e is the base of the natural logarithm.

[0045] Therefore, when the large language model needs to generate a standard and conservative answer based on legal provisions, the temperature parameter will be set to a low temperature value (such as 0.3 - 0.5). The large language model will choose words with higher probabilities to avoid generating irrelevant or legally non-compliant content. When multiple different answers are needed for comparison, especially for complex or open-ended questions, setting the temperature parameter to a high temperature value (such as 0.7 - 1.0) can help the large language model generate more creative or diverse answers.

[0046] After multiple tests and practical applications, the inventors found that: when answering legal-related questions, too low a temperature parameter (such as 0.1) may lead to overly single, uncreative, and rigid answers, making it difficult to cover multiple possibilities of legal provisions; while too high a temperature parameter (such as 1.6 or higher) may lead to overly vague answers, prone to generating inappropriate content or even incorrect inferences; when the temperature parameter is set to 0.9, the accuracy of the answers generated by the large language model is the highest, and it can also meet the need for answer diversity. Therefore, in tasks related to legal provisions, the temperature parameter can be set between 0.7 - 1.5. By setting the temperature within this range, the model can generate different candidate answers, increasing the space and possibility for subsequent ELO-ranking selection. Since the large language model reaches a good balance between performance and quality when the temperature parameter is 0.9, if this value is used to generate legal question-and-answer pairs, it can provide appropriate diversity without having too much impact on the language quality generated by the model.

[0047] After generating multiple answers by adjusting the generation parameters of the large language model, answers that are obviously irrelevant or logically chaotic can be initially filtered to form question-and-answer pairs consisting of candidate questions and answers.

[0048] In the embodiments of the present application, the question-and-answer pairs generated in the above process can be input into the second large language model (AI judge). Through discriminative prompts (Prompts), the question-and-answer pairs are screened for quality. Based on the pre-trial legal professional evaluation criteria, the candidate content is evaluated. The legal professional evaluation criteria can include, but are not limited to: legal professionalism, relevance, conciseness, creativity, integrity, etc.

[0049] Exemplarily, in the embodiments of the present application, the first large language model can include two AI assistants. In the legal professional evaluation, the AI judge first compares the answers generated by the two AI assistants. Subsequently, according to the evaluation results, clear discriminative labels (such as **"Assistant A is better", "Draw", "Assistant B is better"**) are output, and a rationality analysis process for the discriminative labels is provided, which is used as the basis for subsequent optimization. This can ensure that the first large language model accurately quotes legal provisions and uses rigorous legal reasoning to draw conclusions.

[0050] The system improves the screening efficiency of the question-and-answer quality through the intelligent discrimination ability of the large language model, ensuring that the finally output legal questions and answers have higher professionalism, accuracy and practicality.

[0051] Exemplarily, the prompts input to the AI judge can be as follows: <System> Please act as a reviewer with professional legal knowledge and evaluate the quality of the answers provided by two AI assistants to the following user prompt. Your task is to judge which assistant's answer is better based on the understanding of Chinese legal issues.

[0052] Before starting the evaluation, compare the answers of the two assistants. The judgment criteria include considering whether the assistants' answers are legally professional, relevant and concise. Being legally professional means correctly interpreting Chinese legal concepts, quoting appropriate and accurate legal provisions (such as "Article X of the Law of XXX"), and drawing conclusions through rigorous legal reasoning. The answer should reflect an in-depth understanding of the legal system and specific legal issues. Relevance means that all parts of the answer are closely related and fully target the specific legal issue of the user prompt, avoiding irrelevant expansions or misleading information. Conciseness means that the answer is clear and not verbose or excessive.

[0053] Then, consider the creativity and novelty of the assistants' answers if necessary. Finally, identify any important information missing from the assistants' answers that would be beneficial for responding to the user prompt.

[0054] After providing the explanation, you must only output one of the following as the final judgment result and mark the label. The labels can be as the following 3 examples: 1. Assistant A is better: [[A>B]] 2. Tie, equal in quality: [[A=B]] 3. Assistant B is better: [[B>A]] Example output: "My final ruling is a tie: [[A=B]]".

[0055] < / System> <|User Prompt|> {User Prompt} <|The Start of Assistant A’s Answer|> {Answer of A} <|The End of Assistant A’s Answer|> <|The Start of Assistant B’s Answer|> {Answer of B} <|The End of Assistant B’s Answer|> Through the above prompts sent to the AI judge, the AI judge can score the candidate question-and-answer pairs after permutation and combination to obtain a scoring list. For example, if there are four answers, named a, b, c, and d respectively, by permuting and combining them and inputting them into the AI judge in pairs (the order matters), 12 comparison results can be obtained. Suppose the comparison results are [a>b, a>c, a<d, b>c, b<a, b>d, c>a, c<b, c<d, d>a, d>b, d<c]. Then, by inputting this comparison result into the ELO dynamic scoring module based on the ELO-ranking technology, a scoring list can be obtained. Based on this scoring list, the optimal selection strategy for the candidate question-and-answer pairs can be determined to further obtain the optimal answer.

[0056] Using the ELO dynamic scoring module to optimize the candidate questions and answers can dynamically adjust the scores of the question-and-answer pairs by comparing the answers pairwise, ensuring that the finally selected question-and-answer pairs have higher logic, relevance, and language expression quality.

[0057] Exemplarily, for the comparison sequence generated by the question-and-answer quality evaluation system in the AI judge, the ELO dynamic scoring module can assign an initial score to all candidates in the comparison sequence. Suppose the initial score is set to 1500. For example: a = 1500, b = 1500, c = 1500, d = 1500. Then, according to the order of the comparison sequence, compare them pairwise and gradually update the comparison result list after comparison.

[0058] Exemplarily, taking the candidate question-and-answer pairs A and B (abbreviated as candidate A and candidate B) composed of candidate questions and answers as examples, the expected winning rate of each comparison can be calculated according to the following ELO formula, and then the corresponding scores can be updated: (1)Expected winning rate (E) For each question-and-answer pair composed of candidate questions and answers, the expected winning rate E can be calculated through the ELO algorithm, and can be calculated by the following formula:

[0059]

[0060] Among them, R A and R B are the current scores of candidate A and candidate B respectively (which can be understood as the current scores of the candidate question-and-answer pairs), E A and E B are the expected winning rates of candidate A and candidate B. This expected winning rate E adopts a dynamic adjustment based on logarithmic ratio, and evaluates the results through the non-linear mapping between the score gap and the expected winning rate. This dynamic adjustment process of logarithmic ratio can be based on the distribution of the Logistic function of logistic regression, so as to more effectively describe the non-linear relationship of the winning rate.

[0061] If the score of candidate A is higher than that of candidate B, then the probability of candidate A winning will be greater. At this time, the value of E A will be closer to 1, and E B will be close to 0; vice versa.

[0062] If R A =R B , then, E A =E B The expected winning rates of candidate A and candidate B are both 64%; If R A is 400 points higher than R B , then the expected winning rate of candidate A is 90%; If R A is 200 points higher than R B , then the expected winning rate of candidate A is 76%.

[0063] (2)Update score According to the difference between the above score results and the expected winning rate, the scores of both candidate A and candidate B can be updated through the following formula:

[0064]

[0065] Among them, R’ Aand R' B is the updated score, S A and S B can characterize the comparison result between candidate A and candidate B, and the value can be 0 (failure), 0.5 (draw), 1 (victory). If candidate A wins, S A = 1, S B = 0; if candidate B wins, S A = 0, S B = 1; if it's a draw, S A = S B = 0.5. EA and EB can characterize the expected winning rates of candidate A and candidate B, and K is an adjustment factor that controls the magnitude of score changes. A larger K will result in larger score fluctuations, while a smaller K will result in smaller score fluctuations. Exemplarily, the value of K can be but is not limited to being between 10 and 40.

[0066] Assume that the following steps are processed sequentially [a>b, a>c, a<d, b>c, b<a, b>d, c>a, c<b, c<d, d>a, d>b, d<c] and other comparison results, and update the corresponding scores after each calculation. After multiple rounds of iteration, when the scores tend to be stable, a final ranking can be obtained, and the answer with the highest score is taken as the best answer. Here, the threshold of the maximum score change can be set to 0.01. When the score changes of all objects in the current round of iteration are less than the threshold 0.01, the scores tend to be stable and the iteration stops.

[0067] The candidate question and answer selection strategy based on ELO dynamic scoring can automatically and dynamically screen high-quality questions and answers from a large number of candidate contents, ensuring that the finally selected contents meet the optimal standards in terms of logic, relevance, and expression quality.

[0068] After each round of comparison, the candidate questions and the best questions can be correspondingly stored in a structured database, and the fields can be but are not limited to including at least one of the following: question ID, relevant article ID, generated question, generated answer, and score details. This structured database can provide a standardized API interface to support the calling and retrieval of generated contents by other large language models.

[0069] The embodiments of this application can use the sample data in the above-mentioned structured database to train a large language model to obtain an intelligent legal question and answer system, complete the automatic answering of legal questions asked by users, thereby providing legal consultation services, and use the legal data generated by the large language model for system self-optimization to improve the accuracy and relevance of question and answer. This solution can be used to automatically generate relevant legal document drafts (contract drafts, judgment drafts), etc.

[0070] Based on the same inventive concept, an embodiment of the present application further provides an electronic device, such as Figure 2 shown, the electronic device includes at least one processor, at least one memory, an input device, and a display. The memory can be used to store a computer program, and the computer program can include instructions and data to implement the steps of any of the above methods.

[0071] The input device can include, but is not limited to, at least one of the following devices: touch screen, microphone, keyboard, mouse, image sensor, camera, radar; Legal issues can include, but are not limited to, at least one of the following data forms: text data, audio data, images, videos, radar data.

[0072] The memory can be a random access memory, read-only memory, non-volatile, programmable ROM, erasable PROM, electrically erasable, flash memory, optical memory, and registers, etc. The processor can be a general-purpose processor. The general-purpose processor can be a processor that executes specific steps and / or operations by reading and executing the computer program stored in the memory. The general-purpose processor may use the memory stored during the execution of the steps and / or operations. The general-purpose processor can be a central processing unit, ASIC, and FPGA, etc. During implementation, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor or the instructions in software form. The method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware processor, or executed and completed by the combination of the hardware and software modules in the processor.

[0073] Exemplarily, the input device includes, but is not limited to, at least one of a keyboard, a touch panel, a voice input device, and an image sensor.

[0074] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, fiber optic, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a solid-state disk (SSD), etc.

[0075] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0076] Each embodiment in this specification is described in a related manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0077] The above are only the preferred embodiments of the present application and are not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application are included in the protection scope of the present application.

Claims

1. A method for generating sample data, characterized in that: The method comprises: Generate multiple legal questions based on the target legal provisions and prompt words through the first language model; Inputting the plurality of legal questions into a second language model to determine candidate questions; The first language model generates a plurality of answers based on the candidate question, the target legal provision and the prompt word; The multiple answers and the candidate questions are input into the second largest language model to obtain answer scores of the multiple answers, and the answer scores are compared based on ELO-ranking technology to obtain the best question-answer pair as sample data.

2. The method according to claim 1, characterized in that The plurality of legal questions are input into the second language model to determine candidate questions, including: Inputting the plurality of legal issues into the second largest language model, and obtaining problem scores for the plurality of legal issues based on the legal provisions compliance and expression clarity of the plurality of legal issues; The problem scores are compared based on the ELO-ranking technology, and the legal issues with the highest problem scores are selected as candidate issues.

3. The method according to claim 1, characterized in that The first language model adopts the FT-Llama3 framework, and the second language model adopts the RGA framework.

4. The method according to claim 1, characterized in that The first language model generates multiple answers based on the candidate question, the target legal provision and the prompt word, including: By adjusting the generation parameters of the first large language model, the first large language model generates multiple answers according to the candidate question, the target legal provision and the prompt word, wherein the generation parameters include at least one of a temperature parameter, a random seed, a Top-K sampling parameter and a Top-P sampling parameter.

5. The method according to claim 4, characterized in that When the generation parameter includes a temperature parameter, the temperature parameter is set between 0.7 and 1.

5.

6. The method according to claim 1, characterized in that The first language model generates multiple answers based on the candidate question, the target legal provision and the prompt word, including: The value of the temperature parameter in the softmax function of the first language model is 0.

9. In each generation cycle, the random seed in the first large language model is randomly adjusted, and the first large language model combines the candidate question, the target legal provision and the prompt word to generate diversified answers based on the random seed.

7. The method according to claim 1, characterized in that The best question-answer pair is obtained by comparing the answer scores based on the ELO-ranking technology, including: Obtaining the question score of the candidate question and the answer scores of the multiple answers, and integrating the scores of the legal question and the corresponding answers by using preset weights; Combined with the integrated scores, the ELO-ranking technology is used to rank the question-answer pairs consisting of legal questions and corresponding answers, wherein the question scores and answer scores include at least one of: compliance with legal provisions, accuracy of questions and answers, logic of questions and answers, and clarity of expression; The question-answer pair with the highest score is taken as the best question-answer pair.

8. The method according to claim 1, characterized in that Inputting the multiple answers and the candidate questions into the second largest language model to obtain answer scores for the multiple answers includes: The second language model analyzes the consistency of the legal provisions between the multiple answers and the candidate questions according to the target legal provisions; The second language model analyzes the precedent relevance of the multiple answers based on the target legal provision and its own knowledge base; The second largest language model analyzes the logical coherence of the multiple answers according to the contextual association relationship between the multiple answers and the candidate questions; For each answer and the candidate question, the second largest language model is used to integrate the legal consistency, precedent relevance and logical coherence of the question and answer pair to obtain a score for the question and answer pair.

9. A question-answering method based on a large language model, characterized in that: include: Inputting the legal question and the prompt word into the large language model through an input device, wherein the large language model is trained by the sample data generated by the method according to any one of claims 1 to 8; The answer output by the large language model is displayed through a display device.

10. An electronic device, characterized in that: include: Input devices for collecting legal issues; At least one processor and at least one memory, wherein the at least one memory can be used to store a computer program, and the at least one processor can execute the computer program to implement the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method and device for generating question and answer pairs

    CN116881470A

  • Intelligent trademark law question and answer method for enhancing big language model reasoning through knowledge graph

    CN117807202A

  • Basic large model optimization method applied to legal field

    CN118153714A

  • Large language model reliable legal question and answer generation method based on knowledge fine tuning

    CN118210891A

  • Question and answer pair generation method and system based on large language model

    CN118332086A

Cited By

  • LLM-RAG sample construction method and device, and storage medium

    CN120705581A

  • Knowledge graph query method and device based on multi-language model result fusion

    CN121636662A

  • Knowledge graph query method and device based on multi-language model result fusion

    CN121636662B