Model training method, legal issue analysis method and electronic equipment

By reconstructing legal issues into retrieval-style sample data and training an autoregressive base model, combined with QLoRA fine-tuning technology and cross-entropy loss function, the accuracy problem of large language models in legal consultation is solved, achieving more accurate and reliable legal issue analysis.

CN120105108BActive Publication Date: 2025-09-09HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510578269.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-09-09
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

Traditional large language models lack accuracy, reliability, and authority when providing legal consulting services, and are unable to effectively handle highly professional legal issues.

Method used

By reconstructing the generative sample data of legal issues into retrieval sample data, combining QLoRA fine-tuning technology and the cross-entropy loss function with a neglect loss mechanism, training the autoregressive base model and retrieval model, and using the Faiss database to store legal knowledge, active retrieval and accurate analysis are achieved.

Benefits of technology

It improves the accuracy and rigor of large language models in legal issue analysis, reduces misleading and hallucinatory phenomena, and enhances the reliability and authority of responses to legal issues.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105108B_ABST
    Figure CN120105108B_ABST
Patent Text Reader

Abstract

This application discloses a model training method, comprising: inputting generative sample data of legal questions into a large language model and reconstructing it into retrieval sample data; inputting the retrieval sample data into an autoregressive base model and training the autoregressive base model based on the QLoRA fine-tuning technique, wherein the training loss function of the autoregressive base model adopts a cross-entropy loss function with a neglect loss mechanism; and using the output probability distribution of the autoregressive base model as the target distribution, training a retrieval model based on the retrieval sample data. This application can solve the problem of low accuracy of large language models in answering legal questions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a model training method, a legal issue analysis method, and an electronic device. Background Art

[0002] Due to the problems of low efficiency, high cost and narrow coverage of traditional legal consulting services, traditional legal consulting services can no longer meet the growing demand for legal services.

[0003] Therefore, people have also begun to use artificial intelligence technologies, such as LLM (Large Language Model), to provide legal consulting services. However, since legal issues are often highly professional, large language models are prone to hallucinations or incorrect reasoning. Therefore, the current large language models lack accuracy, reliability, and authority. Summary of the Invention

[0004] The purpose of this application is to provide a model training method and a legal issue analysis method to solve the problem that large language models do not provide accurate answers to legal questions.

[0005] In a first aspect, an embodiment of the present application provides a model training method, comprising:

[0006] Generative sample data of legal issues are input into a large language model and reconstructed into retrieval sample data;

[0007] Inputting the retrieval sample data into an autoregressive base model, and training the autoregressive base model based on the QLoRA fine-tuning technology, wherein the training loss function of the autoregressive base model adopts a cross entropy loss function with a neglect loss mechanism;

[0008] The output probability distribution of the autoregressive base model is used as the target distribution, and the retrieval model is trained based on the retrieval sample data.

[0009] Optionally, before inputting the generative sample data of legal questions into the large language model and reconstructing them into retrieval sample data, the method further includes:

[0010] The large language model is combined with a three-dimensional one-hot vector to determine the type of the legal question, wherein the types of the legal question include: generative legal questions and retrieval legal questions.

[0011] Optionally, the large language model combines the three-dimensional one-hot vector to determine the type of the legal issue, including:

[0012] Determine whether the legal issue needs to be searched using the large language model:

[0013] When the legal question does not need to be searched, determining that the type of the legal question is a generative legal question;

[0014] When the legal issue needs to be searched, it is determined based on the position index of the three-dimensional unique-hot vector whether the search type of the legal issue belongs to legal article search and / or similar case search.

[0015] Optionally, reconstructing the generative sample data of legal issues into retrieval sample data includes:

[0016] The generative sample data is reconstructed according to a dictionary format based on the message key, and a backtracking label and a prompt word are added to the generative sample data to obtain the retrieval sample data.

[0017] Optionally, the autoregressive base model is trained based on the QLoRA fine-tuning technology, wherein the training loss function of the autoregressive base model adopts a cross-entropy loss function with a neglect loss mechanism, including:

[0018] Based on the QLoRA fine-tuning technology, the user role model role data and the backtracking label-wrapped data of the retrieval sample data are obtained, and the autoregressive base model is trained to learn the conditional probability distribution of the retrieval sample data for the purpose of generating the minimum cross entropy loss. When the autoregressive base model determines that the representation of the retrieval sample data does not require retrieval, the retrieval model is not trained.

[0019] Optionally, the step of using the output probability distribution of the autoregressive base model as a target distribution and training the retrieval model based on the retrieval formula sample data includes:

[0020] Performing negative example mining on the search-type sample data;

[0021] Using the positive and negative examples of the retrieval formula sample data to perform positive and negative example scoring on the retrieval model to obtain positive and negative example scoring results;

[0022] Based on the positive and negative example scoring results and combined with KL divergence loss, the retrieval model is trained to learn the output probability distribution of the autoregressive base model.

[0023] Optionally, the step of training the retrieval model to learn the output probability distribution of the autoregressive base model based on the positive and negative example scoring results in combination with KL divergence loss includes:

[0024] Combining the results of positive and negative example scores, forward propagation is performed on the positive example samples and the negative example samples;

[0025] Matching the output probability distribution of the retrieval model and the autoregressive base model based on KL divergence loss, retaining the task knowledge learned by the retrieval model through the positive and negative samples;

[0026] The retrieval model and the encoding of the legal provisions in the task knowledge are stored in the Faiss database of the large language model.

[0027] In a second aspect, embodiments of the present application provide a method for analyzing legal issues, including:

[0028] Inputting legal issues into the large language model;

[0029] The large language model encodes the legal issue based on the prompt word and generates a backtracking label, wherein the backtracking label indicates whether the legal issue needs to be retrieved;

[0030] If the backtracking tag represents the legal issue that needs to be searched:

[0031] Invoking a corresponding retrieval model based on the backtracking label to analyze the legal issue;

[0032] If the backtracking tag indicates that the legal issue does not need to be retrieved, the large language model is directly used to analyze the legal issue.

[0033] Optionally, calling a corresponding retrieval model based on the backtracking tag to analyze the legal issue includes:

[0034] Invoke a search model corresponding to the backtracking label, and determine, based on a position index combined with the three-dimensional one-hot vector, whether the search type of the legal issue is a legal article search and / or a similar case search;

[0035] Based on the retrieval type of the legal issue, a corresponding retrieval model is called to analyze the legal issue.

[0036] In a third aspect, an embodiment of the present application provides an electronic device, including:

[0037] Input device for collecting legal issues;

[0038] At least one processor and at least one memory, the at least one memory can be used to store a computer program, and the at least one processor can execute the computer program to implement the above method.

[0039] Optionally, the input device includes at least one of the following devices:

[0040] Touch screen, microphone, keyboard, mouse, image sensor, camera, radar;

[0041] The legal issue includes at least one of the following data forms:

[0042] Text data, audio data, images, videos, radar data.

[0043] The embodiment of the present application constructs a large language model capable of active retrieval through the reconstruction of legal issues and the autoregressive base model, and realizes active retrieval through QLoRA fine-tuning technology and a cross-entropy loss function with an ignore loss mechanism, thereby reducing the probability of the model giving inaccurate or irrelevant answers, reducing the probability of being misled and having hallucinations, and improving the rigor and accuracy of answers to legal issues. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 A schematic diagram of a model training process provided in an embodiment of the present application;

[0045] Figure 2 A schematic diagram of a model generation process provided in an embodiment of the present application;

[0046] Figure 3 A schematic diagram of a sample data partitioning method provided in an embodiment of the present application;

[0047] Figure 4 A schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0048] The present application will be described in detail below in conjunction with the specific embodiments shown in the accompanying drawings, but these embodiments do not limit the present application. Structural, methodological, or functional changes made by ordinary technicians in this field based on these embodiments are included in the scope of protection of the present application.

[0049] If we rely solely on the reasoning ability of the large language model to answer questions related to legal precedents, it will often be difficult to understand the specific precedents, and it will be easy to produce hallucinations or erroneous reasoning. It may even happen that all references and adaptations of legal provisions rely on the reasoning ability of the large language model itself, and cannot be linked to the correct legal provisions.

[0050] Please refer to Figure 1 , an embodiment of the present application provides a model training method, comprising the following steps:

[0051] S101: Inputting the generative sample data of legal issues into the large language model and reconstructing them into retrieval sample data.

[0052] S102: Inputting the retrieval sample data into the autoregressive base model, and training the autoregressive base model based on the QLoRA fine-tuning technology, wherein the training loss function of the autoregressive base model adopts a cross entropy loss function with a neglect loss mechanism;

[0053] S103: Using the output probability distribution of the autoregressive base model as the target distribution, and training the retrieval model based on the retrieval sample data.

[0054] For example, let's assume we're using LawBench's 20 tasks to score positive and negative examples. These 20 LawBench tasks are divided into three areas: legal knowledge memorization, legal knowledge comprehension, and legal knowledge application. Each of these three areas is converted into the training data format of a large language model, where the training data format includes at least a generative data format and a retrieval data format. In other words, the generative sample data in LawBench's 20 tasks uses the generative data format, while the retrieval sample data uses the retrieval data format.

[0055] Exemplarily, the above-mentioned legal knowledge memorization may include: legal article recitation, knowledge question and answer; the above-mentioned legal knowledge comprehension may include: document proofreading, dispute focus identification, problem topic identification, reading comprehension, named entity recognition, argument mining, event detection, trigger word extraction, etc.; the above-mentioned legal knowledge application may include: legal article prediction, case analysis, consultation, etc.

[0056] Exemplarily, the generated data format may adopt a dictionary format, which includes a messages key, and the key value of messages may be a list, representing conversation data.

[0057] For example, the search-based data format can add a backtracking tag to the conversation data based on the generative data format. This backtracking tag can indicate that legal data needs to be retrieved to answer the question during the model training process. The backtracking tag can also carry a prompt word, which is used to inform the large language model whether a search is required and when to initiate the search.

[0058] In the embodiment of the present application, an optional search formula data format may be as follows:

[0059] {

[0060] "messages":[

[0061] {"content":'Role: Legal Consulting Expert\nTool: Use the search model to <retrieval>Tag call, retrieval results are used <retrieval>< / retrieval> Package returned. \nGoals: Provide accurate answers based on authoritative legal information and solve user problems promptly. \nWorkflow:\nUnderstand user questions. \nDetermine whether to search:\nIf you need to search, use <retrieval>Tag calls the search model. \n If not needed, answer directly. \n Answer based on the search results, referencing meta_data <reference>< / reference> Package. \nRepeat the process of determining whether further search is needed until the answer is complete. ',"role":"system"}

[0062] The above messages can be interpreted as follows:

[0063] The content includes: the tool role is a legal consulting expert; the backtracking tag is used to call the external retrieval model and return the call results; nGoals represents the accurate answer based on authoritative legal information; Workflow is used to understand the user's question and determine whether to call the external retrieval model. If the judgment result is that the external retrieval model needs to be called, the external retrieval model is called through the above Retrieval tool and called through meta_data <retrieval>< / retrieval> The package returns the search results, and Workflow then determines whether it needs to further call the external search model based on the returned search results. Until Workflow determines that it is no longer necessary to call the external search model, it combines the call results and the accurate answer of nGoals to give the answer to the question.

[0064] Sample documents can be real legal case documents publicly available online. Based on the prompt words, the sample documents are rewritten into a conversational format suitable for large language model training. Positive and negative examples are collected through positive and negative example mining for model training.

[0065] For example, meta-llama / Meta-Llama-3-8B-Instruct, Qwen / Qwen2.5-7B-Instruct, etc. in the large language model can be used as the autoregressive base model. Based on the retrieval sample data, the QLoRA fine-tuning technology is used to fine-tune the autoregressive base model on the legal field data, and the legal knowledge in the retrieval sample data is learned so that the generated data satisfies the cross-entropy loss function with the neglect loss mechanism, and the training of the autoregressive base model is completed, thereby obtaining the generative model.

[0066] In an embodiment of the present application, the autoregressive base model may be a large language model, which uses autoregression (AR) to generate sample data in the above-mentioned training process, and thus may generate a data sequence of sample data based on probability distribution.

[0067] Assume that the AR method is used to generate the data sequence of sample data, and the sample data sequence is , the goal of training a generative model can be to learn the conditional probability distribution of the sample data sequence ,in In this way, when generating text, the generative model predicts the probability of the next word based on the previous words that have been generated.

[0068] Cross-entropy is a loss function used to measure the difference between two probability distributions. In this embodiment of the present application, the cross-entropy loss function with the neglect loss mechanism can be ,in, One-hot encoding vector format can be used to represent the true value in the i-th sample data, and the length of the sample data is n; The probability distribution representing the i-th question predicted by the large language model has a length of n. For example, if the vocabulary size is V, then at each step, the model has to predict the next word X t , which is equivalent to selecting a category from V categories. In this embodiment of the application, the search formula sample data can be selected as the true value, using the user role data and <retrieval>< / retrieval> The label-wrapped data is used to train the autoregressive base model with minimum cross entropy loss to obtain the generative model.

[0069] The above retrieval model can use BAAI / bge-large-en-v1.5, BAAI / bge-m3, etc. as the base model, using the FlagEmbedding and contrastive learning frameworks. The output probability distribution of the above autoregressive base model is used as the target distribution, and the base model of the retrieval model is trained based on the query sample data.

[0070] In the embodiment of the present application, the retrieval model can be trained through the following model generation process.

[0071] Mining negative example documents from sample documents;

[0072] The pre-trained model is scored for positive and negative examples using the positive and negative examples in the sample documents;

[0073] Combined with the results of positive and negative example scoring, the pre-trained model is trained by output probability distribution matching, retaining the task knowledge learned by the pre-trained model through positive and negative samples to obtain a knowledge model;

[0074] The legal provisions corresponding to the above task knowledge are encoded, and the knowledge model and the corresponding encoding results are stored in the Faiss database of the large language model.

[0075] Reference Figure 2 In an optional embodiment of the present application, the method for analyzing legal issues may include:

[0076] The user inputs a question into the large language model, which transmits the question to the generative model (obtained after training the autoregressive base model) and the retrieval model respectively. The generative model generates a backtracking label, which instructs the retrieval model to call. The retrieval model queries the question code in the Faiss database for question information based on the question. If the question code cannot be found, the text is sent to the generative model.

[0077] Negative example mining can be implemented in the following ways:

[0078] For a given sample data set, a document is selected from all sample data as a positive example, and K documents similar to this document are mined as hard negative examples. Because of their semantic similarity, positive and hard negative examples provide high-quality data comparison for retrieval model training. Through forward propagation, these hard negative examples are provided for retrieval model training, increasing the learning difficulty of the retrieval model and improving its semantic expression capabilities.

[0079] Exemplarily, the retrieval model can be generated by the following process:

[0080] Assume there is a pre-trained model M, which corresponds to a set of positive document D. , forward propagation is performed on all documents in D using M, and a set of n*m dimensional vectors V is obtained. Similarly, the legal question q is converted into an m-dimensional vector q through M. dev For the m-dimensional vector q corresponding to the legal problem q vec and the vectors in the vector group V Calculated using the cosine similarity formula:

[0081]

[0082] in, represents the dot product of two vectors, where represents the L2 norm of the vector. Using the above formula, we can obtain the similarity between legal question q and each document vector in the sample document set D, denoted as S. Next, we find the top K indices with the largest similarity values, which serve as the K negative examples to be mined. The corresponding index sj is the negative example score. Similarly, we can calculate the positive example score for the positive document q and q.

[0083] Finally, for a (q, d) pair, (q, d, N, S) is obtained through negative example mining and positive and negative example scoring, where , , In, n i represents the i-th negative sample, s i Represents the pre-trained retrieval model M for (q,n i ) rating. Represents the score of (q, d) under the pre-training model M.

[0084] Since the embodiment of the present application uses BAAI / bge-large-zh-v1.5 as the base model of the retrieval model, it can align questions and answers well. However, if the model is used directly to continue answering legal questions, it will face a catastrophic forgetting problem (when the model learns a new task, the network parameter update will cause the model to forget the tasks it has learned before). Therefore, the output probability distribution of the old model can be used as the target distribution, and the output probability distribution of the new model under the same input can be used as the basis for fine-tuning to minimize the cross-entropy loss, so that the output probability distribution of the new model is as close as possible to the output probability distribution of the old model, thereby retaining the knowledge about the original task in the old model.

[0085] For example, for a data sample (q, d, N, S), during model training, the KL divergence of the data sample (q, d, N, S) can be calculated as the KL divergence loss for model training in the following manner:

[0086] For a data sample (q, d, N, S), perform Softmax normalization on the positive and negative sample scores, and update S to P so that ,in , s i Represents the pre-trained retrieval model M for (q,n i ) rating. Represents the score of (q, d) under the pre-training model M.

[0087] In this embodiment, the softmax function can be used to convert any real number vector into a probability distribution, that is, each element of the output vector is between 0 and 1, and the sum of all elements is 1, which can ensure the validity of the output probability distribution of the model, and further ensure that the input used to calculate the KL divergence is a legal probability distribution, which satisfies the basic properties of probability, making the calculation of the KL divergence of the model training meaningful; in the calculation process, the softmax function helps to improve numerical stability. By performing exponential operations on each element of the input vector and then normalizing it, it avoids the numerical underflow or overflow problems that may occur when directly calculating the probability distribution. Especially when processing large-scale data or complex models, the numerical stability can ensure the accuracy and reliability of the calculation results; in the model of the embodiment of the present application, the softmax function is used to generate the probability distribution of the category, which can be directly compatible with the output form of the model, and is convenient for calculating the loss function such as KL divergence during the model training and evaluation process, which is used to measure the difference between the model prediction distribution and the true distribution, thereby guiding the parameter update and optimization of the model; after the softmax function converts the original score into a probability distribution, the result has a clear probability meaning, that is, each element represents the probability of the corresponding category or event occurring, so that the softmax-based The KL divergence calculated from the normalized probability distribution is easier to interpret and can intuitively reflect the probability differences between the two distributions in each category, which helps to analyze the model's prediction results and perform model diagnosis.

[0088] The above example uses a pre-trained model to score positive and negative examples. The following example uses the scoring to train the retrieval model.

[0089] Use the retrieval model to perform a forward propagation on q, d, and N respectively, and calculate the similarity between their respective "d, N and q". The obtained similarity is also normalized using Softmax to obtain the normalized similarity Q of the model-inferred d and N with respect to q.

[0090] Calculate the KL divergence loss, the formula is as follows:

[0091] ;

[0092] For example, suppose P is (0.1, 0.2, 0.7) and Q is (0.2, 0.3, 0.5);

[0093] Then D kl = 0.1*log(0.1) + 0.2*log(0.2) + 0.7 * log(0.7) – 0.1 * log(0.2) – 0.2 * log(0.3) – 0.7 * log(0.5).

[0094] In the process of learning new tasks, constraints based on KL divergence can limit the extent of model parameter updates, prevent excessive changes to parameters related to old tasks due to overfitting to new tasks, ensure that model parameters do not change drastically when learning new tasks, maintain the stability of the model, and enable the model to maintain relatively consistent performance in learning different tasks.

[0095] This embodiment of the application utilizes negative example mining, positive and negative example scoring, and a cross-entropy loss function with a neglect loss mechanism to enable large language models to achieve excellent retrieval performance in the legal retrieval vertical. After the retrieval model is trained, the legal provisions corresponding to the legal issues are encoded using the Faiss vector database and stored in the Faiss database, preparing for subsequent legal provision retrieval using the large language model.

[0096] The embodiment of the present application implements a legal problem analysis method, including

[0097] Inputting legal issues into the large language model;

[0098] The large language model encodes the legal issue based on the prompt word and generates a backtracking tag, which indicates whether the legal issue needs to be searched.

[0099] If the backtracking tag represents the legal issue that needs to be searched:

[0100] Calling a corresponding retrieval model based on the backtracking tag to analyze the legal issue;

[0101] If the label represents the legal issue and no retrieval is required, the large language model is directly used to analyze the legal issue.

[0102] Optionally, a corresponding retrieval model is called based on the backtracking tag to analyze the legal issue, including:

[0103] The search model corresponding to the backtracking label is called, and based on the position index of the three-dimensional one-hot vector, it is determined that the search type of the legal issue is a legal article search and / or a similar case search;

[0104] Based on the retrieval type of the legal issue, the corresponding retrieval model is called to analyze the legal issue.

[0105] For example, when the generative model generates <retrieval>When the tokens corresponding to the backtracking tag are retrieved, the large language model calls the retrieval model. The backtracking tag representation needs to call the retrieval model encoding. The retrieval model encodes the user question and uses the encoding result to search the prepared faiss database using the vector cosine similarity. When the retrieval model retrieves information with a cosine similarity that meets the threshold, the legal provisions in the retrieved information are spliced ​​into <retrieval> After the position, and< / retrieval> To end, use the backtracking label and the following< / retrieval> The tag wraps the retrieved legal information. This reconstructs the data generated by the generative model, introducing relevant legal provisions after the backtracking tag. This helps the large language model achieve more legally sound responses to legal questions.

[0106] In the embodiment of the present application, the large language model is combined with the three-dimensional one-hot vector to determine the type of legal issue, which can be achieved in the following way:

[0107] Use expert models and large language models to determine whether the legal issue needs to be searched:

[0108] When the legal question does not need to be searched, determining that the type of the legal question is a generative legal question;

[0109] When the legal issue needs to be searched, it is determined based on the position index of the three-dimensional unique-hot vector whether the search type of the legal issue belongs to legal article search and / or similar case search.

[0110] For example, the sample question can be evaluated by N expert models and M pre-trained large language models, and N+M judgment results can be summarized to determine whether the sample question needs to be retrieved. For example, the N+M judgment results can be summarized in the following way:

[0111]

[0112] Here, E(i) and L(i) represent the scores of the expert model and the large language model, respectively. A score of 1 indicates that a search is necessary, and a score of 0 indicates that a search is not necessary. For a sample question, if the aggregated score is greater than 0.5, it indicates that the question requires a search and the retrieval model is invoked. Otherwise, no further search is required. In this case, the large language model can directly generate the answer to the question.

[0113] If the aggregated score is not greater than 0.5, further search can be performed in the following ways:

[0114] The three-dimensional one-hot vector determines whether the legal question belongs to a search type: a legal article search, a similar case search, or both. For example, the three-dimensional one-hot vector R can use a format such as (0,1,0) or (0,0,1). When a category is determined, the value of that position is 1, and the value of the other positions is 0.

[0115] For example, the search type score of a legal issue can be evaluated in the following way:

[0116]

[0117] Where N represents the number of expert models, M represents the number of large language models, i represents the i-th position index, Indicates that the three-dimensional unique hot vector corresponding to the i-th position index corresponds to a coordinate, represents the three-dimensional vector given by the expert, The three-dimensional vector representation output by the large language model can be used to find the search type to which the question belongs through limitation. The coordinate with the highest score is the type.

[0118] By calculating the score value corresponding to the position index of each retrieval type through the three-dimensional unique-hot vector, and taking the retrieval type corresponding to the retrieval type with the largest score value as the retrieval type to which the legal issue belongs, the type of legal issue can be predicted more accurately, thereby improving the accuracy of subsequent problem analysis.

[0119] In the embodiment of the present application, the prompt word can include one of the following three search types: "legal article search", "similar case search" or "legal article and similar case search". The large language model models the prompt word and question, and constructs , then, predict the next value through the autoregressive base model This value is the value marked by the expert model and the pre-trained large language model as to whether retrieval is required and the corresponding retrieval type.

[0120] For example, the above process can be implemented by the following formula. The learning goal of the autoregressive base model is to maximize the probability P(X), so that the autoregressive base model can reflect on whether the question needs to be retrieved through the question. The maximum probability P(X) is expressed as:

[0121]

[0122] Among them, θ represents the parameters of the generative model, and T represents the Markov sequence inference.

[0123] In the embodiment of the present application, the training data of the large language model can be divided into three parts, represented by role, namely: system, user, and assistant.

[0124] {

[0125] "messages":[

[0126] {"content":'You are a legal assistant with professional legal skills. You need to determine whether a legal question requires a search. If the question does not require a search, please input and answer the question directly; if a search is required, please select one of the following three search types based on the question: "Legal Article Search", "Similar Case Search" or "Legal Article and Similar Case Search". When you input the corresponding search type, the search engine will return the relevant content to you and wrap it with the corresponding end mark. Please use the search engine content to answer the question reasonably. And use the legal content used. <reference> and< / reference> Meta_data of the legal content used by the package',"role":"system"}

[0127] In a large language model, sample data can be divided into Figure 4 The corresponding five parts are prompt words, user questions, retrieval labels, retrieval content and model output.

[0128] The cross entropy loss function with the ignore loss mechanism is:

[0129] ;

[0130] ;

[0131] Among them, V represents all the words in the sample data, t represents the t-th moment of model training, which can be understood as the training time position of the t-th word in the training data. The true value of the sample data at the tth moment is the i-th value in V. The prediction at time t of the representation model is probability. Indicates whether the loss should be masked at the t-th moment. The meaning of the formula can be combined Figure 3 understand, Moment representation: prompt words and user questions do not require calling the retrieval model. The subscript position representing the search content, The sequence length is k. The time step that represents the retrieval label. When reaching this moment, the loss needs to be calculated to allow the model to learn self-reflection. The previous content is used to determine whether the question needs to be retrieved. is the time step of the model output.

[0132] The cross-entropy loss function with a neglect loss mechanism can mitigate the hallucination problem of large language models by shielding the retrieval data and preventing the large language model from modeling the retrieval data. This prevents the large language model from memorizing legal data and subsequently generating untrusted legal provisions during reasoning. Furthermore, it reduces invalid gradient updates. Since the retrieval data is fixed text provided externally, the large language model itself does not need to learn the retrieval data, allowing it to focus more on developing conversational and reasoning capabilities.

[0133] The embodiment of the present application realizes active retrieval through QLoRA fine-tuning technology and a cross-entropy loss function with an ignore loss mechanism, and realizes the integration of legal knowledge through positive and negative example scoring and output probability matching based on KL divergence, thereby reducing the probability of the model giving inaccurate or irrelevant answers, reducing the probability of being misled and having hallucinations, and improving the rigor and accuracy of responses to legal issues.

[0134] Based on the same inventive concept, the embodiment of the present application further provides an electronic device, such as Figure 4 As shown, the electronic device includes at least one processor and at least one memory. The memory can be used to store a computer program. The computer program can include instructions and data to implement the steps of any of the above methods.

[0135] The input device includes at least one of the following devices:

[0136] Touch screen, microphone, keyboard, mouse, image sensor, camera, radar;

[0137] The legal issue includes at least one of the following data forms:

[0138] Text data, audio data, images, videos, radar data.

[0139] Illustratively, in an embodiment of the present application, legal questions and answers may be displayed on a display screen.

[0140] The memory may be a random access memory, a read-only memory, a non-volatile memory, a programmable ROM, an erasable PROM, an electrically erasable memory, a flash memory, an optical memory, a register, and the like. The processor may be a general-purpose processor, which may be a processor that performs specific steps and / or operations by reading and executing a computer program stored in the memory, and the general-purpose processor may use data stored in the memory in the process of performing the steps and / or operations. The general-purpose processor may be a central processing unit, an ASIC, an FPGA, and the like. During implementation, each step of the above method may be completed by an integrated logic circuit of hardware in the processor or by instructions in the form of software. The method disclosed in conjunction with the embodiments of the present application may be directly embodied as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor.

[0141] Exemplarily, the input device includes but is not limited to at least one of a keyboard, a touch panel, a voice input device, and an image sensor.

[0142] In the above embodiments, all or part of the embodiments can be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be magnetic media (e.g., floppy disk, hard disk, tape), optical media (e.g., DVD), or solid-state drive (SSD).

[0143] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprise," "include," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0144] Each embodiment in this specification is described in a related manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiment is generally similar to the method embodiment, so the description is relatively simple. For related parts, refer to the description of the method embodiment.

[0145] The above description is only a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application are included in the scope of protection of the present application.< / retrieval> < / retrieval>

Claims

1. A model training method, characterized in that: include: The large language model is combined with a three-dimensional one-hot vector to determine the type of legal question, wherein the types of legal questions include: generative legal questions and search-based legal questions; Determine whether the legal issue needs to be searched using the large language model: When the legal question does not need to be searched, determining that the type of the legal question is a generative legal question; When the legal issue needs to be searched, determining whether the search type of the legal issue belongs to legal article search and / or similar case search based on the position index of the three-dimensional one-hot vector; The generative sample data of the legal question is input into a large language model and reconstructed into retrieval sample data; the retrieval sample data is input into an autoregressive base model, and the model role data and the data wrapped by the backtracking label of the retrieval sample data are obtained based on the QLoRA fine-tuning technology. With the purpose of generating the minimum cross entropy loss, the autoregressive base model is trained to learn the conditional probability distribution of the retrieval sample data, wherein the training loss function of the autoregressive base model adopts a cross entropy loss function with a neglect loss mechanism, and the calculation formula of the cross entropy loss function with the neglect loss mechanism is: ; ; in, Represents all the words in the sample data, t represents the t-th moment of model training, which can be understood as the training time position of the t-th word in the training data. The true value of the sample data at the tth moment is The i-th value inside, The prediction at time t of the representation model is The probability of Indicates whether the loss should be masked at the t-th moment, Moment representation: prompt words and user questions do not require the use of retrieval models; The subscript position representing the search content, is the sequence length; Characterize the time step of the retrieved label; The time steps that represent the model output; Using the output probability distribution of the autoregressive base model as the target distribution, and training the retrieval model based on the retrieval formula sample data; When the autoregressive base model determines that the retrieval formula sample data representation does not require retrieval, the retrieval model is not trained.

2. The model training method according to claim 1, wherein: The step of inputting the generative sample data of the legal question into a large language model and reconstructing the data into retrieval sample data includes: The generative sample data is reconstructed according to a dictionary format based on the message key, and a backtracking label and a prompt word are added to the generative sample data to obtain the retrieval sample data.

3. The model training method according to claim 1, wherein: The method of using the output probability distribution of the autoregressive base model as the target distribution and training the retrieval model based on the retrieval formula sample data includes: Performing negative example mining on the search-type sample data; Using the positive and negative examples of the retrieval formula sample data to perform positive and negative example scoring on the retrieval model to obtain positive and negative example scoring results; Based on the positive and negative example scoring results and combined with KL divergence loss, the retrieval model is trained to learn the output probability distribution of the autoregressive base model.

4. The model training method according to claim 3, wherein: The step of training the retrieval model to learn the output probability distribution of the autoregressive base model based on the positive and negative example scoring results and in combination with KL divergence loss includes: Combining the results of positive and negative example scores, forward propagation is performed on the positive example samples and the negative example samples; Matching the output probability distribution of the retrieval model and the autoregressive base model based on KL divergence loss, retaining the task knowledge learned by the retrieval model through the positive and negative samples; The retrieval model and the encoding of the legal provisions in the task knowledge are stored in the Faiss database of the large language model.

5. A method for analyzing legal issues, characterized in that: The method comprises: Inputting the legal question into a large language model, wherein the large language model is trained by the method according to claim 1; The large language model encodes the legal issue based on the prompt word and generates a backtracking label, wherein the backtracking label indicates whether the legal issue needs to be retrieved; If the backtracking tag represents the legal issue that needs to be searched: Invoking a corresponding retrieval model based on the backtracking label to analyze the legal issue; If the backtracking tag indicates that the legal issue does not need to be retrieved, the large language model is directly used to analyze the legal issue.

6. The legal problem analysis method according to claim 5, characterized in that: The corresponding retrieval model is called based on the backtracking tag to analyze the legal issue, including: Invoke a search model corresponding to the backtracking label, and determine, based on a position index combined with the three-dimensional one-hot vector, whether the search type of the legal issue is a legal article search and / or a similar case search; Based on the retrieval type of the legal issue, a corresponding retrieval model is called to analyze the legal issue.

7. An electronic device, characterized in that: include: Input device for collecting legal issues; At least one processor and at least one memory, wherein the at least one memory can be used to store a computer program, and the at least one processor can execute the computer program to implement the method according to any one of claims 1 to 6.

8. The electronic device according to claim 7, wherein: The electronic device further includes an input device, wherein the input device includes at least one of the following devices: Touch screen, microphone, keyboard, mouse, image sensor, camera, radar; The legal issue includes at least one of the following data forms: Text data, audio data, images, videos, radar data.

Citation Information

Patent Citations

  • Large language model enhanced question and answer generation method

    CN118013051A

  • Border examination legal question and answer dynamic retrieval enhancement generation method and system

    CN118520093A

  • Intelligent question answering method and system based on multi-module collaborative optimization

    CN119557409A