Legal document generation method and system based on reading comprehension and intention recognition model

Through the legal document generation method based on reading comprehension and intent identification models, the problem of inaccurate answers to legal consultation in the prior art is solved, and high accuracy and professional legal document generation is achieved, and the performance and accuracy are significantly improved.

CN114297342BActive Publication Date: 2025-05-13CHONGQING DANIU COGNITIVE TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111501714.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-09
Publication Date
2025-05-13
Estimated Expiration
2041-12-09

AI Technical Summary

Technical Problem

The prior art is difficult to provide both accurate, professional and satisfactory answers to legal consultation, especially in the legal document generation system, and it is impossible to accurately understand the intentions and intentions of the parties.

Method used

The legal document generation method is adopted based on reading comprehension and intent recognition models. By obtaining long text statements, the first round of legal elements are extracted using the long text reading comprehension model, and the legal scenarios are obtained compared with the legal element library. The missing elements are obtained through multiple rounds of dialogue, and the results are finally entered into the intent recognition model to generate legal advice or contract documents.

Benefits of technology

The parties can obtain accurate, professional and satisfactory legal consultation related answers. Compared with traditional RNN and LSTM models, the performance and accuracy have been significantly improved, and the convergence time of the model has been reduced and the training cost has been reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114297342B_ABST
    Figure CN114297342B_ABST
Patent Text Reader

Abstract

This application proposes a method and system for generating legal documents based on a reading comprehension and intent recognition model, which belongs to the field of legal consulting technology. The method includes: obtaining a long text statement; inputting the long text statement into a long text reading comprehension model to obtain a preliminary round of legal elements; comparing the preliminary round of legal elements with the legal element library to obtain the legal scenario that the consultant wants to consult; initiating multiple rounds of dialogue with the consultant based on the missing elements to obtain the results of the multiple rounds of dialogue; inputting the results of the multiple rounds of dialogue into the intent recognition model to obtain the intent recognition results; making reasoning decisions to automatically generate a legal consulting opinion or contract document. The system includes: a data acquisition module, an element acquisition module, a scenario acquisition module, a multi-round dialogue module, an intent recognition module, and a document generation module. This application reduces the model convergence time, reduces the training cost, and enables the parties to obtain accurate and professional answers to legal consulting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of legal consulting technology, and specifically relates to a method and system for generating legal documents based on reading comprehension and intent recognition models. Background Art

[0002] In modern society, people often encounter legal issues in their daily lives, such as marriage, property inheritance, labor disputes, loans, etc. The parties often consult local lawyers or search for relevant answers online. If the parties consult lawyers, on the one hand, lawyers have different personal qualifications and different fields of in-depth research, so the answers given by each lawyer may be different, which makes it difficult for the parties to make a correct decision; on the other hand, there is a huge gap in the number of lawyers, and the number of existing lawyers is difficult to meet the growing number of legal consultation issues, which makes it difficult for the parties to obtain accurate, professional and satisfactory answers. If the parties search for relevant answers online, it is often difficult for the parties to match the process they are looking for with the process given online, and thus it is impossible to obtain accurate answers.

[0003] In the prior art, there are some legal document generation systems, through which the parties can conduct human-computer dialogue and finally obtain a legal opinion consultation letter or a contract document. However, the legal opinion consultation letter or contract document obtained through the system is not very accurate. The reason is that the system cannot accurately understand the parties' meanings and is not very clear about the parties' intentions. The text that the parties need to input is relatively long. The use of traditional machine learning methods and basic neural network models, such as RNN, LSTM, etc., is far from meeting the requirements in terms of performance and accuracy.

[0004] With regard to the problem that it is difficult for parties to obtain accurate, professional and satisfactory answers to legal consultation in the existing technology, no relevant solution has been found so far. Summary of the invention

[0005] In response to the above technical problems, this application proposes a legal document generation method and system based on reading comprehension and intent recognition model.

[0006] In the first aspect, the present application proposes a method for generating legal documents based on a reading comprehension and intention recognition model, comprising the following steps:

[0007] Get long text statements;

[0008] Inputting the long text statement into a long text reading comprehension model to obtain a preliminary round of legal elements;

[0009] Compare the initial round of legal elements with the legal element database to obtain the legal scenario that the consultant wants to consult;

[0010] Obtain missing elements according to the legal scenario, initiate multiple rounds of dialogue with the consultant based on the missing elements, and obtain the results of the multiple rounds of dialogue;

[0011] Input the results of multiple rounds of conversations into the intent recognition model to obtain the intent recognition results;

[0012] Based on the intention recognition results and the initial round of legal elements, reasoning and decision-making are carried out to automatically generate legal advisory opinions or contract documents.

[0013] The step of inputting the long text statement into the long text reading comprehension model to obtain the initial round of legal elements includes the following steps:

[0014] Inputting the long text statement into the extractive reading comprehension model and the multiple-choice reading comprehension model respectively, and obtaining the extractive reading comprehension answer and the multiple-choice reading comprehension answer respectively;

[0015] The answers to the extractive reading comprehension questions and the answers to the multiple-choice reading comprehension questions are used as the initial round of legal elements.

[0016] The steps of inputting the long text statement into the extractive reading comprehension model to obtain the extractive reading comprehension answer are as follows:

[0017] The long text statement and the question are concatenated and unified to a fixed length to obtain a concatenated result of a fixed length;

[0018] Encode the fixed-length concatenated result using the BERT pre-trained model, and extract the encoded feature vector;

[0019] Inputting the encoded feature vector into a classifier of an extractive reading comprehension model, and finally calculating the logical value of each position;

[0020] According to the logical value, the extractive reading comprehension answer is selected through the softmax function.

[0021] The BERT pre-trained model adopts a transfer learning strategy to train the model, which includes two stages, using two data sets for training, and both data sets are annotated; in the first stage, only the first data set is used for training to obtain a first BERT pre-trained model and a first model weight; in the second stage, the first model weight is used as the initial weight of the second BERT pre-trained model; and then the second data set and a part of the first data set are used for training again to obtain the final BERT pre-trained model.

[0022] The steps of inputting the long paragraph text statement into the multiple-choice reading comprehension model to obtain the multiple-choice reading comprehension answer are as follows:

[0023] The long text statement and the question are concatenated and unified to a fixed length to obtain a concatenated result of a fixed length;

[0024] Encode the fixed-length concatenated result using the ALBERT pre-trained model, and extract the encoded feature vector;

[0025] Input the encoded feature vector into the classifier of the reading comprehension model, and finally calculate the logical value of each position;

[0026] According to the logic value, the answer to the multiple-choice reading comprehension question is selected through a softmax function;

[0027] Set a threshold T, and determine the difference between the threshold T and the largest logical value and the second largest logical value;

[0028] If the difference between the largest logical value and the second largest logical value is not greater than the threshold T, the answer to the multiple-choice reading comprehension question is considered to be an UNKNOW answer;

[0029] If the difference between the largest logical value and the second largest logical value is greater than the threshold value T, the answer to the multiple-choice reading comprehension question is directly output.

[0030] The classifier of the extractive reading comprehension model is specifically a linear classifier with a dimension of 2×d, where d is the hidden layer state dimension;

[0031] The classifier of the reading comprehension model is set to two fully connected layers, the first fully connected layer is a linear layer with a dimension of d×d and a tanh activation function; the second fully connected layer is a linear layer with a dimension of d×1 and no activation function.

[0032] The steps of inputting the results of multiple rounds of dialogue into the intent recognition model to obtain the intent recognition results are as follows:

[0033] The results of the multiple rounds of dialogues are respectively input into a prototype network model based on a hybrid attention mechanism and a word vector similarity comparison model to obtain a first result and a second result respectively;

[0034] Performing a weighted summation of the first result and the second result to obtain a final result, wherein the first result, the second result and the final result all include categories and their corresponding probabilities;

[0035] The final results are sorted, and the sorted categories and their corresponding probabilities are returned in sequence.

[0036] The results of the multi-round dialogue are input into the prototype network model based on the hybrid attention mechanism to obtain the first result. The process steps are as follows:

[0037] Marking the results of the multiple rounds of conversations as text embeddings and position embeddings respectively;

[0038] Concatenating the text embedding with the position embedding;

[0039] The concatenated result is input into the CNN network and maximum pooling is performed, and the maximum pooling result output is the encoding information of the result of the multi-round dialogue;

[0040] Input the encoding information into the support set, extract features using an improved solution formula, and obtain N small class prototype vectors;

[0041] Calculate similarity between the encoded information of the results of the multiple rounds of dialogue and the N small class prototype vectors to obtain N similarity values;

[0042] The N similarity values ​​are converted into each category and the corresponding probability of each category.

[0043] The results of the multi-round dialogue are input into the word vector similarity comparison model to obtain the second result, and the process steps are as follows:

[0044] Divide the actual scene data in the scene library into S categories;

[0045] For each category, the category name is used as the keyword, corresponding to a keyword vector K;

[0046] At the same time, the sentences in each category are segmented and stop words are removed to obtain word vectors. The word vectors are added and averaged to obtain the flag vector V of each category.

[0047] Calculate similarity between the question vector Q and the keyword vector K and the flag vector V respectively to obtain a first similarity calculation result and a second similarity calculation result;

[0048] Performing a weighted summation of the first similarity calculation result and the second similarity calculation result, and finally obtaining the corresponding probability of each category;

[0049] Each category and the corresponding probability of each category are the second result.

[0050] In the second aspect, the present application proposes a legal document generation system based on a reading comprehension and intention recognition model, including: a data acquisition module, an element acquisition module, a scene acquisition module, a multi-round dialogue module, an intention recognition module, and a document generation module;

[0051] The data acquisition module, the element acquisition module, the scene acquisition module, the multi-round dialogue module, the intention recognition module, and the document generation module are connected in sequence, and the element acquisition module and the intention recognition module are respectively connected to the document generation module;

[0052] The data acquisition module is used to acquire long text statements;

[0053] The element acquisition module is used to input the long text statement into the long text reading comprehension model to obtain the initial round of legal elements;

[0054] The scenario acquisition module is used to compare the initial round of legal elements with the legal element library to obtain the legal scenario that the consultant wants to consult;

[0055] The multi-round dialogue module is used to obtain missing elements according to the legal scenario, initiate multi-round dialogues with the consultant according to the missing elements, and obtain the results of the multi-round dialogues;

[0056] The intention recognition module is used to input the results of multiple rounds of dialogue into the intention recognition model to obtain the intention recognition result;

[0057] The document generation module is used to make inference decisions based on the intention recognition results and the initial round of legal elements, and automatically generate a legal advisory opinion or contract document.

[0058] Beneficial technical effects:

[0059] The present application discloses a method and system for generating legal documents based on a reading comprehension and intent recognition model, which enables the parties to obtain accurate, professional and satisfactory answers to legal consultation. The solution based on the pre-trained model in the present application only needs to be iterated and fine-tuned to meet different NLP tasks, which greatly reduces the convergence time of the model and reduces the training cost. At the same time, a bidirectional self-attention mechanism is adopted, and the model effect of this mechanism is far superior to the previous RNN and LSTM model solutions. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 This is a flow chart of the method for generating legal documents according to an embodiment of the present application;

[0061] Figure 2 This is a flowchart of the method for generating legal documents according to an embodiment of the present application;

[0062] Figure 3 This is a schematic diagram of the process of obtaining the initial round of legal elements in an embodiment of the present application;

[0063] Figure 4 A schematic diagram of an extractive reading comprehension model according to an embodiment of the present application;

[0064] Figure 5A schematic diagram of a topic-based reading comprehension model according to an embodiment of the present application;

[0065] Figure 6 This is a schematic diagram of an intent recognition model according to an embodiment of the present application;

[0066] Figure 7 A schematic diagram of a prototype network model of an embodiment of the present application;

[0067] Figure 8 A schematic diagram of the process of extracting prototypes from each subclass in an embodiment of the present application;

[0068] Fig. 9 This is a schematic diagram of the process of constructing N subclass prototype vectors in an embodiment of the present application;

[0069] Fig.10 A schematic diagram of a similarity comparison model according to an embodiment of the present application;

[0070] Fig.11 This is a schematic diagram of the feature-level attention extractor process of an embodiment of the present application;

[0071] Fig.12 This is a flowchart of obtaining the initial round of legal elements in an embodiment of the present application;

[0072] Fig.13 This is a flow chart of the extractive reading comprehension model of an embodiment of the present application;

[0073] Fig.14 This is a flow chart of the topic-based reading comprehension model of an embodiment of the present application;

[0074] Fig.15 This is a flow chart of the intent recognition model of an embodiment of the present application;

[0075] Fig.16 A flowchart of obtaining a first result according to an embodiment of the present application;

[0076] Fig.17 This is a flow chart of the process of constructing N subclass prototype vectors according to an embodiment of the present application;

[0077] Fig.18 A flow chart of obtaining a second result according to an embodiment of the present application;

[0078] Fig.19 This is a functional block diagram of the legal document generation system of an embodiment of the present application. DETAILED DESCRIPTION

[0079] The present application is further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present application.

[0080] In the first aspect, the present application proposes a method for generating legal documents based on a reading comprehension and intention recognition model, such as Figure 1 , Figure 2 As shown, the following steps are included:

[0081] Step S1: Obtain a long text statement;

[0082] Step S2: inputting the long text statement into a long text reading comprehension model to obtain a preliminary round of legal elements;

[0083] Step S3: Compare the initial round of legal elements with the legal element database to obtain the legal scenario that the consultant wants to consult;

[0084] Step S4: Obtaining missing elements according to the legal scenario, initiating multiple rounds of dialogue with the consultant according to the missing elements, and obtaining results of the multiple rounds of dialogue;

[0085] The legal scenarios that the consultant wants to consult include civil consultation, divorce consultation, work injury consultation, etc., and may also contain some other event information. For example, the people involved, the place, the time, etc. After determining the legal scenario, based on the legal knowledge base and the knowledge graph, we can get what information is needed to issue an accurate and complete legal document, that is, the missing elements, and then we will have multiple rounds of dialogue with the consultant based on the missing information. The dialogue method is an active questioning method. The consultant only needs to use voice answers or manual selection to answer relevant questions. The questions are in the form of multiple-choice questions and fill-in-the-blank supplements to facilitate the consultant's understanding. When the core information required for the scenario is obtained through dialogue, the dialogue will end, and the system has fully understood the consultant's consulting intentions. The specific method of multi-round dialogue can use voice answers, or answer text or select questions to conduct multi-round dialogues, which are all prior art and will not be repeated in this application.

[0086] Step S5: input the results of multiple rounds of dialogue into the intention recognition model to obtain the intention recognition result;

[0087] Step S6: Make inference decisions based on the intention recognition results and the initial round of legal elements, and automatically generate a legal advisory opinion or contract document.

[0088] Reasoning decision is to perform reasoning query on key nodes of documents based on the legal elements extracted in the early stage and the key information obtained from the conversation combined with the legal knowledge graph. For example, the relevant local laws and regulations are obtained based on the geographical location information obtained in the early stage (the same legal issue in different regions may have different policies), and the relevant document data (such as the amount of compensation in traffic accidents, etc.) is calculated according to local laws and regulations. After completing the entire reasoning process, an accurate and standard legal document can be output. There are many existing algorithms for reasoning decision, which are not the innovation of this application and will not be repeated in this application.

[0089] The long paragraph statement is input into the long text reading comprehension model to obtain the initial round of legal elements, such as Figure 3 , Fig.12 As shown, the following steps are included:

[0090] Step S2.1: inputting the long text statement into the extractive reading comprehension model and the multiple-choice reading comprehension model respectively, and obtaining the extractive reading comprehension answer and the multiple-choice reading comprehension answer respectively;

[0091] Step S2.2: The extractive reading comprehension answers and the multiple-choice reading comprehension answers are used as the first round of legal elements.

[0092] The text length of multiple-choice reading comprehension and extractive reading comprehension is relatively long. The use of traditional machine learning methods and basic neural network models, such as RNN, LSTM, etc., no longer meets the requirements in terms of accuracy and computing time. Therefore, a pre-trained model is used to complete this application. The pre-trained model is used as the encoder of the model to obtain the representation of the data and train the model by fine-tuning on the pre-trained model. Specifically, a classifier is added after the pre-trained model and a loss function is designed. Finally, for practical applications, a strategy for the model to answer "UNKNOW" is formulated.

[0093] This application names the extractive model TTM-SQ (Transfer Trained Model for SpanQuestion), and uses a transfer learning strategy during training to improve the robustness of the model while accelerating the model training speed.

[0094] The long text statement is input into the extractive reading comprehension model to obtain the extractive reading comprehension answer, such as Figure 4 , Fig.13 As shown, the steps are as follows:

[0095] Step S2.1.11: concatenate the long text statement with the question and unify them to a fixed length to obtain a concatenated result of a fixed length;

[0096] Step S2.1.12: Encode the fixed-length concatenated result using the BERT pre-trained model, and extract the encoded feature vector;

[0097] Step S2.1.13: input the encoded feature vector into the classifier of the extractive reading comprehension model, and finally calculate the logical value of each position;

[0098] Step S2.1.14: According to the logical value, the extractive reading comprehension answer is selected through the softmax function.

[0099] The BERT pre-trained model adopts a transfer learning strategy to train the model, which includes two stages, using two data sets for training, and both data sets are annotated; in the first stage, only the first data set is used for training to obtain the first BERT pre-trained model and the first model weight; in the second stage, the first model weight is used as the initial weight of the second BERT pre-trained model; and then the second data set and a small part of the first data set are used for training again to obtain the final BERT pre-trained model.

[0100] In this embodiment, the model is trained using the strategy of transfer learning. Specifically, there are two stages of training, using two data sets for training. In the first stage, only the Fa Yan Bei data set is used for training. In the second stage, the weights of the pre-trained model trained in the first stage are used as the initial weights of the model. Then the labeled data set and a small part of the Fa Yan Bei data set are used for training again to obtain the final model. This training strategy has the following advantages:

[0101] (1) The first stage of training accelerates the convergence speed of the model when training on the target dataset;

[0102] (2) Improve model robustness;

[0103] (3) The second phase of the Fayan Cup has a smaller data set, which makes the model more suitable for the form of questions required by Party A and improves the accuracy of answering questions.

[0104] The classifier of the extractive reading comprehension model is specifically a linear classifier with a dimension of 2×d, where d is the hidden layer state dimension. r is the length of the input sequence. Features After passing through the classifier, a vector l is generated. s ∈R r and l s ∈R r , l s and l e The values ​​on each dimension represent the logit values ​​(i.e., logical values) of the position as the starting and ending positions, i.e., start_logit and end_logit. The model will judge the answer based on the logit values. When selecting the answer, invalid answers such as the starting position after the ending position will be removed, and softmax will be performed according to start_logit+end_logit. The softmax formula is as follows:

[0105]

[0106] Among them, z iis the output value of the i-th node, C is the value of N similarities,

[0107] The largest value is taken as the predicted answer. At the same time, if the softmax value of UNKNOW is the largest, the answer is judged to be UNKNOW. According to BERT's self-attention mechanism, the encoded [CLS] position can represent the semantics of the entire input sequence, and then judge the relationship between the two input sentences. Therefore, when calculating the logit value of UNKNOW, the model uses the start_logit and end_logit corresponding to the [CLS] position for summation.

[0108] Considering that the input text of multiple-choice reading comprehension is long, if BERT is used as the encoder, the time cost of training prediction is too high. ALBERT has advantages over BERT, such as small size and fast training speed. Therefore, after considering many aspects, the pre-trained model ALBERT is used as the encoder of the TTM-MC model.

[0109] In actual use, the original case text, question and each answer option are concatenated and sent to the ALBERT model for encoding. If the length is different from the set input sequence length, it is truncated or padded. The output of the last layer of the pre-trained model is taken as the input of the next step. For a number of (n) options for a question, n encodings are required.

[0110] The long paragraph text statement is input into the multiple-choice reading comprehension model to obtain the multiple-choice reading comprehension answer, such as Figure 5 , Fig.14 As shown, the steps are as follows:

[0111] Step S2.1.11.1: concatenate the long text statement with the question and unify them to a fixed length to obtain a concatenated result of a fixed length;

[0112] Step S2.1.11.2: Encode the fixed-length concatenated result using the ALBERT pre-trained model, and extract the encoded feature vector;

[0113] Step S2.1.11.3: Input the encoded feature vector into the classifier of the reading comprehension model, and finally calculate the logical value of each position;

[0114] Step S2.1.11.4: Selecting the answer to the multiple-choice reading comprehension question through a softmax function according to the logical value; z i is the output value of the i-th node, C is the value of N similarities, and the obtained value is the probability of each category.

[0115] Step S2.1.11.5: Set a threshold value T, and determine the difference between the threshold value T and the largest logic value and the second largest logic value;

[0116] Step S2.1.11.6: If the difference between the largest logical value and the second largest logical value is not greater than the threshold T, the answer to the multiple-choice reading comprehension question is considered to be an UNKNOW answer;

[0117] Step S2.1.11.7: If the difference between the largest logical value and the second largest logical value is greater than the threshold T, the answer to the multiple-choice reading comprehension question is directly output.

[0118] The method of using the ALBERT pre-trained model to encode the fixed-length concatenated result includes: encoding n times is required for n options of a question.

[0119] When selecting the structure of the classifier, the test results of models using the same encoder and different classifiers were compared experimentally, and finally A two-layer FCNN was selected. MAN is a classifier that can perform multi-step reasoning, which serves as a representative of complex structure classifiers. The classifier used in the model consists of a two-layer feedforward network. The experimental results show that the accuracy difference between the complex classifier (MAN) and the classifier used by the model is very small, and the performance of the two classifiers has its own characteristics according to different types of problems: the complex classifier performs better on problems that require more reasoning, but the effect decreases in simpler problems. Taking into account the slight differences in the distribution of problems and data sets in actual business use, and the impact of the classifier structure on model training time and prediction time and other factors. This application finally selected a simple classifier with a 2-layer feedforward network (A two-layer FCNN) as the final classifier of the TTM-MC model.

[0120] The classifier of the reading comprehension model is set to two fully connected layers, the first fully connected layer is a linear layer with a dimension of d×d and a tanh activation function; the second fully connected layer is a linear layer with a dimension of d×1 and no activation function;

[0121] After obtaining the features encoded by the encoder, the model uses m vectors V∈R of [CLS] bits (CLS is classification for downstream classification tasks) n×d Subsequent calculations are performed, where m represents the number of options, and m options are m vectors. After obtaining V, V is passed through a linear layer with a tanh activation function to obtain a vector V′∈R m×d , where the tanh activation function is:

[0122]

[0123] Among them, x h is the input of the activation function, and then V′ passes through a linear layer with a dimension of d×1 to obtain the logical values ​​of the m options.

[0124] The ALBERT pre-training model adopts a transfer learning strategy to train the model, which includes two stages. The first stage uses a natural language inference dataset for training, and the second stage uses a RACE dataset for training.

[0125] Similar to the extractive model, the transfer learning strategy is also used in the training process of the TTM-MC model. Since the target dataset RACE contains a large number of reasoning questions, the natural language inference dataset is used for training in the first stage, and the RACE dataset is used for training in the second stage. This training strategy has the following advantages:

[0126] (1) The first stage of training accelerates the convergence speed of the model when training on the target dataset;

[0127] (2) Through training on natural language inference tasks, the accuracy of the model in answering questions is improved.

[0128] The loss function is used to evaluate the degree of inconsistency between the model's predicted value and the true value. The smaller the loss function, the better the robustness of the model. The loss function is placed at the end of the forward propagation of the neural network training. The result obtained after passing through the multi-layer network will calculate the loss with the true value, and then the parameters in the network will be updated through back propagation. If the loss function value cannot be reduced after long-term training, the learning rate, activation function or loss function will generally be adjusted, but for the pre-trained model based on this application, it will generally not appear because this application is only doing fine-tuning.

[0129] The cross entropy function of the starting and ending positions and the answer is selected as the loss function of the model.

[0130]

[0131] Where y represents each position in the sentence and the value is the length of the sentence. p(y) represents the position of the answer and the value is 1 if the x position is the correct start / end position, otherwise it is 0. q(y) is the logit value (i.e. logical value) of the x position predicted by the model.

[0132] The final loss function is:

[0133] H=(H s +H e ) / 2

[0134] H s and H eRespectively represent the cross entropy calculated from the starting position and the ending position.

[0135] Example of the experimental process of the embodiment of this application:

[0136] Hardware environment: Linux server, configured as GPU: GTX1080Ti, CPU: Intel(R)Xeon(R) CPU E5-2678.

[0137] Software environment: The server system is Ubuntu 16.04.5, the Python version is 3.7, and the CUDA version is 10.2.

[0138] Extractive Reading Comprehension Dataset

[0139] (1) Law Research Cup Dataset: This dataset is from the legal documents published by the China Judgment Documents Network, mainly involving first-instance judgments of civil and criminal cases. There are about 10,000 data in total, which are divided into training, development and testing in proportion. Each data includes several questions. For the training set, each question contains only one standard answer. For the development and test sets, each question contains 3 standard answers. The answer content can be a fragment of the case, which can be YES or NO, or it can be a refusal to answer, that is, the answer content is empty.

[0140] (2) Party A's labeled dataset (target dataset): The dataset labeled by Party A totaled 13,603 items, including 3,262 items with answers of YES or NO and 10,341 items with answers that can be found in the text. Most of them are question-answer pairs about time, money, etc. Therefore, the model can answer questions about these very accurately during training. However, this dataset lacks question-answer pairs that do not have answers in the text. The data format is shown in Table 1, and the data format diagram is shown in Table 1.

[0141] Table 1 Extractive reading comprehension data format description table

[0142]

[0143]

[0144] Extractive reading comprehension parameter selection Extractive reading comprehension parameter selection is shown in Table 2.

[0145] Table 2 Parameters of the extractive reading comprehension training model

[0146]

[0147]

[0148] Multiple-choice reading comprehension dataset

[0149] NLI dataset: Natural language inference is the task of judging the semantic relationship between sentence pairs. It includes judging the relationship between sentence pairs: neutral, implicated, contradictory; whether sentence pairs are similar, etc. Commonly used datasets include SNLI, Multi-NLI and other datasets. Stanford Natural Language Inference (SNLI) is the most commonly used version of natural language inference. It contains 550,152 training samples, 1,000 verification samples, and 10,000 test samples. Each sample is a sentence pair, and each sentence pair is labeled with one of these three labels: neutral, implicated, and contradictory. Multi-Genre Natural Language Inference (MNLI) collects 433,000 sentence pairs. This corpus is an extension of SNLI, covering a wide range, including spoken and written language, and supports unique cross-genre generalization evaluation.

[0150] RACE dataset: The RACE dataset collects English exams for junior and senior high school students aged 12-18, including nearly 28,000 paragraphs and 100,000 questions asked by human experts. It is divided into middle and high, where middle refers to the English for the junior high school entrance exam and high refers to the English for the college entrance exam. The statistics of the number of articles and questions for train, dev, and high of RACE-M, RACE-H, and RACE are shown in Table 3. The statistics of the length of articles, questions, and options of RACE are shown in Table 4. The statistics of the proportion of various reasoning types in RACE are shown in Table 5. The format description of the dataset is shown in Table 6.

[0151] Table 3 Statistics of the number of RACE articles and questions

[0152]

[0153] Table 4 Statistics of the length of RACE articles, questions and options

[0154]

[0155]

[0156] Table 5 Statistics of the proportion of various reasoning types in RACE

[0157]

[0158] Table 6 Multiple-choice question reading comprehension data format description table

[0159]

[0160]

[0161] Multiple choice reading comprehension parameter selection

[0162] The model of multiple-choice reading comprehension uses a large number of hyperparameters, such as the maximum length of encoding, the number of training times, the storage path of data, the storage path of the model, etc. Therefore, users need to have a certain degree of understanding of the model's parameter table so that the model can play a greater role on the target data set. The parameter selection of multiple-choice reading comprehension is shown in Table 7, which introduces the hyperparameter type, function and corresponding selected value of the model parameters.

[0163] Table 7 Parameters of the multiple choice reading comprehension training model

[0164]

[0165]

[0166] modeling_albert.py is the structure file of the TTM-MC model, which defines the model structure, forward propagation method, loss function and other parts.

[0167] This application adopts a prototype network based on a hybrid attention mechanism (Hybrid Attention-Based Prototypical Networks), which can better solve the impact of sample noise on experimental results compared to ordinary prototype networks (Prototypical Networks).

[0168] Firstly, a prototype network based on a hybrid attention mechanism is used to construct the logo vector of each category of each subdivided law. Then, the semantic distance between the feature vector of the user's sentence and the logo vector is calculated to achieve small sample classification.

[0169] The original prototype network calculates the prototype by taking the average of the instance sentences in suppprtset as the prototype of each relationship. The idea of ​​any prototype network to solve the prototype, but the direct averaging method defaults to the same value for the weight of each input sample, which will significantly affect the prototype solution when there are few input samples and the samples are noisy.

[0170] Sample instance-level attention mechanism: In few-shot learning, if the training process samples contain noise, it will significantly affect the solution of the prototype. This application proposes a sample instance-level attention module, which focuses more attention on the instances related to the query and reduces the impact of noise. This application modifies the formula for solving the prototype.

[0171] Feature-level attention mechanism: The original prototype network uses a simple Euclidean distance as the distance function. Since there are fewer instances in the support set in few-shot learning, the features extracted from the support set have the problem of data sparsity. Therefore, when classifying special relationships in the feature space, some dimensions have stronger distinguishing power. This application adopts a feature-level attention method to alleviate the problem of feature sparsity and measure spatial distance in a more appropriate way. This application replaces the formula d(s1-s2)=(s1-s2) 2 Modified to d(s1-s2) = z1(s1-s2) 2 , where z1 is calculated by the feature-level attention extractor in the figure below.

[0172] The results of multiple rounds of dialogue are input into the intention recognition model to obtain the intention recognition result, such as Figure 6 , Fig.15 The steps shown are as follows:

[0173] Step S2.1.21: inputting the results of the multiple rounds of dialogue into the prototype network model based on the hybrid attention mechanism and the word vector similarity comparison model respectively, to obtain the first result and the second result respectively;

[0174] Step S2.1.22: performing weighted summation of the first result and the second result to obtain a final result, wherein the first result, the second result and the final result all contain categories and their corresponding probabilities;

[0175] Step S2.1.23: Sort the final results, and return the sorted categories and their corresponding probabilities in turn.

[0176] The results of the multi-round dialogue are input into the prototype network model based on the hybrid attention mechanism to obtain the first result, such as Figure 7 , Figure 8 , Fig.16 As shown, the process steps are as follows:

[0177] Step S2.1.22.11: Marking the results of the multiple rounds of dialogues as text embeddings and position embeddings respectively;

[0178] Step S2.1.22.12: concatenating the text embedding and the position embedding;

[0179] Step S2.1.22.13 inputs the concatenated result into the CNN network and performs maximum pooling, and the output maximum pooling result is the encoding information of the result of the multi-round dialogue;

[0180] The encoding process is the process of vectorizing user input.

[0181] Step S2.1.22.14: Input the encoded information into the support set, use the improved solution formula to extract features, and obtain N small class prototype vectors;

[0182] Figure 8 The prototype extraction process for each subcategory is performed using the improved solution formula.

[0183] Step S2.1.22.15: Calculate the similarity between the encoded information of the result of the multi-round dialogue and the N small class prototype vectors to obtain N similarity values; generally, the cosine similarity function is used for calculation;

[0184] Step S2.1.22.16: Convert the N similarity values ​​into each category and the corresponding probability of each category.

[0185] The conversion method is completed through a Softmax function, and the conversion function is z i is the output value of the i-th node, C is the value of N similarities, and the obtained value is the probability of each category.

[0186] Other mainstream classification models usually output a vector of fixed dimension in the last layer. For example, an article may have three categories, such as finance, sports, and entertainment. If an article is input into the model, the output result is a vector like [0.5, 0.3, 0.2]. The three numbers in the vector indicate the probability of the article belonging to finance, sports, and entertainment respectively.

[0187] However, the prototype network classification model of this application is different. Its essence is to calculate the distance between the user's input and the "prototype". The model will extract a "prototype" vector for each subclass based on the text in the provided different subclass txt files, and after adding a new subclass txt file, the model will automatically extract the prototype of the new subclass and add it to the original "prototype" set when the classification function is executed for the first time.

[0188] Figure 8 As shown in the figure, it is easy to know that if you add a new category, then add a txt file, let the model automatically use the text in the txt file to extract a "prototype", and then do one more similarity calculation during the classification process. This is the principle of "adding new categories without retraining".

[0189] The N small class prototype vectors are as follows: Fig. 9 , Fig.17 As shown, the construction process is as follows:

[0190] Step S100: Read N sub-categories of txt files;

[0191] Step S101: converting the txt files of the N subcategories into json files of the N subcategories;

[0192] Step S102: Use the json files of the N subclasses to perform prototype extraction to obtain N subclass prototype vectors.

[0193] Only the required small category of txt files are processed into json files (this process is very fast, just a word segmentation process, which takes about one millisecond), and the processed json files are placed in a specific folder, and then the model is allowed to read the json files from the specific folder. (As for the function of processing into json, this algorithm has been encapsulated, and a simple modification can realize "reading a specific txt" and "storing the formed json in a specific location")

[0194] The prototype vector is extracted by the improved solution formula. The distance formula is d(s1-s2) = (s1-s2) 2 Modified to d(s1-s2)=z1(s1-s2) 2 Therefore, when calculating the distance, we need to first calculate a vector Z through convolution based on the prototype vector. i (like Fig.11 ), and then calculate the distance.

[0195] The formula for solving the prototype is modified from:

[0196]

[0197] Modified to

[0198]

[0199] For relation i, the number of samples is n i , whose prototype eigenvector is c i , j represents the jth sample in the i-th relationship (1≤j≤n i ), α j represents the weight of the jth sample, Represents the feature vector obtained after encoding the jth sample in the i-th relationship.

[0200] Among them, α j Defined as

[0201]

[0202]

[0203] Among them, a j From the Softmax function, we get (e jas the corresponding parameter); x is the feature vector of the sample, g(·) is a linear layer, which is the product of the elements, σ(·) is an activation function, and this application chooses tanh as σ(·), and sumf·g is the sum of all elements of the vector.

[0204] Formula d(s1-s2)=(s1-s2) 2 Modified to d(s1-s2)=z1(s1-s2) 2 , d represents the distance function between two samples, and s represents the feature vector of the sample.

[0205] like Fig.11 The feature-level attention extractor process is as follows: the input is the feature vector of K samples, and then passes through a three-layer convolutional network. The first layer is a 32-channel convolutional layer, the second layer is a 64-channel convolutional layer, and the third layer is a channel convolutional layer to ensure that the result is an independent vector. The activation function in the middle uses the commonly used Relu function. After such a simple three-layer network, an attention vector Z based on sparse features can be obtained. i .

[0206] The effect of the word vector similarity comparison model is far worse than that of the prototype network classification model, but adding this model can form a complementary situation with the prototype network classification model. Sometimes the prototype network model fails to make a good prediction, so the word vector comparison model can be used to supplement the prediction results of the prototype network model. However, since this model does not have a good prediction effect most of the time, this application only gives it a small weight.

[0207] The results of the multi-round dialogue are input into the word vector similarity comparison model to obtain the second result, as shown in Figure 10. Fig.18 As shown, the process steps are as follows:

[0208] Step S2.1.22.21: Divide the actual scene data in the scene library into S categories; the scene library stores a collection of all scenes.

[0209] Step S2.1.22.22: For each category, the category name is used as the keyword, and a keyword vector K is corresponding to each category;

[0210] Step S2.1.22.23: Segment the sentences in each category and remove stop words at the same time to obtain word vectors, add the word vectors and take the average to obtain the logo vector V of each category; the word vectors here are vector conversions completed by the word2vec word vector model.

[0211] Step S2.1.22.24: Calculate the similarity between the question vector Q and the keyword vector K and the logo vector V respectively to obtain the first similarity calculation result and the second similarity calculation result; the question is also vectorized using the word2vec word vector model.

[0212] Step S2.1.22.25: performing weighted summation on the first similarity calculation result and the second similarity calculation result, and finally obtaining the corresponding probability of each category;

[0213] Step S2.1.22.26: Each category and the corresponding probability of each category are the second result.

[0214] When the data in the scene library is updated, the word vector similarity comparison model w2vModel is updated in time. When the data in the scene library is updated frequently, the prototype network model is retrained in time to improve the accuracy of the prototype network model.

[0215] The similarity calculation formula is as follows: Among them, A and B represent two eigenvectors. Fig.10 As shown, the weighted sum is finally performed to obtain the probability p1′ of the category. In this way, each category has a corresponding probability p2′, …, pN′ as the output of w2vModel.

[0216] Experimental examples of the present application:

[0217] Hardware environment: Linux server, configured as GPU: GTX1080Ti*4, CPU: Intel(R)Xeon(R) CPUE5-2678.

[0218] Software environment: The server system is Ubuntu 16.04.5, the Python version is 3.6, and the CUDA version is 10.2. The required environment dependencies are shown in Table 8.

[0219] Table 8 Environment dependency table

[0220]

[0221] The dataset is divided into training datasets for prototype network training and word vector training, scenario datasets for application in actual scenarios, and test datasets for model effect testing.

[0222] The legal consultation dataset contains 46 categories. The file name is the category name, in txt format, and the number of samples in each category is no less than 100. All of them are user consultation sentences with a strong colloquial style. The specific format is shown in Table 9:

[0223] Table 9 Training dataset format

[0224]

[0225] There are different scenario datasets in different environments. The following provides the dataset format of a certain environment. Initially, there are 25 categories in total, and the number of samples in each category is not less than 20. In actual application, the categories of the scenarios can be added or deleted, and the merging of scenario contents is supported. For example, "work injury compensation" and "work injury identification" are merged into the "work injury" category. At the same time, there is a "chat" dataset that can be clearly distinguished from specific legal scenarios. The specific format is as follows:

[0226] Table 10 Scene dataset format

[0227]

[0228] The test dataset is used to test the model effect. The specific format is similar to the scene dataset, and the categories in the test dataset must correspond to the categories in the scene dataset, but there is no limit on the number of samples.

[0229] On the second aspect, this application proposes a legal document generation system based on reading comprehension and intention recognition model, such as Fig.19 As shown, it includes: data acquisition module, element acquisition module, scene acquisition module, multi-round dialogue module, intention recognition module, and document generation module;

[0230] The data acquisition module, the element acquisition module, the scene acquisition module, the multi-round dialogue module, the intention recognition module, and the document generation module are connected in sequence, and the element acquisition module and the intention recognition module are respectively connected to the document generation module;

[0231] The data acquisition module is used to acquire long text statements;

[0232] The element acquisition module is used to input the long text statement into the long text reading comprehension model to obtain the initial round of legal elements;

[0233] The scenario acquisition module is used to compare the initial round of legal elements with the legal element library to obtain the legal scenario that the consultant wants to consult;

[0234] The multi-round dialogue module is used to obtain missing elements according to the legal scenario, initiate multi-round dialogues with the consultant according to the missing elements, and obtain the results of the multi-round dialogues;

[0235] The intention recognition module is used to input the results of multiple rounds of dialogue into the intention recognition model to obtain the intention recognition result;

[0236] The document generation module is used to make inference decisions based on the intention recognition results and the initial round of legal elements, and automatically generate a legal advisory opinion or contract document.

[0237] The applicant of the present invention has made a detailed explanation and description of the implementation examples of the present invention in conjunction with the drawings in the specification. However, those skilled in the art should understand that the above implementation examples are only preferred implementation schemes of the present invention, and the detailed description is only to help readers better understand the spirit of the present invention, but not to limit the scope of protection of the present invention. On the contrary, any improvements or modifications based on the inventive spirit of the present invention should fall within the scope of protection of the present invention.

Claims

1. A method for generating legal documents based on reading comprehension and intention recognition model, characterized in that: The steps include: Get long text statements; Inputting the long text statement into a long text reading comprehension model to obtain a preliminary round of legal elements; Compare the initial round of legal elements with the legal element database to obtain the legal scenario that the consultant wants to consult; Obtain missing elements according to the legal scenario, initiate multiple rounds of dialogue with the consultant based on the missing elements, and obtain the results of the multiple rounds of dialogue; Input the results of multiple rounds of conversations into the intent recognition model to obtain the intent recognition results. The steps are as follows: The results of the multiple rounds of dialogues are respectively input into a prototype network model based on a hybrid attention mechanism and a word vector similarity comparison model to obtain a first result and a second result respectively; Performing a weighted summation of the first result and the second result to obtain a final result, wherein the first result, the second result and the final result all include categories and their corresponding probabilities; Sorting the final results, and returning the sorted categories and their corresponding probabilities in turn; Based on the intention recognition results and the initial round of legal elements, reasoning and decision-making are carried out to automatically generate legal advisory opinions or contract documents.

2. The method for generating legal documents based on a reading comprehension and intention recognition model as claimed in claim 1, characterized in that: The step of inputting the long text statement into the long text reading comprehension model to obtain the initial round of legal elements includes the following steps: Inputting the long text statement into the extractive reading comprehension model and the multiple-choice reading comprehension model respectively, and obtaining the extractive reading comprehension answer and the multiple-choice reading comprehension answer respectively; The answers to the extractive reading comprehension questions and the answers to the multiple-choice reading comprehension questions are used as the initial round of legal elements.

3. The method for generating legal documents based on reading comprehension and intention recognition model as claimed in claim 2, characterized in that: The steps of inputting the long text statement into the extractive reading comprehension model to obtain the extractive reading comprehension answer are as follows: The long text statement and the question are concatenated and unified to a fixed length to obtain a concatenated result of a fixed length; Encode the fixed-length concatenated result using the BERT pre-trained model, and extract the encoded feature vector; Inputting the encoded feature vector into a classifier of an extractive reading comprehension model, and finally calculating the logical value of each position; According to the logical value, the extractive reading comprehension answer is selected through the softmax function.

4. The method for generating legal documents based on a reading comprehension and intention recognition model as claimed in claim 3, characterized in that: The BERT pre-training model adopts the transfer learning strategy to train the model, which includes two stages and uses two datasets for training, and both datasets are annotated; In the first stage, only the first data set is used for training to obtain the first BERT pre-trained model and the first model weights; In the second stage, the first model weights are used as the initial weights of the second BERT pre-trained model; The second data set and a part of the first data set are used for training again to obtain the final BERT pre-training model.

5. The method for generating legal documents based on reading comprehension and intention recognition model as claimed in claim 2, characterized in that: The steps of inputting the long paragraph text statement into the multiple-choice reading comprehension model to obtain the multiple-choice reading comprehension answer are as follows: The long text statement and the question are concatenated and unified to a fixed length to obtain a concatenated result of a fixed length; Encode the fixed-length concatenated result using the BERT pre-trained model, and extract the encoded feature vector; Input the encoded feature vector into the classifier of the reading comprehension model, and finally calculate the logical value of each position; According to the logic value, the answer to the multiple-choice reading comprehension question is selected through a softmax function; Set a threshold T, and determine the difference between the threshold T and the largest logical value and the second largest logical value; If the difference between the largest logical value and the second largest logical value is not greater than the threshold T, the answer to the multiple-choice reading comprehension question is considered to be an UNKNOW answer; If the difference between the largest logical value and the second largest logical value is greater than the threshold value T, the answer to the multiple-choice reading comprehension question is directly output.

6. The method for generating legal documents based on reading comprehension and intention recognition model as claimed in claim 3 or 5, characterized in that: The classifier of the extractive reading comprehension model is specifically a classifier with a dimension of A linear classifier, where is the hidden state dimension; the classifier of the reading comprehension model is set to two fully connected layers, the first fully connected layer has a dimension of With The second fully connected layer is a linear layer with a dimension of A linear layer without an activation function.

7. The method for generating legal documents based on reading comprehension and intention recognition model as claimed in claim 1, characterized in that: The results of multiple rounds of dialogue are input into the prototype network model based on the hybrid attention mechanism to obtain the first result. The process steps are as follows: Marking the results of the multiple rounds of conversations as text embeddings and position embeddings respectively; Concatenating the text embedding with the position embedding; The concatenated result is input into the CNN network and maximum pooling is performed, and the maximum pooling result output is the encoding information of the result of the multi-round dialogue; The encoding information is input into the support set, and the features are extracted using an improved solution formula, which specifically includes: the formula for solving the prototype is: , modified to: , is the feature vector of the prototype; For relationship The number of samples; j represents the jth sample in the i-th relationship, ; represents the weight of the jth sample in the i-th relationship, Represents the feature vector obtained after encoding the jth sample in the i-th relationship; Get N small class prototype vectors; Calculate similarity between the encoded information of the results of the multiple rounds of dialogue and the N small class prototype vectors to obtain N similarity values; The N similarity values ​​are converted into each category and the corresponding probability of each category.

8. The method for generating legal documents based on reading comprehension and intention recognition model as claimed in claim 1, characterized in that: The results of multiple rounds of dialogue are input into the word vector similarity comparison model to obtain the second result. The process steps are as follows: Divide the actual scene data in the scene library into S categories; For each category, the category name is used as the keyword, corresponding to a keyword vector K; At the same time, the sentences in each category are segmented and stop words are removed to obtain word vectors. The word vectors are added and averaged to obtain the flag vector V of each category. Calculate similarity between the question vector Q and the keyword vector K and the flag vector V respectively to obtain a first similarity calculation result and a second similarity calculation result; Performing a weighted summation of the first similarity calculation result and the second similarity calculation result, and finally obtaining the corresponding probability of each category; Each category and the corresponding probability of each category are the second result.

9. A legal document generation system based on reading comprehension and intention recognition model, characterized in that: include: Data acquisition module, factor acquisition module, scene acquisition module, multi-round dialogue module, intention recognition module, document generation module; The data acquisition module, the element acquisition module, the scene acquisition module, the multi-round dialogue module, the intention recognition module, and the document generation module are connected in sequence, and the element acquisition module and the intention recognition module are respectively connected to the document generation module; The data acquisition module is used to acquire long text statements; The element acquisition module is used to input the long text statement into the long text reading comprehension model to obtain the initial round of legal elements; The scenario acquisition module is used to compare the initial round of legal elements with the legal element library to obtain the legal scenario that the consultant wants to consult; The multi-round dialogue module is used to obtain missing elements according to the legal scenario, initiate multi-round dialogues with the consultant according to the missing elements, and obtain the results of the multi-round dialogues; The intention recognition module is used to input the results of multiple rounds of dialogue into the intention recognition model to obtain the intention recognition result. The steps are as follows: The results of the multiple rounds of dialogues are respectively input into a prototype network model based on a hybrid attention mechanism and a word vector similarity comparison model to obtain a first result and a second result respectively; Performing a weighted summation of the first result and the second result to obtain a final result, wherein the first result, the second result and the final result all include categories and their corresponding probabilities; Sorting the final results, and returning the sorted categories and their corresponding probabilities in turn; The document generation module is used to make inference decisions based on the intention recognition results and the initial round of legal elements, and automatically generate a legal advisory opinion or contract document.

Citation Information

Patent Citations

  • Medical dialogue system intention recognition and classification method based on deep learning

    CN110110059A

  • Man-machine interaction method and device based on semantic net and intention recognition and medium

    CN112069298A

  • IPTV terminal legal consultation method and system based on intelligent interaction

    CN113438516A