Deep Event Extraction Method in the Judicial Field Integrating Multi-Task and Multi-Label Learning

By applying a method of fusion of multi-task and multi-label learning in the judicial field, using BERT model and multi-task technology to achieve trigger word extraction, event classification and factor extraction, solving the efficiency and accuracy of in-depth event extraction in the judicial field, and achieving high-performance event extraction effect.

CN114580428BActive Publication Date: 2025-06-20NO 15 INST OF CHINA ELECTRONICS TECH GRP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210078832.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-24
Publication Date
2025-06-20
Estimated Expiration
2042-01-24

AI Technical Summary

Technical Problem

It is difficult for the prior art to achieve in-depth event extraction in the judicial field, especially in the process of identifying event types, trigger words and event elements in text data.

Method used

Using a method of fusion of multi-task and multi-task, trigger word extraction and event classification are implemented based on BERT pre-trained model and multi-task, and event feature extraction is realized through multi-task classification. This method optimizes data in the judicial field, uses task relevance to improve learning performance, and reduces interference between factors through the exclusive network layer.

Benefits of technology

It improves the accuracy and efficiency of event extraction in the judicial field, reduces the need for manual feature pattern construction, saves manpower, and improves the generalization ability and performance of the model through data augmentation and task joint learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114580428B_ABST
    Figure CN114580428B_ABST
Patent Text Reader

Abstract

The present invention discloses a deep event extraction method in the judicial field that integrates multi-task and multi-label learning. It can implement trigger word extraction and event classification based on the BERT pre-trained model and multi-task, and realize event extraction in the judicial field for event element extraction through multi-label classification on enhanced data. Currently, aiming at the characteristics of judicial field texts, an event extraction model based on the pre-trained model BERT is proposed. The BERT is optimized on domain data through the masked LM method to learn feature representations more suitable for domain knowledge; the trigger word extraction and event classification tasks are combined, and the two tasks are unified into a single loss function in the form of multi-task, using the correlation between tasks to promote the improvement of learning performance; the start and end annotations of event elements are used for learning and prediction, and for multiple event elements, corresponding network layers are designed for extraction respectively to reduce the mutual interference between different elements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of event extraction, and particularly to a deep event extraction method in the judicial field that integrates multi-task and multi-label learning. Background Art

[0002] Event extraction is a classic information extraction (IE) task in the field of natural language processing (NLP). It requires us to identify the important elements of events related to our goals from semi-structured or even unstructured data, either manually or automatically. There are five relatively important concepts in the event extraction task: event mention, event type, trigger, argument, and the role of the argument. An event mention refers to a phrase or sentence that describes event information. An event type refers to the type of event, such as "theft event". A trigger is a word that indicates the occurrence of an event, usually a verb. An argument refers to important information such as time, place, and person used to describe an event. The role of the argument is the function of the argument in the process of the event. From the perspective of the text information of event extraction, it can be divided into sentence-based event extraction and document-based event extraction. From the perspective of the event extraction model, event extraction can adopt a pipeline structure model or a joint model. From the perspective of the event extraction goal, it includes event extraction based on a specific schema and event extraction in the open domain. In the past decade, thanks to the rapid improvement of the computing power of graphics processing units (GPUs), deep learning has achieved good results in many fields compared with machine learning, such as automatic translation, image recognition, natural language processing, etc. Currently, there are a large number of application practices in the military, financial, biological and other fields, greatly improving the accuracy on the basis of traditional machine learning methods.

[0003] Currently, the main methods in the field of event extraction are divided into three categories:

[0004] The first category is the event extraction method based on pattern matching. The pattern matching method is to identify and extract events under the guidance of some patterns. Patterns are mainly used to indicate the context constraint environment that constitutes the target information, and centrally reflect the integration of domain knowledge and language knowledge. During extraction, the event text is input, and the information that meets the pattern constraint conditions is found through various pattern matching algorithms (such as regular expressions) as the output.

[0005] The second category is the machine learning-based event extraction method. By manually extracting relevant features, it uses machine learning methods based on pipeline or joint model to identify events, transforming the identification of event categories and event elements into classification problems. Among them, the pipeline-based method transforms the event extraction task into a multi-stage classification problem, sequentially executing multiple classifiers; the joint model-based method jointly learns trigger word recognition and element extraction, making full use of the correlation between event trigger words and elements, effectively improving the performance of the model.

[0006] The third category is the deep learning-based event extraction method. Through word embedding tools such as word2vec and n-gram models, it obtains the word embedding information corresponding to the text. Using the word vector embedding information, it learns the semantic information of the text through a bidirectional long short-term memory network (Bi-LSTM), obtains the feature representation based on the comprehensive context content, and then adds constraint conditions through a conditional random field to obtain the final event extraction result.

[0007] In the pattern matching-based event extraction method, since the patterns are mainly established by manual methods, this method is time-consuming and laborious, and also requires users to have high professional domain skills. The pattern matching-based method can achieve relatively good results in a specific domain, but the portability of the system is poor. When transplanting from one domain to another, the patterns need to be rebuilt. And the construction of patterns is time-consuming and laborious, requiring the guidance of domain experts.

[0008] Although the machine learning-based method does not depend on the content and format of the corpus, it requires a large-scale labeled corpus. Otherwise, there will be a relatively serious data sparsity problem. However, the current corpus scale is difficult to meet the application requirements, and manually labeling the corpus is time-consuming and laborious. Commonly used machine learning methods mainly include the hidden Markov model and the conditional random field. The hidden Markov model is suitable for relatively small data sets. If the data set is relatively complex, its simple feature functions cannot cover the features of the complex data set. The conditional random field has one less assumption than the hidden Markov model. The reduced constraints enable the conditional random field to utilize more features than the hidden Markov model, such as the context information of the observation sequence and the features of the elements of the observation sequence itself. The conditional random field model can utilize the context information, so it is more suitable for Chinese part-of-speech tagging. Its performance depends on feature selection, and the quality of feature selection directly determines the performance of the model.

[0009] Deep learning-based methods, such as RNN, LSTM, and BILSTM models, are very powerful in sequence modeling and can capture long-term context information. In addition, they have the ability of neural networks to fit non-linearities, which are beyond the reach of conditional random fields. For each time step, the output layer is affected by the hidden layer (containing context information) and the input layer (the current input), but the output layers at different time steps are independent of each other. For the current time step, we hope to find an output with the highest probability, but the outputs at other time steps have no influence on the current output. If there is a strong dependence between them (for example, an adjective is usually followed by a noun, there are certain constraints), LSTM cannot model these constraints, and the performance of the LSTM model will be limited.

[0010] Currently, there is no solution for deep event extraction for text data in the judicial field. Summary of the Invention

[0011] In view of this, the present invention provides a deep event extraction method for the judicial field that integrates multi-task and multi-label learning, which can realize trigger word extraction and event classification based on the BERT pre-trained model and multi-task, and realize event element extraction in the judicial field event extraction through multi-label classification on the enhanced data.

[0012] To achieve the above object, the technical solution of the present invention includes the following steps:

[0013] Step 1: Take the data in the judicial field for manual annotation, and the annotated labels include event types and event elements to obtain a judicial field dataset.

[0014] Step 2: Use the Chinese pre-trained language model BERT on the judicial field dataset, and adopt the Masked LM language learning model to optimize the network, learn the network parameters suitable for the judicial field knowledge, so as to obtain the judicial field BERT model, and use the output of the judicial field BERT model as the semantic information of the text.

[0015] Step 3: Construct a multi-task network. The multi-task network uses the semantic information of the text extracted by the judicial field BERT model as the input. The multi-task network defines a loss function jointly defined by three tasks: predicting the start position of the trigger word, predicting the end position of the trigger word, and predicting the event type for optimization. The output of the multi-task network includes the predicted event type, the predicted start position of the trigger word, and the predicted end position of the trigger word.

[0016] Step 4: Determine event elements according to the event type and construct an event element extraction model. The event element extraction model takes the text semantic information extracted by the BERT model in the judicial field as input, learns exclusive network parameters for each event element, and at the last layer of the network corresponding to each event element, predicts whether each token belongs to the start position or the end position of the current event element for each tokenized word.

[0017] Further, use the judicial field dataset to fine-tune the network of the Chinese pre-trained language model BERT on the judicial field dataset by using the Masked LM language learning model. Specifically:

[0018] On the manually annotated judicial field dataset, use Masked LM to fine-tune the parameters of the BERT model. During training, adopt the following strategy: randomly select 15% of the words in the sentence for masking. Among the words selected for masking, 80% are really replaced with [Mask], 10% are not replaced, and the remaining 10% are replaced with a random word.

[0019] Further, Step 2 is specifically as follows:

[0020] The set of judicial field events E = {E1, …, E N}, where E1 to E N are the 1st to the Nth judicial field events; the set of text information corresponding to the judicial field events is S = {S1, …, S N}, where S1 to S N are the text information corresponding to the 1st to the Nth judicial field events respectively; the maximum value of the epoch in the BERT model is Epoches, the number of batches in each epoch is batch_per_epoch; the BERT base model is Bert_base_chinese, and the maximum length of each sentence is max_len;

[0021] For all epochs in the BERT model, execute the following training process to obtain the fine-tuned BERT model parameters:

[0022] For each batch in the epoch, execute S1 to S4:

[0023] S1 Pad the input sentence with zeros or truncate it to a length of max_len to obtain the index I1 of the tokenized sentence;

[0024] S2 Randomly select 15% of the words in the sentence for masking. Among the words selected for masking, 80% are really replaced with [Mask], 10% are not replaced, and the remaining 10% are replaced with a random word;

[0025] The sentence after obtaining the Mask is input into the BERT base model Bert_base_chinese to obtain feature vectors, and then followed by θ0 to predict the index I2 of the word segmentation corresponding to each position of the sentence;

[0026] S4 uses the Adam optimizer to minimize the difference between I1 and I2, which is defined as the first loss function L(θ, θ0); when the first loss function on the validation set no longer decreases within a certain number of epochs, an early stopping strategy is adopted;

[0027] Furthermore, the first loss function L(θ, θ0) is defined as follows:

[0028]

[0029] where θ is the parameter of the Encoder part in the BERT model, the input passes through θ to obtain feature vectors, θ0 is the parameter following θ in the Masked LM task, |V| is the size of the dictionary composed of the masked words; m i represents the masked word; p(m = m i |θ, θ0) represents the probability that the predicted word m is the masked word m given the learned parameters θ and θ0; i ;

[0030] In the above training process, in the first two epochs of the BERT model, θ is fixed, and θ0 is adjusted with a learning rate of lr = 5e -4 In the subsequent epochs, θ and θ0 are adjusted simultaneously with a learning rate of lr = 1e -5 until the stopping condition is reached.

[0031] Furthermore, step 3 is specifically: after tokenizing the judicial data text, the position embedding, segment embedding, and word embedding of each token are obtained, and the three embeddings are input into the fine-tuned BERT model in the judicial field to obtain the feature vector of each token, which is the semantic information of the text;

[0032] The position embedding is the position of the token in the input text; the segment embedding is the paragraph to which the token belongs in the input text; the word embedding is the index position of the token in the BERT dictionary;

[0033] The set of judicial field events E = {E1, …, E N}, E1 to E N are the 1st to the Nth judicial field events; the set of text information corresponding to the judicial field events is S = {S1, …, S N}, S1 to SN The text information corresponding to the 1st to the Nth judicial domain events respectively; the set of trigger words corresponding to the judicial domain events is TR = {Tr1, …, Tr N}, and the set of event types is TY = {Ty1, …, Ty N}, Tr1 to Tr N are the trigger words corresponding to the 1st to the Nth judicial domain events respectively, and Ty1 to Ty N are the event types corresponding to the 1st to the Nth judicial domain events respectively; the maximum value of epoch is Epoches, and the number of batches in each epoch is batch_per_epoch. For the fine-tuned BERT model Bert_fine_tune, the maximum length of each sentence is max_len;

[0034] For all epochs in the BERT model, perform the following training process to obtain the learned model parameters for event element extraction:

[0035] For each batch in the epoch, perform SS1 to SS4:

[0036] SS1. Pad the input sentence with zeros or truncate it to a length of max_len, obtain the one-hot encoding of the event type, the start position and the end position of the trigger word;

[0037] SS2. Input the sentence into Bert_fine_tune to obtain the feature vector

[0038] SS3. The feature vector is followed by θ1 to predict the probability of the event type, followed by θ2 to predict the probability of the start position of the trigger word, and followed by θ3 to predict the probability of the end position of the trigger word;

[0039] SS4. Construct the second loss function L T (θ, θ1, θ2, θ3) = L1(θ, θ1) + L2(θ, θ2) + L3(θ, θ3), and use the Adam optimizer to minimize the second loss function.

[0040] SS5. When the loss on the validation set no longer decreases within a certain number of epochs, adopt the early stopping strategy.

[0041] Furthermore, the second loss function is defined as follows:

[0042] L T (θ, θ1, θ2, θ3) = L1(θ, θ1) + L2(θ, θ2) + L3(θ, θ3)

[0043] Among them, θ is the parameter of the Encoder part in the BERT model, and L1(θ,θ1), L2(θ,θ2), and L3(θ,θ3) respectively correspond to the loss function related to the event type prediction task, the loss function related to the trigger word start position prediction task, and the loss function related to the trigger word end position prediction task

[0044]

[0045] θ1 is the fully connected layer network parameter corresponding to the event type prediction task, Type is the one-hot representation of the input event type, which is a vector of length M, where M is the number of all event types, and Type i is the i-th element of Type. is the probability that the current event predicted by the model belongs to the event type, which is a vector of length M, is the probability that the current event type is predicted as the i-th event type.

[0046] Secondly,

[0047]

[0048] θ2 is the fully connected layer network parameter corresponding to the trigger word start position prediction task, L represents the maximum length of the input, and Start and are both vectors of length L. Specifically, Start is the one-hot representation of the trigger word start position in the input text, is the probability that each position of the current input predicted by the model is the start position of the trigger word. Start i is the i-th element of the vector Start; is the i-th element of the vector .

[0049] Finally,

[0050]

[0051] θ3 is the fully connected layer network parameter corresponding to the trigger word end position prediction task, L represents the maximum length of the input, and End and are both vectors of length L. Specifically, End is the one-hot representation of the trigger word end position in the input text, is the probability that each position of the current input predicted by the model is the end position of the trigger word. End i is the i-th element of the vector End i ; is the i-th element of the vector .

[0052] Further, after tokenizing the judicial data text, the position embedding, segment embedding, and word embedding of each token are obtained; the position embedding of the token is the position of the token in the input text, the segment embedding of the token is the paragraph to which the token belongs in the input text, and the word embedding of the token is the index position of the token in the BERT dictionary. The three embeddings are input into the BERT model in the judicial field obtained after tuning to obtain the feature vector of each token;

[0053] Specifically, step 4 is as follows: construct an event element extraction model, including constructing two exclusive prediction networks for each event element respectively to predict the start position and end position of the event element;

[0054] Among them, the exclusive prediction network for predicting the start position is specifically: for each event element, traverse all tokens, and input the feature vector obtained by the tuned BERT of the corresponding token into the fully connected layer for start position prediction, map it to a low-dimensional vector of length 2, and predict whether the current token is the start position of the event element through softmax; the exclusive prediction network for predicting the end position is the same as the exclusive prediction network for predicting the start position, and is used to predict whether each token is the end position of the corresponding event element.

[0055] The set of events E in the judicial field = {E1, …, E N}, E1 to E N are the 1st to the Nth events in the judicial field; the set of text information corresponding to the events in the judicial field is S = {S1, …, S N}, S1 to S N are the text information corresponding to the 1st to the Nth events in the judicial field respectively; the set of trigger words corresponding to the events in the judicial field is TR = {Tr1, …, Tr N} and the set of event types is TY = {Ty1, …, Ty N}, Tr1 to Tr N are the trigger words corresponding to the 1st to the Nth events in the judicial field respectively, and Ty1 to Ty N are the event types corresponding to the 1st to the Nth events in the judicial field respectively; the maximum value of epoch is Epoches, the number of batches in each epoch is batch_per_epoch. The tuned BERT model is Bert_fine_tune, and the maximum length of each sentence is max_len;

[0056] SSS1. Pad the input sentence with zeros or truncate it to a length of max_len, and obtain the event element list and the start position and end position of each event element;

[0057] SSS1. Input the sentence into Bert_fine_tune to obtain the feature vector

[0058] SSS1. Feature vector Followed by |R| and |R| Predict the start position and end position of each event element respectively;

[0059] SSS1. Construct the third loss function L = L s +L e , and use the Adam optimizer to minimize the third loss function;

[0060] SSS1. When the loss on the validation set no longer decreases within a certain number of epochs, adopt the early stopping strategy.

[0061] Furthermore, use the tuned BERT model in the underlying layer to extract the feature vector of the input event E, and then predict the probability that the token t in the input sentence is the start position s of the element r with the following probability

[0062]

[0063] where are the fully connected layer network parameters for predicting the start position of the element r; θ(E) is the feature vector extracted by the tuned BERT for the input event E.

[0064] And predict the probability that the token t in the input sentence is the end position e of the element r with the following probability

[0065]

[0066] where are the fully connected layer network parameters for predicting the end position of the element r; softmax is the normalized exponential function;

[0067] For the task of event element extraction, the third loss function is defined as follows:

[0068] L = L s +L e

[0069] L s is the loss function related to the start position, specifically as follows:

[0070]

[0071] Where S is the input sentence, |S| is the number of word segments in S, R is the set of all elements, and |R| is the number of all elements. is the binary true value indicating whether the corresponding position is the starting position s of the element r. is the predicted probability indicating whether the corresponding position is the starting position s of the element r; CrossEntropy() is the cross-entropy function.

[0072] Similarly, define L e as follows:

[0073]

[0074] Where is the binary true value indicating whether the corresponding position is the ending position e of the element r. is the predicted probability indicating whether the corresponding position is the ending position e of the element r.

[0075] Beneficial effects:

[0076] 1. The present invention provides a deep event extraction method in the judicial field that integrates multi-task and multi-label learning. Currently, aiming at the characteristics of judicial field texts, an event extraction model based on the pre-trained model BERT is proposed. The BERT is optimized on domain data through the masked LM method to learn feature representations more suitable for domain knowledge; the trigger word extraction and event classification tasks are combined, and the two tasks are unified into a single loss function in the form of multi-task, leveraging the correlation between tasks to promote the improvement of learning performance; the start and end annotations of event elements are used for learning and prediction, and for multiple event elements, corresponding network layers are designed for extraction respectively to reduce the mutual interference between different elements.

[0077] 2. The deep learning method based on the tuning of the pre-trained language model provided by the present invention does not require manual construction of feature patterns, saving a lot of manpower. The BERT model can extract features at the word level and sentence level, and can bring better performance by better feature extraction of semantic vectors. The tuning on judicial data improves the adaptability of the language model to judicial data and the accuracy of low-dimensional feature representation; jointly learn the two closely related tasks of trigger word recognition and event type classification, and use the correlation between tasks to promote the improvement of learning performance; in the form of multi-label classification problems, for multiple elements of the event, the corresponding network layers are designed for extraction, reducing the mutual interference between different elements, especially when there is overlap between the contents of two elements; by introducing [NAT] in the input data, the situation where the roles of some event elements are empty (non-existent) in the schema-based event extraction task is effectively solved, and the schemas of different event types are merged, and the collection of all event elements is extracted each time to achieve data enhancement. Combining the improvements in the above four aspects, the model performance is improved and applied to judicial data.

[0078] 3. The present invention provides a deep event extraction method in the judicial field that integrates multi-task and multi-label learning. It implements trigger word extraction and event classification based on the BERT pre-training model and multi-task, and jointly learns two closely related tasks, trigger word recognition and event type classification. It uses the correlation between tasks to promote the improvement of learning performance. At the same time, it can also alleviate the overfitting of the model in a form similar to regularization and improve the generalization ability of the model.

[0079] 4. The present invention provides a deep event extraction method in the judicial field that integrates multi-task and multi-label learning. In the form of a multi-label classification problem, for multiple elements of an event, corresponding network layers are designed for extraction to reduce mutual interference between different elements, especially when there is overlap between the contents of two elements. By predicting the starting and ending positions of event elements respectively, it avoids the need for additional constraints in the traditional BIO strategy, such as "B" must be followed by "I" instead of "B" and the decoding process, and truly realizes end-to-end learning.

[0080] 5. The present invention provides a deep event extraction method in the judicial field that integrates multi-task and multi-label learning. By introducing [NAT] (representing non-text information, no content) into the input data, it effectively solves the situation where the roles of some event elements are empty (non-existent) in the schema-based event extraction task; at the same time, the method can merge schemas of different event types and extract the collection of all event elements each time. While fully expanding the limited data set, [NAT] can also be used as a "negative sample" to balance the data distribution and improve the model effect. Description of the Drawings

[0081] Figure 1 It is a structural diagram of BERT pre-training and fine-tuning;

[0082] Figure 2 It is a schematic diagram of Input embedding;

[0083] Figure 3 It is a schematic diagram of the BERT pre-training task process;

[0084] Figure 4 It is a schematic diagram of the Mask LM task in the BERT pre-training stage;

[0085] Figure 5 It is a schematic diagram of the hard sharing mechanism;

[0086] Figure 6 It is a schematic diagram of the soft sharing mechanism;

[0087] Figure 7 It is a schematic diagram of domain data tuning based on BERT and Masked LM;

[0088] Figure 8 It is a flow chart of event trigger word extraction based on BERT and multi-task;

[0089] Figure 9 It is a flow chart of event element extraction based on BERT and multi-label classification tasks;

[0090] Figure 10 It is a schematic diagram of data augmentation. Detailed Implementation Manner

[0091] The present invention will be described in detail below in conjunction with the drawings and by way of examples.

[0092] A judicial domain event extraction technology that realizes trigger word extraction and event classification based on the BERT pre-training model and multi-task, and realizes event element extraction through multi-label classification on the augmented data.

[0093] The present invention applies the BERT pre-training model, multi-task and multi-label classification technologies, and the principle is as follows:

[0094] BERT Pre-trained Model: The BERT model comes from the paper "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding". It is a dynamic word vector technology. Different from traditional static word vectors, dynamic word vectors can generate word vectors dynamically according to specific context information. This model uses a bidirectional Transformer. By training on an unlabeled dataset and comprehensively considering context feature information, it further improves the generalization ability of word vectors. The model fully describes the relationship features at the character level, word level, sentence level, and even between sentences through two methods: Masked LM and Next Sentence Prediction, obtains more sufficient semantic information, and can better solve phenomena such as polysemy. It is planned to use the Chinese pre-trained language model BERT on judicial domain data and optimize the network with the MaskedLM language learning model to learn network parameters more suitable for domain knowledge and provide feature representations corresponding to the text for subsequent processing. The judicial domain data targeted by this invention includes judgment documents.

[0095] Structure: The BERT model mainly utilizes the Encoder structure of the Transformer. It uses the most primitive Transformer, but the model structure is deeper than that of the Transformer. The Transformer Encoder contains 6 Encoder blocks, the BERT-base model contains 12 Encoder blocks, and the BERT-large contains 24 Encoder blocks.

[0096] Training: The training is mainly divided into two stages: the pre-training stage and the Fine-tuning stage. The pre-training stage is obtained by training on a large dataset according to some pre-training tasks. The Fine-tuning stage is for fine-tuning when used for some downstream tasks later, such as text classification, part-of-speech tagging, question answering systems, etc. BERT can be fine-tuned on different tasks without adjusting the structure.

[0097] Pre-training Task 1: The first pre-training task of BERT is Masked LM. Randomly mask a part of the words (also called tokens) in the sentence, and then use the context information to predict the masked words at the same time, so that the meaning of the words can be better understood according to the full text. Masked LM is the focus of BERT.

[0098] Pretraining Task 2: The second pretraining task of BERT is Next Sentence Prediction (NSP). This task mainly enables the model to better understand the relationships between sentences.

[0099] (1) The BERT structure is as Figure 1 shown. Figure 1 In it, the left figure represents the pretraining process, and the right figure is the fine-tuning process for specific tasks.

[0100] The input of BERT can contain a pair of sentences (sentence A and sentence B), or it can be a single sentence. At the same time, BERT adds some special-purpose flag bits: The [CLS] flag is placed at the beginning of the first sentence, and the resulting representation vector C obtained by BERT can be used for subsequent classification tasks. The [SEP] flag is used to separate two input sentences. For example, for input sentences A and B, the [SEP] flag should be added after sentences A and B. The [MASK] flag is used to cover some words in the sentence. After covering the word with [MASK], the [MASK] vector output by BERT is then used to predict what the word is. For example, given two sentences "mydog is cute" and "he likes palying" as input samples, BERT will be transformed into "[CLS]my dog is cute[SEP]he likes play##ing[SEP]". The WordPiece method is used in BERT, and words will be split into sub-word units (SubWord), so some words will be split into roots. For example, "palying" will become "paly" + "##ing".

[0101] After BERT gets the sentence to be input, the words of the sentence need to be converted into Embedding, and Embedding is represented by E. Different from Transformer, the input Embedding of BERT is obtained by adding three parts: Token Embedding, Segment Embedding, and Position Embedding.

[0102] Figure 2 It is a schematic diagram of Input embedding.

[0103] Token Embedding: The Embedding of words, such as [CLS]dog, etc., is learned through training.

[0104] Segment Embedding: It is used to distinguish whether each word belongs to sentence A or sentence B. If only one sentence is input, only EA is used, and it is learned through training.

[0105] Position Embedding: Encodes the position where a word appears. Different from Transformer which uses a fixed formula for calculation, the Position Embedding in BERT is also learned. In BERT, it is assumed that the longest sentence is 512.

[0106] (2) Training:

[0107] After the Embedding of the words in the BERT input sentence, the model is trained through a pre-training method. There are two tasks in pre-training.

[0108] The first is Masked LM. Randomly replace some words in the sentence with [MASK], then pass the sentence into BERT to encode the information of each word, and finally use the encoded information T[MASK] of [MASK] to predict the correct word at that position.

[0109] The second is next sentence prediction. Input sentences A and B into BERT and predict whether B is the next sentence of A, using the encoded information C of [CLS] for prediction.

[0110] The process of BERT pre-training can be represented by Figure 3 to represent.

[0111] The BERT model obtained through pre-training can be fine-tuned (Fine-tuning stage) when used for specific NLP tasks later. The BERT model can be applied to a variety of different NLP tasks.

[0112] (3) Masked LM

[0113] When predicting a word, it is necessary to utilize both the left (previous context) and right (next context) information of the word to make the best prediction. Models like ELMo that perform left-to-right and right-to-left separately are called shallow bidirectional models. BERT hopes to train a deep bidirectional model on the Transformer Encoder structure, so the method of Mask LM is proposed for training.

[0114] Mask LM is used to prevent information leakage. For example, when predicting the word "natural", if the "natural" in the input part is not masked, then the information of "natural" can be directly obtained at the prediction output. Figure 4 Schematic diagram of the Mask LM task in the BERT pre-training stage.

[0115] When BERT is trained, it only predicts the words at the [Mask] positions, so that the context information can be utilized simultaneously. However, when it is used subsequently, the [Mask] words do not appear in the sentence, which will affect the performance of the model. Therefore, the following strategy is adopted during training: randomly select 15% of the words in the sentence for masking. Among the words selected for masking, 80% are truly replaced with [Mask], 10% are not replaced, and the remaining 10% are replaced with a random word.

[0116] For example, for the sentence "I like natural language processing", if the word "natural" is selected for masking, then: with a probability of 80%, the sentence "I like natural language processing" is converted into the sentence "I like [Mask] language processing"; with a probability of 10%, the sentence remains "I like natural language processing" unchanged; with a probability of 10%, the word "natural" is replaced with another random word, such as "English", and the sentence "I like natural language processing" is converted into the sentence "I like English language processing". The above is the first pre-training task of BERT, Masked LM.

[0117] (4)Next Sentence Prediction (NSP)

[0118] The second pre-training task of BERT is Next Sentence Prediction (NSP), that is, next sentence prediction. Given two sentences A and B, it is necessary to predict whether sentence B is the next sentence of sentence A.

[0119] The main reason why BERT uses this pre-training task is that many downstream tasks, such as question answering systems (QA) and natural language inference (NLI), require the model to be able to understand the relationship between two sentences, but this cannot be achieved by training a language model.

[0120] When BERT is being trained, there is a 50% probability of selecting two consecutive sentences A and B, and a 50% probability of selecting two non-consecutive sentences A and B. Then, through the output C of the [CLS] flag bit, it is predicted whether the next sentence of sentence A is sentence B.

[0121] multi-task learning

[0122] Multi-task learning is a machine learning method that is opposite to single-task learning. As the name implies, multi-task learning is a machine learning method that learns multiple tasks simultaneously. Multi-task learning simultaneously learns the classifiers for humans and dogs and the gender classifiers for men and women.

[0123] In the field of machine learning, the standard algorithm theory is to learn one task at a time. Complex learning problems are first decomposed into theoretically independent sub-problems, and then each sub-problem is learned separately. Finally, a mathematical model of the complex problem is established by combining the learning results of the sub-problems. Multi-task learning is a form of joint learning where multiple tasks are learned in parallel and the results influence each other.

[0124] There are many forms of multi-task learning, such as Joint Learning, Learning to Learn, Learning with Auxiliary Tasks, etc. These are just some of the aliases. Generally speaking, once it is found that more than one objective function is being optimized, multi-task learning can be used to effectively solve the problem.

[0125] Compared with standard single-task learning, there are two main challenges in training multiple tasks while learning shared representations:

[0126] Loss function (how to balance different tasks): The loss function of multi-task learning assigns weights to the losses of each task. In this process, it is necessary to ensure that all tasks are equally important and not let simple tasks dominate the entire training process. Manually setting weights is inefficient and not optimal. Therefore, it is very necessary and important to automatically learn these weights or design a network that is robust to all weights.

[0127] Network structure (how to achieve network parameter sharing): An efficient multi-task network structure must take into account both the feature-sharing part and the task-specific part. It is necessary to learn the generalization representation between tasks (to avoid overfitting) and also learn the unique features of each task (to avoid underfitting). In multi-task learning based on deep neural networks, two common methods are used: hard sharing and soft sharing of hidden layer parameters.

[0128] (1) Hard sharing mechanism of parameters: As Figure 5 shown, the hard sharing mechanism of parameters is the most common way in multi-task learning of neural networks. Generally speaking, it can be applied to all hidden layers of all tasks, while retaining the task-related output layers. The hard sharing mechanism reduces the risk of overfitting. Intuitively, this is very meaningful. The more tasks are learned simultaneously, the more the model can capture the same representation of multiple tasks, thus reducing the risk of overfitting on the original tasks.

[0129] (2) Soft sharing mechanism of parameters: As Figure 6As shown, in the soft sharing mechanism, each task has its own model and its own parameters. The similarity of the parameters is ensured by regularizing the distance between the model parameters, such as L2 distance regularization, trace norm regularization, etc. The constraints for the soft sharing mechanism in deep neural networks are largely influenced by the regularization techniques in traditional multi-task learning.

[0130] Multi-label classification learning

[0131] First, it should be pointed out that multi-label and multi-class are different. Multi-class is relative to binary classification, meaning that the things to be classified have more than two categories. It could be choosing one from 3 categories (such as iris classification) or one from 10 categories (such as handwritten digit recognition mnist). While multi-label is a more general case where a sample may have more than 1 label. For example, the label of a picture can include both a person and a dog at the same time. In our application, multi-classification means that the entities in a judicial event need to be classified into categories such as time, place, victim, etc. At the same time, the same description "for the purpose of illegal possession, stealing public and private property multiple times with a relatively large amount" can be both a criminal act and a criminal circumstance, that is, the same text description may have multiple labels.

[0132] From the perspective of solving problems, multi-label classification algorithms can be divided into two main categories: one is the method based on problem transformation, and the other is the method based on algorithm adaptation. The method based on problem transformation is to transform the problem data so that existing algorithms can be used; the method based on algorithm adaptation refers to expanding a specific algorithm to be able to process multi-label data, improving the algorithm to be applicable to the data.

[0133] Strategically, the main strategies of multi-label classification can be roughly divided into three categories: First-order strategy: Considering that the labels are independent of each other, the multi-label problem can be transformed into multiple ordinary classification problems. Second-order strategy: This category considers the pairwise correlation between labels, which will result in a significant increase in computational complexity. Higher-order strategy: This is to consider the correlation between multiple labels, and the computational complexity will be even higher. In this work, the semantic correlation between different labels is relatively low, such as the location of the crime and the time of the crime. Therefore, the first-order strategy can be used to transform the multi-label problem into multiple ordinary classification problems.

[0134] Although a two-layer network can theoretically fit all distributions, it is not easy to learn. Therefore, in practice, we usually increase the depth and width of the neural network to enhance its learning ability and facilitate fitting the distribution of the training data. In convolutional neural networks, some experiments have shown that depth is more important than width. However, as the neural network deepens, the number of parameters to be learned also increases, which makes it more likely to cause overfitting. When the dataset is small, too many parameters will fit all the characteristics of the dataset rather than the commonalities between the data. So, what is overfitting? As mentioned in a previous blog, it means that the neural network can highly fit the distribution of the training data, but has a very low accuracy for the test data and lacks generalization ability. Therefore, in this case, data augmentation emerged to prevent overfitting.

[0135] Data augmentation was initially applied more in the field of computer vision. It mainly uses various techniques to generate new training samples, and can create new data by translating, rotating, compressing, adjusting colors, etc. of the image. Although the 'new' samples change their appearance to some extent, the labels of the samples remain unchanged. And the data in NLP is discrete, which makes it impossible to directly and simply transform the input data. Changing one word may change the meaning of the whole sentence. There are generally two ideas for existing NLP data augmentation, one is adding noise, and the other is back-translation, both of which are supervised methods. Adding noise means creating new data similar to the original data by replacing words, deleting words, etc. on the basis of the original data. Back-translation is to translate the original data into another language and then translate it back to the original language. Due to differences in language logical order, etc., the back-translation method can often obtain new data that is quite different from the original data.

[0136] Several commonly used text augmentation techniques for adding noise are Synonyms Replace (SR), Randomly Insert (RI), Randomly Swap (RS), and RandomlyDelete (RD). A simple introduction is as follows:

[0137] (1) Synonyms Replace (SR): Without considering stopwords, randomly select n words in the sentence, and then randomly select synonyms from the synonym dictionary and replace them.

[0138] Eg: “I like this movie very much” -> “I like this film very much”. The sentence still has the same meaning and is very likely to have the same label.

[0139] (2) Randomly Insert (RI): Without considering stopwords, randomly select a word, then randomly choose one from its set of synonyms and insert it at a random position in the original sentence. This process can be repeated n times.

[0140] Eg: “I like this movie very much” -> “Love I like this movie very much”.

[0141] (3) Randomly Swap (RS): In a sentence, randomly select two words and swap their positions. This process can be repeated n times.

[0142] Eg: “How to evaluate the 2017 Zhihu Kanshan Cup Machine Learning Competition?” -> “2017 Machine Learning? How to evaluate the Kanshan Cup in Zhihu?”

[0143] (4) Randomly Delete (RD): For each word in a sentence, randomly delete it with probability p.

[0144] Eg: “How to evaluate the 2017 Zhihu Kanshan Cup Machine Learning Competition?” -> “How 2017 Kanshan Cup Machine Learning”.

[0145] The technical solution of the present invention includes the following four steps:

[0146] Step 1: Take the data in the judicial field for manual annotation. The annotated labels include event types and event elements, and a judicial field dataset is obtained.

[0147] Step 2: Domain data tuning based on BERT and Masked LM: Use Masked LM to tune the parameters of BERT on the manually annotated judgment document dataset of larceny (the dataset contains label system data such as event type - event element, etc.). The following strategy is adopted during training. Randomly select 15% of the words in the sentence for masking. Among the words selected for masking, 80% are really replaced with [Mask], 10% are not replaced, and the remaining 10% are replaced with a random word. Generate word vectors containing context information to provide feature representations corresponding to the text. The flowchart is as Figure 7 .

[0148] Its loss function is defined as follows:

[0149]

[0150] where θ is the parameter of the Encoder part in BERT. The input passes through θ to obtain the feature vector. θ0 is the parameter connected after θ in the Masked LM task, and |V| is the size of the dictionary composed of the masked words.

[0151] During the training process, to prevent excessive perturbation to the BERT parameters, in the first two epochs, θ is fixed, and θ0 is adjusted with a learning rate of lr = 5e -4 and in the subsequent epochs, θ and θ0 are adjusted simultaneously with a learning rate of lr = 1e -5 until the stopping condition is reached. The process of optimizing BERT on judicial domain data using the Masked LM technique is shown in Algorithm 1.

[0152]

[0153]

[0154] Step 3: Event trigger word extraction based on BERT and multi-task: The BERT model is adopted to extract the semantic information of the text, which is used as the input, and the loss function jointly defined by three tasks including event type prediction, trigger word start position prediction, and trigger word end position prediction is used for optimization in the multi-task. The flowchart is as Figure 8 shown.

[0155] Specifically, after tokenizing the judicial data text, the position embedding (i.e., the position of the token in the input text), segment embedding (i.e., which segment of the input text the token belongs to, defaulting to the first segment, i.e., 0 in this work), and word embedding (i.e., the index position of the token in the BERT dictionary) of each token are obtained. The three embeddings are input into the optimized BERT model to obtain the feature vector of each token. Subsequently, for Task 1, the feature vector represented by [CLS] obtained (this vector represents the overall semantic information of the text) is input into the fully connected layer, mapped into a vector with a length equal to the number of event types, and then followed by softmax to predict the event type (e.g., Figure 8 if the probability that task 1 predicts the current event type as Type A is 0.7 in, then the current event type is considered Type A); for Task 2, the feature vector of each token obtained from the optimized BERT is followed by a fully connected layer, mapped into a vector with a length of 2, and then followed by softmax to predict whether the corresponding token is the start position of the event trigger word (e.g., Figure 8 task 2 in believes that the character "steal" is the start position of the trigger word "steal" for the "theft event"); for Task 3, the network design is the same as that of Task 2, predicting whether each token at each position is the end position of the event trigger word (e.g., Figure 8 task 3 in believes that the character "take" is the end position of the trigger word "steal" for the "theft event").

[0156] Its loss function is defined as follows:

[0157] L T (θ, θ1, θ2, θ3) = L1(θ, θ1) + L2(θ, θ2) + L3(θ, θ3)

[0158] Where θ is the parameter of the Encoder part in the BERT model, and L1(θ, θ1), L2(θ, θ2), and L3(θ, θ3) correspond to the loss function related to the event type prediction task, the loss function related to the trigger word start position prediction task, and the loss function related to the trigger word end position prediction task, respectively.

[0159] Where θ is defined as described above. Additionally,

[0160]

[0161] θ1 is the fully connected layer network parameter corresponding to the event type prediction task. Type is the one - hot representation of the input event type, which is a vector of length M, where M is the number of all event types. Type i is the i - th element of Type; is the probability that the current event predicted by the model belongs to the event type, which is a vector of length M, is the probability that the current event type is predicted as the i - th event type.

[0162] θ1 is the fully connected layer network parameter specific to the event type prediction of Task 1, is defined as follows:

[0163]

[0164] N represents the number of event types. Type is the one - hot representation of the type of the input event, is the probability that the current event predicted by the model belongs to N event types.

[0165] Secondly,

[0166]

[0167] θ2 is the fully connected layer network parameter corresponding to the trigger word start position prediction task. L represents the maximum length of the input. Start and are both vectors of length L. Specifically, Start is the one - hot representation of the start position of the trigger word in the input text, is the probability that each position of the current input is the start position of the trigger word predicted by the model; Start i is the i - th element of the vector Start; is the i-th element of the vector in

[0168] θ2 is the fully-connected layer network parameter specific to the prediction of the start position of the trigger word for Task 2, which is defined as follows:

[0169]

[0170] L represents the maximum length of the input. Start is the one-hot representation of the start position of the trigger word in the input text, and

[0171] Finally,

[0172]

[0173] θ3 is the fully-connected layer network parameter corresponding to the prediction of the end position of the trigger word. L represents the maximum length of the input. End and are both vectors of length L. Specifically, End is the one-hot representation of the end position of the trigger word in the input text, and i is the i-th element of the vector End i in is the i-th element of the vector in

[0174] θ3 is the fully-connected layer network parameter specific to the prediction of the start position of the trigger word for Task 3, which is defined as follows:

[0175]

[0176] L represents the maximum length of the input. End is the one-hot representation of the end position of the trigger word in the input text, and

[0177]

[0178]

[0179] Step 4: Event Element Extraction Method Based on BERT and Multi-Label Classification Tasks: The BERT model is adopted to extract the semantic information of the text, which is used as the input. Exclusive network parameters are learned for each event element. At the last layer of the corresponding network, for each token, it is predicted whether it belongs to the start position or the end position of the current event element. The flow chart is as Figure 9 shown.

[0180] Specifically, after tokenizing the judicial data text, the position embedding of each token (i.e., the position of the token in the input text), the segment embedding (i.e., which segment of the input text the token belongs to, defaulting to the first segment, i.e., 0 in this work), and the word embedding (i.e., the index position of the token in the BERT dictionary) are obtained. The three embeddings are input into the fine-tuned BERT model to obtain the feature vector of each token. Subsequently, two exclusive prediction networks are trained for each event element to predict the start position and the end position of the event element respectively. Taking the exclusive prediction network for predicting the start position as an example, for each event element, all tokens are traversed, and the feature vector of the corresponding token obtained by the fine-tuned BERT is input into the fully connected layer for start position prediction, mapped into a low-dimensional vector of length 2, and whether the current token is the start position of the event element is predicted through softmax. The exclusive prediction network for predicting the end position is the same, used to predict whether each token is the end position of the corresponding event element.

[0181] Each event type corresponds to different event elements. For example, for a theft event, the event elements include the occurrence time, location, suspect, and stolen items; the event elements included in an access event are: access time, departure time, access location, access party, accessed party, etc.

[0182] At the bottom layer, the fine-tuned BERT model is used to extract the feature vector of the input event E. Subsequently, the probability that the token t in the input sentence is the start position s of the element r is predicted as follows:

[0183]

[0184] where, is the network parameter of the fully connected layer for start position prediction of the element r;

[0185] And the probability that the token t in the input sentence is the end position e of the element r is predicted as follows:

[0186]

[0187] where, is the network parameter of the fully connected layer for end position prediction of the element r.

[0188] For the task of event feature extraction, the loss function is defined as follows:

[0189]

[0190] is the loss function related to the starting position, as follows:

[0191]

[0192] Where S is the input sentence, |S| is the number of word segments in S, R is the set of all elements, and |R| is the number of all elements. is the binary true value of whether the corresponding position is the starting position s of element r, is the predicted probability of whether the corresponding position is the starting position s of element r. Similarly, we can define as follows:

[0193]

[0194] The learning process of the event feature extraction model based on multi-label learning is shown in Algorithm 3.

[0195]

[0196]

[0197] Introducing [NAT] (representing non-text information, no content) into BERT’s dictionary, and setting the last bit of the input data (i.e., the position embedding, segment embedding, and word embedding of each word segment obtained after the aforementioned judicial data text is segmented) to [NAT], effectively solving the problem that the roles of some event elements are empty (non-existent) in the schema-based event extraction task; at the same time, this method can merge schemas of different event types and extract the collection of all event elements each time. In the case of fully expanding the limited data set, [NAT] can also be used as a "negative sample" to balance the data distribution. Data enhancement such as Figure 10 As shown,

[0198] (1) The BERT pre-trained model is combined with multi-task to implement trigger word extraction and event classification. The two closely related tasks of trigger word recognition and event type classification are jointly learned. The correlation between tasks is used to promote the improvement of learning performance. At the same time, it can also alleviate the overfitting of the model in a form similar to regularization and improve the generalization ability of the model.

[0199] (2) In the form of a multi-label classification problem, for each of the multiple elements of an event, the corresponding network layers are designed for extraction to reduce the mutual interference between different elements, especially when the content of two elements overlaps. By predicting the starting and ending positions of event elements separately, the need for additional constraints such as "B" must be followed by "I" instead of "B" and the decoding process in the traditional BIO strategy is avoided, thus truly achieving end-to-end learning.

[0200] (3) By introducing [NAT] (representing non-text information, no content) into the input data, the problem that the roles of some event elements are empty (non-existent) in the schema-based event extraction task is effectively solved; at the same time, this method can merge the schemas of different event types and extract the collection of all event elements each time. While fully expanding the limited data set, [NAT] can also be used as a "negative sample" to balance the data distribution and improve the model effect.

[0201] Using a deep learning method based on pre-trained language model tuning, there is no need to manually construct feature patterns, saving a lot of manpower. The BERT model can extract features at the word level and sentence level. By better extracting features from semantic vectors to bring better performance, the tuning on judicial data improves the adaptability of the language model to judicial data and the accuracy of low-dimensional feature representation; jointly learning trigger word recognition and event type classification, two closely related tasks, uses the correlation between tasks to promote the improvement of learning performance; in the form of multi-label classification problems, for multiple elements of an event, the corresponding network layers are designed for extraction to reduce mutual interference between different elements, especially when there is overlap between the contents of two elements; by introducing [NAT] in the input data, the situation where the roles of some event elements are empty (non-existent) in the schema-based event extraction task is effectively solved, and the schemas of different event types are merged, and the collection of all event elements is extracted each time to achieve data enhancement. Combining the above four improvements, the model performance is improved and applied to judicial data.

[0202] In summary, the above are only preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A deep event extraction method in the judicial field that integrates multi-task learning and multi-label learning, characterized in that, Including the following steps Step 1: Obtain data in the judicial field for manual annotation. The labels to be annotated include event types and event elements, so as to obtain a judicial field dataset; Step 2: Use the Chinese pre-trained language model BERT on the judicial field dataset and adopt the MaskedLM language learning model to optimize the network, learning network parameters suitable for judicial field knowledge, so as to obtain a judicial field BERT model. Use the output of the judicial field BERT model as the semantic information of the text; Step 3: Construct a multi-task network. The multi-task network uses the semantic information of the text extracted by the judicial field BERT model as input. The multi-task network defines a loss function jointly defined by three tasks: prediction of the start position of the trigger word, prediction of the end position of the trigger word, and prediction of the event type for optimization. The output of the multi-task network includes the predicted event type, the predicted start position of the trigger word, and the predicted end position of the trigger word; Step 4: Determine event elements according to the event type and construct an event element extraction model. The event element extraction model uses the semantic information of the text extracted by the judicial field BERT model as input and learns exclusive network parameters for each event element. At the last layer of the network corresponding to each event element, for each tokenized token, predict whether it belongs to the start position or the end position of the current event element.

2. The method according to claim 1, characterized in that, The use of the MaskedLM language learning model to optimize the network of the Chinese pre-trained language model BERT on the judicial field dataset specifically is as follows: Use MaskedLM to optimize the parameters of the BERT model on the manually annotated judicial field dataset. During training, adopt the following strategy: randomly select 15% of the words in the sentence for masking. Among the words selected for masking, 80% are really replaced with [Mask], 10% are not replaced, and the remaining 10% are replaced with a random word.

3. The method according to claim 1, characterized in that, The specific content of Step 2 is as follows: The set of judicial domain events E = {E1, …, E N}, where E1 to E N are the 1st to the Nth judicial domain events; the set of text information corresponding to the judicial domain events is S = {S1, …, S N}, and S1 to S N are the text information corresponding to the 1st to the Nth judicial domain events respectively; the maximum value of epochs in the BERT model is Epoches, and the number of batches per epoch is batch_per_epoch; the BERT base model is Bert_base_chinese, and the maximum length of each sentence is max_len; For all epochs in the BERT model, execute the following training process to obtain the optimized BERT model parameters: For each batch in the epoch, execute S1~S4: S1 Pad the input sentence with zeros or truncate it to a length of max_len to obtain the index I1 of the tokenized sentence; S2 Randomly select 15% of the words in the sentence for masking. Among the words selected for masking, 80% are really replaced with [Mask], 10% are not replaced, and the remaining 10% are replaced with a random word; S3 Input the masked sentence into the BERT base model Bert_base_chinese to obtain a feature vector, and then connect with θ0 to predict the index I2 of the token corresponding to each position of the sentence; S4 uses the Adam optimizer to minimize the difference between I1 and I2, defined as the first loss function L(θ,θ0); when the first loss function on the validation set no longer decreases within a certain number of epochs, the early stopping strategy is adopted; The first loss function L(θ,θ0) is defined as follows: Among them, θ are the parameters of the Encoder part in the BERT model. The input passes through θ to obtain a feature vector. θ0 are the parameters following θ in the MaskedLM task. |V| is the size of the vocabulary composed of the masked words; m i represents the masked word; p(m = m i |θ, θ0) represents the probability that the predicted word m is the masked word m i given the learned parameters θ and θ0; In the training process, in the first two epochs of the BERT model, fix θ, with a learning rate of lr = 5e -4 Adjust θ0. In the subsequent epochs, adjust both θ and θ0 with a learning rate of lr = 1e -5 until the stopping condition is reached.

4. The method according to claim 3, characterized in that, The specific steps of step 3 are as follows: After tokenizing the judicial data text, the position embedding, segment embedding, and word embedding of each token are obtained. The three embeddings are input into the fine-tuned BERT model in the judicial field to obtain the feature vector of each token, which is the semantic information of the text; The position embedding is the position of the token in the input text; the segment embedding is the paragraph to which the token belongs in the input text; the word embedding is the index position of the token in the BERT dictionary; The set of judicial domain events E = {E1, …, E N}, where E1 to E N are the 1st to the Nth judicial domain events; the set of text information corresponding to the judicial domain events is S = {S1, …, S N}, where S1 to S N are the text information corresponding to the 1st to the Nth judicial domain events respectively; the set of trigger words corresponding to the judicial domain events is TR = {Tr1, …, Tr N} and the set of event types is TY = {Ty1, …, Ty N}, where Tr1 to Tr N are the trigger words corresponding to the 1st to the Nth judicial domain events respectively, and Ty1 to Ty N are the event types corresponding to the 1st to the Nth judicial domain events respectively; the maximum value of epoch is Epoches, the number of batches per epoch is batch_per_epoch; the fine-tuned BERT model is Bert_fine_tune, and the maximum length of each sentence is max_len; For all epochs in the BERT model, the following training process is executed to obtain the model parameters for event element extraction learned: For each batch in the epoch, execute SS1~SS4: SS1. Pad the input sentence with zeros or truncate it to a length of max_len, obtain the one-hot encoding of the event type, the start position and end position of the trigger word; SS2. Input the sentence into Bert_fine_tune to obtain the feature vector SS3. Feature vector Followed by θ1 to predict the probability of the event type, followed by θ2 to predict the probability of the start position of the trigger word, and followed by θ3 to predict the probability of the end position of the trigger word; SS4. Construct the second loss function L T (θ, θ1, θ2, θ3) = L1(θ, θ1) + L2(θ, θ2) + L3(θ, θ3), and use the Adam optimizer to minimize the second loss function; SS5. When the loss on the validation set no longer decreases within a certain number of epochs, the early stopping strategy is adopted.

5. The method according to claim 4, characterized in that, The second loss function is defined as follows: L T (θ, θ1, θ2, θ3) = L1(θ, θ1) + L2(θ, θ2) + L3(θ, θ3) Among them, θ is the parameter of the Encoder part in the BERT model, and L1(θ,θ1), L2(θ,θ2), and L3(θ,θ3) correspond to the loss function related to the event type prediction task, the loss function related to the trigger word start position prediction task, and the loss function related to the trigger word end position prediction task respectively θ1 is the fully connected layer network parameter corresponding to the event type prediction task. Type is the one-hot representation of the input event type, which is a vector of length M, where M is the number of all event types. Type i is the i-th element of Type; is the probability that the current event predicted by the model belongs to the event type, which is a vector of length M, is the probability that the current event type is predicted as the i-th event type; Secondly: θ2 are the fully connected layer network parameters corresponding to the trigger word start position prediction task, L represents the maximum length of the input, Start and are both vectors of length L. Specifically, Start is the one-hot representation of the start position of the trigger word in the input text, is the probability that the current input positions are the start positions of the trigger word predicted by the model; Start i is the i-th element in the vector Start; is the i-th element in the vector ; Finally, θ3 is the fully connected layer network parameter corresponding to the trigger word end position prediction task, L represents the maximum length of the input, End and are both vectors of length L. Specifically, End is the one-hot representation of the end position of the trigger word in the input text, is the probability that each position of the current input predicted by the model is the end position of the trigger word; End i is the i-th element in the vector End i ; is the i-th element in the vector .

6. The method according to claim 5, wherein, After tokenizing the judicial data text, the position embedding, segment embedding, and word embedding of each token are obtained; the position embedding of the token is the position of the token in the input text, the segment embedding of the token is the paragraph to which the token belongs in the input text, the word embedding of the token is the index position of the token in the BERT dictionary, and the three embeddings are input into the fine-tuned BERT model in the judicial field to obtain the feature vector of each token; The specific steps of step 4 are as follows: construct an event element extraction model, including constructing two exclusive prediction networks for each event element to predict the start position and end position of the event element respectively; Among them, the exclusive prediction network for predicting the start position is specifically: for each event element, traverse all tokens, input the feature vector obtained by passing the corresponding token through the fine-tuned BERT into the fully connected layer for start position prediction, map it to a low-dimensional vector of length 2, and predict whether the current token is the start position of the event element through softmax; the exclusive prediction network for predicting the end position is the same as the exclusive prediction network for predicting the start position, and is used to predict whether each token is the end position of the corresponding event element; The set of events in the judicial domain \(E = \{E_1, \ldots, E\) N \}, where \(E_1\) to \(E\) N are the 1st to the \(N\)th events in the judicial domain; the set of text information corresponding to the events in the judicial domain is \(S=\{S_1, \ldots, S\) N \}, and \(S_1\) to \(S\) N are the text information corresponding to the 1st to the \(N\)th events in the judicial domain respectively; the set of trigger words corresponding to the events in the judicial domain is \(TR = \{Tr_1, \ldots, Tr\) N \} and the set of event types is \(TY=\{Ty_1, \ldots, Ty\) N \}, where \(Tr_1\) to \(Tr\) N are the trigger words corresponding to the 1st to the \(N\)th events in the judicial domain respectively, and \(Ty_1\) to \(Ty\) N are the event types corresponding to the 1st to the \(N\)th events in the judicial domain respectively; the maximum value of the epoch is Epoches, the number of batches per epoch is batch_per_epoch; the fine-tuned BERT model is Bert_fine_tune, and the maximum length of each sentence is max_len; SSS1. Pad the input sentence with zeros or truncate it to a length of max_len, and obtain the list of event elements and the start and end positions of each event element; SSS1. Input the sentence into Bert_fine_tune to obtain the feature vector SSS1. Feature vector Followed by |R| And |R| Predict the start position and end position of each event element respectively; SSS1. Construct the third loss function \(L = L\) s + L e , and use the Adam optimizer to minimize the third loss function; SSS1. Adopt the early stopping strategy when the loss on the validation set no longer decreases within a certain number of epochs; where R is the set of all elements, and |R| is the number of all elements. The fully-connected layer network parameters predicted for the start position of element r The fully-connected layer network parameters predicted for the end position of element r, L s is the loss function related to the start position.

7. The method according to claim 6, wherein, At the bottom layer, use the optimized BERT model in the judicial field to extract the feature vector of the input event E, and then predict the probability that the token t in the input sentence is the starting position s of the element r with the following probability Among them, The fully connected layer network parameters predicted for the starting position of element r; θ(E) is the feature vector extracted by the fine-tuned BERT for the input event E; And predict the probability that the word segmentation t in the input sentence is the end position e of the element r with the following probability Among them, is the fully connected layer network parameter predicted for the end position of element r; softmax is the normalization exponential function; For the task of event element extraction, the third loss function is defined as follows: L = L s + L e L s is a loss function related to the starting position, specifically as follows: Among them, S is the input sentence, |S| is the number of word segments in S, R is the set of all elements, and |R| is the number of all elements. is the binary true value indicating whether the corresponding position is the starting position s of the element r. is the predicted probability indicating whether the corresponding position is the starting position s of the element r; CrossEntropy() is the cross-entropy function. Similarly, define L e as follows: Among them, is the binary true value indicating whether the corresponding position is the end position e of the element r, and is the predicted probability indicating whether the corresponding position is the end position e of the element r.

Citation Information

Patent Citations

  • Chinese text key information extraction method based on pre-trained language model

    CN111444721A

  • Case event knowledge graph construction method and related equipment

    CN112632223A