A multi-charge prediction method
The case description and crime were encoded through the UniLM model, and the case-crime attention mechanism was introduced, which solved the problem of the logical relationship and semantic information of the crimes in the prediction of multiple crimes, and improved the accuracy of the prediction of multiple crimes.
Patent Information
- Application Number
- CN202210124252.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-10
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-02-10
AI Technical Summary
The prior art fails to effectively consider the logical relationship and semantic information between crimes in the multi-crime prediction task, resulting in the multi-label classification model performing poorly in multi-crime cases.
UniLM's bidirectional language model is used to encode the case description and crime, introduce the case-crime attention mechanism, calculate the correlation between the hidden state of the current time step of the decoder and all crimes, and integrate it into the decoding process.
The accuracy of multiple crime prediction has been significantly improved, and the accuracy of multiple crime prediction has been improved by obtaining semantic information of each crime and the relationship between the crimes.
Smart Images

Figure CN114519103B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a multi - charge prediction method. Background Art
[0002] In today's society, with the strong promotion of relevant government departments and major information platforms, people's awareness of safeguarding their rights and interests through legal means has been increasing day by day. At present, artificial intelligence technology is showing an expanding trend in various construction fields of intelligent courts. After being trained with a large amount of data, the deep - learning model can quickly learn important information in case facts and make corresponding judgments, thus reducing the burden on judicial practitioners and improving work efficiency.
[0003] The charge prediction task is one of the core tasks in the field of intelligent trial, which aims to automatically predict the charges violated by the criminal subject based on the facts of the crime. Existing research usually models the charge prediction task as a multi - label classification problem for solution. The case description text is compressed into a context vector integrating semantic information through an encoder, and then this vector is input into a classifier to generate the probability of each charge. Finally, a prior threshold is set. If the probability is greater than the threshold, it is considered that the charge corresponds to the case.
[0004] Such methods perform well in single - charge cases, but poorly in multi - charge cases, mainly having the following problems: First, converting multi - charge prediction into multiple single - charge predictions does not consider the potential logical relationships between charges and is not interpretable; Second, differentiating different charges with numbers or symbols not only does not consider the semantic information carried by the charges themselves, but also the Euclidean distance of the charge encoding vectors cannot well reflect the differences between charges. Summary of the Invention
[0005] Object of the Invention: The technical problem to be solved by the present invention is to provide a multi - charge prediction method aiming at the deficiencies of the prior art. From the perspective of natural - language generation, it models the case - description sequence and the charge sequence, encodes the case description and all charges respectively using the bidirectional language model of UniLM (Unified Pre - trained Language Model), and introduces a case - charge attention mechanism to alleviate the deficiencies of the multi - label classification model.
[0006] To achieve the above object, the present invention provides a multi - charge prediction method integrating charge semantics and case - charge attention, which encodes the case description and all charges respectively, calculates the correlation between the hidden state of the decoder at the current time step and all charges, and integrates the result into the current hidden state of the decoder. The specific steps are as follows:
[0007] Input the case description text, obtain the key information in the case description through a text automatic summarization model, splice the case description and the case summary, and the input representation method is the same as that of Bert. Use the bidirectional language model of the UniLM model to encode it, and convert the input text into a vector representation;
[0008] Also use the bidirectional language model of the UniLM model to encode all charges, and the output is the corresponding charge encoding sequence. This sequence can not only contain the semantic information of all charges, reflect the differences between charges, but also learn the mutual relationships between charges;
[0009] Take the output corresponding to the last position of the case encoding module as the current hidden state of the decoder, which contains the relevant important information of the forward sequence including the generated charges. Calculate the correlation coefficient between it and each item in the charge encoder output sequence, and normalize the correlation coefficient through the softmax function of the normalized exponential function, which represents the attention weight. Multiply the output of each charge encoder by its corresponding normalized weight and then sum to obtain the charge attention encoding vector at the current time step;
[0010] Use the unidirectional language model of the UniLM model to complete charge prediction, which is only visible to the forward sequence. The hidden state is the output vector corresponding to the last position of the case encoding module. Splice it with the charge attention encoding vector at the current time step as the input of the current decoder, and map it to the dimension of the output dictionary through the fully connected layer for charge prediction until the end marker is predicted.
[0011] The method of the present invention includes the following steps:
[0012] Step 1, train the text summarization model, input the case description into the trained text summarization model, and obtain the case summary;
[0013] Step 2, splice the case description and the case summary,
[0014] Step 3, streamline all charges into a string with a length of X1 (usually taking the value of 504);
[0015] Step 4, take the case encoding vector obtained in Step 2 as the current hidden state in the decoding process, calculate the correlation coefficient between it and each item in the charge encoding sequence in Step 3, and normalize the correlation coefficient through the softmax function of the normalized exponential function, which represents the attention weight. Multiply the output of each charge encoder by its corresponding normalized weight and then sum to obtain the case and charge attention encoding vector at the current time step;
[0016] Step 5: Concatenate the case coding vector with the case and charge attention coding vector as the input at the current time step in the decoding process. The decoding process is implemented by multiple fully connected neural networks. Map the input through a fully connected layer (consisting of multiple fully connected neural networks) to the dimension of the Chinese dictionary, and the charge corresponding to the position with the highest probability is the charge predicted at the current time step.
[0017] Step 1 includes: The text summarization model also selects the UniLM model, and the training dataset is the THUCNews news dataset. When training the text summarization model, the input content includes: "[CLS]News content[SEP]News summary [SEP]", where "[CLS]" is the start token, "[SEP]" is the end token. Mask the words in the news summary through the masking mechanism of UniLM, and let the UniLM model learn to recover the masked words one by one. The training objective is to maximize the likelihood of the masked words based on the context. The end token "[SEP]" can also be masked, and the model ends the prediction when it predicts the end token.
[0018] Step 2 includes: Add "@@@" as a separator between the case description and the case summary. Perform word segmentation and word embedding respectively through the word-level tokenizer given by Bert official and the Chinese dictionary, convert the input text into a vector sequence, and use the bidirectional language model of the UniLM model to encode this sequence. Take the last vector of the output as the case coding vector at the current time step.
[0019] In Step 2, during training, the case description and the corresponding charge are input in the form of sentence pairs. During testing, only the case description is input, in the format: "[CLS]Case description@@@Case summary[SEP]", where "@@@" is used to distinguish the case description and the case summary, and "[CLS]" and "[SEP]" are special tokens for the start and end of the sentence respectively. The representation of each word is composed of the combination of word embedding, position embedding, and segment embedding.
[0020] In Step 2, first, the word-level tokenizer of Bert is used to tokenize the concatenated case description and case summary, and return an array of the tokenized words. Then, according to the one-to-one correspondence between the words and values in the Chinese dictionary, convert the array of words into an array of values, and use the nn.embedding method of the deep learning pytorch framework to convert the one-hot encoding of each word into a 768-dimensional dense vector.
[0021] The position embedding encodes the position information of the word into a feature vector, thus introducing the word position relationship;
[0022] Step 2 also includes: for an input of length 512, the word vector dimension is 768, the positional embedding is a query table of (512, 768), and the positional embedding at each position of the sequence corresponds to the corresponding row in the table, and the values therein are continuously learned during the model training process. The segment embedding is used to distinguish two sentences, for example, whether B is the context of A (dialogue scenario, question-and-answer scenario, etc.), and can be used as an identifier for what training method the model adopts (unidirectional, bidirectional, sequence-to-sequence). For a sentence pair, the feature value of the first sentence is 0, and the feature value of the second sentence is 1.
[0023] Step 2 also includes: The backbone network of the UniLM model includes 24 layers of Transformer networks (Transformer is an attention-based encoder-decoder framework). After word embedding, the input vector of the UniLM model is transformed into a sequence H0 = [x1,..., x |x| composed of 768-dimensional word vectors, which is fed into 24 layers of Transformer networks to fuse context information at different layers. Each layer of Transformer uses multi-head attention to fuse the vector output by the previous layer. The output encoded by the l-th layer is representing the output of the word vector at the corresponding position encoded by the l-th layer;
[0024] For the l-th layer of Transformer network, the calculation method of the output of the self-attention head in the Transformer network is:
[0025]
[0026]
[0027]
[0028] The output H of the previous layer l-1 is linearly projected into a query vector Q, a key vector K, and a value vector V through three parameter matrices of the l-th layer respectively (all three matrices are randomly initialized when created, and the parameters will be updated during the training process and are part of the model parameters). d k is the dimension of the word vector, and the mask matrix M is used to control whether the information at the corresponding position is visible to the context. M ij represents the value at the i-th row and j-th column in the mask matrix M; if the value of M ij is 0, then allow toattend means that all words can be accessed. If the value of M ijIf the value is -∞, then "prevent from attending" means invisible to the context. The first segmented input is the input for crime prediction, which consists of the case description text and the case summary. When the model encodes the word vectors of this part, it uses the bidirectional language model of UniLM for encoding, that is, all elements of the mask matrix M are set to 0, allowing the model to fully learn the context information. For the second segmented part, that is, the corresponding crime part, a unidirectional language model needs to be used for encoding, and the masking method is the same as that of the Transformer model. To enable the model to have the ability of parallel computing, the sequence is copied N times, where N is the length of the second segmented part, and then added to a mask matrix with 0 in the lower triangle and -∞ in the upper triangle, so as to achieve the purpose of masking the word vectors backward from the current position, so that when the model makes a prediction at the current position, it can only notice the information of the forward sequence, and the last position of the case encoding is taken as the current hidden state of the decoder.
[0029] Step 3 includes: performing word segmentation and word embedding through the word-level tokenizer and Chinese dictionary given by Bert official respectively, converting the input text into a vector sequence, and encoding the vector sequence using the bidirectional language model of the UniLM model, and taking the output vectors of all positions after encoding.
[0030] For the crime encoding in Step 3, the output vectors of all positions after encoding are taken, where the word embedding and encoding method are the same as those in Step 2, and will not be elaborated here.
[0031] Step 4 includes: taking the case encoding vector of Step 2 as the current hidden state in the decoding process, calculating the correlation coefficient between the current hidden state and each item in the crime encoding. For the first calculation of attention, the output vector corresponding to the end marker [SEP] is used as the initial hidden state hidden of the decoder, and dot product operations are performed on the hidden vector and the output sequence [key1, key2,..., key n of the crime encoding in Step 3:
[0032] similarity(hidden, key i ) = hidden · key i
[0033] where key i is the i-th item in the output sequence of the crime encoding, n is the length of the crime sequence, similarity is the correlation coefficient between the current hidden state hidden and all crimes, and the obtained correlation coefficients are normalized using the softmax function of the normalized exponential function:
[0034]
[0035] where similarityi represents the value of the i-th item in similarity, L x represents the total length of similarity, which is also the length of the crime name coding output sequence, a i represents the attention weight of the i-th item in the output sequence of the crime name coding, and calculates the case and crime name attention coding vectors Attention(hidden, keys) at the current time step;
[0036] In step 4, the case and crime name attention coding vectors Attention(hidden, keys) at the current time step are calculated using the following formula:
[0037]
[0038] where keys represents the output sequence of the crime name coding, and key i is the value at the i-th position in keys. Obtain the case-crime name attention coding vector at the current time step, which contains not only the important information of the case description but also the semantic information of the relevant crime names.
[0039] Step 5 includes: concatenating the current hidden state in the decoding process with the case-crime name attention coding vector, and mapping it to the dimension of the dictionary through a fully connected network to obtain the predicted crime name.
[0040] Beneficial effects: The present invention completes the multi-crime name prediction task from the perspective of sequence generation. When encoding the case description, the case summary is incorporated to enhance the key information, and at the same time, all crime names are encoded to obtain the semantic information of each crime name and the mutual relationship between crime names. Finally, in the decoding process, a case-crime name attention mechanism is added to calculate the correlation between the hidden state and all crime names in the decoding process, significantly improving the accuracy of multi-crime name prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The following further specifically describes the present invention in conjunction with the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.
[0042] Figure 1 is the system architecture diagram of the present invention;
[0043] Figure 2 is the data flow diagram for calling the model for prediction;
[0044] Figure 3 is the multi-crime name prediction model structure diagram. DETAILED DESCRIPTION OF THE INVENTION
[0045] Embodiment
[0046] Such as Figure 1The system architecture diagram of the present invention is shown below. The core lies in the design and implementation of the multi - charge prediction system.
[0047] As Figure 2 shown is the data flow diagram of the called model prediction. The core lies in the pre - processing operations of the model input and output.
[0048] Figure 3 It is the structure diagram of the multi - charge prediction model. The core lies in the overall structure of the charge prediction model.
[0049] A multi - charge prediction method that fuses charge semantics and case - charge attention. By pre - processing the case description and training a prediction model based on a large amount of training data, it accurately predicts the corresponding charges of the case to assist court staff in making sentencing references. It includes:
[0050] First, select the Seq2seq model of UniLM to train the text automatic summarization model. The training dataset is the THUCNews news dataset. THUCNews is filtered and generated based on the historical data of the Sina News RSS subscription channel from 2005 to 2011, containing 740,000 news documents. It was originally a text classification dataset, but the first line of each news text is the summary corresponding to that news segment, similar to a big title in traditional news, which contains the key information of the news. Based on this dataset, the input and output of the text automatic summarization model are constructed and trained. The generated summary can not only contain the key information of the input but also have good readability.
[0051] Taking "In the workshop of a certain company in a certain town of a certain city, the defendant Hu had a quarrel with the victim Sun over trivial work matters, and then the defendant Hu injured the left abdomen of the victim Sun with a wooden cushion. After appraisal: The injury to Sun's left abdomen has reached the second degree of serious injury" as the case description in an embodiment, input this case description into the pre - trained text summarization model to obtain the case summary "The defendant Hu injured the left abdomen of the victim with a wooden cushion", and splice the case description and the case summary in the following form to obtain the input of the case encoding: [CLS] In the workshop of a certain company in a certain town of a certain city, the defendant Hu had a quarrel with the victim Sun over trivial work matters, and then the defendant Hu injured the left abdomen of the victim Sun with a wooden cushion. After appraisal: The injury to Sun's left abdomen has reached the second degree of serious injury @@@ The defendant Hu injured the left abdomen of the victim with a wooden cushion [SEP].
[0052] Since directly splicing all the charges involved in the 2018 Legal Research Cup exceeds the maximum input length of the UniLM model, the charges are streamlined, retaining the main information that differentiates each charge from other charges. After streamlining, the spliced charges are: [CLS] Charge [SEP].
[0053] Where [CLS] and [SEP] are the start and end markers of the text respectively, and @@@ is used as the separator between the case description and the case summary.
[0054] Then, the UniLM model is used to encode the above text respectively. The word-level tokenizer splits it into a sequence of words. For example, the case encoding part is tokenized as: [CLS, defendant, …, left abdomen, SEP]. Then, each word is converted into the corresponding value in the Chinese dictionary of Bert to obtain a sequence of values. Each item in the value sequence is converted into a 768-dimensional word vector through the word embedding method of the Bert model, and a sequence of case word vectors [fact1, fact2, …, fact m and a sequence of accusation word vectors [accusation1, accusation2, …, accusatio n are obtained, where m and n are the number of words after tokenization of the case encoding input and the accusation encoding input respectively.
[0055] The sequence of case word vectors and the sequence of accusation word vectors are respectively fed into the 24-layer Transformer network to fuse the context information at different layers. Each layer of the Transformer uses multi-head attention to fuse the vector output by the previous layer. The output of the l-th layer encoding is
[0056] For the l-th layer of the Transformer network, the calculation method of the output of the self-attention head in the Transformer network is:
[0057]
[0058]
[0059]
[0060] The output H l-1 of the previous layer is linearly projected into the query vector Q, the key vector K, and the value vector V respectively through the three parameter matrices of the l-th layer. d k is the dimension of the word vector, softmax is the normalized exponential function, and the mask matrix M is used to control whether the information at the corresponding position is visible to the context. If the value is 0, it means that all words can be accessed. If the value is -∞, it means that it is invisible to the context. The output of the 24th layer of the Transformer is the encoded output vector.
[0061] Take the last vector of the case encoding vector as the current hidden state during the decoding process, calculate the correlation coefficient between the current hidden state and each item in the crime name encoding. For the first calculation of attention, the output vector corresponding to the end marker [SEP] is used as the initial hidden state hidden of the decoder. Perform a dot product operation on the hidden vector and the output sequence [key1, key2, …, key n as follows:
[0062] similarity(hidden, key i ) = hidden · key i
[0063] where key i is the i-th item in the output sequence of the crime name encoding, n is the length of the crime name sequence, similarity is the correlation coefficient between the current hidden state hidden and all crime names, and normalize the obtained correlation coefficient using the softmax function:
[0064]
[0065] where similarity i represents the value of the i-th item in similarity, L x represents the total length of similarity, which is also the length of the output sequence of the crime name encoding, and a i represents the attention weight of each item in the output sequence of the crime name encoding. Calculate the case - crime name attention encoding vector Attention(hidden, keys) at the current time step:
[0066]
[0067] where keys represents the output sequence of the crime name encoding, and key i is the value at the i-th position in keys. Obtain the case - crime name attention encoding vector at the current time step, which contains not only the important information of the case description but also the semantic information of the relevant crime names. Concatenate the current hidden state during the decoding process with the case - crime name attention encoding vector, and map it to the dimension of the dictionary through a fully connected network to obtain the crime name predicted at the current time step. Append the predicted crime name to the back of the input of the case encoding, and repeat the above encoding and decoding processes until the end marker is predicted.
[0068] The overall process of the present invention is as Figure 2 shown: Encode the case description and all crime names respectively, perform an attention operation on the last output of the case encoding and the crime name encoding sequence, and integrate the attention encoding into the hidden state of the decoding process for crime name prediction.
[0069] A multi-charge prediction method that fuses charge semantics and case details - as shown below Figure 3 First, input the case description into a pre-trained text summarization model to obtain the case summary. Incorporating the case summary can increase the proportion of key information in the input text and enrich the important information related to charge prediction contained in the input. Secondly, placing the summary at the end of the input text significantly shortens the distance between the key information and the target charge, which helps to quickly establish a connection between the input and the output. Finally, due to the length limit of the input text, incorporating the summary can effectively alleviate the problem of information loss caused by truncation during data preprocessing due to overly long text. Then, encode the case description and the case summary in the form of "[CLS]Case description@@@Case summary[SEP]", where "@@@" is used to distinguish the case description from the case summary, and "[CLS]" and "[SEP]" are special tokens indicating the start and end of a sentence respectively. Take the output vector corresponding to the position of "[SEP]" as the initial hidden state of the decoder. Next, encode all charges using the bidirectional language model of the UniLM model in the same way, and take the outputs corresponding to all positions [h1, h2, …, h n . Calculate the correlation coefficient score between the decoder hidden state and the charge encoding sequence, normalize the correlation coefficient using the softmax function of the normalized exponential function, which represents the attention weight. Multiply the output of each charge encoder by its corresponding normalized weight and sum them up to obtain the charge attention encoding vector at the current time step. Then, concatenate the case encoding vector and the charge attention vector as the input to the decoder at the current time step. Map this input through a fully connected layer to the dimension of the Chinese dictionary, and take the charge corresponding to the position with the highest probability as the charge predicted at the current time step.
[0070] The present invention provides a multi-charge prediction method. There are many ways to implement this technical solution. The above is only the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented using existing technologies.
Claims
1. A multi-charge prediction method, characterized in that, It includes the following steps: Step 1: Train the text summarization model, input the case description into the trained text summarization model, and obtain the case summary; Step 1 includes: The text summarization model selects the UniLM model. When training the text summarization model, the input content includes: [CLS] news content [SEP] news summary [SEP], where [CLS] is the start token and [SEP] is the end token. Use the masking mechanism of the UniLM model to mask the words in the news summary, and let the UniLM model learn to recover the masked words one by one. The training objective is to maximize the likelihood of the masked words based on the context. The end token [SEP] can also be masked, and the model ends the prediction when it predicts the end token; Step 2: Concatenate the case description and the case summary; Step 3: Condense all charges into a string of length X1; Step 4: Take the case encoding vector obtained in Step 2 as the current hidden state in the decoding process, and calculate the case and charge attention encoding vectors at the current time step; Step 5: Concatenate the case encoding vector and the case and charge attention encoding vectors as the input at the current time step in the decoding process. Map the input through a fully connected layer to the dimension of the Chinese dictionary, and take the charge corresponding to the position with the highest probability as the charge predicted at the current time step; Step 2 includes: Add @@@ as a separator between the case description and the case summary, perform word segmentation and word embedding through a word-level tokenizer and the Chinese dictionary respectively, convert the input text into a vector sequence, and use the bidirectional language model of the UniLM model to encode the vector sequence. Take the last vector of the output as the case encoding vector at the current time step; In Step 2, during training, the case description and the corresponding charges are input in the form of sentence pairs. During testing, only the case description is input, and the format is: [CLS] case description @@@ case summary [SEP], where @@@ is used to distinguish the case description and the case summary. The representation of each word is composed of a combination of word embedding, position embedding, and segment embedding; In Step 2, first, the Bert word-level tokenizer performs word segmentation on the concatenated case description and case summary, and returns an array of the segmented words. Then, according to the one-to-one correspondence between the words and values in the Chinese dictionary, convert the array of words into an array of values, and use the nn.embedding method of the deep learning pytorch framework to convert the one-hot encoding of each word into a 768-dimensional dense vector; Position embedding encodes the position information of words into feature vectors, thereby introducing word position relationships; Step 2 also includes: For an input of length 512, the word vector dimension is 768, and the position embedding is a query table of (512, 768). The position embedding of each position in the sequence corresponds to the corresponding row in the table, and the values in it are continuously learned during the model training process; Step 2 further includes: The backbone network of the UniLM model consists of 24 layers of Transformer networks. After word embedding, the input vector of the UniLM model is transformed into a sequence H0 = [x1,..., x |x| composed of 768-dimensional word vectors, which is fed into the 24-layer Transformer network to fuse context information at different layers. Each layer of Transformer uses multi-head attention to fuse the vector output from the previous layer, and the encoding output of the l-th layer is indicating the output at the corresponding position of the word vector encoded in the l-th layer; For the l-th layer of the Transformer network, the self-attention head A in the Transformer network l is calculated as follows: Q = H l-1 W l Q , K = H l-1 W l K , V = H l-1 W l V The output H of the previous layer l-1 Passed through three parameter matrices W of the l-th layer l Q 、W l K 、W l V Are linearly projected into query vector Q, key vector K, and value vector V respectively. d k Is the dimension of the word vector. The mask matrix M is used to control whether the information at the corresponding position is visible to the context. M ij Represents the value at the i-th row and j-th column in the mask matrix M; if M ij The value is 0, then allow to attend means that all words can be accessed. If M ij The value is -∞, then prevent from attending means it is invisible to the context.
2. The method according to claim 1, wherein Step 3 includes: Perform word segmentation and word embedding through a word-level tokenizer and the Chinese dictionary respectively, convert the input text into a vector sequence, and use the bidirectional language model of the UniLM model to encode the vector sequence. Take the output vectors at all positions after encoding.
3. The method according to claim 2, wherein Step 4 includes: taking the case coding vector in Step 2 as the current hidden state during the decoding process, calculating the correlation coefficient between the current hidden state and each item in the crime name coding. For the first calculation of attention, the output vector corresponding to the end marker [SEP] is used as the initial hidden state hidden of the decoder. Perform a dot product operation on the hidden vector and the output sequence [key1, key2, …, key n similarity(hidden, key i ) = hidden · key i where key i is the i-th item in the output sequence of crime code numbers, n is the length of the crime sequence, similarity is the correlation coefficient between the current hidden state hidden and all crimes, and the obtained correlation coefficients are normalized by the normalized exponential function softmax: where similarity i represents the value of the i-th item in similarity, L x represents the total length of similarity, a i represents the attention weight of the i-th item in the output sequence of the charge code, and calculates the case and charge attention encoding vectors Attention(hidden, keys) at the current time step.
4. The method according to claim 3, wherein In step 4, the case and charge attention encoding vector Attention(hidden, keys) at the current time step is calculated using the following formula: where keys represents the output sequence of crime code encodings, and key i is the value at the i-th position in keys.
Citation Information
Patent Citations
Category prediction method and device based on theme information
CN110162787A
BERT pre-training model-based text abstract generation method
CN113128214A