Training method for Chinese sentence simplification model, Chinese sentence simplification method and device

By introducing the contrastive learning loss function into the encoder-decoder pre-training model and increasing the distance between positive and negative examples, the problems of low sentence simplification efficiency and insufficient fidelity of simplified content in the existing technology are solved, and efficient and controllable sentence simplification is achieved.

CN114757204BActive Publication Date: 2025-09-09BEIJING KNOWLEDGE ATLAS TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210459421.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-27
Publication Date
2025-09-09
Estimated Expiration
2042-04-27

AI Technical Summary

Technical Problem

Existing sentence simplification methods are inefficient when executed in batches, the length of simplified content is uncontrollable, and the fidelity is low.

Method used

A pre-training model based on an encoder-decoder structure is adopted, combined with a contrastive learning method. By introducing a contrastive learning loss function, the distance between positive and negative samples is increased, thereby controlling the length and fidelity of the generated simplified sentences.

Benefits of technology

The training efficiency of the sentence simplification model and the fidelity of generated simplified sentences are improved, and the length and content of generated simplified sentences can be controlled according to actual needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114757204B_ABST
    Figure CN114757204B_ABST
Patent Text Reader

Abstract

This application proposes a training method, a Chinese sentence simplification method, and a device for the Chinese sentence simplification model. The training method for the Chinese sentence simplification model includes: obtaining a dataset of complex sentence-simple sentence pairs containing supervisory signals and a Chinese monolingual pre-training model; selecting a simple sentence from the current complex sentence-simple sentence pair as a positive sample in each training batch, and randomly selecting a preset number of simple sentences from other sentence pairs in the same training batch as negative samples; projecting the complex sentence, positive sample, and negative sample into a vector representation space, and obtaining the hidden layer vectors of the last layer of the encoder respectively; calculating the contrastive learning loss, and calculating the cross-entropy loss of the desired simple sentence through the decoder; and jointly training the Chinese monolingual pre-training model by minimizing the contrastive learning loss and cross-entropy loss of the simple sentence output by the Chinese monolingual pre-training model. The simplified model obtained by this method can improve the controllability and fidelity of the generated simplified sentences.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of natural language processing technology, and in particular to a training method for a Chinese sentence simplification model, a Chinese sentence simplification method and a device. Background Art

[0002] Sentence simplification aims to modify the content and structure of the original sentence, reducing its complexity while retaining its main idea and keeping it close to its original semantics, making it easier to read and understand. Sentence simplification can also be used to improve the performance of other natural language processing (NLP) tasks, such as text summarization, information extraction, and machine translation.

[0003] Early sentence simplification methods used rule-based approaches, primarily focusing on lexical or syntactic simplification, resulting in low accuracy. Current research approaches fall into two main categories: one treating sentence simplification as a monolingual machine translation task; the other treating sentence simplification as a sentence editing task involving deletion, insertion, and retention.

[0004] However, this method of sentence simplification may require many steps when executing batch simplification tasks, resulting in low efficiency. Furthermore, in some specific application scenarios, the length of the simplified content can be uncontrollable, and the simplified content may be less faithful to the original sentence. Summary of the Invention

[0005] The present application aims to solve one of the technical problems in the related art at least to a certain extent.

[0006] To this end, the first purpose of this application is to propose a training method for a Chinese sentence simplification model, which introduces contrastive learning and fine-tunes the encoder-decoder pre-training model. By introducing contrastive learning loss on the basis of the cross-entropy loss function to increase the distance between positive and negative samples, the trained Chinese sentence simplification model can control the generated simplified sentences according to actual needs when simplifying sentences, thereby improving the fidelity of the generated simplified sentences and reducing the complexity of the training process. The trained Chinese sentence simplification model can process different Chinese sentences conveniently and efficiently, solving the problems of low training efficiency of the Chinese sentence simplification model and low efficiency of execution of the simplification task.

[0007] The second object of this application is to provide a method for simplifying Chinese sentences. This method inputs a Chinese sentence to be simplified into a Chinese sentence simplification model obtained by using a training method for a Chinese sentence simplification model, thereby improving the controllability and fidelity of the obtained simplified sentence and solving the problem of uncontrollable length and low fidelity of simplified content.

[0008] The third purpose of this application is to provide a training device for a Chinese sentence simplification model.

[0009] The fourth objective of this application is to provide a Chinese sentence simplification device.

[0010] A fifth object of the present application is to provide a non-transitory computer-readable storage medium.

[0011] To achieve the above objectives, the first embodiment of the present application proposes a method for training a Chinese sentence simplification model, comprising the following steps:

[0012] Obtain a preset dataset of complex sentence-simple sentence pairs containing supervisory signals as training data, and obtain a Chinese monolingual pre-training model based on an encoder-decoder structure;

[0013] Based on the contrastive learning method, in each training batch, the simple sentences in the current complex sentence-simple sentence pair are selected as positive samples, and a preset number of simple sentences in other sentence pairs in the same training batch are randomly selected as negative samples;

[0014] Projecting the complex sentence, the positive sample, and the negative sample in the current complex sentence-simple sentence pair into a vector representation space, and obtaining hidden layer vectors of the complex sentence, the positive sample, and the negative sample in the last layer of the encoder respectively;

[0015] Based on the hidden layer vector, calculating the contrastive learning loss, and calculating the cross entropy loss of generating the expected simple sentence through the decoder;

[0016] The Chinese monolingual pre-training model is jointly trained by minimizing the contrastive learning loss and the cross entropy loss of the simple sentences output by the Chinese monolingual pre-training model, so as to fine-tune the pre-training model to obtain a Chinese sentence simplification model.

[0017] Optionally, in one embodiment of the present application, the cross entropy loss of generating the desired simple sentence is calculated by the following formula:

[0018]

[0019] in,

[0020] in, represents the cross entropy loss, p θ (y (i) |x (i) ) represents the conditional probability of the output sequence, x (i) =x1 (i) ,…,x M(i) , x (i) Represents the i-th complex sentence of input length M, y (i) =y1 (i) ,…,y N (i) ,y (i) Represents the i-th simple sentence of length N generated.

[0021] Optionally, in one embodiment of the present application, the contrastive learning loss is calculated by the following formula:

[0022]

[0023] in, represents the contrastive learning loss, z x (i) Represents the vector representation of the i-th complex sentence, z y (i) represents the vector representation of the ith simple sentence, τ represents the temperature coefficient, S = {z y (j) : j≠i} is a set of randomly sampled negative simple sentences, (·,·) is the cosine similarity function.

[0024] Optionally, in one embodiment of the present application, obtaining a Chinese monolingual pre-training model based on an encoder-decoder structure includes: selecting punctuation marks, numbers, English letters and high-frequency Chinese words commonly used in Chinese sentences as a new vocabulary; replacing the original vocabulary of a preset multilingual pre-training model based on an encoder-decoder structure with the new vocabulary, and updating the representation parameters of the input vector and output vector of the multilingual pre-training model to update the multilingual pre-training model; saving the new vocabulary and the updated pre-training model to prune the multilingual pre-training model to the Chinese monolingual pre-training model.

[0025] To achieve the above-mentioned purpose, the second embodiment of the present application proposes a Chinese sentence simplification method, comprising the following steps:

[0026] Obtaining a Chinese sentence to be simplified, and inputting the Chinese sentence to be simplified into a Chinese sentence simplification model obtained by using the training method of the Chinese sentence simplification model described in the above embodiment to perform sentence simplification;

[0027] Obtain a predicted simplified sentence output by the Chinese sentence simplification model.

[0028] To achieve the above objectives, the third embodiment of the present application proposes a training device for a Chinese sentence simplification model, comprising the following modules:

[0029] The first acquisition module is used to obtain a preset dataset of complex sentence-simple sentence pairs containing supervision signals as training data, and obtain a Chinese monolingual pre-training model based on an encoder-decoder structure;

[0030] A selection module is used to select simple sentences from the current complex sentence-simple sentence pair in each training batch as positive samples based on contrastive learning, and randomly select a preset number of simple sentences from other sentence pairs in the same training batch as negative samples;

[0031] A second acquisition module is configured to project the complex sentence, the positive example, and the negative example in the current complex sentence-simple sentence pair into a vector representation space, and respectively obtain hidden layer vectors of the complex sentence, the positive example, and the negative example in the last layer of the encoder;

[0032] A loss calculation module, configured to calculate a contrastive learning loss based on the hidden layer vector and calculate a cross entropy loss for generating a desired simple sentence through a decoder;

[0033] A minimization calculation module is used to jointly train the Chinese monolingual pre-training model by minimizing the contrastive learning loss and the cross-entropy loss of the simple sentence output by the Chinese monolingual pre-training model, so as to fine-tune the pre-training model to obtain a Chinese sentence simplification model.

[0034] Optionally, in one embodiment of the present application, the loss calculation module is specifically configured to calculate the cross entropy loss of generating the desired simple sentence using the following formula:

[0035]

[0036] in,

[0037] in, represents the cross entropy loss, p θ (y (i) |x (i) ) represents the conditional probability of the output sequence, x (i) =x1 (i) ,…,x M (i) , x (i) Represents the i-th complex sentence of input length M, y (i) =y1 (i) ,…,y N (i) ,y (i) Represents the i-th simple sentence of length N generated.

[0038] Optionally, in one embodiment of the present application, the loss calculation module is further configured to calculate the contrastive learning loss using the following formula:

[0039]

[0040] in, represents the contrastive learning loss, z x (i) Represents the vector representation of the i-th complex sentence, z y (i) represents the vector representation of the ith simple sentence, τ represents the temperature coefficient, S = {z y (j) : j≠i} is a set of randomly sampled negative simple sentences, (·,·) is the cosine similarity function.

[0041] To achieve the above objectives, the fourth embodiment of the present application provides a Chinese sentence simplification device, comprising the following modules:

[0042] a simplification execution module, configured to obtain a Chinese sentence to be simplified, and input the Chinese sentence to be simplified into a Chinese sentence simplification model obtained by using the training method of the Chinese sentence simplification model described in the above embodiment to perform sentence simplification;

[0043] The simplified sentence acquisition module is used to obtain the predicted simplified sentences output by the Chinese sentence simplification model.

[0044] In order to implement the above embodiments, the fifth aspect of the present application also proposes a non-temporary computer-readable storage medium on which a computer program is stored. When the computer program is executed by the processor, it implements the training method of the Chinese sentence simplification model in the above embodiments, or implements the Chinese sentence simplification method in the above embodiments.

[0045] The technical solution provided by the embodiments of the present application brings at least the following beneficial effects: the present application regards the sentence simplification task as an end-to-end conditional generation task, introduces contrastive learning on the basis of the pre-trained model of the encoder-decoder for fine-tuning, and introduces contrastive learning loss on the basis of the cross-entropy loss function to increase the distance between positive and negative samples, so that the trained Chinese sentence simplification model can control the generated simplified sentences according to actual needs when performing sentence simplification, thereby improving the fidelity of the generated simplified sentences, and reducing the complexity of the training process. The trained Chinese sentence simplification model can process different Chinese sentences conveniently and efficiently. The present application not only makes the generation results controllable according to different needs but also improves the fidelity of the target simplified sentences.

[0046] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which

[0048] Figure 1 A flowchart of a method for training a Chinese sentence simplification model proposed in an embodiment of the present application;

[0049] Figure 2 A schematic diagram of a specific principle of introducing contrastive learning to simplify model training proposed in an embodiment of the present application;

[0050] Figure 3 A flowchart of a Chinese sentence simplification method proposed in an embodiment of the present application;

[0051] Figure 4 A schematic diagram of the structure of a training device for a Chinese sentence simplification model proposed in an embodiment of the present application;

[0052] Figure 5 This is a schematic diagram of the structure of a Chinese sentence simplification device proposed in an embodiment of the present application. DETAILED DESCRIPTION

[0053] The following describes embodiments of the present invention in detail, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and are not to be construed as limiting the present invention.

[0054] The following describes the Chinese sentence simplification model training method, Chinese sentence simplification method and device proposed in the embodiments of the present invention with reference to the accompanying drawings.

[0055] Figure 1 This is a flowchart of a method for training a Chinese sentence simplification model proposed in an embodiment of the present application, such as Figure 1 The method comprises the following steps:

[0056] Step S101: obtain a preset data set of complex sentence-simple sentence pairs containing supervisory signals as training data, and obtain a Chinese monolingual pre-training model based on an encoder-decoder structure.

[0057] Among them, a complex sentence-simple sentence pair is a sentence pair consisting of a longer complex sentence and a semantically similar but shorter simple sentence, for example, "He was fired by the company" - "The company fired him."

[0058] Among them, the supervision signal is the ratio information between the complex sentences and the simple sentences in the sentence pair. For example, the supervision signal can include information such as sentence length ratio, edit distance ratio, vocabulary complexity ratio and syntactic tree depth ratio.

[0059] In specific implementation, multiple semantically similar complex sentence-simple sentence pairs can be pre-mined, and then the supervision signal of each sentence pair can be calculated. Each supervision signal is then added to the corresponding complex sentence-simple sentence pair, thereby summarizing multiple complex sentence-simple sentence pairs containing supervision signals and generating a dataset of complex sentence-simple sentence pairs containing supervision signals. This dataset is used as training data for subsequent training of the Chinese sentence simplification model.

[0060] As one possible implementation approach, first, multiple semantically similar complex-simple sentence pairs are mined using an unsupervised learning approach. For example, large-scale, processed Chinese sentences are downloaded from resource libraries such as the CC-100 website. Language tool libraries such as the LASER and faiss libraries are then used to obtain vector representations for each sentence and create sentence indexes. Eight candidate sentences similar to each sentence are then mined using Euclidean distance. These candidate sentences are then further conditionally filtered to mine semantically similar complex-simple sentence pairs. Next, supervisory signals are calculated for each complex-simple sentence pair. For example, for the edit distance ratio in the example above, the Levenshtein distance between the complex and simple sentences can be calculated, as well as the Levenshtein distance ratios for edit operations such as deletion, insertion, and substitution, to obtain four different edit distance ratios. Finally, each supervisory signal is appended as a string to the start of the complex sentence in the corresponding sentence pair, generating a dataset of complex-simple sentence pairs with supervisory signals. For example, the supervision signal of a complex sentence-simple sentence pair is added to the starting position of the original complex sentence in the sentence pair as <supervision signal name_supervision signal ratio>, and a complex sentence-simple sentence dataset with supervision signals is formed on the basis of the original sentence pair. Each row of data is in the form of: <supervision signal name 1_supervision signal ratio 1>...<supervision signal name n_supervision signal ratio n> complex sentence, simple sentence, so that the generated dataset contains multiple rows of data in the above form.

[0061] Furthermore, a Chinese monolingual pre-trained model based on an encoder-decoder structure is obtained so that a sentence simplification model can be obtained by training the Chinese monolingual pre-trained model. In one embodiment of the present application, a relatively mature multilingual pre-trained model in the related art can be first obtained, and then the multilingual pre-trained model can be pruned to obtain a Chinese monolingual pre-trained model. This avoids rebuilding the Chinese language pre-trained model, avoids the complex steps of building a pre-trained model, reduces the computing resources required to obtain the pre-trained model, and is conducive to improving the training efficiency of the Chinese sentence simplification model.

[0062] As one possible implementation method, obtaining a Chinese monolingual pre-training model based on an encoder-decoder structure includes the following steps: first selecting punctuation marks, numbers, English letters and high-frequency Chinese words commonly used in Chinese sentences as a new vocabulary, then replacing the original vocabulary of the preset multilingual pre-training model based on the encoder-decoder structure with the new vocabulary, and updating the representation parameters of the input vector and output vector of the multilingual pre-training model to update the multilingual pre-training model, and finally saving the new vocabulary and the updated pre-training model to prune the multilingual pre-training model to a Chinese monolingual pre-training model.

[0063] In this example, the vocabulary is updated by deleting all words from the original vocabulary except for the new vocabulary. The model's input and output vector representations are then updated, replacing the parameters of the input and output vector representations in the multilingual pre-trained model to update the neural network.

[0064] It should be noted that the multilingual pre-training model that includes Chinese includes multiple other languages. For the training of the Chinese sentence simplification model of this application, it contains too much useless redundant information and consumes too many resources. The vector representation parameters of other languages ​​may account for more than half of the model parameters. This useless information will waste computing resources during training. Therefore, this application prunes the multilingual pre-training model into a single language pre-training model suitable for Chinese, which can further improve the efficiency of model training.

[0065] In step S102 , based on contrastive learning, simple sentences in the current complex sentence-simple sentence pair are selected in each training batch as positive samples, and a preset number of simple sentences are randomly selected from other sentence pairs in the same training batch as negative samples.

[0066] Contrastive learning is a type of self-supervised learning that automatically constructs positive and negative examples to learn a representation learning model. This model makes similar instances closer in projected space and dissimilar instances farther apart. This application, based on the idea of ​​contrastive learning, trains a Chinese sentence simplification model by adding contrasting positive and negative examples, aiming to learn vector representations of simple sentences.

[0067] The training batch size represents the number of parameters passed to the program for training at one time. For example, if there are 1000 training data points in a dataset of complex-simple sentence pairs containing supervisory signals, and the training batch size is set to 100, then the Chinese sentence simplification model will first be trained using the first 100 parameters in the dataset, i.e., data points 1-100. After this training is complete, the model parameters are updated and the parameter weights are adjusted. The next training batch, i.e., data points 101-200, will then be used for training, and the above training process will be repeated until all training data has been called.

[0068] Specifically, in each training batch, the present application selects the simple sentence from the current complex-simple sentence pair as a positive sample, and randomly selects a preset number of simple sentences from other sentence pairs in the same training batch as negative samples. The other sentence pairs refer to sentence pairs in the same training batch other than the current complex-simple sentence pair. The preset number is an arbitrary number determined based on actual needs, for example, it can be all other sentence pairs.

[0069] For example, assuming the batch size is 8 and the currently selected complex-simple sentence pair is the third sample, the complex sentence in the third sample is determined, and the simple sentence in the third sample is used as a positive example. The simple sentences in the remaining seven complex-simple sentence pairs in the training batch will be used as negative examples. In other examples, when the preset number is less than the total number of remaining sentence pairs, a corresponding number of sentence pairs can be randomly selected from the remaining sentence pairs as targets. In this way, negative examples are randomly selected from the same batch as non-target output sequences.

[0070] Step S103 : Project the complex sentence, positive examples, and negative examples in the current complex sentence-simple sentence pair into a vector representation space, and obtain the hidden layer vectors of the complex sentence, positive examples, and negative examples in the last layer of the encoder, respectively.

[0071] Specifically, the original complex sentence and the target simple sentence are projected into a potential vector representation space, where the target simple sentence refers to the simple sentence in the original excavated complex sentence-simple sentence pair, that is, the positive sample in the current sentence pair determined. Since the Chinese monolingual pre-training model of this application is based on the encoder-decoder structure, the complex sentence, positive sample, and negative sample can be input into the encoder to obtain the vector representation during specific implementation. Then, the hidden layer vectors of the last layer of the encoder are obtained for the complex sentence, positive sample, and negative sample respectively.

[0072] Step S104: Calculate the contrastive learning loss based on the hidden layer vector, and calculate the cross entropy loss of the desired simple sentence through the decoder.

[0073] Specifically, the contrastive learning loss is calculated based on the hidden layer vectors of the complex sentence, positive examples, and negative examples obtained in step S103, and then the cross entropy loss of the expected simple sentence generated by the Chinese sentence simplification model is calculated by the decoder, where the expected simple sentence can be the target simple sentence in the above embodiment or the correct simple sentence manually annotated, etc.

[0074] In one embodiment of the present application, the decoder may be controlled to calculate the cross entropy loss of the desired simple sentence using the following formula:

[0075]

[0076] in,

[0077] in, represents the cross entropy loss, p θ (y (i) |x (i) ) represents the conditional probability of the output sequence, x (i) =x1 (i) ,…,x M (i) , x (i) Represents the i-th complex sentence of input length M, y (i) =y1 (i) ,…,y N (i) ,y (i) Represents the i-th simple sentence of length N generated.

[0078] In one embodiment of the present application, the contrastive learning loss is calculated using the following formula:

[0079]

[0080] in, represents the contrastive learning loss, z x (i) Represents the vector representation of the i-th complex sentence, z y (i) represents the vector representation of the ith simple sentence, τ represents the temperature coefficient, S = {z y (j) : j≠i} is a set of randomly sampled negative simple sentences, (·,·) is the cosine similarity function.

[0081] Among them, because the contrastive loss function is a loss function with the property of self-discovery of difficult negative samples, by focusing on difficult samples, it is not necessary to further move away samples that are already far away, and the main focus is on how to move away samples that are not far away, so as to make the representation space more uniform. This application adjusts the degree of attention to difficult samples through the temperature coefficient τ. The temperature coefficient τ value is set according to actual needs to adjust the degree of attention to separate the current sample from the other most similar samples.

[0082] Step S105 , jointly training the Chinese monolingual pre-training model by minimizing the contrastive learning loss and cross entropy loss of the simple sentences output by the Chinese monolingual pre-training model, so as to fine-tune the pre-training model to obtain a Chinese sentence simplification model.

[0083] Specifically, during the training and fine-tuning process of the Chinese sentence simplification model, the Chinese sentence simplification model can be jointly trained and optimized by minimizing the cross-entropy loss and contrastive learning loss of the generated target sequence. Minimizing the generated target sequence refers to the simple sentence predicted by the model output by the Chinese sentence simplification model. This application minimizes the cross-entropy loss and contrastive learning loss of the generated target sequence to reduce the error between the simple sentence output by the Chinese sentence simplification model and the target simple sentence of the above embodiment.

[0084] In the embodiment of this application, minimize Among them, the calculation method of cross entropy loss and contrastive learning loss can refer to the formula in step S104. The present application minimizes the cross entropy loss and contrastive learning loss of the generated target sequence. For the complex sentence data input into the model, the Chinese sentence simplification model is trained to maximize the similarity between the original complex sentence and the target simple sentence, while minimizing the similarity between the complex sentence and the negative sample. It should be noted that the other steps of the present application for training the Chinese sentence simplification model can refer to the steps of model training through contrastive learning in the relevant technology, such as determining the training end conditions through stochastic gradient descent, etc. The implementation principles are similar and will not be repeated here.

[0085] Therefore, the model training method of the present application adds contrastive learning loss on the basis of the original cross-entropy loss, and then fine-tunes the pruned pre-trained model according to the training data in the complex sentence-simple sentence pair dataset with supervised signals. During the fine-tuning process, a contrastive learning loss function is added by automatically generating positive and negative samples, so as to train a model suitable for Chinese sentence simplification during the fine-tuning process.

[0086] To sum up, the training method of the Chinese sentence simplification model in the embodiment of the present application introduces contrastive learning and fine-tunes it on the basis of the pre-trained model of the encoder-decoder, and increases the distance between positive and negative samples by introducing contrastive learning loss on the basis of the cross-entropy loss function, so that the trained Chinese sentence simplification model can control the generated simplified sentences according to actual needs when simplifying sentences, thereby improving the fidelity of the generated simplified sentences and reducing the complexity of the training process. The trained Chinese sentence simplification model can process different Chinese sentences conveniently and efficiently.

[0087] Based on the above embodiments, in order to more clearly describe the specific implementation process of the training method of the Chinese sentence simplification model of the embodiment of the present application, a complex sentence "He was fired by the company." and a simple sentence "The company fired him." are used as examples to describe in a specific embodiment.

[0088] Figure 2This is a schematic diagram of a specific principle of introducing contrastive learning to simplify model training proposed in the embodiment of this application. Figure 2 As shown, in this embodiment, the supervisory signal ratio is first calculated based on the complex sentence and the simple sentence, and the supervisory signal is added to the starting end of the input complex sentence to obtain a data set including the sentence pairs of <supervisory signal name 1_supervisory signal ratio 1>…<supervisory signal name n_supervisory signal ratio n> "He was fired by the company.", "The company fired him." Then, the simple sentence referenced in the example, that is, the simple sentence in the original mined complex sentence-simple sentence pair (the simple sentence of the target in the above embodiment, the simple sentence "The company fired him." in this example) is used as a positive example, and several negative examples are randomly extracted from the current training batch, such as "The weather is so nice today!" and "He walks so fast", etc. Then, the pre-trained language models (that is, Figure 2 PLM) encoder (i.e. Figure 2 The hidden layer vector of the last layer of the Encoder in . Then calculate the contrastive learning loss Then pass through the decoder (i.e. Figure 2 Decoder in Calculate the cross entropy loss of generating the desired simple sentence Finally, by jointly training and optimizing the model by minimizing the contrastive learning loss and the generation loss, when the complex sentence "He was fired by the company" is input into the Chinese sentence simplification model, the predicted simple sentence output by the model has the greatest similarity with "The company fired him."

[0089] In order to implement the above embodiment, the present application also proposes a Chinese sentence simplification method. Figure 3 This is a flow chart of a Chinese sentence simplification method proposed in the embodiment of this application. Figure 3 As shown, the method includes the following steps:

[0090] Step S301: obtain a Chinese sentence to be simplified, and input the Chinese sentence to be simplified into a Chinese sentence simplification model obtained by adopting a training method of a Chinese sentence simplification model to perform sentence simplification.

[0091] Specifically, the Chinese sentences to be simplified refer to relatively complex Chinese sentences that need to be simplified in practical applications. It can be understood that the training method of the Chinese sentence simplification model in this step is the training method of the Chinese sentence simplification model described in the above embodiment.

[0092] Step S302: Obtain the predicted simplified sentence output by the Chinese sentence simplification model.

[0093] In this method, the training method of the Chinese sentence simplification model of the above-mentioned embodiment is used to train a Chinese sentence simplification model for sentence simplification. When performing sentence simplification, the generated simplified sentences can be controlled according to actual needs, thereby improving the fidelity of the generated simplified sentences. As a result, the fidelity of the simplified sentences output by the model obtained in the embodiment of the present application is higher. When running the model, the generated simplified sentences can be controlled according to different needs, thereby solving the problem that the length of the simplified content is uncontrollable and the fidelity is low when simplifying Chinese sentences.

[0094] In order to implement the above embodiment, the present application also proposes a training device for a Chinese sentence simplification model.

[0095] Figure 4 This is a structural diagram of a training device for a Chinese sentence simplification model proposed in an embodiment of the present application.

[0096] like Figure 4 As shown, the training device for the Chinese sentence simplification model includes a first acquisition module 100, a selection module 200, a second acquisition module 300, a loss calculation module 400 and a minimization calculation module 500.

[0097] Among them, the first acquisition module 100 is used to obtain a preset data set of complex sentence-simple sentence pairs containing supervision signals as training data, and obtain a Chinese monolingual pre-training model based on an encoder-decoder structure.

[0098] The selection module 200 is used to select the simple sentences in the current complex sentence-simple sentence pair in each training batch as positive samples based on contrastive learning, and randomly select a preset number of simple sentences in other sentence pairs in the same training batch as negative samples.

[0099] The second acquisition module 300 is used to project the complex sentence, positive examples and negative examples in the current complex sentence-simple sentence pair into the vector representation space, and obtain the hidden layer vectors of the complex sentence, positive examples and negative examples in the last layer of the encoder respectively.

[0100] The loss calculation module 400 is used to calculate the contrastive learning loss based on the hidden layer vector and calculate the cross entropy loss of generating the expected simple sentence through the decoder.

[0101] The minimization calculation module 500 is used to jointly train the Chinese monolingual pre-training model by minimizing the contrastive learning loss and cross entropy loss of the simple sentences output by the Chinese monolingual pre-training model, so as to fine-tune the pre-training model to obtain a Chinese sentence simplification model.

[0102] Optionally, in one embodiment of the present application, the loss calculation module 400 is specifically configured to calculate the cross entropy loss of generating the desired simple sentence using the following formula:

[0103]

[0104] in,

[0105] in, represents the cross entropy loss, p θ (y (i) |x (i) ) represents the conditional probability of the output sequence, x (i) =x1 (i) ,…,x M (i) , x (i) Represents the i-th complex sentence of input length M, y (i) =u1 (i) ,…,y N (i) ,y (i) Represents the i-th simple sentence of length N generated.

[0106] In one embodiment of the present application, the loss calculation module 400 is further configured to calculate the contrastive learning loss using the following formula:

[0107]

[0108] in, represents the contrastive learning loss, z x (i) Represents the vector representation of the i-th complex sentence, z y (i) represents the vector representation of the ith simple sentence, τ represents the temperature coefficient, S = {z y (j) : j≠i} is a set of randomly sampled negative simple sentences, and sim(·,·) is the cosine similarity function.

[0109] Optionally, in one embodiment of the present application, the first acquisition module 100 is specifically used to: select punctuation marks, numbers, English letters and high-frequency Chinese words commonly used in Chinese sentences as a new vocabulary; replace the original vocabulary of the preset multilingual pre-training model based on the encoder-decoder structure with the new vocabulary, and update the representation parameters of the input vector and output vector of the multilingual pre-training model to update the multilingual pre-training model; save the new vocabulary and the updated pre-training model to prune the multilingual pre-training model to a Chinese monolingual pre-training model.

[0110] It should be noted that the explanation of the training method of the Chinese sentence simplification model in the above embodiment is also applicable to the device of this embodiment, and the implementation principle is similar, so it will not be repeated here.

[0111] To sum up, the training device of the Chinese sentence simplification model in the embodiment of the present application introduces contrastive learning and fine-tunes the pre-trained model of the encoder-decoder, and increases the distance between positive and negative samples by introducing contrastive learning loss on the basis of the cross-entropy loss function, so that the trained Chinese sentence simplification model can control the generated simplified sentences according to actual needs when simplifying sentences, thereby improving the fidelity of the generated simplified sentences and reducing the complexity of the training process. The trained Chinese sentence simplification model can process different Chinese sentences conveniently and efficiently.

[0112] In order to implement the above embodiment, the present application also proposes a Chinese sentence simplification device. Figure 5 This is a schematic diagram of the structure of a Chinese sentence simplification device proposed in an embodiment of the present application.

[0113] like Figure 5 As shown, the Chinese sentence simplification device includes:

[0114] The simplification execution module 1000 is used to obtain a Chinese sentence to be simplified, and input the Chinese sentence to be simplified into a Chinese sentence simplification model obtained by using a training method of a Chinese sentence simplification model to perform sentence simplification.

[0115] The simplified sentence acquisition module 2000 is used to obtain the predicted simplified sentences output by the Chinese sentence simplification model.

[0116] It should be noted that the explanation of the Chinese sentence simplification method in the above embodiment is also applicable to the device of this embodiment, and the implementation principle is similar, which will not be repeated here.

[0117] In this device, the training method of the Chinese sentence simplification model of the above embodiment is used to train a Chinese sentence simplification model for sentence simplification. When simplifying sentences, the generated simplified sentences can be controlled according to actual needs, thereby improving the fidelity of the generated simplified sentences. As a result, the fidelity of the simplified sentences output by the model obtained by this device is higher. When running the model, the generated simplified sentences can be controlled according to different needs, thereby solving the problem of uncontrollable length and low fidelity of simplified content when simplifying Chinese sentences.

[0118] In order to implement the above-mentioned embodiments, the present invention also proposes a non-temporary computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, it implements the training method of the Chinese sentence simplification model described in the embodiment of the first aspect of the present application, or, when the computer program is executed by a processor, it implements the Chinese sentence simplification method described in the embodiment of the second aspect of the present application.

[0119] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, if schematic expressions of the above terms are used in multiple embodiments or examples, it does not mean that these embodiments or examples are the same. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples, unless they are contradictory.

[0120] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of such features. Throughout the description of this application, "plurality" means at least two, for example, two, three, etc., unless otherwise specifically defined.

[0121] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application belong.

[0122] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing it in another suitable manner if necessary, and then storing it in a computer memory.

[0123] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0124] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0125] In addition, the functional units in the various embodiments of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0126] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present application. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A training method for a Chinese sentence simplification model, characterized in that: The following steps are involved: A preset dataset of complex sentence-simple sentence pairs containing supervisory signals is obtained as training data, and a Chinese monolingual pre-training model based on an encoder-decoder structure is obtained; the complex sentence-simple sentence pairs are composed of a complex sentence and a simple sentence with similar semantics but shorter length, and the supervisory signals are ratio information between the complex sentence and the simple sentence in the sentence pair, and the supervisory signals include sentence length ratio, edit distance ratio, vocabulary complexity ratio, and syntactic tree depth ratio; Based on the contrastive learning method, in each training batch, the simple sentences in the current complex sentence-simple sentence pair are selected as positive samples, and a preset number of simple sentences in other sentence pairs in the same training batch are randomly selected as negative samples; Projecting the complex sentence, the positive sample, and the negative sample in the current complex sentence-simple sentence pair into a vector representation space, and obtaining hidden layer vectors of the complex sentence, the positive sample, and the negative sample in the last layer of the encoder respectively; Based on the hidden layer vector, calculating the contrastive learning loss, and calculating the cross entropy loss of generating the expected simple sentence through the decoder; Jointly training the Chinese monolingual pre-training model by minimizing the contrastive learning loss and the cross-entropy loss of the simple sentences output by the Chinese monolingual pre-training model, so as to fine-tune the pre-training model to obtain a Chinese sentence simplification model; The method of obtaining a Chinese monolingual pre-trained model based on an encoder-decoder structure includes: Select punctuation marks, numbers, English letters and high-frequency Chinese words commonly used in Chinese sentences as new vocabulary; Replacing an original vocabulary of a preset encoder-decoder structure-based multilingual pre-trained model with the new vocabulary, and updating representation parameters of input vectors and output vectors of the multilingual pre-trained model to update the multilingual pre-trained model; The new vocabulary and the updated pre-trained model are saved to prune the multilingual pre-trained model into the Chinese monolingual pre-trained model.

2. The training method according to claim 1, characterized in that The cross entropy loss of the desired simple sentence is calculated by the following formula: in, in, represents the cross entropy loss, p θ (y (i) |x (i) ) represents the conditional probability of the output sequence, x (i) =x1 (i) ,…,x M (i) , x (i) Represents the i-th complex sentence of input length M, y (i) =y1 (i) ,…,y N (i) ,y (i) Represents the i-th simple sentence of length N generated.

3. The training method according to claim 2, characterized in that The contrastive learning loss is calculated by the following formula: in, represents the contrastive learning loss, z x (i) Represents the vector representation of the i-th complex sentence, z y (i) represents the vector representation of the ith simple sentence, η represents the temperature coefficient, S = {z y (j) : j≠i} is a set of randomly sampled negative simple sentences, and sim(·,·) is the cosine similarity function.

4. A method for simplifying Chinese sentences, characterized in that: The following steps are involved: Obtaining a Chinese sentence to be simplified, and inputting the Chinese sentence to be simplified into a Chinese sentence simplification model obtained by using the training method of the Chinese sentence simplification model according to any one of claims 1 to 3 for sentence simplification; Obtain a predicted simplified sentence output by the Chinese sentence simplification model.

5. A training device for a Chinese sentence simplification model, characterized in that: include: A first acquisition module is configured to acquire a preset dataset of complex sentence-simple sentence pairs containing supervisory signals as training data, and to acquire a Chinese monolingual pre-training model based on an encoder-decoder structure; the complex sentence-simple sentence pairs are composed of a complex sentence and a semantically similar but shorter simple sentence, and the supervisory signals are ratio information between the complex sentence and the simple sentence in the sentence pair, including sentence length ratio, edit distance ratio, lexical complexity ratio, and syntactic tree depth ratio; A selection module is used to select simple sentences from the current complex sentence-simple sentence pair in each training batch as positive samples based on contrastive learning, and randomly select a preset number of simple sentences from other sentence pairs in the same training batch as negative samples; A second acquisition module is configured to project the complex sentence, the positive example, and the negative example in the current complex sentence-simple sentence pair into a vector representation space, and respectively obtain hidden layer vectors of the complex sentence, the positive example, and the negative example in the last layer of the encoder; A loss calculation module, configured to calculate a contrastive learning loss based on the hidden layer vector and calculate a cross entropy loss for generating a desired simple sentence through a decoder; a minimization calculation module, configured to jointly train the Chinese monolingual pre-training model by minimizing the contrastive learning loss and the cross-entropy loss of the simple sentences output by the Chinese monolingual pre-training model, so as to fine-tune the pre-training model to obtain a Chinese sentence simplification model; The obtaining of a Chinese monolingual pre-training model based on an encoder-decoder structure includes: selecting punctuation marks, numbers, English letters, and high-frequency Chinese words commonly used in Chinese sentences as a new vocabulary; The original vocabulary of a preset encoder-decoder structure-based multilingual pre-trained model is replaced with the new vocabulary, and the representation parameters of the input vector and output vector of the multilingual pre-trained model are updated to update the multilingual pre-trained model; the new vocabulary and the updated pre-trained model are saved to prune the multilingual pre-trained model to the Chinese monolingual pre-trained model.

6. The device according to claim 5, characterized in that The loss calculation module is specifically used to calculate the cross entropy loss of generating the desired simple sentence using the following formula: in, in, represents the cross entropy loss, p θ (y (i) |x (i) ) represents the conditional probability of the output sequence, x (i) =x1 (i) ,…,x M (i) , x (i) Represents the i-th complex sentence of input length M, y (i) =y1 (i) ,…,y N (i) ,y (i) Represents the i-th simple sentence of length N generated.

7. The device according to claim 5, characterized in that The loss calculation module is further configured to calculate the contrastive learning loss using the following formula: in, represents the contrastive learning loss, z x (i) Represents the vector representation of the i-th complex sentence, z y (i) represents the vector representation of the ith simple sentence, τ represents the temperature coefficient, S = {z y (j) : j≠i} is a set of randomly sampled negative simple sentences, and sim(·,·) is the cosine similarity function.

8. A Chinese sentence simplification device, characterized in that: The following steps are involved: a simplification execution module, configured to obtain a Chinese sentence to be simplified, and input the Chinese sentence to be simplified into a Chinese sentence simplification model obtained by using the training method of the Chinese sentence simplification model according to any one of claims 1 to 3 for sentence simplification; The simplified sentence acquisition module is used to obtain the predicted simplified sentences output by the Chinese sentence simplification model.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the training method of the Chinese sentence simplification model as described in any one of claims 1 to 3, or when the computer program is executed by a processor, it implements the Chinese sentence simplification method as described in claim 4.

Citation Information

Patent Citations

  • Text statement processing method and device, computer equipment and storage medium

    CN111950269A