A text generation model and a text generation method

By introducing a hierarchical decoder structure, the problems of semantic incoherence and lack of diversity in long text generation are solved, and the generation of long texts with semantic coherence and diversity is realized.

CN114462419BActive Publication Date: 2025-11-07CHEZHI HULIAN BEIJING SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210032379.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-12
Publication Date
2025-11-07
Estimated Expiration
2042-01-12

AI Technical Summary

Technical Problem

Existing network structures cannot dynamically model the complex semantic structure of long texts when generating paragraph-style texts, resulting in semantic incoherence and insufficient diversity in the generated texts, especially in advertising texts where repetition is common.

Method used

A hierarchical decoder structure is introduced. Through clause content planning units and word generation units, the content of the clause is first determined, and then words are generated. Combined with bundle search constraints, semantic information of different granularities is captured.

Benefits of technology

It generates semantically coherent and diverse long texts, avoids repetitive sentences, and improves the quality of long text generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114462419B_ABST
    Figure CN114462419B_ABST
Patent Text Reader

Abstract

The present disclosure discloses a text generation model and a text generation method. The text generation model comprises an encoding module and a decoding module. The encoding module is adapted to process input data to extract a first vector indicating semantic features thereof; the decoding module is adapted to process the first vector to generate at least one sentence vector to compose a long text. Further, the decoding module further comprises a sub-sentence content planning unit coupled with the encoding module, adapted to receive an output of the encoding module, process the first vector output by the encoding module to determine at least one second vector indicating semantic features of a sub-sentence; a word generation unit coupled with the sub-sentence content planning unit, adapted to process the second vector to generate a plurality of word vectors corresponding to words, and combine the word vectors into at least one sentence vector to generate the long text.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of computer networks, and particularly relates to a text generation scheme. BACKGROUND

[0002] Natural language generation (NLG) is an important part of natural language processing (NLP) field, and its main purpose is to reduce the communication gap between human and machine, and convert non-language format data into language format that can be understood by human.

[0003] At present, the academic circle usually uses an encoder-decoder network structure to generate text. The encoder is used to analyze the input data (usually a sequence), and compress the input sequence into a vector of a specified length, which is used to indicate the semantics of the sequence. The decoder is responsible for decoding the semantic vector, attribute feature representation, etc. to generate natural language.

[0004] The existing network structure has good effect on generating coherent short text, but for long text in paragraph form, it cannot dynamically model the input data during generation and capture the complex semantic structure of long text well, resulting in the inability to obtain long-distance information, and thus causing poor quality of generated text.

[0005] In addition, for long text of the advertisement type, the diversity of the generated text is also very important. Generally speaking, a traditional model can only give one or a few fixed outputs corresponding to a set of inputs, but for application scenarios such as advertisements and reviews, repeated text frequently appearing is unacceptable.

[0006] Therefore, a new text generation scheme is needed to solve the problem of incoherent semantic expression and lack of diversity when generating long text in paragraph form. SUMMARY

[0007] The present disclosure provides a text generation model and a text generation method to try to solve or at least alleviate at least one of the above problems.

[0008] According to one aspect of the present disclosure, a decoding module is provided, which is adapted to be arranged in a text generation model and coupled with an encoding module, and includes: a clause content planning unit coupled with the encoding module, adapted to receive the output of the encoding module, and process the first vector output by the encoding module to determine at least one second vector indicating the semantic features of the clause; and a word generation unit coupled with the clause content planning unit, adapted to process the second vector to generate a word vector corresponding to a plurality of words, and combine the word vectors into at least one sentence vector to generate long text.

[0009] Optionally, in the decoding module according to the present disclosure, the word generation unit is further adapted to, at each time step, perform a beam search to obtain probability values of the sentence vectors; select the first number of sentence vectors in sequence as a candidate sequence according to the order of the probability values from large to small; calculate the difference values between each two sentence vectors in the candidate sequence respectively; reconstruct the candidate sequence based on the difference values; repeat the steps of calculating the difference values and reconstructing the candidate sequence until the difference values meet the preset condition, and determine the sentence vectors belonging to the candidate sequence.

[0010] Optionally, in the decoding module according to the present disclosure, the word generation unit is further adapted to, when the difference values of the two sentence vectors are greater than the threshold value, remove one of the two sentence vectors from the candidate sequence; and sequentially add a sentence vector with the largest probability value to the candidate sequence to reconstruct the candidate sequence.

[0011] Optionally, in the decoding module according to the present disclosure, the sub-sentence content planning unit is further adapted to determine a probability distribution of the sub-sentence semantic feature of the current time step based on the first vector output by the encoding module and the sub-sentence semantic feature output at the last time step; and generate the second vector indicating the sub-sentence semantic feature of the current time step based on the probability distribution.

[0012] Optionally, in the decoding module according to the present disclosure, the sub-sentence semantic feature comprises at least one of the following features: entity attribute of the sub-sentence, theme feature, and sentiment feature.

[0013] Optionally, in the decoding module according to the present disclosure, the sub-sentence content planning unit adopts a recurrent neural network or a long short-term memory network.

[0014] According to another aspect of the present disclosure, a text generation model is provided, comprising: an encoding module adapted to process input data to extract a first vector indicating semantic features thereof; and a decoding module as described above, coupled with the encoding module, adapted to process the first vector to generate at least one sentence vector to compose a long text.

[0015] According to still another aspect of the present disclosure, a text generation method is provided, comprising the steps of: extracting a first vector indicating semantic features from input data; determining at least one second vector indicating sub-sentence semantic features based on the first vector; processing the second vector to generate a plurality of word vectors; and combining the word vectors to generate at least one sentence vector to generate a long text.

[0016] According to still another aspect of the present disclosure, a computing device is provided, comprising: one or more processors and memories; and one or more programs, wherein the one or more programs are stored in the memories and configured to be executed by the one or more processors, and the one or more programs comprise instructions for executing any of the methods described above.

[0017] According to yet another aspect of the present disclosure, a computer readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by a computing device, cause the computing device to perform any of the methods as described above.

[0018] In summary, according to the decoding module of the present disclosure, by introducing the clause content planning unit, the decoder structure is divided into levels, not the entire sentence is generated at one time, but the clause expression content is determined first, and then the word generation is performed. When the word generation unit performs the word generation, by setting the restriction condition of the beam search, the diversified text generation is realized. The whole scheme fully utilizes the representation of the sentence level and the word level, fully captures the semantic information of different granularities of long text, to generate long text with semantic coherence and no repetition. BRIEF DESCRIPTION OF DRAWINGS

[0019] To the accomplishment of the foregoing and related ends, certain illustrative aspects are described herein in connection with the following description and the annexed drawings. These aspects are indicative of various ways in which the principles disclosed herein can be practiced and all aspects and equivalents thereof are intended to be within the scope of the claimed subject matter. The foregoing and other objects, features, and advantages of the disclosure will be apparent from the following description of one or more aspects and as illustrated in the accompanying drawings. The same reference numbers in different drawings identify the same components or elements.

[0020] Figure 1 A schematic diagram of a computing device 100 according to some embodiments of the present disclosure is shown;

[0021] Figure 2 A schematic diagram of a text generation model 200 according to some embodiments of the present disclosure is shown;

[0022] Figure 3 A structural schematic diagram of a decoding module 220 according to some embodiments of the present disclosure is shown; and

[0023] Figure 4 A flowchart of a text generation method 400 according to some embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0024] Exemplary embodiments of the present disclosure will be described hereinafter with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be embodied in various forms without being limited by the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure can be more thoroughly and completely understood, and so that the scope of the present disclosure can be conveyed to those skilled in the art.

[0025] As described above, in the encoder-decoder network structure for generating descriptive long text, the natural language characteristics of long text can cause the generated text to have incoherent semantic expression and lack of diversity. Generally, for this kind of scene, the quantity and quality of the training data set can be improved as much as possible to ensure the diversity of the data in the training process. However, in practice, simply relying on diversified training data to improve the generation of long text does not have an ideal effect.

[0026] The disclosure designs a hierarchical decoder structure as a generation network by modifying the neural network of the decoder stage, and fully captures the semantic information between long text sentences.

[0027] The text automatic generation scheme provided by the embodiment of the disclosure is applied to the scene of generating descriptive long text such as advertisements and comments, which can make the output text coherent and ensure the diversity of the text to avoid the occurrence of a large number of repeated sentences in the long text.

[0028] The text automatic generation scheme of the embodiment of the disclosure can be executed in one or more computing devices. Figure 1 is a block diagram of an example computing device 100.

[0029] In the basic configuration 102, the computing device 100 typically includes a system memory 106 and one or more processors 104. A memory bus 108 can be used for communicating between the processor 104 and the system memory 106.

[0030] Depending on the desired configuration, the processor 104 can be any type of processing unit including, but not limited to, a microprocessor (μP), a microcontroller (μC), a digital signal processor (DSP), or any combination thereof. The processor 104 can include one or more levels of cache memory 110 and 112, a processor core 114, and registers 116. An example processor core 114 can include an arithmetic logic unit (ALU), a floating point unit (FPU), a digital signal processing core (DSP Core), or any combination thereof. An example memory controller 118 can be used with the processor 104 or, in some implementations, the memory controller 118 can be an internal part of the processor 104.

[0031] Depending on the desired configuration, the system memory 106 can be of any type including but not limited to volatile memory (such as RAM), non-volatile memory (such as ROM, flash memory, etc.) or any combination thereof. System memory 106 can include an operating system 120, one or more applications 122, and program data 124. In some embodiments, application 122 can be arranged to operate with the operating system 120 on the program data 124.

[0032] The computing device 100 can also include a bus 1 10 or other communication mechanism for communicating information, and a processor 1 12 coupled to the bus 1 10 for processing information. The computing device 100 also includes a main memory 1 14, such as random access memory (RAM) or other dynamic storage device, coupled to the bus 1 10 for storing information and instructions to be executed by the processor 1 12. Main memory 1 14 can also be used for storing temporary variables or other intermediate information during execution of instructions to be executed by the processor 1 12. The computing device 100 further includes a read only memory (ROM) 1 16 or other static storage device coupled to the bus 1 10 for storing static information and instructions for the processor 1 12. A storage device 1 18, such as a magnetic disk or optical disk, is provided and coupled to the bus 1 10 for storing information and instructions.

[0033] The network communication link can be one example of a communication media. Communication media can typically be embodied by computer readable instructions, data structures, program modules, etc., in a modulated data signal, such as a carrier wave or other transport mechanism, and can include any information delivery media. A "modulated data signal" can be a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, communication media can include wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, radio frequency (RF), microwave, infrared (IR) and other wireless media. The term computer readable media as used herein can include both storage media and communication media. In some embodiments, computer readable media stores one or more programs, which are executable by a computer.

[0034] The computing device 100 can be implemented as part of a small-sized portable (or mobile) electronic device, such as a cellular phone, personal digital assistant (PDA), personal media player device, wireless network browsing device, personal headset, application-specific device, or a hybrid device that may include any of the above functions. The computing device 100 can also be implemented as a personal computer, including desktop and laptop computer configurations. The computing device 100 can also be implemented as a server with the above configurations.

[0035] In an embodiment according to the present disclosure, computing device 100 is configured to execute a text generation method, wherein application 122 of computing device 100 includes multiple program instructions for executing text generation method 400 according to the present disclosure, and program data 124 may also store relevant data of text generation model 200 for executing method 400, including but not limited to training data, hyperparameter information, etc.

[0036] Figure 2 A schematic diagram of a text generation model 200 according to some embodiments of the present disclosure is shown. The text generation model 200 is used to process an input sequence to generate text, particularly long paragraph-style text.

[0037] The text generation model 200 includes a coupled encoding module 210 and a decoding module 220. According to some embodiments of this disclosure, both the encoding module 210 and the decoding module 220 employ recurrent neural networks (RNNs).

[0038] like Figure 2 As shown, the encoding module 210 receives input data and processes it to extract semantic features. The input data can be text, numbers, images, etc., and is usually input to the encoding module 210 in the form of a sequence.

[0039] In one embodiment, the encoding module 210 transforms an input sequence of variable length into a semantic variable of fixed length, denoted as a first vector. This first vector indicates the semantic features of the input data.

[0040] Suppose the input sequence is x1,...,x T , where x i It is the i-th word in the input sequence. At time step t, the RNN will input x. t eigenvector x t The hidden state h of the previous time step t-1 Transform into the hidden state h of the current time step t Express the transformation of the hidden layer of an RNN using a function f:

[0041] ht =f(xt,h) t-1 )

[0042] Next, the encoding module 210 transforms the hidden states at each time step into the first vector using a custom function q:

[0043] z = q(h1,...,h) T )

[0044] For example, when choosing q(h1,...,h) T ) = h T At that time, the first vector is the hidden state h of the input sequence at the final time step. T .

[0045] The encoding module 210 described above is a unidirectional RNN, where the hidden state at each time step depends only on the input subsequence before and after that time step. According to embodiments of this disclosure, the encoding module 210 can also be constructed using a bidirectional RNN. In this case, the hidden state of the encoding module 210 at each time step depends simultaneously on the subsequence before and after that time step (including the input at the current time step), and encodes information about the entire sequence. Further details are omitted here.

[0046] The decoding module 220 receives the first vector, processes the first vector, and generates at least one sentence vector to form a long text.

[0047] In existing text generation models, the common decoder structure generates words directly from the latent variable information (i.e., the first vector) obtained from the encoder, generating one word at each time step, and concatenating the words to form a sentence.

[0048] According to the decoding module 220 of this disclosure, the decoder structure is divided into layers. Instead of generating the entire sentence at once, the content of the clause is determined first, and then words are generated. This fully utilizes the representation at the sentence level and word level to capture the semantic information of different granularities of the long text, so as to generate a semantically coherent and non-repetitive long text.

[0049] like Figure 2 As shown, the decoding module 220 further includes a clause content planning unit 222 and a word generation unit 224. Figure 3 A schematic diagram of the structure of a decoding module 220 according to some embodiments of the present disclosure is shown.

[0050] like Figure 3 Clause content planning unit 222 receives the first vector output by encoding module 210 (e.g., ... Figure 3 As shown in z), the first vector is processed to determine at least one second vector indicating the semantic features of the clause, where the semantic features of the clause are... Figure 3Mt+n is denoted.

[0051] According to the embodiments of the present disclosure, one clause can contain multiple clause semantic features. In an embodiment, the clause semantic features represent the attributes and feature values to be expressed by each clause, which at least include one of the following features: entity attributes of the clause, theme features, sentiment attitude features (such as praise or criticism sentiment, etc.), without limitation. The clause semantic features are used to guide the text generation task of the word generation unit 224, limit the range of the text generation task, and thus divide the generation task of long text into subtasks of clause generation.

[0052] In an embodiment, the clause semantic features Mt+n are expressed in the mathematical form of a binary tuple, to adapt to multiple semantic features, and the expression form is as follows, T i denotes all semantic features of the i-th clause:

[0053] T i ={Mt-n,...,Mt,Mt+n}

[0054] T i ={<x t-n ,y t-n >,……<x t ,y t >,……<x t+n ,y t+n >},

[0055] wherein <x t , y t > is a feature vector expressed in mathematics here:

[0056] x t ={M c},c∈{1,2,...c hidden}

[0057] y t ={M c},c∈{1,2,...c hidden}

[0058] In combination Figure 3 , in the clause content planning unit 222, St-1, St, St+1 respectively represent the states of a recurrent neural network at t-1, t and t+1 time. St represents the state of the recurrent neural network at t time, and its input includes the first vector z of the output of the encoding module 210 and the hidden layer output w of St-1. The output w represents the processing of a clause, that is, T i In other words, the clause content planning unit 222 decodes to obtain the clause semantic features at the current time step based on the first vector output by the encoding module 210 and the clause semantic features output at the last time step.

[0059] According to an embodiment of the present disclosure, the clause content planning unit 222 adopts a recurrent neural network (RNN) or a long short-term memory network (LSTM).

[0060] As described above, the second vector T i output by the clause content planning unit 222 is a c hidden dimensional vector, and T i specifically contains several semantic features, which are determined according to a hyperparameter threshold p in the RNN or LSTM network. In an embodiment, the clause content planning unit 222 can use a self-defined output layer and a Softmax operation to determine a probability distribution of the semantic features of the clause at the current time step based on the first vector output by the encoding module 210 and the semantic features of the clause output at the previous time step. Then, based on the probability distribution, the second vector indicating the semantic features of the clause at the current time step is generated. Specifically, the second vector contains c hidden dimensional semantic features as candidate semantic features. The probability value of each candidate semantic feature belonging to the clause at the current time step is determined through the probability distribution, and only when the probability value of the candidate semantic feature belonging to the clause is greater than the threshold p, the semantic feature is output (for example, the number corresponding to the semantic feature in the second vector is output as 1); otherwise, the semantic feature is not output (for example, the number corresponding to the semantic feature in the second vector is output as 0).

[0061] According to an embodiment of the present disclosure, the clause content planning unit 222 can be inserted into different decoder network structures as a general unit without affecting the original network structure, determine the clause generation task, and thus divide the decoder into levels to capture the semantic features of the clause so as to generate long text with semantic coherence.

[0062] Continuing as Figure 3 In addition to the clause content planning unit 222, the decoding module 220 further includes a word generation unit 224. The word generation unit 224 can adopt a general decoder network structure to capture fine-grained semantic features by selecting specific words. In an embodiment, the word generation unit 224 first processes the second vector to generate word vectors corresponding to a plurality of words. Then, the word vectors are combined into at least one sentence vector to generate long text.

[0063] According to an embodiment of the present disclosure, the overall design idea of the word generation unit 224 is to extend the beam search to increase the diversity of the text, and the core idea is to add a diversity constraint condition in the beam search process of the decoder.

[0064] As Figure 3In the word generation unit 224, Ht-p, Ht-2, Ht-1, Ht, Ht+1,... represent the state of the recurrent neural network at t-p, t-2, t-1, t, t+1,... time, respectively. Take Ht as an example, which represents the state of the recurrent neural network at t time, and its input includes the second vector of the corresponding time step of the output of the clause content planning unit 222 and the hidden layer output w of Ht-1; its output Yt and the output Yt+1 at Ht+1 time are combined to obtain a sentence vector corresponding to the clause "However". For more specific description of the decoding module, reference can be made to the description of the encoding module 210 and the existing encoder structure, which will not be described here again. The present disclosure introduces a restriction condition of beam search on the basis of the existing decoding module, which will be emphatically described below.

[0065] The word generation unit 224 first decodes the second vector to obtain a word vector indicating a word. The present disclosure does not make too many restrictions on this, and a general decoder processing manner can be adopted.

[0066] In an embodiment, the word generation unit 224 combines and generates a sentence vector from the word vector in the following manner.

[0067] Firstly, beam search is performed at each time step to obtain the probability value of each sentence vector. According to the order from large to small of the probability value, the first quantity of sentence vectors are sequentially selected as the candidate sequence. Assuming that the sentence vectors are recorded in the order from large to small of the probability value as 1, 2, 3,..., K, K+1,.... The first quantity is K, and then the first K largest probability value sentence vectors (i.e., sequence) are sequentially retained to form the candidate sequence.

[0068] Secondly, the difference value between each two sentence vectors in the candidate sequence is calculated. In an embodiment, the sentence vectors in the candidate sequence are grouped two by two, and the difference value between each two sentence vectors is calculated. The difference value can be defined as the Euclidean distance value of the two sentence vectors, but is not limited thereto.

[0069] Assuming that two sentence vectors are represented as a(x 11 , x 12 ... x 1n ) and b(x 21 , x 22 ... x 2n ), the difference value S ab between the two sentence vectors is determined by the following formula:

[0070]

[0071] In a third step, based on the difference value, the candidate sequence is reconstructed. In an embodiment, a hyperparameter is set: a threshold t for the difference score. When the difference value of two sentence vectors is greater than the threshold t, one of the two sentence vectors is removed from the candidate sequence. Optionally, one of the two sentence vectors is randomly selected and removed. Meanwhile, the sentence vector with the largest probability value (i.e., the K+1th sentence vector) is sequentially added to the candidate sequence to reconstruct the candidate sequence.

[0072] In a fourth step, based on the reconstructed candidate sequence, the steps of iteratively calculating the difference value (i.e., the second step) and reconstructing the candidate sequence (i.e., the third step) are repeated until the calculated difference value satisfies a preset condition, and the sentence vector belonging to the candidate sequence is determined. As described above, the preset condition is that the difference value of two sentence vectors is not greater than the threshold t.

[0073] Finally, if the difference value of two sentence vectors is not greater than the threshold t, the sequence is continuously extended at each time step until the decoding stage is completed.

[0074] According to the word generation unit 224 of the present disclosure, in the process of beam search, the diversity restriction condition is increased by calculating the difference value between sentences, so that the diversity of the generated text is increased while ensuring the difference between the generated texts.

[0075] It should be noted that, whether it is the sub-sentence content planning unit 222, the word generation unit 224, or the decoding module 220 composed of the two, it can be inserted into a general decoder network structure as a general network module or directly used as a decoder network without affecting the structure of the original neural network. The neuron parameters of this network module are introduced into the parameter set and loss function of the neural network, and the entire neural network is continuously learned through gradient backpropagation until the network parameters converge.

[0076] The text generation model 200 of the present disclosure mainly solves the problems of incoherent semantics and insufficient diversity in long text generation. The decoder of the common Encoder-Decoder structure in text generation is divided into levels, and the sub-sentence content planning unit 222 is introduced to capture the semantic information at the sentence level, i.e., the semantic features at a higher level: the theme of the sentence, the expressed praise or criticism emotion, etc. In addition, the word generation unit 224 is designed based on the beam search to generate long texts with diverse expression forms.

[0077] Figure 4 A flowchart of a text generation method 400 according to some embodiments of the present disclosure is shown. The text generation method 400 is performed based on the text generation model 200. It should be noted that the execution flow of the method 400 is complementary to the description of the text generation model 200 and the decoding module 220, and the repeated parts will not be described again.

[0078] In summary, the execution flow of method 400 is as follows: input data is fed into text generation model 200, and after processing, a long text containing multiple sentences is output. The input data can be text, images, audio, etc., and this disclosure does not impose any restrictions on it. The following describes each execution step.

[0079] like Figure 4 As shown, the text generation method 400 begins with step S410. In step S410, a first vector indicating its semantic features is extracted from the input data based on the encoding module 210.

[0080] Subsequently, in step S420, based on the first vector, at least one second vector indicating the semantic features of the clause is determined using the clause content planning unit 222 in the decoding module 220.

[0081] Subsequently, in step S430, the word generation unit 224 in the decoding module 220 is used to process the second vector to generate multiple word vectors.

[0082] Subsequently, in step S440, word vectors are combined into at least one sentence vector to generate long text.

[0083] According to one embodiment, a beam search is performed at each time step to obtain the probability value of each sentence vector. A first number of sentence vectors are selected sequentially as candidate sequences, based on their probability values ​​from largest to smallest. Then, for each candidate sequence, the difference value between every two sentence vectors is calculated. Based on the difference values, the candidate sequence is reconstructed. This process of calculating the difference values ​​and reconstructing the candidate sequences is repeated iteratively until the difference values ​​meet a preset condition, at which point the sentence vectors belonging to the candidate sequence are determined. For details on calculating the difference values ​​and setting the preset conditions, please refer to the preceding descriptions.

[0084] According to the text generation method disclosed herein, step S420 divides the long text generation task into multiple sub-tasks to fully capture the semantic information between clauses, thereby improving the problem of semantic incoherence. Steps S430 and S440 incorporate differential considerations during bundle search to increase text diversity.

[0085] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of this disclosure may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.

[0086] Similarly, it is to be understood that the disclosure of a particular feature or aspect in one example does not imply that the feature or aspect is necessary for or is limiting of the disclosure. Rather, the disclosure of various features and aspects in the specification and the claims should be construed to be illustrative and not restrictive, with the scope of the disclosure being given by the appended claims.

[0087] Those skilled in the art will understand that the modules, or units, or components of the devices in the examples disclosed herein can be arranged in a device as described in the examples, or alternatively can be located in one or more devices different from the devices in the examples. The modules in the foregoing examples can be combined into one module or further divided into multiple sub-modules.

[0088] Those skilled in the art will understand that the modules in the devices in the examples can be adaptively changed and disposed in one or more devices different from the examples. The modules or units or components in the examples can be combined into one module or unit or component, and further can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination of all the features disclosed in the specification (including the accompanying claims, abstract and drawings), and all the processes or units of any method or device disclosed thus can be adopted. Unless explicitly stated otherwise, each feature disclosed in the specification (including the accompanying claims, abstract and drawings) can be replaced by an alternative feature providing the same, equivalent or similar purpose.

[0089] The disclosure also discloses:

[0090] A6. The decoding module of A2, wherein the word generation unit is further adapted to determine a difference value between two sentence vectors by:

[0091]

[0092] wherein two sentence vectors are represented as a(x 11 , x 12 … x 1n ) and b(x 21 , x 22 … x 2n ).

[0093] A7. The decoding module of any of Al-6, wherein the clause content planning unit employs a recurrent neural network or a long short-term memory network.

[0094] Furthermore, to the extent that the terms "comprises", "includes", "has", "contains", "involves", and the like can be used in the detailed description or the claims, these terms are always intended to be inclusive in a manner similar to the term "comprising" as an open transition term without precluding any additional or other elements or limitations.

[0095] Furthermore, some of the embodiments described herein are of a "method" or a "process" that can be embodied in software, firmware or hardware, and when embodied in software, can be implemented with computer- executable instructions (computer-readable code). The computer-readable code comprises a computer program of instructions that is

[0096] As used herein, the terms "first", "second", "third", etc. are used merely as labels, and are not intended to impose numerical requirements on their objects.

[0097] While the present disclosure has been described in connection with limited number of embodiments, those skilled in the art will appreciate that other embodiments can be devised which fall within the scope of the present disclosure. Additionally, it is to be understood that the description of the present disclosure is intended to be illustrative, and not restrictive, and that numerous other modifications and variations are intended to be possible within the scope of the present disclosure. It is intended that the scope of the present disclosure not be limited by any of the foregoing details, but instead be defined by the scope of the claims to follow.

Claims

1. A decoding module, adapted to be arranged in a text generation model, coupled with an encoding module, comprising: a clause content planning unit, coupled with the encoding module, adapted to receive an output of the encoding module, and process a first vector output by the encoding module to determine a second vector indicating semantic features of a clause; a word generation unit, coupled with the clause content planning unit, adapted to process the second vector to generate word vectors corresponding to a plurality of words, and combine the word vectors into at least one sentence vector to generate long text; wherein the word generation unit is further adapted to: perform beam search at each time step to obtain probability values of the sentence vectors; select a first number of sentence vectors in order of probability values from large to small as a candidate sequence; and calculate a difference value between each two sentence vectors in the candidate sequence respectively; reconstruct the candidate sequence based on the difference values; and repeat the steps of calculating the difference value and reconstructing the candidate sequence until the difference value meets a preset condition, to determine the sentence vectors belonging to the candidate sequence.

2. The decoding module of claim 1, wherein, The word generation unit is further adapted to: remove one of the two sentence vectors from the candidate sequence when the difference value between the two sentence vectors is greater than a threshold value; and reconstruct the candidate sequence by sequentially adding a sentence vector with a maximum probability value to the candidate sequence.

3. The decoding module of claim 1 or 2, wherein, The clause content planning unit is further adapted to: determine a probability distribution of the semantic features of a clause at a current time step based on the first vector output by the encoding module and the semantic features of a clause output at a previous time step; and generate the second vector indicating the semantic features of the clause at the current time step based on the probability distribution.

4. The decoding module of claim 1 or 2, wherein, The semantic features of the clause include at least one of the following features: entity attributes, topic features, and sentiment features of the clause.

5. The decoding module of claim 1, wherein, The word generation unit is further adapted to determine the difference value between the two sentence vectors by: wherein two sentence vectors are represented as a(x 11 ,x 12 ……x 1n ) and b(x 21 ,x 22 ……x 2n ). 6.The decoding module of claim 1 or 2, wherein the clause content planning unit employs a recurrent neural network or a long short-term memory network. 7.A text generation model, comprising: an encoding module, adapted to process input data to extract a first vector indicating semantic features thereof; the decoding module of any one of claims 1-6, coupled with the encoding module, adapted to process the first vector to generate at least one sentence vector to compose long text. 8.A text generation method, comprising the steps of: extracting a first vector indicating semantic features of input data; determining a second vector indicating semantic features of a clause based on the first vector; processing the second vector to generate a plurality of word vectors; and combining the word vectors into at least one sentence vector to generate long text. The step of combining the word vectors into at least one sentence vector comprises: performing a beam search at each time step to obtain a probability value of each sentence vector; sequentially selecting a first number of sentence vectors in descending order of the probability values as a candidate sequence; calculating a difference value between each two sentence vectors in the candidate sequence; reconstructing the candidate sequence based on the difference value; and repeating the steps of calculating the difference value and reconstructing the candidate sequence until the difference value meets a preset condition, and determining the sentence vectors belonging to the candidate sequence.

9. A computing device comprising: one or more processors; memory; one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs comprise instructions for performing the method of claim 8.

10. A computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by a computing device, cause the computing device to perform the method of claim 8.

Citation Information

Patent Citations

  • Short text generation method and device and readable storage medium

    CN111126059A

  • Abstract generation method, device and equipment

    CN111723194A