A text generation method, device, storage medium, and electronic device
Through the machine learning model of the seq2seq framework, the encoder and decoder are used to generate target text containing custom text, which solves the problem of lack of personalization of artificial intelligence copy generated and achieves the personalization and outstanding feature effect of copywriting.
Patent Information
- Application Number
- CN202110104025.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-26
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-01-26
AI Technical Summary
In the prior art, the copywriting generated by artificial intelligence lacks personalization and is difficult to highlight the characteristics of the product, resulting in the generated text being the same.
Using the machine learning model of the seq2seq framework, the initial text and custom text are input through the combination of encoder and decoder to generate target text containing user needs.
It realizes personalized copywriting generation, avoids the same phenomenon, meets user needs, and highlights the characteristics of the product.
Smart Images

Figure CN113762523B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the technical field of text generation, and in particular, to a text generation method, apparatus, storage medium, and electronic device. Background Art
[0002] With the continuous improvement of artificial intelligence technology, machines have been able to continuously replace humans to work in various scenarios. Specifically in the scenario of copywriting creation, the creation ability of AI has shined in directions such as news comments, product marketing copywriting, and product reviews. For example, in the scenario of product marketing copywriting, the machine generates a vivid marketing copy based on information such as the description and attributes of the product, thereby impressing consumers to place orders and ultimately increasing the sales volume of the product.
[0003] As a brand-new technology, AI-generated copywriting can indeed provide a smooth description of the product. However, in actual production, this copywriting generation process is uncontrollable, that is, artificial intelligence algorithms are mostly presented in an end-to-end form, and the whole process is close to a black box, resulting in the generated texts being stereotyped and difficult to highlight the characteristics of the product. Summary of the Invention
[0004] The present invention provides a text generation method, apparatus, storage medium, and electronic device to generate a target text including the custom text.
[0005] In a first aspect, an embodiment of the present invention provides a text generation method, including:
[0006] Obtain an initial text and a custom text;
[0007] Input the initial text into the encoder of the text generation model to obtain the encoded feature of the initial text;
[0008] Input the encoded feature and the custom text into the decoder of the text generation model to obtain a target text including the custom text.
[0009] In a second aspect, an embodiment of the present invention further provides a text generation apparatus, including:
[0010] A text acquisition module for obtaining an initial text and a custom text;
[0011] An encoded feature acquisition module for inputting the initial text into the encoder of the text generation model to obtain the encoded feature of the initial text;
[0012] A target text generation module for inputting the encoded feature and the custom text into the decoder of the text generation model to obtain a target text including the custom text.
[0013] In a third aspect, an embodiment of the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the text generation method provided in any embodiment of the present invention is implemented.
[0014] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the text generation method provided in any embodiment of the present invention is implemented.
[0015] The technical solution provided in this embodiment performs text generation through a text generation model configured with an encoder and a decoder. The custom text and the encoded features output by the encoder are used as the input information of the decoder. The decoder decodes the custom text and the encoded features to obtain a targeted text including the custom text, that is, the target text, ensuring that the target text includes the text information required by the user, avoiding the uniformity of the copywriting, and achieving the effect of meeting the user's needs and highlighting the features. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a flowchart of a text generation method provided in Embodiment 1 of the present invention;
[0017] Figure 2 is a schematic structural diagram of a text generation model provided in an embodiment of the present invention;
[0018] Figure 3 is a process view of a text generation method provided in an embodiment of the present invention;
[0019] Figure 4 is a schematic diagram of a decoding process provided in an embodiment of the present invention;
[0020] Figure 5 is a schematic diagram of a decoding process provided in an embodiment of the present invention;
[0021] Figure 6 is a schematic diagram of a decoding process of a candidate text provided in an embodiment of the present invention;
[0022] Figure 7 is a schematic diagram of a decoding process of a candidate text provided in an embodiment of the present invention;
[0023] Figure 8 is a schematic structural diagram of a text generation device provided in an embodiment of the present invention;
[0024] Figure 9 is a schematic structural diagram of an electronic device provided in Embodiment 4 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. Additionally, it should be noted that for the sake of description, only the parts related to the present invention rather than all the structures are shown in the accompanying drawings.
[0026] Embodiment 1
[0027] Figure 1 FIG. is a flowchart of a text generation method provided in Embodiment 1 of the present invention. This embodiment is applicable to the situation of generating a target text based on a text generation model. This method can be executed by a text generation device provided in an embodiment of the present invention. The device can be implemented in software and / or hardware, and can be integrated into an electronic device such as a computer or a mobile phone. The specific steps are as follows:
[0028] S110. Obtain an initial text and a custom text.
[0029] S120. Input the initial text into the encoder of the text generation model to obtain the encoded features of the initial text.
[0030] S130. Input the encoded features and the custom text into the decoder of the text generation model to obtain a target text including the custom text.
[0031] In this embodiment, the text generation model is a machine learning model in the seq2seq framework, including an encoder and a decoder. Exemplarily, refer to Figure 2 , Figure 2 which is a schematic structural diagram of the text generation model provided in an embodiment of the present invention. Figure 2 In, the encoder is connected to the decoder. The encoder is used to encode the input text to obtain encoded features, and input the encoded features and the custom text into the decoder. The decoder decodes the encoded features and the custom text and outputs the target text corresponding to the input text. In some embodiments, the encoder and the decoder are respectively recurrent network models. The network structures of the above encoder and decoder are not limited and can be set according to requirements. Exemplarily, the encoder and the decoder can be recurrent network models based on the self-attention mechanism, or can also be recurrent network models such as LSTM (Long Short-Term Memory). In some embodiments, the encoder and the decoder can be different types of recurrent network models.
[0032] Among them, the initial text is the text from which information is to be extracted. For example, it can be, but is not limited to, product description text, text introducing an object or a scene, news text, or an article, etc. The custom text can be a single word, a phrase, or a sentence, and can be set according to user needs. The target text can be a summary text, a copywriting text, etc. of the initial text, and there is no limitation on this.
[0033] In some embodiments, the encoded features of the initial text and the custom text are used as the input to the decoder. Optionally, the custom text is converted into a text vector, and this text vector and the encoded features are input into the decoder; alternatively, it can also be that the custom text is tokenized to obtain multiple tokens, and each token is converted into a word vector, and each word vector and the encoded features are input into the decoder. The decoder synchronously processes the encoded features and the custom text, breaking the situation of getting uniform texts based only on the encoded features. The target text obtained includes the custom text, achieving the effect of highlighting the information characteristics. It should be noted that the target text includes each word in the custom text, and each word in the custom text can be distributed in the target text or set as a whole in the target text. Exemplarily, if the custom text is "full-screen mobile phone", correspondingly, the target text obtained by the text generation model can be "This is a full-screen mobile phone" or "The Huawei mobile phone is configured with a full-screen".
[0034] In some embodiments, when the custom text includes multiple tokens, each token in the custom text can be input into the decoder at different decoding nodes of the decoder, so that the decoder processes each token separately, improving the diversity of the output target text. Exemplarily, if the decoder is a recurrent network model, multiple tokens in the custom text can be used as the input information for different recursions. Correspondingly, the input information for each recursion can include the output of the previous recursion and the externally input token.
[0035] In some embodiments, the output result of the text generation model includes multiple candidate texts and the confidence probability of each candidate text. The candidate text with the highest confidence probability is determined as the target text. Further, it can be to determine the candidate texts that include the custom text among the candidate texts, and among the candidate texts that include the custom text, the candidate text with the highest confidence probability is determined as the target text.
[0036] In some embodiments, the output result of the text generation model includes multiple candidate texts. The candidate texts that include the custom text are screened, the text quality of the above-screened candidate texts is evaluated, and the candidate text with the highest quality evaluation score is determined as the target text.
[0037] Based on the above embodiments, the text generation model can be trained based on sample texts and the standard texts corresponding to the sample texts. Optionally, the text generation model can be a copywriting generation model. Correspondingly, the copywriting generation model is trained based on sample texts and the standard copywriting corresponding to the sample texts, where the sample texts include initial sample texts and custom sample texts. For example, the sample text is input into the encoder of the text generation model to be trained, and the encoded features output by the encoder and the custom sample text are input into the decoder to obtain a predicted copywriting. A loss function is generated based on the predicted text and the standard copywriting, and the above text generation model is adjusted backward based on this loss function. The above training process is repeated until the training conditions are met, and a trained text generation model is obtained.
[0038] The technical solution provided in this embodiment performs text generation through a text generation model configured with an encoder and a decoder. The custom text and the encoded features output by the encoder are used as the input information for the decoder. The encoder decodes the custom text and the encoded features to obtain a targeted text including the custom text, that is, the target text, ensuring that the target text includes the text information required by the user, avoiding the uniformity of copywriting, and achieving the effect of meeting the user's needs and having prominent features.
[0039] Embodiment 2
[0040] Figure 3 It is a process view of a text generation method provided by an embodiment of the present invention. Based on the above embodiments, the method of inputting custom text into the decoder is refined. The method includes:
[0041] S210. Obtain an initial text and a custom text, perform word segmentation on the custom text to obtain at least one word segment of the custom text.
[0042] S320. Input the initial text into the encoder of the text generation model to obtain the encoded features of the initial text.
[0043] S230. Determine the input nodes of each word segment in sequence according to the order of each word segment in the custom text.
[0044] S240. Input the encoded features into the decoder, and input each word segment into the decoder based on the input nodes of each word segment to obtain the target text output by the decoder.
[0045] In this embodiment, the word segmentation order of each word segment of the custom text is determined. Optionally, the word segmentation order can be the order of each word segment in the custom text. Exemplarily, if the custom text is "full-screen mobile phone", the order of the word segments "full screen" and "mobile phone" is that "full screen" is before and "mobile phone" is after. Optionally, the word segmentation order can also be determined based on the importance of each word segment in the custom text. The higher the importance, the more forward the order.
[0046] The decoder is a recurrent network model. Correspondingly, the input nodes of the word segments can be each recurrent input node. Optionally, the order of the recurrent input nodes of each word segment input to the decoder can be the same as the word segmentation order, that is, the word segmentation order is before, and the recurrent input node of the input decoder is before. Correspondingly, the word segmentation order is after, and the recurrent input node of the input decoder is after. Exemplarily, for the i-th word segment in the word segmentation order, it is input to the decoder at the j-th recurrent input node, and the (i + 1)-th word segment is input to the decoder at the (j + 1)-th recurrent input node, where i and j can be the same or different. In some embodiments, j = i, i is a positive integer greater than or equal to 1, and j is a positive integer greater than or equal to 0.
[0047] In some embodiments, inputting the encoded feature into the decoder and inputting each word segment into the decoder based on the input nodes of each word segment can be inputting the encoded feature into the decoder. The decoder performs the first loop, takes the output result of the first loop and the word segment in the first order as the input information for the second loop of the decoder, takes the output result of the second loop and the word segment in the second order as the input information for the third loop of the decoder, and so on, to obtain the target text.
[0048] In some embodiments, the encoded feature is used as the initial input information of the encoder, that is, the input node j of the initial loop is 1. In each loop, the word segment i corresponding to the input node j of each loop is determined, and this word segment i and the input result of the previous loop are used as the input information of the current input node j, where j = i. Exemplarily, see Figure 4 , Figure 4 is a schematic diagram of a decoding process provided by an embodiment of the present invention.
[0049] Based on the above embodiments, the output result of each loop is obtained through beam search decoding.
[0050] Optionally, inputting the encoded feature into the decoder, and inputting each of the word segments into the decoder based on the input nodes of the word segments, to obtain the target text output by the decoder, including: using the encoded feature as the initial input information of the recurrent network model, and performing the following recurrent process until the recurrent end condition is satisfied, and determining the target text based on the output result of the recurrent network model: obtaining the current candidate text output by the recurrent network module in any recurrent iteration, where the current candidate text includes a first candidate text corresponding to the encoded feature, or the first candidate text and second candidate texts respectively corresponding to each of the input word segments; inputting the current candidate text and the word segment of the current input node into the recurrent network model, to obtain an updated candidate text corresponding to the current candidate text and a new candidate text corresponding to the word segment of the current input node.
[0051] Exemplarily, refer to Figure 5 , Figure 5 which is a schematic diagram of the decoding process provided by an embodiment of the present invention. Figure 5 In [the figure], the encoded feature is used as the initial input information for the first recurrent decoding of the decoder, to obtain a first candidate text corresponding to the encoded feature, that is, the current candidate text. This first candidate text is the candidate text corresponding to the position of the first word in the target text, and this first candidate text may include one or more candidate words. The first candidate text and the word segment of the current input node (the first word segment) are input into the recurrent network model for the second recurrent decoding. Among them, the recurrent network model decodes the first candidate text and the first word segment of the current input node respectively, to obtain an updated candidate text of the first candidate text and a new candidate text of the first word segment. Among them, this updated candidate text is a combined text of candidate words at the positions of the first word and the second word, and this updated candidate text may include one or more combined texts. Similarly, the new candidate text includes a combined text of candidate words at the positions of the first word and the second word. The above updated candidate text is used as the first candidate text in the current candidate text, the above new candidate text is used as the second candidate text (that is, the first word segment is an input word segment), and the first candidate text, the second candidate text, and the word segment of the current input node (the second word segment) are input into the recurrent network model for the third recurrent decoding, to obtain a first updated candidate text corresponding to the first candidate text input above, a second updated candidate text corresponding to the second candidate text, and a new candidate text of the second word segment, and so on, to obtain the target text.
[0052] Among them, each text in the updated candidate text is the text formed by combining each text output in the previous cycle with the next candidate word. Here, the next candidate word can be a word in the initial text and / or a segmented word of the custom text that has been input. Each text in the newly added candidate text is the text formed by combining the newly added analysis with the next candidate word. Here, the next candidate word can be a word in the initial text and / or a segmented word of the custom text that has been input.
[0053] Exemplarily, refer to Figure 6 , Figure 6 which is a schematic diagram of the decoding process of the candidate text provided by the embodiment of the present invention. Figure 6 Decoding starts from . The encoded features are input into the decoder (i.e., the recurrent network model) to obtain the first candidate text. The corresponding first candidate text includes the candidate words "this model" and "this unit". The above first candidate text and the first segmented word "full-screen" are input into the decoder for recurrent decoding to obtain the updated candidate texts "this model of mobile phone" and "this model of cost performance" corresponding to the first candidate text, and the newly added candidate texts "Huawei full-screen" and "this full-screen" corresponding to the first segmented word. The above updated candidate texts are used as the current first candidate information, the newly added candidate texts are used as the current second candidate texts, and the first candidate information, the second candidate texts, and the segmented word "mobile phone" of the current input node are input into the decoder for recurrent decoding to obtain the first updated candidate texts "this model of mobile phone has" and "this model of cost performance is extremely high" corresponding to the first candidate text, the second updated candidate texts "this model of mobile phone with full screen" and "this full-screen is high-definition" corresponding to the second candidate texts, and the newly added candidate information "the full-screen mobile phone is great" and "this full-screen mobile phone" of the second segmented word, and so on. It should be noted that each candidate information is not limited to two text messages, and can also be other preset numbers of text messages, which is not limited herein.
[0054] In this embodiment, by inputting multiple segmented words in the custom text into the decoder at different recurrent input nodes for recurrent decoding to obtain the candidate text including the input segmented words, it is convenient to determine the target text including each segmented word in the above candidate text, achieving the effect of inserting the custom text into the target text. At the same time, the candidate text output from the encoded features and the candidate texts corresponding to each segmented word form different batches of output results. For example, refer to Figure 6, which facilitates determining the target text among candidate texts in each batch and avoids the problem that one or more word segments are missed during the decoding process when processed in a mixed manner in the same batch. Exemplarily, when screening candidate texts based on the prediction probability of the candidate texts, there is a situation where the prediction probability of a candidate text containing a word segment is less than that of a candidate text not containing a word segment. In the case of mixed processing, there is a situation where the candidate text containing the word segment is excluded, resulting in the omission of the above-mentioned word segment. In this embodiment, the candidate texts corresponding to each word segment are screened and processed in layers to ensure that there is candidate information including each word segment. Further, by sorting each word segment, the input node of each word segment is determined, and the insertion position of each word segment is optimized, achieving the effect of optimizing the position of the custom text in the target text.
[0055] Based on the above embodiment, inputting the current candidate text and the word segment of the current input node into the recurrent network model includes: if there is no corresponding word segment for the current input node, inputting the current candidate text into the recurrent network model. Specifically, for each input node, it is determined whether there is a corresponding word segment to be input. If not, the candidate texts decoded in the previous cycle are input into the decoder.
[0056] Based on the above embodiment, there is at least one custom text, and correspondingly, each custom text respectively corresponds to at least one word segment. Exemplarily, the custom text includes "full-screen mobile phone" and "color-changing rear case". Correspondingly, "full-screen mobile phone" corresponds to the word segments "full screen" and "mobile phone", and "color-changing rear case" corresponds to the word segments "color change" and "rear case". It can be to mix and sort the word segments of multiple custom texts, or to sort the word segments of each custom text separately. Correspondingly, each sorting position can correspond to multiple word segments. Exemplarily, the first sorting position corresponds to the word segments "full screen" and "color change", and the second sorting position corresponds to "mobile phone" and "rear case". Among them, the word segments in the same sorting position correspond to the same input node.
[0057] Correspondingly, inputting the current candidate text and the word segment of the current input node into the recurrent network model to obtain the updated candidate text corresponding to the current candidate text and the new candidate text corresponding to the word segment of the current input node includes: inputting the current candidate text and at least one word segment of the current input node into the recurrent network model to obtain the updated candidate text corresponding to the current candidate text and the new candidate texts respectively corresponding to each word segment of the current input node.
[0058] At the input node of any cycle of the decoder, determine the word segmentation corresponding to the current input node. The word segmentation can be none, one, or two or more. Input the determined word segmentation and the current candidate text into the decoder for cyclic decoding to obtain the new candidate texts corresponding to the respective analyses of the input. Exemplarily, see Figure 7 , Figure 7 which is a schematic diagram of the decoding process of the candidate text provided by the embodiment of the present invention. Figure 7 In the input node of the second cycle, the input word segmentations include "full-screen" and "color-changing", and in the input node of the third cycle, the input word segmentation includes "mobile phone".
[0059] In this embodiment, by synchronously processing multiple custom texts, a target text including multiple custom texts is obtained. Synchronously processing multiple custom texts at the same time improves the text processing efficiency and avoids the situation that equally important word segmentations have significant effect differences in the target text due to the decoding order.
[0060] On the basis of the above embodiment, each text in the current candidate text is configured with a corresponding confidence probability. Exemplarily, each text in the candidate text may include at least one candidate word, and each candidate word corresponds to a confidence probability. The product of the confidence probabilities of each candidate word is the confidence probability of the formed text. Exemplarily, see Figure 6 where the confidence probability of "this model" is 0.5 and the confidence probability of "mobile phone" is 0.8, then the confidence probability of "this model mobile phone" is 0.4.
[0061] Optionally, before inputting the current candidate text and the word segmentation of the current input node into the recurrent network model, it further includes: for the current candidate text, screening a first preset number of texts based on the confidence probabilities of the texts in the first candidate text, and screening a first preset number of texts based on the confidence probabilities of the texts in the second candidate text. Each batch of candidate texts output by the decoder (such as the first candidate text and the second candidate text) includes a large number of candidate words and texts formed by a large number of candidate words (i.e., texts formed by combining candidate words). Among them, the candidate words can be the analyses in the initial text. To simplify the decoding process, for each batch of candidate texts output by the decoder each time, screen based on the confidence probabilities of the texts in the candidate text to obtain a first preset number of texts, that is, sort the texts in each batch of candidate texts from large to small based on the confidence probabilities, screen the first preset number of texts before sorting, discard other texts, and the current candidate text after screening, that is, the first candidate text after screening and the second candidate text after screening corresponding to each input word segmentation. Among them, the first preset number for screening is a positive integer greater than 1. For example, it can be 5 or 10, which is not limited thereto and can be set according to user requirements.
[0062] Correspondingly, inputting the current candidate text and the word segmentation of the current input node into the recurrent network model includes: inputting the filtered current candidate text and the word segmentation of the current input node into the recurrent network model. In this embodiment, the candidates for each recurrent input of the decoder are filtered by the confidence probability, and the texts with low confidence probability are discarded, reducing the computational amount while avoiding the interference of low-probability texts and improving the text generation efficiency and accuracy.
[0063] The decoder executes the above process in a loop until the loop end condition is met to obtain the target text. The loop end condition includes: a preset number of loops and / or the end flag is included in the output candidate text. The preset number of loops is the word length of the target text. When the number of loops of the decoder meets the preset number of loops, the decoding ends and the target text is determined. The end flag is pre-set. If the text with the highest confidence probability in the current candidate text includes the end flag, the decoding ends; or, if the text with the highest confidence probability in the second candidate text corresponding to the word segmentation of the last input includes the end flag, the decoding ends. The word segmentation of the last input is the word segmentation input to the decoder last time.
[0064] Based on the output result of the recurrent network model in the above embodiment, determining the target text includes: determining the target text including the custom text from the second candidate texts respectively corresponding to each input word segmentation. Each word segmentation corresponds to a batch of second candidate texts, and each second candidate text includes multiple texts. The texts including the custom text in each second candidate text are determined as the target text, ensuring that the target text includes the custom text. The diversity of the target text is improved by obtaining multiple target texts. In some embodiments, the target text is determined based on the confidence probability of the texts including the custom text in each second candidate text. For example, the text with a confidence probability greater than the preset probability threshold in the texts including the custom text can be determined as the target text, or the text with the highest confidence probability can be determined as the target text.
[0065] Based on the output result of the recurrent network model in the above embodiment, determining the target text includes: in the second candidate text corresponding to the word segmentation of the last input, determining the candidate text with the highest confidence probability as the target text. Among them, the second candidate text corresponding to the word segmentation of the last input includes each word segmentation in the custom text, and there is no need to judge the custom text. Determining the candidate text with the highest confidence probability as the target text ensures the text quality of the target text.
[0066] Embodiment III
[0067] Figure 8 is a schematic structural diagram of a text generation device provided by an embodiment of the present invention. The device includes:
[0068] A text acquisition module 310, configured to acquire an initial text and a custom text;
[0069] An encoding feature acquisition module 320, configured to input the initial text into an encoder of a text generation model to obtain an encoding feature of the initial text;
[0070] A target text generation module 330, configured to input the encoding feature and the custom text into a decoder of the text generation model to obtain a target text including the custom text.
[0071] Based on the above embodiments, the apparatus further includes:
[0072] A word segmentation module, configured to perform word segmentation on the custom text after acquiring the custom text to obtain at least one word segment of the custom text;
[0073] Correspondingly, the target text generation module 330 includes:
[0074] An input node determination unit, configured to sequentially determine input nodes of each word segment according to the order of each word segment in the custom text;
[0075] A target text generation unit, configured to input the encoding feature into the decoder, and input each word segment into the decoder based on the input nodes of each word segment to obtain a target text output by the decoder.
[0076] Based on the above embodiments, the decoder is a recurrent network model;
[0077] Based on the above embodiments, the target text generation unit is configured to:
[0078] A current candidate text acquisition subunit, configured to acquire a current candidate text output by the recurrent network module in any cycle, where the current candidate text includes a first candidate text corresponding to the encoding feature, or the first candidate text and second candidate texts respectively corresponding to each input word segment, and the encoding feature is initial input information of the recurrent network model;
[0079] A candidate text update subunit, configured to input the current candidate text and the word segment of the current input node into the recurrent network model to obtain an updated candidate text corresponding to the current candidate text and a new candidate text corresponding to the word segment of the current input node;
[0080] A target text determination subunit, configured to determine a target text based on an output result of the recurrent network model if a cycle end condition is satisfied.
[0081] Based on the above embodiments, each text in the current candidate text is configured with a corresponding confidence probability.
[0082] Based on the above embodiments, the target text generation unit further includes:
[0083] A text screening subunit, before inputting the current candidate text and the word segmentation of the current input node into the recurrent network model, for the current candidate text, screening a first preset number of texts based on the confidence probabilities of the texts in the first candidate text, and screening a first preset number of texts based on the confidence probabilities of the texts in the second candidate text.
[0084] Correspondingly, the candidate text update subunit is used for:
[0085] Inputting the screened current candidate text and the word segmentation of the current input node into the recurrent network model.
[0086] Based on the above embodiments, the loop end condition includes: a preset number of loops and / or the candidate text output later includes an end identifier;
[0087] Based on the above embodiments, the target text determination subunit is used for:
[0088] Determining a target text including the custom text from the second candidate texts corresponding to each input word segmentation; or,
[0089] In the second candidate text corresponding to the word segmentation input at the end, determining the candidate text with the highest confidence probability as the target text.
[0090] The above product can execute the method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the executed method.
[0091] Embodiment 4
[0092] Figure 9 It is a schematic structural diagram of an electronic device provided by Embodiment 4 of the present invention. Figure 9 It shows a block diagram of an electronic device 12 suitable for implementing the embodiments of the present invention. Figure 9 The shown electronic device 12 is only an example, and should not bring any limitation to the functions and usage scopes of the embodiments of the present invention. The device 12 is typically an electronic device that undertakes the function of image classification.
[0093] As Figure 9 shown, the electronic device 12 is presented in the form of a general-purpose computing device. The components of the electronic device 12 may include, but are not limited to: one or more processors 16, a storage device 28, and a bus 18 connecting different system components (including the storage device 28 and the processor 16).
[0094] Bus 18 represents one or more of several types of bus architectures, including a memory bus or memory controller, a peripheral bus, an Accelerated Graphics Port, a processor, or a local bus using any of the various bus architectures. By way of example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.
[0095] Electronic device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by electronic device 12, including volatile and nonvolatile media, removable and non-removable media.
[0096] Storage device 28 may include computer system readable media in the form of volatile memory, such as Random Access Memory (RAM) 30 and / or cache memory 32. Electronic device 12 may further include other removable / non-removable, volatile / nonvolatile computer system storage media. By way of example only, storage system 34 can be used for reading and writing on non-removable, nonvolatile magnetic media ( Figure 9 not shown, typically referred to as a "hard disk drive"). Although Figure 9 not shown in the figure, a disk drive for reading and writing on a removable nonvolatile disk (such as a "floppy disk") and an optical disk drive for reading and writing on a removable nonvolatile optical disk (such as a Compact Disc-Read Only Memory (CD-ROM), Digital Video Disc-Read Only Memory (DVD-ROM), or other optical media) may be provided. In these cases, each drive may be connected to bus 18 through one or more data media interfaces. Storage device 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present invention.
[0097] A program 36 having a set (at least one) of program modules 26 can be stored, for example, in a storage device 28. Such program modules 26 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment. The program modules 26 generally execute the functions and / or methods in the embodiments described in the present invention.
[0098] The electronic device 12 can also communicate with one or more external devices 14 (such as a keyboard, a pointing device, a camera, a display 24, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 12, and / or communicate with any device that enables the electronic device 12 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface 22. Moreover, the electronic device 12 can also communicate with one or more networks (such as a Local Area Network (LAN), a Wide Area Network (WAN), and / or a public network, such as the Internet) through a network adapter 20. As shown in the figure, the network adapter 20 communicates with other modules of the electronic device 12 through a bus 18. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, Redundant Arrays of Independent Disks (RAID) systems, tape drives, and data backup storage systems, etc.
[0099] The processor 16 executes various functional applications and data processing by running the program stored in the storage device 28, such as implementing the text generation method provided in the above embodiments of the present invention.
[0100] Embodiment Five
[0101] Embodiment Five of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the text generation method provided in the embodiments of the present invention.
[0102] Certainly, the computer program stored on the computer-readable storage medium provided in the embodiments of the present invention is not limited to the method operations as described above, and can also execute the text generation method provided in any embodiment of the present invention.
[0103] The computer storage medium of the embodiments of the present invention may adopt any combination of one or more computer-readable media. The computer-readable media may be computer-readable signal media or computer-readable storage media. The computer-readable storage media may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage media may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device.
[0104] The computer-readable signal media may include data signals propagated in a baseband or as part of a carrier wave, which carry computer-readable source code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal media may also be any computer-readable media other than the computer-readable storage media, and the computer-readable media may send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.
[0105] The source code contained on the computer-readable media may be transmitted by any appropriate media, including but not limited to wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0106] The computer source code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The source code may be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0107] Note that the above is only a preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, re-adjustments and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments only. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. A text generation method, characterized in that, including: obtaining an initial text and a custom text, where the custom text includes at least one word segment; inputting the initial text into an encoder of a text generation model to obtain an encoded feature of the initial text; inputting the encoded feature and the custom text into a decoder of the text generation model to obtain a target text including the custom text; the target text includes each of the word segments in the custom text; the decoder is a recurrent network model; the input information of the first cycle of the recurrent network model is the encoded feature; the input information of other cycles except the first cycle of the recurrent network model includes the current candidate text output in the previous cycle and the word segment of the current input node; the output information of the decoder includes an updated candidate text corresponding to the current candidate text and a new candidate text corresponding to the word segment of the current input node.
2. The method according to claim 1, characterized in that, After obtaining the custom text, the method further includes: performing word segmentation processing on the custom text to obtain at least one word segment of the custom text; the inputting the encoded feature and the custom text into the decoder of the text generation model to obtain a target text including the custom text includes: sequentially determining input nodes of each word segment according to the order of each word segment in the custom text; inputting the encoded feature into the decoder, and inputting each of the word segments into the decoder based on the input nodes of the word segments to obtain the target text output by the decoder.
3. The method according to claim 2, It is characterized in that wherein, the inputting the encoded feature into the decoder, and inputting each of the word segments into the decoder based on the input nodes of the word segments to obtain the target text output by the decoder includes: taking the encoded feature as the initial input information of the recurrent network model, and performing the following loop process until a loop end condition is met, and determining the target text based on the output result of the recurrent network model: obtaining a current candidate text output in any cycle of the recurrent network module, where the current candidate text includes a first candidate text corresponding to the encoded feature, or the first candidate text and second candidate texts respectively corresponding to each input word segment; inputting the current candidate text and the word segment of the current input node into the recurrent network model to obtain an updated candidate text corresponding to the current candidate text and a new candidate text corresponding to the word segment of the current input node.
4. The method according to claim 3, wherein The inputting the current candidate text and the word segment of the current input node into the recurrent network model includes: if there is no corresponding word segment for the current input node, inputting the current candidate text into the recurrent network model.
5. The method according to claim 3, wherein the custom text is at least one; the inputting the current candidate text and the word segment of the current input node into the recurrent network model to obtain an updated candidate text corresponding to the current candidate text and a new candidate text corresponding to the word segment of the current input node includes: Input the current candidate text and at least one word segment of the current input node into the recurrent network model to obtain an updated candidate text corresponding to the current candidate text and new candidate texts corresponding to the respective word segments of the current input node.
6. The method according to any one of claims 3-5, characterized in that Each text in the current candidate text is configured with a corresponding confidence probability. Before inputting the current candidate text and the word segments of the current input node into the recurrent network model, it further includes: For the current candidate text, screen a first preset number of texts based on the confidence probabilities of the texts in the first candidate text, and screen a first preset number of texts based on the confidence probabilities of the texts in the second candidate text. Correspondingly, the step of inputting the current candidate text and the word segments of the current input node into the recurrent network model includes: Input the screened current candidate text and the word segments of the current input node into the recurrent network model.
7. The method according to claim 3, characterized in that, The loop end condition includes: a preset number of loops and / or an end flag included in the candidate texts output later. Determining the target text based on the output result of the recurrent network model includes: Determining a target text that includes the custom text from the second candidate texts corresponding to the respective input word segments; or, Among the second candidate texts corresponding to the word segments input at the end, determining the candidate text with the highest confidence probability as the target text.
8. A text generation device, characterized in that, It includes: A text acquisition module for acquiring an initial text and a custom text, where the custom text includes at least one word segment. An encoding feature acquisition module for inputting the initial text into the encoder of the text generation model to obtain the encoding feature of the initial text. A target text generation module for inputting the encoding feature and the custom text into the decoder of the text generation model to obtain a target text that includes the custom text; the target text includes each of the word segments in the custom text. The decoder is a recurrent network model; the input information for the first loop of the recurrent network model is the encoding feature, and the input information for other loops of the recurrent network model except the first loop includes the current candidate text output in the previous loop and the word segments of the current input node; the output information of the decoder includes the updated candidate text corresponding to the current candidate text and the new candidate texts corresponding to the word segments of the current input node.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the text generation method as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the text generation method as described in any one of claims 1-7.
Citation Information
Patent Citations
Method for training description text generation model, method and apparatus for generating description text
CN109062937A
Machine translation method and device and storage medium
CN111310485A