Controllable text generation method and device and electronic equipment
Through the method of fusion and small sample learning in the pre-trained language model, the problem of text fluency reduction and attribute interference in multi-attribute controllable text generation is solved, and high-quality and multi-attribute controllable text generation is achieved.
Patent Information
- Application Number
- CN202510231799.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-13
AI Technical Summary
The existing multi-attribute controllable text generation method causes text fluency to decrease, the splicing order affects the controllability of attributes, and there is a problem of single-attribute prompt vectors interfering with each other, causing attribute degradation.
By inputting the prompt text into the pre-trained language model, multiple attribute prompt vectors are fused based on aspect fusion vectors in the model, and associated knowledge is fused based on small sample learning to generate controllable text.
It effectively improves text fluency, alleviates attribute interference problems, and realizes a smooth transition from single-target attribute control to multiple-target attribute control. The generated text has multiple attributes and improves the practicality and flexibility of generation.
Smart Images

Figure CN120145304A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and particularly to a controllable text generation method, apparatus, and electronic device. Background Art
[0002] Multi-attribute controllable text generation refers to the ability to control the generated content according to multiple specified attributes or conditions when generating text. This method enables the generated text to be customized according to different requirements and contexts, increasing the flexibility and applicability of text generation.
[0003] Currently, multi-attribute controllable text generation is mainly achieved by directly concatenating multiple single-attribute prompt vectors into a multi-attribute prompt vector, and then generating multi-attribute controllable text based on the multi-attribute prompt vector.
[0004] However, using this generation method will result in a decrease in the fluency of the generated text, and the concatenation order will also affect the controllability of the attributes. It will also cause interference between single-attribute prompt vectors during the text generation stage, resulting in attribute degradation. Summary of the Invention
[0005] In view of this, the present application provides a controllable text generation method, apparatus, and electronic device, mainly aiming to improve the technical problems in the existing technology that the fluency of the generated text will decrease, the concatenation order will also affect the controllability of the attributes, and it will also cause interference between single-attribute prompt vectors during the text generation stage, resulting in attribute degradation.
[0006] In a first aspect, the present application provides a controllable text generation method, including:
[0007] Obtaining a prompt text for generating a target controllable text;
[0008] Inputting the prompt text into a pre-trained language model, performing attribute fusion on multiple attribute prompt vectors included in the prompt text based on an aspect fusion vector in the pre-trained language model, and performing associated knowledge fusion on the multiple attribute prompt vectors based on few-shot learning, and generating a controllable text based on the attribute fusion result and the associated knowledge fusion result;
[0009] Determining the generated controllable text as the target controllable text, where the target controllable text has the same attributes as the prompt text.
[0010] In a second aspect, the present application provides a controllable text generation apparatus, including:
[0011] An obtaining module, configured to obtain a prompt text for generating a target controllable text;
[0012] An input module, configured to input the prompt text into a pre-trained language model, perform attribute fusion on multiple attribute prompt vectors included in the prompt text based on an aspect fusion vector in the pre-trained language model, and perform associated knowledge fusion on the multiple attribute prompt vectors based on few-shot learning, and perform controllable text generation based on the attribute fusion result and the associated knowledge fusion result;
[0013] A determination module, configured to determine the generated controllable text as the target controllable text, where the target controllable text has the same attributes as the prompt text.
[0014] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the controllable text generation method of the first aspect is implemented.
[0015] In a fourth aspect, the present application provides an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, and when the processor executes the computer program, the controllable text generation method of the first aspect is implemented.
[0016] In a fifth aspect, the present application provides a computer program product, where the computer program product includes a computer program, and when the computer program is executed by a processor, the controllable text generation method of the first aspect is implemented.
[0017] With the above technical solutions, a controllable text generation method, apparatus, and electronic device provided by the present application, wherein the method includes: obtaining a prompt text for generating a target controllable text; inputting the prompt text into a pre-trained language model, performing attribute fusion on multiple attribute prompt vectors included in the prompt text based on an aspect fusion vector in the pre-trained language model, and performing associated knowledge fusion on the multiple attribute prompt vectors based on few-shot learning, and generating a controllable text based on the attribute fusion result and the associated knowledge fusion result; determining the generated controllable text as the target controllable text, where the target controllable text has the same attributes as the prompt text. Compared with the current existing technologies, by inputting the prompt text into the pre-trained language model, performing attribute fusion on multiple attribute prompt vectors included in the prompt text based on the aspect fusion vector in the pre-trained language model, and performing associated knowledge fusion on the multiple attribute prompt vectors based on few-shot learning, and generating a controllable text based on the attribute fusion result and the associated knowledge fusion result, the pre-trained language model in the present application can effectively utilize limited data resources, and can still generate texts that meet the requirements even in the case of lack of sufficient multi-attribute training corpora; the present application also alleviates the attribute interference problem that occurs after simply combining single-attribute continuous prompt vectors, realizes a smooth transition from single-target attribute control to multi-target attribute control, enables the attribute prompt vectors to be better fused in terms of both attributes and associated knowledge based on the aspect fusion vector, avoids interference between attribute vectors, and then can generate controllable texts based on the fusion result, obtaining controllable texts containing multiple attributes, and also improves the practicability and flexibility of controllable text generation.
[0018] The above description is only an overview of the technical solutions of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present application more obvious and understandable, the following specifically gives the specific embodiments of the present application. Brief Description of the Drawings
[0019] The accompanying drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0021] Figure 1 The flowchart showing a controllable text generation method provided by an embodiment of the present application;
[0022] Figure 2 shows a schematic flow chart of a controllable text generation method provided by an embodiment of the present application;
[0023] Figure 3 shows a schematic diagram of an example provided by an embodiment of the present application;
[0024] Figure 4 shows a schematic structural diagram of a controllable text generation device provided by an embodiment of the present application. Detailed implementation manners
[0025] Embodiments of the present application will be described in more detail below with reference to the accompanying drawings. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.
[0026] In order to improve the technical problems in the current existing technologies that may lead to a decrease in the fluency of the generated text, the splicing order may also affect the controllability of attributes, and it may also cause interference between single-attribute prompt vectors during the text generation stage, resulting in attribute degradation. This embodiment provides a controllable text generation method, as Figure 1 shown, the method includes:
[0027] Step 101, obtain a prompt text for generating a target controllable text.
[0028] In the embodiments of the present application, controllable text generation can be that when generating text, the generation process can be guided according to specific control parameters or conditions, so as to obtain text that meets the expected attributes. This method enables the generated text to be customized according to different requirements and contexts, increasing the flexibility and applicability of text generation.
[0029] In some examples, the prompt text plays a crucial role in natural language processing (NLP), especially in text generation tasks based on large-scale pre-trained models. A well-designed prompt can help the model better understand the task requirements and generate high-quality text that meets the expectations.
[0030] For this embodiment, the target controllable text can be a controllable text that needs to be generated based on the prompt text and has the attributes of the prompt text.
[0031] Step 102, input the prompt text into a pre-trained language model, perform attribute fusion on multiple attribute prompt vectors included in the prompt text based on an aspect fusion vector in the pre-trained language model, and perform associated knowledge fusion on the multiple attribute prompt vectors based on few-shot learning, and generate controllable text based on the attribute fusion result and the associated knowledge fusion result.
[0032] In the embodiments of the present application, pre-trained language models (PLMs) are an important technology in the field of natural language processing (NLP). By pre-training on large-scale text data, they learn rich language representations and can thus perform well in various downstream tasks. These models are typically based on deep learning architectures, especially the Transformer architecture, which can capture long-range dependencies and complex semantic information in the text.
[0033] In some examples, the attribute prompt vector contains various attribute information about the expected output text, that is, the attribute prompt vector in the embodiments of the present application contains the attribute information in the prompt text; correspondingly, the aspect fusion vector can integrate information from different sources or different aspects into a unified vector representation to better utilize this information for decision-making, prediction, or generation tasks. By effectively fusing information from different aspects, the expressiveness and accuracy of the model can be enhanced.
[0034] Specifically, in the embodiments of the present application, multiple attribute prompt vectors can be fused through the aspect fusion vector so that the resulting fusion vector has the attributes of multiple attribute vectors.
[0035] For this embodiment, few-shot learning is an important branch in the field of machine learning, aiming to train a model with a small number of labeled samples so that it can effectively solve new tasks or new categories; specifically, in the embodiments of the present application, few-shot learning can enable the vector obtained after fusion to have the associated knowledge of multiple attribute prompt vectors.
[0036] As an optional method, controllable text generation based on the attribute fusion result and the associated knowledge fusion result can generate controllable text with all the attributes of the prompt text through the vector after attribute fusion and the vector after associated knowledge fusion.
[0037] Step 103: Determine the generated controllable text as the target controllable text.
[0038] Among them, the target controllable text has the same attributes as the prompt text.
[0039] Compared with the current existing technologies, in this embodiment, by inputting the prompt text into the pre-trained language model, performing attribute fusion on multiple attribute prompt vectors included in the prompt text based on the aspect fusion vector in the pre-trained language model, and performing associated knowledge fusion on the multiple attribute prompt vectors based on few-shot learning, and performing controllable text generation based on the attribute fusion result and the associated knowledge fusion result, the pre-trained language model in this embodiment can effectively utilize limited data resources, and can still generate texts that meet the requirements even in the case of lack of sufficient multi-attribute training corpora; this embodiment also alleviates the attribute interference problem that occurs after simply combining single-attribute continuous prompt vectors, realizes a smooth transition from single-target attribute control to multi-target attribute control, enables the attribute prompt vectors to be better fused in terms of attributes and associated knowledge based on the aspect fusion vector, avoids interference between attribute vectors, and then can generate controllable texts based on the fusion result, obtain controllable texts containing multiple attributes, and also improves the practicability and flexibility of controllable text generation.
[0040] To further illustrate the specific implementation process of the method in this embodiment, this embodiment provides a specific method as shown in Figure 2 below, and this method includes:
[0041] Step 201, obtain the prompt text for generating the target controllable text.
[0042] In the embodiment of the present application, it is necessary to guide the pre-trained language model to generate a text that simultaneously satisfies multiple specific attributes based on the prompt text. Exemplarily, if the given prompt text is X1:t = {X1, X2,..., Xt}, and a set of target attributes corresponding to the prompt text is Ta1, Ta2,..., Tak, then the embodiment of the present application can generate a continuation text Xt+1:n = {Xt+1, Xt+2,..., Xn} based on the prompt text, so that the entire text conforms to all target attributes at the same time. Specifically, the formal definition of the task of generating the target controllable text can be in the form of Formula 1 as follows, and Formula 1 is specifically as follows:
[0043] P(X|T a1 ,T a2 ,...,T ak ) = M(X 1+t:n |X 1:t ,T a1 ,T a2 ,...,T ak )(Formula 1)
[0044] In Formula 1, M represents the pre-trained language model, Ta1, Ta2,..., Tak represent k different target attributes, and the task is to generate a text that simultaneously has multiple target attributes.
[0045] Step 202: Input the prompt text into the pre-trained language model, parse the prompt text in the pre-trained language model to obtain multiple attribute information included in the prompt text, and generate multiple attribute prompt vectors based on the multiple attribute information.
[0046] Optionally, the training process of the pre-trained language model includes: obtaining a first sample set containing multiple single-attribute prompt vectors, a second sample set corresponding to the first sample set containing multiple multi-attribute prompt vectors, and a prompt text sample set; based on the first sample set, the second sample set, and the prompt text sample set, performing attribute fusion training, attribute information verification training, and associated knowledge fusion training on the model to be trained to obtain the pre-trained language model.
[0047] In the embodiments of the present application, the single-attribute prompt vector can be a prompt vector containing only one type of attribute. Correspondingly, the multi-attribute prompt vector can be a prompt vector that fuses multiple attributes; the first sample set in the embodiments of the present application can be a sample set containing multiple single-attribute prompt vectors; correspondingly, the second sample set can be a sample set containing multiple multi-attribute prompt vectors obtained by fusing the attributes in the first sample set.
[0048] Exemplarily, if the first sample set contains attribute A, attribute B, attribute C, and attribute D; then the second sample set can contain attribute AB, attribute ABC, attribute BC, attribute BD, and so on, and so forth. No further examples will be given here.
[0049] Optionally, the attribute fusion training process in the training process of the pre-trained language model includes: based on the prompt text sample set and the aspect fusion vector, performing vector attribute fusion on the multiple single-attribute prompt vectors included in the first sample set to obtain a set of multi-attribute fusion vectors to be verified after fusion; calculating the fusion error between the set of multi-attribute fusion vectors to be verified and the second sample set, and when the fusion error reaches a predetermined error threshold, performing attribute information verification training on the model to be verified.
[0050] In the embodiments of the present application, as Figure 3 shown, the left and right elliptical regions in the figure can respectively represent the probability distributions of one attribute and another attribute generating text under the guidance of the aspect prompt vector. For example, specifically, it can be the probability distributions of attribute A and attribute B generating text under the guidance of the aspect prompt vector. Further, when these attribute vectors overlap, the formed region can be mapped to generate text that satisfies multiple attributes. The middle circular region marks the probability distribution of the text directly generated using two single-attribute prompt vectors. The distance from the middle circular region to the overlapping region ( Figure 3(indicated by a dashed line in the figure), which can be regarded as the error generated during the attribute fusion process. The aspect fusion vector proposed in this embodiment is to make the probability distribution of the fused prompt vector coincide with the overlapping region as much as possible by means of guiding the fusion of the prompt vectors.
[0051] In some examples, due to the complexity of the deep neural network and the high-dimensional attribute space, it is difficult to accurately define the error between the fused prompt vector and the prompt vector representing multiple attributes in the ideal state. In the embodiments of this application, the fusion error is defined by the difference in the hidden states of the pre-trained language model under the guidance of different prompt vectors. Specifically, as shown in Formula 2, Formula 2 shows that if the text generated under the guidance of the multi-attribute continuous prompt vector obtained by fusion performs better in various indicators, the error generated during fusion is lower. In the same attribute space, assume that the prompt vector representing multiple attributes in the ideal state is P ideal , and the prompt vector obtained by fusing multiple single-attribute prompt vectors is P fused , the prompt text is x, and the output of the last Transformer layer of the pre-trained language model is defined as O(x, P), and the fusion error (Fusion Loss, FL) is obtained through Formula 2. Formula 2 is specifically as follows:
[0052] FL≈||O(x,P ideal )-O(x,P fused )|| (Formula 2)
[0053] It should be noted that the predetermined error threshold can be set according to requirements. In the embodiments of this application, the specific value of the predetermined error threshold is not limited.
[0054] Optionally, the attribute information verification training process during the training of the pre-trained language model includes: determining the text constraint information corresponding to the prompt text sample set; inputting the constraint information into the first pre-trained model, and performing embedding processing on the prompt text sample set and the text constraint information in the first pre-trained model to obtain the word sequence matrix corresponding to the prompt text sample set; splicing the word sequence matrix with the aspect fusion model to obtain the comprehensive embedding information of the prompt text and the aspect fusion vector. When it is determined that the comprehensive embedding information meets the predetermined attribute information verification conditions, perform associated knowledge fusion training on the model to be verified based on few-shot learning.
[0055] In the embodiments of the present application, the training of the aspect fusion vector is still based on prompt fine-tuning. However, different from the single-attribute-based prompt vector, the aspect fusion vector is trained on data with multiple attributes. Therefore, the embodiments introduce constraints in text form and concatenate them, aiming to prompt the specific attributes of the subsequent training text of the vector and the aspects to which the attributes belong, so that the aspect fusion vector can better capture the relationships between aspects. It should be noted that if the attention mechanism of the pre-trained language model adopted in this embodiment is unidirectional, for example, specifically it can be GPT-2, then it is necessary to maintain the input order of the constraint text, the training text, and the aspect prompt vector. Specifically, first initialize a continuous prompt Pasp with a length of l, and the formal representation is shown in Formula III.
[0056]
[0057] In Formula III, d emb is the word embedding dimension of the pre-trained language model.
[0058] For this embodiment, after the prompt text and the constraints in text form are text-embedded, a word sequence matrix is obtained Then it is concatenated with the aspect fusion vector as the input H, as shown in Formula IV, where the function E(·) represents the text embedding operation.
[0059] H = [E(c1); E(x); P asp (Formula IV)
[0060] In Formula IV, H represents the comprehensive embedding of the text and the fusion vector.
[0061] Optionally, the associated knowledge fusion training process in the training process of the pre-trained language model includes: performing linear transformation processing on multiple attribute prompt vectors and aspect fusion vectors; determining the similarity scores between the transformed aspect fusion vectors and the mapped single-attribute prompt vectors based on a predetermined attention mechanism, and performing weighted pooling processing on the mapped single-attribute prompt vectors based on the similarity scores to obtain a weighted fusion vector; training the model to be trained based on the weighted fusion vector to obtain a pre-trained language model.
[0062] As an optional method, in order to learn the associated knowledge between multiple single-attribute prompt vectors in the fusion stage and thus reduce the fusion error, this embodiment introduces an attention mechanism to perform a weighted pooling operation on multiple single-attribute prompt vectors. Specifically, given m single-attribute prompt vectors attr 1 , attr 2 , ……, attr m , where each Represents a specific attribute. In this embodiment, each single-attribute prompt vector is mapped to a new space through a linear transformation layer to make it compatible with the aspect fusion vector. Although the continuous prompt vectors of single attributes are consistent with the aspect fusion vector in dimension, the continuous prompt vectors representing different attributes may be highly correlated in some dimensions. Through linear transformation, the language model can learn a decorrelated representation, which helps to reduce the redundancy between features. The above process is specifically shown in Formula Five as follows:
[0063] q i =W attr ·attr i +b attr (Formula Five)
[0064] In Formula Five, represents the weight matrix, represents the bias term, is the i-th attribute vector after transformation.
[0065] In some examples, in order to further combine the single-attribute prompt vectors, the embodiment of the present application transforms the aspect fusion vector P asp , and obtains the transformed aspect fusion vector P' asp , which is specifically shown in Formula Six as follows:
[0066]
[0067] In Formula Six, and are learnable parameters.
[0068] For this embodiment, for each transformed attribute vector q i , the embodiment of the present application also needs to apply the dot product attention mechanism to calculate the attention score α asp between it and P' i , which is specifically shown in Formula Seven as follows:
[0069]
[0070] As an alternative, the embodiment of the present application also needs to perform a weighted pooling operation on all single-attribute continuous prompt vectors using the calculated attention score α i so that the weighted fusion vector P mult. can effectively capture the mutual relationship between each single-attribute prompt vector, which is specifically shown in Formula Eight as follows:
[0071]
[0072] Step 203: Based on the aspect fusion vector, perform attribute fusion processing on multiple attribute prompt vectors to obtain a fused first multi-attribute prompt vector.
[0073] For the embodiments of the present application, in the pre-trained language model, it is necessary to fuse multiple attribute prompt vectors based on the aspect fusion vector and obtain a fused multi-attribute prompt vector.
[0074] In the embodiments of the present application, the first multi-attribute prompt vector is obtained by performing fusion processing on multiple attribute prompt vectors. For example, if the multiple attribute prompt vectors are attribute prompt vector A and attribute prompt vector B, the first multi-attribute prompt vector can be obtained by fusing attribute prompt vector A and attribute prompt vector B.
[0075] Step 204: Based on few-shot learning, perform associated knowledge fusion on multiple attribute prompt vectors, and perform controllable text generation based on the attribute fusion result and the associated knowledge fusion result.
[0076] For this embodiment, the attention mechanism performs a weighted pooling operation on multiple single-attribute prompt vectors. Specifically, given m single-attribute prompt vectors attr 1 , attr 2 , ……, attr m , where each represents a specific attribute. In this embodiment, each single-attribute prompt vector is mapped to a new space through a linear transformation layer to make it compatible with the aspect fusion vector.
[0077] Step 205: Determine the generated controllable text as the target controllable text.
[0078] Among them, the target controllable text has the same attributes as the prompt text.
[0079] It should be noted that the embodiments of this application disclose a multi-attribute controllable text generation method based on prompt fusion and few-shot learning. First, the training methods of prompt fusion technology and few-shot learning are used to ensure that the model can fully consider and fuse multiple attribute information when generating text, and support learning the key laws and attribute features of text generation even with only a small number of training samples. This method fuses multi-attribute prompt vectors with the overlapping region guiding to the attribute probability space to achieve a more semantically deep prompt learning method. Then, the method realizes the fine-tuning of the fusion vector training and the prompt vector fusion stage, introduces constraints in text form and concatenates them with data of multiple attributes, so that the fusion vector can better capture the relationships between aspects. Finally, experiments are designed, and controllable text generation tasks with two attributes and three attributes are designed, and the experimental results are analyzed by combining automatic evaluation metrics and human evaluation. This multi-attribute controllable text generation method based on prompt fusion and few-shot learning realizes a smooth transition from single-attribute control to multi-attribute control, effectively addresses the problem of the scarcity of large-scale multi-attribute label data, and at the same time maintains the high efficiency of continuous prompt parameters and the convenience of plug-and-play. In the field of controllable text generation in natural language processing, this method demonstrates technological innovation and breakthroughs in methodology, opening up a new perspective and path for future multi-attribute text generation research.
[0080] In some examples, the embodiments of this application achieve the simultaneous control of multiple target attributes by fusing multiple single-attribute continuous prompt vectors, so as to be able to generate text with complex and diverse attributes. Traditional continuous prompt technologies often can only control a single attribute and are difficult to meet the needs of multi-attribute text generation. By introducing aspect fusion vectors, the embodiments of this application can reduce the interference between attributes and improve the quality of the generated text when fusing multiple single-attribute prompt vectors. The embodiments of this application retain the high-efficiency feature of single-attribute prompt vector parameters. In the training stage of the aspect fusion vector, the number of parameters required is only 0.15% of the GPT-2 model, significantly reducing the parameter dependence and having a smaller demand for computing resources compared to other methods. In the training stage of the aspect fusion vector, only the training text with single attributes is used, and only a small amount of multi-attribute training text is relied on for fine-tuning in the prompt vector fusion stage, greatly reducing the difficulty and cost of obtaining training materials.
[0081] The embodiments of this application mainly relate to the technical field of natural language processing (NLP), especially in the sub-field of controllable text generation with multiple attributes. In addition, it uses a few-shot learning training method to weight and fuse multiple target single-attribute continuous prompts, overcoming the problem of the lack of training corpus with multi-target attributes. By introducing aspect fusion vectors and attention mechanisms, combined with few-shot learning, the multi-attribute prompt vectors obtained through training are used to guide the pre-trained language model to generate text, thus achieving a smooth transition from single-attribute to multi-attribute control, overcoming the challenge of the scarcity of large-scale multi-attribute label data, while maintaining the characteristics of high parameter efficiency and plug-and-play of continuous prompts. This work has achieved technological innovation and methodological breakthroughs in the direction of controllable text generation in the field of natural language processing, providing new ideas and directions for future multi-attribute text generation research. By explicitly controlling specific attributes of the generated text, controllable text generation not only improves the flexibility and richness of text generation, but also greatly expands the dimension of human-computer interaction. Currently, ultra-large language models represented by GPT-4 have achieved remarkable achievements in various tasks in the field of natural language processing, capable of generating more natural, accurate, and in-depth text, providing users with a more human-like communication experience.
[0082] It should be noted that compared with manually crafted prompt templates, the method based on continuous prompts in the embodiments of this application is simpler and can make more effective use of the rich knowledge accumulated by the pre-trained language model during its pre-training process. The embodiments of this application can also improve the flexibility and diversity of the language model. Compared with discrete text prompts that rely on preset templates, continuous prompts provide a more flexible input processing method for the language model, enabling the model to directly extract and understand complex instructions from a high-dimensional space, so as to generate text that meets user needs more naturally without being limited by a predefined structure or format. And compared with fine-tuning the entire pre-trained language model, the method based on continuous prompts further reduces the demand for computing resources. At the same time, it retains the plug-and-play function, enabling the directly use of the trained continuous prompt vectors to guide the unmodified original pre-trained language model to generate text during the inference stage, which is easy to integrate and expand, and its practicality has been significantly improved.
[0083] The embodiments of this application approach from the perspective of prompt fusion, avoiding the problem of multiple prompt vectors interfering with each other during the text generation stage. By introducing a pre-trained aspect fusion vector, it provides guidance during the pooling stage of multiple single-attribute continuous vectors, assisting the fused multi-attribute continuous prompt vector to tend towards the overlapping part of multiple attributes in the attribute probability space. During the prompt fusion stage, the embodiments of this application first map the single-attribute vector to a new space through a linear transformation layer, aiming to improve the compatibility of the attribute vector and reduce its redundancy. Then, the dot product attention mechanism is used to calculate the similarity score between the single-attribute continuous prompt vector and the aspect fusion vector, and based on this score, the multi-attribute fusion prompt vector is further weighted and synthesized. In addition, this method is trained using few-shot learning, overcoming the challenge of the lack of multi-attribute labeled text in this field. The successfully trained multi-attribute continuous prompt vector not only maintains the characteristic of high parameter efficiency but also demonstrates good scalability, providing a new perspective for efficient text generation.
[0084] The embodiments of this application also make full use of the text segmentation method to expand the length of the text sequence that can be processed, and at the same time use the cyclic Transformer structure to solve the problem of semantic fragmentation of long documents caused by the previous single segmentation method, so as to improve the ability to model the overall long document representation. At the same time, it can be applied to a variety of natural language processing tasks, and excellent performance has been achieved, providing support for modeling user satisfaction for accurate complaint prediction.
[0085] Compared with the current existing technologies, in this embodiment, by inputting the prompt text into the pre-trained language model, attribute fusion of multiple attribute prompt vectors included in the prompt text is performed in the pre-trained language model based on the aspect fusion vector, and associated knowledge fusion of multiple attribute prompt vectors is performed based on few-shot learning. Based on the attribute fusion result and the associated knowledge fusion result, controllable text generation is performed, enabling the pre-trained language model in this embodiment to effectively utilize limited data resources and still be able to generate text that meets requirements in the absence of sufficient multi-attribute training corpus. This embodiment also alleviates the attribute interference problem that appears after simply combining single-attribute continuous prompt vectors, realizes a smooth transition from single-target attribute control to multi-target attribute control, enables the attribute prompt vector to be better fused in terms of both attributes and associated knowledge based on the aspect fusion vector, avoids interference between attribute vectors, and thus can perform controllable text generation based on the fusion result to obtain controllable text containing multiple attributes, and also improves the practicality and flexibility of controllable text generation.
[0086] Further, as Figure 1 and Figure 2 a specific implementation of the method shown, this embodiment provides a controllable text generation device, as Figure 4As shown in the figure, the device includes: an acquisition module 31, an input module 32, and a determination module 33.
[0087] The acquisition module 31 is configured to acquire a prompt text for generating a target controllable text.
[0088] The input module 32 is configured to input the prompt text into a pre-trained language model, perform attribute fusion on multiple attribute prompt vectors included in the prompt text based on an aspect fusion vector in the pre-trained language model, and perform associated knowledge fusion on the multiple attribute prompt vectors based on few-shot learning, and generate a controllable text based on the attribute fusion result and the associated knowledge fusion result.
[0089] The determination module 33 is configured to determine the generated controllable text as the target controllable text, and the target controllable text has the same attributes as the prompt text.
[0090] In some examples of this embodiment, a training module is further included. The training module is specifically configured to acquire a first sample set including multiple single-attribute prompt vectors, a second sample set corresponding to the first sample set including multiple multi-attribute prompt vectors, and a prompt text sample set; perform attribute fusion training, attribute information verification training, and associated knowledge fusion training on the model to be trained based on the first sample set, the second sample set, and the prompt text sample set to obtain the pre-trained language model.
[0091] In some examples of this embodiment, the training module is specifically further configured to perform vector attribute fusion on the multiple single-attribute prompt vectors included in the first sample set based on the prompt text sample set and the aspect fusion vector to obtain a set of multi-attribute fusion vectors to be verified after fusion; calculate the fusion error between the set of multi-attribute fusion vectors to be verified and the second sample set, and perform attribute information verification training on the model to be verified when the fusion error reaches a predetermined error threshold.
[0092] In some examples of this embodiment, the training module is specifically further configured to determine the text constraint information corresponding to the prompt text sample set; input the constraint information into the first pre-trained model, perform embedding processing on the prompt text sample set and the text constraint information in the first pre-trained model to obtain a word sequence matrix corresponding to the prompt text sample set; splice the word sequence matrix with the aspect fusion model to obtain comprehensive embedding information of the prompt text and the aspect fusion vector, and perform associated knowledge fusion training based on few-shot learning on the model to be verified when it is determined that the comprehensive embedding information meets the predetermined attribute information verification condition.
[0093] In some examples of this embodiment, the training module is further specifically configured to perform a linear transformation process on the multiple attribute prompt vectors and the aspect fusion vector; determine similarity scores between the transformed aspect fusion vector and the mapped single-attribute prompt vectors respectively based on a predetermined attention mechanism, and perform weighted pooling processing on the mapped single-attribute prompt vectors based on the similarity scores to obtain a weighted fusion vector; and train the model to be trained based on the weighted fusion vector to obtain the pre-trained language model.
[0094] In some examples of this embodiment, the input module 32 is further configured to parse the prompt text in the pre-trained language model to obtain multiple attribute information included in the prompt text, and generate the multiple attribute prompt vectors based on the multiple attribute information; correspondingly, the input module 32 is specifically configured to perform attribute fusion processing on the multiple attribute prompt vectors based on the aspect fusion vector to obtain a first multi-attribute prompt vector after fusion.
[0095] In some examples of this embodiment, the input module 32 is further specifically configured to perform a linear transformation process on the multiple attribute prompt vectors and the aspect fusion vector; determine target similarity scores between the transformed aspect fusion vector and the transformed multiple attribute prompt vectors respectively; and perform weighted pooling processing on the mapped multiple attribute prompt vectors based on the target similarity scores to obtain a second multi-attribute prompt vector.
[0096] In some examples of this embodiment, the input module 32 is further specifically configured to determine a target multi-attribute prompt vector based on the first multi-attribute prompt vector and the second multi-attribute prompt vector; and generate controllable text according to the multi-attribute prompt vector.
[0097] It should be noted that other corresponding descriptions of each functional unit involved in the controllable text generation device provided in this embodiment can be referred to Figure 1 and Figure 2 for the corresponding descriptions therein, which will not be elaborated here.
[0098] Based on the method as shown in Figures 1 to 2 above, correspondingly, this embodiment further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method as shown in Figures 1 to 2 above is implemented.
[0099] Based on such understanding, the technical solution of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, mobile hard disk, etc.), and includes several instructions for causing a computer device (such as a personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of the present application.
[0100] Based on the above-mentioned Figures 1 to 2 shown method, and Figure 4 the virtual device embodiment shown, in order to achieve the above object, the embodiment of the present application further provides an electronic device, such as intelligent terminals such as personal computers, servers, laptops, smart phones, and intelligent robots. The device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to implement the above-mentioned Figures 1 to 2 shown method.
[0101] Optionally, the above-mentioned physical device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, etc. The user interface may include a display screen (Display), an input unit such as a keyboard (Keyboard), etc. Optionally, the user interface may further include a USB interface, a card reader interface, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface), etc.
[0102] Those skilled in the art can understand that the above-mentioned physical device structure provided by this embodiment does not limit the physical device, and it may include more or fewer components, or combine some components, or have different component arrangements.
[0103] The storage medium may further include an operating system and a network communication module. The operating system is a program for managing the hardware and software resources of the above-mentioned physical device, and supports the operation of information processing programs and other software and / or programs. The network communication module is used to implement communication between components inside the storage medium, and communication between other hardware and software in the information processing physical device.
[0104] Based on the above-mentioned Figures 1 to 2 shown method, the embodiment of the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the above-mentioned Figures 1 to 2 shown method. The method implemented when the computer program is executed by the processor can refer to various embodiments of the present application, which will not be elaborated here.
[0105] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform, or can also be implemented by hardware. Compared with the current existing technologies, in this embodiment, by inputting the prompt text into the pre-trained language model, performing attribute fusion on multiple attribute prompt vectors included in the prompt text based on the aspect fusion vector in the pre-trained language model, and performing associated knowledge fusion on the multiple attribute prompt vectors based on few-shot learning, and performing controllable text generation based on the attribute fusion result and the associated knowledge fusion result, the pre-trained language model in this embodiment can effectively utilize limited data resources, and can still generate texts that meet the requirements even in the case of lack of sufficient multi-attribute training corpora; this embodiment also alleviates the attribute interference problem that occurs after simply combining single-attribute continuous prompt vectors, realizes a smooth transition from single-target attribute control to multi-target attribute control, enables the attribute prompt vectors to be better fused in terms of attributes and associated knowledge on the basis of the aspect fusion vector, avoids interference between attribute vectors, and then can perform controllable text generation based on the fusion result to obtain controllable texts containing multiple attributes, and also improves the practicability and flexibility of controllable text generation.
[0106] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including the said element.
[0107] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments described herein, but will be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A controllable text generation method, characterized in that: include: Get the hint text used to generate the target controllable text; Inputting the prompt text into a pre-trained language model, performing attribute fusion on multiple attribute prompt vectors contained in the prompt text based on aspect fusion vectors in the pre-trained language model, performing associated knowledge fusion on the multiple attribute prompt vectors based on small sample learning, and generating controllable text based on attribute fusion results and associated knowledge fusion results; The generated controllable text is determined as the target controllable text, and the target controllable text has the same attributes as the prompt text.
2. The method according to claim 1, characterized in that The training process of the pre-trained language model includes: Acquire a first sample set including a plurality of single-attribute prompt vectors, a second sample set including a plurality of multi-attribute prompt vectors corresponding to the first sample set, and a prompt text sample set; Based on the first sample set, the second sample set and the prompt text sample set, attribute fusion training, attribute information verification training and associated knowledge fusion training are performed on the model to be trained to obtain the pre-trained language model.
3. The method according to claim 2, characterized in that The method of performing attribute fusion training, attribute information verification training and associated knowledge fusion training on the model to be trained based on the first sample set, the second sample set and the prompt text sample set to obtain the pre-trained language model includes: Based on the prompt text sample set and the aspect fusion vector, performing vector attribute fusion on the multiple single-attribute prompt vectors included in the first sample set to obtain a fused multi-attribute fusion vector set to be verified; A fusion error between the multi-attribute fusion vector set to be verified and the second sample set is calculated, and when the fusion error reaches a predetermined error threshold, attribute information verification training is performed on the model to be verified.
4. The method according to claim 3, characterized in that The method further comprises: performing attribute fusion training, attribute information verification training and associated knowledge fusion training on the model to be trained based on the first sample set, the second sample set and the prompt text sample set to obtain the pre-trained language model; Determining text constraint information corresponding to the prompt text sample set; Inputting the constraint information into the first pre-training model, embedding the prompt text sample set and the text constraint information in the first pre-training model, and obtaining a word sequence matrix corresponding to the prompt text sample set; The word sequence matrix is concatenated with the aspect fusion model to obtain comprehensive embedding information of the prompt text and the aspect fusion vector. When it is determined that the comprehensive embedding information meets the predetermined attribute information verification condition, the model to be verified is trained on associated knowledge fusion based on small sample learning.
5. The method according to claim 4, characterized in that The method further comprises: performing attribute fusion training, attribute information verification training and associated knowledge fusion training on the model to be trained based on the first sample set, the second sample set and the prompt text sample set to obtain the pre-trained language model; Performing linear transformation processing on the multiple attribute prompt vectors and the aspect fusion vector; Determine the similarity scores between the converted aspect fusion vectors and the mapped single attribute prompt vectors based on a predetermined attention mechanism, and perform weighted pooling processing on the mapped single attribute prompt vectors based on the similarity scores to obtain a weighted fusion vector; The model to be trained is trained based on the weighted fusion vector to obtain the pre-trained language model.
6. The method according to claim 1, characterized in that Before performing attribute fusion on a plurality of attribute prompt vectors contained in the prompt text based on the aspect fusion vector in the pre-trained language model, the method further includes: Parsing the prompt text in the pre-trained language model to obtain a plurality of attribute information contained in the prompt text, and generating the plurality of attribute prompt vectors based on the plurality of attribute information; The performing attribute fusion on a plurality of attribute prompt vectors contained in the prompt text based on the aspect fusion vector in the pre-trained language model comprises: The attribute fusion processing is performed on the multiple attribute prompt vectors based on the aspect fusion vector to obtain a fused first multi-attribute prompt vector.
7. The method according to claim 6, characterized in that The performing associated knowledge fusion on the multiple attribute prompt vectors based on small sample learning includes: Performing linear transformation processing on the multiple attribute prompt vectors and the aspect fusion vector; determining target similarity scores between the transformed aspect fusion vector and the transformed multiple attribute prompt vectors respectively; A weighted pooling process is performed on the mapped multiple attribute prompt vectors based on the target similarity score to obtain a second multiple attribute prompt vector.
8. The method according to claim 7, characterized in that Controllable text generation based on attribute fusion results and associated knowledge fusion results includes: determining a target multi-attribute prompt vector based on the first multi-attribute prompt vector and the second multi-attribute prompt vector; Controllable text generation is performed based on the multi-attribute prompt vector.
9. A controllable text generation method, characterized in that: include: An acquisition module, configured to acquire a prompt text for generating a target controllable text; An input module is configured to input the prompt text into a pre-trained language model, perform attribute fusion on multiple attribute prompt vectors contained in the prompt text based on aspect fusion vectors in the pre-trained language model, perform associated knowledge fusion on the multiple attribute prompt vectors based on small sample learning, and generate controllable text based on the attribute fusion result and the associated knowledge fusion result; The determination module is configured to determine the generated controllable text as the target controllable text, and the target controllable text has the same attributes as the prompt text.
10. An electronic device comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.