Attribute-controllable text generation method based on prefix reinforcement control
By introducing a prefix reinforcement control strategy into attribute-controllable text generation, and dynamically adjusting attribute distribution and attention mechanisms, the problem of unstable target attribute control in long text generation is solved, and the semantic consistency of text generation and the stability of attribute expression are improved.
Patent Information
- Application Number
- CN202510842675.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-12-02
AI Technical Summary
Existing attribute-controlled text generation methods suffer from unstable target attribute control when generating long texts, resulting in diminished attribute expressiveness and difficulty in maintaining semantic consistency and attribute control as the length of generated text increases.
We adopt an attribute-controlled text generation method based on prefix reinforcement control. By using a basic language model and a target attribute adjustment model, and leveraging a conditional language model and prefix control strategy, we dynamically adjust the attribute distribution and attention mechanism during the text generation process. We insert prefix vectors with fixed time steps to maintain the target attribute control effect.
While maintaining language fluency, it improves the stability and expressive precision of target attributes, making it suitable for long text and multi-attribute generation tasks.
Smart Images

Figure CN121052253A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to natural language processing technology, specifically to attribute-controllable text generation technology. Background Technology
[0002] Controllable Text Generation (CTG) is an important research direction in the field of natural language generation. Its goal is to control certain predefined attributes of the output content during text generation, such as sentiment, topic, or style and tone. With the widespread development of social media, news generation, and automated marketing, the demand for natural language generation technologies capable of expressing predefined attributes is constantly growing. These predefined attributes are called target attributes. Applications such as comment generation, personalized copywriting, and emotion-guided responses all require controllable target attributes as their underlying support.
[0003] To achieve effective control over target attributes during text generation, various methods have been proposed, including training language models from scratch, fine-tuning some model parameters, and post-processing methods that adjust the output distribution of the language model during generation. While these methods have made progress in attribute representation, they also face significant challenges. Firstly, existing PLMs primarily learn the statistical regularities of natural text during unsupervised training, without explicitly modeling the attribute information required for controllable generation. Directly applying standard decoding strategies often fails to guarantee that the generated text strictly conforms to the target attributes. Secondly, although some current methods have achieved good results in controlling target attributes, as the length of the generated text gradually increases, problems such as high computational cost or decreased generation quality often arise. Attribute information is easily diluted or interfered with, causing the final generated text to deviate from the target attributes.
[0004] The main reasons for this phenomenon include: firstly, as the sequence length increases, the attention mechanism needs to allocate more attention points at each time step, leading to a dilution of attention to key attribute words; secondly, text generation is a word-by-word construction process, and small deviations in the early stages may accumulate in subsequent steps, eventually causing semantic drift or inconsistencies in the generated content. To alleviate these problems, an attribute control mechanism with dynamic adjustment capabilities, strong scalability, and stable generation quality is needed to adapt to the control requirements of long text and multi-attribute generation tasks. Summary of the Invention
[0005] The technical problem to be solved by this invention is to provide an attribute-controlled text generation scheme that can maintain stable target attribute control, addressing the issues of unstable target attribute control and diminished attribute expressiveness in existing attribute-controlled text generation methods when dealing with long text generation.
[0006] The technical solution adopted by this invention to solve the above-mentioned technical problems is an attribute-controllable text generation method based on prefix enhancement control, comprising the following steps:
[0007] The base language model (baseLM) receives the input cue sequence and outputs the predicted probability distribution of the original text.
[0008] The target attribute adjustment model CC-LMs receives the input cue sequence and target attribute. The class conditional language model CC-LM corresponding to each attribute in CC-LMs performs context encoding based on the input cue sequence and outputs the CC-LM output probability distribution corresponding to each attribute. During the attention calculation process, a prefix control strategy is used to determine whether to insert a prefix corresponding to the attribute at fixed time steps to maintain the target attribute control effect.
[0009] The target attribute adjustment model CC-LMs divides the output probability distribution of the CC-LM corresponding to the target attribute by the sum of the output probability distributions of the CC-LM corresponding to all attributes to obtain the attribute distribution of the target attribute output by the target attribute adjustment model.
[0010] The probability distribution of the original text is adjusted by utilizing the attribute distribution of the target attribute, thereby increasing the probability of words related to the target attribute, and finally outputting subsequent text content that satisfies the target attribute.
[0011] Furthermore, when adjusting the original text probability distribution using the attribute distribution of the target attribute, the control intensity is adjusted by the control coefficient ω.
[0012] Furthermore, before calculating the attribute distribution of the target attribute, the output probability distribution of all CC-LMs is reconstructed to ensure text fluency.
[0013] Specifically, the target attribute is a single attribute or a combination of multiple attributes.
[0014] When the target attribute is a combination of multiple attributes, the prefix parameter θ of the CC-LM corresponding to the multiple attributes multi The prefix of the CC-LM for all the individual attributes contained in the multi-attribute is obtained by linearly combining the parameters of the prefix of the CC-LM for all the individual attributes contained in the multi-attribute; the prefix of the CC-LM for the multi-attribute is obtained by linearly combining the prefix of the CC-LM for all the individual attributes contained in the multi-attribute.
[0015] Specifically, the method for inputting the prefix corresponding to the attribute during the attention calculation process is as follows: In the self-attention module of each layer of CC-LM, the prefix is inserted as an insertion column and concatenated to the fixed time step of the key matrix and the value matrix respectively.
[0016] The beneficial effect of this invention is that by introducing a lightweight prefix learning mechanism and a reinforcement learning-driven dynamic control strategy, it effectively improves the stability of control over target attributes and the accuracy of expression while maintaining language fluency. Attached Figure Description
[0017] Figure 1 Schematic diagram of the model framework Detailed Implementation
[0018] The system implementing the attribute-controlled text generation method includes a base language model (base LM) and target attribute adjustment models (CC-LMs). The target attribute adjustment models consist of several class-conditional language models (CC-LMs), a prefix control strategy module, and an attribute distribution reconstruction module. Assume that the attribute set A consists of k attributes, A = {a1, a2, ..., a...}. i ,...a k Given a clue sequence X 1:T-1 =(x1,x2,...,x T-1 ), and a target attribute a i We require the system to generate a value that satisfies the target attribute a. i Follow-up content X T:N =(x T ,x T+1 ,...,x N N represents the total number of tokens to be generated, T-1 represents the number of tokens already generated, and T:N represents T to N. Each attribute corresponds to a CC-LM.
[0019] Target attribute a i It can be a single emotional attribute, a combination of a single emotional attribute and a single thematic attribute, or a combination of multiple non-mutually exclusive emotional attributes and multiple thematic attributes.
[0020] The base LM outputs the original text probability distribution P(x). t |x <t The target attribute is adjusted by adjusting the model output target attribute a. i The attribute distribution P(a) i |x 1:t ), where x t The text information generated for time step t, x <t x represents the text information generated before time step t. 1:t Let P(A|B) represent the text information generated from time step 1 to t, and let P(A|B) represent the probability of A given B.
[0021] The system is based on the attribute distribution P(a)i |x 1:t ) for the original text probability distribution P(x t |x <t Adjustments are made to increase the probability of words related to the target attribute, i.e., by combining the original text probability distribution P(x) with the target attribute probability distribution P(x). t |x <t ) and attribute distribution P(a i |x 1:t The control strength is adjusted by the control coefficient ω, and the output is the text probability distribution P(x) adjusted by the attribute distribution. t |x <t ,a i ), for example P(x t |x <t ,a i )=P(a i |x 1:t ) ω P(x t |x <t P(A|B,C) represents the probability of A given B and C. Finally, the system determines the probability of the token in the segmentation unit corresponding to the current time step t. P(x t |x <t ,a i Select the generated text information x for the current time step t. t Ultimately, the system outputs the content X that satisfies the target attribute. T:N =(x T ,x T+1 ,...,x N ),
[0022] CC-LM is based on the language model used by base LM, and inserts the target attribute 'a' according to the pre-control strategy during context encoding. i The corresponding prefix vector is used to guide the CC-LM in text generation, increasing the output probability values of attribute-related words and suppressing the output probability values of attribute-irrelevant words. The language model parameters used in the base LM can be obtained using currently available training methods. This invention focuses on the training of the CC-LM in the system. Optionally, the base LM uses a language model that has already been trained on large-scale general data, i.e., a pre-trained language model (PLM). The model parameters of the CC-LM are obtained from the frozen parameters θ of the pre-trained language model (PLM). G and learnable attribute a i The corresponding Prefix parameter It consists of two parts. In the CC-LM model, attribute a... i The corresponding prefix vector is represented as:
[0023] The prefix vector (Prefix) for a single attribute (single sentiment attribute or single topic attribute) is obtained through training.
[0024] Specifically, the learning of single-attribute prefix vectors employs a frozen fine-tuning strategy. In the self-attention module of each layer of the pre-trained language model PLM, an additional prefix vector is concatenated between the key K and the value V. During training, only the parameters are updated. Do not update the weights θ of the language model G Let each single attribute a i Each dataset has a corresponding training dataset with attribute labels for a single attribute 'a'. i The prefix is used to select a complete text sequence X = (x1, x2, ..., x...). t ,...,x T The training objective is to minimize the language model loss function for text generation under the target attribute.
[0025]
[0026] After completing the CC-LM training for a single attribute, the single attribute a can be obtained. i The corresponding prefix and parameters in CC-LM
[0027] Specifically, for multi-attribute CC-LMs, separate training is not required. The prefix vector of a multi-attribute CC-LM can be obtained by linearly combining the individual attributes it contains. Similarly, the prefix parameters of a multi-attribute CC-LM can be obtained by linearly combining the prefix parameters of the individual attributes it contains. For example, a multi-attribute CC-LM contains a single sentiment attribute (sentimen) and a single topic attribute (topic), and the prefix parameters θ of this multi-attribute CC-LM... multi This can be achieved by fusing the hyperparameters β1 (which controls the relative influence of emotion) and β2 (which controls the relative influence of topic attributes): θ multi =β1·θ sentiment +β2·θ topic Similarly, the prefix of this multi-attribute CC-LM can also be obtained in the same way: Thus, the set A = {a1, a2, ..., a...} is obtained, consisting of all possible attributes. i ,...a k} and the corresponding CC-LM for all attributes.
[0028] Optionally, those skilled in the art can also obtain multi-attribute prefixes and their corresponding CC-LMs through training.
[0029] Specifically, the target attribute adjusts the output P(a) of the model. i |x 1:t Based on the output probabilities of all CC-LMs in attribute set A, we obtain:
[0030]
[0031] in, For target attribute a i The probability of the token in the segmentation unit corresponding to time step j in the CC-LM output. For any attribute a in attribute set A k The probability distribution is the probability of the token in the segmentation unit corresponding to time step j of the CC-LM output. The above formula uses the mutual exclusivity between different attributes to adjust the output probability distribution. It divides the output probability distribution of the CC-LM corresponding to the target attribute by the sum of the output probability distributions of all CC-LMs, thereby increasing the probability of words related to the target attribute.
[0032] Preferably, in CC-LMs, the attribute distribution reconstruction module, in order to avoid excessive adjustments that could disrupt text fluency, calculates P(a i |x 1:t Previously, to make the attribute distribution of CC-LMs output more stable, the output distribution of each CC-LM was further reconstructed:
[0033]
[0034] Target attribute adjustment model output
[0035] Preferably, the prefix control policy module in CC-LMs is used to determine whether each CC-LM inserts the corresponding prefix at each time step nS based on the reinforcement learning policy of online policy gradient. That is, every S time steps, a check is performed to determine whether an exception should be inserted into the Prefix.
[0036] The prefix control strategy module employs a three-layer nonlinear multilayer perceptron, with each layer having 1024, 1024, and 2 nodes respectively, and ReLU activation functions between layers. It will use the currently generated text information x... 1:nS As input, and at every S time step from the action set {ι + ,ι o Select an action ι to output, ι +This indicates that the corresponding prefix should be inserted at the current time step. o This indicates that no action will be taken.
[0037] Model parameters θ of the prefix control strategy module C The strategy is iteratively updated online to maximize the total expected reward obtainable from the sampling trajectory. Specifically, during text generation, the current strategy... At each time step nS, action i is selected to generate subsequent content. Based on this process, the i-th CC-LM can obtain a set of sampled trajectories τ. i The trajectory is a ternary form consisting of state (the text generated at the nth fixed time step), action (the action selected at the nth fixed time step), and reward (the reward obtained after selecting the action at the nth fixed time step). N i This represents the sequence number of the last token in the sequence. For each sampling trajectory τ i In fact, a corresponding reference trajectory was also sampled. Its reference trajectory τ i' Initial input sequence and τ i Similarly, the prefix control strategy employed is to take ι at each time step nS. o Action, i.e., not inserting a prefix. τ i The return R i From the sampling trajectory τ i and reference trajectory τ i' The two generated texts and The performance metrics differences between the two texts, which correspond to trajectories τ respectively. i and τ i' The goal of the optimal strategy is to maximize the expected total return across all trajectories.
[0038] Specifically, the return R i The calculation method is as follows:
[0039]
[0040] Here, `attr_score(·)` is an external attribute classifier that takes a text sequence as input and outputs a score of attribute relevance. `ppl_score(·)` is based on GPT-2. Medium Calculate the perplexity of the input text. α1 and α2 are hyperparameters that control the relative influence of attribute scores and perplexity.
[0041] The goal of the optimal strategy is to maximize the expected total reward J across all trajectories, where the total reward J is:
[0042]
[0043] in This represents the average return value of the trajectory sampled in the current batch. Indicating in strategy Lower trajectory τ i Expectations;
[0044] Then, the policy parameter θ is optimized using the policy gradient approach. C :
[0045]
[0046] Represents the gradient. This is for rounding down.
[0047] The technical solution of this invention will be clearly and completely described below with reference to the accompanying drawings. The attribute-controllable text generation method based on prefix reinforcement control proposed in this invention consists of two stages and four steps. The first stage includes prefix vector training and attribute distribution adjustment; the second stage includes prefix control strategy training and outputting the generated result.
[0048] Overall framework diagram as follows Figure 1 As shown, in the binary emotion control task, the attribute set A = {negtaive, positive} includes two types of emotions: positive and negative. The two attributes are linearly combined using parameters β1 = β2 = 1. The language model used is a specific version of the GPT-2 series, GPT-2. Medium The prefix length is set to 20, the batch size to 4, the learning rate to 5e-5, and the training epochs to 10. In the policy module training phase of step 3, the prefix insertion interval S = 32, the reward control factors α1 = 10, α2 = 0.1, the training batch size is set to 8, the learning rate to 5e-5, and the training epochs to 10.
[0049] Let the target set be A = {negative, positive}, given a cue sequence X 1:T-1 =(x1,x2,...,x T-1 ), and a target attribute a i a i =negtaive or a i =positive. We require the system to generate subsequent content X that satisfies the target attribute. T:N =(x T ,x T+1 ,...,x N This process can be represented as:
[0050] The following steps will be used to generate subsequent content for the target attribute:
[0051] Step 1: Prefix Vector Training
[0052] For systems that implement attribute-controllable text generation methods, a freeze-fine-tuning strategy is adopted to learn the prefix corresponding to each target attribute.
[0053] For the two CC-LMs in the system, Negative CC-LM and Positive CC-LM, when learning their respective prefixes, only the prefix part (Neg prefix) parameters are updated. Do not update the weights θ of the language model GPT-2 (frozen) G ,
[0054] Suppose that each attribute has a labeled training dataset for the corresponding attribute's prefix, and a complete text sequence X = (x1, x2, ..., x...) is selected from this dataset. T The training objective is to minimize the language model loss function for text generation under the target attribute. After training, the two CC-LMs determine the specific representation of the prefix corresponding to their respective emotion attributes that needs to be inserted into the attention matrix when the self-attention mechanism performs context encoding.
[0055] When the target attribute a i It involves combining multiple attributes. For example, when combining a single emotion attribute with a single theme attribute, the emotion attribute can be either positive or negative, and the theme attribute can be either political, cultural, technological, or military. This results in eight possible combinations of emotion and theme attributes, requiring eight CC-LMs in the CC-LMs. The prefixes corresponding to these eight CC-LMs can be obtained by fusing the prefixes of the single attributes. For example, a... i For a positive cultural theme, a i The corresponding prefixes and CCLM parameters include single-attribute positive sentiment and cultural theme prefixes, respectively. and and the parameter θ of CC-LM positive and θ topic-cultue get: θ multi =β1·θ positive +β2·θ topic-cultue .
[0056] Step 2: Adjusting Attribute Distribution
[0057] Based on the CC-LMs obtained in step 1, the system uses the mutual exclusion between different attributes to adjust the output attribute distribution P(a i |x 1:t For the attribute distribution P(a) in this embodiment, i |x 1:t There are two types: negative emotion distribution and positive emotion distribution.
[0058] According to Bayes' condition formula, we can express P(x) as... t |x <t ,a i )break down:
[0059]
[0060] Where ω represents the added control strength, and P(a) i |x 1:t The outputs of each CC-LM obtained in step 1 can be used to... To avoid disrupting text fluency with excessive adjustments, when calculating P(a)... i |x 1:t Before that, first... Attribute DistributionReconstruction is performed to obtain This includes reconstructing the negative sentiment distribution and the positive sentiment distribution, and then using the reconstructed data... To calculate P(a) i |x 1:t This makes the attribute distribution of CC-LMs output more stable:
[0061]
[0062] Through P(a) i |x 1:t ) ω This allows us to determine the output probability distribution of the raw PLM, P(x). t |x <t ,a i Adjustments were made to increase the probability values of words related to the target attribute and suppress the probability values of words unrelated to the target attribute.
[0063] when a i For positive cultural themes, CC-LMs will utilize the reconstruction results output by CC-LM corresponding to the positive cultural theme. Calculate the numerator of the above equation and use the reconstruction results from the CC-LM output corresponding to each of the eight multi-attribute combinations. Calculate the denominator of the above expression.
[0064] Step 3: Training the prefix control strategy module
[0065] The prefix control policy module in the system employs a policy gradient-based reinforcement learning strategy to determine whether to insert an additional prefix at each time step nS. The set of natural numbers is used. Specifically, the prefix control policy module is a three-layer nonlinear multilayer perceptron, with each layer having 1024, 1024, and 2 nodes respectively, and ReLU activation functions between layers. It will use the currently generated text information x 1:nS As input, the output action1=1 or action2=0 is used to insert the prefix Insert Prefix state, which corresponds to the state from the action set {ι at each time step S. + ,ι o Output an action ι, whose expression is as follows:
[0066]
[0067] Where Policy represents the prefix control policy, θ C These are parameters of the prefix control strategy module, ι + This indicates the action (action1=1) that inserts the Prefix into the corresponding CC-LM at the current time step, while ι o This indicates that no action will be taken (action2 = 0). The action output by the prefix control strategy module is defined as follows:
[0068]
[0069] in These represent the original bond matrices inserted into the i-th CC-LM. Sum matrix The prefix, where i is the attribute index corresponding to the target attribute, is concatenated to obtain the key matrix after inserting the prefix according to the prefix control strategy. Sum matrix The matrix d is the hidden state dimension of CC-LM.
[0070] We train the prefix control policy module using an online policy gradient approach, which iteratively updates θ. C The goal is to maximize the total expected reward obtainable from the sampling trajectory. Specifically, in the text generation process, the current strategy... At each time step nS, action ι is selected to generate subsequent content. Based on this process, a set of sampled trajectories can be obtained, each trajectory being in the form of: For each sampling trajectory τ i In fact, a corresponding trajectory was also sampled. Its initial input sequence and τ i Similarly, at each time step nS, ι will be taken each time. o Action, that is, doing nothing. τ i The return R i From two generated texts and The performance metrics differences between the two texts, which correspond to trajectories τ respectively. i and τ i' The goal of the optimal strategy is to maximize the expected total return across all trajectories.
[0071] Step 4: Generate Results
[0072] After the above training and adjustments, for the current input text, at the current time step t, each CC-LM is guided by the prefix corresponding to the attribute during the attention calculation process. Then, the prefix control policy decides whether to insert an additional prefix into the CC-LMs at each time step nS. Finally, the P(a) constructed by the CC-LMs will be... i |x 1:t ) ω The original output probability distribution P(x) of the language model t |x <t By combining these factors and adjusting the control intensity via ω as needed, the final output probability distribution P(x) is determined based on the target attribute. t |x <t ,a i Ultimately, we select the final output content by random sampling and based on the probability of each token.
[0073] This invention combines prefix learning, weighted decoding, and policy gradient-based reinforcement learning to simultaneously model language coherence and attribute consistency during the generation process. By utilizing a conditional language model to construct attribute probability distributions and dynamically adjusting the timing of control signal insertion, it effectively balances the dual requirements of semantic fluency and attribute control in generated text. This method improves the stability and accuracy of attribute representation in text generation tasks, and is particularly suitable for complex generation scenarios such as long texts and multi-attribute control.
[0074] Although specific illustrative embodiments of the present invention have been described above to facilitate understanding by those skilled in the art, it should be understood that the present invention is not limited to the scope of the specific embodiments. All variations are obvious and any invention utilizing the concept of the present invention is protected.
Claims
1. A method for generating attribute-controllable text based on prefix reinforcement, characterized in that, Includes the following steps: The base language model (baseLM) receives the input cue sequence and outputs the predicted probability distribution of the original text. The target attribute adjustment model CC-LMs receives the input cue sequence and target attribute. The class conditional language model CC-LM corresponding to each attribute in CC-LMs performs context encoding based on the input cue sequence and outputs the CC-LM output probability distribution corresponding to each attribute. During the attention calculation process, a prefix control strategy is used to determine whether to insert a prefix corresponding to the attribute at fixed time steps in order to maintain the control effect of the target attribute. The target attribute adjustment model CC-LMs divides the output probability distribution of the CC-LM corresponding to the target attribute by the sum of the output probability distributions of the CC-LM corresponding to all attributes to obtain the attribute distribution of the target attribute output by the target attribute adjustment model. The probability distribution of the original text is adjusted by utilizing the attribute distribution of the target attribute, thereby increasing the probability of words related to the target attribute, and finally outputting subsequent text content that satisfies the target attribute.
2. The method as described in claim 1, characterized in that, When adjusting the original text probability distribution using the attribute distribution of the target attribute, the control strength is adjusted by the control coefficient ω, and the output is the text probability distribution P(x) adjusted by the attribute distribution. t |x <t ,a i ) is represented as: P(x t |x <t ,a i )∝P(a i |x 1:t ) ω P(x t |x <t ); Wherein, P(x t |x <t Let P(a) be the original text probability distribution. i |x 1:t ) represents the attribute distribution of the target attribute, a i For the target attribute, ω is the control coefficient for adjusting the control intensity, and ∝ represents the direct proportional relationship; x t The text information generated for time step t, x <t x represents the text information generated before time step t. 1:t Let P(A|B) represent the text information generated from time step 1 to t, and let P(A|B) represent the probability of A given B.
3. The method as described in claim 2, characterized in that, Before calculating the attribute distribution of the target attribute, the output probability distribution of all CC-LMs is reconstructed to ensure text fluency.
4. The method as described in claim 3, characterized in that, The method for reconstructing the output probability distribution of each CC-LM is as follows: For any attribute a in attribute set A k The probability distribution corresponding to time step j in the CC-LM output; For the prefix parameter, θ G For language model parameters; For attribute a k The corresponding prefix vector, j is the time step variable, and ln is the natural logarithm; Let P(A|B,C) be the reconstructed probability distribution, representing the probability of A given B and C. Target attribute adjustment model output target attribute distribution 5. The method as described in claim 1, characterized in that, The target attribute is a single attribute or a combination of multiple attributes.
6. The method as described in claim 1, characterized in that, When the target attribute is a combination of multiple attributes, the prefix parameter θ of the CC-LM corresponding to the multiple attributes. multi The prefix of the CC-LM for all the individual attributes contained in the multi-attribute is obtained by linearly combining the parameters of the prefix of the CC-LM for all the individual attributes contained in the multi-attribute; the prefix of the CC-LM for the multi-attribute is obtained by linearly combining the prefix of the CC-LM for all the individual attributes contained in the multi-attribute.
7. The method as described in claim 1, characterized in that, The specific method for inputting the prefix corresponding to the attribute during the attention calculation process is as follows: In the self-attention module of each layer of CC-LM, the prefix is inserted as an insertion column and concatenated to the fixed time step of the key matrix and the value matrix respectively.
8. The method as described in claim 1, characterized in that, In the target attribute adjustment models CC-LMs, each CC-LM learns its respective attribute's prefix using a frozen fine-tuning strategy, updating only the parameters during training. Do not update the language model weights θ G The training objective is to minimize the loss function of the language model for text generation under the target attribute.
9. The method as described in claim 1, characterized in that, The model parameters θ of the prefix control strategy module implementing the prefix control strategy C The online iterative update is optimized to maximize the expected total reward that can be obtained from the sampling trajectory. The sampling trajectory consists of text generated at each fixed time step, the action of choosing whether to insert a prefix, and the reward.