Stylized text generation method and device based on potential modulation, equipment and medium

Through a stylized text generation method based on latent modulation, using temporal coding networks and style conditional projectors for iterative optimization and multi-dimensional style parameter mapping, the problem of low accuracy in stylized text generation is solved, and the accuracy and semantic coherence of stylized text are improved.

CN120706375APending Publication Date: 2025-09-26PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510881705.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing technologies have low accuracy when generating stylized text, make it difficult to achieve multiple rounds of progressive optimization, lack a closed-loop feedback mechanism, have rough style adjustments, have low semantic coupling, and are difficult to dynamically adapt to diverse target needs.

Method used

A stylized text generation method based on latent modulation is adopted. Through the temporal coding network and style conditional projector, variational inference and temporal coding network prediction mechanism are combined to predict the latent space adjustment amount in real time, perform iterative optimization and collaborative mapping of multi-dimensional style parameters, and gradually approach the target style.

Benefits of technology

The accuracy of stylized text generation is improved, semantic coherence is maintained, and it dynamically adapts to diverse target needs, solving the problem of low accuracy in stylized text generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120706375A_ABST
    Figure CN120706375A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of natural language processing, can be applied to business system platforms of financial science and technology, medical health and the like, and discloses a stylized text generation method, device, equipment and medium based on potential modulation. Performing vector conversion on the target style parameter to obtain a potential style vector; inputting the initial potential representation and the potential style vector into a time sequence coding network, and outputting an initial space adjustment amount corresponding to the initial potential representation; performing iterative optimization on the initial potential representation according to the initial space adjustment amount and the potential style vector to obtain a target potential representation; mapping the target potential representation into a preset target style space; and decoding the target potential representation mapped to the target style space to obtain a stylized text corresponding to the input text. And the generation accuracy of the stylized text is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing technology, and in particular to a method, device, equipment and medium for generating stylized text based on latent modulation. Background Art

[0002] In response to the personalized needs of users, stylized text corresponding to the user needs should be input based on the user input content. In order to meet the personalized needs of users, the text generation style needs to be adjusted according to the target audience, and the text, therefore, needs to be accurately analyzed.

[0003] In the field of healthcare, medical records must be rigorously described along a timeline. Medical records that are too colloquial, or patient education materials that use too many professional terms, can make it difficult for patients to understand. Alternatively, the texts are mechanical and lack humanistic care, which may increase patient anxiety. Rule-based generation models have difficulty in handling personalized descriptions of complex conditions and have a single style. As a result, the stylized texts generated for user needs are too one-sided and have low accuracy.

[0004] In the field of financial technology (FinTech), it is often necessary to automatically generate text content that conforms to a specific style based on different user identities, business scenarios, and compliance requirements. This can include highly formal claims instructions, emotionally neutral responses to financial Q&A, or cross-language legal text polishing. However, existing systems often only support a small number of preset style templates and lack flexible adaptation and fine-grained control of user-specified styles. This results in a single generated content style and low accuracy in generating style text tailored to user needs.

[0005] The existing technology generates stylized text through static latent space modulation, which makes it difficult to achieve multiple rounds of progressive optimization, resulting in rough style adjustment and low fidelity. It also only supports a single style injection model, lacks a closed-loop feedback mechanism, has a low degree of coupling between style and semantics, and is difficult to dynamically adapt to diverse target needs, resulting in low accuracy when generating stylized text. Summary of the Invention

[0006] The present invention provides a method, apparatus, device and medium for generating stylized text based on latent modulation, so as to solve the technical problem of low accuracy when generating stylized text.

[0007] In a first aspect, a method for generating stylized text based on latent modulation is provided, comprising:

[0008] Obtaining input text and target style parameters from a target user, converting the input text into an initial latent representation, and performing vector conversion on the target style parameters to obtain a latent style vector;

[0009] Inputting the initial latent representation and the latent style vector into a preset temporal coding network, and outputting an initial spatial adjustment corresponding to the initial latent representation;

[0010] Iteratively optimizing the initial latent representation according to the initial spatial adjustment amount and the latent style vector to obtain a target latent representation;

[0011] Mapping the target latent representation into a preset target style space through a preset style condition projector;

[0012] The target latent representation mapped to the target style space is decoded to obtain a stylized text corresponding to the input text.

[0013] In a second aspect, a stylized text generation device based on latent modulation is provided, comprising:

[0014] A target style parameter conversion module is used to obtain the input text and target style parameters of the target user, convert the input text into an initial latent representation, and perform vector conversion on the target style parameters to obtain a latent style vector;

[0015] an initial spatial adjustment value output module, configured to input the initial latent representation and the latent style vector into a preset temporal coding network, and output an initial spatial adjustment value corresponding to the initial latent representation;

[0016] an initial latent representation iterative optimization module, configured to iteratively optimize the initial latent representation according to the initial spatial adjustment amount and the latent style vector to obtain a target latent representation;

[0017] a target latent representation mapping module, configured to map the target latent representation into a preset target style space through a preset style condition projector;

[0018] The stylized text generation module is used to decode the target latent representation mapped to the target style space to obtain the stylized text corresponding to the input text.

[0019] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the above-mentioned method for generating stylized text based on latent modulation are implemented.

[0020] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned method for generating stylized text based on latent modulation are implemented.

[0021] In the scheme implemented by the above-mentioned method, device, equipment and medium for generating stylized text based on latent modulation, the input text and target style parameters of the target user can be obtained through the client, the input text can be converted into an initial latent representation, the target style parameters can be vectorized to obtain a latent style vector; the initial latent representation and the latent style vector can be input into a preset temporal coding network, and the initial spatial adjustment amount corresponding to the initial latent representation can be output; the initial latent representation can be iteratively optimized according to the initial spatial adjustment amount and the latent style vector to obtain a target latent representation; the target latent representation can be mapped to a preset target style space through a preset style condition projector; the target latent representation mapped to the target style space can be decoded to obtain the input text The corresponding stylized text is fed back to the client. In the present invention, based on the original encoder-decoder framework, a temporal prediction coding module and a staged optimization scheduler are introduced, combined with variational inference and temporal coding network prediction mechanism. The latent space adjustment amount is predicted in real time through the temporal coding network, which can dynamically adjust the optimization direction according to the difference between the current state and the target style, realize closed-loop optimization of the latent space and dynamic control of the style trajectory. At the same time, the style condition projector of the Transformer architecture is used to realize the collaborative mapping and progressive adjustment of multi-dimensional style parameters, so that in the text generation process, the target style is gradually approached and semantic coherence is maintained, thereby solving the technical problem of low accuracy in generating stylized text. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0023] Figure 1 2 is a schematic diagram of an application environment of a method for generating stylized text based on latent modulation according to an embodiment of the present invention;

[0024] Figure 2 1 is a flow chart of a method for generating stylized text based on latent modulation according to an embodiment of the present invention;

[0025] Figure 3 yes Figure 2 A schematic flow chart of a specific implementation of step S1;

[0026] Figure 4 yes Figure 2 A schematic flow chart of a specific implementation of step S5;

[0027] Figure 51 is a schematic structural diagram of a stylized text generation device based on latent modulation according to an embodiment of the present invention;

[0028] Figure 6 is a structural diagram of a computer device in one embodiment of the present invention;

[0029] Figure 7 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0030] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0031] The stylized text generation method based on latent modulation provided by the embodiment of the present invention can be applied in the following fields: Figure 1 In an application environment, the client communicates with the server through a network. The server can obtain the input text and target style parameters of the target user through the client, convert the input text into an initial latent representation, perform vector conversion on the target style parameters, and obtain a latent style vector; input the initial latent representation and the latent style vector into a preset temporal coding network, and output the initial spatial adjustment amount corresponding to the initial latent representation; iteratively optimize the initial latent representation according to the initial spatial adjustment amount and the latent style vector to obtain a target latent representation; map the target latent representation to a preset target style space through a preset style condition projector; decode the target latent representation mapped to the target style space to obtain a stylized text corresponding to the input text, and descaling the stylized text Feedback to the client, in the present invention, based on the original encoder-decoder framework, a temporal prediction coding module and a staged optimization scheduler are introduced, combined with variational inference and temporal coding network prediction mechanism, and the temporal coding network predicts the potential space adjustment amount in real time, which can dynamically adjust the optimization direction according to the difference between the current state and the target style, realize closed-loop optimization of the potential space and dynamic control of the style trajectory, and at the same time, utilize the style condition projector of the Transformer architecture to realize the collaborative mapping and progressive adjustment of multi-dimensional style parameters, so that in the text generation process, the target style is gradually approached and semantic coherence is maintained, thereby solving the technical problem of low accuracy when generating stylized text. Among them, the client can be but not limited to various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server can be implemented with an independent server or a server cluster composed of multiple servers. The present invention is described in detail below through specific embodiments.

[0032] See also Figure 2 As shown, Figure 2 A schematic flow chart of a method for generating stylized text based on latent modulation provided in an embodiment of the present invention includes the following steps:

[0033] S1. Obtain input text and target style parameters from a target user, convert the input text into an initial latent representation, and perform vector conversion on the target style parameters to obtain a latent style vector.

[0034] In this embodiment of the present invention, the input text refers to the original natural language content provided by the target user, serving as the semantic basis for the style transfer or generation task. For example, it is a sentence, paragraph, or document whose style needs to be transformed, such as "This movie is great, I love it so much." The target style parameter is a formalized representation of the desired style of the generated text, such as the target style of {Formality: High, Sentiment: Neutral}.

[0035] In detail, the target user's input text and target style parameters can be obtained from a pre-stored storage area through a computer statement with data capture function (such as Java statement, Python statement, etc.), where the storage area includes but is not limited to a database and a blockchain.

[0036] Furthermore, the input text is converted into a form that can be processed by the model in order to analyze the semantic representation of the user text, providing a high-quality semantic foundation for style modulation and text generation.

[0037] In this embodiment of the present invention, the initial potential representation is the matrix z0 output by the encoder, z0∈R L×768 , where L is the sequence length (including special tokens), each token corresponds to a 768-dimensional vector, the representation of each token incorporates global context information, and z0 is the distributed representation of the text, which converts discrete symbols into points in a continuous vector space to facilitate subsequent model processing.

[0038] In the embodiment of the present invention, converting the input text into an initial latent representation includes:

[0039] Performing a word segmentation operation on the input text to obtain a subword sequence;

[0040] Adding a position code to each subword in the subword sequence to obtain a target input sequence;

[0041] Analyzing the semantic association between each subword in the target input sequence and all subwords in the target input sequence using multi-head self-attention in a preset encoder to obtain an association sequence;

[0042] The associated sequence is nonlinearly transformed using a feedforward neural network in the encoder to obtain an initial potential representation corresponding to the input text.

[0043] Specifically, the input text x is of length L and is in the form of token IDs. The WordPiece / BPE word segmenter is used, and the maximum length supported is 512. The input text is split into subword units that the model can process, and [CLS] and [SEP] are added at the beginning and end of the sequence. However, the Transformer model lacks temporal perception and needs to explicitly encode token position information. Therefore, a position code is added to each subword, and the subword embedding is added to the position code to obtain the final target input sequence.

[0044] Specifically, the representation of each token is updated to the weighted sum of all token representations, where the weight reflects the strength of semantic association. The semantic association is calculated through multi-head self-attention, and the feature expression capability is enhanced through high-dimensional mapping. Nonlinear transformation is introduced to capture the complex semantic interactions between tokens, such as logical relations, metaphors, etc., thereby obtaining the initial potential representation corresponding to the input text, that is, using BERT to encode the text: z0 = BERT(x)∈R L×768 , z0 is the representation of all tokens in the sentence.

[0045] For example, in a medical scenario, the original text is "The patient is a 56-year-old male who complained of polydipsia, polyphagia, and polyuria for one month, accompanied by weight loss. He has a history of hypertension for five years and is currently taking amlodipine to control blood pressure." The corresponding initial latent representation is to convert the attention output into a 768-dimensional vector through two layers of linear transformation and ReLU activation, highlighting the characteristics of medical entities (such as symptoms, drugs, and medical history); for example, in a fintech scenario, the original text is "The company's revenue in Q1 2024 increased by 15.2% year-on-year, and the net profit margin increased by 2.3 percentage points, mainly due to cost control and market expansion of new product lines." The corresponding initial latent representation highlights the characteristics of financial indicators: the representation of numerical values ​​such as 15.2% and 2.3 percentage points is enhanced through nonlinear transformation, and mapped to the latent space required for financial analysis.

[0046] Furthermore, in order to unify the representation space of the target style parameters and the input text, the discrete style parameters need to be converted into a continuous vector space to facilitate model calculation and optimization.

[0047] In the embodiment of the present invention, the potential style vector s is obtained from a Gaussian distribution The random vector sampled from , whose distribution parameters μ and σ are generated by style embedding through MLP.

[0048] In the embodiment of the present invention, referring to Figure 3As shown, the target style parameters are vector-converted to obtain a potential style vector, including:

[0049] S31, extracting emotional dimension data and formality dimension data from the target style parameters;

[0050] S32, converting the emotion dimension data into an emotion embedding vector, and converting the formality dimension data into a formality embedding vector;

[0051] S33, concatenating the emotion embedding vector and the formality embedding vector to obtain a style embedding vector;

[0052] S34, using a preset multi-layer perceptron to analyze probability distribution parameters of the latent space corresponding to the style embedding vector;

[0053] S35. Perform a sampling operation on the style embedding vector according to the probability distribution parameter to obtain a latent style vector.

[0054] Specifically, the target style parameter is a multidimensional style description specified by the user, such as the sentiment dimension: {"intensity": 3, "polarity": "positive"} (quantized as a numerical value, such as 1-5); the formality dimension: {"level": "high"} (quantized as 1-3, such as informal, neutral, formal), and then the sentiment intensity value (such as 3) is mapped to the embedding table Embed emotion The corresponding vector in the formality level (such as "high" → quantized to 3) is mapped to the embedding table Emb formality , and then embed the emotion and formality into dimension splicing, that is, mapping the two options of emotion and formality into vectors, and splicing these two vectors. After a layer of full connection, the style vector s is obtained, with a dimension of d_s, and then the probability distribution parameters μ and logσ are analyzed by the multi-layer perceptron. 2 , then μ represents the center position of the target style in the latent space (e.g., the point corresponding to strong positive + high formality); σ represents the uncertainty or diversity of the style (e.g., the same positive style may be expressed in different ways), and linear interpolation between two points in the latent space can generate a continuous style transition, such as a gradient from formal to informal.

[0055] Specifically, the potential style vector s is obtained as (μ, logσ 2 )=MLP ψ (concat(Emb emotion ,Emb formality ,…)), where each sub-style is an independent embedding vector (e.g., sentiment intensity is quantized at 5 levels, formality is quantized at 3 levels), where μ, logσ 2 The dimensions are all d_s, and the potential style variables are finally sampled: The dimension of s is d_s, noise ∈ is sampled from the standard normal distribution, and the noise scale is adjusted by μ and σ to generate the final potential style vector s∈R d ,

[0056] For example, in the medical scenario, the style parameters are sentiment = neutral and formality = high. The latent style vector s is used to generate medical record summaries that comply with medical standards to avoid emotional tendencies interfering with professional judgment. In the financial scenario, the style parameters are sentiment = positive and formality = medium. The latent style vector s guides the generation of financial product promotion copy to convey optimistic expectations within the compliance range.

[0057] Furthermore, by adjusting μ and σ, the style strength and diversity can be precisely controlled. For example, increasing μ makes the generated text style more creative, and fixing σ ensures style consistency. The latent style vector s and the text semantic representation z0 are in the same vector space, which facilitates subsequent fusion.

[0058] S2. Input the initial latent representation and the latent style vector into a preset temporal coding network, and output an initial spatial adjustment corresponding to the initial latent representation.

[0059] In the embodiment of the present invention, the initial spatial adjustment value is a correction vector of the initial latent representation by the temporal coding network, reflecting the modulation result of the style information on the original semantics.

[0060] In an embodiment of the present invention, inputting the initial latent representation and the latent style vector into a preset temporal coding network and outputting an initial spatial adjustment corresponding to the initial latent representation includes:

[0061] Performing vector concatenation on the initial latent representation and the latent style vector to obtain a concatenated vector;

[0062] Calculating update gate parameters and reset gate parameters for each subword position in the concatenated vector using a gated recurrent unit in a preset temporal coding network;

[0063] Calculate the candidate hidden state of each subword position in the splicing vector according to the reset gate parameters;

[0064] Update the target hidden state of each subword position in the concatenated vector according to the update gate parameter and the candidate hidden state;

[0065] The target hidden states are combined into a hidden state sequence, and an initial spatial adjustment amount corresponding to the initial potential representation is determined according to the hidden state sequence.

[0066] In detail, the initial latent representation z0∈R L×768 , where L is the sequence length, 768 is the feature dimension, and the potential style vector s∈R d, where d is the style vector dimension, the potential style vector s is copied L times along the sequence length to form a matrix of the same shape as z0, and then spliced ​​according to the Token dimension to obtain the splicing vector, so as to calculate the update gate and reset gate parameters based on the gated recurrent unit (GRU), and the update gate calculates the splicing of the hidden state of the previous moment and the current input s. The reset gate parameter determines how much information of the previous state is forgotten, so the reset gate parameter is element-wise multiplied with the previous state, and the mixing ratio of the previous state and the candidate state is controlled by the update gate to obtain the final hidden state sequence after GRU processing. The initial spatial adjustment amount is the difference between the current potential representation and the potential representation of the previous state. The initial spatial adjustment amount represents the adjustment direction and amplitude of the semantic representation of each Token under style guidance.

[0067] Specifically, if the target style is positive, the spatial adjustment may enhance the representation of positive words in the initial latent representation, such as adjusting "not bad" to "great"; if the target style is formal, the spatial adjustment may adjust colloquial expressions to written terms, such as adjusting "done" to "done", and the GRU's gating mechanism adaptively controls the injection intensity of style information by updating the gate and resetting the gate to avoid semantic distortion caused by excessive adjustment.

[0068] Furthermore, the initial spatial adjustment amount is generated based on shallow features and may not be able to accurately cover all semantic details. A single adjustment may only change individual words, while iterative optimization can adjust the emotional intensity of the deep semantics layer by layer. It can gradually approach the optimal solution through multiple iterations and ultimately achieve a balanced state that is both semantically consistent and matches the style.

[0069] S3. Iteratively optimize the initial latent representation according to the initial spatial adjustment amount and the latent style vector to obtain a target latent representation.

[0070] In an embodiment of the present invention, the target latent representation is a final vector representation that integrates the semantic information of the input text and the target style parameters after iterative optimization. It not only retains the content semantics of the original text, but also injects style features such as emotion and formality through the style vector.

[0071] In an embodiment of the present invention, iteratively optimizing the initial latent representation according to the initial spatial adjustment amount and the latent style vector to obtain a target latent representation includes:

[0072] performing a residual update on the initial spatial adjustment amount and the initial latent representation to obtain a latent updated representation;

[0073] fusing the latent update representation with the latent style vector to obtain a fused vector;

[0074] Analyzing the target space adjustment amount corresponding to the fusion vector using a gated recurrent unit in a preset temporal coding network;

[0075] Updating the target spatial adjustment amount by the initial spatial adjustment amount, updating the potential update representation to the initial potential representation, and returning to the step of performing residual updating on the initial spatial adjustment amount and the initial potential representation to obtain the potential update representation, until the number of iterations reaches the number of units of the gated recurrent unit;

[0076] When the number of iterations reaches the number of units of the gated recurrent unit, the output potential update representation is the target potential representation.

[0077] Specifically, the initial spatial adjustment is residually concatenated (i.e., added) with the initial latent representation to obtain the updated latent representation. The essence of the residual concatenation is to preserve the original semantic information while adding incremental style adjustments to avoid content loss caused by direct modification. The updated latent representation is then fused with the latent style vector to form a fused vector, or vector concatenation. This fused vector is then analyzed using a gated recurrent unit (GRU), and the new target spatial adjustment is calculated using update and reset gates. The GRU can capture temporal dependencies within a sequence and exploit deep features of the content-style association within the fused vector.

[0078] Specifically, the target space adjustment amount replaces the initial adjustment amount, and the latent updated representation is used as the new initial latent representation. The process of residual update → fusion → generation of adjustment amount is repeated until the number of iterations is equal to the number of GRU units, that is, the number of network layers. Each layer of GRU corresponds to the feature extraction of one abstract level. When the number of iterations is consistent with the number of layers, the model can complete the content-style fusion from shallow to deep layers, ensuring that the adjustment amount covers all semantic levels. That is, each iteration is equivalent to a fine-tuning, so that the latent representation continues to move toward the target position in the content-style space, and finally reaches a balanced state that is both semantically consistent and matches the style.

[0079] Furthermore, during the iteration process, the hyperparameters of the optimization process are dynamically controlled by the staged optimization scheduler, and during training, the learning rate η is controlled by different time periods t(step) t , and the weight of the total loss function loss, the scheduler adopts a dynamic adjustment strategy: η t =η max exp(-λt / T), β t =β min +(β max -β min )·sigmoid(tT / 2)where η t is the learning rate of step t; λ controls the decay rate of the learning rate; T is the total number of optimization steps; β tis the dynamic KL weight in loss; β min and β max are the minimum and maximum values ​​of KL weights, which are hyperparameters.

[0080] For example, in a medical scenario, the input text is "The patient's blood pressure is 160 / 95 mmHg and the heart rate needs to be monitored." The target style is "Emotional dimension: Urgent; Formality: High." The iterative process is to add urgent keywords to the initial adjustment → [Urgent] The patient's blood pressure is 160 / 95 mmHg and the heart rate needs to be monitored immediately. After iterative optimization, the semantic tendency is adjusted → Warning: The patient's blood pressure reaches 160 / 95 mmHg and the heart rate must be monitored and reported immediately. In a fintech scenario, the input text is "A company's Q1 revenue increased by 15%." The target style is "Investment analysis report" (Emotional dimension: Neutral; Formality: High). The iterative process is to add professional terms to the initial adjustment → A company's revenue increased by 15% year-on-year in the first quarter of 2025. After iterative optimization, in-depth analysis is added → According to financial report data, a company's revenue in Q1 2025 achieved a year-on-year growth of 15%, mainly due to market share expansion and cost control optimization, reflecting its strong operating resilience.

[0081] Furthermore, in order to achieve the collaborative control of multi-dimensional styles (such as emotion and formality) and support the independent or joint adjustment of different style attributes such as emotion and formality, it is necessary to map the target latent representation into the target style space.

[0082] S4. Mapping the target latent representation into a preset target style space through a preset style condition projector.

[0083] In an embodiment of the present invention, the target style space is a predefined standardized vector space. The representation in the target style space has the decoupling characteristics of semantics and style. The semantic information is retained in the token-level representation, and the style information is encoded in the overall distribution of the vector. By adjusting specific dimensions in the target style space, the style can be modified in a targeted manner.

[0084] In an embodiment of the present invention, mapping the target latent representation into a preset target style space using a preset style condition projector includes:

[0085] Performing vector concatenation on the target latent representation and the latent style vector to obtain a latent concatenated vector;

[0086] Calculating the self-attention parameters of the latent splicing vector through a multi-head self-attention layer in a preset style condition projector;

[0087] Performing residual connection and layer normalization processing on the self-attention parameter and the potential splicing vector to obtain a target output parameter;

[0088] Analyzing update layer parameters corresponding to the target output parameters using a feedforward neural network layer in a preset style condition projector;

[0089] The network layer in the style condition projector is updated according to the update layer parameters, and the updated network layer is used to output the target style representation of the target latent representation in the target style space.

[0090] Specifically, the target latent representation and the latent style vector are concatenated along the sequence dimension to obtain a latent concatenated vector. Global correlations are captured through a multi-head self-attention layer, i.e., a QKV matrix is ​​generated. Single-head attention is calculated based on the QKV matrix, and multiple attention heads are calculated in parallel and concatenated to output self-attention parameters. This allows the representation of each token to interact with all tokens in the sequence (including itself). The self-attention mechanism evenly distributes the features of the style vector to all tokens, avoiding local style shifts. For example, the features of the style vector are diffused to the entire sequence to ensure global style consistency. Residual connections are performed on the self-attention parameters to retain the information of the original concatenated vector, avoiding feature loss in deep networks, and normalizing the feature dimensions of each sample to accelerate convergence and improve model stability.

[0091] Specifically, nonlinear transformation is performed through a feedforward neural network, and the feature expression capability is enhanced through high-dimensional mapping. The output of the feedforward network is used as the update layer parameter, which is back-propagated to all trainable parameters of the projector to optimize the semantic-style mapping relationship. Finally, the target style representation is output. The representation of each token has been adapted to the target style space and can be directly used by the decoder to generate stylized text.

[0092] Furthermore, the target latent representation is an abstract continuous vector that cannot be directly understood by humans. Therefore, decoding is the only way to convert the target latent representation into natural language, corresponding to the semantic-style point in the vector space to a sentence in natural language.

[0093] S5. Decode the target latent representation mapped to the target style space to obtain a stylized text corresponding to the input text.

[0094] In an embodiment of the present invention, the stylized text refers to a natural language output that integrates target style parameters, such as emotion, formality, and domain characteristics, while retaining the core semantics of the input text. The style of the entire text is unified and the content is not distorted due to style adjustments.

[0095] In the embodiment of the present invention, referring to Figure 4 As shown, decoding the target latent representation mapped to the target style space to obtain the stylized text corresponding to the input text includes:

[0096] S41, using the cross-attention layer in the preset decoder architecture to analyze the target output subword corresponding to the target latent representation mapped to the target style space;

[0097] S42, forming a subword sequence according to the target output subwords, and analyzing the lexical probability distribution of the subword sequence according to a preset vocabulary;

[0098] S43: Generate stylized text corresponding to the input text according to the vocabulary probability distribution.

[0099] In detail, the Transformer decoder is used, and each layer contains masked self-attention, encoder-decoder cross attention and feedforward network. The cross attention is calculated based on the QKV matrix, where Q comes from the self-attention output of the decoder and the generated sub-word representation, and K and V come from the latent representation of the target style space. When generating each sub-word, the decoder can dynamically focus on the stylized semantic features related to the current context in the latent representation of the target style space. For example, when generating medical text, the cross attention will focus on the urgent style features in the latent representation of the target style space, giving priority to words such as "immediately" and "must".

[0100] Specifically, the decoder inputs the start token [BOS] and generates a probability distribution for the first subword through cross-attention. Each subword is generated and appended to the input sequence until the end token [EOS] is generated or the maximum length is reached. This maps the output sequence to the vocabulary dimension, obtaining a vocabulary probability distribution. The decoder then selects the subword with the highest probability at each step to generate stylized text corresponding to the input text. For example, given the input text: "This movie is great! I love it!", the target style is {Formality: High, Sentiment: Neutral}, and the output text is: "The production quality of this film is commendable, and the overall performance is impressive."

[0101] Furthermore, we conduct end-to-end joint training for TPCM, projector, and decoder, that is, we optimize the TPCM module, projector, and decoder as a whole, and back-propagate the gradient to all module parameters at the same time. The modules evolve collaboratively to avoid the error accumulation of staged training and improve the consistency of style transfer and text generation. That is, through the total loss function L = Lgen + β t ·L KL +L style_match Optimize stylized text, where Lgen is the text generation loss (CrossEntropy), which measures the difference in probability distribution between the model-generated text and the real text; L KL L is the latent variable regularization term (distance from the Gaussian prior), which measures the difference between the two distributions. This loss forces the distribution of the latent style vector s to approach the standard normal, enhancing the generalization and semantic decoupling ability of the latent variable, similar to the regularization idea of ​​VAE; style_matchis the style similarity loss (the style classifier output matches the target), ensuring that the style of the generated text is consistent with the latent style vector s to avoid style drift, where: This formula requires that the mean and variance of the latent style vector s tend to the standard normal distribution.

[0102] For example, in a medical scenario, the target potential representation includes the semantics of a patient's fast heart rate and emergency style features, so cross-attention focuses on the emergency features, giving priority to words such as immediate and urgent, and increasing the probability of medical terms such as monitoring and processing in the probability distribution. The generated stylized text is: the patient's heart rate reaches 120 beats / minute, and electrocardiogram monitoring is required immediately and the emergency treatment process is initiated; in a fintech scenario, the target potential representation includes the semantics of a stock price drop of more than 5% and risk warning style features, cross-attention strengthens the activation of dimensions such as risk and vigilance, and beam search retains professional expressions such as increased volatility and cautious operation, and the generated stylized text is: a company's stock price fell by 5.2% today, and volatility may intensify in the short term. Investors are advised to remain cautious and reasonably control their positions.

[0103] It can be seen that in the above scheme, on the basis of the original encoder-decoder framework, a temporal prediction coding module and a staged optimization scheduler are introduced, combined with variational inference and temporal coding network prediction mechanism, and the temporal coding network predicts the latent space adjustment amount in real time. It can dynamically adjust the optimization direction according to the difference between the current state and the target style, realize closed-loop optimization of the latent space and dynamic control of the style trajectory, and at the same time, use the style condition projector of the Transformer architecture to realize the collaborative mapping and progressive adjustment of multi-dimensional style parameters, so that in the text generation process, the target style can be gradually approached and semantic coherence can be maintained, thereby solving the technical problem of low accuracy in generating stylized text.

[0104] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0105] In one embodiment, a stylized text generation device based on latent modulation is provided, and the stylized text generation device based on latent modulation corresponds to the stylized text generation method based on latent modulation in the above embodiment. Figure 5 As shown, the stylized text generation apparatus based on latent modulation includes a target style parameter conversion module 101, an initial spatial adjustment value output module 102, an initial latent representation iterative optimization module 103, a target latent representation mapping module 104, and a stylized text generation module 105. The functional modules are described in detail as follows:

[0106] A target style parameter conversion module 101 is configured to obtain input text and target style parameters from a target user, convert the input text into an initial latent representation, and perform vector conversion on the target style parameters to obtain a latent style vector;

[0107] An initial spatial adjustment value output module 102 is configured to input the initial latent representation and the latent style vector into a preset temporal coding network and output an initial spatial adjustment value corresponding to the initial latent representation;

[0108] an initial latent representation iterative optimization module 103, configured to iteratively optimize the initial latent representation according to the initial spatial adjustment amount and the latent style vector to obtain a target latent representation;

[0109] a target latent representation mapping module 104, configured to map the target latent representation into a preset target style space using a preset style condition projector;

[0110] The stylized text generation module 105 is configured to decode the target latent representation mapped to the target style space to obtain stylized text corresponding to the input text.

[0111] In one embodiment, the target style parameter conversion module 101, when converting the input text into the initial latent representation, is configured to:

[0112] Performing a word segmentation operation on the input text to obtain a subword sequence;

[0113] Adding a position code to each subword in the subword sequence to obtain a target input sequence;

[0114] Analyzing the semantic association between each subword in the target input sequence and all subwords in the target input sequence using multi-head self-attention in a preset encoder to obtain an association sequence;

[0115] The associated sequence is nonlinearly transformed using a feedforward neural network in the encoder to obtain an initial potential representation corresponding to the input text.

[0116] In one embodiment, the target style parameter conversion module 101, when performing vector conversion on the target style parameter to obtain a latent style vector, is further configured to:

[0117] Extracting emotional dimension data and formality dimension data from the target style parameters;

[0118] Converting the sentiment dimension data into a sentiment embedding vector, and converting the formality dimension data into a formality embedding vector;

[0119] Concatenate the emotion embedding vector and the formality embedding vector to obtain a style embedding vector;

[0120] Analyzing the probability distribution parameters of the latent space corresponding to the style embedding vector using a preset multi-layer perceptron;

[0121] A sampling operation is performed on the style embedding vector according to the probability distribution parameters to obtain a latent style vector.

[0122] In one embodiment, the initial spatial adjustment value output module 102, when inputting the initial latent representation and the latent style vector into a preset temporal coding network and outputting the initial spatial adjustment value corresponding to the initial latent representation, is configured to:

[0123] Performing vector concatenation on the initial latent representation and the latent style vector to obtain a concatenated vector;

[0124] Calculating update gate parameters and reset gate parameters for each subword position in the concatenated vector using a gated recurrent unit in a preset temporal coding network;

[0125] Calculate the candidate hidden state of each subword position in the splicing vector according to the reset gate parameters;

[0126] Update the target hidden state of each subword position in the concatenated vector according to the update gate parameter and the candidate hidden state;

[0127] The target hidden states are combined into a hidden state sequence, and an initial spatial adjustment amount corresponding to the initial potential representation is determined according to the hidden state sequence.

[0128] In one embodiment, the initial latent representation iterative optimization module 103, when iteratively optimizing the initial latent representation using the initial spatial adjustment amount and the latent style vector to obtain a target latent representation, is configured to:

[0129] performing a residual update on the initial spatial adjustment amount and the initial latent representation to obtain a latent updated representation;

[0130] fusing the latent update representation with the latent style vector to obtain a fused vector;

[0131] Analyzing the target space adjustment amount corresponding to the fusion vector using a gated recurrent unit in a preset temporal coding network;

[0132] Updating the target spatial adjustment amount by the initial spatial adjustment amount, updating the potential update representation to the initial potential representation, and returning to the step of performing residual updating on the initial spatial adjustment amount and the initial potential representation to obtain the potential update representation, until the number of iterations reaches the number of units of the gated recurrent unit;

[0133] When the number of iterations reaches the number of units of the gated recurrent unit, the output potential update representation is the target potential representation.

[0134] In one embodiment, the target latent representation mapping module 104, when mapping the target latent representation to a preset target style space using a preset style condition projector, is configured to:

[0135] Performing vector concatenation on the target latent representation and the latent style vector to obtain a latent concatenated vector;

[0136] Calculating the self-attention parameters of the latent splicing vector through a multi-head self-attention layer in a preset style condition projector;

[0137] Performing residual connection and layer normalization processing on the self-attention parameter and the potential splicing vector to obtain a target output parameter;

[0138] Analyzing update layer parameters corresponding to the target output parameters using a feedforward neural network layer in a preset style condition projector;

[0139] The network layer in the style condition projector is updated according to the update layer parameters, and the updated network layer is used to output the target style representation of the target latent representation in the target style space.

[0140] In one embodiment, the stylized text generation module 105 , when decoding the target latent representation mapped to the target style space to obtain the stylized text corresponding to the input text, is configured to:

[0141] Analyze the target output subwords corresponding to the target latent representation mapped to the target style space using the cross-attention layer in the preset decoder architecture;

[0142] Forming a subword sequence according to the target output subwords, and analyzing the lexical probability distribution of the subword sequence according to a preset vocabulary;

[0143] Generate stylized text corresponding to the input text according to the vocabulary probability distribution.

[0144] The present invention provides a stylized text generation device based on latent modulation. On the basis of the original encoder-decoder framework, it introduces a temporal prediction coding module and a staged optimization scheduler, combines variational inference with the temporal coding network prediction mechanism, and predicts the latent space adjustment amount in real time through the temporal coding network. It can dynamically adjust the optimization direction according to the difference between the current state and the target style, realize closed-loop optimization of the latent space and dynamic control of the style trajectory. At the same time, it uses the style condition projector of the Transformer architecture to realize the collaborative mapping and progressive adjustment of multi-dimensional style parameters, so that in the text generation process, the target style is gradually approached and semantic coherence is maintained, thereby solving the technical problem of low accuracy when generating stylized text.

[0145] The specific limitations of the latent modulation-based stylized text generation device can be found in the limitations of the latent modulation-based stylized text generation method described above and will not be elaborated upon here. The various modules in the latent modulation-based stylized text generation device described above can be implemented in whole or in part through software, hardware, or a combination thereof. Each of the modules can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a computer device memory in software form, so that the processor can call and execute the corresponding operations of each of the modules.

[0146] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a stylized text generation method based on potential modulation.

[0147] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 7As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the client side of a method for generating stylized text based on potential modulation.

[0148] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:

[0149] Obtaining input text and target style parameters from a target user, converting the input text into an initial latent representation, and performing vector conversion on the target style parameters to obtain a latent style vector;

[0150] Inputting the initial latent representation and the latent style vector into a preset temporal coding network, and outputting an initial spatial adjustment corresponding to the initial latent representation;

[0151] Iteratively optimizing the initial latent representation according to the initial spatial adjustment amount and the latent style vector to obtain a target latent representation;

[0152] Mapping the target latent representation into a preset target style space through a preset style condition projector;

[0153] The target latent representation mapped to the target style space is decoded to obtain a stylized text corresponding to the input text.

[0154] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0155] Obtaining input text and target style parameters from a target user, converting the input text into an initial latent representation, and performing vector conversion on the target style parameters to obtain a latent style vector;

[0156] Inputting the initial latent representation and the latent style vector into a preset temporal coding network, and outputting an initial spatial adjustment corresponding to the initial latent representation;

[0157] Iteratively optimizing the initial latent representation according to the initial spatial adjustment amount and the latent style vector to obtain a target latent representation;

[0158] Mapping the target latent representation into a preset target style space through a preset style condition projector;

[0159] The target latent representation mapped to the target style space is decoded to obtain a stylized text corresponding to the input text.

[0160] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0161] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0162] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0163] It should be noted that if software tools or components other than those of our company appear in the embodiments of this application, they are only used for illustration and do not represent actual use.

[0164] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A method for generating stylized text based on latent modulation, characterized in that: include: Obtaining input text and target style parameters from a target user, converting the input text into an initial latent representation, and performing vector conversion on the target style parameters to obtain a latent style vector; Inputting the initial latent representation and the latent style vector into a preset temporal coding network, and outputting an initial spatial adjustment corresponding to the initial latent representation; Iteratively optimizing the initial latent representation according to the initial spatial adjustment amount and the latent style vector to obtain a target latent representation; Mapping the target latent representation into a preset target style space through a preset style condition projector; The target latent representation mapped to the target style space is decoded to obtain a stylized text corresponding to the input text.

2. The method for generating stylized text based on latent modulation according to claim 1, wherein: The converting the input text into an initial latent representation comprises: Performing a word segmentation operation on the input text to obtain a subword sequence; Adding a position code to each subword in the subword sequence to obtain a target input sequence; Analyzing the semantic association between each subword in the target input sequence and all subwords in the target input sequence using multi-head self-attention in a preset encoder to obtain an association sequence; The associated sequence is nonlinearly transformed using a feedforward neural network in the encoder to obtain an initial potential representation corresponding to the input text.

3. The method for generating stylized text based on latent modulation according to claim 1, wherein: The performing vector conversion on the target style parameter to obtain a potential style vector includes: Extracting emotional dimension data and formality dimension data from the target style parameters; Converting the sentiment dimension data into a sentiment embedding vector, and converting the formality dimension data into a formality embedding vector; Concatenate the emotion embedding vector and the formality embedding vector to obtain a style embedding vector; Analyzing the probability distribution parameters of the latent space corresponding to the style embedding vector using a preset multi-layer perceptron; A sampling operation is performed on the style embedding vector according to the probability distribution parameters to obtain a latent style vector.

4. The method for generating stylized text based on latent modulation according to claim 1, wherein: Inputting the initial latent representation and the latent style vector into a preset temporal coding network and outputting an initial spatial adjustment corresponding to the initial latent representation includes: Performing vector concatenation on the initial latent representation and the latent style vector to obtain a concatenated vector; Calculating update gate parameters and reset gate parameters for each subword position in the concatenated vector using a gated recurrent unit in a preset temporal coding network; Calculate the candidate hidden state of each subword position in the splicing vector according to the reset gate parameters; Update the target hidden state of each subword position in the concatenated vector according to the update gate parameter and the candidate hidden state; The target hidden states are combined into a hidden state sequence, and an initial spatial adjustment amount corresponding to the initial potential representation is determined according to the hidden state sequence.

5. The method for generating stylized text based on latent modulation according to claim 1, wherein: Iteratively optimizing the initial latent representation according to the initial spatial adjustment amount and the latent style vector to obtain a target latent representation includes: performing a residual update on the initial spatial adjustment amount and the initial latent representation to obtain a latent updated representation; fusing the latent update representation with the latent style vector to obtain a fused vector; Analyzing the target space adjustment amount corresponding to the fusion vector using a gated recurrent unit in a preset temporal coding network; Updating the target spatial adjustment amount by the initial spatial adjustment amount, updating the potential update representation to the initial potential representation, and returning to the step of performing residual updating on the initial spatial adjustment amount and the initial potential representation to obtain the potential update representation, until the number of iterations reaches the number of units of the gated recurrent unit; When the number of iterations reaches the number of units of the gated recurrent unit, the output potential update representation is the target potential representation.

6. The method for generating stylized text based on latent modulation according to claim 1, wherein: Mapping the target latent representation to a preset target style space using a preset style condition projector includes: Performing vector concatenation on the target latent representation and the latent style vector to obtain a latent concatenated vector; Calculating the self-attention parameters of the latent splicing vector through a multi-head self-attention layer in a preset style condition projector; Performing residual connection and layer normalization processing on the self-attention parameter and the potential splicing vector to obtain a target output parameter; Analyzing update layer parameters corresponding to the target output parameters using a feedforward neural network layer in a preset style condition projector; The network layer in the style condition projector is updated according to the update layer parameters, and the updated network layer is used to output the target style representation of the target latent representation in the target style space.

7. The method for generating stylized text based on latent modulation according to claim 1, wherein: The decoding of the target latent representation mapped to the target style space to obtain the stylized text corresponding to the input text includes: Analyze the target output subwords corresponding to the target latent representation mapped to the target style space using the cross-attention layer in the preset decoder architecture; Forming a subword sequence according to the target output subwords, and analyzing the lexical probability distribution of the subword sequence according to a preset vocabulary; Generate stylized text corresponding to the input text according to the vocabulary probability distribution.

8. A stylized text generation device based on latent modulation, characterized in that: include: A target style parameter conversion module is used to obtain the input text and target style parameters of the target user, convert the input text into an initial latent representation, and perform vector conversion on the target style parameters to obtain a latent style vector; an initial spatial adjustment value output module, configured to input the initial latent representation and the latent style vector into a preset temporal coding network, and output an initial spatial adjustment value corresponding to the initial latent representation; an initial latent representation iterative optimization module, configured to iteratively optimize the initial latent representation according to the initial spatial adjustment amount and the latent style vector to obtain a target latent representation; a target latent representation mapping module, configured to map the target latent representation into a preset target style space through a preset style condition projector; The stylized text generation module is used to decode the target latent representation mapped to the target style space to obtain the stylized text corresponding to the input text.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method for generating stylized text based on latent modulation according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for generating stylized text based on latent modulation according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • AIGC-based tourist attraction personalized propaganda copywriting automatic generation method and system

    CN121835623A