A multi-style text transfer method based on a shared neuron activation modulation network

By constructing a neuron binary classification and shared neuron activation modulation network (SNAR), the problems of style conflict and semantic shift in multi-style text transfer are solved, achieving the effect of accurate overlay of multi-style texts while keeping the content unchanged.

CN122433690APending Publication Date: 2026-07-21BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-05
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing text style transfer methods are prone to style conflicts, semantic shifts, and inconsistent style expressions during multi-style collaborative regulation. Furthermore, existing neuron regulation methods lead to style coupling and content distortion in the generated results in multi-style scenarios.

Method used

By constructing a binary classification of neurons, differential activation regulation, and an embedded shared neuron activation modulation network (SNAR), we can achieve multi-style collaborative transfer of text, directly intervene in the activation state of neurons, reduce dependence on large-scale parallel labeled corpora, and avoid activation conflicts when multiple styles are superimposed.

Benefits of technology

It achieves multi-style text generation with precise style overlay and unchanged content, effectively solving the problems of style conflict and semantic deviation in multi-style collaborative control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122433690A_ABST
    Figure CN122433690A_ABST
Patent Text Reader

Abstract

The application discloses a multi-style text migration method based on a shared neuron activation modulation network, and belongs to the technical field of natural language processing, large language model neuron regulation and text style migration. The method first constructs a style dimension and neuron mapping relationship, and divides exclusive neurons and shared neurons; directional activation modulation is performed on the exclusive neurons, and a embedded shared neuron activation modulation network (SNAR) is constructed; through style interaction embedding, joint coding, context modeling and content cross attention decoding, the shared neuron activation distribution is dynamically reorganized, and multi-style representation conflicts are eliminated; finally, the activation values after modulation of the two types of neurons are fused, and multi-style text is generated through model reasoning. The application can efficiently realize cross-dimension multi-style collaborative migration without fine-tuning the large language model, solve the style conflict, semantic deviation and activation competition problems of existing methods, and is suitable for human-computer interaction, content creation and other scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of natural language processing, text style transfer and neuron modulation technology of large language models, and particularly relates to the fields of multi-dimensional text style collaborative transfer, construction of shared neuron activation modulation networks and differential activation modulation. Background Technology

[0002] Existing text style transfer methods primarily employ a single-style, single-task approach, adjusting only a single style attribute during a single generation process. When text involves multiple style dimensions such as sentiment, formality, style, or toxicity, multiple models are typically required to process them sequentially. Existing methods are prone to style conflicts, semantic shifts, and inconsistent style expressions during multi-style collaborative regulation. Current neuron modulation methods mainly target single-style scenarios, achieving style intervention by directly modifying the activation values ​​of hidden layer neurons. However, in multi-style scenarios, different style dimensions share some neurons, leading to competition among multiple style control signals on these shared neurons. This results in style coupling, style overlay, and content distortion in the generated results. Therefore, a multi-style text transfer method capable of dynamically modulating shared neurons is needed. Summary of the Invention

[0003] This invention proposes a multi-style text transfer method based on a shared neuron activation modulation network (SNAR). Through neuron binary classification, differential activation modulation, and the construction of an embedded shared neuron activation modulation network (SNAR), it achieves collaborative transfer of multiple text styles. This method directly intervenes in neuron activation states during the model inference stage, eliminating the need for fine-tuning the underlying large language model, reducing reliance on large-scale parallel labeled corpora, and simultaneously achieving the collaborative generation of two or more cross-dimensional target styles. This effectively avoids activation conflicts when multiple styles are superimposed, achieving accurate style superposition while maintaining content integrity.

[0004] This invention proposes a multi-style text transfer method based on a shared neuron activation modulation network, comprising the following steps:

[0005] 1) Neuron-style mapping relationship identification: Establish a mapping matrix between style dimension and neuron, and divide the set of exclusive neurons E and the set of shared neurons S;

[0006] 2) Training of Shared Neuron Activation Modulation Network (SNAR): Construct a lightweight Transformer-structured SNAR network and train it using source text and style-labeled data;

[0007] 3) Target Style Text Generation: Input the original text, extract the hidden layer activations, modulate the dedicated neurons and shared neurons respectively, fuse them to restore the model's inference, and output multi-style text. Further, in step 1) neuron-style mapping relationship recognition, a style dimension-neuron mapping matrix is ​​first established based on the single-style neuron recognition results. Let the set of neurons in the Lth layer of the model be:

[0008] Where, n i Let represent the i-th neuron, and J represent the total number of neurons in the L-th layer. For any style dimension s, the corresponding set of style neurons is represented as: in, This represents the set of associated neuron indices corresponding to the k-th single style dimension. Let represent the style response score of the i-th neuron in the k-th style dimension calculated according to the style neuron recognition method, and τ represent the threshold. The neurons are divided into a dedicated neuron set E and a shared neuron set S, as shown in formulas (3) and (4):

[0009]

[0010]

[0011] in, Let E represent the mapping value between the k-th single style dimension and the i-th neuron, and let S represent the set of dedicated neurons and S represent the set of shared neurons. Further, in step 2), the shared neurons activate the modulation network SNAR training. First, statistical feature vectors are extracted from the shared neurons: in, This represents the mean activation of shared neurons. This represents the standard deviation of shared neuron activation. Encoding the target style set into a style mask vector yields formula (6): in, Represents the target style mask vector. This represents the binary indicator value of the k-th style dimension, where K represents the number of target style dimensions. Further, the style mask is input into the style interaction embedding module to obtain the style embedding representation:

[0012] in, This represents the style embedding representation after isolating the target style information; This indicates a preset style of interactive embedded function; This represents the Hadamard product, used to perform element-wise multiplication of matrices. Represents the style mask vector The transpose of ; The resulting style embedding represents the high-dimensional real space to which it belongs, the spatial dimension of which is constrained by the batch size B, the text sequence length F, and the feature dimension D. Style interaction modeling is then completed using a multilayer perceptron.

[0013] Furthermore, the statistical features X of the shared neurons s With style context representation H s By concatenating the features, we obtain the joint features: Joint feature input to Transformer Encoder for deep dependency modeling:

[0014] in, This indicates a dimensionality increase operation, used to perform joint features on a specified feature axis. Expanding the single dimension to adapt to network input requirements; ( Let represent the forward propagation function of the multi-layer Transformer encoder. Furthermore, a content-based cross-attention mechanism is employed in the decoding stage, utilizing the source text content features C as keys and values, and the context features H... c As a query, the content alignment features are obtained: in, d represents the transpose of the source text content feature matrix C. k This represents the scaling dimension of the query and key across the attention mechanism, used to prevent gradient vanishing due to excessively large dot product values. Furthermore, the target statistical features of shared neurons under the target style combination are output through a two-layer MLP:

[0015] Furthermore, in step 3) of target style text generation, an activation value-oriented modulation strategy is adopted for the dedicated neurons. Let the original activation vector of the dedicated neuron cluster in the Lth layer feedforward network of the model be... The modulation result of the specific neuron is then generated. Represented as formula (13): in, This represents the set of target style-specific neurons in the original activation tensor of layer L. The activation component, This indicates the corresponding activated component after modulation. Indicates the activation modulation intensity; determined by the modulated activation component. Constituting the results of specific neuronal regulation Furthermore, the original activation values ​​of the shared neurons are corrected based on the target statistical characteristics: in, The activation value of the i-th shared neuron carrying the target style feature after modification. This is used to construct shared neuronal regulatory outcomes. The original activation value of the i-th shared neuron. The original statistical characteristics are represented by ε, which is a constant to avoid division by zero. , The target mean and target standard deviation are predicted by the SNAR network; the corrected activation values ​​of each shared neuron are... Constituting the regulatory results of shared neurons Furthermore, the regulation results of the specific neurons are combined with the regulation results of the shared neurons to form a new hidden layer activation tensor: Where H E H represents the activation result of a specific neuron. S This represents the activation results of shared neurons. Finally, the model's forward propagation is resumed, and the target style text is decoded and output. Attached Figure Description

[0016] Figure 1 This is a diagram illustrating the overall framework of the multi-style text transfer method based on shared neuron activation modulation networks of this invention. Detailed Implementation

[0017] To make the technical solution of the present invention clearer, the present invention will be described in detail below with reference to embodiments. Step 101: Input source text and target style combination; Step 102: Input the source text into a pre-trained language model and extract the hidden state of the target layer; Step 103: Identify the dedicated neuron set E and the shared neuron set S according to the pre-constructed style dimension-neuron mapping matrix; Step 104: Perform targeted activation value modulation on the dedicated neurons to generate the dedicated neuron modulation result H. E Step 105: Extract the shared neuron activation statistical features X s And generate the target style mask vector r t Step 106: Place X s With r t Input the shared neuron activation modulation network SNAR; Step 107: Within SNAR, execute the following five module logics sequentially: First, the style interaction embedding module will use the one-hot style mask r t Encoding as style combination embedding representation Subsequently, the joint feature encoding module embeds the obtained style with the shared neuron statistical features X. s By concatenating along the channel dimension and performing a hybrid mapping, joint features are obtained. Next, the Transformer context modeling module processes the joint features. Perform self-attention computation to extract shared activation context representations Subsequently, the content cross-attention fusion module uses contextual features... To measure query volume, using the source text content feature C as the key and value, cross-attention computation is performed to achieve content alignment fusion, resulting in content alignment features. Finally, the content alignment features are output through the target statistical feature output module. Mapped to the final target statistical feature T s Step 108: Output the shared neuron target statistical features T s Step 109, according to T s By correcting the original activation values ​​of the shared neurons, the regulation result H of the shared neurons is obtained. S Step 110, Fusion H E With H S Forming a new hidden layer activation tensor H final Step 111: H final Replace the original hidden layer activation and restore the model forward propagation; Step 112: Generate the target style text through the language model decoder.

[0018] During implementation, the target style combination includes any one or more combinations of sentiment style, formality style, stylistic style, and toxic style. The Shared Neuron Modulation Network (SNAR) employs a lightweight Transformer architecture, comprising a style interaction embedding layer, a joint encoding layer, a Transformer Encoder layer, and a content cross-attention decoding layer. Specifically, the style interaction embedding layer encodes the interaction relationships between target styles; the joint encoding layer fuses the statistical features of shared neurons with style embedding features; the Transformer Encoder establishes the contextual dependencies between shared neurons; and the content cross-attention decoding layer preserves the content features of the source text during generation.

[0019] The following is a specific example:

[0020] A multistyle transfer method based on shared neuron activation modulation, comprising the following steps:

[0021] 1. Binary division of neurons

[0022] Selecting two styles, sentiment and formality, the neurons in the last layer of the model were divided as follows: 348 neurons specific to sentiment; 348 neurons specific to formality; and 61 shared neurons.

[0023] Construct a mapping matrix M to obtain sets E and S.

[0024] 2. Specific neuron modulation

[0025] Pick Based on the target style combination, the set of dedicated neurons corresponding to the target style is extracted from the style dimension-neuron mapping matrix. An activation modulation offset of intensity α is applied to the corresponding activation component according to formula (13) to obtain the dedicated neuron modulation result.

[0026] 3. SNAR-modulated shared neurons

[0027] Style mask: Positive + Formal; Extracting statistical features of shared neurons Construct a target style binary mask r t =[1, 1, 0...]; This sets X... s With r t The input is a pre-trained SNAR module; within SNAR, a content cross-attention fusion module performs cross-attention computation to align content using contextual features as queries and source text content features C as keys and values. Finally, the output module predicts the target statistical features. Substitute these values ​​into formula (14) to update the original activation values ​​of the shared neurons, thus obtaining the regulation results of the shared neurons. .

[0028] 4. Activate the fusion to generate text

[0029] Substituting into formula (15), the regulation result H of the specific neuron is... E With shared neuron regulation results Combine to form a new hidden layer activation tensor .

[0030] Input model decoding, output positive + formal style text.

[0031] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A multi-style text transfer method based on a shared neuron activation modulation network, characterized in that, include: A. Based on a given set of multi-target styles to be transferred, identify the dedicated neuron index set and the shared neuron index set in the pre-trained large language model, and construct a mapping matrix between the target style and neurons; B. Input the source style sample set into the pre-trained large language model and extract the activation statistical features of shared neurons under the source style; Using the extracted activation statistics and target style mask as input, and the target statistics of shared neurons under the target style combination as supervision, a shared neuron activation modulation network is trained to learn the mapping from source style activation statistics and target style mask to target statistics. C. Input the source style samples into the pre-trained large language model, activate and modulate the dedicated neurons according to the target style set, and then input the shared neuron activation and modulation network with the shared neuron statistical features of the source style text and the target style mask to predict the target statistical features of the shared neurons and generate the target text after multi-style transfer.

2. The multi-style text transfer method based on a shared neuron activation modulation network according to claim 1, characterized in that, Step A involves identifying the specific neuron index set and the shared neuron index set, and constructing a mapping matrix between the target style and neurons. Specifically, this includes: A1. For each single style dimension in the multi-objective style set, input the corresponding style sample into the pre-trained large language model to obtain the set of associated neuron indices corresponding to each single style dimension. The calculation method is as follows: in, Indicates the first A set of associated neuron indexes corresponding to a single style dimension. This represents the i-th neuron in the target layer. This indicates that the i-th neuron, calculated according to the style neuron recognition method, is in the i-th position. Style response scores on each style dimension, where τ represents the threshold; A2. Construct a style dimension-neuron mapping matrix based on the associated neuron index set, the mapping matrix being represented as follows: in, This represents the mapping value between the k-th single style dimension and the i-th neuron; A3. Based on the style dimension-neuron mapping matrix, perform binary functional partitioning of the target layer neurons, the partitioning method is as follows: in, Represents a set of dedicated neuron indexes. This represents the set of shared neuron indices, where K represents the number of target style dimensions. Indicates the first The number of style dimensions associated with each neuron determines whether neurons associated with only one style dimension are assigned to a dedicated neuron index set, or neurons associated with at least two style dimensions are assigned to a shared neuron index set.

3. The multi-style text transfer method based on a shared neuron activation modulation network according to claim 1, characterized in that, Step B involves training the shared neuron activation modulation network, specifically including: B1. Input the source style sample set into the pre-trained large language model, extract the activation components corresponding to the shared neuron index set S in the target layer, calculate the shared neuron activation statistical feature vector, and encode the target style set into a target style mask vector. The calculation method is as follows: in, This represents the statistical feature vector of shared neuron activation. This represents the mean activation of shared neurons. Indicates the standard deviation of shared neuron activation. Represents the target style mask vector. Represents the binary indicator value of the k-th style dimension; B2. Convert the target style mask vector The input style interaction embedding layer is used to obtain the style embedding representation, and style interaction modeling is completed through a multilayer perceptron. The calculation method is as follows: in, This represents the target style embedding vector. This represents the preset style interaction embedding function, and ⊙ represents the Hadamard product. This represents the transpose of the target style mask vector. The style context representation is given, and MLP(⋅) represents the multilayer perceptron mapping function; B3, the shared neuron activation statistical feature vector is used. With style context representation The features are concatenated to obtain joint features, which are then expanded into a sequence form and input into a TransformerEncoder for deep dependency modeling. The calculation method is as follows: in, This represents the joint features obtained by concatenation. Concat(⋅) represents the feature concatenation operation, Unsqueeze(⋅) represents the dimensionality increase operation, and TransformerEncoder(⋅) represents the forward propagation function of a multi-layer Transformer encoder. This represents the contextual features obtained through deep dependency modeling; B4. In the decoding stage of the shared neuron activation modulation network, a content-cross-attention mechanism is employed to integrate contextual features. As a query, the source text content feature C is used as the key and value to obtain the content alignment feature. The target statistical feature output module then outputs the target statistical feature of the shared neurons under the target style combination. The calculation method is as follows: , in, Indicates the characteristics of the source text content. This represents the transpose of the source text content feature matrix. This indicates the scaling dimension of the content across the attention mechanism for queries and keys. This represents the content alignment feature, and Softmax(⋅) represents the normalization exponential function. MLP represents the statistical characteristics of shared neuron targets. out (⋅) represents the target statistical feature output module; during training, the real statistical features of shared neurons extracted from samples under the target style combination are used as supervision signals to calculate the loss and update the network parameters until convergence, thus obtaining the trained shared neuron activation modulation network.

4. The multi-style text transfer method based on a shared neuron activation modulation network according to claim 1, characterized in that, Step C generates the target text after multi-style transfer, specifically including: C1, obtaining the source text to be transferred. List of target style names Post-trained shared neuron activation modulation network, and dedicated neuron index set and shared neuron index set The source text to be migrated Input the pre-trained large language model and obtain its rank in the large language model. The original activation tensor of the layer and create the original activation tensor copy As the activation tensor to be modulated, where... This indicates the source text to be migrated. This represents a list of target style names. Indicates the target layer for performing activation modulation. Represents the original activation tensor. C2 represents the activation tensor to be modulated; C3 represents the list of target style names. And the style dimension—neuron mapping matrix—determines the specific set of neuron indices corresponding to the target style. and the activation tensor to be modulated The intensity applied to the corresponding dedicated neuron is The activation modulation offset is calculated as follows: in, This represents the set of neuron indices specific to the target style. Represents the set of neuron indices in the original activation tensor. The corresponding activation component, Represents the set of specific neuron indices in the activation tensor to be modulated. The corresponding activation component, Indicates the activation modulation intensity, the Constituting the activation results of specific neurons C3, from the original activation tensor Extract the shared neuron index set Calculate the shared neuron activation statistical feature vector for the corresponding activation components. List of target style names Convert to target style mask and will and Input the trained shared neuron activation modulation network to obtain the target statistical features of the shared neurons. The original activation value of the shared neuron is corrected based on the target statistical characteristics of the shared neuron. The calculation method is as follows: in, This represents the original activation value of the i-th shared neuron. Let represent the modified activation value of the i-th shared neuron carrying the target style feature, and ε represent a constant to avoid division by zero. This represents the target mean in the target statistical features of shared neurons. The target standard deviation represents the target statistical characteristics of shared neurons, and the standard deviation of each shared neuron is... Constituting the activation results of shared neurons ; C4. Combine the activation results of the dedicated neurons with the activation results of the shared neurons to form a new hidden layer activation tensor. The calculation method is as follows: in, This represents the merged hidden layer activation tensor, where Merge(⋅) denotes the merging operation; C5, using the new hidden layer activation tensor The original hidden layer activations are replaced, the forward computation after the Lth layer of the pre-trained large language model is restored, and the model output is decoded to generate the final target style text. The calculation method is as follows: in, This indicates the pre-trained large language model. Forward computation after the layer, Indicates a decoding operation. This represents the target text after multidimensional style transformation.