Multi-modal sentiment analysis method based on depth decoupling and cross-modal semantic alignment
Through the cross-modal semantic alignment method of HSIC criterion decoupling and text-guided cross-modal semantic alignment methods, the problems of redundancy and conflict between modes in multimodal sentiment analysis are solved, and more efficient emotion prediction is achieved.
Patent Information
- Application Number
- CN202510432171.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-22
Smart Images

Figure CN120354348A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence, and particularly relates to a multi-modal sentiment analysis method combining feature decoupling and cross-modal alignment, which is applicable to the joint semantic modeling and robustness prediction scenarios of text, speech, and visual data. Background Art
[0002] The statements in this section only provide background technical information related to the present disclosure, and these statements may constitute prior art. In the process of implementing the present invention, the inventors found that there are at least the following problems in the prior art.
[0003] With the wide application of sentiment analysis technology, especially in social media and customer feedback analysis, multi-modal sentiment analysis (MSA) has gradually become an important direction for improving the accuracy of sentiment recognition. Traditional sentiment analysis methods mainly rely on a single modality (such as text analysis, speech analysis, or visual analysis), but a single modality often cannot comprehensively reflect the multi-dimensional characteristics of sentiment. Multi-modal sentiment analysis can provide more abundant information by combining text, audio, and visual modalities, thereby improving the accuracy of sentiment prediction.
[0004] Existing multi-modal sentiment analysis methods mainly fall into two categories: representation learning-based models and fusion-based models. Representation learning methods focus on enhancing the feature representation ability of each modality. However, due to significant distribution differences between different modalities, cross-modal alignment is difficult, leading to redundant information and conflicts. Fusion methods directly fuse information from different modalities for sentiment prediction. However, redundant information and semantic conflicts between modalities are still important factors limiting the effectiveness of these methods.
[0005] To overcome these challenges, some decoupling and alignment methods have been proposed in recent years. For example, the patent with application number 202211628659.2 titled "A Multi-modal Sentiment Analysis Method Combining Multiple Features and Attention Mechanisms", or the patent with application number 202410778237.6 titled "A Multi-modal Sentiment Analysis Method Based on Multi-agent Collaboration", uses alignment or decoupling analysis methods. However, existing decoupling methods often have difficulty retaining the key information of modalities, and alignment methods also have problems with poor cross-modal alignment effects. Therefore, how to achieve effective cross-modal alignment on the basis of decoupling and improve the accuracy of sentiment analysis has become a key issue in the field of multi-modal sentiment analysis. Summary of the Invention
[0006] Aiming at the above problems, the purpose of the present invention is to solve some problems in the prior art, or at least alleviate these problems.
[0007] A multi-modal sentiment analysis method based on deep decoupling and cross-modal semantic alignment, comprising the following steps:
[0008] Consider a sentiment analysis system that includes three modalities: text (T), audio (A), and video (V). Preprocess and extract features from the input video data to obtain input features X that include visual features, acoustic features, and word vector representations. m Among them, m ∈ T, A, V represents different modalities.
[0009] The feature decoupling mechanism based on the HSIC criterion decouples the three modalities of the input feature X m to obtain unique features and shared features
[0010] Align the semantics of the three modalities across modalities: Use the text-guided attention mechanism to perform semantic alignment on the shared features to obtain the aligned feature h after semantic alignment. m At the same time, for the unique features and the shared features perform positional encoding first to obtain and Subsequently, perform self-attention learning to obtain the shared features and unique features For modality m, its aligned feature is calculated as:
[0011] h m = Attention(Q T , K m , V m )
[0012] where Q T is the text query matrix, and K m and V m are the key-value matrices of modality m.
[0013] Input the aligned feature h m , the shared features after self-attention learning, and the unique features into the hierarchical prediction module for prediction, and weighted fusion to obtain the final prediction label
[0014]
[0015] where, are respectively and auxiliary prediction labels obtained by pooling and prediction; is the main prediction label, which is composed of h m , and It is obtained by splicing after pooling and then making a prediction; μ1, μ2, and μ3 are weight coefficients; finally, the prediction loss is obtained:
[0016]
[0017] where N is the number of samples, y i is the true label of the i-th sample, is the model prediction result for the probability value of the i-th sample in
[0018] The feature decoupling mechanism based on the HSIC criterion decouples the three modalities of the input feature X m and includes the following steps:
[0019] Extract the unique features m and shared features of the three modalities of the input feature X
[0020] Introduce the HSIC loss to obtain the minimization of the correlation between the unique features and the shared features:
[0021]
[0022] where tr(·) represents the trace operation of the matrix, that is, the sum of the main diagonal elements of the square matrix, and are the kernel matrices, H is the centering matrix, and N is the number of samples; is the decoupling loss of each modality, and after adding them up, the complete decoupling loss L HSIC is obtained.
[0023] Introduce the reconstruction loss to ensure information integrity:
[0024]
[0025] where is the decoupling reconstruction loss of each modality, and after adding them up, the final complete reconstruction loss L r is obtained to ensure the information integrity after decoupling; g m is the reconstruction function.
[0026] The multi-modal sentiment analysis method based on deep decoupling and cross-modal semantic alignment also includes model optimization to minimize the prediction loss; the complete loss L total is composed of the prediction loss, the decoupling loss, and the reconstruction loss, and the formula is as follows:
[0027] L total = L pred + λ1L HSIC + λ2L r
[0028] where λ1 and λ2 are weight coefficients.
[0029] Furthermore, the complete loss L total can be transformed into the formula:
[0030]
[0031] The λ1 and λ2 are respectively set to 0.1 and 0.01; the Adam optimizer is used for model training, the initial learning rate is set to 0.0001, the number of samples per batch is 32, and the number of training epochs is 50.
[0032] Furthermore, the specific features and the shared features are extracted through two independent Transformer encoders.
[0033] A text-guided attention mechanism is adopted to perform semantic alignment on the shared features including the following steps:
[0034] Perform positional encoding on the shared features :
[0035]
[0036] where PE is the sine positional encoding, is the feature after positional encoding;
[0037] Calculate the attention between the text modality and other modalities (including the text modality itself):
[0038]
[0039] where is a learnable projection matrix, d is the attention dimension; softmax is the normalized exponential function;
[0040] Finally, obtain the aligned feature h m :
[0041]
[0042] where LayerNorm is layer normalization.
[0043] Input the aligned feature h m , the shared feature after self-attention learning and the specific features into the hierarchical prediction module for prediction, and perform weighted fusion to obtain the final prediction label including the following steps:
[0044] The shared features after self-attention learning and the unique features are obtained after pooling and predictions are made to obtain the auxiliary prediction labels
[0045]
[0046] The aligned feature h m is obtained after pooling
[0047] Concatenated and predictions are made together to obtain the main prediction label
[0048]
[0049] where FCLayer is the fully connected layer;
[0050] The three prediction labels are weighted and fused to obtain the final prediction label:
[0051]
[0052] The input video data is preprocessed and feature extracted, including the following steps:
[0053] The input video data is divided into N frames per second to obtain the audio segment and text description corresponding to the video frame sequence;
[0054] For the video modality, the pre-trained FACET model is used to extract the visual features of each frame;
[0055] For the audio modality, the COVAREP toolkit is used to extract the acoustic features;
[0056] For the text modality, the BERT model is used to convert the text into word vector representations;
[0057] Finally, the preprocessed input features are obtained:
[0058]
[0059] where m ∈ T, A, V represents different modalities, N m is the sequence length, d m is the feature dimension.
[0060] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the multi-modal sentiment analysis method based on deep decoupling and cross-modal semantic alignment are implemented.
[0061] The present invention has the following beneficial effects:
[0062] 1. Information decoupling and redundancy removal: Through deep decoupling technology, the unique attributes of modalities are separated from the shared emotional semantics, avoiding the influence of information redundancy and conflicts between different modalities, and improving the accuracy of sentiment analysis.
[0063] 2. Cross-modal alignment: Through the text-guided cross-modal alignment mechanism, the semantic consistency between modalities is effectively enhanced, the fusion effect of multi-modal features is improved, and the model can better understand the relevance of cross-modal information during sentiment prediction.
[0064] 3. Hierarchical fusion strategy: Through the hierarchical fusion strategy, shared features and unique features are combined, further improving the accuracy of sentiment analysis and enhancing the robustness of the model.
[0065] 4. Comprehensive loss function: A comprehensive loss function is designed, which can optimize the decoupling, alignment, and feature fusion processes simultaneously, ensuring the maximization of the model's performance in multi-modal sentiment analysis.
[0066] Through the synergistic effect of the above technical solutions, the present invention realizes effective cross-modal alignment on the basis of decoupling, thereby improving the accuracy of sentiment analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 is a flowchart of a multi-modal sentiment analysis method based on deep decoupling and cross-modal semantic alignment of the present invention;
[0068] Figure 2 is the overall architecture diagram of the model of the present invention;
[0069] Figure 3 is the cross-modal semantic alignment module diagram of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0070] The following further describes the present invention with reference to the accompanying drawings. The embodiments of the present invention are only used to illustrate the present invention and do not limit the present invention. Without departing from the technical idea of the present invention, various substitutions and changes made according to ordinary technical knowledge and customary means in the art shall all be included within the scope of the present invention.
[0071] To solve the problems existing in the above-mentioned prior art, the present invention proposes a multi-modal sentiment analysis method based on deep decoupling and cross-modal semantic alignment, which reduces redundant information between modalities through feature decoupling and improves the quality of feature representation by using text-guided semantic alignment to achieve accurate prediction of multi-modal sentiment analysis. To achieve the above object, the present invention adopts the following technical solutions, as Figure 1 or shown in FIG. 2.
[0072] Step 1: Preprocessing of original feature data
[0073] Build the input model of the multi-modal sentiment analysis system. As Figure 2 shown, preprocess and extract features from the input video data. In the data preprocessing stage, first sample the video at a frame rate of 30 frames per second to obtain the audio segments and text descriptions corresponding to the video frame sequence. (1) For the video modality, use the FACET toolkit to extract facial behavior features, including facial action units (Action Units), facial landmarks (Landmarks), etc., to obtain a 35-dimensional visual feature representation X V ∈R (T×35) , where F is the number of video frames; (2) For the audio modality, use the COVAREP toolkit to extract the audio signal at a sampling rate of 16 kHz, and calculate the acoustic features including fundamental frequency (F0), voice quality, harmonic features, etc., to obtain X A ∈R (F×74) ; (3) For the text modality, use the BERT model to convert the text into a 768-dimensional word vector representation, to obtain X T ∈R (L×768) , where L is the length of the text sequence.
[0074] Consider a sentiment analysis system that includes three modalities: text (T), audio (A), and video (V). The input feature representation after preprocessing the original input data is:
[0075]
[0076] where m ∈ T, A, V represents different modalities, N m is the sequence length, and d m is the feature dimension.
[0077] FACET tool: A facial expression analysis tool developed by iMotions, Facial Action CodingSystem Estimator Tool, used to automatically analyze visual features such as facial expression action units, head pose, and gaze direction in videos, and is a commonly used face analysis tool in emotion recognition.
[0078] COVAREP toolkit: Collaborative Voice Analysis Repository for SpeechTechnologies, an open-source speech analysis toolkit that can extract a series of low-level acoustic features for emotion recognition, such as pitch, formants, glottal frequency, etc.
[0079] BERT model: The full name is Bidirectional Encoder Representations from Transformers. It is a pre-trained language model based on the Transformer structure, which can capture semantic information bidirectionally from the context and transform the input text into a high-dimensional dense semantic representation.
[0080] Step 2: Feature decoupling
[0081] Construct a feature decoupling module. To separate modality-specific information and shared semantic information, the present invention designs a feature decoupling mechanism based on the HSIC criterion (Hilbert-Schmidt Independence Criterion), as Figure 1 shown. For modality m ∈ {T, A, V}, first perform feature transformation through two independent Transformer encoders:
[0082]
[0083] where TransformerEnc s and TransformerEnc u are respectively used to extract shared features and specific features
[0084] Transformer encoder: Transformer is a deep neural network structure based on the self-attention mechanism, which was first used to process natural language tasks.
[0085] To ensure the independence of the two types of features, introduce the HSIC loss:
[0086]
[0087] where tr(·) represents the trace operation of the matrix, that is, the sum of the elements on the main diagonal of the square matrix, is the kernel matrix, H = I - (1 / N)11 T is the centering matrix, N is the number of samples, and here T represents the transpose of the matrix. That is is the decoupling loss for each modality, and the complete decoupling loss L HSIC .
[0088] At the same time, ensure the integrity of the features through the reconstruction loss:
[0089]
[0090] where The decoupled reconstruction loss for each modality is added to obtain the final complete reconstruction loss L r , ensuring the integrity of the decoupled information; g m is the reconstruction function, implemented through a multi-layer perceptron.
[0091] Step 3: Semantic alignment
[0092] Construct a text-guided semantic alignment module. Utilize the dominant role of the text modality to achieve cross-modal alignment of shared features through the attention mechanism, and obtain the aligned feature h after semantic alignment m . For modality m, its aligned feature is calculated as:
[0093] h m = Attention(Q T , K m , V m )
[0094] where Q T is the text query matrix, and K m and V m are the key-value matrices of modality m.
[0095] Specifically, as Figure 3 shown, it includes the following steps:
[0096] Considering that the text modality has the strongest semantic expression ability, the present invention adopts a text-guided attention mechanism to achieve semantic alignment of shared features. Specifically, because the purpose of decoupling to obtain the shared features of the three modalities is to obtain features that can be aligned, the present invention needs to perform semantic alignment on the shared features.
[0097] First, perform positional encoding on the shared features:
[0098]
[0099] where PE is the sine positional encoding, is the feature after positional encoding.
[0100] Then calculate the attention between the text modality and other modalities (including the text modality itself):
[0101]
[0102] where is the learnable projection matrix, d is the attention dimension. softmax is the normalized exponential function. The finally obtained aligned feature is h m , representing the new feature obtained after the shared features of modality m and modality T are aligned. Among them, the m modality includes T itself, and h m The calculation method is:
[0103]
[0104] Among them, LayerNorm is layer normalization, that is, normalization is performed on all samples.
[0105] In addition to cross-modal semantic alignment of the shared features, we also perform self-attention learning on the shared features themselves to obtain to improve the quality of the decoupled shared features, thereby improving the quality of semantic alignment. At the same time, the shared features themselves also need to participate in subsequent multi-level predictions, and self-attention learning can improve the quality of the shared features. Therefore:
[0106]
[0107] In addition, the present invention uses the unique features of each modality as supplementary information for the subsequent multi-level prediction module. After self-attention learning, the unique features can not only improve the quality of the decoupled unique features, thereby improving the information quality, but also play an effect of increasing the generalization ability of the model in the subsequent multi-level prediction module. Therefore, after position encoding, they also enter the encoder self-attention learning to obtain
[0108]
[0109] Step 4: Hierarchical prediction
[0110] Construct a hierarchical prediction module. In the previous step, we obtained three types of features h that finally participated in the hierarchical prediction m , We need to use these features to obtain multiple prediction results respectively.
[0111] First, for the shared features, perform temporal pooling:
[0112]
[0113] Then, concatenate the shared features of each modality and perform prediction through a multi-layer perceptron:
[0114]
[0115] For the unique features, perform pooling and prediction in the same way:
[0116]
[0117] For the alignment feature h m , perform pooling in the same way to obtain but do not perform prediction:
[0118]
[0119] In the hierarchical prediction module, is the label for auxiliary prediction, while for the main prediction label is concatenated and predicted together:
[0120]
[0121] FCLayer is a fully connected layer.
[0122] The complete sentiment prediction label is obtained by weighted fusion of the three:
[0123]
[0124] μ is the fusion weight, and finally the prediction loss is obtained:
[0125]
[0126] where N is the number of samples, y i is the true label of the i-th sample, is the model prediction result for the probability value of the i-th sample in the result. This is a binary cross-entropy loss function. By weighted fusion of the three prediction results, the validity of the main prediction result is ensured, and effective information is also provided for the other two labels. The possible interference among them is also reduced by small weights, which is equivalent to adding noise to improve the generalization ability of the model.
[0127] Step Five: Model Optimization
[0128] The main optimization objective of the present invention is to minimize the prediction loss, L HSIC is to minimize the correlation between the specific features and the shared features, L r is to ensure the integrity of the decoupled information. The overall loss function of the present invention consists of the prediction loss, the decoupling loss, and the reconstruction loss.
[0129]
[0130] where L pred is the prediction loss, λ1 and λ2 are weight coefficients, which are set to 0.1 and 0.01 respectively. The Adam optimizer is used for model training. The initial learning rate is set to 0.0001, the number of samples per batch is 32, and the number of training epochs is 50. The Adam optimizer is adaptive moment estimation, which is a widely used deep learning optimization algorithm.
[0131] The present invention proposes a multi-modal sentiment analysis method (Disentanglement Alignment Model, DAM) based on deep disentanglement and cross-modal semantic alignment. This method decouples the features of text, audio, and visual modalities, separates the unique attributes of each modality from the shared sentiment semantics, and solves the problems of redundant and conflicting information between modalities in existing multi-modal sentiment analysis methods. On this basis, a text-guided cross-modal alignment mechanism is proposed, and the shared features are aligned through a cross-modal attention mechanism, enhancing the semantic consistency between different modalities. At the same time, through a hierarchical fusion strategy, the decoupled unique features and the aligned shared features are effectively fused to improve the accuracy of sentiment prediction.
[0132] Compared with the prior art, the present invention has the following advantages: (1) achieving deep feature disentanglement through the HSIC criterion, effectively reducing the information redundancy between modalities, and improving the independence of feature representation; (2) introducing a text-guided semantic alignment mechanism, making full use of the semantic information of the text modality to guide feature alignment, and enhancing the semantic consistency of cross-modal features; (3) adopting a hierarchical prediction strategy, using unique features and shared features for prediction respectively, and achieving more accurate sentiment analysis through weighted fusion. The experimental results of the method of the present invention on public standard datasets such as CMU-MOSI and CMU-MOSEI show that it has significant performance advantages compared with existing methods.
[0133] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the multi-modal sentiment analysis method based on deep disentanglement and cross-modal semantic alignment.
[0134] The conventional techniques and the solutions not described in detail in the above embodiments are well known in the art, so they will not be elaborated here. The above embodiments have described in detail the preferred embodiments of the present invention. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various simple modifications can be made to the technical solutions of the present invention, and these simple modifications all fall within the protection scope of the present invention.
Claims
1. A multi-modal sentiment analysis method based on deep decoupling and cross-modal semantic alignment, characterized in that It includes the following steps: Consider a sentiment analysis system that includes three modalities: text (T), audio (A), and video (V). Preprocess and extract features from the input video data to obtain the input feature X, which includes visual features, acoustic features, and word vector representations. m Among them, m ∈ T, A, V represents different modalities. The feature decoupling mechanism based on the HSIC criterion decouples the three modalities of the input feature X m to obtain the unique features and the shared features Cross-modal alignment of the semantics of three modalities: Using a text-guided attention mechanism to perform semantic alignment on the shared features to obtain the aligned feature h after semantic alignment m ; At the same time, for the unique features and the shared features perform positional encoding first to obtain and Subsequently, perform self-attention learning to obtain the shared feature and the unique feature For modality m, its aligned feature is calculated as: h m = Attention(Q T , K m , V m ) Among them, Q T is the text query matrix, K m and V m are the key-value matrices of modality m; Align the feature h m , the shared feature after self-attention learning and the unique feature are input into the hierarchical prediction module for prediction, and weighted fusion is performed to obtain the final prediction label Among them, are respectively and auxiliary prediction labels obtained by pooling and prediction; is the main prediction label, which is obtained by splicing h m , and after pooling and then performing prediction together; μ1, μ2, and μ3 are weight coefficients; finally, the prediction loss is obtained: Among them, N is the number of samples, and y i is the true label of the i-th sample, and is the probability value of the i-th sample in the model prediction result.
2. The multimodal sentiment analysis method based on deep decoupling and cross-modal semantic alignment according to claim 1, wherein The feature decoupling mechanism based on the HSIC criterion decouples the three modalities of the input feature X m and includes the following steps: Extract the input feature X separately m The unique features of the three modalities and the shared features Introduce the HSIC loss to minimize the correlation between the specific features and the shared features: where tr(·) represents the trace operation of a matrix, that is, the sum of the elements on the main diagonal of a square matrix, and is the kernel matrix, H is the centering matrix, and N is the number of samples; is the decoupling loss for each modality, and the complete decoupling loss L is obtained after summing them up HSIC ; Introduce the reconstruction loss to ensure information integrity: where is the decoupled reconstruction loss for each modality, and the sum is used to obtain the final complete reconstruction loss L r , ensuring the integrity of the decoupled information; g m is the reconstruction function.
3. The multimodal sentiment analysis method based on deep decoupling and cross-modal semantic alignment according to claim 2, wherein It also includes model optimization to minimize the prediction loss; the complete loss L total is composed of a prediction loss, a decoupling loss, and a reconstruction loss, and the formula is as follows: L total = L pred + λ1L HSIC + λ2L r where λ1 and λ2 are weight coefficients.
4. The multimodal sentiment analysis method based on deep decoupling and cross-modal semantic alignment according to claim 3, wherein Complete loss L total The formula of which can be transformed into: The λ1 and λ2 are set to 0.1 and 0.01 respectively; the Adam optimizer is used for model training, the initial learning rate is set to 0.0001, the number of samples per batch is 32, and the number of training epochs is 50.
5. The multimodal sentiment analysis method based on deep decoupling and cross-modal semantic alignment according to claim 2, characterized in that The specific features and shared features Feature extraction is performed by two independent Transformer encoders.
6. The multimodal sentiment analysis method based on deep decoupling and cross-modal semantic alignment according to claim 1, characterized in that Adopt a text-guided attention mechanism to perform semantic alignment on the shared features The steps are as follows: Perform positional encoding on the shared feature : where PE is the sine positional encoding, is the feature after positional encoding; Calculate the attention between the text modality and other modalities (including the text modality itself): where is a learnable projection matrix, d is the attention dimension; softmax is the normalized exponential function; Finally obtain the aligned feature h m : where LayerNrom is layer normalization.
7. The multimodal sentiment analysis method based on deep decoupling and cross-modal semantic alignment according to claim 1, characterized in that Align the alignment feature h m , the shared feature after self-attention learning and the unique feature are input into the hierarchical prediction module for prediction, and weighted fusion is performed to obtain the final prediction label including the following steps: The shared features after self-attention learning and the unique features are obtained after pooling and predictions are made to obtain the labels for auxiliary prediction Align the feature h m Obtained after pooling Splicing Perform predictions together to obtain the main prediction label where FCLayer is a fully connected layer; Fuse the three predicted labels with weights to obtain the final predicted label:
8. The multimodal sentiment analysis method based on deep decoupling and cross-modal semantic alignment according to claim 1, wherein Preprocess and extract features from the input video data, including the following steps: Divide the input video data into N frames per second to obtain the audio segments and text descriptions corresponding to the video frame sequence; For the video modality, use the pre-trained FACET model to extract the visual features of each frame; For the audio modality, use the COVAREP toolkit to extract acoustic features; For the text modality, use the BERT model to convert the text into word vector representations; Finally, obtain the preprocessed input features: where \(m\in T\), \(A\), \(V\) represent different modalities, \(N\) m is the sequence length, and \(d\) m is the feature dimension.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the multi-modal sentiment analysis method based on deep decoupling and cross-modal semantic alignment according to any one of claims 1 to 8.
Citation Information
Patent Citations
Multi-modal sentiment analysis method fusing multiple features and attention mechanism
CN116028846A
Multi-modal sentiment analysis method based on multi-agent cooperation
CN118673406A
Cited By
Space-time decoupling sentiment analysis method and system based on multi-modal data
CN120930072A
Multi-modal sentiment analysis method and system based on main modal two-stage guidance
CN121145882A
Multi-modal sentiment analysis method and system based on principal modal two-stage guidance
CN121145882B
Cross-modal semantic alignment method, device and equipment based on fine-grained semantic decoupling
CN121350651A
Cross-modal semantic alignment method, device and equipment based on fine-grained semantic decoupling
CN121350651B