Method and system for endowing large language model with multi-modal sentiment analysis capability
By employing modal preprocessing and feature alignment, fusion, and projection processing of the MSA adapter, the problems of modal noise, distribution differences, and high computational costs in multimodal sentiment analysis are solved, achieving efficient and accurate sentiment analysis suitable for consumer devices.
Patent Information
- Application Number
- CN202511185294.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-11-21
AI Technical Summary
Existing technologies in multimodal sentiment analysis suffer from problems such as modal data noise, distribution differences, information imbalance, high computational cost of large language models, and impaired generalization ability, resulting in low accuracy and efficiency of analysis.
Employing a modality preprocessing module, an MSA adapter, and a frozen large language model body, pseudo-token sequences are generated for sentiment prediction by feature extraction, alignment, fusion, and projection. Only the adapter parameters are updated while the body parameters are frozen. Dynamic attention and multi-scale feature fusion are used to enhance analytical capabilities.
It reduces computational costs, maintains the generalization ability of large language models, improves the accuracy and efficiency of multimodal sentiment analysis, adapts to different datasets, and supports deployment on consumer devices.
Smart Images

Figure CN120995400A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of sentiment analysis technology, specifically to a method and system for endowing large language models with multimodal sentiment analysis capabilities. Background Technology
[0002] The statements in this section are merely background information relating to this disclosure and do not necessarily constitute prior art.
[0003] The development of social media (such as Twitter, TikTok, and YouTube) has led to an explosive growth in multimodal video data, incorporating text, acoustic, and visual elements. This data contains rich emotional information, and traditional single-text sentiment analysis techniques cannot fully extract the emotions contained within the data. Therefore, multimodal sentiment analysis has become the main research direction in the field of sentiment analysis, achieving more accurate sentiment recognition by integrating multimodal data.
[0004] Specifically, text sentiment analysis focuses on the identification and classification of sentiment in text data (such as social media comments and customer feedback), using natural language processing techniques (sentiment dictionaries, machine learning / deep learning models) to categorize sentiment as positive, negative, or neutral.
[0005] Audio sentiment analysis extracts emotional features from audio signals. It first preprocesses the audio signal (denoising and standardization), then extracts prosodic features (speech rate, pitch, and volume changes) and timbre features (vocalization style and resonance characteristics), and then associates them with emotions (e.g., fast speech rate and rising pitch may correspond to excitement).
[0006] Visual sentiment analysis combines computer vision and emotion recognition technologies to extract information such as facial expressions and body language from keyframes of videos and identify the dominant emotion (such as happiness or sadness).
[0007] Multimodal sentiment analysis comprehensively utilizes text, audio, and visual data, extracting and fusing features from each modality to overcome the limitations of a single modality and capture human emotions more comprehensively (e.g., the same text combined with different tones or facial expressions may convey different emotions). It is widely used in scenarios such as social media analysis and human-computer interaction.
[0008] Despite some progress in multimodal sentiment analysis, the following technical challenges still exist: (1) Each modal data contains a lot of irrelevant noise, and the representation distributions of text, audio and visual data are significantly different. Existing domain separation methods (dividing modality invariant data into specific subspaces) cannot completely solve the problem of representation deentanglement, which affects the accuracy of analysis.
[0009] (2) The distribution of emotional information is uneven, and the quality of information in different modalities is inconsistent. Non-textual modalities (audio, visual) often suffer from a lack of emotional information. Existing methods for extracting information from non-textual modalities to enhance textual representation cannot fundamentally solve this problem.
[0010] (3) Although large language models perform well in natural language processing, they have problems such as high computational cost (requiring large-scale pre-training and adaptation) and impaired generalization ability (task training may destroy the original semantic understanding ability) when directly processing multimodal sentiment analysis. Furthermore, the multimodal information fusion is insufficient and the time scale difference is not considered, which further affects the performance. Summary of the Invention
[0011] To address the aforementioned issues, this invention provides a method and system for endowing large language models with multimodal sentiment analysis capabilities. The system primarily comprises a modality preprocessing module, an MSA adapter, and a frozen large language model body. The modality preprocessing module is responsible for converting text, audio, and visual multimodal inputs into unified feature representations: text modalities generate sequence embeddings using the large language model's built-in text embedder; audio and visual modalities extract feature sequences using pre-training tools (such as OpenFace and COVAREP) and convert them into embedding vectors that meet dimensionality requirements.
[0012] The first aspect of this invention provides a method for endowing large language models with multimodal sentiment analysis capabilities, comprising: Acquire text, audio, and visual data of the sample to be tested; Feature extraction is performed on text, audio, and visual data respectively to generate text embedding sequences. T Audio embedding vector A and visual embedding vectors V ; The audio embedding vector A and visual embedding vectors V The input MSA adapter generates a pseudo-token sequence through feature alignment, feature fusion, and projection processing. P ; The pseudo-token sequence is combined with the text embedding sequence and the task prompt embedding sequence. Concatenate into the input sequence S Sentiment prediction is performed using a large language model with frozen parameters.
[0013] Furthermore, the feature alignment step includes: Compress text embedding sequences using global average pooling. T Generate global feature vectors The audio embedding vectors are processed separately using a unidirectional long short-term memory network. A and visual embedding vectors V Output hidden vector and ; Will 、 , Projecting onto a unified dimensional space, we obtain , ; The dynamic weights of text to audio and visual features are calculated based on the intermodal attention mechanism. and And achieve feature alignment according to the following formula: , .
[0014] Furthermore, the feature fusion step includes: Aligned features and Element-wise summation is performed to obtain the initial fusion features. U ; The initial fused features were processed using multilayer perceptrons at three different scales. U Perform multi-scale feature extraction to generate multi-scale features. ; The FiLM-Gate mechanism is introduced to fuse multi-scale features, and then a unified fused feature is generated by compression using 1×1 convolution. .
[0015] Furthermore, the projection processing steps include: Adjusting the fusion features through the first linear layer The dimension that makes it consistent with the text embedding sequence. T Consistent dimensions; The feature dimension is expanded by a second linear layer, generating a preset number of pseudo-token sequences. P It satisfies the formula: .
[0016] Furthermore, feature extraction is performed on the text, audio, and visual data respectively to generate text embedding sequences. T Audio embedding vector A and visual embedding vectors V ,include: Text Embedded Sequence T Generated using the text embedder built into the large language model; Audio embedding vector A 74-dimensional prosodic features were extracted using the COVAREP tool; Visual embedding vectors V 35-dimensional temporal features of facial key points were extracted using the OpenFace tool.
[0017] Furthermore, it also includes the training optimization of the large language model, the steps of which include: The predicted loss for the next token is calculated using an autoregressive approach. Only the parameters of the MSA adapter are updated via backpropagation; the parameters of the main language model remain frozen.
[0018] Furthermore, the input sequence S The splicing order is: ; Sentiment prediction results are generated by a large language model in an autoregressive manner.
[0019] A second aspect of the present invention provides a system for endowing large language models with multimodal sentiment analysis capabilities, comprising: The data acquisition module is used to acquire text, audio, and visual data of the sample to be tested; The feature extraction module is used to extract features from text, audio, and visual data respectively, and generate text embedding sequences. T Audio embedding vector A and visual embedding vectors V ; The pseudo-token sequence generation module is used to embed the audio vector. A and visual embedding vectors V The input MSA adapter generates a pseudo-token sequence through feature alignment, feature fusion, and projection processing. P ; The sentiment prediction module is used to combine the pseudo-token sequence with the text embedding sequence and the task prompt embedding sequence. Concatenate into the input sequence S Sentiment prediction is performed using a large language model with frozen parameters.
[0020] A third aspect of the present invention provides an apparatus for endowing a large language model with multimodal sentiment analysis capabilities, the apparatus comprising a memory and a processor; the memory for storing a computer program; and the processor for implementing the above-described method for endowing a large language model with multimodal sentiment analysis capabilities when the computer program is executed.
[0021] A fourth aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for endowing large language models with multimodal sentiment analysis capabilities.
[0022] Compared with existing technologies, the method and system provided by this invention, which endow large language models with multimodal sentiment analysis capabilities, have the following beneficial effects: (1) Addressing the technical problem of high computational cost and difficulty in deployment on consumer devices when using large language models for multimodal sentiment analysis in existing technologies, this invention achieves the connection between non-textual modalities and large language models through the MSA adapter. During training, only the adapter parameters are updated, while the main parameters of the large language model are frozen. This eliminates the need for full pre-training of the large language model, optimizing only a small number of adapter parameters, significantly reducing computational load. The small training parameter size allows deployment on consumer-grade GPUs, significantly reducing computational resource requirements.
[0023] (2) To maintain the generalization ability of the large language model, this invention freezes the main parameters of the large language model to prevent its generalization ability from being impaired by task training, thus achieving "plug and play". By freezing the main parameters of the large language model, its semantic space is adapted only through the MSA adapter, thus preventing the large language model from being "contaminated" by task training and preserving its original cross-task semantic understanding ability.
[0024] (3) The feature alignment module provided by this invention realizes explicit association between text and non-text features through dynamic attention weights; the feature fusion module adopts multi-scale MLP and FiLM-Gate mechanism, combined with 1×1 convolution compression, to improve the efficiency of multi-scale information fusion. The dynamic alignment module reduces the modality distribution gap, and the multi-scale fusion module fully explores the deep interaction relationship of each modality to reduce information loss.
[0025] (4) Text embedding relies on the built-in tools of the large language model, while audio / visual feature extraction uses general pre-training tools (COVAREP, OpenFace) and supports Chinese and English datasets (such as MOSEI, CHERMA). The modular design reduces the dependence on specific data formats and adapts to different modality preprocessing tools and datasets. Attached Figure Description
[0026] The accompanying drawings, which form part of this disclosure, are used to provide a further understanding of this disclosure. The illustrative embodiments of this disclosure and their descriptions are used to explain this disclosure and do not constitute an undue limitation of this disclosure.
[0027] Figure 1 This is a flowchart of the steps in the method for endowing a large language model with multimodal sentiment analysis capabilities, as provided in Embodiment 1 of the present invention. Figure 2 This is a schematic diagram of the method for endowing large language models with multimodal sentiment analysis capabilities provided in Embodiment 1 of the present invention; Figure 3 yes Figure 2 A schematic diagram of the MSA adapter in the diagram; Figure 4 yes Figure 2 A diagram illustrating the text embedder in the text; Figure 5 yes Figure 2 A schematic diagram of the structure of the large language model in China; Figure 6 This is a schematic diagram of a system that endows a large language model with multimodal sentiment analysis capabilities, as provided in Embodiment 2 of the present invention. Detailed Implementation
[0028] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0029] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments of the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. Furthermore, it should be understood that the terms “comprising” and “having”, and any variations thereof, are intended to cover non-exclusive inclusion, for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0030] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0031] All data acquisition in this embodiment is carried out in accordance with laws and regulations and with user consent, and the data is used legally.
[0032] Example 1 like Figure 1 As shown, this invention provides a method for endowing large language models with multimodal sentiment analysis capabilities, including: Acquire text, audio, and visual data of the sample to be tested; Feature extraction is performed on text, audio, and visual data respectively to generate text embedding sequences. T Audio embedding vector A and visual embedding vectors V ; The audio embedding vector A and visual embedding vectors V The input MSA adapter generates a pseudo-token sequence through feature alignment, feature fusion, and projection processing. P ; The pseudo-token sequence is combined with the text embedding sequence and the task prompt embedding sequence. Concatenate into the input sequence S Sentiment prediction is performed using a large language model with frozen parameters.
[0033] This invention defines a comprehensive process for endowing large language models with multimodal sentiment analysis capabilities, including modal feature extraction (text, audio, and visual), pseudo-token generation (MSA adapter processing), large language model inference (concatenating the input sequence and making predictions), and training optimization (updating adapter parameters only). Through modular design, modal preprocessing, MSA adapter, and freezing of the large language model, seamless integration of multimodal data and the large language model is achieved, preserving the generalization ability of the large language model while reducing computational costs, and providing a basic framework for subsequent feature alignment and fusion.
[0034] Specifically, the feature alignment step includes: Compress text embedding sequences using global average pooling. T Generate global feature vectors The audio embedding vectors are processed separately using a unidirectional long short-term memory network. A and visual embedding vectors V Output hidden vector and ; Will 、 , Projecting onto a unified dimensional space, we obtain , ; The dynamic weights of text to audio and visual features are calculated based on the intermodal attention mechanism. and And achieve feature alignment according to the following formula: , .
[0035] This invention, based on the processing logic of the feature alignment module in the MSA adapter, includes compressing text features through global average pooling, modeling audio / visual temporal features using a unidirectional long short-term memory network (sLSTM), unifying dimensions through linear projection, and calculating dynamic weights through an intermodal attention mechanism, ultimately achieving dynamic alignment between audio / visual features and text features. This solves the technical problem of "large modal distribution gaps" in the background technology by enhancing the explicit association between text and non-text features through dynamic weights (adjusted according to context), reducing the interference of modal differences on analysis, and improving cross-modal feature consistency.
[0036] Specifically, the feature fusion steps include: Aligned features and Element-wise summation is performed to obtain the initial fusion features. U ; The initial fused features were processed using multilayer perceptrons at three different scales.U Perform multi-scale feature extraction to generate multi-scale features. ; The FiLM-Gate mechanism is introduced to fuse multi-scale features, and then a unified fused feature is generated by compression using 1×1 convolution. .
[0037] The defined feature fusion module's processing flow includes element-wise summation of aligned audio / visual features (initial fusion), feature extraction via multi-scale MLP, FiLM-Gate fusion, and 1×1 convolution compression to generate unified fused features. This addresses the issue of insufficient multimodal information fusion. Multi-scale MLP captures information at different granularities, the FiLM-Gate mechanism balances fusion stability and detail preservation, and 1×1 convolution optimizes feature dimensions, making the fused information more comprehensive and adaptable to subsequent processing, thus improving the accuracy of sentiment analysis.
[0038] Specifically, the projection processing steps include: Adjusting the fusion features through the first linear layer The dimension that makes it consistent with the text embedding sequence. T Consistent dimensions; The feature dimension is expanded by a second linear layer, generating a preset number of pseudo-token sequences. P It satisfies the formula: .
[0039] This invention defines the processing logic of the projection module, which sequentially adjusts the fusion feature dimension (consistent with text embedding) and expands the generated pseudo-token sequence through two linear layers to achieve the conversion of non-text features into a format understandable by a large language model. This solves the problem that "non-textual modalities cannot be directly recognized by large language models." The pseudo-token acts as a "bridge" to "translate" audio / visual emotional information into a form interpretable by the large language model, ensuring that multimodal information is effectively input into the model and providing support for accurate prediction.
[0040] Specifically, the step involves extracting features from text, audio, and visual data respectively to generate text embedding sequences. T Audio embedding vector A and visual embedding vectors V ,include: Text Embedded Sequence T Through the text embeddings built into the large language model (such as...) Figure 4 (As shown) generated; Audio embedding vector A 74-dimensional prosodic features were extracted using the COVAREP tool; Visual embedding vectors V35-dimensional temporal features of facial key points were extracted using the OpenFace tool.
[0041] This invention specifies the tools and parameters for text, audio, and visual feature extraction. Text embeddings are generated using tools built into a large language model; audio features are extracted using COVAREP (74-dimensional prosodic features); and visual features are extracted using OpenFace (35-dimensional temporal features of facial key points). A standardized modality preprocessing workflow ensures the stability and compatibility of feature extraction, supports different datasets (such as MOSEI and CHERMA), and improves the method's versatility and reproducibility.
[0042] Specifically, it also includes the training optimization of large language models, the steps of which include: The predicted loss for the next token is calculated using an autoregressive approach. Only the parameters of the MSA adapter are updated via backpropagation; the parameters of the main language model remain frozen.
[0043] This invention defines the training optimization logic, uses an autoregressive approach to calculate the prediction loss for the next token, updates only the MSA adapter parameters, and freezes the main parameters of the large language model. This addresses the problem of "impaired generalization ability of large language models" by freezing the main parameters to prevent the semantic space from being corrupted, optimizing only the adapter parameters to reduce computational costs, and ensuring that the adapter is compatible with the semantic space of the large language model, achieving "plug and play".
[0044] Specifically, the input sequence S The splicing order is: ; Sentiment prediction results are generated by a large language model in an autoregressive manner.
[0045] This invention clarifies the concatenation order of the input sequence S (pseudo-token P; text embedding T; task prompt embedding Tprompt) and restricts sentiment prediction to be generated by a large language model in an autoregressive manner. It achieves a standardized input format, ensuring that multimodal information (non-text pseudo-tokens, text, and task instructions) is input into the model in an orderly manner. The autoregressive prediction method adapts to the generation logic of the large language model, improving the coherence and accuracy of sentiment prediction.
[0046] In one specific embodiment, the input consists of information from audio, text, and video modalities, along with a specific prompt, and the output is the sentiment predicted by a large language model. Its core lies in achieving efficient processing of multimodal data and seamless integration with a large language model through modular design. The entire process covers the complete chain from raw data input to sentiment result output, with each stage working in close collaboration to ensure the accuracy and computational efficiency of sentiment analysis.
[0047] In the modal feature preprocessing stage, the system first standardizes the input text, visual, and audio data to lay the foundation for subsequent cross-modal processing. For the text modality, the system directly utilizes the built-in text embedding tool of the large language model (such as a word segmentation tool based on SentencePiece) to perform tokenization, converting natural language into fixed-dimensional sequence embedding vectors T and T'. prompt The visual modality was processed using the OpenFace toolkit to extract facial keypoint features, generating a 35-dimensional temporal feature sequence. The audio modality was processed using the COVARE toolkit to extract 74-dimensional prosodic features. After preprocessing, the features of all three modalities were stored in a unified tensor format to ensure direct access by subsequent modules.
[0048] like Figure 3 The MSA adapter, as the core of this invention, achieves feature alignment and fusion between non-textual and textual modalities through a three-layer progressive processing, and generates pseudo-tokens that can be understood by a large language model. The three modules of the MSA adapter are described in detail below: feature alignment module, feature fusion module, and projection module.
[0049] I. Feature Alignment Module The feature alignment module is crucial for achieving multimodal feature alignment, and its core lies in establishing explicit associations between text and non-text features. First, the text embedding T is compressed using global average pooling (GAP) to generate a global feature vector. The vector integrates the overall sentiment of the text; then comes the temporal modeling stage using a one-way long short-term memory network. The system deploys independent one-way long short-term memory networks (sLSTM) for the temporal characteristics of visual and audio features respectively. The visual sLSTM parses the embedded vector V frame by frame, capturing the changes in facial expressions over time and outputting the hidden vector. The audio sLSTM performs temporal encoding on the embedded vector A to extract dynamic features such as intonation and speech rate, generating a hidden vector. Finally, through three independent linear projection layers, the... , , Mapping to a unified dimension yields feature vectors. , , .
[0050]
[0051]
[0052]
[0053] Where the projection matrix W t Wv W a The dimensions are h×d respectively. t h×h v h×h a This ensures that features of different modalities are comparable in the same dimensional space.
[0054] Then embed the text sequence Context-aware text features are generated using a single Transformer encoder (with self-attention at both ends). And obtained through global average pooling It is used to capture dynamic changes in the semantics of text.
[0055] Then, through an intermodal attention mechanism, the dynamic weights of the text on visual and audio features are calculated:
[0056]
[0057] in , Here, σ is a learnable parameter, and σ is the sigmoid function with output weights ranging from [0,1]. Similarly, , It is also a learnable parameter.
[0058] Finally, dynamic feature alignment is achieved based on dynamic weights:
[0059]
[0060] in and It will adaptively adjust to changes in the context of the input sample: when the text semantics are strongly correlated with visual features, A value close to 1 increases the contribution of visual features; conversely, a value far less than 1 decreases the weight.
[0061] II. Feature Fusion Module Features processed by the feature alignment module enter the feature fusion module, firstly... and Element-wise summation is performed to obtain the initial fused feature U; then, multi-scale feature extraction is performed on U through three parallel MLP networks:
[0062] This step involves extracting and fusing features of U at different scales. Different values allow the MLP to capture information at different granularities. This module also introduces the FiLM-Gate mechanism in feature fusion, which combines the fine-grained modulation capability of FiLM with the stability and information preservation capability of residual gating, thereby achieving information-rich, stable feature fusion that preserves the original details.
[0063] Finally, the fusion results from the three different scales are stacked together to obtain... Then, a 1×1 convolution (Conv) is used to compress the information, finally yielding the fused result. This makes it more suitable for the requirements of subsequent large language models.
[0064] III. Projection Module The projection module, acting as a bridge between the MSA adapter and the large language model, is responsible for converting fused features into pseudo-tokens that conform to the input format of the large language model, thereby enabling effective interaction between the non-textual modality and the large language model. After passing through the feature alignment and feature fusion modules, although key emotional features from both visual and audio sources have been integrated and a certain degree of alignment with the textual modality has been achieved at the feature level, their dimensions and forms differ from the text embeddings that the large language model can process, making them unrecognizable directly. Therefore, further processing and transformation are required through the projection module.
[0065] The projection module consists of two linear layers, which function sequentially. The primary function of the first linear layer is to adjust the fused information. The feature dimension is adjusted to have the same dimension as the sequence embedding T of the text modality. After dimension adjustment, the second linear layer expands the processed features into n pseudo-tokens P according to the preset number of pseudo-tokens.
[0066]
[0067] These pseudo-tokens are a form of "language" that the large language model can understand. They condense emotional information from the visual and audio modalities, essentially "translating" non-textual content into a token sequence that the large language model can interpret. Through this process, visual and audio information that is originally difficult for the large language model to process directly is successfully transformed into an input form that the large language model can accept, thanks to the pseudo-tokens generated by the projection module. This provides the necessary non-textual information support for the large language model to perform multimodal sentiment analysis and sentiment recognition tasks. The structure of the large language model is as follows: Figure 5 As shown.
[0068] The final generated pseudo-token P will be combined with the sequence embedding T of the text modality and the sequence embedding T of the task prompt. promptThe sequences are concatenated to form a sequence that is input into the large language model. This enables large language models with frozen parameters to utilize these inputs that integrate multimodal information to complete the autoregressive generation task of sentiment labels, thereby achieving accurate analysis of multimodal sentiment.
[0069] Example 2 like Figure 6 As shown, this embodiment provides a system that endows large language models with multimodal sentiment analysis capabilities, including: The data acquisition module is used to acquire text, audio, and visual data of the sample to be tested; The feature extraction module is used to extract features from text, audio and visual data respectively, and generate text embedding sequence T, audio embedding vector A and visual embedding vector V; The pseudo token sequence generation module is used to input the audio embedding vector A and the visual embedding vector V into the MSA adapter, and generate a pseudo token sequence P through feature alignment, feature fusion and projection processing in sequence; The sentiment prediction module is used to concatenate the pseudo-token sequence with the text embedding sequence and the task prompt embedding sequence to form an input sequence S, and input the frozen parameters to a large language model for sentiment prediction.
[0070] Example 3 This embodiment provides a device for endowing a large language model with multimodal sentiment analysis capabilities. The device includes a memory and a processor. The memory is used to store computer programs. The processor is used to implement the above-described method for endowing a large language model with multimodal sentiment analysis capabilities when the computer programs are executed.
[0071] The processor is connected to the memory, and one or more computer programs are stored in the memory. When the electronic device is running, the processor executes one or more computer programs stored in the memory to cause the electronic device to perform the method described in Embodiment 1.
[0072] It should be understood that in this embodiment, the processor can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0073] Memory may include read-only memory and random access memory, and provides instructions and data to the processor. A portion of memory may also include non-volatile random access memory. For example, memory may also store information about the device type.
[0074] In the implementation process, each step of the above method can be completed by the integrated logic circuits in the processor hardware or by software instructions.
[0075] The method in Embodiment 1 can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor. The software modules can reside in readily available storage media in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, a detailed description is not provided here.
[0076] Those skilled in the art will recognize that the units and algorithm steps described in connection with the various examples of this embodiment can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this invention.
[0077] Example 4 In another embodiment of the present invention, a computer-readable storage medium stores a computer program, which, when executed by a processor, implements the method described above for endowing a large language model with multimodal sentiment analysis capabilities.
[0078] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc. In this application, the units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of the present invention according to actual needs. Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units can be implemented in hardware or as software functional units.
[0079] While the present invention has been disclosed above, its scope of protection is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention, and all such changes and modifications will fall within the scope of protection of the present invention.
Claims
1. A method for endowing large language models with multimodal sentiment analysis capabilities, characterized in that, include: Acquire text, audio, and visual data of the sample to be tested; Feature extraction is performed on text, audio, and visual data respectively to generate text embedding sequences. T Audio embedding vector A and visual embedding vectors V ; The audio embedding vector A and visual embedding vectors V The input MSA adapter generates a pseudo-token sequence through feature alignment, feature fusion, and projection processing. P ; The pseudo-token sequence is combined with the text embedding sequence and the task prompt embedding sequence. Concatenate into the input sequence S Sentiment prediction is performed using a large language model with frozen parameters.
2. The method as described in claim 1, characterized in that, The feature alignment step includes: Compress text embedding sequences using global average pooling. T Generate global feature vectors The audio embedding vectors are processed separately using a unidirectional long short-term memory network. A and visual embedding vectors V Output hidden vector and ; Will 、 , Projecting onto a unified dimensional space, we obtain , ; The dynamic weights of text to audio and visual features are calculated based on the intermodal attention mechanism. and And achieve feature alignment according to the following formula: , 。 3. The method as described in claim 2, characterized in that, The feature fusion steps include: Aligned features and Element-wise summation is performed to obtain the initial fusion features. U ; The initial fused features were processed using multilayer perceptrons at three different scales. U Perform multi-scale feature extraction to generate multi-scale features. ; The FiLM-Gate mechanism is introduced to fuse multi-scale features, and then a unified fused feature is generated by compression using 1×1 convolution. .
4. The method as described in claim 3, characterized in that, The projection processing steps include: Adjusting the fusion features through the first linear layer The dimension that makes it consistent with the text embedding sequence. T Consistent dimensions; The feature dimension is expanded by a second linear layer, generating a preset number of pseudo-token sequences. P It satisfies the formula: 。 5. The method as described in claim 1, characterized in that, The text, audio, and visual data are each subjected to feature extraction to generate a text embedding sequence. T Audio embedding vector A and visual embedding vectors V ,include: Text Embedded Sequence T Generated using the text embedder built into the large language model; Audio embedding vector A 74-dimensional prosodic features were extracted using the COVAREP tool; Visual embedding vectors V 35-dimensional temporal features of facial key points were extracted using the OpenFace tool.
6. The method as described in claim 1, characterized in that, It also includes training optimization of large language models, the steps of which include: The predicted loss for the next token is calculated using an autoregressive approach. Only the parameters of the MSA adapter are updated via backpropagation; the parameters of the main language model remain frozen.
7. The method as described in claim 1, characterized in that, The input sequence S The splicing order is: ; Sentiment prediction results are generated by a large language model in an autoregressive manner.
8. A system that endows large language models with multimodal sentiment analysis capabilities, characterized in that, include: The data acquisition module is used to acquire text, audio, and visual data of the sample to be tested; The feature extraction module is used to extract features from text, audio, and visual data respectively, and generate text embedding sequences. T Audio embedding vector A and visual embedding vectors V ; The pseudo-token sequence generation module is used to embed the audio vector. A and visual embedding vectors V The input MSA adapter generates a pseudo-token sequence through feature alignment, feature fusion, and projection processing. P ; The sentiment prediction module is used to combine the pseudo-token sequence with the text embedding sequence and the task prompt embedding sequence. Concatenate into the input sequence S Sentiment prediction is performed using a large language model with frozen parameters.
9. A device for endowing large language models with multimodal sentiment analysis capabilities, characterized in that, The device includes a memory and a processor; the memory is used to store a computer program; the processor is used to implement, when executing the computer program, the method for endowing a large language model with multimodal sentiment analysis capabilities as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program that, when executed by a processor, implements the method for endowing a large language model with multimodal sentiment analysis capabilities as described in any one of claims 1 to 7.