Emotion recognition method, multi-modal large model training method, equipment and storage medium
By processing multimodal dialogue data using a large multimodal model to obtain emotion labels and emotion reasons, and training with independent low-rank adaptive fine-tuning layers, the problem of the large multimodal model affecting language interaction during emotion recognition is solved, thereby improving emotional expression and empathy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HONOR DEVICE CO LTD
- Filing Date
- 2024-10-25
- Publication Date
- 2026-04-28
AI Technical Summary
Existing multimodal large models affect language interaction capabilities during emotion recognition, and can only recognize emotion tags under specific prompts, failing to achieve emotional expression and resulting in a poor user experience.
Multimodal dialogue data is processed by a large multimodal model to obtain emotion labels and emotional reasons. Emotional expression is then performed based on the emotion labels. The model is trained using an independent low-rank adaptive fine-tuning layer to ensure that language interaction capabilities are not affected.
It improves the user experience, enables accurate emotion recognition and emotional expression during language interaction, and enhances the empathy capabilities of electronic devices.
Smart Images

Figure CN121935673A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of smart terminals, and in particular to an emotion recognition method, a multimodal large model training method, a device, and a storage medium. Background Technology
[0002] With the development of electronic device technology, artificial intelligence assistants (such as voice assistants) in electronic devices have become an indispensable part of people's lives. To provide users with high-quality emotional interaction and companionship through AI assistants, these assistants need not only strong language interaction capabilities but also empathy, i.e., the ability to recognize emotions. Currently, the main approach is to fine-tune large-scale language models using Low-Rank Adaptation of Large Language Models (LoRA) on the multimodal models used for language interaction, in order to achieve emotion recognition simultaneously with language interaction.
[0003] However, the method of fine-tuning the multimodal large model using LoRA not only affects or even severely damages the language interaction ability of the multimodal large model, but can only identify the emotion labels used to represent the user's emotions under specific prompts. It cannot achieve emotion recognition while interacting with the user, and therefore cannot express emotions based on the identified user emotions to achieve empathetic language interaction, resulting in a poor user experience. Summary of the Invention
[0004] This application provides an emotion recognition method, a multimodal large model training method, a device, and a storage medium. By processing multimodal dialogue data through a multimodal large model, emotion tags and emotion reasons can be obtained. Then, based on the emotion tags, the dialogue response content can be expressed emotionally to provide users with better emotional interaction and emotional companionship, thereby improving the user experience.
[0005] To achieve the above objectives, the embodiments of this application adopt the following technical solutions:
[0006] Firstly, this application provides an emotion recognition method applied to an electronic device, the emotion recognition method comprising:
[0007] Acquire multimodal dialogue data;
[0008] The multimodal dialogue data is processed based on the encoder and adapter in the multimodal large model to obtain the input features of the large language model in the multimodal large model;
[0009] The input features are processed by multiple decoding layers in the large language model to obtain the dialogue response text and multiple intermediate features; multiple low-rank adaptive fine-tuning layers in the large language model are used to process the multiple intermediate features to obtain emotion labels to represent user emotions and emotion reason text to represent the reasons for emotions.
[0010] The emotion recognition method provided in this application, because the multimodal large language model has the function of identifying emotion labels and emotion causes, can accurately determine emotion labels and emotion causes by processing the acquired multimodal data through the multimodal large model. Furthermore, based on the emotion labels, the dialogue response text can be expressed emotionally to provide users with better emotional interaction and emotional companionship, thereby improving the user experience. At the same time, the emotion labels and emotion cause text can guide the generation of subsequent dialogue response text, emotion labels, and emotion causes, which is beneficial to further improving the empathy capabilities of electronic devices.
[0011] In one possible design approach of the first aspect, the input features are processed based on multiple decoding layers in a large language model to obtain dialogue response text and multiple intermediate features, including: processing the input features based on multiple first decoding layers to obtain a first intermediate feature among the intermediate features; processing the first intermediate feature based on multiple second decoding layers to obtain multiple second intermediate features among the intermediate features; performing text conversion processing on the last second intermediate feature among the multiple second intermediate features to obtain the dialogue response text; wherein the last second intermediate feature corresponds to the last second decoding layer among the multiple second decoding layers.
[0012] Based on the above technical solution, since the weight parameters of the multiple first and second decoding layers in the large speech model are the same as those in the large language model before adding the low-rank adaptive fine-tuning layer, processing the input features through the multiple first and second decoding layers in the large language model will not affect the first intermediate features output by the multiple first decoding layers and the multiple second intermediate features output by the multiple second decoding layers. Furthermore, performing text conversion processing on the last of the multiple second intermediate features can accurately obtain the dialogue response text.
[0013] In one possible design approach of the first aspect, multiple intermediate features are processed based on multiple low-rank adaptive fine-tuning layers in a large language model to obtain an emotion label representing the user's emotion and an emotion reason text representing the emotion reason. This includes: processing a first intermediate feature and multiple second intermediate features based on multiple low-rank adaptive fine-tuning layers in a large language model to obtain multiple fine-tuning features; performing text transformation processing on the last fine-tuning feature among the multiple fine-tuning features to obtain the emotion label and the emotion reason text; wherein, the last fine-tuning feature corresponds to the last low-rank adaptive fine-tuning layer among the multiple low-rank adaptive fine-tuning layers.
[0014] Based on the above scheme, since the multiple low-rank adaptive fine-tuning layers in the large language model are independent neural network layers, processing the first intermediate feature and multiple second intermediate features through multiple low-rank adaptive fine-tuning layers does not affect the processing of the first intermediate feature by multiple second decoding layers, nor is it affected by the processing of the first intermediate feature by multiple second decoding layers. Therefore, the multiple fine-tuned features obtained by processing the first intermediate feature and multiple second intermediate features through multiple low-rank adaptive fine-tuning layers are not coupled with the multiple second intermediate features, i.e., they are completely unrelated. Furthermore, performing text conversion processing on the last fine-tuned feature among the multiple fine-tuned features can yield accurate emotion labels and emotion reason text.
[0015] In one possible design of the first aspect, the plurality of second decoding layers include the first second decoding layer to the mth second decoding layer; the mth second decoding layer is the last second decoding layer among the plurality of second decoding layers; the plurality of second intermediate features include the first second intermediate feature to the mth second intermediate feature; the first intermediate feature is processed based on the plurality of second decoding layers among the plurality of decoding layers to obtain the plurality of second intermediate features among the intermediate features, including: processing the first intermediate feature based on the first second decoding layer to obtain the first second intermediate feature; processing the (i-1)th second intermediate feature based on the i-th second decoding layer to obtain the i-th second intermediate feature; where i is an integer greater than 1 and less than or equal to m.
[0016] Based on the above scheme, since the weight parameters of multiple second decoding layers in the large speech model are the same as the weight parameters of multiple second decoding layers in the large language model before the low-rank adaptive fine-tuning layer is added, multiple second intermediate features can be accurately obtained when the first intermediate features are processed sequentially by multiple second decoders.
[0017] In one possible design approach of the first aspect, multiple low-rank adaptive fine-tuning layers include the first low-rank adaptive fine-tuning layer to the p-th low-rank adaptive fine-tuning layer; the p-th low-rank adaptive fine-tuning layer is the last low-rank adaptive fine-tuning layer among the multiple low-rank adaptive fine-tuning layers; multiple fine-tuning features include the first fine-tuning feature to the j-th fine-tuning feature; where j is an integer greater than 1 and less than or equal to p; the first intermediate feature and multiple second intermediate features are processed based on the multiple low-rank adaptive fine-tuning layers in the large language model to obtain multiple fine-tuning features, including: processing the first intermediate feature based on the first low-rank adaptive fine-tuning layer to obtain the first fine-tuning feature; determining the j-1th comprehensive feature based on the (j-1)th fine-tuning feature and the (j-1)th second intermediate feature; and processing the (j-1)th comprehensive feature based on the j-th low-rank adaptive fine-tuning layer to obtain the j-th fine-tuning feature.
[0018] Based on the above scheme, since the multiple low-rank adaptive fine-tuning layers in the large language model are independent neural network layers, when processing the first intermediate feature through multiple low-rank adaptive fine-tuning layers, multiple fine-tuning features that are not related to multiple second intermediate features can be accurately obtained.
[0019] In one possible design approach of the first aspect, the (j-1)th comprehensive feature is determined based on the (j-1)th fine-tuning feature and the (j-1)th second intermediate feature, including: performing an addition operation on the (j-1)th fine-tuning feature and the (j-1)th second intermediate feature to obtain the (j-1)th feature sum; and determining the (j-1)th feature sum as the (j-1)th comprehensive feature.
[0020] Based on this scheme, comprehensive features related to the second intermediate feature can be obtained, which can facilitate the processing of comprehensive features to accurately determine the emotion label and the emotion reason text.
[0021] Secondly, this application provides a multimodal large model training method, which includes: acquiring multiple sets of multimodal sample data and sample emotion labels and sample emotion reason texts corresponding to each set of multimodal sample data; processing the multimodal sample data based on the encoder and adapter in the initial multimodal large model to obtain sample input features of the initial large language model in the initial multimodal large model; processing the sample input features based on multiple decoding layers in the initial large language model to obtain multiple prediction intermediate features; processing the multiple prediction intermediate features based on multiple initial low-rank adaptive fine-tuning layers in the initial large language model to obtain predicted emotion labels and predicted emotion reason texts; using the predicted emotion labels and predicted emotion language texts as the initial training output of the initial multimodal large model, and the sample emotion labels and sample emotion reason texts as supervision information, iteratively training the initial multimodal large model to obtain the trained multimodal large model.
[0022] The multimodal large model training method provided in this application uses predicted emotion labels and predicted emotion language text as the initial training output of the initial multimodal large model, and sample emotion labels and sample emotion reason text as supervision information to iteratively train the initial multimodal large model. Therefore, it is possible to obtain a multimodal large model with emotion and emotion reason recognition functions.
[0023] In one possible design approach of the first aspect, the predicted emotion label and the predicted emotion language text are used as the initial training output of the initial multimodal large model, and the sample emotion label and the sample emotion reason text are used as supervision information. The initial multimodal large model is iteratively trained to obtain the trained multimodal large model, including: determining the loss value of the multimodal large model based on the predicted emotion label, the predicted emotion language text, the sample emotion label, and the sample emotion reason text; and iteratively updating the weight parameters in multiple initial low-rank adaptive fine-tuning layers based on the loss value to obtain the trained multimodal large model.
[0024] Based on this scheme, the loss value of the multimodal large model is determined only based on the predicted sentiment labels, predicted sentiment reason text, sample sentiment labels, and sample sentiment reason text corresponding to multiple low-rank adaptive fine-tuning layers. The weight parameters of only the multiple low-rank adaptive fine-tuning layers are iteratively updated based on the loss value; the weight parameters of other neural network layers in the multimodal large model are not updated. That is, other neural network layers besides the multiple low-rank adaptive fine-tuning layers are not trained. Therefore, the computational cost of model training can be reduced, and the efficiency of model training can be improved.
[0025] Thirdly, this application provides an electronic device comprising: a display screen, a memory, and one or more processors; the display screen, the memory, and the processors are coupled; wherein the memory stores computer program code, the computer program code including computer instructions, which, when executed by the processor, cause the electronic device to perform the emotion recognition method provided in the first aspect and any of its possible design embodiments.
[0026] Fourthly, this application provides a training device, comprising:
[0027] Processor; memory used to store processor-executable instructions;
[0028] The processor is configured to execute executable instructions to implement the multimodal large model training method provided in the second aspect and any of its possible design schemes.
[0029] Fifthly, this application provides a computer-readable storage medium including computer instructions that, when executed on an electronic device, cause the electronic device to perform the emotion recognition method provided in the first aspect and any of its possible design embodiments.
[0030] In a sixth aspect, this application provides a computer-readable storage medium including computer instructions that, when executed on a training device, cause the training device to perform the multimodal large model training method provided in the second aspect and any of its possible design embodiments.
[0031] In a seventh aspect, this application provides a computer program product comprising executable instructions that, when the computer program product is run on an electronic device, cause the electronic device to perform the emotion recognition method provided in the first aspect and any of its possible design embodiments.
[0032] Eighthly, an apparatus (e.g., a system-on-a-chip) is provided, comprising a processor for supporting an electronic device in performing the functions described in the second aspect above. In one possible design, the apparatus further comprises a memory for storing program instructions and data necessary for the electronic device. When the apparatus is a system-on-a-chip, it may be composed of chips or may include chips and other discrete devices.
[0033] It is understood that the beneficial effects that the technical solutions provided in the second to eighth aspects above can achieve can be referred to the beneficial effects in the first aspect and any of its possible design methods, and will not be repeated here. Attached Figure Description
[0034] Figure 1 A schematic diagram of the hardware architecture of an electronic device provided in an embodiment of this application;
[0035] Figure 2 A flowchart illustrating an emotion recognition method provided in an embodiment of this application;
[0036] Figure 3 A schematic diagram of the display desktop of an electronic device provided in an embodiment of this application;
[0037] Figure 4 This is a schematic diagram of the structure of a multimodal large model provided in an embodiment of this application;
[0038] Figure 5 A schematic diagram illustrating the principle of a large speech model provided in an embodiment of this application;
[0039] Figure 6 A flowchart illustrating a multimodal large model training method provided in this application embodiment;
[0040] Figure 7 A flowchart illustrating another emotion recognition method provided in an embodiment of this application;
[0041] Figure 8A flowchart illustrating another multimodal large model training method provided in this application embodiment;
[0042] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0043] Figure 10 This is a schematic diagram of the structure of a training device provided in an embodiment of this application;
[0044] Figure 11 This is a schematic diagram of a chip system provided in an embodiment of this application. Detailed Implementation
[0045] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that “ / ” means “or,” for example, A / B can mean A or B; “and / or” in the text is merely a description of the relationship between related objects, indicating that three relationships can exist, for example, A and / or B can mean: A alone, A and B simultaneously, and B alone.
[0046] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application can be combined with other embodiments.
[0047] The terms "first" and "second" in the following embodiments of this application are for descriptive purposes only and should not be construed as implying relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.
[0048] To facilitate understanding of this application, the terms used in this application are explained below.
[0049] Multimodal large models refer to artificial intelligence models that can process and understand multiple types of data, typically including data in different modalities such as text, images, and audio.
[0050] Large Language Models (LLMs) are an important technology in the field of artificial intelligence. They are trained with large amounts of data and are designed to understand and generate natural language text.
[0051] Low-Rank Adapation (LoRA) is a method for fine-tuning LLMs. It approximates the model's weight updates by applying a low-rank decomposition matrix to the LLM's weight matrix, rather than directly updating the original high-dimensional weight matrix. This significantly reduces the number of parameters in the LLM, thereby lowering computational complexity and memory requirements.
[0052] A text-to-speech (TTS) module is a hardware device or software component that can convert digital information or text content into speech signals.
[0053] In related technologies, electronic devices can converse with users through voice assistants. Taking electronic devices, including mobile phones, as an example, users can input a wake-up voice message for the voice assistant, such as "Hello, YOYO!" The electronic device can respond to this wake-up voice message, activating the voice assistant and outputting a response voice message, "Hello, how can I help you?", thus initiating a human-computer dialogue or chat. Furthermore, to achieve emotional interaction between humans and computers, i.e., empathy, it is necessary to recognize the user's emotions during language interaction, generate language interaction content based on the recognized user emotions, and express the language interaction content emotionally. For example, based on different user emotions (e.g., sadness, happiness, worry, etc.), different interaction strategies (e.g., comfort, encouragement, advice, humor, etc.) can be provided for the user's input dialogue content, corresponding to different dialogue responses, expressed through different timbre or tone of voice.
[0054] Currently, empathic interaction between humans and computers is mainly achieved by fine-tuning the multimodal large model used for language interaction using LoRA. However, this method of fine-tuning the multimodal large model using LoRA not only affects or even severely damages the language interaction capabilities of the multimodal large model, but also can only identify emotion tags used to express emotions under specific prompts. It cannot achieve emotion recognition while interacting with language, and therefore cannot express emotions based on the identified user emotions to achieve empathic language interaction, resulting in a poor user experience.
[0055] To address the aforementioned technical issues, this application provides an emotion recognition method applied to an electronic device that provides a voice assistant. Upon activating the voice assistant, in response to user-input dialogue (such as voice chat), the electronic device acquires data from different modalities and processes this data using a multimodal large model. This allows it to determine emotion tags and emotional causes, generate appropriate dialogue responses based on these tags, and express these responses emotionally based on the emotion tags. This provides users with higher-quality emotional interaction and companionship, enhancing the user experience.
[0056] For example, the aforementioned electronic devices may be mobile phones, tablets, desktop computers, laptops, handheld computers, notebook computers, ultra-mobile personal computers (UMPCs), netbooks, as well as voice assistant devices such as cellular phones, personal digital assistants (PDAs), augmented reality (AR) / virtual reality (VR) devices. This application does not impose special limitations on the specific form of the electronic device. The following description uses a mobile phone as an example.
[0057] Reference Figure 1 As shown, the electronic device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, antenna 1, antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a display screen 193, a subscriber identification module (SIM) card interface 194, and a camera 195, etc. The sensor module 180 may include pressure sensors, gyroscope sensors, barometric pressure sensors, magnetic sensors, accelerometers, distance sensors, proximity sensors, fingerprint sensors, temperature sensors, touch sensors, ambient light sensors, bone conduction sensors, etc.
[0058] Processor 110 may include one or more processing units, such as an access point (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). These different processing units may be independent devices or integrated into one or more processors.
[0059] A controller can be the nerve center and command center of an electronic device. Based on the instruction opcode and timing signals, the controller generates operation control signals to control the fetching and execution of instructions.
[0060] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can directly retrieve it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system. In this embodiment, the processor can be a System-on-a-Chip (SoC).
[0061] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an I2C interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0062] The charging management module 140 is used to receive charging input from a power supply device (e.g., a charger, a laptop, etc.). The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 can receive charging input from the wired charger via a USB interface 130. In some wireless charging embodiments, the charging management module 140 can receive wireless charging input via the wireless charging coil of the electronic device.
[0063] While charging the battery 142, the charging management module 140 can also supply power to the electronic device through the power management module 141. Specifically, the battery 142 can be composed of multiple batteries connected in series. The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110.
[0064] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, display screen 193, camera 195, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as the voltage, current, battery cycle count, and battery health status (leakage current, impedance) of the battery 142. In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device; for example, the power management module 141 and the charging management module 140 may be different functional modules within the same chip.
[0065] In this embodiment, the power management module 141 and / or the charging management module 140 may include a fast charging chip. The power management module 141 and / or the charging management module 140 can utilize the fast charging chip to achieve high-power charging of the battery. The fast charging chip has multiple registers for storing corresponding ADC data for different types of data.
[0066] The external memory interface 120 can be used to connect to external non-volatile memory, thereby expanding the storage capacity of the electronic device. The external non-volatile memory communicates with the processor 110 through the external memory interface 120 to perform data storage functions. For example, music, video, and other files can be stored in the external non-volatile memory.
[0067] Internal memory 121 may include one or more random access memory (RAM) and one or more non-volatile memory (NVM). The RAM can be directly read and written by the processor 110 and can be used to store executable programs (e.g., machine instructions) of the operating system or other running programs, as well as user and application data. The NVM can also store executable programs and user and application data, and can be pre-loaded into the RAM for direct read and write operations by the processor 110.
[0068] A touch sensor, also known as a "touch device," can be located on the display screen 193. The touch sensor and the display screen 193 together form a touchscreen, also called a "touchscreen." The touch sensor detects touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 193. In other embodiments, the touch sensor may also be located on the surface of the electronic device, in a different position than the display screen 193.
[0069] A pressure sensor is used to sense pressure signals and convert them into electrical signals. In some embodiments, the pressure sensor may be located on the display screen 193. There are many types of pressure sensors, such as resistive pressure sensors, inductive pressure sensors, and capacitive pressure sensors. When a touch operation is applied to the display screen 193, the electronic device monitors the intensity of the touch operation based on the pressure sensor. The electronic device can also calculate the touch location based on the monitoring signal from the pressure sensor. In some embodiments, touch operations applied to the same touch location but with different intensities can correspond to different operation commands. For example, when a touch operation with an intensity less than a first pressure threshold is applied to the SMS application icon, a command to view an SMS message is executed. When a touch operation with an intensity greater than or equal to the first pressure threshold is applied to the SMS application icon, a command to create a new SMS message is executed.
[0070] A temperature sensor is a device used to measure the temperature of an environment or object. It converts temperature into an electrical or digital signal, allowing the temperature to be read and processed by a computer or other electronic device.
[0071] A gyroscope (GYRO-sensor), also known as a ground sensor or gyroscope sensor, traditionally consists of an internal gyroscope. A three-axis gyroscope can simultaneously measure position, trajectory, and acceleration in six directions. A single-axis gyroscope can only measure quantities in two directions, meaning a system typically requires three gyroscopes, while a single three-axis gyroscope can replace three single-axis gyroscopes. The working principle of a three-axis gyroscope is to measure the angle between the vertical axis of the gyroscope rotor and the device in a three-dimensional coordinate system, and calculate the angular velocity. The angle and angular velocity are used to determine the object's motion state in three-dimensional space. A three-axis gyroscope can simultaneously measure six directions: up, down, left, right, forward, and backward (the composite direction can also be decomposed into three-axis coordinates), ultimately determining the device's trajectory and acceleration. In other words, a three-axis gyroscope determines the device's current motion state by measuring its own rotation, such as forward, backward, up, down, left, or right; and whether it is accelerating (angular velocity) or decelerating (angular velocity). In this embodiment, the gyroscope sensor in the electronic device is a three-axis gyroscope sensor. Based on the detection data from the gyroscope sensor, the electronic device can determine the location or area on the electronic device corresponding to the user's thermal feedback operation.
[0072] The electronic device implements display functions through a GPU, a display screen 193, and an application processor. The GPU is a microprocessor for image editing, connected to the display screen 193 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0073] Electronic devices can achieve shooting functions through ISP, camera 195, video codec, GPU, display 193 and application processor.
[0074] The Information Service Provider (ISP) is used to process data fed back from the camera 195. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization on image noise and brightness. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 195. The camera 195 is used to capture still images or videos. In some embodiments, the electronic device may include one or N cameras, where N is a positive integer greater than 1. The camera 195 can be a front-facing camera or a rear-facing camera.
[0075] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when an electronic device is selecting a frequency, a DSP can perform a Fourier transform on the frequency energy.
[0076] Display screen 193 is used to display images, videos, etc. Display screen 193 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniled LED, a microLED, a micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device may include one or N displays 193, where N is a positive integer greater than 1.
[0077] In this embodiment of the application, the display screen 193 can be used to display the interface of an electronic device (e.g., a camera preview interface, a video preview interface, a final preview interface, etc.), and display images captured by any one or more cameras 195 in the interface.
[0078] The wireless communication function of electronic devices can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem, and baseband processor.
[0079] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in an electronic device can be used to cover one or more communication frequency bands. Different antennas can also be reused to improve antenna utilization.
[0080] The mobile communication module 150 can provide wireless communication solutions, including 2G / 3G / 4G / 5G, for use in electronic devices. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 150 can be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 can be housed in the same device.
[0081] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through audio devices (not limited to speaker 170A, receiver 170B, etc.) or displays images or videos through the display screen 193. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 110 and may be housed in the same device as the mobile communication module 150 or other functional modules.
[0082] The wireless communication module 160 can provide solutions for wireless communication applications in electronic devices, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.
[0083] The SIM card interface 194 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 194 to make contact with and detach from the electronic device. The electronic device can support one or more SIM card interfaces. The SIM card interface 194 supports Nano SIM cards, Micro SIM cards, and other SIM cards. Multiple cards can be inserted into the same SIM card interface 194 simultaneously. The SIM card interface 194 is also compatible with external memory cards. The electronic device interacts with the network through the SIM card to achieve functions such as calls and data communication. One SIM card corresponds to one user number.
[0084] It is understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are merely illustrative and do not constitute a limitation on the structure of the electronic device. In other embodiments of this application, the electronic device may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0085] Of course, it is understandable that the above diagram is merely an illustrative example when the electronic device is a mobile phone. If the electronic device is a tablet, handheld computer, personal computer (PC), PDA, wearable device (such as a smartwatch, smart bracelet), or other device form factor, the structure of the electronic device may include more... Figure 1 The fewer structures shown can also include more than Figure 1 The structures shown are not limited here.
[0086] It is understandable that, generally speaking, the implementation of electronic device functions requires not only hardware support but also software cooperation. The software system of an electronic device can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application embodiment does not limit the architecture adopted by the software system of the electronic device. This application embodiment uses a layered architecture of the electronic device's software system as an example for illustrative purposes. Exemplarily, this layered architecture can be, from top to bottom, an application layer, a framework layer (or application framework layer), a HAL layer (hardware abstraction layer), and a driver layer (or kernel layer).
[0087] The methods described in the following embodiments can all be implemented in electronic devices having the above-described hardware structure and any of the layered software system architectures.
[0088] In this application embodiment, the voice assistant wake-up methods include keyword wake-up, power-on wake-up, and breath wake-up. This application embodiment does not limit the wake-up method of the voice assistant; this disclosure uses keyword wake-up as an example for illustrative explanation. For example, an electronic device can wake up the voice assistant when it recognizes user input including a keyword wake-up voice, such as "Hello YOYO".
[0089] Taking keyword wake-up as an example, the electronic device can wake up the voice assistant in response to an input wake-up voice message containing a keyword, whether on the desktop, lock screen, video application interface, or chat application interface. This application embodiment does not limit the interface on which the electronic device is located when waking up the voice assistant; this embodiment uses the desktop as an example for illustrative purposes.
[0090] The following combination Figure 2 This application provides a detailed description of an emotion recognition method based on an embodiment. (Refer to...) Figure 2 As shown, the emotion recognition method provided in this application embodiment can be executed after the electronic device wakes up the voice assistant, when the voice assistant is in the awake state, and may include steps S201 to S209:
[0091] Step S201: Obtain dialogue data.
[0092] For example, dialogue data can be data generated by an electronic device in response to a user's input dialogue operation to initiate a chat. Dialogue data may include dialogue content in voice or text format, as well as the user's image. This application embodiment does not limit the format of the dialogue content; this application embodiment uses voice format as an example for illustrative purposes.
[0093] In some examples, if the dialogue operation includes a user's voice input, the electronic device can respond to that voice input by converting it into speech-formatted dialogue content. For example, the speech-formatted dialogue content could be a voice message to start a chat, such as "I'm feeling rather sad today!" or "I'm especially happy today because I finally finished a long-running task," or a question-and-answer voice message such as "How to guide children to control their emotions" or "What gift should I give my mother for Mother's Day?"
[0094] In other examples, if the dialogue involves the user typing text via a keyboard, the electronic device receives the dialogue content in that text format. For instance, consider an electronic device displaying a desktop. Figure 3 As shown, the interface 301 of the electronic device includes a prompt card 3011 and a keyboard control 3012. In response to the user's trigger operation (such as a click operation) on the keyboard control 3012, the electronic device can provide a keyboard for the user to input dialogue content in the form of text.
[0095] In some other examples, if the dialogue operation includes the user's voice, the electronic device can respond to the voice by converting it into speech content, and simultaneously trigger the camera to turn on and take a picture, so that the electronic device can acquire the user's image through the camera.
[0096] For example, an electronic device may automatically control a camera to rotate to capture environmental images, select images including the user from the captured environmental images, and identify the images including the user as user images.
[0097] In step S202, in response to the dialogue data including user image and dialogue audio, audio-to-text processing is performed on the dialogue audio to obtain dialogue text.
[0098] For example, the dialogue image may correspond to a user image, and the dialogue audio may correspond to dialogue content in voice format. Step S202 may include: when the electronic device acquires the user image and the dialogue audio, it uses Automatic Speech Recognition (ASR) in the voice assistant to perform processes such as acquisition, preprocessing, feature extraction, acoustic model matching, language model decoding, and semantic understanding on the dialogue audio to obtain the text corresponding to the dialogue audio, and then identifies the text as the dialogue text.
[0099] Understandably, after obtaining the dialogue image, dialogue audio, and dialogue text, the electronic device can input the dialogue image, dialogue audio, and dialogue text into a pre-trained multimodal large model downloaded from the cloud. The multimodal large model then processes the user image, dialogue audio, and dialogue text to obtain the dialogue response text.
[0100] In other examples, if the dialogue data includes text-formatted dialogue content (e.g., input text) and user images, but excludes dialogue audio, then speech-to-text processing is unnecessary; instead, the text-formatted dialogue content is directly identified as the dialogue text. Furthermore, the dialogue text and user images can be input into a multimodal large model, which processes the user images and dialogue text to obtain the dialogue response text.
[0101] The multimodal large model in this embodiment can be a trained multimodal large model obtained by training an initial multimodal large model after adding multiple low-rank adaptive fine-tuning layers.
[0102] In some examples, the initial multimodal large model can be a trained multimodal large model with language interaction capabilities. Furthermore, multiple low-rank adaptive fine-tuning layers can be added to the multiple decoding layers of the large language model in the initial multimodal large model, and only these added low-rank adaptive layers are trained to obtain the multimodal large model. For example, as... Figure 4 As shown, the multimodal large model 40 may include a multimodal encoder 41, an adapter 42, and a large language model 43 connected in series. The large language model 43 may include multiple first decoding layers 431, multiple second decoding layers 432, and multiple low-rank adaptive fine-tuning layers 433. The initial multimodal large model may include a multimodal encoder 41, an adapter 42, multiple first decoding layers 431, and multiple second decoding layers 432. The multiple low-rank adaptive fine-tuning layers 433 may be multiple trained low-rank adaptive fine-tuning layers 433.
[0103] In some examples, dialogue response data may include the dialogue response content in text format, emotion tags to indicate the user's emotional state, and text summarizing the reasons for the user's emotion.
[0104] Step S203: Based on the multimodal encoder, feature extraction processing is performed on the user image, dialogue audio and dialogue text to obtain the first image feature, the first audio feature and the first text feature.
[0105] For example, the multimodal encoder 41 may include a visual encoder 411 for encoding images, an audio encoder 412 for encoding audio, and a text encoder 413 for encoding text. The visual encoder 411, audio encoder 412, and text encoder 413 are independent of each other and can operate simultaneously.
[0106] The visual encoder 411 may include at least one convolutional neural network (CNN) layer and / or a transformer layer. Similarly, like the visual encoder 411, the audio encoder 412 may also include at least one CNN layer and / or a transformer layer. The text encoder 413 may include a first tokenizer module 4131 and a token embedding module 4132. The token embedding module 4132 may also include at least one CNN layer and / or a transformer layer.
[0107] For example, the first image feature, the first audio feature, and the first text feature can all be represented by vectors. Therefore, the first image feature, the first audio feature, and the first text feature in the embodiments of this application can also be referred to as the first image feature vector, the first audio feature vector, and the first text feature vector, respectively.
[0108] In some examples, the first image feature, the first audio feature, and the first text feature can all be feature vectors of arbitrary size. The first dimension corresponding to the first image feature can be represented by L1×D1, the second dimension corresponding to the first audio feature can be represented by L2×D2, and the third dimension corresponding to the first text feature can be represented by L3×D3. Here, L1 represents the first sequence length, D1 represents the first feature dimension; L2 represents the second sequence length, D2 represents the second feature dimension; L3 represents the third sequence length, and D3 represents the third feature dimension. This application does not limit the size of the first dimension corresponding to the first image feature, the second dimension corresponding to the first audio feature, and the third dimension corresponding to the first text feature. However, it is understood that the third dimension corresponding to the first text feature corresponds to the size of the input feature of the large language model. For example, the third feature dimension D3 corresponding to the first text feature is the same as the feature dimension of the input feature of the large language model.
[0109] For example, the first size corresponding to the first image feature, the second size corresponding to the first audio feature, and the third size corresponding to the first text feature may each be different. In some examples, L1, L2, and L3 may be 90, 80, and 100, respectively; D1, D2, and D3 may be 1024, 3200, and 6144, respectively. That is, the first size may be 90×1024, the second size may be 80×3200, and the third size may be 100×6144.
[0110] In some examples, step 203 may include: extracting features from the user image using a visual encoder 411 to obtain first image features; simultaneously extracting features from the dialogue audio using an audio encoder 412 to obtain first audio features; and simultaneously segmenting and embedding the dialogue text using a text encoder 413 to obtain first text features.
[0111] For example, the process of segmenting and embedding the first text by the text encoder 413 to obtain the first text features may include: determining prompt words, the first tokenizer module 4131 segmenting the first text based on the prompt words to obtain multiple words (also known as tokens), and then embedding the multiple tokens by the token embedding module 4132 to obtain the first text features.
[0112] In this embodiment, if the first size corresponding to the first image feature, the second size corresponding to the first audio feature, and the third size corresponding to the first text feature are all different, then it is impossible to effectively fuse the first image feature, the first audio feature, and the first text feature with different sizes. Therefore, it is necessary to perform feature alignment on the first image feature, the first audio feature, and the first text feature, and to achieve effective fusion between multimodal data through the feature-aligned first image feature, the first audio feature, and the first text feature.
[0113] Taking feature alignment of the first image feature, the first audio feature, and the first text feature as an example. Since the D3 corresponding to the first text feature is the same as the feature dimension requirement of the large language model for the input data, the D1 of the first image feature, the D2 corresponding to the first audio feature, and the D3 corresponding to the first text feature can be aligned, thereby ensuring that the feature dimensions of the first image feature, the first audio feature, and the first text feature are the same, and all are D3.
[0114] In some examples, the first image feature, the first audio feature, and the first text feature can be transformed by adapter 42 to align the feature dimensions of the first image feature, the first audio feature, and the first text feature, i.e., S204 is executed.
[0115] Step S204: Based on the adapter, feature transformation processing is performed on the first image features, the first speech features, and the first text features to obtain the input features of the large language model.
[0116] For example, step S204 may include: performing feature mapping on the first image features and the first audio features based on the adapter to obtain the second image features and the second audio features; and then determining the input features of the large language model based on the second image features, the second audio features, and the first text features.
[0117] Specifically, the fourth feature dimension D4 corresponding to the second image feature and the fifth feature dimension D5 corresponding to the second audio feature are the same as the third feature dimension D3. The fourth sequence length L4 corresponding to the second image feature is the same as the first sequence length L1, and the fifth sequence length L5 corresponding to the second image feature is the same as the second sequence length L2.
[0118] For example, if L1, L2, and L3 are 90, 80, and 100 respectively, and D1, D2, and D3 are 1024, 3200, and 6144 respectively, then L4 is 90, L5 is 80, and D4 and D5 are both 6144. That is, the size of the second image feature can be 90×6144, and the size of the second audio feature can be 80×6144.
[0119] For example, such as Figure 4 As shown, adapter 42 may include an image adapter 421 and an audio adapter 422. Based on the adapter, feature mapping is performed on the first image features and the first audio features to obtain second image features and second speech features. This may include: performing feature mapping on the first image features through image adapter 421, aligning the feature dimensions of the first image features from the first feature dimension to the third feature dimension to obtain the second image features; and performing feature mapping on the first audio features through audio adapter 422, aligning the feature dimensions of the first audio features from the second feature dimension to the third feature dimension to obtain the second audio features.
[0120] Determining the input features of a large language model based on second image features, second audio features, and first text features may include concatenating the second image features, second audio features, and first text features to obtain the input features of the large language model.
[0121] For example, if the fourth size corresponding to the second image feature is 90×6144, the fifth size corresponding to the second audio feature is 80×6144, and the first text feature is 100×6144, then the size of the input feature of the large language model is 270×6144.
[0122] For example, such as Figure 4As shown, the large language model 43 may include multiple first decoding layers 431, multiple second decoding layers 432, and multiple low-rank adaptive fine-tuning layers 433. Each second decoding layer 432 corresponds to one low-rank adaptive fine-tuning layer 433.
[0123] In some examples, the number of first decoding layers 431 in multiple first decoding layers 431, the number of second decoding layers 432 in multiple second decoding layers 432, and the number of low-rank adaptive fine-tuning layers 433 in multiple low-rank adaptive fine-tuning layers 433 are all related to the structure of the large language model 43. In some examples, the number of first decoding layers 431 can be any integer greater than or equal to 1, and the number of second decoding layers 432 is the same as the number of low-rank adaptive fine-tuning layers 433, and both can be any positive integer greater than or equal to 1. This application does not limit the number of first decoding layers 431, the number of second decoding layers 432, and the number of low-rank adaptive fine-tuning layers 433 in its embodiments. Figure 4 As shown, this embodiment of the application takes as an example that the number of the first decoding layer 431 is n-1, the number of the second decoding layer 432 and the number of the low-rank adaptive fine-tuning layer 433 are both Nn.
[0124] Step S205: Process the input features based on multiple first decoding layers in the large language model to obtain the first intermediate features.
[0125] Exemplarily, each of the multiple first decoding layers can be a Transformer-structured neural network layer (Transformer layer) or a feedforward neural network layer. This application embodiment does not limit the structure of the first decoding layer. This application embodiment uses a Transformer layer as an example for illustrative explanation.
[0126] Taking the input feature of the large language model as X1, the multiple first decoding layers 431 include n-1 first decoding layers 431 connected in sequence, and the weight parameters of the n-1 first decoding layers 431 connected in sequence are W1 to W1 respectively. n-1 For example. (Reference) Figure 5 As shown, the output features of the n-1 first decoding layers 431 can be X1*W1, X1*W1*W2, ..., X1*W1*W2*...W n-1 .
[0127] For example, the first intermediate feature can be the feature output by the last first decoding layer among multiple first decoding layers. The last first decoding layer can be adjacent to the second decoding layer and the low-rank adaptive fine-tuning layer. Taking the input feature of the large language model as X1, the number of first decoding layers 431 is n-1, and the weight parameters of the n-1 first decoding layers 431 are W1 to W1 respectively. n-1For example, the first intermediate feature can be the output feature of the (n-1)th first decoding layer 431, that is, the first intermediate feature is X1*W1*W2……*W n-1 .
[0128] For example, the size of the first intermediate feature is related to the structure of the plurality of first decoding layers 431 in the large language model 43. In some examples, the size of the first intermediate feature is larger than the size of the input features of the large language model. In other examples, the size of the first intermediate feature may be smaller than or equal to the size of the input features of the large language model. This application embodiment does not limit the size of the first intermediate feature.
[0129] For example, by processing the first intermediate feature based on multiple second decoding layers 432, multiple second intermediate features can be obtained. Each second decoding layer outputs one second intermediate feature. Similarly, by processing the first intermediate feature based on multiple low-rank adaptive fine-tuning layers 433, multiple fine-tuned features can be obtained. Each low-rank adaptive fine-tuning layer 433 outputs one fine-tuned feature.
[0130] In some examples, the number of multiple second decoding layers is Nn. The multiple second decoding layers 432 may include a first second decoding layer to an Nn-th second decoding layer connected in series. The multiple second intermediate features may include a first second intermediate feature to an Nn-th second intermediate feature corresponding to each of the first to Nn-th second decoding layers.
[0131] Taking Nn as an example, the multiple low-rank adaptive fine-tuning layers include the first low-rank adaptive fine-tuning layer to the Nn low-rank adaptive fine-tuning layer connected in series, and the multiple fine-tuning features may include the first fine-tuning feature to the Nn fine-tuning feature corresponding to the first low-rank adaptive fine-tuning layer to the Nn low-rank adaptive fine-tuning layer, respectively.
[0132] Step S206: Process the first intermediate feature based on the first second decoding layer among multiple second decoding layers to obtain the first second intermediate feature; process the first intermediate feature based on the first low-rank adaptive fine-tuning layer among multiple low-rank adaptive fine-tuning layers to obtain the first fine-tuned feature.
[0133] For example, since the structure of each of the multiple second decoding layers is similar to the structure of each of the multiple first decoding layers 431, the embodiments of this application will not be described again here.
[0134] In some examples, reference Figure 4 As shown, the first second decoding layer 432 can be one of the multiple second decoding layers 432 that is adjacent to the multiple first decoding layers 431. For example, see reference... Figure 5 As shown, the first second decoding layer 432 can be the weight parameter W.n The corresponding decoding layer, that is, the first second decoding layer 432, can be related to the weight parameter W. n-1 After the corresponding first decoding layer, and with the weight parameter W n-1 The second decoding layer 432 is adjacent to the first decoding layer.
[0135] For example, the first second intermediate feature can be a feature output by the first second decoding layer. In some examples, the first second intermediate feature can be determined based on the input features of the first second decoding layer and the weight parameters of the first second decoding layer. For example, the first second intermediate feature can be determined by the following formula (1).
[0136] X3 = W n *X2 (1);
[0137] Where X3 is the first second intermediate feature, X2 is the first intermediate feature, that is, the input feature of the first second decoding layer, W n These are the weight parameters for the first second decoding layer.
[0138] In some examples, if X2 = X1 * W 1* W 2* ...*W n-1 Then X3 = X1 * W1 * W 2* ...*W n-1 *W n .
[0139] Each low-rank adaptive fine-tuning layer in a multi-layer low-rank adaptive fine-tuning system can be a neural network layer whose weight parameters are represented by low-rank matrices. In some examples, the weight parameters of each low-rank adaptive fine-tuning layer can be represented by the product of low-rank matrices A and B, i.e., A*B. Taking Nn low-rank adaptive fine-tuning layers as an example, the weight parameter W corresponding to the k-th low-rank adaptive fine-tuning layer among these Nn low-rank adaptive fine-tuning layers... Ψk It can be A ψk *B Ψk , where k is greater than or equal to 1 and less than or equal to Nn.
[0140] For example, A Ψk Dimensions and B Ψk The dimension of each is the same as the weight parameter W corresponding to the kth second decoding layer 432. (k+n-1) Related. In some examples, A Ψk and B Ψk The dimension of the product and W (k+n-1) The dimensions are the same. For example, taking a matrix with a rank of 128 as an example, if W (k+n-1) If the dimension is (6144, 8192), then B Ψk and Aψk The dimensions can be (6144, 128) and (128, 8192) respectively; for example, taking a matrix with a rank of 218 as an example, if W (k+n-1) If the dimensions are (6144, 8192), then B ψk and A ψk The dimensions can be (6144, 218) and (218, 8192), respectively. Among them, the k-th second decoding layer 432 corresponds to the k-th low-rank adaptive fine-tuning layer 433.
[0141] In some examples, reference Figure 4 As shown, the first low-rank adaptive fine-tuning layer 433 can be one of the multiple low-rank adaptive fine-tuning layers 433 that is adjacent to the multiple first decoding layers 431. For example, refer to... Figure 5 As shown, the first low-rank adaptive fine-tuning layer 433 can be related to the weight parameter W. n-1 After the corresponding first decoding layer 431, and with the weight parameter W n The corresponding second decoding layer 432 corresponds to the low-rank adaptive fine-tuning layer 433.
[0142] In some examples, the first fine-tuned feature can be a feature output by the first low-rank adaptive fine-tuning layer. In some examples, the first fine-tuned feature can be determined based on the input features of the first low-rank adaptive fine-tuning layer and the weight parameters of the first low-rank adaptive fine-tuning layer. For example, the first fine-tuned feature can be determined by the following formula (2).
[0143] X4 = W ψ1 *X2 (2);
[0144] Where X4 is the first fine-tuning feature, X2 is the first intermediate feature, i.e., the input feature of the first low-rank adaptive fine-tuning layer, and W ψ1 These are the weight parameters for the first low-rank adaptive fine-tuning layer.
[0145] In some examples, if X2 = X 1* W 1* W 2* ...*W n-1 X4 = X 1* W 1* W 2* ...*W n-1 *W ψ1 .
[0146] In this embodiment, since the weight parameters corresponding to the first second decoding layer have the same dimension as the weight parameters corresponding to the first low-rank adaptive fine-tuning layer, the size of the first second intermediate feature obtained by processing the first intermediate feature based on the first second decoding layer is the same as the size of the first fine-tuned feature obtained by processing the first intermediate feature based on the first low-rank adaptive fine-tuning layer.
[0147] Step S207: Process the (i-1)th second intermediate feature based on the i-th second decoding layer to obtain the i-th second intermediate feature; process the (i-1)th comprehensive feature based on the i-th low-rank adaptive fine-tuning layer to obtain the i-th fine-tuned feature.
[0148] Here, the (i-1)th comprehensive feature is a feature that combines the (i-1)th second intermediate feature and the (i-1)th fine-tuned feature. In some examples, the (i-1)th comprehensive feature can be obtained by adding the (i-1)th second intermediate feature and the (i-1)th fine-tuned feature. i is an integer greater than 1 and less than or equal to Nn.
[0149] For example, consider a plurality of second decoding layers comprising Nn second decoding layers and a plurality of low-rank adaptive fine-tuning layers comprising Nn low-rank adaptive fine-tuning layers. The Nn-th second decoding layer among the plurality of second decoding layers can correspond to the last second decoding layer. The Nn-th second intermediate feature among the plurality of second intermediate features can correspond to the last second intermediate feature. The Nn-th low-rank adaptive fine-tuning layer can correspond to the last low-rank adaptive fine-tuning layer. The Nn-th fine-tuned feature can correspond to the last fine-tuned feature.
[0150] Since the implementation method of processing the (i-1)th second intermediate feature among multiple second intermediate features based on the i-th second decoding layer among multiple second decoding layers to obtain the i-th second intermediate feature is similar to the implementation method of processing the first intermediate feature based on the first second decoding layer to obtain the first second intermediate feature, the embodiments of this application will not be described again here. Let the multiple second decoding layers include Nn second decoding layers, and the weight parameter corresponding to the i-th second decoding layer among these Nn second decoding layers be W. k+n-1 For example. (Reference) Figure 5 As shown, if the first second intermediate feature is X 1* W 1* W 2* ...*W n-1 *W n Then the second second intermediate special is X. 1* W 1* W 2* ...*W n-1 *W n *W n+1 The third second intermediate feature is X.1* W 1* W 2* ...*W n-1 *W n *W n+1 *W n+2 And so on, the Nn-th second intermediate feature is X 1* W 1* W 2* ...*W n-1 *W n *W n+1 *W n+2 *……W N .
[0151] The i-th low-rank adaptive fine-tuning layer processes the (i-1)-th comprehensive feature to obtain the (i-1)-th fine-tuned feature. This process may include: performing an addition operation on the (i-1)-th second intermediate feature and the (i-1)-th fine-tuned feature to obtain the (i-1)-th comprehensive feature; and processing the (i-1)-th comprehensive feature based on the i-th low-rank adaptive fine-tuning layer to obtain the i-th fine-tuned feature.
[0152] Taking i = 2, i.e., the second low-rank adaptive fine-tuning layer, as an example, since the first second intermediate feature has the same size as the first fine-tuned feature, we can directly add the first second intermediate feature and the first fine-tuned feature to obtain the first comprehensive feature (the first comprehensive feature), and use the first comprehensive feature as the input feature of the second low-rank adaptive fine-tuning layer. For example, if the first second intermediate feature is X... 1* W 1* W 2* ...*W n-1 *W n The first fine-tuning feature is X. 1* W 1* W 2* ...*W n-1 *W ψ1 Then the first comprehensive feature is X 1* W 1* W 2* ...*W n-1 *W n +X 1* W 1* W 2* ...*W n-1* W ψ1 .
[0153] The implementation method of processing the (i-1)th comprehensive feature based on the i-th low-rank adaptive fine-tuning layer to obtain the i-th fine-tuned feature is similar to the implementation method of processing the first intermediate feature based on the first second decoding layer to obtain the first second intermediate feature. Therefore, the embodiments of this application will not be described in detail here. Multiple low-rank adaptive fine-tuning layers include Nn low-rank adaptive fine-tuning layers, and the weight parameter corresponding to the i-th low-rank adaptive fine-tuning layer is W. ψk For example, if the first comprehensive feature X 1* W 1* W 2* ...*W n-1 *W n +X 1* W 1* W 2* ...*W n-1 *W ψ1 Then the second fine-tuning feature is (X) 1* W 1* W 2* ...*W n-1 *W n +X 1* W 1* W 2* ...*W n-1* W ψ1 )*W ψ2 Furthermore, the second comprehensive feature is (X) 1* W 1* W 2* ...*W n-1 *W n +X 1* W 1* W 2* ...*W n-1* W ψ1 )*W ψ2 +X 1* W 1* W 2* ...*W n-1 *W n *W n+1 The third fine-tuning feature is [(X 1* W 1* W 2* ...*W n-1 *W n +X 1* W 1* W 2* ...*W n-1* W ψ1 )*W ψ2 +X 1* W 1* W 2* ...*W n-1 *Wn *W n+1 ]*W ψ3 .
[0154] Step S208: Perform text conversion processing on the last second intermediate feature to obtain the dialogue response text; perform text conversion processing on the last fine-tuning feature to obtain the emotion label and emotion reason text.
[0155] For example, refer to Figure 5 As shown, the last second intermediate feature can be processed into text by the second tokenizer module 434 to obtain the dialogue response text. In some examples, the second tokenizer module 434 performs vector-to-text processing on the first output feature to obtain the dialogue response text. This may include: finding the correspondence between the first output feature and the text through the second tokenizer module 434, determining the dialogue response in text form based on the first output feature and the correspondence, and defining the dialogue response in text form as the dialogue response text.
[0156] For example, refer to Figure 5 As shown, the last fine-tuned feature can be processed into text by the third Tokenizer module 435 to obtain the emotion label and the emotion reason text. Since the implementation of the text conversion processing of the last fine-tuned feature by the third Tokenizer module 435 to obtain the emotion label and the emotion reason text is similar to the implementation of the vector-to-text processing of the first output feature by the second Tokenizer module 434 to obtain the dialogue response text, this embodiment will not be described in detail.
[0157] The emotion recognition method provided in this disclosure utilizes a multimodal large language model, which is a model with emotion labeling and emotion cause recognition capabilities. Therefore, by processing the acquired multimodal data through this model, emotion labels and emotions can be determined relatively accurately. Furthermore, based on these emotion labels, the dialogue response text can be expressed emotionally to provide users with better emotional interaction and companionship, thereby improving the user experience. Simultaneously, the emotion labels and emotions used in the text can guide the generation of subsequent dialogue response texts, emotion labels, and emotions, further enhancing the empathy capabilities of electronic devices.
[0158] After the multimodal large model outputs the dialogue response text, emotion tag, and emotion reason text, it can also generate an emotional response speech based on the emotion tag and dialogue response text, and update the historical dialogue, user profile, and / or personal knowledge graph based on the emotion tag and emotion reason. In some examples, at least one of the following steps S209, S210, and S211 can be performed after step S208.
[0159] Step S209: Process the dialogue response text and emotion tags based on the TTS module to generate an emotional response voice.
[0160] For example, the emotional response speech can be a speech answer rich in emotional expression. In some examples, the emotion expressed by the emotional response speech is the same as the emotion corresponding to the emotion label, and the speech content of the emotional response speech can be the same as the content of the dialogue response text. For example, when the emotion label is a smiley face, the emotional response speech can be a speech answer expressed in a cheerful tone. For example, when the emotion label is crying, the emotional response speech can be a speech answer expressed in a sad tone.
[0161] For example, step S209 may include: the TTS module determining the voice expression style based on the emotion tag, converting the dialogue response text into voice of the voice expression style, and using the voice of the voice expression style as the emotion response voice.
[0162] Step S210: The emotion tag and the text of the emotion reason are used as historical dialogues, and prompt words are generated based on the historical dialogues; when processing the first audio of the next frame, the first text corresponding to the next frame audio is feature extracted based on the prompt words to obtain the first text features corresponding to the next frame audio.
[0163] For example, cue words can be identified from historical dialogues that include emotion tags and text explaining the reasons for the emotion. Then, when processing the first audio frame of the next frame, feature extraction is performed based on the cue words.
[0164] Because historical dialogues include emotion tags and emotional reason text corresponding to the current dialogue data, the prompts generated based on historical dialogues are more consistent with the actual dialogue scenario. Therefore, based on these prompts, the current dialogue can be accurately guided, improving the accuracy of dialogue responses and thus enhancing the user experience.
[0165] Step S211: Store the emotion tags and emotion reasons in the database, update the prompt words based on the data in the database, and extract features from the first text based on the updated prompt words to obtain the updated features of the first text.
[0166] For example, the database could be a user profile or a personal knowledge graph. Updating prompts based on data in the database could include: re-identifying data stored in the database of emotion tags and emotion reasons to obtain updated prompts.
[0167] Because the database contains a wealth of emotion tags and text explaining the reasons for those emotions, the prompts generated based on the data in the database are more likely to reflect the user's personal habits. As a result, these prompts can more accurately guide the next conversation, improve the accuracy of dialogue responses, and thus enhance the user experience.
[0168] The embodiments of this application do not limit the order of steps S209, S210, and S211. In some examples, steps S209, S210, and S211 can be performed simultaneously.
[0169] Figure 6 This is a flowchart illustrating a multimodal large-scale model training method provided in an embodiment of this application. This multimodal large-scale model training method can be performed in the cloud. Figure 6 As shown, the multimodal large model training method includes the following steps S601 to S610.
[0170] Step S601: Obtain multiple sets of multimodal sample data.
[0171] The multimodal sample data includes at least one of sample images, sample audio, and sample text.
[0172] For example, images of people, audio of conversations, and corresponding text of conversations can be acquired by multiple different users engaging in dialogues using electronic devices. The acquired images are used as sample images, the acquired audio of conversations are used as sample audio, and the acquired text of conversations are used as sample text. In some examples, a single frame of image, a single frame of audio, and the corresponding text of that single frame of audio acquired by any user while using an electronic device can be used as a set of multimodal sample data.
[0173] In some examples, multiple sets of sample data can be pre-determined before training a multimodal large model.
[0174] Step S602: Obtain the sample emotion labels and sample emotion reason text corresponding to each group of multimodal sample data.
[0175] For example, a sample emotion label and a sample emotion reason text can be determined using a set of multimodal sample data.
[0176] In this embodiment, sample emotion labels and sample emotion text can be used as the ground truth of the multimodal large model, i.e., the label information of the multimodal large model. In some examples, the sample emotion labels and sample emotion reason text corresponding to each group of sample data can be predetermined before training the multimodal large model.
[0177] Step S603: Based on the multimodal encoder in the initial multimodal large model, perform feature extraction processing on the multimodal sample data to obtain multimodal sample features.
[0178] For example, the initial multimodal large model can be a neural network model that adds multiple initial low-rank adaptive fine-tuning layers (untrained low-rank adaptive fine-tuning layers) to a trained multimodal large model with language interaction capabilities to output emotion labels and emotion reason text. In some examples, the initial multimodal large model may include a trained multimodal encoder, a trained adapter, multiple trained first decoding layers, multiple trained second decoding layers, and multiple initial low-rank adaptive fine-tuning layers to be trained. The trained multimodal encoder, the trained adapter, the multiple trained first decoding layers, and the multiple trained second decoding layers can be used to generate dialogue response text for voice interaction.
[0179] In this embodiment, when training a multimodal large model, the multimodal encoder, adapter, multiple first decoding layers, and multiple second decoding layers that have been trained in the initial multimodal large model are no longer trained. That is, the weight parameters corresponding to the multimodal encoder, adapter, multiple first decoding layers, and multiple second decoding layers that have been trained are not changed. Only multiple low-rank adaptive fine-tuning layers are trained, and the weight parameters of multiple low-rank adaptive fine-tuning layers are adjusted during the training process.
[0180] For example, taking multimodal sample data including sample user images, sample dialogue audio, and sample dialogue text as an example, feature extraction processing of the multimodal sample data based on the multimodal encoder in the initial multimodal large model to obtain multimodal sample features may include: performing feature extraction processing on the sample user images, sample dialogue audio, and sample dialogue text based on the multimodal encoder in the initial multimodal large model to obtain first sample image features, first sample audio features, and first sample text features.
[0181] For example, the implementation method of performing feature extraction processing on sample user images, sample dialogue audio, and sample dialogue text based on the multimodal encoder in the initial multimodal large model to obtain the first sample image features, first sample audio features, and first sample text features can be found in [reference missing]. Figure 2 Step S203 of the illustrated embodiment will not be repeated here.
[0182] Step S604: Based on the adapter in the initial multimodal large model, perform feature transformation processing on the multimodal sample features to obtain the sample input features of the initial large language model in the initial multimodal large model.
[0183] For example, the implementation of step S604 can be found in the following example. Figure 2 Step S204 of the illustrated embodiment will not be repeated here.
[0184] Step S605: Process the sample input features based on multiple first decoding layers in the initial large language model to obtain the first prediction intermediate features.
[0185] For example, the implementation of step S605 can be found in the following example: Figure 2 Step S205 of the illustrated embodiment will not be repeated in this application embodiment.
[0186] Step S606: Process the first prediction intermediate feature based on the first second decoding layer among multiple second decoding layers in the initial large language model to obtain the first second prediction intermediate feature; process the first prediction intermediate feature based on the first initial low-rank adaptive fine-tuning layer among multiple initial low-rank adaptive fine-tuning layers in the initial large language model to obtain the first prediction fine-tuning feature.
[0187] For example, the implementation of step S606 can be found in the following example: Figure 2 Step S206 of the illustrated embodiment will not be repeated in this application embodiment.
[0188] Step S607: Based on the i-th second decoding layer among multiple second decoding layers, process the (i-1)-th second prediction intermediate feature among multiple second prediction features to obtain the i-th second prediction intermediate feature; based on the i-th low-rank adaptive fine-tuning layer, process the (i-1)-th prediction comprehensive feature to obtain the i-th prediction fine-tuning feature.
[0189] Here, the (i-1)th prediction integrated feature is a feature that integrates the (i-1)th second prediction intermediate feature and the (i-1)th prediction fine-tuning feature. i is an integer greater than 1 and less than or equal to m.
[0190] In some examples, reference Figure 4 or Figure 5 As shown, m can be Nn.
[0191] For example, the implementation of step S607 can be found in the following example: Figure 2 Step S207 of the illustrated embodiment will not be repeated in this application embodiment.
[0192] Step S608: Perform text conversion processing on the last second prediction intermediate feature to obtain the predicted dialogue response text; perform text conversion processing on the last prediction fine-tuning feature to obtain the predicted emotion label and the predicted emotion reason text.
[0193] For example, the implementation of step S608 can be found in the following example: Figure 2 Steps S208 and S209 of the illustrated embodiment will not be repeated in this application embodiment.
[0194] Step S609: Determine the loss value of the multimodal large model based on the predicted emotion label, the predicted emotion reason text, the sample emotion label, and the sample emotion reason text.
[0195] For example, step S609 may include: determining a first loss value based on the predicted emotion label and the sample emotion label; determining a second loss value based on the predicted emotion cause text and the sample emotion cause text; and determining the loss value of a multimodal large model based on the first loss value and the second loss value.
[0196] In some instances, determining the loss value of a multimodal large model based on a first loss value and a second loss value may include: determining a first weighting coefficient for the first loss value and a first product of the first loss value and the first weighting coefficient; determining a second weighting coefficient for the second loss value and a second product of the second loss value and the second weighting coefficient; and determining the loss value of the multimodal large model based on the sum of the first product and the second product.
[0197] Step S610: Iteratively update the weight parameters in multiple low-rank adaptive fine-tuning layers based on the loss value until the updated multimodal large model satisfies the convergence condition, thus obtaining the trained multimodal large model.
[0198] For example, step S610 may include: updating the weight parameters in multiple initial multimodal adaptation fine-tuning layers using the loss value until the updated multimodal large model satisfies the convergence condition. If the updated model satisfies the convergence condition, the trained multimodal large model is obtained.
[0199] For example, convergence conditions may include at least one of the following: the loss value output by the model is less than a first preset value, the change in weights between two adjacent iterations is less than a second preset value, and the number of iterations reaches a preset number. In some examples, both the first and second preset values can be determined based on the model accuracy; however, this embodiment does not limit the magnitude of the first and second preset values.
[0200] The multimodal large model training method provided in this application determines the loss value of the multimodal large model based solely on the predicted sentiment labels, predicted sentiment reason text, sample sentiment labels, and sample sentiment reason text corresponding to multiple low-rank adaptive fine-tuning layers. It iteratively updates the weight parameters of only the multiple low-rank adaptive fine-tuning layers based on the loss value, without updating the weight parameters of other neural network layers in the multimodal large model besides the multiple low-rank adaptive fine-tuning layers. That is, it does not train other neural network layers besides the multiple low-rank adaptive fine-tuning layers. Therefore, it can reduce the computational load of model training and improve the efficiency of model training.
[0201] Figure 7 This is a flowchart illustrating another emotion recognition method provided in the application embodiment. Figure 7As shown, the emotion recognition method includes steps S701 to S703.
[0202] Step S701: Obtain multimodal dialogue data.
[0203] Multimodal dialogue data can correspond Figure 2 The dialogue data in the illustrated embodiment. The implementation of step S703 can be found in... Figure 2 Step S201 of the illustrated embodiment.
[0204] Step S702: Process the multimodal dialogue data based on the encoder and adapter in the multimodal large model to obtain the input features of the large language model in the multimodal large model.
[0205] In some embodiments of this application, step S702 may include: extracting features from the multimodal dialogue data based on the encoder in the multimodal large model to obtain multimodal dialogue features; and processing the multimodal dialogue features based on the adapter in the multimodal large model to obtain the input features of the large language model in the multimodal large model.
[0206] For example, the implementation of step S702 can be found in the following example. Figure 2 Steps S202 to S204 of the illustrated embodiment will not be repeated here.
[0207] Step S703: The input features are processed based on multiple decoders in the large speech model to obtain the dialogue response text and multiple intermediate features; the multiple intermediate features are processed based on multiple low-rank adaptive fine-tuning layers in the large speech model to obtain emotion labels for representing user emotions and emotion reason text for representing emotion reasons.
[0208] The implementation method of step S703 can be found in the following example. Figure 2 Steps S205 to S207 of the illustrated embodiment will not be repeated here.
[0209] In some embodiments of this application, the input features are processed based on multiple decoders in a large speech model to obtain dialogue response text and multiple intermediate features, including: processing the input features based on multiple first decoding layers among multiple decoding layers to obtain a first intermediate feature among intermediate features; processing the first intermediate feature based on multiple second decoding layers among multiple decoding layers to obtain multiple second intermediate features among multiple intermediate features; performing text conversion processing on the last second intermediate feature among multiple second intermediate features to obtain dialogue response text; wherein, the last second intermediate feature corresponds to the last second decoding layer among multiple second decoding layers.
[0210] In some embodiments of this application, multiple intermediate features are processed based on multiple low-rank adaptive fine-tuning layers in a large language model to obtain an emotion label representing a user's emotion and an emotion reason text representing the emotion reason. This includes: processing a first intermediate feature and multiple second intermediate features based on multiple low-rank adaptive fine-tuning layers in a large language model to obtain multiple fine-tuning features; performing text conversion processing on the last fine-tuning feature among the multiple fine-tuning features to obtain an emotion label and an emotion reason text; wherein, the last fine-tuning feature corresponds to the last low-rank adaptive fine-tuning layer among the multiple low-rank adaptive fine-tuning layers.
[0211] In some embodiments of this application, the plurality of second decoding layers include a first second decoding layer to a m-th second decoding layer; the m-th second decoding layer is the last second decoding layer among the plurality of second decoding layers; the plurality of second intermediate features include a first second intermediate feature to a m-th second intermediate feature; processing the first intermediate feature based on the plurality of second decoding layers among the plurality of decoding layers to obtain the plurality of second intermediate features among the intermediate features includes: processing the first intermediate feature based on the first second decoding layer to obtain the first second intermediate feature; processing the (i-1)-th second intermediate feature based on the i-th second decoding layer to obtain the i-th second intermediate feature; where i is an integer greater than 1 and less than or equal to m.
[0212] In some embodiments of this application, the multiple low-rank adaptive fine-tuning layers include the first low-rank adaptive fine-tuning layer to the p-th low-rank adaptive fine-tuning layer; the p-th low-rank adaptive fine-tuning layer is the last low-rank adaptive fine-tuning layer among the multiple low-rank adaptive fine-tuning layers; the multiple fine-tuning features include the first fine-tuning feature to the j-th fine-tuning feature; where j is an integer greater than 1 and less than or equal to p. Based on the multiple low-rank adaptive fine-tuning layers in the large language model, the first intermediate feature and multiple second intermediate features are processed to obtain multiple fine-tuning features, including: processing the first intermediate feature based on the first low-rank adaptive fine-tuning layer to obtain the first fine-tuning feature; determining the (j-1)-th comprehensive feature based on the (j-1)-th fine-tuning feature and the (j-1)-th second intermediate feature; and processing the (j-1)-th comprehensive feature based on the j-th low-rank adaptive fine-tuning layer to obtain the j-th fine-tuning feature.
[0213] In some embodiments of this application, determining the (j-1)th comprehensive feature based on the (j-1)th fine-tuning feature and the (j-1)th second intermediate feature includes: performing an addition operation on the (j-1)th fine-tuning feature and the (j-1)th second intermediate feature to obtain the (j-1)th feature sum; and determining the (j-1)th feature sum as the (j-1)th comprehensive feature.
[0214] Figure 8 This is a flowchart illustrating another multimodal large model training method provided in an embodiment of this application. For example... Figure 8As shown, the multimodal large model training method includes the following steps S801 to S804.
[0215] Step S801: Obtain multiple sets of multimodal sample data and the sample emotion labels and sample emotion reason texts corresponding to each set of multimodal sample data.
[0216] For example, the implementation of step S801 can be found in the following example. Figure 6 Steps S601 and S602 of the illustrated embodiment will not be repeated here.
[0217] Step S802: Based on the encoder and adapter in the initial multimodal large model, process the multimodal sample data to obtain the sample input features of the initial large language model in the initial multimodal large model.
[0218] For example, the implementation of step S802 can be found in the following example. Figure 6 Steps S603 and S604 of the illustrated embodiment will not be repeated in this application embodiment.
[0219] Step S803: Process the sample input features based on multiple decoding layers in the initial large language model to obtain multiple intermediate prediction features; process the multiple intermediate prediction features based on multiple initial low-rank adaptive fine-tuning layers in the initial large language model to obtain the predicted sentiment label and the predicted sentiment reason text.
[0220] In some examples, the sample input features are processed based on multiple decoding layers in the initial large language model to obtain multiple intermediate prediction features. This may include: processing the sample input features based on multiple first decoding layers in the multiple decoding layers to obtain the first intermediate prediction feature among the multiple intermediate prediction features.
[0221] The first prediction intermediate features are processed by multiple second decoding layers in multiple decoding layers to obtain multiple second prediction intermediate features among multiple prediction intermediate features.
[0222] In some examples, multiple intermediate prediction features are processed based on multiple initial low-rank adaptive fine-tuning layers in the initial large language model to obtain predicted sentiment labels and predicted sentiment reason text. This may include: processing the first intermediate prediction feature and multiple second intermediate prediction features based on multiple initial low-rank adaptive fine-tuning layers in the large language model to obtain multiple predicted fine-tuning features; performing text transformation processing on the last predicted fine-tuning feature among the multiple predicted fine-tuning features to obtain predicted sentiment labels and predicted sentiment reason text; wherein, the last predicted fine-tuning feature corresponds to the last initial low-rank adaptive fine-tuning layer among the multiple initial low-rank adaptive fine-tuning layers.
[0223] The implementation method of step S803 can be found in the following example. Figure 6Steps S605 to S607 of the illustrated embodiment will not be repeated in this application embodiment.
[0224] Step S804: Using the predicted emotion label and the predicted emotion language text as the initial training output of the initial multimodal large model, and the sample emotion label and the sample emotion reason text as supervision information, the initial multimodal large model is iteratively trained to obtain the trained multimodal large model.
[0225] In some embodiments, step S804 may include: determining the loss value of the multimodal large model based on the predicted emotion label, the predicted emotion language text, the sample emotion label, and the sample emotion reason text; iteratively updating the weight parameters in multiple initial low-rank adaptive fine-tuning layers based on the loss value to obtain the trained multimodal large model.
[0226] This application embodiment can divide the above-described electronic device into functional modules based on the method example described above. For example, each function can be divided into its own functional modules, or two or more functions can be integrated into one processing module. The integrated modules can be implemented in hardware or as software functional modules. It should be noted that the module division in this embodiment is illustrative and only represents one logical functional division; in actual implementation, there may be other division methods.
[0227] When dividing each function into modules according to its corresponding function, refer to Figure 9 As shown in the figure, this application provides an electronic device 900 that can implement the emotion recognition method provided in the foregoing embodiments. The electronic device 900 may include a first acquisition module 901, a first processing module 902, and a second processing module 903.
[0228] The first acquisition module 901 is used to acquire multimodal dialogue data;
[0229] The first processing module 902 is used to process multimodal dialogue data based on the encoder and adapter in the multimodal large model to obtain the input features of the large language model in the multimodal large model;
[0230] The second processing module 903 is used to process the input features based on multiple decoding layers in the large language model to obtain the dialogue response text and multiple intermediate features; and to process the multiple intermediate features based on multiple low-rank adaptive fine-tuning layers in the large language model to obtain emotion labels for representing user emotions and emotion reason text for representing emotion reasons.
[0231] In some embodiments, the second processing module 903 is specifically configured to process the input features based on multiple first decoding layers among multiple decoding layers to obtain a first intermediate feature among intermediate features; process the first intermediate feature based on multiple second decoding layers among multiple decoding layers to obtain multiple second intermediate features among multiple intermediate features; and perform text conversion processing on the last second intermediate feature among multiple second intermediate features to obtain dialogue response text; wherein the last second intermediate feature corresponds to the last second decoding layer among multiple second decoding layers.
[0232] In some embodiments, the second processing module 903 is specifically used to process the first intermediate feature and multiple second intermediate features based on multiple low-rank adaptive fine-tuning layers in the large language model to obtain multiple fine-tuning features; and to perform text conversion processing on the last fine-tuning feature among the multiple fine-tuning features to obtain an emotion label and emotion reason text; wherein, the last fine-tuning feature corresponds to the last low-rank adaptive fine-tuning layer among the multiple low-rank adaptive fine-tuning layers.
[0233] In some embodiments, the plurality of second decoding layers include a first second decoding layer to a m-th second decoding layer; the m-th second decoding layer is the last second decoding layer among the plurality of second decoding layers; the plurality of second intermediate features include a first second intermediate feature to a m-th second intermediate feature; the second processing module 903 is specifically used to process the first intermediate feature based on the first second decoding layer to obtain the first second intermediate feature; and to process the (i-1)-th second intermediate feature based on the i-th second decoding layer to obtain the i-th second intermediate feature; wherein i is an integer greater than 1 and less than or equal to m.
[0234] In some embodiments, the plurality of low-rank adaptive fine-tuning layers include a first low-rank adaptive fine-tuning layer to a p-th low-rank adaptive fine-tuning layer; the p-th low-rank adaptive fine-tuning layer is the last low-rank adaptive fine-tuning layer among the plurality of low-rank adaptive fine-tuning layers; the plurality of fine-tuning features include a first fine-tuning feature to a (j-1)-th fine-tuning feature; where j is an integer greater than 1 and less than or equal to p; the second processing module 903 is specifically used to process the first intermediate feature based on the first low-rank adaptive fine-tuning layer to obtain the first fine-tuning feature; to determine the (j-1)-th comprehensive feature based on the (j-1)-th fine-tuning feature and the (j-1)-th second intermediate feature; and to process the (j-1)-th comprehensive feature based on the j-th low-rank adaptive fine-tuning layer to obtain the j-th fine-tuning feature.
[0235] In some embodiments, the second processing module 903 is specifically used to perform an addition operation on the (j-1)th fine-tuning feature and the (j-1)th second intermediate feature to obtain the (j-1)th feature sum; and to determine the (j-1)th feature sum as the (j-1)th comprehensive feature.
[0236] Regarding the electronic devices in the above embodiments, the specific methods by which each module performs its operations have been described in detail in the embodiments of the information display method described above, and will not be elaborated here. The related beneficial effects can also be referred to the related beneficial effects of the aforementioned information display method, and will not be repeated here.
[0237] This application also provides an electronic device, which includes: a display screen, a memory, and one or more processors; the display screen, the memory, and the processors are coupled; wherein, the memory stores computer program code, which includes computer instructions, and when the computer instructions are executed by the processor, the electronic device performs the depth estimation method provided in the foregoing embodiments. The specific structure of this electronic device can be referred to... Figure 1 The structure of the electronic device shown is illustrated.
[0238] This application also provides a training device, such as... Figure 10 As shown, the training device 1000 includes: a second acquisition module 1001, a third processing module 1002, a fourth processing module 1003, and a training module 1004.
[0239] The second acquisition module 1001 is used to acquire multiple sets of multimodal sample data and the sample sentiment labels and sample sentiment reason text corresponding to each set of multimodal sample data;
[0240] The third processing module 1002 is used to process multimodal sample data based on the encoder and adapter in the initial multimodal large model to obtain the sample input features of the initial large language model in the initial multimodal large model;
[0241] The fourth processing module 1003 is used to process the sample input features based on multiple decoding layers in the initial large language model to obtain multiple intermediate prediction features; and to process the multiple intermediate prediction features based on multiple initial low-rank adaptive fine-tuning layers in the initial large language model to obtain the predicted sentiment label and the predicted sentiment reason text.
[0242] Training module 1004 is used to train the initial multimodal large model with predicted emotion labels and predicted emotion language text as the initial training output of the initial multimodal large model, and sample emotion labels and sample emotion reason text as supervision information. The initial multimodal large model is trained iteratively to obtain the trained multimodal large model.
[0243] In some examples, the training module 1004 is specifically used to determine the loss value of the multimodal large model based on the predicted emotion label, the predicted emotion language text, the sample emotion label, and the sample emotion reason text; and to iteratively update the weight parameters in multiple initial low-rank adaptive fine-tuning layers based on the loss value to obtain the trained multimodal large model.
[0244] This application also provides a chip system, such as... Figure 11 As shown, the chip system 1100 includes at least one processor 1101 and at least one interface circuit 1102. The processor 1101 and the interface circuit 1102 are interconnected via lines. For example, the interface circuit 1102 can be used to receive signals from other devices (e.g., the memory of an electronic device). As another example, the interface circuit 1102 can be used to send signals to other devices (e.g., the processor 1101).
[0245] For example, interface circuit 1102 can read instructions stored in memory and send those instructions to processor 1101. When the instructions are executed by processor 1101, the electronic device / training device can perform the steps in the above embodiments. Of course, the chip system may also include other discrete devices, and this application embodiment does not specifically limit this.
[0246] This application also provides a computer-readable storage medium including computer instructions that, when executed on an electronic device, cause the electronic device to perform the emotion recognition method provided in the foregoing embodiments.
[0247] This application also provides a computer program product containing executable instructions that, when run on an electronic device, cause the electronic device to perform the emotion recognition method provided in the foregoing embodiments.
[0248] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0249] In the several embodiments provided in this application, it should be understood that the disclosed apparatus / device and method can be implemented in other ways. For example, the apparatus / device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0250] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0251] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0252] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially or in other words, the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0253] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. An emotion recognition method, characterized in that, Applied to electronic devices, the method includes: Acquire multimodal dialogue data; The multimodal dialogue data is processed based on the encoder and adapter in the multimodal large model to obtain the input features of the large language model in the multimodal large model; The input features are processed by multiple decoding layers in the large language model to obtain dialogue response text and multiple intermediate features; the intermediate features are processed by multiple low-rank adaptive fine-tuning layers in the large language model to obtain emotion labels representing user emotions and emotion reason text representing the cause of emotions.
2. The method according to claim 1, characterized in that, The input features are processed based on multiple decoding layers in the large language model to obtain the dialogue response text and multiple intermediate features, including: The input features are processed based on the multiple first decoding layers among the multiple decoding layers to obtain the first intermediate feature among the multiple intermediate features; The first intermediate feature is processed based on multiple second decoding layers among the multiple decoding layers to obtain multiple second intermediate features among the multiple intermediate features; The last of the plurality of second intermediate features is subjected to text conversion processing to obtain the dialogue response text; wherein, the last second intermediate feature corresponds to the last second decoding layer among the plurality of second decoding layers.
3. The method according to claim 2, characterized in that, The process of processing the intermediate features using multiple low-rank adaptive fine-tuning layers in the large language model to obtain emotion tags representing user emotions and emotion reason text representing the reasons for emotions includes: The first intermediate feature and the multiple second intermediate features are processed based on multiple low-rank adaptive fine-tuning layers in the large language model to obtain multiple fine-tuned features; The last fine-tuning feature among the plurality of fine-tuning features is subjected to text conversion processing to obtain the emotion label and the emotion reason text; wherein, the last fine-tuning feature corresponds to the last low-rank adaptive fine-tuning layer among the plurality of low-rank adaptive fine-tuning layers.
4. The method according to claim 3, characterized in that, The plurality of second decoding layers include a first second decoding layer to a m-th second decoding layer; the m-th second decoding layer is the last second decoding layer among the plurality of second decoding layers; the plurality of second intermediate features include a first second intermediate feature to a m-th second intermediate feature; the step of processing the first intermediate feature based on the plurality of second decoding layers among the plurality of decoding layers to obtain the plurality of second intermediate features among the intermediate features includes: The first intermediate feature is processed based on the first second decoding layer to obtain the first second intermediate feature; The i-th second intermediate feature is processed based on the i-th second decoding layer to obtain the i-th second intermediate feature; where i is an integer greater than 1 and less than or equal to m.
5. The method according to claim 4, characterized in that, The plurality of low-rank adaptive fine-tuning layers include the first low-rank adaptive fine-tuning layer to the p-th low-rank adaptive fine-tuning layer; the p-th low-rank adaptive fine-tuning layer is the last low-rank adaptive fine-tuning layer among the plurality of low-rank adaptive fine-tuning layers; the plurality of fine-tuning features include the first fine-tuning feature to the j-th fine-tuning feature; Where j is an integer greater than 1 and less than or equal to p; the first intermediate feature and the multiple second intermediate features are processed by multiple low-rank adaptive fine-tuning layers in the large language model to obtain multiple fine-tuned features, including: The first intermediate feature is processed based on the first low-rank adaptive fine-tuning layer to obtain the first fine-tuned feature; Based on the (j-1)th fine-tuning feature and the (j-1)th second intermediate feature, the (j-1)th comprehensive feature is determined; The j-th comprehensive feature is processed by the j-th low-rank adaptive fine-tuning layer to obtain the j-th fine-tuned feature.
6. The method according to claim 5, characterized in that, The determination of the (j-1)th comprehensive feature based on the (j-1)th fine-tuned feature and the (j-1)th second intermediate feature includes: The (j-1)th fine-tuning feature and the (j-1)th second intermediate feature are added together to obtain the (j-1)th feature sum. The (j-1)th feature is determined as the (j-1)th comprehensive feature.
7. A method for training a multimodal large model, characterized in that, include: Obtain multiple sets of multimodal sample data and the corresponding sample sentiment labels and sample sentiment reason text for each set of multimodal sample data; The multimodal sample data is processed based on the encoder and adapter in the initial multimodal large model to obtain the sample input features of the initial large language model in the initial multimodal large model; The sample input features are processed based on multiple decoding layers in the initial large language model to obtain multiple intermediate prediction features; Based on the multiple initial low-rank adaptive fine-tuning layers in the initial large language model, the multiple intermediate prediction features are processed to obtain the predicted sentiment label and the predicted sentiment reason text. The predicted emotion labels and the predicted emotion language text are used as the initial training output of the initial multimodal large model, and the sample emotion labels and the sample emotion reason text are used as supervision information. The initial multimodal large model is iteratively trained to obtain the trained multimodal large model.
8. The method according to claim 7, characterized in that, The process of using the predicted emotion label and the predicted emotion text as the initial training output of the initial multimodal large model, and the sample emotion label and the sample emotion reason text as supervision information, iteratively training the initial multimodal large model to obtain the trained multimodal large model includes: Based on the predicted sentiment label, the predicted sentiment language text, the sample sentiment label, and the sample sentiment reason text, the loss value of the multimodal large model is determined; Based on the loss value, the weight parameters in the multiple initial low-rank adaptive fine-tuning layers are iteratively updated to obtain the trained multimodal large model.
9. An electronic device, characterized in that, The device includes a display screen, a memory, and one or more processors; the display screen, the memory, and the processors are coupled; wherein the memory stores computer program code, the computer program code including computer instructions, which, when executed by the processor, cause the electronic device to perform the emotion recognition method as described in any one of claims 1-6.
10. A training device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the executable instructions to implement the multimodal large model training method as described in claim 7 or 8.
11. A computer-readable storage medium, characterized in that, Includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the emotion recognition method as described in any one of claims 1-6.
12. A computer-readable storage medium, characterized in that, Includes computer instructions that, when executed on a training device, cause the training device to perform the multimodal large model training method as described in claim 7 or 8.