Emotion recognition method and device based on multi-modal information and related equipment

By acquiring facial images and voice data in real time, and combining self-attention mechanisms and emotion classification models, the system identifies and soothes the negative emotions of the parties involved, thus solving the problem of low communication efficiency and declining quality caused by the accumulation of emotions in legal practice, and improving communication efficiency and service quality.

CN121786725APending Publication Date: 2026-04-03BEIJING TAIXIN TIANCHENG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In legal practice, the accumulation of negative emotions between users and lawyers leads to low communication efficiency and a decline in the quality of legal services, with a lack of effective emotional guidance and intervention mechanisms.

Method used

By acquiring real-time facial images and voice data of the parties involved, and combining self-attention mechanisms and pre-trained emotion classification models, negative emotions are identified and positive emotional guidance information is displayed to lawyer users to soothe their emotions.

Benefits of technology

It improves the accuracy of emotion recognition, prevents lawyers from having their professional judgment affected by negative emotions, and enhances communication efficiency and the quality of legal services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121786725A_ABST
    Figure CN121786725A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an emotion recognition method and device based on multi-modal information and related equipment. The method comprises the following steps: respectively acquiring expression features and voice features of a party user in real time; state feedback of the lawyer user to the party user is obtained, a state evaluation label of the party user is obtained, the state evaluation label is coded, and auxiliary features are obtained; splicing the expression features, the voice features and the auxiliary features to obtain spliced features, and performing attention weighting processing on the spliced features based on a preset self-attention mechanism to obtain attention features; performing emotion recognition based on the attention features and a pre-trained emotion classification model to obtain an emotion recognition result; judging whether the emotion recognition result belongs to a negative emotion type or not; and if yes, displaying preset forward emotion guide information to the lawyer user. According to the method, when the party is in the negative emotion, the preset positive emotion guide information is displayed to the lawyer, and the lawyer is prevented from influencing professional judgment due to the negative emotion of the party.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of artificial intelligence technology, and in particular to an emotion recognition method, apparatus and related equipment based on multimodal information. Background Technology

[0002] In legal practice, core processes such as legal consultation and case communication have long suffered from both efficiency and quality issues due to emotional imbalances. This is essentially a one-way accumulation and two-way impact of negative emotions between users and lawyers, and an effective mechanism for emotional guidance and intervention has yet to be established. Specifically:

[0003] From the perspective of the users seeking legal advice, core scenarios (such as divorce disputes, debt collection, labor arbitration, and criminal defense) often involve the loss of core interests, infringement of personal rights, or significant life changes. Users are often already in a state of negative emotions such as tension, anxiety, anger, and helplessness before seeking consultation. Once the consultation begins, these negative emotions are easily triggered and amplified by case details, leading to problems such as disordered logical thinking and fragmented expression. Users may be unable to clearly outline key case points (such as the evidence formation process, details of rights infringement, and the course of communication and negotiation) chronologically or due to emotional resistance, and may even deliberately conceal core information (such as concealing transaction records in debt disputes or avoiding key points of conflict in divorce disputes), or even experience emotional breakdowns leading to communication breakdowns. This forces lawyers to spend a significant amount of time calming users and organizing fragmented information, greatly reducing the efficiency and accuracy of case information gathering, and consequently affecting the timeliness and relevance of subsequent case analysis and strategy development.

[0004] From the perspective of legal services, lawyers, as professional service providers, not only need to handle the transmission of negative emotions from users, but also need to deal with various conflict-ridden and tragic cases (such as domestic violence, serious torts, and criminal offenses) on a long-term basis, continuously receiving negative information from the cases themselves and negative emotional feedback from users. Due to the lack of effective emotional buffering and channeling mechanisms, lawyers are prone to problems such as emotional exhaustion and decreased empathy when exposed to a high-pressure, high-negative-emotional work environment for a long time. On the one hand, negative emotions may affect the lawyer's professional judgment, leading to a decrease in sensitivity to case details and biased legal analysis; on the other hand, long-term accumulated emotional pressure can lead to "professional burnout," manifested as decreased work enthusiasm and insufficient service patience, thereby affecting the overall quality of legal services and user experience. Summary of the Invention

[0005] This invention provides an emotion recognition method, apparatus, and related equipment based on multimodal information, aiming to solve the technical problem that traditional technologies cannot accurately identify the emotions of the parties involved and assist lawyers in emotional reassurance.

[0006] In a first aspect, embodiments of the present invention provide an emotion recognition method based on multimodal information, comprising:

[0007] The system acquires facial images of the user in real time and extracts features from the facial images to obtain expression features.

[0008] The voice data of the user is acquired in real time, and the voice features of the voice data are extracted. The voice features include at least speech rate, volume, fundamental frequency, number of pauses and pause duration.

[0009] Based on preset client status tags, obtain the lawyer user's status feedback to the client user, obtain the lawyer user's status evaluation tag for the client user, and encode the status evaluation tag to obtain auxiliary features;

[0010] The facial expression features, speech features, and auxiliary features are concatenated to obtain concatenated features, and attention-weighted processing is applied to the concatenated features based on a preset self-attention mechanism to obtain attention features;

[0011] Emotion recognition is performed based on the attention features and the pre-trained emotion classification model to obtain the emotion recognition result;

[0012] Determine whether the emotion recognition result belongs to the negative emotion type;

[0013] If so, pre-set positive emotional guidance information will be displayed to the lawyer user.

[0014] Secondly, embodiments of the present invention provide an emotion recognition device based on multimodal information, comprising:

[0015] The facial expression feature extraction module is used to acquire the user's facial image in real time and extract features from the facial image to obtain facial expression features;

[0016] The speech feature extraction module is used to acquire the speech data of the user in real time and extract the speech features of the speech data. The speech features include at least speech rate, volume, fundamental frequency, number of pauses and pause duration.

[0017] The auxiliary feature extraction module is used to obtain the lawyer user's status feedback to the client user based on the preset client status labels, obtain the lawyer user's status evaluation label to the client user, and encode the status evaluation label to obtain auxiliary features;

[0018] The feature fusion module is used to concatenate the facial expression features, speech features, and auxiliary features to obtain concatenated features, and to perform attention weighting processing on the concatenated features based on a preset self-attention mechanism to obtain attention features.

[0019] An emotion recognition module is used to perform emotion recognition based on the attention features and a pre-trained emotion classification model to obtain the emotion recognition result.

[0020] The judgment module is used to determine whether the emotion recognition result belongs to the negative emotion type;

[0021] The first display module is used to display preset positive emotional guidance information to the lawyer user if the situation is as described.

[0022] Thirdly, embodiments of the present invention provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the emotion recognition method based on multimodal information described in the first aspect.

[0023] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the emotion recognition method based on multimodal information described in the first aspect.

[0024] This invention provides an emotion recognition method, apparatus, and related equipment based on multimodal information. The method acquires facial images of the client user in real time and extracts features from these images to obtain facial expression features; it also acquires voice data of the client user in real time and extracts voice features from the voice data; based on preset client status labels, it acquires the lawyer user's status feedback to the client user to obtain a lawyer user's status evaluation label, and encodes this label to obtain auxiliary features; it concatenates the facial expression features, voice features, and auxiliary features to obtain concatenated features, and performs attention-weighted processing on the concatenated features based on a preset self-attention mechanism to obtain attention features; it performs emotion recognition based on the attention features and a pre-trained emotion classification model to obtain the emotion recognition result; it determines whether the emotion recognition result belongs to a negative emotion type; if so, it displays preset positive emotion guidance information to the lawyer user. This method combines visual information, voice information, and lawyers' feedback on the client's state to identify the client's emotions, improving the accuracy of emotion identification. When the client is in a negative emotion, it promptly displays pre-set positive emotion guidance information to the lawyer user, preventing the lawyer user from being affected by the client's emotions in legal services, thus avoiding a decrease in the lawyer user's professional judgment, reduced sensitivity to case details, and biased legal analysis. Attached Figure Description

[0025] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 This is a flowchart illustrating an embodiment of the emotion recognition method based on multimodal information provided by the present invention.

[0027] Figure 2 This is a schematic block diagram of an emotion recognition device based on multimodal information provided in an embodiment of the present invention. Detailed Implementation

[0028] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0029] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0030] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0031] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0032] Please see Figure 1 This is a flowchart illustrating an emotion recognition method based on multimodal information provided in an embodiment of the present invention. The method includes steps S110 to S170.

[0033] Step S110: Acquire the facial image of the user in real time, and extract features from the facial image to obtain expression features;

[0034] Step S120: Acquire the voice data of the user in real time and extract the voice features of the voice data. The voice features include at least speech rate, volume, fundamental frequency, number of pauses and pause duration.

[0035] Step S130: Based on the preset client status tags, obtain the lawyer user's status feedback to the client user, obtain the lawyer user's status evaluation tag to the client user, and encode the status evaluation tag to obtain auxiliary features;

[0036] In this embodiment, facial images of the user are acquired in real time via multimedia (such as a camera). These images are preprocessed by adjusting the resolution of each frame to a uniform size. A neural network is used to detect the facial region, outputting facial coordinates and cropping a partial facial image. The cropped facial image is then aligned using affine transformations (rotation, translation, scaling) based on 68 facial key points (corners of the eyes, tip of the nose, corners of the mouth, etc.) to ensure the eyes are level and the tip of the nose is centered, eliminating the impact of pose differences on AU detection. Furthermore, histogram equalization is used to enhance the contrast of facial textures, and pixel values ​​are normalized by subtracting the mean of the pre-trained model to ensure the input data conforms to the model's training distribution, resulting in a standard image.

[0037] In one embodiment, based on the AU definition of FaceReader (a facial expression analysis system), a CNN hybrid model is used to extract facial expression features from facial images, such as the intensity (0-5 levels) of 44 action units, including eyelid closure (AU43), frowning (AU4), and widening (AU5). FaceReader is the world's first commercially developed automatic facial expression analysis tool, allowing users to objectively assess their emotional changes and identify emotions by combining different AUs. The following are basic emotion-related AUs:

[0038] Anger: Frowning (AU4), eyelids tightening (AU7), lower lip raised (AU17);

[0039] Sadness: Drooping eyebrows (AU3+AU4), stretched corners of the mouth (AU20);

[0040] Surprise: Eyelids lift (AU5), eyebrows raise (AU1);

[0041] Fear: Eyelids tighten (AU7), eyebrows raised (AU1);

[0042] Disgust: wrinkled nose (AU10), raised upper lip (AU12);

[0043] Joy: Upturned corners of the mouth (AU12), lifted cheeks (AU6).

[0044] Furthermore, the CNN hybrid model adopts a hybrid structure of "main trunk feature extraction + branch AU regression", which is divided into the following modules:

[0045] I. Backbone Network:

[0046] ResNet50 was selected as the backbone network. The top classification layer was removed, and one convolutional block, four residual blocks and one global average pooling layer were retained to extract global facial features (such as facial contours and muscle distribution). The Conv1 layer outputs 256-dimensional shallow features (capturing edges and textures), and the Conv4 layer outputs 2048-dimensional deep features (capturing complex muscle movement patterns, such as the contraction of the glabella muscles when frowning).

[0047] II. AU Branch Network (Feature Refinement):

[0048] For the anatomical locations of 44 AUs (e.g., AU43 near the eye, AU4 between the eyebrows, and AU5 near the eye), 12 local branch networks were designed:

[0049] Eye branch: Focusing on the peri-eye region (cropped from 1 / 4 of a 224×224 face image), local features of eye AU such as eyelid closure (AU43) and staring (AU5) are extracted through two 3×3 convolutional layers;

[0050] Glabellar branch: Focusing on the glabellar region, local features of frowning (AU4) are extracted through two 3×3 convolutional layers;

[0051] Mouth branch: Focusing on the perioral region, extracting mouth AU features such as corner stretching (AU12) and pouting (AU15);

[0052] Each branch network outputs 512-dimensional local features, which are then concatenated with the deep features of the backbone network to form a fused feature of "global + local" (2048 + 512 = 2560 dimensions).

[0053] III. AU Intensity Regression

[0054] The spliced ​​and fused features are processed by three fully connected layers (FC layers, 2560→1024→512→44) to output 44 continuous intensity values ​​(range 0-5) to obtain facial expression features.

[0055] In one embodiment, the user's voice data is acquired in real time, and voice features are extracted based on the voice data. The voice features include speech rate (syllables / second), volume (dB value), fundamental frequency (F0) fluctuation, and number / duration of pauses. Specific features are designed for the target emotion: anxiety corresponds to "speech rate increase ≥30% + increase in fundamental frequency standard deviation", and anger corresponds to "peak volume ≥85dB + increase in mean fundamental frequency".

[0056] In one embodiment, structured feedback labels are designed, and lawyer users mark six categories of labels through the terminal, such as "vague user expression", "emotionally agitated", "logical contradiction", "unclear needs", "missing key information", and "redundant expression", which are then converted into auxiliary feature vectors encoded by 0-1.

[0057] Step S140: The facial expression features, speech features and auxiliary features are concatenated to obtain concatenated features, and attention-weighted processing is performed on the concatenated features based on a preset self-attention mechanism to obtain attention features;

[0058] In this embodiment, the 44-dimensional facial AU intensity vector, 12-dimensional speech feature vector, and 6-dimensional lawyer feedback vector are first concatenated into a 62-dimensional basic feature. Then, the correlation weight between each feature and the target emotion is calculated through a self-attention mechanism, and the contribution of each feature is dynamically adjusted to obtain the attention feature. Self-attention is a mechanism that allows each element (such as a word) in the input sequence to calculate its correlation weight with other elements through query, key, and value vectors, and finally aggregates the global information in a weighted manner.

[0059] Step S150: Perform emotion recognition based on the attention features and the pre-trained emotion classification model to obtain the emotion recognition result;

[0060] Step S160: Determine whether the emotion recognition result belongs to the negative emotion type;

[0061] Step S170: If so, display preset positive emotion guidance information to the user in question.

[0062] In this embodiment, attention features are input into a pre-trained emotion classification model for emotion recognition, yielding the emotion recognition result. To prevent the client's emotions from remaining consistently negative, leading to low communication efficiency, the system determines whether the emotion recognition result belongs to a negative emotion type. When a negative emotion is detected, pre-set positive emotion guidance information is displayed to the lawyer user on the screen—such as a soft blue background with dynamic gradient (color psychology confirms that blue can reduce anxiety) and a slowly looping "breathing guidance animation" (such as an undulating wave icon to guide the client's breathing). This visually conveys a sense of "calm" to the lawyer user, preventing the lawyer's professional judgment from being affected by the client's negative emotions, thus reducing sensitivity to case details and causing biases in legal analysis. Furthermore, to prevent the client from remaining immersed in negative emotions, the system monitors the number of consecutive negative emotion types in the emotion recognition result. When the number reaches a preset threshold (e.g., 3 times), pre-set emotion guidance scripts are displayed to the lawyer user to help soothe the client's emotions, prevent the cumulative spread of negative emotions, and generate an emotion status report to remind the lawyer user to take a break when necessary.

[0063] In one embodiment, the training process of the emotion classification model includes: acquiring a sample facial expression feature set, a sample speech dataset, and a corresponding sample auxiliary feature set; concatenating the sample facial expression feature set, the sample speech dataset, and the corresponding sample auxiliary feature set one by one to construct a sample concatenated feature set; inputting the sample concatenated feature set and the corresponding real emotion label into an initial Transformer classification model for emotion classification training to obtain the model classification result; calculating the classification loss between the model classification result and the corresponding real emotion label according to a preset loss function; and backpropagating the model parameters of the initial Transformer classification model according to the classification loss until the model parameters of the initial Transformer classification model converge to obtain the emotion classification model.

[0064] This method acquires facial images of the client in real time and extracts features from these images to obtain facial expression features; it also acquires voice data of the client in real time and extracts voice features from the voice data; based on preset client status labels, it obtains the lawyer's status feedback to the client to obtain the lawyer's status evaluation label, and encodes the status evaluation label to obtain auxiliary features; it concatenates the facial expression features, voice features, and auxiliary features to obtain concatenated features, and applies attention weighting processing to the concatenated features based on a preset self-attention mechanism to obtain attention features; it performs emotion recognition based on the attention features and a pre-trained emotion classification model to obtain the emotion recognition result; it determines whether the emotion recognition result belongs to the negative emotion type; if so, it displays preset positive emotion guidance information to the lawyer. This method combines visual information, voice information, and the lawyer's status feedback to the client to identify the client's emotions, improving the accuracy of emotion recognition. It also promptly displays preset positive emotion guidance information to the lawyer when the client is in a negative emotion, preventing the lawyer from being influenced by the client's emotions in legal services, thus avoiding a decrease in the lawyer's sensitivity to case details and deviations in legal analysis.

[0065] This invention also provides an emotion recognition device based on multimodal information, which is used to execute any of the aforementioned embodiments of the emotion recognition method based on multimodal information. Specifically, please refer to... Figure 2 , Figure 2 This is a schematic block diagram of an emotion recognition device based on multimodal information provided in an embodiment of the present invention. The emotion recognition device 100 based on multimodal information can be configured in a server.

[0066] like Figure 2 As shown, the emotion recognition device 100 based on multimodal information includes an expression feature extraction module 110, a voice feature extraction module 120, an auxiliary feature extraction module 130, a feature fusion module 140, an emotion recognition module 150, a judgment module 160, and a first display module 170.

[0067] The facial expression feature extraction module 110 is used to acquire the facial image of the user in real time and extract features from the facial image to obtain facial expression features;

[0068] The speech feature extraction module 120 is used to acquire the speech data of the user in real time and extract the speech features of the speech data. The speech features include at least speech rate, volume, fundamental frequency, number of pauses and pause duration.

[0069] The auxiliary feature extraction module 130 is used to obtain the lawyer user's status feedback to the client user based on the preset client status labels, obtain the lawyer user's status evaluation label to the client user, and encode the status evaluation label to obtain auxiliary features;

[0070] The feature fusion module 140 is used to concatenate the facial expression features, speech features and auxiliary features to obtain concatenated features, and to perform attention weighting processing on the concatenated features based on a preset self-attention mechanism to obtain attention features;

[0071] The emotion recognition module 150 is used to perform emotion recognition based on the attention features and the pre-trained emotion classification model to obtain the emotion recognition result;

[0072] The judgment module 160 is used to determine whether the emotion recognition result belongs to the negative emotion type;

[0073] The first display module 170 is used to display preset positive emotional guidance information to the lawyer user if the situation is as described.

[0074] In one embodiment, the emotion recognition device 100 based on multimodal information further includes:

[0075] The monitoring module is used to monitor the number of times the emotion recognition result continuously belongs to the negative emotion type;

[0076] The second display module is used to display preset emotional guidance scripts to the lawyer user if the number of times reaches a preset threshold.

[0077] This invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the emotion recognition method based on multimodal information as described above.

[0078] In another embodiment of the invention, a computer-readable storage medium is provided. This computer-readable storage medium may be a non-volatile computer-readable storage medium. The computer-readable storage medium stores a computer program that, when executed by a processor, causes the processor to perform the emotion recognition method based on multimodal information as described above.

[0079] Those skilled in the art will readily understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this invention.

[0080] In the embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Units with the same function may be grouped into one unit. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, devices, or units, or it may be an electrical, mechanical, or other form of connection.

[0081] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of the present invention, depending on actual needs.

[0082] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0083] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks.

[0084] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. An emotion recognition method based on multimodal information, characterized in that, include: The system acquires facial images of the user in real time and extracts features from the facial images to obtain expression features. The voice data of the user is acquired in real time, and the voice features of the voice data are extracted. The voice features include at least speech rate, volume, fundamental frequency, number of pauses and pause duration. Based on preset client status tags, obtain the lawyer user's status feedback to the client user, obtain the lawyer user's status evaluation tag for the client user, and encode the status evaluation tag to obtain auxiliary features; The facial expression features, speech features, and auxiliary features are concatenated to obtain concatenated features, and attention-weighted processing is applied to the concatenated features based on a preset self-attention mechanism to obtain attention features; Emotion recognition is performed based on the attention features and the pre-trained emotion classification model to obtain the emotion recognition result; Determine whether the emotion recognition result belongs to the negative emotion type; If so, pre-set positive emotional guidance information will be displayed to the lawyer user.

2. The emotion recognition method based on multimodal information as described in claim 1, characterized in that, After determining whether the emotion recognition result belongs to the negative emotion type, the process includes: Monitor the number of consecutive negative emotion types identified by the emotion recognition result; If the number of occurrences reaches a preset threshold, a pre-set emotional guidance script will be displayed to the lawyer user.

3. The emotion recognition method based on multimodal information as described in claim 1, characterized in that, The training process of the emotion classification model includes: Obtain the sample facial expression feature set, sample speech dataset, and corresponding sample auxiliary feature set, and concatenate the sample facial expression feature set, sample speech dataset, and corresponding sample auxiliary feature set one by one to construct the sample concatenated feature set; The sample concatenation feature set and the corresponding real emotion label are input into the initial Transformer classification model for emotion classification training to obtain the model classification result; The classification loss between the model classification result and the corresponding real emotion label is calculated according to the preset loss function, and the model parameters of the initial Transformer classification model are backpropagated according to the classification loss until the model parameters of the initial Transformer classification model converge, thus obtaining the emotion classification model.

4. The emotion recognition method based on multimodal information as described in claim 1, characterized in that, The step of extracting features from the facial image to obtain expression features includes: The facial image is preprocessed to obtain a standard image; The standard image is subjected to multi-dimensional deep feature extraction through the backbone network in the pre-set CNN hybrid model to obtain multi-dimensional deep features. The standard image is subjected to multi-dimensional shallow feature extraction through the branch network in the CNN hybrid model to obtain multi-dimensional shallow features. The facial expression features are obtained by splicing and fusing each deep feature with its corresponding shallow feature.

5. The emotion recognition method based on multimodal information as described in claim 4, characterized in that, The backbone network adopts the ResNet50 network, which includes one convolutional block, four residual blocks and one global average pooling layer.

6. An emotion recognition device based on multimodal information, characterized in that, include: The facial expression feature extraction module is used to acquire the user's facial image in real time and extract features from the facial image to obtain facial expression features; The speech feature extraction module is used to acquire the speech data of the user in real time and extract the speech features of the speech data. The speech features include at least speech rate, volume, fundamental frequency, number of pauses and pause duration. The auxiliary feature extraction module is used to obtain the lawyer user's status feedback to the client user based on the preset client status labels, obtain the lawyer user's status evaluation label to the client user, and encode the status evaluation label to obtain auxiliary features; The feature fusion module is used to concatenate the facial expression features, speech features, and auxiliary features to obtain concatenated features, and to perform attention weighting processing on the concatenated features based on a preset self-attention mechanism to obtain attention features. An emotion recognition module is used to perform emotion recognition based on the attention features and a pre-trained emotion classification model to obtain the emotion recognition result. The judgment module is used to determine whether the emotion recognition result belongs to the negative emotion type; The first display module is used to display preset positive emotional guidance information to the lawyer user if the situation is as described.

7. The emotion recognition device based on multimodal information as described in claim 6, characterized in that, Also includes: The monitoring module is used to monitor the number of times the emotion recognition result continuously belongs to the negative emotion type; The second display module is used to display preset emotional buffer information to the lawyer user if the number of times reaches a preset threshold.

8. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the emotion recognition method based on multimodal information as described in any one of claims 1 to 5.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, causes the processor to perform the emotion recognition method based on multimodal information as described in any one of claims 1 to 5.