Emotion recognition method and system based on large model

By collecting user's active and passive emotion data and combining it with scene information, and embedding semantic guidance content to dynamically update the emotion recognition model, the problem of insufficient accuracy and adaptability of emotion recognition in existing technologies is solved, and efficient and accurate user emotion recognition is achieved.

CN121661694APending Publication Date: 2026-03-13JIANGSU ZHUODUN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-10
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing emotion recognition technologies struggle to quickly and accurately identify users' true emotions, especially for new users or those with low usage frequency. Furthermore, they lack adaptability and flexibility to changes in users' emotional expression patterns, resulting in recognition results lagging behind users' actual emotional states.

Method used

By collecting users' active and passive emotion data, combining it with scene context information, embedding semantic guidance content in real time to obtain unconscious annotations, dynamically updating the emotion recognition model, and using a personalized expression preference adapter to fuse with the standard emotion recognition model, the operating status is dynamically adjusted to improve recognition accuracy and adaptability.

Benefits of technology

The model achieves efficient and accurate emotion recognition, adapting to users' personalized and subtle facial expression changes, improving recognition accuracy and adaptability, reducing initial training costs and time, and enhancing the reliability of recognition results and resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121661694A_ABST
    Figure CN121661694A_ABST
Patent Text Reader

Abstract

The invention provides an emotion recognition method and system based on a large model, and relates to the technical field of emotion recognition. The method comprises the following steps: after receiving an initial starting instruction of a user side, collecting active emotion data of a user to finish initial training of a model; after a user starts a function, passive emotion data containing natural facial expressions are collected in real time, passive candidate data are generated in combination with scene contexts such as use scene types and operation behavior records, and to-be-labeled data are obtained through feature screening. When the to-be-labeled data reaches a trigger condition, semantic guidance content associated with expression features is embedded in natural interaction, user feedback is obtained, unconscious labeling is determined in combination with facial expressions during feedback, a model is dynamically updated accordingly, finally, real-time facial expression data flow is analyzed through the updated model, and an emotion recognition result is obtained. According to the method, active and passive data and scene information are combined, the accuracy and adaptability of emotion recognition are improved, and the real emotion of the user can be recognized more accurately and quickly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of emotion recognition technology, and in particular to an emotion recognition method and system based on a large model. Background Technology

[0002] With the rapid development of artificial intelligence technology, emotion recognition has become an important research area in human-computer interaction. Existing emotion recognition technologies primarily collect users' facial expressions, speech, and text information, and then process and analyze it using cloud-based, general-purpose emotion recognition models to identify the user's emotional state. However, this general-purpose model-based emotion recognition method has significant limitations: because everyone's emotional expression is different, and there are significant individual differences, general-purpose models struggle to accurately capture the emotional characteristics of different users, resulting in low recognition accuracy. This is especially true when users are in complex or subtle emotional states, where the recognition performance is even less than ideal.

[0003] To address these issues, some researchers have proposed emotion recognition methods based on personalized learning. This method collects historical emotion data from specific users to build a personal emotion model library. Then, it combines transfer learning techniques to specifically adjust and optimize general emotion models, making them more closely reflect the individual user's emotional expression characteristics. This approach can improve the accuracy of emotion recognition to a certain extent, especially for users with high usage frequency; the recognition effect gradually improves as data accumulates.

[0004] However, existing methods struggle to quickly and accurately determine a user's true emotions. For new users or those with low usage frequency, the system requires a long period of data accumulation to achieve ideal recognition results. Furthermore, much of the collected data is gathered while the user is consciously aware of the situation, which may not reflect the user's most genuine and natural feelings. Moreover, existing methods lack sufficient adaptability and flexibility to adjust the model promptly when a user's emotional expression patterns change, resulting in recognition results lagging behind the user's actual emotional state and failing to meet the demands for real-time and efficient emotional interaction. Summary of the Invention

[0005] This application provides a large-model-based emotion recognition method and system for efficiently and accurately identifying user emotions.

[0006] Firstly, this application provides a large-scale model-based emotion recognition method, which includes: upon receiving an initial activation command from the user terminal, collecting the user's active emotion data for initial training of the emotion recognition model; after the user activates the emotion recognition function, collecting passive emotion data containing the user's natural facial expression data stream in real time, and generating passive candidate data by combining the user's current scene context information, including the usage scene type and operation behavior records; performing feature filtering on the passive candidate data to obtain data to be labeled; when the data to be labeled meets the triggering condition, embedding semantic guidance content in the natural interaction scene between the user and the user terminal, the semantic guidance content being associated with the facial expression features corresponding to the data to be labeled; obtaining the user's feedback information on the semantic guidance content, and determining the unconscious label corresponding to the data to be labeled by combining the facial expression features at the time of feedback; dynamically updating the emotion recognition model based on the unconscious label; and dynamically analyzing the acquired real-time facial expression data stream through the updated emotion recognition model to obtain the emotion recognition result.

[0007] By employing the aforementioned technical solution, passive emotion data is collected during natural interactions and combined with scene context to accurately filter out "data to be labeled" that the model struggles to understand. Subsequently, an innovative approach of embedding semantic guidance is used to acquire "unconscious annotations" with real-world contextual meaning without disrupting the user's normal experience. This high-quality, real-world-derived labeled data is used to dynamically update the model, enabling it to continuously learn and adapt to the user's personalized and subtle facial expression changes. The entire process achieves a closed loop from initial model development to dynamic optimization. By combining active and passive data with scene information, the accuracy and adaptability of emotion recognition are improved, allowing for a more precise reflection of the user's true emotions.

[0008] In conjunction with some embodiments of the first aspect, in some embodiments, after receiving the initial activation instruction from the user terminal, the step of collecting the user's active emotion data for the initial training of the emotion recognition model includes: sending an emotion guidance instruction to the user terminal to guide the user to make corresponding facial expressions according to a preset emotion type, and collecting it as a first expression sample; presenting the user with immersive incentive content related to the preset emotion type, and collecting a second expression sample; constructing an expression calibration pair between the first expression sample and the second expression sample, and training to obtain a personalized expression preference adapter; and integrating the personalized expression preference adapter with a standard emotion recognition model pre-deployed in the cloud to determine the final emotion recognition model, wherein the standard emotion recognition model is pre-trained in the cloud using machine learning based on different facial expressions and corresponding emotion labels of multiple different users.

[0009] By employing the aforementioned technical solution, the differences in emotional expression of the same user under two states—"deliberate expression" and "natural expression"—were successfully captured. This calibration pair is key to training a personalized facial expression preference adapter, which learns the user's unique facial expression habits. Ultimately, this lightweight adapter is integrated with a powerful standard emotion recognition model, rather than being trained from scratch. This leverages the extensive knowledge of the cloud-based model while injecting precise individualized features. This not only improves the accuracy of the initial model and the depth of understanding of individual user expressions but also reduces the cost and time of initial training.

[0010] In conjunction with some embodiments of the first aspect, in some embodiments, the step of dynamically analyzing the acquired real-time facial expression data stream using an updated emotion recognition model to obtain emotion recognition results specifically includes: generating a temporal feature sequence containing a multidimensional emotion probability distribution and confidence score based on the real-time facial expression data stream; determining three indicators based on the temporal feature sequence within a preset time window, including an emotion confidence change trend, a dominant emotion stability index, and an emotion change gradient. The emotion confidence change trend is obtained by calculating the mean and variance of the confidence score within the preset time window; the dominant emotion stability index is obtained by analyzing the duration and switching frequency of the dominant emotion type; and the emotion change gradient is obtained by calculating the difference between the vectors of the emotion probability distribution at adjacent time points. When the emotion confidence change trend is stable and higher than a first threshold, When the dominant emotion stability index is higher than the second threshold, a low-power monitoring state is entered, reducing the frequency of facial expression data collection and local computing resource allocation. When all three indicators are within the standard fluctuation range, the standard local analysis state is maintained at the user end, maintaining the preset standard facial expression data collection frequency and computing resource allocation. When the emotion confidence change trend is continuously lower than the third threshold, or the dominant emotion stability index is lower than the fourth threshold, or the emotion change gradient is higher than the fifth threshold, a cloud collaborative analysis state is triggered, wherein the third threshold is lower than the first threshold, and the fourth threshold is lower than the third threshold. In the cloud collaborative analysis state, the facial expression data of the corresponding emotion event and the historical data within the previous time window are packaged and uploaded to the cloud for analysis. The cloud analysis results are fused and calibrated with the local analysis results to obtain the final emotion recognition result.

[0011] By adopting the above technical solution, a time-series feature sequence is generated and three indicators are determined. Different operating states are switched based on these indicators. When stable, a low-power state is entered to save resources; during standard fluctuations, a standard state is maintained to ensure recognition quality; and during anomalies, cloud collaboration is triggered to improve analysis accuracy. Dynamically adjusting the state balances recognition accuracy and resource consumption. Combined with cloud-based fusion calibration, the reliability and operational efficiency of emotion recognition results are further improved.

[0012] In conjunction with some embodiments of the first aspect, in some embodiments, when the data to be labeled meets the triggering condition, the step of embedding semantic guidance content in the natural interaction scenario between the user and the terminal specifically includes: extracting the facial expression features corresponding to the data to be labeled and establishing an unlabeled facial expression feature library; during the interaction between the user and the terminal, monitoring the user's current facial expression in real time, extracting the current feature information of the current facial expression, and matching it with the feature information in the unlabeled facial expression feature library; if the matching degree between the current feature information and the target feature information in the unlabeled facial expression feature library exceeds a preset matching degree threshold, locking the continuous time period in which the current facial expression appears, and recording the current interaction scenario information context within the continuous time period; generating interactive statements related to the scenario task based on the current interaction scenario information context, and embedding the semantic guidance content into the interactive statements to present to the user.

[0013] By adopting the above technical solution, an unlabeled facial expression feature library is established. This library matches the user's current facial expression features in real time, locks the time period and records the scene when a threshold is exceeded, and generates relevant interactive statements to embed semantic guidance. This allows for precise triggering of guidance within natural interactions, generating content based on the scene, avoiding abrupt interruptions to the user, improving user cooperation, and making the obtained feedback more natural and authentic, thus improving the quality of labeled data.

[0014] In conjunction with some embodiments of the first aspect, in some embodiments, the steps for when the data to be labeled reaches the trigger condition specifically include: performing preliminary emotion recognition on the data to be labeled to obtain a preliminary emotion probability distribution and confidence score; determining that the data to be labeled has reached the trigger condition when at least one of the following preset conditions is met: the confidence score is lower than a preset confidence threshold; in the preliminary emotion probability distribution, the difference between the highest probability value and the second highest probability value is lower than a preset discrimination threshold; the distance between the facial expression feature corresponding to the data to be labeled and the existing features in the user's historical facial expression feature library exceeds a preset novelty threshold.

[0015] By adopting the above technical solution, this method defines clear and quantifiable trigger boundaries for the selection of "data to be labeled." It is no longer a vague judgment, but based on three specific indicators: low confidence scores directly identify inaccurate model recognition results; low discrimination thresholds accurately capture the model's "hesitant" state between two or more similar emotions; and high novelty thresholds ensure that the system can discover and learn novel emotional expressions never before seen by the user. These three dimensions constitute a comprehensive selection matrix, making the triggered labeling highly targeted. This ensures that each interactive labeling acts on the key points where the model most needs to learn and improve, thereby enhancing the efficiency and effectiveness of model dynamic updates.

[0016] In conjunction with some embodiments of the first aspect, in some embodiments, the step of obtaining user feedback information on the semantic guidance content and determining the unconscious annotation corresponding to the data to be annotated based on facial expression features during feedback specifically includes: collecting comprehensive feedback signals from the user during feedback interaction, including voice replies, text input, body movements, or facial expression changes; classifying the feedback based on the comprehensive feedback signal: if it is determined to be direct affirmative or negative feedback, it is directly associated with the corresponding preset emotion tag and determined to be unconscious annotation; if it is determined to be ambiguous or no feedback, it is marked as a low-confidence tag or the current annotation operation is ignored; using the facial expression features during user feedback as a calibration factor, when the facial expression features during user feedback match the comprehensive feedback signal, the weight of the unconscious annotation is increased according to the degree of matching; if the facial expression features during user feedback do not match the comprehensive feedback signal, the weight of the unconscious annotation is decreased according to the degree of matching or a second confirmation is performed.

[0017] By adopting the above technical solutions, multi-dimensional feedback combined with facial expression calibration makes unconscious annotation more accurate, reduces the impact of erroneous annotation, improves the effectiveness of annotation data, provides high-quality basis for model updates, and enhances the dynamic optimization effect of the model.

[0018] In some embodiments of the first aspect, after the step of dynamically updating the emotion recognition model using the unconscious annotation, the method further includes: periodically calculating the feature deviation between the active emotion data and the data corresponding to the unconscious annotation; if the feature deviation exceeds a preset deviation threshold, triggering lightweight active training to obtain supplementary facial expression data, wherein the active training includes sending targeted emotion guidance instructions to the user terminal, guiding the user to supplement and display facial expressions corresponding to the emotion types with feature deviations; and correcting the emotion recognition model by combining the supplementary facial expression data and the unconscious annotation data of the corresponding emotion types.

[0019] By employing the aforementioned technical solution, the model, during its continuous learning of unconsciously labeled data, may gradually deviate from its initial cognitive baseline established by actively generated emotion data due to data bias. This solution checks the model's health by periodically calculating feature bias. Once the bias exceeds a threshold, the emotion recognition system triggers a lightweight, targeted active training. Unlike the initial training, this is not comprehensive; instead, it only requires the user to supplement the display of emotional expressions where cognitive bias has occurred. This mechanism acts as a calibration anchor, periodically pulling the model back onto the correct track without frequently disturbing the user, ensuring the accuracy and stability of the emotion recognition model during its long-term self-evolution.

[0020] In a second aspect, this application provides an emotion recognition system, which includes: one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code, the computer program code including computer instructions, and the one or more processors call the computer instructions to cause the emotion recognition system to perform the method described in the first aspect and any possible implementation thereof. Attached Figure Description

[0021] Figure 1 This is a flowchart illustrating an emotion recognition method based on a large model in an embodiment of this application. Figure 2 This is another flowchart illustrating the emotion recognition method based on a large model in this application embodiment; Figure 3 This is a schematic diagram of the physical device structure of an emotion recognition system in the embodiments of this application. Detailed Implementation

[0022] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used herein, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the listed items.

[0023] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.

[0024] For ease of understanding, the method provided in this implementation is described in process below. Please refer to [link / reference]. Figure 1 This is a flowchart illustrating an emotion recognition method based on a large model in an embodiment of this application.

[0025] S101. After receiving the initial activation command from the user terminal, collect the user's active emotion data for the initial training of the emotion recognition model. "User device" refers to the terminal device used by the user, such as smartphones, tablets, smartwatches, smart home robots, etc., used to run emotion recognition-related functions; "active emotion data" refers to the data generated when the user consciously displays a specific emotion under guidance, including corresponding facial expression information.

[0026] When a user first activates the emotion recognition function, the system currently lacks any emotion data about that user. Therefore, it needs to actively collect basic data to initialize the emotion recognition model. Typical application scenarios include users using the emotion recognition function of a related application or device for the first time, such as opening a health app or smart educational device with emotion sensing capabilities for the first time.

[0027] Once the emotion recognition system receives the initial activation command from the user, it begins the process of collecting proactive emotion data and initial model training. First, the system guides the user to provide proactive emotion data. This is because different users have different facial features and expression habits, and directly applying a generic model may result in low recognition accuracy. The proactive data collection process involves a series of guided steps, such as prompting the user to make facial expressions of common emotions like "happy," "angry," "sad," and "surprised," while simultaneously recording image or video data of these expressions using the user's camera or other devices.

[0028] After collecting sufficient proactive emotion data, the system preprocesses this data, including facial feature point extraction and expression region segmentation, transforming the raw data into feature vectors that the model can recognize. Subsequently, this proactive data with clear emotion labels is used to train and adjust the base model. The base model may be a pre-trained general emotion recognition model. Through initial training, the user's personalized facial expression features are incorporated, enabling the model to more accurately identify the user's emotions. For example, for the emotion of "happiness," different users may exhibit different behaviors; some users may raise their lips significantly and show their teeth, while others may only smile slightly. The initial training aims to help the model memorize the typical facial expression features of the user under specific emotions, laying the foundation for subsequent real-time emotion recognition.

[0029] In some embodiments, this step can be implemented as follows: The system sends emotion guidance instructions to the user. These instructions are designed based on preset emotion types to ensure that the user can understand and make the corresponding facial expressions. Corresponding expression images may be attached as references. After receiving the instructions, the user will convey them to the user through screen display or voice broadcast, guiding the user to actively make the corresponding facial expressions. Simultaneously, the system uses the user's camera to capture facial images or videos during this process, forming a first expression sample. The first expression sample is characterized by the user consciously making expressions according to the instructions, clearly corresponding to the preset emotion type, but may have a certain degree of deliberate intent.

[0030] The system then presents immersive, motivating content related to preset emotion types to the user through the app. This content is carefully designed based on the emotion type to maximize the natural expression of the corresponding emotion. For example, for the emotion type "sadness," a touching movie clip or sad music might be presented; for "anger," an infuriating news report might be shown. Users often unconsciously express genuine emotional reactions while experiencing this content, and the facial expression data collected by the system at this time becomes the second expression sample. The second expression sample is closer to the user's natural emotional expression and can compensate for any deliberate flaws in the first expression sample. Then, the first and second expression samples are combined to construct an expression calibration pair. Each preset emotion type corresponds to such a calibration pair, where the first sample provides a standard reference for the emotion, and the second sample provides the characteristics of how the user naturally expresses that emotion. By training these expression calibration pairs, the system can learn the differences and commonalities in how the user expresses the same emotion in different contexts, thereby generating a personalized expression preference adapter. This adapter can capture the user's unique facial expression habits, such as some users habitually squinting when happy, while others habitually open their mouths to laugh.

[0031] Finally, the personalized expression preference adapter is integrated into the standard emotion recognition model pre-deployed in the cloud to form the final emotion recognition model. The standard emotion recognition model is trained on a large amount of facial expression data from different users and has the ability to recognize common emotional characteristics, but it may not be able to accurately adapt to the individual differences of each user. By integrating the personalized expression preference adapter, the standard model can be fine-tuned, allowing it to better recognize the emotional expression characteristics of specific users while retaining its universality. For example, the standard model might have a recognition threshold of 15 degrees for "smiling," while the personalized adapter, through learning, discovers that the user's mouth corners typically rise 10 degrees when smiling. After integration, the model will adjust this recognition threshold, improving the accuracy of recognizing the user's smiling expression.

[0032] S102. After the user enables the emotion recognition function, passive emotion data containing the user's natural facial expression data stream is collected in real time, and passive candidate data is generated by combining the user's current scene context information, which includes the usage scene type and operation behavior records. Among them, "natural facial expression data stream" refers to the continuous data sequence formed by the changes in facial expressions over time when users are in a natural state without deliberately displaying emotions, such as the changes in facial expressions such as smiling and frowning when using the device in daily life; "passive emotion data" refers to emotion-related data generated by users in a natural state without being deliberately guided by the system, as opposed to active emotion data; "scene context information" refers to the environment and operating background information in which users are using the device.

[0033] Once the user enables the emotion recognition function, the system begins to collect passive emotion data in real time. Unlike the active collection during the initial training, the user is in a natural state and is not deliberately guided by the system. First, the system begins to capture the user's facial images in real time using the image acquisition device (such as a camera). Then, it extracts facial feature point information from each frame of the image using facial feature extraction algorithms (such as a deep learning-based facial keypoint detection model), including the shape and position parameters of areas such as eyebrows, eyes, nose, mouth, and cheeks. These parameters change over time to form a natural facial expression data stream.

[0034] Next, the system synchronously collects the user's current context information. The scenario type can be determined in various ways, such as based on the type of application the user is using (e.g., office software for work scenarios, games for entertainment scenarios), the device's geographical location (office, home), and time information (weekday daytime, weekend evening). User behavior records are obtained by recording the user's device actions, including clicked interface elements, entered text, viewed pages, and dwell time. For example, repeatedly swiping social media pages or repeatedly viewing product details in a shopping app will be recorded. Finally, the system fuses the collected passive emotion data with the current context information to generate passive candidate data. This fusion process is not a simple information overlay but establishes a correlation between the two. For example, it associates "a user's laughing expression while watching a comedy video" with "entertainment scenarios and repeatedly clicking to play the next video," allowing subsequent emotion analysis to combine specific scenarios and user behavior, improving the accuracy and relevance of emotion recognition. For example, the same frown might indicate confusion in a learning context or frustration in a gaming context. Combining contextual information allows for a more accurate judgment of the emotional meaning.

[0035] S103. Perform feature filtering on the passive candidate data to obtain the data to be labeled; Among them, "feature filtering" refers to extracting valuable feature information from the data that can effectively reflect the user's emotional state and removing redundant, irrelevant or interfering data.

[0036] First, feature extraction is performed on the passive candidate data. Key facial features are extracted from the natural facial expression data stream, such as eyebrow shape, eye opening and closing, mouth corner curvature, and facial muscle movement amplitude. These features directly reflect the user's emotional state. Simultaneously, emotion-related features are extracted from the contextual information, such as the emotional tendency of the usage scenario (office scenarios may show more focus or fatigue, while entertainment scenarios may show more pleasure or excitement) and the emotional association of operational behaviors (rapid swiping may indicate impatience, repeated clicking may indicate confusion). Then, preliminary feature screening is performed, eliminating features clearly unrelated to emotion, such as background information in facial images and operation records with extremely low emotional relevance in the context. Next, feature importance evaluation methods are used to rank the extracted features by importance, retaining those with higher importance. Common evaluation methods include feature importance scoring based on machine learning models, such as using a random forest model to calculate the contribution of each feature to the emotion recognition results, retaining features with high contributions. Furthermore, feature correlation needs to be considered; if two features are highly correlated, it may cause information redundancy, in which case the more representative feature can be retained. Furthermore, the current state of the emotion recognition model is used for filtering. Feature data corresponding to emotions that the model can already accurately identify may have their filtering priority reduced, while feature data related to emotions with lower model recognition accuracy will be prioritized. Finally, the filtered feature data is integrated to form the data to be labeled. This data to be labeled contains key features that can effectively reflect user emotions while eliminating redundant and irrelevant information, laying a solid foundation for subsequent labeling and model updates.

[0037] S104. When the data to be labeled meets the triggering condition, semantic guidance content is embedded in the natural interaction scenario between the user and the terminal. The semantic guidance content is associated with the facial expression features corresponding to the data to be labeled. Among them, "natural interaction scenarios between users and the device" refers to situations in which users interact with the device naturally without deliberate intervention during their daily use of the device, such as browsing web pages, using applications, and interacting with the device via voice.

[0038] First, it's necessary to determine if the data to be labeled meets the triggering conditions. Data that the model struggles to accurately identify, is ambiguous, or possesses novelty needs to be filtered out. This allows for obtaining more accurate emotion labels through user feedback, thereby optimizing the emotion recognition model. Specifically, preliminary emotion recognition is performed on the acquired data. This process involves inputting the data into the current emotion recognition model. The model analyzes facial expression features and relevant scene information in the data, outputting preliminary emotion judgment results, namely a preliminary emotion probability distribution and corresponding confidence scores. The preliminary emotion probability distribution shows the likelihood of the data belonging to each preset emotion type. For example, in a model containing four emotion types—happiness, sadness, anger, and calm—the probability distribution of a certain data to be labeled might be: happiness 40%, calm 35%, sadness 15%, and anger 10%. The confidence score is the model's reliability assessment of this distribution result based on its internal algorithm, typically represented by a value between 0 and 1, where 0 indicates completely unreliable and 1 indicates completely reliable.

[0039] Next, the system determines whether the data to be labeled meets the trigger conditions based on three preset conditions. The first condition is whether the confidence score is lower than the preset confidence threshold. A low confidence score indicates insufficient reliability of the model's emotion recognition results for the current data to be labeled, potentially indicating a large error, requiring further processing. For example, if the confidence threshold is set to 0.7, and the confidence score of a certain data to be labeled is 0.65, which is lower than the threshold, then the trigger condition is met. The second condition is whether the difference between the highest and second-highest probability values ​​in the preliminary emotion probability distribution is lower than the preset discrimination threshold. A small difference indicates that the model has difficulty clearly distinguishing between the two emotion types, resulting in ambiguous recognition results. For example, if the discrimination threshold is set to 0.15, and the difference between the highest probability (40% happy) and the second-highest probability (35% calm) of a certain data to be labeled is 0.05, which is lower than the threshold, then the trigger condition is met. The third condition is to calculate the distance between the facial expression features corresponding to the data to be labeled and existing features in the user's historical facial expression feature library, and determine whether this distance exceeds the preset novelty threshold. The distance here is typically calculated using methods such as Euclidean distance and cosine distance of feature vectors. A larger distance indicates a greater difference between the expression feature and historical features, potentially representing a new emotional expression or a new emotional state of the user. For example, if the novelty threshold is set to 0.8 and the calculated distance is 0.9, exceeding the threshold, the trigger condition is met. As long as the data to be labeled meets at least one of the above three conditions, it is determined to have met the trigger condition, requiring subsequent semantic guidance and feedback collection.

[0040] If the triggering conditions are met, semantic guidance content needs to be embedded in the natural interaction between the user and the application without disrupting the user's normal workflow. The specific process will be described in detail in subsequent steps S201 to S204, and will not be repeated here.

[0041] S105. Obtain user feedback information on the semantic guidance content, and determine the unconscious annotation corresponding to the data to be annotated by combining the facial expression features during the feedback. Unconscious labeling refers to emotion labels that are indirectly inferred from the analysis of users' comprehensive feedback signals in non-intrusive natural interaction scenarios, and that can reflect the user's true emotional state.

[0042] After embedding semantic guidance content into users, comprehensive feedback signals can be collected and analyzed during and after the user's feedback process. Combined with facial expression features during feedback, accurate and reliable unconscious annotations can be generated, providing high-quality annotation data for the dynamic updating of emotion recognition models.

[0043] First, comprehensive feedback signals from the user during the interaction are collected. This requires the user to utilize multiple sensors and input devices, such as recording the user's voice responses in real time via microphone, recording the user's text input via input devices, and capturing the user's body movements (such as head movements and hand gestures) and facial expressions (such as raising the corners of the mouth and furrowing the brow) via camera. The collection of these signals needs to be synchronized in time to ensure that different types of signals can be mapped to the same feedback moment during subsequent analysis, thereby accurately reflecting the user's overall state during the feedback process.

[0044] Next, the feedback is categorized based on the comprehensive feedback signal. For direct positive or negative feedback, the system extracts the explicit emotional tendency using natural language processing (for speech and text) or action recognition (for body movements). For example, when a user replies "I was indeed very angry just now," the system recognizes the explicit emotional direction of "anger" and directly associates the preset "anger" emotional label with the feedback, classifying it as an unconscious label. For ambiguous feedback, such as a user replying "more or less," since the corresponding emotional type cannot be clearly determined, the system will mark it as a low-confidence label. Such labels may be assigned lower weights in subsequent model updates. In the case of no feedback, to avoid introducing invalid information, the system will choose to ignore the current labeling operation and try to obtain user feedback again at a suitable time later.

[0045] The system then incorporates facial expression features from user feedback as a calibration factor to adjust the weight of unconscious annotations. It analyzes the collected facial expression features during feedback, extracting emotional features (such as smiling features when happy, frowning features when angry), and calculates the matching degree with the emotional tendency reflected in the overall feedback signal. When the matching degree is high, for example, if the user inputs "I am happy" and a smiling facial expression is detected simultaneously, the feedback signal is considered highly credible, and the system will increase the weight of the unconscious annotation accordingly, allowing it to play a greater role in model training. When the matching degree is low, for example, if the user's voice reply is "I am calm" but their facial expression shows sadness, the feedback signal may be contradictory or the user's expression may be insincere, and the system will decrease the weight of the unconscious annotation accordingly. If the matching degree is extremely low, such as in a completely contradictory situation, a secondary confirmation mechanism will be triggered, embedding semantic guidance content again (such as "You look a little sad, is that right?") to confirm with the user and obtain more accurate feedback information.

[0046] Optionally, firstly, multimodal fusion technology is used to fuse the collected speech, text, body movements, and facial expression change signals, transforming the signals from different modalities into a unified feature vector. Then, a trained classification model is used to classify the fused features and determine the feedback type (direct affirmation / negation, ambiguous, no feedback). For direct feedback, the corresponding preset emotion label is matched as an initial unconscious label, and then the similarity (i.e., matching degree) between the facial expression features at the time of feedback and the standard expression features corresponding to the emotion label is calculated, and the labeling weight is adjusted according to the similarity. For ambiguous feedback, it is directly marked as a low-confidence label. For no feedback, the timestamp is recorded and the current labeling process ends. Optionally, a feedback signal-emotion tag mapping library is first established to store the association between common direct feedback signals and their corresponding emotion tags. When a comprehensive feedback signal is received, a search is first performed in the mapping library. If a matching direct feedback signal is found, the corresponding emotion tag is extracted. At the same time, facial expression features during feedback are analyzed in real time to generate an expression emotion probability distribution. The probability value corresponding to the retrieved emotion tag in this probability distribution is calculated as the degree of matching. If the probability value is higher than a preset threshold, the weight is increased; if it is lower than the threshold, the weight is decreased or a second confirmation is triggered. For signals that are not matched in the mapping library, they are judged as fuzzy feedback and marked with a low confidence tag. If there is no signal input, it is judged as no feedback and ignored.

[0047] S106. Dynamically update the emotion recognition model based on the unconscious annotation; After identifying the unconscious annotations corresponding to the data to be labeled, the model performance can be continuously optimized by integrating new labeled data to ensure the accuracy and personalization of emotion recognition results. First, the unconsciously labeled data needs to be preprocessed. The preprocessed data with unconscious annotations is then used as new training samples and combined with the model's original training data to form an updated training set. During this process, the weighting of the new data needs to be considered. Generally, recently acquired unconsciously labeled data is more relevant to the user's current emotional expression habits and may therefore be given higher weights, while the weight of earlier historical data can be appropriately reduced to improve the model's adaptability to the user's recent emotional characteristics.

[0048] Then, incremental learning is used to update the emotion recognition model. Incremental learning can adjust the model using only new data without retraining the entire model, saving computational resources and preventing the model from forgetting previously learned knowledge. Specifically, new training samples are input into the model, and the model parameters are adjusted through the backpropagation algorithm to minimize the prediction error on the new data. At the same time, to prevent the model from overfitting to the new data, a regularization mechanism is needed, such as adding a regularization term to the loss function to constrain the range of variation of the model parameters.

[0049] In addition, the updated model needs to be evaluated. The model's accuracy and recall are tested using a validation set (facial expression data containing known emotion labels) to determine if the update is effective. If the evaluation results meet preset criteria (such as improved or stable accuracy), the updated model is accepted; if the results are unsatisfactory, the update strategy may need to be adjusted, such as increasing the number of training iterations or adjusting data weights, and the model may need to be updated again.

[0050] S107. Dynamically analyze the acquired real-time facial expression data stream using the updated emotion recognition model to obtain emotion recognition results.

[0051] First, the model receives a real-time facial expression data stream. This data stream is typically input in the form of video frames. The model needs to preprocess each frame. Then, the updated model analyzes the preprocessed feature vectors and outputs the emotion probability distribution and confidence score for each time point. The emotion probability distribution represents the likelihood that the user's current expression belongs to various preset emotion types (such as happy, sad, angry, etc.), while the confidence score reflects the model's trustworthiness of that judgment. For example, the model might output "Happy: 60%, Calm: 30%, Confidence: 0.85". Next, the model performs dynamic analysis incorporating the time dimension. By observing changes in the emotion probability distribution and confidence score over a period of time (such as the past 30 seconds), the dynamic trends of emotions are identified.

[0052] Furthermore, the model also incorporates contextual information about the user's current situation (such as the application being used and the user's actions) to refine the recognition results. For example, a laughing emoji while playing a game is more likely to correspond to "excitement," while a frowning emoji while reading the news may correspond to "confusion" or "dissatisfaction." By fusing contextual information, the accuracy of emotion recognition is further improved. Finally, the model outputs the final emotion recognition result. The presentation format of the result can be adjusted according to the needs of the application scenario, and different responses can be controlled on different user devices based on different emotions.

[0053] In some embodiments, during the analysis of real-time facial expression data streams using an updated emotion recognition model, the device's operating status can be adaptively adjusted by dynamically analyzing the temporal changes in emotion features, thereby optimizing resource consumption while ensuring recognition accuracy, and ultimately generating reliable emotion recognition results.

[0054] First, a temporal feature sequence is generated by combining real-time facial expression data streams. The system continuously receives facial expression video streams captured by a camera, extracting one frame at regular intervals (e.g., 50 milliseconds). Each frame is analyzed using an updated emotion recognition model, outputting the multidimensional emotion probability distribution (e.g., happy 30%, calm 50%, sad 20%) and confidence score (e.g., 0.8) corresponding to that time point. This data is arranged chronologically to form a temporal feature sequence containing timestamps, multidimensional emotion probabilities, and confidence scores, providing foundational data for subsequent trend analysis. Then, three indicators are calculated based on the temporal feature sequence within a preset time window. For the trend of emotion confidence changes, the arithmetic mean (reflecting the overall credibility level) and variance (reflecting score fluctuations) of all confidence scores within the window are calculated. A higher mean and smaller variance indicate more stable confidence changes. For the dominant emotion stability index, first determine the dominant emotion (the emotion type with the highest probability) at each time point within the window, count the duration of each dominant emotion, and calculate the number of times the dominant emotion switches per unit time (switching frequency). Integrate these two factors into a stability index using a formula (e.g., weighted sum of durations divided by switching frequency plus 1). A higher index indicates a more stable dominant emotion. For the emotion change gradient, treat the multidimensional emotion probability distributions at adjacent time points as high-dimensional vectors, and calculate the Euclidean distance or cosine distance between them as the vector difference. A larger difference indicates a more drastic emotion change.

[0055] Then, based on the comparison results of the three indicators with preset thresholds, the device's operating state is switched. When the trend of emotion confidence is stable (variance below a certain value) and the mean is higher than the first threshold, while the dominant emotion stability index is higher than the second threshold, it indicates that the user's emotional state is stable and the model recognition is reliable. At this time, it enters a low-power monitoring state, reducing the camera's acquisition frame rate (e.g., from 30 frames / second to 5 frames / second) and reducing the allocation of local processor computing resources to save power and computing power. When all three indicators are within their respective standard fluctuation ranges (e.g., the mean confidence is between the first and third thresholds, the stability index is between the second and fourth thresholds, and the change gradient is lower than the fifth threshold), it indicates that the user's emotional state is normal and the model recognition effect is stable. It maintains a standard local analysis state, keeping the preset standard acquisition frequency and resource allocation. When the mean of the change trend of emotion confidence is continuously lower than the third threshold (indicating low recognition confidence), or the dominant emotion stability index is lower than the fourth threshold (indicating frequent emotion switching), or the emotion change gradient is higher than the fifth threshold (indicating drastic emotion change), the cloud collaborative analysis state is triggered. At this time, the facial expression data corresponding to the triggering event (such as the moment of sudden emotion change) and the historical data within the preset time window (such as 1 minute) before the event occur are packaged and uploaded to the cloud system via the network.

[0056] Finally, in the cloud-based collaborative analysis mode, the cloud leverages its more powerful computing resources and more comprehensive training data to perform in-depth analysis on the uploaded data, outputting cloud-based analysis results. The system then fuses and calibrates the cloud results with the local analysis results, using methods such as weighted averaging (assigning weights based on the confidence levels of both) or voting mechanisms (selecting the majority result) to eliminate potential discrepancies and ultimately generate emotion recognition results that are both accurate and reliable. This allows for the allocation of more resources to improve the accuracy of emotion recognition when user emotions change rapidly, while reducing resource consumption through a low-power monitoring mode when emotions are stable, achieving dynamic allocation and efficient utilization of computing resources. Through the powerful computing capabilities and rich training data of the cloud, complex feature relationships in emotion data can be deeply explored, compensating for the limitations of local devices in terms of computing power and data when processing highly dynamic emotion changes, effectively avoiding recognition delays or misjudgments caused by rapid emotion fluctuations. The fusion and calibration mechanism of cloud and local results retains the responsiveness advantage of local real-time analysis while integrating the accuracy of cloud-based in-depth analysis, ensuring that the system outputs high-quality recognition results regardless of whether the user's emotions are stable or undergoing drastic changes.

[0057] In the above embodiment, a complete closed loop of "initialization-data collection-screening and labeling-model update-accurate recognition" is formed. It can quickly build an initial model that fits the user through active data, capture natural emotions based on passive data and scene information, and continuously optimize the model with unconscious labeling. It effectively solves the problems in traditional emotion recognition, such as the inability of general models to adapt to individual differences, the lag of emotion recognition in the user's actual state, and the disconnect of labeled data from real scenes. Thus, it realizes personalized iteration and dynamic optimization of emotion recognition model, improves the accuracy, adaptability and real-time performance of emotion recognition, and can more accurately and quickly capture and identify the user's real emotions, meeting the needs of efficient emotional interaction in diverse scenarios.

[0058] Based on the above, the following section provides a more detailed process for assigning emotion tags to users' facial expressions without their conscious awareness or interference. Please refer to [link / reference]. Figure 2 This is another flowchart illustrating the emotion recognition method based on a large model in this application embodiment.

[0059] S201. Extract the facial features corresponding to the data to be labeled and establish an unlabeled facial feature library; Among them, facial expression features refer to key information extracted from facial expression data that reflects emotional state, such as the amplitude of facial muscle movements, eye status (open / closed eyes, pupil changes), the curvature of the corners of the mouth, and the position of the eyebrows. The unlabeled facial expression feature library is a collection used to store the features corresponding to these unlabeled facial expression data, which are then used for subsequent matching with facial expression features acquired in real time. For example, the data to be labeled may be a video clip of a user's mouth slightly raised, but it is difficult to determine whether it is a smile or another emotion. The angle and duration of the raised corners of the mouth extracted from this clip are facial expression features, and these features will be stored in the unlabeled facial expression feature library.

[0060] After filtering passive candidate data to obtain the data to be labeled, preparations are made for accurately identifying similar expressions and embedding semantic guidance content during user interaction. First, expression features need to be extracted from the data to be labeled. This process typically utilizes computer vision technology to process images or video frames containing facial expressions. For static images, face detection algorithms are used to locate key facial regions, such as eyes, nose, mouth, and eyebrows, and then the geometric features (such as distances and angles between organs) and texture features (such as changes in skin texture) of these regions are calculated. For video sequences, dynamic features are also extracted, such as the rate and magnitude of change in facial expressions over time. For example, when processing a video of a user frowning, features such as the descent of the eyebrows, the change in distance between the eyebrows, and the duration of this change are extracted. After extracting these expression features, an unlabeled expression feature library needs to be established. This library organizes and stores features according to certain rules, and may perform preliminary clustering based on the similarity of expression features, grouping similar features together to facilitate subsequent matching operations. Meanwhile, to improve the efficiency of subsequent matching, features will be standardized to eliminate the influence of different data collection conditions (such as lighting and angle), ensuring the consistency and comparability of features. The establishment of an unlabeled facial expression feature library enables the system to effectively manage and retrieve these facial expression features to be confirmed, laying the foundation for real-time monitoring and matching of facial expressions during user-end interactions.

[0061] S202. During the interaction between the user and the terminal, monitor the user's current facial expression in real time, extract the current feature information of the current facial expression, and match it with the feature information in the unlabeled expression feature library. First, the system needs to monitor the user's current facial expressions in real time. This is typically achieved using image acquisition devices such as cameras on the user's device. These devices continuously capture facial images or video streams and transmit them to the processing module. The processing module processes the captured images or video streams in real time, starting with face detection to ensure accurate location of the user's facial area and avoid interference from other objects in the background. After detecting a face, it analyzes the facial expressions in each frame of the image or video, capturing dynamic changes in the expressions. Next, it extracts the current feature information of the current facial expression. The extraction method is consistent with the method used to extract the facial expression features of the data to be labeled in S201 to ensure the comparability of features, including geometric features, texture features, and dynamic features. For example, when a user suddenly frowns while watching a video, the system extracts features such as the change in eyebrow position and the intensity of the frown. Then, the extracted current feature information is matched with the feature information in the unlabeled expression feature library. During the matching process, appropriate similarity calculation algorithms, such as Euclidean distance and cosine similarity, are used. For each current feature, the system iterates through all features in the unlabeled emoji feature library and calculates the similarity score between them. If a current feature has a high similarity to a feature in the unlabeled emoji feature library, it means that the user's current emoji is similar to a previously unlabeled emoji, and the system can then take further action.

[0062] S203. If the matching degree between the current feature information and the target feature information in the unlabeled expression feature library exceeds the preset matching degree threshold, then lock the continuous time period in which the current facial expression appears, and record the current interaction scene information context within the continuous time period. After matching the current feature information with the feature information in the unlabeled expression feature library, and the matching degree between the two exceeds the preset matching degree threshold, when the system determines that the matching degree between the current feature information and the target feature information exceeds the preset matching degree threshold, it indicates that the user's current expression is highly similar to the previously unlabeled expression and requires further attention. At this time, the system will lock the continuous time period of the current facial expression. Because emotions have a certain continuity, the locking process needs to accurately determine the start and end times of the expression, which is usually achieved by monitoring changes in expression features: when the expression feature first reaches the matching degree threshold, it is recorded as the start time; then, continuous monitoring continues, and when the expression feature is below the matching degree threshold multiple times consecutively (e.g., more than 3 frames), it is recorded as the end time. The time period between these two time points is the continuous time period. For example, when a user is watching a program using video software, if the matching degree of a certain expression feature exceeds the threshold from frame 100, and falls below the threshold from frame 150, and remains below the threshold for 5 consecutive frames, then the continuous time period is the time range corresponding to frame 100 to frame 150. While locking the continuous time period, the system will also record the current interaction scene information context within that time period. This requires the system to interact with various applications and system modules on the user's end to collect relevant information, such as the name of the application the user is using, the operation path within the application (e.g., which buttons were clicked, what content was entered), the timestamp of the interaction, the device's network status (e.g., Wi-Fi or mobile data connection), and the device's geographical location (if authorized by the user). For example, if a user is chatting with someone on a social application within a continuous period of time and sends a message saying "This movie is amazing," while the device is connected to home Wi-Fi, this information will be recorded as contextual information for the interaction scenario. The purpose of recording this information is to ensure that the subsequently generated interactive statements are closely integrated with the user's current operation and environment, making the semantic guidance content more natural and appropriate, avoiding any abruptness for the user, and thus improving the effectiveness of user feedback.

[0063] S204. Generate interactive statements related to the scene task based on the current interactive scene information context, and embed the semantic guidance content into the interactive statements and present it to the user.

[0064] The purpose of this step is to naturally obtain user feedback on relevant facial expressions without interfering with normal user interaction, in order to determine the unconscious annotation of the data to be labeled. First, the system needs to determine the scenario task based on the context of the current interaction. This requires analyzing and understanding the collected scenario information, such as analyzing the type of application the user uses (e.g., shopping, social, entertainment) and their actions (e.g., browsing, searching, purchasing, sending messages) to determine the user's ongoing task. For example, if a user searches for "birthday gift" in a shopping app and views the details of multiple gifts, the scenario task can be determined as "choosing a birthday gift." After clarifying the scenario task, the system generates interactive statements related to that task. The generation of interactive statements typically uses natural language generation technology, combining the characteristics of the scenario task and the user's operating habits to make the statements consistent with the current context and avoid awkwardness. For example, for the scenario task of "choosing a birthday gift," the interactive statement could be, "It looks like you're busy choosing a birthday gift. Do you have any special ideas?" Next, semantic guidance content needs to be embedded into the interactive statements. Semantic guidance content needs to be associated with the facial features corresponding to the data to be labeled. For example, if the facial features of the data to be labeled might correspond to confusion, then the semantic guidance content could be, "You seem hesitant about this choice. Do you have any questions?" The embedding process must ensure that the semantic guidance content blends naturally with the interactive statement, without disrupting the coherence and readability of the statement. Finally, the interactive statement containing the semantic guidance content is presented to the user. The presentation method needs to be selected based on the type of device and the current scenario. For example, in mobile applications, pop-ups or message prompts can be used; in voice interaction devices such as smart speakers, voice broadcasts can be used. At the same time, attention should be paid to the timing of presentation. Avoid disturbing the user when performing critical operations (such as entering a password or confirming payment). Presentation should be done during gaps in user operations or at relatively quiet moments to improve user attention and willingness to respond.

[0065] In this embodiment, facial expression features are first extracted from the data to be labeled and an unlabeled facial expression feature library is established, providing a precise comparison basis for subsequent matching. Then, facial expressions are monitored in real time during user interaction and matched with features in the library. When the matching degree meets the standard, a continuous time period is locked and the scene context is recorded. Finally, interactive sentences containing semantic guidance content are generated in combination with the scene, realizing the accurate association and natural embedding of semantic guidance content with the user's current scene and facial expression features. Therefore, user feedback related to the data to be labeled can be obtained efficiently without interfering with the user's normal interactive experience. This effectively solves the problems of high manual labeling costs and the disconnect between labeled data and actual scenes in traditional emotion recognition. In turn, it realizes the dynamic supplementation of training data for emotion recognition models and the continuous optimization of models, improving the accuracy and personalization of emotion recognition.

[0066] The emotion recognition system in the embodiments of this invention is described below from the perspective of hardware processing. Please refer to [link / reference]. Figure 3 This is a schematic diagram of the physical device structure of an emotion recognition system in this application embodiment.

[0067] It should be noted that, Figure 3 The structure of the emotion recognition system shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.

[0068] like Figure 3 As shown, the emotion recognition system includes a Central Processing Unit (CPU) 301, which can perform various appropriate actions and processes based on a program stored in Read-Only Memory (ROM) 302 or a program loaded from storage portion 308 into Random Access Memory (RAM) 303, such as performing the methods described in the above embodiments. The RAM 303 also stores various programs and data required for the operation of the emotion recognition system. The CPU 301, ROM 302, and RAM 303 are interconnected via a bus 304. An Input / Output (I / O) interface 305 is also connected to the bus 304.

[0069] The following components are connected to I / O interface 305: input section 306 including audio input devices, push-button switches, etc.; output section 307 including a liquid crystal display (LCD) and audio output devices, indicator lights, etc.; storage section 308 including a hard disk, etc.; and communication section 309 including a network interface card such as a LAN (Local Area Network) card, modem, etc. Communication section 309 performs communication processing via a network such as the Internet. Drive 310 is also connected to I / O interface 305 as needed. Removable media 311, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 310 as needed so that computer programs read from them can be installed into storage section 308 as needed.

[0070] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing computer programs for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 309, and / or installed from removable medium 311. When the computer program is executed by central processing unit (CPU) 301, it performs the various functions defined in the present invention.

[0071] It should be noted that specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an emotion recognition system, apparatus, or device.

[0072] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of emotion recognition systems, methods, and computer program products according to various embodiments of the present invention. Each block in the flowchart or block diagram may represent a module, program segment, or portion of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that shown in the drawings.

[0073] Specifically, the emotion recognition system in this embodiment includes a processor and a memory. The memory stores a computer program, and when the computer program is executed by the processor, it implements the emotion recognition method based on a large model provided in the above embodiment.

[0074] In another aspect, the present invention also provides a computer-readable storage medium, which may be included in the emotion recognition system described in the above embodiments; or it may exist independently and not incorporated into the emotion recognition system. The storage medium carries one or more computer programs that, when executed by a processor of the emotion recognition system, cause the emotion recognition system to implement the large-model-based emotion recognition method provided in the above embodiments.

[0075] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0076] As used in the above embodiments, depending on the context, the term "when..." can be interpreted as meaning "if...", "after...", "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if (the stated condition or event) is interpreted as meaning "if determining...", "in response to determining...", "when (the stated condition or event) is detected", or "in response to detecting (the stated condition or event)".

[0077] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A large-model-based emotion recognition method, characterized in that, The method includes: After receiving the initial activation command from the user terminal, the system collects the user's active emotion data for the initial training of the emotion recognition model. After the user enables the emotion recognition function, passive emotion data containing the user's natural facial expression data stream is collected in real time. Passive candidate data is generated by combining the user's current scene context information, which includes the usage scene type and operation behavior records. The passive candidate data is subjected to feature filtering to obtain the data to be labeled; When the data to be labeled meets the triggering condition, semantic guidance content is embedded in the natural interaction scenario between the user and the terminal, and the semantic guidance content is associated with the facial expression features corresponding to the data to be labeled. Obtain user feedback on the semantic guidance content, and determine the unconscious annotation corresponding to the data to be annotated by combining the facial expression features during the feedback; The emotion recognition model is dynamically updated by incorporating the unconscious annotations. The updated emotion recognition model is used to dynamically analyze the acquired real-time facial expression data stream to obtain emotion recognition results.

2. The method according to claim 1, characterized in that, After receiving the initial activation command from the user, the steps for collecting the user's proactive emotion data for the initial training of the emotion recognition model include: Send an emotion guidance instruction to the user terminal to guide the user to make corresponding facial expressions according to preset emotion types, and collect them as the first expression sample; By presenting users with immersive, motivating content related to the preset emotion type, and collecting second expression samples; The first expression sample and the second expression sample are used to construct an expression calibration pair, and a personalized expression preference adapter is trained to obtain the expression calibration pair. The personalized expression preference adapter is integrated with a standard emotion recognition model pre-deployed in the cloud to determine the final emotion recognition model. The standard emotion recognition model is obtained by machine learning training in the cloud based on different facial expressions and corresponding emotion labels of multiple different users.

3. The method according to claim 1, characterized in that, The step of dynamically analyzing the acquired real-time facial expression data stream using the updated emotion recognition model to obtain emotion recognition results specifically includes: The real-time facial expression data stream is combined to generate a time-series feature sequence containing a multidimensional emotion probability distribution and confidence score; Based on the temporal feature sequence within a preset time window, three indicators are determined, including the trend of change in emotional confidence, the stability index of dominant emotion, and the gradient of emotional change. The trend of change in emotional confidence is obtained by calculating the mean and variance of the confidence score within the preset time window. The stability index of dominant emotion is obtained by analyzing the duration and switching frequency of the dominant emotion type. The gradient of emotional change is obtained by calculating the difference between the vectors of the probability distribution of emotions at adjacent time points. When the trend of the emotion confidence change is stable and higher than the first threshold, and the dominant emotion stability index is higher than the second threshold, then enter the low-power monitoring state, reduce the frequency of facial expression data collection and reduce the allocation of local computing resources. When all three indicators are within the standard fluctuation range, the standard local analysis state of the user end is maintained, and the preset standard facial expression data collection frequency and computing resource allocation are maintained. When the trend of the change in the emotional confidence level is continuously lower than the third threshold, or the dominant emotional stability index is lower than the fourth threshold, or the emotional change gradient is higher than the fifth threshold, a cloud-based collaborative analysis state is triggered, wherein the third threshold is lower than the first threshold and the fourth threshold is lower than the third threshold. In the cloud-based collaborative analysis mode, facial expression data that triggers the corresponding emotional event and historical data within the preceding time window are packaged and uploaded to the cloud for analysis; The cloud-based analysis results are fused and calibrated with the local analysis results to obtain the final emotion recognition result.

4. The method according to claim 1, characterized in that, When the data to be labeled meets the triggering condition, the step of embedding semantic guidance content in the natural interaction scenario between the user and the terminal specifically includes: Extract the facial expression features corresponding to the data to be labeled and establish an unlabeled facial expression feature library; During the interaction between the user and the terminal, the user's current facial expression is monitored in real time, the current feature information of the current facial expression is extracted, and it is matched with the feature information in the unlabeled expression feature library; If the matching degree between the current feature information and the target feature information in the unlabeled expression feature library exceeds a preset matching degree threshold, then the continuous time period in which the current facial expression appears is locked, and the current interaction scene information context within the continuous time period is recorded. Based on the context of the current interactive scenario information, generate interactive statements related to the scenario task, and embed the semantic guidance content into the interactive statements to present them to the user.

5. The method according to claim 1, characterized in that, The steps for when the data to be labeled reaches the trigger condition specifically include: Preliminary emotion recognition is performed on the data to be labeled to obtain a preliminary emotion probability distribution and confidence score. The data to be labeled is determined to have reached the trigger condition when at least one of the following preset conditions is met: the confidence score is lower than the preset confidence threshold. In the preliminary emotion probability distribution, the difference between the highest probability value and the second highest probability value is lower than the preset discrimination threshold. The distance between the facial expression features corresponding to the data to be labeled and the existing features in the user's historical facial expression feature library exceeds a preset novelty threshold.

6. The method according to claim 1, characterized in that, The steps of obtaining user feedback on the semantically guided content and determining the unconscious annotations corresponding to the data to be annotated based on facial expression features during feedback specifically include: Collect comprehensive feedback signals from users during feedback interactions, including voice responses, text input, body movements, or facial expressions. Based on the comprehensive feedback signal, the feedback is classified: if it is determined to be a direct affirmative or negative feedback, it is directly associated with the corresponding preset emotion tag and determined to be an unconscious label; If the label is determined to be ambiguous or without feedback, mark it as a low-confidence label or ignore the current labeling operation; The facial expression features during user feedback are used as calibration factors. When the facial expression features during user feedback match the comprehensive feedback signal, the weight of the unconscious annotation is increased according to the degree of matching. If the facial expression features of the user do not match the overall feedback signal, the weight of the unconscious annotation will be reduced or a second confirmation will be required, depending on the degree of matching.

7. The method according to claim 1, characterized in that, After the step of dynamically updating the emotion recognition model by incorporating the unconscious annotations, the method further includes: Periodically calculate the feature deviation between active emotion data and the corresponding unconsciously labeled data; If the feature deviation exceeds a preset deviation threshold, lightweight active training is triggered to obtain supplementary facial expression data. The active training includes sending targeted emotion guidance instructions to the user terminal, guiding the user to supplement and display facial expressions corresponding to the emotion types with feature deviations. The emotion recognition model was improved by combining supplementary facial expression data and unconsciously labeled data of corresponding emotion types.

8. An emotion recognition system, characterized in that, The emotion recognition system includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code including computer instructions, and the one or more processors call the computer instructions to cause the emotion recognition system to perform the method as described in any one of claims 1-7.