Facial composite emotion recognition method and system based on scene induction and medium

By acquiring facial videos and utilizing a composite emotion recognition model and multi-dimensional elicited scenarios, the problem of insufficient accuracy in composite emotion recognition in existing technologies has been solved. This enables accurate identification of composite emotion categories and intensities, enhancing the application value of emotion recognition technology in real-world scenarios.

CN121096005APending Publication Date: 2025-12-09GUANGZHOU EMOTION CALCULATION TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511332630.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-18
Publication Date
2025-12-09

AI Technical Summary

Technical Problem

Existing emotion recognition technologies struggle to accurately identify complex emotions, resulting in insufficient recognition accuracy in scenarios such as mental health assessment, human-computer interaction, and intelligent video surveillance, thus failing to provide precise feedback and analysis.

Method used

By acquiring facial videos of target users and using a pre-trained composite emotion recognition model to extract facial micro-expression sequences, and combining multi-dimensional immersive evoked scenarios and dynamic evoked control models, a composite emotion sample database is constructed to achieve accurate identification of composite emotion categories and intensities.

Benefits of technology

It significantly improves the accuracy and practicality of complex emotion recognition, enabling it to provide more relevant feedback and analysis in various application scenarios and meet practical needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121096005A_ABST
    Figure CN121096005A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of emotion recognition, and discloses a facial composite emotion recognition method and system based on scene induction and a medium. According to the method provided by the invention, the face video of the target user is acquired, the face micro-expression sequence of the target user is extracted, the sequence is input into the pre-trained composite emotion recognition model, and the composite emotion evaluation information of the target user is acquired. Compared with common expressions, the micro-expression sequence can better reflect the real emotional state of the user, the combination condition of basic emotions in the composite emotions can be accurately captured in combination with the specially-trained composite emotion recognition model, and the refinement degree and precision of composite emotion recognition are remarkably improved. Meanwhile, the acquisition of the composite emotion intensity can realize more detailed description instead of staying at a category level, so that the method can provide a reliable emotion recognition basis for various actual application scenes needing to accurately grasp the emotion of the user, and the requirement of the actual scenes for accurate recognition of the composite emotion is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of emotion recognition technology, specifically to a method, system, and medium for scene-induced facial complex emotion recognition. Background Technology

[0002] With the rapid development of artificial intelligence technology, emotion recognition, as a core technology in fields such as human-computer interaction, mental health assessment, intelligent education, and video surveillance, is seeing its application scenarios expand and its demand become increasingly urgent. Currently, most mainstream emotion recognition technologies focus on recognizing basic emotions, such as anger, happiness, sadness, surprise, fear, disgust, and neutral emotions. Related research is relatively mature and can achieve high recognition accuracy in specific scenarios.

[0003] However, in real life, human emotional expression is often far more rich and complex. Individuals frequently experience complex emotional states, which are mixtures of two or more basic emotions, such as feeling both "nervous and excited" or "loving and hating." Psychological research further indicates that these complex emotions are not static combinations, but rather, under the cognitive guidance and dynamic modulation of facial muscle movement patterns in specific scenarios, they generate rich expressive forms highly characteristic of those scenarios. For example, the research of Cowen's team clearly confirms that emotions such as "amazement" and "focus" are scenario-based products that integrate multiple basic emotional characteristics. It can be said that this complex emotion, deeply bound to a specific scenario, has become the mainstream form of genuine human emotional expression.

[0004] Existing technologies have significant limitations when dealing with complex emotions. Most of them are based on traditional basic emotion recognition methods, which can only identify a single emotion category. This is fundamentally contradictory to the "multi-emotion combination" of complex emotions. Therefore, they often can only forcibly classify complex emotions into a certain basic emotion, such as misjudging "confusion + frustration" as simply "sadness", or they may not even be able to give an effective recognition result.

[0005] This deficiency in recognition capabilities severely limits the in-depth application of emotion recognition technology in real-world scenarios. For example, in mental health assessments, it is difficult to accurately identify complex emotions such as anxiety and depression present simultaneously in an individual, thus affecting the accuracy of the assessment; in human-computer interaction, it is also unable to provide accurate and appropriate feedback based on the user's complex emotional responses (such as "expectation + tension"), reducing the naturalness and effectiveness of the interaction; and in scenarios such as intelligent video surveillance, grasping the true emotional state of individuals in complex social situations also becomes difficult.

[0006] It is evident that existing emotion recognition technologies have significant shortcomings when dealing with complex human emotions, and there is an urgent need to develop more effective and accurate methods for recognizing complex emotions to meet the growing demand for complex emotion recognition in real-world application scenarios. Summary of the Invention

[0007] To overcome the problems existing in related technologies, this disclosure provides a method, system and medium for facial complex emotion recognition based on scene-induced emotions, aiming to improve the accuracy and practicality of complex emotion recognition, so as to better adapt to the diverse needs of complex emotion recognition in actual application scenarios.

[0008] According to a first aspect of the present disclosure, a method for recognizing complex facial emotions based on scene-induced emotions is provided, the method comprising:

[0009] Obtain facial video of the target user;

[0010] The target user's facial video is input into a pre-trained composite emotion recognition model. Based on the model, the composite emotion assessment information of the target user is identified, wherein the composite emotion assessment information includes composite emotion category and composite emotion intensity.

[0011] Optionally, the construction of the pre-trained composite emotion recognition model includes:

[0012] Acquire facial videos of participants under complex emotions and corresponding complex emotion assessment information for each video;

[0013] Based on the facial videos of the test subjects under complex emotions, extract facial micro-expression sequences;

[0014] The facial micro-expression sequences are correlated with the corresponding composite emotion assessment information of the test subjects to construct a composite emotion sample database;

[0015] The composite emotion sample database is divided into a training set and a test set. Based on the training set, an initial composite emotion recognition model is trained to obtain a trained composite emotion recognition model. The trained composite emotion recognition model is then validated based on the test set. The initial composite emotion recognition model takes facial micro-expression sequences as input and composite emotion evaluation information as output.

[0016] The trained composite emotion recognition model is obtained after verification if the indicators meet the preset threshold.

[0017] Optionally, acquiring facial videos of test subjects under complex emotions and corresponding complex emotion assessment information for each video includes:

[0018] Determine the categories of complex emotions to be induced and the intensity of the complex emotions to be induced for each category, wherein the complex emotion categories include at least two different basic emotions;

[0019] Based on the categories of complex emotions to be induced, an immersive induction scenario is constructed;

[0020] Based on the immersive evoked scenario and the intensity of the complex emotions to be evoked, the experimental participants were evoked to induce complex emotions.

[0021] Acquire facial videos of the test subjects in the immersive induced scenario, as well as composite emotion assessment information of the test subjects corresponding to each video.

[0022] Optionally, acquiring the facial videos of the test subjects in the immersive induced scenario, and the composite emotion assessment information of the test subjects corresponding to each video, includes:

[0023] During the process of inducing the complex emotions, facial videos of the test subjects were collected simultaneously;

[0024] The experimenters obtained their self-assessment results under the immersive induced scenario, and the self-assessment results included: the category of complex emotion and the intensity of complex emotion;

[0025] The consistency between the self-assessment results and the categories and intensities of the complex emotions to be induced was verified.

[0026] The facial video corresponding to the consistent verification results is used as the facial video of the test subject in the immersive induced scenario, and the self-assessment result is used as the corresponding composite emotional assessment information of the test subject.

[0027] Optionally, the immersive evoked scenarios include: visual stimulation scenarios, auditory stimulation scenarios, somatosensory stimulation scenarios, cognitive task stimulation scenarios, and social situational stimulation scenarios.

[0028] The process of inducing complex emotions in test subjects based on the immersive evoked scenario and the intensity of the complex emotions to be induced includes: inducing complex emotions in test subjects according to a dynamic evoked control model, wherein the dynamic evoked control model is...

[0029] (1);

[0030] Where Ie is the intensity of the complex emotion at time t, and its range is determined by the intensity of the complex emotion to be induced. Let be the intensity of the i-th scene stimulus at time t; Assign a weight to the contribution of the stimulus to the complex emotion. satisfy n is the number of stimulus types, and n satisfies n≥2.

[0031] Optionally, the step of extracting facial micro-expression sequences from the facial video of the test subject under complex emotions includes:

[0032] The facial video is preprocessed to generate a spatial-temporal paired face image sample set, wherein the spatial-temporal paired face image sample set includes a static face image sequence and a set of inter-frame optical flow sequences of face images;

[0033] The static face image sequence is input into a two-dimensional convolutional neural network to extract spatial features and generate a static face image appearance feature vector sequence.

[0034] The optical flow sequence set of the face image is input into a three-dimensional convolutional neural network to extract temporal features, analyze the motion patterns between image frames, capture the spatiotemporal evolution of muscle movement, and generate a dynamic feature vector sequence of the face image.

[0035] The static facial image appearance feature vector sequence and the dynamic facial image feature vector sequence are weighted and fused to generate a temporal feature curve of the facial image;

[0036] The temporal feature curve of the face image is denoised, and the micro-expression sequence of the face is extracted based on the denoised temporal feature curve of the face image.

[0037] Optionally, the step of preprocessing the facial video to generate a spatially-temporally paired set of facial image samples includes:

[0038] Face detection is performed frame by frame on the facial video to locate facial regions;

[0039] Based on the located facial region, facial pose estimation and correction are performed to generate an initial adjusted facial region.

[0040] Key facial feature points of the initially adjusted facial region are extracted and geometrically normalized. The facial pose is then fine-tuned to obtain a pose-normalized facial image sequence.

[0041] Sliding window sampling is performed on the pose-normalized face image sequence to generate a fixed-length pose-normalized face image continuous subsequence.

[0042] Optical flow field calculations are performed on continuous subsequences of the pose-normalized face images to obtain an inter-frame optical flow sequence set of face images.

[0043] Extract representative frames from each pose-normalized face image continuous subsequence to obtain a static face image sequence;

[0044] Based on the static face image sequence and the set of inter-frame optical flow sequences of face images, a spatial-temporal paired face image sample set is generated.

[0045] Optionally, validating the trained composite emotion recognition model based on the test set includes:

[0046] The facial micro-expression sequences of the test set are input into the trained composite emotion recognition model to obtain the predicted composite emotion category set Ci and the predicted composite emotion intensity set Ii;

[0047] Based on the real composite sentiment assessment information of the test set, the following validation metrics were calculated:

[0048] For each composite emotion category Ci, calculate the composite emotion category accuracy Acci:

[0049] (2);

[0050] Where TPi is the number of samples that are both true and predicted as Ci, TNi is the number of samples that are neither true nor predicted as Ci, and N is the total number of samples in the test set.

[0051] The arithmetic mean of the accuracy (Acci) for all composite emotion categories is used to obtain the macro accuracy (MA) for the composite emotion category:

[0052] MA= (3);

[0053] Where K is the number of composite emotion categories in the test set;

[0054] Calculate the root mean square error (RMSE) of the composite emotion intensity based on each composite emotion intensity Ii:

[0055] (4);

[0056] Where Ii represents the predicted composite emotion intensity of the i-th sample. Let represent the true composite emotion intensity of the i-th sample.

[0057] According to a second aspect of the present disclosure, a scene-induced facial complex emotion recognition system is provided, the system comprising:

[0058] The video acquisition module is used to acquire facial videos of the target user.

[0059] The composite emotion recognition module is used to input the target user's facial video into a pre-trained composite emotion recognition model, and based on the model, to identify the target user's composite emotion assessment information, wherein the composite emotion assessment information includes composite emotion category and composite emotion intensity.

[0060] According to a third aspect of the present disclosure, the computer-readable storage medium stores computer instructions, which, when invoked, are used to execute the above-described scene-induced facial complex emotion recognition method.

[0061] In summary, this invention provides a method, system, and medium for scene-induced facial complex emotion recognition. The facial complex emotion recognition method provided by this invention acquires a facial video of a target user and inputs it into a trained complex emotion recognition model. Internally, the model extracts and identifies the micro-expression sequences of the target user's face based on the facial video, thereby obtaining the target user's complex emotion assessment information. Since micro-expression sequences reflect a user's true emotional state more accurately than ordinary expressions, combined with a specially trained complex emotion recognition model, it can accurately capture the combination of basic emotions in complex emotions, significantly improving the refinement and accuracy of complex emotion recognition. Simultaneously, by acquiring the intensity of complex emotions, a more detailed description of complex emotions can be provided, rather than merely remaining at the category level. This allows the method to provide more appropriate feedback based on emotion intensity in human-computer interaction scenarios, effectively improving the interaction effect. The complex emotion recognition method provided by this invention can provide reliable emotion recognition basis for various practical application scenarios that require accurate understanding of user emotions, meeting the needs of practical scenarios for accurate recognition of complex emotions. The complex emotion recognition system provided by this invention, through the collaborative cooperation of the video acquisition module and the complex emotion recognition module, forms a complete complex emotion recognition process. The two modules have a clear division of labor and are closely connected, enabling efficient completion of the entire process from facial video capture to the acquisition of complex emotion assessment information. This system facilitates rapid deployment and application in practical situations, providing stable and efficient system support for various scenarios requiring real-time acquisition of users' complex emotions. The computer-readable storage medium provided by this invention stores computer instructions for executing the aforementioned complex emotion recognition method or system, allowing the relevant methods and systems to be conveniently stored, transmitted, and invoked in the form of computer programs; facilitating the flexible application of complex emotion recognition technology to various hardware platforms and real-world scenarios.

[0062] Other features and advantages disclosed in this invention will be described in detail in the following detailed description section. Attached Figure Description

[0063] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the following detailed description to explain the present disclosure, but do not constitute a limitation thereof. In the drawings:

[0064] Figure 1 is an emotion classification table illustrated according to an exemplary embodiment.

[0065] Figure 2 is a flowchart illustrating a scene-induced facial complex emotion recognition method according to an exemplary embodiment;

[0066] Figure 3 is a flowchart illustrating a method for constructing a composite emotion recognition model according to an exemplary embodiment;

[0067] Figure 4 is a flowchart illustrating a method for obtaining facial videos of test subjects under complex emotions and the test subject's complex emotion assessment information corresponding to each video, according to an exemplary embodiment.

[0068] Figure 5 is a flowchart illustrating a method for obtaining facial videos of test subjects in an immersive induced scenario, and composite emotion assessment information of the test subjects corresponding to each video, according to an exemplary embodiment.

[0069] Figure 6 is a flowchart illustrating a method for extracting facial micro-expression sequences from facial videos of test subjects under complex emotions, according to an exemplary embodiment.

[0070] Figure 7 is a flowchart illustrating a method for preprocessing facial video data to generate a spatially-temporally paired set of facial image samples according to an exemplary embodiment.

[0071] Figure 8 is a block diagram illustrating a scene-induced facial complex emotion recognition system according to an exemplary embodiment. Detailed Implementation

[0072] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the scope of the present disclosure.

[0073] Figure 1 is an emotion classification table illustrated according to an exemplary embodiment. Figure 1 As shown, in the field of emotion research, Paul Ekman's basic emotion theory defines single and independent basic emotion types such as happiness, sadness, anger, and surprise. These emotions have clear physiological characteristics and expression patterns, and are the basic units constituting emotions. It should be noted that the subordinate manifestations (specific states) of each basic emotion type vary in different scenarios. The subordinate manifestations listed in Figure 1 are only typical examples, not all possible forms. For example, the basic emotion of happiness, in addition to... Figure 1 In addition to the "joy, satisfaction, and pleasure" listed above, it may also manifest as specific states such as "excitement, relief, and ecstasy"; angry subordinates may express "dissatisfaction, hostility, and irritability" as well as "indignation, annoyance, and rage," depending on the triggering scenario and the intensity of the emotion. A large body of literature and theoretical work has already systematically elaborated on basic emotion theories, so they will not be repeated here.

[0074] In contrast, complex emotions are a complex emotional state formed by the interweaving and fusion of two or more basic emotions, such as "both anxious and depressed" (integrating basic emotional components such as fear and anger) or "both impatient and expectant" (integrating basic emotions such as anger and happiness). Their expression is often deeply tied to specific scenarios and has become the mainstream form of human emotional expression.

[0075] Currently, most existing technologies in the field of emotion recognition are still limited to recognizing the aforementioned single basic emotions. These technical solutions are based on basic emotion theory and can only judge independent emotion categories such as happiness and sadness, outputting corresponding category information. However, in practical applications, when an individual is in a complex state of multiple intertwined emotions, existing technologies often struggle to accurately capture such complex emotions: they either forcibly classify them into a certain basic emotion or output invalid results due to their inability to process them.

[0076] This technological limitation leads to significant constraints in many scenarios: in mental health assessments, it struggles to distinguish between coexisting complex emotions such as anxiety and depression, resulting in distorted assessment results; in human-computer interaction, it cannot provide accurate feedback based on the user's complex state of "both impatience and anticipation," impacting the user experience; and in educational tutoring, it cannot identify the student's learning state of "both tension and focus" (integrating basic emotions such as fear and surprise) to adjust teaching strategies. Furthermore, existing technologies typically cannot quantify the intensity of emotions, only providing rough category labels.

[0077] Therefore, although basic emotion recognition has certain application value in specific scenarios, its capabilities are clearly insufficient when faced with scenarios that require a deep understanding of users' true feelings, the implementation of refined interactions, or the assessment of complex psychological states. It cannot meet the growing demand for depth and breadth of emotion understanding, nor can it adapt to real emotional expression scenarios dominated by complex emotions.

[0078] To overcome the above-mentioned shortcomings, this invention provides a method for facial complex emotion recognition based on scene-induced factors, such as... Figure 2 As shown, the method includes the following steps:

[0079] Step S101: Obtain the target user's facial video.

[0080] Step S102: Input the target user's facial video into a pre-trained composite emotion recognition model. Based on this model, identify the target user's composite emotion assessment information. The composite emotion assessment information includes the composite emotion category and the composite emotion intensity.

[0081] In step S101, there are no limitations on the method or format for acquiring the target user's facial video. The acquisition method can utilize various devices with camera capabilities, such as webcams and camcorders; acquisition via wired connection, wireless transmission, or local storage is applicable. The video format can include, but is not limited to, common formats such as MP4, AVI, and MOV. The frame rate is also not limited, and can be, but is not limited to, 24fps, 30fps, and 60fps. The video only needs to completely record the dynamic facial information of the target user.

[0082] In step S101, dynamic facial video of the target user is acquired to provide the raw data foundation for subsequent extraction of micro-expression sequences. Compared with static images, video can completely record the temporal changes of facial muscle movements, ensuring the capture of subtle and transient dynamic features in complex emotional expressions, and providing data support for accurate identification of complex emotions.

[0083] In step S102, the trained model outputs composite emotion assessment information containing category and intensity, overcoming the limitation of existing technologies that can only identify a single basic emotion. On the one hand, composite emotion category identification can clearly identify combinations of multiple basic emotions, such as "tension + excitement," avoiding errors caused by forced classification; on the other hand, intensity quantification can reflect the proportion and dynamic changes of each emotional component, achieving a refined description of composite emotions.

[0084] In an exemplary embodiment of this invention, the facial complex emotion recognition method provided by this invention inputs a target user's facial video into a pre-trained complex emotion recognition model to obtain complex emotion assessment information including category and intensity. This enables it to achieve precise adaptation in real-world scenarios: In mental health assessment, it can output complex emotion assessment information such as "anxiety-depression (anxiety intensity 0.6, depression intensity 0.4)" based on the patient's actual situation, providing quantitative basis for physicians to differentiate comorbid states and avoiding misjudgment based on a single emotion. In human-computer interaction, it can target the user's current emotional state, such as "anxiety-expectation (anxiety intensity 0.5, expectation intensity 0.4)". The device can dynamically adjust its response logic, such as accelerating core operations while displaying progress feedback. In intelligent education, after identifying a student's "confusion-interest (confusion intensity 0.5, interest intensity 0.6)," the teaching system can accurately push tiered explanations, balancing explanation and interest maintenance. In job interviews, analyzing a candidate's "nervousness-confidence (nervousness intensity 0.3, confidence intensity 0.8)" can help interviewers more comprehensively assess their psychological qualities. By simultaneously outputting category and intensity, this method overcomes the limitation of existing technologies that can only identify a single emotion, providing more realistic emotional data support for various scenarios and significantly enhancing application value.

[0085] For example, in an exemplary embodiment disclosed in this invention, the invention also provides a method for constructing a pre-trained composite emotion recognition model, such as... Figure 3 As shown, the specific steps include:

[0086] Step S201: Obtain facial videos of the test subjects under complex emotions and the corresponding complex emotion assessment information of the test subjects for each video.

[0087] In step S201, facial videos and complex emotion assessment information of test subjects under complex emotional states are collected in a targeted manner to ensure that the data is directly related to complex emotional scenarios. This solves the problem that the training data in the prior art is mostly single emotion and lacks complex emotion samples, laying the foundation for the model to learn complex emotion features.

[0088] Step S202: Extract facial micro-expression sequences from the facial videos of the test subjects under complex emotions.

[0089] In step S202, this method extracts facial micro-expression sequences from the facial videos of test subjects, transforming the raw video data into core feature data that can be used for model training. Facial micro-expression sequences can accurately capture the dynamic superposition and instantaneous changes of various basic emotions in complex emotions, thereby enabling the model to learn the emotional essence closer to real-life situations, and thus significantly improving the model's ability and accuracy in recognizing complex emotions.

[0090] Step S203: Establish a correlation between the facial micro-expression sequences and the corresponding composite emotion assessment information of the test subjects to construct a composite emotion sample database.

[0091] In step S203, facial micro-expression sequences are associated with corresponding composite emotion assessment information to construct a composite emotion sample database for model training. This process aligns unstructured video data with structured composite emotion assessment information, forming a large-scale, high-quality, and highly traceable labeled dataset. This provides a solid data foundation for subsequent model training, effectively avoiding the bottleneck in model generalization ability caused by messy data and inconsistent labeling, thus ensuring the final effect of model training.

[0092] Step S204: Divide the composite emotion sample database into a training set and a test set. Based on the training set, train an initial composite emotion recognition model to obtain a trained composite emotion recognition model. Validate the trained composite emotion recognition model based on the test set. The initial composite emotion recognition model takes facial micro-expression sequences as input and composite emotion evaluation information as output.

[0093] In step S204, the model's architecture is defined as taking facial micro-expression sequences as input and outputting complex emotion assessment information. It leverages deep learning's ability to fit complex features to achieve accurate mapping from micro-expressions to complex emotions. Compared to traditional machine learning models, this architecture better captures the spatiotemporal evolution characteristics of micro-expressions, improving the complexity adaptability of complex emotion recognition.

[0094] By dividing the database into training and testing sets, model training and independent validation are separated, avoiding overfitting. Validation based on the testing set objectively evaluates model performance, ensuring that the model can stably identify complex emotions even on unseen data, thus solving the problem of weak generalization ability in existing technologies.

[0095] Step S205: Obtain a well-trained composite emotion recognition model whose validation metrics meet the preset threshold.

[0096] In step S205, qualified models are screened and verified using a preset threshold to ensure that the model's recognition accuracy meets the requirements of practical applications. This step provides a reliable model for subsequent complex emotion recognition in real-world scenarios, ensuring the practicality and stability of the technical solution.

[0097] For example, such as Figure 4 As shown, in an exemplary embodiment of the present invention, step S201: acquiring facial videos of test subjects under complex emotions and the test subject's complex emotion assessment information corresponding to each video, specifically includes the following steps:

[0098] Step S301: Determine the categories of complex emotions to be induced and the intensity of the complex emotions to be induced for each category. The complex emotion categories include at least two different basic emotions.

[0099] In step S301, by pre-defining the category of the complex emotion to be induced and the range of the complex intensity to be induced, the targeting and controllability of emotion induction are achieved, solving the problems of no clear target and ambiguous sample labels in the prior art, and ensuring that the collected data is directly related to the specific complex emotion.

[0100] Step S302: Construct an immersive triggering scenario based on the category of complex emotions to be induced.

[0101] In step S302, an immersive scenario is designed for the complex emotions to be induced, which can more closely resemble the real situation to stimulate the complex emotions of the test subjects. Compared with a single stimulus method, it can enhance the effectiveness and stability of emotion induction and ensure that the collected facial videos truly reflect the complex emotional state.

[0102] Step S303: Based on the immersive evoked scenario and the intensity of the complex emotion to be evoked, induce complex emotions in the test subjects.

[0103] In step S303, the immersive scene is combined with the intensity of the complex emotion to be induced. Through multi-dimensional synergistic stimulation, the generation process of complex emotion is precisely controlled, so that the emotional state of the test subjects is more in line with the target category and intensity, avoiding insufficient or biased emotional expression caused by single stimulation, and improving the relevance of sample data.

[0104] Step S304: Obtain facial videos of the test subjects in immersive induced scenarios, as well as composite emotion assessment information of the test subjects corresponding to each video.

[0105] In step S304, since both the facial video and the evaluation information are acquired within the same immersive evoked scenario, they naturally carry scenario-related features (such as scenario stimulus type and intensity). This means that the subsequently constructed composite emotion sample database not only contains the mapping relationship between "micro-expressions and composite emotion evaluation information" but also implicitly contains the correlation pattern between "scenario stimuli and emotional responses." This correlated data can provide richer contextual information for model training, thereby improving the model's adaptability to recognizing composite emotions in different scenarios.

[0106] Specifically, such as Figure 5 As shown, in an exemplary embodiment of the present invention, step S304: acquiring facial videos of the test subjects in an immersive induced scenario, and composite emotion assessment information of the test subjects corresponding to each video, specifically includes the following steps:

[0107] Step S401: During the process of inducing complex emotions, facial videos of the test subjects are collected simultaneously;

[0108] In step S401, facial videos are simultaneously acquired during the emotion induction process, and the dynamic changes of facial micro-expressions corresponding to complex emotions are fully recorded, providing raw data for subsequent extraction of facial micro-expression sequences and ensuring the time synchronization of data with emotional states.

[0109] Step S402: Obtain the self-assessment results of the test subjects in the immersive induced scenario. The self-assessment results include: the category of complex emotions and the intensity of complex emotions.

[0110] In step S402, self-assessment results of the participants can be obtained through emotion assessment tools. By using standardized emotion assessment tools, subjective emotional feedback from participants is collected, linking objective video data with subjective emotional experiences. This overcomes the limitations of relying solely on behavioral data and provides a direct basis for generating composite emotion assessment information. In the exemplary embodiments disclosed in this invention, standardized emotion assessment tools can be, but are not limited to, the SAM Emotion Self-Assessment Scale, the PANAS scale, the SAM scale, and the PAD model, etc. These standardized emotion assessment tools, through structured questionnaires or scales, can effectively capture an individual's subjective emotional state. In step S305, using these tools can overcome the limitations of relying solely on behavioral data and provide a direct basis for generating composite emotion assessment information. These tools have wide applications in psychology, clinical medicine, marketing, and other fields, ensuring that subjective emotional feedback is combined with objective video data, thus improving data quality.

[0111] Step S403: Verify the consistency between the self-assessment results and the categories and intensities of the complex emotions to be induced.

[0112] Step S404: Use the facial video corresponding to the consistent verification results as the facial video of the test subject in the immersive induced scenario, and use the self-assessment results as the corresponding composite emotion assessment information of the test subject.

[0113] In step S403, the test subject's self-assessment results are compared bidirectionally with the pre-determined categories and intensity ranges of the complex emotions to be induced. For example, if the category and intensity of the complex emotions to be induced are "tension-expectation (tension intensity 0.3-0.5, expectation intensity 0.6-0.8)," the verification process must simultaneously confirm whether the self-assessed emotion category is indeed "tension-expectation," and whether the intensity of both emotions falls within the preset range. This multi-dimensional verification not only eliminates misjudgments of emotion categories but also filters out samples with excessive intensity deviations.

[0114] In step S404, the facial video corresponding to the verification result is rigidly bound to the self-assessment result. This binding mechanism establishes a one-to-one spatiotemporal correlation between the dynamic sequence information of facial micro-expressions recorded in the facial video and the complex emotion category and intensity of the self-assessment. This ensures the authenticity of the mapping relationship between "facial micro-expression features - complex emotion assessment information" and provides a complete data chain of "scene stimulus - emotion expression - subjective experience" for subsequent model training. Compared with unverified data binding methods, this step effectively avoids the problem of "incorrect association between facial video and self-assessment result" and significantly improves the purity and data correlation of the sample database.

[0115] In steps S403 and S404, the consistency between the self-assessment results and the preset composite emotion categories and intensities to be induced is verified to select real and valid composite emotion samples, and invalid data that fails to induce emotions or has biased composite emotion assessment information is excluded. This ensures the accuracy and reliability of the final composite emotion assessment information and avoids interference from low-quality sample data on model training.

[0116] Specifically, in the exemplary embodiments disclosed in this invention, the immersive evoked scenarios include: visual stimulation scenarios, auditory stimulation scenarios, somatosensory stimulation scenarios, cognitive task stimulation scenarios, and social situational stimulation scenarios. Step S303: Based on the immersive evoked scenarios and the intensity of the complex emotions to be evoked, the experimental subjects' complex emotions are evoked, specifically including: evoking the experimental subjects' complex emotions according to a dynamic evoked control model, wherein the dynamic evoked control model is...

[0117] (1);

[0118] Where Ie is the intensity of the complex emotion at time t, and its range is determined by the intensity of the complex emotion to be induced. Let be the intensity of the i-th scene stimulus at time t; The weight of the stimulus's contribution to the complex emotion; satisfy n is the number of stimulus types, and n satisfies n≥2.

[0119] The method for inducing complex emotions based on scenarios provided by this invention is not limited to single-sensory stimulation, but comprehensively utilizes multiple stimulation scenarios such as visual, auditory, somatosensory, cognitive tasks, and social situations. These scenarios are not simply piled up, but flexibly combined according to the type of complex emotion to be induced, forming a multimodal stimulus combination. This multimodal stimulus combination simulates the complex environment in which emotions arise in the real world, activating multiple sensory channels simultaneously, making the induced emotional experience closer to the real feeling under natural conditions, avoiding the stereotypical or unrealistic emotional reactions that may result from single stimuli. Simultaneously, a dynamic induction control model is introduced into the emotion induction process. The core of this model lies in its ability to dynamically and in real-time regulate the intensity of complex emotions. By adjusting the intensity of each stimulation scenario... and their corresponding weights It can precisely control the intensity of the induced emotion Ie within a preset target range, solving the problem that the intensity of emotion induction is difficult to accurately grasp in existing technologies.

[0120] In this invention, such as Figure 1As shown, the induction and identification of complex emotions are based on seven basic emotions: anger, happiness, sadness, surprise, fear, disgust, and neutral emotions. Specific complex emotion categories are formed by combining two or more different basic emotions, and the intensity range corresponding to each basic emotion category is determined simultaneously, thereby determining the overall intensity range of complex emotions. The intensity range of complex emotions can be determined by superimposing the intensity ranges of two or more basic emotions.

[0121] Specifically, for a given category of complex emotion, a suitable combination of immersive evoked scenarios is first selected. Then, a dynamic evoked control model is used to adjust the intensity and weight of stimuli in each scenario. For example, the weight of visual stimuli is increased to strengthen the "fear" component, and the intensity of auditory stimuli is adjusted to match the preset intensity range of "sadness." Ultimately, this achieves precise induction of the complex emotion category and intensity, providing standardized complex emotion samples for subsequent model training.

[0122] To create a comprehensive and representative database of complex emotions, an exemplary embodiment of this disclosure employs a phased induction process. In the initial phase, binary complex emotions composed of any two basic emotions are induced, and facial video data of the participants is simultaneously collected. At the same time, each video segment is precisely labeled with its corresponding binary complex emotion category and intensity. After inducing and collecting data for all binary combinations, ternary basic emotion combinations are induced, again acquiring video data and generating corresponding complex emotion assessment information. This process continues, gradually increasing the number of basic emotions involved to conduct more complex complex emotion induction experiments. Through this progressive and exhaustive approach, a comprehensive and clearly labeled database of complex emotions can ultimately be constructed.

[0123] For example, such as Figure 6 As shown, in an exemplary embodiment of the present invention, step S202: extracting facial micro-expression sequences from the facial video of the test subject under complex emotions, specifically includes the following steps:

[0124] Step S501: Perform video data preprocessing on the facial video to generate a spatial-temporal paired face image sample set, wherein the spatial-temporal paired face image sample set includes a static face image sequence and a set of inter-frame optical flow sequences of face images.

[0125] In step S501, video data preprocessing simultaneously generates a set of static face image sequences and an optical flow sequence set, transforming the original video into a complementary data structure with spatial (static image) and temporal (optical flow) dimensions. The static images preserve the spatial distribution features of facial muscle morphology, while the optical flow sequence records the motion trajectories between frames, laying the foundation for subsequent extraction of spatial and temporal features.

[0126] Step S502: Input the static face image sequence into a two-dimensional convolutional neural network to extract spatial features and generate a static face image appearance feature vector sequence.

[0127] In step S502, a two-dimensional convolutional neural network is used to process static facial images, automatically extracting spatial-dimensional facial features such as eyebrow shape and mouth corner curvature. The local perception capability and weight sharing mechanism of the two-dimensional convolutional neural network enable it to efficiently capture the spatial relationships of key facial points. The generated facial feature vector sequence retains the detailed features of micro-expressions in the spatial domain, providing spatial-dimensional semantic information for subsequent fusion with dynamic features.

[0128] Step S503: Input the optical flow sequence set of face images into a three-dimensional convolutional neural network, extract temporal features, analyze the motion patterns between image frames, capture the spatiotemporal evolution of muscle movement, and generate a dynamic feature vector sequence of face images.

[0129] In step S503, a three-dimensional convolutional neural network is used to process the optical flow sequence. The motion patterns between frames are analyzed through the third-dimensional convolutional kernel (time dimension) to capture the temporal changes in facial micro-expressions from their generation to their disappearance. This step encodes the spatiotemporal evolution features of facial micro-expressions into a dynamic feature vector sequence, breaking through the limitations of traditional methods that only analyze single frames or adjacent frames, and significantly improving the ability to capture brief and subtle micro-expressions.

[0130] Step S504: Weighted fusion of the static face image appearance feature vector sequence and the face image dynamic feature vector sequence to generate a face image temporal feature curve.

[0131] In step S504, a weighted fusion mechanism organically combines spatial appearance features with temporal dynamic features, adaptively adjusting the weights based on the importance of spatial and temporal features in different complex emotional expressions. For example, the spatial feature of raised eyebrows has a higher weight in surprise, while the temporal feature of pupil constriction has a higher weight in fear. The temporal feature curve generated by this fusion method preserves both the morphological and temporal evolution features of micro-expressions, enabling subsequent recognition models to analyze complex emotions from multiple dimensions.

[0132] Step S505: Denoise the temporal feature curve of the face image, and extract the micro-expression sequence of the face based on the denoised temporal feature curve of the face image.

[0133] In step S505, the temporal feature curve of the facial image is susceptible to various noise interferences during acquisition and processing, such as natural facial tremors, lighting changes, and device noise. These noises cause irregular fluctuations in the feature curve, masking the true features of micro-expressions. Through targeted denoising algorithms, such as wavelet transform and Kalman filtering, these useless information caused by lighting changes, head movements, and sensor noise can be accurately identified and removed. This makes the denoised curve smoother and cleaner, truly reflecting the subtle movement trajectories of facial muscles, fundamentally improving the purity and signal-to-noise ratio of micro-expression features.

[0134] Specifically, such as Figure 7 As shown, in an exemplary embodiment of the present invention, step S501: preprocessing the facial video to generate a spatially-temporally paired face image sample set includes the following steps:

[0135] Step S601: Perform face detection frame by frame on the facial video and locate the facial region.

[0136] Step S602: Based on the located facial region, perform facial pose estimation and correction to generate an initial adjusted facial region.

[0137] Step S603: Extract key facial feature points from the initial facial region, perform geometric normalization on them, fine-tune the facial pose, and obtain a pose-normalized facial image sequence.

[0138] In steps S601 to S603, face detection and pose estimation are performed frame by frame to locate facial regions and correct pose deviations such as tilt and rotation, ensuring spatial consistency in subsequent feature extraction. Based on key facial feature points, such as the corners of the eyes and the tip of the nose, geometric transformations are performed to eliminate individual differences in facial structure, aligning the same expressions across different individuals in space. This process is used to reduce interference from irrelevant factors.

[0139] Step S604: Perform sliding window sampling on the pose-normalized face image sequence to generate a fixed-length pose-normalized face image continuous subsequence.

[0140] In step S604, the face image sequence is segmented into fixed-length continuous subsequences of face images, ensuring that each subsequence contains a complete cycle of facial expression changes. This operation balances the integrity of temporal features with computational efficiency, providing stable data units for subsequent optical flow calculations and feature extraction.

[0141] Step S605: Perform optical flow field calculation on the continuous subsequence of pose-normalized face images to obtain the inter-frame optical flow sequence set of face images.

[0142] In step S605, the motion trajectory of facial muscles is quantified by calculating the optical flow vectors between adjacent frames, generating an optical flow sequence set that reflects the dynamic changes in micro-expressions. Optical flow features are highly sensitive to rapid and subtle muscle movements, effectively capturing temporal information that traditional static images cannot express.

[0143] Step S606: Extract representative frames from each pose-normalized face image continuous subsequence to obtain a static face image sequence.

[0144] In step S606, the most representative frame is selected from each subsequence as a static face image, preserving the spatial morphological features of micro-expressions. Compared with traditional full-frame processing, this method reduces redundant data and improves feature extraction efficiency.

[0145] Step S607: Generate a spatial-temporal paired face image sample set based on the static face image sequence and the face image inter-frame optical flow sequence set.

[0146] In step S607, the static face image sequence (spatial feature carrier) is paired with the optical flow sequence set (temporal feature carrier) to form a spatiotemporally complementary sample structure. This pairing mechanism enables the subsequent model to simultaneously learn the spatial morphology and temporal evolution of facial expressions, significantly enhancing the ability to capture multidimensional features in complex emotions.

[0147] For example, in an exemplary embodiment disclosed in this invention, step S205: validating the trained composite emotion recognition model based on a test set includes the following steps:

[0148] The facial micro-expression sequences from the test set are input into the trained composite emotion recognition model to obtain the predicted composite emotion category set Ci and the predicted composite emotion intensity set Ii.

[0149] Based on the real composite sentiment assessment information from the test set, the following validation metrics were calculated:

[0150] For each composite emotion category Ci, calculate the composite emotion category accuracy Acci:

[0151] (2);

[0152] Where TPi is the number of samples that are both true and predicted as Ci, TNi is the number of samples that are neither true nor predicted as Ci, and N is the total number of samples in the test set.

[0153] The arithmetic mean of the accuracy (Acci) for all composite emotion categories is used to obtain the macro accuracy (MA) for the composite emotion category:

[0154] MA= (3);

[0155] Where K is the number of composite emotion categories in the test set;

[0156] Calculate the root mean square error (RMSE) of the composite emotion intensity based on each composite emotion intensity Ii:

[0157] (4);

[0158] Where Ii represents the predicted composite emotion intensity of the i-th sample. Let represent the true composite emotion intensity of the i-th sample.

[0159] Formula (2) includes the contributions of both true positives (TPi) and true negatives (TNi), evaluating both the model's ability to correctly identify a specific category Ci and its ability to exclude non-Ci categories. TPi reflects the model's correct detection rate for that category, while TNi reflects the model's ability to avoid misclassifying other composite emotion categories as Ci. By calculating the accuracy of each category Ci separately, evaluation bias caused by uneven sample distribution is avoided. When there are few samples of a certain composite emotion category in the test set, Acci can still independently reflect the model's ability to identify that category, preventing it from being masked by the high accuracy of the majority class. Acci measures the accuracy of the composite emotion recognition model in identifying a specific composite emotion category, with a value range of [0,1]. The closer the value is to 1, the better the recognition effect of that category.

[0160] In formula (3), by calculating the arithmetic mean of the accuracy of all categories, MA provides the model's overall performance across all complex emotion categories, reflecting the model's overall ability to recognize and balance various complex emotions. The range of MA is [0,1]. When MA is close to 1, it indicates that the model's ability to recognize various complex emotions is balanced and efficient; if MA is significantly lower than the Acci of some categories, it indicates that the model has a weakness in category recognition and needs to be optimized accordingly.

[0161] Formula (4) calculates the root mean square error (RMSE) between the predicted and actual composite emotion intensity, thus accurately assessing the model's ability to quantify composite emotion intensity. A smaller RMSE indicates better model performance in this area. Compared to indicators that only assess category accuracy, RMSE evaluates the continuous variable of composite emotion intensity, making it more aligned with the actual needs of composite emotion recognition.

[0162] Specifically, in the exemplary embodiments disclosed in this invention, the preset threshold for the class accuracy (Acci) of all composite emotion categories is 0.7, ensuring that the composite emotion recognition model has a stable recognition ability for each specific composite emotion category, avoiding the impact of insufficient recognition accuracy of individual categories on the overall performance. The preset threshold for the macro-accuracy (MA) of composite emotion categories is 0.75, ensuring that the model's comprehensive recognition performance across all composite emotion categories reaches a high level, reflecting its universal adaptability to multiple types of composite emotions. The preset threshold for the root mean square error (RMSE) of composite emotion intensity is 0.15, ensuring that the model's quantitative prediction accuracy for composite emotion intensity meets the needs of practical applications, and effectively capturing subtle differences in emotion intensity.

[0163] When a composite emotion recognition model simultaneously meets or exceeds the above three metrics on the test set, it indicates that the model has met the expected performance requirements and can accurately and reliably predict composite emotion categories and intensities. At this point, a composite emotion recognition model that satisfies the preset thresholds for the above three metrics is selected.

[0164] Figure 8 This is a block diagram illustrating a scene-induced facial complex emotion recognition system according to an exemplary embodiment, such as... Figure 7 As shown, the system 700 includes:

[0165] Video acquisition module 701 is used to acquire facial videos of the target user;

[0166] The composite emotion recognition module 702 is used to input the target user's facial video into a pre-trained composite emotion recognition model. Based on the model, it identifies the target user's composite emotion assessment information, which includes composite emotion category and composite emotion intensity.

[0167] In the exemplary embodiments disclosed in this invention, the video acquisition module 701, as a front-end data acquisition unit, adapts to various camera devices to ensure complete capture of the user's facial dynamic information in different scenarios. The composite emotion recognition module 702 is connected to the video acquisition module 701, acquiring the target user's facial video from the video acquisition module 701 and extracting the facial micro-expression sequence. This module integrates video preprocessing, spatiotemporal feature extraction, and noise reduction functions, accurately extracting subtle micro-expression dynamic features reflecting composite emotions from the video, transforming unstructured video data into a high-dimensional feature sequence that can be processed by the model, providing core input for subsequent emotion recognition. The composite emotion recognition module 702 inputs the extracted facial micro-expression sequence into a pre-trained composite emotion recognition model and outputs composite emotion assessment information containing category and intensity. This module, relying on the pre-trained composite emotion recognition model, achieves a direct mapping from facial micro-expression features to composite emotion assessment information, breaking through the limitations of traditional single-emotion category recognition and quantifying the output of composite emotion intensity, improving the refinement of emotion description, and ensuring accurate emotion data support for scenarios such as mental health assessment and intelligent education.

[0168] The scene-induced facial complex emotion recognition system provided by this invention achieves functional decoupling and efficient collaboration through modular design. The video acquisition module ensures data integrity and continuity, while the complex emotion recognition module is responsible for extracting facial micro-expression sequences from facial videos and completing the final recognition task, outputting complex emotion assessment information. This architecture ensures professional processing at each stage while avoiding redundancy in data transmission and processing through orderly connection between modules. This enables the system to complete the complex emotion recognition process end-to-end, significantly improving the practicality and feasibility of the technical solution and meeting the needs for rapid and accurate recognition of complex emotions in real-world scenarios.

[0169] For example, in an exemplary embodiment disclosed in this invention, this application further proposes a computer-readable storage medium storing computer instructions that, when invoked, are used to execute the scene-induced facial complex emotion recognition method or system described in the foregoing embodiments.

[0170] Specifically, this application transforms the logical flow of the aforementioned scene-induced facial complex emotion recognition method and system into computer-executable code and embeds it in a computer-readable storage medium, achieving standardized encapsulation and cross-platform deployment of complex emotion recognition technology. Thus, the trained complex emotion recognition model can be directly adapted to terminal devices such as mobile devices and embedded systems, independent of the development environment. For example, in mental health monitoring scenarios, users can invoke instructions from the storage medium via mobile terminals to analyze their own or monitored subjects' complex emotional states in real time; in the field of intelligent security, edge computing devices can directly read data from the storage medium and perform emotion recognition, significantly reducing reliance on cloud services. This technical solution, through standardized carriers and cross-platform adaptability, effectively improves the engineering implementation efficiency of complex emotion recognition technology and expands its coverage in multiple application scenarios.

[0171] The preferred embodiments of this disclosure have been described in detail above with reference to the accompanying drawings. However, this disclosure is not limited to the specific details of the above embodiments. Within the scope of the technical concept of this disclosure, various simple modifications can be made to the technical solutions of this disclosure, and these simple modifications all fall within the protection scope of this disclosure.

[0172] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. In order to avoid unnecessary repetition, this disclosure will not describe the various possible combinations separately.

[0173] Furthermore, various different embodiments of this disclosure can be combined in any way, as long as they do not violate the spirit of this disclosure, they should also be regarded as the content disclosed in this disclosure.

Claims

1. A method for recognizing complex facial emotions based on scene-induced emotions, characterized in that, The method includes: Obtain facial video of the target user; The target user's facial video is input into a pre-trained composite emotion recognition model. Based on the model, the composite emotion assessment information of the target user is identified, wherein the composite emotion assessment information includes composite emotion category and composite emotion intensity.

2. The scene-induced facial complex emotion recognition method according to claim 1, characterized in that, The construction of the pre-trained composite emotion recognition model includes: Acquire facial videos of participants under complex emotions and corresponding complex emotion assessment information for each video; Based on the facial videos of the test subjects under complex emotions, extract facial micro-expression sequences; The facial micro-expression sequences are correlated with the corresponding composite emotion assessment information of the test subjects to construct a composite emotion sample database; The composite emotion sample database is divided into a training set and a test set. Based on the training set, an initial composite emotion recognition model is trained to obtain a trained composite emotion recognition model. The trained composite emotion recognition model is then validated based on the test set. The initial composite emotion recognition model takes facial micro-expression sequences as input and composite emotion evaluation information as output. The trained composite emotion recognition model is obtained after verification if the indicators meet the preset threshold.

3. The scene-induced facial complex emotion recognition method according to claim 2, characterized in that, The acquisition of facial videos of test subjects under complex emotions and the corresponding complex emotion assessment information for each video includes: Determine the categories of complex emotions to be induced and the intensity of the complex emotions to be induced for each category, wherein the complex emotion categories include at least two different basic emotions; Based on the categories of complex emotions to be induced, an immersive induction scenario is constructed; Based on the immersive evoked scenario and the intensity of the complex emotions to be evoked, the experimental participants were evoked to induce complex emotions. Acquire facial videos of the test subjects in the immersive induced scenario, as well as composite emotion assessment information of the test subjects corresponding to each video.

4. The scene-induced facial complex emotion recognition method according to claim 3, characterized in that, The acquisition of facial videos of test subjects in the immersive induced scenario, and composite emotion assessment information of the test subjects corresponding to each video, includes: During the process of inducing the complex emotions, facial videos of the test subjects were collected simultaneously; The experimenters obtained their self-assessment results under the immersive induced scenario, and the self-assessment results included: the category of complex emotion and the intensity of complex emotion; The consistency between the self-assessment results and the categories and intensities of the complex emotions to be induced was verified. The facial video corresponding to the consistent verification results is used as the facial video of the test subject in the immersive induced scenario, and the self-assessment result is used as the corresponding composite emotional assessment information of the test subject.

5. The scene-induced facial complex emotion recognition method according to claim 3, characterized in that, The immersive evoked scenarios include: visual stimulation scenarios, auditory stimulation scenarios, somatosensory stimulation scenarios, cognitive task stimulation scenarios, and social situational stimulation scenarios. The process of inducing complex emotions in test subjects based on the immersive evoked scenario and the intensity of the complex emotions to be induced includes: inducing complex emotions in test subjects according to a dynamic evoked control model, wherein the dynamic evoked control model is... (1); Where Ie is the intensity of the complex emotion at time t, and its range is determined by the intensity of the complex emotion to be induced. Let be the intensity of the i-th scene stimulus at time t; Assign a weight to the contribution of the stimulus to the complex emotion. satisfy n is the number of stimulus types, and n satisfies n≥2.

6. The scene-induced facial complex emotion recognition method according to claim 2, characterized in that, The step of extracting facial micro-expression sequences from the facial videos of the test subjects under complex emotions includes: The facial video is preprocessed to generate a spatial-temporal paired face image sample set, wherein the spatial-temporal paired face image sample set includes a static face image sequence and a set of inter-frame optical flow sequences of face images; The static face image sequence is input into a two-dimensional convolutional neural network to extract spatial features and generate a static face image appearance feature vector sequence. The optical flow sequence set of the face image is input into a three-dimensional convolutional neural network to extract temporal features, analyze the motion patterns between image frames, capture the spatiotemporal evolution of muscle movement, and generate a dynamic feature vector sequence of the face image. The static facial image appearance feature vector sequence and the dynamic facial image feature vector sequence are weighted and fused to generate a temporal feature curve of the facial image; The temporal feature curve of the face image is denoised, and the micro-expression sequence of the face is extracted based on the denoised temporal feature curve of the face image.

7. The scene-induced facial complex emotion recognition method according to claim 6, characterized in that, The step of preprocessing the facial video to generate a spatially-temporally paired set of facial image samples includes: Face detection is performed frame by frame on the facial video to locate facial regions; Based on the located facial region, facial pose estimation and correction are performed to generate an initial adjusted facial region. Key facial feature points of the initially adjusted facial region are extracted and geometrically normalized. The facial pose is then fine-tuned to obtain a pose-normalized facial image sequence. Sliding window sampling is performed on the pose-normalized face image sequence to generate a fixed-length pose-normalized face image continuous subsequence. Optical flow field calculations are performed on continuous subsequences of the pose-normalized face images to obtain an inter-frame optical flow sequence set of face images. Extract representative frames from each pose-normalized face image continuous subsequence to obtain a static face image sequence; Based on the static face image sequence and the set of inter-frame optical flow sequences of face images, a spatial-temporal paired face image sample set is generated.

8. The scene-induced facial complex emotion recognition method according to claim 2, characterized in that, The validation of the trained composite emotion recognition model based on the test set includes: The facial micro-expression sequences of the test set are input into the trained composite emotion recognition model to obtain the predicted composite emotion category set Ci and the predicted composite emotion intensity set Ii; Based on the real composite sentiment assessment information of the test set, the following validation metrics were calculated: For each composite emotion category Ci, calculate the composite emotion category accuracy Acci: (2); Where TPi is the number of samples that are both true and predicted as Ci, TNi is the number of samples that are neither true nor predicted as Ci, and N is the total number of samples in the test set. The arithmetic mean of the accuracy (Acci) for all composite emotion categories is used to obtain the macro accuracy (MA) for the composite emotion category: AND= (3); Where K is the number of composite emotion categories in the test set; Calculate the root mean square error (RMSE) of the composite emotion intensity based on each composite emotion intensity Ii: (4); Where Ii represents the predicted composite emotion intensity of the i-th sample. Let represent the true composite emotion intensity of the i-th sample.

9. A scene-induced facial complex emotion recognition system, characterized in that, The system includes: The video acquisition module is used to acquire facial videos of the target user. The composite emotion recognition module is used to input the target user's facial video into a pre-trained composite emotion recognition model, and based on the model, to identify the target user's composite emotion assessment information, wherein the composite emotion assessment information includes composite emotion category and composite emotion intensity.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, which, when invoked, are used to execute the scene-induced facial complex emotion recognition method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • A video sequence expression recognition method based on mixed deep learning

    CN109190479A

  • Real-time emotion recognition method and system fusing pupil data and facial expression

    CN113837153A

  • Model training method and device, face emotion recognition method and device, equipment and storage medium

    CN119495114A

  • Emotion evaluation system for old people based on multi-modal physiological and behavior signals

    CN120154337A

  • Micro-expression recognition method based on multi-scale spatiotemporal feature neural network

    US20220269881A1