Emotional pressure recognition method and system based on action form fusion graph neural network

By adopting the emotional stress recognition method of the action morphological fusion graph neural network in campus scenarios, the problem of inaccurate emotion judgment in the existing technology is solved, and accurate identification and real-time monitoring of campus personnel's emotions are achieved, and the effectiveness of safety management is improved.

CN120198976AActive Publication Date: 2025-06-24延安大学西安创新学院

Patent Information

Application Number
CN202510523337.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-06-24
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

The emotional discrimination of personnel in the existing technology in campus scenarios is inaccurate, and the spatial and temporal structure and complementary relationship of the movement forms of various parts of the human body are not fully captured and analyzed, resulting in insufficient timeliness of emotional recognition, misunderstanding or neglect.

Method used

The emotional stress recognition method based on the action morphological fusion graph neural network is adopted, and video image preprocessing is performed through deep learning algorithms to build a dual-graph data spatial structure. The segmentation features of emotional reinforcement are extracted using the T-GCN model combined with the graph partitioning strategy, and the recognition accuracy is improved through transfer learning and small sample learning.

Benefits of technology

It realizes accurate identification of emotions of campus personnel, improves the accuracy and reliability of emotional stress judgment, enhances the real-time monitoring capabilities of campus safety management, and has a wide range of practical application prospects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198976A_ABST
    Figure CN120198976A_ABST
Patent Text Reader

Abstract

The invention discloses an emotion pressure recognition method and system based on an action form fusion graph neural network, and belongs to the field of image processing and machine learning. A double-graph data structure is constructed, wherein graph data with human body key points as nodes and graph data with image frames as nodes are processed through T-GCN and PCSN respectively. The T-GCN model fuses a GCN layer and a Mama encoder, the cohesion and coupling ability of emotion category characterization is enhanced through a graph partitioning strategy, and the spatial-temporal characteristic relation of human body actions is effectively captured. The PCSN adopts a three-branch parallel structure including one-dimensional global average pooling, 3 * 3 convolution and 5 * 5 convolution, efficient extraction of non-homogeneous features of the human body, the face and the hand is achieved, and the long and short distance dependence and local cross-channel interaction relation of the global features is established. According to the system, through combination of transfer learning and small sample learning, a public data set is utilized to initialize a network, and secondary training is performed on a small sample data set of campus emotion pressure, so that the recognition precision is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to an emotion stress recognition method and system based on an action form fusion graph neural network. Background Art

[0002] Emotion stress is a psychological stress state formed by an individual under the stimulation of external things; short-term emotion stress helps an individual to adapt to environmental changes by adjusting their own physiology or psychology; however, under long-term emotion stress, if an individual cannot reasonably relieve it, it may evolve into physical and mental health problems and have extensive and even long-term effects. In the current situation of accelerating information globalization, social competition is becoming increasingly fierce and the educational environment has undergone great changes. Physical and mental health has become an important cornerstone for personal growth, interpersonal relationships and career development. Paying attention to the physical and mental health development of students is of utmost importance; at present, a number of measures have been taken to comprehensively strengthen and improve the mental health work of students, making the mental health work system of students more complete in terms of health education, monitoring and early warning, counseling services, and intervention and treatment, and the mental health work pattern of students with coordinated linkage among schools, families, society and relevant departments more perfect. Among them, monitoring and early warning can diagnose students' emotion stress in a timely manner, prevent the formation of individual psychological problems, and provide sufficient preparation time for intervention and treatment, which is a key link in the mental health work system of students.

[0003] According to different monitoring target elements, emotion stress monitoring can be divided into physiological monitoring and non-physiological monitoring. Physiological monitoring methods require wearing monitoring equipment and having professional knowledge reserves to judge emotion stress, and are not conducive to real-time emotion monitoring scenarios for the school population; in non-physiological monitoring methods, the mental health assessment scale may contain deceptive information; emotion recognition relying on natural language processing technology may have certain difficulties in capturing information such as group audio and text, resulting in possible deficiencies in the timeliness of emotion analysis; while relying on image processing technology, its characteristics of easy group monitoring, key tracking, and real-time image capture have received extensive attention in the field of emotion recognition.

[0004] Currently, the methods for discriminating human emotions based on images still have relatively serious defects. Firstly, the dominant data type is single. Most of the existing methods only consider the representation of emotions by the morphology of a single part of the human body. For example, they only focus on facial expressions and ignore the emotional information conveyed by other parts such as body postures and limb movements. Secondly, the spatio-temporal structure and complementary relationship of the action forms of various parts of the human body are not fully considered. The human body is an organic whole, and the actions and postures of each part are interrelated in different time and space dimensions and jointly express emotions. However, the current methods fail to comprehensively capture and analyze these complex relationships, resulting in a lack of accurate discrimination basis and comprehensive and overall recognition ability when facing the complex psychological process of the emotions of people in the campus scene, leading to misunderstandings or neglect of the emotional states of campus personnel. In the campus environment, if the abnormal emotional states of campus personnel cannot be recognized in time, the best intervention opportunity may be missed, increasing the risk of campus safety management. Summary of the Invention

[0005] The purpose of the present invention is to overcome the problem of inaccurate discrimination of the emotions of people in the campus scene, and proposes an emotion stress recognition method and system based on an action form fusion graph neural network.

[0006] To achieve the above object, the present invention adopts the following technical solutions: In the first aspect, the present invention provides an emotion stress recognition method based on an action form fusion graph neural network, including the following steps: S1. Preprocess the video images of campus personnel, perform live detection through a deep learning algorithm, and generate images of campus personnel; S2. Perform human, hand, and face entity segmentation on the images of campus personnel, and use a feature extraction model to locate the key points of the images based on the dimensions of non-physiological feature data; S3. Construct a dual-graph data space structure, and use a T-GCN model combined with a graph partitioning strategy to extract emotion-enhanced segmentation features; The dual-graph data space structure includes: A graph data space structure with human key points as nodes, and the edge weights are dynamically allocated based on the spatial configuration partitioning strategy according to the distance relationship between the joints and the root node; A graph data space structure with image frames as nodes, and the edge weights are allocated based on the similarity between image feature vectors; S4. Through the transfer learning method, the T-GCN model is first trained using a public emotion dataset, and then secondarily trained and tested using a campus emotion stress small sample dataset; S5. One-hot encode the personal information of the personnel, construct graph data nodes using the image frames, extract non-homogeneous features of the human body, hands, and face based on PCSN, and fuse the key point data with the personal information encoding to complete the comprehensive recognition of emotional stress and obtain the basis for emotional stress discrimination; S6. Pool and fuse the emotion-enhanced segmentation features of the human body, hands, and face with the basis for emotional stress discrimination, and output the recognition result and classification probability of the emotional stress category through the KAN network classifier.

[0007] Furthermore, the T-GCN model includes a GCN layer and a Mamba encoder. The GCN layer is used to capture the spatio-temporal features of human actions, and the Mamba encoder processes the sequence data in parallel through SSMs.

[0008] Furthermore, the PCSN adopts a multi-branch parallel structure to collect image feature information within different scale fields of view, including a one-dimensional global average pooling encoding branch, a 3×3 convolution branch, and a 5×5 convolution branch.

[0009] Furthermore, the one-dimensional global average pooling encoding branch includes performing one-dimensional global average pooling on the feature vector in parallel and vertical directions using parallel routing to obtain a one-dimensional global average pooling feature vector in the horizontal direction and a global average pooling feature vector in the vertical direction; after transposing the one-dimensional global average pooling feature vector in the horizontal direction, it is merged with the global average pooling feature vector in the vertical direction to obtain a global average pooling feature vector; after performing a non-linear transformation on the global average pooling feature vector, an activated feature vector is obtained, and a max pooling operation and an average pooling operation are performed on the activated feature vector; The 3×3 convolution branch includes performing convolution processing on the input feature vector using a convolution kernel of 3 and then performing BN normalization to obtain a normalized feature vector, and performing max pooling and average pooling operations on the normalized feature vector; The 5×5 convolution branch includes performing convolution processing on the input feature vector using a convolution kernel of 5 and then performing BN normalization to obtain a normalized feature vector, and performing max pooling and average pooling operations on the normalized feature vector; The PCSN obtains the parallel local cross-space features through matrix dot product operation on the result vectors of max pooling and average pooling of the three-branch structure, and the parallel local cross-space features are output through the KAN network layer; The KAN network of the KAN network classifier and the KAN network layer of the PCSN both adopt a network structure with 3 grid intervals and 3 orders.

[0010] Furthermore, the dimension of the non-physiological feature data is: [N, C, T, V, M] Among them, N is the number of videos input in the same batch, C is the joint feature, T is the number of image frames in a single video, V is the number of key points, and M is the maximum number of recognized people in a single-frame image; The key points of the campus personnel image include human key points, facial key points, and hand key points; The facial key points include internal key points and external contour key points. The internal key points include the positions and opening / closing states of the eyebrows, eyes, mouth, and nose; The hand key points include hand key points and background data; The human key points include human key points. The human key points include human degrees of freedom joints, and the human degrees of freedom joints include the neck, shoulders, elbows, wrists, waist, knees, and ankles; The campus personnel information includes gender, age, major, and grades; The emotional stress categories include happiness, sadness, fear, surprise, anger, and jealousy.

[0011] Furthermore, the preprocessing uses Instruct-IPT and FSRCNN. The FSRCNN uses a transposed convolutional layer, and the convolutional kernel sizes of the transposed convolutional layer include 9×9, 81×81, and 27×27; The live detection uses a cascade detector, and the width of the target recognition area is 1000 pixels and the height is 1000 pixels; The entity segmentation uses the cascade model of the cascade detector and an image segmentation algorithm; the cascade model of the cascade detector includes a full-body human feature cascade classifier, a hand feature cascade classifier, and a frontal face feature classifier; In the segmentation mask of the image segmentation algorithm, the background area is marked as 0, and the foreground area is marked as 255. The cascade detector samples the Haar cascade detector in the open-source library OpenCV, and the image segmentation algorithm uses the GrabCut algorithm; The width of the minimum recognition area of the full-body human feature cascade classifier is 1000 pixels and the height is 1000 pixels. The width of the minimum recognition area of the hand feature cascade classifier is 500 pixels and the height is 500 pixels. The width of the minimum recognition area of the frontal face feature classifier is 500 pixels and the height is 500 pixels.

[0012] In a second aspect, the present invention provides an emotional stress recognition system based on an action form fusion graph neural network, including: A live detection module, configured to preprocess the video images of campus personnel, perform live detection through a deep learning algorithm, and generate images of campus personnel; The key point positioning module is used to perform human body, hand and face entity segmentation on the images of school personnel, and use the feature extraction model to locate the key points of the images based on the dimensions of non-physiological feature data; The dual-graph data construction module is used to construct the spatial structure of dual-graph data, and adopt the T-GCN model combined with the graph partitioning strategy to extract the segmentation features with enhanced emotions; the spatial structure of the dual-graph data includes: the graph data spatial structure with human key points as nodes, and the edge weights are dynamically allocated based on the spatial configuration partitioning strategy according to the distance relationship between the joints and the root node; the graph data spatial structure with image frames as nodes, and the edge weights are allocated based on the similarity between the image feature vectors; The T-GCN model training module is used to, through the transfer learning method, first train the T-GCN model using the public emotion dataset, and then perform secondary training and testing on the small sample dataset of campus emotion stress; The non-homogeneous feature extraction module is used to perform one-hot encoding on the personal information of the personnel, use the image frames to construct graph data nodes, extract the non-homogeneous features of the human body, hands and face based on PCSN, and fuse the key point data with the personal information encoding to complete the comprehensive recognition of emotion stress and obtain the basis for emotion stress discrimination; The classified emotion recognition result module is used to pool and fuse the emotion-enhanced segmentation features of the human body, hands and face with the basis for emotion stress discrimination, and output the emotion stress category recognition result and classification probability through the KAN network classifier.

[0013] Furthermore, it further includes: The perception control module is used to capture the image information of school personnel and the time information attached to the image information of school personnel; The user module participates in system management, use and data sharing, as well as the interaction between users; The service providing module is used to interact through the service interface. The user module includes personal information management services, image data management services, mental health information management services, and emotion stress warning and push services; The resource exchange module is used to perform data and service interactions with the external environment for resource exchange; The operation and maintenance control module is used to monitor the system operation and maintenance services and the operation and maintenance of equipment and network resources.

[0014] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the emotion stress recognition method based on the action form fusion graph neural network as described above is implemented.

[0015] Fourthly, the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the emotion stress recognition method based on the action morphology fusion graph neural network as described above.

[0016] Compared with the prior art, the present invention has the following beneficial technical effects: The emotion stress recognition method based on the action morphology fusion graph neural network proposed by the present invention aggregates human body morphological actions such as facial actions, human postures, and gesture actions, and emotion stress categories to respectively construct graph data with human key points and image frames as graph nodes as the judgment method. The graph data constructed by the human key points uses the graph partitioning strategy and T-GCN processing to strengthen the cohesion and coupling ability of emotion category representation. The graph data constructed by the image frames meets the non-homogeneous feature extraction requirements of the human body, hands, and face through the PCSN (Parallel Cross Spatial Multi Scale Feature Extraction Network), and at the same time establishes the long-range and short-range dependencies of global feature information and local cross-channel interaction relationships. This method not only constructs the internal association of the multi-level airspace cross-fusion of human body morphological actions, but also improves the influence of the temporal changes of multi-level airspace morphological action features on emotion stress, thereby enriching the basis for emotion stress discrimination and improving the accuracy and reliability of the results. For the first time, multi-dimensional joint reasoning of human postures, local limb actions, and micro-expressions is realized, and the requirement for the amount of model training data is reduced through the small sample transfer learning mechanism, making it more feasible to implement in the campus scenario. The system combines transfer learning and small sample learning, initializes the network using a public dataset, and then performs secondary training on the campus emotion stress small sample dataset, significantly improving the recognition accuracy.

[0017] Furthermore, the method of the present invention breaks through the limitations of traditional emotion monitoring based on single physiological characteristics, comprehensively integrates body, facial, and gesture information, and effectively models the spatio-temporal structure of action morphology. By combining small sample learning and transfer learning, the application efficiency of the system in the campus scenario is significantly improved, with strong real-time monitoring ability and high recognition accuracy, especially having broad practical application prospects in the fields of education, health, and safety management. The present invention provides effective technical support for emotion stress monitoring and intervention, especially showing unique advantages in group real-time monitoring and emotion early warning.

[0018] The emotional stress recognition system based on the action form fusion graph neural network proposed by the present invention improves the ability of coordinated sharing of data and business value; uses image perception and acquisition devices such as campus and classroom cameras to achieve group perception, key tracking, and real-time capture of the morphological action image information of school personnel, and integrates personal information. The combination of MGCN and PCSN realizes the ability to perceive and analyze the emotional stress of personnel in an associative and memory-based manner with the action form as a reference, improving the credibility of the measurement and analysis results of emotional stress; and can recommend activities such as personal exercise, education, diet, sleep, and environmental selection according to the emotional stress results to adjust and maintain positive energy emotions. In particular, for personnel with abnormal emotional stress, early intervention and guidance can be carried out and retrospective monitoring can be strengthened. By real-time monitoring the emotional stress of school personnel, it helps to accelerate the recovery of abnormal personnel and reduce the local impact. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings described herein are for illustrative purposes only and are not intended to limit the scope of the disclosure of the present invention in any way. Additionally, the shapes and proportional dimensions of the various components in the drawings are only schematic and are used to assist in understanding the present invention, rather than specifically defining the shapes and proportional dimensions of the various components of the present invention. In the drawings: Figure 1 is a flowchart of the emotional stress recognition method based on the action form fusion graph neural network of the present invention.

[0020] Figure 2 is a structural diagram of the emotional stress recognition system based on the action form fusion graph neural network of the present invention.

[0021] Figure 3 is an electronic device diagram of the emotional stress recognition method based on the action form fusion graph neural network of the present invention.

[0022] Figure 4 is a network framework diagram of the emotional stress recognition model based on the action form feature fusion graph convolutional network in the embodiment.

[0023] Figure 5 is a schematic diagram of the parallel cross-space multi-scale feature extraction network.

[0024] Figure 6 is a schematic diagram of the spatio-temporal graph convolutional neural network model.

[0025] Figure 7 is the campus emotional Internet of Things monitoring system architecture of the emotional recognition model network based on the action form feature fusion graph convolutional network in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0027] Embodiment 1 Refer to Figure 1 , a method for identifying emotional stress based on an action form fusion graph neural network, comprising the following steps: Preprocess the video images of campus personnel, perform live detection through a deep learning algorithm to generate images of on-campus personnel; perform human body, hand, and face entity segmentation on the images of on-campus personnel, and use a feature extraction model to locate the key points of the images based on the dimensions of non-physiological feature data; construct a dual graph data space structure, and use the T-GCN model combined with a graph partitioning strategy to extract emotion-enhanced segmentation features; the dual graph data space structure includes: a graph data space structure with human key points as nodes, and the edge weights are dynamically allocated based on the spatial configuration partitioning strategy according to the distance relationship between the joints and the root node; a graph data space structure with image frames as nodes, and the edge weights are allocated based on the similarity between image feature vectors; through transfer learning, the T-GCN model is first trained using a public emotion dataset, and then secondarily trained and tested on a campus emotion stress small sample dataset; perform one-hot encoding on the personal information of the personnel, use the image frames to construct graph data nodes, extract non-homogeneous features of the human body, hands, and face based on PCSN, and fuse the key point data with the personal information encoding to complete the comprehensive identification of emotional stress and obtain the basis for emotional stress discrimination; pool and fuse the emotion-enhanced segmentation features of the human body, hands, and face with the basis for emotional stress discrimination, and output the emotion stress category recognition result and classification probability through the KAN network classifier.

[0028] In this embodiment, the graph data used includes key point graph data and image frame graph data as the data basis for emotion recognition. Among them, the image frame data, as global information features, includes personal overall features, interaction features with personnel, etc., and the key point graph data is local features including facial features, human body forms, and subtle hand movement features, constructing a network structure model for graph data processing with multi-layer perception, improving the accuracy of emotion representation and the basis for discrimination reference.

[0029] Embodiment 2 Refer to Figure 2 , a system for identifying emotional stress based on an action form fusion graph neural network, comprising: The in vivo detection module is used to preprocess the video images of campus personnel, perform in vivo detection through deep learning algorithms, and generate images of on-campus personnel; The key point localization module is used to perform human body, hand, and face entity segmentation on the on-campus personnel images, and use a feature extraction model to locate the key points of the images based on the dimensions of non-physiological feature data; The dual-graph data construction module is used to construct a dual-graph data spatial structure, and use the T-GCN model combined with a graph partitioning strategy to extract emotion-enhanced segmentation features; the dual-graph data spatial structure includes: a graph data spatial structure with human key points as nodes, and the edge weights are dynamically allocated based on the spatial configuration partitioning strategy according to the distance relationship between joints and the root node; a graph data spatial structure with image frames as nodes, and the edge weights are allocated based on the similarity between image feature vectors; The T-GCN model training module is used to, through transfer learning, first train the T-GCN model using a public emotion dataset, and then perform secondary training and testing using a small sample dataset of campus emotion stress; The non-homogeneous feature extraction module is used to perform one-hot encoding on the personal information of personnel, use image frames to construct graph data nodes, extract non-homogeneous features of the human body, hands, and face based on PCSN, and fuse key point data with personal information encoding to complete the comprehensive recognition of emotion stress and obtain the basis for emotion stress discrimination; The classified emotion recognition result module is used to pool and fuse the emotion-enhanced segmentation features of the human body, hands, and face with the basis for emotion stress discrimination, and output the emotion stress category recognition result and classification probability through a KAN network classifier.

[0030] See Figure 7 The perception control module is used to capture on-campus personnel image information through high-definition cameras on campus, in classrooms, and at gates, and the on-campus personnel image information is attached with time information; The user module is used to associate involved institutions and personnel, participate in system management, use, and data sharing, as well as interactions between users. The involved institutions and personnel include teachers, student management offices, mental health rooms, as well as teachers and students individually; The service providing module is used to interact through service interfaces. The user module includes personal information management services, image data management services, mental health information management services, as well as emotion stress warning and push services; The resource exchange module is used to perform data and service interactions with the external environment for resource exchange. The interfaces for resource exchange include mental health guidance interfaces, psychosocial research interfaces, online education interfaces, and game entertainment business interfaces; The operation and maintenance control module is used to monitor system operation and maintenance services and the operation and maintenance of equipment and network resources.

[0031] The emotional stress recognition system is established based on the six-domain model, which improves the ability to coordinate and share data and business value. The six-domain model is an important concept in the field of the Internet of Things. By systematically combing the application-related elements of the Internet of Things industry, six major domains are set to build a collaborative ecosystem for the Internet of Things. Using campus and classroom cameras and other image perception and acquisition equipment to achieve group perception, key tracking and real-time capture of the image information of the shape and action of people in the school and integrate personal information, a spatiotemporal graph convolutional network algorithm model is constructed, which realizes the ability to perceive and analyze the emotional stress of people in an associative and memory-based manner with the action form as a reference, and improves the credibility of the results of emotional stress measurement and analysis; and based on the emotional stress results, personal sports, education, diet, sleep and environmental selection activities can be recommended to adjust and maintain positive emotions, especially for people with abnormal emotional stress, early intervention guidance and enhanced traceability monitoring.

[0032] Embodiment 3 See also Figure 3 An electronic device comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for identifying emotional stress based on an action morphology fusion graph neural network is implemented.

[0033] Embodiment 4 A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, an emotional stress recognition method based on an action morphology fusion graph neural network is described.

[0034] Embodiment 5 This embodiment provides an emotional stress recognition method based on action morphology fusion graph neural network, see Figure 1 , the specific workflow is as follows: The first step is image capture and data cleaning. Use campus and classroom cameras to capture personnel image data. Apply Instruct-IPT (Instruct-Image Processing Transformer, an instruction image processing transformer model) to the personnel image data captured by campus and classroom cameras for denoising, deblurring, de-raining, de-fogging, and de-snowing. Instruct-IPT is an image processing technology that solves various image restoration tasks, such as denoising, deblurring, de-raining, de-fogging, and de-snowing, through weight modulation. And use the FSRCNN (Fast Super-Resolution Convolutional Neural Network) with 9×9, 81×81, and 27×27 transposed convolutional layers to upsample the video images to generate high-resolution images. Due to problems such as no target, blurred target, too small or overly obscured, too large angle offset, and overly dim night scene shooting during image capture, the information content is less during the model training process, unable to update the model parameters and have a significant impact, or due to less information, the model over-focuses on such features, resulting in overfitting problems. The Haar cascade detector in OpenCV (Open Source Computer Vision Library) can be used for live detection, and set the target recognition area size to (1000, 1000) to achieve an automated data cleaning process. (1000, 1000) means the width is 1000 pixels and the height is 1000 pixels.

[0035] Step 2, Instance segmentation of human body parts. Use the Haar cascade detector and GrabCut algorithm (an image segmentation algorithm) in OpenCV to perform entity segmentation on the face, body, and hand in the image. Select the cascade models: haarcascade_fullbody.xml (a file containing a Haar feature cascade classifier for detecting the whole body in an image), haarcascade_hand.xml (a file containing a Haar feature cascade classifier for detecting hands), and haarcascade_frontalface_default.xml (a file containing a Haar feature cascade classifier for detecting frontal faces). The minimum recognition area pixels in the detectMultiScale (an object detection function) of the three are (1000, 1000), (500, 500), and (500, 500). The background area in the segmentation mask is marked as 0, and the foreground area is marked as 255. (1000, 1000) means the width is 1000 pixels and the height is 1000 pixels, and (500, 500) means the width is 500 pixels and the height is 500 pixels.

[0036] Step 3: Key point localization and recognition. During the process of human emotion change, the spatio-temporal changes of non-physiological features are an important manifestation of emotions, such as facial expressions, body postures, and hand movements. One way to identify the categories corresponding to non-physiological features is to construct the key point localization and recognition of body parts. At the same time, the key points in this example are also used as nodes in the graph data for emotion stress recognition. In this example, Heatmaps are used to construct key points. For facial key point detection, shape_predictor_68_face_landmarks.dat (a facial feature point detection model) in the OpenCV+Dlib library (Dlib is a machine learning algorithm and tool library, famous for its efficient and accurate face detection and 68-point feature point detector) is used for 68-point calibration. 51 internal key points can represent the positions and opening / closing states of eyebrows, eyes, mouth, and nose, and 17 external contour key points; for body and hand key point detection, body_pose_deploy.prototxt (a definition file for extracting features from input images and estimating the joint positions of the human body) and body_pose_model.pth (a definition file for the neural network weights for estimating human joint positions), hand_pose_deploy.prototxt (a neural network architecture definition file for gesture pose recognition), and hand_pose_model.pth (a parameter value file obtained after training the neural network for gesture pose recognition) models in OpenPose (a deep learning-based human pose estimation framework) are used to extract key points. There are 22 in total for the hand, including 21 hand key nodes and 22 points as background data; there are 18 key points for the human body, corresponding to the degrees of freedom joints of the human body, such as the neck, shoulders, elbows, wrists, waist, knees, and ankles.

[0037] The dimension of non-physiological feature data is [N, C, T, V, M], where N is the number of videos input in a batch; C is the joint feature, that is, the horizontal and vertical axis positions of the key points and the ACC (Accuracy) confidence; T is the number of image frames in a single video; V is the number of key points; M is the maximum number of people recognized in a single-frame image.

[0038] Step 4: Establishment of the graph data spatial structure. This example constructs two types of graph data respectively for Figure 4For parts 1 and 2, one type uses human key points as nodes, and the edges are the connections between the nodes. The edge weight assignment strategy adopts spatial configuration partitioning, that is, the distance between the root node and the center of gravity of different parts of the human body is used as the reference distance, and then the relationship between the distance between other adjacent nodes and the root node and the reference distance is judged. When it is greater than the reference distance, its weight is calibrated to 3, when it is less than the reference distance, the weight is 0, and when it is equal to the reference, the weight is 1, giving the idea that the farther the joint is from the center of gravity, the greater its movement, and the greater its impact on the behavioral characteristics, so more considerations should be given; the second type uses image frames as graph nodes, and the edges are the connections between different frames, and its weight is the similarity between image feature vectors.

[0039] Figure 4 In it, time dimension represents the time dimension, spatial dimension represents the spatial dimension, fullbody is the human body posture, hand is the hand, face is the face, and POOL is pooling.

[0040] Step 5: Construct a small sample dataset of emotional stress. The possible actions corresponding to the human body forms under different emotions can be referred to the following table. Use the target detection annotation tool LabelImg (an open-source image annotation tool) to annotate the face, human body, hand joints and emotional categories and associate them with personal information.

[0041]

[0042] Step 6: Personal information acquisition and encoding. In campus mental health statistics, it shows that emotional stress is often highly correlated with gender, age, major, and academic or teaching and research achievements. According to mental health statistics, men have greater emotional fluctuations during extreme events and are more inclined to express emotions through body movements. When angry or excited, men may show more obvious body languages, such as clenching fists and stamping feet.

[0043] In this example, the personnel information is obtained through OpenCV face recognition linkage, and the information including gender, age, major, grades, etc. is one-hot encoded into a one-dimensional feature vector, with an age interval of 2; the average academic (research) performance is divided into three categories: fail, pass, and excellent; For example, [Zhang San, male, 20, software engineering, average score of 82] corresponds to the encoded information [0101000101], From left to right, the 0-1st bit corresponds to gender, the 2-4th bit corresponds to age, the 5-7th bit corresponds to major, and the 8-9th bit corresponds to academic performance.

[0044] Step 7: Normalization of human key-point features. In the time and space dimensions, the relative positions of key points in images under actions in different emotional states vary greatly. According to statistical principles, the positions of key points in different batches and image frames follow the random distribution characteristics in the spatial distribution pattern. By performing BN (Batch Normalization) normalization on the key-point features, it will not cause a significant fluctuation in the prediction accuracy, but can accelerate the convergence of the model.

[0045] Step 8: Use the KAN network to construct the PCSN to meet the requirements for extracting non-homogeneous features of the face, human body, and hand.

[0046] 8.1, as Figure 5 shown, perform one-dimensional global average pooling on the input feature vector X[C, H, W] in the parallel and vertical directions. Here, C is the number of channels, H and W are the spatial dimensions of the input, H is the height, and W is the width. The encoding expressions for one-dimensional global average pooling in the horizontal and vertical directions are:

[0047] M is the size of the output vector after one-dimensional average pooling of the X feature vector. M(H) is the output vector calculated in the vertical direction. is an entirety, which is to take the vector data of the i-th row in the horizontal direction, and its range is from 0 to W, that is, to scan the entire width of the picture.

[0048]

[0049] M(W) is the output vector calculated in the vertical direction. is to take the vector data of the i-th column in the vertical direction, and its range is from 0 to H, that is, to scan the entire height of the picture.

[0050] One-dimensional global average pooling in the horizontal and vertical directions is the output feature vector after one-dimensional global processing of the image, mainly to obtain the global shallow features of the image. In order to fuse with the later deep semantic features and more accurately represent emotional stress.

[0051] Figure 5 In, Input represents the input, X Avg Pool represents one-dimensional global average pooling in the parallel direction, Y Avg Pool represents one-dimensional global average pooling in the vertical direction, Concat represents concatenation, Max Pool represents max pooling, Avg Pool represents average pooling, matmul represents the matmul function (a function that performs matrix multiplication), KANConv represents the convolution of the KAN network, BN represents batch normalization, and Output represents the output; 8.2 After transposing the one-dimensional global average pooling encoded vector in the horizontal direction and merging it with the global average pooling vector in the vertical direction, we get , , where T is the vector transpose. The one-dimensional pooling shares the convolutional kernel parameters, thus correlating local cross-influences and global similarities; 8.3 Use the ELU non-saturating activation function (Exponential Linear Unit is a non-saturating activation function) to process the global average pooling encoded vector and the feature vector X, accurately preserving the spatial structure information of the input vector in the channels and preventing it from falling into dead neurons, avoiding the problems of gradient explosion or vanishing.

[0052] 8.4 Perform max pooling and average pooling on the above vectors respectively, retaining the dominant features and local features.

[0053] 8.5 Process the input feature vector X with convolutional kernels of 3 and 5 respectively, then perform BN normalization and the same max pooling and average pooling operations as above, and use ELU as the activation function for feature merging processing, and form a three-branch structure with the one-dimensional global average pooling encoding to collect image feature information within different scale fields of view.

[0054] 8.6 Use the matmul function for matrix dot product operation on the result vectors of max pooling and average pooling of the three-branch structure to achieve parallel local cross-spatial feature extraction, and establish the long-range and short-range dependencies of global feature information and local cross-channel interaction relationships.

[0055] 8.7 Use a KAN network with 3 grid intervals and 3 orders to improve the network's non-linear fitting as a PCSN network.

[0056] In the ninth step, PCSN is used to process different entity segmentation objects respectively. Since the face, body, and hand have different biological structures and require different attention-based feature extraction network structures, body feature extraction mainly involves joint position and body pose key point detection, and a large-area feature extraction network is needed to accurately mark the positions of each key point; hand feature extraction focuses on finger pose, palm pose, joint completeness, etc. At the same time, due to the small size and high flexibility of the hand relative to the body size, a fine-grained feature extraction network is required; especially for face feature extraction, features such as facial contours, eyes, nose, mouth, and their mutual relationships can best represent the emotional characteristics of a person. Compared with the body and hand, the face has rich texture feature information, so more attention is paid to the extraction of feature vectors at different levels. During the model training process, PCSN is randomly initialized with parameters to expand its feature extraction ability for the face, body, and hand. The internal association of multi-level spatial domain cross-fusion in PCSN, combined with the mechanism of the KAN network (Kolmogorov - Arnold network, a neural network architecture), enhances the nonlinear fitting ability of the model and improves the model accuracy.

[0057] In the tenth step: Spatiotemporal graph convolutional neural network model. Mamba, through SSMs (Selective State Space Models), can not only process time-domain sequence data in parallel, but also the structured state space model has fewer parameters compared to the Transformer (a deep learning model based on the attention mechanism). At the same time, it has a hardware-aware algorithm optimization process, which is beneficial to further improving the real-time monitoring of the system, optimizing the network structure, and enhancing the real-time response of the system. A spatiotemporal graph convolutional neural network is constructed based on Mamba (a simplified state space model architecture) and GCN (Graph Convolutional Network). Considering the spatiotemporal structure and complementary relationship of morphological actions, and the hardware-aware computational parameter tuning mechanism of Mamba, it can improve the real-time response of the system and achieve real-time monitoring. By leveraging the complementary advantages of the MGCN (Mamba Graph Convolutional Network, spatiotemporal graph convolutional neural network model) and the GCN model, that is, the spatiotemporal processing paradigm of data, the understanding of emotion images and videos is realized. As Figure 6 shown, the model input X is a video of human morphological actions with T frames, Figure 6 where X0, X T-1 , X T are the 0th frame, the (T - 1)th frame, and the Tth frame of the image respectively. GCN is used to capture the spatial feature relationships between nodes in the two types of graph data constructed, while Mamba can capture the global feature relationships in the time dimension over long distances; Figure 6The Mamba encoder in Chinese represents the Mamba encoder. By taking the human body morphological action feature vector output by the GCN in the time domain as the input of the Mamba encoder, the model is enabled to simultaneously possess the ability to learn spatio-temporal dimension information. Finally, the output feature vector Y of the MGCN model has the comprehensive expression ability for emotion stress recognition.

[0058] In the MGCN model of this example, the number of GCN layers is 4. By increasing the number of layers, multi-hop node features can be aggregated to obtain higher-level semantic information and improve the model recognition accuracy.

[0059] The eleventh step is emotion stress classification. As Figure 4 shown, finally, the processing results of the two types of graph data are fused to achieve the deep semantic representation of multi-level spatial and temporal morphological action features. After being processed by pooling and the KAN network classifier, the emotion stress category and classification probability are output, where the KAN network classifier is a network structure using 3 grid intervals and 3 orders. By invoking the KAN network, a dynamic activation function is used in the classifier, enhancing the model's ability to extract non-linear features and further improving the model discrimination accuracy.

[0060] The twelfth step is model training.

[0061] 1) The L-Softmax loss (large margin Softmax loss function) is adopted as the loss function in the emotion recognition training process. It is more compact within the same-class feature space, enhancing the significance between emotion stress categories. The L-Softmax loss increases the classification difficulty by adding an angular margin to the Softmax (normalized exponential function), thereby improving the model's generalization ability.

[0062] 2) Transfer learning is adopted. The publicly available datasets Fer2013 (Facial Expression Recognition Challenge 2013, facial emotion dataset), E-Gait (body pose emotion dataset), and SMG (micro-gesture dataset) are used to train only Figure 4 the part of the MGCN network model in the dashed box 1. At this time, the network in the dashed box 2 is disconnected. Figure 4 The hyperparameter settings include the learning rate set to 0.0001, the batch size set to 100, the number of iterations to 10000, and the optimizer selected as Adam (Adaptive Moment Estimation). The Adam optimizer is an optimization algorithm based on the gradient descent method, combining the momentum method and the method of adaptive learning rate, and is widely used in the training of deep learning models.

[0063] 3) Load the model parameters obtained from the training in step 2 into Figure 4In the virtual box 1 network structure of the model network structure, the small sample data of the emotional stress in this instance is divided into 7:2:1 as the training set, validation set and test set. After training, the model is saved for system testing and deployment.

[0064] Step 8, system construction. As Figure 7 shown, the six-domain model of the Internet of Things is used to construct an emotional stress monitoring system. This model constructs the system based on resource and service characteristics, effectively avoiding the repeated deployment of software and hardware devices, systems, etc., and improving the coordinated sharing ability of data and business value.

[0065] The system in this instance includes a perception control module, a user module, a service provision module, a resource exchange module, and an operation and maintenance management and control module.

[0066] The perception control module, as a collection of device entities, transmits and interacts the information of real-world entity resources with the network environment. In this instance, it includes capturing the image information of people on campus through high-definition cameras in the campus, classrooms, and gates, along with the time information; the service provision module is a collection of system applications and services, associated with other domains for control and interaction through service interfaces. In this instance, it includes personal information management services, image data management services, mental health information management services, and emotional stress warning and push services; the user domain is associated with the institutions and personnel involved in this system, participating in system management, use, data sharing, and interactions between users. In this instance, it includes teachers, student management offices, mental health rooms, and individual teachers and students; the resource exchange domain is for the system to interact with the external environment in terms of data and services, establish business relationships, and promote the improvement of each other's information and service capabilities. In this instance, it includes resource exchange through business interfaces such as providing mental health guidance, psychosocial research, online education, and game entertainment; the operation and maintenance management and control module, as a collection of entities to ensure the normal operation of the overall system, covers the operation and maintenance services of the stress monitoring system and the operation and maintenance of equipment and network resources in this instance.

[0067] In this instance, the interaction relationships among the modules are as follows: After the perception control module collects data through devices with image shooting functions such as high-definition cameras, mobile phones, and computers, it uploads the data to the server where the service provision module is deployed for image storage, processing, and recognition, and can push the results to the campus safety supervision, mental health room, or specific personnel. The mental health room can also monitor the status of key personnel in real time through the emotional stress monitoring service for monitoring and early warning or tracking the status after intervention and treatment; on the one hand, the resource exchange domain provides services to the outside through business interfaces such as mental health guidance, psychosocial research, online education, and game entertainment, and on the other hand, obtains external service experience and data, and uses the reinforcement learning mechanism to improve the accuracy of the models and systems in the service provision domain and iteratively optimize the service quality; the operation and maintenance management and control domain monitors the software, hardware, and network devices in the system through information interaction with the perception control domain and the service provision domain to ensure the normal operation of the entire system.

[0068] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0069] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0070] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implement the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0071] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks. Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: still can modify the specific implementation manners of the present invention or make equivalent replacements, and any modification or equivalent replacement without departing from the spirit and scope of the present invention shall be covered by the protection scope of the present invention.

Claims

1. A method for identifying emotional stress based on action morphology fusion graph neural network, characterized in that: The following steps are involved: S1. Preprocess the video images of campus personnel, perform liveness detection through deep learning algorithms, and generate images of campus personnel; S2. Perform body, hand and facial entity segmentation on the images of people in the school, and use the feature extraction model to locate the key points of the image based on the dimensions of non-physiological feature data; S3, construct a dual graph data space structure, and use the T-GCN model combined with the graph partitioning strategy to extract the segmentation features of emotion reinforcement; The dual graph data space structure comprises: The spatial structure of graph data with human key points as nodes, and the edge weights are dynamically allocated based on the distance relationship between joints and root nodes based on the spatial configuration partition strategy; A graph data space structure with image frames as nodes, and edge weights are assigned based on the similarity between image feature vectors; S4. Through the transfer learning method, the T-GCN model is first trained using the public emotion dataset, and then trained and tested again using the campus emotion stress small sample dataset; S5. Perform unique-hot encoding on the personal information of the personnel, use the image frame to construct the graph data node, extract the non-homogeneous features of the human body, hands, and face based on PCSN, and integrate the key point data with the personal information coding to complete the comprehensive identification of emotional stress and obtain the basis for emotional stress discrimination; S6. The emotion enhancement segmentation features of the human body, hands, and face are pooled and fused with the basis for emotional stress discrimination, and the emotional stress category recognition results and classification probabilities are output through the KAN network classifier.

2. According to the method for emotional stress recognition based on action morphology fusion graph neural network in claim 1, the T-GCN model includes a GCN layer and a Mamba encoder, the GCN layer is used to capture the spatiotemporal characteristics of human body movements, and the Mamba encoder processes sequence data in parallel through SSMs.

3. According to the method for emotional stress recognition based on action morphology fusion graph neural network in claim 1, the PCSN adopts a multi-branch parallel structure to collect image feature information in different scales of field of view, including a one-dimensional global average pooling encoding branch, a 3×3 convolution branch and a 5×5 convolution branch.

4. The emotional stress recognition method based on action morphology fusion graph neural network according to claim 3 is characterized in that: The one-dimensional global average pooling encoding branch includes using parallel routing to perform one-dimensional global average pooling on the feature vector in parallel and vertical directions to obtain a one-dimensional global average pooling feature vector in the horizontal direction and a global average pooling feature vector in the vertical direction; after transposing the one-dimensional global average pooling feature vector in the horizontal direction, merging it with the global average pooling feature vector in the vertical direction to obtain a global average pooling feature vector; performing a nonlinear transformation on the global average pooling feature vector to obtain an activated feature vector, and performing a maximum pooling operation and an average pooling operation on the activated feature vector; The 3×3 convolution branch includes performing convolution processing with a convolution kernel number of 3 on the input feature vector and then performing BN normalization to obtain a normalized feature vector, and performing maximum pooling and average pooling operations on the normalized feature vector; The 5×5 convolution branch includes performing convolution processing on the input feature vector with a convolution kernel number of 5 and then performing BN normalization to obtain a normalized feature vector, and performing maximum pooling and average pooling operations on the normalized feature vector; The PCSN obtains parallel local cross-space features by performing matrix dot product operations on the result vectors of the maximum pooling and average pooling of the three-branch structure, constructs the long- and short-range dependencies of the global feature information and the local cross-channel interaction relationship, and the parallel local cross-space features are output through the KAN network layer; The KAN network classifier and the KAN network of the KAN network layer of PCSN both adopt a network structure with 3 grid intervals and 3 orders.

5. The emotional stress recognition method based on action morphology fusion graph neural network according to claim 1 is characterized in that: The non-physiological feature data dimensions are: [N, C, T, V, M] Among them, N is the number of videos input in the same batch, C is the joint feature, T is the number of image frames in a single video, V is the number of key points, and M is the maximum number of people that can be recognized in a single frame image; The key points of the campus personnel image include human body key points, facial key points and hand key points; The facial key points include internal key points and external contour key points, and the internal key points include the position and opening and closing state of eyebrows, eyes, mouth and nose; The hand key points include hand key points and background data; The human body key points include human body key points, human body key points include human body degree of freedom joints, and human body degree of freedom joints include neck, shoulder, elbow, wrist, waist, knee and ankle; The campus personnel information includes gender, age, major and grades; The emotional stress categories include happiness, sadness, fear, surprise, anger, and jealousy.

6. The emotional stress recognition method based on action morphology fusion graph neural network according to claim 1 is characterized in that: The preprocessing adopts Instruct-IPT and FSRCNN, the FSRCNN adopts a transposed convolution layer, and the convolution kernel sizes of the transposed convolution layer include 9×9, 81×81 and 27×27; The liveness detection adopts a cascade detector, and the width and height of the target recognition area are 1000 pixels and 1000 pixels respectively; The entity segmentation adopts a cascade model of a cascade detector and an image segmentation algorithm; the cascade model of the cascade detector includes a whole body human feature cascade classifier, a hand feature cascade classifier and a frontal face feature classifier; In the segmentation mask of the image segmentation algorithm, the background area is marked as 0, and the foreground area is marked as 255. The cascade detector samples the Haar cascade detector in the open source library OpenCV, and the image segmentation algorithm adopts the GrabCut algorithm; The minimum recognition area of ​​the whole body feature cascade classifier has a width of 1000 pixels and a height of 1000 pixels, the minimum recognition area of ​​the hand feature cascade classifier has a width of 500 pixels and a height of 500 pixels, and the minimum recognition area of ​​the frontal face feature classifier has a width of 500 pixels and a height of 500 pixels.

7. An emotional stress recognition system based on action morphology fusion graph neural network, characterized in that: include: The liveness detection module is used to pre-process the video images of campus personnel, perform liveness detection through deep learning algorithms, and generate images of campus personnel; The key point positioning module is used to segment the human body, hands and faces of the images of people in the school, and locate the key points of the image based on the dimensions of non-physiological feature data using the feature extraction model; A dual graph data construction module is used to construct a dual graph data space structure, and use the T-GCN model combined with a graph partitioning strategy to extract segmentation features of emotion reinforcement; the dual graph data space structure includes: a graph data space structure with human key points as nodes, and edge weights are dynamically allocated based on the distance relationship between joints and root nodes based on a spatial configuration partitioning strategy; a graph data space structure with image frames as nodes, and edge weights are allocated based on the similarity between image feature vectors; The T-GCN model training module is used to train the T-GCN model using a public emotion dataset through transfer learning, and then conduct secondary training and testing on a small sample dataset of campus emotion stress; The non-homogeneous feature extraction module is used to perform unique-hot encoding on the personal information of personnel, construct graph data nodes using image frames, extract non-homogeneous features of the human body, hands, and face based on PCSN, and integrate key point data with personal information coding to complete the comprehensive identification of emotional stress and obtain the basis for emotional stress discrimination; The classification emotion recognition result module is used to pool and fuse the emotion enhancement segmentation features of the human body, hands, and face with the basis for emotional stress discrimination, and output the emotional stress category recognition results and classification probability through the KAN network classifier.

8. The emotional stress recognition system based on action morphology fusion graph neural network according to claim 7 is characterized in that: Also includes: A perception control module, used to capture image information of people in the school, and the image information of people in the school with time information; User module, involved in system management, usage and data sharing, and interaction between users; The service provision module is used to interact through the service interface. The user module includes personal information management services, image data management services, mental health information management services, and emotional stress warning and push services; Resource exchange module, used to interact with the external environment for data and service exchange; The operation and maintenance control module is used to monitor system operation and maintenance services and equipment and network resource operation and maintenance.

9. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for identifying emotional stress based on an action morphology fusion graph neural network as described in any one of claims 1 to 6 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for identifying emotional stress based on an action morphology fusion graph neural network described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Behavior abnormality crowd detection method based on gestures and facial expressions

    CN112084922A

  • Multi-modal emotion recognition method based on graph convolutional network

    CN116229225A

  • Driver smoking detection algorithm based on ST-GCN and YOLOv5n

    CN117746399A

  • Emotion recognition method and system based on multi-modal feature and hierarchical feature fusion

    CN118709094A

  • Emotion recognition method and system based on multi-branch fusion attention mechanism

    WO2025065808A1

Cited By

  • Strip mine transportation road segmentation method and device, electronic equipment and storage medium

    CN120543864A

  • Emotion detection method based on spatial-temporal feature fusion of facial key points

    CN120877354A

  • An emotion detection method based on facial key point space-time feature fusion

    CN120877354B

  • Human body language emotion recognition method and device, equipment and medium

    CN120932290A

  • Multi-modal image fusion method and system based on three-domain collaborative nonlinear reconstruction

    CN122415354A