Facial emotion analysis method based on spatial-temporal feature network construction
By constructing a spatiotemporal feature network of facial expressions and identifying dynamic key points, the lack of dynamic correlation modeling of facial expression analysis in the prior art is solved, and more accurate emotion recognition and interpretability assessment are achieved.
Patent Information
- Application Number
- CN202510773390.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-18
AI Technical Summary
Existing facial expression analysis technology is difficult to fully capture the complex spatial and temporal dynamic characteristics of expressions, and lacks dynamic correlation modeling and dynamic key point recognition capabilities, resulting in insufficient accuracy and interpretability of emotional recognition.
A facial emotion analysis method based on spatiotemporal feature network is constructed, and a dynamic facial feature network is established through video frame segmentation and feature extraction, a dynamic facial feature is identified, and a dynamic network feature and dynamic graph neural network are used for emotional calculation.
It achieves a more accurate and interpretable assessment of facial expressions, can have a deeper understanding of the generation mechanism and emotional meaning of facial expressions, and improves the accuracy and reliability of emotional recognition.
Smart Images

Figure CN120340098A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the cross - field of computer vision and mental health assessment, and specifically relates to the innovative application of affective computing, machine learning, and complex network theory in facial expression analysis. Background Art
[0002] Current facial expression analysis technology has experienced significant development. Early research mainly relied on hand - designed local feature descriptors, such as Local Binary Patterns (LBP), Histogram of Oriented Gradients (HOG), etc., combined with traditional machine learning algorithms such as Support Vector Machines (SVM) and Bayesian classifiers for classification. Such methods performed moderately well in controlled environments, but their performance often dropped significantly when facing real - world challenges such as illumination changes, pose diversity, and facial occlusion, and usually ignored the fact that facial expression is a dynamic and continuous process. In recent years, the rise of deep learning technology has greatly promoted the development of this field. Methods based on Convolutional Neural Networks (CNN) can automatically learn multi - level feature representations from low - level to high - level in images, and the attention mechanism enables the model to focus on regions crucial for facial expression (such as eyes and mouth). To capture the dynamic characteristics of facial expressions, Recurrent Neural Networks (RNN) and their variants, especially Long Short - Term Memory Networks (LSTM), have been widely used to process time series of facial features. At the same time, strategies for fusing multi - modal data such as facial visual information, speech intonation, and physiological signals have also received increasing attention in order to improve the comprehensiveness and robustness of emotion recognition.
[0003] With the development of facial expression analysis technology, although certain achievements have been made, there are still many challenges in practical applications. Existing methods have poor robustness in natural environments, and factors such as illumination changes, pose differences, and occlusion can seriously interfere with the recognition of facial expressions. Different illumination intensities and angles can change the brightness and contrast of facial images, resulting in inaccurate feature extraction; people have diverse poses in daily life, and head rotation, tilting, etc. can change the position and shape of facial features; wearing occluders such as glasses and masks also increases the difficulty of recognition.
[0004] Facial expressions themselves are highly subjective and complex. Different people may have different ways of expressing the same emotion, and even the same person may have different facial expressions when expressing the same emotion in different situations. This makes the label consistency among different annotators poor during the data annotation process, thereby affecting the training quality and evaluation effect of the model.
[0005] Traditional facial expression analysis methods usually follow a fixed process, generally including several steps such as facial data capture, data preprocessing, feature extraction, and classification. In the data capture stage, facial images or video data are collected through devices such as cameras; in the preprocessing process, operations such as denoising and normalization are performed on the collected data to improve data quality; in the feature extraction link, pre-designed algorithms are used to extract facial features; finally, classification algorithms are used to classify the extracted features to determine the expression category. Although this process is systematic and standardized, it has obvious deficiencies. It often treats each facial feature as an independent individual for processing, and fails to fully consider the complex interactions between various facial features. In fact, the changes in facial expressions are the result of the collaborative action of multiple features. For example, the raising of eyebrows and the opening of the mouth may occur simultaneously to jointly express surprise, while traditional methods cannot effectively capture the associations between such features. After analysis by the applicant, the specific defects are as follows: Lack of comprehensive capture of complex spatio-temporal dynamic characteristics: Facial expressions are a continuous and dynamic process, which contains rich spatio-temporal information. Existing methods, including many deep learning-based methods, although they can process time series data, often have difficulty comprehensively and finely capturing the complex dynamic changes of facial features in both the time and space dimensions, and fail to deeply model the collaborative effects and mutual influences between different facial features over time. Simple time series models or static graph models are difficult to effectively capture such dynamic inter-feature collaborative relationships.
[0006] Ignoring the modeling and analysis of the dynamic association network between facial features: Although some methods begin to consider the spatial relationships between facial features, they usually model this relationship as a static network. The generation of facial expressions is a process in which multiple features form a dynamic association network and evolve over time. For example, during the process of an expression changing from calm to angry, the correlation intensity and pattern between eye, eyebrow, and mouth features will change. Existing methods fail to explicitly construct and analyze such a dynamically evolving inter-feature relationship network.
[0007] Difficulty in identifying dynamic key feature points: During the facial expression process, the importance of different feature points or their roles in the inter-feature relationship network are dynamically changing. Some feature points may play key roles in certain dynamic patterns of the network (for example, their connectivity fluctuates violently or remains stable continuously), and these "dynamic key points" may better reflect the essence of the expression than static key points. Existing technologies lack a systematic method to identify the feature points that are significant during the evolution of the dynamic network.
[0008] Limited interpretability: Although many end-to-end deep learning-based models have powerful performance, their internal mechanisms are "black boxes", making it difficult to reveal which specific dynamic behaviors or collaborative patterns of facial features play key roles in the final emotion judgment, which limits the application and further optimization of the models.
[0009] In summary, when analyzing facial expressions, existing methods often have difficulty in comprehensively and continuously capturing the dynamic change characteristics of expressions in both the time and space dimensions. In particular, they fail to effectively model and analyze the dynamic association network of facial features over time and identify the dynamic key points therein, resulting in limited accuracy and reliability of emotion recognition and insufficient interpretability. Summary of the Invention
[0010] In view of the problems existing in the background art, the present invention proposes a facial emotion analysis method based on the construction of a spatio-temporal feature network. The core of the present invention is to model the interaction between different features during the facial expression process as a network that dynamically evolves over time, and to achieve a more accurate and interpretable assessment of an individual's emotional state by analyzing the structural attributes, evolution laws of the network, and identifying the dynamic key points therein. The present invention aims to more deeply understand the generation mechanism and emotional meaning of facial expressions and overcome the deficiencies of the prior art in capturing complex spatio-temporal dynamic information and identifying dynamic important features. Technical Solution
[0011] A facial emotion analysis method based on the construction of a spatio-temporal feature network, which includes the following steps: S1. Video frame segmentation and feature extraction, including: S1.1. Obtain video data and perform video frame segmentation; S1.2. Obtain facial features in the video frame, including: facial feature points, head pose, gaze direction, facial action units AUs, S1.3. Construct a feature time series for each feature obtained in S1.2; S1.4. Perform normalization processing on each extracted feature time series; S2. According to the set time window length and time step, divide the video into multiple time periods. For each time period, based on the feature sequences of all frames within that time period, construct a facial feature network representing the association between facial features. Synthesize the networks calculated for all time periods to obtain a dynamic network sequence; S3. Perform dynamic network feature extraction and dynamic key point analysis based on the facial feature dynamic network; S4. Perform emotion calculation or psychological state assessment based on the dynamic network (features) or dynamic key points.
[0012] Specifically, in S1: Use a deep learning model to extract facial feature points, and the deep learning model is Openface, InsightFace or DeepFace; The facial feature points include the feature points of eyes, nose, and mouth; The head posture includes pitch angle, yaw angle, and roll angle, which are directly extracted by Openface; The gaze direction information is directly extracted by Openface.
[0013] Preferably, in S1, Openface is used to extract facial features of each frame, and the corresponding coordinates are obtained by extracting facial feature points, and based on this, the shape of the eyebrows, the degree of eye opening, and the degree of mouth opening and closing are calculated as facial feature values.
[0014] Specifically, in S2, the facial feature dynamic network is constructed through the following steps: S2.1. Define all the extracted facial features as nodes of the facial feature network; S2.2. Use the Pearson correlation coefficient to calculate the linear correlation between different facial feature time series; for each pair of facial feature time series ( x i , x j ), calculate its Pearson correlation coefficient through the following formula r ij :
[0015] where, and respectively represent the feature vectors of the i rd and j th facial feature time series at the t th frame, and respectively represent the means of the feature vectors of the i th and j th facial feature time series; Obtain the Pearson correlation coefficient matrix R, R ij =r ij ; S2.3. Adjacency matrix construction and graph percolation heuristic: Based on the calculated Pearson correlation coefficient matrix R, construct the adjacency matrix A of the facial feature network for the current time period; to obtain a dynamic network sequence.
[0016] Specifically, in S3, dynamic network feature extraction and dynamic key point analysis specifically include the following steps: S3.1. For the network Gt constructed for each time period, calculate its structural properties, including node degree, clustering coefficient, betweenness centrality, and eigenvector centrality; S3.2. Dynamic network feature design: Design new dynamic network features to capture the dynamic evolution pattern of the network; S3.3. Dynamic key point recognition: Based on the dynamic network features designed above, identify the dynamic key nodes of the network; the dynamic key nodes of the network are the feature points whose network attributes or inter - correlation patterns exhibit specific dynamic patterns during the evolution of facial expressions. S3.4. Network feature analysis: Analyze the topological structure features of the network.
[0017] Specifically, in S3.2, the new dynamic network features specifically include: Statistical features of node attribute time series: Calculate the mean, standard deviation, maximum value, minimum value, and range of change of the degree and centrality attributes of each node in the entire video sequence. Dynamic features of node attribute time series: Calculate the change rate, acceleration, and fluctuation frequency of the node attribute time series. Persistence / stability features of nodes in network evolution: Calculate the proportion of time periods in which nodes maintain high connectivity or high centrality. Dynamicity of nodes in community structure: Partition the network for each time period, and calculate the stability or frequency of switching of node community membership. Dynamic importance based on percolation analysis: During the simulation of network percolation, evaluate the degree of influence on the network connectivity in subsequent time periods after nodes are removed. Node control role and dynamic features of network controllability: Calculate the control centrality of nodes, the dynamic patterns of serving as key control roles (such as Control Hub, driving nodes), and the time - evolution features of the overall network controllability index.
[0018] Specifically, in S3.4, the topological structure features include: Degree distribution: Describes the distribution of nodes with different degree values in the network. Clustering coefficient: Measures the degree of aggregation of nodes in the network. Average path length: Represents the average shortest path length between any two nodes in the network.
[0019] Specifically, S4 specifically includes the following steps: S4.1. Construct key - point feature vectors. S4.2. Use the key - point feature vectors or dynamic network sequences as model inputs, and the classification result regression values of normal / depressed as model outputs, construct a dataset and label it, and divide it into a training set and a validation set. S4.3. Conduct model training and evaluation to determine the best model. S4.4. Obtain the key - point feature vectors or dynamic network sequences of the facial image to be predicted, and input them into the best model to obtain the facial emotion analysis result.
[0020] Specifically, in S4.1, the key-point feature vectors are constructed in one of the following ways: ① Concatenate or statistically process the original feature time series of the selected dynamic key nodes as input features; ② Concatenate the dynamic network features of the selected dynamic key nodes as input features; ③ Concatenate the dynamic network features of all nodes as input features; ④ Use the dynamic network sequence itself as input features.
[0021] Specifically, in S4.2: When the model input is the key-point feature vector, a traditional classifier or sequence model is used for classification or regression; When the model input is the dynamic network sequence, a dynamic graph neural network is used as the model for classification training.
[0022] Advantages of the present invention: The technical solution of the present invention solves the problems that existing methods are difficult to comprehensively capture the complex dynamic characteristics of facial expressions, lack the ability of dynamic association modeling and dynamic key-point recognition by constructing and analyzing the spatio-temporal dynamic network of facial features and identifying the dynamic key points therein. The logic and mechanism for solving the problems are as follows: Facial expressions are the result of the dynamic cooperation of multiple features, and this dynamic cooperation mode is the key to emotional expression. By segmenting the video data into time periods or windows and constructing an association network of facial features within each time period, a dynamic network sequence is obtained, which explicitly models the spatio-temporal evolution of the associations between features. Through in-depth analysis of this dynamic network sequence, especially by designing and calculating network features (such as variability and persistence of connectivity) that capture the dynamic characteristics of nodes and identifying dynamic key points based on these dynamic features, it is possible to reveal which features and their dynamic behaviors in the network are most important for emotional expression. Utilizing these dynamic network features, dynamic key-point information, or directly processing the dynamic graph sequence using a dynamic graph neural network model can make more effective use of the discriminant information contained in the dynamic associations. Brief Description of the Drawings
[0023] Figure 1 It is a flowchart of the facial emotion analysis method of the present invention.
[0024] Figure 2 It is a comparison diagram of network features in the embodiment. Detailed Embodiments
[0025] The logic for the present invention to solve the problem lies in that the generation and change of facial expressions are not independent behaviors of a single feature, but the result of the coordinated and collaborative change of multiple facial features. This collaborative pattern and its change over time (dynamicity) are the keys to understanding expressions and emotions. In this solution, facial features are regarded as network nodes, the correlation between features is regarded as edges, and different networks are constructed at different time periods, thereby explicitly modeling the spatio-temporal correlation network between facial features. The analysis of this dynamic network sequence, especially the identification of dynamic key points (those nodes with special behaviors during the network evolution), can reveal which features and their dynamic associations are most important for expressing specific emotions. For example, when expressing happiness, the upward curl of the mouth corners and the appearance of crow's feet at the corners of the eyes are coordinated; while when expressing sadness, the collaborative pattern of the downward droop of the eyebrows and the downward pull of the mouth corners may be different. By capturing these changing collaborative patterns and the dynamic key features that play a dominant role in these patterns, the emotional state can be understood more deeply and accurately than by only analyzing independent features or static associations. The dynamic network features extracted based on dynamic key points (such as the time variability and persistence of degree and centrality) directly quantify the manifestation of the dynamic behaviors of these nodes in the network, providing more discriminative information for subsequent classification tasks. Using models such as Dynamic GCN to directly process dynamic graphs further utilizes the complete spatio-temporal structure information of the network.
[0026] Combined with Figure 1 , the present invention discloses a facial emotion analysis method based on the construction of a spatio-temporal feature network, which includes the following steps: S1. Video frame segmentation and feature extraction, including: S1.1. Obtain video data and perform video frame segmentation; S1.2. Obtain facial features in the video frame, including: facial feature points, head pose, gaze direction, facial action units AUs, S1.3. Construct a feature time series for each feature obtained in S1.2; S1.4. Perform normalization processing on each extracted feature time series; S2. According to the set time window length and time step, divide the video into multiple time periods. For each time period, based on the feature sequences of all frames within that time period, construct a facial feature network representing the association between facial features. Synthesize the networks calculated for all time periods to obtain a dynamic network sequence; S3. Perform dynamic network feature extraction and dynamic key point analysis based on the facial feature dynamic network; S4. Perform emotion calculation or psychological state assessment based on the dynamic network (features) or dynamic key points.
[0027] In a preferred solution: Extract facial feature points using a deep learning model, where the deep learning model is Openface, InsightFace, or DeepFace; The facial feature points include the feature points of the eyes, nose, and mouth; The head pose includes the pitch angle, yaw angle, and roll angle, which are directly extracted by Openface; The gaze direction information is directly extracted by Openface.
[0028] To implement the extraction of relevant information based on Openface, the inventor provides the following pseudocode: Initialization Set: - Video source directory TEMP_VIDEO_DIR - Feature output directory OUTPUT_DIR - OpenFace executable file FEATURE_EXTRACTION_BIN If OUTPUT_DIR does not exist, create it Define the function extract_features(video_path, video_output_dir) Construct the command: FEATURE_EXTRACTION_BIN -f video_path -out_dir video_output_dir -all Execute the command and return whether it is successful Main process Collect all video files with extensions.mp4 / .avi / .mov / .wmv / .flv in TEMP_VIDEO_DIR → video_list If video_list is empty, end the program For each video in video_list: Take the file name (without extension) as the subdirectory name Create the corresponding subdirectory video_output_dir under OUTPUT_DIR Call extract_features(video_path, video_output_dir) Report that the processing is completed after finishing.
[0029] In S1, Openface is used to extract the facial features of each frame. The corresponding coordinates are obtained by extracting the facial feature points, and the shape of the eyebrows, the degree of eye opening, and the degree of mouth opening and closing are calculated based on this as the facial feature values.
[0030] To accurately capture and quantify facial expressions and dynamics from a video, the OpenFace toolkit is used to extract the detailed facial features of each frame. This process begins with processing each input frame image: First, face detection is performed to locate the face region in the image, and a unique face_id is assigned to each detected face for tracking in the video sequence. Once the face is located, the next crucial step is facial key point detection. At this stage, a set of predefined 2D key points (e.g., the standard 68 or more key points including eye details) are accurately calibrated within the detected face region. These points (such as x_0 to x_67, y_0 to y_67) indicate the contours and positions of important facial components such as eyebrows, eyes, nose, and mouth in the image pixel coordinate system.
[0031] Based on these 2D key points, higher-level facial information is further inferred. By matching these 2D points with a general 3D face model and combining the camera parameters (estimated by OpenFace), the three-dimensional pose of the head (such as translation pose_Tx, pose_Ty, pose_Tz and rotation pose_Rx, pose_Ry, pose_Rz) can be calculated, which describes the spatial position and orientation of the head in the camera coordinate system. With the head pose, the 2D key points can be back-projected or converted into 3D world coordinates (such as X_0 to X_67, Y_0 to Y_67, Z_0 to Z_67) through the fitted 3D model, thus providing the true scale information of the facial structure.
[0032] For the eye region, more refined eye key points (such as eye_lmk_x_0 to eye_lmk_x_55 and their corresponding Y and 3D coordinates) are used to estimate the gaze direction. This includes calculating a three-dimensional gaze vector for each eye (such as gaze_0_x, gaze_0_y, gaze_0_z for the left eye and gaze_1_x, gaze_1_y, gaze_1_z for the right eye), as well as an average gaze angle (gaze_angle_x, gaze_angle_y) that fuses the information of both eyes. These gaze features reveal the focus of an individual's visual attention.
[0033] The description of facial expressions is mainly achieved by analyzing the geometric changes of facial key points. For example, the shape of the eyebrows (such as raising, lowering, squeezing), the degree of eye opening, and the opening and closing degree and shape of the mouth (such as the amplitude of the upward curl of the mouth corners when smiling, the size of the mouth opening when surprised), etc., can all be quantified by calculating the distances, angles, or relative position changes between specific key points. Further, these geometric changes or the texture information of local image patches are input into a pre-trained machine learning model to identify and quantify a series of standard facial action units (AUs), such as AU01_r (inner brow raise intensity) or AU12_c (classification result of mouth corner stretch). In addition, OpenFace also fits a parametric 3D face model and outputs parameters describing the specific shape and current expression of the face.
[0034] Therefore, through a series of processing steps from the bottom layer to the top layer - from basic face and key point detection, to 3D pose and shape reconstruction, and then to the analysis of specific gaze and expression action units - the visual information is transformed into a set of structured numerical features. These features not only contain precise descriptions of specific facial regions (eyebrows, eyes, nose, mouth, etc.), but also can more intuitively and quantitatively reflect the subtle changes in facial expressions, providing a crucial data basis for subsequent research on emotion recognition, mood analysis, etc.
[0035] In S2, the facial feature dynamic network is constructed through the following steps: S2.1: Define all the extracted facial features as the nodes of the facial feature network; S2.2: Use the Pearson correlation coefficient to calculate the linear correlation between different facial feature time series; for each pair of facial feature time series ( x i , x j ), calculate its Pearson correlation coefficient through the following formula r ij :
[0036] where, and respectively represent the feature vectors of the i th and the j th facial feature time series at the t th frame, and respectively represent the means of the feature vectors of the i th and the j th facial feature time series; Obtain the Pearson correlation coefficient matrix R, R ij =r ij ; S2.3, Adjacency Matrix Construction and Graph Percolation Heuristic: Based on the calculated Pearson correlation coefficient matrix R, construct the adjacency matrix A of the facial feature network for the current time period; to obtain a dynamic network sequence.
[0037] In S3, the dynamic network feature extraction and dynamic key point analysis specifically include the following steps: S3.1, For each network Gt constructed for each time period, calculate its structural properties, including node degree, clustering coefficient, betweenness centrality, and eigenvector centrality; S3.2, Dynamic Network Feature Design: Design new dynamic network features to capture the dynamic evolution pattern of the network; S3.3, Dynamic Key Point Identification: Based on the above-designed dynamic network features, identify the dynamic key nodes of the network; the dynamic key nodes of the network are the feature points whose network properties or interconnection patterns show specific dynamic patterns during the facial expression evolution process; S3.4, Network Feature Analysis: Analyze the topological structure features of the network.
[0038] In S3.2, the new dynamic network features specifically include: Statistical Features of Node Attribute Time Series: Calculate the average value, standard deviation, maximum value, minimum value, and change range of the degree and centrality attributes of each node in the entire video sequence; Dynamic Features of Node Attribute Time Series: Calculate the change rate, acceleration, and fluctuation frequency of the node attribute time series; Persistence / Stability Features of Nodes in Network Evolution: Calculate the proportion of time periods in which nodes maintain high connectivity or high centrality; Dynamicity of Nodes in Community Structure: Perform community partitioning on the network for each time period, and calculate the stability or frequency of switching of node community membership; Dynamic Importance Based on Percolation Analysis: During the simulation of network percolation, evaluate the degree of influence on the network connectivity of subsequent time periods after nodes are removed; Node Control Role and Dynamic Features of Network Controllability: Calculate the control centrality of nodes, the dynamic patterns of serving as key control roles (such as Control Hub, driving nodes), and the time evolution features of the overall network controllability index.
[0039] In S3.4, the topological structure features include: Degree Distribution: Describes the distribution of nodes with different degree values in the network; Clustering Coefficient: Measures the degree of aggregation of nodes in the network; Average Path Length: Represents the average shortest path length between any two nodes in the network.
[0040] S4 specifically includes the following steps: S4.1. Construct a key-point feature vector; S4.2. Use the key-point feature vector or the dynamic network sequence as the model input, and use the regression value of the normal / depressed classification result as the model output. Construct a data set, label it, and divide it into a training set and a validation set; S4.3. Conduct model training and evaluation to determine the best model; S4.4. Obtain the key-point feature vector or the dynamic network sequence of the facial image to be predicted, and input it into the best model to obtain the facial emotion analysis result.
[0041] In S4.1, the key-point feature vector is constructed in one of the following ways: ① Concatenate or statistically process the original feature time series of the selected dynamic key nodes as the input feature; ② Concatenate the dynamic network features of the selected dynamic key nodes as the input feature; ③ Concatenate the dynamic network features of all nodes as the input feature; ④ Use the dynamic network sequence itself as the input feature.
[0042] In S4.2: When the model input is the key-point feature vector, a traditional classifier or sequence model is used for classification or regression; When the model input is the dynamic network sequence, a dynamic graph neural network is used as the model for classification training.
[0043] The steps of the present invention are described in detail below with a specific embodiment as follows:
[0044] 1. Video data preprocessing and feature extraction.
[0045] In this embodiment, first, the video data to be analyzed is selected. The video data is the facial videos of primary and middle school students collected. Each video has a corresponding label: The DASS21 scale contains three sub-scales: the depression scale, the anxiety scale, and the stress scale. For the depression scale, a score of ≤9 is normal, 10 - 13 is mild, 14 - 20 is moderate, 21 - 27 is severe, and ≥28 is very severe; for the anxiety scale, a score of ≤7 is normal, 8 - 9 is mild, 10 - 14 is moderate, 15 - 19 is severe, and ≥20 is very severe; for the stress scale, a score of ≤14 is normal, 15 - 18 is mild, 19 - 25 is moderate, 26 - 33 is severe, and ≥34 is very severe. In this embodiment, the subsequent classification task is completed according to the scores of the depression scale. The labels are binarized according to the scores. And in order to reduce the interference between stress, anxiety, and depression and the unclear distinction between mild cases and normal people caused by inaccurate test results, those with scores on all three scales less than the corresponding boundary scores are set as normal people with a label of 0, and those with a depression scale score greater than 14 are the depressed population with a label of 1. Then the video is extracted into video frame data, and the professional facial feature extraction tool OpenFace is used to process it. This tool can detect the faces in the video frame by frame and automatically extract rich facial feature information. The continuous frame sequence is divided into sliding windows with a window size of 50 frames and a step size of 30 frames to obtain a series of time periods. The extracted features include but are not limited to: o Facial key point coordinates: Record the spatial position information of each key point on the face (such as eyes, mouth, nose, etc.); o Head pose: Describe the rotation and displacement of the head in three-dimensional space; o Facial Action Unit (AU): Quantify the activities of facial muscles, providing a basis for subsequent facial expression analysis.
[0046] All the extracted features are finally output in the form of a CSV file. Each CSV file represents the facial feature time series of a person. Each feature is standardized by Z-score so that its mean is 0 and the standard deviation is 1. Such processing helps to eliminate the scale differences of different features.
[0047] 2. Construction of the facial feature dynamic network.
[0048] For each time period \(t\) (\(t = 1,\ldots,T\), where \(T\) is the total number of time periods), construct the facial feature network \(G_t\) within that time period. Each node in the network \(G_t\) represents a time series of normalized facial features (a sequence of 50 frames within the current time period). For the current time period \(t\), calculate the Pearson correlation coefficient \(r_{ij}(t)\) between any two feature time series \(i\) and \(j\) (\(i,j\in\{1,\ldots,156\}\)) within these 50 frames. Construct the correlation coefficient matrix \(R_t\) for the current time period, where \(R_t[i][j]=r_{ij}(t)\). For the current time period \(t\), we calculate the proportion of the size of the largest connected component of the network \(G_t(|r|>\theta)\) to the total number of nodes for different correlation thresholds \(\theta\in[0,1]\). Plot the curve of this proportion versus \(\theta\) (the percolation curve). Select the threshold corresponding to the maximum slope of the percolation curve (i.e., the rapid growth of the largest connected component) as the dynamic threshold \(\theta_t\) for the current time period. For any node pair \((i,j)\), if \(|r_{ij}(t)|>\theta_t\), then \(A_t[i][j]=|r_{ij}(t)|\); otherwise, \(A_t[i][j]=0\). Set the diagonal elements to 0. Repeat this process to obtain a series of sparse weighted adjacency matrices \(\{A_1,A_2,\ldots,A_T\}\), representing the dynamic evolution of the facial feature network over time.
[0049] 3. Dynamic network feature extraction and dynamic key point analysis.
[0050] Analyze the dynamic network sequence \(\{A_1,\ldots,A_T\}\) to extract features that capture the dynamic characteristics of the nodes.
[0051] For each node \(i\) (\(i = 1,\ldots,n\)), calculate the following dynamic network features: Average degree (AvgDeg\(_i\)): The average degree of node \(i\) in the networks over all time periods.
[0052] AvgDeg\(_i=(1 / T)*\sum_{t = 1}^T\) Degree\((i,t)\), where Degree\((i,t)\) is the degree of node \(i\) in the network \(A_t\).
[0053] Degree standard deviation (DegStd\(_i\) - dynamic variability): The temporal standard deviation of the degrees of node \(i\) in the networks over all time periods.
[0054] DegStd\(_i = \) StdDev(\(\{\)Degree\((i,1),\ldots,\)Degree\((i,T)\}\)). DegStd\(_i\) captures the dynamic activity or instability of the node connectivity.
[0055] High Degree Persistence (HighDegPersisti - Dynamic Persistence): The proportion of time periods during which the degree of node i exceeds its average degree (or the average degree of all nodes).
[0056] HighDegPersisti = (1 / T) * ∑{t=1}^T I(Degree(i, t)>AvgDegi), where I() is the indicator function. HighDegPersisti measures the persistence of the node as an important connection point.
[0057] Standard Deviation of Degree Change Rate (DegRateStdi - Dynamic Activity Burstiness): The standard deviation of the rate of change (difference) of the degree of node i between consecutive time periods.
[0058] DegRatei(t) = Degree(i, t+1) - Degree(i, t). DegRateStdi = StdDev({DegRatei(1), ..., DegRatei(T-1)}). DegRateStdi captures the intensity of the change in the node's connectivity.
[0059] The mean, standard deviation, and persistence characteristics of other centrality metrics such as betweenness centrality and eigenvector centrality can be calculated similarly.
[0060] Dynamic Key Point Identification: In this embodiment, we identify nodes that simultaneously meet the following conditions as dynamic key points: DegStdi is higher than the 75th percentile of DegStd of all nodes and HighDegPersisti is higher than the 50th percentile of HighDegPersist of all nodes. This method identifies nodes with both dynamic activity and a certain degree of persistence in connectivity. Alternatively, directly use the above-calculated dynamic network characteristics of all nodes as input features for a classification model.
[0061] 4. Network Feature Analysis.
[0062] Conduct an in-depth analysis of the above-obtained network representation, such as: • Use a community detection algorithm (such as the Louvain algorithm) to identify the functional modules of facial features; • Calculate metrics such as the degree distribution, clustering coefficient, and average path length of the entire network to understand the topological characteristics of the network; • Combine the dynamic changes in the time window to explore the key nodes and evolution trends of the network at different stages.
[0063] These analyses not only help to identify the interrelationships and division of labor among facial features, but also provide data support for subsequent studies on the correlations between expressions and factors such as psychology and emotions. The comparison chart of network feature analysis is as shown in Figure 2 which shows that there are differences between the networks constructed by different labeled populations.
[0064] 5. Classification tasks based on dynamic network features.
[0065] After completing the above network construction and feature extraction, this embodiment further uses a Dynamic Graph Neural Network (DGNN) to perform a classification task on the facial feature network. The process includes: • Network input preparation: For each individual, construct a dynamic graph sequence {G1, ..., GT}. Each graph Gt is an undirected weighted graph with n nodes, and the adjacency matrix is At. Each node can be assigned a node feature vector, such as including the average AU intensity, average pose change, etc. of the node at the current time period, as well as the dynamic network feature of the node (such as DegStdi); • Model selection and training: According to specific task requirements, select a suitable GNN model (such as DynamicGCN), and construct the corresponding loss function. For classification tasks, use cross-entropy loss. The constructed DynamicGCN model adopts a temporal graph convolutional network architecture, which mainly includes the following components: Node feature embedding layer: Receives the original node features (including topological features such as degree and clustering coefficient), and maps them to a high-dimensional latent space through linear transformation. The initial dimension of each node is 2-3, and it is mapped to a 64-dimensional latent space. Temporal graph convolutional layer: Processes the graph structures of multiple time steps, and each time step's graph is processed by a separate GCN layer. The GCN layer applies a message passing mechanism to aggregate neighbor node information. Temporal attention mechanism: Assigns weights to the graph representations of different time steps and calculates the weighted combination. Graph pooling layer: Adopts global mean pooling to aggregate node-level features into graph-level representations. Prediction layer: Designs different output layers according to the task type (regression or classification): Classification task: A linear layer followed by a Softmax activation outputs class probabilities. Network parameter settings The key parameters of the DynamicGCN model include: Hidden channel number: 64, Activation function: ReLU, Dropout rate: 0.5, Learning rate: 0.001, Weight decay: 5e-4 (L2 regularization); • Model prediction and evaluation: After training, input the test set or the new facial feature network into the model to output the corresponding classification labels. Subsequently, evaluate the model performance according to indicators such as accuracy and precision. Specifically, select accuracy, precision, recall, and F1 metrics to evaluate the classification effect; Through the above steps, this embodiment not only realizes the construction and spatio-temporal dynamic analysis of the dynamic facial feature network, but also further classifies the network using graph neural networks, providing an efficient and scalable solution for a variety of downstream application scenarios (such as emotion recognition, individual recognition, facial action intensity prediction, etc.). Through the complete process of this embodiment, facial features can be extracted, network constructed and analyzed in a video, and graph neural networks can be used to carry out classification and regression tasks, which can not only capture the details of instantaneous features, but also examine the dynamic patterns evolving over time, providing a relatively complete technical means for facial expression recognition, emotion research and other related applications. The detailed steps and processing methods of this embodiment can be adjusted and extended according to actual requirements and environments.
[0066] The facial emotion analysis method based on the construction of spatio-temporal feature network proposed by the present invention has a wide range of application scenarios:
[0067] 1. Mental health assessment and assisted diagnosis and treatment: In the screening and diagnosis of mental diseases (such as depression, anxiety disorder) and psychological disorders, by analyzing the dynamic network and dynamic key points of the patient's facial expressions, provide objective and quantitative assessment basis for doctors. For example, identifying abnormal dynamic association patterns of specific facial features related to depression (such as around the eyes, corners of the mouth) or abnormal behaviors in the network evolution (such as low connectivity variability, low activity), can assist doctors in judging the severity and type of the condition.
[0068] 2. Real-time effect detection and evaluation: Real-time monitor the dynamic changes of the patient's facial expressions during the treatment process. Dynamic network analysis can reveal whether the treatment has improved the collaborative pattern between facial features and whether the dynamic key points have restored normal dynamic behaviors, so as to evaluate the effectiveness of the treatment plan and make adjustments.
[0069] 3. Psychological counseling and treatment: Help psychological counselors capture the emotional dynamics and subtle changes of the clients more accurately. By visualizing (although the visualization details are removed in this article, but conceptually the analysis results can be visualized) and analyzing the dynamic network and key points of facial features, counselors can understand the inner state of the clients more deeply and identify the trigger points and patterns of their emotional changes.
[0070] The specific embodiments described in this article are only illustrative of the spirit of the present invention. Those skilled in the art of the present invention can make various modifications or supplements to the described specific embodiments or use similar ways to replace them, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.
Claims
1. A facial emotion analysis method based on the construction of a spatio-temporal feature network, characterized in that It includes the following steps: S1. Video frame segmentation and feature extraction, including: S1.
1. Obtain video data and perform video frame segmentation; S1.
2. Obtain facial features in the video frame, including: facial feature points, head pose, gaze direction, facial action units AUs; S1.
3. Construct a feature time series for each feature obtained in S1.2; S1.
4. Perform normalization processing on each extracted feature time series; S2. According to the set time window length and time step, divide the video into multiple time periods; for each time period, construct a facial feature network representing the association between facial features based on the feature sequences of all frames within that time period; synthesize the networks calculated for all time periods to obtain a dynamic network sequence; S3. Perform dynamic network feature extraction and dynamic key point analysis based on the facial feature dynamic network; S4. Perform emotion calculation or psychological state assessment based on the dynamic network features or dynamic key points.
2. The method according to claim 1, wherein In S1: Use a deep learning model to extract facial feature points, and the deep learning model is Openface, InsightFace or DeepFace; The facial feature points include the feature points of eyes, nose and mouth; The head pose includes pitch angle, yaw angle and roll angle, which are directly extracted by Openface; The gaze direction information is directly extracted by Openface.
3. The method according to claim 2, wherein In S1, use Openface to extract the facial features of each frame, obtain the corresponding coordinates by extracting facial feature points, and calculate the shape of eyebrows, the opening degree of eyes, and the opening and closing degree of the mouth as facial feature values based on this.
4. The method according to claim 1, wherein In S2, the facial feature dynamic network is constructed through the following steps: S2.
1. Define all the extracted facial features as the nodes of the facial feature network; S2.
2. Calculate the linear correlation between different facial feature time series using the Pearson correlation coefficient; for each pair of facial feature time series ( x i , x j ), calculate their Pearson correlation coefficient through the following formula r ij : Among them, and respectively represent the feature vectors of the i -th and j -th facial feature time series at the t -th frame, and respectively represent the means of the feature vectors of the i -th and j -th facial feature time series; Obtain the Pearson correlation coefficient matrix R, R ij =r ij ; S2.
3. Adjacency matrix construction and graph percolation heuristic: Based on the calculated Pearson correlation coefficient matrix R, construct the adjacency matrix A of the facial feature network for the current time period; to obtain a dynamic network sequence.
5. The method according to claim 1, wherein In S3, the dynamic network feature extraction and dynamic key point analysis specifically include the following steps: S3.
1. For the network Gt constructed for each time period, calculate its structural properties, including node degree, clustering coefficient, betweenness centrality, eigenvector centrality; S3.
2. Design of dynamic network features: Design new dynamic network features to capture the dynamic evolution pattern of the network; S3.
3. Identification of dynamic key points: Based on the above-designed dynamic network features, identify the dynamic key nodes of the network; the dynamic key nodes of the network are the feature points whose network properties or mutual association patterns show specific dynamic patterns during the evolution of facial expressions; S3.
4. Analysis of network feature: Analyze the topological structure features of the network.
6. The method according to claim 5, wherein In S3.2, the new dynamic network features specifically include: Statistical features of the node attribute time series: Calculate the average value, standard deviation, maximum value, minimum value, and change range of the degree and centrality attributes of each node in the entire video sequence; Dynamic features of the node attribute time series: Calculate the change rate, acceleration, and fluctuation frequency of the node attribute time series; Persistence / stability characteristics of nodes in network evolution: Calculate the proportion of time periods during which nodes maintain high connectivity or high centrality; Dynamics of nodes in community structure: Partition the network for each time period, and calculate the stability or frequency of switching of the node community membership; Dynamic importance based on percolation analysis: During the simulation of network percolation, evaluate the degree of influence on the network connectivity in subsequent time periods after a node is removed; Node control roles and dynamic characteristics of network controllability: Calculate the control centrality of nodes, the dynamic patterns of playing key control roles, and the time evolution characteristics of the overall network controllability index.
7. The method according to claim 5, wherein In S3.4, the topological structure characteristics include: Degree distribution: Describes the distribution of nodes with different degree values in the network; Clustering coefficient: Measures the degree of aggregation of nodes in the network; Average path length: Represents the average shortest path length between any two nodes in the network.
8. The method according to claim 1, wherein S4 specifically includes the following steps: S4.
1. Construct a key-point feature vector; S4.
2. Use the key-point feature vector or the dynamic network sequence as the model input, and use the classification result regression value of normal / depressed as the model output to construct a data set and label it, and divide it into a training set and a validation set; S4.
3. Conduct model training and evaluation to determine the best model; S4.
4. Obtain the key-point feature vector or the dynamic network sequence of the facial image to be predicted, and input it into the best model to obtain the facial emotion analysis result.
9. The method according to claim 8, characterized in that In S4.1, the key-point feature vector is constructed in one of the following ways: ① Concatenate or statistically process the original feature time series of the selected dynamic key nodes as input features; ② Concatenate the dynamic network features of the selected dynamic key nodes as input features; ③ Concatenate the dynamic network features of all nodes as input features; ④ Use the dynamic network sequence itself as input features.
10. The method according to claim 8, wherein In S4.2: When the model input is the key-point feature vector, a traditional classifier or sequence model is used for classification or regression; When the model input is the dynamic network sequence, a dynamic graph neural network is used as the model for classification training.
Citation Information
Cited By
Emotional state dynamic monitoring system and method based on video image spatial-temporal characteristics
CN122024014A
Emotion state dynamic monitoring system and method based on spatio-temporal features of video images
CN122024014B
Method and system for analyzing dynamic facial features of psychological state based on causal graph learning
CN122135415A