Pain Expression Detection Methods and Systems
By acquiring facial video streams, performing facial key point localization and frame alignment, dividing muscle regions, constructing a muscle dynamics model, and fusing multi-scale spatiotemporal features, the problem of insufficient dynamic feature capture in existing pain expression detection is solved, achieving high-accuracy pain level detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN HUAANTAI INTELLIGENT TECH CO LTD
- Filing Date
- 2026-03-03
- Publication Date
- 2026-06-02
AI Technical Summary
Existing methods for detecting pain expressions rely on static image analysis, which cannot effectively capture facial dynamics and temporal features, resulting in low detection accuracy and difficulty in meeting practical application needs.
The system acquires facial video streams, performs facial landmark localization and frame alignment, divides facial muscle regions and extracts texture features, constructs a facial muscle dynamics model, generates adaptive candidate spatiotemporal windows through multi-scale spatiotemporal feature and attention-weighted fusion, verifies consistency by combining muscle state vectors, and finally outputs the pain level.
By accurately capturing facial dynamics and temporal changes, and combining muscle dynamics models with attention-weighted fusion, the detection accuracy is improved to meet practical application needs.
Smart Images

Figure CN122135416A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and more specifically, to a method and system for detecting pain expressions. Background Technology
[0002] With the development of computer vision technology, automatic pain detection based on facial expressions has become a research hotspot. However, most existing pain expression detection methods rely on static image analysis, which cannot capture the dynamic changes in facial expressions. Pain expressions are usually manifested as subtle changes and rapid contractions of facial muscles. Static images are difficult to effectively extract these temporal features, resulting in low detection accuracy and failing to meet the needs of practical applications. Summary of the Invention
[0003] The main purpose of this application is to provide a pain expression detection method and system, which aims to solve the technical problem that existing pain expression detection methods rely on static image analysis, cannot capture facial dynamics and temporal features, resulting in low detection accuracy and difficulty in meeting the needs of practical applications.
[0004] The first aspect of this application proposes a method for detecting pain expression, including: Acquire facial video streams, perform facial key point localization and frame alignment, segment facial muscle regions and extract texture features to obtain aligned frame sequences and muscle region texture feature sets. Motion energy is calculated based on texture feature set. Micro-expression temporal segments are extracted and enhanced from aligned frame sequence. Multi-scale spatiotemporal features are extracted, facial muscle dynamics model is constructed and state estimation is completed, resulting in enhanced micro-expression segments, multi-scale spatiotemporal features and muscle state vectors. Based on multi-scale spatiotemporal features, spatiotemporal attention weighted fusion is performed to obtain weighted spatiotemporal features. Combined with muscle state vector analysis, muscle synergy relationship is generated to generate an adaptive candidate spatiotemporal window set. After dynamic fusion of weighted spatiotemporal features, temporal modeling is used to obtain the probability distribution of pain level. Multi-scale micro-expression fusion features are obtained by fusing enhanced micro-expression fragment features, and the optimal spatiotemporal window is selected. Muscle dynamic consistency is verified based on muscle state vectors, micro-expression temporal consistency is verified based on multi-scale micro-expression fusion features, and the effectiveness of the optimal spatiotemporal window is verified. If any verification fails, the corresponding step is returned for reprocessing. If all verifications pass, the final pain level is output.
[0005] Further, the steps of acquiring facial video streams, performing facial landmark localization and inter-frame alignment, segmenting facial muscle regions and extracting texture features to obtain aligned frame sequences and muscle region texture feature sets include: Collect bimodal facial video streams and resting segments, construct individual muscle feature baselines and pre-train muscle dynamics models, combine trajectories and baselines to locate facial key points in the video streams, and obtain a set of bimodal key point coordinates; Based on the coordinate set, the dual-modal fusion weight is calculated to perform inter-frame predictive joint alignment. At the same time, optical flow field features are extracted to realize cross-frame association of key points, resulting in a predicted aligned frame sequence. Based on the aligned frame sequence, combined with facial anatomical atlas and muscle dynamics model state parameters, refined muscle regions are divided, and muscle region division results are generated. Based on the segmentation results, bimodal texture features are extracted and differiated from the individual muscle feature baseline. Cross-frame associated texture features are then fused to obtain a set of texture features for the muscle region.
[0006] Furthermore, the steps of calculating motion energy based on texture feature sets, extracting and enhancing micro-expression temporal segments from aligned frame sequences, extracting multi-scale spatiotemporal features, constructing a facial muscle dynamics model and completing state estimation to obtain enhanced micro-expression segments, multi-scale spatiotemporal features, and muscle state vectors include: A multi-scale feature pyramid is constructed based on the texture feature set. The motion energy at each scale is calculated and the optical flow field features are fused to obtain a multi-scale motion energy set with enhanced optical flow. By combining spatial attention mechanisms to locate key pain areas, and extracting micro-expression temporal segments containing peak frames based on the hierarchical extraction of motion energy sets, key features are enhanced through attention enhancement networks to obtain multi-scale enhanced micro-expression segments; Spatial features, temporal dependence features, and optical flow temporal features at various scales are extracted from the fragments and fused to form multi-scale spatiotemporal features; By combining bimodal information and attention weights, a facial muscle dynamics model is constructed, and the state transition matrix and noise covariance matrix are optimized. Input the texture feature set to complete the state estimation and obtain the muscle state vector.
[0007] Furthermore, the step of generating an adaptive candidate spatiotemporal window set by performing spatiotemporal attention weighted fusion based on multi-scale spatiotemporal features to obtain weighted spatiotemporal features, combining muscle state vectors to analyze muscle synergy relationships, includes: Multi-scale channel decomposition is performed on multi-scale spatiotemporal features. Spatial attention weights are calculated independently for each group and spatial attention maps are generated. Pain-related regions are identified through dynamic thresholding to obtain spatial weighted features. Multi-head self-attention module and causal convolution module are constructed for temporal features. The outputs of the two modules are weighted and fused to obtain temporal weighted features. Then, spatial weighted features and temporal weighted features are fused to obtain weighted spatiotemporal features. Based on muscle state vectors, a vector autoregression model is constructed to analyze the causal relationship between muscles and construct a muscle synergy graph. Facial muscle function groups are divided to calculate the synergy within each group. A directed graph of muscle propagation is constructed and key propagation paths are selected. Candidate spatiotemporal windows are generated based on micro-expression fragments, muscle coordination maps, functional group coordination, and key propagation paths. The window quality score is calculated and the feature consistency is verified. If the score is met, it is determined as an adaptive candidate spatiotemporal window set; otherwise, it is returned to adjust the dynamic coefficients and reprocessed.
[0008] Furthermore, the steps of constructing a vector autoregression model based on muscle state vectors, analyzing causal relationships between muscles and constructing a muscle synergy graph, dividing facial muscle function groups to calculate intra-group synergy, constructing a directed muscle propagation graph and screening key propagation paths include: A regularized vector autoregression model is constructed for the muscle state vector. After fitting the model, the corresponding muscle pairs are extracted, the causal relationship between muscles is analyzed and the relevant statistics are calculated, and a sparse muscle synergy graph is constructed. A nonlinear fractional vector autoregressive model is constructed to process the time series of muscle states and analyze the relevant causal relationships. A corresponding synergy graph is constructed, and various synergy graphs are weighted and fused to obtain a fused muscle synergy graph. Based on facial anatomy, facial muscle function groups are divided, hierarchical clustering is performed on each function group to divide it into sub-function groups, multi-scale cross-correlation functions of muscles within the group are calculated, strong synergistic function groups are labeled, and the synergistic degree of function groups is obtained. By integrating muscle synergy graphs and functional group synergy, a hierarchical fractional directed muscle propagation graph is constructed to extract the main propagation paths and screen key propagation paths.
[0009] Furthermore, the steps of obtaining a pain level probability distribution by dynamic fusion of weighted spatiotemporal features and time-series modeling, obtaining multi-scale micro-expression fusion features by fusing enhanced micro-expression fragment features, and selecting the optimal spatiotemporal window include: Pain contribution weights and motion energy curves are calculated based on muscle state vectors. Weighted spatiotemporal features are then subjected to adaptive weighting and a candidate window set is generated to obtain weighted adaptive spatiotemporal features. The weighted adaptive spatiotemporal features are input into the temporal modeling network, and local temporal modeling is performed based on the candidate window set to output a pain level probability distribution sequence. Temporal transformation is performed on the enhanced micro-expression fragment features to extract the intensity spectrum. Multi-scale fusion is then performed based on the pain level probability distribution sequence and the intensity spectrum to obtain multi-scale micro-expression fusion features. Based on the confidence ranking of the pain level probability distribution sequence and the energy distribution of the intensity spectrum, the optimal spatiotemporal window is selected from the multi-scale micro-expression fusion features.
[0010] Furthermore, the step of inputting the weighted adaptive spatiotemporal features into the temporal modeling network, performing local temporal modeling based on the candidate window set, and outputting a pain level probability distribution sequence includes: Residual supplementation, spatiotemporal decoupling and pain semantic alignment are performed on the weighted adaptive spatiotemporal features. The muscle activation intensity of the candidate window is calculated by combining the muscle state vector. The modeling depth is determined according to the motion energy gradient. The feature interval is divided by windowing to obtain the spatiotemporal feature set of window decoupling semantic residuals and window modeling configuration information. The feature set is input into the temporal modeling network according to the window modeling configuration information. Hierarchical local temporal modeling is performed on each feature interval to extract the deep temporal features of each candidate window, thus obtaining a semantic deep temporal feature set. Integrate semantic deep temporal feature sets, perform residual correction based on decoupled features after temporal fusion, and combine muscle activation intensity and motion energy gradient for multi-scale attention weighted aggregation to obtain fused aggregated probability feature sequence; The spatiotemporal dimension aggregation of the fused and aggregated probability feature sequence is processed, global temporal regularization optimization is performed, and the pain level probability distribution sequence is output.
[0011] Further, the step of inputting the feature set into the temporal modeling network according to the window modeling configuration information, performing hierarchical local temporal modeling on each feature interval, extracting the deep temporal features of each candidate window, and obtaining the semantic deep temporal feature set includes: Frame-level temporal alignment is performed on the feature set, and semantic feature enhancement, residual feature completion and multi-scale feature mapping are completed simultaneously. The muscle activation intensity of the candidate window is calculated by combining the muscle state vector and spatiotemporal decoupling is performed to obtain the hierarchical window configuration and decoupled alignment semantic residual multi-scale feature interval set. The feature interval set is configured into a hierarchical window and input into the temporal modeling network. The spatiotemporal decoupling local temporal modeling is performed by combining hierarchical and scale-based methods. The initial decoupling depth temporal features are extracted to obtain a multi-scale hierarchical decoupling temporal feature set. Semantic association extraction and muscle semantic enhancement are performed on the feature set. Feature hierarchical aggregation, multi-scale fusion, residual supplementation and temporal feature correction are completed in sequence to obtain a semantically enhanced and corrected deep temporal feature set. By integrating this feature set and performing global regularization according to the temporal relationship of the window, semantic scale adaptation and inter-level feature fusion are completed to obtain a semantic deep temporal feature set.
[0012] Furthermore, the steps of verifying muscle dynamic consistency based on muscle state vectors, verifying micro-expression temporal consistency based on multi-scale micro-expression fusion features, and verifying the effectiveness of the optimal spatiotemporal window, where if any verification fails, the corresponding step is returned for reprocessing; and if all verifications pass, the final pain level is output, include: Based on the muscle state vector, the consistency of muscle dynamics is verified. The prediction error is calculated by filtering, a muscle co-evaluation graph is constructed, and node embeddings are learned by graph neural network. Relevant parameters are calculated to construct a joint consistency index, and the prediction error matrix, node embedding vector and joint consistency score are output. Based on the multi-scale micro-expression fusion features, the temporal consistency of micro-expressions is verified. A probabilistic model is constructed and predicted samples are generated by sampling. Log likelihood is calculated. A generative adversarial network is used to generate samples and calculate relevant losses. A temporal consistency index is constructed and the corresponding score and sample-related features are output. Verify the effectiveness of the optimal spatiotemporal window, calculate the relevant attention weight distribution and sample diversity within the window, integrate relevant indicators to obtain the window quality score, combine the two types of consistency scores to construct an overall verification framework, and output the window effectiveness result and the comprehensive consistency score. If the overall validation fails, joint feedback optimization is performed using methods such as gradient descent and model parameter updates to recalculate each indicator and verify the optimization effect; if all indicators meet the criteria, the final pain level and confidence index are output; if they still do not meet the criteria, the process is returned to adjust the parameters and reprocess.
[0013] A second aspect of this application is also proposed.
[0014] The first aspect of this plan brings the following benefits: This application acquires facial video streams and aligns them frame by frame, overcoming the limitations of static image analysis. It extracts enhanced micro-expression temporal segments and multi-scale spatiotemporal features, accurately capturing facial dynamics and temporal changes. By combining muscle dynamics models and attention-weighted fusion, it strengthens feature correlation. Through double consistency verification and closed-loop optimization, it significantly improves detection accuracy, effectively solves the pain points of existing technologies, and meets the needs of practical applications. Attached Figure Description
[0015] Figure 1 This is a flowchart illustrating a pain expression detection method according to an embodiment of this application; Figure 2 This is a schematic diagram of the structure of a pain expression detection system according to an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a computer device according to an embodiment of this application; The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0017] Those skilled in the art will understand that, unless explicitly stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in the specification of this application means the presence of features, integers, steps, operations, elements, modules, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, modules, components, and / or groups thereof. It should be understood that when an element is “connected” or “coupled” to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein may include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any modules and all combinations of one or more associated listed items.
[0018] Those skilled in the art will understand that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.
[0019] Reference Figure 1 This application provides a method for detecting pain expression, including: S1: Acquire facial video stream, perform facial key point localization and inter-frame alignment, divide facial muscle regions and extract texture features to obtain aligned frame sequence and muscle region texture feature set; S2: Calculate motion energy based on texture feature set, extract and enhance micro-expression temporal segments from aligned frame sequence, extract multi-scale spatiotemporal features, construct facial muscle dynamics model and complete state estimation, and obtain enhanced micro-expression segments, multi-scale spatiotemporal features and muscle state vectors; S3: Based on multi-scale spatiotemporal features, perform spatiotemporal attention weighted fusion to obtain weighted spatiotemporal features, combine muscle state vectors to analyze muscle synergy, and generate an adaptive candidate spatiotemporal window set; S4: After dynamic fusion of weighted spatiotemporal features, time-series modeling is used to obtain the probability distribution of pain level. Enhanced micro-expression fragment features are fused to obtain multi-scale micro-expression fusion features, and the optimal spatiotemporal window is selected. S5: Verify muscle dynamic consistency based on muscle state vector, verify micro-expression temporal consistency based on multi-scale micro-expression fusion features, and verify the validity of the optimal spatiotemporal window. If any verification fails, return to the corresponding step for reprocessing. If all verifications pass, output the final pain level.
[0020] In step S1, the core is to acquire a bimodal facial video stream and complete facial key point localization and inter-frame alignment. After segmenting refined muscle regions, fusion differential texture features are extracted to obtain aligned frame sequences and muscle region texture feature sets. This is the basic data acquisition and preprocessing step in the entire pain expression detection process, providing standardized and personalized feature data sources for subsequent motion energy calculation, muscle dynamics modeling, and full-process feature analysis in S2. First, a bimodal acquisition benchmark system needs to be established: acquire a bimodal video stream of visible light + near-infrared light on the face, simultaneously acquire a 10-second resting state video segment, extract multi-scale fusion features to construct a fusion-type individual muscle feature baseline, pre-train a muscle dynamics model that fuses pain attention, output multi-scale pre-motion trajectories of 12 muscle regions such as the frontalis and zygomaticus major, construct a 3-layer image pyramid for cross-modal early fusion of the bimodal video stream, and obtain an initial set of 68 facial key points through scale-weighted initial localization. Verify that the cross-modal matching degree is ≥0.7 and the multi-scale consistency is ≥0.8. After achieving the standards, a set of personalized bimodal key point coordinates with motion prediction is obtained. Next, actual feature processing and extraction were carried out: based on the coordinate set, a dual-modal fusion weight of 0.6 for visible light and 0.4 for near-infrared light was calculated, and inter-frame predictive joint alignment was performed. Optical flow field features of adjacent frames (modulus length ≥ 0.45) were extracted to achieve cross-frame association of key points, compensate for micro-expression motion deviations, and obtain a cross-modal associated predictive aligned frame sequence. Combining facial anatomical atlas and model state parameters, the face was divided into three functional groups (upper, middle, and lower) with a total of 12 refined muscle regions. Dual-modal texture features were extracted and differiated with the resting baseline. The cross-frame associated texture features were fused to obtain a fused differential muscle region texture feature set. The feature and model prediction deviation was verified to be ≤ 0.1 and the cross-frame similarity ≥ 0.8. After meeting the standards, S1 processing was completed. The obtained feature set will be directly used as the core input for S2 motion energy calculation and muscle dynamics modeling, forming a closed loop. S1, through dual-modal acquisition, personalized baseline construction, and refined muscle region division, provides highly robust standardized feature data for subsequent detection steps, effectively compensates for micro-expression motion deviations, and ensures the effectiveness and uniqueness of the features.
[0021] In step S2, the core is to calculate motion energy based on the muscle region texture feature set output from S1, extract and enhance micro-expression temporal segments, extract multi-scale spatiotemporal features, and simultaneously construct a facial muscle dynamics model to complete state estimation. This is the core link in the feature enhancement and dynamic modeling of pain expression detection, providing key features and quantitative muscle state basis for the spatiotemporal attention fusion and muscle coordination analysis in S3. First, based on the fusion differential texture feature set from S1, a three-layer multi-scale feature pyramid is constructed with a bottom layer of 2×2, a middle layer of 4×4, and a top layer of 8×8. The motion energy of visible light and near-infrared dual modes at each scale is calculated, and the optical flow field features extracted from S1 (modulus length ≥ 0.45) are fused to obtain an optical flow-enhanced multi-scale motion energy set. Combined with the spatial attention mechanism, regions such as the zygomaticus major muscle and the corrugator supercilii muscle with a weight ≥ 0.7 are identified as key pain regions. Next, based on the motion energy set, micro-expression temporal segments containing peak frames are extracted hierarchically. The bottom layer takes high-energy frames with energy ≥0.6, and the top layer takes medium-energy frames with energy 0.3-0.6. The top 10% of peak frames in the sequence are retained. Key features are enhanced by an attention enhancement network. Then, spatial features, temporal dependency features, and optical flow temporal features at each scale are extracted and fused into multi-scale spatiotemporal features. Combining the dual-modal fusion weights of S1 (0.6 for visible light and 0.4 for near-infrared light) and attention weights, a multi-scale adapted facial muscle dynamics model is constructed. The state transition matrix and noise covariance matrix are optimized. The texture feature set of S1 is input into the model to complete the state estimation, and muscle state vectors of 12 muscle regions are obtained. The cross-scale feature consistency is verified to be ≥0.8 and the estimation accuracy is ≥0.9. After meeting the standards, the enhanced micro-expression segments, multi-scale spatiotemporal features, and muscle state vectors are output, which are directly used as the core inputs for the spatiotemporal attention weighted fusion and muscle synergy analysis of S3, forming a closed loop process. S2 achieves noise reduction enhancement of micro-expression features and precise quantification of muscle state, highlighting key spatiotemporal features of pain and providing a reliable kinetic basis for subsequent muscle synergy analysis.
[0022] In step S3, the core of step S3 is to perform spatiotemporal attention weighted fusion based on the multi-scale spatiotemporal features output by S2 to obtain weighted spatiotemporal features. Combined with the muscle state vector, the facial muscle coordination relationship is analyzed to generate an adaptive candidate spatiotemporal window set. This is the key link between feature fusion and window selection, providing a weighted feature basis and multiple sets of candidate window samples for the temporal modeling and optimal window selection in S4. First, the multi-scale spatiotemporal features of S2 are decomposed into multi-scale channels, dividing the feature channels into four groups. Spatial attention weights are calculated independently for each group. Spatial attention maps are generated using 7×7 convolutional kernels. Binarization is performed using the feature mean plus 1.5 times the standard deviation as a dynamic threshold to identify pain-related regions such as the corrugator supercilii and depressor anguli oris muscles with weights ≥0.7, thus obtaining spatially weighted features. Simultaneously, a multi-head self-attention module with eight attention heads and a dimension of 64 is constructed for the temporal convolutional network features. This module is paired with a causal convolutional module with three layers of kernels of 3, 5, and 7. The outputs of the two modules are weighted and fused to obtain temporally weighted features. Finally, the weighted spatiotemporal features are obtained by fusing the spatial and temporal weighted features. Next, based on the state vectors of the 12 muscle regions in S2, a second-order regularized vector autoregression model is constructed. Granger causality is determined with a significance level of 0.05 and an F-test ≥ 3.89, and a sparse muscle co-operation graph is constructed. According to the facial muscle classification rules in S1, the 12 muscles are divided into three functional groups: upper, middle, and lower. The muscle cross-correlation function within each group is calculated (maximum latency 10 frames). The group with an average synchronicity ≥ 0.7 is marked as a strong co-operation functional group. Then, a directed muscle propagation graph is constructed to screen critical propagation paths with a path length < 3. Finally, based on the enhanced micro-expression fragments and muscle co-operation graph information in S2, a greedy algorithm is used to generate 5 candidate windows of 10-60 frames. Each window must contain at least one micro-expression fragment, ≥ 1 strong co-operation group, and ≥ 2 critical paths. The window quality score is calculated to be ≥ 0.6 to meet the standard. The weighted spatiotemporal features and the candidate window set are output as the core input for the S4 temporal modeling, forming a closed-loop process. S3 achieves precise weighted fusion of spatiotemporal features, quantifies the synergistic relationship and propagation path between muscles, and the generated candidate windows provide a scientific and reliable sample basis for the optimal selection of S4.
[0023] In step S4, the core is to dynamically fuse the weighted spatiotemporal features output from S3 and then perform temporal modeling to obtain the pain level probability distribution. This involves fusing the enhanced micro-expression fragment features from S2 to obtain multi-scale micro-expression fusion features, and selecting the optimal spatiotemporal window from the candidate windows. This is the core step in pain level quantification and precise window selection, providing a clear verification object for the consistency verification in S5. First, based on the 12 muscle region state vectors from S2, the pain contribution weights of key pain areas such as the corrugator supercilii and depressor anguli oris muscles are calculated. Combined with the facial muscle motion energy curve, the weighted spatiotemporal features from S3 are subjected to adaptive weighting processing. Simultaneously, five adaptive candidate spatiotemporal windows of 10-60 frames generated in S3 are retrieved to obtain the weighted adaptive spatiotemporal features. Next, this feature is input into a temporal modeling network. Hierarchical local temporal modeling is performed according to the window modeling configuration. First, residual supplementation, spatiotemporal decoupling, and pain semantic alignment are performed on the feature. The muscle activation intensity of the candidate window is calculated by combining the muscle state vector. The modeling depth is determined according to the motion energy gradient. After extracting the depth temporal features by window, they are integrated and residual correction is performed. Then, multi-scale attention weighted aggregation and global temporal regularization optimization are performed to output the probability distribution sequence of four pain levels: no pain, mild, moderate, and severe. Then, the intensity spectrum is extracted by performing temporal Fourier transform on the enhanced micro-expression fragment features of S2. Multi-scale feature fusion is performed based on the pain level probability distribution sequence and the intensity spectrum to obtain multi-scale micro-expression fusion features. Finally, based on the confidence ranking of the probability distribution sequence (high confidence ≥ 0.8) and the energy distribution of the intensity spectrum (core energy ratio ≥ 60%), the optimal spatiotemporal window is selected from 5 candidate windows. This window will serve as the sole object for verifying the dynamics, temporal consistency, and validity of S5, forming a closed loop process. S4 quantifies pain levels and integrates micro-expression features to improve the accuracy of window filtering, providing a clear and high-quality verification object for S5's consistency verification.
[0024] In step S5, the core is to verify the consistency of muscle dynamics based on the muscle state vectors in S2, verify the temporal consistency of micro-expressions based on the multi-scale micro-expression fusion features in S4, and verify the effectiveness of the optimal spatiotemporal window selected in S4. After closed-loop optimization, the final pain level is output. This is the final verification and result output link of the entire pain expression detection process, which determines the accuracy and reliability of the detection results. First, muscle dynamic consistency is verified based on the 12 facial muscle region state vectors output by S2. Kalman filtering is used to calculate the muscle movement prediction error. Combined with the muscle co-graph in S3, a graph neural network is constructed and node embeddings are learned. The prediction error matrix and node embedding smoothness are calculated. A joint consistency index is constructed according to preset rules, and a score threshold of ≥0.6 is set to output the corresponding joint consistency score. Next, based on the multi-scale micro-expression fusion features of S4, the temporal consistency of micro-expressions is verified. A Bayesian Hidden Markov Joint Probability Model with 5 temporal states is constructed. The posterior distribution is estimated through variational inference, and a recurrent generative adversarial network is used to generate real micro-expression samples. A comprehensive temporal consistency index is constructed by fusing observational likelihood, diffusion loss, etc., with a threshold ≥ 0.7, and the temporal consistency score is output. Then, the effectiveness of the optimal spatiotemporal window selected by S4 is verified. The attention weight distribution of the graph neural network and the sample diversity of the GAN within the window are calculated. The window quality score is obtained by fusing attention entropy and diversity index, with a threshold ≥ 0.6. A comprehensive consistency verification framework is constructed. Verification passes if all three scores meet the criteria. If the muscle dynamics consistency verification fails, the process returns to the step of building the facial muscle dynamics model for reprocessing. If the micro-expression temporal consistency verification fails, the process returns to the step of extracting and enhancing micro-expression temporal segments for reprocessing. If the optimal spatiotemporal window validity verification fails, the process returns to the step of generating an adaptive candidate spatiotemporal window set for reprocessing. If any one fails, joint feedback optimization is performed, updating the state transition matrix of the muscle dynamics model using gradient descent and updating the Bayesian model parameters using the expectation-maximization algorithm, recalculating all indicators until the criteria are met. After successful verification, based on the pain level probability distribution sequence in S4, the final pain level and confidence index are output, such as moderate pain with a confidence level of 0.88, forming a closed-loop detection process. S5, through triple consistency verification and closed-loop feedback optimization, verifies the detection results from multiple dimensions, significantly improving the accuracy of pain detection and ensuring the reliability and precision of the output results.
[0025] In one embodiment, the steps of acquiring a facial video stream, performing facial landmark localization and inter-frame alignment, segmenting facial muscle regions and extracting texture features to obtain an aligned frame sequence and a set of texture features of muscle regions include: S10: Collect facial bimodal video streams and resting segments, construct individual muscle feature baselines and pre-train muscle dynamics models, combine trajectories and baselines to locate facial key points in the video streams, and obtain a set of bimodal key point coordinates; S11: Calculate the dual-modal fusion weights based on the coordinate set, perform inter-frame predictive joint alignment, and extract optical flow field features to realize cross-frame association of key points, thereby obtaining the predicted aligned frame sequence; S12: Based on the aligned frame sequence, combined with facial anatomical atlas and muscle dynamics model state parameters, the muscle regions are divided into refined regions, and muscle region division results are generated. S13: Based on the segmentation results, extract bimodal texture features and subtract them from the individual muscle feature baselines. Then, fuse cross-frame associated texture features to obtain a set of texture features for the muscle region.
[0026] In this embodiment, firstly, a bimodal video stream and resting segment of the face are acquired, an individual muscle feature baseline is constructed and a muscle dynamics model is pre-trained, and the facial key point localization is completed by combining the trajectory and the baseline to obtain a set of bimodal key point coordinates. A dual-modal acquisition device using visible light and near-infrared light was employed to acquire facial video streams of the subjects at 30fps, simultaneously acquiring 10-second video clips of the subjects in a calm, expressionless state. Multi-scale fusion features of the calm clips were extracted to construct a dynamically updated fusion-based individual muscle feature baseline. A multi-scale muscle dynamics model fusing pain and attention was pre-trained. This pre-trained muscle dynamics model provides the initial parameters and state transition matrix for the subsequent construction of a facial muscle dynamics model. The multi-scale pre-motion trajectories and attention weights of 12 muscle regions, including the frontalis and corrugator supercilii, were output. Combining these trajectories, the individual baseline, and the attention weights, a three-layer image pyramid was constructed from the dual-modal video streams for cross-modal early fusion. Scale-weighted initial localization of 68 facial key points was achieved. After cross-scale fusion and cross-modal reverse mapping to verify and correct deviations, and optimization of the pain region localization accuracy, the cross-modal matching degree was verified to be ≥0.7 and the multi-scale consistency ≥0.8. Upon achieving these standards, a set of personalized dual-modal key point coordinates with motion prediction was obtained, providing a precise coordinate reference for subsequent inter-frame alignment.
[0027] Then, based on the coordinate set, the dual-modal fusion weight is calculated, and inter-frame predictive joint alignment is performed. Simultaneously, optical flow field features are extracted to achieve cross-frame association of key points, resulting in a predicted aligned frame sequence. Based on the feature sharpness and noise resistance of the dual-modal video stream, a visible light modal fusion weight of 0.6 and a near-infrared light modal fusion weight of 0.4 are calculated. Based on these weights, inter-frame predictive joint alignment of the dual-modal key point coordinate set is performed to eliminate inter-frame displacement deviation. Simultaneously, optical flow field features of adjacent frames are extracted, and features with an optical flow modulus ≥ 0.45 are enhanced to achieve cross-frame association of key points, effectively compensating for subtle movement deviations in micro-expressions and preserving the continuity of facial muscle movements. Finally, a cross-modal associated predicted aligned frame sequence is obtained. After verifying an inter-frame coherence of ≥ 0.8, the process proceeds to the next stage.
[0028] Secondly, based on the aligned frame sequence, combined with facial anatomical atlases and muscle dynamics model state parameters, refined muscle regions are divided, generating muscle region division results. Using the predicted aligned frame sequence as a basis, and referring to standard facial anatomical atlases, combined with the state parameters of the muscle dynamics model and the pre-motion trajectories of 12 muscle regions, facial muscles are divided into three functional groups according to their movement function: upper, middle, and lower. The upper group includes four muscle regions such as the frontalis and corrugator supercilii; the middle group includes four muscle regions such as the zygomaticus major and buccinator; and the lower group includes four muscle regions such as the depressor anguli oris and orbicularis oris. This completes the division of 12 refined muscle regions, adapting to individual muscle distribution and matching muscle movement trajectories, ensuring that the division results are consistent with the actual muscle movement patterns.
[0029] Finally, based on the segmentation results, bimodal texture features are extracted and differiated from the individual muscle feature baseline. Cross-frame associated texture features are then fused to obtain a set of muscle region texture features. For the segmentation results of 12 refined muscle regions, visible light blood oxygenation texture features and near-infrared temperature texture features are extracted for each region. These two types of texture features are then differiated from the previously constructed individual muscle feature baseline to eliminate interference from individual basic facial features. The differiated bimodal texture features are then weighted and fused with the cross-frame associated texture features to generate fused differential texture features. The feature deviation from the model prediction is verified to be ≤0.1, and the cross-frame similarity is ≥0.8. After meeting these standards, a set of muscle region texture features is obtained. The muscle region texture feature set obtained by differiating the bimodal texture features from the individual muscle feature baseline, along with the predicted aligned frame sequence, serves as the core input for motion energy calculation and micro-expression fragment extraction in step S2, forming a closed loop. This embodiment achieves standardized preprocessing of bimodal video streams, constructs personalized muscle baselines, accurately segments muscle regions, and provides highly robust feature data for subsequent detection.
[0030] In one embodiment, the steps of acquiring a bimodal facial video stream and resting segments, constructing an individual muscle feature baseline and pre-training a muscle dynamics model, and combining the trajectory and baseline to perform facial key point localization on the video stream to obtain a set of bimodal key point coordinates include: S101: Acquire facial dual-modal video streams and multi-time resting segments, extract multi-scale fusion features from resting segments, construct a fusion-type individual muscle feature baseline, pre-train a muscle dynamics model, and output muscle pre-motion trajectories and attention weights; S102: Combining baseline, trajectory and attention weights, construct a multi-scale image pyramid for the video stream and perform cross-modal early fusion. Scale-weighted initial localization is used to obtain the initial set of key points. S103: Perform cross-scale fusion and cross-modal reverse mapping verification on the initial key point set, correct the fusion and scale bias, and then optimize the pain area localization accuracy based on attention weights to obtain the optimized key point set. S104: Verify the cross-modal matching degree and multi-scale consistency of the optimized key points. If the standard is not met, return to adjust the fusion weight and sub-scale parameters. If the standard is met, obtain the set of bimodal personalized key point coordinates.
[0031] In this embodiment, firstly, a dual-modal facial video stream and multiple resting segments are acquired. Multi-scale fusion features of the resting segments are extracted to construct a fusion-based individual muscle feature baseline. A muscle dynamics model is pre-trained, and muscle pre-motion trajectories and attention weights are output. A visible light + near-infrared light dual-modal acquisition device is used to acquire the subject's facial video stream at a frame rate of 30fps. Three 10-second expressionless resting segments are simultaneously acquired, covering different slight head postures. Texture, blood oxygen, and temperature fusion features of the resting segments at 2×2, 4×4, and 8×8 scales are extracted. A fusion-based individual muscle feature baseline adapted to the subject is constructed through feature normalization and dynamic weighting. At the same time, a multi-scale muscle dynamics model that integrates pain attention is pre-trained. For 12 facial muscle regions, such as the frontalis and corrugator supercilii, the multi-scale muscle pre-motion trajectories of each region are output, as well as pain attention weights—weights for key pain areas such as corrugator supercilii and depressor anguli oris are ≥0.7, and weights for non-core areas such as zygomaticus major and buccinator are 0.3-0.6, providing personalized benchmarks and weights for subsequent key point localization. Then, combining baseline, trajectory, and attention weights, a multi-scale image pyramid is constructed for the video stream, and cross-modal early fusion is performed. Initial localization is obtained by weighted scaling at different scales to obtain an initial set of key points. Based on the constructed fusion-type individual muscle feature baseline, the pre-motion trajectories of 12 muscle regions, and attention weights, a three-layer multi-scale image pyramid (bottom 2×2, middle 4×4, top 8×8) is constructed for the bimodal video stream. Weighting coefficients are assigned according to the feature sharpness at each scale (bottom 0.5, middle 0.3, top 0.2). Cross-modal early fusion is performed for the visible light and near-infrared modalities. Then, combined with pain attention weights, the fused feature map is weighted at different scales. Initial localization is completed for the anatomical locations of 68 facial key points, obtaining an initial set of coordinates for the 68 key points, thus achieving preliminary spatial localization of the key points.
[0032] Secondly, cross-scale fusion and cross-modal back-mapping verification were performed on the initial keypoint set to correct fusion and scale bias. Then, the localization accuracy of the pain region was optimized based on attention weights to obtain the optimized keypoint set. Cross-scale fusion was performed on the initial keypoint sets at three scales using a weighted average method to eliminate localization bias between scales. At the same time, cross-modal back-mapping verification was performed to calculate the coordinate bias of the corresponding keypoints in the visible light and near-infrared modes. If the bias was greater than 2 pixels, it was corrected by coordinate interpolation to complete the overall correction of fusion and scale bias. Then, based on a pain attention weight of ≥0.7, the coordinates of keypoints around the corrugator supercilii and depressor anguli oris muscles were finely adjusted to improve the localization accuracy of pain-related areas. Finally, the optimized set of 68 facial keypoint coordinates was obtained.
[0033] Finally, the cross-modal matching degree and multi-scale consistency of the optimized keypoints are verified. If the standards are not met, the process returns to adjust the fusion weights and scale parameters. If the standards are met, a bimodal personalized keypoint coordinate set is obtained. The verification indicators are set as cross-modal matching degree ≥ 0.7 and multi-scale consistency ≥ 0.8. The optimized keypoint set is quantitatively verified. If either indicator is not met, the process returns to adjust the bimodal fusion weights (visible light, near-infrared) and the scale weighting coefficients of the multi-scale image pyramid, and the localization and optimization are completed again. If both indicators are met, the final bimodal personalized keypoint coordinate set with motion prediction is determined. This set serves as the core coordinate reference for subsequent inter-frame predictive joint alignment, forming a closed loop with the previous inter-frame alignment steps. This embodiment improves the accuracy and personalized adaptability of keypoint localization through multi-time resting baseline and multi-scale localization optimization, providing a reliable coordinate reference for subsequent preprocessing.
[0034] In one embodiment, the steps of calculating motion energy based on texture feature sets, extracting and enhancing micro-expression temporal segments from aligned frame sequences, extracting multi-scale spatiotemporal features, constructing a facial muscle dynamics model and completing state estimation to obtain enhanced micro-expression segments, multi-scale spatiotemporal features, and muscle state vectors include: S20: Construct a multi-scale feature pyramid based on the texture feature set, calculate the motion energy at each scale and fuse the optical flow field features to obtain a multi-scale motion energy set with enhanced optical flow. S21: Combine spatial attention mechanism to locate key pain areas, extract micro-expression time segments containing peak frames based on the hierarchical extraction of motion energy sets, and enhance key features through attention enhancement network to obtain multi-scale enhanced micro-expression segments; S22: Extract spatial features, temporal dependence features, and optical flow temporal features from the fragments and fuse them to form multi-scale spatiotemporal features; S23: Combine bimodal information with attention weights to construct a facial muscle dynamics model and optimize the state transition matrix and noise covariance matrix; S24: Input the texture feature set to complete the state estimation and obtain the muscle state vector.
[0035] In this embodiment, a multi-scale feature pyramid is first constructed based on the texture feature set. Motion energy at each scale is calculated and fused with optical flow field features to obtain an optically enhanced multi-scale motion energy set. Based on the fused differential texture feature set of 12 facial muscle regions output by S1, a three-layer multi-scale feature pyramid is constructed: a bottom layer of 2×2, a middle layer of 4×4, and a top layer of 8×8. Weights are allocated as follows: bottom layer 0.5, middle layer 0.3, and top layer 0.2. Muscle motion energy in both visible light (fusion weight 0.6) and near-infrared light (fusion weight 0.4) modes at each scale is calculated. Then, the optical flow field features extracted from S1 (modulus length ≥ 0.45) are fused. The energy features of subtle micro-expression movements are compensated through weighted superposition, ultimately obtaining an optically enhanced multi-scale motion energy set. After verifying that the signal-to-noise ratio of the set is ≥ 30dB, it provides energy feature basis for subsequent micro-expression fragment extraction. Then, a spatial attention mechanism is used to locate key pain areas. Based on the motion energy set, micro-expression temporal fragments containing peak frames are extracted hierarchically. Key features are enhanced through an attention enhancement network to obtain multi-scale enhanced micro-expression fragments. Spatial attention mechanisms were used to calculate the attention weights of each muscle region. Regions with weights ≥0.7, such as the corrugator supercilii, depressor anguli oris, and frontalis muscles, were identified as key pain regions. Based on a multi-scale motion energy set, hierarchical extraction was performed. The bottom pyramid extracted high-energy frames with energy ≥0.6, and the top pyramid extracted medium-energy frames with energy ≥0.3-0.6. The top 10% of peak frames in each sequence were retained. For short segments, the first 3 core frames were retained, and for long segments, the first 5 core frames were retained, resulting in a multi-scale hierarchical segment set. This set was input into an attention enhancement network, and resources were allocated according to "spatial weight × temporal weight × optical flow modulus × hierarchical weight". Regions with a product ≥0.6 were enhanced with a coefficient of 1.3, and regions <0.3 were suppressed for noise. After verifying that the inter-frame coherence of the segments was ≥0.8 and the multi-scale feature coverage was ≥80%, multi-scale enhanced micro-expression segments were obtained.
[0036] Secondly, spatial features, temporal dependency features, and optical flow temporal features at various scales are extracted from the fragments and fused to form multi-scale spatiotemporal features. For multi-scale enhanced micro-expression fragments, spatial texture features of muscle regions are extracted through convolution operations at 2×2, 4×4, and 8×8 scales, temporal dependency features between frames are extracted through a multi-head self-attention module, and optical flow temporal features of muscle motion are extracted through optical flow field sequences. Then, the three types of features are weighted and fused according to the weights of each scale to highlight the spatiotemporal correlation features of key pain areas. Finally, a multi-scale spatiotemporal feature containing spatial, temporal, and motion dimensions is formed, that is, the features of each scale are fused with a weight of 0.5 for the bottom layer, 0.3 for the middle layer, and 0.2 for the top layer to form a multi-scale spatiotemporal feature, providing a core feature source for subsequent spatiotemporal attention fusion. Furthermore, by combining bimodal information and attention weights, a facial muscle dynamics model is constructed, and the state transition matrix and noise covariance matrix are optimized. Combining the dual-modal fusion weights of S1 (0.6 for visible light and 0.4 for near-infrared) with the attention weights for pain regions in this step, a multi-scale adaptive facial muscle dynamics model is constructed for 12 facial muscle regions. Based on the temporal pattern of muscle movement, the gradient descent method is used to optimize the state transition matrix of the model and reduce the bias of muscle movement state prediction. At the same time, based on the deviation value between texture features and model prediction, the noise covariance matrix is optimized to improve the noise resistance of the model and make the model more suitable for the individual muscle movement pattern of the test subject.
[0037] Finally, the texture feature set is input to complete the state estimation, resulting in a muscle state vector. The fused differential texture feature set of the 12 muscle regions output from S1 is input into the optimized facial muscle dynamics model to quantify and estimate the state parameters such as motion amplitude, motion rate, and contraction frequency of each muscle region, outputting a multi-dimensional muscle state vector for each muscle region. The model verifies that cross-scale feature consistency is ≥0.8 and model estimation accuracy is ≥0.9. Once these standards are met, the final muscle state vector is determined. This vector, along with multi-scale spatiotemporal features and enhanced micro-expression fragments, serves as the core input for S3 spatiotemporal attention-weighted fusion and muscle co-analysis, forming a closed-loop process. This embodiment achieves noise reduction enhancement of micro-expression features and multi-scale spatiotemporal feature extraction, accurately quantifying muscle movement states and providing reliable features and model basis for subsequent analysis.
[0038] In one embodiment, the steps of locating key pain regions using spatial attention mechanisms, extracting micro-expression temporal segments containing peak frames based on a hierarchical extraction of motion energy sets, and enhancing key features through an attention enhancement network to obtain multi-scale enhanced micro-expression segments include: S210: Construct a multi-level attention pyramid, integrate spatial attention and temporal attention and superimpose optical flow field features to locate key pain areas and output a multi-scale spatiotemporal-optical flow fusion weight matrix; S211: Based on the fusion weight matrix, extract micro-expression time segments according to the pyramid hierarchy, select different energy frames in different regions, retain peak frames and core frames, and output a multi-scale hierarchical segment set. S212: Input the multi-scale hierarchical segment set into the attention enhancement network, allocate network resources according to the product of multi-dimensional weights, enhance key region features and suppress noisy region features, and output fused enhanced segments; S213: Verify the inter-frame coherence and multi-scale feature coverage of the fused enhancement segment. If both criteria are met, the fused enhancement segment is determined to be a multi-scale enhanced micro-expression segment. If the criteria are not met, return to adjust the optical flow fusion coefficient and hierarchical weight for reprocessing.
[0039] In this embodiment, a multi-level attention pyramid is first constructed, fusing spatial attention and temporal attention and superimposing optical flow field features to locate key pain regions and output a multi-scale spatiotemporal-optical flow fusion weight matrix. Based on the previously described 3-layer attention pyramid architecture (bottom layer 2×2, middle layer 4×4, top layer 8×8), for the texture features and motion energy features of 12 facial muscle regions, the spatial attention weights of visible light and near-infrared dual modes are first calculated. Regions with a consistent region (difference ≤ 0.1) and a weight ≥ 0.7 are selected, and regions with a difference > 0.1 are assigned a high-weight mode with a weight ≥ 0.7. At the same time, the weight of regions with a muscle coordination coefficient ≥ 0.6 is increased by 0.1 to initially determine key pain regions. Then, the temporal parameters are dynamically updated in a 30-second window, and the inter-frame energy changes of consistent coordination regions are analyzed. With a rate ≥ 0.55 and a single difference region ≥ 0.45, the optical flow field features extracted earlier (illumination / motion blur scene modulus ≥ 0.45) are superimposed. For consistent regions, the fusion value is calculated according to "bimodal average weight × 0.5 + temporal weight × 0.25 + optical flow modulus × 0.15 + synergy coefficient × 0.1". For differing regions, the calculation is dominated by the high-weight mode. After verifying that the bimodal consistency of the fusion matrix is ≥ 0.7, the synergy consistency is ≥ 0.7, and the signal-to-noise ratio is ≥ 30dB, a multi-scale spatiotemporal-optical flow fusion weight matrix is output, providing accurate weight basis for fragment extraction.
[0040] Then, based on the fusion weight matrix, micro-expression time-series segments are extracted hierarchically according to the pyramid structure. Different energy frames are selected for different regions, and peak frames and core frames are retained in all cases, outputting a multi-scale hierarchical segment set. Referring to the hierarchical weight allocation of the fusion weight matrix (bottom layer 0.5, middle layer 0.3, top layer 0.2), differentiated segment extraction is performed for each level. The bottom pyramid selects high-energy frames with motion energy ≥0.6 for key pain areas, and the top pyramid selects medium-energy frames with energy between 0.3 and 0.6. Non-key areas are processed by lowering the energy threshold according to the hierarchical weight. The top 10% of peak frames of the sequence are retained in all levels. For short micro-expression segments with a duration ≤10 frames, the first 3 core frames are retained, and for long micro-expression segments with a duration >10 frames, the first 5 core frames are retained. Frame retention rate is improved for key pain areas such as the corrugator supercilii and depressor anguli oris muscles. Finally, a multi-scale hierarchical segment set covering each level and muscle region is obtained. Next, the multi-scale hierarchical fragment set is input into the attention enhancement network. Network resources are allocated according to the product of multi-dimensional weights to enhance key region features and suppress noisy region features, outputting a fused enhanced fragment. The multi-scale hierarchical fragment set is input into the attention enhancement network, and network computing resources are allocated according to the product of multi-dimensional weights: "spatial attention weight × temporal attention weight × optical flow modulus × hierarchical weight". Features of key pain regions with a product ≥ 0.6 are enhanced with an enhancement coefficient of 1.3. Features of non-key regions and noisy regions with a product < 0.3 are suppressed. Features of transition regions with a product of 0.3-0.6 are smoothed. Emphasis is placed on enhancing the micro-expression features of core pain regions such as the corrugator supercilii, frontalis, and depressor anguli oris muscles, while weakening background noise in irrelevant facial areas. After feature enhancement, a fused enhanced fragment is output.
[0041] Finally, the inter-frame coherence and multi-scale feature coverage of the fused and enhanced segment are verified. If both meet the criteria, the fused and enhanced segment is determined to be a multi-scale enhanced micro-expression segment; otherwise, the process returns to adjust the optical flow fusion coefficient and hierarchical weights for reprocessing. The verification quantification indicators are set as an average inter-frame coherence similarity ≥ 0.8 and a multi-scale feature coverage ≥ 80%. A full-dimensional verification is performed on the fused and enhanced segment. If both indicators meet the criteria, the segment is directly determined to be a multi-scale enhanced micro-expression segment; if either indicator fails to meet the criteria, the process returns to the previous steps of adjusting the optical flow field feature fusion coefficient and pyramid hierarchical weights, and the pyramid construction, segment extraction, and feature enhancement processes are repeated. The finally qualified multi-scale enhanced micro-expression segment will serve as the core input for subsequent extraction of multi-scale spatiotemporal features, forming a closed loop with the previously mentioned motion energy set and subsequent feature extraction steps. This embodiment achieves precise pain region localization through multi-dimensional attention fusion, hierarchically extracts and enhances micro-expression features, effectively suppresses noise, and improves the effectiveness and feature recognition of micro-expression segments.
[0042] In one embodiment, the steps of constructing a multi-level attention pyramid, fusing spatial attention and temporal attention, superimposing optical flow field features, locating key pain regions, and outputting a multi-scale spatiotemporal-optical flow fusion weight matrix include: S2101: Construct a multi-level attention pyramid, calculate bimodal spatial attention weights, take the average weight for consistent regions and the high-weight modal weight for dissimilar regions, and adjust the corresponding region weights in combination with the muscle synergy coefficient to complete the determination of key pain regions. S2102: Dynamically update temporal parameters according to the temporal window, extract temporal features based on the type of key pain region, extract optical flow features in noisy scenes, and obtain a set of noise-resistant temporal and optical flow features; S2103: Combining bimodal spatial attention weights, noise-resistant temporal and optical flow feature sets, and muscle synergy features, consistent regions are calculated according to preset rules for fusion values, while different regions are fused according to the high-weight mode, forming a multi-scale spatiotemporal-optical flow fusion matrix.
[0043] In this embodiment, a multi-level attention pyramid is first constructed, and bimodal spatial attention weights are calculated. Consistent regions are assigned average weights, while discrepancy regions are assigned high-weight modal weights. The weights of corresponding regions are adjusted based on muscle synergy coefficients to determine the key pain areas. Based on the previously established 3-layer attention pyramid architecture (bottom layer 2×2, middle layer 4×4, top layer 8×8), for 12 facial muscle regions covering the upper, middle, and lower functional groups, bimodal spatial attention weights for visible light (fusion weight 0.6) and near-infrared light (fusion weight 0.4) are calculated. Modal consistency is first determined by pixel feature similarity. Regions with a bimodal weight difference ≤ 0.1 are considered consistent regions, and the average bimodal weight is used as the base weight for these regions. Regions with a difference > 0.1 are considered discrepancy regions, and the higher-weight modal weight with clearer features is directly selected as the base weight. Combining the muscle synergy coefficients calculated earlier, a weight increment of 0.1 is added to the basic weights of muscle regions with synergy coefficients ≥0.6 (such as corrugator supercilii and depressor supercilii, depressor anguli oris and mentalis). Finally, regions with weights ≥0.7 are identified as key pain regions, mainly including core regions such as corrugator supercilii, depressor anguli oris, and frontalis muscle, thus defining a precise scope of focus for subsequent feature extraction.
[0044] Then, the temporal parameters are dynamically updated according to the temporal window, and temporal features are extracted based on the type of key pain regions. Optical flow features are extracted in noisy scenes to obtain a set of noise-resistant temporal and optical flow features. A 30-second sliding temporal window is set to dynamically update the temporal parameters of the video stream. Based on the modal attributes of the key pain regions, consistent and cooperative regions and differentially isolated regions are divided. For consistent and cooperative regions, the inter-frame energy change rate threshold is set to ≥0.55, and for differentially isolated regions, the threshold is lowered to ≥0.45. Temporal motion features of each region are extracted according to the corresponding threshold. For lighting noise and motion blur noise scenes that are prone to occur in actual acquisition, the optical flow field features extracted above are retrieved, and the optical flow modulus threshold is increased to ≥0.45 for noise-resistant screening. Optical flow temporal features that meet the threshold are extracted. Finally, the temporal features of different types of key regions are integrated with the noise-resistant optical flow features to obtain a set of noise-resistant temporal and optical flow features, which effectively avoids the interference of environmental noise on feature extraction.
[0045] Finally, combining the bimodal spatial attention weights, noise-resistant temporal and optical flow feature sets, and muscle coordination features, the fusion value is calculated according to preset rules for consistent regions, while the fusion of dissimilar regions is dominated by the high-weight mode, forming a multi-scale spatiotemporal-optical flow fusion matrix. First, the preset calculation rules for feature fusion are defined. For consistent regions, the fusion value calculation formula is: bimodal average weight × 0.5 + temporal attention weight × 0.25 + optical flow modulus × 0.15 + muscle coordination coefficient × 0.1. For dissimilar regions, the spatial attention weight of the high-weight mode is used as the core, and the fusion value is obtained by successively multiplying it by the temporal attention weight, optical flow modulus, and muscle coordination coefficient of that region. Following a three-tiered attention pyramid, fusion values are calculated for each muscle region in the bottom, middle, and top layers according to corresponding rules. Fixed weights are assigned to each layer (bottom 0.5, middle 0.3, top 0.2) for weighted fusion. The fusion values of each layer and region are arranged in a matrix to generate a multi-scale spatiotemporal-optical flow fusion matrix. The matrix is verified to have bimodal consistency ≥ 0.7, coherence consistency ≥ 0.7, and signal-to-noise ratio ≥ 30dB before the final result is determined. This matrix will serve as the core weight basis for subsequent hierarchical extraction of micro-expression time-series fragments, forming a complete workflow loop with the previously described motion energy set and subsequent fragment extraction steps. This embodiment accurately locates key pain areas through multi-dimensional weight calculation and noise-resistant feature extraction. The generated fusion matrix provides reliable weights for micro-expression extraction, improving the targeting and noise resistance of feature selection.
[0046] In one embodiment, the step of generating an adaptive candidate spatiotemporal window set by performing spatiotemporal attention weighted fusion based on multi-scale spatiotemporal features to obtain weighted spatiotemporal features, combining muscle state vectors to analyze muscle synergy, includes: S30: Perform multi-scale channel decomposition on multi-scale spatiotemporal features, independently calculate spatial attention weights for each group and generate spatial attention maps, identify pain-related regions through dynamic threshold processing, and obtain spatial weighted features; S31: Construct a multi-head self-attention module and a causal convolution module for temporal features, and perform weighted fusion of the outputs of the two modules to obtain temporal weighted features. Then, fuse the spatial weighted features and temporal weighted features to obtain weighted spatiotemporal features. S32: Construct a vector autoregression model based on muscle state vectors, analyze the causal relationship between muscles and construct a muscle synergy graph, divide facial muscle function groups to calculate the synergy within the groups, construct a directed graph of muscle propagation and screen key propagation paths; S33: Generate candidate spatiotemporal windows based on micro-expression fragments, muscle coordination maps, functional group coordination degree and key propagation paths, calculate window quality scores and verify feature consistency. If the scores are met, they are determined as an adaptive candidate spatiotemporal window set; otherwise, return to adjust dynamic coefficients and reprocess.
[0047] In this embodiment, the multi-scale spatiotemporal features are first decomposed into multi-scale channels. Spatial attention weights are independently calculated for each group, and spatial attention maps are generated. Pain-related regions are identified through dynamic thresholding to obtain spatially weighted features. Based on the multi-scale spatiotemporal features of the upper, middle, and lower muscle function groups output by S2, the feature channels are divided into four independent groups. Spatial attention weights are independently calculated for each group of spatiotemporal features at scales of 2×2, 4×4, and 8×8. Spatial attention maps are generated by convolution operations on the features of each group using a 7×7 convolution kernel. The attention maps are then binarized using a dynamic threshold of "feature mean + 1.5 times the standard deviation". Regions such as the corrugator supercilii, depressor anguli oris, and frontalis muscles with attention weights ≥ 0.7 are identified as pain-related regions. The features of these regions are weighted and enhanced, while non-pain regions retain their basic weights. Finally, the processing results of the four groups are integrated to obtain spatially weighted features, laying the spatial dimension feature foundation for spatiotemporal feature fusion. Then, a multi-head self-attention module and a causal convolutional module are constructed for the temporal features. The outputs of the two modules are weighted and fused to obtain temporally weighted features. Then, spatial weighted features and temporally weighted features are fused to obtain weighted spatiotemporal features. For the temporal dimension features in the multi-scale spatiotemporal features, a multi-head self-attention module with 8 attention heads and 64 dimensions per head is constructed to capture long-distance temporal dependencies between frames. At the same time, a causal convolutional module with 3 convolutional kernels of sizes 3, 5, and 7 is constructed. Masking is used to ensure that the current time step depends only on historical information to extract local temporal features. The outputs of the multi-head self-attention module and the causal convolutional module are weighted and fused using trainable fusion weights to obtain temporally weighted features. Then, the spatial weighted features and temporally weighted features are deeply fused according to a 1:1 fusion weight for spatial and temporal dimensions to highlight the spatiotemporal correlation characteristics of pain-related areas. Finally, the weighted spatiotemporal features are obtained, providing the fused core features for subsequent muscle coordination analysis.
[0048] Secondly, a vector autoregression model was constructed based on muscle state vectors to analyze the causal relationships between muscles and to build a muscle synergy graph. Facial muscle functional groups were divided to calculate the synergy within each group, and a directed muscle propagation graph was constructed to screen key propagation paths. Based on the muscle state vectors of the 12 facial muscle regions output by S2, a second-order regularized vector autoregression model was constructed. Using a significance level of 0.05 and an F-test statistic ≥3.89 as the criteria, Granger causal relationships between muscles were analyzed, and muscle pairs with causal relationships were extracted to construct a sparse muscle synergy graph. The 12 muscle regions were divided into three functional groups (upper, middle, and lower) according to the rules mentioned above. The cross-correlation function of muscle pairs within each group was calculated (maximum latency 10 frames). Functional groups with an average synchronicity ≥0.7 within the group were marked as strongly synergistic functional groups. Based on the synergy strength of the muscle synergy graph, a directed muscle propagation graph was constructed. A breadth-first search algorithm was used to calculate the shortest path between nodes, and key propagation paths with a path length <3 were screened to complete the full-dimensional quantitative analysis of muscle synergy relationships.
[0049] Finally, candidate spatiotemporal windows are generated based on micro-expression fragments, muscle coordination maps, functional group coordination, and key propagation paths. Window quality scores are calculated, and feature consistency is verified. If the score meets the standard, the window is selected as an adaptive candidate spatiotemporal window set; otherwise, the process is returned to adjust the dynamic coefficients and reprocess. Using the multi-scale enhanced micro-expression fragments output from S2, the aforementioned sparse muscle coordination maps, strongly coordinated functional groups, and key propagation paths, a greedy algorithm generates five candidate spatiotemporal windows with durations ranging from 10 to 60 frames. Each window must contain at least one micro-expression fragment, one strongly coordinated functional group, and two key propagation paths. A window quality score is calculated by comprehensively considering pain probability, muscle movement amplitude, functional group coordination, and the number of key propagation paths, with a score threshold of ≥0.6. Feature consistency within the window is also verified. If both requirements are met, the window is selected as an adaptive candidate spatiotemporal window set; otherwise, the process is returned to adjust the dynamic coefficients of spatiotemporal feature fusion, and feature fusion and window generation are repeated. The final candidate spatiotemporal window set will serve as the core input for S4 temporal modeling and optimal window selection, forming a closed loop with the preceding feature processing and subsequent pain level quantification steps. This embodiment achieves accurate weighted fusion of multi-scale spatiotemporal features, quantifies muscle synergy and propagation paths, and generates candidate windows that provide scientific analysis samples for subsequent detection.
[0050] In one embodiment, the steps of constructing a vector autoregression model based on muscle state vectors, analyzing causal relationships between muscles and constructing a muscle synergy graph, dividing facial muscle function groups to calculate intra-group synergy, constructing a directed muscle propagation graph and screening key propagation paths include: S320: Construct a regularized vector autoregression model for the muscle state vector, fit the model and extract the corresponding muscle pairs, analyze the causal relationship between muscles and calculate the relevant statistics, and construct a sparse muscle synergy graph. S321: Construct a nonlinear fractional vector autoregressive model, process the time series of muscle states and analyze the relevant causal relationships, construct the corresponding synergistic graph, and weight and fuse various synergistic graphs to obtain a fused muscle synergistic graph. S322: Based on facial anatomy, facial muscle functional groups are divided, hierarchical clustering is performed on each functional group to divide it into sub-functional groups, multi-scale cross-correlation functions of muscles within the group are calculated, strong synergistic functional groups are labeled, and functional group synergy is obtained. S323: Based on the fusion of muscle synergy graph and functional group synergy, a hierarchical fractional-order directed muscle propagation graph is constructed to extract the main propagation paths and screen key propagation paths.
[0051] In this embodiment, a regularized vector autoregressive model is first constructed for the muscle state vectors. After fitting the model, corresponding muscle pairs are extracted, the causal relationship between muscles is analyzed, and relevant statistics are calculated to construct a sparse muscle collaboration graph. Based on the muscle state vectors of 12 facial muscle regions, such as the frontalis and corrugator supercilii, output by S2, a regularized vector autoregressive model is constructed. The optimal model order is determined to be second order using the Akaike information content criterion. Lasso regression is used to sparsify the model coefficients. The regularization parameter in the range of 0.01-0.1 is determined through five-fold cross-validation. The model is fitted to the state vectors of the 12 muscle regions, and the muscle pairs corresponding to non-zero coefficients in the model are extracted as objects with potential associations. Then, the Granger causality statistics of the muscle pairs are calculated. The significance level is set to 0.05. If the F-test statistic is ≥3.89, it is determined that there is a Granger causal relationship between the two. Only these muscle pairs with actual causal relationships are used as edges, and the 12 muscle regions are used as nodes to construct a sparse muscle collaboration graph. Redundant edges without association are removed, making the quantification of muscle collaboration relationships more accurate.
[0052] Then, a nonlinear fractional vector autoregressive model is constructed to process the muscle state time series and analyze the relevant causal relationships. Corresponding synergy graphs are constructed, and various synergy graphs are weighted and fused to obtain a fused muscle synergy graph. To address the nonlinear characteristics of muscle movement, a nonlinear fractional vector autoregressive (VAR) model is constructed. A neural network with two hidden layers and 128 neurons per layer approximates the nonlinear function, using the hyperbolic tangent function as the activation function. A fractional difference operator is used to process the muscle state time series, and a grid search within the range of 0.2-1.8 is conducted with a step size of 0.2 to determine the optimal fractional parameters. Based on this model, nonlinear Granger causality and fractional Granger causality between muscles are calculated. Using a significance level of 0.05 as the criterion, nonlinear Granger causality graphs and fractional muscle synergy graphs are constructed respectively. Then, a weighted fusion is performed according to the weights of the sparse muscle synergy graph (0.4), the nonlinear Granger causality graph (0.3), and the fractional muscle synergy graph (0.3) to obtain a fused muscle synergy graph that takes into account linear, nonlinear, and fractional characteristics. Specifically, the fused muscle synergy graph is obtained by weighting the sparse muscle synergy graph (0.4) corresponding to the regularized VAR model and the synergy graph (0.6) corresponding to the nonlinear fractional VAR model, thus fully reconstructing the complex synergistic relationships between muscles.
[0053] Secondly, facial muscle function groups are divided based on facial anatomy. Each function group is then hierarchically clustered into sub-function groups. The multi-scale cross-correlation function of muscles within each group is calculated, and strongly synergistic function groups are labeled to obtain the synergistic degree of the function groups. Referring to facial anatomical atlases, 12 muscle regions were divided into three functional groups based on their motor function: upper (frontalis, corrugator supercilii, etc.), middle (zygomaticus major, buccinator, etc.), and lower (depressor anguli oris, orbicularis oris, etc.). Hierarchical clustering was performed on each functional group: first, fuzzy C-means clustering (2-4 clusters, fuzzy index 1.5) was used, combined with density peak clustering to determine cluster centers, and then the average connectivity method was used to complete 2-4 layers of hierarchical clustering, further dividing each functional group into sub-functional groups; then, the multi-scale cross-correlation function of muscle pairs within the group was calculated using Morlet wavelet transform, with 8 wavelet scales set, ranging from 2 to 16 frames. If the synchronicity of a sub-functional group is ≥0.8 or the synchronicity of two or more scales is ≥0.7, it is marked as a strongly cohesive sub-group. If the proportion of strongly cohesive sub-groups in a functional group exceeds 60%, it is marked as a strongly cohesive functional group. At the same time, the average synchronicity value of each group was calculated as the functional group synergy degree.
[0054] Finally, relying on the fused muscle co-operation graph and the functional group synergy, a hierarchical fractional-order directed muscle propagation graph is constructed to extract the main propagation paths and screen key propagation paths. Based on the synergy strength of the fused muscle co-operation graph and the synergy of each functional group, a hierarchical fractional-order directed muscle propagation graph is constructed. The absolute value of the Pearson correlation coefficient and the fractional-order cross-correlation are weighted at 0.5:0.5 as the edge weights. The minimum spanning tree algorithm is used to extract the main propagation paths from complex muscle associations. Then, the breadth-first search algorithm is used to calculate the shortest path between nodes in the graph, and muscle pairs with path length <3 are screened as key propagation paths. The key propagation paths are then layered into upper, middle, and lower layers according to the upper, middle, and lower muscle functional groups. After the path screening is completed, the sparsity, nonlinearity, and hierarchical functional combination rationality of the fused co-operation graph are verified. Once the criteria are met, the final key propagation path is determined. This path, together with the fused muscle co-operation graph and the functional group synergy, serves as the core basis for S3 to generate adaptive candidate spatiotemporal windows, forming a closed loop with the spatiotemporal feature fusion and subsequent window generation steps. This embodiment realizes multi-dimensional modeling of muscle synergy, accurately quantifies intra-group synergy and muscle propagation path, and provides a scientific basis for muscle association for candidate spatiotemporal window generation.
[0055] In one embodiment, the steps of dividing facial muscle functional groups based on facial anatomy, performing hierarchical clustering to divide each functional group into sub-functional groups, calculating the multi-scale cross-correlation function of muscles within the group, labeling strongly synergistic functional groups, and obtaining the functional group synergy include: S3220: Based on facial anatomy, the muscle region is divided into facial muscle functional groups. The cross-correlation function of muscle pairs within each functional group is calculated and a cross-correlation matrix is constructed. The cross-correlation matrix is weighted by combining anatomical weights to obtain a weighted cross-correlation matrix. S3221: Construct a similarity map for each functional group based on the weighted cross-correlation matrix, calculate the Laplacian matrix of the map and extract feature vectors to construct spectral embedding features, and simultaneously calculate the local density and relative distance of muscle pairs to construct a decision map, thus obtaining the spectral embedding features and the decision map; S3222: Based on spectral embedding features and decision graphs, hierarchical clustering is performed on each functional group. The cluster centers are determined by combining fuzzy C-means clustering and density peak clustering. Sub-functional groups are divided and the fuzzy compactness and density compactness of the sub-functional groups are calculated. Strongly co-located sub-groups are labeled to obtain the sub-functional group division results. S3223: Based on the sub-functional group division results, calculate the multi-scale cross-correlation function of muscle pairs within the group, extract the cross-correlation coefficients at different time scales through wavelet transform, verify the multi-scale synchronicity, and label the strongly cooperating functional groups to obtain the functional group cooperation degree.
[0056] In this embodiment, the muscle region is first divided into facial muscle functional groups based on facial anatomy. The cross-correlation function of muscle pairs within each functional group is calculated and a cross-correlation matrix is constructed. The cross-correlation matrix is then weighted by combining anatomical weights to obtain a weighted cross-correlation matrix. Referring to facial anatomical atlases, the 12 facial muscle regions mentioned above are divided into three functional groups—upper, middle, and lower—based on their motor function and anatomical location. The upper group includes four muscle regions, such as the frontalis and corrugator supercilii; the middle group includes four muscle regions, such as the zygomaticus major and buccinator; and the lower group includes four muscle regions, such as the depressor anguli oris and orbicularis oris. For each muscle pair within a functional group, a cross-correlation function with a maximum time delay of 10 frames is calculated. Based on the function values, a muscle cross-correlation matrix for each functional group is constructed. Then, an anatomical weight matrix is introduced to classify muscle pairs according to their degree of association: strong connection (weight 1.2, such as corrugator supercilii and depressor supercilii), moderate connection (weight 1.0, such as the zygomaticus major and buccinator), and weak connection (weight 0.8, such as the orbicularis oris and mentalis). The anatomical weights and cross-correlation matrices are fused using Hadamard multiplication to obtain a weighted cross-correlation matrix for each functional group, laying a quantitative foundation for subsequent similarity analysis.
[0057] Then, a similarity map is constructed for each functional group based on the weighted cross-correlation matrix. The Laplacian matrix of the map is calculated, and feature vectors are extracted to construct spectral embedding features. Simultaneously, the local density and relative distance of muscle pairs are calculated to construct a decision map, resulting in spectral embedding features and a decision map. A similarity map is constructed based on the weighted cross-correlation matrix of each functional group, where nodes represent muscle regions, and edge weights are 1 minus the weighted cross-correlation function value of the corresponding muscle pair. The Laplacian matrix of the similarity map is calculated, and feature vectors corresponding to the first three feature values are extracted to construct global spectral embedding features. Simultaneously, the local density and relative distance of muscle pairs are calculated, and the standard deviation of the cross-correlation function is set as the neighborhood radius to determine the density distribution of each muscle pair. A decision map is constructed based on the density distribution and relative distance, visually presenting the degree of association between muscle regions and achieving dual extraction of muscle association features.
[0058] Secondly, based on spectral embedding features and decision graphs, hierarchical clustering is performed on each functional group. Fuzzy C-means clustering and density peak clustering are combined to determine cluster centers, sub-functional groups are divided, and the fuzzy compactness and density compactness of each sub-functional group are calculated. Strongly cooperating sub-groups are marked, resulting in the sub-functional group division results. By integrating the global association information of spectral embedding features and the local density information of the decision graph, hierarchical clustering is performed on the three functional groups. First, fuzzy C-means clustering (with 2-4 clusters and a fuzzy index of 1.5) is used to initially divide the clusters. Then, density peak clustering is combined to accurately determine the cluster centers. The average connectivity method is used to complete 2-4 levels of hierarchical clustering, further dividing each functional group into several sub-functional groups. Fuzzy compactness and density compactness are calculated for each sub-functional group. If both indices exceed 0.6, it is marked as a strongly cooperating sub-group. For example, the lower functional group can be divided into the orbicularis oris muscle sub-group. The orbicularis oris muscle sub-group meets the double compactness standard and is marked as a strongly cooperating sub-group. Finally, the sub-functional group division results for each functional group are obtained.
[0059] Finally, based on the sub-functional group division results, the multi-scale cross-correlation function of muscle pairs within each group is calculated. Wavelet transform is used to extract cross-correlation coefficients at different time scales to verify multi-scale synchronicity and label strongly cooperating functional groups, thus obtaining the functional group cooperability. For each functional group's sub-functional group division results, the multi-scale cross-correlation function is calculated for all muscle pairs within the group. Mollett wavelet transform is used to extract cross-correlation coefficients at 8 scales (2 to 16 frames) to verify the multi-scale synchronicity of the sub-functional groups. If the sub-functional group synchronicity is ≥0.8 or the synchronicity at two or more scales is ≥0.7, it is classified as a highly synchronized sub-functional group. If the proportion of strongly cooperating sub-groups within a functional group is ≥60%, the functional group is labeled as a strongly cooperating functional group. Simultaneously, the average multi-scale cross-correlation coefficient of all muscle pairs within the functional group is calculated as the cooperability value for that functional group. The final functional group division results, strong cooperability labels, and cooperability values will serve as the core basis for constructing the hierarchical fractional-order directed muscle propagation graph, forming a closed-loop process with the muscle cooperation graph construction and key propagation path selection steps. This embodiment combines anatomical and quantitative analysis to achieve a refined division of muscle functional groups and accurately label strong synergistic groups, providing a reliable group-level synergistic basis for muscle propagation path analysis.
[0060] In one embodiment, the steps of constructing a similarity map for each functional group based on a weighted cross-correlation matrix, calculating the Laplacian matrix of the map and extracting feature vectors to construct spectral embedding features, and simultaneously calculating the local density and relative distance of muscle pairs to construct a decision map, thereby obtaining the spectral embedding features and the decision map, include: S32210: Calculate the weighted cross-correlation function for muscle pairs within each functional group, introduce the anatomical weight matrix and fuse multimodal features, and use specified operations to fuse the anatomical weights with each modal cross-correlation matrix to construct a multimodal weighted cross-correlation matrix; S32211: Tensor decomposition is used to identify higher-order associations between muscles, construct a hypergraph association matrix, calculate the local density and relative distance of muscle pairs and determine the adaptive density threshold, and construct a hypergraph similarity structure and a multimodal local density decision graph. S32212: Based on the adaptive density threshold, the muscle edges are enhanced and suppressed, the Laplacian matrix of the multimodal hypergraph is constructed, the hypermodality is calculated and verified, and if it does not meet the standard, the decomposition parameters and density threshold are adjusted to obtain the enhanced multimodal hypergraph. S32213: Calculate the eigenvalues and eigenvectors of the Laplacian matrix of the multimodal hypergraph, construct and fuse higher-order and local spectral embedding features, verify the feature reconstruction error, and output the multimodal hypergraph and fused spectral embedding features if the target is met.
[0061] In this embodiment, firstly, a weighted cross-correlation function is calculated for muscle pairs within each functional group. An anatomical weight matrix is introduced and multimodal features are fused. A specified operation is used to fuse the anatomical weights with the cross-correlation matrices of each modality to construct a multimodal weighted cross-correlation matrix. For the three facial muscle functional groups (upper, middle, and lower) as described above, the weighted cross-correlation function for muscle pairs within each group with a maximum latency of 10 frames is calculated. An anatomical weight matrix is introduced, and muscle pairs are divided into strongly connected (weight 1.2), moderately connected (weight 0.9), and weakly connected (weight 0.7) according to their anatomical association. Simultaneously, three modal features—RGB texture, infrared temperature, and optical flow—are fused, with fusion weights set to 0.5, 0.3, and 0.2, respectively. First, a muscle pair cross-correlation matrix is constructed separately for each modality. Then, the Hadamard multiplication method is used to fuse the anatomical weight matrix with each modal cross-correlation matrix one by one. Finally, the fused matrices of each modality are weighted and superimposed according to their modal fusion weights to obtain the multimodal weighted cross-correlation matrix for each functional group, laying a multidimensional quantitative foundation for high-order association recognition. Then, tensor decomposition was used to identify higher-order associations between muscles, constructing a hypergraph association matrix. Local density and relative distance of muscle pairs were calculated, and an adaptive density threshold was determined. A hypergraph similarity structure and a multimodal local density decision graph were then constructed. For the multimodal weighted cross-correlation matrix of each functional group, CP tensor decomposition was used to identify higher-order associations between muscles. The decomposition rank was set to 5, and the reconstruction error threshold was set to 0.1. The number of identified higher-order associations was used as the number of hyperedges to construct the hypergraph association matrix, overcoming the limitations of traditional pairwise association analysis. Simultaneously, the local density and relative distance of each muscle pair were calculated, and the neighborhood radius was determined using the five nearest neighbors method. The median of the local density was set as the adaptive density threshold. Based on the hypergraph association matrix, a hypergraph similarity structure was constructed. Combining the distribution characteristics of local density and relative distance, a multimodal local density decision graph was constructed, intuitively presenting the higher-order associations and density distribution characteristics between muscles.
[0062] Secondly, muscle edges are enhanced and suppressed based on an adaptive density threshold. A multimodal hypergraph Laplacian matrix is constructed, and the hypermodule degree is calculated and validated. If the threshold is not met, the decomposition parameters and density threshold are adjusted to obtain the enhanced multimodal hypergraph. Muscle edges in the hypergraph are differentiated according to the adaptive density threshold. Edges with local densities higher than the threshold are enhanced with an enhancement coefficient of 1.5, while edges with densities lower than the threshold are suppressed with a suppression coefficient of 0.5 to reduce noise interference. A multimodal hypergraph Laplacian matrix is constructed based on the processed hypergraph. The hypermodule degree of the graph is calculated and validated with a threshold of 0.3. If the hypermodule degree is lower than the threshold, the CP tensor decomposition rank and adaptive density threshold are adjusted, and the higher-order association identification and hypergraph construction are repeated until the target is met, resulting in an enhanced multimodal hypergraph that makes the higher-order association features of muscles more prominent.
[0063] Finally, the eigenvalues and eigenvectors of the Laplacian matrix of the multimodal hypergraph are calculated, and higher-order and local spectral embedding features are constructed and fused. The feature reconstruction error is verified, and if the threshold is met, the multimodal hypergraph and fused spectral embedding features are output. Eigenvalue decomposition is performed on the enhanced multimodal hypergraph Laplacian matrix, and the eigenvectors corresponding to the three smallest non-zero eigenvalues are extracted to construct higher-order spectral embedding features. Simultaneously, local spectral embedding features are calculated based on the five nearest neighbor subgraphs of muscles. The two types of features are weighted and fused with a weight of 0.7 for higher-order features and 0.3 for local features to obtain fused spectral embedding features. A feature reconstruction error threshold of 0.15 is set, and the fused spectral embedding features are reconstructed and verified. If the error exceeds the threshold, the feature fusion weights are adjusted. If the threshold is met, the final multimodal hypergraph and fused spectral embedding features are output. This result will serve as the core feature basis for the hierarchical clustering described above, providing support for accurately dividing sub-functional groups and determining muscle synergy, forming a closed-loop process with the muscle functional group clustering analysis steps. This embodiment achieves accurate identification of high-order associations between muscles, optimizes the hypergraph structure and spectral embedding features, and provides high-dimensional feature basis for hierarchical clustering of muscle functional groups.
[0064] In one embodiment, the steps of obtaining a pain level probability distribution by dynamic fusion of weighted spatiotemporal features and time-series modeling, obtaining multi-scale micro-expression fusion features by fusing enhanced micro-expression fragment features, and selecting the optimal spatiotemporal window include: S40: Calculate the pain contribution weight and motion energy curve based on the muscle state vector, perform weight adaptive weighting on the weighted spatiotemporal features and generate a candidate window set to obtain weight adaptive spatiotemporal features; S41: Input the weighted adaptive spatiotemporal features into the temporal modeling network, perform local temporal modeling based on the candidate window set, and output the pain level probability distribution sequence; S42: Temporal transformation is performed on the enhanced micro-expression fragment features to extract the intensity spectrum. Multi-scale fusion is then performed based on the pain level probability distribution sequence and the intensity spectrum to obtain multi-scale micro-expression fusion features. S43: Based on the confidence ranking of the pain level probability distribution sequence and the energy distribution of the intensity spectrum, the optimal spatiotemporal window is selected from the multi-scale micro-expression fusion features.
[0065] In this embodiment, firstly, pain contribution weights and motion energy curves are calculated based on muscle state vectors. The weighted spatiotemporal features are then subjected to adaptive weighting, and a candidate window set is generated to obtain weighted adaptive spatiotemporal features. Based on the muscle state vectors of 12 facial muscle regions, such as the frontalis and corrugator supercilii, output by S2, core parameters such as motion amplitude, contraction rate, and activation frequency of each muscle are extracted. Weight coefficients are assigned according to pain correlation, with the corrugator supercilii and depressor anguli oris muscles having a weight coefficient of 1.2, and non-core regions having a weight coefficient of 0.8. The pain contribution weights of each muscle region are then calculated. Simultaneously, the overall motion energy curve of the facial muscles is plotted according to the video stream frame sequence to capture energy peaks and trends. Based on the gradient changes of the pain contribution weights and motion energy curves, the weighted spatiotemporal features output by S3 are subjected to adaptive weighting. Features corresponding to energy peak frames and key pain regions are weighted and enhanced. Then, five adaptive candidate spatiotemporal window sets with durations of 10-60 frames generated by S3 are retrieved, and the weighted spatiotemporal features are matched with the windows in terms of dimension, ultimately obtaining weighted adaptive spatiotemporal features, providing accurate feature input for temporal modeling.
[0066] Then, the weighted adaptive spatiotemporal features are input into the temporal modeling network. Local temporal modeling is performed based on the candidate window set, and the pain level probability distribution sequence is output. The weighted adaptive spatiotemporal features are input into the temporal modeling network, and residual supplementation, spatiotemporal decoupling, and pain semantic alignment are performed on the features. Muscle activation intensity within each candidate window is calculated based on the muscle state vector. Differentiated modeling depth is determined for different windows based on the motion energy gradient, and the five candidate windows are divided into feature intervals. Hierarchical local temporal modeling is then performed on each feature interval, extracting the depth temporal features of each window. After integrating all features, residual correction is performed based on the decoupled features. Multi-scale attention weighted aggregation is performed by combining muscle activation intensity and motion energy gradient. Finally, spatiotemporal dimension aggregation and global temporal regularization optimization are performed on the aggregated features, outputting a pain level probability distribution sequence containing four levels: no pain, mild, moderate, and severe, achieving a preliminary quantitative determination of pain level.
[0067] Secondly, the intensity spectrum of the enhanced micro-expression fragment features is extracted by temporal transformation. Based on the pain level probability distribution sequence and the intensity spectrum, multi-scale fusion is performed to obtain multi-scale micro-expression fusion features. The multi-scale enhanced micro-expression fragment features output by S2 are retrieved and subjected to temporal Fourier transform, decomposing them into temporal signals of three scales: 2×2, 4×4, and 8×8. The micro-expression intensity spectrum at each scale is extracted to capture the energy distribution and change characteristics of micro-expressions at different scales. Then, the pain level probability distribution sequence and the intensity spectrum at each scale are deeply fused according to multi-scale weights of 0.5 for the bottom layer, 0.3 for the middle layer, and 0.2 for the top layer. The regions with high confidence in the probability distribution and large proportion of energy in the intensity spectrum are enhanced, while the noise in low-confidence and low-energy regions is weakened. Finally, multi-scale micro-expression fusion features that fuse pain probability and micro-expression features are obtained, making the features more in line with the core requirements of pain detection.
[0068] Finally, based on the confidence ranking of the pain level probability distribution sequence and the energy distribution of the intensity spectrum, the optimal spatiotemporal window is selected from the multi-scale micro-expression fusion features. A quantitative judgment standard is set: regions with a confidence level ≥0.8 in the pain level probability distribution sequence are marked as high-confidence regions, and regions with a core energy proportion ≥60% in the micro-expression intensity spectrum are marked as high-energy regions. Each of the five candidate spatiotemporal windows is verified, and the window that simultaneously contains both high-confidence and high-energy regions and has the best feature continuity is selected as the optimal spatiotemporal window. If multiple windows meet the criteria, the window with the highest score is selected based on the weighted score of the comprehensive confidence level and energy proportion. The finally selected optimal spatiotemporal window will serve as the core verification object for S5 muscle dynamics consistency and micro-expression temporal consistency verification, forming a complete closed-loop process with the preceding feature processing and subsequent consistency verification steps. This embodiment achieves quantitative modeling of pain levels and deep fusion of micro-expression features, accurately selecting the optimal spatiotemporal window and providing a high-quality analysis object for subsequent verification.
[0069] In one embodiment, the step of inputting weighted adaptive spatiotemporal features into a temporal modeling network, performing local temporal modeling based on a candidate window set, and outputting a pain level probability distribution sequence includes: S410: Perform residual supplementation, spatiotemporal decoupling and pain semantic alignment on weighted adaptive spatiotemporal features, calculate the muscle activation intensity of candidate windows by combining muscle state vectors, determine the modeling depth according to motion energy gradient, and divide feature intervals by windowing to obtain window decoupling semantic residual spatiotemporal feature set and window modeling configuration information. S411: Input the feature set into the temporal modeling network according to the window modeling configuration information, perform hierarchical local temporal modeling on each feature interval, extract the deep temporal features of each candidate window, and obtain the semantic deep temporal feature set; S412: Integrate semantic deep temporal feature sets, perform residual correction based on decoupled features after temporal fusion, and combine muscle activation intensity and motion energy gradient for multi-scale attention weighted aggregation to obtain fused aggregated probability feature sequence; S413: Perform spatiotemporal dimension aggregation on the fused aggregated probability feature sequence, perform global temporal regularization optimization, and output the pain level probability distribution sequence.
[0070] In this embodiment, the weighted adaptive spatiotemporal features are first supplemented with residuals, decoupled spatiotemporally, and aligned with the semantics of pain. The muscle activation intensity of the candidate window is calculated by combining the muscle state vector. The modeling depth is determined according to the motion energy gradient. The feature interval is divided into windows to obtain the spatiotemporal feature set of window decoupling semantic residuals and window modeling configuration information. Based on the weighted adaptive spatiotemporal features obtained above, a residual network is first used to complete the missing subtle textures and temporal features in feature acquisition and fusion. Then, spatiotemporal decoupling is performed to separate and extract the spatial muscle region features and the temporal frame sequence features. At the same time, semantic alignment of pain-related features is completed based on pain semantic labels to enhance the feature expression of key pain regions such as the corrugator supercilii and depressor anguli oris muscles. Combining the 12 muscle region state vectors output by S2, the muscle activation intensity within the 5 candidate spatiotemporal windows (durations of 15, 20, 30, 40, and 50 frames) generated by S3 is calculated. Differentiated modeling depths are determined for each window according to the motion energy gradient. Windows with an energy gradient ≥ 0.7 are set to 3 layers of modeling depth, and those with a gradient < 0.7 are set to 2 layers. Then, the decoupled features are divided into window intervals according to the window duration and modeling depth to obtain the window decoupled semantic residual spatiotemporal feature set. At the same time, the configuration information such as the modeling depth, feature dimension, and scale weight of each window is recorded to provide standardized input for subsequent temporal modeling.
[0071] Then, the feature set is input into the temporal modeling network according to the window modeling configuration information. Hierarchical local temporal modeling is performed on each feature interval to extract the deep temporal features of each candidate window, resulting in a semantic deep temporal feature set. The spatiotemporal feature set of window decoupling semantic residuals is input into the temporal modeling network according to the modeling configuration information. For the feature interval of each candidate window, hierarchical local temporal modeling is performed at multiple scales of 2×2, 4×4, and 8×8. The bottom layer performs fine-grained temporal feature extraction on high-energy frames, and the top layer performs overall temporal trend capture on the frame sequence. The causal convolution module is combined to extract local temporal dependencies, and the multi-head self-attention module captures long-distance temporal associations. Deep temporal features are extracted for each of the five candidate windows. The extracted features are fused and enhanced with pain semantic features to eliminate the interference of non-pain features. Finally, a semantic deep temporal feature set containing the spatiotemporal semantic features of each window is obtained, realizing the hierarchical and semantic expression of features. Secondly, the semantic deep temporal feature set is integrated, and residual correction is performed based on the decoupled features after temporal fusion. Multi-scale attention-weighted aggregation is then performed by combining muscle activation intensity and motion energy gradient to obtain a fused aggregated probability feature sequence. The semantic deep temporal feature sets of the five candidate windows are fused in an ordered manner according to the frame temporal sequence of the video stream. Residual correction is performed based on the previously decoupled spatial and temporal features, and the deviation between the fused features and the original features is calculated and corrected to reduce feature errors in the modeling process. Then, multi-scale attention weights are set by combining the muscle activation intensity and motion energy gradient of each window. The feature regions with activation intensity ≥0.8 and energy gradient ≥0.7 are weighted at 1.3, the medium intensity and gradient regions are weighted at 1.0, and the low-level regions are weighted at 0.7. The fused features are then weighted and aggregated to generate a fused aggregated probability feature sequence containing pain level probability information, highlighting the core temporal features related to pain.
[0072] Finally, the fused and aggregated probability feature sequence is subjected to spatiotemporal aggregation processing, and global temporal regularization optimization is performed to output a pain level probability distribution sequence. The fused and aggregated probability feature sequence is aggregated with spatial muscle region features and temporal frame sequence features, integrating multi-window and multi-scale probability features. Then, through a global temporal regularization algorithm, temporal deviations between frames are aligned, local fluctuations in probability features are smoothed, and probability jumps caused by changes in the frame rate of micro-expression segments are eliminated. The final output is a pain level probability distribution sequence containing four levels: no pain, mild, moderate, and severe. This sequence will serve as the core basis for subsequent fusion with the micro-expression intensity spectrum, forming a closed loop with the multi-scale micro-expression fusion feature extraction and optimal spatiotemporal window selection steps. This embodiment achieves refined processing of spatiotemporal features and hierarchical temporal modeling, improving the accuracy of pain level probability features and providing a reliable temporal basis for subsequent feature fusion.
[0073] In one embodiment, the step of inputting the feature set into a temporal modeling network according to window modeling configuration information, performing hierarchical local temporal modeling on each feature interval, extracting deep temporal features of each candidate window, and obtaining a semantic deep temporal feature set includes: S4110: Perform frame-level temporal alignment on the feature set, simultaneously complete semantic feature enhancement, residual feature completion and multi-scale feature mapping, calculate the muscle activation intensity of the candidate window by combining the muscle state vector and perform spatiotemporal decoupling, and obtain the hierarchical window configuration and decoupled aligned semantic residual multi-scale feature interval set; S4111: The feature interval set is configured into the time series modeling network according to the hierarchical window, and the spatiotemporal decoupling local time series modeling combining hierarchical and scale-based methods is performed to extract the initial decoupling depth time series features and obtain the multi-scale hierarchical decoupling time series feature set. S4112: Perform semantic association extraction and muscle semantic enhancement on the feature set, and sequentially complete feature hierarchical aggregation, multi-scale fusion, residual supplementation and temporal feature correction to obtain a semantically enhanced and corrected deep temporal feature set; S4113: Integrate the feature set, perform global regularization according to the window temporal relationship, complete semantic scale adaptation and inter-level feature fusion, and obtain a semantic deep temporal feature set.
[0074] In this embodiment, the feature set is first aligned at the frame level, and semantic feature enhancement, residual feature completion and multi-scale feature mapping are completed simultaneously. The muscle activation intensity of the candidate window is calculated by combining the muscle state vector and spatiotemporal decoupling is performed to obtain the hierarchical window configuration and decoupled aligned semantic residual multi-scale feature interval set. For the spatiotemporal feature set of decoupled semantic residuals obtained above, frame-level temporal alignment is first performed according to the video stream time step to eliminate the inter-frame offset deviation caused by rapid micro-expression movements. Simultaneously, semantic feature enhancement is performed on key pain areas such as the corrugator supercilii and depressor anguli oris muscles based on pain semantic labels. The missing subtle muscle movement features during feature acquisition are supplemented by residual networks. Then, the features are uniformly mapped to three scales: 2×2, 4×4, and 8×8 to form multi-scale feature mapping results. Combined with the 12 facial muscle region state vectors output by S2, the muscle activation intensity in 5 candidate spatiotemporal windows (15, 20, 30, 40, and 50 frames) is calculated. Spatiotemporal decoupling processing is performed on the features in the spatial (muscle region) and temporal (frame sequence) dimensions. At the same time, the modeling depth, scale weight, and other parameters are updated according to the motion energy gradient of each window to generate a hierarchical window configuration. Finally, the decoupled and aligned semantic residual multi-scale feature interval set is obtained, which prepares standardized features for temporal modeling.
[0075] Then, the feature interval set is configured into a hierarchical window and input into the temporal modeling network. The spatiotemporal decoupling local temporal modeling is performed by combining hierarchical and scale-based methods. The initial decoupling depth temporal features are extracted to obtain a multi-scale hierarchical decoupling temporal feature set. The decoupled and aligned semantic residual multi-scale feature interval set is input into the temporal modeling network in a hierarchical window configuration. For high-energy windows with motion energy gradient ≥ 0.7 and a modeling depth of 3 layers, more granular hierarchical processing is performed, while the hierarchical extraction is simplified for windows with gradient < 0.7 and a modeling depth of 2 layers. At the same time, combined with the scale-based processing logic, the bottom layer extracts subtle temporal features of high-energy frames at a 2×2 scale, and the top layer captures the overall temporal trend of the frame sequence at an 8×8 scale. Local temporal dependencies are extracted through causal convolution modules (3, 5, and 7 convolutional kernels), and long-distance temporal correlations are captured through multi-head self-attention modules (8 heads, 64 dimensions). The modeling logic of spatiotemporal decoupling is maintained throughout the process. Initial decoupling depth temporal features are extracted for each scale and each level. After integrating all features, a multi-scale hierarchical decoupled temporal feature set is obtained, realizing hierarchical and scale-based accurate extraction of temporal features.
[0076] Secondly, semantic association extraction and muscle semantic enhancement are performed on the feature set. This involves sequentially completing feature hierarchical aggregation, multi-scale fusion, residual supplementation, and temporal feature correction to obtain a semantically enhanced and corrected deep temporal feature set. From the multi-scale hierarchical decoupled temporal feature set, semantic association features of pain-related muscle regions are extracted, such as the synergistic movement association between the corrugator supercilii and depressor supercilii, and between the depressor anguli oris and mentalis muscles. Based on muscle semantic labels, these pain-coordinating regions are feature-enhanced. First, the temporal features at each level are weighted and aggregated. Then, multi-scale feature fusion is performed with weights of 0.5 for the bottom layer, 0.3 for the middle layer, and 0.2 for the top layer. A residual network is used to supplement the feature bias generated during the fusion process. Finally, a temporal correction algorithm is used to eliminate inter-frame temporal jitter errors and correct feature temporal offset problems, ultimately obtaining a semantically enhanced and corrected deep temporal feature set, making the features more aligned with the semantic and temporal requirements of pain detection.
[0077] Finally, the feature set is integrated, and global normalization is performed according to the temporal relationship of the windows to complete semantic scale adaptation and inter-level feature fusion, resulting in a semantic deep temporal feature set. The semantic enhancement and correction deep temporal feature sets corresponding to the five candidate windows are integrated, and global temporal normalization is performed strictly according to the frame temporal relationship of the video stream to align temporal deviations between windows and eliminate feature fragmentation caused by window division. Simultaneously, based on the expression requirements of pain semantic features, semantic scale adaptation at different scales is completed to ensure consistent semantic expression of features at each scale. Then, deep fusion is performed on features at different modeling levels to eliminate feature conflicts between levels and strengthen the expression of key pain temporal features, ultimately obtaining a semantic deep temporal feature set. This feature set will serve as the core input for generating the previously described fusion and aggregation probability feature sequence, forming a closed loop with the residual correction and multi-scale attention weighted aggregation steps. This embodiment achieves multi-dimensional refined processing and semantic enhancement of temporal features, improving the expressive power of deep temporal features and providing high-quality feature basis for pain probability modeling.
[0078] In one embodiment, the step of configuring the feature interval set into a hierarchical window and inputting it into a temporal modeling network, performing hierarchical and scale-based spatiotemporal decoupling local temporal modeling, extracting initial decoupling depth temporal features, and obtaining a multi-scale hierarchical decoupling temporal feature set includes: S41110: Perform scale stratification and hierarchical division on the feature interval set according to the hierarchical window configuration, and simultaneously complete the window hierarchy mapping and spatiotemporal decoupling calibration. Combine the muscle state vector to calculate the muscle activation intensity, extract the micro-expression intensity spectrum through temporal Fourier transform, and obtain the decoupling spatial feature set and decoupling temporal feature set through spatiotemporal decoupling. S41112: Input the decoupled spatial feature set into the frequency domain branch of the time-series modeling network according to the configuration, perform frequency domain subscale modeling, and extract the frequency domain spatial feature set; combine the decoupled temporal feature set with the micro-expression intensity spectrum, perform frequency domain subscale modeling, and extract the frequency domain temporal feature set; S41113: Perform deep feature mining on two types of frequency domain feature sets, enhance the deep feature representation of candidate windows, and obtain a scale-level hierarchical subset of frequency domain features; S41114: Integrate all feature subsets, and after frequency domain time-series fusion, sort them by scale and hierarchy to obtain a multi-scale hierarchical decoupled time-series feature set.
[0079] In this embodiment, the feature interval set is first scaled and hierarchically divided according to a hierarchical window configuration. Simultaneously, window hierarchy mapping and spatiotemporal decoupling calibration are completed. Muscle activation intensity is calculated using muscle state vectors, and micro-expression intensity spectra are extracted via temporal Fourier transform. Spatiotemporal decoupling yields decoupled spatial and temporal feature sets. For the decoupled alignment semantic residual multi-scale feature interval set mentioned earlier, based on the hierarchical window configuration, scale stratification is first completed using 2×2, 4×4, and 8×8 methods. Then, hierarchical division is performed according to motion energy gradients—candidate windows with energy gradients ≥ 0.7 at frames 30, 40, and 50 are set with 3 modeling layers, while windows with gradients < 0.7 at frames 15 and 20 are set with 2 modeling layers. Simultaneously, accurate scale and hierarchy mapping for each window is completed, and spatiotemporal features are decoupled and calibrated to eliminate inter-frame offset and muscle region feature aliasing bias. Combining the 12 facial muscle region state vectors output by S2, the muscle activation intensity of regions such as the frontalis muscle and corrugator supercilii muscle within 5 candidate windows is calculated. Pain-critical regions with activation intensity ≥0.8 are marked with features. Then, a temporal Fourier transform is performed on the feature set to extract micro-expression intensity spectra at 8 scales to capture temporal energy changes. Finally, strict spatiotemporal decoupling is achieved, resulting in a decoupling spatial feature set in terms of muscle regions and a decoupling temporal feature set in terms of frame sequences, laying a feature-separated foundation for frequency domain modeling.
[0080] Then, the decoupled spatial feature set is input into the frequency domain branch of the temporal modeling network according to the configuration, and frequency domain sub-scale modeling is performed to extract the frequency domain spatial feature set. The decoupled temporal feature set is combined with the micro-expression intensity spectrum, and frequency domain sub-scale modeling is performed to extract the frequency domain temporal feature set. The decoupled spatial feature set is configured according to scale layering and hierarchy, and input into the frequency domain branch of the temporal modeling network. Fine-grained frequency domain analysis is performed on the 2×2 scale to extract the local frequency domain features of each muscle region, and global frequency domain analysis is performed on the 8×8 scale to capture the frequency domain correlation features of the overall facial muscles. Frequency domain spatial feature sets of each scale and level are extracted through convolution operation. The decoupled temporal feature set is dimension-matched with the micro-expression intensity spectrum, and input into the frequency domain branch to perform sub-scale modeling. Combined with the high-energy region marking of the intensity spectrum, the frequency domain features of the pain micro-expression temporal nodes are enhanced, and a frequency domain temporal feature set that can reflect the temporal pattern of the frame sequence is extracted, realizing the accurate extraction of the frequency domain dimension of spatiotemporal features.
[0081] Secondly, deep feature mining was performed on the two types of frequency domain feature sets to enhance the deep feature representation of candidate windows, resulting in scale-level hierarchical frequency domain feature subsets. Deep feature mining was conducted on both the frequency domain spatial and temporal feature sets. The frequency domain spatial features of key pain regions such as the corrugator supercilii and depressor anguli oris muscles were weighted and enhanced with a coefficient of 1.3, while feature smoothing was performed on non-core regions. The frequency domain temporal features corresponding to high-energy frames of the micro-expression intensity spectrum were extracted with a focus, while the noise features of low-energy frames were weakened. Simultaneously, according to the rules of scale stratification (2×2, 4×4, 8×8) and hierarchical division (2 layers, 3 layers), corresponding frequency domain feature content was extracted for each of the five candidate windows, forming independent scale-level hierarchical frequency domain feature subsets. Each subset was matched with the modeling configuration of the corresponding window to ensure accurate correspondence between features and windows.
[0082] Finally, all feature subsets are integrated, and after frequency domain temporal fusion, they are sorted by scale and hierarchy to obtain a multi-scale hierarchical decoupled temporal feature set. The scale-level hierarchical frequency domain feature subsets of the five candidate windows are globally integrated. First, frequency domain temporal fusion is performed according to the frame temporal relationship of the video stream to eliminate feature fragmentation caused by window division. Then, scale weighting is performed according to scale weights (0.5 for the bottom layer 2×2, 0.3 for the middle layer 4×4, and 0.2 for the top layer 8×8), and hierarchical regularization is performed by modeling level (2 layers, 3 layers). The fused features are smoothed to eliminate feature conflicts between scales and levels, ultimately yielding a multi-scale hierarchical decoupled temporal feature set. This feature set will serve as the core input for the preceding semantic association extraction and muscle semantic enhancement, forming a complete closed loop with subsequent feature enhancement and correction steps. This embodiment mines the frequency domain patterns of spatiotemporal features through frequency domain scale-level modeling, strengthening the expression of pain-related features and providing high-quality temporal feature basis for subsequent semantic processing.
[0083] In one embodiment, the steps of verifying muscle dynamic consistency based on muscle state vectors, verifying micro-expression temporal consistency based on multi-scale micro-expression fusion features, verifying the effectiveness of the optimal spatiotemporal window, returning to the corresponding step for reprocessing if any verification fails, and outputting the final pain level if all verifications pass, include: S50: Based on muscle state vectors, verify muscle dynamic consistency, calculate prediction error through filtering, construct muscle co-construction graph and use graph neural network to learn node embedding, calculate relevant parameters to construct joint consistency index, and output prediction error matrix, node embedding vector and joint consistency score. S51: Verify the temporal consistency of micro-expressions based on multi-scale micro-expression fusion features, construct a probabilistic model and generate predicted samples through sampling, calculate the log likelihood, use a generative adversarial network to generate samples and calculate the relevant loss, construct a temporal consistency index and output the corresponding score and sample-related features. S52: Verify the effectiveness of the optimal spatiotemporal window, calculate the relevant attention weight distribution and sample diversity within the window, integrate relevant indicators to obtain the window quality score, combine the two types of consistency scores to construct the overall verification framework, and output the window effectiveness result and the comprehensive consistency score. S53: If the overall validation fails, perform joint feedback optimization by means of gradient descent, model parameter update, etc., recalculate each index and verify the optimization effect; if all standards are met, output the final pain level and confidence index; if they are still not met, continue to return to adjust the parameters and reprocess.
[0084] In this embodiment, muscle dynamics consistency is first verified based on muscle state vectors. Prediction errors are calculated through filtering, a muscle co-location graph is constructed, and a graph neural network is used to learn node embeddings. Relevant parameters are calculated to construct a joint consistency index, and the prediction error matrix, node embedding vectors, and joint consistency score are output. Based on the muscle state vectors of 12 facial muscle regions, such as the frontalis and corrugator supercilii, output in S2, a Kalman filter algorithm is used to predict the motion state of each muscle. The deviation between the actual muscle state vector and the predicted vector is calculated to obtain the prediction error matrix. Combined with the sparse muscle co-location graph constructed in S3, a simple graph neural network is built. The muscle state vectors and co-location graph node information are input, and the node embedding vectors of each muscle region are learned. While calculating the smoothness of the embedding vectors, the mean square error of the prediction error matrix is combined to construct a joint consistency index, with a score threshold of ≥0.6. For example, the mean square error of the prediction error for key pain areas such as the corrugator supercilii and depressor anguli oris muscles is ≤0.05, and the node embedding smoothness is ≥0.8. The final calculated joint consistency score is 0.75, meeting the requirements. The prediction error matrix, node embedding vectors, and the score are output simultaneously, completing the initial verification of muscle dynamics consistency.
[0085] Then, based on the multi-scale micro-expression fusion features, the temporal consistency of micro-expressions is verified. A probabilistic model is constructed, and predicted samples are generated through sampling. The log-likelihood is calculated. A generative adversarial network is used to generate samples and calculate relevant losses. A temporal consistency index is constructed, and the corresponding score and sample-related features are output. The multi-scale micro-expression fusion features output by S4 are retrieved, and a Bayesian Hidden Markov Joint Probabilistic Model containing 5 temporal states is constructed. The posterior distribution of the model is estimated through variational inference algorithm. Gibbs sampling is used to generate 100 sets of micro-expression prediction samples, and the log-likelihood value between the samples and the original multi-scale micro-expression fusion features is calculated. At the same time, a recurrent generative adversarial network is built to generate highly realistic micro-expression samples by inputting micro-expression features. The observation likelihood, diffusion loss, and adversarial loss are calculated. The log-likelihood value and various loss values are fused to construct a temporal consistency index, with a threshold of ≥0.7. Combining the micro-expression segments within the optimal spatiotemporal window, a temporal consistency score of 0.78 is calculated, which meets the standard. The score, predicted samples, and related features of generated samples are output, completing the micro-expression temporal consistency verification, which echoes the micro-expression feature extraction steps described above.
[0086] Next, the validity of the optimal spatiotemporal window is verified. The distribution of relevant attention weights and sample diversity within the window are calculated, and relevant indicators are fused to obtain a window quality score. An overall verification framework is constructed by combining two types of consistency scores, and the window validity result and comprehensive consistency score are output. For the optimal spatiotemporal window selected by S4 (assumed to be 30 frames in length, including high-confidence pain probability regions and high-energy micro-expression regions), the attention weight distribution of the graph neural network within the window is calculated, and the attention weight ratio of key pain regions is statistically analyzed. Simultaneously, the diversity index of samples generated by the generative adversarial network is calculated. The attention entropy and sample diversity index are fused to construct a window quality score, with a threshold of ≥0.6. The calculated window quality score is 0.65, which meets the requirements. Combining the joint consistency score (0.75) and temporal consistency score (0.78) obtained previously, an overall verification framework is constructed. The comprehensive consistency score is the weighted average of the three (with weight ratios of 0.3, 0.4, and 0.3 respectively), resulting in a comprehensive score of 0.74. The optimal spatiotemporal window is deemed valid, and the validity result and comprehensive score are output.
[0087] Finally, if the overall validation fails, joint feedback optimization is performed using gradient descent and model parameter updates to recalculate all indicators and verify the optimization effect. If all indicators meet the standards, the final pain level and confidence index are output; otherwise, the process returns to adjust parameters and reprocess. If the overall consistency score does not meet the standard (e.g., below 0.6), the state transition matrix of the muscle dynamics model is updated using the gradient descent algorithm to reduce prediction error. The Bayesian Hidden Markov Model parameters are updated using the expectation-maximization algorithm to optimize sample generation. Joint consistency, temporal consistency, and window quality scores are recalculated to verify whether the optimized indicators meet the standards. In this case, all scores meet the standards. Based on the pain level probability distribution sequence output by S4, the level with the highest probability is selected as the final result, such as moderate pain. The confidence index for this level is calculated to be 0.88 and output simultaneously. This final result is the output of the entire pain expression detection process, forming a complete closed loop with all processing steps from S1 to S4, ensuring the reliability of the detection results. This embodiment uses triple consistency verification and joint feedback optimization to verify the rationality of the detection process from multiple dimensions, greatly improving the accuracy of pain detection and ensuring the accuracy and reliability of the output results.
[0088] In one embodiment, the steps of verifying the temporal consistency of micro-expressions based on multi-scale micro-expression fusion features, constructing a probabilistic model and generating predicted samples through sampling, calculating the log-likelihood, generating samples using a generative adversarial network and calculating relevant losses, constructing a temporal consistency index and outputting the corresponding score and sample-related features include: S510: Construct a Bayesian Hidden Markov Joint Probability Model based on Multi-Scale Micro-Expression Fusion Features, define the temporal state of micro-expressions, use Bayesian inference to estimate the state transition matrix, use variational inference to estimate the posterior distribution, construct the joint probability distribution, and record the state transition probability and observation likelihood. S511: Use the normalized flow model and the diffusion probability model to jointly generate prediction samples, generate intermediate samples by sampling and inverse transformation of the standard normal distribution, gradually denoise the intermediate samples, calculate the flow likelihood and diffusion loss, construct the prediction sample distribution, and record the state label and noise level of the samples. S512: Construct a recurrent generative adversarial network and a self-attention mechanism. The generator adopts an encoder and decoder structure, and the discriminator uses a convolutional network. Cyclic consistency is calculated through cyclic mapping to obtain attention-weighted temporal features. S513: Integrate observational likelihood, flow likelihood, diffusion loss, attention loss, and cyclic consistency loss to construct a comprehensive temporal consistency index. Calculate the Euclidean distance between the prediction and the true features and verify the effectiveness of the index. If the index is met, output the predicted sample distribution, attention weight, and comprehensive temporal consistency score. If the index is not met, return to adjust the model parameters and reprocess.
[0089] In this embodiment, a Bayesian Hidden Markov Joint Probability Model is first constructed based on multi-scale micro-expression fusion features. The temporal states of micro-expressions are defined, and the state transition matrix is estimated using Bayesian inference. The posterior distribution is estimated using variational inference, and a joint probability distribution is constructed. The state transition probabilities and observation likelihoods are recorded. The multi-scale micro-expression fusion features output by S4 (covering three scales: 2×2, 4×4, and 8×8, fusing pain level probability and micro-expression intensity spectrum features) are retrieved. Combining the movement patterns of 12 facial muscle regions such as the frontalis and corrugator supercilii, five micro-expression temporal states (resting state, slight contraction state, moderate contraction state, severe contraction state, and recovery state) are defined, corresponding to the temporal change in pain sensation from no pain to severe pain. Using a Bayesian inference algorithm, based on the time-series sequence of muscle state vectors output by S2, the transition matrix between five states is estimated. The transition probability from slight contraction to moderate contraction is set to 0.6, from moderate to severe contraction to 0.4, and from recovery state to resting state to 0.7. Variational inference is used to approximate the posterior distribution of the model, reducing computational complexity. A joint probability distribution of micro-expression time-series states and features is constructed, and the transition probabilities and observed likelihoods of each state are recorded simultaneously (e.g., the observed likelihood of moderate contraction of the corrugator supercilii is ≥0.8), laying a probabilistic foundation for subsequent sample generation and consistency determination.
[0090] Then, a normalized flow model and a diffusion probability model are jointly used to generate predicted samples. Intermediate samples are generated through sampling and inverse transformation of the standard normal distribution. The intermediate samples are progressively denoised, and the flow likelihood and diffusion loss are calculated to construct the predicted sample distribution. The state labels and noise levels of the samples are recorded. Based on the above joint probability distribution, a joint generation framework of the normalized flow model and the diffusion probability model is built. 100 groups of noise vectors are randomly sampled from the standard normal distribution. The noise vectors are mapped to micro-expression intermediate samples through inverse transformation. A 10-step denoising process is set. Each step adopts a linear denoising strategy to reduce sample noise and ensure that the samples fit the real micro-expression temporal features. The flow likelihood (measuring the fit between the sample distribution and the real distribution) and diffusion loss (controlling the deviation of the denoising process) of each group of samples are calculated. The flow likelihood is set to ≥0.8 and the diffusion loss to ≤0.15. A complete predicted sample distribution is constructed. Each group of samples is labeled with a corresponding temporal state label (such as moderate contraction state) and noise level (noise level ≤0.05 after denoising). The generated predicted samples will be used for consistency comparison with the original micro-expression features.
[0091] Secondly, a recurrent generative adversarial network (Cycle-GAN) and a self-attention mechanism are constructed. The generator adopts an encoder-decoder structure, while the discriminator uses a convolutional network. Cyclic consistency is calculated through cyclic mapping to obtain attention-weighted temporal features. A Cycle-GAN is built where the generator uses an encoder-decoder dual-branch structure. The encoder extracts temporal correlation information from multi-scale micro-expression fusion features (outputting a 64-dimensional feature vector), and the decoder maps the feature vector to highly realistic micro-expression samples. The discriminator uses a 3-layer convolutional network to distinguish generated samples from original micro-expression features, optimizing the realism of the generated samples. A self-attention mechanism is introduced to strengthen the feature weights of key pain areas such as the corrugator supercilii and depressor anguli oris muscles (attention weight ≥ 0.7 for key areas) and weaken background noise in non-core areas. The generated samples are reverse-mapped to feature vectors through cyclic mapping, and the cyclic consistency error of the features before and after mapping is calculated (error ≤ 0.05). Finally, attention-weighted micro-expression temporal features are obtained, highlighting the temporal correlation of pain-related micro-expressions.
[0092] Finally, a comprehensive temporal consistency index is constructed by integrating observational likelihood, flow likelihood, diffusion loss, attention loss, and cycle consistency loss. The Euclidean distance between the predicted and true features is calculated, and the index's effectiveness is verified. If the index meets the standard, the predicted sample distribution, attention weights, and comprehensive temporal consistency score are output. If the index does not meet the standard, the process returns to adjust the model parameters and reprocess. The weight allocation for the comprehensive temporal consistency index is set as follows: observational likelihood 0.3, flow likelihood 0.2, diffusion loss 0.2, attention loss 0.15, and cycle consistency loss 0.15. The index score is calculated by integrating the parameters, and the Euclidean distance between the predicted sample features and the original multi-scale micro-expression fusion features (distance ≤ 0.1) is also calculated to double verify the index's effectiveness. The threshold for the comprehensive temporal consistency score is set to ≥ 0.7. In this calculation, the score is 0.78 and the Euclidean distance is 0.08, both meeting the requirements. The predicted sample distribution, attention weight distribution, and comprehensive score are output. If the score does not meet the standard, the process returns to adjust the state transition parameters of the Bayesian Hidden Markov Model and the number of denoising steps in the diffusion model, and the sample generation and index calculation are repeated. The comprehensive temporal consistency score output in this study will be used in conjunction with the muscle dynamics consistency score and window quality score mentioned earlier to form a closed-loop process with the muscle dynamics consistency verification and optimal window validity verification steps. This embodiment accurately verifies the temporal consistency of micro-expressions by jointly generating samples using multiple models and fusing multi-dimensional losses, thereby improving the reliability of temporal features and providing strong support for the verification of overall detection results.
[0093] refer to Figure 2 A pain expression detection system, comprising: The feature acquisition module 100 is used to acquire facial video streams, perform facial key point localization and inter-frame alignment, divide facial muscle regions and extract texture features, and obtain an aligned frame sequence and a set of texture features of muscle regions. The micro-expression enhancement module 200 is used to calculate motion energy based on texture feature set, extract and enhance micro-expression temporal segments from aligned frame sequence, extract multi-scale spatiotemporal features, construct facial muscle dynamics model and complete state estimation, and obtain enhanced micro-expression segments, multi-scale spatiotemporal features and muscle state vectors. The fusion window generation module 300 is used to perform spatiotemporal attention weighted fusion based on multi-scale spatiotemporal features to obtain weighted spatiotemporal features, and combine muscle state vectors to analyze muscle synergy relationships to generate an adaptive candidate spatiotemporal window set. The temporal modeling and filtering module 400 is used to obtain the pain level probability distribution by dynamic fusion of weighted spatiotemporal features, fuse enhanced micro-expression fragment features to obtain multi-scale micro-expression fusion features, and filter out the optimal spatiotemporal window. The verification result output module 500 is used to verify the consistency of muscle dynamics based on muscle state vectors, verify the consistency of micro-expression temporal sequence based on multi-scale micro-expression fusion features, and verify the validity of the optimal spatiotemporal window. If any verification fails, the corresponding step is returned for reprocessing. If all verifications pass, the final pain level is output.
[0094] Furthermore, the aforementioned feature acquisition module 100 includes: The dual-modal acquisition and key point localization unit is used to acquire facial dual-modal video streams and resting segments, construct individual muscle feature baselines and pre-train muscle dynamics models, and perform facial key point localization on the video stream by combining the trajectory and baseline to obtain a set of dual-modal key point coordinates. The inter-frame alignment and cross-frame association unit is used to calculate the dual-modal fusion weight based on the coordinate set, perform inter-frame predictive joint alignment, and extract optical flow field features to realize cross-frame association of key points, thereby obtaining the predicted aligned frame sequence. The muscle region fine segmentation unit is used to segment the muscle region finely based on the aligned frame sequence, combined with facial anatomical atlas and muscle dynamics model state parameters, and generate muscle region segmentation results; The dual-modal texture feature extraction unit is used to extract dual-modal texture features from the segmentation results, differ with the individual muscle feature baseline, and fuse cross-frame associated texture features to obtain a set of texture features for the muscle region.
[0095] Furthermore, the aforementioned dual-modal acquisition and key point localization unit includes: The dual-modal data and resting segment acquisition unit is used to acquire facial dual-modal video streams and multi-time resting segments, extract multi-scale fusion features of resting segments, construct a fusion-type individual muscle feature baseline, pre-train a muscle dynamics model, and output muscle pre-movement trajectories and attention weights. The multi-scale cross-modal early fusion initial localization unit is used to combine baseline, trajectory and attention weights to construct a multi-scale image pyramid for the video stream and perform cross-modal early fusion. The initial key point set is obtained by the scale-weighted initial localization. The key point localization optimization unit is used to perform cross-scale fusion and cross-modal reverse mapping verification on the initial key point set, correct fusion and scale bias, and then optimize the pain area localization accuracy based on attention weights to obtain the optimized key point set. The key point localization verification and output unit is used to verify the cross-modal matching degree and multi-scale consistency of the optimized key points. If the standard is not met, the unit returns to adjust the fusion weight and sub-scale parameters. If the standard is met, the unit obtains a set of bimodal personalized key point coordinates.
[0096] Furthermore, the aforementioned micro-expression enhancement module 200 includes: A multi-scale motion energy set generation unit is used to construct a multi-scale feature pyramid based on a texture feature set, calculate the motion energy at each scale, and fuse optical flow field features to obtain an optical flow-enhanced multi-scale motion energy set. The multi-scale enhanced micro-expression fragment extraction unit is used to locate key pain areas by combining spatial attention mechanisms, extract micro-expression temporal fragments containing peak frames based on the hierarchical extraction of motion energy sets, and enhance key features through attention enhancement network to obtain multi-scale enhanced micro-expression fragments. The multi-scale spatiotemporal feature fusion unit is used to extract spatial features, temporal dependence features and optical flow temporal features at each scale from the fragment and fuse them to form multi-scale spatiotemporal features. The facial muscle dynamics model construction and optimization unit is used to combine bimodal information and attention weights to construct a facial muscle dynamics model and optimize the state transition matrix and noise covariance matrix. The muscle state vector estimation and output unit is used to input a texture feature set to complete state estimation and obtain a muscle state vector.
[0097] Furthermore, the aforementioned multi-scale enhanced micro-expression fragment extraction unit includes: Attention pyramid building unit is used to construct a multi-level attention pyramid, integrate spatial attention and temporal attention and superimpose optical flow field features to locate key pain areas and output a multi-scale spatiotemporal-optical flow fusion weight matrix. The hierarchical micro-expression extraction unit is used to extract micro-expression time segments according to the pyramid level based on the fusion weight matrix, select different energy frames in different regions, retain peak frames and core frames, and output a multi-scale hierarchical segment set. The attention feature enhancement unit is used to input the multi-scale hierarchical fragment set into the attention enhancement network, allocate network resources according to the product of multi-dimensional weights, enhance the features of key regions and suppress the features of noisy regions, and output the fused enhanced fragment. The micro-expression segment verification unit is used to verify the inter-frame coherence and multi-scale feature coverage of the fused and enhanced segment. If both criteria are met, the fused and enhanced segment is determined to be a multi-scale enhanced micro-expression segment. If the criteria are not met, the process is returned to adjust the optical flow fusion coefficient and hierarchical weight for reprocessing.
[0098] Furthermore, the aforementioned attention pyramid building blocks include: The spatial weight calculation unit is used to construct a multi-level attention pyramid, calculate the bimodal spatial attention weight, take the average weight for consistent regions and take the high-weight modal weight for different regions, and adjust the weight of the corresponding region in combination with the muscle synergy coefficient to complete the determination of the key pain region. The temporal optical flow extraction unit is used to dynamically update temporal parameters according to the temporal window, extract temporal features based on the type of pain key region, extract optical flow features in noisy scenes, and obtain a noise-resistant temporal and optical flow feature set; The fusion matrix generation unit is used to combine bimodal spatial attention weights, noise-resistant temporal and optical flow feature sets, and muscle coordination features. Consistent regions are calculated according to preset rules, while different regions are fused according to the high-weight mode, forming a multi-scale spatiotemporal-optical flow fusion matrix.
[0099] Furthermore, the aforementioned fusion window generation module 300 includes: Spatial weighted feature units are used to perform multi-scale channel decomposition on multi-scale spatiotemporal features. Each group independently calculates spatial attention weights and generates a spatial attention map. Pain-related regions are identified through dynamic threshold processing to obtain spatial weighted features. The spatiotemporal feature fusion unit is used to construct a multi-head self-attention module and a causal convolution module for temporal features. The outputs of the two types of modules are weighted and fused to obtain temporal weighted features. Then, the spatial weighted features and temporal weighted features are fused to obtain weighted spatiotemporal features. The muscle synergy analysis unit is used to construct a vector autoregression model based on muscle state vectors, analyze the causal relationship between muscles and construct a muscle synergy graph, divide facial muscle function groups to calculate the synergy degree within the group, construct a directed graph of muscle propagation and screen key propagation paths. The candidate window generation unit is used to generate candidate spatiotemporal windows based on micro-expression fragments, muscle synergy maps, functional group synergy, and key propagation paths. It calculates the window quality score and verifies feature consistency. If the score is met, it is determined as an adaptive candidate spatiotemporal window set; otherwise, it returns to adjust the dynamic coefficients and reprocess.
[0100] Furthermore, the aforementioned muscle synergy analysis unit includes: The sparse co-operation graph construction unit is used to construct a regularized vector autoregression model for muscle state vectors, extract corresponding muscle pairs after fitting the model, analyze the causal relationship between muscles and calculate relevant statistics, and construct a sparse muscle co-operation graph. The fusion synergy graph generation unit is used to construct a nonlinear fractional vector autoregressive model, process muscle state time series and analyze relevant causal relationships, construct corresponding synergy graphs, and weightedly fuse various synergy graphs to obtain a fused muscle synergy graph. The functional group collaborative computing unit is used to divide facial muscle functional groups based on facial anatomy, perform hierarchical clustering of each functional group to divide sub-functional groups, calculate the multi-scale cross-correlation function of muscles within the group, label strong collaborative functional groups and obtain the functional group synergy degree. The critical path screening unit is used to construct a hierarchical fractional directed muscle propagation graph based on the fusion of muscle synergy graph and functional group synergy, extract the main propagation paths and screen the critical propagation paths.
[0101] Furthermore, the aforementioned functional group collaborative computing unit includes: The weighted correlation matrix unit is used to divide the muscle region into facial muscle functional groups based on facial anatomy, calculate the cross-correlation function of muscle pairs within each functional group and construct the cross-correlation matrix, and then weight the cross-correlation matrix by combining anatomical weights to obtain the weighted cross-correlation matrix. The spectral embedding decision graph unit is used to construct a similarity graph for each functional group based on the weighted cross-correlation matrix, calculate the Laplacian matrix of the graph and extract feature vectors to construct spectral embedding features, and at the same time calculate the local density and relative distance of muscle pairs to construct a decision graph, thus obtaining spectral embedding features and decision graph; The sub-functional group clustering unit is used to perform hierarchical clustering of each functional group based on spectral embedding features and decision maps. It combines fuzzy C-means clustering and density peak clustering to determine the cluster center, divide the sub-functional groups, calculate the fuzzy compactness and density compactness of the sub-functional groups, label the strong cooperating sub-groups, and obtain the sub-functional group division results. The synergy calculation unit is used to calculate the multi-scale cross-correlation function of muscle pairs within a group based on the sub-functional group division results. It extracts the cross-correlation coefficients at different time scales through wavelet transform, verifies multi-scale synchronicity, and labels strongly synergistic functional groups to obtain the functional group synergy.
[0102] Furthermore, the aforementioned spectral embedding decision graph unit includes: The multimodal matrix construction unit is used to calculate the weighted cross-correlation function of muscle pairs within each functional group. It introduces the anatomical weight matrix and integrates multimodal features. It uses specified operations to fuse the anatomical weights with the cross-correlation matrices of each modality to construct the multimodal weighted cross-correlation matrix. The hypergraph decision graph construction unit is used to identify higher-order associations between muscles using tensor decomposition, construct a hypergraph association matrix, calculate the local density and relative distance of muscle pairs and determine an adaptive density threshold, and construct a hypergraph similarity structure and a multimodal local density decision graph. The hypergraph enhancement and verification unit is used to enhance and suppress muscle edges based on an adaptive density threshold, construct the Laplacian matrix of the multimodal hypergraph, calculate and verify the hypermodality, and if it fails to meet the standard, it returns to adjust the decomposition parameters and density threshold to obtain the enhanced multimodal hypergraph. The spectral embedding feature output unit is used to calculate the eigenvalues and eigenvectors of the Laplacian matrix of the multimodal hypergraph, construct and fuse higher-order and local spectral embedding features, verify the feature reconstruction error, and output the multimodal hypergraph and fused spectral embedding features if the target is met.
[0103] Furthermore, the aforementioned time series modeling and filtering module 400 includes: The weighted adaptive unit is used to calculate the pain contribution weight and motion energy curve based on the muscle state vector, perform weighted adaptive weighting on the weighted spatiotemporal features and generate a candidate window set to obtain the weighted adaptive spatiotemporal features. The temporal modeling output unit is used to input the weighted adaptive spatiotemporal features into the temporal modeling network, perform local temporal modeling based on the candidate window set, and output the pain level probability distribution sequence. The micro-expression fusion unit is used to extract the intensity spectrum by performing temporal transformation on the enhanced micro-expression fragment features, and to perform multi-scale fusion based on the pain level probability distribution sequence and the intensity spectrum to obtain multi-scale micro-expression fusion features. The optimal window selection unit is used to select the optimal spatiotemporal window from multi-scale micro-expression fusion features based on the confidence ranking of the pain level probability distribution sequence and the energy distribution of the intensity spectrum.
[0104] Furthermore, the aforementioned time-series modeling output unit includes: The feature preprocessing unit is used to perform residual supplementation, spatiotemporal decoupling and pain semantic alignment on the weighted adaptive spatiotemporal features. It calculates the muscle activation intensity of the candidate window by combining the muscle state vector, determines the modeling depth according to the motion energy gradient, and divides the feature interval by window to obtain the spatiotemporal feature set of window decoupling semantic residual and window modeling configuration information. The temporal feature extraction unit is used to input the feature set into the temporal modeling network according to the window modeling configuration information, perform hierarchical local temporal modeling on each feature interval, extract the deep temporal features of each candidate window, and obtain the semantic deep temporal feature set. The probabilistic feature aggregation unit is used to integrate semantic deep temporal feature sets. After temporal fusion, residual correction is performed based on decoupled features. Multi-scale attention-weighted aggregation is performed by combining muscle activation intensity and motion energy gradient to obtain the fused aggregated probabilistic feature sequence. The probability sequence output unit is used to perform spatiotemporal dimension aggregation processing on the fused and aggregated probability feature sequence, perform global temporal regularization optimization, and output the pain level probability distribution sequence.
[0105] Furthermore, the aforementioned temporal feature extraction unit includes: The frame-level alignment and decoupling unit is used to perform frame-level temporal alignment of the feature set, and simultaneously complete semantic feature enhancement, residual feature completion and multi-scale feature mapping. It combines the muscle state vector to calculate the muscle activation intensity of the candidate window and performs spatiotemporal decoupling to obtain the hierarchical window configuration and decoupling alignment semantic residual multi-scale feature interval set. The hierarchical temporal modeling unit is used to input the feature interval set into the temporal modeling network according to the hierarchical window configuration, perform spatiotemporal decoupling local temporal modeling that combines hierarchical and scale-based methods, extract the initial decoupling depth temporal features, and obtain a multi-scale hierarchical decoupling temporal feature set. The semantic feature correction unit is used to extract semantic associations and enhance muscle semantics on the feature set. It sequentially completes feature hierarchical aggregation, multi-scale fusion, residual supplementation and temporal feature correction to obtain a semantically enhanced and corrected deep temporal feature set. The temporal feature integration unit is used to integrate the feature set, perform global regularization according to the temporal relationship of the window, complete semantic scale adaptation and inter-level feature fusion, and obtain a semantic deep temporal feature set.
[0106] Furthermore, the aforementioned hierarchical time series modeling unit includes: The scale-level decoupling unit is used to perform scale-level stratification and hierarchical division of the feature interval set according to the hierarchical window configuration, and simultaneously complete the window hierarchy mapping and spatiotemporal decoupling calibration. It calculates the muscle activation intensity by combining the muscle state vector, extracts the micro-expression intensity spectrum through temporal Fourier transform, and obtains the decoupling spatial feature set and decoupling temporal feature set through spatiotemporal decoupling. The frequency domain subscale modeling unit is used to input the decoupled spatial feature set into the frequency domain branch of the time-series modeling network according to the configuration, perform frequency domain subscale modeling, and extract the frequency domain spatial feature set; and combine the decoupled temporal feature set with the micro-expression intensity spectrum to perform frequency domain subscale modeling and extract the frequency domain temporal feature set. The frequency domain feature mining unit is used to perform deep feature mining on two types of frequency domain feature sets, enhance the deep feature expression of candidate windows, and obtain a scale-level hierarchical frequency domain feature subset. The frequency domain feature integration unit is used to integrate all feature subsets. After frequency domain time-series fusion, the features are sorted by scale and hierarchy to obtain a multi-scale hierarchical decoupled time-series feature set.
[0107] Furthermore, the aforementioned verification result output module 500 includes: The dynamic consistency unit is used to verify muscle dynamic consistency based on muscle state vectors. It calculates prediction error through filtering, constructs muscle co-operation graph and uses graph neural network to learn node embedding, calculates relevant parameters to construct joint consistency index, and outputs prediction error matrix, node embedding vector and joint consistency score. The temporal consistency unit is used to verify the temporal consistency of micro-expressions based on multi-scale micro-expression fusion features. It constructs a probabilistic model and generates predicted samples through sampling, calculates the log likelihood, uses a generative adversarial network to generate samples and calculates the relevant loss, constructs a temporal consistency index and outputs the corresponding score and sample-related features. The window validity unit is used to verify the validity of the optimal spatiotemporal window. It calculates the relevant attention weight distribution and sample diversity within the window, integrates relevant indicators to obtain the window quality score, combines the two types of consistency scores to construct the overall verification framework, and outputs the window validity result and the comprehensive consistency score. The result optimization output unit is used to perform joint feedback optimization by means of gradient descent, model parameter update, etc. if the overall validation fails, recalculate each index and verify the optimization effect; if all criteria are met, the final pain level and confidence index are output; if they are still not met, the process continues to return to adjust the parameters and reprocess.
[0108] Furthermore, the aforementioned timing consistency unit includes: The probability model construction unit is used to construct a Bayesian hidden Markov joint probability model based on multi-scale micro-expression fusion features, define the temporal state of micro-expressions, estimate the state transition matrix using Bayesian inference, estimate the posterior distribution using variational inference, construct the joint probability distribution, and record the state transition probability and observation likelihood. The prediction sample generation unit is used to jointly generate prediction samples using a normalized flow model and a diffusion probability model, generate intermediate samples through standard normal distribution sampling and inverse transformation, progressively denoise the intermediate samples, calculate the flow likelihood and diffusion loss, construct the prediction sample distribution, and record the state label and noise level of the samples. Adversarial network building units are used to construct recurrent generative adversarial networks and self-attention mechanisms. The generator adopts an encoder and decoder structure, and the discriminator uses a convolutional network. Cyclic consistency is calculated through cyclic mapping to obtain attention-weighted temporal features. The consistency index output unit is used to integrate observation likelihood, flow likelihood, diffusion loss, attention loss, and cyclic consistency loss to construct a comprehensive temporal consistency index. It calculates the Euclidean distance between the prediction and the true features and verifies the effectiveness of the index. If the index is met, it outputs the predicted sample distribution, attention weight, and comprehensive temporal consistency score. If the index is not met, it returns to adjust the model parameters and reprocess.
[0109] Reference Figure 3 This application also provides a computer device, which may be a server, and its internal structure may be as follows: Figure 3 As shown, this computer device includes a processor, memory, network interface, and database connected via a bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores operations, computer programs, and the database. The internal memory provides an environment for the operation of the operations and computer programs stored in the non-volatile storage media. The database stores data such as pain expression detection methods. The network interface is used for communication with external terminals via a network connection. When executed by a processor, this computer program implements a pain expression detection method, including: acquiring facial video streams, performing facial key point localization and frame alignment, segmenting facial muscle regions and extracting texture features to obtain an aligned frame sequence and a set of texture features for muscle regions; calculating motion energy based on the texture feature set, extracting and enhancing micro-expression temporal segments from the aligned frame sequence, extracting multi-scale spatiotemporal features, constructing a facial muscle dynamics model and performing state estimation to obtain enhanced micro-expression segments, multi-scale spatiotemporal features, and muscle state vectors; performing spatiotemporal attention weighted fusion based on multi-scale spatiotemporal features to obtain weighted spatiotemporal features, analyzing muscle synergy relationships in conjunction with muscle state vectors, and generating an adaptive candidate spatiotemporal window set; performing temporal modeling on the dynamic fusion of weighted spatiotemporal features to obtain a pain level probability distribution, fusing enhanced micro-expression segment features to obtain multi-scale micro-expression fusion features, and selecting the optimal spatiotemporal window; verifying muscle dynamics consistency based on muscle state vectors, verifying micro-expression temporal consistency based on multi-scale micro-expression fusion features, and verifying the validity of the optimal spatiotemporal window; if any verification fails, returning to the corresponding step for reprocessing; if all verifications pass, outputting the final pain level.
[0110] One embodiment of this application also provides a computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements a pain expression detection method, including: acquiring a facial video stream, performing facial key point localization and inter-frame alignment, dividing facial muscle regions and extracting texture features to obtain an aligned frame sequence and a set of texture features of muscle regions; calculating motion energy based on the texture feature set, extracting and enhancing micro-expression temporal segments from the aligned frame sequence, extracting multi-scale spatiotemporal features, constructing a facial muscle dynamics model and completing state estimation to obtain enhanced micro-expression segments, multi-scale spatiotemporal features, and muscle state vectors. Based on multi-scale spatiotemporal features, spatiotemporal attention-weighted fusion is performed to obtain weighted spatiotemporal features. Muscle synergy is analyzed in conjunction with muscle state vectors to generate an adaptive candidate spatiotemporal window set. After dynamic fusion of weighted spatiotemporal features, temporal modeling is performed to obtain the probability distribution of pain level. Enhanced micro-expression fragment features are fused to obtain multi-scale micro-expression fusion features, and the optimal spatiotemporal window is selected. Muscle dynamics consistency is verified based on muscle state vectors, and micro-expression temporal consistency is verified based on multi-scale micro-expression fusion features. The validity of the optimal spatiotemporal window is verified. If any verification fails, the corresponding step is returned for reprocessing. If all verifications pass, the final pain level is output.
[0111] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media provided in this application and used in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual-speed SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0112] The above description is only a preferred embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural changes made based on the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A method for detecting pain expression, characterized in that, include: Acquire facial video streams, perform facial key point localization and frame alignment, segment facial muscle regions and extract texture features to obtain aligned frame sequences and muscle region texture feature sets. Motion energy is calculated based on texture feature set. Micro-expression temporal segments are extracted and enhanced from aligned frame sequence. Multi-scale spatiotemporal features are extracted, facial muscle dynamics model is constructed and state estimation is completed, resulting in enhanced micro-expression segments, multi-scale spatiotemporal features and muscle state vectors. Based on multi-scale spatiotemporal features, spatiotemporal attention weighted fusion is performed to obtain weighted spatiotemporal features. Combined with muscle state vector analysis, muscle synergy relationship is generated to generate an adaptive candidate spatiotemporal window set. After dynamic fusion of weighted spatiotemporal features, temporal modeling is used to obtain the probability distribution of pain level. Multi-scale micro-expression fusion features are obtained by fusing enhanced micro-expression fragment features, and the optimal spatiotemporal window is selected. Muscle dynamic consistency is verified based on muscle state vectors, micro-expression temporal consistency is verified based on multi-scale micro-expression fusion features, and the effectiveness of the optimal spatiotemporal window is verified. If any verification fails, the corresponding step is returned for reprocessing. If all verifications pass, the final pain level is output.
2. The pain expression detection method according to claim 1, characterized in that, The steps of acquiring facial video streams, performing facial landmark localization and inter-frame alignment, segmenting facial muscle regions and extracting texture features to obtain aligned frame sequences and muscle region texture feature sets include: Collect bimodal facial video streams and resting segments, construct individual muscle feature baselines and pre-train muscle dynamics models, combine trajectories and baselines to locate facial key points in the video streams, and obtain a set of bimodal key point coordinates; Based on the coordinate set, the dual-modal fusion weight is calculated to perform inter-frame predictive joint alignment. At the same time, optical flow field features are extracted to realize cross-frame association of key points, resulting in a predicted aligned frame sequence. Based on the aligned frame sequence, combined with facial anatomical atlas and muscle dynamics model state parameters, refined muscle regions are divided, and muscle region division results are generated. Based on the segmentation results, bimodal texture features are extracted and differiated from the individual muscle feature baseline. Cross-frame associated texture features are then fused to obtain a set of texture features for the muscle region.
3. The pain expression detection method according to claim 1, characterized in that, The steps of calculating motion energy based on texture feature sets, extracting and enhancing micro-expression temporal segments from aligned frame sequences, extracting multi-scale spatiotemporal features, constructing a facial muscle dynamics model and completing state estimation to obtain enhanced micro-expression segments, multi-scale spatiotemporal features, and muscle state vectors include: A multi-scale feature pyramid is constructed based on the texture feature set. The motion energy at each scale is calculated and the optical flow field features are fused to obtain a multi-scale motion energy set with enhanced optical flow. By combining spatial attention mechanisms to locate key pain areas, and extracting micro-expression temporal segments containing peak frames based on the hierarchical extraction of motion energy sets, key features are enhanced through attention enhancement networks to obtain multi-scale enhanced micro-expression segments; Spatial features, temporal dependence features, and optical flow temporal features at various scales are extracted from the fragments and fused to form multi-scale spatiotemporal features; By combining bimodal information and attention weights, a facial muscle dynamics model is constructed, and the state transition matrix and noise covariance matrix are optimized. Input the texture feature set to complete the state estimation and obtain the muscle state vector.
4. The pain expression detection method according to claim 1, characterized in that, The steps of obtaining weighted spatiotemporal features by performing spatiotemporal attention weighted fusion based on multi-scale spatiotemporal features, combining muscle state vectors to analyze muscle synergy, and generating an adaptive candidate spatiotemporal window set include: Multi-scale channel decomposition is performed on multi-scale spatiotemporal features. Spatial attention weights are calculated independently for each group and spatial attention maps are generated. Pain-related regions are identified through dynamic thresholding to obtain spatial weighted features. Multi-head self-attention module and causal convolution module are constructed for temporal features. The outputs of the two modules are weighted and fused to obtain temporal weighted features. Then, spatial weighted features and temporal weighted features are fused to obtain weighted spatiotemporal features. Based on muscle state vectors, a vector autoregression model is constructed to analyze the causal relationship between muscles and construct a muscle synergy graph. Facial muscle function groups are divided to calculate the synergy within each group. A directed graph of muscle propagation is constructed and key propagation paths are selected. Candidate spatiotemporal windows are generated based on micro-expression fragments, muscle coordination maps, functional group coordination, and key propagation paths. The window quality score is calculated and the feature consistency is verified. If the score is met, it is determined as an adaptive candidate spatiotemporal window set; otherwise, it is returned to adjust the dynamic coefficients and reprocessed.
5. The pain expression detection method according to claim 4, characterized in that, The steps of constructing a vector autoregression model based on muscle state vectors, analyzing causal relationships between muscles and constructing a muscle synergy graph, dividing facial muscle function groups and calculating intra-group synergy, constructing a directed muscle propagation graph and screening key propagation paths include: A regularized vector autoregression model is constructed for the muscle state vector. After fitting the model, the corresponding muscle pairs are extracted, the causal relationship between muscles is analyzed and the relevant statistics are calculated, and a sparse muscle synergy graph is constructed. A nonlinear fractional vector autoregressive model is constructed to process the time series of muscle states and analyze the relevant causal relationships. A corresponding synergy graph is constructed, and various synergy graphs are weighted and fused to obtain a fused muscle synergy graph. Based on facial anatomy, facial muscle function groups are divided, hierarchical clustering is performed on each function group to divide it into sub-function groups, multi-scale cross-correlation functions of muscles within the group are calculated, strong synergistic function groups are labeled, and the synergistic degree of function groups is obtained. By integrating muscle synergy graphs and functional group synergy, a hierarchical fractional directed muscle propagation graph is constructed to extract the main propagation paths and screen key propagation paths.
6. The pain expression detection method according to claim 1, characterized in that, The steps of dynamically fusing weighted spatiotemporal features to obtain a pain level probability distribution through temporal modeling, fusing enhanced micro-expression fragment features to obtain multi-scale micro-expression fusion features, and selecting the optimal spatiotemporal window include: Pain contribution weights and motion energy curves are calculated based on muscle state vectors. Weighted spatiotemporal features are then subjected to adaptive weighting and a candidate window set is generated to obtain weighted adaptive spatiotemporal features. The weighted adaptive spatiotemporal features are input into the temporal modeling network, and local temporal modeling is performed based on the candidate window set to output a pain level probability distribution sequence. Temporal transformation is performed on the enhanced micro-expression fragment features to extract the intensity spectrum. Multi-scale fusion is then performed based on the pain level probability distribution sequence and the intensity spectrum to obtain multi-scale micro-expression fusion features. Based on the confidence ranking of the pain level probability distribution sequence and the energy distribution of the intensity spectrum, the optimal spatiotemporal window is selected from the multi-scale micro-expression fusion features.
7. The pain expression detection method according to claim 6, characterized in that, The steps of inputting weighted adaptive spatiotemporal features into a temporal modeling network, performing local temporal modeling based on a candidate window set, and outputting a pain level probability distribution sequence include: Residual supplementation, spatiotemporal decoupling and pain semantic alignment are performed on the weighted adaptive spatiotemporal features. The muscle activation intensity of the candidate window is calculated by combining the muscle state vector. The modeling depth is determined according to the motion energy gradient. The feature interval is divided by windowing to obtain the spatiotemporal feature set of window decoupling semantic residuals and window modeling configuration information. The feature set is input into the temporal modeling network according to the window modeling configuration information. Hierarchical local temporal modeling is performed on each feature interval to extract the deep temporal features of each candidate window, thus obtaining a semantic deep temporal feature set. Integrate semantic deep temporal feature sets, perform residual correction based on decoupled features after temporal fusion, and combine muscle activation intensity and motion energy gradient for multi-scale attention weighted aggregation to obtain fused aggregated probability feature sequence; The spatiotemporal dimension aggregation of the fused and aggregated probability feature sequence is processed, global temporal regularization optimization is performed, and the pain level probability distribution sequence is output.
8. The pain expression detection method according to claim 7, characterized in that, The steps of inputting the feature set into the temporal modeling network according to the window modeling configuration information, performing hierarchical local temporal modeling on each feature interval, extracting the deep temporal features of each candidate window, and obtaining the semantic deep temporal feature set include: Frame-level temporal alignment is performed on the feature set, and semantic feature enhancement, residual feature completion and multi-scale feature mapping are completed simultaneously. The muscle activation intensity of the candidate window is calculated by combining the muscle state vector and spatiotemporal decoupling is performed to obtain the hierarchical window configuration and decoupled alignment semantic residual multi-scale feature interval set. The feature interval set is configured into a hierarchical window and input into the temporal modeling network. The spatiotemporal decoupling local temporal modeling is performed by combining hierarchical and scale-based methods. The initial decoupling depth temporal features are extracted to obtain a multi-scale hierarchical decoupling temporal feature set. Semantic association extraction and muscle semantic enhancement are performed on the feature set. Feature hierarchical aggregation, multi-scale fusion, residual supplementation and temporal feature correction are completed in sequence to obtain a semantically enhanced and corrected deep temporal feature set. By integrating this feature set and performing global regularization according to the temporal relationship of the window, semantic scale adaptation and inter-level feature fusion are completed to obtain a semantic deep temporal feature set.
9. The pain expression detection method according to claim 1, characterized in that, The steps for verifying muscle dynamic consistency based on muscle state vectors, verifying micro-expression temporal consistency based on multi-scale micro-expression fusion features, and checking the effectiveness of the optimal spatiotemporal window, where failure in any verification results in reprocessing of the corresponding step, and success in all verifications results in the output of the final pain level, include: Based on the muscle state vector, the consistency of muscle dynamics is verified. The prediction error is calculated by filtering, a muscle co-evaluation graph is constructed, and node embeddings are learned by graph neural network. Relevant parameters are calculated to construct a joint consistency index, and the prediction error matrix, node embedding vector and joint consistency score are output. Based on the multi-scale micro-expression fusion features, the temporal consistency of micro-expressions is verified. A probabilistic model is constructed and predicted samples are generated by sampling. Log likelihood is calculated. A generative adversarial network is used to generate samples and calculate relevant losses. A temporal consistency index is constructed and the corresponding score and sample-related features are output. Verify the effectiveness of the optimal spatiotemporal window, calculate the relevant attention weight distribution and sample diversity within the window, integrate relevant indicators to obtain the window quality score, combine the two types of consistency scores to construct an overall verification framework, and output the window effectiveness result and the comprehensive consistency score. If the overall validation fails, joint feedback optimization is performed using methods such as gradient descent and model parameter updates to recalculate each indicator and verify the optimization effect; if all indicators meet the criteria, the final pain level and confidence index are output; if they still do not meet the criteria, the process is returned to adjust the parameters and reprocess.
10. A pain expression detection system, characterized in that, include: The feature acquisition module is used to acquire facial video streams, perform facial key point localization and inter-frame alignment, divide facial muscle regions and extract texture features, and obtain aligned frame sequences and muscle region texture feature sets. The micro-expression enhancement module is used to calculate motion energy based on texture feature set, extract and enhance micro-expression temporal segments from aligned frame sequence, extract multi-scale spatiotemporal features, construct facial muscle dynamics model and complete state estimation, and obtain enhanced micro-expression segments, multi-scale spatiotemporal features and muscle state vectors. The fusion window generation module is used to perform spatiotemporal attention weighted fusion based on multi-scale spatiotemporal features to obtain weighted spatiotemporal features, and combine muscle state vectors to analyze muscle synergy relationships to generate an adaptive candidate spatiotemporal window set. The temporal modeling and filtering module is used to obtain the pain level probability distribution by dynamically fusing weighted spatiotemporal features, fusing enhanced micro-expression fragment features to obtain multi-scale micro-expression fusion features, and filtering out the optimal spatiotemporal window. The verification result output module is used to verify the consistency of muscle dynamics based on muscle state vectors, verify the consistency of micro-expression temporal sequence based on multi-scale micro-expression fusion features, and verify the validity of the optimal spatiotemporal window. If any verification fails, the corresponding step is returned for reprocessing. If all verifications pass, the final pain level is output.