Roadside driver state recognition method and system fusing visual and infrared data

By integrating visual and infrared data, a roadside driver status recognition method has been developed, which solves the problems of insufficient environmental adaptability and recognition accuracy in existing technologies. This enables all-weather driver status monitoring and early warning, and improves the robustness and accuracy of the recognition system.

CN121121710BActive Publication Date: 2026-03-03BEIJING YIZHUANG INTELLIGENT CITY RES INST GRP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511654473.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-12
Publication Date
2026-03-03
Estimated Expiration
2045-11-12

AI Technical Summary

Technical Problem

Existing roadside driver status recognition technologies show a significant decrease in recognition accuracy under changing lighting conditions, at night, or in adverse weather conditions, lacking environmental adaptability. Furthermore, traditional methods do not extract driver facial features comprehensively enough, resulting in low accuracy in judging abnormal states, especially with insufficient sensitivity to subtle changes in state.

Method used

Image data is collected by roadside visual cameras and infrared cameras, a feature pyramid is established, cross-modal correlation coefficient registration is performed, a complementary weight mapping table is constructed, the linkage motion features of facial functional areas are extracted, an anomaly discrimination index is generated, and adaptive fusion and accurate recognition are achieved.

Benefits of technology

It enables all-weather driver status monitoring, improves the stability and accuracy of identification, and can promptly identify dangerous driving conditions such as fatigue and inattention, providing early warning information and enhancing road traffic safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121121710B_ABST
    Figure CN121121710B_ABST
Patent Text Reader

Abstract

The application provides a roadside driver state recognition method and system fusing visual and infrared data, relates to the technical field of intelligent traffic safety monitoring, and comprises the following steps: collecting data through a visual camera and an infrared camera, establishing a feature pyramid and performing cross-modal registration, constructing a complementary weight mapping table for adaptive fusion, extracting driver facial features to construct a connection graph and a deformation mode, comparing normal driving benchmark features to generate an abnormality discrimination index, and realizing accurate recognition and early warning of the driver state, so that the accuracy and robustness of driver state recognition under complex lighting and environmental conditions can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent traffic safety monitoring technology, and in particular to a roadside driver status recognition method and system that integrates visual and infrared data. Background Technology

[0002] With the continuous development of intelligent transportation systems, roadside driver status recognition technology is playing an increasingly important role in road safety monitoring and vehicle management. Driver status includes abnormal states such as fatigued driving, inattentiveness, and drunk driving, which can lead to traffic accidents. Traditional driver status recognition mainly relies on in-vehicle monitoring systems, such as steering wheel sensors and biosensors, but these systems often require vehicle modifications, resulting in high implementation costs. Roadside monitoring, as a non-intrusive monitoring method, deploys monitoring equipment on both sides of the road to monitor the driver status of passing vehicles in real time without requiring any modifications to the vehicles themselves, and has great application potential.

[0003] Currently, roadside driver status recognition technology faces the following shortcomings: Most existing roadside monitoring systems rely on single-modal data, such as analyzing only visible light visual images. Recognition accuracy drops significantly under varying lighting conditions, at night, or inclement weather, lacking adaptability to different environmental conditions. Traditional image processing methods are not comprehensive or accurate enough in extracting driver facial features. Most methods focus on static features while ignoring the interconnected motion features between different facial functional areas, resulting in low accuracy in judging abnormal driver states, especially insufficient sensitivity to subtle changes. Existing technologies generally employ simple data overlay or averaging methods for feature fusion, failing to fully consider the complementarity and redundancy between different modal data. They cannot adaptively adjust the weights of different features according to changes in the actual scene, thus limiting the system's robustness and accuracy in complex environments. Summary of the Invention

[0004] This invention provides a roadside driver status recognition method and system that integrates visual and infrared data, which can solve the problems in the prior art.

[0005] A first aspect of the present invention provides a roadside driver status recognition method that integrates visual and infrared data, comprising:

[0006] Images of the driver's cab area of ​​the target vehicle are acquired by roadside vision cameras and infrared cameras, and preprocessed to obtain visual image data and infrared image data.

[0007] A feature pyramid is constructed from visual image data and infrared image data. A feature parameter association table is built using cross-modal correlation coefficients and spatiotemporal registration is performed to obtain a registration feature set. A complementary weight mapping table is constructed based on the feature parameter association table, and the registration feature set is adaptively fused to obtain a complementary enhanced feature map.

[0008] The driver's facial region is extracted from the complementary enhanced feature map, and the corresponding three-dimensional geometric parameters and depth parameters are calculated. The feature parameters are then combined to generate a multi-dimensional feature sequence.

[0009] A feature point connection map is established based on facial feature points in a multidimensional feature sequence. Functional regions are divided by deformation patterns, and the linkage motion features of each functional region of the face are extracted.

[0010] Based on the linkage motion features, a baseline feature for normal driving behavior is constructed. The multi-dimensional difference value of the current feature is calculated through feature mapping rules to generate an anomaly discrimination index for the driver's state.

[0011] The driver's status is determined based on the anomaly detection index, and a warning message is generated when the anomaly detection index exceeds a preset anomaly threshold.

[0012] In one optional embodiment, a feature pyramid is constructed from visual image data and infrared image data. A feature parameter association table is built using cross-modal correlation coefficients, and spatiotemporal registration is performed to obtain a registration feature set including:

[0013] A feature pyramid is constructed for visual image data and infrared image data. Texture features, edge features, gradient features and statistical features are extracted from different scale layers of the feature pyramid. The features from different scale layers are combined to generate visual feature sets and infrared feature sets.

[0014] The visual feature set and the infrared feature set are standardized, and the local mutual information and global mutual information between the visual feature set and the infrared feature set are calculated to generate cross-modal correlation coefficients.

[0015] Based on the cross-modal correlation coefficient, the similarity of feature parameters is calculated. The visual feature set and the infrared feature set are grouped according to the similarity of feature parameters to determine the feature groups and construct a feature parameter association table.

[0016] Spatiotemporal registration of the visual feature set and the infrared feature set is performed to obtain the registered feature set.

[0017] In one optional embodiment, a complementary weight mapping table is constructed based on a feature parameter association table, and the registered feature set is adaptively fused to obtain a complementary enhanced feature map, including:

[0018] Calculate the information entropy and dispersion of each feature group in the feature parameter association table, and construct a complementary weight mapping table based on the information entropy and the dispersion. The complementary weight mapping table includes inter-group weights and intra-group weights.

[0019] The complementary weight mapping table is used to perform inter-group weighting and intra-group weighting on the registration feature set to generate weighted features;

[0020] Based on the distribution characteristics of the weighted features, an adaptive fusion function is constructed, which includes linear fusion components and nonlinear fusion components.

[0021] The weighted features are fused using the adaptive fusion function to generate an initial feature map. The initial feature map is then subjected to spatial domain filtering and frequency domain enhancement to obtain a complementary enhanced feature map.

[0022] In one optional embodiment, a feature point connection map is established based on facial feature points in a multidimensional feature sequence. Functional regions are divided by deformation patterns, and the linkage motion features of each functional region of the face are extracted, including:

[0023] A feature point connection graph is established based on the spatial location of facial feature points. Each node in the feature point connection graph corresponds to a facial feature point, and the connection weight between nodes is determined according to the degree of anatomical correlation.

[0024] Calculate the displacement vector of each node in the feature point connection graph, determine the deformation mode of the feature point connection graph based on the displacement vector, and extract the deformation direction and deformation amplitude in the deformation mode;

[0025] Based on the deformation pattern of the feature point connection map, the facial feature points are clustered, and the facial feature points whose deformation direction angle is less than a preset angle and whose deformation amplitude difference is less than a preset amplitude difference are divided into the same functional region, and the deformation features of the functional region are established.

[0026] Construct a deformation transfer network between the functional regions, including forward transfer connections and reverse inhibition connections, to obtain the transfer paths between the functional regions;

[0027] Based on the transmission path and the deformation features, the motion correlation between adjacent functional regions in the deformation transmission network is calculated, the temporal sequence and amplitude change relationship of the functional regions are extracted, and the linkage motion features of each functional region of the face are determined.

[0028] In one optional embodiment, a deformation transfer network is constructed between the functional regions, including forward transfer connections and reverse inhibition connections, to obtain the transfer paths between the functional regions, including:

[0029] Based on the boundary feature points of adjacent functional regions, the deformation gradient between adjacent functional regions is calculated. The deformation gradient includes the deformation direction gradient and the deformation amplitude gradient at the boundary feature points.

[0030] Collect the change data of the deformation gradient within a continuous time window, and calculate the stable duration and fluctuation range of the deformation gradient based on the change data;

[0031] When the deformation gradient is greater than or equal to a preset positive gradient threshold and the stability duration is greater than a preset time threshold, a positive transit connection is established between adjacent functional regions. The weight of the positive transit connection is positively correlated with the deformation gradient and the stability duration.

[0032] When the deformation gradient is less than a preset negative gradient threshold and the fluctuation range is less than a preset fluctuation threshold, an inverse suppression connection is established between adjacent functional regions. The weight of the inverse suppression connection is positively correlated with the absolute value of the deformation gradient.

[0033] A deformation propagation network is constructed by combining the weights of forward propagation connections and reverse inhibition connections. The deformation propagation intensity is determined based on the weights of each connection in the deformation propagation network, and propagation paths whose deformation propagation intensity exceeds a preset regional association threshold are extracted.

[0034] In one optional embodiment, a baseline feature for normal driving behavior is constructed based on the linkage motion features. The multi-dimensional difference value of the current feature is calculated using feature mapping rules to generate an anomaly discrimination index for the driver's state, including:

[0035] The linkage motion features under standard driving scenarios are collected, and feature dimensionality reduction is performed using principal component analysis to obtain feature component groups. Based on the feature component groups, motion amplitude distribution parameters, relative motion relationship distribution parameters, and time series change law distribution parameters are calculated, and feature fusion is performed to obtain the baseline features of normal driving behavior.

[0036] A contrastive loss function is determined for the baseline features of normal driving behavior, including a feature vector distance calculation term, a feature distribution divergence calculation term, and a time series consistency constraint calculation term. Based on the contrastive loss function, a feature encoder is iteratively trained to obtain feature mapping rules.

[0037] Collect the linkage motion features of the current driving scenario, and convert them into the current feature representation through feature mapping rules. Based on the current feature representation and the normal driving behavior benchmark features, determine the distance metric value, distribution deviation value, and time series anomaly value.

[0038] The abnormal scores for single features, feature combinations, and time-series patterns are calculated separately, and then weighted according to preset abnormal weights to generate an abnormality discrimination index for the driver's state.

[0039] A second aspect of the present invention provides a roadside driver status recognition system that integrates visual and infrared data, comprising:

[0040] The first unit is used to acquire images of the driver's cab area of ​​the target vehicle through roadside vision cameras and infrared cameras, and to preprocess them to obtain visual image data and infrared image data.

[0041] The second unit is used to build feature pyramids from visual image data and infrared image data, construct feature parameter association tables through cross-modal correlation coefficients and perform spatiotemporal registration to obtain a registration feature set; based on the feature parameter association tables, a complementary weight mapping table is constructed, and the registration feature set is adaptively fused to obtain a complementary enhanced feature map.

[0042] The third unit is used to extract the driver's facial region from the complementary enhanced feature map, calculate the corresponding three-dimensional geometric parameters and depth parameters, and combine the feature parameters to generate a multi-dimensional feature sequence.

[0043] The fourth unit is used to build a feature point connection map based on facial feature points in a multidimensional feature sequence, divide functional regions by deformation patterns, and extract the linkage motion features of each functional region of the face.

[0044] The fifth unit is used to construct benchmark features for normal driving behavior based on linkage motion features, calculate the multi-dimensional difference values ​​of the current features through feature mapping rules, and generate an anomaly discrimination index for the driver's state.

[0045] The sixth unit is used to determine the driver's status based on the anomaly discrimination index, and to generate a warning message when the anomaly discrimination index exceeds a preset anomaly threshold.

[0046] A third aspect of the present invention provides an electronic device, comprising:

[0047] processor;

[0048] Memory used to store processor-executable instructions;

[0049] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0050] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0051] In this embodiment of the invention, a roadside driver state recognition method that integrates visual and infrared data achieves all-weather driver state monitoring, effectively solving the problem of poor monitoring performance of a single visual sensor in complex environments and improving the stability and accuracy of driver state recognition. A feature parameter association table constructed using feature pyramids and cross-modal correlation coefficients enables precise registration and adaptive fusion of visual and infrared data, enhancing the extraction capability of key features and effectively addressing the challenges of complex road conditions such as lighting changes and occlusion, ensuring the robustness of the recognition system. Based on the analysis of linked motion features of facial functional regions and an anomaly discrimination index mechanism, accurate identification of dangerous driving states such as driver fatigue and inattention is achieved, providing timely early warning information and possessing significant application value for road traffic safety. Attached Figure Description

[0052] Figure 1 This is a flowchart illustrating the roadside driver status recognition method that integrates visual and infrared data according to an embodiment of the present invention.

[0053] Figure 2 Flowchart for constructing an adaptive deformation transfer network. Detailed Implementation

[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0055] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0056] Figure 1 This is a flowchart illustrating the roadside driver state recognition method that integrates visual and infrared data according to an embodiment of the present invention. Figure 1 As shown, the method includes:

[0057] Images of the driver's cab area of ​​the target vehicle are acquired by roadside vision cameras and infrared cameras, and preprocessed to obtain visual image data and infrared image data.

[0058] A feature pyramid is constructed from visual image data and infrared image data. A feature parameter association table is built using cross-modal correlation coefficients and spatiotemporal registration is performed to obtain a registration feature set. A complementary weight mapping table is constructed based on the feature parameter association table, and the registration feature set is adaptively fused to obtain a complementary enhanced feature map.

[0059] The driver's facial region is extracted from the complementary enhanced feature map, and the corresponding three-dimensional geometric parameters and depth parameters are calculated. The feature parameters are then combined to generate a multi-dimensional feature sequence.

[0060] A feature point connection map is established based on facial feature points in a multidimensional feature sequence. Functional regions are divided by deformation patterns, and the linkage motion features of each functional region of the face are extracted.

[0061] Based on the linkage motion features, a baseline feature for normal driving behavior is constructed. The multi-dimensional difference value of the current feature is calculated through feature mapping rules to generate an anomaly discrimination index for the driver's state.

[0062] The driver's status is determined based on the anomaly detection index, and a warning message is generated when the anomaly detection index exceeds a preset anomaly threshold.

[0063] In one specific implementation, images of the driver's cab area of ​​the target vehicle are acquired using a roadside-mounted visual camera and an infrared camera. The visual camera acquires images in the visible light band with a resolution of 1920×1080 pixels and a frame rate of 30 frames per second; the infrared camera acquires thermal images with a resolution of 640×480 pixels and a frame rate of 25 frames per second. The acquired raw images undergo preprocessing, including denoising, correction, and normalization. Denoising employs a bilateral filtering method to preserve image edge information while reducing noise; correction includes geometric correction and color correction to eliminate lens distortion and ambient lighting differences; normalization adjusts the image pixel values ​​to the range of 0-1, facilitating subsequent feature extraction.

[0064] Feature pyramids were constructed for both preprocessed visual and infrared image data. Each feature pyramid consisted of five levels, with a resolution ratio of 1:2 between adjacent levels. At each level, corner features, edge features, and texture features were extracted. Corner features were obtained by calculating gradient changes in local image regions; edge features were obtained by detecting abrupt changes in pixel grayscale values; and texture features were obtained by analyzing the statistical characteristics of pixel distribution. After feature extraction, a feature parameter association table was constructed using cross-modal correlation coefficients. The correlation between corresponding regions in the visual and infrared images was calculated, and feature points with a correlation coefficient greater than 0.75 were paired and recorded in the association table. Each pairing included feature point coordinates, feature type, feature value, and correlation coefficient.

[0065] Spatiotemporal registration based on a feature parameter association table addresses the resolution differences and viewpoint deviations between visual and infrared images. A combination of rigid and non-rigid transformations maps feature points from the infrared image to the visual image coordinate system, resulting in a registration feature set. Rigid transformations handle overall displacement and rotation, while non-rigid transformations handle local deformation. Registration errors are controlled within 3 pixels.

[0066] A complementary weight mapping table is constructed based on the feature parameter association table. For each pair of registered features, the differences in their performance under different lighting and temperature conditions are analyzed and quantified as complementary weight values. The complementary weight values ​​range from 0 to 1, representing the degree of complementarity of the feature in the visual image and the infrared image. For example, under low lighting conditions, the feature weight in the infrared image is higher; when the temperature in the cab is uniform, the feature weight in the visual image is higher.

[0067] Adaptive fusion of the registered feature sets yields a complementary enhanced feature map. During fusion, the contribution of each feature point is dynamically adjusted according to a complementarity weight mapping table, preserving complementary information while suppressing redundant information. The fused feature map maintains the same resolution as the visual image but includes thermal information provided by the infrared image.

[0068] The driver's facial region is extracted from the complementary enhanced feature map, and the window region is detected. Then, the driver's head and face are located within the window region. For the detected facial region, corresponding 3D geometric and depth parameters are calculated. The geometric parameters include the facial contour point set, the position of facial features, and facial angles; the depth parameters are obtained through disparity estimation, representing the depth information of each part of the face. A multi-dimensional feature sequence is generated by combining the feature parameters, including temporal variation data of 68 key facial points.

[0069] A feature point connection map is constructed based on facial feature points in a multidimensional feature sequence. Sixty-eight key facial points are connected according to anatomical relationships to form a graph structure, with each edge representing the relationship between two feature points. By analyzing the deformation patterns of the feature point connection map, facial functional regions are divided, including the eye region, mouth region, forehead region, and jaw region. Linked motion features are extracted for each functional region to describe the coordinated change patterns of feature points within the region. These linked motion features include deformation amplitude, deformation speed, and deformation periodicity.

[0070] Benchmark features for normal driving behavior are constructed based on coordinated motion characteristics. Typical movement patterns of various facial functional areas are statistically analyzed from a large amount of normal driving data to establish reference standards. Benchmark features include normal range values ​​for indicators such as blink frequency (average 15-20 times per minute), mouth movement amplitude, and head posture stability. Multi-dimensional differences between the current features and benchmark features are calculated using feature mapping rules. These differences include frequency differences, amplitude differences, coordination differences, and duration differences.

[0071] A weighted combination of multi-dimensional difference values ​​generates an anomaly discrimination index for driver status. The weights are determined through training with a large amount of driving behavior data, reflecting the importance of each dimension's difference in judging an abnormal state. The anomaly discrimination index ranges from 0 to 100, with higher values ​​indicating a higher degree of anomaly. Driver status is judged based on the anomaly discrimination index. When the index exceeds a preset anomaly threshold (usually set to 75), the system considers the driver to be in a fatigued, distracted, or abnormal state, generating a warning message. Warning messages are divided into three levels: alert level (anomaly index 75-85), warning level (anomaly index 85-95), and emergency level (anomaly index above 95).

[0072] In one optional implementation, a feature pyramid is constructed from visual image data and infrared image data. A feature parameter association table is built using cross-modal correlation coefficients, and spatiotemporal registration is performed to obtain a registration feature set including:

[0073] A feature pyramid is constructed for visual image data and infrared image data. Texture features, edge features, gradient features and statistical features are extracted from different scale layers of the feature pyramid. The features from different scale layers are combined to generate visual feature sets and infrared feature sets.

[0074] The visual feature set and the infrared feature set are standardized, and the local mutual information and global mutual information between the visual feature set and the infrared feature set are calculated to generate cross-modal correlation coefficients.

[0075] Based on the cross-modal correlation coefficient, the similarity of feature parameters is calculated. The visual feature set and the infrared feature set are grouped according to the similarity of feature parameters to determine the feature groups and construct a feature parameter association table.

[0076] Spatiotemporal registration of the visual feature set and the infrared feature set is performed to obtain the registered feature set.

[0077] In one specific implementation, the feature pyramid can be constructed using the difference of Gaussian operator for visual image data and infrared image data. Specifically, the original visual image and infrared image are decomposed into four scale layers, each with an image size half that of the previous layer. Based on the original image (640×480 pixels), images of 320×240, 160×120, and 80×60 pixels are generated sequentially, forming a complete feature pyramid structure.

[0078] When extracting features at different scales of the feature pyramid, texture features are extracted using the Local Binary Pattern (LBP) operator, with a neighborhood radius of 2 and 8 sampling points, resulting in a 256-dimensional feature vector. Edge features are extracted using the Canny edge detector, with a low threshold set to 0.4 times the image mean and a high threshold set to 2.5 times the low threshold, resulting in a binary edge map. Gradient features are calculated using the Sobel operator to compute horizontal and vertical gradients, and a gradient direction histogram with 16 bins is constructed based on the gradient magnitude and direction. Statistical features include 10 statistical measures such as mean, variance, skewness, and kurtosis, calculated within a 9×9 local window. After extraction at each scale, the different features are concatenated to form the feature descriptor for that scale. Taking a 320×240 scale as an example, the extracted feature dimensions are: 256-dimensional texture features, 1-dimensional (binary) edge features, 16-dimensional gradient features, and 10-dimensional statistical features, totaling 283 dimensions. Subsequently, the features from the four scales are combined to generate visual and infrared feature sets.

[0079] When standardizing the visual and infrared feature sets, the Z-score standardization method is used, which involves subtracting the mean from each feature dimension and then dividing by the standard deviation. Taking the texture features of the infrared feature set as an example, the original feature values ​​range from [0, 255], and after standardization, they are distributed in the range of [-2.3, 2.7], with a mean of 0 and a standard deviation of 1.

[0080] When calculating the local and global mutual information between the visual and infrared feature sets, the local mutual information is calculated within a 15×15 window with a window sliding step of 5 pixels; the global mutual information is calculated based on the entire image. Taking a 160×120 scale layer as an example, the extracted local mutual information forms a 30×24 feature map, while the global mutual information is a single value. The calculation results of the local and global mutual information are normalized to the [0, 1] interval to generate a cross-modal correlation coefficient matrix.

[0081] When calculating feature parameter similarity based on cross-modal correlation coefficients, the Euclidean distance is calculated for each pair of visual and infrared features, with a threshold of 0.35. Features are considered to have high similarity when the normalized distance between them is less than the threshold. At an 80×60 scale layer, the similarity matrix size is 4800×4800, representing the similarity between features at each pixel location and features at other locations. When grouping the visual and infrared feature sets according to feature parameter similarity, a hierarchical clustering algorithm is used, with a set of 8 categories, dividing the features in both sets into 8 feature groups. Each feature group contains a set of highly similar features.

[0082] When constructing the feature parameter association table, a mapping matrix is ​​created, with a size equal to the number of visual feature groups multiplied by the number of infrared feature groups. Each element in the matrix represents the matching weight between the corresponding visual feature group and the infrared feature group. Taking actual data as an example, in the constructed 8×8 association table, the matching weights are distributed in the interval [0, 1], with larger values ​​indicating a higher degree of matching. For example, the matching weight between visual feature group 1 and infrared feature group 3 is 0.87, indicating a high correlation; while the matching weight between visual feature group 1 and infrared feature group 7 is only 0.12, indicating a low correlation.

[0083] When performing spatiotemporal registration of the visual and infrared feature sets, temporal registration is achieved through timestamp alignment, while spatial registration employs feature matching and affine transformation based on an association table. For feature matching, feature pairs with a matching weight greater than 0.7 are selected for precise matching. The RANSAC algorithm is used to remove erroneous matches, with a confidence level of 0.99 and 1000 iterations. Taking the original 640×480 resolution image as an example, approximately 2000 feature points are detected, and approximately 650 pairs are successfully matched. Based on the matched point pairs, the affine transformation matrix is ​​estimated, transforming the infrared image to the visual image coordinate system to achieve precise registration. The average pixel error after registration is less than 1.5 pixels.

[0084] The resulting registration feature set contains spatiotemporally aligned visual and infrared features, which can be directly used for subsequent visual tasks such as target detection and scene understanding. The registration feature set retains the original feature dimensions and adds a registration quality index, ranging from [0, 1], to represent the confidence level of the registration accuracy.

[0085] In one optional implementation, a complementary weight mapping table is constructed based on the feature parameter association table, and the registered feature set is adaptively fused to obtain a complementary enhanced feature map, including:

[0086] Calculate the information entropy and dispersion of each feature group in the feature parameter association table, and construct a complementary weight mapping table based on the information entropy and the dispersion. The complementary weight mapping table includes inter-group weights and intra-group weights.

[0087] The complementary weight mapping table is used to perform inter-group weighting and intra-group weighting on the registration feature set to generate weighted features;

[0088] Based on the distribution characteristics of the weighted features, an adaptive fusion function is constructed, which includes linear fusion components and nonlinear fusion components.

[0089] The weighted features are fused using the adaptive fusion function to generate an initial feature map. The initial feature map is then subjected to spatial domain filtering and frequency domain enhancement to obtain a complementary enhanced feature map.

[0090] In one specific implementation, it is assumed that the acquired feature parameter association table contains multiple sets of feature combinations, each set containing several feature parameters. For a given feature set, its information entropy and dispersion are calculated. Information entropy is calculated by the probability distribution of each feature parameter in the feature set, reflecting the richness of information within the feature set. For example, for a feature set containing three parameters—pixel brightness, texture complexity, and edge strength—with probability distributions of 0.3, 0.5, and 0.2 respectively, the calculated information entropy is 1.485. Dispersion is measured by the deviation of each parameter within the feature set from the mean, reflecting the degree of concentration or dispersion of the feature distribution. For the feature set with parameter values ​​of 85, 120, and 65, the calculated dispersion is 22.5.

[0091] Based on the calculated information entropy and dispersion, a complementarity weight mapping table is constructed. This table includes two parts: between-group weights and within-group weights. The between-group weight represents the degree of complementarity between one set of features and another set of features. It is calculated as a function of the ratio of the information entropy of the two sets of features and the difference in dispersion. Specifically, the closer the information entropy ratio is to 1 and the greater the difference in dispersion, the stronger the complementarity, and the higher the weight is assigned. For example, if the information entropy ratio of feature group 1 to feature group 2 is 0.92 and the dispersion difference is 35, the resulting between-group weight is 0.78. The within-group weight represents the importance of a feature within a group, determined by the ratio of the information content of that feature to the average information content within the group. For example, if the information contents of three features within a group are 0.7, 0.5, and 0.9, and the average information content is 0.7, then the within-group weights of the three features are 1.0, 0.71, and 1.29, respectively.

[0092] The pre-constructed complementary weight mapping table is used to perform inter-group and intra-group weighting on the registration feature set to generate weighted features. Assume the registration feature set contains three sets of features, each with four feature parameters. Inter-group weighting adjusts the overall contribution of different feature groups; for example, the overall weight of the first feature group is 0.35, the second is 0.4, and the third is 0.25. Intra-group weighting refines the contribution of each parameter within each group; for example, the weights of the four parameters in the first group are 1.2, 0.8, 1.1, and 0.9, respectively. Through these two levels of weighting (inter-group and intra-group), a weighted feature set that comprehensively considers complementarity is generated. For example, a parameter with an original feature value of 75 becomes 36 after being weighted with an inter-group weight of 0.4 and an intra-group weight of 1.2.

[0093] An adaptive fusion function is constructed based on the distribution characteristics of the weighted features. The adaptive fusion function includes linear and nonlinear fusion components. The distribution histogram of the weighted features is analyzed, and its statistical properties such as kurtosis, skewness, and variance are calculated. When the feature distribution approximates a normal distribution, the linear fusion component dominates; when the distribution exhibits multimodal or skewed characteristics, the nonlinear fusion component dominates. The linear fusion component is implemented using a weighted average, while the nonlinear fusion component is implemented using an exponential or logarithmic transformation. For example, for a feature distribution with a kurtosis of 3.2 and a skewness of 0.3, the weight of the linear component is set to 0.65, and the weight of the nonlinear component is set to 0.35. The linear part is weighted and summed, while the nonlinear part is summed after an exponential transformation, and the two parts are then combined proportionally.

[0094] An initial feature map is generated by fusing the weighted features using a constructed adaptive fusion function. Taking image processing as an example, assuming the image size is 640×480 pixels, the weighted features include three sets: edge features, texture features, and color features. For each pixel location, the adaptive fusion function is applied to calculate its fused feature value. For example, if the three weighted feature values ​​for a certain pixel location are (120, 85, 65), after calculation by a linear fusion component (weight 0.7) and a non-linear fusion component (weight 0.3), a fused value of 98 is obtained, which is used as the value of that pixel in the initial feature map.

[0095] Spatial domain filtering and frequency domain enhancement are performed on the initial feature map to obtain a complementary enhanced feature map. Spatial domain filtering uses a Gaussian filter with an adaptive kernel size. The kernel size is dynamically adjusted based on the complexity of local features; a small kernel (e.g., 3×3) is used to preserve details in complex regions, while a large kernel (e.g., 7×7) is used to suppress noise in flat regions. Frequency domain enhancement transforms the initial feature map to the frequency domain using a Fast Fourier Transform, enhancing the mid-to-high frequency components. The enhancement coefficient is adaptively set according to the spectral distribution characteristics of the feature map; typically, the mid-frequency enhancement coefficient is 1.5-2.5, and the high-frequency enhancement coefficient is 1.2-1.8. After spatial domain filtering removes noise, frequency domain enhancement further improves feature contrast, ultimately resulting in a complementary enhanced feature map that is rich in detail and has suppressed noise.

[0096] In one optional implementation, a feature point connection map is established based on facial feature points in a multidimensional feature sequence. Functional regions are divided by deformation patterns, and the linkage motion features of each functional region of the face are extracted, including:

[0097] A feature point connection graph is established based on the spatial location of facial feature points. Each node in the feature point connection graph corresponds to a facial feature point, and the connection weight between nodes is determined according to the degree of anatomical correlation.

[0098] Calculate the displacement vector of each node in the feature point connection graph, determine the deformation mode of the feature point connection graph based on the displacement vector, and extract the deformation direction and deformation amplitude in the deformation mode;

[0099] Based on the deformation pattern of the feature point connection map, the facial feature points are clustered, and the facial feature points whose deformation direction angle is less than a preset angle and whose deformation amplitude difference is less than a preset amplitude difference are divided into the same functional region, and the deformation features of the functional region are established.

[0100] Construct a deformation transfer network between the functional regions, including forward transfer connections and reverse inhibition connections, to obtain the transfer paths between the functional regions;

[0101] Based on the transmission path and the deformation features, the motion correlation between adjacent functional regions in the deformation transmission network is calculated, the temporal sequence and amplitude change relationship of the functional regions are extracted, and the linkage motion features of each functional region of the face are determined.

[0102] In one specific implementation, facial feature point data containing a multi-dimensional feature sequence is acquired. These feature points typically include feature points at key locations such as the eyes, eyebrows, nose, mouth, and facial contours, totaling 68 points. The spatial position of each feature point is represented by three-dimensional coordinates, where the first and second coordinates represent the position on a two-dimensional image plane, and the third coordinate represents depth information. The feature point data can be acquired using a depth camera or a multi-view camera system, with a sampling frequency of 30 frames per second to ensure the capture of subtle facial movements.

[0103] Based on the acquired facial feature point data, a feature point connectivity graph was constructed. In this graph, each node represents a facial feature point, and the connection weights between nodes are determined according to their anatomical correlation. Specifically, for feature points belonging to the same muscle control area, such as those around the corner of the eye, the connection weight is set to 0.8; for adjacent feature points belonging to different muscle control areas, such as those between the corner of the eye and the end of the eyebrow, the connection weight is set to 0.5; and for feature points with no direct anatomical correlation, such as those between the forehead and the corner of the mouth, the connection weight is set to 0.1. This constructed feature point connectivity graph reflects the anatomical correlations between facial feature points, providing a foundation for subsequent analysis.

[0104] For facial feature points in consecutive frames, the displacement vectors of each node in the graph connecting the feature points are calculated. Specifically, for facial feature points in frame t and frame (t-1), the displacement changes of each feature point in the x, y, and z directions are calculated to obtain the displacement vectors. For example, for the i-th feature point, its displacement vector is (Δx_i, Δy_i, Δz_i), where Δx_i = x_i(t) - x_i(t-1), and Δy_i and Δz_i are calculated in a similar way.

[0105] Based on the calculated displacement vectors, the deformation patterns of the feature point connection diagram are determined, and the deformation direction and magnitude are extracted from the deformation patterns. The deformation direction is determined by the direction of the displacement vector and can be represented by the unit vector of the displacement vector. The deformation magnitude is represented by the magnitude of the displacement vector, reflecting the distance the feature point moves. For example, for the displacement vector (Δx_i, Δy_i, Δz_i) of the i-th feature point, its deformation magnitude is (Δx_i... 2 + Δy_i 2 + Δz_i 2 ) 1 / 2 The deformation direction is the displacement vector divided by its modulus.

[0106] Facial feature points are clustered based on the deformation patterns of the feature point connection map. Feature points whose deformation direction angle is less than a preset angle and whose deformation amplitude difference is less than a preset amplitude difference are grouped into the same functional region. In this embodiment, the preset angle is set to 15 degrees, and the preset amplitude difference is set to 0.2 mm. For example, if the deformation direction angle between two feature points is 10 degrees and the deformation amplitude difference is 0.15 mm, then these two feature points are grouped into the same functional region. In this way, facial feature points are divided into functional regions such as the eyebrow region, upper eyelid region, lower eyelid region, nasal wing region, corner of mouth region, upper lip region, and lower lip region.

[0107] For each functional region, its deformation characteristics are established. These characteristics include the average deformation direction and average deformation amplitude of the feature points within the region. The average deformation direction is calculated by the vector sum of the displacement vectors of all feature points within the region, while the average deformation amplitude is the average of the deformation amplitudes of all feature points within the region. For example, the average deformation direction of the eyebrow region might be upward and slightly outward, with an average deformation amplitude of 0.5 mm.

[0108] A deformation transfer network is constructed between functional regions, including forward transitive connections and reverse inhibitory connections, to obtain the transfer paths between functional regions. A forward transitive connection indicates that the movement of one functional region promotes the movement of another, and the connection weight is set to a positive value, ranging from 0 to 1. A reverse inhibitory connection indicates that the movement of one functional region inhibits the movement of another, and the connection weight is set to a negative value, ranging from -1 to 0. For example, the forward transitive connection weight between the eyebrow region and the upper eyelid region is set to 0.7, indicating that when the eyebrow is raised, the upper eyelid has a high probability of raising simultaneously; while the reverse inhibitory connection weight between the upper lip region and the lower lip region is set to -0.3, indicating that when the upper lip is raised, the lower lip has a certain probability of drooping.

[0109] Based on the transmission path and deformation characteristics, the motion correlation between adjacent functional regions in the deformation transmission network is calculated. Motion correlation is calculated by analyzing the temporal changes in the deformation characteristics of functional regions across multiple consecutive frames of data. Specifically, for two adjacent functional regions, Region 1 and Region 2, their deformation characteristic sequences across 30 consecutive frames of data are calculated. If the deformation of Region 1 always precedes that of Region 2, and the magnitudes of their deformations show a regular proportional relationship, then the motion of Region 1 is considered to trigger the motion of Region 2, indicating a high correlation between the two. For example, analysis reveals that the upward movement of the eyebrow region precedes the upward movement of the upper eyelid region by approximately 0.1 seconds, and that for every 1 mm increase in the eyebrow region, the upper eyelid region increases by an average of 0.7 mm. Therefore, a strong motion correlation is determined between the eyebrow region and the upper eyelid region.

[0110] The temporal sequence and amplitude variations of functional areas are extracted to determine the coordinated motion characteristics of various facial functional areas. These coordinated motion characteristics describe the collaborative movement patterns between functional areas during facial expression changes, including the sequence of motion triggers, the proportion of motion amplitude, and the duration of motion. For example, the coordinated motion characteristics of a surprised expression might be: raising the eyebrows triggers raising the upper eyelids, simultaneously causing the forehead to rise, followed by lowering the mouth. Throughout this process, the amplitude of movement in the eyebrow area is 1.4 times that of the upper eyelid area and 2.1 times that of the forehead area.

[0111] The above methods can comprehensively and meticulously describe the coordinated movement characteristics of various functional areas of the face, providing important technical support for applications such as facial expression analysis, virtual character animation, and human-computer interaction.

[0112] In one optional implementation, a deformation transfer network is constructed between the functional regions, including forward transfer connections and reverse inhibition connections, to obtain the transfer paths between the functional regions, including:

[0113] Based on the boundary feature points of adjacent functional regions, the deformation gradient between adjacent functional regions is calculated. The deformation gradient includes the deformation direction gradient and the deformation amplitude gradient at the boundary feature points.

[0114] Collect the change data of the deformation gradient within a continuous time window, and calculate the stable duration and fluctuation range of the deformation gradient based on the change data;

[0115] When the deformation gradient is greater than or equal to a preset positive gradient threshold and the stability duration is greater than a preset time threshold, a positive transit connection is established between adjacent functional regions. The weight of the positive transit connection is positively correlated with the deformation gradient and the stability duration.

[0116] When the deformation gradient is less than a preset negative gradient threshold and the fluctuation range is less than a preset fluctuation threshold, an inverse suppression connection is established between adjacent functional regions. The weight of the inverse suppression connection is positively correlated with the absolute value of the deformation gradient.

[0117] A deformation propagation network is constructed by combining the weights of forward propagation connections and reverse inhibition connections. The deformation propagation intensity is determined based on the weights of each connection in the deformation propagation network, and propagation paths whose deformation propagation intensity exceeds a preset regional association threshold are extracted.

[0118] In constructing the deformation transfer network, the deformation gradient is calculated based on the boundary feature points of adjacent functional regions. This is achieved by scanning the boundaries of adjacent functional regions and extracting a set of feature points. For any two adjacent functional regions, boundary feature points are uniformly sampled along their common boundary, typically one feature point every 5 millimeters. The displacement changes of these feature points under stress are collected, and the deformation direction gradient and deformation magnitude gradient of each feature point are calculated. The deformation direction gradient represents the rate of change of the angle between the displacement direction of the feature point and the normal direction of the boundary, while the deformation magnitude gradient represents the rate of change of the displacement magnitude per unit distance. For example, for two feature points p1 and p2 10 millimeters apart on the boundary, if p1 has a displacement of 2 millimeters with an angle of 30 degrees to the normal, and p2 has a displacement of 3 millimeters with an angle of 45 degrees to the normal, then the deformation magnitude gradient of this region is 0.1 millimeters / millimeter, and the direction gradient is 1.5 degrees / millimeter.

[0119] Deformation gradient change data are collected within a continuous time window, typically set to 10 seconds. During this period, the deformation gradient change at each boundary feature point is recorded at a frequency of 100 Hz. The stable duration refers to the cumulative time the deformation gradient remains within a certain numerical range, while the fluctuation range refers to the difference between the maximum and minimum deformation gradient values ​​within the observation window. For example, if the deformation amplitude gradient at a boundary feature point remains within the range of 0.08-0.12 mm / mm for 7 seconds out of 10 seconds, its stable duration is 7 seconds; if the maximum deformation amplitude gradient is 0.15 mm / mm and the minimum is 0.05 mm / mm throughout the entire observation window, then the fluctuation range is 0.1 mm / mm.

[0120] When the deformation gradient is greater than or equal to a preset positive gradient threshold and the stabilization duration is greater than a preset time threshold, a positive transitive connection is established between adjacent functional regions. In practical applications, the positive gradient threshold is typically set to 0.05 mm / mm, and the time threshold is set to 5 seconds. The weight calculation for the positive transitive connection considers the combined effect of the deformation gradient and the stabilization duration; the weight value is equal to the product of the deformation gradient and the stabilization duration multiplied by an adjustment coefficient. The adjustment coefficient is typically set to 0.2 to standardize the weight value to the range of 0-1. For example, for a deformation gradient of 0.1 mm / mm and a stabilization duration of 6 seconds, the weight of the positive transitive connection is 0.1 × 6 × 0.2 = 0.12.

[0121] When the deformation gradient is less than a preset negative gradient threshold and the fluctuation range is less than a preset fluctuation threshold, an inverse suppression connection is established between adjacent functional regions. The negative gradient threshold is typically set to -0.03 mm / mm, and the fluctuation threshold is set to 0.02 mm / mm. The weight of the inverse suppression connection is positively correlated with the absolute value of the deformation gradient, and is calculated by multiplying the absolute value of the deformation gradient by an adjustment coefficient, which is typically set to 3.0. For example, if the deformation amplitude gradient of a boundary feature point is -0.04 mm / mm and the fluctuation range is 0.01 mm / mm, then the weight of the inverse suppression connection is |-0.04|×3.0=0.12.

[0122] The deformation transfer network is constructed by combining the weights of forward transitive connections and reverse inhibition connections. When both forward transitive and reverse inhibition connections exist between pairs of functional regions, the weight difference is used as the final connection weight. After the network is constructed, the deformation transfer strength is determined based on the weights of each connection in the deformation transfer network. The transfer strength is calculated as the ratio of the connection weight to the path length, where the path length refers to the number of connections traversed from the starting functional region to the target functional region. The system extracts transfer paths whose deformation transfer strength exceeds a preset regional association threshold, which is typically set to 0.08.

[0123] In a specific example, a structure containing five functional regions (labeled R1-R5) was analyzed. Deformation analysis revealed that the deformation gradient from R1 to R2 was 0.09 mm / mm, with a stable duration of 7.5 seconds, establishing a forward transitive connection with a weight of 0.09 × 7.5 × 0.2 = 0.135; the deformation gradient from R2 to R3 was -0.05 mm / mm, with a fluctuation range of 0.015 mm / mm, establishing an inverse suppression connection with a weight of 0.15; the deformation gradient from R3 to R4 was 0.12 mm / mm, with a stable duration of 6 seconds, establishing a forward transitive connection with a weight of 0.144; the deformation gradient from R4 to R5 was 0.07 mm / mm, with a stable duration of 4 seconds. Since the stable duration was less than the threshold of 5 seconds, no connection was established. Finally, the transmission path R1→R2→R3→R4 was extracted, and the transmission strength was (0.135+0.15+0.144) / 3=0.143, which exceeded the regional association threshold of 0.08. Therefore, it was identified as a valid deformation transmission path.

[0124] The deformation transfer network constructed in this way can accurately reflect the deformation transfer relationship between functional areas, providing a data foundation for subsequent structural optimization and coordinated control of functional areas.

[0125] like Figure 2 The diagram shown illustrates the construction flowchart of the adaptive deformation transfer network.

[0126] In one optional implementation, a baseline feature for normal driving behavior is constructed based on the linkage motion features. Multi-dimensional difference values ​​of the current features are calculated using feature mapping rules to generate an anomaly discrimination index for the driver's state, including:

[0127] The linkage motion features under standard driving scenarios are collected, and feature dimensionality reduction is performed using principal component analysis to obtain feature component groups. Based on the feature component groups, motion amplitude distribution parameters, relative motion relationship distribution parameters, and time series change law distribution parameters are calculated, and feature fusion is performed to obtain the baseline features of normal driving behavior.

[0128] A contrastive loss function is determined for the baseline features of normal driving behavior, including a feature vector distance calculation term, a feature distribution divergence calculation term, and a time series consistency constraint calculation term. Based on the contrastive loss function, a feature encoder is iteratively trained to obtain feature mapping rules.

[0129] Collect the linkage motion features of the current driving scenario, and convert them into the current feature representation through feature mapping rules. Based on the current feature representation and the normal driving behavior benchmark features, determine the distance metric value, distribution deviation value, and time series anomaly value.

[0130] The abnormal scores for single features, feature combinations, and time-series patterns are calculated separately, and then weighted according to preset abnormal weights to generate an abnormality discrimination index for the driver's state.

[0131] In one specific implementation, when collecting linkage motion features under standard driving scenarios, roadside-mounted visual cameras and infrared cameras are used for simultaneous acquisition, with a sampling frequency of 25 frames per second and a continuous acquisition duration of 120 minutes. The acquisition process covers various standard driving behaviors, including normal driving, lane changing, steering, acceleration, and deceleration scenarios, collecting a total of 500 sets of valid data samples. The acquired linkage motion features include the spatial coordinates of 68 facial key points and their trajectory data changing over time, forming a high-dimensional feature space.

[0132] Principal component analysis (PCA) is used to reduce the dimensionality of the collected motion features, retaining principal components with a cumulative contribution rate of 95%, typically reducing the dimensions to 12-15. During dimensionality reduction, the original data is standardized to eliminate the influence of dimensions. The dimensionality-reduced feature component groups retain the main information of the original features while reducing redundant dimensions, thus improving the efficiency of subsequent calculations. For example, blinking in the eye region may be reduced to 2-3 principal components, representing the frequency, amplitude, and symmetry characteristics of blinking.

[0133] Motion amplitude distribution parameters were calculated based on the dimensionality-reduced feature component sets. These parameters describe the statistical characteristics of motion amplitude in various facial functional areas, including mean, standard deviation, skewness, and kurtosis. Under normal driving conditions, the mean blink amplitude in the eye area was 3.2 mm, with a standard deviation of 0.5 mm; the mean opening and closing amplitude in the mouth area was 5.7 mm, with a standard deviation of 1.2 mm; and the mean head posture change amplitude was 3.8 degrees, with a standard deviation of 0.9 degrees.

[0134] The distribution parameters of relative motion relationships are calculated to describe the coordinated relationship of movements between different functional areas of the face. The linkage characteristics between areas are quantified by calculating the correlation, phase difference, and amplitude ratio between the motion trajectories of different functional areas. For example, under normal driving conditions, the correlation coefficient between blinking movements of the left and right eyes is 0.92, indicating high synchronicity; the correlation coefficient between eye and mouth movements is 0.35, indicating a certain degree of independence; and the correlation coefficient between head turning and eye movements is 0.78, indicating relatively strong coordination.

[0135] The distribution parameters of temporal variation patterns are calculated to describe the dynamic characteristics of facial movements over time. Periodicity, trend, and abrupt change features are extracted through time-frequency analysis of the motion sequences. Under normal driving conditions, the blinking frequency is distributed in the range of 15-20 times per minute; the frequency of minor head adjustments is distributed in the range of 8-12 times per minute; and the duration of facial expression changes is distributed in the range of 0.5-2 seconds.

[0136] The distribution parameters of motion amplitude, relative motion relationship, and temporal variation are fused to construct a baseline feature for normal driving behavior. The fusion process employs a multi-layer feature fusion strategy, with lower-layer features capturing local detail changes and higher-layer features expressing global semantic information. The fused baseline feature contains 120 feature dimensions, covering all aspects of normal driving behavior and forming a complete representation of driving behavior features.

[0137] A contrastive loss function is determined based on the baseline features of normal driving behavior, comprising three core calculation terms. The feature vector distance calculation term uses Euclidean distance to measure the difference between two feature vectors; when the distance value exceeds a preset threshold of 5.0, it is considered an abnormal feature. The feature distribution divergence calculation term uses divergence to measure the difference between two feature distributions, calculating the relative entropy between the distribution parameters; when the divergence value exceeds a preset threshold of 2.5, it is considered an abnormal distribution. The temporal consistency constraint calculation term evaluates the temporal continuity and smoothness of the feature sequence by calculating the rate of change of features at adjacent time points; when the rate of change exceeds a preset threshold of 0.4, it is considered an abnormal change.

[0138] The feature encoder is trained iteratively using a contrastive loss function, with parameters optimized using gradient descent. The initial learning rate is set to 0.01 and gradually reduced to 0.001 during training. The training dataset contains 400 normal driving samples and 100 abnormal driving samples, with the abnormal samples covering various scenarios such as fatigued driving, distracted driving, and emotionally unstable driving. The training process uses a batch size of 32 and 200 training epochs. Training is terminated early when the validation set loss does not show a significant decrease for 10 consecutive epochs. After training, the feature encoder can convert the input linkage motion features into standardized feature representations, forming feature mapping rules.

[0139] The system collects the linkage motion features of the current driving scenario. Real-time image data of the driver's facial region is acquired using roadside vision cameras and infrared cameras, and linkage motion features are extracted. A trained feature encoder is used to apply feature mapping rules to convert the currently collected linkage motion features into a current feature representation. This representation is then compared with the baseline features of normal driving behavior, and three difference values ​​are calculated: Distance metric: The Euclidean distance between the current feature vector and the baseline feature vector; normal range 0-5, exceeding 5 is considered abnormal. Distribution deviation metric: The divergence between the current feature distribution and the baseline feature distribution; normal range 0-2.5, exceeding 2.5 is considered abnormal. Temporal anomaly metric: The difference between the temporal continuity of the current feature sequence and the baseline features; normal range 0-0.4, exceeding 0.4 is considered abnormal.

[0140] Calculate the anomaly scores for single features, feature combinations, and time-series patterns separately. Single feature anomaly scores reflect the degree of abnormality in an individual feature dimension; for example, excessively low blinking frequency may indicate fatigue, and head posture deviation may indicate distraction. Feature combination anomaly scores reflect the degree of abnormality in the relationship between multiple features; for example, incoordination between eye movements and head posture may indicate inattention. Time-series pattern anomaly scores reflect the degree of abnormality in feature changes over time; for example, a gradually decreasing blinking frequency may indicate increasing fatigue.

[0141] An anomaly detection index for driver status is generated through weighted calculation based on preset anomaly weights. The weight for a single feature anomaly score is set to 0.3, the weight for a feature combination anomaly score is set to 0.4, and the weight for a time-series pattern anomaly score is set to 0.3. The anomaly detection index ranges from 0 to 100. Under normal driving conditions, the index value is typically below 50; for slightly abnormal conditions, the index value is between 50 and 75; and for severely abnormal conditions, the index value is above 75. In practical applications, the roadside monitoring system continuously calculates the anomaly detection index to achieve real-time monitoring of driver status, meeting traffic safety supervision requirements.

[0142] The roadside driver status recognition system integrating visual and infrared data according to embodiments of the present invention includes:

[0143] The first unit is used to acquire images of the driver's cab area of ​​the target vehicle through roadside vision cameras and infrared cameras, and to preprocess them to obtain visual image data and infrared image data.

[0144] The second unit is used to build feature pyramids from visual image data and infrared image data, construct feature parameter association tables through cross-modal correlation coefficients and perform spatiotemporal registration to obtain a registration feature set; based on the feature parameter association tables, a complementary weight mapping table is constructed, and the registration feature set is adaptively fused to obtain a complementary enhanced feature map.

[0145] The third unit is used to extract the driver's facial region from the complementary enhanced feature map, calculate the corresponding three-dimensional geometric parameters and depth parameters, and combine the feature parameters to generate a multi-dimensional feature sequence.

[0146] The fourth unit is used to build a feature point connection map based on facial feature points in a multidimensional feature sequence, divide functional regions by deformation patterns, and extract the linkage motion features of each functional region of the face.

[0147] The fifth unit is used to construct benchmark features for normal driving behavior based on linkage motion features, calculate the multi-dimensional difference values ​​of the current features through feature mapping rules, and generate an anomaly discrimination index for the driver's state.

[0148] The sixth unit is used to determine the driver's status based on the anomaly discrimination index, and to generate a warning message when the anomaly discrimination index exceeds a preset anomaly threshold.

[0149] A third aspect of the present invention provides an electronic device, comprising:

[0150] processor;

[0151] Memory used to store processor-executable instructions;

[0152] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0153] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0154] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.

[0155] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for fusing visual and infrared data for roadside driver state recognition, characterized in that, The method comprises the following steps: Collecting images of the cab area of a target vehicle through a roadside visual camera and an infrared camera, and preprocessing the images to obtain visual image data and infrared image data; Building a feature pyramid for the visual image data and the infrared image data, constructing a feature parameter correlation table through cross-modal correlation coefficients, and performing spatio-temporal registration to obtain a registered feature set; Based on the feature parameter correlation table, a complementarity weight mapping table is constructed, and the registered feature set is adaptively fused to obtain a complementary enhanced feature map; From the complementary enhanced feature map, the driver's face region is extracted, the corresponding three-dimensional geometric parameters and depth parameters are calculated, and a multi-dimensional feature sequence is generated in combination with the feature parameters; Based on the facial feature points in the multi-dimensional feature sequence, a feature point connection graph is established, and the functional regions of the face are divided by a deformation mode, and the linkage motion features of each functional region of the face are extracted, including: Based on the spatial positions of the facial feature points, a feature point connection graph is established, each node in the feature point connection graph corresponds to a facial feature point, and the connection weights between the nodes are determined according to the degree of anatomical association; The displacement vectors of each node in the feature point connection graph are calculated, the deformation mode of the feature point connection graph is determined based on the displacement vectors, and the deformation direction and deformation amplitude in the deformation mode are extracted; According to the deformation mode of the feature point connection graph, the facial feature points are clustered, the facial feature points with an included angle of the deformation direction less than a preset included angle and a difference in the deformation amplitude less than a preset amplitude difference are divided into the same functional region, and the deformation features of the functional regions are established; A deformation transmission network between the functional regions is constructed, including forward transmission connections and reverse inhibition connections, to obtain transmission paths between the functional regions; Based on the transmission paths and the deformation features, the motion correlation between adjacent functional regions in the deformation transmission network is calculated, the motion sequence and amplitude change relationship of the functional regions in time sequence are extracted, and the linkage motion features of each functional region of the face are determined; Based on the linkage motion features, a normal driving behavior benchmark feature is constructed, a multi-dimensional difference value of the current feature is calculated through a feature mapping rule, and an abnormality discrimination index of the driver state is generated; According to the abnormality discrimination index, the state of the driver is judged, and when the abnormality discrimination index exceeds a preset abnormal threshold, a warning information is generated.

2. The method of claim 1, wherein, Building a feature pyramid for the visual image data and the infrared image data, constructing a feature parameter correlation table through cross-modal correlation coefficients, and performing spatio-temporal registration to obtain a registered feature set includes: Building a feature pyramid for the visual image data and the infrared image data, extracting texture features, edge features, gradient features and statistical features at different scale layers of the feature pyramid, and combining the features at different scale layers to generate a visual feature set and an infrared feature set; Standardizing the visual feature set and the infrared feature set, calculating the local mutual information and the global mutual information between the visual feature set and the infrared feature set, and generating cross-modal correlation coefficients; Based on the cross-modal correlation coefficients, the feature parameter similarity is calculated, the visual feature set and the infrared feature set are grouped according to the feature parameter similarity, the feature groups are determined, and the feature parameter correlation table is constructed; The visual feature set and the infrared feature set are registered in the spatio-temporal domain to obtain a registered feature set.

3. The method of claim 2, wherein, The complementary weight mapping table is constructed based on the feature parameter correlation table, and the complementary enhanced feature map is obtained by adaptively fusing the registered feature set, including: The information entropy and the dispersion degree of each feature group in the feature parameter correlation table are calculated, and the complementary weight mapping table is constructed based on the information entropy and the dispersion degree, and the complementary weight mapping table includes inter-group weight and intra-group weight; The inter-group weight and the intra-group weight of the registered feature set are calculated by using the complementary weight mapping table, and a weighted feature is generated; According to the distribution characteristics of the weighted feature, an adaptive fusion function is constructed, and the adaptive fusion function includes a linear fusion component and a nonlinear fusion component; The feature fusion is performed on the weighted feature by using the adaptive fusion function, and an initial feature map is generated, and the complementary enhanced feature map is obtained by performing spatial domain filtering and frequency domain enhancement on the initial feature map.

4. The method of claim 1, wherein, A deformation transmission network between the functional areas is constructed, including forward transmission connection and reverse inhibition connection, and a transmission path between the functional areas is obtained, including: Based on the boundary feature points of adjacent functional areas, the deformation gradient between adjacent functional areas is calculated, and the deformation gradient includes the deformation direction gradient and the deformation amplitude gradient at the boundary feature points; The change data of the deformation gradient in a continuous time window is collected, and the stable duration and the fluctuation range of the deformation gradient are calculated based on the change data; When the deformation gradient is greater than or equal to a preset positive gradient threshold and the stable duration is greater than a preset time threshold, a forward transmission connection is established between adjacent functional areas, and the weight of the forward transmission connection is positively correlated with the deformation gradient and the stable duration; When the deformation gradient is less than a preset negative gradient threshold and the fluctuation range is less than a preset fluctuation threshold, a reverse inhibition connection is established between adjacent functional areas, and the weight of the reverse inhibition connection is positively correlated with the absolute value of the deformation gradient; The weights of the forward transmission connection and the reverse inhibition connection are combined to construct a deformation transmission network, the deformation transmission strength is determined according to the weight of each connection in the deformation transmission network, and the transmission path whose deformation transmission strength exceeds a preset area correlation threshold is extracted.

5. The method of claim 1, wherein, The normal driving behavior benchmark feature is constructed based on the linkage motion feature, the multi-dimensional difference value of the current feature is calculated by using the feature mapping rule, and the abnormality discrimination index of the driver state is generated, including: The linkage motion feature in the standard driving scene is collected, the feature dimension is reduced by using the principal component analysis method to obtain a feature component group, the motion amplitude distribution parameter, the relative motion relationship distribution parameter and the time sequence change rule distribution parameter are calculated based on the feature component group, and the feature fusion is performed to obtain the normal driving behavior benchmark feature; The contrast loss function is determined for the normal driving behavior benchmark feature, including the feature vector distance calculation item, the feature distribution divergence calculation item and the time sequence consistency constraint calculation item, the feature encoder is iteratively trained based on the contrast loss function, and the feature mapping rule is obtained; The linkage motion characteristics of the current driving scene are collected, and are mapped and converted into current feature representations through feature mapping rules. The current feature representations are compared with the normal driving behavior benchmark features to determine distance metric values, distribution deviation values, and time sequence abnormality values. Single feature abnormality scores, feature combination abnormality scores, and time sequence pattern abnormality scores are calculated respectively, and an abnormality discrimination index of the driver state is generated through weighted calculation according to preset abnormality weights.

6. A roadside driver state recognition system fusing vision and infrared data for implementing the method of any of the preceding claims 1-5, characterized in that, The method comprises the following steps: A first unit is configured to collect images of a cab area of a target vehicle through a roadside vision camera and an infrared camera, and to obtain vision image data and infrared image data through preprocessing; A second unit is configured to establish a feature pyramid for the vision image data and the infrared image data, to construct a feature parameter correlation table through a cross-modal correlation coefficient, and to perform spatio-temporal domain registration to obtain a registered feature set; A third unit is configured to extract a driver face region from the complementary enhanced feature map, to calculate corresponding three-dimensional geometric parameters and depth parameters, and to combine the feature parameters to generate a multi-dimensional feature sequence; A fourth unit is configured to establish a feature point connection graph based on facial feature points in the multi-dimensional feature sequence, to divide functional regions through a morphing mode, and to extract linkage motion characteristics of each functional region of the face; A fifth unit is configured to construct a normal driving behavior benchmark feature based on the linkage motion characteristics, to calculate multi-dimensional difference values of current features through a feature mapping rule, and to generate an abnormality discrimination index of the driver state; A sixth unit is configured to determine the driver state according to the abnormality discrimination index, and to generate a warning information when the abnormality discrimination index exceeds a preset abnormality threshold. The method comprises the following steps:

7. An electronic device, comprising: A processor; A memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method of any one of claims 1 to 5. The computer program instructions are executed by the processor to implement the method of any one of claims 1 to 5.

8. A computer-readable storage medium having stored thereon computer program instructions, wherein, ​

Citation Information

Patent Citations

  • Non-invasive driver driving fatigue state identification method and system

    CN118051810A

  • Fatigue driving monitoring and early warning method based on multi-dimensional feature fusion

    CN119495164A