Attention correlation estimation method and system fusing driver state and scene understanding
By adopting dynamic adaptive filtering and global context information sharing mechanisms in complex dynamic environments, the end-to-end cross-frame mapping of driver attention and scenarios is realized, solving the problems of inaccurate correlation and dependence on high-cost sensors in the prior art, and improving the robustness and ease of deployment of the system.
Patent Information
- Application Number
- CN202411953603.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art is difficult to accurately correlate driver attention and scenarios in complex dynamic environments, and the two-stage mapping method relies on high-cost sensors and complex data processing, resulting in device and data dependence, reducing the reliability of the system and the possibility of wide deployment.
Dynamic adaptive filtering and global context information sharing mechanism are adopted to realize cross-domain feature fusion through W-Net architecture, deeply model the complex relationship between cross-view information in and out of the car, and realize end-to-end cross-frame mapping of attention.
Improves the accuracy of driver's attention and scene association, overcomes device dependence and data dependence, and enhances the robustness and easy-to-deployment features of the system.
Smart Images

Figure CN120047925A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field related to assisted driving systems, and in particular, relates to an attention correlation estimation method and system that integrates driver state and scene understanding. Background Art
[0002] In recent years, when realizing the correlation estimation of driver attention and scene, some works focus on mapping driver attention to different cockpit areas, or directly estimating the driver's fixation point on the road by detecting the driver's state. These methods ignore the complementary information between the two views and lack interpretability. Other studies use depth information and camera calibration to map the estimated line of sight to the scene, and use this two-stage method to achieve the correlation between attention and scene. However, these two-stage methods rely on high-cost sensors and complex data processing processes, resulting in equipment and data dependence, which hinders wide deployment. In addition, when mapping to a distant driving scene, the fixation estimation error will be amplified, and the mapping result will be further affected by the loss of depth data, thereby reducing the overall reliability.
[0003] In order to accurately correlate driver attention and scene, by integrating cross-domain information of the driver's face and traffic scene at the feature level, and using the implicit connection between these perspectives to end-to-end estimate the attention distribution of the driver in the scene, the problems existing in the above two-stage method can be avoided. However, when dealing with the task of correlating driver attention and scene in a complex dynamic environment, the cross-view information from the driver and the driving scene is very complex, making a simple fusion module insufficient to fully capture the basic features and relationships; the current methods prioritize the fusion itself, but cannot ensure high-quality independent feature representations before fusion, which will damage the effectiveness of the fusion; the fusion process will introduce semantic inconsistencies, reducing the robustness and accuracy of model prediction. Summary of the Invention
[0004] In order to solve the above problems, the present invention proposes an attention correlation estimation method and system that integrates driver state and scene understanding. Through a dynamic adaptive filtering and global context information sharing mechanism, the challenges brought by complex dynamic environments and non-fixed driver postures are overcome. Based on the proposed W-Net architecture, cross-domain feature fusion is realized, and the complex relationship between cross-view information inside and outside the vehicle is deeply modeled, achieving end-to-end cross-frame mapping of attention and improving the accuracy of attention estimation related to the scene.
[0005] In a first aspect, the present invention provides an attention correlation estimation method that integrates driver state and scene understanding;
[0006] The overall method is designed as an end-to-end method based on dual-domain fusion. A novel "independent encoding-independent partial decoding-fusion decoding" W-shaped network architecture is designed to perform high-quality encoding and fusion of cross-domain images respectively, and directly generate the attention heat map of the driver.
[0007] In the independent encoding stage, the in-vehicle driver's facial video information and the synchronized out-of-vehicle scene video information are obtained; the synchronized adjacent frame images of the in-vehicle driver's facial video information and the out-of-vehicle scene video information are extracted; double-branch feature extraction is performed on the two adjacent frame images respectively to obtain corresponding multi-level features; the multi-level feature representation of the scene branch is processed by multi-layer convolution to reduce the number of channels, so that the multi-level features of the scene branch have the same number of channels as the corresponding levels of the driver face branch.
[0008] Furthermore, joint adaptive dynamic filtering in the frequency domain and time domain is performed on the multi-level features between frames of the two branches respectively;
[0009] The adjacent frame features are divided into several non-overlapping sub-regions. After independently performing correlation calculations within each sub-region, weighted averaging is performed with the pixel grid to determine the optimal local motion vector of each pixel. The dynamic feature map of the sub-region is obtained by calculating the difference, and the overall dynamic feature estimation result is obtained by integrating the dynamic feature predictions of all local blocks;
[0010] The features are transformed to the frequency domain through two-dimensional fast Fourier transform, and frequency domain filtering is performed through the convolutional layer to adjust the intensity of different frequency components of the image to filter out the self-motion recorded indiscriminately;
[0011] The features are transformed back to the time domain, and a spatial attention mask is obtained through spatial attention for time domain filtering of the features to highlight important spatial elements.
[0012] Furthermore, global context information sharing is performed on the multi-level features of the facial feature branch to adaptively process key features of different scales in non-fixed postures; multi-scale information is extracted by using dilated convolutional branches with different dilation rates for intra-layer features, and then the feature maps of different branches are fused using a hierarchical feature fusion structure; among them, channel mixing attention is used to fuse adjacent branches;
[0013] Features at different levels are used as reference anchor feature maps in turn; the spatial dimensions of adjacent levels are adjusted to be the same as the reference anchor feature map through adaptive average pooling or bilinear upsampling operations, and the convolutional layer is applied to transform the features to the feature space of the same dimension, and the transformed features are fused with the reference anchor feature map through element-by-element addition to achieve the interaction of multi-level information. Furthermore, in the independent part decoding stage, the multi-level features of the two branches are decoded level by level in the domain respectively through the channel-space hybrid attention mechanism to align the semantics of the features in the domain. The channel-space hybrid attention mechanism balances the cross-channel and spatial contextual relationships through the designed channel integration unit CIU and spatial integration unit SIU, balances the contribution of features at different levels in the decoding process, and achieves high-quality decoding.
[0014] Furthermore, in the fusion decoding stage, the same-level features of the two branches are fused and decoded step by step through the cross-attention mechanism and the channel-space hybrid attention mechanism, and the driver's attention heat map associated with the scene is obtained through decoding prediction. Furthermore, multiple frames of images in a time segment centered on the current frame are selected, and the fixation points in each frame are mapped to the coordinates of the current frame through homography transformation; the Gaussian function is used to simulate the probability distribution of the fixation point, and the spatiotemporal weight mechanism is introduced to assign weights according to the spatiotemporal proximity between the fixation point and the current frame. By integrating the weighted Gaussian distribution, the true value of the driver's visual attention heat map is generated.
[0015] In a second aspect, the present invention further provides an attention association estimation system integrating driver status and scene understanding, comprising:
[0016] The data acquisition module is configured to: obtain the facial video information of the driver in the vehicle and the video information of the scene outside the vehicle;
[0017] The independent encoding module is configured to: extract synchronous adjacent frame images of the driver's facial video information inside the vehicle and the scene video information outside the vehicle; perform dual-branch feature extraction on the two adjacent frame images to obtain corresponding multi-level features; perform frequency domain and time domain joint adaptive dynamic filtering on the inter-frame multi-level features of the two branches; and perform global context information sharing on the multi-level features of the facial feature branch to adaptively process key features of different scales under non-fixed postures;
[0018] The independent part decoding module is configured to: decode the multi-level features of the two branches level by level in the domain through the hybrid attention mechanism, and align the semantics of the features in the domain;
[0019] The fusion decoding module is configured to: fuse and decode the same-level features of the two branches step by step through the cross-attention mechanism and the hybrid attention mechanism, and obtain the driver's attention heat map associated with the scene through decoding prediction.
[0020] In a third aspect, the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and capable of running on the processor. When the processor executes the program, the steps of the driver attention prediction method based on state tracking and scenario fusion described in the first aspect are implemented.
[0021] In a fourth aspect, the present invention further provides a computer program product. The computer program product includes a computer program. When the computer program is executed by a processor, the steps of the driver attention prediction method based on state tracking and scenario fusion described in the first aspect are implemented.
[0022] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0023] 1. In the present invention, first, synchronous adjacent frame images of in-vehicle driver facial video information and out-of-vehicle scene video information are extracted; dual-branch feature extraction is respectively performed on the two types of adjacent frame images to obtain corresponding multi-level features; frequency-domain and time-domain joint adaptive dynamic filtering is respectively performed on the inter-frame multi-level features of the two branches; and global context information sharing is performed on the multi-level features of the facial feature branch to adaptively process key features of different scales in non-fixed postures. Then, through a hybrid attention mechanism, the multi-level features of the two branches are respectively decoded step by step within the domain to align the feature semantics within the domain. Finally, through a cross-attention mechanism and a hybrid attention mechanism, the same-level features of the two branches are gradually fused and decoded, and a driver attention heat map associated with the scenario is obtained through decoding prediction. Through the dynamic adaptive filtering and global context information sharing mechanism, the challenges brought by complex dynamic environments and non-fixed driver postures are overcome. Based on cross-domain feature fusion, the complex relationship between in-vehicle and out-of-vehicle cross-perspective information is deeply modeled, realizing end-to-end cross-frame mapping of attention and improving the accuracy of attention estimation associated with the scenario.
[0024] 2. In response to the challenge of attention cross-frame mapping, the present invention uses an end-to-end method of cross-domain feature fusion to model the implicit relationship between the two perspectives at the feature level. Compared with existing two-stage methods, it overcomes device dependence and data dependence, has stronger interpretability; does not rely on complex processes such as depth information and camera parameters, avoiding the introduction of additional errors and having good robustness; does not require complex processes such as camera calibration and depth estimation, has low system complexity and is easy to deploy widely; does not rely on expensive additional devices such as depth sensors, but only requires two cameras facing the driver and the driving scenario to capture the states under the two perspectives, which is economically feasible.
[0025] 3. In view of the semantic inconsistency problem existing in the existing cross-domain fusion methods, the present invention proposes a W-Net network architecture for cross-domain fusion, which adopts an overall process of "independent encoding - independent partial decoding - fusion decoding", focuses on the core two-step decoding strategy, ensures semantic consistency within and between domains, optimizes the feature quality throughout the link, and demonstrates efficient cross-domain information fusion capabilities for dual-source heterogeneous data.
[0026] 4. In view of the challenge of the interweaving of self-motion and independent motion in a highly complex dynamic traffic environment, the present invention designs a dynamic adaptive filtering module. By extracting the inter-frame dynamic features and performing frequency-domain - time-domain joint filtering, the self-motion is filtered out and the independent motion is separated. Compared with the previous method of indiscriminately extracting all motions based on the optical flow branch, the key motion features are effectively enhanced, and the feature quality and information utilization efficiency are improved.
[0027] 5. In view of the problem of state capture of multi-scale feature extraction and aggregation for drivers in non-fixed postures, the present invention proposes a global context information sharing module, which performs intra-layer multi-scale feature aggregation and inter-layer multi-scale feature semantic alignment on facial features, takes into account both local details and global structural features, optimizes the consistency and comprehensiveness of feature representation, and adaptively processes non-fixed postures and multi-views. Compared with the existing method that obtains a combined feature set through cumbersome preprocessing such as key point detection and head pose detection, the problems of information redundancy and inefficiency are overcome. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The accompanying drawings forming a part of this embodiment are used to provide a further understanding of this embodiment. The schematic embodiments and descriptions thereof are used to explain this embodiment and do not constitute an improper limitation of this embodiment.
[0029] Figure 1 It is a schematic flow chart of the end-to-end association method between driver attention and driving scenarios in Embodiment 1 of the present invention;
[0030] Figure 2 It is a schematic flow chart of the dynamic adaptive filtering in Embodiment 1 of the present invention;
[0031] Figure 3 It is a schematic flow chart of the global context information sharing in Embodiment 1 of the present invention;
[0032] Figure 4 It is a schematic flow chart of the W-Net architecture in Embodiment 1 of the present invention.
[0033] Figure 5 It is a schematic flow chart of the channel spatial hybrid attention in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0035] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0036] Embodiment 1:
[0037] In recent years, autonomous driving technology has made significant progress, but the driver is still crucial in the human-machine collaboration at a lower automation level, especially in complex scenarios where the system's capabilities are insufficient and manual driver intervention is required. The driver's attention directly affects the effectiveness and safety of the collaboration, so attention monitoring is crucial for response capabilities. With the progress of autonomous driving technology, even in the case of fully automated driving, studying the attention patterns of human drivers will continue to provide necessary support for enhancing the human-like perception, planning, and decision-making capabilities of autonomous driving systems.
[0038] Therefore, driver monitoring systems have attracted great interest because they can real-time sense the driver's attention state and adjust behaviors accordingly. Among various research topics in this field, the association between driver attention and driving scenarios is particularly valuable, with a focus on estimating the area or target that the driver is currently focusing on in the current environment. Such research has great potential applications in enhancing road safety, improving vehicle-human interaction, and supporting the decision-making process of autonomous driving.
[0039] This study addresses the complex problem of understanding driving scenarios, capturing driver attention, and mapping this attention to the scenario. The main challenges in this field come from different perspectives of driver attention and road conditions: the internal (inside the vehicle) view and the external (road) view. There is no clear overlapping information between them, making the integration of cross-domain information a difficult task. The inherent complex causal relationships between these views further exacerbate this challenge. Specifically, the driver's attention is largely affected by traffic conditions. However, due to the driver's subjective initiative, their attention may shift multiple times even before obvious changes occur in the driving environment. This makes it extremely difficult to clearly define the relationship between driver attention transfer and driving environment changes. In addition, during the driving task, both the driver's behavior and road conditions are continuously and unpredictably changing.
[0040] Specifically, drivers often make a large number of eye movements and head rotations to collect information, which makes it complex to precisely track their attention due to these dynamic postures. In addition, sudden events in the driving environment require immediate response, such as the appearance of pedestrians or lane changes by other vehicles, which poses a challenge to the adaptability of the monitoring model.
[0041] Many studies have focused on analyzing a single perspective. For example, attention estimation from the driver's perspective captures the driver's rapidly changing state by analyzing a combination of facial expressions, head movements, and other features. On the other hand, some studies predict the areas or targets that drivers should focus on by analyzing key traffic elements from the road perspective. However, these studies are limited to single-perspective analysis and fail to establish a direct association between the driver's actual attention and the scene.
[0042] In recent years, some studies have attempted to simulate this association at the result level. Some work focuses on mapping the driver's attention to different cockpit areas or directly estimating the driver's gaze points on the road by detecting the driver's state. These methods ignore the complementary information between the two perspectives and lack interpretability. Some studies use depth information and camera calibration to map the estimated line of sight to the scene, achieving the association between attention and the scene using this two-stage method.
[0043] However, these two-stage methods rely on high-cost sensors and complex data processing pipelines, resulting in device and data dependencies that hinder wide deployment. In addition, their accuracy and robustness are limited: when mapped to distant driving scenes, the gaze estimation error is amplified, and the mapping results are further affected by the loss of depth data, reducing the overall reliability.
[0044] A potential way to accurately associate the driver's attention with the scene is to end-to-end estimate the driver's attention distribution in the scene by integrating cross-domain information of the driver's face and the traffic scene at the feature level and leveraging the implicit connection between these perspectives, which can avoid the problems existing in the above two-stage method. To aggregate information from cross-domain data sources, existing cross-domain integration methods usually adopt a Y-shaped network structure, with the focus on combining information through a feature fusion module. However, several key challenges arise when applying these methods to the task of associating the driver's attention with the scene in a complex dynamic environment: the cross-view information from the driver and the driving scene is very complex, making a simple fusion module insufficient to fully capture the basic features and relationships. These methods prioritize the fusion itself but cannot ensure high-quality individual feature representations before fusion, which may undermine the effectiveness of the fusion. The fusion module process may introduce semantic inconsistencies, reducing the robustness and accuracy of the model prediction. These problems highlight the inherent challenges of cross-domain information integration.
[0045] In summary, in the prior art, little attention has been paid to the accurate association between driver attention and scenarios. Existing two-stage mapping methods have many device limitations and data dependencies, with limited accuracy and poor robustness. Therefore, the present invention provides an end-to-end attention association estimation method based on driver state tracking and dynamic scenario understanding, focusing on end-to-end attention association estimation based on driver state tracking and dynamic scenario understanding. By fusing complementary information from two perspectives inside and outside the vehicle, the potential relationship between attention and scenarios is modeled at the feature level, thereby end-to-end mapping the driver's attention to the scenario and adaptively handling the challenges brought by highly complex dynamic environments and the driver's non-fixed postures, overcoming device dependence and improving system robustness.
[0046] As Figures 1 to 5 shown, in order to improve the quality of feature representation and ensure semantic consistency, a W-Net architecture is proposed in this example. This architecture systematically integrates information from two different domains to achieve end-to-end mapping of driver attention to the scenario. Its significant feature lies in the W-shaped framework of "independent encoding - independent partial decoding - fusion decoding", which uses skip connections to promote multi-scale feature integration throughout the encoding and decoding stages. The two-step decoding strategy is the essence of W-Net. Its main idea is to perform intra-domain feature alignment before feature fusion decoding, which ensures that the features extracted from each domain are semantically consistent, thus ensuring that the subsequent fusion process operates on well-aligned and high-quality feature representations. The method in this implementation can include the following steps:
[0047] S1. Independent encoding stage:
[0048] Optionally, in the independent encoding stage, synchronized videos of the driver inside the vehicle and the scenario outside the vehicle are obtained, and adjacent frames are respectively obtained. The adjacent frames of the driver image and the scenario image are independently feature-extracted through a dual-branch, generating corresponding multi-level features. Through dimensionality reduction processing with channel adjustment, multiple same-level features of the two branches are made to have the same number of channels.
[0049] In one of the embodiments, optionally, the model is trained based on the publicly available dataset Look-Both-Ways (LBW). The LBW dataset is a publicly available dataset, including driving scenario videos associated with driver attention, driver face videos, and driver three-dimensional line-of-sight ground truth. This dataset is captured using synchronized and calibrated cameras, a monocular camera facing the driver and a stereo image facing the scenario, as well as an eye-tracking glasses worn by the driver to capture 3D gaze data aligned with the real driving environment.
[0050] S101. In the dual-branch independent encoding stage, feature extraction is respectively performed on the driver face branch and the driving scenario branch to obtain corresponding multi-level feature representations.
[0051] In this embodiment, adjacent two-frame images of the driver face branch and the driving scene branch are respectively subjected to feature extraction through the Resnet18 and Resnet50 networks to obtain corresponding multi-level feature representations.
[0052] Specifically, the driver face feature extraction branch selects adjacent face image frames and as inputs, with its H F =W F =224, and uses Resnet-18 as the backbone network to extract features (removing the last fully connected layer). The backbone network is divided into four stacked convolutional units (P 1 , P 2 , P 3 , P 4 ), and the output features of each unit are correspondingly represented as:
[0053]
[0054] Among them, C Fi =64×2 (i-1) ,
[0055] The driving scene branch selects adjacent two-frame face image frames and as inputs, with its H F =W F =224, and uses Resnet-18 as the backbone network to extract features (removing the last fully connected layer). The backbone network is divided into four stacked convolutional units (P 1 , P 2 , P 3 , P 4 ), and the output features of each unit are correspondingly represented as:
[0056]
[0057] Among them, C Fi =64×2 (i-1) ,
[0058] Subsequently, they are sent into the corresponding four-channel adjustment units to make C Si =C Fi =64×2 (i-1) , i∈{1, 2, …4}, and obtain more context information. The channel adjustment process can be represented as the following operation:
[0059] a o =Conv 3×3 (Cat[Conv 1×1(a i ), Conv 3×3 (Conv 1×1 (a i )))
[0060] Among them, a i represents the input feature, a o represents the output feature, Cat(·) represents feature concatenation, Conv 1×1 (·) and Conv 2×3 (·) respectively represent the convolution operations with convolution kernels of 1×1 and 3×3.
[0061] So far, the multi-level feature extraction of the input dual-branch image has been completed.
[0062] S102. Perform frequency-domain and time-domain joint adaptive dynamic filtering on the inter-frame features of the two branches respectively to filter out self-motion and separate the independent motion enhanced feature representation. Specifically, calculate the dynamic features through the inter-frame feature consistency, transform them to the frequency domain through the fast Fourier transform, perform frequency-domain filtering through the convolutional layer, transform them back to the time domain through the inverse fast Fourier transform, and then perform time-domain filtering through the spatial attention filtering to obtain enhanced dynamic cues.
[0063] In a real driving scenario, the driver and the driving scenario are always in motion. To improve the model's ability to perceive key dynamic information in a highly complex dynamic environment, in this embodiment, a dynamic adaptive filtering DynamicAdaptive Filter (DAF) method is designed to guide the model to focus on the key motion features in the temporal images.
[0064] Specifically, divide the features of the input adjacent frames into a series of non-overlapping feature sub-regions After independently performing the correlation calculation in each sub-region, perform weighted averaging with the pixel grid to determine the optimal local motion vector of each pixel Finally, by calculating the difference from P Ln , the dynamic feature map D Ln of the sub-region can be obtained. This process is as shown in the formula:
[0065]
[0066] Among them, H L and W L respectively represent the height and width of the sub-region, each element in Corr represents the correlation value between the coordinates in L n1 and l n2 , and M L represents the matching probability distribution.
[0067] Next, by integrating the dynamic feature predictions of all local blocks, the overall dynamic feature estimation result D is obtained H ×W×2 . Encoding is performed to obtain a high-level semantic representation
[0068] D e = Conv 3×3 (Conv 1×1 (D))
[0069] In the above process of extracting the dynamic feature map, the motion attributes of all regions in the image are recorded without discrimination. In the frequency domain, different frequency components usually correspond to motion changes of different scales or speeds. To obtain the significant dynamic information in the image, in this example, the features are filtered in the frequency domain: transformed to the frequency domain through the two-dimensional fast Fourier transform (FFT(·)), and then, the frequency domain features are convolved and filtered to adjust the intensities of different frequency components of the image, enabling the model to adaptively filter out redundant dynamic information and retain the key feature D fe . This process can be expressed as:
[0070] D fe = IFFT(Conv 1×1 (FFT(D S )))
[0071] Next, D fe is input into the spatial attention unit, and this process is called SA(·). The specific representation of this process is as follows: this unit comprehensively extracts the spatial information of the input feature map using the mean and max pooling operations, generates attention weights through a convolutional layer, and passes these weights through the Sigmoid function to generate the spatial attention mask A s , and each element of the attention map can be interpreted as the importance score of each position, used to highlight the important regions therein. This process can be expressed by the following formula
[0072]
[0073] where, F Avg (·) and F Max (·) represent adaptive average pooling and adaptive max pooling, F Avg (D e ) ∈ R 1 ×H×W , F Max (D e ) ∈ R 1×H×W , σ(·) represents the sigmoid function, A s represents the spatial attention mask, and D S represents the feature after spatial attention
[0074] So far, the spatial attention mask A generated under the guidance of dynamic features has been obtained. s . Then, these two attention masks are applied to the corresponding original features. The spatial attention mask is multiplied with the original features in matrix form, and the importance of the elements in each region of the original features is changed based on the weights defined in the mask, further adjusting the spatial information in the feature map.
[0075] Apply dynamic adaptive filtering to the driver's facial features and scene features respectively:
[0076]
[0077] S103. Share global context information among the multi-level features of the facial feature branch to refine the feature representation, so as to adaptively process key features of different scales such as the head and eyes in non-fixed postures. Use multi-branch dilated convolutions and hierarchical feature fusion to integrate information for intra-level multi-scale feature extraction; use inter-level multi-scale features to aggregate the anchor feature maps from the base layer and the relevant feature sets from adjacent layers.
[0078] The global context information sharing aims to extract and aggregate multi-scale driver facial features to establish a global and refined feature representation, which includes two processes: intra-level multi-scale feature aggregation (ILA) and cross-level feature semantic alignment (CLA).
[0079] In order to mine the potential multi-scale information in the features at all levels of the encoder, an intra-level multi-scale feature aggregation unit ILA is designed in this example. For the original input features Project them onto k groups of features through grouped point convolutions where k represents the number of parallel branches, and then use dilated convolution branches with different dilation rates, and these branches learn in parallel from receptive fields of different sizes to obtain In order to eliminate the grid artifacts that may be caused by large receptive fields, a computationally efficient hierarchical feature fusion (HFF) structure is used to fuse the feature maps; a channel mixing attention (CHA) is designed to combine the feature maps from adjacent branches in turn, avoiding the problem of information redundancy or dilution that may be brought by simple addition. The specific process of the channel mixing attention CHA is as follows:
[0080]
[0081] where F Avg (·) represents a two-dimensional average pooling operation, w n-1 , wn represents the weight of the corresponding channel feature, σ represents the sigmoid function, and CHatt represents the channel mixing attention mask. represents the fused feature. After layer-by-layer fusion, the fusion result is further mapped through point convolution and residual connection to generate a comprehensive feature representation.
[0082]
[0083] To effectively fuse the different semantic information of features at different layers and enhance the consistency of feature representation, this embodiment designs a cross-level feature semantic alignment unit CLA.
[0084] CLA aims to aggregate the global multi-scale information contained in the feature representations at different levels. In this example, the reference anchor feature map and the feature set at the adjacent level are aggregated to create a more coherent and efficient information flow. Specifically, this unit takes a reference anchor feature map and the adjacent-level feature set N i as inputs, where:
[0085]
[0086] First, each feature in the feature set N i is adjusted to the same spatial size as f mi through adaptive average pooling or bilinear upsampling operations. Then, a convolutional layer is applied to transform these features into a feature space of the same dimension, and they are fused with the reference anchor feature map f mi by element-wise addition to achieve the interaction of multi-level information. This semantic alignment process is called CLA i .
[0087]
[0088] So far, the input driver face features have been processed through global context information sharing to obtain an enhanced feature representation
[0089] S2. In the independent partial decoding stage, the multi-level features of the two branches are decoded step by step within the domain through the hybrid attention mechanism to achieve in-domain feature semantic alignment and improve the feature quality.
[0090] To address the semantic inconsistency problem in the feature aggregation of driver face and road scene images, the independent partial decoding step decodes each feature type separately before fusion to ensure semantic alignment within each domain. The Channel-Spatial Hybrid Attention (CSHA) aims to better recover the lost detailed information and enhance the feature representation.
[0091] In this example, is input into CSHA together with the shallower feature maps from the skip connection. Input CSHA. The shape becomes through the upsampling operation and linear transformation. Next, it and are fed into the CHA unit for channel mixing attention to emphasize the important cross-channel relationships. The specific process of CHA is as follows:
[0092]
[0093] To recover finer spatial information during the decoding process, spatial mixing attention is designed for fine-grained spatial fusion. The Spatial Hybrid Attention (SHA) captures the spatial context relationships between the two features, enabling the model to more accurately reconstruct the key regions in the image. The specific process is as follows:
[0094]
[0095] where Mean(·) represents calculating the average in the channel dimension, SHatt represents the spatial mixing attention, represents the output result of the feature after passing through the channel-spatial hybrid attention.
[0096] For the face features, independent partial decoding is performed on g si , i ∈ {1, 2, 3, 4}. By cascading 3 CSHA units, the corresponding output features f pi , i ∈ {1, 2, 3, 4} are obtained.
[0097]
[0098] Similarly, for the driving scene features after independent partial decoding, the output features
[0099]
[0100] S3. In the fusion decoding stage, the same-level features of the two branches are gradually fused and decoded through the cross-attention mechanism and the channel-spatial hybrid attention mechanism to generate the final feature representation, and the final predicted driver attention heat map associated with the scene is obtained through the decoding layer.
[0101] To deeply fuse the driver's face and scene features, a cross-domain fusion decoding step is designed in this example, and the long-distance dependence capture ability of the multi-head cross-attention is used to achieve feature fusion.
[0102] The multi-head cross-attention unit mainly consists of three main parts: feature encoding, cross-attention, and the feed-forward neural network FFN. This unit takes the features of the driving scene and the driver's face features as inputs, where C si = C Fi = 64 × 2 (i-1) .
[0103] First, the two input features are encoded separately, and the formula is as follows:
[0104]
[0105] For the scene features, Pos si represents the position encoding corresponding to s ei , Conv si×si represents the convolution operation with a convolution kernel and a stride both of si, and Flatten represents the flattening operation. The operation for the face features is the same.
[0106] Next, the multi-head cross-attention is used for feature fusion. The scene features and the face features are linearly transformed separately and the input sequence is divided into multiple subsequences, then the attention of the sequences is calculated in different subspaces, and then the results of each subspace are concatenated. The process of the multi-head cross-attention is as follows:
[0107]
[0108] Among them, n represents the number of heads, which is set to 4 in this paper, and d k represents the dimension of the subspace. Among them, the linear mapping of s ei is used as Q (query), and the linear mappings of f ei are used as K (key) and V (value). Concat represents feature concatenation, and Linear represents the linear transformation layer. MHCA o is the output of MHCA.
[0109] Next, the outputs of all attention heads are integrated through a linear layer. Finally, the fused features are reconstructed back to be the same as s eiThe same spatial dimensions.
[0110]
[0111] Among them, Upsample represents the upsampling operation, and Conv represents the convolutional layer. After the scene features and driver features are fused layer by layer, the fused feature d is obtained i , i ∈ {1, 2, 3, 4}.
[0112] Next, 3 times of hierarchical decoding are performed through the CSHA module, and the shape fusion feature is obtained at the last level Next, through the head, finally an attention heat map with the same resolution as the input scene graph is obtained, and this process is expressed as follows:
[0113]
[0114] Among them, σ represents the sigmoid function, F Avg (·) represents adaptive average pooling, ConvT represents 2 consecutive transposed convolutional layers, CR(·) represents a channel number reduction unit, and after passing through each such unit, the number of channels will be reduced to the original represents the finally predicted driver attention map, which has the same resolution as the input driving scene.
[0115] So far, the model has estimated the attention heat map associated with the scene.
[0116] In this example, the attention ground truth map is divided into two types: discrete fixation map and continuous fixation map.
[0117] For the discrete fixation map, in the work of the LBW dataset, the measure of the driver's attention visual saliency ground truth s g (x) is expressed as a function of the 3D gaze direction g, and can be predicted through the facial appearance image I g . Given the depth estimation of the stereo scene camera in X, 3D points are reconstructed and projected onto the eye center e to form the direction s. The angular difference between s and the gaze direction g is used to model the projected scene saliency s g for modeling.
[0118]
[0119] Among them, R ∈ SO(3) is the rotation matrix that transforms the scene image coordinate system into the face image coordinate system, K is the camera intrinsic parameter, d(x) ∈ R+ is the depth, X is the 3D point corresponding to x, e ∈ R3 is the 3D position of the eye center, s(x) is the direction vector formed by the 3D point of the eye position e and X seen, g is the true gaze direction, I sIt is a scene image. The resulting ground truth map of discrete fixation points.
[0120] For the continuous fixation map, in the experiment, when generating the discrete fixation map, there was a problem that the attention map was inaccurate due to the missing part of the depth map. To improve this problem and considering the continuity of the driver's line of sight, in this example, a method similar to the DR(eye)VE dataset was adopted to produce the "continuous fixation map" within 1 second. Multiple frames of images within the time segment centered on the current frame were selected, and the fixation points in each frame were mapped to the coordinates of the current frame through homography transformation. A Gaussian function with fixed parameters was used to simulate the fixation point probability distribution, and a spatio-temporal weight mechanism was introduced to assign weights according to the spatio-temporal proximity between the fixation points and the current frame. By integrating these weighted Gaussian distributions, an accurate hot map of the driver's visual attention was generated. The formula for calculating the weights of different frames is as follows:
[0121]
[0122] where, spatial_distance k represents the spatial distance, time_distance k represents the time distance, which represents the number of frames between any frame within the time window and the current frame. Since the number of frames reflects the length of time, in this embodiment, the difference in the number of frames is also used to represent the distance in time. a ∈ R represents an adjustable real number parameter. distance k represents the overall distance, probability k represents the corresponding weight, and σ = 200.
[0123] When predicting the driver's attention based on the driver's attention hot map, it can be achieved by methods such as the observation method, the comparison method, or a trained network model, etc., which will not be elaborated here.
[0124] S4: Loss function and evaluation metrics:
[0125] Since different models may have their own advantages and disadvantages in different metrics, in this example, the average results of multiple metrics were used to comprehensively evaluate the model more comprehensively. The metrics based on position include the Normalized Scanpath Saliency (NSS), and the metrics based on distribution include the Kullback-Leibler divergence (KL), the Linear Correlation Coefficient (CC), and the Similarity Index (SIM). When the model is trained, the loss function is set as the combination of the above 4 evaluation metrics to optimize the parameters of the network, and its calculation formula is as follows:
[0126] L(E, G) = ρ 1 MAE(E, G) + ρ 2NSS(E, G) + ρ 3 KL(E, G) + ρ 4 SIM(E, G) + ρ 5 CC(E, G)
[0127] Where E is the network's estimated driver attention result map; G is the ground truth map of driver attention; G is the ground-truth, which is the standard driver attention map after processing the attention map; μ is the weight of each loss (by experience, ρ 1 is 1, ρ 2 is -0.005, ρ 3 is 1, ρ 4 is -0.2, ρ 5 is -0.1); MAE, NSS, KL, SIM, and CC represent the mean absolute error, NSS loss, KL loss, SIM loss, and CC loss respectively, and their calculation formulas are as follows:
[0128]
[0129]
[0130] Where E i and G i represent the points on the estimated driver attention map and the ground truth map respectively, δ represents the regularization constant to prevent division by zero, Cov(·) represents the covariance function, and σ and μ represent the standard deviation and the mean respectively.
[0131] During the training process, AdamW is used as the optimizer and CyclicLR is used as the learning rate scheduler, and its learning rate boundaries are 2e-6 to 1e-5. This example is implemented in the Python 3.7.2 and Pytorch 1.12.1 environments, equipped with 2 NVIDIA A800 GPUs and an Intel(R) Xeon(R) Gold 6326 CPU @ 2.90GHz. Using mixed-precision training to balance speed and accuracy, the batch size is set to 64, and a total of 40 epochs are trained.
[0132] Experimental results:
[0133] Under the above evaluation metrics, as shown in Table 1, when the discrete fixation point map is used as the ground truth, the final output results of the model after training on the dataset are KL = 0.9299, CC = 0.6624, SIM = 0.5485, and NSS = 2.6273.
[0134] Table 1 Final output results of the model after training on the dataset when the discrete fixation point map is used as the ground truth
[0135]
[0136] As shown in Table 2, when the continuous fixation point map is used as the ground truth, KL = 2.4054, CC = 0.3496, SIM = 0.2422, and NSS = 2.7883.
[0137] Table 2 Final output results of the model after training on the dataset with the continuous fixation point map as the ground truth
[0138]
[0139] In summary, in this embodiment, end-to-end estimation is achieved through the "independent encoding - independent partial decoding - fusion decoding" W-shaped network architecture. Specifically, in the independent encoding stage, synchronous adjacent frames of the driver and the scene are obtained, and dual-branch independent feature extraction is performed to generate corresponding multi-level features; the inter-frame features of the two branches are respectively subjected to joint adaptive dynamic filtering in the frequency domain and time domain to separate independent motions and enhance feature representation; global context information sharing is performed on the multi-level features of the facial feature branch to refine the feature representation for adaptively processing key features of different scales such as the head and eyes in non-fixed postures; in the independent partial decoding stage, the multi-level features of the two branches are respectively decoded step by step within the domain through a hybrid attention mechanism to achieve semantic alignment of intra-domain features and improve feature quality; in the fusion decoding stage, the same-level features of the two branches are gradually fused and decoded through a cross-attention mechanism and a channel-space hybrid attention mechanism to generate the final feature representation, and the final predicted driver attention heat map associated with the scene is obtained through the decoding layer. The present invention can avoid data and device dependencies of traditional methods, overcome the cross-frame mapping challenges in highly complex dynamic driving scenarios, and achieve end-to-end estimation of the environmental area that the driver is currently focusing on.
[0140] Embodiment 2:
[0141] This embodiment provides an attention correlation estimation system that fuses driver state and scene understanding, including:
[0142] A data acquisition module, configured to: acquire in-vehicle driver facial video information and out-of-vehicle scene video information;
[0143] An independent encoding module, configured to: extract synchronous adjacent frame images of in-vehicle driver facial video information and out-of-vehicle scene video information; perform dual-branch feature extraction on the two types of adjacent frame images respectively to obtain corresponding multi-level features; perform joint adaptive dynamic filtering in the frequency domain and time domain on the inter-frame multi-level features of the two branches respectively; and perform global context information sharing on the multi-level features of the facial feature branch to adaptively process key features of different scales in non-fixed postures;
[0144] An independent partial decoding module is configured to: perform intra-domain step-by-step decoding on multi-level features of two branches respectively through a hybrid attention mechanism, align the intra-domain feature semantics, and the channel-spatial hybrid attention mechanism balances the cross-channel and spatial context relationships through a designed channel integration unit CIU and a spatial integration unit SIU, balances the contributions of features at different levels during the decoding process, and achieves high-quality decoding;
[0145] A fusion decoding module is configured to: perform step-by-step fusion and decoding on the same-level features of two branches through a cross-attention mechanism and a hybrid attention mechanism, and obtain a driver attention heat map associated with the scene through decoding prediction.
[0146] The working method of the system is the same as the attention correlation estimation method for fusing driver state and scene understanding in Embodiment 1, and will not be elaborated here.
[0147] Embodiment 3:
[0148] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored on the memory and capable of running on the processor. When the processor executes the program, the steps of the attention correlation estimation method for fusing driver state and scene understanding described in Embodiment 1 are implemented.
[0149] Embodiment 4:
[0150] This embodiment provides a computer program product. The computer program product includes a computer program. When the computer program is executed by a processor, the steps of the attention correlation estimation method for fusing driver state and scene understanding described in Embodiment 1 are implemented.
Claims
1. An attention correlation estimation method integrating driver status and scene understanding, characterized in that: include: Obtain the driver's facial video information inside the car and the scene video information outside the car; Extracting synchronous adjacent frame images of the driver's facial video information inside the vehicle and the scene video information outside the vehicle; Perform dual-branch feature extraction on two adjacent frame images to obtain corresponding multi-level features; The multi-level features between frames of the two branches are subjected to joint adaptive dynamic filtering in the frequency domain and time domain respectively; and the multi-level features of the facial feature branch are subjected to global context information sharing to adaptively process key features of different scales under non-fixed postures; Through the channel-spatial hybrid attention mechanism, the multi-level features of the two branches are decoded level by level in the domain and the semantics of the features in the domain are aligned; Through the cross attention mechanism and channel-spatial mixed attention mechanism, the same-level features of the two branches are fused and decoded step by step, and the driver's attention heat map associated with the scene is estimated through the decoding head.
2. The attention association estimation method integrating driver status and scene understanding as claimed in claim 1, characterized in that: The overall design of the method is end-to-end attention estimation based on the dual-domain fusion of driver facial features and driving scene features; through the W-shaped network architecture of independent encoding-independent partial decoding-fusion decoding, cross-domain images are encoded and fused, and the driver's attention heat map associated with the scene is directly generated by integrating the input dual-domain information.
3. The attention association estimation method integrating driver status and scene understanding as claimed in claim 1, characterized in that: The driver's facial features and driving scene features are encoded independently by two branches to obtain a multi-level feature representation. Multi-layer convolution is used to reduce the number of channels so that the multi-level features of the scene branch and the corresponding levels of the driver's facial branch maintain the same number of channels.
4. The attention association estimation method integrating driver status and scene understanding as claimed in claim 1, characterized in that: The frequency-domain and time-domain joint adaptive dynamic filtering is performed on the multi-level features between frames of the two branches respectively; the features of adjacent frames are divided into several non-overlapping sub-regions, and after independent correlation calculation is performed in each sub-region, weighted average is performed with the pixel grid to determine the optimal local motion vector of each pixel, and the dynamic feature map of the sub-region is obtained by calculating the difference, and the overall dynamic feature estimation result is obtained by integrating the dynamic feature predictions of all local blocks; The features are converted to the frequency domain through a two-dimensional fast Fourier transform, and frequency domain filtering is performed through a convolutional layer to adjust the intensity of different frequency components of the image to filter out the self-motion that is recorded indiscriminately; The features are transformed back to the time domain, and the spatial attention mask is obtained through spatial attention, which is used to perform time domain filtering on the features and highlight important spatial elements.
5. The attention association estimation method integrating driver status and scene understanding as claimed in claim 1, characterized in that: The multi-level features of the facial feature branch are shared with global context information to adaptively process key features of different scales under non-fixed postures; multi-scale information is extracted by using dilated convolution branches with different dilation rates for intra-layer features, and then the feature maps of different branches are fused using a hierarchical feature fusion structure; channel mixed attention is used to fuse adjacent branches; The features of different levels are used as the benchmark anchor feature map in turn; the spatial dimensions of adjacent levels are adjusted to be the same as the benchmark anchor feature map through adaptive average pooling or bilinear upsampling operations, the convolutional layer is applied to transform the features to the feature space of the same dimension, and the converted features are fused with the benchmark anchor feature map through element-by-element addition to realize the interaction of multi-level information.
6. The attention association estimation method integrating driver status and scene understanding as claimed in claim 2, characterized in that: In the independent decoding stage, the multi-level features of the two branches are decoded step by step in the domain through the hybrid attention mechanism to align the semantics of the features in the domain. The channel-space hybrid attention mechanism balances the cross-channel and spatial contextual relationships through the designed channel integration unit and spatial integration unit, and balances the contribution of features at different levels in the decoding process to achieve decoding.
7. The attention association estimation method integrating driver status and scene understanding as claimed in claim 1, characterized in that: When performing step-by-step fusion and decoding, the scene features and facial features are fused through multi-head cross attention, and the multi-level features of the two branches are decoded step by step through the channel-spatial mixed attention mechanism, aggregating the dual-domain heterogeneous information from the driver's facial features and driving scene features.
8. An attention correlation estimation system integrating driver status and scene understanding, characterized in that: include: The data acquisition module is configured to: obtain the facial video information of the driver in the vehicle and the video information of the scene outside the vehicle; The independent encoding module is configured to: extract synchronous adjacent frame images of the driver's face video information inside the vehicle and the scene video information outside the vehicle; perform dual-branch feature extraction on the two adjacent frame images respectively to obtain corresponding multi-level features; The multi-level features between frames of the two branches are subjected to joint adaptive dynamic filtering in the frequency domain and time domain respectively; and the multi-level features of the facial feature branch are subjected to global context information sharing to adaptively process key features of different scales under non-fixed postures; The independent part decoding module is configured to: decode the multi-level features of the two branches level by level in the domain through the channel-spatial hybrid attention mechanism, and align the semantics of the features in the domain; The fusion decoding module is configured to: fuse and decode the same-level features of the two branches step by step through the multi-head cross attention mechanism and the channel-space mixed attention mechanism, and obtain the driver's attention heat map associated with the scene through decoding prediction.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: When the processor executes the program, the steps of the attention association estimation method integrating driver status and scene understanding as described in any one of claims 1 to 7 are implemented.
10. A computer program product, characterized in that The computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the attention association estimation method integrating driver status and scene understanding as described in any one of claims 1 to 7 are implemented.