Rapid fusion network driver attention prediction method in dangerous driving scene

By combining DeeplabV3, multi-path 3D coding architecture, convolutional long short-term memory network and AC-mix module, the accuracy and computing efficiency of driver attention prediction in complex driving scenarios are solved, and efficient driver attention prediction is achieved.

CN120564149APending Publication Date: 2025-08-29ZHONGBEI UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510499620.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-01-17
Filing Date
2025-04-21
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

The prior art is difficult to accurately predict driver attention in complex driving scenarios, and the computing resource demand is high, and the semantic context information is lacking sufficient consideration, resulting in a high misjudgment rate.

Method used

The semantic images of video frames are extracted by DeeplabV3 semantic segmentation method, combined with multi-path 3D encoding architecture and convolutional long and short-term memory network, and used the AC-mix module to fuse spatiotemporal and semantic context features, model the scene relationship through the graph convolution network, and used the attention graph decoding module to generate the driver's attention map.

Benefits of technology

Improves the accuracy and computing efficiency of driver attention prediction, reduces computational complexity, and enables the rapid identification of key objects or areas that attract driver attention in complex driving environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564149A_ABST
    Figure CN120564149A_ABST
Patent Text Reader

Abstract

A fast fusion network driver attention prediction method in a dangerous driving scene comprises the following steps: step 1, segmenting an RGB set in a traffic accident video set into semantic images with different semantic features frame by frame, and extracting spatio-temporal features and semantic features of the semantic images; 2, fusing the spatio-temporal features and the semantic context features of the image extracted in the step 1 by using an attention strategy; step 3, constructing an attention fusion module by using an AC-mix module to quickly identify a key object or region attracting the attention of the driver; and 4, converting the potential driver attention map obtained in the step 3 into a final driver attention map by using an attention map decoding module. According to the method, the semantic context related to the driving scene is comprehensively fused, so that the prediction accuracy is improved; meanwhile, an AC-mix module is integrated, and the global perception capability and the local feature extraction capability are combined, so that the path calculation efficiency is improved, and the calculation complexity in the prediction process is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of safe driving technology, and specifically provides a method for rapidly integrating network driver attention prediction in dangerous driving scenarios. Background Art

[0002] Advanced driver assistance systems (ADAS) are crucial to the development of autonomous driving. Understanding driver intent is a key research area for both ADA systems and collaborative human-machine autonomous driving technologies. A statistical survey shows that drivers' tendency to focus their attention on unusual areas during driving is a significant contributing factor to accidents. Therefore, among this vast amount of driving information, ADAS must prioritize key information directly related to driving behavior and accurately predict the driver's primary focus.

[0003] The traffic driving environment is a complex and dynamically changing scene, where numerous objective and subjective factors converge to automatically control the driver's field of vision and attention. These factors can range from bottom-up sensory stimuli, such as posted speed limit signs or traffic lights, to top-down goals or experiences, such as pedestrians and oncoming vehicles. While driving in traffic, drivers typically allocate their attention to the most important and salient areas or targets at the moment. Understanding how drivers allocate their potential attention and the areas of primary focus is crucial for driver assistance systems.

[0004] In the field of visual saliency detection, traditional methods typically use low-level image attributes such as brightness, color, edges, contrast, foreground-background differences, and discontinuities to identify salient regions, resulting in poor robustness and a large number of false positives in detection results. Following the application of deep learning to the field of saliency detection, convolutional neural networks and other methods have been integrated into the task, enabling the learning of richer feature representations and saliency patterns during the detection process. Furthermore, reinforcement learning has been combined with top-down attention models to produce more accurate attention prediction results. While existing research has optimized model architectures and adopted multimodal or reward mechanisms, achieving some improvement in accuracy, it lacks sufficient consideration of the semantic contextual information in complex driving scenarios, resulting in high randomness in real-world driving scenarios and high computational resource requirements. Summary of the Invention

[0005] The present invention provides a method for rapidly fusion-network driver attention prediction in dangerous driving scenarios to address the deficiencies in the prior art.

[0006] The present invention is achieved through the following technical solutions:

[0007] A fast fusion network driver attention prediction method in dangerous driving scenarios includes the following steps:

[0008] Step 1: Segment the RGB set of the traffic accident video set into semantic images with different semantic features frame by frame, and extract their spatiotemporal features and semantic features;

[0009] Step 2: Use the attention strategy to fuse the spatiotemporal features of the image extracted in step 1 with the semantic context features;

[0010] Step 3: Use the attention fusion module built with the AC-mix module to quickly identify key objects or areas that attract the driver's attention;

[0011] Step 4: Use the attention map decoding module to convert the potential driver attention map obtained in step 3 into the final driver attention map.

[0012] In the method for rapidly fusion-network driver attention prediction in dangerous driving scenarios as described above, the DeeplabV3 semantic segmentation method is used in step 1 to obtain a semantic image of the video clip; at the same time, the size of the input continuous frames is adjusted to 256×192, and the spatiotemporal information of the video frames is extracted using 3D convolution in a multi-path 3D coding architecture.

[0013] In the above-mentioned method for rapidly fusion-network driver attention prediction in dangerous driving scenarios, in step 1, after normalizing the semantic graph in the multi-path 3D coding architecture, the relationship between different semantic categories of the scene is modeled through a graph convolutional network to extract the semantic context features of the driving scene.

[0014] In the above-mentioned method for rapidly fusion-network driver attention prediction in dangerous driving scenarios, after obtaining the spatiotemporal features and semantic context features of the traffic accident video set in step 2, a convolutional long short-term memory network is used to learn and transfer the fusion details in consecutive T frames to the T+1 frame.

[0015] In the aforementioned method for rapidly fusion-network driver attention prediction in dangerous driving scenarios, in step three, when fusing details within consecutive T frames, AC-mix utilizes the characteristic of using depthwise separable convolutions instead of less efficient tensor shift operations. This optimizes the computational path and reduces the number of repeated computations to achieve the effect of rapidly identifying key objects or areas that attract the driver's attention.

[0016] In the above-mentioned method for rapidly fusion-network driver attention prediction in dangerous driving scenarios, the method for constructing the attention fusion module in step 3 includes the following steps:

[0017] Step 1: Split the selected dataset into 3:1:1 ratio for training, validation, and testing respectively;

[0018] Step 2: Use the DeeplabV3 semantic segmentation method to obtain the semantic image of the video clip. DeeplabV3 uses dilated convolution and ASPP structure to improve the ability to segment objects of different scales;

[0019] Step 3: 3D convolution in the multi-path 3D coding architecture extracts the spatiotemporal information of the video frame. Each path of the multi-path 3D coding architecture has the same structure, consisting of three interleaved blocks, namely 3D convolution blocks, 3D batch normalization + rectified linear unit blocks, and 3D maximum pooling blocks. The 3D batch normalization + rectified linear unit blocks are used to accelerate the convergence of network training and resist gradient disappearance. Each path of the multi-path 3D coding architecture has a total of 10 3D convolution blocks, 10 3D batch normalization + rectified linear unit blocks, and 3 3D maximum pooling blocks;

[0020] Step 4: After steps 2 and 3, the spatiotemporal features and semantic context features of the traffic accident video set are obtained. The convolutional long short-term memory network is used to learn and transfer the fusion details in consecutive T frames to the T+1 frame, using the memory unit C t and hidden state unit H t The time t is used to control the memory update and output H t , and sequentially transmits spatiotemporal scene features. The convolutional long short-term memory network module realizes the transition of the fusion details of T consecutive frames to t+1, which is defined as:

[0021] [C t+1 ,H t+1 ]=ConvLSTM(Z 1:t ,W c )

[0022] Among them, Z 1:t is the input sequence before time t, W c is the weight parameter to be learned;

[0023] Step 5: When fusing details within consecutive T frames, AC-mix uses depthwise separable convolutions instead of less efficient tensor shift operations. This optimizes the computational path and reduces repeated computations, enabling rapid identification of key objects or areas that attract the driver's attention.

[0024] Step 6: Use the attention map decoding module to convert the potential driver attention map obtained in step 3 into the final driver attention map;

[0025] Step 7: Measure the performance of the network in this study. Relative entropy, Pearson correlation coefficient, similarity coefficient, standardized scan path significance, and area under the ROC curve are selected as significance evaluation indicators.

[0026] In the above-mentioned method for predicting driver attention using a fast fusion network in dangerous driving scenarios, the semantic context features of the driving scene are extracted by modeling the relationship between different semantic categories of the scene through a graph convolutional network. The operation of extracting the semantic context features of the driving scene includes the following steps:

[0027] Step (1): Frame image construction based on the features of each semantic image, semantic features Among them, H, W, and C represent The height, width and number of channels of We reformulate the semantic features of the t-th frame as a matrix Where N = H × W represents the number of nodes in the frame, which is the measurement The pairwise node similarity within is defined as:

[0028]

[0029] in, and express The linear transformation of is the node similarity matrix, and then Normalize to get the correlation matrix Correlation Matrix is constructed as a frame graph;

[0030] Step (2): After obtaining the frame graph, use graph convolution to calculate the relationship between nodes, which is defined as:

[0031]

[0032] in, represents the weight of the graph convolution layer, is the output of each graph convolution layer, for each Perform graph convolution calculation to obtain the semantic context features of T frames In the multi-path 3D coding architecture, H, W, and C are 32, 24, and 512, respectively.

[0033] In the above-mentioned method for rapidly fusion-network driver attention prediction in dangerous driving scenarios, after obtaining the hidden driver attention map in step 6, it is input into the attention map decoding module to generate the final driver attention map of the (T+1)th frame. The attention map decoding module is implemented as upsampling (×4) → convolution (3×3,128) → batch normalization + rectification linearization → upsampling (×2) → convolution (3×3,1) → Sigmold, where the Sigmoid function is used to limit the output value of the driver attention map to [0,1].

[0034] In the aforementioned method for rapidly fusion-based driver attention prediction in dangerous driving scenarios, the relative entropy, also known as Kullback-Leibler divergence, can be used to determine the degree of difference between the probability distribution of the predicted attention map and the true probability distribution of the attention map. The greater the similarity between the two distributions, the lower the relative entropy, while the greater the difference, the greater the relative entropy. The mathematical expression is:

[0035]

[0036] Among them, ε represents a very small regularization coefficient, i represents the i-th pixel, Q represents the probability distribution of the real image, and P represents the probability distribution of the predicted image;

[0037] In the above-mentioned fast fusion network driver attention prediction method under dangerous driving scenarios, the Pearson correlation coefficient represents the degree of linear correlation between two attention maps, and the result is usually between -1 and 1. When using this metric, the continuous distribution of the importance prediction result P and the authenticity of human eye attention Q is considered to be a random variable, and its mathematical expression is:

[0038]

[0039] Among them, CoV(,) represents the covariance,

[0040] In the above-mentioned fast fusion network driver attention prediction method under dangerous driving scenarios, the similarity coefficient is used to evaluate the similarity between two saliency maps. The continuous distribution of the saliency prediction result P and the true value of human eye attention Q is regarded as a probability distribution. The closer it is to 1, the more accurate the prediction result is and the closer it is to the true value. Its mathematical expression is:

[0041] SIM(P,Q)=∑ i min(P' i ,Q' I ),∑ i P' i =1,∑ i Q' i =1;

[0042] In the above-mentioned fast fusion network driver attention prediction method for dangerous driving scenarios, the standardized scan path saliency is defined as the average value of the normalized saliency at the human eye focus position. The prediction ability of the attention fusion module for the driver's gaze point position can be evaluated by determining the saliency probability value of each gaze point position to determine the quality of the model prediction result. Its mathematical expression is:

[0043]

[0044] Where N represents the number of gaze points of all drivers, σ and μ are the mean and standard deviation of the prediction results;

[0045] In the above-mentioned method for predicting driver attention using a fast fusion network in dangerous driving scenarios, the area under the ROC curve is constructed by plotting the false positive probability on the horizontal axis and the true positive probability on the vertical axis. The significance detection result P can be binarized to obtain the ROC curve by sliding the threshold on [0, 1]. When a smaller threshold is used, the overall similarity between the two probability distributions can be calculated; conversely, when a larger threshold is set, the similarity between the two distributions at the peak can be determined, and the AUC index can be calculated through the ROC curve. The larger the AUC value, the better the algorithm performance. When the AUC value is close to 1, it indicates that the significance estimate is completely consistent with the true value calibration. According to the definition of the ROC curve, the AUC index is mainly affected by the high threshold, and its mathematical expression is:

[0046]

[0047] According to the method for rapidly fusion network driver attention prediction in dangerous driving scenarios as described above, after obtaining the hidden driver attention map H in step 4, it is input into the attention map decoding module to generate the final driver attention map of the (T+1)th frame.

[0048] The advantages of the present invention are: the present invention fully incorporates the semantic context related to the driving scenario, thereby improving the accuracy of the prediction; at the same time, the present invention integrates the AC-mix module, combines global perception capabilities and local feature extraction capabilities, improves the efficiency of the calculation path, and reduces the computational complexity in the prediction process. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0050] Figure 1 is a flow chart of the present invention;

[0051] Figure 2 It is a schematic diagram of the network model structure constructed by the present invention;

[0052] Figure 3 This is the attention prediction map under the “travel through” condition of the verification experiment of the present invention;

[0053] Figure 4is the attention prediction graph under the “collision” condition of the verification test of the present invention;

[0054] Figure 5 is the attention prediction map under the “influence” condition of the verification experiment of the present invention;

[0055] Figure 6 is a driver attention distribution diagram of the verification test of the present invention;

[0056] Figure 7 This is the attention prediction graph comparing the present invention with the SENet model. DETAILED DESCRIPTION

[0057] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0058] like Figure 1 As shown in FIG, a fast fusion network driver attention prediction method in dangerous driving scenarios includes the following steps:

[0059] Step 1: Segment the RGB set of the traffic accident video set into semantic images with different semantic features frame by frame, and extract their spatiotemporal features and semantic features;

[0060] Step 2: Use the attention strategy to fuse the spatiotemporal features of the image extracted in step 1 with the semantic context features;

[0061] Step 3: Use the attention fusion module built with the AC-mix module to quickly identify key objects or areas that attract the driver's attention;

[0062] Step 4: Use the attention map decoding module to convert the potential driver attention map obtained in step 3 into the final driver attention map.

[0063] Preferably, in step 1 described in this embodiment, the DeeplabV3 semantic segmentation method is used to obtain the semantic image of the video clip; at the same time, the size of the input continuous frames is adjusted to 256×192, and the 3D convolution in the multi-path 3D coding architecture is used to extract the spatiotemporal information of the video frames.

[0064] Preferably, in step 1 described in this embodiment, after the semantic graph is normalized in the multi-path 3D coding architecture, the relationship between different semantic categories of the scene is modeled through a graph convolutional network to extract the semantic context features of the driving scene.

[0065] Preferably, in step 2 of this embodiment, after obtaining the spatiotemporal features and semantic context features of the traffic accident video set, a convolutional long short-term memory network is used to learn and transfer the fusion details in T consecutive frames to the T+1 frame.

[0066] Preferably, in step three of this embodiment, when fusing the details within consecutive T frames, AC-mix can use depthwise separable convolution to replace the less efficient tensor shift operation, thereby optimizing the calculation path and reducing the number of repeated calculations to achieve the effect of quickly identifying key objects or areas that attract the driver's attention.

[0067] Preferably, the method for constructing the attention fusion module in step three of this embodiment includes the following steps:

[0068] Step 1: Split the selected dataset into 3:1:1 ratios for training, validation, and testing, i.e., 598 sequences (about 214k frames), 198 sequences (about 64k frames), and 222 sequences (about 70k frames), respectively.

[0069] Step 2: Use the DeeplabV3 semantic segmentation method to obtain the semantic image of the video clip. DeeplabV3 uses dilated convolution and ASPP structure to improve the ability to segment objects of different scales. Dilated convolution allows feature maps to propagate at different scales, while ASPP structure allows the network to perform parallel processing at multiple depths of feature maps, thereby obtaining semantic context information of driving scenes at different scales.

[0070] Step 3: The 3D convolution in the multi-path 3D coding architecture extracts the spatiotemporal information of the video frame. Each path of the multi-path 3D coding architecture has the same structure, consisting of three interlaced blocks, namely 3D convolution blocks, 3D batch normalization + rectified linear unit blocks, and 3D maximum pooling blocks. The 3D batch normalization + rectified linear unit blocks are used to accelerate the convergence of network training and resist gradient disappearance. Each path of the multi-path 3D coding architecture has a total of 10 3D convolution blocks, 10 3D batch normalization + rectified linear unit blocks, and 3 3D maximum pooling blocks, as shown in Figure 2 As shown, Figure 2 The size of each block is shown, where the values ​​of each block specify the block height, block width, and block channel respectively; 3D convolution extracts temporal information by sliding in the time dimension and capturing the dynamic changes between previous and next frames;

[0071] Step 4: After steps 2 and 3, the spatiotemporal features and semantic context features of the traffic accident video set are obtained. The convolutional long short-term memory network is used to learn and transfer the fusion details in consecutive T frames to the T+1 frame, using the memory unit Ct and hidden state unit H t Time t (i.e. frame index) is used to control memory update and output H t , and sequentially transmits spatiotemporal scene features (the present invention ignores the difference between visual path and semantic path); the convolutional long short-term memory network module realizes the transition of fusion details of T consecutive frames to t+1, which is defined as:

[0072] [C t+1 ,H t+1 ]=ConvLSTM(Z 1:t ,W c )

[0073] Among them, Z 1:t is the input sequence before time t, W c is the weight parameter to be learned;

[0074] Step 5: When fusing details within T consecutive frames, AC-mix utilizes depthwise separable convolutions instead of less efficient tensor shift operations. By optimizing computation paths and reducing repeated computations, it can quickly identify key objects or areas that attract the driver's attention. AC-mix introduces an attention mechanism that learns the importance (i.e., weight) of each neighboring node, allowing the network to focus more on neighboring nodes that are more important to the current task, thereby reducing unnecessary computations. Furthermore, by introducing a hybrid computation path strategy, the propagation paths of node information are divided into different categories or levels according to certain criteria. By rationally dividing computation paths, AC-mix can reduce redundant computations between nodes, thereby improving computational efficiency.

[0075] Step 6: Use the attention map decoding module to convert the potential driver attention map obtained in step 3 into the final driver attention map;

[0076] Step 7: Measure the performance of the network in this study. Relative entropy, Pearson correlation coefficient, similarity coefficient, standardized scan path significance, and area under the ROC curve are selected as significance evaluation indicators.

[0077] Preferably, the extraction of the semantic context features of the driving scene described in this embodiment is obtained by modeling the relationship between different semantic categories of the scene through a graph convolutional network. The extraction operation of the semantic context features of the driving scene includes the following steps:

[0078] Step (1): Frame image construction based on the features of each semantic image, semantic features Among them, H, W, and C represent The height, width and number of channels of We reformulate the semantic features of the t-th frame as a matrix Where N = H × W represents the number of nodes in the frame, which is the measurement The pairwise node similarity within is defined as:

[0079]

[0080] in, and express The linear transformation of is the node similarity matrix, and then Normalize to get the correlation matrix Correlation Matrix is constructed as a frame graph;

[0081] Step (2): After obtaining the frame graph, use graph convolution to calculate the relationship between nodes, which is defined as:

[0082]

[0083] in, represents the weight of the graph convolution layer, is the output of each graph convolution layer, for each Perform graph convolution calculation to obtain the semantic context features of T frames In the multi-path 3D coding architecture, H, W, and C are 32, 24, and 512, respectively.

[0084] Preferably, after obtaining the hidden driver attention map in step 6 described in this embodiment, it is input into the attention map decoding module to generate the final driver attention map of the (T+1)th frame. The attention map decoding module is implemented as upsampling (×4) → convolution (3×3,128) → batch normalization + rectification linearization → upsampling (×2) → convolution (3×3,1) → Sigmold, where the Sigmoid function is used to limit the output value of the driver attention map to [0,1].

[0085] Preferably, the relative entropy described in this embodiment is also called Kullback-Leibler divergence, which can be used to determine the degree of difference between the probability distribution of the attention prediction map and the true probability distribution of the attention map. The greater the similarity between the two distributions, the lower the relative entropy, and the greater the difference, the greater the relative entropy. The mathematical expression is:

[0086]

[0087] Among them, ε represents a very small regularization coefficient, i represents the i-th pixel, Q represents the probability distribution of the real image, and P represents the probability distribution of the predicted image;

[0088] Preferably, the Pearson correlation coefficient described in this embodiment represents the degree of linear correlation between two attention maps, and the result is usually between -1 and 1. When using this metric, the continuous distribution of the importance prediction result P and the authenticity of human eye attention Q is considered to be a random variable, and its mathematical expression is:

[0089]

[0090] Among them, CoV(,) represents the covariance,

[0091] Preferably, the similarity coefficient described in this embodiment is used to evaluate the degree of similarity between two saliency maps. The continuous distribution of the saliency prediction result P and the true value of human eye attention Q is regarded as a probability distribution. The closer it is to 1, the more accurate the prediction result is and the closer it is to the true value. Its mathematical expression is:

[0092] SIM(P,Q)=∑ i min(P' i ,Q' I ),∑ i P' i =1,∑ i Q' i =1;

[0093] Preferably, the standardized scan path saliency described in this embodiment is defined as the average value of the normalized saliency (average value is 0 and normalized standard deviation) at the human eye focus position. The ability of the attention fusion module to predict the driver's gaze point position can be evaluated by judging the saliency probability value of each gaze point position to determine the quality of the model prediction result. Its mathematical expression is:

[0094]

[0095] Where N represents the number of gaze points of all drivers, σ and μ are the mean and standard deviation of the prediction results;

[0096] Preferably, the area under the ROC curve described in this embodiment is constructed by plotting the false positive probability (FPR) on the horizontal axis and the true positive probability (TPR) on the vertical axis. The significance detection result P can be binarized to obtain the ROC curve by sliding the threshold on [0, 1] under the ROC curve. When a smaller threshold is used, the overall similarity between the two probability distributions can be calculated; conversely, when a larger threshold is set, the similarity of the two distributions at the peak can be determined, and the AUC index can be calculated by the ROC curve. The larger the AUC value, the better the algorithm performance. When the AUC value is close to 1, it indicates that the significance estimate is completely consistent with the true value calibration. According to the definition of the ROC curve, the AUC index is mainly affected by the high threshold, and its mathematical expression is:

[0097]

[0098] Preferably, after obtaining the hidden driver attention map H in step 4 described in this embodiment, it is input into the attention map decoding module to generate the final driver attention map of the (T+1)th frame.

[0099] Verification test

[0100] First, we divided the video dataset into three parts: training, validation, and test. During training, we used a learning rate of 0.0001 and the Adam optimization algorithm for parameter optimization, with a momentum decay coefficient of 0.9 and a weight decay factor of 10⁻4 to prevent overfitting. To accelerate training, the video frames were uniformly resized to 320×192 pixels.

[0101] The SCFF-Net model is implemented on the TensorFlow framework, enabling a full end-to-end training process from input to output. Furthermore, the model is trained on a computing platform equipped with a high-performance NVIDIA RTX 3060Ti GPU, significantly improving training efficiency and computing power.

[0102] The present invention trains the SCFF-Net network model, selects the optimal model file to test the test set, and finally visualizes the predicted binary image and superimposes it on the original image (such as Figure 3-6 for easy observation.

[0103] Depend on Figure 3-6 It can be seen that Figure 3-6 The first column is the original image, the second column is the driver's eye movement data diagram, and the third column is the driver's attention diagram predicted by the attention prediction model of the present invention.

[0104] pass Figure 3-6 It can be seen from the comparison that the present invention can more accurately predict the driver's line of sight. It can be observed from the prediction graph that when a pedestrian crosses the road (such as Figure 3 As shown in the figure, the driver's attention is often scattered in a tail shape because pedestrians often cross the road at a relatively fast speed. The driver needs to observe a large area in front of his sight to avoid collision. When the driver collides with a pedestrian (such as Figure 4 The model can predict factors that may hinder the driver's attention (such as Figure 5 Driver assistance systems can use this information to alert the driver to potential hazards and help remove obstacles in time to ensure driving safety.

[0105] At the same time, our prediction model is able to detect bottom-up driving-related information, such as pedestrian information and nearby vehicles (such as Figure 6 (a) and (b) in the figure), as well as important information from top to bottom, such as the left front of the road (such as Figure 6 As shown in row (c) of the figure, our model's predictions are highly correlated with human eye tracking data. This indicates that the model accurately predicts the driver's attention distribution. Essentially, the model provides insight into and predicts the driver's focus on key targets and areas, while also including estimates of attention on less important points of focus. These predictions are highly consistent with drivers' actual driving experience and attention allocation patterns, fully demonstrating the potential and practicality of this research model in promoting road safety.

[0106] Figure 7 For comparison of experimental images (the first column is the original image, the second column is the driver's eye movement data map, the third column is the driver's attention map predicted by the attention prediction model in this paper, and the fourth column is the driver's attention map predicted by the SENet network).

[0107] The present invention also compares the deep learning-based saliency model, the Squeeze and Excitation Network (SENet). The SENet model was initially trained on the DADA-2000 dataset and then used to predict driver attention. The prediction results of the SENet model were compared with the attention model proposed in the present invention. Figure 7 (The first column is the original image, the second column is the driver's eye movement data map, the third column is the driver's attention map predicted by the attention prediction model in this paper, and the fourth column is the driver's attention map predicted by the SENet network). Figure 7 It can be seen that the prediction range of the SENet model is significantly larger than the prediction range of the model in this paper (such as Figure 7 (a), (b), (d) and (e) in Figure 3). In comparison, the prediction results of the present invention are more consistent with the real driver's eye movement map. Therefore, the SENet model's interpretation of attention is not as accurate as the present invention. At the same time, the driver's attention map predicted by the SENet model will have branches (such as Figure 7 (c) in the ), and by Figure 7 In the prediction graph in row (f) of Figure 3, we can see that the area predicted by the SENet model is biased to the right, which does not match the actual driver's attention distribution. Therefore, the prediction ability of this model is better than the existing advanced attention model in this field, namely the SENet model.

[0108] At the same time, the evaluation parameters are compared, and the results are shown in Table 1.

[0109]

[0110] Table 1

[0111] As can be seen from the data in Table 1, network evaluation parameters show that the integration of the AC-Mix module in this invention significantly improves both the number of floating-point operations and the network training time, but the increase in network parameters is minimal. Therefore, the overall computational efficiency of this invention is significantly enhanced compared to the classic model, while improving model performance without significantly increasing the computational burden. In summary, this invention offers significant advantages in computational efficiency and reliability, significantly improving upon existing state-of-the-art methods in this field.

[0112] The various significant performance indicators in the present invention are statistically analyzed, and the results are shown in Table 2.

[0113]

[0114] Table 2

[0115] As can be seen from the data in Table 2, all parameters of the present invention are within reasonable ranges and demonstrate good performance. This demonstrates the rationality and effectiveness of the present invention in both feature extraction and information transfer driven by the attention mechanism.

[0116] The experimental results are divided into four cases: "crossing", "collision", "impact" and "attention distribution". Then, the comparative results of driver attention prediction between the network architecture of this paper and the SENet network architecture are presented. Finally, the usefulness of the evaluation parameters and significance indicators of the present invention are demonstrated.

[0117] In summary, this invention cleverly integrates the visual features of RGB images in driving scenes with deep semantic contextual information, accurately identifying and focusing on key elements and areas that influence driver attention. This enables highly accurate prediction of driver attention distribution in complex and ever-changing driving environments. This invention not only deepens our understanding of driver attention mechanisms but also provides a novel technical approach for improving the intelligence and safety of driver assistance systems.

[0118] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A fast network-based driver attention prediction method for dangerous driving scenarios, characterized by: The steps include: Step 1: Segment the RGB set of the traffic accident video set into semantic images with different semantic features frame by frame, and extract their spatiotemporal features and semantic features; Step 2: Use the attention strategy to fuse the spatiotemporal features of the image extracted in step 1 with the semantic context features; Step 3: Use the attention fusion module built with the AC-mix module to quickly identify key objects or areas that attract the driver's attention; Step 4: Use the attention map decoding module to convert the potential driver attention map obtained in step 3 into the final driver attention map.

2. The method for rapidly fusion-based driver attention prediction in dangerous driving scenarios according to claim 1 is characterized by: In the step 1, the DeeplabV3 semantic segmentation method is used to obtain the semantic image of the video clip; at the same time, the size of the input continuous frames is adjusted to 256×192, and the 3D convolution in the multi-path 3D coding architecture is used to extract the spatiotemporal information of the video frames.

3. The method for rapidly fusion-based driver attention prediction in dangerous driving scenarios according to claim 2 is characterized by: In the step 1, after normalizing the semantic graph in the multi-path 3D coding architecture, the relationship between different semantic categories of the scene is modeled through a graph convolutional network to extract the semantic context features of the driving scene.

4. The method for rapidly fusion-based driver attention prediction in dangerous driving scenarios according to claim 1 is characterized by: In the step 2, after obtaining the spatiotemporal features and semantic context features of the traffic accident video set, a convolutional long short-term memory network is used to learn and transfer the fusion details in T consecutive frames to the T+1 frame.

5. The method for rapidly fusion-based driver attention prediction in dangerous driving scenarios according to claim 1 is characterized by: In step 3, when fusing the details within consecutive T frames, AC-mix utilizes the characteristic that depthwise separable convolution can replace the less efficient tensor shift operation. By optimizing the calculation path and reducing the number of repeated calculations, the key objects or areas that attract the driver's attention can be quickly identified.

6. The method for rapidly fusion-based driver attention prediction in dangerous driving scenarios according to claim 1 is characterized by: described. Step 1: Split the selected dataset into 3:1:1 ratio for training, validation, and testing respectively; Step 2: Use the DeeplabV3 semantic segmentation method to obtain the semantic image of the video clip. DeeplabV3 uses dilated convolution and ASPP structure to improve the ability to segment objects of different scales; Step 3: 3D convolution in the multi-path 3D coding architecture extracts the spatiotemporal information of the video frame. Each path of the multi-path 3D coding architecture has the same structure, consisting of three interleaved blocks, namely 3D convolution blocks, 3D batch normalization + rectified linear unit blocks, and 3D maximum pooling blocks. The 3D batch normalization + rectified linear unit blocks are used to accelerate the convergence of network training and resist gradient disappearance. Each path of the multi-path 3D coding architecture has a total of 10 3D convolution blocks, 10 3D batch normalization + rectified linear unit blocks, and 3 3D maximum pooling blocks; Step 4: After steps 2 and 3, the spatiotemporal features and semantic context features of the traffic accident video set are obtained. The convolutional long short-term memory network is used to learn and transfer the fusion details in consecutive T frames to the T+1 frame, using the memory unit C t and hidden state unit H t The time t is used to control the memory update and output H t , and sequentially transmits spatiotemporal scene features. The convolutional long short-term memory network module realizes the transition of the fusion details of T consecutive frames to t+1, which is defined as: [C t+1 ,H t+1 ]=ConvLSTM(Z 1:t ,W c ) Among them, Z 1:t is the input sequence before time t, W c is the weight parameter to be learned; Step 5: When fusing details within consecutive T frames, AC-mix uses depthwise separable convolutions instead of less efficient tensor shift operations. This optimizes the computational path and reduces repeated computations, enabling rapid identification of key objects or areas that attract the driver's attention. Step 6: Use the attention map decoding module to convert the potential driver attention map obtained in step 3 into the final driver attention map; Step 7: Measure the performance of the network in this study. Relative entropy, Pearson correlation coefficient, similarity coefficient, standardized scan path significance, and area under the ROC curve are selected as significance evaluation indicators.

7. The method for rapidly fusion-based driver attention prediction in dangerous driving scenarios according to claim 6 is characterized by: The extraction of the semantic context features of the driving scene is obtained by modeling the relationship between different semantic categories of the scene through a graph convolutional network. The extraction operation of the semantic context features of the driving scene includes the following steps: Step (1): Frame image construction based on the features of each semantic image, semantic features Among them, H, W, and C represent The height, width and number of channels of We reformulate the semantic features of the t-th frame as a matrix Where N = H × W represents the number of nodes in the frame, which is the measurement The pairwise node similarity within is defined as: in, and express The linear transformation of is the node similarity matrix, and then Normalize to get the correlation matrix Correlation Matrix is constructed as a frame graph; Step (2): After obtaining the frame graph, use graph convolution to calculate the relationship between nodes, which is defined as: in, represents the weight of the graph convolution layer, is the output of each graph convolution layer, for each Perform graph convolution calculation to obtain the semantic context features of T frames In the multi-path 3D coding architecture, H, W, and C are 32, 24, and 512, respectively.

8. The method for rapidly fusion-based driver attention prediction in dangerous driving scenarios according to claim 6 is characterized by: After obtaining the hidden driver attention map in step 6, it is input into the attention map decoding module to generate the final driver attention map of the (T+1)th frame. The attention map decoding module is implemented as upsampling (×4) → convolution (3×3,128) → batch normalization + rectification linearization → upsampling (×2) → convolution (3×3,1) → Sigmold, where the Sigmoid function is used to limit the output value of the driver attention map to [0,1].

9. The method for rapidly fusion-based driver attention prediction in dangerous driving scenarios according to claim 6 is characterized by: The relative entropy is also called Kullback-Leibler divergence, which can be used to determine the degree of difference between the probability distribution of the attention prediction map and the true probability distribution of the attention map. The greater the similarity between the two distributions, the lower the relative entropy, and the greater the difference, the greater the relative entropy. The mathematical expression is: Among them, ε represents a very small regularization coefficient, i represents the i-th pixel, Q represents the probability distribution of the real image, and P represents the probability distribution of the predicted image; The Pearson correlation coefficient represents the degree of linear correlation between two attention maps. The result is usually between -1 and 1. When using this metric, the continuous distribution of the importance prediction result P and the authenticity of human eye attention Q is considered to be a random variable. Its mathematical expression is: Among them, CoV(,) represents the covariance, The similarity coefficient is used to evaluate the similarity between two saliency maps. The continuous distribution of the saliency prediction result P and the true value of human eye attention Q is regarded as a probability distribution. The closer it is to 1, the more accurate the prediction result is and the closer it is to the true value. Its mathematical expression is: SIM(P,Q)=∑ i min(P' i ,Q' I ),∑ i P' i =1,∑ i Q' i =1; The standardized scan path saliency is defined as the average of the normalized saliency at the human eye focus position. The ability of the attention fusion module to predict the driver's gaze point position can be evaluated by determining the saliency probability value of each gaze point position to determine the quality of the model prediction result. Its mathematical expression is: Where N represents the number of gaze points of all drivers, σ and μ are the mean and standard deviation of the prediction results; The area under the ROC curve is constructed by plotting the false positive probability on the horizontal axis and the true positive probability on the vertical axis. The significance test result P can be binarized to obtain the ROC curve by sliding the threshold on [0, 1] under the ROC curve. When a smaller threshold is used, the overall similarity between the two probability distributions can be calculated; conversely, when a larger threshold is set, the similarity between the two distributions at the peak can be determined, and the AUC index can be calculated by the ROC curve. The larger the AUC value, the better the algorithm performance. When the AUC value is close to 1, it indicates that the significance estimate is completely consistent with the true value calibration. According to the definition of the ROC curve, the AUC index is mainly affected by the high threshold, and its mathematical expression is:

10. The method for rapidly fusion-based driver attention prediction in dangerous driving scenarios according to claim 1 is characterized by: After obtaining the hidden driver attention map H in step 4, it is input into the attention map decoding module to generate the final driver attention map of the (T+1)th frame.

Citation Information

Cited By

  • Conditional automatic driving takeover prompting method and system

    CN122244842A