A method, system, device and medium for predicting driver visual attention
By combining convolutional neural network and Transformer's attention prediction network, the problem of low visual attention prediction accuracy of drivers in the prior art is solved, and more accurate driver attention distribution analysis is achieved, improving the safety and reliability of advanced driving assistance and unmanned driving technology.
Patent Information
- Application Number
- CN202111258136.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-27
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2041-10-27
AI Technical Summary
The deep learning method based on convolution in the prior art has limited receptive field in driver visual attention prediction, low prediction results accuracy, and it is difficult to accurately analyze and predict driver attention allocation.
Combining the convolutional neural network and the Transformer's attention prediction network, by obtaining driving scene images and training data, preprocessing, training the model, and outputting color visual attention prediction result diagrams, combining local and global information.
It significantly improves the accuracy and accuracy of driver visual attention prediction, can better analyze the driver's attention distribution when performing driving tasks, and provides important reference information for advanced driving assistance systems and unmanned driving technology.
Smart Images

Figure CN114092900B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a method, system, device and medium for predicting driver visual attention. Background Art
[0002] With the development of algorithms, computing power, data, and other related technologies, the field of automotive driving technology has ushered in a wave of intelligence, with the emergence of advanced driver assistance systems and autonomous driving technologies. Among these, methods for analyzing driver attention, a key area of research, have seen rapid progress. In driving scenarios, drivers allocate their attention based on both top-down driving task dynamics and bottom-up driving environment stimuli. Research has shown that the primary cause of traffic accidents is driver inattention. Therefore, studying, analyzing, and predicting driver attention allocation is of great significance to the safety of advanced driver assistance systems and autonomous driving technologies. Existing technologies primarily use convolution-based deep learning methods to predict driver visual attention. However, these methods suffer from limited receptive fields and low prediction accuracy. Summary of the invention
[0003] In view of this, embodiments of the present invention provide a method, system, device, and medium for predicting driver visual attention, so as to achieve simple and accurate prediction of driver visual attention.
[0004] In one aspect, the present invention discloses a method for predicting driver visual attention, comprising:
[0005] Obtain initial images and training data for driving scenarios;
[0006] Performing a preprocessing operation on the initial image to obtain a preprocessed image;
[0007] Inputting the training data into an attention prediction network combining a convolutional neural network and a Transformer to train an attention prediction model;
[0008] Inputting the preprocessed image into the attention prediction model, and outputting an attention prediction probability map;
[0009] The attention prediction probability map is subjected to color visualization processing to obtain an attention prediction result map.
[0010] Optionally, performing a preprocessing operation on the initial image to obtain a preprocessed image includes:
[0011] Performing image scaling processing on the initial image to obtain a first processed image;
[0012] Perform image random flipping on the first processed image to obtain a second processed image;
[0013] Perform image normalization on the second processed image to obtain a preprocessed image.
[0014] Optionally, the step of inputting the training data into the attention prediction network combining a convolutional neural network and a Transformer to train an attention prediction model includes:
[0015] Shuffle the training data and input it into the attention prediction model to be trained, and use the backpropagation algorithm to train the attention prediction network to obtain a trained attention prediction model. The attention prediction model is trained using cross-entropy as the loss function.
[0016] Optionally, the step of inputting the preprocessed image into the attention prediction model and outputting an attention prediction probability map includes:
[0017] Input the preprocessed image into an initial convolutional layer for feature extraction to obtain initial features;
[0018] Input the initial features into a Transformer encoder to obtain global information and output initial global information features;
[0019] Input the initial global information features into a hybrid feature extraction module to output hybrid feature information;
[0020] Input the hybrid feature information into a decoder to output an attention prediction probability map.
[0021] Optionally, the step of performing color visualization on the attention prediction probability map to obtain an attention prediction result map includes:
[0022] Enlarge the pixel range of the attention prediction probability map to obtain a pixel image;
[0023] Perform pseudo-color processing on the pixel image to obtain an attention prediction result map.
[0024] Optionally, the step of inputting the hybrid feature information into a decoder to output an attention prediction probability map includes:
[0025] Input the hybrid feature information into a decoder, where the decoder includes a convolutional layer, an upsampling layer, and an output layer;
[0026] Alternately process the hybrid feature information through the convolutional layer and the upsampling layer, and output an attention prediction probability map through the output layer.
[0027] Optionally, the method further includes:
[0028] Select evaluation indicators based on probability distribution to evaluate the attention prediction probability map, and obtain an evaluation result.
[0029] On the other hand, an embodiment of the present invention also discloses a driver's visual attention prediction system, including:
[0030] A first module for obtaining an initial image and training data in a driving scenario;
[0031] A second module for performing preprocessing operations on the initial image to obtain a preprocessed image;
[0032] A third module for inputting the training data into an attention prediction network that combines a convolutional neural network and a Transformer, and training to obtain an attention prediction model;
[0033] A fourth module for inputting the preprocessed image into the attention prediction model and outputting an attention prediction probability map;
[0034] A fifth module for performing color visualization processing on the attention prediction probability map to obtain an attention prediction result map.
[0035] On the other hand, an embodiment of the present invention also discloses an electronic device, including a processor and a memory;
[0036] The memory is used to store programs;
[0037] The processor executes the program to implement the method described above.
[0038] On the other hand, an embodiment of the present invention also discloses a computer-readable storage medium, and the storage medium stores a program, and the program is executed by a processor to implement the method described above.
[0039] On the other hand, an embodiment of the present invention also discloses a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method described above.
[0040] Compared with the prior art, the present invention adopting the above technical solution has the following technical effects: The present invention obtains an initial image and training data in a driving scenario; performs preprocessing operations on the initial image to obtain a preprocessed image; inputs the training data into an attention prediction network combining a convolutional neural network and a Transformer to train an attention prediction model; inputs the preprocessed image into the attention prediction model, and outputs an attention prediction probability map; performs color visualization processing on the attention prediction probability map to obtain an attention prediction result map; and can well obtain and combine local and global information, and accurately predict the visual attention distribution of a driver when performing a driving task. Description of the Drawings
[0041] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0042] Figure 1 It is a flowchart of a method for predicting a driver's visual attention in an embodiment of the present invention;
[0043] Figure 2 It is a structural diagram of an attention prediction model in an embodiment of the present invention;
[0044] Figure 3 It is a structural diagram of an encoder of an attention prediction model in an embodiment of the present invention;
[0045] Figure 4 It is a comparison diagram of evaluation results in an embodiment of the present invention. Detailed Embodiments
[0046] In order to make the purpose, technical solutions and advantages of the present application clearer, the following further details the present application in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0047] Refer to Figure 1 , an embodiment of the present invention provides a method for predicting a driver's visual attention, including:
[0048] S1. Obtain an initial image and training data in a driving scenario;
[0049] S2. Perform preprocessing operations on the initial image to obtain a preprocessed image;
[0050] S3. Input the training data into the attention prediction network that combines a convolutional neural network and a Transformer, and train to obtain an attention prediction model;
[0051] S4. Input the preprocessed image into the attention prediction model, and output an attention prediction probability map;
[0052] S5. Perform color visualization processing on the attention prediction probability map to obtain an attention prediction result map.
[0053] Further as a preferred embodiment, in the above step S2, the preprocessing operation on the initial image to obtain a preprocessed image includes:
[0054] Perform image scaling processing on the initial image to obtain a first processed image;
[0055] Perform random image flipping processing on the first processed image to obtain a second processed image;
[0056] Perform image normalization processing on the second processed image to obtain a preprocessed image.
[0057] Among them, in the embodiment of the present invention, an initial image is obtained by shooting an RGB original image in a driving scene from the first-person perspective. The first-person perspective can be a driving recorder or other imaging instruments. Then, image scaling processing is performed to scale the obtained initial image to a first processed image with a pixel size of 256 pixels in length and 256 pixels in width, so as to reduce the calculation amount and meet the subsequent network input requirements. The first processed image is randomly flipped with a probability of 50% to obtain a second processed image, so as to alleviate the center bias problem. The center bias problem can be specifically divided into the center bias problem of the first-person driving scene dataset and the center bias problem of the driver attention prediction method. Their relationship is: the center bias problem of the first-person driving scene dataset will lead to the center bias problem of the driver attention prediction method. The center bias problem of the first-person driving scene dataset refers to that the vast majority of visual attention data on the dataset are concentrated near the center point of the image, which is mainly caused by the following two reasons: (1) In the driving scene, the driving recorder often takes the vanishing point of the road ahead as the center of the captured image; (2) In most normal driving scenes, the driver will concentrate their visual attention on the vanishing point of the road ahead. The mainstream open-source datasets in the current driving scene generally have the center bias problem. The center bias problem of the driver attention prediction method refers to that most of the visual attention results predicted by the method are always concentrated near the center point of the image. The reason for this phenomenon is that the driver attention prediction method uses the first-person driving scene dataset for network training. Due to the inherent distribution of the dataset, most of the prediction results of the driver attention prediction method will also tend to be concentrated near the center point of the image. Finally, the pixel values of the second processed image are subtracted by the pixel average value of the initial image set and divided by the pixel variance of the initial image set for normalization to obtain a preprocessed image. By performing preprocessing operations on the initial image, data enhancement of the driving scene image can be achieved.
[0058] Further as a preferred embodiment, in the above step S3, the inputting the training data into the attention prediction network combining a convolutional neural network and a Transformer to train an attention prediction model includes:
[0059] The training data is shuffled and then input into the attention prediction model to be trained, and the attention prediction network is trained using the backpropagation algorithm to obtain a trained attention prediction model. The attention prediction model is trained using cross-entropy as the loss function.
[0060] Among them, in the embodiments of the present invention, the BDD-A dataset and the Driving-diverse dataset, which are open-source in the industry, are used as the training data of the model. To avoid the adverse impact of the original arrangement order of the dataset on model training, the training set in the dataset can be randomly shuffled and then input into the model for training. The attention prediction network uses cross-entropy as the loss function. Compared with the basic mean square error, using cross-entropy as the loss function can obtain a larger error value to accelerate the training of the model. The backpropagation algorithm is used in the training process, and the Adam optimizer is selected. The learning rate is set to 1×10^-4 to update and iterate the weights of the model. After training 100 times, the weights of the attention prediction network are saved, and the attention prediction model is obtained through training. The advantage of using the Adam optimizer is that it can dynamically change the learning rate for different parameters. The calculation formula of the cross-entropy loss function H(p,q) is as follows:
[0061]
[0062] In the formula, p represents the probability distribution of the model prediction result, q represents the probability distribution of the true label, n represents the total number of samples, x i represents the i-th sample, and i represents a positive integer.
[0063] Further as a preferred implementation manner, in the above step S4, the step of inputting the preprocessed image into the attention prediction model and outputting an attention prediction probability map includes:
[0064] Input the preprocessed image into the initialization convolutional layer for feature extraction to obtain initialization features;
[0065] Input the initialization features into the Transformer encoder to obtain global information and output initialization global information features;
[0066] Input the initialization global information features into the hybrid feature extraction module and output hybrid feature information;
[0067] Input the hybrid feature information into the decoder and output the attention prediction probability map.
[0068] Among them, referring to Figure 2 , Figure 2The attention prediction model is a structural diagram of an attention prediction model based on Transformer and CNN, which can be divided into two major parts: an encoder and a decoder. The encoder consists of 1 initialized 7*7 convolutional layer, 3 3*3 convolutional layers, and 4 pure Transformer encoder structures. The decoder consists of 4 upsampling layers and 4 convolutional layers. It should be noted that the encoder is designed with a residual structure, which can reduce the probability of gradient explosion and gradient dispersion, ensuring that information is not lost as the network deepens. In the embodiment of the present invention, a preprocessed image with a size of 256*256 pixels is input into the initialized convolutional layer in the attention prediction model for preliminary feature extraction, obtaining initialized features with a dimension of 64*64*64. The initialized features are input into the Transformer encoder to preliminarily obtain global information and output initialized global information features with a dimension of 64*64*64. The initialized global information features are input into a hybrid feature extraction module composed of an encoding convolutional layer and a Transformer encoder. This hybrid feature extraction model needs to be repeated three times, and finally outputs hybrid feature information with a dimension of 16*16*512. The finally output hybrid feature information with a dimension of 16*16*512 is input into the decoder of the attention prediction model. Through the continuous alternation of the decoding convolutional layer and the upsampling layer, the dimensionality reduction of the feature image size and depth is achieved, and finally, the driver attention prediction probability map is output through the output layer. Refer to Figure 3 , the initialized convolutional layer consists of a convolutional operation with a size of 7*7, batch normalization, a ReLU activation function, and a max pooling operation with a size of 2*2. The Transformer encoder consists of an image serialization processing module, 4 Transformer layers, and a deformation layer. The image serialization processing module clips the initialized features into multiple image patches through a linear projection function, and then converts each image patch into a one-dimensional sequence. In the image serialization processing, the relative position information between pixels is actually lost, so a spatial position encoding p (specific embedding) needs to be added to each image patch e to form the final one-dimensional sequence E as shown in the following formula:
[0069] E = {e1 + p1, e2 + p2, …, e N + p N};
[0070] In the formula, N represents the number of image patches.
[0071] The Transformer layer normalizes the weights through layer normalization and performs self-attention calculations on the one-dimensional sequence E in parallel. The self-attention calculation consists of the calculation of three parts of the one-dimensional sequence E: Q (query, query), K (key, key), and V (value, value). The input of the l-th layer is the output of the l-1-th layer. The calculation formulas for Q, K, and V in the l-th layer are:
[0072] Q l = EW Q ;
[0073] K l = EW K ;
[0074] V l = EW V ;
[0075] Wherein, W Q ∈R C×d , W K ∈R C×d , W K ∈R C×d respectively represent three learnable parameter matrices of the linear layer, C represents the number of channels of the image, and d represents the dimension.
[0076] After calculating Q l , K l , V l , the calculation formula for the self-attention value SA(E) of this layer is:
[0077]
[0078] Wherein, represents the weight value, represents the dimension of K l .
[0079] The self-attention values of this layer are concatenated to obtain the multi-head self-attention value. The multi-head attention can obtain global features more comprehensively and improve the performance of the attention prediction model. The multi-head attention value is subjected to layer normalization processing and sent to a multi-layer perceptron MLP for calculation. The multi-head attention result and the multi-layer perceptron result are added as the output of the Transformer layer. The deformation layer rearranges the one-dimensional sequence E output by the Transformer layer into a two-dimensional sequence M in the original pixel order. If the total length of the one-dimensional sequence E is D, the length H and width W of the two-dimensional sequence M have the following relationship:
[0080] D = H × W;
[0081] H = W.
[0082] Further as a preferred embodiment, in the above step S5, the color visualization process of the attention prediction probability map to obtain the attention prediction result map includes:
[0083] Enlarge the pixel range of the attention prediction probability map to obtain a pixel image;
[0084] Perform pseudo-color processing on the pixel image to obtain an attention prediction result map.
[0085] Among them, in the embodiment of the present invention, the pixel values of the attention prediction result map are multiplied by 255 to expand the pixel value range to [0, 255], and pseudo-color operations are performed through the open-source image processing library OpenCV to obtain the driver attention prediction result map. Optional pseudo-color operation modes include JET, SPRING, COOL, HSV, PINK, HOT, etc.
[0086] Further as a preferred embodiment, the inputting the mixed feature information into the decoder to output an attention prediction probability map includes:
[0087] Input the mixed feature information into the decoder, and the decoder includes a convolutional layer, an upsampling layer, and an output layer;
[0088] Perform alternating processing on the mixed feature information through the convolutional layer and the upsampling layer, and output an attention prediction probability map through the output layer.
[0089] Among them, the first eight layers of the decoder are continuously alternated through 4 convolutional layers and 4 upsampling layers to achieve dimensionality reduction of the image size and depth, and the output dimension is 256*256*16. Among them, the convolutional layer consists of a convolutional operation with a size of 3*3 and a stride of 1, batch normalization, and a ReLU activation function. The upsampling layer is a bilinear interpolation upsampling operation (bilinear sample) with a size of 2*2 and a stride of 1. The ninth layer of the decoder is the output layer, which consists of a convolutional operation with a size of 3*3 and a stride of 1, batch normalization, and a Sigmoid activation function.
[0090] Further as a preferred embodiment, the method further includes:
[0091] Select an evaluation index based on the probability distribution to evaluate the attention prediction probability map to obtain an evaluation result.
[0092] Among them, the attention prediction probability map is a grayscale map with a size of 256*256 pixels. The value of each pixel is between [0,1], which represents the probability value of each pixel being noticed. The higher the value, the more likely the pixel is to be noticed and the more likely it is an important area affecting the execution of driving tasks. For example, when there are pedestrians crossing the zebra crossing, cars suddenly changing lanes, traffic signs such as traffic lights, etc. on the current road surface, the pixel values corresponding to these areas will be significantly higher than those of other areas to indicate their importance for the execution of driving tasks. The attention prediction probability map can provide important references for the advanced driver assistance system to assist the driver in driving safely. When there are multiple significant areas in the attention prediction probability map or significant changes occur in the context, the system will give a reminder. By detecting the important areas affecting the execution of driving tasks, the attention prediction probability map can provide important references for the unmanned driving technology to improve the perception ability and decision-making accuracy of the unmanned driving technology. The attention prediction result map is a three-channel RGB color map with a size of 256*256, which can be displayed as the visualization result of the unmanned driving technology decision-making to improve the interpretability of the unmanned driving technology. In the embodiments of the present invention, two evaluation indicators based on probability distribution are selected: Kullback–Leibler divergence (KLD) and Pearson’s Correlation Coefficient (CC). At the same time, the test sets of the mainstream BDD-A dataset and the Driving-diverse dataset are selected for test evaluation. Based on the BDD-A dataset, the KLD of the method of the present invention is 1.1 and the correlation coefficient is 0.63. Based on the Driving-diverse dataset, the KLD of the method of the present invention is 0.86 and the correlation coefficient is 0.69.
[0093] Referring to Figure 4 , Figure 4 is a comparison chart of the evaluation results of the embodiments of the present invention, as Figure 4As shown in the figure, the scenario is a typical crossroads. Such a complex intersection is challenging for both humans and machines. Many human drivers often cause traffic accidents due to lack of proper concentration. In this scenario, a red light appears ahead, and at the same time, a bicycle approaches from the right side of the road. As a result, the human driver decelerates and stops the vehicle. By comparing the visual attention of the human driver and the present invention, it is not difficult to find that the visual attention predicted by the present invention is more concentratedly distributed at important positions such as the red light ahead, the approaching bicycle, and the vanishing point of the road ahead, which affect safe driving. However, the visual attention of humans is more dispersed, and some of the attention is not concentrated on the above important positions. This may be because the human driver thinks the vehicle speed is slow and the state is safe during the deceleration and stop process, so he is distracted. In addition, for the bicycle approaching from the right side, the method of the present invention has noticed the bicycle since the first frame and continuously allocates visual attention according to the position of the bicycle. When the bicycle gets closer to the vehicle, more attention is allocated, and when the bicycle moves away from the vehicle, the attention allocation decreases. On the contrary, the visual attention of the human driver does not have such an obvious process of allocation change. In summary, it can be considered that compared with human drivers, the method of the present invention can allocate visual attention more reasonably and correctly in some complex crossroads scenarios.
[0094] The process of the present invention specifically includes: The embodiment of the present invention takes the initial image in the driving scenario from the first-person perspective and uses the test set of the data set as the training data. By performing image scaling, random image flipping, and image normalization on the initial image, the initial image is enhanced to obtain a preprocessed image. The training data is input into the attention prediction network that combines a convolutional neural network and a Transformer. By combining the shuffled training data with the backpropagation algorithm for training, an attention prediction model is obtained. The attention prediction model is a model based on a convolutional neural network and a Transformer network, which can better obtain and combine local and global information. By inputting the preprocessed image into the attention prediction model, an attention prediction probability map is output; the attention prediction probability map is subjected to color visualization processing to obtain an attention prediction result map.
[0095] The embodiment of the present invention also discloses a driver's visual attention prediction system, including:
[0096] The first module is used to obtain the initial image and training data in the driving scenario;
[0097] The second module is used to perform preprocessing operations on the initial image to obtain a preprocessed image;
[0098] The third module is configured to input the training data into an attention prediction network that combines a convolutional neural network and a Transformer, and train an attention prediction model;
[0099] The fourth module is configured to input the preprocessed image into the attention prediction model and output an attention prediction probability map;
[0100] The fifth module is configured to perform color visualization processing on the attention prediction probability map to obtain an attention prediction result map.
[0101] Corresponding to Figure 1 the method of, an embodiment of the present invention further provides an electronic device, including a processor and a memory; the memory is used to store a program; the processor executes the program to implement the method as described above.
[0102] Corresponding to Figure 1 the method of, an embodiment of the present invention further provides a computer-readable storage medium, where the storage medium stores a program, and the program is executed by a processor to implement the method as described above.
[0103] An embodiment of the present invention also discloses a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes Figure 1 the method shown.
[0104] In summary, the embodiments of the present invention have the following advantages: The visual attention prediction method proposed by the present invention effectively combines a Transformer network and a convolutional neural network. Compared with existing methods, it significantly improves the model's ability to obtain image information of driving scenarios, strengthens the model's ability to analyze the top-down and bottom-up attention of drivers when performing driving tasks, and can accurately predict the visual attention distribution of drivers when performing driving tasks.
[0105] In some alternative embodiments, the functions / operations mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the functions / operations involved, two consecutive blocks shown can actually be executed substantially simultaneously or the blocks can sometimes be executed in the reverse order. In addition, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are expected, in which the order of various operations is changed and the sub-operations described as part of a larger operation are executed independently.
[0106] In addition, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the functions and / or features described may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for an understanding of the present invention. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skills of an engineer. Thus, those skilled in the art can implement the present invention as set forth in the claims without undue experimentation. It should also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.
[0107] If the described functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0108] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with such instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0109] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.
[0110] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, the multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0111] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0112] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.
[0113] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.
Claims
1. A method for predicting a driver's visual attention, characterized in that Including: Obtain an initial image and training data in a driving scenario; Perform preprocessing operations on the initial image to obtain a preprocessed image; Input the training data into an attention prediction network combining a convolutional neural network and a Transformer, and train to obtain an attention prediction model; Input the preprocessed image into the attention prediction model and output an attention prediction probability map; Perform color visualization processing on the attention prediction probability map to obtain an attention prediction result map; The step of inputting the preprocessed image into the attention prediction model and outputting an attention prediction probability map includes: Input the preprocessed image into an initial convolutional layer for feature extraction to obtain initial features; Input the initial features into a Transformer encoder to obtain global information and output initial global information features; Input the initial global information features into a hybrid feature extraction module and output hybrid feature information; Input the hybrid feature information into a decoder and output an attention prediction probability map; The Transformer encoder includes an image serialization processing module, a Transformer layer, and a deformation layer; The image serialization processing module is used to perform image block shearing processing on the initial features and convert them into an initial one-dimensional sequence; The Transformer layer is used to perform self-attention calculation on the initial one-dimensional sequence to obtain a target one-dimensional sequence; The deformation layer is used to arrange the target one-dimensional sequence in pixel order to obtain the initial global information features.
2. The driver visual attention prediction method according to claim 1, characterized in that, The step of performing preprocessing operations on the initial image to obtain a preprocessed image includes: Perform image scaling processing on the initial image to obtain a first processed image; Perform random image flipping processing on the first processed image to obtain a second processed image; Perform image normalization processing on the second processed image to obtain a preprocessed image.
3. The driver visual attention prediction method according to claim 1, wherein The step of inputting the training data into an attention prediction network combining a convolutional neural network and a Transformer and training to obtain an attention prediction model includes: Shuffle the training data and input it into the attention prediction model to be trained, and use the backpropagation algorithm to train the attention prediction network to obtain a trained attention prediction model. The attention prediction model is trained using cross-entropy as the loss function.
4. A method for predicting a driver's visual attention according to claim 1, characterized in that The step of performing color visualization processing on the attention prediction probability map to obtain an attention prediction result map includes: Enlarge the pixel range of the attention prediction probability map to obtain a pixel image; Perform pseudo-color processing on the pixel image to obtain an attention prediction result map.
5. A method for predicting a driver's visual attention according to claim 1, characterized in that, The step of inputting the hybrid feature information into a decoder and outputting an attention prediction probability map includes: Input the hybrid feature information into a decoder, and the decoder includes a convolutional layer, an upsampling layer, and an output layer; Alternately process the hybrid feature information through the convolutional layer and the upsampling layer, and output an attention prediction probability map through the output layer.
6. The driver visual attention prediction method according to claim 1, characterized in that The method further includes: Select evaluation metrics based on probability distribution to evaluate the attention prediction probability map and obtain an evaluation result.
7. A driver's visual attention prediction system, characterized in that, It includes: The first module is used to obtain the initial image and training data in the driving scenario; The second module is used to perform preprocessing operations on the initial image to obtain a preprocessed image; The third module is used to input the training data into an attention prediction network that combines a convolutional neural network and a Transformer, and train to obtain an attention prediction model; The fourth module is used to input the preprocessed image into the attention prediction model and output an attention prediction probability map; The fifth module is used to perform color visualization processing on the attention prediction probability map to obtain an attention prediction result map; The third module, which is used to input the preprocessed image into the attention prediction model and output an attention prediction probability map, includes: Input the preprocessed image into an initialized convolutional layer for feature extraction to obtain initialized features; Input the initialized features into a Transformer encoder to obtain global information and output initialized global information features; Input the initialized global information features into a hybrid feature extraction module and output hybrid feature information; Input the hybrid feature information into a decoder and output an attention prediction probability map; The Transformer encoder includes an image serialization processing module, a Transformer layer, and a deformation layer; The image serialization processing module is used to perform image block shearing processing on the initialized features and convert them into an initial one-dimensional sequence; The Transformer layer is used to perform self-attention calculation on the initial one-dimensional sequence to obtain a target one-dimensional sequence; The deformation layer is used to arrange the target one-dimensional sequence in pixel order to obtain the initialized global information features.
8. An electronic device, characterized in that, It includes a processor and a memory; The memory is used to store programs; The processor executes the program to implement the method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The storage medium stores a program, and the program is executed by the processor to implement the method according to any one of claims 1-6.
Citation Information
Patent Citations
Driving early warning method based on driver visual attention prediction
CN112699821A