Driver Attention Prediction Method and System Based on On-vehicle Forward-looking Image
By extracting and fusing visual and semantic features in the vehicle front-view image, and combining deep learning algorithms, driver attention maps are generated, which solves the problem of difficult to identify and predict driver attention in the existing technology, and improves the information acquisition ability of the autonomous driving system.
Patent Information
- Application Number
- CN202410846366.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-27
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-06-27
AI Technical Summary
The prior art is difficult to effectively identify and predict driver attention, especially in complex traffic scenarios, affecting the understanding and information acquisition of autonomous driving systems.
Visual features and semantic features are extracted through vehicle front-view images, combined with deep learning algorithms, feature fusion and spatial-temporal sequence prediction are performed, and driver attention maps are generated.
It realizes rapid identification and accurate prediction of drivers' attention, and improves the understanding and information acquisition capabilities of autonomous driving systems for driving scenarios.
Smart Images

Figure CN118609104B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of driver attention prediction, and particularly relates to assisting drivers in observing road conditions and providing expert experience support for obtaining information for autonomous driving. Specifically, it is a method and system for predicting driver attention based on on-vehicle front-view images. Background Art
[0002] In recent years, advanced driver assistance systems such as autonomous driving systems and vehicle-to-everything (V2X) have promoted the development of intelligent vehicles and evolved the driving experience towards being more comfortable, safer, and easier. The autonomous driving system is an important component of intelligent vehicles and has important application prospects in multiple fields such as private cars, public transportation, and logistics transportation. To fully achieve autonomous driving, that is, to reach the fifth level (L5) defined in SAE J3016 "Levels of Driving Automation", it is necessary to effectively identify external stimulus information. Therefore, learning human visual attention, focusing experience, information prediction, and tracking during driving can promote the understanding of driving scenarios and information acquisition by autonomous driving systems. Summary of the Invention
[0003] To overcome the deficiencies of the above-mentioned prior art, the present invention provides a method and system for predicting driver attention based on on-vehicle front-view images. Through the data support provided by on-vehicle front-view images and combined with deep learning algorithms, effective prediction of driver attention is achieved.
[0004] According to one aspect of the specification of the present invention, there is provided a method for predicting driver attention based on on-vehicle front-view images, including:
[0005] Performing visual feature and semantic feature extraction based on the acquired on-vehicle front-view image;
[0006] Performing visual feature and semantic feature fusion based on the extracted visual features and semantic features;
[0007] Performing spatio-temporal sequence prediction and generating driver attention based on the fused features.
[0008] As a further technical solution, performing visual feature extraction based on the acquired on-vehicle front-view image includes:
[0009] Dividing the on-vehicle front-view image into blocks and performing position encoding on each block;
[0010] Calculating the correlation between blocks and obtaining visual features based on the mean value of the calculated correlations.
[0011] As a further technical solution, performing semantic feature extraction based on the acquired on-vehicle front-view image includes:
[0012] Performing semantic segmentation of the on-vehicle front-view image in a traffic environment to obtain a number of semantic images;
[0013] Convert a plurality of the semantic graphs into a plurality of semantic nodes;
[0014] Encode a plurality of the semantic nodes and convert them into a semantic node adjacency matrix, and the semantic node adjacency matrix and the semantic nodes form a semantic graph;
[0015] Encode the semantic graph to obtain semantic features.
[0016] As a further technical solution, based on the extracted visual features and semantic features, perform visual feature and semantic feature fusion, including:
[0017] Perform a Hadamard product on the extracted visual features and semantic features to obtain fused features;
[0018] Use a fully connected neural network to re-extract the fused features in a way of first reducing the dimension and then increasing the dimension to obtain the re-extracted fused features;
[0019] Add the re-extracted fused features to the visual features to obtain the final fused features.
[0020] As a further technical solution, based on the fused features, perform spatio-temporal sequence prediction and driver attention generation, including:
[0021] Based on the fused features, use an LSTM neural network to memorize, reason, and generate the fused features of the predicted frame;
[0022] Decode the fused features of the predicted frame into driver attention, and under the supervision of the neural network, output a driving attention map that meets the accuracy requirements.
[0023] As a further technical solution, the method further includes:
[0024] Supervise the generation of driver attention from the fused features of the predicted frame, and use the real driver attention as the true value of the discriminator, and cycle training until the generation accuracy of the driver attention reaches the preset accuracy requirements.
[0025] According to one aspect of the specification of the present invention, there is provided a driver attention prediction device based on an in-vehicle front view image, including:
[0026] A feature extraction module for extracting visual features and semantic features based on the acquired in-vehicle front view image;
[0027] A feature fusion module for performing visual feature and semantic feature fusion based on the extracted visual features and semantic features;
[0028] A prediction and generation module for performing spatio-temporal sequence prediction and driver attention generation based on the fused features.
[0029] As a further technical solution, the device further includes: a supervision module, configured to supervise the generation of the driver's attention and perform cyclic training until the generation accuracy of the driver's attention reaches a preset requirement.
[0030] According to one aspect of the specification of the present invention, there is provided an electronic device, including: at least one processor, at least one memory, and a communication interface; wherein, the processor, the memory, and the communication interface communicate with each other; the memory stores program instructions executed by the processor, and the processor calls the program instructions to execute the method described above.
[0031] According to one aspect of the specification of the present invention, there is provided a non-transitory computer-readable storage medium, where the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the method described above.
[0032] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0033] In view of the variability of the driver's attention and traffic scenarios, the present invention proposes a method for predicting the driver's attention based on on-vehicle forward-looking images. The method adopts a deep learning framework, extracts visual features such as color and texture and semantic features such as roads, pedestrians, vehicles, and traffic facilities from on-vehicle forward-looking images; obtains fused features by combining visual features and semantic features, and maps the attracting factors of on-vehicle forward-looking images to the driver's attention; memorizes and reasons about the fused features of multiple frames to obtain the fused features of the predicted frame, and decodes the fused features of the predicted frame and uses a generative adversarial neural network to supervise to obtain the driver's attention.
[0034] The present invention realizes the rapid identification of key information in time-series images, which is helpful for designing autonomous driving tasks, optimizing visual scene understanding, and providing decision-making knowledge. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings used in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0036] Figure 1 It is a technical roadmap of the method for predicting the driver's attention based on on-vehicle forward-looking images provided by the embodiment of the present invention;
[0037] Figure 2 It is a schematic diagram of extracting visual features and semantic features of on-vehicle forward-looking images provided by the embodiment of the present invention;
[0038] Figure 3 Schematic diagram of the fusion of visual features and semantic features provided by the embodiments of the present invention;
[0039] Figure 4 Schematic diagram of the prediction of the spatio-temporal sequence of the fusion features and the generation of driver attention provided by the embodiments of the present invention;
[0040] Figure 5 Schematic diagram of the structure of the driver attention prediction device based on on-vehicle front view images provided by the embodiments of the present invention;
[0041] Figure 6 Schematic diagram of the structure of the electronic device provided by the embodiments of the present invention. Specific embodiments
[0042] It should be noted that:
[0043] The on-vehicle front view images have high temporal resolution and spatial resolution, providing data support for the driver attention prediction method of the present invention. With the improvement of computer performance, deep learning algorithms can extract more multi-source information from the images, laying a foundation for the driving attention prediction technology.
[0044] Under the pytorch deep learning framework, the present invention provides a driver attention prediction method based on on-vehicle front view images for different traffic scenarios, including: extracting visual features and semantic features of on-vehicle front view images, fusing visual features and semantic features, predicting the spatio-temporal sequence of the fusion features and generating driver attention. Through the interaction of these steps, the rapid recognition of key information in sequential images is achieved, the accuracy of driver attention generation is improved, which helps to design autonomous driving tasks, optimize visual scene understanding and provide decision-making knowledge.
[0045] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. In addition, the technical features in each embodiment or a single embodiment provided by the present invention can be combined with each other arbitrarily to form a new technical solution. This combination is not restricted by the order of steps and / or the pattern of structural composition, but must be based on what can be achieved by those of ordinary skill in the art. When the combination of technical solutions appears to be contradictory or unable to be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.
[0046] An embodiment of the present invention provides a method for predicting driver attention based on in-vehicle forward-looking images. Refer to Figure 1 . The method includes: extracting image visual features and semantic features, fusing visual features and semantic features, predicting the spatio-temporal sequence of the fused features, and driver attention.
[0047] Based on the content of the above method embodiment, as a preferred embodiment, in the method for predicting driver attention based on in-vehicle forward-looking images provided by the embodiment of the present invention, the extraction of image visual features and semantic features is as follows. Refer to Figure 2 . The method includes: using a ViT (Vision Transformer) neural network to divide the in-vehicle forward-looking image into blocks. The length and width of the blocks are one-eighth of the original in-vehicle forward-looking image. Position encoding is performed on the blocks, but classification encoding is not performed. The correlation between the blocks is calculated through a multi-head attention module, and the mean value of the calculated correlation is the result of visual feature extraction. It should be noted that in this embodiment, relative visual features in different traffic scenarios can be obtained, realizing global perception of the in-vehicle forward-looking image and extraction of visual features such as color and texture, avoiding problems such as limited receptive fields caused by directly extracting visual features by a convolutional neural network. Since this embodiment only needs to extract features such as color and texture without knowing the probability of which category these features belong to, only position encoding is performed on the blocks in this embodiment, and classification encoding is not performed.
[0048] As a preferred embodiment, the extraction of image visual features and semantic features further includes: using a DeepLabv3 semantic segmentation neural network to perform semantic segmentation on the in-vehicle forward-looking image to obtain semantic images such as roads, pedestrians, vehicles, and traffic facilities, reducing misjudgment of semantic information; using a convolutional neural network to perform perception and encoding on the semantic image to convert the semantic map into semantic nodes; using a fully connected neural network to construct the association between semantic nodes to obtain a semantic node adjacency matrix; taking the obtained semantic node features and semantic node adjacency matrix as the expression form of the graph to form a semantic graph, improving the association strength between different traffic elements in different traffic scenarios; using a 3-layer graph convolutional GCN neural network to encode the semantic graph to obtain the semantic features of the in-vehicle forward-looking image. Compared with directly extracting semantic features using neural networks such as Transformer and CNN, the semantic features in this embodiment have stronger expression ability for the association between different traffic elements and higher acquisition accuracy of semantic elements.
[0049] Based on the content of the above method embodiments, as a preferred embodiment, in the embodiment of the present invention, the driver attention prediction method based on in-vehicle forward-looking images, the fusion of visual features and semantic features is shown in FIG. 3. The method includes: performing a Hadamard product on the visual features and semantic features obtained in the previous step, that is, multiplying each corresponding position one by one, with the dimension remaining unchanged, to obtain fused features; using a fully connected neural network to perform deep feature extraction on the fused features, and re-extracting the fused features in a way of first reducing the dimension and then increasing the dimension to prevent overfitting; using a residual network structure to add the fused features and the visual features, that is, adding each corresponding position one by one, with the dimension remaining unchanged, and the final fused features after addition express the joint effect of the visual and semantic elements in the in-vehicle forward-looking image on the driver's attention, mapping the attracting elements of the in-vehicle forward-looking image to the driver's attention. The feature fusion method in this embodiment avoids directly using methods such as feature addition and feature splicing, reduces information loss in the feature fusion process, and uses a residual network structure and a fully connected neural network to first reduce the dimension and then increase the dimension, improving the expression ability of the fused features for different weights of visual features and semantic features.
[0050] Based on the content of the above method embodiments, as a preferred embodiment, in the embodiment of the present invention, the driver attention prediction method based on in-vehicle forward-looking images, the prediction of the fused feature spatio-temporal sequence and the generation of driver attention are shown in Figure 4 . The method includes: using an LSTM neural network to memorize, reason, and generate the fused features of the prediction frame for several sequences of fused features; using a multi-layer 2D transposed convolutional neural network to decode the fused features to generate the driver attention with the same dimension as the in-vehicle forward-looking image; using a generative adversarial neural network GAN to supervise the generation of the driver attention from the fused features of the prediction frame, and using the real driver attention as the true value of the discriminator, and training in a loop to improve the generation accuracy of the driver attention.
[0051] The implementation basis of each embodiment of the present invention is realized through programmed processing by a device with a processor function. Therefore, in engineering practice, the technical solutions and functions of each embodiment of the present invention are encapsulated into various modules. Based on this actual situation, on the basis of the above embodiments, the embodiment of the present invention provides a driver attention prediction device based on in-vehicle forward-looking images, and this device is used to execute the driver attention prediction method based on in-vehicle forward-looking images in the above method embodiments.
[0052] See Figure 5 , this device includes: a feature extraction module, which is used to extract visual features and semantic features based on the obtained in-vehicle forward-looking images; a feature fusion module, which is used to fuse visual features and semantic features based on the extracted visual features and semantic features; a prediction and generation module, which is used to perform spatio-temporal sequence prediction and driver attention generation based on the fused features.
[0053] The driver attention prediction device based on on-vehicle front-view images provided by the embodiments of the present invention adopts Figure 5 several modules in
[0054] It should be noted that the device embodiments provided by the present invention, in addition to being used to implement the methods in the above method embodiments, are also used to implement the methods in other method embodiments provided by the present invention. The difference is only in setting corresponding functional modules. The principle is basically the same as that of the above device embodiments provided by the present invention. As long as those skilled in the art, on the basis of the above device embodiments, refer to the specific technical solutions in other method embodiments, obtain corresponding technical means by combining technical features, and the technical solutions constituted by these technical means, and improve the system in the above device embodiments on the premise of ensuring the practicability of the technical solutions, corresponding device-like embodiments can be obtained for implementing the methods in other method-like embodiments. For example:
[0055] Based on the content of the above device embodiments, as a preferred embodiment, the driver attention prediction device based on on-vehicle front-view images provided by the embodiments of the present invention further includes:
[0056] The first feature extraction sub-module is used to divide the on-vehicle front-view image into blocks, and perform position encoding on each block; calculate the correlation between the blocks, and obtain visual features according to the mean value of the calculated correlation.
[0057] The second feature extraction sub-module is used to perform semantic segmentation on the on-vehicle front-view image in a traffic environment to obtain several semantic images; convert the several semantic images into several semantic nodes; encode the several semantic nodes and convert them into a semantic node adjacency matrix, and the semantic node adjacency matrix and the semantic nodes form a semantic graph; encode the semantic graph to obtain semantic features.
[0058] Based on the content of the above device embodiments, as a preferred embodiment, the driver attention prediction device based on on-vehicle front-view images provided by the embodiments of the present invention further includes:
[0059] The first feature fusion sub-module is used to perform Hadamard product on the extracted visual features and semantic features to obtain fused features;
[0060] The second feature fusion sub-module is used to re-extract the fused features by first reducing the dimension and then increasing the dimension using a fully connected neural network, and obtain the re-extracted fused features;
[0061] The third feature fusion sub-module is used to add the re-extracted fused features to the visual features to obtain the final fused features.
[0062] Based on the content of the above device embodiment, as a preferred embodiment, the driver attention prediction device based on in-vehicle front view images provided in the embodiments of the present invention further includes:
[0063] The prediction sub-module is used to memorize, reason and generate the fused features of the prediction frame based on the fused features by using an LSTM neural network;
[0064] The generation sub-module is used to decode the fused features of the prediction frame into driver attention, and under the supervision of the neural network, output a driving attention map that meets the accuracy requirements.
[0065] Based on the content of the above device embodiment, as a preferred embodiment, the driver attention prediction device based on in-vehicle front view images provided in the embodiments of the present invention further includes:
[0066] The supervision sub-module is used to supervise the generation of driver attention from the fused features of the prediction frame, and use the real driver attention as the true value of the discriminator, and cycle training until the generation accuracy of the driver attention reaches the preset accuracy requirement.
[0067] The method of the embodiments of the present invention is implemented relying on an electronic device. Therefore, it is necessary to introduce the relevant electronic device. For this purpose, an embodiment of the present invention provides an electronic device, as Figure 6 shown, the electronic device includes: at least one processor, a communication interface, at least one memory, and a communication bus. Among them, at least one processor, the communication interface, and at least one memory complete communication with each other through the communication bus. At least one processor calls the logic instructions in at least one memory to execute all or part of the steps of the methods provided in the foregoing method embodiments.
[0068] In addition, when the logic instructions in at least one of the above memories are implemented in the form of software functional units and sold or used as independent products, they are stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present invention essentially, or the part that contributes to the prior art, or a part of this technical solution is embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which is a personal computer, a server, or a network device) to execute all or part of the steps of the methods described in various method embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical disks, and various media for storing program codes.
[0069] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, located in one place, or distributed to multiple network units. Select some or all of the modules according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0070] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can be implemented in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can be implemented in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0071] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more flows or multiple flows and / or blocks Figure 1 one or more blocks or multiple blocks.
[0072] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 in one block or multiple blocks.
[0073] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 in one block or multiple blocks.
[0074] In summary of the above embodiments, in view of the variability of driver attention and traffic scenarios, the present invention proposes a method for predicting driver attention based on on-vehicle forward vision images. The method uses the pytorch deep learning framework and includes the following steps: extracting visual features such as color and texture of the on-vehicle forward vision images; extracting semantic features of the on-vehicle forward vision images, including roads, pedestrians, vehicles, and traffic facilities, etc.; combining the visual features and semantic features to obtain fusion features, and mapping the attracting factors of the on-vehicle forward vision images to driver attention; memorizing and reasoning about the fusion features of multiple frames to obtain the fusion features of the prediction frame, and decoding the fusion features of the prediction frame and using a generative adversarial neural network for supervision to obtain driver attention. The present invention realizes the rapid identification of key information in sequential images, which helps to design autonomous driving tasks, optimize visual scene understanding, and provide decision-making knowledge.
[0075] It should be noted that the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0076] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features. And these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.
Claims
1. A driver attention prediction method based on vehicle-mounted forward-looking images, characterized in that: include: Based on the acquired vehicle-mounted forward-view image, visual features and semantic features are extracted, and only position encoding is performed when the visual features are extracted using a VIT neural network; Extracting the semantic features includes: using the DeepLabv3 semantic segmentation neural network to perform semantic segmentation on the vehicle-mounted forward-view image in a traffic environment to obtain a number of semantic images; using a convolutional neural network to perceive and encode the semantic images, and converting a number of the semantic graphs into a number of semantic nodes; using a fully connected neural network to build associations between semantic nodes and obtain a semantic node adjacency matrix; using the obtained semantic node features and the semantic node adjacency matrix as a graph expression to form a semantic graph; using a 3-layer graph convolution GCN neural network to encode the semantic graph to obtain semantic features; Based on the extracted visual features and semantic features, the visual features and semantic features are fused; Based on the fused features, spatiotemporal sequence prediction and driver attention generation are performed.
2. The method for predicting driver attention based on vehicle-mounted forward-looking images according to claim 1, characterized in that: Based on the acquired vehicle-mounted forward-view image, visual feature extraction is performed, including: Divide the vehicle-mounted front view image into blocks and perform position coding on each block; The correlation between the blocks is calculated, and the visual features are obtained according to the average value of the calculated correlation.
3. The driver attention prediction method based on vehicle-mounted forward-looking images according to claim 1 is characterized in that: Based on the extracted visual features and semantic features, the visual features and semantic features are fused, including: Perform Hadamard product on the extracted visual features and semantic features to obtain fusion features; Re-extracting the fused features by using a fully connected neural network in a manner of first reducing the dimension and then increasing the dimension to obtain re-extracted fused features; The re-extracted fusion features are added to the visual features to obtain the final fusion features.
4. The method for predicting driver attention based on vehicle-mounted forward-looking images according to claim 1, characterized in that: Based on the fused features, spatiotemporal sequence prediction and driver attention generation are performed, including: Based on the fused features, the LSTM neural network is used to memorize, infer and generate the fused features of the prediction frame; The fused features of the predicted frame are decoded into driver attention, and under the supervision of a neural network, a driving attention map that meets accuracy requirements is output.
5. The method for predicting driver attention based on vehicle-mounted forward-looking images according to claim 4, characterized in that: The method further comprises: The fused features of the supervised prediction frame are used to generate the driver's attention, and the real driver's attention is used as the true value of the discriminator. The training is repeated until the generation accuracy of the driver's attention reaches the preset accuracy requirement.
6. A driver attention prediction device based on vehicle-mounted forward-looking images, characterized in that: include: A feature extraction module, used for extracting visual features and semantic features based on the acquired vehicle-mounted forward-view image, and only performing position encoding when extracting the visual features using a VIT neural network; Extracting the semantic features includes: using the DeepLabv3 semantic segmentation neural network to perform semantic segmentation on the vehicle-mounted front view image in a traffic environment to obtain a number of semantic images; using a convolutional neural network to perceive and encode the semantic images, and converting the semantic graphs into a number of semantic nodes; using a fully connected neural network to build associations between semantic nodes and obtain a semantic node adjacency matrix; using the obtained semantic node features and the semantic node adjacency matrix as a graph expression to form a semantic graph; using a 3-layer graph convolution GCN neural network to encode the semantic graph to obtain semantic features; A feature fusion module is used to fuse visual features and semantic features based on the extracted visual features and semantic features; The prediction and generation module is used to perform spatiotemporal sequence prediction and driver attention generation based on the fused features.
7. The driver attention prediction device based on vehicle-mounted forward-looking images according to claim 6, characterized in that: The device also includes: a supervision module, which is used to supervise the generation of driver attention and perform cyclic training until the generation accuracy of the driver's attention reaches a preset requirement.
8. An electronic device, characterized in that: include: At least one processor, at least one memory and a communication interface; wherein the processor, memory and communication interface communicate with each other; The memory stores program instructions executed by the processor, and the processor calls the program instructions to execute the method according to any one of claims 1 to 5.
9. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores computer instructions, which cause the computer to execute the method of any one of claims 1 to 5.
Citation Information
Patent Citations
Driving early warning method based on driver visual attention prediction
CN112699821A
End-to-end automatic driving behavior planning method based on graph attention
CN117341727A
Infrared and visible light image fusion method and system based on channel attention
CN117745560A