End-to-end automatic driving method and system based on cognitive enhancement, and vehicle
By introducing an end-to-end autonomous driving method that simulates embodied cognitive data, and simulating the decision-making process of human drivers, the problem of insufficient intelligence of autonomous driving is solved and a safer and more accurate driving trajectory generation is achieved.
Patent Information
- Application Number
- CN202510173415.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-07-04
AI Technical Summary
The existing autonomous driving technology is not very intelligent and cannot achieve true autonomous driving. The driver still needs to actively control it, and the safety and driving trajectory accuracy are insufficient.
Simulated embodied cognitive data is introduced, and the decision-making process of human drivers is simulated through pre-trained end-to-end autonomous driving model and embodied cognitive data generation model, and cognitively enhanced brain-like autonomous driving decisions are generated.
Significantly reduce the collision rate, more accurate driving trajectory generation, and improve the safety performance and intelligence of autonomous driving.
Smart Images

Figure CN120258096A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving, and particularly to an autonomous driving enhancement method, system, and vehicle based on human cognitive data. Background Art
[0002] Autonomous driving technology is crucial for transferring driving authority from human drivers to sensors and artificial intelligence, and is expected to improve traffic efficiency and safety. Some auxiliary driving technologies are commonly applied in the current automotive industry, such as a vehicle distance monitoring system for safe driving and a path planning technology applicable to simple road sections. These technologies have alleviated the driver's operation burden to a certain extent. However, in most cases, the driver still needs to actively control the vehicle, and true autonomous driving cannot be achieved, resulting in low intelligence. Summary of the Invention
[0003] The present invention provides an end-to-end autonomous driving method, system, and vehicle based on embodied cognition enhancement to solve the defect of low intelligence in existing autonomous driving technologies. By introducing simulated embodied cognition data, the autonomous driving system of the present invention can simulate the decision-making process of human drivers, significantly reduce the collision rate, and generate more accurate driving trajectories.
[0004] The present invention provides an end-to-end autonomous driving method based on embodied cognition enhancement, including: determining current driving environment information; inputting the current driving environment information and simulated embodied cognition data into a pre-trained end-to-end autonomous driving model to obtain a brain-like autonomous driving decision with enhanced cognition output by the pre-trained end-to-end autonomous driving model; the simulated embodied cognition data is a physiological simulation signal related to cognition generated by calling a pre-trained embodied cognition data generation model based on the current driving environment information; wherein, the pre-trained end-to-end autonomous driving model is trained based on an end-to-end autonomous driving data set; the end-to-end autonomous driving data set includes a driving environment information set, simulated embodied cognition data samples, and end-to-end autonomous driving decision samples; the driving environment information set includes driving environment information samples; the simulated embodied cognition data samples are physiological simulation signal samples related to cognition generated by calling the pre-trained embodied cognition data generation model; the pre-trained embodied cognition data generation model is trained based on the driving environment information samples and real embodied cognition data samples; the real embodied cognition data samples are physiological real signals related to cognition.
[0005] An end-to-end autonomous driving method based on embodied cognition enhancement provided by the present invention further includes a training method for the pre-trained embodied cognition data generation model: obtaining the driving environment information sample and the real embodied cognition data sample; preprocessing and data segmentation of the driving environment information sample to obtain driving environment training data; encoding the driving environment training data through a pre-trained visual model to obtain an initial visual embedding vector; mapping the initial visual embedding vector to a driving cognition feature space through a preset projection layer according to the real embodied cognition data sample to obtain a driver visual embedding vector; inputting the driver visual embedding vector into a pre-trained neural network for deep semantic feature extraction to obtain a driver cognition embedding vector output by the neural network; and training the embodied cognition data generation model according to the driver visual embedding vector and the driver cognition embedding vector.
[0006] An end-to-end autonomous driving method based on embodied cognition enhancement provided by the present invention further includes: segmenting and aligning the sampling rate of the driving environment information sample and the sampling rate of the real embodied cognition data sample.
[0007] In an end-to-end autonomous driving method based on embodied cognition enhancement provided by the present invention, the neural network is a spatio-temporal convolutional network, and the spatio-temporal convolutional network is used to extract spatio-temporal signal features of the driver visual embedding vector.
[0008] In an end-to-end autonomous driving method based on embodied cognition enhancement provided by the present invention, the step of inputting the current driving environment information and simulated embodied cognition data into a pre-trained end-to-end autonomous driving model to obtain a brain-like autonomous driving decision with enhanced cognition output by the pre-trained end-to-end autonomous driving model includes: performing a BEV query, an agent query, and a map query according to the current driving environment information to obtain a BEV query result, an agent query result, and a map query result; performing cross-attention calculation according to the simulated embodied cognition data, the BEV query result, and the agent query result to obtain a self-agent interaction result; performing cross-attention calculation according to the simulated embodied cognition data, the BEV query result, and the map query result to obtain a self-map interaction result; and obtaining a brain-like autonomous driving trajectory planning with enhanced cognition according to the self-agent interaction result and the self-map interaction result.
[0009] In an end-to-end autonomous driving method based on embodied cognition enhancement provided by the present invention, the loss function of the pre-trained end-to-end autonomous driving model is: + ; The loss function of the pre-trained embodied cognition data generation model is: ; Among them, is the loss function of the pre-trained embodied cognitive data generation model, is the map loss function, is the agent loss function, is the imitation learning loss function, is the scene constraint loss function, is the corresponding weight, is the corresponding weight, is the corresponding weight, is the corresponding weight, is the corresponding weight, i is the number of samples of the target EEG, j is the number of electrodes, N s is the number of sample points, D e is the number of electrodes, is the i th sample of the target EEG and the EEG signal of the j th electrode, is the i th sample of the generated EEG and the EEG signal of the j th electrode.
[0010] The present invention also provides an end-to-end autonomous driving system based on embodied cognitive enhancement, including: a determination module for determining current driving environment information; a decision-making module for inputting the current driving environment information and simulated embodied cognitive data into a pre-trained end-to-end autonomous driving model to obtain a brain-like autonomous driving decision with cognitive enhancement output by the pre-trained end-to-end autonomous driving model; the simulated embodied cognitive data is a physiological simulation signal related to cognition generated by calling a pre-trained embodied cognitive data generation model based on the current driving environment information; wherein, the pre-trained end-to-end autonomous driving model is trained based on an end-to-end autonomous driving data set; the end-to-end autonomous driving data set includes a driving environment information set, simulated embodied cognitive data samples and end-to-end autonomous driving decision samples; the driving environment information set includes driving environment information samples; the simulated embodied cognitive data samples are physiological simulation signal samples related to cognition generated by calling the pre-trained embodied cognitive data generation model; the pre-trained embodied cognitive data generation model is trained based on the driving environment information samples and real embodied cognitive data samples; the real embodied cognitive data samples are physiological real signals related to cognition.
[0011] The present invention also provides an autonomous vehicle, which is characterized by including the above-mentioned end-to-end autonomous driving system enhanced by embodied cognition.
[0012] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method for end-to-end autonomous driving enhanced by embodied cognition as described in any one of the above is implemented.
[0013] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for end-to-end autonomous driving enhanced by embodied cognition as described in any one of the above is implemented.
[0014] The present invention provides a method, a system, and a vehicle for end-to-end autonomous driving enhanced by embodied cognition. The method includes: determining current driving environment information; inputting the current driving environment information and simulated embodied cognition data into a pre-trained end-to-end autonomous driving model to obtain a brain-like autonomous driving decision enhanced by cognition output by the pre-trained end-to-end autonomous driving model; the simulated embodied cognition data is a physiological simulation signal related to cognition generated by calling a pre-trained embodied cognition data generation model based on the current driving environment information; wherein, the pre-trained end-to-end autonomous driving model is trained based on an end-to-end autonomous driving data set; the pre-trained embodied cognition data generation model is trained based on driving environment information samples and real embodied cognition data samples. By introducing simulated embodied cognition data, the present invention enables the autonomous driving system to simulate the decision-making process of human drivers, significantly reducing the collision rate and making the generation of driving trajectories more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 FIG. is a schematic flowchart of a method for end-to-end autonomous driving enhanced by embodied cognition provided by the present invention.
[0017] Figure 2 FIG. is a schematic principle diagram of a method for end-to-end autonomous driving enhanced by embodied cognition provided by the present invention.
[0018] Figure 3 FIG. is a schematic structural diagram of an end-to-end autonomous driving system enhanced by embodied cognition provided by the present invention.
[0019] Figure 4 It is a schematic structural diagram of the electronic device provided by the present invention. Specific embodiments
[0020] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0021] Please refer to Figure 1 , Figure 1 which is a schematic flow diagram of an end-to-end autonomous driving method based on embodied cognition enhancement provided by the present invention.
[0022] The present invention provides an end-to-end autonomous driving method based on embodied cognition enhancement, including: 101: Determine the current driving environment information; 102: Input the current driving environment information and simulated embodied cognition data into a pre-trained end-to-end autonomous driving model to obtain a brain-like autonomous driving decision with enhanced cognition output by the pre-trained end-to-end autonomous driving model; the simulated embodied cognition data is cognitive-related physiological simulation signals generated by calling a pre-trained embodied cognition data generation model based on the current driving environment information; wherein, the pre-trained end-to-end autonomous driving model is trained based on an end-to-end autonomous driving data set; the end-to-end autonomous driving data set includes a driving environment information set, simulated embodied cognition data samples and end-to-end autonomous driving decision samples; the driving environment information set includes driving environment information samples; the simulated embodied cognition data samples are cognitive-related physiological simulation signal samples generated by calling a pre-trained embodied cognition data generation model; the pre-trained embodied cognition data generation model is trained based on driving environment information samples and real embodied cognition data samples; the real embodied cognition data samples are cognitive-related physiological real signals.
[0023] In the pursuit of autonomous driving, the emergence of end-to-end autonomous driving has recently attracted increasing attention. The end-to-end approach represents a paradigm shift by integrating the subtasks of autonomous driving, namely 3D perception, map semantic segmentation, motion prediction, and planning, into a multi-task learning framework. This approach enhances safety by addressing perception errors that may not be correctable during the planning phase and improves the integrity of the framework, thereby making the optimization objectives and performance evaluations more consistent. Humans, as highly capable embodied agents, possess unparalleled reasoning and understanding abilities. For example, a driver can navigate safely even in unfamiliar scenarios. This invention hypothesizes that the human brain contains a special world model that enables scene-level reasoning to assist in decision-making. For instance, when turning on a mountain road, despite limited visibility, a driver can predict oncoming vehicles and make informed decisions to ensure safety. By understanding the brain's scene reasoning and understanding mechanisms, the end-to-end approach can better predict scene evolution and driving intentions. This invention utilizes embedded cognition to enhance end-to-end autonomous driving performance. First, the current driving environment information can be determined through the external sensors of the autonomous driving system. Then, the current driving environment information and simulated embodied cognition data are input into a pre-trained end-to-end autonomous driving model, and a brain-like autonomous driving decision enhanced by cognition generated by the pre-trained end-to-end autonomous driving model can be obtained. The simulated embodied cognition data is cognitive-related physiological simulation signals (such as electroencephalogram, magnetoencephalogram, nuclear magnetic resonance, heart rate, electromyogram, skin conductance, etc.) generated by invoking a pre-trained embodied cognition data generation model based on the current driving environment information. By introducing the simulated embodied cognition data, this invention enables the autonomous driving system to better simulate the decision-making process of human drivers, significantly reducing the collision rate, generating more accurate driving trajectories, and substantially improving the overall safety performance.
[0024] In addition, the pre-trained embodied cognition data generation model aligns the visual model and cognitive-related physiological signals (EEG signals) during the driving process with a self-collected multi-modal human driving dataset, aiming to mimic the changes in the driver's brain signals in the driving scenario. The samples required for training include driving environment information samples and real embodied cognition data samples. It can also include sensor data, operation behaviors, etc. The driving environment information set includes driving environment information samples. The driving environment information sample can be a part (video) of the driving environment information set. The real embodied cognition data sample is cognitive-related physiological real signals (including electroencephalogram, eye movement, heart rate, skin conductance, etc.).
[0025] The end-to-end autonomous driving dataset required for training a pre-trained end-to-end autonomous driving model includes a collection of driving environment information, simulated embodied cognition data samples, and end-to-end autonomous driving decision samples. The simulated embodied cognition data samples are physiological simulation signal samples related to cognition generated by invoking a pre-trained embodied cognition data generation model. Through empirical experiments on a large-scale end-to-end autonomous driving dataset, the performance of the pre-trained end-to-end autonomous driving model is verified.
[0026] As a preferred embodiment, it further includes a training method for the pre-trained embodied cognition data generation model: obtaining driving environment information samples and real embodied cognition data samples; preprocessing and data splitting the driving environment information samples to obtain driving environment training data; encoding the driving environment training data through a pre-trained visual model to obtain an initial visual embedding vector; mapping the initial visual embedding vector to the driving cognition feature space through a preset projection layer according to the real embodied cognition data samples to obtain a driver visual embedding vector; inputting the driver visual embedding vector into a pre-trained neural network for deep semantic feature extraction to obtain a driver cognition embedding vector output by the neural network; training the embodied cognition data generation model according to the driver visual embedding vector and the driver cognition embedding vector.
[0027] Please refer to Figure 2 , Figure 2 which is a schematic diagram of the principle of an end-to-end autonomous driving method based on embodied cognition enhancement provided by the present invention.
[0028] As a preferred embodiment, it further includes: splitting and aligning the sampling rate of the driving environment information samples with the sampling rate of the real embodied cognition data samples.
[0029] To train the embodied cognition data generation model, in this embodiment, real electroencephalogram (EEG) signal samples (real embodied cognition data samples) and driving environment video samples (driving environment information samples) are taken as examples to illustrate the generation of cognitive representations through supervised learning. Specifically, the EEG data is first preprocessed. The signals are rereferenced to the M1 and M2 electrodes and band-pass filtered between 0 and 5 Hz. Independent component analysis (ICA) is performed to remove artifacts related to eye movements, channel noise, heartbeats, and limb movements. Then, training samples of the embodied cognition data generation model are constructed in the form of time series data pairs (video, electroencephalogram) from the driver's video data and EEG signals. The duration varies depending on the participant and driving conditions. The data is segmented into 8-second segments. Data segments shorter than 8 seconds are also included as single samples. The video data is encoded through a pre-trained visual model (such as ResNet50) to obtain an initial visual embedding vector; the initial visual embedding vector is mapped to the driving cognitive feature space through a preset projection layer to obtain the driver's visual embedding vector. The EEG signal is represented as a two-dimensional matrix , where is the target EEG signal, N s is the number of sample points, D e is the number of electrodes (i.e., the dimension of the EEG).
[0030] Since the sampling rates of the EEG signals (real embodied cognition data samples) and video frames (driving environment information samples) are inconsistent, multiple EEG samples need to be generated for each visual embedding output of ResNet50 to align the sampling rate of the driving environment information samples with that of the real embodied cognition data samples.
[0031] The relationship between the sampling rate of the driving environment video samples and the sampling rate of the EEG signal samples is as follows: , , where f eeg is the sampling rate of the EEG signal samples, f video is the sampling rate of the driving environment video samples, N spf is the number of samples to be generated per frame, is the assignment operator.
[0032] To ensure f eeg can be divided by f videoDivide to adjust the sampling rate of EEG. Since it is difficult to change the sampling rate of the video, interpolation is applied to modify the sampling rate of EEG. Filter the EEG signal into a lower frequency band to minimize the loss of cognitive information. Therefore, interpolation can be performed without affecting the data quality. When The reason for adding '1' is to prevent the reduction of the EEG sampling rate, which may lead to the loss of important information.
[0033] After being processed by ResNet50, the input video is encoded into an embedding , where N f is the number of frames, D v is the embedding length of the video frame. The projection layer maps each frame embedding to the feature space of the EEG representation . Then, through N f and N spf are multiplied to obtain N s , and the EEG embedding can be reshaped to match EEG tgt 's shape. In the implementation, these cognitive data are processed in mini-batches, so the first dimension of the tensor is set to the batch size, denoted by B and . To describe the method of the present invention more clearly and in detail, the batch size dimension of the tensor is retained in the following parts.
[0034] As a preferred embodiment, the neural network is a spatio-temporal convolutional network, which is used to extract the spatio-temporal signal features of the driver's visual embedding vector.
[0035] To further extract the driver's cognitive embedding vector, after extracting , the potential latent factors are rearranged to maintain the time dimension, denoted as . The brain is a form of efficient and fast embodied intelligence. In the driving scenario, the driver will call on potential cognition for complex and fast reasoning. This process corresponds to the brain's reasoning during the duration of , rather than being limited to the moment of capturing the image. Therefore, convolution is applied in the spatio-temporal dimensions N spf D e to extract the driving cognitive embedding in the dense spatio-temporal features. The driving cognition embodies deep semantics. To extract the spatio-temporal signal features of the driver's visual embedding vector, the spatio-temporal convolutional network consists of a convolutional module and three residual modules. The convolutional module maintains the size of the spatio-temporal feature map while increasing the number of channels to Q egoDimension D q One - quarter of, where the ego query Q ego is used for end - to - end method to learn implicit scene features. The initial residual module includes a series of stacked 3x3 convolutional layers with a stride of 1, normalization layers, and activation function layers. The number of output channels remains unchanged, keeping the feature size unchanged. In contrast, the second and third residual modules use 3x3 convolutions with a stride of 2 and double the channels. It is worth noting that this design minimizes the information loss of pooling operations and uses residual connections for self - mapping interactions informed by driving cognition. After interacting with the agent, it prevents the vanishing gradient near the output layer. The spatio - temporal convolution process finally extracts the potential , as the driver cognition embedding vector.
[0036] As a preferred embodiment, input the current driving environment information and simulated embodied cognition data into a pre - trained end - to - end autonomous driving model to obtain the cognition - enhanced brain - like autonomous driving decision output by the pre - trained end - to - end autonomous driving model, including: performing BEV query, agent query, and map query according to the current driving environment information to obtain BEV query results, agent query results, and map query results; performing cross - attention calculation according to the simulated embodied cognition data, BEV query results, and agent query results to obtain self - agent interaction results; performing cross - attention calculation according to the simulated embodied cognition data, BEV query results, and map query results to obtain self - map interaction results; obtaining the cognition - enhanced brain - like autonomous driving trajectory planning according to the self - agent interaction results and self - map interaction results.
[0037] In this embodiment, perception is crucial for autonomous driving, especially regarding map elements and traffic agents. Perform BEV query, agent query, and map query according to the current driving environment information to obtain BEV query results, agent query results, and map query results. Specifically, the bird's - eye view (BEV) can effectively integrate multi - modal data and has important prospects for perception tasks. Use deformable attention to extract agent - level features from the shared BEV feature map. The attributes of the agent, such as position, class score, and orientation, are decoded from the agent query through an MLP - based decoder head. Similarly, map queries, including map vectors and classification scores, can be extracted using this method. Use the initialized BEV query to learn implicit scene features, and then derive the planned trajectory through interaction with map and agent elements.
[0038] To learn the spatial and motion features of traffic agents, first extract the empirical cognition about traffic agents from latent variables. It uses adaptive pooling and a single - layer MLP CEDistillation is performed to form proxy cognitive embeddings, which encapsulate the driver's experience of the movement and behavior patterns of various traffic proxy categories. This cognition is integrated into the BEV query as an important enhancement of the potential information of the scene. Although the positions of the ego vehicle and other proxies are encoded using a single-layer MLP PE This encoding serves as the query position embedding and the key position embedding, capturing the relative position relationship. Then, the BEV query and the proxy query are processed by the transformer decoder, and cross-attention calculations are performed based on the simulated embodied cognitive data to obtain the ego-proxy interaction results.
[0039] After interacting with the proxy, the updated BEV query further interacts with the map query. To provide insights into the crossability and hazard levels associated with each map element (e.g., lane dividers, road boundaries, and crosswalks). Similar to the interaction with the proxy, a single-layer MLP is designed to encode the relative position between the map element and the ego vehicle, and cross-attention calculations are performed based on the simulated embodied cognitive data, the BEV query results, and the map query results to obtain the ego-map interaction results, which encapsulate comprehensive information about the driving scene. Then, the ego-proxy interaction results and the ego-map interaction results are fed into a planner (the VAD Planner in the present invention) to generate a cognition-enhanced brain-like autonomous driving trajectory plan.
[0040] As a preferred embodiment, the loss function of the pre-trained end-to-end autonomous driving model is:[[]] + ; The loss function of the pre-trained embodied cognitive data generation model is:[[]] ; Wherein, is the loss function of the pre-trained embodied cognitive data generation model, is the map loss function, is the proxy loss function, is the imitation learning loss function, is the scene constraint loss function, is the corresponding weight, is the corresponding weight, is the corresponding weight, is the corresponding weight, is the corresponding weight, i is the number of samples of the target EEG, j is the number of electrodes, N s is the number of sample points,D e is the number of electrodes, is the i th sample of the target EEG and the EEG signal of the j th electrode, is the i th sample of the generated EEG and the EEG signal of the j th electrode.
[0041] In this embodiment, the pre-trained end-to-end autonomous driving model requires a BEVFormer network to perceive the elements in the driving scene and predict their respective motion states, so as to provide basic information for the subsequent planning process.
[0042] The map loss, denoted as , includes the regression loss between the predicted map points and the ground truth, as well as the focal loss of the map elements.
[0043] Meanwhile, the agent loss, denoted as , includes the 3D perception loss and the prediction loss. The 3D perception loss includes a classification component based on the focal loss and a bounding box regression component based on the loss. For the motion prediction task, from the N k predicted future trajectories of each agent, the trajectory representing the minimum final displacement error (minFDE) is selected; then, the loss is calculated with the ground truth trajectory. The focal loss is used as the multi-modal motion classification loss.
[0044] For autonomous driving, it is important to prioritize the understanding of the driving scene and the acquisition of a safe driving mode, rather than just replicating the expert's driving trajectory. Therefore, while retaining the imitation learning loss component , a lower weight is assigned to it, enabling the model to initially adopt a safer driving paradigm without overfitting to the expert driving trajectory.
[0045] Avoiding collisions is crucial for ensuring the safety of autonomous driving, and the driver's perception plays a vital role in this context. To address this issue, a scene constraint loss with a higher weight is implemented. This loss includes three components: the ego-agent collision constraint, the ego-boundary crossing constraint, and the ego-lane direction constraint.
[0046] These components are used to limit the interaction between the ego-agent and other agents, prevent crossing the road boundary, and maintain the correct driving direction respectively.
[0047] The end-to-end autonomous driving model can be trained end-to-end using the proposed loss functions above.
[0048] The end-to-end autonomous driving system based on embodied cognition enhancement provided by the present invention will be described below. The end-to-end autonomous driving system based on embodied cognition enhancement described below can be referred to correspondingly with the end-to-end autonomous driving method based on embodied cognition enhancement described above.
[0049] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of an end-to-end autonomous driving system based on embodied cognition enhancement provided by the present invention.
[0050] The present invention also provides an end-to-end autonomous driving system based on embodied cognition enhancement, including: a determination module 301 for determining current driving environment information; a decision-making module 302 for inputting the current driving environment information and simulated embodied cognition data into a pre-trained end-to-end autonomous driving model to obtain a brain-like autonomous driving decision with enhanced cognition output by the pre-trained end-to-end autonomous driving model; the simulated embodied cognition data is a physiological simulation signal related to cognition generated by calling a pre-trained embodied cognition data generation model based on the current driving environment information; wherein, the pre-trained end-to-end autonomous driving model is trained based on an end-to-end autonomous driving data set; the end-to-end autonomous driving data set includes a driving environment information set, simulated embodied cognition data samples, and end-to-end autonomous driving decision samples; the driving environment information set includes driving environment information samples; the simulated embodied cognition data samples are physiological simulation signal samples related to cognition generated by calling a pre-trained embodied cognition data generation model; the pre-trained embodied cognition data generation model is trained based on driving environment information samples and real embodied cognition data samples; the real embodied cognition data samples are physiological real signals related to cognition.
[0051] The autonomous driving vehicle provided by the present invention will be described below. The autonomous driving vehicle described below can be referred to correspondingly with the end-to-end autonomous driving method based on embodied cognition enhancement described above.
[0052] The present invention also provides an autonomous driving vehicle, characterized in that it includes the above-mentioned end-to-end autonomous driving system based on embodied cognition enhancement.
[0053] Figure 4 which exemplifies a schematic structural diagram of an electronic device, as Figure 4As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communications interface 420, and the memory 430 complete communication with each other through the communication bus 440. The processor 410 may call logic instructions in the memory 430 to execute an end-to-end autonomous driving method based on embodied cognition enhancement. The method includes: determining current driving environment information; inputting the current driving environment information and simulated embodied cognition data into a pre-trained end-to-end autonomous driving model to obtain an end-to-end autonomous driving decision with enhanced cognition output by the pre-trained end-to-end autonomous driving model; the simulated embodied cognition data is a physiological simulation signal related to cognition generated by calling a pre-trained embodied cognition data generation model based on the current driving environment information; among them, the pre-trained end-to-end autonomous driving model is trained based on an end-to-end autonomous driving data set; the end-to-end autonomous driving data set includes a driving environment information set, simulated embodied cognition data samples, and end-to-end autonomous driving decision samples; the driving environment information set includes driving environment information samples; the simulated embodied cognition data samples are physiological simulation signal samples related to cognition generated by calling a pre-trained embodied cognition data generation model; the pre-trained embodied cognition data generation model is trained based on driving environment information samples and real embodied cognition data samples; the real embodied cognition data samples are physiological real signals related to cognition.
[0054] In addition, when the logic instructions in the above-mentioned memory 430 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0055] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the end-to-end autonomous driving method based on embodied cognition enhancement provided by the above-mentioned various methods. The method includes: determining current driving environment information; inputting the current driving environment information and simulated embodied cognition data into a pre-trained end-to-end autonomous driving model to obtain a brain-like autonomous driving decision with enhanced cognition output by the pre-trained end-to-end autonomous driving model; the simulated embodied cognition data is a cognition-related physiological simulation signal generated by calling a pre-trained embodied cognition data generation model based on the current driving environment information; wherein, the pre-trained end-to-end autonomous driving model is trained based on an end-to-end autonomous driving data set; the end-to-end autonomous driving data set includes a driving environment information set, simulated embodied cognition data samples, and end-to-end autonomous driving decision samples; the driving environment information set includes driving environment information samples; the simulated embodied cognition data samples are cognition-related physiological simulation signal samples generated by calling a pre-trained embodied cognition data generation model; the pre-trained embodied cognition data generation model is trained based on driving environment information samples and real embodied cognition data samples; the real embodied cognition data samples are cognition-related physiological real signals.
[0056] In yet another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the end-to-end autonomous driving method based on embodied cognition enhancement provided by the above-mentioned various methods. The method includes: determining current driving environment information; inputting the current driving environment information and simulated embodied cognition data into a pre-trained end-to-end autonomous driving model to obtain a brain-like autonomous driving decision with enhanced cognition output by the pre-trained end-to-end autonomous driving model; the simulated embodied cognition data is a cognition-related physiological simulation signal generated by calling a pre-trained embodied cognition data generation model based on the current driving environment information; wherein, the pre-trained end-to-end autonomous driving model is trained based on an end-to-end autonomous driving data set; the end-to-end autonomous driving data set includes a driving environment information set, simulated embodied cognition data samples, and end-to-end autonomous driving decision samples; the driving environment information set includes driving environment information samples; the simulated embodied cognition data samples are cognition-related physiological simulation signal samples generated by calling a pre-trained embodied cognition data generation model; the pre-trained embodied cognition data generation model is trained based on driving environment information samples and real embodied cognition data samples; the real embodied cognition data samples are cognition-related physiological real signals.
[0057] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.
[0058] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0059] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An end-to-end autonomous driving method based on embodied cognition enhancement, characterized in that, Including: Determine the current driving environment information; Input the current driving environment information and simulated embodied cognition data into a pre-trained end-to-end autonomous driving model to obtain a brain-inspired autonomous driving decision with enhanced cognition output by the pre-trained end-to-end autonomous driving model; the simulated embodied cognition data is cognitive-related physiological simulation signals generated by invoking a pre-trained embodied cognition data generation model based on the current driving environment information; Among them, the pre-trained end-to-end autonomous driving model is trained based on an end-to-end autonomous driving dataset; the end-to-end autonomous driving dataset includes a driving environment information set, simulated embodied cognition data samples, and end-to-end autonomous driving decision samples; the driving environment information set includes driving environment information samples; the simulated embodied cognition data samples are cognitive-related physiological simulation signal samples generated by invoking the pre-trained embodied cognition data generation model; the pre-trained embodied cognition data generation model is trained based on the driving environment information samples and real embodied cognition data samples; the real embodied cognition data samples are cognitive-related physiological real signals.
2. The end-to-end autonomous driving method based on embodied cognition enhancement according to claim 1, wherein It also includes the training method of the pre-trained embodied cognition data generation model: Obtain the driving environment information samples and the real embodied cognition data samples; Preprocess and segment the driving environment information samples to obtain driving environment training data; Encode the driving environment training data through a pre-trained visual model to obtain an initial visual embedding vector; According to the real embodied cognition data samples, map the initial visual embedding vector to a driving cognition feature space through a preset projection layer to obtain a driver visual embedding vector; Input the driver visual embedding vector into a pre-trained neural network for deep semantic feature extraction to obtain a driver cognitive embedding vector output by the neural network; Train the embodied cognition data generation model according to the driver visual embedding vector and the driver cognitive embedding vector.
3. The end-to-end autonomous driving method based on embodied cognition enhancement according to claim 1, wherein It also includes: Perform segmentation alignment on the sampling rate of the driving environment information samples and the sampling rate of the real embodied cognition data samples.
4. The end-to-end autonomous driving method based on embodied cognition enhancement according to claim 2, wherein The neural network is a spatio-temporal convolutional network, and the spatio-temporal convolutional network is used to extract the spatio-temporal signal features of the driver visual embedding vector.
5. The end-to-end autonomous driving method based on embodied cognition enhancement according to claim 1, wherein The step of inputting the current driving environment information and simulated embodied cognition data into a pre-trained end-to-end autonomous driving model to obtain a brain-inspired autonomous driving decision with enhanced cognition output by the pre-trained end-to-end autonomous driving model includes: Perform BEV query, agent query, and map query according to the current driving environment information to obtain BEV query results, agent query results, and map query results; Perform cross-attention calculation according to the simulated embodied cognition data, the BEV query results, and the agent query results to obtain a self-agent interaction result; Perform cross-attention calculation according to the simulated embodied cognition data, the BEV query results, and the map query results to obtain a self-map interaction result; Based on the self-agent interaction result and the self-map interaction result, a cognitive-enhanced brain-inspired autonomous driving trajectory planning is obtained.
6. The end-to-end autonomous driving method based on embodied cognition enhancement according to any one of claims 1 to 5, characterized in that The loss function of the pre-trained end-to-end autonomous driving model is: + ; The loss function of the pre-trained embodied cognition data generation model is: ; Among them, is the loss function of the pre-trained embodied cognitive data generation model, is the map loss function, is the agent loss function, is the imitation learning loss function, is the scene constraint loss function, is the corresponding weight, is the corresponding weight, is the corresponding weight, is the corresponding weight, is the corresponding weight, i is the number of samples of the target EEG, j is the number of electrodes, N s is the number of sample points, D e is the number of electrodes, is the i th sample of the target EEG and the EEG signal of the j th electrode, is the i th sample of the generated EEG and the EEG signal of the j th electrode.
7. An end-to-end autonomous driving system based on embodied cognition enhancement, characterized in that, Including: A determination module for determining current driving environment information; A decision-making module for inputting the current driving environment information and simulated embodied cognition data into the pre-trained end-to-end autonomous driving model to obtain a cognitive-enhanced brain-inspired autonomous driving decision output by the pre-trained end-to-end autonomous driving model; the simulated embodied cognition data is a cognition-related physiological simulation signal generated by calling the pre-trained embodied cognition data generation model based on the current driving environment information; Wherein, the pre-trained end-to-end autonomous driving model is trained based on an end-to-end autonomous driving data set; the end-to-end autonomous driving data set includes a driving environment information set, simulated embodied cognition data samples, and end-to-end autonomous driving decision samples; the driving environment information set includes driving environment information samples; the simulated embodied cognition data samples are cognition-related physiological simulation signal samples generated by calling the pre-trained embodied cognition data generation model; the pre-trained embodied cognition data generation model is trained based on the driving environment information samples and real embodied cognition data samples; the real embodied cognition data samples are cognition-related physiological real signals.
8. An autonomous vehicle, characterized in that, Including the end-to-end autonomous driving system based on embodied cognition enhancement described in claim 7.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the end-to-end autonomous driving method based on embodied cognition enhancement described in any one of claims 1 to 6 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the end-to-end autonomous driving method based on embodied cognition enhancement described in any one of claims 1 to 6 is implemented.