End-to-end intelligent driving method, system and device based on artificial intelligence and medium

By adopting an end-to-end model in an intelligent driving system, the perceived data is directly mapped to driving actions, solving the problem of performance degradation in the face of rare or complex scenarios, achieving higher adaptability and performance.

CN120135211APending Publication Date: 2025-06-13GAC HONDA AUTOMOBILE CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510305754.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Existing intelligent driving systems may not be able to handle them accurately when facing rare or complex new scenarios outside of training data, resulting in system performance degradation or failure.

Method used

The end-to-end intelligent driving method based on artificial intelligence is adopted, and the vehicle perception data is directly mapped to the driving actions through the end-to-end model, reducing the system complexity and debugging difficulty and improving adaptability.

Benefits of technology

It realizes the direct mapping of vehicle perceived data to driving actions, simplifies system design and debugging, and improves the system's adaptability and performance in different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120135211A_ABST
    Figure CN120135211A_ABST
Patent Text Reader

Abstract

The invention discloses an end-to-end intelligent driving method, system and device based on artificial intelligence and a medium. The method comprises the following steps: acquiring camera sensing data and radar sensing data of a target vehicle; inputting the camera sensing data and the radar sensing data into a pre-trained end-to-end intelligent driving model to obtain a target driving action; and controlling the target vehicle to execute the target driving action. According to the embodiment of the invention, direct mapping from the vehicle sensing data to the vehicle driving action is realized by using the end-to-end model, the complexity and debugging difficulty of the vehicle intelligent driving system are reduced, the adaptability of the vehicle intelligent driving system is improved, and the method can be widely applied to the technical field of intelligent driving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent driving technology, and in particular to an end-to-end intelligent driving method, system, device and medium based on artificial intelligence. Background Art

[0002] Existing intelligent driving solutions are mainly divided into the following parts:

[0003] 1) Perception system. It includes a variety of sensors, such as cameras, millimeter-wave radars and ultrasonic radars, which are used to perceive the surrounding environment and combine the input information of various sensors through sensor fusion technology.

[0004] 2) Decision-making system. According to the high-precision map and real-time traffic information, a path planning algorithm is used to plan a route that not only conforms to traffic rules but also can reach the destination efficiently. Based on the data obtained by the perception system and the high-precision map information, the decision-making algorithm will make a judgment on the driving behavior of the vehicle. For example, when the vehicle in front suddenly brakes, the decision-making algorithm will comprehensively consider factors such as the speed of the vehicle itself and the distance from the vehicle in front, and decide whether to take emergency braking or lane change avoidance, and will calculate appropriate braking or steering operations according to the dynamic performance of the vehicle and the surrounding road environment.

[0005] 3) Execution system. It includes a chassis control system and a vehicle body control system. The chassis control system realizes reasonable steering, lane change and other actions by controlling the steering, braking and power systems. The vehicle body control system adjusts the braking force and driving force to prevent the vehicle from skidding or spinning.

[0006] The above intelligent driving solutions have the following problems:

[0007] 1) Each module of the existing intelligent driving system is usually designed for specific scenarios and tasks. When encountering rare scenarios or complex new scenarios outside the training data, such as special-shaped obstacles or non-standard traffic signs on the road, each module may not be able to process accurately, resulting in a decline or even failure of the overall system performance.

[0008] 2) The perception module transmits the processed information to the decision-making module, and the decision-making module then transmits the instruction to the control module. In this process, the information may be lost or distorted.

[0009] 3) Each module is designed and developed independently, and it is very difficult to achieve the global optimum of the entire system. Each module may perform well in its own tasks, but when combined together, there may be incoordination and conflicts between modules.

[0010] In summary, the existing intelligent driving solutions have the disadvantages of complex design, difficult debugging and poor adaptability. Summary of the Invention

[0011] The object of the present invention is to solve at least to some extent one of the technical problems existing in the prior art.

[0012] To this end, an object of an embodiment of the present invention is to provide an end-to-end intelligent driving method based on artificial intelligence. This method realizes the direct mapping from vehicle perception data to vehicle driving actions by using an end-to-end model, reduces the complexity and debugging difficulty of the vehicle intelligent driving system, and improves the adaptability of the vehicle intelligent driving system.

[0013] Another object of an embodiment of the present invention is to provide an end-to-end intelligent driving system based on artificial intelligence.

[0014] In order to achieve the above technical object, the technical solutions adopted in the embodiments of the present invention include:

[0015] In a first aspect, an embodiment of the present invention provides an end-to-end intelligent driving method based on artificial intelligence, including the following steps:

[0016] Obtain the camera perception data and radar perception data of the target vehicle;

[0017] Input the camera perception data and the radar perception data into a pre-trained end-to-end intelligent driving model to obtain a target driving action;

[0018] Control the target vehicle to execute the target driving action.

[0019] Further, in an embodiment of the present invention, the obtaining of the camera perception data and radar perception data of the target vehicle specifically includes:

[0020] Obtain the first image data captured by the camera of the target vehicle and the first point cloud data detected by the radar;

[0021] Perform normalization, cropping, and data augmentation processing on the first image data to obtain the camera perception data;

[0022] Perform filtering and coordinate transformation processing on the first point cloud data to obtain the radar perception data.

[0023] Further, in an embodiment of the present invention, the end-to-end intelligent driving model is trained through the following steps:

[0024] Obtain the camera sample data and radar sample data of the test vehicle under different road conditions and different weather conditions, and determine the corresponding driving action labels through manual annotation;

[0025] Input the camera sample data and the radar sample data into a pre-constructed multi-branch hybrid neural network to obtain a predicted driving action;

[0026] Determine a loss value according to the predicted driving action and the driving action label;

[0027] Update the parameters of the multi-branch hybrid neural network according to the loss value through the stochastic gradient descent algorithm to obtain the trained end-to-end intelligent driving model.

[0028] Further, in an embodiment of the present invention, the multi-branch hybrid neural network includes an input layer, a CNN branch network, an LSTM branch network, a feature fusion layer, and a fully connected layer. The CNN branch network includes multiple convolutional layers and a spatial pyramid pooling layer for extracting the spatial features of the camera sample data. The LSTM branch network includes a one-dimensional convolutional layer and a bidirectional LSTM layer for extracting the temporal features of the radar sample data.

[0029] Further, in an embodiment of the present invention, the step of inputting the camera sample data and the radar sample data into a pre-constructed multi-branch hybrid neural network to obtain a predicted driving action specifically includes:

[0030] Input the camera sample data into the CNN branch network through the input layer, and input the radar sample data into the LSTM branch network;

[0031] Perform multi-scale convolution processing on the camera sample data through the multiple convolutional layers to obtain multi-scale feature subgraphs, and perform feature fusion on the multi-scale feature subgraphs through the spatial pyramid pooling layer to obtain the spatial features;

[0032] Perform local temporal pattern extraction on the radar sample data through the one-dimensional convolutional layer to obtain local temporal features, and capture the forward and backward temporal dependencies of the local temporal features through the bidirectional LSTM layer to obtain the temporal features;

[0033] Perform feature fusion on the spatial features and the temporal features through the feature fusion layer to obtain fused features;

[0034] Map the fused features to the sample label space through the fully connected layer to obtain the predicted driving action.

[0035] Further, in an embodiment of the present invention, the step of performing feature fusion on the spatial features and the temporal features through the feature fusion layer to obtain fused features specifically includes:

[0036] Determine the attention weight matrices of the spatial features and the temporal features through the feature fusion layer based on the multi-head self-attention mechanism;

[0037] Perform multi-head splicing and linear projection on each dimension of the spatial feature and the temporal feature according to the attention weight matrix to obtain the fusion feature.

[0038] Further, in an embodiment of the present invention, the controlling the target vehicle to execute the target driving action specifically includes:

[0039] Generate a corresponding driving control instruction according to the target driving action;

[0040] Execute the driving control instruction through the driving control system of the target vehicle, so that the target vehicle believes in the target driving action.

[0041] In a second aspect, an embodiment of the present invention provides an end-to-end intelligent driving system based on artificial intelligence, including:

[0042] A data acquisition module, configured to acquire camera perception data and radar perception data of a target vehicle;

[0043] A driving decision module, configured to input the camera perception data and the radar perception data into a pre-trained end-to-end intelligent driving model to obtain a target driving action;

[0044] An action execution module, configured to control the target vehicle to execute the target driving action.

[0045] In a third aspect, an embodiment of the present invention provides an end-to-end intelligent driving device based on artificial intelligence, including:

[0046] At least one processor;

[0047] At least one memory, configured to store at least one program;

[0048] When the at least one program is executed by the at least one processor, the at least one processor implements the above-mentioned end-to-end intelligent driving method based on artificial intelligence.

[0049] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, in which a processor-executable program is stored, and the processor-executable program is used to execute the above-mentioned end-to-end intelligent driving method based on artificial intelligence when executed by a processor.

[0050] The advantages and beneficial effects of the present invention will be partially given in the following description, partially will become obvious from the following description, or will be understood through the practice of the present invention:

[0051] In an embodiment of the present invention, camera perception data and radar perception data of a target vehicle are obtained, the camera perception data and the radar perception data are input into a pre-trained end-to-end intelligent driving model to obtain a target driving action, and the target vehicle is controlled to execute the target driving action. The embodiment of the present invention realizes the direct mapping from vehicle perception data to vehicle driving actions by using an end-to-end model, reduces the complexity and debugging difficulty of the vehicle intelligent driving system, and improves the adaptability of the vehicle intelligent driving system. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following introduces the drawings required to be used in the embodiments of the present invention. It should be understood that the drawings in the following introduction are only for conveniently and clearly presenting some embodiments of the technical solutions in the present invention, and those skilled in the art can also obtain other drawings based on these drawings without creative efforts.

[0053] Figure 1 It is a flowchart of steps of an end-to-end intelligent driving method based on artificial intelligence provided by an embodiment of the present invention;

[0054] Figure 2 It is a schematic structural diagram of a multi-branch hybrid neural network provided by an embodiment of the present invention;

[0055] Figure 3 It is a structural block diagram of an end-to-end intelligent driving system based on artificial intelligence provided by an embodiment of the present invention;

[0056] Figure 4 It is a structural block diagram of an end-to-end intelligent driving device based on artificial intelligence provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] The embodiments of the present invention are described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present invention, and should not be construed as a limitation to the present invention. For the step numbers in the following embodiments, they are only set for the convenience of explanation and illustration, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0058] In the description of the present invention, "a plurality of" means two or more. If the first and second are described, it is only for the purpose of distinguishing technical features and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features or implicitly specifying the sequence of the indicated technical features. In addition, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field of the present technology.

[0059] Referring to Figure 1 , an end-to-end intelligent driving method based on artificial intelligence is provided in an embodiment of the present invention, which specifically includes the following steps:

[0060] S101. Obtain the camera perception data and radar perception data of the target vehicle;

[0061] S102. Input the camera perception data and radar perception data into a pre-trained end-to-end intelligent driving model to obtain a target driving action;

[0062] S103. Control the target vehicle to execute the target driving action.

[0063] Specifically, the embodiment of the present invention realizes the direct mapping from vehicle perception data (camera data, radar data) to vehicle driving actions (such as steering, accelerating, braking) by using an end-to-end model, skipping the intermediate complex feature engineering and path planning steps, reducing the complexity and debugging difficulty of the vehicle intelligent driving system, and improving the adaptability of the vehicle intelligent driving system.

[0064] Further as an optional implementation manner, obtaining the camera perception data and radar perception data of the target vehicle specifically includes:

[0065] S1011. Obtain the first image data captured by the camera of the target vehicle and the first point cloud data detected by the radar;

[0066] S1012. Perform normalization, cropping, and data augmentation processing on the first image data to obtain camera perception data;

[0067] S1013. Perform filtering and coordinate transformation processing on the first point cloud data to obtain radar perception data.

[0068] Specifically, perform normalization (normalize the pixel values to the 0-1 interval), cropping (remove irrelevant parts in the image), and data augmentation (increase the diversity of data by means of rotation, flipping, brightness adjustment, etc.) on the first image data collected by the camera to obtain camera perception data; perform operations such as filtering (remove noise) and coordinate transformation on the first point cloud data detected by the radar to obtain radar perception data, so that the data can be better utilized by the algorithm.

[0069] Further as an optional implementation manner, the end-to-end intelligent driving model is trained through the following steps:

[0070] S201. Obtain the camera sample data and radar sample data of the test vehicle under different road conditions and different weather conditions, and determine the corresponding driving action labels through manual annotation;

[0071] S202. Input the camera sample data and radar sample data into a pre-constructed multi-branch hybrid neural network to obtain the predicted driving actions;

[0072] S203. Determine the loss value according to the predicted driving actions and the driving action labels;

[0073] S204. Update the parameters of the multi-branch hybrid neural network through the stochastic gradient descent algorithm according to the loss value to obtain the trained end-to-end intelligent driving model.

[0074] Specifically, collect the camera sample data and radar sample data under different road conditions (such as urban roads, highways, rural roads, etc.) and different weather conditions (such as sunny days, rainy days, snowy days, etc.), and manually determine the corresponding correct driving actions to obtain the driving action labels. It should be noted that the camera sample data and radar sample data also need to be preprocessed, and the specific process is similar to the real-time camera perception data and radar perception data, which will not be elaborated here.

[0075] Input the camera sample data and radar sample data into a pre-constructed multi-branch hybrid neural network. Feature extraction is performed on the camera sample data and radar sample data through different branch networks respectively, so as to learn the mapping logic of the camera sample data and radar sample data and output the corresponding predicted driving actions; measure the difference between the model output and the correct driving action through a preset loss function to obtain the loss value. By minimizing the loss value, the output of the model gets closer and closer to the correct driving action. Use the stochastic gradient descent (SGD) or other algorithms such as AdaGrad, Adadelta, Adam, etc. to update according to the gradient of the model parameters with respect to the loss value, so that the model gradually converges to a better state; adjust some hyperparameters (such as learning rate, number of network layers, number of neurons, etc.) to optimize the performance of the model. At the same time, use a validation set (a part of the data divided from the dataset) to verify the performance of the model to prevent overfitting. Finally, the trained end-to-end intelligent driving model can be obtained.

[0076] The multi-branch hybrid neural network of the embodiment of the present invention will be further introduced and described below.

[0077] As a further optional implementation, the multi-branch hybrid neural network includes an input layer, a CNN branch network, an LSTM branch network, a feature fusion layer, and a fully connected layer. The CNN branch network includes multiple convolutional layers and a spatial pyramid pooling layer for extracting spatial features of camera sample data. The LSTM branch network includes a one-dimensional convolutional layer and a bidirectional LSTM layer for extracting temporal features of radar sample data.

[0078] Specifically, as Figure 2 shown in the schematic diagram of the structure of the multi-branch hybrid neural network provided by the embodiment of the present invention. It can be seen that the multi-branch hybrid neural network includes an input layer, a CNN branch network, an LSTM branch network, a feature fusion layer, and a fully connected layer. Among them, the input layer is used to receive preprocessed sensor data. For example, image data is input in the form of a pixel matrix, and radar data is input in the form of a point cloud or a distance matrix. The CNN branch network consists of multiple convolutional layers (specifically, it can be 3 layers, each containing 32 / 64 / 128 3x3 convolutional kernels) and a spatial pyramid pooling layer, which is used to extract the spatial features of camera sample data and output a 128-dimensional feature vector. The LSTM branch network consists of a one-dimensional convolutional layer and a bidirectional LSTM layer, which is used to extract the temporal features of radar sample data and output a 256-dimensional feature vector (concatenated forward and backward). The feature fusion layer is used to fuse the spatial features and temporal features to obtain fused features. The fully connected layer is used to map the fused features to the sample label space, thereby generating driving decisions, such as steering, accelerating, or braking actions.

[0079] Based on the above architecture, a multi-branch hybrid neural network is constructed, and parameters such as the number of layers of the network, the number of neurons in each layer, and the activation function are determined and initialized. Moreover, initial values are assigned to the weights and biases of the network, and thus a multi-branch hybrid neural network for end-to-end intelligent driving model training can be obtained.

[0080] As a further optional implementation, the step of inputting camera sample data and radar sample data into a pre-constructed multi-branch hybrid neural network to obtain predicted driving actions specifically includes:

[0081] S2021. Input the camera sample data into the CNN branch network through the input layer, and input the radar sample data into the LSTM branch network.

[0082] S2022. Perform multi-scale convolution processing on the camera sample data through multiple convolutional layers to obtain multi-scale feature subgraphs, and perform feature fusion on the multi-scale feature subgraphs through the spatial pyramid pooling layer to obtain spatial features.

[0083] S2023. Extract local temporal patterns from the radar sample data through a one-dimensional convolutional layer to obtain local temporal features, and capture the forward and backward temporal dependencies of the local temporal features through a bidirectional LSTM layer to obtain temporal features;

[0084] S2024. Perform feature fusion on the spatial features and temporal features through a feature fusion layer to obtain fused features;

[0085] S2025. Map the fused features to the sample label space through a fully connected layer to obtain the predicted driving actions.

[0086] Specifically, in the embodiments of the present invention, the features in the camera data and radar data are automatically learned through the stochastic gradient descent algorithm without the need for manual design of the feature extraction steps. Among them, the CNN branch network can automatically identify spatial features such as edges, textures, and objects in the camera data, and the LSTM branch network can automatically identify temporal features such as change trends and forward and backward dependencies in the radar data. Therefore, the model can directly map from the original data to the driving decision, skipping the traditional feature extraction and path planning steps. This end-to-end learning method simplifies the system design, reduces the debugging difficulty, and improves the adaptability of the model to different scenarios.

[0087] Further as an optional implementation manner, perform feature fusion on the spatial features and temporal features through a feature fusion layer to obtain fused features, which specifically includes:

[0088] S20241. Determine the attention weight matrices of the spatial features and temporal features based on the multi-head self-attention mechanism through the feature fusion layer;

[0089] S20242. Perform multi-head splicing and linear projection on each dimension of the spatial features and temporal features according to the attention weight matrices to obtain fused features.

[0090] Specifically, feature fusion can adopt various methods such as splicing, attention fusion, and gated fusion. In the embodiments of the present invention, considering the need to dynamically adjust the importance of the spatial features and temporal features according to the real-time scenario, the attention fusion method is adopted for feature fusion.

[0091] First, perform feature alignment and embedding on the spatial features output by the CNN branch network and the temporal features output by the LSTM branch network to obtain a joint feature matrix; then perform multi-head division, that is, split the joint feature matrix into h independent heads, X = [Head 1 ; Head 2 ;...; Head h , and each independent head focuses on different spatio-temporal interaction patterns; then calculate Q i = Head i WQ , K i = Head i W K , V i = Head i W V , generate corresponding triples (Q i , K i , V i ) for each independent head, where the parameter matrices W Q , W K , W V are learnable; calculate the spatio-temporal attention scores for each independent head based on the triples and a preset mask matrix, and the formula is as follows:

[0092]

[0093] where, Attention i represents the spatio-temporal attention score of the i-th independent head, and M represents the mask matrix.

[0094] Perform intra-head feature weighting on each independent head according to the spatio-temporal attention scores, and output Z i = Attention i V i , and then perform multi-head concatenation and linear projection to obtain the output fused feature Z = Concat(Z 1 ,..., Z h )W O , where, W O represents the projection matrix.

[0095] The above describes the training process of the end-to-end intelligent driving model. Input the camera perception data and radar perception data of the target vehicle into the trained end-to-end intelligent driving model, and the target driving actions inferred by the model can be obtained.

[0096] Further, as an optional implementation manner, control the target vehicle to execute the target driving action, which specifically includes:

[0097] S1031. Generate a corresponding driving control instruction according to the target driving action;

[0098] S1032. Execute the driving control instruction through the driving control system of the target vehicle, so that the target vehicle believes in the target driving action.

[0099] Specifically, the trained end-to-end intelligent driving model is deployed on the vehicle's computing platform, which receives sensor data in real time and generates target driving actions (such as steering, accelerating, braking). Then, driving control instructions executable by the motor are generated based on the target driving actions and executed through the vehicle's driving control system, thereby realizing autonomous driving.

[0100] The method steps of the embodiments of the present invention are described above. It can be understood that the embodiments of the present invention utilize an end-to-end model to achieve a direct mapping from vehicle perception data (camera data, radar data) to vehicle driving actions (such as steering, accelerating, braking), skipping the intermediate complex feature engineering and path planning steps, reducing the complexity and debugging difficulty of the vehicle intelligent driving system, and improving the adaptability of the vehicle intelligent driving system.

[0101] Compared with the prior art, the embodiments of the present invention also have the following advantages:

[0102] 1) Simplify system complexity: Regarding the entire driving task as a direct mapping from input (such as sensor data) to output (such as vehicle driving actions) greatly simplifies the system architecture, reduces the coupling degree between modules, and lowers the difficulty of system design and debugging.

[0103] 2) Better adaptability and flexibility: Without the need to manually design the rules and algorithms for each module, in the face of new driving scenarios, road conditions, or traffic rules, the system can quickly adapt by updating data and retraining.

[0104] 3) Have self-learning and evolution capabilities: New data can be continuously collected during actual operation, and the model can be updated through online learning or incremental learning, thereby continuously evolving and optimizing its own driving capabilities to adapt to the changing traffic environment and driving requirements.

[0105] Referring to Figure 3 , the embodiments of the present invention provide an end-to-end intelligent driving system based on artificial intelligence, including:

[0106] A data acquisition module for acquiring camera perception data and radar perception data of the target vehicle;

[0107] A driving decision-making module for inputting the camera perception data and radar perception data into a pre-trained end-to-end intelligent driving model to obtain target driving actions;

[0108] An action execution module for controlling the target vehicle to execute the target driving actions.

[0109] The content in the above method embodiments is applicable to the system embodiments of the present invention. The functions specifically implemented in the system embodiments of the present invention are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0110] Referring to Figure 4 , an end-to-end intelligent driving device based on artificial intelligence is provided in an embodiment of the present invention, including:

[0111] At least one processor;

[0112] At least one memory for storing at least one program;

[0113] When the above at least one program is executed by the above at least one processor, the above at least one processor implements the above-mentioned end-to-end intelligent driving method based on artificial intelligence.

[0114] The content in the above method embodiments is applicable to the device embodiments of the present invention. The functions specifically implemented in the device embodiments of the present invention are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0115] An embodiment of the present invention further provides a computer-readable storage medium, in which a program executable by a processor is stored. The program executable by the processor is used to execute the above-mentioned end-to-end intelligent driving method based on artificial intelligence when executed by the processor.

[0116] A computer-readable storage medium according to an embodiment of the present invention can execute an end-to-end intelligent driving method provided in an embodiment of the method of the present invention, can execute any combination of implementation steps of the method embodiment, and has the corresponding functions and beneficial effects of the method.

[0117] An embodiment of the present invention also discloses a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes Figure 1 the method shown.

[0118] In some alternative embodiments, the functions / operations recited in the block diagrams may not occur in the order presented in the operational illustrations. For example, depending on the functions / operations involved, two blocks shown in succession may actually be executed substantially simultaneously or the blocks may sometimes be executed in reverse order. Further, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for the purpose of providing a more thorough understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated in which the order of various operations is altered and in which sub-operations described as part of a larger operation are performed independently.

[0119] In addition, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the above functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for an understanding of the present invention. Rather, given the attributes, functions and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skill of an engineer. Thus, those of ordinary skill in the art can implement the present invention as set forth in the claims without undue experimentation. It should also be understood that the particular concepts disclosed are illustrative only and are not intended to limit the scope of the present invention, the scope of which is determined by the full scope of the appended claims and their equivalents.

[0120] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product stored in a storage medium, including several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the above methods of the various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as a USB flash drive, a portable hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.

[0121] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definable sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in conjunction with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0122] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the above program can be printed, because the above program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.

[0123] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having suitable combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0124] In the above description of this specification, the descriptions referring to terms such as "one embodiment / Example", "another embodiment / Example", or "certain embodiments / Examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0125] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art will understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.

[0126] The above is a specific description of the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included in the scope defined by the claims of this application.

Claims

1. An end-to-end intelligent driving method based on artificial intelligence, characterized in that: The following steps are involved: Obtain camera perception data and radar perception data of the target vehicle; Inputting the camera perception data and the radar perception data into a pre-trained end-to-end intelligent driving model to obtain a target driving action; The target vehicle is controlled to perform the target driving action.

2. The end-to-end intelligent driving method based on artificial intelligence according to claim 1, characterized in that: The step of acquiring the camera perception data and radar perception data of the target vehicle specifically includes: Acquire first image data captured by a camera of the target vehicle and first point cloud data detected by a radar; Normalizing, cropping, and data enhancement processing are performed on the first image data to obtain the camera perception data; The first point cloud data is filtered and coordinate transformed to obtain the radar perception data.

3. The end-to-end intelligent driving method based on artificial intelligence according to claim 1, characterized in that: The end-to-end intelligent driving model is trained by the following steps: Obtain camera sample data and radar sample data of the test vehicle under different road conditions and weather conditions, and determine the corresponding driving action labels through manual annotation; Inputting the camera sample data and the radar sample data into a pre-built multi-branch hybrid neural network to obtain a predicted driving action; Determining a loss value according to the predicted driving action and the driving action label; The parameters of the multi-branch hybrid neural network are updated according to the loss value through a stochastic gradient descent algorithm to obtain the trained end-to-end intelligent driving model.

4. The end-to-end intelligent driving method based on artificial intelligence according to claim 3, characterized in that: The multi-branch hybrid neural network includes an input layer, a CNN branch network, an LSTM branch network, a feature fusion layer and a fully connected layer. The CNN branch network includes multi-layer convolution layers and spatial pyramid pooling layers for extracting the spatial features of the camera sample data. The LSTM branch network includes a one-dimensional convolution layer and a bidirectional LSTM layer for extracting the temporal features of the radar sample data.

5. The end-to-end intelligent driving method based on artificial intelligence according to claim 4, characterized in that: The step of inputting the camera sample data and the radar sample data into a pre-built multi-branch hybrid neural network to obtain a predicted driving action specifically includes: Input the camera sample data into the CNN branch network through the input layer, and input the radar sample data into the LSTM branch network; Performing multi-scale convolution processing on the camera sample data through the multi-layer convolution layer to obtain a multi-scale feature subgraph, and performing feature fusion on the multi-scale feature subgraph through the spatial pyramid pooling layer to obtain the spatial feature; Extracting local time series patterns from the radar sample data through the one-dimensional convolution layer to obtain local time series features, and capturing forward and backward time series dependencies of the local time series features through the bidirectional LSTM layer to obtain the time series features; The spatial feature and the temporal feature are fused by the feature fusion layer to obtain a fused feature; The fused features are mapped to a sample label space through the fully connected layer to obtain the predicted driving action.

6. The end-to-end intelligent driving method based on artificial intelligence according to claim 5, characterized in that: The feature fusion layer is used to fuse the spatial feature and the temporal feature to obtain a fused feature, which specifically includes: Determine the attention weight matrix of the spatial feature and the temporal feature based on the multi-head self-attention mechanism through the feature fusion layer; According to the attention weight matrix, multiple stitching and linear projection are performed on each dimension of the spatial feature and the temporal feature to obtain the fusion feature.

7. An end-to-end intelligent driving method based on artificial intelligence according to any one of claims 1 to 6, characterized in that: The controlling the target vehicle to perform the target driving action specifically includes: generating a corresponding driving control instruction according to the target driving action; The driving control instruction is executed by a driving control system of the target vehicle so that the target vehicle believes in the target driving action.

8. An end-to-end intelligent driving system based on artificial intelligence, characterized in that: include: A data acquisition module, used to acquire camera perception data and radar perception data of the target vehicle; A driving decision module, used for inputting the camera perception data and the radar perception data into a pre-trained end-to-end intelligent driving model to obtain a target driving action; An action execution module is used to control the target vehicle to execute the target driving action.

9. An end-to-end intelligent driving device based on artificial intelligence, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements an end-to-end intelligent driving method based on artificial intelligence as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a program executable by a processor, characterized in that: The processor-executable program is used to execute an end-to-end intelligent driving method based on artificial intelligence as described in any one of claims 1 to 7 when executed by the processor.

Citation Information

Cited By

  • Vehicle control method, device, equipment and medium

    CN120792858A