Intelligent driving behavior method based on machine learning

Through the intelligent driving behavior method based on machine learning, sensor data and reinforcement learning algorithms are used to optimize environmental perception and decision-making, the problem of insufficient flexibility and adaptability of traditional autonomous driving systems in complex environments is solved, and the accuracy and response speed of the intelligent driving system are improved.

CN120246009AActive Publication Date: 2025-07-04DALIAN UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510300241.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-04
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

Traditional autonomous driving systems are difficult to achieve the flexibility and adaptability of human drivers in dynamically changing real scenarios. Especially in complex urban driving environments, AI systems are difficult to effectively identify and process multiple situations, resulting in insufficient algorithm accuracy and response speed.

Method used

Using an intelligent driving behavior method based on machine learning, we use the sensor data and labeled data of real road scenes to enhance image and point cloud data, build feature vectors, and use reinforcement learning algorithms to select driving scenarios and generate driving behaviors, and optimize environmental perception and decision-making capabilities.

Benefits of technology

It improves the adaptability and reliability of intelligent driving systems in complex environments, and achieves higher algorithm accuracy and response speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120246009A_ABST
    Figure CN120246009A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent driving behavior method based on machine learning, and the method comprises the steps: obtaining basic data, including driving behavior data in a real road scene, calculating a feature vector of a driving scene according to the basic data, and selecting the driving scene based on the feature vector through a reinforcement learning algorithm. Driving behaviors in the driving scene are generated through a reinforcement learning algorithm based on the feature vectors, and finally strategy updating is carried out to obtain an optimal model method. The purposes of optimizing the environment perception and decision-making capability of the intelligent driving system and improving the adaptability and reliability in an actual driving scene are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of information technology, and specifically relates to an intelligent driving behavior method based on machine learning. Background Art

[0002] In the context of the rapid development of the automotive industry today, intelligent driving technology is attracting increasing attention. With the continuous progress of artificial intelligence (AI) technology, especially the breakthroughs in computer vision, deep learning, and sensor technology, automotive manufacturers, technology companies, and research institutions are competing to develop intelligent vehicles with autonomous driving capabilities. These technological advancements enable vehicles to make real-time decisions in complex driving environments, improving driving safety and efficiency. However, despite the broad prospects of intelligent driving technology, there are still many challenges in practical applications.

[0003] The complexity of driving scenarios is one of the main difficulties faced by intelligent driving technology. Drivers need to pay attention to various environmental factors during driving, including road conditions, traffic signs, pedestrians, other vehicles, and emergencies. These factors pose severe challenges to the accuracy and timeliness of driving decisions. Although traditional autonomous driving systems have made some progress in perceiving and understanding the surrounding environment, they are still difficult to achieve the flexibility and adaptability of human drivers in dynamic real-world scenarios. In complex urban driving environments, AI systems must be able to effectively identify and process multiple situations, which places higher requirements on the accuracy and response speed of algorithms.

[0004] Therefore, in view of the above problems, the present invention proposes a relatively general intelligent driving behavior learning method based on machine learning algorithms, aiming to optimize the environmental perception and decision-making capabilities of intelligent driving systems, thereby improving their adaptability and reliability in actual driving scenarios.

[0005] The neural network in this method refers to Bacon, P.-L., Harb, J.,&Precup, D. (2017). The Option-Critic Architecture. Proceedings of the AAAI Conference on Artificial Intelligence, 31(1). https: / / arxiv.org / abs / 1609.05140. Summary of the Invention

[0006] The objective of the present invention is to solve the problem that traditional autonomous driving systems are difficult to achieve the flexibility and adaptability of human drivers in dynamically changing real-world scenarios. In complex urban driving environments, AI systems must be able to effectively identify and process various situations, which poses higher requirements for the accuracy and response speed of algorithms.

[0007] To solve the above problems, the present invention proposes an intelligent driving behavior method based on machine learning, including: S1: Obtain basic data; The basic data includes driving behavior data in real road scenarios, which contains 1000 driving scenarios, each driving scenario lasts for 20 seconds, and each driving scenario includes sensor data and annotation data; S2: Calculate the feature vectors of driving scenarios based on the basic data; S2-1: Load data; Load image data, load point cloud data, load the position of the vehicle itself; S2-2: Augment image data; Random cropping, color jittering; S2-3: Augment point cloud data; Random rotation, random translation; S2-4: Construct feature vectors; Extract image features, extract point cloud features; S2-5: Represent feature vectors; Concatenate image features, point cloud features, the state of the vehicle itself, and the states of surrounding objects into the final feature vector s; S3: Select driving scenarios based on the feature vectors through a reinforcement learning algorithm; S4: Generate driving behaviors in driving scenarios based on the feature vectors through a reinforcement learning algorithm; S5: Update the policy to obtain the optimal model method.

[0008] In a preferred embodiment, it includes: S1: Obtain basic data; The basic data includes driving behavior data in real road scenarios, which contains 1000 driving scenarios, each driving scenario lasts for 20 seconds, and each driving scenario includes sensor data and annotation data; The sensor data includes: Cameras: 6 cameras, namely front, rear, left, right, left front, and right front. The cameras provide RGB images I B H*W*3 B H*W*3 represents a three-dimensional tensor set of H*W*3, where the height H = 900 and the width W = 1600; LiDAR: One 32-line LiDAR, providing point cloud data P B N*3 , where N is the number of points, and B N*3 represents a set of 3D tensors of N * 3. Each point is (x, y, z, intensity, ring), where x, y, and z represent the coordinates of each point in 3D space, intensity represents the signal intensity reflected after the laser beam emitted by the LiDAR hits an object, and ring represents the beam number of the LiDAR; Radar: Five radars in the front, rear, left, right, and rear left, providing target detection data; GPS and IMU: Providing the ego-vehicle position p = [p x, p y , p z and speed v = [v x , v y , v z , representing the position p and speed v of each point in the three directions of the x, y, and z coordinates in 3D space; The annotation data includes: 3D bounding box: 3D annotations of 2 types of objects. Each object is represented as b = [x, y, z, w, l, h, θ], where x, y, and z are the 3D space coordinates of the center point of the 3D annotation, w represents the width of the bounding box, l represents the length of the bounding box, h represents the height of the bounding box, and θ represents the orientation angle; Attributes: The motion state S of the object, including: stationary, moving, behavior, and behavior includes parking and turning; Trajectory: The motion trajectory T of the object = [p1, p2,..., p i , where p i = [x i , y i , z i represents the 3D space coordinates of the object's location at the i-th moment; Map information: High-precision map data, including lanes and traffic signs; S2: Calculating the feature vector of the driving scenario based on the basic data; S2-1: Loading data; Loading image data: Obtaining the image I from the camera, with a shape of (H * W * 3); Loading point cloud data: Obtaining the point cloud data P from the LiDAR, with a shape of (N, 5), and each point is (x, y, z, intensity, ring); Loading the ego-vehicle position: The ego-vehicle position p = [p x, p y , p z and speed v = [v x , v y , vz ; S2-2: Image data augmentation; Random cropping: Randomly crop a region I from image I crop , with size (H C , W C ), and the formula is: I crop (x, y) = I(x + Δx, y + Δy) Δx ~ u(0, W - W C ) Δy ~ u(0, H - H C ) where Δx and Δy represent the coordinate offsets of the upper left corner of the cropping region, and the values are randomly sampled from a uniform distribution. u represents a uniform distribution, and W C、 H C represent the width and height of the cropped image; Color jittering: Randomly adjust the brightness, contrast, and saturation of image I to obtain enhanced image data I jitter , and the formula is: I jitter = T color (I crop ) where T color represents the color transformation function, including adjustments to brightness, contrast, and saturation. The specific derivation process includes: where α brightness 、 α contrast , α saturation represent randomly sampled values from a uniform distribution, and β brightness , β contrast , β saturation represent the adjustment ranges of brightness, contrast, and saturation. HSV2RGB and RGB2HSV represent the conversion functions between RGB and HSV; S2-3: Point cloud data augmentation; Random rotation: Randomly rotate the point cloud data P by an angle θ ~ u(-θ max, θ max ), and the formula is: Random translation: Randomly translate the point cloud data P rot to obtain enhanced point cloud data P trans , and the formula is: P trans= P rot + Δt Δt = [Δt x , Δty , Δt z where Δt ~ u(-t max , t max ) represents the translation amount, t max represents the maximum range of the translation operation, and Δt x , Δt y , Δt z represent the translation amounts in three directions in three-dimensional space respectively; S2-4: Construct the feature vector; Image feature extraction: Use a convolutional neural network to extract the image feature f image , and the formula is: f image = CNN(I jitter ) Point cloud feature extraction: Use PointNet to extract the point cloud feature f pointcloud , and the formula is: f pointcloud = PointNet(P trans ) The state s of the ego vehicle ego includes the position p and v, and the formula is: s ego = [p x, p y , p z , v x , v y , v z The state s of surrounding objects obj includes the 3D bounding box and attributes of the object, and the formula is: s obj = [b1, b2,..., b M b i = [x i , y i , z i , w i , l i , h i , θ i , v xi , v xi where M represents the number of objects in the surrounding space; S2-5: Feature vector representation; Concatenate the image feature, point cloud feature, ego vehicle state, and surrounding object state into the final feature vector s, and the formula is: s = [f image , f pointcloud , s​​​​ego , s obj S3: Select the driving scenario based on the feature vector through the reinforcement learning algorithm; S3-1: Define the driving scenario o ∈ {o1, o2, o3, o4}, where o1 is the urban road, o2 is the highway, o3 is the rural road, and o4 is the tunnel; Based on π option policy, select the most suitable current driving scenario o, and this policy is parameterized by using the neural network π option (o|s; θ option ) where θ option represents the neural network parameters of the upper-level policy. The training objective of the neural network is to maximize the cumulative reward, and the formula is: where, J option represents the cumulative reward, E represents the expected value, t represents the time step, T represents the maximum time step, γ represents the discount factor, and R represents the reward function. The specific formula of the reward function R is: where, w1, w2, and w3 represent the weight coefficients, R safety represents the safety reward, M represents the number of objects around the host vehicle, d i represents the distance between the host vehicle and the i-th object, d min represents the safety distance threshold, and R efficiency represents the efficiency reward, v represents the current vehicle speed, and v target represents the target speed, and R confort represents the comfort reward, and acceleration represents the current acceleration; S4: Generate the driving behavior in the driving scenario based on the feature vector through the reinforcement learning algorithm; Define the driving behavior a ∈ {a1, a2, a3, a4, a5}, where a1 represents the acceleration, a2 represents the braking force, a3 represents the throttle opening, a4 represents the steering angle, and a5 represents the steering speed; The objective of the policy π action is to select the optimal driving behavior a in the current driving scenario o. Use the neural network π action (a|s, o; θ action ) parameterized driving behavior policy, and output the optimal driving behavior a as the actual driving behavior of intelligent driving. Here, θ action represents the neural network parameters of the lower-level policy. The training objective of the neural network is to maximize the cumulative reward, and the formula is: where, J action ​C represents the cumulative reward, E represents the expected value, t represents the time step, T represents the maximum time step, γ represents the discount factor, and R represents the reward function; the termination function β determines whether the current driving scenario o terminates, and outputs termination as the basis for judging whether the current driving scenario terminates. If it is True, it terminates; if it is False, it continues the current driving scenario. The neural network β(termination|s; θ β parameterizes the termination function θ β Neural network parameters. The training objective is to maximize the cumulative reward, and the formula is: where J β represents the cumulative reward, E represents the expected value, t represents the time step, T represents the maximum time step, γ represents the discount factor, and R represents the reward function; S5: Policy update; According to the feature vector s, driving scenario o, driving behavior a, and reward R, use the gradient ascent method to update the parameters of the upper-level policy, lower-level policy, and termination function. The formula is: where α represents the learning rate, represents the gradient of the upper-level policy objective function with respect to the parameter J option , represents the gradient of the lower-level policy objective function with respect to the parameter J action , represents the gradient of the termination function objective function with respect to the parameter J β . Repeat data sampling and policy learning until the model converges to obtain the optimal model method.

[0009] In the preferred mode, the two types of objects include: vehicles, pedestrians, and bicycles.

[0010] The beneficial effects of the present invention: By analyzing the driving scenario, the decomposition of the intelligent driving task is realized, which is beneficial to the optimal allocation of resources for intelligent driving, accurate product delivery, and improvement of customer satisfaction; by using hierarchical reinforcement learning to decompose the intelligent driving scenario into multiple intelligent driving subtasks, using the upper-level policy to complete the selection of the driving task for the current scenario, using the lower-level policy to complete the solution of the driving task, and using the basic data to assist the training process, the utilization efficiency of data and the training effect of the model are improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 is a schematic diagram of the working process of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0012] Example 1: An intelligent driving behavior method based on machine learning includes: S1: Obtain basic data; The basic data includes driving behavior data in real road scenarios, which contains 1000 driving scenarios, each driving scenario lasts for 20 seconds, and each driving scenario includes sensor data and annotation data; S2: Calculate the feature vector of the driving scenario according to the basic data; S2-1: Load data; Load image data, load point cloud data, load the position of the vehicle itself; S2-2: Image data augmentation; Random cropping, color jittering; S2-3: Point cloud data augmentation; Random rotation, random translation; S2-4: Construct the feature vector; Image feature extraction, point cloud feature extraction; S2-5: Feature vector representation; Concatenate the image feature, point cloud feature, the state of the vehicle itself and the state of surrounding objects into the final feature vector s; S3: Select the driving scenario based on the feature vector through the reinforcement learning algorithm; S4: Generate the driving behavior in the driving scenario based on the feature vector through the reinforcement learning algorithm; S5: The method of obtaining the optimal model by policy update.

[0013] Including: S1: Obtain basic data; The basic data includes driving behavior data in real road scenarios, which contains 1000 driving scenarios, each driving scenario lasts for 20 seconds, and each driving scenario includes sensor data and annotation data; The sensor data includes: Camera: 6 cameras in the front, rear, left, right, left front, and right front. The camera provides RGB images I B H*W*3 and B H*W*3 represents a three-dimensional tensor set of H*W*3, where the height H = 900 and the width W = 1600; LiDAR: 1 32-line LiDAR, providing point cloud data P B N*3 and N is the number of points. B N*3 represents a three-dimensional tensor set of N*3, and each point is (x, y, z, intensity, ring). x, y, and z represent the coordinates of each point in three-dimensional space, intensity represents the signal intensity reflected back after the laser beam emitted by the LiDAR hits an object, and ring represents the beam number of the LiDAR; Radar: Five radars, namely front, rear, left, right, and rear left, provide target detection data; GPS and IMU: Provide the ego vehicle position p = [p x, p y p z and speed v = [v x v y v z , representing the position p and speed v of each point in the three directions of the x, y, and z coordinates in three-dimensional space; The labeled data includes: 3D bounding box: 3D labels of two types of objects, and each object is represented as b = [x, y, z, w, l, h, θ], where x, y, and z are the three-dimensional space coordinates of the center point of the 3D label, w represents the width of the bounding box, l represents the length of the bounding box, h represents the height of the bounding box, and θ represents the orientation angle; Attributes: The motion state S of the object, including: stationary, moving, behavior, and behavior includes parking and turning; Trajectory: The motion trajectory T of the object = [p1, p2,..., p i p i = [x i y i z i represents the three-dimensional space coordinates of the object's position at the i-th moment; Map information: High-precision map data, including lanes and traffic signs; S2: Calculate the feature vector of the driving scenario based on the basic data; S2-1: Load data; Load image data: Obtain the image I from the camera, with the shape of (H*W*3); Load point cloud data: Obtain the point cloud data P from the lidar, with the shape of (N, 5), and each point is (x, y, z, intensity, ring); Load ego vehicle position: The ego vehicle position p = [p x, p y p z and speed v = [v x v y v z ; S2-2: Image data augmentation; Random cropping: Randomly crop a region I crop from the image I, with the size of (H C W C ), and the formula is: I crop (x, y) = I(x + Δx, y + Δy) Δx ~ u(0, W - WC ) Δy ~ u(0, H - H C ) where Δx and Δy represent the coordinate offset of the upper - left corner of the cropping area, and the values are randomly sampled from a uniform distribution. u represents a uniform distribution, and W C、 H C represent the width and height of the cropped image; Color jittering: Randomly adjust the brightness, contrast, and saturation of the image I to obtain the image enhancement data I jitter , and the formula is: I jitter = T color (I crop ) where T color represents the color transformation function, including the adjustment of brightness, contrast, and saturation. The specific derivation process is as follows: where α brightness 、 α contrast 、α saturation represent randomly sampled values from a uniform distribution, and β brightness 、β contrast 、β saturation represent the adjustment ranges of brightness, contrast, and saturation. HSV2RGB and RGB2HSV represent the conversion functions between RGB and HSV; S2 - 3: Point cloud data enhancement; Random rotation: Randomly rotate the point cloud data P, and the rotation angle θ ~ u(-θ max, θ max ), and the formula is: Random translation: Randomly translate the point cloud data P rot to obtain the enhanced point cloud data P trans , and the formula is: P trans= P rot +Δt Δt = [Δt x , Δt y , Δt z where Δt ~ u(-t max , t max ) represents the translation amount, and t max represents the maximum range of the translation operation. Δt x 、Δt y 、Δt z represent the translation amounts in three directions in the three - dimensional space respectively; ​S2-4: Construct feature vectors; Image feature extraction: Use a convolutional neural network to extract the image feature f image , the formula is: f image = CNN(I jitter ) Point cloud feature extraction: Use PointNet to extract the point cloud feature f pointcloud , the formula is: f pointcloud = PointNet(P trans ) The state s of the ego vehicle ego includes the position p and v, the formula is: s ego = [p x, p y , p z , v x , v y , v z The state s of surrounding objects obj includes the 3D bounding box and attributes of the objects, the formula is: s obj = [b1, b2,..., b M b i = [x i , y i , z i , w i , l i , h i , θ i , v xi , v xi where M represents the number of objects in the surrounding space; S2-5: Feature vector representation; Concatenate the image feature, point cloud feature, ego vehicle state, and surrounding object state into the final feature vector s, the formula is: s = [f image , f pointcloud , s ego , s obj S3: Select the driving scenario based on the feature vector through the reinforcement learning algorithm; S3-1: Define the driving scenarios o ∈ {o1, o2, o3, o4}, where o1 is the urban road, o2 is the highway, o3 is the rural road, and o4 is the tunnel; Based on π option policy, select the most suitable current driving scenario o, and this policy is to use the neural network πoption (o|s; θ option ) is parameterized, where θ option represents the neural network parameters of the upper-level policy. The training objective of the neural network is to maximize the cumulative reward, and the formula is: where J option represents the cumulative reward, E represents the expected value, t represents the time step, T represents the maximum time step, γ represents the discount factor, and R represents the reward function. The specific formula of the reward function R is: where w1, w2, and w3 represent the weight coefficients, and R safety represents the safety reward, M represents the number of objects around the ego vehicle, d i represents the distance between the ego vehicle and the i-th object, and d min represents the safety distance threshold, and R efficiency represents the efficiency reward, v represents the current vehicle speed, and v target represents the target speed, and R confort represents the comfort reward, and acceleration represents the current acceleration; S4: Generate driving behaviors in the driving scenario through a reinforcement learning algorithm based on the feature vector; Define driving behavior a ∈ {a1, a2, a3, a4, a5}, where a1 represents the acceleration, a2 represents the braking force, a3 represents the throttle opening, a4 represents the steering angle, and a5 represents the steering speed; The objective of the policy π action is to select the optimal driving behavior a in the current driving scenario o, and use the neural network π action (a|s, o; θ action ) to parameterize the driving behavior policy, and output the optimal driving behavior a as the actual driving behavior of intelligent driving. Here, θ action represents the neural network parameters of the lower-level policy. The training objective of the neural network is to maximize the cumulative reward, and the formula is: where J action represents the cumulative reward, E represents the expected value, t represents the time step, T represents the maximum time step, γ represents the discount factor, and R represents the reward function; the termination function β determines whether the current driving scenario o terminates, and outputs termination as the basis for judging whether the current driving scenario terminates. If it is True, it terminates; if it is False, it continues the current driving scenario. Use the neural network β(termination|s; θ β ) to parameterize the termination function θ βThe neural network parameters, with the training objective of maximizing the cumulative reward, and the formula is: where J β represents the cumulative reward, E represents the expected value, t represents the time step, T represents the maximum time step, γ represents the discount factor, and R represents the reward function; S5: Policy update; According to the feature vector s, driving scenario o, driving behavior a, and reward R, use the gradient ascent method to update the parameters of the upper-level policy, lower-level policy, and termination function. The formula is: where α represents the learning rate, represents the gradient of the upper-level policy objective function with respect to the parameter J option , represents the gradient of the lower-level policy objective function with respect to the parameter J action , represents the gradient of the termination function objective function with respect to the parameter J β . Repeat data sampling and policy learning until the model converges to obtain the optimal model method. The two types of objects include: vehicles, pedestrians, and bicycles.

[0014] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.

Claims

1. An intelligent driving behavior method based on machine learning, characterized in that, Including: S1: Obtain basic data; The basic data includes driving behavior data in real road scenarios, which contains 1000 driving scenarios, each driving scenario lasts for 20 seconds, and each driving scenario includes sensor data and annotation data; S2: Calculate the feature vector of the driving scenario according to the basic data; S2-1: Load data; Load image data, load point cloud data, load the position of the ego vehicle; S2-2: Augment the image data; Random cropping, color jittering; S2-3: Augment the point cloud data; Random rotation, random translation; S2-4: Construct the feature vector; Extract image features, extract point cloud features; S2-5: Represent the feature vector; Concatenate the image features, point cloud features, ego vehicle state, and surrounding object states into the final feature vector s; S3: Select the driving scenario based on the feature vector through the reinforcement learning algorithm; S4: Generate the driving behavior in the driving scenario based on the feature vector through the reinforcement learning algorithm; S5: Method for updating the policy to obtain the optimal model.

2. The intelligent driving behavior method based on machine learning according to claim 1, characterized in that, Including: S1: Obtain basic data; The basic data includes driving behavior data in real road scenarios, which contains 1000 driving scenarios, each driving scenario lasts for 20 seconds, and each driving scenario includes sensor data and annotation data; The sensor data includes: Camera: six cameras, namely front, rear, left, right, left front, and right front cameras, which provide RGB images I B H*W*3 , B H*W*3 represents a set of three-dimensional tensors of H*W*3, where the height H = 900 and the width W = 1600; LiDAR: One 32-line LiDAR, providing point cloud data P B N*3 , where N is the number of points, and B N*3 represents a set of three-dimensional tensors of N * 3. Each point is (x, y, z, intensity, ring). x, y, and z represent the coordinates of each point in three-dimensional space. Intensity represents the signal intensity of the laser beam emitted by the LiDAR after hitting an object and reflecting back. Ring represents the line number of the LiDAR; Radar: Five radars in the front, rear, left, right, and rear left, providing target detection data; GPS and IMU: Provide the position p = [p x, p y ,p z and velocity v = [v x ,v y ,v z , representing the position p and velocity v of each point in the three directions of the x, y, and z coordinates in three-dimensional space; The annotation data includes: 3D bounding box: 3D annotations of 2 types of objects, each object is represented as b = [x, y, z, w, l, h, θ], where x, y, z are the three-dimensional spatial coordinates of the 3D annotation center point, w represents the width of the bounding box, l represents the length of the bounding box, h represents the height of the bounding box, and θ represents the orientation angle; Attribute: The motion state S of the object, including: stationary, moving, behavior, and behavior includes parking, turning; Trajectory: The motion trajectory T of an object = [p1, p2,..., p i , where p i = [x i , y i , z i represents the three-dimensional spatial coordinates of the position of the object at the i-th moment; Map information: High-precision map data, including lanes, traffic signs; S2: Calculate the feature vector of the driving scenario according to the basic data; S2-1: Load data; Load image data: Obtain the image I from the camera, with the shape of (H*W*3); Load point cloud data: Obtain the point cloud data P from the lidar, with the shape of (N, 5), and each point is (x, y, z, intensity, ring); Loaded ego vehicle position: ego vehicle position p = [p x, p y ,p z and speed v = [v x ,v y ,v z ; S2-2: Augment the image data; Random cropping: Randomly crop a region I from image I crop , with size (H C , W C ), and the formula is: I crop (x, y) = I(x + Δx, y + Δy) Δx ~ u(0, W - W C ) Δy ~ u(0, H - H C ) Among them, Δx and Δy represent the coordinate offsets of the upper left corner of the cropping area, and the values are randomly sampled from a uniform distribution. u represents a uniform distribution, and W C、 H C represent the width and height of the cropped image; Color jittering: randomly adjust the brightness, contrast, and saturation of image I to obtain image enhancement data I jitter , and the formula is: I jitter =T color (I crop ) Among them, T color represents a color transformation function, including adjustments of brightness, contrast, and saturation. The specific derivation process includes: Among them, α brightness 、 α contrast and α saturation represent uniformly distributed random sampling, β brightness and β contrast and β saturation represent the adjustment ranges of brightness, contrast, and saturation, and HSV2RGB and RGB2HSV represent the conversion functions between RGB and HSV; S2-3: Augment the point cloud data; Random rotation: Randomly rotate the point cloud data P, and the rotation angle θ ~ u(-θ max, θ max ), and the formula is: Random translation: For the point cloud data P rot perform random translation to obtain the enhanced point cloud data P trans , and the formula is: P trans= P rot +Δt Δt = [Δt x , Δt y , Δt z ​ where, Δt ~ u(-t max , t max ) represents the translation amount, t max represents the maximum range of the translation operation, and Δt x , Δt y , Δt z represent the translation amounts in three directions in the three-dimensional space respectively; S2-4: Construct the feature vector; Image feature extraction: Use a convolutional neural network to extract the image feature f image , and the formula is: f image =CNN(I jitter ) Point cloud feature extraction: Use PointNet to extract point cloud feature f pointcloud , the formula is: f pointcloud =PointNet (P trans ) Self-vehicle state s ego including position p and v, with the formula: s ego =[p x, p y ,p z ,v x ,v y ,v z ​ Surrounding object state s obj It includes the 3D bounding box and attributes of the object, and the formula is: s obj = [b1, b2,..., b M ​ b i =[x i ,y i ,z i ,w i ,l i ,h i ,θ i ,v xi ,v xi ​ Where M represents the number of objects in the surrounding space; S2-5: Represent the feature vector; Concatenate the image features, point cloud features, ego vehicle state, and surrounding object states into the final feature vector s, and the formula is: s = [f image , f pointcloud , s ego , s obj ​ S3: Select the driving scenario based on the feature vector through the reinforcement learning algorithm; S3-1: Define the driving scenario o ∈ {o1, o2, o3, o4}, o1 is the urban road, o2 is the highway, o3 is the rural road, o4 is the tunnel; Based on π option The policy selects the most suitable current driving scenario o, and this policy is to use the neural network π option (o|s; θ option ) is parameterized, where θ option represents the neural network parameters of the upper-level policy. The training objective of the neural network is to maximize the cumulative reward, and the formula is: Among them, J option represents the cumulative reward, E represents the expected value, t represents the time step, T represents the maximum time step, γ represents the discount factor, and R represents the reward function. The specific formula of the reward function R is as follows: Among them, w1, w2, and w3 represent weight coefficients, and R safety represents the safety reward, M represents the number of objects around the vehicle itself, and d i represents the distance between the vehicle itself and the i-th object, and d min represents the safety distance threshold, and R efficiency represents the efficiency reward, v represents the current vehicle speed, and v target represents the target speed, and R confort represents the comfort reward, and acceleration represents the current acceleration; S4: Generate the driving behavior in the driving scenario based on the feature vector through the reinforcement learning algorithm; Define the driving behavior a ∈ {a1, a2, a3, a4, a5}, a1 represents the acceleration, a2 represents the braking force, a3 represents the throttle opening, a4 represents the steering angle, a5 represents the steering speed; Policy π action aims to select the optimal driving behavior a in the current driving scenario o, using the neural network π action (a|s, o; θ action ) to parameterize the driving behavior policy, and output the optimal driving behavior a as the actual driving behavior of intelligent driving, where θ action represents the neural network parameters of the lower-level policy. The training objective of the neural network is to maximize the cumulative reward, and the formula is: Among them, J action represents the cumulative reward, E represents the expected value, t represents the time step, T represents the maximum time step, γ represents the discount factor, and R represents the reward function; the termination function β determines whether the current driving scenario o terminates, and outputs termination as the basis for judging whether the current driving scenario terminates. If it is True, it terminates. If it is False, it continues the current driving scenario. The neural network β(termination|s; θ β ) parameterizes the termination function θ β Neural network parameters. The training objective is to maximize the cumulative reward, and the formula is: Among them, J β represents the cumulative reward, E represents the expected value, t represents the time step, T represents the maximum time step, γ represents the discount factor, and R represents the reward function; S5: Policy update; Based on the feature vector s, driving scenario o, driving behavior a, and reward R, use the gradient ascent method to update the parameters of the upper-layer policy, lower-layer policy, and termination function. The formula is as follows: where α represents the learning rate, represents the gradient of the upper-layer policy objective function with respect to parameter J option of, represents the gradient of the lower-layer policy objective function with respect to parameter J action of, represents the gradient of the termination function objective function with respect to parameter J β of. Repeat data sampling and policy learning until the model converges to obtain the optimal model method.

3. The intelligent driving behavior method based on machine learning according to claim 2, wherein The two types of objects include: vehicles, pedestrians, and bicycles.

Citation Information

Patent Citations

  • Multi-mode reinforcement learning vehicle decision planning method with compensation feedback

    CN118917179A

  • Behavior decision optimization method and system based on intelligent driving

    CN119389224A

  • Method for combating stop-and-go wave problem using deep reinforcement learning based autonomous vehicles, recording medium and device for performing the method

    US20220363279A1