Automatic driving method and device based on multi-dimensional reward function

By constructing a multi-dimensional reward function autonomous driving method, combining imitation learning and reinforcement learning frameworks, multi-dimensional evaluation of driving behavior and action value judgment are achieved, the shortcomings of traditional autonomous driving systems in terms of safety and reliability are solved, and the safety and reliability of autonomous driving systems are improved.

CN120396996APending Publication Date: 2025-08-01ZHEJIANG WUWEN ZHIXING TECHNOLOGY CO LTD
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510370398.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing autonomous driving control methods rely on a single imitation learning or reinforcement learning framework and lack a multi-dimensional evaluation mechanism, resulting in insufficient safety and compliance of autonomous driving systems, and the design of reward function is not comprehensive enough, making it difficult to accurately judge the long-term benefits of driving actions.

Method used

Build an autonomous driving method based on multi-dimensional reward function. Through imitation learning and reinforcement learning frameworks, multiple cameras are used to collect environmental information, build a strategy generation network and discriminator network, design a multi-dimensional reward function model, conduct real-time detection and evaluation of behaviors such as running red lights, pressing lines, deviating from lanes, collisions, etc., build a driving behavior reward and punishment matrix, and generate the optimal driving action based on the actor's judgment network and evaluate the action value.

Benefits of technology

It significantly improves the safety and reliability of the autonomous driving system, can accurately evaluate driving behavior, generate reasonable reward signals to guide strategy learning, adapt to complex traffic environments, and achieve stable and reliable driving control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120396996A_ABST
    Figure CN120396996A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an automatic driving method and device based on a multi-dimensional reward function, and the method and device achieve the control of a driving motion through the combination of an imitation learning framework and a reinforcement learning framework, collection of environment information through a plurality of cameras, and construction of a strategy generation network and a discriminator network. A multi-dimensional reward function model is designed, behaviors such as red light running, line pressing, lane departure and collision are detected and evaluated in real time, and a driving behavior reward and punishment matrix is constructed. Based on an actor evaluation network architecture, environment information and a navigation instruction are input into an actor network to generate an optimal driving action, and the action value is evaluated through the evaluation network to realize dynamic parameter optimization. According to the method, the defects of the traditional technology in the aspects of driving behavior evaluation, action value judgment and the like are effectively overcome, and the safety and reliability of the automatic driving system are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of autonomous driving, and specifically to an autonomous driving method and device based on a multi-dimensional reward function. Background Art

[0002] Existing autonomous driving control methods have obvious deficiencies. Traditional methods mainly rely on a single imitation learning or reinforcement learning framework, lacking a multi-dimensional evaluation mechanism for driving behaviors, and it is difficult to ensure the safety and compliance of autonomous driving.

[0003] In addition, there are bottlenecks in the design of the reward function in the existing technology. Most systems use simple distance or time targets as reward signals, failing to fully consider multiple dimensions such as compliance with traffic rules and driving behavior norms, which affects the learning effect of the model.

[0004] Existing systems have technical shortcomings in action evaluation and policy optimization. Lack of an effective action value evaluation mechanism, it is difficult to accurately judge the long-term benefits of different driving actions. Solving these problems is of great significance for improving the safety and practicality of autonomous driving systems. Summary of the Invention

[0005] In view of the problems in the existing technology, this application provides an autonomous driving method and device based on a multi-dimensional reward function, which can effectively solve the deficiencies of traditional technologies in aspects such as driving behavior evaluation and action value judgment, and significantly improve the safety and reliability of autonomous driving systems.

[0006] To solve at least one of the above problems, this application provides the following technical solutions:

[0007] In a first aspect, this application provides an autonomous driving method based on a multi-dimensional reward function, including:

[0008] Construct an imitation learning autonomous driving model, input the environmental image information, vehicle navigation instruction information, and vehicle speed information collected by the in-vehicle front camera and left and right wide-angle cameras into the policy generation network. The policy generation network generates a driving action control signal, and input the driving action control signal and the current environmental state into the discriminator network. The discriminator network outputs the probability score that the action-state pair belongs to the expert data set, and trains the policy generation network based on the probability score; the driving action control signal includes a throttle control amount, a steering control amount, and a braking control amount;

[0009] Build a multi-dimensional reward function model, collect the state data during vehicle driving, detect the vehicle's red-light running behavior, lane-changing and line-crossing behavior, deviation from the motor vehicle lane behavior, and vehicle collision behavior, record the distance the vehicle has traveled and the remaining distance, construct a driving behavior reward and punishment matrix based on the state detection results, and input the driving behavior reward and punishment matrix into the multi-dimensional reward function model to obtain the reward score of the vehicle's current driving state;

[0010] Build an actor-critic network, input the environmental image information, navigation instruction information, and vehicle speed information into the actor network, the actor network outputs the optimal driving action in the current state, input the optimal driving action and the current state into the critic network, the critic network outputs the action value evaluation score, optimize the parameters of the actor network based on the reward score and the action value evaluation score, and use the optimized actor network to generate an autonomous driving control instruction.

[0011] Further, it also includes: collecting on-vehicle image information, installing a forward camera at the front of the vehicle, installing wide-angle cameras on both the left and right sides of the vehicle, calibrating the lens focal length, field of view angle, image resolution, and sampling frequency of the cameras, constructing the environmental image data collected by the calibrated cameras, the navigation instruction data generated by the on-vehicle navigation system, and the speed information collected by the vehicle speed sensor into state input data, and performing normalization processing on the state input data;

[0012] Input the normalized state input data into the policy generation network of the neural network structure, use the convolutional layer to extract image features, use the fully connected layer to fuse image features, navigation features, and speed features, generate an action output vector of the throttle control amount, steering control amount, and braking control amount based on the activation function, input the action output vector and the vehicle's current state into the discriminant model in the discriminator network, calculate the expert probability score of the state-action pair, and perform backpropagation optimization on the parameters of the policy generation network according to the expert probability score.

[0013] Further, it also includes: collecting demonstration data of expert drivers, recording the environmental state information and corresponding driving action data during expert driving, marking the expert driving data as positive samples, marking the driving action data output by the policy generation network as negative samples, constructing a training data set for the discriminator network, performing normalization processing on the throttle control amount, steering control amount, and braking control amount in the training data set, and splicing the normalized action data with the corresponding environmental state data to construct a state-action pair;

[0014] Input the state-action pair into the discriminator network. Use a multi-layer perceptron structure to perform feature mapping on the state-action pair, calculate the similarity score between the state-action pair and the expert data distribution, calculate the prediction error of the discriminator network based on the cross-entropy loss function, optimize the parameters of the discriminator network using the gradient descent method, and use the similarity score output by the discriminator network as the training reward signal for the policy generation network to perform iterative optimization training on the policy generation network.

[0015] Further, it also includes: constructing a vehicle state detection sub-model, collecting the traffic light state information of the lane where the vehicle is located, calculating the relative distance between the vehicle position and the traffic light position based on the electronic map data, collecting the relative position information between the vehicle contour and the lane line, collecting the distribution information of obstacles around the vehicle based on the lidar sensor, calculating the remaining distance between the vehicle and the target end point according to the global positioning system data, and constructing the vehicle state detection data into a state vector.

[0016] Extract features from the state vector, judge whether the vehicle runs a red light based on the traffic light state and the relative distance, judge whether there is a behavior of crossing the line or deviating from the lane based on the relative position between the vehicle contour and the lane line, judge whether a collision occurs based on the obstacle distribution information, calculate the proportion of the distance the vehicle has traveled to the total travel distance, construct a violation behavior label vector, and input the state vector and the violation behavior label vector into a multi-dimensional reward function model for training.

[0017] Further, it also includes: constructing a reward and punishment matrix according to the vehicle state detection results, assigning a violation punishment weight for the vehicle running a red light behavior, assigning a violation punishment weight for the vehicle crossing the line and changing lanes behavior, assigning a violation punishment weight for the vehicle deviating from the motor vehicle lane behavior, assigning a violation punishment weight for the vehicle collision behavior, assigning a task completion reward weight for the vehicle reaching the target position, and constructing a driving behavior reward and punishment matrix based on the value of the task completion reward weight.

[0018] Input the driving behavior reward and punishment matrix and the vehicle state detection vector into a multi-dimensional reward function model. Use a deep neural network structure to extract features from the state detection vector, calculate the reward and punishment scores of each driving behavior based on the weight values in the reward and punishment matrix, perform weighted summation on each reward and punishment score, and generate a comprehensive reward score reflecting the current driving state of the vehicle.

[0019] Further, it also includes: constructing an actor network model, using a multi-layer convolutional neural network to extract the visual features of the environmental image, using a fully connected layer to process the navigation instruction information and the vehicle speed information, performing feature fusion on the visual features, navigation features, and vehicle speed features, and mapping the fused features into an action space distribution through a policy head network to sample the optimal driving actions of the throttle control amount, steering control amount, and braking control amount based on the action space distribution.

[0020] Construct a judgment network model, splice the features of the optimal driving action and the current environmental state, perform a non-linear transformation on the state-action features using a multi-layer perceptron network, map the transformed features to a scalar value evaluation score based on the value head network, perform weighted fusion on the value evaluation score and the reward score output by the reward function model, construct a loss function for the actor network, and optimize and update the parameters of the actor network based on the policy gradient method.

[0021] Further, it also includes: constructing a parameter optimization model for the actor network using the proximal policy optimization algorithm, calculating the advantage function value by comparing the reward score output by the reward function model and the value evaluation score output by the judgment network, constructing a policy gradient based on the advantage function value, constraining the difference between the action distribution output by the actor network and the target policy distribution, using the trust region policy optimization method to limit the step size of each parameter update, and iteratively optimizing the parameters of the actor network based on the stochastic gradient descent method;

[0022] Deploy the optimized actor network to the vehicle-mounted computing platform, collect vehicle environmental image information, navigation instruction information, and vehicle speed information to construct a state input vector, input the state input vector into the optimized actor network, generate driving control instructions for the throttle control amount, steering control amount, and braking control amount based on the action output layer of the actor network, and send the driving control instructions to the vehicle execution system.

[0023] In a second aspect, the present application provides an autonomous driving device based on a multi-dimensional reward function, including:

[0024] A driving model construction module, used to construct an imitation learning autonomous driving model, input the environmental image information, vehicle navigation instruction information, and vehicle speed information collected by the vehicle-mounted front camera and left and right wide-angle cameras into the policy generation network, the policy generation network generates a driving action control signal, input the driving action control signal and the current environmental state into the discriminator network, the discriminator network outputs the probability score that the action-state pair belongs to the expert data set, and train the policy generation network based on the probability score; the driving action control signal includes the throttle control amount, the steering control amount, and the braking control amount;

[0025] A reward function construction module, used to construct a multi-dimensional reward function model, collect the state data during the vehicle driving process, detect the vehicle's red-light running behavior, lane-changing and line-crossing behavior, deviation from the motor vehicle lane behavior, and vehicle collision behavior, record the distance the vehicle has traveled and the remaining driving distance, construct a driving behavior reward and punishment matrix based on the state detection results, and input the driving behavior reward and punishment matrix into the multi-dimensional reward function model to obtain the reward score of the vehicle's current driving state;

[0026] An autonomous driving execution module is used to construct an actor-critic network, input the environmental image information, navigation instruction information, and vehicle speed information into the actor network. The actor network outputs the optimal driving action in the current state, inputs the optimal driving action and the current state into the critic network, and the critic network outputs an action value evaluation score. Based on the reward score and the action value evaluation score, the parameters of the actor network are optimized, and an autonomous driving control instruction is generated using the optimized actor network.

[0027] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the autonomous driving method based on the multi-dimensional reward function are implemented.

[0028] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the autonomous driving method based on the multi-dimensional reward function are implemented.

[0029] In a fifth aspect, the present application provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the autonomous driving method based on the multi-dimensional reward function are implemented.

[0030] As can be seen from the above technical solutions, the present application provides an autonomous driving method and device based on a multi-dimensional reward function. By combining an imitation learning and a reinforcement learning framework, environmental information is collected through multiple cameras, and a policy generation network and a discriminator network are constructed to achieve driving action control. A multi-dimensional reward function model is designed to detect and evaluate behaviors such as running a red light, crossing a line, deviating from a lane, and collision in real time, and a driving behavior reward and punishment matrix is constructed. Based on the actor-critic network architecture, environmental information and navigation instructions are input into the actor network to generate the optimal driving action, and the action value is evaluated through the critic network to achieve dynamic parameter optimization. This method effectively solves the deficiencies of traditional technologies in aspects such as driving behavior evaluation and action value judgment, and significantly improves the safety and reliability of the autonomous driving system. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0032] Figure 1 It is a flowchart of the autonomous driving method based on the multi-dimensional reward function in the embodiments of the present application;

[0033] Figure 2 It is a structural diagram of an autonomous driving device based on a multi-dimensional reward function in an embodiment of the present application;

[0034] Figure 3 It is a schematic structural diagram of an electronic device in an embodiment of the present application.

[0035] Reference numerals:

[0036] Electronic device 9600, central processing unit 9100, memory 9140, communication module 9110, input unit 9120, audio processor 9130, display 9160, power supply 9170, buffer memory 9141, application / function storage unit 9142, data storage unit 9143, driver program storage unit 9144, antenna 9111, speaker 9131, microphone 9132. Detailed implementation manners

[0037] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the scope of protection of the present application.

[0038] The acquisition, storage, use, processing, etc. of data in the technical solutions of the present application all comply with the relevant provisions of national laws and regulations.

[0039] Considering the problems existing in the prior art, the present application provides an autonomous driving method and device based on a multi-dimensional reward function. By combining an imitation learning and a reinforcement learning framework, environmental information is collected through multiple cameras, and a policy generation network and a discriminator network are constructed to achieve driving action control. A multi-dimensional reward function model is designed to detect and evaluate behaviors such as running a red light, crossing a line, deviating from a lane, and collision in real time, and a driving behavior reward and punishment matrix is constructed. Based on the actor-critic network architecture, environmental information and navigation instructions are input into the actor network to generate optimal driving actions, and the action value is evaluated through the critic network to achieve dynamic parameter optimization. This method effectively solves the deficiencies of the traditional technology in aspects such as driving behavior evaluation and action value judgment, and significantly improves the safety and reliability of the autonomous driving system.

[0040] In order to effectively solve the deficiencies of the traditional technology in aspects such as driving behavior evaluation and action value judgment, and significantly improve the safety and reliability of the autonomous driving system, the present application provides an embodiment of an autonomous driving method based on a multi-dimensional reward function. Refer to Figure 1, the autonomous driving method based on the multi-dimensional reward function specifically includes the following content:

[0041] Step S101: Construct an imitation learning autonomous driving model. Input the environmental image information, vehicle navigation instruction information, and vehicle speed information collected by the on-vehicle front camera and left and right wide-angle cameras into the policy generation network. The policy generation network generates a driving action control signal. Input the driving action control signal and the current environmental state into the discriminator network. The discriminator network outputs the probability score that the action-state pair belongs to the expert data set. Train the policy generation network based on the probability score; the driving action control signal includes the throttle control amount, steering control amount, and braking control amount.

[0042] Optionally, in this embodiment, an imitation learning autonomous driving model is constructed based on a deep learning framework. Install a high-definition camera with a 120-degree field of view in the front of the vehicle to collect the information of the front road environment; install a 170-degree ultra-wide-angle camera on each of the left and right sides to expand the field of view and improve the perception ability of adjacent lanes and blind spots. These three cameras form a multi-view perception system to jointly construct a comprehensive perception of the driving environment. The resolution of the image data collected by the camera is 1920×1080, and the sampling frequency is 30Hz to ensure the timeliness and clarity of the image information.

[0043] In this embodiment, an innovative policy generation network structure is designed. This network uses ResNet-50 as the backbone network for visual feature extraction, which includes 5 residual blocks. The number of convolutional layers in each residual block is 64, 128, 256, 512, and 1024 in sequence. The network input is three-way image data. After passing through independent feature extraction branches, feature fusion is performed in the feature fusion layer. The fusion method uses an attention mechanism, and the calculation formula is: F = Σ(wi*Fi), where Fi is the feature of each branch and wi is the dynamically calculated attention weight. This design can adaptively adjust the importance of information from each perspective according to different scenarios.

[0044] In this embodiment, the navigation instruction information provided by the on-vehicle navigation system is encoded. The navigation instruction includes the steering information, distance information, and lane change requirement of the next key intersection, etc. Use the embedding layer to convert the discrete navigation instruction into a continuous feature vector, and the vector dimension is 64. At the same time, after normalizing the vehicle speed information, it is concatenated with the navigation feature to form the state information related to control. This design enables the model to perceive and plan complex driving behaviors such as turning and lane changing in advance.

[0045] This embodiment implements an efficient feature fusion mechanism. The visual features and control state features are fused through a multi-layer perceptron, which consists of three fully connected layers with the number of neurons being 512, 256, and 128 in sequence, and the ReLU activation function is adopted. The fused features are mapped to the action space through the policy head network, and the throttle control amount, steering control amount, and braking control amount are output. The value range of the control amount is normalized to the interval [-1, 1] through the tanh function to facilitate the docking of the actual execution system.

[0046] This embodiment innovatively designs the discriminator network structure. The discriminator adopts an improved WGAN architecture, which includes 6 fully connected layers with 256 neurons in each layer. The input is the concatenated vector of the current environmental state and driving actions, and the output is the probability score that the state-action pair belongs to the expert data. To improve the discrimination accuracy, a gradient penalty term is added to the network to prevent the problems of gradient disappearance and mode collapse. The loss function of the discriminator is: where x is the state, a is the expert action, and a' is the generated action.

[0047] This embodiment establishes a complete expert data acquisition process. Professional drivers with more than 10 years of driving experience are invited to conduct demonstration drives, covering various scenarios such as urban roads and highways. The collected data includes environmental image sequences, vehicle speed curves, steering wheel angles, throttle and brake pedal positions, etc. These data are preprocessed and screened to construct a high-quality expert dataset. Data augmentation methods include random cropping, horizontal flipping, and brightness adjustment, etc., to improve the generalization ability of the model.

[0048] This embodiment implements a stable training optimization strategy. The training process adopts an alternating optimization method. In each round, the discriminator is updated 5 times first, and then the policy generation network is updated 1 time. The optimization goal of the policy network is to maximize the probability output of the discriminator, and at the same time, an entropy regularization term is added in the action space to encourage the policy to explore more possible driving behaviors. The Adam algorithm is used as the optimizer, and the learning rate adopts the cosine annealing strategy, gradually decreasing from 1e-4 to 1e-6. To prevent overfitting, a dropout mechanism is introduced, and the dropout rate is set to 0.3.

[0049] Through the above technological innovations, this embodiment effectively solves the perception and control problems of the autonomous driving system in complex environments. In practical applications, this solution can imitate the driving style of human drivers and generate smooth and natural control instructions. It is particularly suitable for complex scenarios such as urban roads. Through multi-perspective perception and deep imitation learning, the safety and comfort of the autonomous driving system are significantly improved. The adaptive characteristics of this solution enable it to handle various road conditions and weather conditions, and through continuous online learning and optimization, stable and reliable autonomous driving control is achieved.

[0050] Step S102: Construct a multi-dimensional reward function model, collect the state data during the vehicle driving process, detect the vehicle's red-light running behavior, lane-changing and line-crossing behavior, behavior of deviating from the motor vehicle lane, and vehicle collision behavior, record the distance the vehicle has traveled and the remaining driving distance, construct a driving behavior reward and punishment matrix based on the state detection results, and input the driving behavior reward and punishment matrix into the multi-dimensional reward function model to obtain the reward score of the vehicle's current driving state;

[0051] Optionally, in this embodiment, a multi-dimensional reward function model is designed, and a complete state acquisition system is first constructed. A high-precision GPS positioning module is installed on the vehicle to obtain the geographical coordinate information of the vehicle in real time, and the positioning accuracy is better than 0.1 meter. At the same time, an environmental perception network is constructed by millimeter-wave radar and lidar. The scanning frequency of the radar is 20Hz, covering a 360-degree area around the vehicle, and is used to detect the positions and motion states of surrounding vehicles and obstacles.

[0052] This embodiment realizes an accurate traffic light state detection mechanism. The on-vehicle camera is used to collect the images of the traffic lights ahead, and an improved YOLOv5 object detection model is used to identify the positions and states of the traffic lights. The backbone network of the model adopts the CSPDarknet53 structure, and a spatial attention module is added to improve the detection accuracy. At the same time, the traffic light position information in the high-precision map is fused with the real-time detection results to calculate the relative distance and passing time between the vehicle and the traffic lights, and determine whether there is a red-light running behavior.

[0053] This embodiment innovatively designs a lane line detection algorithm. An improved LaneNet model is used for semantic segmentation of the lane lines. The network structure includes two parts: an encoder and a decoder. The encoder uses EfficientNet-B4 to extract features, and the decoder improves the segmentation accuracy through multi-scale feature fusion. The model outputs the pixel-level annotation results of the lane lines, and combines the vehicle's attitude information to calculate the relative position relationship between the vehicle and the lane lines. The determination of lane-changing and line-crossing behavior is based on the minimum distance threshold between the wheels and the lane lines, and the continuity of the lane-changing process is also considered.

[0054] This embodiment designs a reliable lane departure detection method. Based on the lane line detection results, a lane keeping warning system is established. The lateral offset distance and deviation angle between the vehicle center line and the lane center line are calculated. When the offset exceeds the safety threshold, it is determined as a lane departure. The system also considers the vehicle speed factor and dynamically adjusts the warning threshold as the vehicle speed increases. The determination formula for the deviation warning is: W = d + kvθ, where d is the lateral offset distance, v is the vehicle speed, θ is the deviation angle, and k is the adjustment coefficient.

[0055] This embodiment implements a comprehensive collision risk assessment mechanism. By fusing millimeter-wave radar and lidar data, a dynamic obstacle map around the vehicle is constructed. The time to collision (TTC) is calculated for each detected obstacle, where TTC = d / v_r, with d being the relative distance and v_r being the relative speed. Different risk levels are set according to the TTC value and obstacle type to evaluate the collision risk in real time. At the same time, the movement trend of the obstacle is considered to predict the possible collision trajectory.

[0056] This embodiment establishes an innovative reward and punishment matrix structure. Different penalty weights are set for different violations: the weight for running a red light is -1.0, the weight for lane-changing and crossing the line is -0.6, the weight for lane departure is -0.4, and the weight for collision risk is -0.8. At the same time, positive rewards related to task completion are set: the basic reward for driving along the planned route is 0.1, the stage reward for reaching the waypoint is 0.3, and the target reward for completing the entire journey is 1.0. The design of the reward and punishment matrix fully considers the danger level and task importance of various behaviors.

[0057] This embodiment implements a deep reward function calculation process. The state detection results and the reward and punishment matrix are input into a multi-layer perceptron network, which contains three hidden layers with the number of neurons being 256, 128, and 64 in sequence. Each layer uses the LeakyReLU activation function to avoid gradient vanishing. The input features of the network include the detection results of various violations, the current task completion degree, and other information. The output layer uses the tanh activation function to map the reward score to the interval [-1, 1]. The final reward calculation formula is: R = Σ(wisi) + βp, where wi is the behavior weight, si is the state score, β is the progress coefficient, and p is the task completion degree.

[0058] Through the above technological innovations, this embodiment effectively solves several key problems in the design of reward signals for autonomous driving systems, such as sparse rewards, delayed feedback, and multi-objective trade-offs. In practical applications, this solution can accurately evaluate the driving behavior of the vehicle and generate reasonable reward signals to guide policy learning. It is particularly suitable for autonomous driving tasks in complex traffic environments. Through multi-dimensional state detection and refined reward and punishment design, it significantly improves the driving safety and task completion efficiency of the vehicle. The adaptive characteristics of this solution enable it to handle various driving scenarios and achieve stable and reliable policy optimization through real-time state evaluation and reward calculation.

[0059] Step S103: Construct an actor-critic network. Input the environmental image information, navigation instruction information, and vehicle speed information into the actor network. The actor network outputs the optimal driving action in the current state. Input the optimal driving action and the current state into the critic network. The critic network outputs an action value evaluation score. Optimize the parameters of the actor network based on the reward score and the action value evaluation score. Use the optimized actor network to generate an autonomous driving control instruction.

[0060] Optionally, in this embodiment, an actor-critic network is constructed based on a deep reinforcement learning framework. The actor network adopts a hierarchical design. The input layer receives three-way environmental image data, navigation instruction data, and vehicle speed data. The image processing branch uses EfficientNet-B3 as a feature extractor to extract environmental visual information through multi-scale feature fusion. The navigation instruction is converted into a 64-dimensional vector through an encoding layer and input into the control information processing branch together with the vehicle speed information. The features of the two branches are fused through a cross-modal attention mechanism to generate a comprehensive feature representation.

[0061] This embodiment implements an innovative action generation mechanism. The policy head of the actor network uses a Gaussian mixture model (GMM) to represent the action distribution, which contains 8 Gaussian components. Each component is described by a mean vector μ and a covariance matrix Σ. The probability density function of the distribution is: p(a) = Σπi * N(a|μi,Σi), where πi is the mixing weight and N represents the Gaussian distribution. This design enables the model to generate a multi-modal action distribution to adapt to the requirements of different driving scenarios. The policy network outputs an action vector in three dimensions: throttle control amount, steering control amount, and braking control amount.

[0062] This embodiment designs a double Q critic network structure. The critic network adopts a double-network architecture, including a main network and a target network. Each network uses a multi-layer perceptron structure, which contains four hidden layers, and the number of neurons is 512, 256, 128, and 64 in sequence. The input feature is the concatenation of the state vector and the action vector, which is mapped to a Q value through a non-linear transformation. The double-network design effectively alleviates the overestimation problem of Q value estimation and improves the stability of the evaluation. The calculation formula of the Q value is: Q(s,a) = r + γ * min(Q1(s',a'),Q2(s',a')), where r is the immediate reward, γ is the discount factor, and s' and a' are the next state and action respectively.

[0063] This embodiment realizes a stable network training mechanism. The Soft Actor-Critic (SAC) algorithm is used for training, and the objective function consists of two parts: the policy objective and the value function objective. The policy objective function is: J(π) = E[Q(s,a) - α * log(π(a|s))], where α is the temperature parameter used to adjust the balance between exploration and exploitation. The value function objective adopts the mean squared error loss: L(Q) = E[(Q(s,a) - y)^2], where y is the target value. The temperature parameter α is updated through an adaptive adjustment mechanism to meet the preset policy entropy constraint.

[0064] This embodiment establishes an innovative experience replay mechanism. A prioritized experience replay buffer is designed with a capacity of 100,000 transition samples. Each sample contains information such as state, action, reward, next state, etc. The priority of the sample is calculated based on the TD error: p = |r + γ * V(s') - V(s)|^α, where α is the priority exponent. The sampling probability is proportional to the priority, ensuring that important samples are used for training more frequently. At the same time, importance weight correction is introduced to prevent sampling bias.

[0065] This embodiment realizes an efficient parameter optimization strategy. The Polyak averaging method is used to update the target network parameters: θ' = τ * θ + (1 - τ) * θ', where τ is the soft update coefficient. The policy network uses the reparameterization trick for gradient calculation to reduce the variance caused by sampling. The optimizer adopts the Adam algorithm with a learning rate of 3e-4 and a weight decay coefficient of 1e-5. To improve the training efficiency, a parallel environment sampling and asynchronous update mechanism is realized.

[0066] This embodiment designs a complete control instruction generation process. The optimized actor network is deployed on the in-vehicle computing platform, and the NVIDIA Xavier processor is used for real-time inference. The network input frame rate is synchronized with the camera sampling frequency to ensure the real-time nature of the control signal. The generated control instructions are sent to the execution system through the CAN bus to achieve precise control of the throttle, steering, and brake actuators. At the same time, the smoothing process of the control instructions is realized to avoid discomfort caused by sudden changes.

[0067] Through the above technological innovations, this embodiment effectively solves multiple key problems in the policy optimization of the autonomous driving system: unstable policy learning, low exploration efficiency, discontinuous control, etc. In practical applications, this solution can generate stable and safe driving control instructions. It is particularly suitable for autonomous driving tasks in complex traffic environments. Through deep reinforcement learning and multi-modal perception, the robustness and adaptability of the driving policy are significantly improved. The lifelong learning feature of this solution enables it to continuously adapt to new driving scenarios and realizes the continuous evolution of intelligent driving control through online optimization and experience accumulation.

[0068] As can be seen from the above description, the autonomous driving method based on a multi-dimensional reward function provided by the embodiments of the present application can combine the imitation learning and reinforcement learning frameworks, collect environmental information through multiple cameras, and construct a policy generation network and a discriminator network to achieve driving action control. Design a multi-dimensional reward function model to detect and evaluate behaviors such as running red lights, crossing the line, deviating from the lane, and collisions in real time, and construct a driving behavior reward and punishment matrix. Based on the actor-critic network architecture, input environmental information and navigation instructions into the actor network to generate optimal driving actions, and evaluate the action value through the critic network to achieve dynamic optimization of parameters. This method effectively solves the deficiencies of traditional technologies in aspects such as driving behavior evaluation and action value judgment, and significantly improves the safety and reliability of the autonomous driving system.

[0069] In an embodiment of the autonomous driving method based on a multi-dimensional reward function of the present application, the following content may also be specifically included:

[0070] Step S201: Collect vehicle-mounted image information. Install a forward camera at the front of the vehicle and wide-angle cameras on both the left and right sides of the vehicle. Calibrate the lens focal length, field of view angle, image resolution, and sampling frequency of the cameras. Construct the environmental image data collected by the calibrated cameras, the navigation instruction data generated by the vehicle-mounted navigation system, and the speed information collected by the vehicle speed sensor into state input data, and perform normalization processing on the state input data;

[0071] Step S202: Input the normalized state input data into the policy generation network of the neural network structure. Use the convolutional layer to extract image features, use the fully connected layer to fuse image features, navigation features, and speed features, and generate an action output vector of throttle control amount, steering control amount, and braking control amount based on the activation function. Input the action output vector and the current state of the vehicle into the discriminant model in the discriminator network, calculate the expert probability score of the state-action pair, and perform backpropagation optimization on the parameters of the policy generation network according to the expert probability score.

[0072] Optionally, in this embodiment, a precise multi-camera installation layout is first carried out. Install a high-definition forward camera in the middle of the front bumper of the vehicle. The lens uses a fixed-focus lens with a focal length of 6mm and a field of view angle of 120 degrees, covering the key driving area in front. Install a wide-angle camera at each of the left and right rearview mirror positions. Use a fisheye lens with a focal length of 3.6mm, and the field of view angle reaches 170 degrees, effectively expanding the observation range of the side blind area. The installation height and inclination angle of the three cameras are precisely adjusted to ensure the continuity of the overlapping area of the field of view.

[0073] This embodiment realizes a strict camera calibration process. The Zhang's calibration method is used to calibrate the internal parameters of the camera. A 9×6 checkerboard calibration board is used, and calibration images are collected at different distances and angles. Through corner detection and reprojection error optimization, the focal length matrix K and distortion coefficient D of the camera are obtained. The external parameter calibration adopts a multi-target joint optimization method to establish the geometric transformation relationship between three cameras. The reprojection error after calibration is controlled within 1 pixel, ensuring the accuracy of image measurement.

[0074] This embodiment designs an efficient image acquisition mechanism. The camera uses a global exposure CMOS sensor with a resolution of 1920×1080 and a sampling frequency of 30Hz. The image data undergoes real-time color correction, dynamic range compression, and noise suppression by the ISP processing unit. To adapt to different lighting conditions, adaptive exposure control is implemented, and the exposure time range is from 0.1ms to 33ms, ensuring the balance of image clarity and dynamic range.

[0075] This embodiment innovatively designs a multi-source data fusion scheme. The instruction data provided by the in-vehicle navigation system includes path planning information, the type and distance of the next key action point, etc. These discrete navigation instructions are converted into vector form through one-hot encoding. The vehicle speed information is collected by a high-precision Hall sensor with a sampling frequency of 100Hz, and the sampling noise is eliminated through Kalman filtering. The construction of the state input data adopts a time series window method, including the historical information of 3 consecutive frames, providing time series context.

[0076] This embodiment realizes complete data normalization processing. The image data is processed by contrast-limited adaptive histogram equalization (CLAHE) to enhance local details. The pixel values are normalized to the [-1, 1] interval by subtracting the mean and dividing by the variance. The navigation feature vector undergoes L2 regularization processing to ensure the consistency of the feature magnitude. The speed information adopts the maximum-minimum normalization method and is mapped to the [0, 1] interval. The normalization parameters are determined through large-scale data statistics and are dynamically updated during online operation.

[0077] This embodiment establishes a deep policy generation network. ResNet-34 is used as the backbone network for visual feature extraction, which contains multiple residual blocks, effectively alleviating the gradient vanishing problem of deep networks. The convolutional layer uses a 3×3 convolutional kernel, and the number of channels is 64, 128, 256, and 512 in sequence. The feature map passes through the spatial pyramid pooling module to generate multi-scale feature representations. The feature fusion layer adopts an attention mechanism to dynamically adjust the weights of features from different sources, and the fusion formula is: F = Σ(wi * Fi), where wi is the attention weight and Fi is the feature of each modality.

[0078] In this embodiment, an accurate control signal generation mechanism is designed. The output layer of the policy network uses a dual activation function structure. The throttle and brake control amounts are mapped to the interval [0, 1] through the sigmoid function, and the steering control amount is mapped to the interval [-1, 1] through the tanh function. To increase the smoothness of the output, a temporal smoothing constraint is introduced, and the change in control amounts at adjacent times is restricted. The action generation process takes into account physical constraints such as steering angular velocity limits and acceleration / deceleration limits.

[0079] This embodiment implements an innovative discriminator training scheme. The discriminant model adopts the Wasserstein - GAN architecture, and the training stability is improved through the gradient penalty term. The calculation formula for the expert probability score is: D(s,a) = V(s) + λ * A(s,a), where V(s) is the state value function, A(s,a) is the advantage function, and λ is the balance factor. The Adam optimizer is used for backpropagation optimization, with a learning rate of 1e - 4, and gradient clipping is introduced to prevent gradient explosion.

[0080] Through the above - mentioned technological innovations, this embodiment effectively solves multiple key problems in the environmental perception and control generation of the autonomous driving system: incomplete visual information, misalignment of multi - source data, non - smooth control signals, etc. In practical applications, this solution can accurately perceive complex road environments and generate safe and stable control instructions. It is particularly suitable for scenarios such as urban roads. Through multi - perspective perception and deep - learning strategies, it significantly improves the environmental adaptability and control accuracy of the autonomous driving system. The data - driven characteristics of this solution enable it to continuously learn and optimize from expert data, realizing the continuous evolution of intelligent driving control.

[0081] In an embodiment of the autonomous driving method based on a multi - dimensional reward function of the present application, the following content may also be specifically included:

[0082] Step S301: Collect demonstration data of expert drivers, record the environmental state information and corresponding driving action data during the expert driving process, mark the expert driving data as positive samples, mark the driving action data output by the policy generation network as negative samples, construct a training data set for the discriminator network, perform normalization processing on the throttle control amount, steering control amount, and brake control amount in the training data set, and splice the normalized action data with the corresponding environmental state data to construct state - action pairs;

[0083] Step S302: Input the state-action pair into the discriminator network. Use a multi-layer perceptron structure to perform feature mapping on the state-action pair, calculate the similarity score between the state-action pair and the expert data distribution, calculate the prediction error of the discriminator network based on the cross-entropy loss function, optimize the parameters of the discriminator network using the gradient descent method, and use the similarity score output by the discriminator network as the training reward signal for the policy generation network to perform iterative optimization training on the policy generation network.

[0084] Optionally, in this embodiment, a professional driving data acquisition scheme is first designed. Professional drivers with an advanced driver's license and more than 10 years of driving experience are selected for demonstration driving, and the acquisition routes cover various scenarios such as urban roads, highways, and rural roads. During the acquisition process, high-precision sensors are used to record environmental state information, including three-way camera image data, lidar point cloud data, GPS positioning data, etc. At the same time, the driver's control data is recorded, and signals such as the steering wheel angle, accelerator pedal position, and brake pedal position are collected through the CAN bus, with a sampling frequency of 100Hz.

[0085] This embodiment implements a high-quality data preprocessing mechanism. Synchronize and denoise the collected raw data, and use timestamps to align the data streams of different sensors. The environmental image data is processed by de-distortion and illumination normalization to improve the image quality. The vehicle state data is smoothed by a Kalman filter to eliminate the influence of sensor noise. The driving action data is filtered using a moving window average method, with a window size of 5 frames, to retain the smoothness and continuity of the actions.

[0086] This embodiment innovatively constructs a positive and negative sample annotation strategy. Mark the expert driving data as positive samples, and each sample contains the current environmental state and the corresponding driving action. For the construction of negative samples, a method of randomly sampling by the policy generation network is used to generate action sequences that are somewhat different from the expert actions. To increase the diversity of negative samples, noise perturbations are added during the sampling process, and the perturbation intensity is controlled by a Gaussian distribution: ε~N(0,σ2), where σ is the standard deviation. At the same time, a sample screening mechanism is designed to filter out unreasonable negative samples.

[0087] This embodiment designs a complete data normalization process. Perform Min-Max normalization on the throttle control amount, steering control amount, and brake control amount, and map the values to the [0,1] interval. The normalization formula is: x'=(x - xmin) / (xmax - xmin), where x is the original value, and xmin and xmax are the minimum and maximum values of the historical data respectively. To ensure the stability of normalization, the normalization parameters are updated using a moving window statistics method, with a window size of 10000 frames.

[0088] This embodiment realizes an innovative construction method for state-action pairs. The normalized action data is concatenated with the corresponding environmental state data in terms of features to construct high-dimensional state-action pairs. The state features include three parts: image features, vehicle state features, and environmental features, and feature fusion is performed through an attention mechanism. The dimension of the fused features is 512, which contains the key information of the driving scenario. The feature fusion formula is: F = Attention(Ws[s;a]), where s is the state feature, a is the action feature, and Ws is a learnable weight matrix.

[0089] This embodiment constructs a deep discriminator network structure. A multi-layer perceptron structure is adopted, which includes 5 hidden layers, and the number of neurons is 512, 256, 128, 64, and 32 in sequence. Each layer uses the LeakyReLU activation function to prevent gradient vanishing. The input of the network is the state-action pair feature, and the output is the probability score that the state-action pair belongs to the expert data. To improve the robustness of discrimination, a Dropout layer is added to the network with a dropout rate of 0.2.

[0090] This embodiment designs a stable network training strategy. The cross-entropy loss function is used to calculate the prediction error of the discriminator: L = -Σ(y*log(p)+(1-y)log(1-p)), where y is the label and p is the predicted probability. The Adam optimizer is used in the optimization process, and the initial value of the learning rate is 1e-4, and the cosine annealing strategy is adopted for dynamic adjustment. To prevent mode collapse, a gradient penalty term is introduced: where λ is the penalty coefficient. At the same time, an experience replay mechanism is implemented to improve the stability of training.

[0091] This embodiment realizes an efficient policy optimization mechanism. The similarity score output by the discriminator is used as the reward signal for the policy network, and the policy gradient method is used for optimization. The calculation formula of the policy gradient is: where π is the policy function and r is the discriminator score. To improve the exploration efficiency, an entropy regularization term is added to the policy to encourage the policy to generate diverse actions. At the same time, importance sampling technology is adopted to improve the sample utilization efficiency.

[0092] Through the above technological innovations, this embodiment effectively solves several key problems in autonomous driving imitation learning: unstable quality of expert data, insufficient discriminator training, low efficiency of policy optimization, etc. In practical applications, this solution can accurately learn the driving style of expert drivers and generate safe and stable driving actions. It is particularly suitable for autonomous driving tasks in complex traffic environments. Through high-quality data collection and deep imitation learning, the robustness and generalization ability of the driving strategy are significantly improved. The adaptive characteristics of this solution enable it to continuously learn from expert data, and through iterative optimization and online learning, the continuous evolution of the driving strategy is achieved.

[0093] In an embodiment of the automatic driving method based on a multi-dimensional reward function of the present application, the following content may also be specifically included:

[0094] Step S401: Construct a vehicle state detection sub-model, collect the traffic light status information of the lane where the vehicle is located, calculate the relative distance between the vehicle position and the traffic light position based on the electronic map data, collect the relative position information between the vehicle contour and the lane line, collect the distribution information of obstacles around the vehicle based on the lidar sensor, calculate the remaining distance between the vehicle and the target end point according to the global positioning system data, and construct the vehicle state detection data into a state vector;

[0095] Step S402: Extract features from the state vector, judge whether the vehicle runs a red light based on the traffic light status and the relative distance, judge whether there is a behavior of pressing the line or deviating from the lane based on the relative position between the vehicle contour and the lane line, judge whether a collision occurs based on the obstacle distribution information, calculate the proportion of the distance the vehicle has traveled to the total driving distance, construct a violation behavior label vector, and input the state vector and the violation behavior label vector into a multi-dimensional reward function model for training.

[0096] Optionally, in this embodiment, a multi-dimensional state detection system is first designed. A high-resolution camera is installed in front of the vehicle, which is specifically used for traffic light recognition, and an improved YOLOv5 target detection model is used for real-time detection. The backbone network of the model uses CSPDarknet53, and an attention mechanism is added to enhance the detection ability for small targets. The detection confidence threshold of the traffic light status is set to 0.85 to ensure the reliability of the detection results. At the same time, the accurate coordinates of the traffic lights are obtained through a high-precision electronic map, and the relative distance is calculated in combination with the GPS positioning information of the vehicle.

[0097] This embodiment realizes an accurate lane line detection scheme. The on-vehicle camera is used to collect the front road surface image, and the improved LaneNet model is used for lane line segmentation. The model adopts an encoder-decoder structure. The encoder uses EfficientNet-B3 to extract features, and the decoder improves the segmentation accuracy through multi-scale feature fusion. At the same time, the vehicle attitude estimation algorithm is used to calculate the relative position relationship between the vehicle contour and the lane line, and the position calculation formula is: d = |ax + by + c| / √(a2 + b2), where (a, b, c) are the parameters of the lane line equation, and (x, y) are the coordinates of the vehicle contour points.

[0098] This embodiment constructs an all-round obstacle perception system. A 32-line lidar is installed on the top of the vehicle, with a scanning frequency of 10 Hz, covering a 360-degree area around the vehicle. Surrounding obstacles are identified through point cloud segmentation and clustering algorithms, and the position, speed, and size information of each obstacle are calculated. Obstacle tracking uses an improved Kalman filter, and the state vector includes position, speed, and acceleration information. The collision risk assessment is based on the time to collision (TTC) metric, TTC = d / v_r, where d is the relative distance and v_r is the relative speed.

[0099] This embodiment designs an innovative state vector construction method. The data collected by each sensor are synchronized in time and registered in space to construct a unified state representation. The state vector includes traffic light state features (4 dimensions), relative distance features (3 dimensions), lane position features (6 dimensions), obstacle distribution features (16 dimensions), and navigation information features (8 dimensions). Feature fusion uses an attention mechanism, and the calculation formula is: F = Attention(Ws[f1; f2; f3; f4; f5]), where fi are various types of features and Ws is a learnable weight matrix.

[0100] This embodiment implements a strict violation determination mechanism. For the behavior of running a red light, both the signal light state and the distance from the vehicle to the stop line are considered, and the determination rule is designed: when the red light is on, if the vehicle is less than the safe distance from the stop line and does not stop, it is determined to run a red light. The determination of the line pressing behavior is based on the minimum distance threshold between the wheel and the lane line, and the lane change intention signal is also considered. The lane departure determination is based on the lateral offset and deviation angle between the vehicle center line and the lane center line.

[0101] This embodiment designs a complete collision detection algorithm. A collision risk assessment model is constructed based on the obstacle distribution information, and the minimum TTC value with surrounding obstacles is calculated. When the TTC is less than the safety threshold and the relative distance is less than the warning distance, it is determined that there is a collision risk. The safety threshold is dynamically adjusted according to the vehicle speed, and the adjustment formula is: Ts = T0 + k*v, where T0 is the basic threshold, v is the vehicle speed, and k is the adjustment coefficient.

[0102] This embodiment establishes an innovative behavior label encoding scheme. Various types of violation behaviors are encoded into binary label vectors, with a vector dimension of 5, corresponding to running a red light, line pressing, lane deviation, collision risk, and task progress respectively. The task progress is calculated by the ratio of the traveled distance to the total distance and is converted into a scalar value through piecewise linear mapping. The generation of the label vector considers the persistence and severity of the behavior, and the stability of the label is improved through time series smoothing processing.

[0103] This embodiment implements an efficient reward function training mechanism. A deep neural network structure is used to construct the reward function model. The network contains 4 hidden layers, and the number of neurons is 256, 128, 64, and 32 in sequence. The input layer receives the concatenated features of the state vector and the violation behavior label vector, and the output layer generates a scalar reward value. The training process adopts the supervised learning method, and the loss function is a weighted combination of the mean square error and cross entropy. To improve the generalization ability of the model, dropout and L2 regularization are introduced.

[0104] Through the above technological innovations, this embodiment effectively solves multiple key problems in the state detection and behavior evaluation of the autonomous driving system: sensor data asynchronization, ambiguous violation behavior determination, inaccurate reward signals, etc. In practical applications, this solution can accurately detect various violation behaviors and generate reasonable reward signals to guide policy learning. It is particularly suitable for autonomous driving tasks in complex traffic environments. Through multi-dimensional state detection and precise behavior evaluation, the safety and standardization of the driving strategy are significantly improved. The adaptive characteristics of this solution enable it to handle various driving scenarios and achieve reliable driving behavior constraints through real-time state monitoring and evaluation.

[0105] In an embodiment of the autonomous driving method based on a multi-dimensional reward function of the present application, the following content may also be specifically included:

[0106] Step S501: Construct a reward and punishment matrix according to the vehicle state detection results. Assign a violation punishment weight for the vehicle running a red light, a violation punishment weight for the vehicle crossing the line and changing lanes, a violation punishment weight for the vehicle deviating from the motor vehicle lane, a violation punishment weight for the vehicle collision, and a task completion reward weight for the vehicle reaching the target position. Based on the value of the task completion reward weight, construct a driving behavior reward and punishment matrix;

[0107] Step S502: Input the driving behavior reward and punishment matrix and the vehicle state detection vector into the multi-dimensional reward function model. Use a deep neural network structure to extract features from the state detection vector, calculate the reward and punishment scores of each driving behavior based on the weight values in the reward and punishment matrix, and perform weighted summation on the reward and punishment scores to generate a comprehensive reward score reflecting the current driving state of the vehicle.

[0108] Optionally, this embodiment first designs a multi-level reward and punishment weight system. For the behavior of running a red light, a dynamic punishment weight is designed according to the vehicle speed and the timing of running the red light. The weight calculation formula is: w1 = -α * (v / v_max) * (1 - d / d_safe), where v is the current vehicle speed, v_max is the speed limit value, d is the distance from the stop line, d_safe is the safety distance, and α is the basic weight coefficient. When the vehicle is closer to the stop line and the vehicle speed is faster in the red light state, the punishment weight is greater, effectively suppressing dangerous driving behaviors.

[0109] This embodiment realizes an accurate lane - change behavior evaluation mechanism. The penalty weight for lane - crossing lane - change behavior considers the necessity of lane - change and the normativity of execution. The weight calculation formula is: w2 = -β*(1 - p_nav)*(d_line / d_safe), where p_nav is the probability of lane - change necessity given by the navigation system, d_line is the minimum distance from the lane line, d_safe is the safety - distance threshold, and β is the basic weight coefficient. This design allows necessary lane - change operations while ensuring lane - change safety.

[0110] This embodiment constructs an adaptive lane - departure penalty mechanism. The penalty weight for deviating from the motor vehicle lane is dynamically adjusted according to the deviation degree and duration. The calculation formula is: w3 = -γ*(d_offset / d_max)(1 + λt), where d_offset is the lateral offset distance, d_max is the lane width, t is the deviation duration, and γ and λ are adjustment coefficients. By introducing the time factor, a greater penalty is imposed on continuous deviation behaviors, prompting the vehicle to return to the correct lane in a timely manner.

[0111] This embodiment designs an innovative collision - risk assessment scheme. The penalty weight for collision behavior is comprehensively calculated based on time - to - collision (TTC) and relative distance. The weight formula is: w4 = -δ*(1 / TTC)*(d_min / d_warning), where TTC is the time - to - collision value, d_min is the minimum safety distance, d_warning is the warning distance, and δ is the basic weight coefficient. This design considers both dynamic collision risks and static safety - distance maintenance.

[0112] This embodiment realizes a progressive task - completion reward mechanism. The reward weight for reaching the target position is designed using a piece - wise function. The weight calculation formula is: w5 = ε*(d_complete / d_total)^k, where d_complete is the completed distance, d_total is the total distance, k is the shape parameter of the reward curve, and ε is the basic reward coefficient. This design enables the reward to grow non - linearly with the task - completion degree, motivating the vehicle to continuously advance towards the target.

[0113] This embodiment constructs an innovative reward - punishment matrix structure. The weights of various behaviors are organized in the form of a 5×5 matrix. The diagonal elements of the matrix represent the basic weights of each behavior, and the non - diagonal elements represent the interaction effects between behaviors. The calculation of the matrix considers the temporal correlation of behaviors and updates the historical weights through an exponential moving average: W_t = μ*W_(t - 1)+(1 - μ)*W_current, where μ is the smoothing factor.

[0114] In this embodiment, a deep feature extraction network is designed. A multi-layer perceptron structure is used to process the state detection vector. The network contains four hidden layers, and the number of neurons is 256, 128, 64, and 32 in sequence. A batch normalization layer and a ReLU activation function are connected after each layer to improve the stability of feature extraction. The correlation between state variables is considered in the feature extraction process, and the importance of different features is dynamically adjusted through an attention mechanism.

[0115] In this embodiment, an accurate method for calculating reward and punishment scores is implemented. Based on the extracted state features and the weight values of the reward and punishment matrix, a bilinear mapping is used to calculate the reward and punishment scores of each behavior: s_i = F^TW_iF, where F is the state feature vector and W_i is the weight matrix corresponding to the behavior. To handle the scale difference between features and weights, a normalization layer is introduced for numerical adjustment.

[0116] In this embodiment, a reliable comprehensive scoring mechanism is established. The reward and punishment scores of each item are adaptively weighted and summed, and the weight coefficients are calculated through a soft attention mechanism: a_i = softmax(v^Ttanh(Ws_i)), where s_i is the score of each item, and W and v are learnable parameters. The final comprehensive reward score is obtained through weighted combination: R = Σ(a_i*s_i). This design can dynamically adjust the importance of each behavior according to the current state.

[0117] Through the above technological innovations in this embodiment, multiple key problems in the behavior evaluation and reward calculation of the autonomous driving system are effectively solved: inconsistent reward and punishment scales, ignored behavior interactions, unstable evaluation results, etc. In practical applications, this solution can accurately evaluate the normativity of driving behaviors and generate reasonable reward signals to guide policy learning. It is particularly suitable for autonomous driving tasks in complex traffic environments. Through multi-dimensional behavior evaluation and accurate reward calculation, the safety and efficiency of driving strategies are significantly improved. The adaptive characteristics of this solution enable it to handle various driving scenarios and achieve reliable driving behavior guidance through dynamic weight adjustment and comprehensive evaluation.

[0118] In an embodiment of the autonomous driving method based on a multi-dimensional reward function in this application, the following content may also be specifically included:

[0119] Step S601: Construct an actor network model. Use a multi-layer convolutional neural network to extract the visual features of the environmental image, use a fully connected layer to process the navigation instruction information and vehicle speed information, fuse the visual features, navigation features, and vehicle speed features, and map the fused features to the action space distribution through a policy head network. Based on the action space distribution sampling, generate the optimal driving actions of the throttle control amount, steering control amount, and braking control amount.

[0120] Step S602: Construct a judgment network model, concatenate the optimal driving actions with the current environmental state, perform non-linear transformation on the state-action features using a multi-layer perceptron network, map the transformed features to a scalar value evaluation score based on the value head network, perform weighted fusion on the value evaluation score and the reward score output by the reward function model, construct the loss function of the actor network, and optimize and update the parameters of the actor network based on the policy gradient method.

[0121] Optionally, in this embodiment, a deep actor network architecture is designed. The visual feature extraction branch uses EfficientNetV2-B3 as the backbone network, which includes multiple MBConv blocks and attention modules. The input is three-channel image data with a resolution of 640×480, and multi-scale features are extracted through the backbone network. The feature pyramid structure includes five scale layers from P3 to P7, and the perceptual ability is enhanced through top-down feature fusion. The features of each scale layer are enhanced through channel attention and spatial attention modules to improve the perception ability of key regions.

[0122] This embodiment implements an efficient navigation feature processing mechanism. The navigation instruction information is converted into a 64-dimensional vector through the Embedding layer, including path planning, steering indication, and distance information. The vehicle speed information is processed through a three-layer fully connected network, and the number of neurons is 64, 32, and 16 in sequence, using the ReLU activation function. The fusion of navigation and speed features adopts a gating mechanism, and the fusion formula is: F = σ(Wg[fn; fv]) * tanh(Wf[fn; fv]), where fn is the navigation feature, fv is the speed feature, and Wg and Wf are learnable weight matrices.

[0123] This embodiment innovatively designs a feature fusion strategy. The multi-head cross-attention mechanism is used to fuse visual features, navigation features, and speed features. The number of attention heads is 8, and each head independently learns the correlation between different features. The calculation formula for the fused features is: F = MultiHead(Q, K, V), where Q, K, and V are obtained by linear transformation of different features respectively. To enhance the expression ability of the features, residual connections and layer normalization are introduced to effectively prevent feature degradation.

[0124] This embodiment implements a flexible policy head network structure. The Gaussian mixture model (GMM) is used to represent the action space distribution, which includes 8 Gaussian components. Each component is described by a mean vector μ, a covariance matrix Σ, and a mixing weight π. The probability density function of the distribution is: p(a) = Σπi * N(a|μi, Σi), where N represents the Gaussian distribution. The policy network outputs three action dimensions: throttle control amount, steering control amount, and braking control amount, and the value range of each dimension is normalized to the [-1, 1] interval through the tanh function.

[0125] This embodiment designs a dual evaluation network structure. The dual Q-network architecture is adopted, and each network contains 6 fully connected layers, with the number of neurons being 512, 256, 128, 64, 32, and 16 in sequence. The input feature is the concatenation of the state vector and the action vector, and the state-action value function is generated through a non-linear transformation. To enhance the feature extraction ability, residual connections are introduced in the intermediate layer. The design of the dual network effectively alleviates the problem of overestimation of Q-values and improves the stability of evaluation.

[0126] This embodiment establishes an innovative value evaluation mechanism. The value head network adopts the Dueling architecture, which decomposes the Q-value into the state value function V(s) and the advantage function A(s,a). The final Q-value calculation formula is: Q(s,a) = V(s) + A(s,a) - mean(A(s,·)), where mean(A(s,·)) is the mean value of the advantage function. This design enables the network to learn the intrinsic value of the state and the relative advantage of the action separately, improving the accuracy of evaluation.

[0127] This embodiment implements an efficient reward fusion strategy. The value evaluation score output by the evaluation network and the reward score of the reward function model are softly weighted and fused. The fusion weight is adaptively adjusted through meta-learning, and the weight update formula is: w = softmax(Wm[v;r]), where v is the value score, r is the reward score, and Wm is the parameter of the meta-learner. This design can dynamically adjust the importance of the two signals according to different scenarios.

[0128] This embodiment designs a stable loss function structure. The loss function of the actor network includes a policy gradient term, an entropy regularization term, and a value constraint term. The proximal policy optimization (PPO) algorithm is used for the policy gradient, and the objective function is: L = E[min(ρ*A, clip(ρ, 1 - ε, 1 + ε)*A)], where ρ is the probability ratio of the old and new policies, A is the advantage function, and ε is the clipping parameter. Entropy regularization prevents the policy from converging prematurely by controlling the exploration degree of the policy.

[0129] This embodiment implements an innovative parameter optimization strategy. The Actor-Critic framework is used for training. The PPO algorithm is used for policy optimization, and the TD(λ) algorithm is used for value function optimization. The Adam optimizer is selected, and the learning rate adopts the cosine annealing strategy, gradually decreasing from 1e-4 to 1e-6. To improve the training efficiency, a parallel environment sampling and asynchronous update mechanism are implemented. At the same time, an experience replay buffer is introduced to improve the sample utilization efficiency.

[0130] Through the above technological innovations, this embodiment effectively solves multiple key problems in the policy learning of the autonomous driving system: insufficient feature extraction, unstable action generation, inaccurate value evaluation, etc. In practical applications, this solution can generate safe and stable driving actions and adapt to complex and changeable traffic environments. It is particularly suitable for scenarios such as urban roads. Through deep reinforcement learning and multi-modal perception, the decision-making ability and control accuracy of the autonomous driving system are significantly improved. The adaptive characteristics of this solution enable it to continuously learn and optimize. Through online training and experience accumulation, continuous evolution of driving strategies is achieved.

[0131] In an embodiment of the autonomous driving method based on a multi-dimensional reward function in the present application, the following content may also be specifically included:

[0132] Step S701: Use the proximal policy optimization algorithm to construct a parameter optimization model for the actor network. Compare and calculate the advantage function value by comparing the reward score output by the reward function model with the value evaluation score output by the critic network. Construct a policy gradient based on the advantage function value, constrain the difference between the action distribution output by the actor network and the target policy distribution, use the trust region policy optimization method to limit the step size of each parameter update, and iteratively optimize the parameters of the actor network based on the stochastic gradient descent method;

[0133] Step S702: Deploy the optimized actor network to the vehicle-mounted computing platform, collect vehicle environment image information, navigation instruction information, and vehicle speed information to construct a state input vector, input the state input vector into the optimized actor network, generate driving control instructions for the throttle control amount, steering control amount, and braking control amount based on the action output layer of the actor network, and send the driving control instructions to the vehicle execution system.

[0134] Optionally, this embodiment first designs an innovative parameter optimization mechanism. Based on the proximal policy optimization (PPO) algorithm, an optimization model is constructed. The generalized advantage estimation (GAE) method is used to calculate the advantage function: A(s,a) = Σ(γλ)^i * δ_t+i, where δ_t = r_t + γV(s_(t+1)) - V(s_t), γ is the discount factor, and λ is the GAE parameter. This design effectively balances the bias and variance of the estimation and provides more stable gradient information. The advantage function reflects the quality of the current action relative to the average performance and provides reliable guidance for policy updates.

[0135] This embodiment implements an accurate policy gradient calculation method. Based on the calculated advantage function values, a PPO-Clip objective function is constructed: L = min(r_t(θ) * A_t, clip(r_t(θ), 1 - ε, 1 + ε)A_t), where r_t(θ) is the probability ratio of the old and new policies, and ε is the clipping range parameter. Through the clipping operation of the probability ratio, it prevents performance collapse caused by excessive policy updates. At the same time, an entropy regularization term is introduced: L_entropy = -αΣπ(a|s) * logπ(a|s), which encourages the policy to explore new action spaces.

[0136] This embodiment constructs a strict trust region constraint. The KL divergence is used to measure the distribution difference before and after the policy update: KL(π_old||π_new), and an adaptive KL divergence threshold β is set. When the KL divergence caused by the policy update exceeds the threshold, the update step size is adjusted through a penalty term: L_total = L_clip - c * max(0, KL - β), where c is the penalty coefficient. This design ensures the stability of the policy update and avoids excessive distribution changes.

[0137] This embodiment designs an innovative parameter update strategy. The stochastic gradient descent method is used for iterative optimization, and the learning rate is updated through an adaptive adjustment mechanism: η_t = η_0 * sqrt(1 - β_2^t) / (1 - β_1^t), where β_1 and β_2 are momentum parameters. To improve the training efficiency, a parallel data collection and asynchronous parameter update mechanism is implemented. Each training batch contains trajectories from multiple parallel environments, and importance sampling is performed through an experience replay buffer.

[0138] This embodiment implements an efficient network deployment solution. NVIDIA Xavier is selected as the in-vehicle computing platform, which has an 8-core ARM processor and a 512-core Volta GPU. The trained actor network is optimized through the TensorRT framework, including techniques such as weight quantization, computation graph optimization, and dynamic batching. The inference latency of the model is controlled within 10 ms, meeting the requirements of real-time control.

[0139] This embodiment designs a reliable state acquisition mechanism. The environmental image information is collected in real time by three cameras with a resolution of 1920×1080 and a frame rate of 30 fps. The image data is preprocessed by an ISP processing unit, including denoising, white balance, and exposure compensation. The navigation instruction information is obtained from the in-vehicle navigation system, including the type and distance of the next key action point. The vehicle speed information is collected by a high-precision Hall sensor with a sampling frequency of 100 Hz.

[0140] In this embodiment, an innovative control instruction generation method is constructed. The action output layer of the actor network adopts a dual-head design, and the mean head and variance head respectively generate the parameters of the action distribution. The throttle control amount and the brake control amount are mapped to the interval [0, 1] through the sigmoid function, and the steering control amount is mapped to the interval [-1, 1] through the tanh function. To ensure the smoothness of control, an action filter is introduced: a_t = α * a_(t - 1) + (1 - α) * a_new, where α is the smoothing coefficient.

[0141] In this embodiment, a reliable execution system docking is achieved. The generated control instructions are sent to the vehicle execution system through the CAN bus, and the communication protocol follows the SAE J1939 standard. The throttle and brake controls drive the actuator through a linear motor, and the response time is less than 50 ms. The steering control is executed through an electric power steering system to achieve precise angle control. At the same time, a safety inspection mechanism for control instructions is established to filter out unreasonable control values.

[0142] Through the above technological innovations, this embodiment effectively solves multiple key problems in the strategy optimization and control execution of the autonomous driving system: unstable parameter update, non-smooth control instructions, large execution delay, etc. In practical applications, this solution can generate safe and smooth driving control instructions. It is particularly suitable for autonomous driving tasks in complex traffic environments. Through stable strategy optimization and precise control execution, it significantly improves the safety and comfort of driving. The real-time performance and reliability of this solution enable it to handle various driving scenarios, and through continuous optimization and feedback, a closed-loop optimization of intelligent driving control is achieved.

[0143] In order to effectively solve the deficiencies of traditional technologies in aspects such as driving behavior evaluation and action value judgment, and significantly improve the safety and reliability of the autonomous driving system, this application provides an embodiment of an autonomous driving device based on a multi-dimensional reward function for implementing all or part of the content of the autonomous driving method based on the multi-dimensional reward function. Refer to Figure 2 The autonomous driving device based on the multi-dimensional reward function specifically includes the following contents:

[0144] A driving model construction module 10, which is used to construct an imitation learning autonomous driving model. The environmental image information, vehicle navigation instruction information, and vehicle speed information collected by the on-vehicle front camera and left and right wide-angle cameras are input into the policy generation network. The policy generation network generates a driving action control signal. The driving action control signal and the current environmental state are input into the discriminator network. The discriminator network outputs the probability score of the action state pair belonging to the expert data set. The policy generation network is trained based on the probability score. The driving action control signal includes the throttle control amount, the steering control amount, and the brake control amount.

[0145] The reward function construction module 20 is used to construct a multi-dimensional reward function model, collect the state data during the vehicle driving process, detect the vehicle running a red light, the vehicle changing lanes and crossing the line, the vehicle deviating from the motor vehicle lane, and the vehicle collision behavior, record the distance the vehicle has traveled and the remaining driving distance, construct a driving behavior reward and punishment matrix based on the state detection results, and input the driving behavior reward and punishment matrix into the multi-dimensional reward function model to obtain the reward score of the current driving state of the vehicle;

[0146] The automatic driving execution module 30 is used to construct an actor-critic network, input the environmental image information, navigation instruction information, and vehicle speed information into the actor network. The actor network outputs the optimal driving action in the current state, input the optimal driving action and the current state into the critic network, and the critic network outputs the action value evaluation score. Optimize the parameters of the actor network based on the reward score and the action value evaluation score, and generate an automatic driving control instruction using the optimized actor network.

[0147] As can be seen from the above description, the automatic driving device based on a multi-dimensional reward function provided by the embodiment of the present application can combine the imitation learning and reinforcement learning frameworks, collect environmental information through multiple cameras, construct a policy generation network and a discriminator network to achieve driving action control. Design a multi-dimensional reward function model to detect and evaluate behaviors such as running a red light, crossing the line, deviating from the lane, and collision in real time, and construct a driving behavior reward and punishment matrix. Based on the actor-critic network architecture, input environmental information and navigation instructions into the actor network to generate the optimal driving action, and evaluate the action value through the critic network to achieve dynamic parameter optimization. This method effectively solves the deficiencies of traditional technologies in aspects such as driving behavior evaluation and action value judgment, and significantly improves the safety and reliability of the automatic driving system.

[0148] From the hardware level, in order to effectively solve the deficiencies of traditional technologies in aspects such as driving behavior evaluation and action value judgment, and significantly improve the safety and reliability of the automatic driving system, the present application provides an embodiment of an electronic device for implementing all or part of the content of the automatic driving method based on a multi-dimensional reward function. The electronic device specifically includes the following content:

[0149] A processor, a memory, a communications interface, and a bus; wherein, the processor, the memory, and the communications interface complete communication with each other through the bus; the communications interface is used to implement information transmission between the autonomous driving device based on a multi-dimensional reward function and related devices such as a core business system, a user terminal, and a related database, etc.; the logic controller may be a desktop computer, a tablet computer, a mobile terminal, etc., and this embodiment is not limited thereto. In this embodiment, the logic controller may be implemented with reference to the embodiments of the autonomous driving method based on a multi-dimensional reward function and the embodiments of the autonomous driving device based on a multi-dimensional reward function, the content of which is incorporated herein, and the repeated parts will not be elaborated.

[0150] It can be understood that the user terminal may include a smart phone, a tablet electronic device, a network set-top box, a portable computer, a desktop computer, a personal digital assistant (PDA), a vehicle-mounted device, a smart wearable device, etc. Among them, the smart wearable device may include smart glasses, a smart watch, a smart bracelet, etc.

[0151] In practical applications, part of the autonomous driving method based on a multi-dimensional reward function may be executed on the electronic device side as described above, or all operations may be completed in the client device. Specifically, it can be selected according to the processing capacity of the client device and the limitations of the user usage scenario, etc. This application does not make any limitations in this regard. If all operations are completed in the client device, the client device may further include a processor.

[0152] The above-mentioned client device may have a communication module (i.e., a communication unit), and may be communicatively connected to a remote server to implement data transmission with the server. The server may include a server on the task scheduling center side, and in other implementation scenarios, it may also include a server of an intermediate platform, such as a server of a third-party server platform communicatively linked to the task scheduling center server. The server may include a single computer device, or may include a server cluster composed of multiple servers, or a server structure of a distributed device.

[0153] Figure 3 This is a schematic block diagram of the system composition of the electronic device 9600 according to an embodiment of the present application. As Figure 3 shown, the electronic device 9600 may include a central processing unit 9100 and a memory 9140; the memory 9140 is coupled to the central processing unit 9100. It should be noted that this Figure 3 is exemplary; other types of structures may also be used to supplement or replace this structure to implement telecommunication functions or other functions.

[0154] In one embodiment, the function of the autonomous driving method based on the multi-dimensional reward function can be integrated into the central processing unit 9100. Among them, the central processing unit 9100 can be configured to perform the following controls:

[0155] Step S101: Construct an imitation learning autonomous driving model, input the environmental image information, vehicle navigation instruction information, and vehicle speed information collected by the in-vehicle forward camera and left and right wide-angle cameras into the policy generation network. The policy generation network generates a driving action control signal, inputs the driving action control signal and the current environmental state into the discriminator network. The discriminator network outputs the probability score that the action-state pair belongs to the expert data set, and trains the policy generation network based on the probability score; the driving action control signal includes the throttle control amount, the steering control amount, and the braking control amount;

[0156] Step S102: Construct a multi-dimensional reward function model, collect the state data during the vehicle driving process, detect the vehicle's red-light running behavior, lane-changing and line-crossing behavior, deviation from the motor vehicle lane behavior, and vehicle collision behavior, record the distance the vehicle has traveled and the remaining driving distance, construct a driving behavior reward and punishment matrix based on the state detection results, and input the driving behavior reward and punishment matrix into the multi-dimensional reward function model to obtain the reward score of the vehicle's current driving state;

[0157] Step S103: Construct an actor-critic network, input the environmental image information, navigation instruction information, and vehicle speed information into the actor network. The actor network outputs the optimal driving action in the current state, inputs the optimal driving action and the current state into the critic network. The critic network outputs the action value evaluation score, and optimizes the parameters of the actor network based on the reward score and the action value evaluation score, and uses the optimized actor network to generate an autonomous driving control instruction.

[0158] As can be seen from the above description, the electronic device provided in the embodiment of the present application combines the imitation learning and reinforcement learning frameworks, collects environmental information through multiple cameras, constructs a policy generation network and a discriminator network to achieve driving action control. Design a multi-dimensional reward function model to detect and evaluate behaviors such as running red lights, crossing lines, deviating from lanes, and collisions in real time, and construct a driving behavior reward and punishment matrix. Based on the actor-critic network architecture, input environmental information and navigation instructions into the actor network to generate the optimal driving action, and evaluate the action value through the critic network to achieve dynamic parameter optimization. This method effectively solves the deficiencies of traditional technologies in aspects such as driving behavior evaluation and action value judgment, and significantly improves the safety and reliability of the autonomous driving system.

[0159] In another embodiment, the autonomous driving device based on the multi-dimensional reward function can be separately configured from the central processor 9100. For example, the autonomous driving device based on the multi-dimensional reward function can be configured as a chip connected to the central processor 9100, and the functions of the autonomous driving method based on the multi-dimensional reward function can be realized through the control of the central processor.

[0160] As Figure 3 shown, the electronic device 9600 may further include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It should be noted that the electronic device 9600 does not necessarily have to include Figure 3 all the components shown in Figure 3 ; in addition, the electronic device 9600 may further include

[0161] As Figure 3 shown, the central processor 9100 is sometimes also referred to as a controller or an operation control, and may include a microprocessor or other processor devices and / or logic devices. The central processor 9100 receives inputs and controls the operations of the various components of the electronic device 9600.

[0162] Among them, the memory 9140 can be, for example, one or more of a buffer, a flash memory, a hard drive, a removable medium, a volatile memory, a non-volatile memory, or other suitable devices. The above information related to failures can be stored, and in addition, programs for executing relevant information can also be stored. And the central processor 9100 can execute the programs stored in the memory 9140 to implement information storage or processing, etc.

[0163] The input unit 9120 provides inputs to the central processor 9100. The input unit 9120 is, for example, a key or a touch input device. The power supply 9170 is used to supply power to the electronic device 9600. The display 9160 is used to display display objects such as images and texts. The display can be, for example, an LCD display, but is not limited thereto.

[0164] The memory 9140 can be a solid-state memory. For example, a read-only memory (ROM), a random access memory (RAM), a SIM card, etc. It can also be such a memory that stores information even when powered off, can be selectively erased and has more data stored. Examples of such a memory are sometimes referred to as EPROMs, etc. The memory 9140 can also be some other type of device. The memory 9140 includes a buffer memory 9141 (sometimes referred to as a buffer). The memory 9140 can include an application / function storage unit 9142, and the application / function storage unit 9142 is used to store application programs and function programs or the processes for operating the electronic device 9600 through the central processor 9100.

[0165] The memory 9140 may further include a data storage unit 9143 for storing data such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 9144 of the memory 9140 may include various drivers of the electronic device for communication functions and / or for performing other functions of the electronic device (such as a messaging application, an address book application, etc.).

[0166] The communication module 9110 is a transmitter / receiver that transmits and receives signals via the antenna 9111. The communication module 9110 (transmitter / receiver) is coupled to the central processor 9100 to provide input signals and receive output signals, which may be the same as in the case of a conventional mobile communication terminal.

[0167] Based on different communication technologies, multiple communication modules 9110 may be provided in the same electronic device, such as a cellular network module, a Bluetooth module, and / or a wireless local area network module, etc. The communication module 9110 (transmitter / receiver) is also coupled to the speaker 9131 and the microphone 9132 via the audio processor 9130 to provide an audio output via the speaker 9131 and receive an audio input from the microphone 9132, so as to implement normal telecommunication functions. The audio processor 9130 may include any suitable buffer, decoder, amplifier, etc. In addition, the audio processor 9130 is also coupled to the central processor 9100, so that recording can be performed on the local machine through the microphone 9132, and the sound stored on the local machine can be played through the speaker 9131.

[0168] An embodiment of the present application also provides a computer-readable storage medium capable of implementing all steps of the multi-dimensional reward function-based autonomous driving method with the execution subject being a server or a client in the above embodiments. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, all steps of the multi-dimensional reward function-based autonomous driving method with the execution subject being a server or a client in the above embodiments are implemented. For example, when the processor executes the computer program, the following steps are implemented:

[0169] Step S101: Construct an imitation learning autonomous driving model, input the environmental image information, vehicle navigation instruction information, and vehicle speed information collected by the in-vehicle front camera and the left and right wide-angle cameras into the policy generation network. The policy generation network generates a driving action control signal, input the driving action control signal and the current environmental state into the discriminator network. The discriminator network outputs the probability score that the action-state pair belongs to the expert data set, and trains the policy generation network based on the probability score; the driving action control signal includes a throttle control amount, a steering control amount, and a braking control amount;

[0170] Step S102: Construct a multi-dimensional reward function model, collect the state data during vehicle driving, detect the vehicle's red-light running behavior, lane-changing and line-crossing behavior, deviation from the motor vehicle lane behavior, and vehicle collision behavior, record the distance the vehicle has traveled and the remaining driving distance, construct a driving behavior reward and punishment matrix based on the state detection results, and input the driving behavior reward and punishment matrix into the multi-dimensional reward function model to obtain the reward score of the vehicle's current driving state;

[0171] Step S103: Construct an actor-critic network, input the environmental image information, navigation instruction information, and vehicle speed information into the actor network. The actor network outputs the optimal driving action in the current state, input the optimal driving action and the current state into the critic network, and the critic network outputs the action value evaluation score. Optimize the parameters of the actor network based on the reward score and the action value evaluation score, and use the optimized actor network to generate an autonomous driving control instruction.

[0172] As can be seen from the above description, the computer-readable storage medium provided in the embodiments of the present application combines the imitation learning and reinforcement learning frameworks, collects environmental information through multiple cameras, constructs a policy generation network and a discriminator network to realize driving action control. Design a multi-dimensional reward function model to detect and evaluate behaviors such as running red lights, crossing lines, deviating from lanes, and collisions in real time, and construct a driving behavior reward and punishment matrix. Based on the actor-critic network architecture, input environmental information and navigation instructions into the actor network to generate the optimal driving action, and evaluate the action value through the critic network to realize dynamic parameter optimization. This method effectively solves the deficiencies of traditional technologies in aspects such as driving behavior evaluation and action value judgment, and significantly improves the safety and reliability of the autonomous driving system.

[0173] An embodiment of the present application also provides a computer program product that can implement all the steps of the autonomous driving method based on a multi-dimensional reward function whose execution subject in the above embodiments is a server or a client. When the computer program / instructions are executed by a processor, the steps of the autonomous driving method based on the multi-dimensional reward function are implemented. For example, the computer program / instructions implement the following steps:

[0174] Step S101: Construct an imitation learning autonomous driving model, input the environmental image information, vehicle navigation instruction information, and vehicle speed information collected by the on-vehicle front camera and left and right wide-angle cameras into the policy generation network. The policy generation network generates a driving action control signal, input the driving action control signal and the current environmental state into the discriminator network, and the discriminator network outputs the probability score that the action-state pair belongs to the expert data set. Train the policy generation network based on the probability score; the driving action control signal includes a throttle control amount, a steering control amount, and a braking control amount;

[0175] Step S102: Construct a multi-dimensional reward function model, collect the state data during the vehicle driving process, detect the vehicle's red-light running behavior, lane-changing and line-crossing behavior, deviation from the motor vehicle lane behavior, and vehicle collision behavior, record the distance the vehicle has traveled and the remaining driving distance, construct a driving behavior reward and punishment matrix based on the state detection results, and input the driving behavior reward and punishment matrix into the multi-dimensional reward function model to obtain the reward score of the vehicle's current driving state;

[0176] Step S103: Construct an actor-critic network, input the environmental image information, navigation instruction information, and vehicle speed information into the actor network, the actor network outputs the optimal driving action in the current state, input the optimal driving action and the current state into the critic network, the critic network outputs the action value evaluation score, optimize the parameters of the actor network based on the reward score and the action value evaluation score, and use the optimized actor network to generate an autonomous driving control instruction.

[0177] As can be seen from the above description, the computer program product provided by the embodiments of the present application combines the imitation learning and reinforcement learning frameworks, collects environmental information through multiple cameras, constructs a policy generation network and a discriminator network to achieve driving action control. Design a multi-dimensional reward function model to detect and evaluate behaviors such as running red lights, crossing lines, deviating from lanes, and collisions in real time, and construct a driving behavior reward and punishment matrix. Based on the actor-critic network architecture, input the environmental information and navigation instructions into the actor network to generate the optimal driving action, and evaluate the action value through the critic network to achieve dynamic parameter optimization. This method effectively solves the deficiencies of traditional technologies in aspects such as driving behavior evaluation and action value judgment, and significantly improves the safety and reliability of the autonomous driving system.

[0178] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, an apparatus, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0179] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (devices), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce a means for implementing the functions specified in one or more flows and / or blocks Figure 1 in one or more flows and / or blocks Figure 1 or in one or more blocks.

[0180] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the functions specified in one or more flows and / or blocks Figure 1 in one or more flows and / or blocks Figure 1 or in one or more blocks.

[0181] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows and / or blocks Figure 1 in one or more flows and / or blocks Figure 1 or in one or more blocks.

[0182] Specific embodiments are applied in the present invention to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, based on the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. An autonomous driving method based on a multi-dimensional reward function, characterized in that, The method includes: Constructing an imitation learning autonomous driving model, inputting the environmental image information, vehicle navigation instruction information, and vehicle speed information collected by the in-vehicle front camera and left and right wide-angle cameras into a policy generation network. The policy generation network generates a driving action control signal, inputs the driving action control signal and the current environmental state into a discriminator network. The discriminator network outputs the probability score of the action-state pair belonging to the expert data set, and trains the policy generation network based on the probability score; the driving action control signal includes a throttle control amount, a steering control amount, and a braking control amount; Constructing a multi-dimensional reward function model, collecting the state data during vehicle driving, detecting the vehicle's red-light running behavior, lane-changing and line-crossing behavior, deviation from the motor vehicle lane behavior, and vehicle collision behavior, recording the distance the vehicle has traveled and the remaining driving distance, constructing a driving behavior reward and punishment matrix based on the state detection results, and inputting the driving behavior reward and punishment matrix into the multi-dimensional reward function model to obtain the reward score of the vehicle's current driving state; Constructing an actor-critic network, inputting the environmental image information, navigation instruction information, and vehicle speed information into the actor network. The actor network outputs the optimal driving action in the current state, inputs the optimal driving action and the current state into the critic network. The critic network outputs an action value evaluation score, optimizes the parameters of the actor network based on the reward score and the action value evaluation score, and uses the optimized actor network to generate an autonomous driving control instruction.

2. The automatic driving method based on a multi-dimensional reward function according to claim 1, wherein The constructing of the imitation learning autonomous driving model, inputting the environmental image information, vehicle navigation instruction information, and vehicle speed information collected by the in-vehicle front camera and left and right wide-angle cameras into the policy generation network, and the policy generation network generating a driving action control signal includes: Collecting in-vehicle image information, installing a front camera at the front of the vehicle and wide-angle cameras on both the left and right sides of the vehicle, calibrating the lens focal length, field of view angle, image resolution, and sampling frequency of the cameras, constructing the environmental image data collected by the calibrated cameras, the navigation instruction data generated by the in-vehicle navigation system, and the speed information collected by the vehicle speed sensor as state input data, and normalizing the state input data; Inputting the normalized state input data into the policy generation network of the neural network structure, using a convolutional layer to extract image features, using a fully connected layer to fuse the image features, navigation features, and speed features, generating an action output vector of the throttle control amount, steering control amount, and braking control amount based on an activation function, inputting the action output vector and the current state of the vehicle into the discriminant model in the discriminator network, calculating the expert probability score of the state-action pair, and performing backpropagation optimization on the parameters of the policy generation network according to the expert probability score.

3. The autonomous driving method based on a multi-dimensional reward function according to claim 1, wherein Inputting the driving action control signal and the current environmental state into the discriminator network, the discriminator network outputs the probability score of the action-state pair belonging to the expert data set, and trains the policy generation network based on the probability score; The driving action control signal includes a throttle control amount, a steering control amount, and a braking control amount, and includes: Collect the demonstration data of expert drivers, record the environmental state information and corresponding driving action data during the expert driving process, label the expert driving data as positive samples, label the driving action data output by the policy generation network as negative samples, construct the training data set of the discriminator network, normalize the throttle control amount, steering control amount and braking control amount in the training data set, and splice the normalized action data with the corresponding environmental state data to construct state-action pairs; Input the state-action pairs into the discriminator network, use a multi-layer perceptron structure to perform feature mapping on the state-action pairs, calculate the similarity score between the state-action pairs and the expert data distribution, calculate the prediction error of the discriminator network based on the cross-entropy loss function, use the gradient descent method to optimize the discriminator network parameters, and use the similarity score output by the discriminator network as the training reward signal of the policy generation network to perform iterative optimization training on the policy generation network.

4. The autonomous driving method based on a multi-dimensional reward function according to claim 1, wherein The construction of the multi-dimensional reward function model, collect the state data during the vehicle driving process, detect the vehicle's red light running behavior, lane-changing and line-crossing behavior, deviation from the motor vehicle lane behavior and vehicle collision behavior, and record the distance the vehicle has traveled and the remaining distance, including: Construct a vehicle state detection sub-model, collect the traffic light state information of the lane where the vehicle is located, calculate the relative distance between the vehicle position and the traffic light position based on the electronic map data, collect the relative position information between the vehicle contour and the lane line, collect the distribution information of obstacles around the vehicle based on the lidar sensor, and calculate the remaining distance between the vehicle and the target end point according to the global positioning system data, and construct the vehicle state detection data into a state vector; Extract features from the state vector, judge whether the vehicle runs a red light based on the traffic light state and relative distance, judge whether there is a line-crossing or lane deviation behavior based on the relative position between the vehicle contour and the lane line, judge whether a collision occurs based on the obstacle distribution information, calculate the proportion of the distance the vehicle has traveled to the total driving distance, construct a violation behavior label vector, and input the state vector and the violation behavior label vector into the multi-dimensional reward function model for training.

5. The autonomous driving method based on a multi-dimensional reward function according to claim 1, wherein The construction of the driving behavior reward and punishment matrix based on the state detection results, input the driving behavior reward and punishment matrix into the multi-dimensional reward function model to obtain the reward score of the vehicle's current driving state, including: Construct a reward and punishment matrix according to the vehicle state detection results, assign a violation punishment weight for the vehicle's red light running behavior, assign a violation punishment weight for the vehicle's line-crossing and lane-changing behavior, assign a violation punishment weight for the vehicle's deviation from the motor vehicle lane behavior, assign a violation punishment weight for the vehicle's collision behavior, assign a task completion reward weight for the vehicle's arrival at the target position, and construct a driving behavior reward and punishment matrix based on the value of the task completion reward weight; Input the driving behavior reward and punishment matrix and the vehicle state detection vector into the multi-dimensional reward function model, use a deep neural network structure to extract features from the state detection vector, calculate the reward and punishment scores of each driving behavior based on the weight values in the reward and punishment matrix, and perform weighted summation on each reward and punishment score to generate a comprehensive reward score reflecting the vehicle's current driving state.

6. The automatic driving method based on a multi-dimensional reward function according to claim 1, characterized in that, Construct the actor evaluation network. Input the environmental image information, navigation instruction information, and vehicle speed information into the actor network. The actor network outputs the optimal driving action in the current state. Input the optimal driving action and the current state into the evaluation network. The evaluation network outputs an action value evaluation score, including: Construct an actor network model. Use a multi-layer convolutional neural network to extract the visual features of the environmental image. Use a fully connected layer to process the navigation instruction information and vehicle speed information. Perform feature fusion on the visual features, navigation features, and vehicle speed features. Map the fused features to the action space distribution through the policy head network. Sample based on the action space distribution to generate the optimal driving actions of the throttle control amount, steering control amount, and braking control amount; Construct an evaluation network model. Concatenate the features of the optimal driving action and the current environmental state. Use a multi-layer perceptron network to perform a non-linear transformation on the state-action features. Map the transformed features to a scalar value evaluation score based on the value head network. Perform weighted fusion on the value evaluation score and the reward score output by the reward function model. Construct a loss function for the actor network, and optimize and update the parameters of the actor network based on the policy gradient method.

7. The autonomous driving method based on a multi-dimensional reward function according to claim 1, wherein Optimize the parameters of the actor network based on the reward score and the action value evaluation score. Use the optimized actor network to generate an autonomous driving control instruction, including: Use the proximal policy optimization algorithm to construct a parameter optimization model for the actor network. Compare and calculate the advantage function value between the reward score output by the reward function model and the value evaluation score output by the evaluation network. Construct a policy gradient based on the advantage function value to constrain the difference between the action distribution output by the actor network and the target policy distribution. Use the trust region policy optimization method to limit the step size of each parameter update. Iteratively optimize the parameters of the actor network based on the stochastic gradient descent method; Deploy the optimized actor network to the vehicle-mounted computing platform. Collect the vehicle environmental image information, navigation instruction information, and vehicle speed information to construct a state input vector. Input the state input vector into the optimized actor network. Generate driving control instructions for the throttle control amount, steering control amount, and braking control amount based on the action output layer of the actor network. Send the driving control instructions to the vehicle execution system.

8. An autonomous driving device based on a multi-dimensional reward function, characterized in that, The device includes: A driving model construction module, used to construct an imitation learning autonomous driving model. Input the environmental image information, vehicle navigation instruction information, and vehicle speed information collected by the vehicle-mounted front camera and left and right wide-angle cameras into the policy generation network. The policy generation network generates a driving action control signal. Input the driving action control signal and the current environmental state into the discriminator network. The discriminator network outputs the probability score that the action-state pair belongs to the expert data set. Train the policy generation network based on the probability score; the driving action control signal includes the throttle control amount, steering control amount, and braking control amount; A reward function construction module, which is used to construct a multi-dimensional reward function model, collect state data during the vehicle driving process, detect the vehicle's red-light running behavior, lane-changing and line-crossing behavior, behavior of deviating from the motor vehicle lane, and vehicle collision behavior, record the distance the vehicle has traveled and the remaining driving distance, construct a driving behavior reward and punishment matrix based on the state detection results, and input the driving behavior reward and punishment matrix into the multi-dimensional reward function model to obtain the reward score of the vehicle's current driving state; An autonomous driving execution module, which is used to construct an actor-critic network, input the environmental image information, navigation instruction information, and vehicle speed information into the actor network, the actor network outputs the optimal driving action in the current state, input the optimal driving action and the current state into the critic network, the critic network outputs the action value evaluation score, optimize the parameters of the actor network based on the reward score and the action value evaluation score, and generate an autonomous driving control instruction using the optimized actor network.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the autonomous driving method based on a multi-dimensional reward function according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the autonomous driving method based on a multi-dimensional reward function according to any one of claims 1 to 7.

Citation Information

Cited By

  • Emergency vehicle priority and queuing optimization cooperative control method

    CN121214701A

  • Artificial intelligence model training method and device, computer equipment, medium and product

    CN121525780A

  • Artificial intelligence model training method and apparatus, computer device, medium, and product

    CN121525780B

  • End-to-end intelligent driving training system, method and device and storage medium

    CN122021363A

  • End-to-end intelligent driving training system, method, device and storage medium

    CN122021363B