Intelligent driving simulation test method and system for internal combustion locomotive
By building a virtual simulation environment and a multi-sensor fusion system, combined with a deep reinforcement learning network, the problem of insufficient authenticity of simulation environments and perceived data in the existing technology is solved, and efficient simulation testing and driving strategy optimization of the intelligent driving system of internal combustion locomotives is realized.
Patent Information
- Application Number
- CN202510355217.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-25
AI Technical Summary
The existing simulation test methods for intelligent driving of internal combustion locomotives have insufficient authenticity and complexity in the simulation environment, the perceived data simulation is not accurate enough, and it is difficult to fully simulate the actual driving scenarios, which affects the reliability of the test results.
By building a virtual simulation environment, the driving status information and scene complexity data of the internal diesel locomotive are collected, and the three-dimensional point cloud data, distance information, relative speed information, and image feature information of obstacles are obtained by using virtual lidar, virtual millimeter wave radar and virtual vision sensors, and multi-modal fusion processing is carried out to generate standardized perception data. At the same time, a deep reinforcement learning network is built, including prediction branch networks and dual-channel control networks, and driving control parameters are generated.
It realizes efficient simulation testing of the intelligent driving system of internal combustion locomotives, improves the accuracy of obstacle perception and scenario understanding, enhances the dynamic optimization ability of driving strategies, and ensures driving safety and adaptability in complex environments.
Smart Images

Figure CN120178700A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to intelligent driving technology, and in particular to an intelligent driving simulation test method and system for diesel locomotives. Background Art
[0002] Currently, the research and development and testing of intelligent driving systems for diesel locomotives mainly rely on real vehicle road tests. Although this method can obtain real driving data, it has problems such as high cost, low efficiency, and high safety risks. Especially in the testing of complex scenarios and extreme conditions, it is often difficult to fully cover in the actual environment. The existing intelligent driving simulation test methods for diesel locomotives mainly have the following deficiencies: the authenticity and complexity of the simulation environment are insufficient, and it is difficult to comprehensively simulate various complex situations in the actual driving scenario. The simulation of perception data is not accurate enough, especially in the aspect of multi-sensor fusion, which affects the reliability of the simulation test results. The generation of driving strategies lacks the ability to dynamically adapt to the complexity of the scenario and it is difficult to make optimal decisions for different scenarios. Summary of the Invention
[0003] Embodiments of the present invention provide an intelligent driving simulation test method and system for diesel locomotives, which can solve the problems in the prior art.
[0004] In the first aspect of the embodiments of the present invention, An intelligent driving simulation test method for diesel locomotives is provided, including: Constructing a virtual simulation environment; collecting the driving state information of the diesel locomotive; obtaining scene complexity data, where the scene complexity data includes the road curvature change rate, the number of obstacles in the scene, and the obstacle movement complexity; In the virtual simulation environment, obtaining three-dimensional point cloud data of obstacles through a virtual lidar; obtaining the distance information and relative speed information of obstacles through a virtual millimeter-wave radar and simulating the measurement error caused by multipath effects; obtaining the image feature information of obstacles through a virtual vision sensor and simulating the imaging quality fluctuation caused by light changes; performing multi-modal fusion processing on the three-dimensional point cloud data, distance information, relative speed information, and image feature information to generate standardized perception data; Constructing a deep reinforcement learning network, including a prediction branch network for generating multi-modal prediction trajectories and a dual-channel control network for generating driving strategies; wherein, for conventional driving strategies, the safety index is set as the network constraint condition, and the weights of the comfort and smoothness indexes increase with the increase of the scene complexity, and the weight of the efficiency index decreases accordingly; for emergency avoidance driving strategies, safety constraint optimization is performed based on collision risk assessment; Inputting the driving state information, standardized perception data, and scene complexity data into the deep reinforcement learning network to generate driving control parameters.
[0005] In an alternative embodiment, the generation of the standardized perception data includes: obtaining three-dimensional point cloud data of obstacles through a virtual lidar, filtering the three-dimensional point cloud data according to the local point cloud center and standard deviation to obtain filtered point cloud data; obtaining distance information and relative velocity information of obstacles by transmitting a frequency-modulated continuous wave signal through a virtual millimeter-wave radar, and compensating for multipath effect errors of the distance information and relative velocity information according to the time delay, attenuation coefficient, and Doppler frequency shift of multiple propagation paths; obtaining image feature information of obstacles through a virtual vision sensor, compensating for illumination changes of the image feature information according to the local illumination intensity to obtain image feature data; performing spatio-temporal alignment on the filtered point cloud data, compensated distance information and relative velocity information, and image feature data, and uniformly converting the data to a global coordinate system according to timestamps; performing multi-modal fusion on the spatio-temporally aligned data, performing weighted feature combination according to the weights of each modal feature to obtain a fusion feature, and performing information fusion on the fusion feature according to the basic probability assignment and conflict factor; performing standardized processing and quality assessment on the fused data, calculating information entropy according to the feature probability distribution, and generating standardized perception data.
[0006] In an alternative embodiment, the deep reinforcement learning network includes: The prediction branch network receives the position, velocity, and acceleration information of obstacles in the standardized perception data, extracts trajectory features through a multi-head self-attention encoder, and generates a multi-modal prediction trajectory based on a recurrent neural network decoder; constructing the multi-modal prediction trajectory, driving state information, and scene complexity data into a state space vector and inputting it into a dual-channel control network, where the dual-channel control network includes a regular driving channel and an emergency avoidance channel; The regular driving channel generates a regular driving strategy based on the Actor-Critic network structure, sets the safety index as a constraint condition of the Actor-Critic network, the weights of the comfort index and the smoothness index increase with the increase of the scene complexity, the weight of the efficiency index decreases accordingly, and optimizes the regular driving strategy through the temporal difference error; the emergency avoidance channel generates an avoidance driving strategy based on collision risk assessment and optimizes the avoidance driving strategy through safety constraints; evaluating the confidence levels of the regular driving strategy and the avoidance driving strategy based on spatio-temporal collision risk and minimum safety distance, mapping the confidence levels into dynamic fusion weights; and performing weighted fusion on the regular driving strategy and the avoidance driving strategy according to the dynamic fusion weights to generate driving control parameters.
[0007] In an alternative embodiment, Evaluating the confidence levels of the normal driving strategy and the emergency avoidance driving strategy based on the spatio-temporal collision risk and the minimum safety distance, and mapping the confidence levels to dynamic fusion weights includes: Calculating the collision probability in the time dimension based on the multi-modal prediction trajectory, substituting the deviation between the time collision point of the multi-modal prediction trajectory and the expected collision time into the Gaussian distribution function, and generating a collision risk value in the time dimension according to the trajectory weight; calculating the collision probability in the space dimension based on the position probability density function and the speed constraint function, and performing an integral operation on the position probability density function and the speed constraint function in the collision danger area to generate a collision risk value in the space dimension; Calculating the dynamic safety distance according to the vehicle's own speed, the leading vehicle's speed, the reaction time, the maximum deceleration and the minimum deceleration, and determining the ratio of the actual vehicle distance to the dynamic safety distance as the safety margin coefficient; Performing a weighted combination of the inverse values of the collision risk value in the time dimension and the collision risk value in the space dimension and the safety margin coefficient to generate the confidence level of the normal driving strategy; using the weighted sum of the exponentially decaying value of the safety margin coefficient and the spatio-temporal dimension collision risk value as the confidence level of the emergency avoidance driving strategy; Substituting the confidence level of the normal driving strategy and the confidence level of the emergency avoidance driving strategy into the Softmax mapping function for normalization processing to generate the initial fusion weight; performing an exponential moving average process on the initial fusion weight to generate the dynamic fusion weight.
[0008] In an alternative embodiment, The training of the deep reinforcement learning network includes: Constructing a training curriculum based on the scene complexity index, normalizing the road curvature change rate, the number of obstacles in the scene and the obstacle movement complexity to generate a scene difficulty score, grading the training samples according to the scene difficulty score, and establishing a curriculum sequence from simple to complex; Adopting a two-stage training strategy, performing behavior cloning training based on the human driver demonstration data in the pre-training stage to enable the network to obtain basic driving capabilities, and introducing an experience replay mechanism during the training based on the curriculum sequence in the reinforcement learning stage; Constructing a dynamic curriculum adjustment mechanism, setting training efficiency evaluation indicators including the convergence speed of the policy network and the improvement amplitude of the evaluation index, adaptively adjusting the curriculum difficulty growth step and the number of training rounds for each difficulty level based on the training efficiency evaluation results, and automatically reverting to the training scenario with a difficulty level reduced by 1 for supplementary training when the training efficiency is lower than the preset efficiency threshold; Set a multi-dimensional scoring standard to evaluate the control smoothness and operation efficiency of the regular driving lane under different scenario complexities, and evaluate the timeliness and safety margin of the emergency avoidance strategy for the emergency avoidance lane. When the weighted result of the multi-dimensional scoring is greater than the preset scoring threshold, switch to the training scenario with the difficulty level increased by 1.
[0009] In an alternative embodiment, The dynamic curriculum adjustment mechanism includes: Construct a multi-dimensional training efficiency evaluation system, take the convergence speed and the improvement amplitude of the evaluation index as short-term efficiency indicators, and construct long-term efficiency indicators based on the generalization performance and anti-interference ability of the policy network in the validation scenario; perform weighted combination on the short-term efficiency indicators and the long-term efficiency indicators to generate a comprehensive training efficiency value; Establish a distributed training architecture, allocate the training tasks of the curriculum sequence to multiple parallel training environments, set different random perturbation parameters in each training environment, perform parallel training based on the asynchronous gradient update method, compare the performance of the policy network in different parallel training environments, and generate a stability evaluation value of the policy network; Determine the curriculum difficulty adjustment direction according to the comprehensive training efficiency value and the stability evaluation value. When the comprehensive training efficiency value is greater than the preset efficiency upper limit and the stability evaluation value is greater than the preset stability threshold, increase the curriculum difficulty level according to the adaptive step size; when the comprehensive training efficiency value is less than the preset efficiency threshold, decrease the curriculum difficulty level by 1; Construct a hierarchical storage structure for the experience replay mechanism, store training samples of different difficulty levels in the corresponding experience pool levels; during the supplementary training process of decreasing the difficulty level by 1, perform importance sampling from the experience pool of the corresponding difficulty level to generate a supplementary training sequence; Set a curriculum difficulty smooth transition mechanism. During the process of difficulty level conversion, calculate the mixing weights of adjacent difficulty levels based on the exponential decay function, and dynamically combine the training samples of different difficulty levels according to the mixing weights to generate a transition training sample set.
[0010] In an alternative embodiment, Constructing a virtual simulation environment includes: Determine the road geometric features and slope features based on the road parameters; set the type parameters and layout position parameters of the road facility features; set the environmental parameters of the weather features; Establish a feature combination constraint rule library, including: matching rules for road geometric features and slope features, layout rules for road facility features, and combination rules for weather features; Based on the feature combination constraint rule library, screen out the feature combinations that meet the constraint conditions, construct a feature fitness matrix, and the fitness value of the feature fitness matrix represents the matching degree of the corresponding feature combination. Automatically combine features using a hierarchical combination method, including: selecting the road geometric feature and slope feature with the highest fitness value to generate a basic road section, selecting the road facility feature with the highest fitness value on the basic road section for layout, and selecting the weather feature with the highest fitness value for superposition to generate a complete scene; Use the Monte Carlo method to randomly sample the feature combination result, and apply the Markov chain for feature transfer to generate a composite test scene.
[0011] In an alternative embodiment, Using the Monte Carlo method to randomly sample the feature combination result and applying the Markov chain for feature transfer to generate a composite test scene includes: Calculating the sum of the product of the scene coverage rate and the first preset weight and the product of the rare event probability and the second preset weight to obtain the scene importance; calculating the sampling temperature parameter according to the scene importance, and substituting the scene importance and the sampling temperature parameter into the Softmax function to generate the scene sampling probability; Performing Monte Carlo sampling on the feature combination result according to the scene sampling probability to obtain scene features; inputting the scene features and scene condition constraints into a conditional variational autoencoder to generate a mean parameter and a variance parameter; Performing Gaussian distribution sampling based on the mean parameter and the variance parameter to obtain a latent variable; constructing a Markov transition matrix for the latent variable, and performing feature migration on the latent variable based on the Markov transition matrix to generate a migrated latent variable; Decoding the migrated latent variable and the scene condition constraints to generate migrated scene features; calculating the KL divergence between the migrated scene features and the original scene features, and when the KL divergence is less than a preset divergence threshold, determining the migrated scene features as an effective composite test scene.
[0012] In the second aspect of the embodiments of the present invention, Provide an intelligent driving simulation test system for diesel locomotives, including: A first unit for constructing a virtual simulation environment; collecting the driving state information of a diesel locomotive; obtaining scene complexity data, where the scene complexity data includes the road curvature change rate, the number of obstacles in the scene, and the obstacle movement complexity; A second unit is configured to, in the virtual simulation environment, obtain three-dimensional point cloud data of obstacles through a virtual lidar; obtain distance information and relative speed information of obstacles through a virtual millimeter-wave radar, and simulate measurement errors caused by multipath effects; obtain image feature information of obstacles through a virtual vision sensor, and simulate fluctuations in imaging quality caused by changes in illumination; perform multi-modal fusion processing on the three-dimensional point cloud data, distance information, relative speed information, and image feature information to generate standardized perception data; A third unit is configured to construct a deep reinforcement learning network, including a prediction branch network for generating multi-modal prediction trajectories and a dual-channel control network for generating driving strategies; wherein, for conventional driving strategies, a safety index is set as a network constraint condition, and the weights of comfort and smoothness indexes increase with the increase of scene complexity, while the weight of the efficiency index decreases accordingly; for emergency avoidance driving strategies, safety constraint optimization is performed based on collision risk assessment; the driving state information, standardized perception data, and scene complexity data are input into the deep reinforcement learning network to generate driving control parameters.
[0013] By constructing a virtual simulation environment and a multi-sensor fusion system, the present invention can comprehensively and accurately obtain various information of the operating environment of diesel locomotives, improve the accuracy of obstacle perception and scene understanding, and provide a reliable data basis for intelligent driving decisions.
[0014] The present invention adopts a deep reinforcement learning network architecture. Through the coordinated cooperation of the prediction branch network and the dual-channel control network, it realizes the accurate prediction of multi-modal trajectories and the dynamic optimization of driving strategies, enabling the system to adaptively adjust control parameters according to different scene complexities, and taking into account comfort and efficiency while ensuring safety.
[0015] The present invention dynamically adjusts the control strategy weights based on scene complexity, and introduces a collision risk assessment mechanism for safety constraint optimization, significantly improving the driving safety and adaptability of diesel locomotives in complex environments, while ensuring the smoothness of the driving experience, and achieving a comprehensive balance in multiple aspects such as safety, comfort, and efficiency of the intelligent driving system. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a schematic flow chart of the intelligent driving simulation test method for a diesel locomotive according to an embodiment of the present invention; Figure 2 is a simulation effect diagram of the automatic driving system of the present invention in a complex scene; Figure 3 is a comparison diagram of safety performance under different scenes; Figure 4 is a comparison diagram of the reward curve and the difficulty adjustment path during the training process. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0018] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments may be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0019] Figure 1 is a schematic flow chart of the intelligent driving simulation test method for diesel locomotives in the embodiments of the present invention. As Figure 1 shown, the method includes: Constructing a virtual simulation environment; collecting the driving state information of the diesel locomotive; obtaining scene complexity data, where the scene complexity data includes the road curvature change rate, the number of obstacles in the scene, and the obstacle movement complexity; In the virtual simulation environment, obtaining the three-dimensional point cloud data of the obstacle through a virtual lidar; obtaining the distance information and relative speed information of the obstacle through a virtual millimeter-wave radar and simulating the measurement error caused by multipath effects; obtaining the image feature information of the obstacle through a virtual vision sensor and simulating the imaging quality fluctuation caused by light changes; performing multi-modal fusion processing on the three-dimensional point cloud data, distance information, relative speed information, and image feature information to generate standardized perception data; Constructing a deep reinforcement learning network, including a prediction branch network for generating multi-modal prediction trajectories and a dual-channel control network for generating driving strategies; among them, for the conventional driving strategy, the safety index is set as the network constraint condition, and the weights of the comfort and smoothness indexes increase with the increase of the scene complexity, and the weight of the efficiency index decreases accordingly; for the emergency avoidance driving strategy, safety constraint optimization is performed based on the collision risk assessment. Inputting the driving state information, standardized perception data, and scene complexity data into the deep reinforcement learning network to generate driving control parameters.
[0020] In an alternative embodiment, the generation of the standardized perception data includes: Obtaining the three-dimensional point cloud data of the obstacle through a virtual lidar, and filtering the three-dimensional point cloud data according to the local point cloud center and standard deviation to obtain the filtered point cloud data; The virtual millimeter-wave radar transmits a frequency-modulated continuous wave signal to obtain the distance information and relative velocity information of the obstacle, and compensates the multi-path effect error of the distance information and relative velocity information according to the time delay, attenuation coefficient, and Doppler frequency shift of multiple propagation paths; The virtual vision sensor is used to obtain the image feature information of the obstacle, and the image feature information is compensated for illumination changes according to the local illumination intensity to obtain image feature data; Perform spatio-temporal alignment on the filtered point cloud data, compensated distance information and relative velocity information, and image feature data, and uniformly convert the data to the global coordinate system according to the time stamp; Perform multi-modal fusion on the spatio-temporally aligned data, perform weighted feature combination according to the weights of each modal feature to obtain a fusion feature, and perform information fusion on the fusion feature according to the basic probability assignment and conflict factor; Perform normalization processing and quality assessment on the fused data, calculate the information entropy according to the feature probability distribution, and generate normalized perception data.
[0021] Exemplarily, the virtual lidar is used to obtain the three-dimensional point cloud data of the obstacle. The virtual lidar can simulate the scanning process of the actual lidar and generate point cloud data containing information such as the position and shape of the obstacle. For example, for a moving car, point cloud data containing features such as the body contour, windows, and wheels can be obtained. Then, the obtained point cloud data is filtered. Specifically, the central coordinates and standard deviation of the local point cloud are calculated, and the abnormal points that deviate too far from the center are removed according to the set threshold. For example, twice the standard deviation can be set as the threshold, and the points that deviate from the center by more than this threshold are removed. This can effectively remove the noise points caused by measurement errors or environmental interference and improve the quality of the point cloud data.
[0022] The virtual millimeter-wave radar is used to obtain the distance and relative velocity information of the obstacle. The virtual millimeter-wave radar simulates the transmission of a frequency-modulated continuous wave signal and measures the distance and speed of the target by receiving the echo signal. For example, for a vehicle approaching at a speed of 60 km / h at a distance of 100 meters ahead, measurement results of a distance of 100 m and a relative velocity of -16.7 m / s can be obtained. Compensate the multi-path effect error of the obtained distance and velocity information. Since electromagnetic waves will be reflected and scattered during propagation, multiple propagation paths are formed, resulting in measurement errors. When compensating, consider the time delay, attenuation coefficient, and Doppler frequency shift of each path, and correct the original measurement results. For example, the main reflection path can be estimated according to the environmental model, the error introduced by it can be calculated, and this error can be subtracted from the measured value.
[0023] Obtain the image feature information of obstacles through a virtual vision sensor. The virtual vision sensor simulates a camera to capture images of the surrounding environment. For obstacles, extract their color, texture, edge and other features as image feature information. For example, for a red car, features such as its body color, window position, and headlight shape can be extracted. Compensate for the illumination change of the extracted image features. Since the change of natural illumination conditions will affect the brightness and contrast of the image, compensation is required to improve the stability of the features. Specifically, the brightness equalization process can be performed on the image according to the average brightness of the local area. For example, the image can be divided into blocks, and the brightness of each block is normalized so that its mean value remains within a fixed range.
[0024] Perform spatio-temporal alignment on the processed point cloud data, distance and speed information, and image feature data. First, align the data to the same moment according to the timestamps of each sensor. Then, according to the installation positions and postures of each sensor relative to the vehicle body, convert the data into the global coordinate system. For example, the center of the rear axle of the vehicle can be selected as the origin to establish a right-handed coordinate system, and all data is converted to this coordinate system.
[0025] Perform multi-modal data fusion. Determine the weights according to the reliability and accuracy of each modal data. For example, a weight of 0.5 can be given to the point cloud data, 0.3 to the distance and speed information, and 0.2 to the image features according to experience. Then combine the weighted features to obtain the fusion features.
[0026] Perform information fusion on the fusion features using the basic probability assignment method. Calculate the basic probability according to the confidence of each feature, and introduce a conflict factor to handle the contradiction of information from different sources. For example, when the point cloud shows an obstacle ahead while the radar does not detect it, the contradiction can be mediated according to their reliability.
[0027] Finally, perform standardization processing and quality assessment on the fused data. Normalize the features of each dimension to a unified numerical range, such as the interval [0, 1]. Then calculate the information entropy according to the probability distribution of the features to evaluate the uncertainty of the data. The lower the information entropy, the higher the data quality. For example, data with an entropy value less than 0.1 can be marked as high quality, between 0.1 and 0.5 as medium quality, and greater than 0.5 as low quality.
[0028] Through the above steps, finally generate standardized perception data, including comprehensive information such as the position, speed, shape, and appearance of obstacles, as well as the quality assessment results of the data.
[0029] Through multi-sensor virtual simulation and data fusion, the present invention realizes all-round and multi-angle perception of obstacles, improving the comprehensiveness and accuracy of perception data. The virtual lidar provides accurate three-dimensional geometric information, the virtual millimeter-wave radar provides reliable distance and speed measurements, and the virtual vision sensor provides rich appearance features. The fusion of multi-source information greatly enhances the perception ability. A series of data processing and optimization techniques are adopted to effectively improve the quality and reliability of perception data. Point cloud filtering removes abnormal points, multipath effect compensation improves radar measurement accuracy, and illumination change compensation enhances the stability of visual features. These processes ensure that the subsequent fused data has high quality. Through spatio-temporal alignment, multi-modal fusion, and normalization processing, unified expression of heterogeneous data is achieved, facilitating the development and application of subsequent algorithms. Quality assessment provides a quantitative index for the reliability of data, contributing to improving the accuracy and robustness of decision-making. The standardized data format also facilitates data exchange and interoperability between different systems.
[0030] In an alternative embodiment, the deep reinforcement learning network includes: The prediction branch network receives the position, speed, and acceleration information of the obstacles in the standardized perception data, extracts trajectory features through a multi-head self-attention encoder, and generates multi-modal prediction trajectories based on a recurrent neural network decoder; Construct the multi-modal prediction trajectories, driving state information, and scene complexity data into a state space vector and input it into the dual-channel control network, where the dual-channel control network includes a regular driving channel and an emergency avoidance channel; The regular driving channel generates a regular driving strategy based on the Actor-Critic network structure, sets the safety index as a constraint condition of the Actor-Critic network, the weights of the comfort index and the smoothness index increase with the increase of the scene complexity, and the weight of the efficiency index decreases accordingly. Optimize the regular driving strategy through the temporal difference error; the emergency avoidance channel generates an avoidance driving strategy based on the collision risk assessment and optimizes the avoidance driving strategy through safety constraints; Evaluate the confidence of the regular driving strategy and the avoidance driving strategy based on the spatio-temporal collision risk and the minimum safety distance, map the confidence to a dynamic fusion weight; perform weighted fusion on the regular driving strategy and the avoidance driving strategy according to the dynamic fusion weight to generate driving control parameters.
[0031] The present invention proposes an autonomous driving control method based on deep reinforcement learning, which includes a prediction branch network, a dual-channel control network, and a policy fusion module.
[0032] The prediction branch network receives the normalized perception data, including the position, speed, and acceleration information of surrounding obstacles. This network uses an encoder with a multi-head self-attention mechanism to extract the trajectory features of the obstacles. Specifically, the input data is divided into multiple subspaces, and the attention weights are calculated independently for each subspace, and then the results of each subspace are combined. This can capture feature information at different scales. Then, a recurrent neural network decoder based on long short-term memory (LSTM) units is used to generate multiple possible predicted trajectories. For example, for a certain obstacle, multiple trajectory prediction results such as going straight, turning left, and turning right may be generated.
[0033] The multi-modal predicted trajectories are combined with the driving state information of the current vehicle (such as speed, acceleration, steering wheel angle, etc.) and the scene complexity data (such as road type, traffic flow, obstacle motion complexity determined by the weighted standard deviation of the curvature, speed, and acceleration changes of the obstacle trajectory and the inverse ratio of the distance) to construct a state space vector. This vector is passed as an input to a dual-channel control network, including a regular driving channel and an emergency avoidance channel.
[0034] The regular driving channel adopts an Actor-Critic network structure. The Actor network generates action policies according to the current state, such as accelerating, decelerating, steering, etc. The Critic network evaluates the value of the actions. During the training process, safety indicators (such as the minimum distance from other vehicles) are set as constraints. As the scene complexity increases, the weights of comfort indicators (such as the acceleration change rate) and smoothness indicators (such as the steering angle change rate) increase accordingly, while the weight of the efficiency indicator (such as the average speed) decreases. By calculating the temporal difference error between the actual obtained reward and the expected reward, the regular driving strategy is continuously optimized.
[0035] The emergency avoidance channel evaluates the collision risk based on the current state and generates an emergency avoidance driving strategy. Specifically, the trajectory generated by the prediction branch network is used to calculate the minimum spatio-temporal distance from the surrounding obstacles. When this distance is less than a preset threshold, the generation of the emergency avoidance strategy is triggered. The optimization goal of the emergency avoidance strategy is to maximize the safety distance while considering vehicle dynamics constraints.
[0036] Through the prediction branch network that combines the multi-head self-attention mechanism and the LSTM structure, the present invention realizes the multi-scale extraction of obstacle motion features and multi-modal trajectory prediction. This network can simultaneously focus on feature information at different spatio-temporal scales and generate various possible motion trajectories including going straight, turning, etc. For example, in a complex intersection scenario, the system can simultaneously predict the possible turning intentions of multiple vehicles, providing more comprehensive information support for subsequent decision-making and control. The dual-channel control network architecture is adopted, and the generation and evaluation of conventional driving strategies are realized through the Actor-Critic structure, and the weights of various indicators are dynamically adjusted according to the scene complexity. When the scene complexity increases, the system will automatically increase the weights of comfort and smoothness indicators and reduce the weights of efficiency indicators, making the control more conservative. At the same time, an independent emergency avoidance channel is set up to ensure that a safety-oriented avoidance strategy can be quickly generated in case of danger. An innovative strategy fusion mechanism based on spatio-temporal collision risk is proposed.
[0037] In an alternative embodiment, the confidence levels of the conventional driving strategy and the avoidance driving strategy are evaluated based on spatio-temporal collision risk and the minimum safety distance, and mapping the confidence levels to dynamic fusion weights includes: Calculating the collision probability in the time dimension based on the multi-modal prediction trajectory, substituting the deviation between the time collision point of the multi-modal prediction trajectory and the expected collision time into the Gaussian distribution function, and generating a collision risk value in the time dimension according to the trajectory weight; calculating the collision probability in the space dimension based on the position probability density function and the speed constraint function, and performing an integral operation on the position probability density function and the speed constraint function in the collision danger area to generate a collision risk value in the space dimension; Calculating the dynamic safety distance according to the speed of the host vehicle, the speed of the preceding vehicle, the reaction time, the maximum deceleration and the minimum deceleration, and determining the ratio of the actual vehicle distance to the dynamic safety distance as the safety margin coefficient; Performing a weighted combination of the inverse values of the collision risk value in the time dimension, the collision risk value in the space dimension and the safety margin coefficient to generate the confidence level of the conventional driving strategy; using the weighted sum of the exponentially decaying value of the safety margin coefficient and the collision risk value in the spatio-temporal dimension as the confidence level of the avoidance driving strategy; Substituting the confidence level of the conventional driving strategy and the confidence level of the avoidance driving strategy into the Softmax mapping function for normalization processing to generate the initial fusion weight; performing an exponentially weighted moving average process on the initial fusion weight to generate the dynamic fusion weight.
[0038] Exemplarily, the collision risk in the time dimension is evaluated. By predicting the trajectories of surrounding vehicles, multiple predicted trajectories with weights are generated. For each predicted trajectory, the possible collision time point with the host vehicle is calculated, and the time difference between this time point and the expected collision time is substituted into the Gaussian distribution function. When specifically implemented, the mean of the Gaussian distribution can be set to 3 seconds and the standard deviation to 1 second. Multiply the collision probability of each trajectory by the corresponding weight and sum them up to obtain the collision risk value in the time dimension. For example, if the predicted collision time of a certain trajectory is 2.5 seconds and the weight is 0.6, the calculated collision risk of this trajectory is 0.8.
[0039] Calculate the collision risk in the space dimension. Based on the current state of the vehicle, establish the position probability density function of the two-dimensional Gaussian distribution, and at the same time consider the vehicle kinematic constraints to establish the speed constraint function. Numerically integrate the two functions within the predefined collision danger area to obtain the collision risk value in the space dimension. In practical applications, the danger area can be set as a rectangular area 20 meters in front of the vehicle and 2 meters on each side. When a vehicle is detected in this area, the risk calculation is performed.
[0040] Calculate the dynamic safety distance. According to parameters such as the speed of the host vehicle at 80 kilometers per hour, the speed of the leading vehicle at 60 kilometers per hour, the driver's reaction time of 1 second, the maximum deceleration of 8 meters per square second, and the minimum deceleration of 2 meters per square second, the calculated safe following distance is 50 meters. Take the ratio of the actual vehicle distance to the safety distance as the safety margin coefficient. For example, when the actual vehicle distance is 40 meters, the safety margin coefficient is 0.8.
[0041] Generate the confidence of the driving strategy. Combine the collision risk value of 0.8 in the time dimension, the inverse value of 0.3 of the collision risk value in the space dimension, and the safety margin coefficient of 0.8 with weights of 0.4, 0.3, and 0.3 respectively to obtain the confidence of the conventional driving strategy of 0.64. Perform exponential decay processing on the safety margin coefficient and perform weighted summation with the collision risk values in the space-time dimension to obtain the confidence of the risk avoidance driving strategy of 0.75.
[0042] Finally, perform weight mapping. Substitute the confidence of the two strategies into the Softmax function for normalization to obtain the initial fusion weights of 0.45 and 0.55 respectively. Use the exponential moving average method with a decay coefficient of 0.8 to smooth the initial weights, and finally obtain the dynamic fusion weights of 0.42 and 0.58.
[0043] Figure 2The simulation results of the automatic driving control method of the present invention in a complex intersection scenario are shown. During the simulation process, the vehicle on the left (with a darker color) represents the host vehicle, and the vehicle below (with a lighter color) represents surrounding traffic participants. The multi-modal trajectory prediction results are shown in the upper left corner of the figure, including different possible driving paths and their corresponding probabilities (going straight: 0.65, turning left: 0.25, turning right: 0.10). The risk assessment results in the spatio-temporal dimension are shown in the upper right corner. The collision risk value in the time dimension is 0.32, the collision risk value in the space dimension is 0.28, and the comprehensive risk score is 0.30. The output of the dual-channel control network and the fusion result are shown at the bottom. The weight of the conventional driving control is 0.72, and the weight of the emergency avoidance control is 0.28. The fused control command includes an acceleration of -0.5 m / s² and a steering wheel angle of 2.3°. The simulation results show that the method of the present invention can accurately identify potential risks in complex scenarios and generate safe and stable control strategies to effectively cope with the challenges of various driving scenarios.
[0044] Figure 3 The safety performance comparison of three different automatic driving control methods in various scenarios is shown, and it is evaluated through a key indicator of the collision accident rate (percentage). The chart clearly shows the safety advantages of the dual-channel control architecture (square markers) of the present invention compared with the single control strategy of DDPG (triangle markers) and the single control strategy of SAC (round markers) in three typical scenarios. In the highway environment, the collision accident rate of the control architecture of the present invention is only 0.8%, significantly lower than 2.5% of DDPG and 2.1% of SAC. As the scenario complexity increases, this advantage becomes more obvious. Under urban road conditions, the present invention method maintains a relatively low collision rate of 1.5%, while DDPG and SAC are 4.2% and 3.8% respectively. In the most challenging complex intersection scenario, the collision accident rate of the present invention method is 2.1%, only 28.8% of the DDPG method (7.3%) and 32.3% of the SAC method (6.5%).
[0045] The present invention realizes the accurate quantification of spatio-temporal collision risks through multi-modal trajectory prediction and probability density analysis, improving the accuracy and reliability of risk assessment; based on the calculation method of dynamic safety distance and safety margin coefficient, it realizes the adaptive assessment of the safety state of the driving environment, enhancing the adaptability of the system to complex traffic scenarios; adopting a strategy confidence mapping and dynamic weight fusion mechanism, it realizes the smooth switching between conventional driving and emergency avoidance driving strategies, improving the safety and comfort of the automatic driving system.
[0046] In an alternative embodiment, the training of the deep reinforcement learning network includes: Construct training courses based on scene complexity metrics, normalize the road curvature change rate, the number of obstacles in the scene, and the obstacle motion complexity to generate a scene difficulty score, classify the training samples according to the scene difficulty score, and establish a course sequence from simple to complex; Adopt a two-stage training strategy. In the pre-training stage, conduct behavior cloning training based on human driver demonstration data to enable the network to obtain basic driving capabilities. In the reinforcement learning stage, conduct training based on the course sequence and introduce an experience replay mechanism at the same time; Construct a dynamic course adjustment mechanism, set training efficiency evaluation metrics including the convergence speed of the policy network and the improvement amplitude of the evaluation metrics, adaptively adjust the course difficulty growth step and the number of training rounds for each difficulty level based on the training efficiency evaluation results, and automatically fallback to the training scenario with a difficulty level reduced by 1 for supplementary training when the training efficiency is lower than the preset efficiency threshold; Set multi-dimensional scoring criteria, evaluate the control smoothness and operation efficiency of the regular driving lane under different scene complexities, evaluate the timeliness and safety margin of the emergency avoidance strategy for the emergency avoidance lane, and switch to the training scenario with a difficulty level increased by 1 when the weighted result of the multi-dimensional scoring is greater than the preset scoring threshold.
[0047] Exemplarily, construct a scene complexity evaluation system. The road curvature change rate is calculated by the standard deviation of the curvature within every 100-meter section, and its value range is from 0 to 1, where the curvature change rate of a straight road section is close to 0, and that of a sharp turning section is close to 1. The number of obstacles in the scene is normalized based on the average number of obstacles per kilometer, with a value of 0 when the number of obstacles is 0 and a value of 1 when the number of obstacles reaches 20 per kilometer. The obstacle motion complexity is calculated based on the speed change and trajectory curvature of the obstacles, with the complexity of a stationary obstacle being 0 and that of a fast-moving and frequently changing-direction obstacle being close to 1. The above three metrics are weighted and summed to obtain the scene difficulty score, with weights of 0.3, 0.4, and 0.3 respectively.
[0048] Classify the training samples into five difficulty levels according to the scene difficulty score. The difficulty score of 0 - 0.2 is the primary scene, including basic scenes such as straight roads without obstacles; 0.2 - 0.4 is the intermediate scene, including simple curves and stationary obstacles; 0.4 - 0.6 is the advanced scene, including continuous curves and low-speed moving obstacles; 0.6 - 0.8 is the expert scene, including sharp turns and fast-moving obstacles; 0.8 - 1.0 is the challenge scene, including complex road conditions and multiple interacting obstacles.
[0049] In the pre-training stage, 500 hours of human driver demonstration data are collected for behavior cloning. The demonstration data cover various typical scenarios, including standard driving behaviors such as following a vehicle on a straight road, changing lanes to overtake, and avoiding pedestrians. Through imitation learning, the network masters the basic vehicle control ability and reaches a completion rate of over 90% in the primary scenario test after pre-training.
[0050] In the reinforcement learning stage, the difficulty is gradually increased based on the curriculum sequence. The initial number of training rounds for each difficulty level is set to 1000 rounds, and the capacity of the experience replay pool is 100,000 frames of data. The training efficiency evaluation period is 100 rounds, and the evaluation metrics include the average return improvement rate of the policy network on the validation set and the improvement amplitude of the successful completion rate. When the training efficiency in three consecutive evaluation periods is lower than the preset threshold, the training difficulty is automatically reduced and the number of training rounds for the current difficulty level is increased.
[0051] Set the scoring criteria for three dimensions: safety, smoothness, and efficiency. The safety score is based on the minimum distance from obstacles, requiring it to be greater than 3 meters; the smoothness score is based on the lateral acceleration, requiring it to be less than 2 m / s²; the efficiency score is based on the average driving speed, requiring it to be not less than 80% of the target speed. When the weighted result of the scores in the three dimensions is greater than 0.85, the training difficulty level is increased.
[0052] Through the scene complexity quantitative evaluation and curriculum learning strategy, the present invention enables the deep reinforcement learning network to gradually master driving skills from simple to complex, avoiding the training instability problem when directly facing high-difficulty scenarios, and significantly improving the training efficiency and success rate; adopting a two-stage training and dynamic curriculum adjustment mechanism, making full use of human driving experience and adaptively adjusting the difficulty according to the training effect, overcoming the problems of slow convergence and easy getting stuck in local optima in reinforcement learning, and ensuring the stability of the training process and the reliability of the final performance; the training scheme based on multi-dimensional scoring criteria comprehensively considers the performance of the autonomous driving system in aspects such as safety, comfort, and efficiency, ensuring that the trained driving strategy can not only handle complex working conditions but also meet the actual application requirements, and has strong practical value.
[0053] In an optional implementation manner, the dynamic curriculum adjustment mechanism includes: Construct a multi-dimensional training efficiency evaluation system, take the convergence speed and the improvement amplitude of the evaluation index as short-term efficiency indicators, and construct long-term efficiency indicators based on the generalization performance and anti-interference ability of the policy network in the validation scenario; perform weighted combination on the short-term efficiency indicators and the long-term efficiency indicators to generate a comprehensive training efficiency value; Establish a distributed training architecture, allocate the training tasks of the course sequence to multiple parallel training environments, set different random perturbation parameters in each training environment, perform parallel training based on the asynchronous gradient update method, compare the performance of the policy network under different parallel training environments, and generate a stability evaluation value of the policy network; Determine the adjustment direction of the course difficulty according to the comprehensive training efficiency value and the stability evaluation value. When the comprehensive training efficiency value is greater than the preset efficiency upper limit and the stability evaluation value is greater than the preset stability threshold, increase the course difficulty level according to the adaptive step size; when the comprehensive training efficiency value is less than the preset efficiency threshold, decrease the course difficulty level by 1; Construct a hierarchical storage structure for the experience replay mechanism, store training samples of different difficulty levels into the corresponding experience pool levels; during the supplementary training process of decreasing the difficulty level by 1, perform importance sampling from the experience pool of the corresponding difficulty level to generate a supplementary training sequence; Set a course difficulty smooth transition mechanism. During the process of difficulty level conversion, calculate the mixing weights of adjacent difficulty levels based on the exponential decay function, and dynamically combine the training samples of different difficulty levels according to the mixing weights to generate a transition training sample set.
[0054] Exemplarily, construct a multi-dimensional training efficiency evaluation system. This evaluation system includes short-term efficiency indicators and long-term efficiency indicators. The short-term efficiency indicators mainly consider the convergence speed and the improvement amplitude of the evaluation indicators. Specifically, the convergence speed can be measured by calculating the change rate of the policy network parameters in each training cycle, and at the same time record the improvement percentage of the key evaluation indicators (such as reward value, success rate, etc.). The long-term efficiency indicators are based on the generalization performance and anti-interference ability of the policy network in the validation scenario. Multiple validation scenarios with different environmental parameters can be set, test the average performance score of the policy network in these scenarios, and the degree of performance fluctuation after adding random interference. The short-term efficiency indicators and the long-term efficiency indicators are combined to generate a comprehensive training efficiency value through weighted averaging. For example, the weight of the short-term indicators can be set to 0.4, and the weight of the long-term indicators can be set to 0.6 to obtain the final comprehensive training efficiency value.
[0055] Establish a distributed training architecture. Allocate the training tasks of the course sequence to multiple parallel training environments, and set different random perturbation parameters in each environment. For example, 5 parallel training environments can be set, and 5%, 10%, 15%, 20%, and 25% random noise can be added respectively. Use the asynchronous gradient update method for parallel training, and each environment independently calculates the gradient and asynchronously updates the global policy network parameters. Regularly compare the performance of the policy network under different parallel training environments, and calculate the performance variance as the stability evaluation value. The smaller the stability evaluation value, the more stable the performance of the policy network in different environments.
[0056] Determine the direction of course difficulty adjustment based on the comprehensive training efficiency value and the stability evaluation value. Set a preset efficiency upper limit of 0.8 and a preset stability threshold of 0.1. When the comprehensive training efficiency value is greater than 0.8 and the stability evaluation value is less than 0.1, increase the course difficulty level according to the adaptive step size. The adaptive step size can be dynamically adjusted according to the gap between the current efficiency value and the upper limit, and the larger the gap, the larger the step size. When the comprehensive training efficiency value is less than the preset efficiency lower limit (such as 0.5), decrease the course difficulty level by 1.
[0057] Construct a hierarchical storage structure for the experience replay mechanism. Classify and store the training samples into different experience pool levels according to the difficulty level. For example, 5 difficulty levels can be set, corresponding to 5 experience pool levels. During the supplementary training process of decreasing the difficulty level by 1, perform importance sampling from the experience pool corresponding to the difficulty level to generate a supplementary training sequence. Importance sampling can be weighted based on the time decay coefficient and the reward value of the sample.
[0058] Set a course difficulty smooth transition mechanism. During the process of difficulty level conversion, calculate the mixing weights of adjacent difficulty levels based on the exponential decay function. For example, the decay coefficient can be set to 0.9, then the mixing weight of the nth training batch is 0.9 n and 1 - 0.9 n . Dynamically combine the training samples of different difficulty levels according to the mixing weights to generate a transitional training sample set. This can achieve a smooth transition of the difficulty level and avoid drastic fluctuations in the performance of the policy network.
[0059] Figure 4 Figure for comparing the reward curve of the training process with the difficulty adjustment path. The upper figure shows the performance comparison of three course learning methods during the reinforcement learning training process. From the reward curve, it can be seen that the dynamic course adjustment mechanism of the present invention always maintains a higher learning efficiency during the training process and finally reaches a verification reward of 210, which is significantly better than 95 of the fixed step size course learning and 120 of the course learning based on the performance threshold. In terms of the convergence speed, the present invention only needs 124k steps to converge, while the fixed step size and performance threshold methods require 356k and 298k steps respectively. The difficulty level change curve in the lower figure shows that the present invention can dynamically adjust the course difficulty according to the training efficiency to form a reasonable difficulty increase curve, while the fixed step size method mechanically adjusts at a preset interval, and the performance threshold method stagnates at the medium difficulty stage. The stability index of the present invention reaches 0.87, which is much higher than 0.52 and 0.61 of the other two methods, indicating that it has stronger robustness in different training environments.
[0060] The dynamic curriculum adjustment mechanism of the present invention can effectively improve the efficiency and stability of reinforcement learning training. Through a multi-dimensional training efficiency evaluation system, it comprehensively measures the short-term and long-term training effects of the policy network, providing a reliable basis for curriculum difficulty adjustment. The distributed training architecture and stability evaluation mechanism help enhance the robustness and generalization ability of the policy network. The adaptive curriculum difficulty adjustment strategy can dynamically adjust the difficulty level according to the actual situation during the training process, avoiding the training from falling into local optima or divergence. The hierarchical experience replay and supplementary training mechanism can specifically strengthen the performance of the policy network at different difficulty levels. The curriculum difficulty smooth transition mechanism can achieve a progressive adjustment of the difficulty level, reducing the drastic fluctuations in the performance of the policy network and improving the continuity and stability of training. Overall, this mechanism can significantly improve the training efficiency, generalization ability, and stability of the reinforcement learning algorithm, providing strong support for intelligent decision-making in complex environments.
[0061] In an alternative embodiment, constructing a virtual simulation environment includes: Determining road geometric features and slope features based on road parameters; setting type parameters and layout position parameters of road facility features; setting environmental parameters of weather features; Establishing a feature combination constraint rule library, including: matching rules for road geometric features and slope features, layout rules for road facility features, and combination rules for weather features; Screening out feature combinations that meet the constraint conditions based on the feature combination constraint rule library, and constructing a feature fitness matrix, where the fitness value of the feature fitness matrix represents the matching degree of the corresponding feature combination; Automatically combining features using a hierarchical combination method, including: selecting the combination of road geometric features and slope features with the highest fitness value to generate a basic road section, selecting the road facility feature with the highest fitness value to be laid out on the basic road section, and selecting the weather feature with the highest fitness value for superposition to generate a complete scene; Randomly sampling the feature combination results using the Monte Carlo method and applying a Markov chain for feature transfer to generate a composite test scene.
[0062] The present invention provides a method for constructing a virtual simulation environment, which creates a complex and realistic road scene through systematic steps. First, determine the geometric features and slope features of the road based on road parameters. Road geometric features include road width, number of lanes, curve radius, etc.; slope features include longitudinal slope, cross slope, etc. For example, a four-lane highway in both directions can be set, with a width of 28 meters, a curve radius of 1000 meters, a longitudinal slope of 2%, and a cross slope of 1.5%.
[0063] Set the type parameters and layout position parameters of road facility features. Road facilities include traffic signs, guardrails, street lights, etc. For example, a speed limit sign of 80 km / h can be set 500 meters from the road starting point, and a street light can be set every 50 meters on both sides of the road. At the same time, set the environmental parameters of weather features, such as temperature, humidity, visibility, etc. For example, the temperature can be set to 25 °C, the humidity to 60%, and the visibility to 1000 meters.
[0064] To ensure that the generated scenarios are reasonable and credible, establish a feature combination constraint rule library. This rule library includes the matching rules between road geometric features and slope features, the layout rules of road facility features, and the combination rules of weather features. For example, the maximum longitudinal slope of a highway should not exceed 4%, and the cross slope at a curve should match the curve radius; the setting of traffic signs should comply with traffic rules, such as a curve warning sign should be set before a curve; the visibility in foggy weather should be less than 200 meters.
[0065] Based on the feature combination constraint rule library, select the feature combinations that meet the constraint conditions and construct a feature fitness matrix. The fitness value represents the matching degree of the corresponding feature combination, ranging from 0 to 1. For example, for a highway scenario, the fitness of a 28-meter width and a two-way four-lane may be 0.9, while the fitness with a two-way six-lane may be 0.6.
[0066] Use a hierarchical combination method to automatically combine features. First, select the combination of road geometric features and slope features with the highest fitness value to generate a basic road section. For example, select the combination of a two-way four-lane, 28-meter width, 1000-meter curve radius, 2% longitudinal slope, and 1.5% cross slope. Then, select the road facility features with the highest fitness value on the basic road section for layout, such as setting a speed limit sign 500 meters from the starting point and a street light every 50 meters on both sides. Finally, select the weather features with the highest fitness value for superposition, such as a temperature of 25 °C, a humidity of 60%, and a visibility of 1000 meters, to generate a complete scenario.
[0067] To increase the diversity and authenticity of the scenarios, use the Monte Carlo method to randomly sample the feature combination results. For example, randomly sample the road width between 26 meters and 30 meters and the longitudinal slope between 1% and 3%. Then apply a Markov chain for feature transfer to generate composite test scenarios. The Markov chain can simulate the transition probability between features, such as the transition probability from sunny to cloudy is 0.3, and the transition probability from cloudy to rainy is 0.4. In this way, a series of continuously changing scenarios can be generated, such as the process from sunny to cloudy and then to rainy.
[0068] The present invention standardizes the management of road geometric features, slope features, road facility features, and weather features by establishing a feature combination constraint rule base. Based on this rule base, the rationality test of feature combinations is realized to ensure that the generated scenarios meet the engineering reality and traffic specifications. For example, the rule base limits that the maximum longitudinal slope of expressways does not exceed 4%, and the cross slope at the curve matches the curve radius and other constraint conditions. The construction of a feature fitness matrix realizes the quantitative evaluation of feature combinations, and guides the optimal combination of features through fitness values. The hierarchical combination method is adopted to preferentially select feature combinations with high fitness to construct the basic road sections, and then superimpose road facility and weather features, ensuring the hierarchy and rationality of scenario construction.
[0069] In an alternative embodiment, the Monte Carlo method is used to randomly sample the feature combination results, and the Markov chain is applied for feature transfer to generate composite test scenarios, including: Calculate the sum of the product of the scenario coverage rate and the first preset weight and the product of the rare event probability and the second preset weight to obtain the scenario importance; calculate the sampling temperature parameter according to the scenario importance, and substitute the scenario importance and the sampling temperature parameter into the Softmax function to generate the scenario sampling probability; Perform Monte Carlo sampling on the feature combination results according to the scenario sampling probability to obtain scenario features; input the scenario features and scenario condition constraints into a conditional variational autoencoder to generate the mean parameter and the variance parameter; Perform Gaussian distribution sampling based on the mean parameter and the variance parameter to obtain latent variables; construct a Markov transition matrix for the latent variables, and perform feature migration on the latent variables based on the Markov transition matrix to generate the migrated latent variables; Decode the migrated latent variables and the scenario condition constraints to generate the migrated scenario features; calculate the KL divergence between the migrated scenario features and the original scenario features, and when the KL divergence is less than the preset divergence threshold, determine that the migrated scenario features are valid composite test scenarios.
[0070] Exemplarily, calculate the sum of the product of the scenario coverage rate and the first preset weight and the product of the rare event probability and the second preset weight to obtain the scenario importance. A rare event refers to a special scenario type with a frequency lower than the preset threshold in the historical test scenario data, such as abnormal working conditions like extreme weather and sudden obstacles. Specifically, the first preset weight can be set to 0.6 and the second preset weight can be set to 0.4. Assume that the coverage rate of a certain scenario is 0.8 and the rare event probability is 0.7, then the importance calculation of this scenario is: 0.8×0.6 + 0.7×0.4 = 0.76.
[0071] Calculate the sampling temperature parameter according to the scene importance. The sampling temperature parameter can be determined by a mapping function. For example, the scene importance of 0 - 0.3 can be mapped to the temperature parameter of 1.5 - 2.0, 0.3 - 0.6 can be mapped to 1.0 - 1.5, and 0.6 - 1.0 can be mapped to 0.5 - 1.0. In this example, the scene importance is 0.76, and the corresponding sampling temperature parameter can be 0.8.
[0072] Substitute the scene importance and the sampling temperature parameter into the Softmax function to generate the scene sampling probability. The Softmax function can convert the input into a probability distribution. Assume that after being calculated by the Softmax function, the obtained scene sampling probability is 0.65.
[0073] Perform Monte Carlo sampling on the feature combination result according to the scene sampling probability to obtain the scene features. The Monte Carlo method approximates the true distribution through multiple random samplings. Assume that the sampled scene features are [0.9, 0.3, 0.7, 0.5], representing the feature values of four different dimensions.
[0074] Input the scene features and the scene condition constraints into the conditional variational autoencoder to generate the mean parameter and the variance parameter. The scene condition constraints refer to the restrictive conditions imposed on scene generation, including the range of road geometric parameters, traffic rules, physical laws, and other constraints that must be satisfied. The conditional variational autoencoder is a generative model that can learn the latent representation of data. Assume that the generated mean parameter is [0.8, 0.2, 0.6, 0.4], and the variance parameter is [0.1, 0.05, 0.08, 0.06].
[0075] Perform Gaussian distribution sampling based on the mean parameter and the variance parameter to obtain the latent variable. Gaussian distribution sampling can increase the diversity of the generated results. Assume that the sampled latent variable is [0.85, 0.18, 0.65, 0.45].
[0076] Construct a Markov transition matrix for the latent variable, and perform feature migration on the latent variable based on the Markov transition matrix to generate the migrated latent variable. The Markov transition matrix describes the transition probability between states. Assume that after the feature migration, the obtained new latent variable is [0.82, 0.25, 0.68, 0.42].
[0077] Decode the migrated latent variable and the scene condition constraints to generate the migrated scene features. The decoding process is the inverse process of encoding, which maps the latent variable back to the original feature space. Assume that the newly obtained scene features after decoding are [0.88, 0.28, 0.72, 0.48].
[0078] The present invention determines the importance of scenarios through weighted calculation of scenario coverage rate and rare event probability, and adaptively adjusts the sampling temperature parameter based on the importance, effectively improving the sampling efficiency for rare scenarios. By introducing the combination of conditional variational autoencoder and Gaussian distribution sampling, efficient encoding and diverse generation of scenario features are achieved. The Markov transition matrix is used for feature migration, and the generation quality is controlled by the KL divergence threshold, ensuring the continuity and effectiveness of the generated scenarios. This method enables the scenarios after feature migration to maintain the relevance with the original scenarios while generating new effective changes.
[0079] In the second aspect of the embodiments of the present invention, an intelligent driving simulation test system for diesel locomotives is provided. The system includes: A first unit for constructing a virtual simulation environment; collecting the driving state information of the diesel locomotive; obtaining scenario complexity data, where the scenario complexity data includes the road curvature change rate, the number of obstacles in the scenario, and the obstacle movement complexity; A second unit for, in the virtual simulation environment, obtaining the three-dimensional point cloud data of obstacles through a virtual lidar; obtaining the distance information and relative speed information of obstacles through a virtual millimeter-wave radar and simulating the measurement errors caused by multipath effects; obtaining the image feature information of obstacles through a virtual vision sensor and simulating the imaging quality fluctuations caused by light changes; performing multimodal fusion processing on the three-dimensional point cloud data, distance information, relative speed information, and image feature information to generate standardized perception data; A third unit for constructing a deep reinforcement learning network, including a prediction branch network for generating multimodal prediction trajectories and a dual-channel control network for generating driving strategies; wherein, for conventional driving strategies, the safety index is set as the network constraint condition, and the weights of the comfort and smoothness indexes increase with the increase of scenario complexity, while the weight of the efficiency index decreases accordingly; for emergency avoidance driving strategies, safety constraint optimization is performed based on collision risk assessment; inputting the driving state information, standardized perception data, and scenario complexity data into the deep reinforcement learning network to generate driving control parameters.
[0080] In the third aspect of the embodiments of the present invention, An electronic device is provided, including: A processor; A memory for storing instructions executable by the processor; wherein the processor is configured to call the instructions stored in the memory to execute the method described above.
[0081] In the fourth aspect of the embodiments of the present invention, A computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0082] The present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present invention.
[0083] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A diesel locomotive intelligent driving simulation test method, characterized in that: include: Build a virtual simulation environment; Collecting driving status information of diesel locomotives; Obtaining scene complexity data, wherein the scene complexity data includes a rate of change of road curvature, a number of obstacles in the scene, and a complexity of obstacle motion; In the virtual simulation environment, three-dimensional point cloud data of obstacles are obtained through virtual laser radar; distance information and relative speed information of obstacles are obtained through virtual millimeter wave radar, and measurement errors caused by multipath effects are simulated; image feature information of obstacles is obtained through virtual visual sensors, and imaging quality fluctuations caused by changes in illumination are simulated; the three-dimensional point cloud data, distance information, relative speed information and image feature information are subjected to multimodal fusion processing to generate standardized perception data; A deep reinforcement learning network is constructed, including a prediction branch network for generating multimodal prediction trajectories and a dual-channel control network for generating driving strategies; wherein, for conventional driving strategies, safety indicators are set as network constraints, the weights of comfort and smoothness indicators increase with the increase of scene complexity, and the weight of efficiency indicators decreases accordingly; for risk-avoidance driving strategies, safety constraints are optimized based on collision risk assessment; the driving state information, standardized perception data and scene complexity data are input into the deep reinforcement learning network to generate driving control parameters.
2. The method according to claim 1, characterized in that The generation of the standardized perception data includes: Acquire three-dimensional point cloud data of obstacles through a virtual laser radar, and filter the three-dimensional point cloud data according to the local point cloud center and standard deviation to obtain filtered point cloud data; The virtual millimeter wave radar transmits a frequency modulated continuous wave signal to obtain the distance information and relative speed information of the obstacle, and performs multipath effect error compensation on the distance information and relative speed information according to the time delay, attenuation coefficient and Doppler frequency shift of multiple propagation paths; Obtain image feature information of obstacles through a virtual vision sensor, and perform illumination change compensation on the image feature information according to local illumination intensity to obtain image feature data; Performing spatiotemporal alignment on the filtered point cloud data, the compensated distance information and relative speed information, and the image feature data, and uniformly converting the data into a global coordinate system according to the timestamp; Perform multimodal fusion on the spatiotemporally aligned data, perform weighted feature combination according to the weights of each modality feature to obtain fused features, and perform information fusion on the fused features according to basic probability distribution and conflict factors; The fused data is standardized and quality evaluated, and the information entropy is calculated based on the feature probability distribution to generate standardized perception data.
3. The method according to claim 1, characterized in that The deep reinforcement learning network includes: The prediction branch network receives the position, velocity and acceleration information of obstacles in the standardized perception data, extracts trajectory features through a multi-head self-attention encoder, and generates a multimodal prediction trajectory based on a recurrent neural network decoder; The multimodal predicted trajectory, driving state information and scene complexity data are constructed as a state space vector input to a dual-channel control network, wherein the dual-channel control network includes a conventional driving channel and an emergency avoidance channel; The conventional driving channel generates a conventional driving strategy based on the Actor-Critic network structure, sets the safety index as the constraint condition of the Actor-Critic network, and the weight of the comfort index and the smoothness index increases with the increase of the scene complexity, and the weight of the efficiency index decreases accordingly, and the conventional driving strategy is optimized through the temporal difference error; the emergency avoidance channel generates an avoidance driving strategy based on the collision risk assessment, and optimizes the avoidance driving strategy through safety constraints; The confidence of the conventional driving strategy and the risk-avoidance driving strategy is evaluated based on the spatiotemporal collision risk and the minimum safe distance, and the confidence is mapped into a dynamic fusion weight; the conventional driving strategy and the risk-avoidance driving strategy are weightedly fused according to the dynamic fusion weight to generate driving control parameters.
4. The method according to claim 3, characterized in that Evaluating the confidence of the conventional driving strategy and the risk-avoidance driving strategy based on the spatiotemporal collision risk and the minimum safety distance, and mapping the confidence into a dynamic fusion weight includes: The time dimension collision probability is calculated based on the multimodal predicted trajectory, the deviation between the time collision point of the multimodal predicted trajectory and the expected collision time is substituted into the Gaussian distribution function, and the time dimension collision risk value is generated according to the trajectory weight; the space dimension collision probability is calculated based on the position probability density function and the speed constraint function, the position probability density function and the speed constraint function are integrated in the collision risk area to generate the space dimension collision risk value; Calculating a dynamic safety distance according to the speed of the own vehicle, the speed of the preceding vehicle, the reaction time, the maximum deceleration and the minimum deceleration, and determining the ratio of the actual vehicle distance to the dynamic safety distance as a safety margin coefficient; The inverse values of the time dimension collision risk value and the space dimension collision risk value are weighted and combined with the safety margin coefficient to generate the conventional driving strategy confidence; the weighted sum of the exponential decay value of the safety margin coefficient and the time and space dimension collision risk value is used as the risk avoidance driving strategy confidence; Substitute the conventional driving strategy confidence and the risk-avoidance driving strategy confidence into the Softmax mapping function for normalization processing to generate initial fusion weights; perform exponential sliding average processing on the initial fusion weights to generate dynamic fusion weights.
5. The method according to claim 1, characterized in that The training of the deep reinforcement learning network includes: The training course is constructed based on the scene complexity index. The road curvature change rate, the number of obstacles in the scene, and the complexity of obstacle movement are normalized to generate a scene difficulty score. The training samples are graded according to the scene difficulty score to establish a course sequence from simple to complex. A two-stage training strategy is adopted. In the pre-training stage, behavior cloning training is performed based on human driver demonstration data to enable the network to acquire basic driving capabilities. In the reinforcement learning stage, training is performed based on course sequences while introducing an experience replay mechanism. Build a dynamic course adjustment mechanism, set training efficiency evaluation indicators including the convergence speed of the policy network and the improvement of the evaluation indicators, and adaptively adjust the course difficulty increase step and the training rounds for each difficulty level based on the training efficiency evaluation results. When the training efficiency is lower than the preset efficiency threshold, it will automatically fall back to the training scenario with difficulty level minus 1 for supplementary training. A multi-dimensional scoring standard is set to evaluate the control smoothness and operating efficiency of conventional driving channels under different scenario complexities, and to evaluate the timeliness and safety margin of the risk avoidance strategy of emergency risk avoidance channels. When the weighted result of the multi-dimensional scoring is greater than the preset scoring threshold, switch to the training scenario with the difficulty level increased by 1.
6. The method according to claim 5, characterized in that The dynamic course adjustment mechanism includes: Construct a multi-dimensional training efficiency evaluation system, use the convergence speed and the improvement of evaluation indicators as short-term efficiency indicators, and construct long-term efficiency indicators based on the generalization performance and anti-interference ability of the policy network in the verification scenario; perform a weighted combination of the short-term efficiency indicators and the long-term efficiency indicators to generate a comprehensive training efficiency value; Establish a distributed training architecture, distribute the training tasks of the course sequence to multiple parallel training environments, set different random perturbation parameters in each training environment, perform parallel training based on asynchronous gradient update, compare the performance of the policy network in different parallel training environments, and generate a stability evaluation value of the policy network; Determine the course difficulty adjustment direction according to the comprehensive training efficiency value and the stability evaluation value, when the comprehensive training efficiency value is greater than the preset efficiency upper limit and the stability evaluation value is greater than the preset stability threshold, increase the course difficulty level according to the adaptive step length; when the comprehensive training efficiency value is less than the preset efficiency threshold, reduce the course difficulty level by 1; A hierarchical storage structure is constructed for the experience replay mechanism, and training samples of different difficulty levels are stored in corresponding experience pool levels; during the supplementary training process of the difficulty level minus 1, importance sampling is performed from the experience pool of the corresponding difficulty level to generate a supplementary training sequence; A smooth transition mechanism for course difficulty is set up. During the difficulty level conversion process, the mixing weights of adjacent difficulty levels are calculated based on the exponential decay function. The training samples of different difficulty levels are dynamically combined according to the mixing weights to generate a transition training sample set.
7. The method according to claim 1, characterized in that Building a virtual simulation environment includes: Determine road geometry and slope characteristics based on road parameters; set type parameters and layout location parameters of road facility characteristics; set environmental parameters of weather characteristics; Establish a feature combination constraint rule library, including: matching rules between road geometry and slope features, layout rules for road facility features, and combination rules for weather features; Based on the feature combination constraint rule library, feature combinations that meet the constraint conditions are screened out, and a feature fitness matrix is constructed, wherein the fitness value of the feature fitness matrix represents the matching degree of the corresponding feature combination; The features are automatically combined using a hierarchical combination method, including: selecting the road geometry feature and slope feature with the highest fitness value to generate a basic road section, selecting the road facility feature with the highest fitness value to lay out on the basic road section, and selecting the weather feature with the highest fitness value to overlay to generate a complete scene; The Monte Carlo method is used to randomly sample the feature combination results, and the Markov chain is applied to perform feature transfer to generate composite test scenarios.
8. The method according to claim 7, characterized in that The Monte Carlo method is used to randomly sample the feature combination results, and the Markov chain is used for feature transfer to generate composite test scenarios including: Calculate the sum of the product of the scene coverage and the first preset weight and the product of the rare event probability and the second preset weight to obtain the scene importance; calculate the sampling temperature parameter according to the scene importance, substitute the scene importance and the sampling temperature parameter into the Softmax function to generate the scene sampling probability; Monte Carlo sampling is performed on the feature combination result according to the scene sampling probability to obtain scene features; the scene features and scene condition constraints are input into a conditional variational autoencoder to generate mean parameters and variance parameters; Gaussian distribution sampling is performed based on the mean parameter and the variance parameter to obtain a latent variable; a Markov transfer matrix is constructed for the latent variable, and feature migration is performed on the latent variable based on the Markov transfer matrix to generate a migrated latent variable; The migrated latent variables and the scene condition constraints are decoded to generate migrated scene features; the KL divergence between the migrated scene features and the original scene features is calculated, and when the KL divergence is less than a preset divergence threshold, the migrated scene features are determined to be a valid composite test scene.
9. A diesel locomotive intelligent driving simulation test system, used to implement the method according to any one of claims 1 to 8, characterized in that: include: The first unit is used to build a virtual simulation environment and collect the driving status information of the diesel locomotive; Obtaining scene complexity data, wherein the scene complexity data includes a rate of change of road curvature, a number of obstacles in the scene, and a complexity of obstacle motion; The second unit is used to obtain three-dimensional point cloud data of obstacles through virtual laser radar in the virtual simulation environment; obtain distance information and relative speed information of obstacles through virtual millimeter wave radar, and simulate measurement errors caused by multipath effects; obtain image feature information of obstacles through virtual visual sensors, and simulate imaging quality fluctuations caused by illumination changes; perform multimodal fusion processing on the three-dimensional point cloud data, distance information, relative speed information and image feature information to generate standardized perception data; The third unit is used to construct a deep reinforcement learning network, including a prediction branch network for generating multimodal prediction trajectories and a dual-channel control network for generating driving strategies; wherein, for conventional driving strategies, the safety index is set as a network constraint condition, the weights of comfort and smoothness indicators increase with the increase of scene complexity, and the weight of efficiency indicators decreases accordingly; for risk-avoidance driving strategies, safety constraint optimization is performed based on collision risk assessment; the driving state information, standardized perception data and scene complexity data are input into the deep reinforcement learning network to generate driving control parameters.
Citation Information
Patent Citations
Anti-collision unmanned driving simulation system
CN113110106A
Control system and control method for intelligent connected vehicle
CN113272195A
Vehicle risk early warning method and device based on neural network and automatic driving vehicle
CN114333416A
Rapid trunk road vehicle trajectory optimization control method in vehicle-road cooperative mixed traffic flow environment
CN116502529A
Driving assistance device and driving assistance system
CN116895175A
Cited By
Immersive rail transit simulation training system and method based on virtual reality
CN120412369A
Immersive rail transit simulation training system and training method based on virtual reality
CN120412369B
VR-HIL multi-sensor closed-loop test platform for remote driving
CN120428597A
Bidding certificate image text fuzzy evaluation method and device, medium and electronic equipment
CN120599624A
Method and device for determining end-to-end perception decision regulation and control architecture
CN120726596A