A method for generating a rain and fog test scene for an intelligent vehicle based on reinforcement learning

By introducing a risk field-based reward function and a diffusion sine/cosine Q-learning algorithm, the problem of sparse rewards in the generation of intelligent vehicle test scenarios is solved, achieving efficient test scenario generation in rain and fog environments and improving learning efficiency and model convergence speed.

CN120874982BActive Publication Date: 2025-11-28JILIN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511403595.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2025-11-28
Estimated Expiration
2045-09-29

AI Technical Summary

Technical Problem

Existing reinforcement learning-based methods for generating test scenarios for intelligent vehicles suffer from sparse rewards in complex weather conditions such as rain and fog, resulting in slow learning speeds or even failure to learn. Furthermore, existing methods lack consideration of vehicle interaction processes when generating rain and fog scenarios, leading to low testing efficiency.

Method used

A risk field-based reward function setting method is adopted, combined with the diffusion sine and cosine Q-learning algorithm. By introducing continuous risk potential field rewards and a decaying greedy strategy, a diffusion sine and cosine Q-learning process is designed. The sine and cosine algorithms are used for initialization and reverse curriculum generation to alleviate the sparse reward problem and improve learning efficiency.

Benefits of technology

In rainy and foggy environments, the iterative convergence speed of intelligent vehicle test scenario generation is accelerated, the learning efficiency of test scenario generation and the stability of the generation model are improved, and the safety testing of intelligent vehicles under complex weather conditions is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120874982B_ABST
    Figure CN120874982B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of intelligent vehicle test scene generation methods, in particular to a kind of intelligent vehicle rain and fog test scene generation method based on reinforcement learning, steps include: test scene modeling, introduce value function, set reward function, diffusion positive sine Q learning.The present application introduces continuous risk potential reward remodeling method to describe the continuous reward process of intelligent vehicle in rain and fog environment;Design attenuation greed strategy to improve the randomness of scene generation direction;And by using positive sine algorithm to initialize Q-learning, while increasing the expansion positive sine algorithm process as reverse course generation, diffusion positive sine Q learning algorithm does not increase the total number of SCA iteration under the condition, by adjusting each SCA module parameter makes its attention range gradually expand and update Q value as the reverse course of reinforcement learning, further accelerate the iteration convergence speed of the whole test scene generation algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to an intelligent automobile test scene generation method, in particular to an intelligent automobile rain and fog test scene generation method based on reinforcement learning. BACKGROUND

[0002] With the development of intelligent automobiles, a scientific and perfect test technology system is the basis and premise for the landing of the intelligent automobile industry. Compared with the traditional test method based on mileage, the test method based on scene has better repeatability and higher test efficiency, and has become the core technology of the existing intelligent automobile test. Test scene generation is a technology for generating a corresponding test scene parameter according to the output of the measured object by using a generation model. The existing research methods include combinatorial testing, optimization theory, adversarial generation network and reinforcement learning. Among them, the test scene generation method based on combinatorial testing and optimization theory is affected by the dimension collapse problem caused by the high dimension of the test scene parameter. The method based on the adversarial generation network is easy to fall into the mode collapse, resulting in the single category of the generated test scene, and the lack of consideration of the interaction process between the other traffic vehicle behavior and the vehicle directly leading to the traffic accident of the measured vehicle. The research goal of the existing scene generation method based on reinforcement learning is mostly concentrated in the good traffic environment, and there is few test scene generation method for complex weather environments such as rain and fog. In fact, the rain and fog weather environment has stronger weather randomness and complexity, and will produce more challenging test scenes for intelligent automobiles, for example, the following test scene in good weather may become a dangerous scene under a certain rainfall intensity or fog visibility. Testing intelligent automobiles by using rain and fog scenes can improve the test efficiency and reduce the test cost. Q-learning is a classical value-based reinforcement learning method, which has been widely used in intelligent automobile test scene generation. However, the existing scene generation method based on reinforcement learning still has some problems, and the sparsity of rewards is one of the core problems to be solved. Reinforcement learning stimulates the agent to learn through rewards, and the sparsity of rewards means that the distribution of positive rewards in the state space is sparse, that is, the existing scene generation method mostly only rewards 1 at the moment of collision and rewards 0 at other moments, and the agent often needs to go through a complex exploration process to obtain rewards, resulting in slow learning speed of the scene generation model or even unable to learn. SUMMARY

[0003] In order to solve the above-mentioned problems of the test scene generation method, the application provides an intelligent automobile rain and fog test scene generation method based on reinforcement learning, which comprises the following steps:

[0004] Step 1, test scene modeling: the test scene is modeled as a five-tuple defined random process, wherein: is the set of possible motion states of the vehicle in the test scene. a set of actions of the vehicle in the test scenario; a state transition probability function describing the probability of transitioning from one state to another after performing an action; a reward function for evaluating the goodness of each state transition; a reward discount factor for balancing the weight of current reward and future reward.

[0005] Further, the motion state of the vehicle is the spatial position and speed of the measured vehicle and surrounding traffic vehicles under a specific rainfall intensity or fog visibility. Motion state sequence wherein represents the state of the ego vehicle (EV) at the time instant; represents the state of the th background vehicle (BV) at the time instant.

[0006] The action of the vehicle is the lateral and longitudinal acceleration of the measured vehicle and surrounding traffic vehicles under a specific rainfall intensity or fog visibility; the action of the ego vehicle (EV) is calculated by constructing an ego vehicle (EV) agent model. represents the action of the ego vehicle (EV) at the time instant.

[0007] The action of the background vehicle (BV) is obtained by a diffusion cosine Q-learning method. represents the action of the th background vehicle (BV) at the time instant, including lateral and longitudinal acceleration.

[0008] Further, the specific form of the state transition probability function is :

[0009]

[0010] The specific form of the reward function is :

[0011]

[0012] wherein, and both represent the state quantity in the set , is the current state, is the next state; represents a certain action in the action set ; represents the state transition probability function; represents the reward function. The environment then transitions to the next time step The immediate reward given, which measures the current state And the action That comes with it; And The state and action random variables at time And The state variable at the next time step ; Presents the probability that the random variable Takes the value , And The expected value of

[0013] Step 2, Introducing the Value Function: Q-learning is a typical value-based reinforcement learning algorithm. Further, in reinforcement learning, for a given reinforcement learning policy , the state value function And the action value function Are introduced, the state value function is used to measure the long-term return of executing the policy In state , and the action value function is used to measure the expected return of taking action After state Follow the policy , its formal definition is as follows:

[0014]

[0015] At time , let Be the discounted cumulative return at this time; Indicates the expected operation under the policy Indicates the immediate reward and its discounted cumulative item obtained at each future time.

[0016] The value is updated by one step through the Temporal-Difference (TD) algorithm to obtain the optimal policy, and the updated state value function The formula is as follows:

[0017]

[0018] In the formula, Is the learning rate; Is the temporal difference target; Temporal difference bias.

[0019] ​Step 3, setting reward function: in the test scene generation, the application proposes a reward function setting method based on risk field to obtain the reward function, so as to provide continuous and instructive incentive in a wider state space. The application proposes a reward function setting method based on risk field, which constructs a reward function composed of environmental risk field The reward function is introduced into the reinforcement learning process as an immediate reward part, and the specific expression is as follows:

[0020]

[0021] In the formula, represents the risk field deterioration degree caused by rain and fog weather, which is modeled by rainfall intensity or fog visibility; is the motion risk field formed by vehicle motion; is the road risk field formed by static road, and the specific expressions of the two are as follows:

[0022]

[0023] In the formula, represents the real distance between EV and BV, and the unit is m; is the speed of vehicle motion; represents the road condition coefficient; represents the virtual mass of the target object; represents the importance of road line type, and the value of single solid line is greater than that of dotted line; represents the lateral distance between the axis of EV and the lane line; is a constant value, which determines the speed at which the safety risk increases when the EV approaches the lane line.

[0024] Step 4, spread sine and cosine Q learning: the Q-learning reinforcement learning algorithm describes the expected return that the agent can obtain by performing each action in each state by introducing Q value, and constantly updates the Q value table according to the received reward or punishment in the process of interaction between the agent and the environment; After the training of test scene generation is completed, the agent can select the optimal action that can maximize the long-term cumulative return according to the Q value in any state.

[0025] Further, the update of Q value adopts time difference (TD) method, and the update formula of Q value is as follows:

[0026]

[0027] In the formula, represents the value estimation corresponding to the action performed in the current state . point to the next state The maximum Q value that can be achieved among all possible actions in the next state.

[0028] Q-learning agent adopts a greedy policy to determine actions, to strike a balance between exploring new actions and exploiting known optimal actions:

[0029]

[0030] where, represents the probability of taking action in state ; represents the total number of possible actions in state , is the greed coefficient. The greedy policy introduces an exponential decay, i.e., gradually reduces the value of as the training progresses, so that the algorithm has sufficient exploration ability at the early stage, and relies more on the learned good strategy at the later stage, so as to ensure sufficient exploration and promote stable convergence;

[0031] Further, the expression of the greed coefficient is:

[0032]

[0033] where, and represent the set start and end greed coefficients, respectively; represents the current round number; and represent the decay coefficient and decay step, respectively, for adjusting the speed of decay.

[0034] The present application further proposes a "diffusion sine cosine Q-learning" method, which introduces a diffusion sine cosine process in the standard Q-learning framework as the core of the inverse curriculum generation. The inverse curriculum generation is a strategy for alleviating the sparse reward problem by sorting the difficulty of the learning task without adjusting the original reward function. The diffusion sine cosine process is composed of a plurality of sine cosine algorithm (SCA) modules with different parameters in sequence, and each module focuses on a wider search area than the previous one, starting from the target neighborhood and constantly diffusing outward.

[0035] ​The Sine Cosine Algorithm (SCA) is a population search-based optimization method. It first generates a set of initial solutions randomly, and then uses the oscillating properties of sine and cosine functions to make each solution move back and forth around the current global optimum, thereby achieving a balance between global exploration and local development.

[0036] Furthermore, the iterative formula for SCA individuals is:

[0037]

[0038] In the formula, Indicates the current iteration round number. for In the nth iteration Individuals on the dimension Coordinates; in the first In the next iteration For the first Individuals on the dimension coordinate, The global optimal solution obtained in the current iteration is at the th iteration. Coordinates in dimensional space; and Four random numbers are used to control the algorithm's exploration-expansion tradeoff, update stride, the global optimum's pull on the individual, and the choice of sine / cosine operations, respectively. Typically, Uniform sampling from [0, 2π] Uniform sampling from [0,2] Uniform sampling from [0,1]; The value of is calculated according to the following formula:

[0039]

[0040] In the formula, It is a constant. Indicates the current iteration round number. This is the total number of iterations;

[0041] Diffusion-cosine Q-learning consists of three stages:

[0042] Initial SCA phase: The target location is searched using the standard SCA method, and the Q-value table is initially initialized at the same time;

[0043] The expanded SCA phase consists of multiple SCA modules with different parameters, which gradually expand the search radius in sequence, guiding the agent to learn from easy to difficult in the form of a reverse course.

[0044] Q-learning phase: After completing the reverse course, switch to the Q-learning process and use reinforcement learning strategies to finally generate test scenarios.

[0045] In addition to the expansion phase, the parameters of each SCA module are dynamically adjusted to expand the exploration range of the population.

[0046] The updating strategy of the step length coefficient in SCA is optimized, and the specific updating formula is as follows:

[0047]

[0048] In the formula, is the updating parameter of the i-th SCA module, which is used to change the updating direction and step length of the individual, indicates the total number of modules; is the constant of the i-th SCA module, is designed to gradually decrease so that the step length of the individual moving towards the target decreases with the increase of the iteration number, and then the individual is more inclined to gradually explore the area near the target; is a scaling coefficient, which is used to adjust the final value of the updating. The updating formula (14) is used to update , and is randomly updated, and the iterative formula (12) of the SCA individual is used to iterate the population individual, the action selection, action taking, state observation and reward are performed on each individual, and the Q value is updated as the reverse course generation.

[0049] The above process is repeated until the iteration number of the current SCA module reaches the set maximum iteration number, and all SCA modules are executed to complete the spread of the positive and negative sine process. The beneficial effects of the present application are as follows: The present application proposes a new intelligent vehicle rain and fog test scene generation method based on reinforcement learning to solve the problems of the existing test scene generation method.

[0050] The continuous risk potential reward remodeling method is introduced to describe the continuous reward process of the intelligent vehicle in the rain and fog environment. The decay

[0051] greedy strategy is designed to improve the randomness of the scene generation direction. The positive and negative sine algorithm is used to initialize Q-learning, and the spread positive and negative sine algorithm process is added as the reverse course generation. The spread positive and negative sine Q-learning algorithm gradually expands the attention range of each SCA module by adjusting the parameters of the SCA modules without increasing the total number of SCA iterations, and updates the Q value as the reverse course of reinforcement learning, which further accelerates the iteration convergence speed of the whole test scene generation algorithm.Attached Figure Description

[0052] Figure 1 This is a schematic diagram of the overall architecture of the present invention. Detailed Implementation

[0053] This embodiment provides a method for generating rain and fog test scenarios for intelligent vehicles based on reinforcement learning, which includes the following steps:

[0054] The method proposed in this invention is as follows: Figure 1 As shown.

[0055] The test scenario of this invention is modeled as a quintuple. A defined random process, wherein: This is the set of possible vehicle motion states in the test scenario; This is the set of actions that can be selected for vehicles in the test scenario. Let be the state transition probability function, which describes the probability of transitioning from one state to another after performing an action; This is the reward function, used to evaluate the quality of each state transition; This is a reward discount factor used to balance the weight of current and future rewards.

[0056] The motion state of the vehicle refers to the spatial position and speed of the tested vehicle and surrounding traffic vehicles under specific rainfall intensity or fog visibility conditions, and the motion state sequence. ,in Representing the The state of the ego vehicle (EV) at any given moment; Representing the At the [time]th moment The status of surrounding traffic vehicles (BV); Indicates the first A set of actions selected by an agent at any given moment.

[0057] The test scene generation method of this invention considers the state sequence of the tested vehicle and surrounding traffic vehicles under specific rainfall intensity or fog visibility conditions. It satisfies the requirements of a Markov Decision Process (MDP): that is, given the current and all historical states, the conditional probability of the next state depends only on the current state.

[0058] The vehicle's motion refers to the lateral and longitudinal accelerations of the tested vehicle and surrounding traffic vehicles under specific rainfall intensities or fog visibility conditions; the motion of the main vehicle EV is calculated by constructing a main vehicle EV proxy model. Representing the The actions of the main EV at that moment;

[0059] action of the surrounding traffic vehicle BV obtained by the diffusion cosine Q learning method, represent the first action of the surrounding traffic vehicle BV at the first moment, including lateral and longitudinal acceleration;

[0060] The specific form of the state transition probability function is:

[0061]

[0062] The specific form of the reward function is:

[0063]

[0064] wherein, and both represent state quantities in the set , is the current state, is the next state; represents a certain action in the action set ; represents the reward signal given by the environment at the next moment after performing the action , used to measure the benefits brought by the current state and the action ; and are the state and action random variables at moment , respectively, and corresponds to the state variable at the next moment ; represents the probability that the random variable takes the value , and is the expected value of ;

[0065] In reinforcement learning, for a given reinforcement learning strategy , the state value function and the action value function are introduced, the state value function is used to measure the long-term benefits of executing the strategy in the state , and the action value function is used to measure the expected benefits of taking action in the state and then following the strategy , which is defined as follows: ​

[0066]

[0067] Let be the discounted cumulative return at time ; denote the expected operation under policy ; denote the immediate reward and its discounted cumulative item obtained at each future time, respectively.

[0068] In order to obtain the optimal policy, the state value function and the action value function can be estimated in various ways, and the policy can be selected and improved accordingly. The temporal difference (TD) algorithm combines the ideas of Monte Carlo estimation and dynamic programming, and can use the real experience obtained in the sample to update the state value function The formula is as follows:

[0069]

[0070] In the formula, is the learning rate; is the temporal difference target; is the temporal difference bias.

[0071] In the test scene generation, reinforcement learning often faces the problem of sparse rewards. Sparse rewards refer to the fact that in a vast state space, only at the moment of collision can positive feedback be obtained, and during the exploration of the collision scene, the agent hardly gets any immediate incentive. Therefore, it needs to undergo a large number of blind trials to accidentally obtain a reward - to find those high-value collision scenes - which not only significantly slows down the learning process, but may even make it difficult for the model to converge. To solve this problem, the present invention proposes a reward function setting method based on a risk field to provide continuous and instructive incentives in a wider state space, significantly improving the learning efficiency of scene generation. The reward function setting method based on the risk field proposed by the present invention constructs a reward function composed of an environmental risk field The reward function is introduced into the reinforcement learning process as an immediate reward part, and the specific expression is as follows:

[0072]

[0073] In the formula, represents the risk field deterioration degree due to rain and fog weather, which is modeled by rainfall intensity or fog visibility; is the motion risk field formed by the vehicle motion; is the road risk field formed by static road, and the specific expressions of the two are:

[0074]

[0075] In the formula, represents the real distance between EV and BV, with the unit of m; is the speed of vehicle movement; represents the road condition coefficient; represents the virtual mass of the target object; represents the importance of road line type, and the value of single solid line is greater than the value of dotted line; represents the lateral distance between the EV center axis and the lane line; is a constant value, which determines the speed at which the safety risk increases when the EV approaches the lane line.

[0076] Q-learning is a typical value function-based reinforcement learning algorithm. It introduces Q value to describe the expected return that the agent can obtain by performing each action in each state, and constantly corrects the Q value table according to the rewards or punishments received in the process of interaction between the agent and the environment. After completing the training of the test scene generation, the agent can select the optimal action that can maximize the long-term cumulative return according to the Q value in any state. The update of Q value still uses the time difference (TD) method, and the update formula of Q value is as follows:

[0077]

[0078] In the formula, represents the value estimate corresponding to the action performed in the current state ; and represents the maximum Q value that can be obtained in the next state from all available actions. The Q-learning agent uses -greedy strategy to determine the action to balance between exploring new actions and using known optimal actions:

[0079]

[0080] In the above formula, represents the probability of taking action in state ; represents the total number of available actions in state , is the greedy coefficient. The greedy strategy balances the trade-off between exploration and exploitation by switching between random exploration and exploiting the current optimal action: exploration can discover potentially high-value actions, but too much of it makes it difficult to converge; while over-exploiting existing knowledge easily falls into local optima. To strike a balance between the two, the present invention introduces an exponentially decaying greedy scheme, i.e., gradually reducing the value as training progresses, so that the algorithm has enough exploration ability at the beginning, and relies more on learned good strategies at the later stage, thus ensuring sufficient exploration and promoting stable convergence; the expression of the greedy coefficient is:

[0081]

[0082] where and represent the set starting and ending greedy coefficients, respectively; represents the current round number; and represent the decay coefficient and decay step, respectively, for adjusting the speed of decay.

[0083] The Sine Cosine Algorithm (SCA) is a swarm-based optimization method. It first randomly generates a set of initial solutions, and then uses the oscillation characteristics of sine and cosine functions to move each solution around the current global optimal solution, thus achieving a balance between global exploration and local exploitation. The iteration formula of SCA individuals is:

[0084]

[0085] where represents the current iteration round, is the th iteration of the individual coordinate in the th dimension; in the th iteration, is the individual coordinate in the th dimension, is the coordinate of the global optimal solution obtained in the current iteration in the th dimension; and is four random numbers, respectively responsible for controlling the exploration-exploitation trade-off of the algorithm, the update step, the pulling force of the global optimal solution on the individual, and the selection of sine / cosine operation. Usually, uniformly sampled from [0, 2π], uniformly sampled from [0, 2], uniformly sampled from [0, 1]; ​The value of is calculated according to the following formula:

[0086]

[0087] In the formula, It is a constant. Indicates the current iteration round number. It represents the total number of iterations.

[0088] Based on this, this invention further proposes a "diffusion sine and cosine Q-learning" method. This method introduces a diffusion sine and cosine process into the standard Q-learning framework as the core module of inverse curriculum generation to improve learning efficiency. Inverse curriculum generation is a strategy that alleviates the sparse reward problem by ranking the difficulty of learning tasks without adjusting the original reward function. Its basic idea is to break down the training process into multiple stages from easy to difficult, allowing the agent to quickly accumulate experience in simple scenarios first, and then gradually transition to more complex scenarios, thereby accelerating convergence. The diffusion sine and cosine process is designed according to this concept: it is composed of several SCA modules with different parameters combined sequentially, and each module focuses on a wider search area than the previous one—starting from the vicinity of the target and continuously expanding outward. In this way, the agent can focus on the vicinity area most likely to generate high rewards in the early stage. As the course progresses, the search radius gradually increases, helping it to explore the entire state space more comprehensively. As the number of iterations increases, the concentration of the population on the optimal solution gradually decreases, while its scope of focus continues to expand.

[0089] Diffusion-cosine Q-learning can be divided into three main stages: Initial SCA stage: The target location is searched using the standard SCA method, and the Q-value table is initially initialized. Expansion SCA stage: Composed of multiple SCA modules with different parameters, the search radius is gradually expanded sequentially, guiding the agent to learn from easy to difficult in a reverse learning format. Q-learning stage: After completing the reverse learning, the process switches to Q-learning, using reinforcement learning strategies to finally generate the test scenario. Except for the expansion stage, which requires dynamically adjusting the parameters of each SCA module to broaden the population's exploration range, the initial SCA and Q-learning steps are the same as the original diffusion-cosine Q-learning process. The core improvement of this method lies in the step size coefficient in SCA. The update strategy has been optimized, specifically... The updated formula is as follows:

[0090]

[0091] In the formula, It is the first The update parameters of each SCA module are used to change the individual update direction and step size. Indicates the total number of modules; It is the first Constants of each SCA module, The design is to gradually decrease the step length of the individual moving towards the target. It increases and decreases, thus tending to explore the surrounding area gradually; This is a scaling factor used to adjust... The updated final value. The diffusion sine / cosine process inherits the optimal individual from the SCA process; however, unlike SCA-Q, the optimal individual in the population... The solution has already been iterated to the global optimum, meaning no further iterations are needed. Use Update formula (14) Update At the same time, update randomly The population individuals are iterated using the SCA individual iteration formula (12). For each individual, action selection is performed, an action is taken, and then the state and reward are observed and the Q value is updated as a reverse course generation. The above process is repeated until the current SCA module reaches the set maximum number of iterations. The diffusion sine and cosine process ends when all SCA modules have been executed.

[0092] The diffusion sine and cosine Q-learning algorithm, without increasing the total number of SCA iterations, can further accelerate the overall iterative convergence speed of the test scenario generation algorithm by adjusting the parameters of each SCA module to gradually expand its scope of interest and updating the Q value as the inverse course of reinforcement learning.

Claims

1. A method for generating rain and fog test scenarios for intelligent vehicles based on reinforcement learning, characterized in that: Includes the following steps: Step 1, Test Scenario Modeling: Model the scenario as a stochastic process defined by a set of possible vehicle motion states in the test scenario, a set of selectable vehicle actions in the test scenario, a state transition probability function, a reward function, and a reward discount factor quintuple. Step 2: Introduce the value function: Based on reinforcement learning, for a given reinforcement learning policy, introduce the state value function and the action value function; and perform single-step correction of the value using the temporal difference algorithm. Step 3: Set the reward function: In test scenario generation, the reward function is obtained by combining the reward function setting method based on the risk field; the reward function setting method based on the risk field is as follows: Constructing an environmental risk field The reward function, which is incorporated as the immediate reward component, is introduced into the reinforcement learning process. Its specific expression is: ; In the formula, It represents the degree of risk field deterioration caused by rain and fog weather, and is modeled by rainfall intensity or fog visibility; It refers to the motion risk field formed by the movement of vehicles; This refers to the road risk field formed by static roads, and the specific expressions for both are as follows: ; ; In the formula, This represents the actual distance between the main vehicle (EV) and the surrounding traffic vehicles (BV). The speed of the vehicle; Represents road condition coefficients; The virtual mass representing the target object; The importance of road alignment is represented by a single solid line. The value is greater than the value of the dotted line; This represents the lateral distance between the EV's centerline and the lane line. This is a constant value that determines the rate at which the safety risk increases when the EV approaches the lane line; Step 4, Diffusion-cosine Q-learning: The reinforcement learning algorithm introduces Q-values ​​to characterize the expected reward that the agent can obtain by performing each action in each state, and continuously updates the Q-value table; After training to generate the test scenario, the agent selects the optimal action that maximizes the long-term cumulative reward based on the Q value in any state. The diffusion sine and cosine process is introduced into the standard reinforcement learning framework. The diffusion sine and cosine process is composed of several sine and cosine algorithm (SCA) modules with different parameters combined in sequence. Diffusion-cosine Q-learning consists of three stages: initial SCA stage, expansion SCA stage, and Q-learning stage. The parameters of each SCA individual are dynamically adjusted during the expansion SCA phase. The iteration of the SCA individuals is used to iterate the population individuals. For each individual, action selection is performed, the action is taken, and then the state and reward are observed and the Q value is updated as a reverse course generation. The above process is repeated until the number of iterations of the current SCA module reaches the set maximum number of iterations. The diffusion sine and cosine process ends when all SCA modules have been executed.

2. The method for generating rain and fog test scenarios for intelligent vehicles based on reinforcement learning according to claim 1, characterized in that: The motion state of the vehicle refers to the spatial position and speed of the tested vehicle and surrounding traffic vehicles under specific rainfall intensity or fog visibility conditions; motion state sequence. ,in This represents the state of the main vehicle EV at the k-th moment; Represents the k-th time. The status of the surrounding traffic vehicles (BV). The vehicle's motion refers to the lateral and longitudinal accelerations of the tested vehicle and surrounding traffic vehicles under specific rainfall intensities or fog visibility conditions; the motion of the main vehicle EV is calculated by constructing a main vehicle EV proxy model. This represents the action of the main vehicle EV at the k-th moment; The motion of surrounding traffic vehicles (BV) is obtained through a diffusion sine / cosine Q-learning method. Represents the k-th time. The movement of surrounding traffic vehicles (BV) includes lateral and longitudinal acceleration.

3. The method for generating rain and fog test scenarios for intelligent vehicles based on reinforcement learning according to claim 1, characterized in that: The specific form of the state transition probability function for: ; The specific form of the reward function for: ; in, and Both represent sets The state variables in This is the current state. The next state; Represents a set of actions A certain action in; Indicates the execution of an action Then, the environment in the next moment The reward given is used to measure the current state. With action The benefits brought about; and They are time points Random variables of state and action, Corresponding to the next moment State variables; Represents random variables Values The probability, That is The expected value.

4. The method for generating rain and fog test scenarios for intelligent vehicles based on reinforcement learning according to claim 1, characterized in that: The state value function The action value function is used to measure the long-term benefit of executing policy π in state s. Used to measure the action taken in state s The expected return of then following strategy π is formally defined as follows: ; ; At any moment ,remember Accumulate the discount reward for that moment; This represents the expected operation under strategy π; …represent the instant rewards and their cumulative discounts at various future moments; Indicates the execution of an action Then, the environment in the next moment The reward given; The value is corrected step-by-step using a time-difference algorithm, updating the state-value function. The formula is as follows: ; In the formula, It is the learning rate; It is a time-difference objective; This refers to the timing difference bias.

5. The method for generating rain and fog test scenarios for intelligent vehicles based on reinforcement learning according to claim 1, characterized in that: In the diffusion sine and cosine Q-learning process, the Q-value is updated using the temporal difference method, and the Q-value update formula is as follows: ; In the formula, This indicates that an action should be performed in the current state s. The corresponding value estimate; Indicates the next state The maximum Q value that can be achieved among all available actions; Indicates the execution of an action Then, the environment in the next moment The reward given.

6. The method for generating rain and fog test scenarios for intelligent vehicles based on reinforcement learning according to claim 1, characterized in that: Reinforcement learning agents adopt - A greedy strategy is used to determine actions, striking a balance between exploring new actions and utilizing known optimal actions: ; In the formula, This indicates taking an action in state s. The probability of; This represents the total number of actions that can be performed in state s. The greedy coefficient; The greedy strategy introduces exponential decay, meaning that the value is gradually reduced as training progresses. value.

7. The method for generating rain and fog test scenarios for intelligent vehicles based on reinforcement learning according to claim 6, characterized in that: The expression for the greedy coefficient is: ; In the formula, and These represent the greedy coefficients set at the start and end points, respectively. Indicates the current round number; and These represent the attenuation coefficient and attenuation step size, respectively, used for adjustment. The rate of decay.

8. The method for generating rain and fog test scenarios for intelligent vehicles based on reinforcement learning according to claim 1, characterized in that: In step 4, the iterative formula for the SCA individual is: ; In the formula, Indicates the current iteration round number; for In the nth iteration Individuals on the dimension Coordinates; in the first In the next iteration For the first Individuals on the dimension coordinate; The global optimal solution obtained in the current iteration is at the th iteration. Coordinates in dimensional space; Four random numbers, The value is determined according to the formula. Calculate, where, It is a constant. Indicates the current iteration round number. It represents the total number of iterations.

9. The method for generating rain and fog test scenarios for intelligent vehicles based on reinforcement learning according to claim 8, characterized in that: In step 4, the parameters of each SCA individual are dynamically adjusted during the expansion SCA phase, including the step size coefficient in the SCA. Update the coefficients randomly. ; random numbers Uniform sampling from [0, 2π] Uniform sampling from [0,2] Uniform sampling from [0,1]; Step size coefficient The updated formula is as follows: ; In the formula, It is the first The update parameters of each SCA module are used to change the individual update direction and step size; Indicates the total number of modules; It is the first Constants of each SCA module, The design is to gradually decrease the step length of the individual moving towards the target. Increase and decrease; This is a scaling factor used to adjust... The final updated value.

Citation Information

Patent Citations

  • SCA-QL-based path planning method

    CN115016499A

  • Line power flow control method based on deep reinforcement learning

    CN116470511A