Intelligent driving simulation method and system based on multi-agent reinforcement learning
Through the intelligent driving simulation method of multi-agent reinforcement learning, combining safety, efficiency and comfort, the reward function is designed, and the battery energy consumption allocation and vehicle speed are dynamically adjusted, which solves the problems of environmental adaptability and energy waste in the traditional method, and achieves efficient and automated driving optimization.
Patent Information
- Application Number
- CN202510650414.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-05-20
AI Technical Summary
Traditional path planning methods are difficult to adapt to the dynamically changing traffic environment, and allocation strategies lead to energy waste and shortening of range. The selection of attenuation method in reinforcement learning affects the convergence speed and stability of the algorithm.
Using an intelligent driving simulation method based on multi-agent reinforcement learning, the design reward function combines safety, efficiency and comfort, uses the Bellman equation to update the state value table, combines the time series model to predict future vehicle status, dynamically adjusts the battery energy consumption allocation, and adjusts the vehicle speed through Q-learning to maintain a safe distance.
Multi-objective optimization in complex traffic environments has been achieved, driving performance and energy utilization efficiency have been improved, energy waste has been reduced, and driving automation level and robustness have been improved.
Smart Images

Figure CN120540387A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent driving technology, and more specifically, to an intelligent driving simulation method and system based on multi-agent reinforcement learning. Background Art
[0002] Intelligent control technology is primarily used to automate complex systems and processes, enabling efficient, reliable, and intelligent operation. It is widely used in various fields, including industrial production, transportation, and energy management. Its application in rail transit turnout snow-melting systems can improve snow-melting efficiency and reduce track risks.
[0003] The existing technology has the following deficiencies:
[0004] Traditional path planning methods are typically based on static maps and rules, making them difficult to adapt to dynamically changing traffic environments. They neglect global efficiency and safety, and are prone to leading to local optimal solutions. Allocation strategies are often based on fixed rules or simple heuristic algorithms, resulting in wasted energy and reduced range. In reinforcement learning, the choice of decay method significantly impacts the algorithm's convergence speed and stability, making it difficult to adapt to the demands of complex tasks.
[0005] In view of the above problems, the present invention proposes a solution. Summary of the Invention
[0006] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide an intelligent driving simulation method and system based on multi-agent reinforcement learning to solve the problems raised in the above-mentioned background technology.
[0007] To achieve the above object, the present invention provides the following technical solutions:
[0008] An intelligent driving simulation method and system based on multi-agent reinforcement learning includes the following steps:
[0009] Step S1: Design a reward function based on the agent's safety, efficiency, and comfort, interact with the environment to obtain immediate rewards, and use the Bellman equation to update the state value table to gradually find the optimal path;
[0010] Step S2: Using the time series combined with safety, efficiency, and comfort to predict the future vehicle state, perform predictive analysis on power distribution, and dynamically adjust the battery energy consumption distribution strategy based on the prediction results;
[0011] Step S3: Decide whether to use the linear decay method or the exponential decay method based on the complexity of the task and the smoothness of the reward level;
[0012] Step S4: Based on Q-learning and combining safety and power distribution, the vehicle speed is adjusted to maintain a safe distance.
[0013] In a preferred embodiment, step S1 includes the following contents:
[0014] The reward function is obtained by comprehensively considering the three factors of safety, efficiency and comfort, and an immediate reward R is assigned to each action of the agent:
[0015] Define the safety distance threshold D safe , calculate the distance D between the agent and the nearest obstacle current , the safety factor is expressed as:
[0016] Define the remaining distance D from the agent to the target point goal ; Estimate the time required for the agent to reach the target point at the current speed V
[0017] The efficiency coefficient is expressed as: Among them: k1 is a tuning parameter, T goal is the ideal time for the agent to reach the goal point;
[0018] Calculate the acceleration change rate Δa and steering angle change rate Δθ of the intelligent body, and the comfort coefficient is expressed as: R comfort =-(w a |Δa|+w θ |Δθ|), |Δa|: change in acceleration; |Δθ|: change in steering angle; w a and w θ is the weight used to balance the effects of acceleration and steering angle on comfort;
[0019] The weighted sum of the rewards for safety, efficiency, and comfort is used to form the final immediate reward: R = w s ·R safety +w e ·R efficiency +w c ·R comfort , where: w s ,w e ,w c is the corresponding weight of safety, efficiency and comfort, which is used to balance the importance of safety, efficiency and comfort; R safety : is the security reward; R efficiency is the efficiency reward; R comfort It is a comfort bonus;
[0020] The agent performs actions in the environment, and the environment updates its state based on the agent's actions and returns the next state s' and the data of the immediate reward R;
[0021] The state value table V(s) records the value of each state s, the expected value of the long-term cumulative reward that can be obtained according to a certain strategy;
[0022] The Bellman equation is used to update the state value table. The process is expressed as: V(s)←V(s)+α·[R+γ·V(s′)-V(s)], where: V(s) is the value of the current state s, α is the learning rate, which controls the weight of new information, R is the immediate reward, γ is the discount factor, which controls the weight of future rewards, V(s ′ ) is the next state s ′ the value of
[0023] Initialize the state value table V(s), set the initial values of all states to 0 or random values, select an action according to the current strategy, randomly select an action, select the current optimal action with probability 1-∈, execute the action in the environment, observe the next state s′ and immediate reward R, and use the Bellman equation to update the value V(s) of the current state. Repeat the above steps until the state value table converges or the maximum number of iterations is reached.
[0024] In a preferred embodiment, step S2 includes the following:
[0025] Set a threshold for the immediate reward R and record it as R th , when R>R th , the solution will be marked as unreasonable;
[0026] Obtain the judgment results, conduct power distribution forecast analysis through time series models, and further refine and optimize the initialization distribution plan;
[0027] The judgment results include the unreasonable results of the initialization allocation plan;
[0028] Furthermore, the specific steps for power distribution forecast analysis using a time series model are as follows:
[0029] Step B1, obtaining prediction data;
[0030] Step B2, establishing an ARIMAX model;
[0031] Step B3, using the maximum likelihood estimation (MLE) method to estimate the ARIMAX model parameters;
[0032] Step B4: verify the effect of the fitted model and check the goodness of fit of the model by residual analysis;
[0033] Step B5: Use the fitted model to predict the future distribution plan, and use the maximum value of the predicted result as the highest point of power demand in the future time period;
[0034] Power demand data refers to historical power demand data, and its data serves as the main variable of the time series;
[0035] Furthermore, the basic form of the ARIMAX model is:
[0036]
[0037] Where y t is the power demand at the current time point, α is a constant term, is the i-th order autoregressive parameter, p is the order of the autoregressive term, θ j is the jth order moving average parameter, q is the order of the moving average term, ∈ t is a white noise term, representing random error, β k is an exogenous variable X t-k The coefficient of m is the lag order of the exogenous variable, X t-k is an exogenous variable lagged k periods;
[0038] Error term ∈ t Obey the normal distribution N(0,σ 2 ), then the likelihood function is:
[0039]
[0040] Taking the logarithm of the likelihood function gives the log-likelihood function:
[0041]
[0042] By maximizing the log-likelihood function, we can obtain the parameter estimates μ, θ1, β1, β2;
[0043] By predicting the maximum, minimum and average values of the results, the initialization battery energy consumption allocation strategy is dynamically adjusted.
[0044] In a preferred embodiment, step S3 includes the following contents:
[0045] The probability distribution of state transition is expressed as: P(s′|s,a). In a deterministic environment, the state transition is expressed as:
[0046] In a random environment, state transition has a certain probability distribution: P(s′|s,a)=[p1,p2,...,p n ], where: p1 is the probability of taking action a from state s to state s1, and all probabilities satisfy ∑ i P i =1;
[0047] Environmental complexity C rand It is quantified by the entropy of the state transition probability distribution, expressed as: C rand =H(P)=-∑ i p i log(p i );
[0048] The stability S of the reward degree is calculated and the stability score is defined as: Where ε is a small constant, σ R is the standard deviation of the rewards;
[0049] Record the cumulative rewards R1, R2, ..., R obtained by the agent in multiple rounds n , calculate the standard deviation of the reward: Where: R i is the cumulative reward of the i-th round, is the average of the cumulative rewards;
[0050] According to the complexity of the environment C rand And the smoothness S of the reward signal, the linear decay method or the exponential decay method is selected based on the comprehensive judgment, and the comprehensive score is expressed as: T = w C ·C rand +w S (1-S), where w C 、w S is the weight, and the threshold θ is set for the comprehensive score. When, choose exponential decay method: ∈=∈ min +(∈ max -∈ min )·e -λ·current_step Otherwise, the linear decay method is selected:
[0051] In a preferred embodiment, step S4 includes the following contents:
[0052] The state is defined as the current vehicle's speed, the distance to the preceding vehicle, the preceding vehicle's speed, and its own dynamic state. The state space represents the discretization of the above states to form a finite set of states.
[0053] Actions are defined as actions that a vehicle can take;
[0054] Reward function design guides the learning process to ensure safe vehicle distance and energy-efficient driving, including safety rewards and efficiency rewards;
[0055] Initialize the Q-table, assign initial values, learning rate α and discount factor γ to the Q values of all state-action pairs;
[0056] Collect experience through interaction with the environment and store it in the experience pool;
[0057] Sampling and updating: Where: s is the current state, a is the action taken, R is the immediate reward obtained, s ′ is the next state;
[0058] After selecting an action based on the Q-learning algorithm, the power distribution is dynamically adjusted based on the current vehicle power state and external environmental conditions;
[0059] Continuously monitor the distance and relative speed to the vehicle in front, and adjust the vehicle speed in real time based on the decisions made by the Q-learning algorithm to ensure a safe distance.
[0060] An intelligent driving simulation method and system based on multi-agent reinforcement learning includes: a data acquisition module, a data classification module, a data matching module, and a feedback adjustment module, with signal connections between the modules;
[0061] Data acquisition module: This module designs a reward function based on the agent's safety, efficiency, and comfort, interacts with the environment to obtain immediate rewards, and uses the Bellman equation to update the state value table and gradually find the optimal path.
[0062] Data classification module: This module uses time series to combine safety, efficiency, and comfort to predict future vehicle status, conducts predictive analysis on power distribution, and dynamically adjusts the battery energy consumption distribution strategy based on the prediction results.
[0063] Data matching module: decides when to use linear decay or exponential decay based on the complexity of the task and the smoothness of the reward level;
[0064] Feedback adjustment module: Based on Q-learning, it combines safety and power distribution to adjust vehicle speed and maintain a safe distance.
[0065] The technical effects and advantages of the intelligent driving simulation method and system based on multi-agent reinforcement learning of the present invention are as follows:
[0066] Through multi-agent reinforcement learning, the system enables intelligent path planning, power distribution, and speed regulation, reducing human intervention and improving the level of automated driving. Taking into account safety, efficiency, and comfort, it can achieve multi-objective optimization in complex traffic environments, improving overall driving performance. Dynamically adjusting strategies based on environmental changes, it exhibits strong robustness and adaptability, enabling it to excel in diverse driving scenarios. Dynamically adjusting battery energy allocation strategies improves energy efficiency and reduces energy waste, meeting the requirements of sustainable development. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1This is a structural diagram of the intelligent driving simulation method based on multi-agent reinforcement learning of the present invention.
[0068] Figure 2 This is a schematic diagram of the structure of the intelligent driving simulation system based on multi-agent reinforcement learning in the present invention. DETAILED DESCRIPTION
[0069] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0070] Example 1
[0071] The present invention discloses an intelligent driving simulation method and system based on multi-agent reinforcement learning, comprising the steps of:
[0072] Step S1: Design a reward function based on the agent's safety, efficiency, and comfort, interact with the environment to obtain immediate rewards, and use the Bellman equation to update the state value table to gradually find the optimal path;
[0073] Step S2: Using the time series combined with safety, efficiency, and comfort to predict the future vehicle state, perform predictive analysis on power distribution, and dynamically adjust the battery energy consumption distribution strategy based on the prediction results;
[0074] Step S3: Decide whether to use the linear decay method or the exponential decay method based on the complexity of the task and the smoothness of the reward level;
[0075] Step S4: Based on Q-learning and combining safety and power distribution, the vehicle speed is adjusted to maintain a safe distance.
[0076] In step S1, the reward function is designed based on the agent's safety, efficiency, and comfort. The agent interacts with the environment to obtain immediate rewards. The Bellman equation updates the state value table and gradually finds the optimal path. The specific contents include:
[0077] The reward function is the core of guiding the behavior of the agent in reinforcement learning. It influences the learning of the strategy by quantifying the quality of the agent's behavior. The reward function is obtained by comprehensively considering three factors: safety, efficiency, and comfort. An immediate reward R is assigned to each action of the agent, allowing the agent to learn the optimal strategy by maximizing the cumulative reward:
[0078] Safety: Avoid collisions and ensure that the agent maintains a safe distance from other vehicles or obstacles: define a safe distance threshold D safe; Calculate the distance D between the agent and the nearest other vehicle or obstacle current When the distance is less than the safety threshold, a larger negative reward is given, indicating an increased risk; when the distance is greater than the safety threshold, a smaller positive reward is given, indicating a higher level of safety. The safety factor is expressed as: When D current <D safe When D current ≥D safe When , a positive reward is given smoothly. The farther the distance, the higher the reward, but the growth gradually tends to be smooth. k is a regulation parameter used to control the growth rate of the function.
[0079] Efficiency: reach the target point as quickly as possible and reduce travel time: define the remaining distance D from the agent to the target point goal ; Estimate the time required for the agent to reach the target point at the current speed V When the agent is close to the target point, it is given a larger positive reward; when the agent is far away from the target point or the speed is too low, it is given a smaller reward or penalty. The efficiency coefficient is expressed as: When T remaining When the agent speed is >0, a positive reward is given smoothly, and the shorter the remaining time, the higher the reward. When the agent speed is zero, a large negative reward is given, indicating extremely low efficiency. k1 is a tuning parameter used to control the growth rate of the function. T goal is the ideal time for the agent to reach the goal point.
[0080] Comfort: Reduce sudden acceleration and steering to improve passenger experience: Calculate the acceleration change rate Δa and steering angle change rate Δθ of the agent. When the acceleration or steering angle changes greatly, a negative reward is given. When the acceleration and steering angle change is small, a positive reward is given. The comfort coefficient is expressed as: R comfort =-(w a |Δa|+w θ |Δθ|), |Δa|: change in acceleration; |Δθ|: change in steering angle; w a and w θ is the weight used to balance the effects of acceleration and steering angle on comfort.
[0081] The weighted sum of the rewards for safety, efficiency, and comfort is used to form the final immediate reward: R = w s ·R safety +w e ·R efficiency +w c ·R comfort , where: w s ,w e ,wc is the corresponding weight of safety, efficiency and comfort, which is used to balance the importance of safety, efficiency and comfort; R safety : is the security reward; R efficiency is the efficiency reward; R comfort It's a comfort bonus.
[0082] The agent performs actions in the environment, and the environment updates its state based on the agent's actions and returns the next state s' and the data of the immediate reward R;
[0083] The state value table V(s) records the value of each state s, which represents the expected value of the long-term cumulative reward that can be obtained according to a certain strategy starting from the current state;
[0084] The Bellman equation is used to update the state value table. The process is expressed as: V(s)←V(s)+α·[R+γ·V(s′)-V(s)], where: V(s) is the value of the current state s, α is the learning rate, which controls the weight of new information, R is the immediate reward, γ is the discount factor, which controls the weight of future rewards, V(s ′ ) is the next state s ′ the value of
[0085] Initialize the state value table V(s), set the initial values of all states to 0 or random values, select an action according to the current strategy, randomly select an action, select the current optimal action with probability 1-∈, execute the action in the environment, observe the next state s′ and immediate reward R, and use the Bellman equation to update the value V(s) of the current state. Repeat the above steps until the state value table converges or the maximum number of iterations is reached.
[0086] In step S2, the future vehicle state is predicted using time series combined with safety, efficiency, and comfort. Power distribution is analyzed and predicted, and the battery energy consumption distribution strategy is dynamically adjusted based on the prediction results. The specific contents include:
[0087] Set a threshold for the immediate reward R and record it as R th , when R>R th , the solution will be marked as unreasonable;
[0088] Obtain the judgment results, conduct power distribution forecast analysis through time series models, and further refine and optimize the initialization distribution plan;
[0089] The judgment results include the unreasonable results of the initialization allocation plan;
[0090] It should be noted that the time series model used in this embodiment is an ARIMAX model. Refining and optimizing the initialization allocation plan means that if the current initialization allocation plan is determined to be unreasonable, further analysis and adjustment are performed to make the power allocation plan more scientific and reasonable, thereby achieving control of the internal energy consumption of the battery.
[0091] Furthermore, the specific steps for power distribution forecast analysis using a time series model are as follows:
[0092] Step B1, obtaining prediction data;
[0093] Step B2, establishing an ARIMAX model;
[0094] Step B3, using the maximum likelihood estimation (MLE) method to estimate the ARIMAX model parameters;
[0095] Step B4: verify the effect of the fitted model and check the goodness of fit of the model by residual analysis;
[0096] Step B5: Use the fitted model to predict the future distribution plan, and use the maximum value of the prediction result as the highest point of power demand in the future time period.
[0097] Specifically, the prediction data includes safety, efficiency, comfort, and power demand data corresponding to each time point. The safety, efficiency, and comfort data corresponding to each time point have been exemplified in Example 1 and will not be repeated here.
[0098] Power demand data refers to historical power demand data, and its data serves as the main variable of the time series;
[0099] Furthermore, the basic form of the ARIMAX model is:
[0100]
[0101] Where y t is the power demand at the current time point, α is a constant term, is the i-th order autoregressive parameter, p is the order of the autoregressive term, θ j is the jth order moving average parameter, q is the order of the moving average term, ∈ t is a white noise term, representing random error, β k is an exogenous variable X t-k The coefficient of m is the lag order of the exogenous variable, X t-k is an exogenous variable lagged k periods;
[0102] It should be noted that the exogenous variable part can integrate the impact of other relevant variables (such as safety, efficiency, and comfort) on power demand;
[0103] It should be noted that, in step B3, α, θ1, β1, and β2 are calculated using the maximum likelihood estimation (MLE) method. The specific steps are as follows:
[0104] Error term ∈ t Obey the normal distribution N(0,σ 2 ), then the likelihood function is:
[0105]
[0106] Taking the logarithm of the likelihood function gives the log-likelihood function:
[0107]
[0108] By maximizing the log-likelihood function, we can obtain the parameter estimates μ, θ1, β1, β2;
[0109] Dynamically adjust the initialization battery energy consumption allocation strategy by predicting the maximum, minimum and average values of the results;
[0110] This embodiment gives the following examples:
[0111] The maximum value of the predicted result Represents the highest point of power demand in the future time period, formulates peak distribution strategy, and allocates power during the predicted power demand peak period (i.e. close to or equal to At the same time, the battery output power is increased by 20% to ensure the smooth operation of the vehicle and avoid power shortage;
[0112] According to the lowest value of the predicted results or valley interval, formulate valley period allocation strategy, in the valley period of power demand (i.e. close to or below At the same time, reduce the battery output power by 15% to extend the battery life and efficiency;
[0113] According to the average value of the predicted results Formulate a balance period allocation strategy to maintain the normal output of the battery during the period when the power demand is close to the average value, ensuring the continuous and stable operation of the vehicle. time period to maintain the normal output of the battery and ensure the continuous and stable operation of the vehicle.
[0114] In step S3, the linear decay method or the exponential decay method is selected based on the complexity of the task and the smoothness of the reward level. The exploration rate ∈ is dynamically adjusted using the selected screening subtraction method. The specific contents include:
[0115] The probability distribution of state transition is expressed as: P(s′|s,a). In a deterministic environment, the state transition is completely determined, which is expressed as:
[0116] In a random environment, state transition has a certain probability distribution: P(s′|s,a)=[p1,p2,...,p n ], where: p1 is the probability of taking action a from state s to state s1, and all probabilities satisfy ∑ i P i =1;
[0117] Environmental complexity C rand It is quantified by the entropy of the state transition probability distribution, expressed as: C rand =H(P)=-∑ i p i log(p i ), the higher the entropy, the more complex the environment; when the entropy is 0, the environment is completely determined.
[0118] The stability S of the reward degree is calculated and the stability score is defined as: Where ε is a small constant used to avoid the denominator from being zero. R Larger, indicating that the reward signal fluctuates more and S is smaller; σ R The smaller it is, the more stable the reward signal is, and S is larger.
[0119] During the training process, record the cumulative rewards R1, R2, ..., R obtained by the agent in multiple rounds n , calculate the standard deviation of the reward: Where: R i is the cumulative reward of the i-th round, is the average of the cumulative rewards.
[0120] According to the complexity of the environment C rand And the smoothness S of the reward signal, the linear decay method or the exponential decay method is selected based on the comprehensive judgment, and the comprehensive score is expressed as: T = w C ·C rand +w S (1-S), where w C 、w S is a weight used to balance the influence of complexity and smoothness, C rand The larger the value of T, the more complex the task is; the smaller the value of T, the steeper the reward is. The larger the value of T, the more complex the task is or the steeper the reward is. The threshold θ is set for the comprehensive score. When, choose exponential decay method: ∈=∈ min +(∈max -∈ min )·e -λ·current_step Otherwise, the linear decay method is selected:
[0121] In step S4, the updated ∈ value is used to adjust the vehicle speed to maintain a safe distance based on Q-learning combined with safety and power distribution. The specific contents include:
[0122] Step D1: State definition: This is defined as the current vehicle's speed, distance to the preceding vehicle, the preceding vehicle's speed, and its own power state. The state space representation discretizes the above states into a finite set of states.
[0123] Step D2: Action definition: Defined as the actions the vehicle can take, such as accelerating, decelerating, or maintaining the current speed. The action space includes all possible action options;
[0124] Step D3: Reward Function Design: Design a reward function to guide the learning process, ensuring safe distance and energy-efficient driving. Safety Reward: Give positive rewards when maintaining an appropriate distance from the vehicle in front, and large negative rewards when the distance is too close. Efficiency Reward: Give positive rewards when approaching the target point at an appropriate speed while ensuring safety, and negative rewards for excessive deceleration or acceleration.
[0125] Step D4: Q-learning algorithm application: Initialize the Q-table, assign initial values to the Q values of all state-action pairs, learning rate α and discount factor γ: Set appropriate learning rate and discount factor to balance the importance of immediate rewards and future rewards.
[0126] Step D5: Training process: Collect experience (state, action, reward, next state) through interaction with the environment and store it in the experience pool;
[0127] Sampling and updating: Where: s is the current state, a is the action taken, R is the immediate reward obtained, s ′ It is the next state.
[0128] Step D6: Power distribution adjustment: After selecting the action according to the Q-learning algorithm, the power distribution is dynamically adjusted in combination with the current vehicle power state and external environmental conditions to optimize energy consumption and driving efficiency.
[0129] Step D7: Maintaining a safe distance: During driving, the system continuously monitors the distance and relative speed to the vehicle ahead, and adjusts the vehicle speed in real time based on the decisions made by the Q-learning algorithm to ensure a safe distance.
[0130] An intelligent driving simulation method and system based on multi-agent reinforcement learning includes: a data acquisition module, a data classification module, a data matching module, and a feedback adjustment module, with signal connections between the modules;
[0131] Data acquisition module: This module designs a reward function based on the agent's safety, efficiency, and comfort, interacts with the environment to obtain immediate rewards, and uses the Bellman equation to update the state value table and gradually find the optimal path.
[0132] Data classification module: This module uses time series to combine safety, efficiency, and comfort to predict future vehicle status, conducts predictive analysis on power distribution, and dynamically adjusts the battery energy consumption distribution strategy based on the prediction results.
[0133] Data matching module: decides when to use linear decay or exponential decay based on the complexity of the task and the smoothness of the reward level;
[0134] Feedback adjustment module: Based on Q-learning, it combines safety and power distribution to adjust vehicle speed and maintain a safe distance.
[0135] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formulas are set by technicians in this field according to actual conditions.
[0136] The above embodiments may be implemented in whole or in part through software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments may be implemented in whole or in part in the form of a computer program product.
[0137] Those skilled in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application of the technical solution and the invention constraints. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0138] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0139] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0140] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. An intelligent driving simulation method and system based on multi-agent reinforcement learning, characterized by: Including steps: Step S1: Design a reward function based on the agent's safety, efficiency, and comfort, interact with the environment to obtain immediate rewards, and use the Bellman equation to update the state value table to gradually find the optimal path; Step S2: Using the time series combined with safety, efficiency, and comfort to predict the future vehicle state, perform predictive analysis on power distribution, and dynamically adjust the battery energy consumption distribution strategy based on the prediction results; Step S3: Decide whether to use the linear decay method or the exponential decay method based on the complexity of the task and the smoothness of the reward level; Step S4: Based on Q-learning and combining safety and power distribution, the vehicle speed is adjusted to maintain a safe distance.
2. The intelligent driving simulation method and system based on multi-agent reinforcement learning according to claim 1, characterized in that: The reward function is obtained by comprehensively considering the three factors of safety, efficiency and comfort, and an immediate reward R is assigned to each action of the agent: Define the safety distance threshold D safe , calculate the distance D between the agent and the nearest obstacle current , the safety factor is expressed as: Define the remaining distance D from the agent to the target point goal ; Estimate the time required for the agent to reach the target point at the current speed V The efficiency coefficient is expressed as: Among them: k1 is a tuning parameter, T goal is the ideal time for the agent to reach the goal point; Calculate the acceleration change rate Δa and steering angle change rate Δθ of the intelligent body, and the comfort coefficient is expressed as: R comfort =-(w a |Δa|+w θ |Δθ|), |Δa|: change in acceleration; |Δθ|: change in steering angle; w a and w θ is the weight used to balance the effects of acceleration and steering angle on comfort; The weighted sum of the rewards for safety, efficiency, and comfort is used to form the final immediate reward: R = w s ·R safety +w e ·R efficiency +w c ·R comfort , where: w s ,w e ,w c is the corresponding weight of safety, efficiency and comfort, which is used to balance the importance of safety, efficiency and comfort; R safety : is the security reward; R efficiency is the efficiency reward; R comfort It's a comfort bonus.
3. The intelligent driving simulation method and system based on multi-agent reinforcement learning according to claim 2, characterized in that: The agent performs actions in the environment, and the environment updates its state based on the agent's actions and returns the next state s' and the data of the immediate reward R; The state value table V(s) records the value of each state s, the expected value of the long-term cumulative reward that can be obtained according to a certain strategy; The Bellman equation is used to update the state value table. The process is expressed as: V(s)←V(s)+α·[R+γ·V(s′)-V(s)], where: V(s) is the value of the current state s, α is the learning rate, which controls the weight of new information, R is the immediate reward, γ is the discount factor, which controls the weight of future rewards, and V(s′) is the value of the next state s′; Initialize the state value table V(s), set the initial values of all states to 0 or random values, select an action according to the current strategy, randomly select an action, select the current optimal action with probability 1-∈, execute the action in the environment, observe the next state s′ and immediate reward R, and use the Bellman equation to update the value V(s) of the current state. Repeat the above steps until the state value table converges or the maximum number of iterations is reached.
4. The intelligent driving simulation method and system based on multi-agent reinforcement learning according to claim 3, characterized in that: Set a threshold for the immediate reward R and record it as R th , when R>R th , the solution is marked as unreasonable.
5. The intelligent driving simulation method and system based on multi-agent reinforcement learning according to claim 4, characterized in that: Obtain the judgment results, conduct power distribution forecast analysis through time series models, and further refine and optimize the initialization distribution plan; The judgment results include the unreasonable results of the initialization allocation plan; Furthermore, the specific steps for power distribution forecast analysis using a time series model are as follows: Step B1, obtaining prediction data; Step B2, establishing an ARIMAX model; Step B3, using the maximum likelihood estimation (MLE) method to estimate the ARIMAX model parameters; Step B4: verify the effect of the fitted model and check the goodness of fit of the model by residual analysis; Step B5: Use the fitted model to predict the future distribution plan, and use the maximum value of the predicted result as the highest point of power demand in the future time period; Power demand data refers to historical power demand data, and its data serves as the main variable of the time series; Furthermore, the basic form of the ARIMAX model is: Where y t is the power demand at the current time point, α is a constant term, is the i-th order autoregressive parameter, p is the order of the autoregressive term, θ j is the jth order moving average parameter, q is the order of the moving average term, ∈ t is a white noise term, representing random error, β k is an exogenous variable X t-k The coefficient of m is the lag order of the exogenous variable, X t-k is an exogenous variable lagged k periods away.
6. The intelligent driving simulation method and system based on multi-agent reinforcement learning according to claim 5, characterized in that: Error term ∈ t Obey the normal distribution N(0,σ 2 ), then the likelihood function is: Taking the logarithm of the likelihood function gives the log-likelihood function: By maximizing the log-likelihood function, we can obtain the parameter estimates μ, θ1, β1, β2; By predicting the maximum, minimum and average values of the results, the initialization battery energy consumption allocation strategy is dynamically adjusted.
7. The intelligent driving simulation method and system based on multi-agent reinforcement learning according to claim 6, characterized in that: The probability distribution of state transition is expressed as: P(s′|s,a). In a deterministic environment, the state transition is expressed as: In a random environment, state transition has a certain probability distribution: P(s′|s,a)=[p1,p2,...,p n ], where: p1 is the probability of taking action a from state s to state s1, and all probabilities satisfy ∑ i P i =1; Environmental complexity C rand It is quantified by the entropy of the state transition probability distribution, expressed as: C rand =H(P)=-∑ i p i log(p i ); The stability S of the reward degree is calculated and the stability score is defined as: Where ε is a small constant, σ R is the standard deviation of the reward; Record the cumulative rewards R1, R2, ..., R obtained by the agent in multiple rounds n , calculate the standard deviation of the reward: Where: R i is the cumulative reward of the i-th round, is the average of the cumulative rewards.
8. The intelligent driving simulation method and system based on multi-agent reinforcement learning according to claim 7, characterized in that: According to the complexity of the environment C rand And the smoothness S of the reward signal, the linear decay method or the exponential decay method is selected based on the comprehensive judgment, and the comprehensive score is expressed as: T = w C ·C rand +w S (1-S), where w C 、w S is the weight, and the threshold θ is set for the comprehensive score. When, choose exponential decay method: ∈=∈ min +(∈ max -∈ min )·e -λ·current_step ; Otherwise, choose the linear decay method:
9. The intelligent driving simulation method and system based on multi-agent reinforcement learning according to claim 8, characterized in that: The state is defined as the current vehicle's speed, the distance to the preceding vehicle, the preceding vehicle's speed, and its own dynamic state. The state space represents the discretization of the above states to form a finite set of states. Actions are defined as actions that a vehicle can take; Reward function design guides the learning process to ensure safe vehicle distance and energy-efficient driving, including safety rewards and efficiency rewards; Initialize the Q-table, assign initial values, learning rate α and discount factor γ to the Q values of all state-action pairs; Collect experience through interaction with the environment and store it in the experience pool; Sampling and updating: Where: s is the current state, a is the action taken, R is the immediate reward obtained, and s′ is the next state; After selecting an action based on the Q-learning algorithm, the power distribution is dynamically adjusted based on the current vehicle power state and external environmental conditions; Continuously monitor the distance and relative speed to the vehicle in front, and adjust the vehicle speed in real time based on the decisions made by the Q-learning algorithm to ensure a safe distance.
10. The intelligent driving simulation method and system based on multi-agent reinforcement learning according to claims 1-9, characterized in that: include: Data acquisition module, data classification module, data matching module and feedback adjustment module, and signal connections between each module; Data acquisition module: This module designs a reward function based on the agent's safety, efficiency, and comfort, interacts with the environment to obtain immediate rewards, and uses the Bellman equation to update the state value table and gradually find the optimal path. Data classification module: This module uses time series to combine safety, efficiency, and comfort to predict future vehicle status, conducts predictive analysis on power distribution, and dynamically adjusts the battery energy consumption distribution strategy based on the prediction results. Data matching module: decides when to use linear decay or exponential decay based on the complexity of the task and the smoothness of the reward level; Feedback adjustment module: Based on Q-learning, it combines safety and power distribution to adjust vehicle speed and maintain a safe distance.
Citation Information
Patent Citations
Hybrid electric vehicle control method based on multi-agent deep reinforcement learning
CN115793445A
Robot obstacle avoidance reinforcement learning method and device based on obstacle function
CN118170142A
Automatic driving key scene generation method based on multi-agent reinforcement learning
CN118468700A
Multi-agent deep reinforcement learning path planning method based on improved A*heuristic
CN118759846A
Automatic driving freight formation collaborative decision-making method, device and equipment and storage medium
CN119105496A
Cited By
Large-scale intelligent driving beam transporting vehicle obstacle avoidance control method and system based on accurate positioning
CN120972981A