Multi-vehicle cooperative drift control method based on self-state sensing of body-driven-by-wire chassis
By using the embodied drive-by-wire chassis self-state perception network (ESSA network) and the multi-agent reinforcement learning (E-MARL) framework, the complexity and uncertainty in multi-vehicle cooperative drift control are solved, achieving cooperative formation control and stability of multiple vehicles under extreme conditions, and improving the system's adaptive capability and collision avoidance capability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIHANG UNIV
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-17
AI Technical Summary
Existing single-vehicle drift control technology is prone to instability under nonlinear conditions, and traditional methods are difficult to cope with the complexity and uncertainty of multi-vehicle cooperative drift, resulting in low system efficiency and limited application scenarios.
Employing an embodied self-state perception (ESSA) network for drive-by-wire chassis, combined with the multi-agent reinforcement learning (E-MARL) framework, collaborative drift control of multiple vehicles is achieved through graph models and action guidance functions. The embodied self-state perception module infers the internal state in real time and dynamically adjusts the strategy.
It enables collaborative formation control of multiple autonomous vehicles under extreme conditions, enhances the adaptive capability to transient nonlinear dynamic characteristics, ensures consistency of relative position and heading, and effectively avoids collisions, achieving complete closed-loop control from theoretical model to engineering implementation.
Smart Images

Figure CN121884575A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of autonomous driving and vehicle dynamic control technology, specifically to a multi-vehicle cooperative drift control method based on the self-state perception of an embodied drive-by-wire chassis. Background Technology
[0002] Driven by advanced computing power and intelligent algorithms, autonomous driving technology is becoming a key force in transforming human transportation. Drift control technology plays a crucial role in enhancing vehicle agility and safety in autonomous driving. However, existing drift technologies, even in single-vehicle applications, suffer from low fault tolerance and redundancy, are prone to instability under nonlinear conditions, have extremely limited application scenarios, and offer little room for optimization in system efficiency and overall performance. These issues severely restrict the widespread application of drift control in real-world scenarios. Furthermore, real-world applications often require consideration of multi-vehicle drifting, which exacerbates these limitations, and currently, there is a significant lack of technologies for coordinated drift control.
[0003] Therefore, cooperative drift control has become a key breakthrough for the practical application of autonomous driving technology. Collision avoidance, speed matching, and swarm centralization are crucial elements for achieving coordinated motion. When applied to cooperative drift, the swarm principle must be extended to maintain spatial alignment and speed coordination under extreme control conditions. However, traditional reinforcement learning-based methods struggle to cope with the complexity of cooperative drift, exhibiting problems such as agent heterogeneity, poor adaptability to dynamic environments, scalability limitations, and insufficient stability of group cooperation. Therefore, a new method is needed to address the transient nonlinearities and unpredictable uncertainties in drift operations. Summary of the Invention
[0004] In view of the above problems, this invention proposes a multi-vehicle cooperative drift control method based on embodied self-state perception of a drive-by-wire chassis. By introducing an embodied self-state perception (ESSA) network, each vehicle can infer its internal state in real time and dynamically adjust its strategy, thus forming an embodied multi-agent reinforcement learning (E-MARL) framework. The aim is to achieve cooperative drift control of multiple autonomous vehicles under extreme conditions.
[0005] This invention provides a multi-vehicle cooperative drift control method based on the self-state perception of an embodied drive-by-wire chassis, comprising: Step S1: Construct a graph model of the spatial relationships between vehicles participating in cooperative drifting; establish action guidance functions; Step S2: Determine the cooperative drift control task of the multi-vehicle cooperative system based on the graph model and action guidance function, and establish the corresponding cooperative drift control model of the multi-agent Markov decision process. Step S3, let t =1, whent =1 indicates the initial time. Step S4: Determine the time t The advantage value is input into the cooperative drift control model; Step S5, let i =1, when i When =1, it indicates the first vehicle; Step S6: Sensor module based on vehicle chassis A t,i Embodied Self-State Awareness Module B t,i , obtain the moment t No. i Vehicle status memory ; Step S7, based on time t No. i Vehicle status memory, obtaining time t No. i The probability distribution of the predicted vehicle state; Step S8: According to the policy network D t,i and time t No. i The probability distribution of the predicted vehicle state is obtained from the time step. t No. i The movement of the vehicle; Step S9, based on time t No. i The car's movements, obtaining the moment t No. i Multi-objective reward value of vehicles and policy network D t,i Update as of time t+ 1st i Vehicle strategy network; Step S10, Traversal I For each vehicle, repeat steps S6-S9 to obtain the time. t The multi-objective reward value for each vehicle further increases the time available. t The multi-objective joint reward value of the multi-agent system is determined, and the process proceeds to step S11. Step S11, based on time t Multi-agent system multi-objective joint reward value per time t The advantage value is updated, and the time is obtained. t Update the advantage value as a time. t +1 advantage value; Step S12, Judgment t Is it greater than or equal to? T , TIndicates the total number of moments. If so, complete the multi-vehicle cooperative drift control and obtain the final cooperative drift control model; otherwise, let... t = t +1, return to step S4; Step S13: Perform multi-vehicle cooperative drift control based on the final cooperative drift control model.
[0006] Optionally, step S1 includes the following specific steps: Step S101. Treat each vehicle participating in the drift as a vertex to obtain a vertex set; Define the location of each vehicle and obtain the set of spatially neighboring vehicles for each vehicle. A graph model is constructed based on the vertex set and the spatial neighbor vehicle set of each vehicle. Step S102. Determine the elliptical formation parameters; Design elliptical mesh formation constraints based on elliptical formation parameters; Step S103. Establish the action guidance function based on the elliptical mesh formation constraints.
[0007] Optionally, the expression for the action-guided function is:
[0008] in, For the first i Car and the j The distance of the car on the elliptical grid, the first i Car and the j The cars are adjacent to each other. For a smooth window function, For distance metric parameters, To adjust the parameter one of the function shape, To adjust the second parameter of the function shape, For the target distance, To adjust the function shape, parameter three, (.) is the guiding action function.
[0009] Optionally, step S6 includes the following specific steps: Sensor modules based on vehicle chassis B t,i Get the time t No. i Local observation information of the vehicle; Time t No. i Vehicle local observation information is input into the Embodied Self-State Awareness (ESSA) module. C t,i , obtain the moment t No. iKey information from partial vehicle observation Key information about vehicle status ; In the state memory module D t,i Based on time t No. i Key local information of the vehicle Key information about vehicle status , obtain the moment t No. i Vehicle status memory .
[0010] Optionally, the key local information is a dynamic state. The relative position with respect to the trajectory, the relative information with neighboring vehicles, and its own historical actions; The state memory refers to the vehicle at a given time. t The internal state representation is generated by integrating key local information and key information about the vehicle's state. State memory enables the vehicle system to have memory capabilities, allowing it to retain and recall historical information across multiple time steps, thus avoiding redundant decisions.
[0011] Optionally, time t No. i Vehicle status memory The expression is:
[0012] in, Indicates time t No. i Vehicle status memory, For a moment t No. i Vehicle strategy network, For a moment t No. i Key local information of the vehicle For a moment t No. i Key information about the vehicle's status. For a moment t -1st i Vehicle status memory.
[0013] Optionally, step S9 includes the following specific steps: Based on time t No. i The car's movements, obtaining the moment t No. i The multi-objective reward value of a vehicle; Based on time t No. iThe vehicle's multi-objective reward value and vehicle loss function, for time... t No. i The weight and bias parameters of the vehicle's policy network are updated using gradients to obtain the time step. t No. i The vehicle update strategy network, as a time-based... t+ 1st i Vehicle strategy network.
[0014] Optionally, step S11 specifically includes: based on time t The joint reward value of multi-objective systems in multi-agent systems, obtained at time points. t State value and state-action value in multi-intelligence systems; Based on time t The state value and state-action value of a multi-intelligence system are used to update the advantage value at the current moment to obtain the moment. t Update the advantage value as a time. t +1 advantage value.
[0015] Optionally, step 11 also includes time-based... t State-action value of multi-intelligence systems, at any given moment t Critics Network E t,i parameters Make adjustments to obtain the time. t Update the critics network as a time t +1 commentator network.
[0016] Optionally, the multi-target joint reward value includes drift reward, swarm motion reward, and collective navigation reward.
[0017] Optionally, the cooperative drift control model includes a sensor module for acquiring real-time vehicle status information; The embodied self-state awareness module is used to extract key information from local observations; The state memory module is used to generate dynamic contexts; Critics network, used to compute state value function; A policy network is used to generate vehicle action commands. A multi-objective reward function module is used to evaluate the performance of cooperative drift control.
[0018] Compared with the prior art, the present invention has at least the following beneficial effects: (1) This invention realizes for the first time the cooperative formation control of multiple autonomous vehicles in drifting conditions, breaking through the limitation of existing technologies that can only handle single-vehicle drifting or conventional cooperative driving, and extending the control boundary of autonomous driving to extreme scenarios. (2) By introducing an embodied self-state perception network module into each intelligent agent, the present invention enables it to implicitly infer key internal states and uncertainties from historical observation sequences, dynamically form a cognition of the vehicle dynamic context, thereby significantly enhancing the adaptive capability of the control system to transient and nonlinear dynamic characteristics. (3) This invention designs a multi-component reward function that integrates drift, swarm motion and collective navigation, and introduces an elliptical lattice formation topology to guide the agent to maintain relative position and heading consistency while maintaining large sideslip angle drift, and can effectively avoid collisions. (4) This invention establishes a system paradigm of “centralized training and decentralized execution”. After training, vehicles can carry out distributed self-organization and collaboration, realizing a complete closed loop from theoretical model to engineering implementation. Attached Figure Description
[0019] The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of the invention.
[0020] Figure 1 This is a schematic diagram of the geometric shape of the convoy formation in an embodiment of the present invention.
[0021] Figure 2 In this embodiment of the invention, the multilayer perceptron (MLP) network is used to implement the central commentator and distributed... Figure 3 This is a schematic diagram of the internal network module of the E-MARL algorithm in an embodiment of the present invention; Figure 4 This is a schematic diagram in an embodiment of the present invention. Detailed Implementation
[0022] To better understand the above-described objectives, features, and advantages of the present invention, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments of the present invention and the features thereof can be combined with each other. Furthermore, the present invention can be implemented in other ways different from those described herein; therefore, the scope of protection of the present invention is not limited to the specific embodiments disclosed below.
[0023] A specific embodiment of the present invention, such as Figure 1-4 A multi-vehicle cooperative drift control method based on the self-state perception of an embodied drive-by-wire chassis is disclosed, including: Step S1: Construct a graph model of the spatial relationships between the vehicles participating in cooperative drift; Establish an action guidance function; It is understandable that the action guidance function is used to guide drifting vehicles to form a low-energy lattice structure, and is also used as an important basis for designing complex reward functions.
[0024] Optionally, step S1 includes the following specific steps: Step S101. Treat each vehicle participating in the drift as a vertex to obtain a vertex set; Define the location of each vehicle and obtain the set of spatially neighboring vehicles for each vehicle, expressed as:
[0025] in, Represents a specific distance norm; Represents the set of vehicle vertices; Indicates the first i The location of the car , Let m be a real number, and m=2 represent the two-dimensional spatial coordinates, that is, the position of the vehicle in the plane. Indicates the first j The location of the car Indicates the nearest distance threshold. For the first i The set of neighboring vehicles of a vehicle, representing the vehicle with respect to the first vehicle. i The set of all vehicles whose distance to each other is within a threshold range is the set of autonomous vehicles participating in the cooperative drift control task.
[0026] A graph model is constructed based on the vertex set and the spatial neighbor vehicle set of each vehicle. Optionally, the graph model is an acyclic, undirected graph. ; Step S102. Determine the elliptical formation parameters; Based on the parameters of the elliptical formation, the elliptical mesh formation constraint is designed, and the expression is:
[0027] in, Indicates the first j The location of the car The length of the semi-major axis of the elliptical grid; For the first i The first focal point of the elliptical formation of vehicles. For the first i The second focus of the elliptical formation of vehicles can be based on the first i The vehicle's coordinates and velocity direction, and Perform calculations. The length of the semi-minor axis of the elliptical grid. Indicates the first j The vehicle is the first i The vehicle is any vehicle in the set of vehicles adjacent to it.
[0028] Step S103. Establish the action guidance function based on the elliptical mesh formation constraints.
[0029] Optionally, the expression for the action-guided function is:
[0030] in, For the first i Car and the j The distance of the car on the elliptical grid, the first i Car and the j The cars are adjacent to each other. For a smooth window function, For distance metric parameters, To adjust the parameter one of the function shape, To adjust the second parameter of the function shape, For the target distance, To adjust the function shape, parameter three, (.) is the guiding action function.
[0031] Optionally, the drift is a sideslip through a corner.
[0032] Step S2: Determine the cooperative drift control task of the multi-vehicle cooperative system based on the graph model and action guidance function, and establish the corresponding multi-agent Markov decision process. The cooperative drift control model; in, Indicates participation in the cooperative drift control task I A collection of autonomous vehicles. ={1,..., n}, Indicates the first i The vehicle's operating space ∈ ; Representing the state space, Indicates joint rewards, ; This indicates the initial position and status of the convoy at the start; To coordinate rules for the fleet ; Indicates the discount factor. ∈[0,1).
[0033] Optionally, the cooperative drift control model includes a sensor module for acquiring real-time vehicle status information; The embodied self-state awareness module is used to extract key information from local observations; The state memory module is used to generate dynamic contexts; Critics network, used to compute state value function; A policy network is used to generate vehicle action commands. A multi-objective reward function module is used to evaluate the performance of cooperative drift control.
[0034] The global state includes the position, speed, and relative distance of each vehicle; This invention is based on the multi-agent Markov decision process model in step S2. It adopts a centralized training and decentralized execution framework to implement multi-agent reinforcement learning, so that the cooperative drift control enables multiple vehicles to maintain formation while achieving cooperative drift along a given trajectory under the control limits and avoiding collisions.
[0035] The global state includes local observations of all vehicles; Local observations are mapped to corresponding control actions, consisting of lateral control commands and longitudinal control commands. The local observations include their dynamic states. The relative position of the trajectory, the relative information of neighboring vehicles, and its own historical actions.
[0036] in, Represents the velocity vector. For lateral velocity, Indicates the sideslip angle. yaw rate, It is the lateral angle.
[0037] Step S3, let t =1, when t =1 indicates the initial time. Step S4: Determine the time t The advantage value is input into the cooperative drift control model. A t,i ; Step S5, let i =1, when i When =1, it indicates the first vehicle; Step S6: Sensor module based on vehicle chassis B t,i Get the time t No. i Local observation information of a vehicle; the local observation includes dynamic status, relative position with respect to the trajectory, relative information with neighboring vehicles, and its own historical actions; In this invention, both state observation and control input are based on the vehicle chassis, including the state perception of a single vehicle, vehicle position, vehicle speed, center of gravity sideslip angle, etc., as well as the relative positions between different vehicles. The input control quantity also acts on the chassis, specifically manifested as the input command ultimately acting as the command of the controller on the chassis.
[0038] Time t No. iVehicle local observation information is input into the Embodied Self-State Awareness (ESSA) module. C t,i , obtain the moment t No. i Key information from partial vehicle observation Key information about vehicle status ; In the state memory module D t,i Based on time t No. i Key local information of the vehicle Key information about vehicle status Generation time t No. i The dynamic context of the vehicle; Based on time t No. i The dynamic context of the vehicle, obtaining the moment t No. i Vehicle status memory ; Optionally, the key local information is a dynamic state. The relative position with respect to the trajectory, the relative information with neighboring vehicles, and its own historical actions; The state memory refers to the vehicle at a given time. t The internal state representation is generated by integrating key local information and key information about the vehicle's state. State memory enables the vehicle system to have memory capabilities, allowing it to retain and recall historical information across multiple time steps, thus avoiding redundant decisions.
[0039] Optionally, time t No. i Vehicle status memory The expression is:
[0040] in, Indicates time t No. i Vehicle status memory, For a moment t No. i Vehicle strategy network, For a moment t No. i Key local information of the vehicle For a moment t No. i Key information about the vehicle's status. For a moment t -1st i Vehicle status memory.
[0041] Step S7, according to time t Local observation information and time t No. i Vehicle status memory, obtaining time t No. i The probability distribution of the predicted vehicle state is expressed as:
[0042] in, For a moment t No. i The probability distribution of the predicted vehicle state. For the first i The encoder function for the vehicle.
[0043] Based on time t No. i Vehicle status memory, time acquisition t No. i The vehicle's predicted local observation information, as time... t +1th i The local observation information of the vehicle is expressed as:
[0044] in, For a moment t No. i Predictive local observation information of vehicles, This is the decoder function.
[0045] Step S8, Set the time t No. i The probability distribution of the predicted vehicle state of the vehicle input number 1 i Vehicle strategy network E t,i Output time t No. i The vehicle's movements include lateral control commands for steering angle and longitudinal control commands for acceleration; Step S9, based on time t No. i The car's movements, obtaining the moment t No. i The multi-objective reward value of a vehicle; Based on time t No. i The vehicle's multi-objective reward value and vehicle loss function, for time... t No. i The weight and bias parameters of the vehicle's policy network are updated using gradients to obtain the time step. t No. iThe vehicle update strategy network, as a time-based... t+ 1st i Vehicle strategy network; Step S10, Traversal I For each vehicle, repeat steps S6-S9 to obtain the time. t The actions and timing of each vehicle t No. i Vehicle update strategy network and time t The multi-objective reward value for each vehicle further increases the time available. t The multi-objective joint reward value of the multi-agent system is determined, and the process proceeds to step S11. Step S11, based on time t The joint reward value of multi-objective systems in multi-agent systems, obtained at time points. t State value and state-action value in multi-intelligence systems; Based on time t The state value and state-action value of a multi-intelligence system are used to update the advantage value at the current moment to obtain the moment. t Update the advantage value as a time. t +1 advantage value; Based on time t State-action value of multi-intelligence systems, at any given moment t Critics Network E t,i parameters Make adjustments to obtain the time. t Update the critics network as a time t +1 commentator network; Step S12, Judgment t Is it greater than or equal to? T , T Indicates the total number of moments. If so, complete the multi-vehicle cooperative drift control to obtain the final cooperative drift control model; otherwise, let... t = t +1, return to step S4. Step S13: Based on the final cooperative drift control model, obtain the actions of each vehicle at the current moment and perform multi-vehicle cooperative drift control. Optionally, the time... t The expression for the state value of a multi-intelligence system is:
[0046] in, In the policy network π, at time... t Multi-intelligent systems in global state s The state value represents the expected value of the cumulative future returns starting from this global state. Let be the mathematical expectation function. This represents the sequence of actions collected in the policy network π; To be based on the transition probability The obtained state sequence, As a discount factor, For the initial sequence, For a moment t The instant reward received.
[0047] Optionally, the state-action value expression for a multi-agent system is:
[0048] in, This represents the global state under the policy network π. s Execute action at time a Then, the expected value of the cumulative reward. Indicates the initial action. This indicates that the action sequence from time t=1 to the future is collected in the policy network π.
[0049] Alternatively, the expression for the dominance value is:
[0050] in, Represents the dominance value function. This indicates that, under the policy network π, the measurement of the global state... s The average performance of performing action a, i.e., the advantages or disadvantages of choosing action a.
[0051] Optionally, the action includes: lateral control commands and longitudinal control commands; Optionally, the observation information includes the distance to the vehicle in front, the distance to the vehicle behind, the vehicle's speed, the curvature of the curve, and the movement of other vehicles; The local observations include dynamic status, relative position to the trajectory, relative information to neighboring vehicles, and its own historical actions; The global state is a set of local observations of all vehicles, including the position, speed, direction, relative position, and neighbor information of each vehicle; Optionally, a central critic-decentralized policy architecture is constructed using a multilayer perceptron (MLP), comprising a critic network and a policy network; The critic network is an MLP critic network, which serves as a shared critic network for all vehicles. Each vehicle has an MLP policy network; Optionally, the multi-objective joint reward value includes drift reward, swarm motion reward and collective navigation reward, which jointly guide the behavior of multiple vehicles through a collaborative optimization mechanism.
[0052] Optionally, a drift reward is designed to encourage the vehicle to achieve and maintain a desired drift state, wherein the expression for the drift reward is:
[0053] in, Indicates drift bonus, Represents a constant. Represents weight factor one, Represents weight factor two, Represents the constant one. Represents a constant two. Represents the constant three. Represents the constant four. Indicates the maximum speed. Indicates vehicle speed.
[0054] Optionally, a cluster motion reward is designed to promote coordinated movement and heading consistency among participating vehicles, expressed as:
[0055] in, Indicates a reward for clustered movement. Indicates weighting factor one, This indicates a weighting factor of two. It is a natural exponential function. It is a constant of five. For the first j The vehicle's heading angle, For the first i The heading angle of the vehicle.
[0056] Optionally, a collective navigation reward can be designed to prioritize ensuring that fast-moving convoys can traverse the environment in a coordinated and efficient manner; The collective navigation reward expression is:
[0057] in, As a reward for collective navigation, This is the positional consistency weighting coefficient. The constant is six. Geometric center of multi-agent system x coordinate, Geometric center of multi-agent system y coordinate, Indicates the key point on the path that is closest to the center of the multi-intelligence system. x coordinate, Indicates the key point on the path that is closest to the center of the multi-intelligence system. y coordinate, Indicates the directional consistency weight. (.) denotes the inverse sine cofunction. Indicates the direction of travel for the target at the nearest key point on the path. It represents the sum of the velocity vectors of all agents, reflecting the actual direction of movement of the group.
[0058]
[0059]
[0060] in, (.) denotes the loss function of the Critic network. The parameters representing the Critic network, [.] indicates the expected value. In this invention, the Embodied Self-State Awareness (ESSA) module can implicitly infer key vehicle state parameters and uncertainties based on historical observation sequences, thereby enabling vehicles to dynamically adapt to changing operating conditions.
[0061] The parameters of the ESSA module are optimized using gradient descent through the standard backpropagation algorithm.
[0062] This invention constructs a dedicated reinforcement learning simulation environment that supports cooperative drift task execution in various scenarios. For continuous control problems, it introduces action smoothing filtering to suppress command jitter; normalizes state and action signals to enhance numerical stability and convergence performance; and finely adjusts the reward function scale to improve learning efficiency. The entire framework is as follows: Figure 3 As shown.
[0063] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A multi-vehicle cooperative drift control method based on embodiment of the self-state perception of the drive-by-wire chassis, characterized in that, include: Step S1: Construct a graph model of the spatial relationships between the vehicles participating in cooperative drift; Establish an action guidance function; Step S2: Determine the cooperative drift control task of the multi-vehicle cooperative system based on the graph model and action guidance function, and establish the corresponding cooperative drift control model of the multi-agent Markov decision process. Step S3, let t = 1, when t = 1 represents the initial moment; Step S4: Determine the time t The advantage value is input into the cooperative drift control model; Step S5, let i =1, when i When =1, it indicates the first vehicle; Step S6: Sensor module based on vehicle chassis A t,i Embodied Self-State Awareness Module B t,i , obtain the moment t No. i Vehicle status memory ; Step S7, based on time t No. i Vehicle status memory, obtaining time t No. i The probability distribution of the predicted vehicle state; Step S8: According to the policy network D t,i and time t No. i The probability distribution of the predicted vehicle state is obtained from the time step. t No. i The movement of the vehicle; Step S9, based on time t No. i The car's movements, obtaining the moment t No. i Multi-objective reward value of vehicles and policy network D t,i Update as of time t+ 1st i Vehicle strategy network; Step S10, Traversal I For each vehicle, repeat steps S6-S9 to obtain the time. t The multi-objective reward value for each vehicle is represented by time step. t The multi-objective reward value of a multi-agent system further yields the time-sharing results. t The multi-objective joint reward value of the multi-agent system is calculated, and the process proceeds to step S11. Step S11, based on time t Multi-agent system multi-objective joint reward value per time t The advantage value is updated, and the time is obtained. t Update the advantage value as a time. t +1 advantage value; Step S12, Judgment t Is it greater than or equal to? T , T Indicates the total number of moments. If so, complete the multi-vehicle cooperative drift control and obtain the final cooperative drift control model; otherwise, let... t = t +1, return to step S4; Step S13: Perform multi-vehicle cooperative drift control based on the final cooperative drift control model.
2. The multi-vehicle cooperative drift control method based on the self-state perception of an embodied drive-by-wire chassis according to claim 1, characterized in that, The specific steps of step S1 include: Step S101. Treat each vehicle participating in the drift as a vertex to obtain a vertex set; Define the location of each vehicle and obtain the set of spatially neighboring vehicles for each vehicle. A graph model is constructed based on the vertex set and the spatial neighbor vehicle set of each vehicle. Step S102. Determine the elliptical formation parameters; Design elliptical mesh formation constraints based on elliptical formation parameters; Step S103. Establish the action guidance function based on the elliptical mesh formation constraints.
3. The multi-vehicle cooperative drift control method based on the self-state perception of an embodied drive-by-wire chassis according to claim 2, characterized in that, The expression for the action-guided function is: in, For the first i Car and the j The distance of the car on the elliptical grid, the first i Car and the j The cars are adjacent to each other. For a smooth window function, For distance metric parameters, To adjust the parameter one of the function shape, To adjust the second parameter of the function shape, For the target distance, To adjust the function shape, parameter three, (.) is the guiding action function.
4. The multi-vehicle cooperative drift control method based on the self-state perception of an embodied drive-by-wire chassis according to claim 1, characterized in that, The specific steps of step S6 include: Sensor modules based on vehicle chassis B t,i Get the time t No. i Local observation information of the vehicle; Time t No. i Vehicle local observation information is input into the Embodied Self-State Awareness (ESSA) module. C t,i , obtain the moment t No. i Key information from partial vehicle observation Key information about vehicle status ; In the state memory module D t,i Based on time t No. i Key local information of the vehicle Key information about vehicle status , obtain the moment t No. i Vehicle status memory .
5. The multi-vehicle cooperative drift control method based on the self-state perception of an embodied drive-by-wire chassis according to claim 4, characterized in that, time t No. i Vehicle status memory The expression is: in, Indicates time t No. i Vehicle status memory, For a moment t No. i Vehicle strategy network, For a moment t No. i Key local information of the vehicle For a moment t No. i Key information about the vehicle's status. For a moment t -1st i Vehicle status memory.
6. The multi-vehicle cooperative drift control method based on the self-state perception of an embodied drive-by-wire chassis according to claim 1, characterized in that, The specific steps of step S9 include: Based on time t No. i The car's movements, obtaining the moment t No. i Multi-objective reward value of a vehicle; Based on time t No. i The vehicle's multi-objective reward value and vehicle loss function, for time... t No. i The weight and bias parameters of the vehicle's policy network are updated using gradients to obtain the time step. t No. i The vehicle update strategy network, as a time-based... t+ 1st i Vehicle strategy network.
7. The multi-vehicle cooperative drift control method based on the self-state perception of an embodied drive-by-wire chassis according to claim 1, characterized in that, Step S11 specifically includes: based on time t The joint reward value of multi-objective systems in multi-agent systems, obtained at time points. t State value and state-action value in multi-intelligence systems; Based on time t The state value and state-action value of a multi-intelligence system are used to update the advantage value at the current moment to obtain the moment. t Update the advantage value as a time. t +1 advantage value.
8. The multi-vehicle cooperative drift control method based on the self-state perception of an embodied drive-by-wire chassis according to claim 1, characterized in that, Step 11 also includes time-based... t State-action value of multi-intelligence systems, at any given moment t Critics Network E t,i parameters Make adjustments to obtain the time. t Update the critics network as a time t +1 commentator network.
9. The multi-vehicle cooperative drift control method based on the self-state perception of an embodied drive-by-wire chassis according to claim 1, characterized in that, The multi-objective joint reward value includes drift reward, swarm motion reward and collective navigation reward.
10. The multi-vehicle cooperative drift control method based on self-state perception of an embodied drive-by-wire chassis according to claim 1, wherein the cooperative drift control model includes a sensor module for acquiring real-time state information of the vehicles; The embodied self-state awareness module is used to extract key information from local observations; The state memory module is used to generate dynamic contexts; Critics network, used to compute state value function; A policy network is used to generate vehicle action commands. A multi-objective reward function module is used to evaluate the performance of cooperative drift control.