A double-layer game multi-agent cooperative control method for a distributed drive electric vehicle

CN122525958APending Publication Date: 2026-08-07JILIN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JILIN UNIVERSITY
Filing Date
2026-07-10
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0005]针对上述现有技术中多执行器非线性耦合冲突、传统物理模型在极限工况下容易失真、以及传统博弈与优化算法在线求解实时性差的技术痛点,本发明提供一种分布式驱动电动汽车双层博弈多智能体协同控制方法

Benefits of technology

该方法通过引入强化学习技术,构建“上层协作博弈分配、下层非合作协同执行”的双层博弈架构,不仅能够实现转向与制动系统在不同工况下的解耦与帕累托最优平衡,而且利用深度强化学习的离线训练与在线快速前向推断能力,彻底摆脱对高精度轮胎模型的依赖,将底盘协同决策延迟控制在毫秒级,显著提升极端工况下车辆的路径跟踪精度与失稳极限挽救能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122525958A_ABST
    Figure CN122525958A_ABST
Patent Text Reader

Abstract

A double-layer game multi-agent cooperative control method for distributed drive electric vehicles belongs to the technical field of vehicle control and regulation. It solves the technical problems of multi-actuator nonlinear coupling conflict in the prior art, distortion of the traditional physical model under extreme working conditions, and poor real-time performance of traditional game and optimization algorithms in online solution. The double-layer game architecture of "upper-layer cooperative game distribution and lower-layer non-cooperative cooperative execution" not only realizes the decoupling and Pareto optimal balance of the steering and braking system under different working conditions, but also uses the offline training and online fast forward inference capability of deep reinforcement learning to completely eliminate the dependence on high-precision tire models, delay the chassis cooperative decision-making control to the millisecond level, and significantly improve the path tracking accuracy and instability limit rescue capability of the vehicle under extreme working conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of vehicle control and regulation technology, specifically relating to a two-layer game-theoretic multi-agent cooperative control method for distributed drive electric vehicles. Background Technology

[0002] With the rapid development of drive-by-wire chassis and four-wheel independent drive / steering technology, distributed drive electric vehicles (DDEVs), due to their ability to independently and precisely control the torque and steering angle of each wheel, provide a broad platform for improving vehicle active safety and handling limits. Currently, the industry commonly uses an integrated approach of Active Front-wheel / Four-wheel Steering (AFS / 4WS) and Direct Yaw Moment Control (DYC) to improve vehicle lateral stability. However, in practical engineering applications and extreme extreme conditions (such as low-traction surfaces and emergency double lane change obstacle avoidance), existing technologies still face the following key technical challenges that urgently need to be addressed: 1. Coupling conflicts and ground adhesion competition among multiple actuators: 4WS (generating yaw torque through lateral tire force) and DYC (generating yaw torque through longitudinal wheel-end drive / braking difference) are physically highly coupled. Traditional control methods (such as centralized MPC or fuzzy control based on scalar weights) often use fixed weight ratios, leading to frequent control conflicts between steering and braking actuators when entering the instability edge, and even mutual cancellation in the nonlinear limit region, causing chassis vibration and adhesion overload.

[0003] 2. Constraints of Multi-Height Nonlinear Tire Models on Control Accuracy: Traditional collaborative methods based on Model Predictive Control (MPC) or Sliding Mode Control (SMC) rely heavily on simplified or linearized tire mechanical models. When a vehicle enters extreme conditions with large center of gravity sideslip angles and high lateral accelerations, the tire's sideslip force and slip ratio exhibit highly nonlinear saturation characteristics. At this point, traditional mathematical models become severely distorted, leading to controller failure or a sharp decline in stability.

[0004] 3. Real-time bottleneck of traditional game theory / crowd optimization algorithms for online solutions: While existing technologies have attempted to introduce non-cooperative game theory (such as solving the Nash equilibrium matrix) or particle swarm optimization (PSO) algorithms for dynamic weight allocation in multi-system systems, these methods require large-scale matrix iteration or crowd-driven evolution optimization online within each vehicle control cycle. The computational load increases exponentially with the increase in control dimensions, failing to meet the millisecond-level (typically ≤10ms) real-time control response requirements of vehicle ECUs / MDCs. Summary of the Invention

[0005] To address the technical pain points of the prior art, such as nonlinear coupling conflicts among multiple actuators, the tendency of traditional physical models to distort under extreme conditions, and the poor real-time performance of traditional game theory and optimization algorithms in online solutions, this invention provides a two-layer game theory multi-agent cooperative control method for distributed drive electric vehicles.

[0006] The method is specifically as follows: S1. State Perception and Representation: Constructing the state equation and observation equation of the nonlinear system based on tire slip ratio and physical adhesion limit, and assessing the road adhesion coefficient. Estimate; define the real-time boundary instability risk index. A meta-learning feature network is introduced to extract driver features, resulting in the final guidance state space. ; S2. Dynamic Allocation of Upper-Level Collaborative Weights: Constructing a cooperative game agent at the vehicle's upper-level control layer, and allocating weights through factors... The torque distribution between the lower-level active four-wheel steering control system and the direct yaw torque control system is regulated; a Pareto multi-objective joint reward function, which includes physical boundary penalties and kinetic energy losses, is designed to guide the motion control of the upper-level cooperative game agent. S3, Lower-level active four-wheel steering cooperative execution: Construct non-cooperative game-theoretic agents at the vehicle's bottom execution layer, including an additional front-wheel steering agent responsible for coordinating front axle steering and an active rear-wheel steering agent responsible for controlling rear axle steering; Design specialized reward functions for the two agents respectively to guide their motion control. S4. Physical Limit Constraint Interception and Anti-Saturation Control Allocation Execution: Before the lower-level intelligent agent network outputs actions to the physical hardware, a physical hard constraint boundary interception layer is set up; the macroscopic expectation additional yaw moment is solved and the underlying torque control allocation is performed; based on the underlying torque control allocation, the underlying hardware drive execution is performed.

[0007] Furthermore, the state equation of the nonlinear system is: ; The observation equations for the nonlinear system are: ; in, and Represents the discrete time step. Represents the system state vector. Represents the observed input vector. Represents the state transition function of a nonlinear system. Represents the observation mapping function of a nonlinear system. This refers to the physical quantity directly measured and input by the vehicle's onboard sensors. and These are the process noise covariance matrix and the measurement noise covariance matrix, respectively. road surface adhesion coefficient The estimation process involves generating a symmetric Sigma point set with weighted coefficients using an unscented transformation within each control cycle. This point set is then substituted into the nonlinear system state equation and the nonlinear system observation equation for nonlinear recursive propagation. The filter gain is updated in real time using the observed residuals, thereby enabling the estimation of the current road surface adhesion coefficient. Online estimates.

[0008] Furthermore, a real-time boundary instability risk index is defined. Specifically: Based on real-time estimated road adhesion coefficient With longitudinal speed The critical centroid sideslip angle boundary at which tire nonlinear slip occurs is calculated using a two-degree-of-freedom vehicle dynamics model. Boundary of angular velocity with center of mass deflection : ; ; For gravitational acceleration; establish a double asymptote nonlinear phase plane safety boundary, and vehicle state points. The algebraic transformation formula for the physical distance to the boundary is: ; Define the real-time boundary instability risk index The following segmentation mapping is satisfied: ; in, This characterizes the linear dynamic zone in which the vehicle is in an absolutely safe position. This indicates that the vehicle has fully reached or exceeded the nonlinear instability limit boundary.

[0009] Furthermore, guide the state space pass Obtain, among which, The task-representation context vector is obtained as follows: A meta-learning feature network is introduced to extract driver features, and a latent intent extraction network containing an encoder layer and a feature mapping layer is constructed. With sliding time window The serialized steering wheel input and path error within the dataset are used to construct the meta-task dataset. Task representation context vectors in the latent space are extracted through online scrolling using grid gradient optimization. : ; in, This indicates the current sampling time, representing the absolute time reference point for the vehicle controller to execute the current control loop. express The value range is 0 to , Indicates the steering wheel angle. Indicates the angular velocity of the steering wheel rotation. This represents the lateral distance tracking error between the vehicle's center of gravity and the expected planned path. This represents the yaw angle error between the vehicle's longitudinal centerline and the tangent direction of the desired path.

[0010] Furthermore, the upper-level cooperative game agent employs a flexible actor-commentator algorithm, in which the actor network guides the state space. Input, output motion space Limited to a continuous interval: ; Weighting factor As a time-varying resource allocation parameter, the allocation ratio between the lower-level active four-wheel steering control system and the direct yaw moment control system is defined when counteracting the total yaw moment demand of the vehicle.

[0011] Furthermore, the Pareto multi-objective joint reward function is as follows: ; in, The tracking error is the lateral distance between the vehicle's center of gravity and the desired planned path. The sideslip angle is the angle between the vehicle's center of gravity and its body. The yaw rate is angular velocity. The nominal weight is a positive coefficient. For the ideal reference centroid sideslip angle, For the ideal reference yaw rate, This is a penalty term for kinetic energy loss. This is a penalty function for exponentially dangerous cliff-like constraints based on the physical boundary of the phase plane. This is a negative feedback term for manipulation burden based on meta-intention.

[0012] Furthermore, the additional front-wheel steering agent and active rear-wheel steering agent will utilize the cooperative factors allocated by the upper layer. And the local observation states, pieced together from the local vehicle motion states, serve as the inputs for their respective network decisions. The additional front wheel steering agent outputs an action of additional front wheel steering angle. The active rear-wheel steering agent outputs the active rear-wheel steering angle. The additional front-wheel steering agent and the active rear-wheel steering agent adopt the Actor-Critic neural network structure corresponding to the multi-agent proximal strategy optimization algorithm.

[0013] Furthermore, a specialized reward function is added to the front wheel steering agent. for: ; in, , , and These are the specialization performance weighted penalty factors for the additional front wheel steering agent; Specialized reward function for active rear-wheel steering agent for: ; in, , , and These are the specialization performance weighted penalty factors for the active rear-wheel steering agent.

[0014] Furthermore, based on the wheel-end vertical force estimation model, the current vertical load on all four wheels is calculated in real time, according to the road surface adhesion coefficient. The physical adhesion limit friction circle of each tire is constructed. Combined with the nonlinear single-wheel tire brush model, the maximum side slip angle limit that each tire can withstand is derived. Combined with the vehicle motion state, the four-wheel independent steering angle limit is obtained. When the additional front wheel steering angle and the active rear wheel steering angle exceed the obtained corresponding four-wheel independent steering angle limit, the interceptor starts the cut-off correction.

[0015] Furthermore, the macroscopic expectation adds a yaw moment through: get, in, To ultimately issue the desired additional yaw moment target command to the underlying torque distribution actuator; and These are the real-time yaw rate of the actual vehicle and the ideal reference yaw rate, respectively. and These are the real-time yaw acceleration of the actual vehicle and the ideal reference yaw acceleration, respectively. and These are the proportional gain coefficient and the differential gain coefficient, respectively. The underlying torque control allocation is as follows: Establish the control distribution matrix for the chassis overdrive system, and apply the desired additional yaw moment. Total driving force with driver's longitudinal expectation As a control target allocation vector The longitudinal force output vector of the four-wheel independent hub motor is defined as Construct a linear matrix relationship for control allocation: ; Among them, the effective control matrix Defined as: ; in, These are the front track and the rear track, respectively. These are the actual steering angles of the front wheels and the actual steering angles of the rear wheels, respectively.

[0016] The beneficial effects of the method described in this invention are as follows: This method introduces reinforcement learning technology to construct a two-layer game architecture of "upper-layer cooperative game allocation and lower-layer non-cooperative collaborative execution". It can not only achieve decoupling and Pareto optimal balance between steering and braking systems under different working conditions, but also completely get rid of the dependence on high-precision tire models by utilizing the offline training and online fast forward inference capabilities of deep reinforcement learning. It controls the chassis collaborative decision delay to the millisecond level, significantly improving the vehicle's path tracking accuracy and instability limit recovery capability under extreme working conditions. Attached Figure Description

[0017] Figure 1 The overall architecture and control logic flowchart of a two-layer game-theoretic multi-agent cooperative control method for a distributed drive electric vehicle provided in this embodiment of the invention; Figure 2 This is a flowchart illustrating the process of constructing the state space in an embodiment of the present invention; Figure 3 This is the evolution topology of the multi-agent reinforcement learning (SAC-MAPPO) interactive network for "upper-level cooperative game - lower-level non-cooperative Nash equilibrium game" in steps S2 and S3 of this embodiment of the invention. Figure 4 This is a topology diagram of the physical friction circle hard interception and trimming and the underlying torque control distribution closed-loop structure considering nonlinear Anti-Windup constraints described in steps S4-1 and S4-3 of this embodiment of the invention. Detailed Implementation

[0018] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.

[0019] Example 1 This embodiment provides a two-layer game-theoretic multi-agent cooperative control method for distributed drive electric vehicles. The overall flow of the method is as follows: Figure 1 As shown, the method includes the following steps: S1: State perception, physical risk assessment, and driver intent meta-representation To achieve precise game-theoretic control based on vehicle state and maneuvering intent, a perception layer with multi-source information fusion and deep physical cognition capabilities must first be established. Step S1 constructs the global state matrix upon which the control method depends by synchronously demodulating the underlying bus data and inertial signals. Based on this, to quantitatively deconstruct the vehicle dynamics limits and complex human-machine-vehicle dynamic game characteristics, the method further branches into steps S1-1 to S1-3, deeply extracting and reconstructing the strongly guided motion features required for control from three dimensions: multi-dimensional space construction, phase plane physical evolution, and meta-learning-based nonlinear behavioral representation.

[0020] S1-1: Construction of Global State Space and Guiding State Space Signals from the four-wheel wheel-end encoders, the six-axis inertial measurement unit (IMU), and the onboard high-precision positioning and GPS / INS integrated navigation system are acquired via the vehicle bus (CAN / CAN FD). The vehicle's dynamic state is collected in real time to construct an initial global state space vector. : ; in, The lateral distance tracking error between the vehicle's center of gravity and the expected planned path; The yaw angle error is the difference between the vehicle's longitudinal centerline and the tangent direction of the desired path. The longitudinal speed of the vehicle body; The sideslip angle is the angle between the vehicle's center of gravity and its body. This refers to the yaw rate; and These are the first derivatives of the sideslip angular velocity and the yaw angular velocity, respectively. The road adhesion coefficient is estimated online based on unscented Kalman filtering (UKF). For the driver's steering wheel angle, This refers to the angular velocity of the steering wheel rotation.

[0021] The specific implementation method of the online estimation of the unscented Kalman filter (UKF) is as follows: A state estimator is constructed based on a three-degree-of-freedom (longitudinal, lateral, and yaw) vehicle dynamics model and a two-parameter nonlinear tire model; the vehicle's longitudinal velocity, lateral velocity, yaw rate, and the road adhesion coefficients of the four wheels are selected as the system state vectors. ; Wheel speed signals acquired from four-wheel wheel-end encoders, longitudinal / lateral accelerations acquired from a six-axis inertial measurement unit, and steering wheel angle are used as the observation input vectors for the filter. ; Establish the state equation and observation equation of the nonlinear system based on tire slip ratio and physical adhesion limit: ; ; In the formula, and Represents the discrete time step (number of control cycle steps). This refers to the physical quantity directly measured and input by the vehicle's onboard sensors, namely the steering wheel angle. and four-wheel drive / braking torque ; This represents the state transition function of a nonlinear system (i.e., a function that predicts the current vehicle motion state using the vehicle dynamics equations based on the vehicle state at the previous moment and the current drive / steering input). This represents the observation mapping function of a nonlinear system (i.e., the function that establishes the physical and geometric relationship between the internal state variables of the vehicle and the measurements of external sensors such as wheel speed and accelerometer).

[0022] and These are the process noise covariance matrix and the measurement noise covariance matrix, respectively.

[0023] Within each control cycle, an unscented transformation is used to generate a symmetric Sigma point set with weighted coefficients. This point set is then substituted into the nonlinear state equation and observation equation for nonlinear recursive propagation. The filter gain is updated in real time using the observed residuals, thereby achieving the adjustment of the current road surface adhesion coefficient. Millisecond-level high-fidelity online estimation.

[0024] The Sigma point set is not generated for a single variable, but for the entire state vector. (Including vehicle speed, yaw rate and road surface adhesion coefficient) A set of physical state combinations generated by deterministic sampling around the mean and covariance of its current probability distribution.

[0025] By substituting this set of simulation points (Sigma points) with the assumption of road adhesion coefficient into the brush model and dynamic function and In the process, the nominal measured value is calculated and then compared with the actual longitudinal / lateral acceleration and wheel speed measured by the sensor. Only then can the most accurate road adhesion coefficient be corrected in reverse. .

[0026] Within each control cycle, an unscented transformation is used to generate a symmetric Sigma point set with weighted coefficients. The specific association mechanism is as follows: using the current state vector... The posterior mean is centered, and combined with the square root of its covariance matrix, a set of Sigma state points is generated by symmetrical sampling in the multidimensional state space to represent the potential distribution of the current vehicle speed, yaw rate and nominal road adhesion coefficient. Each combination of states in the Sigma state set is substituted into the nonlinear state transition function in parallel. Perform nonlinear time-domain propagation and map the results through a nonlinear observation function. Demodulate the corresponding theoretical nominal observation set; finally, use the actual observation input vector measured by the vehicle-mounted sensors. By comparing the residuals with the theoretical nominal observation set and performing dynamic gain correction using Kalman gain, the adhesion coefficient of the current driving road surface can be determined. High-fidelity online estimation in milliseconds under nonlinear extreme conditions.

[0027] The state vector constructed by the unscented Kalman filter (UKF) Observation input vector The global state space vector constructed in step S1-1 It is not a one-to-one correspondence.

[0028] The optimal estimate calculated by the UKF (United Kingdom Functions) is used as basic metadata and input to the upper layer to be physically synthesized with other sensor signals to form a global state space vector. When upper-level reinforcement learning (SAC / MAPPO) makes game weight decisions and action outputs, it needs to know not only the vehicle's current absolute physical state (such as longitudinal speed) but also the vehicle's current speed. centroid side slip angle yaw rate Furthermore, it is necessary to know the relative spatial geometric error of the vehicle relative to the desired target path (such as lateral position error). Heading angle error ) and the driver's transient maneuvering dynamics (such as In terms of topology, the state estimator, as the underlying core data sensing source of the chassis, utilizes a nonlinear physical transmission mechanism to online rolling demodulate and obtain high-precision longitudinal vehicle speed. yaw rate , centroid side slip angle Implicit mapping fundamental quantities and four-wheel time-varying road adhesion coefficient Subsequently, the optimal physical state estimation vector and the temporal-spatial relative geometric error captured by the onboard visual perception circuit are... Steering wheel angle sensor collects driving dynamics In the time domain, high-frequency alignment and cross-confluence are used to jointly and physically synthesize the global amplified guidance state space vector required for the upper-level decision-making. This allows for the rolling evolution and updating of the upper and lower layer game agents through a data closed loop.

[0029] S1-2: Phase plane physical boundary risk quantification Based on real-time estimated road adhesion coefficient With longitudinal speed The critical centroid sideslip angle boundary at which tire nonlinear slip occurs is calculated using a two-degree-of-freedom vehicle dynamics model. Boundary of angular velocity with center of mass deflection : ; ; Establish a double asymptote nonlinear phase plane safety boundary. Vehicle state points. The algebraic transformation formula for the physical distance to the boundary is: ; Define the real-time boundary instability risk index The following segmentation mapping is satisfied: ; in, This characterizes the linear dynamic zone in which the vehicle is in an absolutely safe position. This indicates that the vehicle has fully reached or exceeded the nonlinear instability limit boundary.

[0030] S1-3: Meta-Learning Feature Representation of Driver's Control Intent A meta-learning feature network is introduced to extract driver features. A latent intent extraction network is constructed, comprising an encoder layer and a feature mapping layer. Using a sliding time window The serialized steering wheel input and path error within the dataset are used to construct the meta-task dataset. Task representation context vectors in the latent space are extracted through online scrolling using grid gradient optimization. : ; Context vector The numerical value represents the driver's current driving behavior (aggressive, steady, or conservative). The meta-learning feature network updates multi-task parameters through a meta-training loss function; Will and The states are concatenated and merged into the initial global state space to form the final bootstrapping state space: ; The process of constructing the guiding state space is as follows: Figure 2 As shown.

[0031] Meaning: Current sampling time This represents the absolute time reference point from which the vehicle controller executes the current control cycle, i.e., the "now." It is the time axis sampling index (step offset) of the time series data, representing the time from the current moment. How many sampling periods were "retrospected" to the past historical time? In the formula... The value range is 0 to The entire expression represents the number of times the expression has been passed from the past... From the start of the current moment, records will continue continuously until the current moment. A set of discrete time series at each point in time. Represents a The real space of dimension 1 Overall meaning: It is a collection real number elements A column vector. This... It is determined by the number of output neurons (feature dimension) of the last layer of the designed meta-learning neural network.

[0032] S2: Dynamic allocation mechanism of upper-level collaborative weights based on cooperative game theory In step S1, the expanded guidance state space, which includes physical risks and human-machine intentions, was completed. After solving, the high-dimensional feature flow is input as the core data driving source into the macroscopic decision-making layer of the system (hardware entity system: including chassis domain controller (domain manager / MDC), execution-level ECU (such as steer-by-wire controller, hub motor inverter MCU), and vehicle communication bus (CANFD / Automotive Ethernet)). Step S2, as the "brain" of the entire chassis control architecture, has the core task of handling the macroscopic hedging of global resource allocation for the multi-actuator system (steering and braking). Since the control boundary and physical mechanism of the system undergo drastic changes when the vehicle evolves from the linear region to the extreme nonlinear region, the method in the following steps S2-1 and S2-2 uses a maximum entropy reinforcement learning strategy to map the complex time-varying vehicle state online into a continuous action slip factor, and designs a joint reward function by introducing a Pareto optimal mechanism to fundamentally solve the global optimization problem of multi-objective conflict in chassis control.

[0033] The upper layer refers to the macro-level collaborative decision-making and resource allocation layer (or central decision-making level controller) in the Chassis Domain Controller (CDC).

[0034] In the vehicle hierarchical control architecture, chassis control is vertically decomposed into different levels: Nominal layer (human-machine interaction input): responsible for receiving the original throttle, brake, and steering wheel signals from the driver (or the autonomous driving planning layer).

[0035] The upper layer (central decision-making layer, referred to as the "upper layer" in this invention): It doesn't directly control the movement of a single motor, but rather acts as a "commander-in-chief" from the perspective of the overall vehicle dynamics (longitudinal, lateral, and yaw). It uses the SAC algorithm to determine which motor should contribute more or less power based on the current vehicle stability and the driver's intentions. In this invention, it specifically outputs a weighting allocation factor. It is used to macroscopically coordinate the proportion of ground adhesion resources used by steering (4WS) and braking (DYC).

[0036] The lower layer (micro execution layer) includes AFS (front wheel steering), ARS (rear wheel steering), and the in-wheel motor controllers (MCUs) for each wheel. These controllers are responsible for converting the macro-level objectives from the upper layer into specific actuation instructions for each piece of hardware.

[0037] S2-1: Construction of Macro-level Collaborative Decision-making Network and Continuous Action Sliding The upper-level cooperative game-theoretic agent, Agent-Upper, employs a maximum entropy deep reinforcement learning network structure—the Soft Actor-Critic (SAC) algorithm. The algorithm's actor network guides the state space. As input, the output action space is limited to a continuous interval: ; The weight allocation factor assigned by the upper layer As a time-varying resource allocation parameter, the allocation ratio between the lower-level active four-wheel steering control (4WS) and direct yaw moment control (DYC) systems in counteracting the vehicle's total yaw moment demand is defined.

[0038] Output motion space This does not refer to the steering wheel's turning angle or the physical deflection space of the tires, but rather to the macroscopic control decision-making action space of the chassis domain controller (central decision-making level). Physically, it is the virtual action by which the central decision-making algorithm macroscopically adjusts the chassis control tendency. Because its output is strictly limited to a continuous range... Inside, this action space actually represents the continuous and seamless sliding space of the two core control functions of the chassis (4WS lateral force control and DYC yaw moment control) relative to the total ground adhesion resources.

[0039] By introducing policy entropy into the objective function To prevent the strategy from prematurely converging to a local optimum, its soft-state value objective function is: ; In the formula, For temperature adaptive adjustment parameters, This represents the maximum simulation or control time step of a single game reinforcement learning training scenario. The mathematical expectation operator represents the state variable. and control actions In the current policy network Guided state-action trajectory distribution topology Generated from sampling, used for statistical averaging of multiple random operating conditions; Indicates in Always according to the state and perform actions The Pareto multi-objective joint reward function value calculated in real time. Indicates the current environmental state Under certain conditions, the Shannon Entropy value of the probability distribution of the output action of the policy network is used in engineering as a mathematical penalty to encourage the algorithm to spontaneously maintain the diversity of control actions.

[0040] Upper-level cooperative game intelligence agent (Agent-Upper): Specific responsibilities: The core task of Agent-Upper is to determine the current physical risk status of the entire vehicle (phase plane instability risk index). ) and the driver's intention to operate (meta-learned feature vector) This approach balances the conflicting control objectives of path tracking accuracy and vehicle lateral stability at a global level. It doesn't require calculating specific steering angles or motor torques; instead, it dynamically and adaptively solves for the macroscopic allocation of chassis resources between the steering and braking systems (i.e., the weighting factor). ).

[0041] On which specific control circuit of the vehicle is it mounted? In the vehicle's physical hardware architecture, the Agent-Upper algorithm serves as the core software control logic, which is mounted on the main central processing unit (such as a SoC chip or a high-performance automotive-grade MCU) of the Chassis Domain Controller (CDC).

[0042] Physical wiring connection method: It receives status messages sent from the data perception layer (sensors and integrated navigation) as algorithm input through the automotive Ethernet or high-performance CAN FD bus line; at the same time, the calculated control commands are sent unidirectionally and in parallel to the bottom chassis steering controller (steering ECU / MDC) line and brake / wheel hub motor controller (drive MCU) line through the same chassis communication bus line, thereby commanding the lower-level actuators to work together.

[0043] S2-2: Optimization Design of Multi-Objective Pareto Joint Reward Function To guide the agent to spontaneously weigh "precise path tracking" against "extreme instability control," a Pareto multi-objective joint reward function is designed, which includes physical boundary penalties and kinetic energy losses. : ; in, The nominal weight is a positive coefficient; the nominal reference state and ( For the ideal reference centroid sideslip angle, (Assuming an ideal reference yaw rate) is obtained from a real-time solution using an ideal two-degree-of-freedom linear rigid body model of the vehicle: ;

[0044] In the formula, This refers to the steering gear ratio; The distance between the front and rear axles of the center of gravity; For the overall vehicle weight; This refers to the total wheelbase; This represents the linear lateral stiffness of the front and rear wheels, defaulting to negative values. Kinetic energy loss penalty term. It is defined as the algebraic sum of power losses due to single-wheel differential braking: ; in, For the corresponding angular velocity of the wheel, This is a braking state indication function. This means that the optimization requires a global summation of the motor power consumption of the four wheels: front left, front right, rear left, and rear right. Indicates the first The real-time electromagnetic drive / braking torque of the hub motor of each wheel (i.e., the corresponding wheel end), when and When traversing the set, there are 4 possible combinations, which perfectly correspond to the four independent wheels of a distributed drive vehicle: Front-left wheel Front right wheel Left rear wheel Right rear wheel; This is an exponentially dangerous cliff-like constraint penalty function based on the phase plane physical boundary, used to achieve a smooth transition from soft constraints to hard control: ; in, This is the base control coefficient. This is a sensitivity index. When the vehicle's risk factor changes abruptly ( When this happens, the penalty expands rapidly, forcibly lowering the output factor. This completely transfers control priority to the DYC system. The formula for the negative feedback term of the manipulation burden based on meta-intention is: ; The preset sensitivity weighting coefficient for maneuvering intent represents the weighted penalty coefficient for driving style's tendency to intervene in control strategies.

[0045] In the flexible actor-commentator (SAC) framework constructed in this invention, the multi-objective joint reward function The parsed values ​​implicitly depend on the current environment's boot state. With actors' online output of actions .

[0046] The specific mapping mechanism is as follows: state variables , , , , These are respectively the system states at the current moment. The intuitive physical representation or state evolution derived value constitutes the state constraint domain in the reward function for the vehicle's trajectory tracking capability and nonlinear stability boundary; while the power loss term... This deeply couples the current output action of the actor network. (i.e., weighting factor) ),action The slip dynamic directly reconstructs the torque distribution command of the underlying Direct Yaw Torque Control (DYC) system to the four independent hub motors, thereby changing the transient slip energy consumption and electromagnetic loss of the four wheels.

[0047] In summary, this formula, by comprehensively weighting state constraints and action power consumption, achieves a theoretical level of reinforcement learning. The quantitative transformation of functions into underlying physical parameter equations for vehicle dynamics.

[0048] S3: 4WS Multi-Agent Cooperative Execution Based on Lower-Level Non-Cooperative Nash Equilibrium Game Macroscopic time-varying resource allocation parameters demodulated in step S2 The macroscopic game boundary between the Active Four-Wheel Steering (4WS) control system and the Direct Yaw Conduction (DYC) control system was successfully established. However, since the 4WS contains two steering mechanisms that interfere with each other at the physical level (front and rear axles), direct centralized actuation can easily lead to control signal cancellation or chassis vibration due to the nonlinear coupling of the lateral forces of the front and rear tires. Therefore, step S3 builds upon the parameter tuning results from the upper layer and constructs a multi-agent reinforcement learning (MARL) architecture based on non-cooperative game theory at the lower execution layer. Through the following steps S3-1 to S3-3, the algorithm solves the Nash equilibrium between the two agents, AFS and ARS, which are in a state of local conflict and global cooperation, achieving millisecond-level non-cooperative coordination inference at the microscopic level.

[0049] S3-1: Establishment of the Non-Cooperative Game Structure Domain A multi-agent reinforcement learning (MARL) mechanism is introduced into the underlying execution layer. The 4WS control problem is deconstructed into a two-agent non-cooperative game: an additional front-wheel steering agent (Agent AFS) responsible for coordinating front axle steering, and an active rear-wheel steering agent (Agent ARS) responsible for controlling rear axle steering. The two agents will use the cooperative factors allocated by the upper layer. And the local observation states, pieced together from the local vehicle motion states, serve as the inputs for their respective network decisions. .

[0050] Both agents employ the Actor-Critic neural network structure corresponding to the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm, and the network parameters of the two agents are independent and specific. Furthermore, to support the rolling solution of the Nash equilibrium in the distributed non-cooperative game, the front axle additional steering agent Agent-AFS and the active rear wheel steering agent Agent-ARS do not use the flexible actor-critic (SAC) network of the single-agent paradigm. Instead, based on the MAPPO algorithm framework, two sets of actor deep feedforward neural network structures with completely decoupled parameters and physically specific mappings are constructed: the actor policy network of Agent-AFS... It is constructed using a multilayer perceptron (MLP), whose input layer seamlessly maps the local observation states. Nonlinear deep extraction of state features is performed through a fully connected layer, and the output layer is mapped to a one-dimensional continuous scalar through a saturated activation function, which is the expected additional front wheel steering angle. The Agent-ARS actor policy network Homogeneous but parameter-independent multilayer perceptrons are constructed, with the input layer mapping the local observation states. The output layer directly maps the output to a one-dimensional continuous scalar, which is the desired active rear wheel steering angle. During the offline centralized training phase, the system additionally constructs a global centralized critic network in parallel to break the non-stationary deadlock in the multi-agent environment and assist in gradient backpropagation. During the in-vehicle online inference phase, the centralized critic network is forcibly removed, and only the forward matrix multiplication calculation is performed by the independent lightweight actor policy networks of the front and rear axles. This enables millisecond-level Nash equilibrium cooperative action output of heterogeneous steering actuators without encroaching on the computing resources of the chassis domain controller (CDC).

[0051] The additional front-wheel steering agent Agent-AFS and the active rear-wheel steering agent Agent-ARS, as micro-level collaborative control algorithm units, are distributed and integrated into the control circuitry of the vehicle chassis drive-by-wire hardware system. Specifically, the Agent-AFS control program is deployed in the front-wheel control core of the front-wheel steering controller (SBW-ECU) or chassis domain controller, and its output is connected to the front axle steering actuator signal via a motor drive control circuit. The Agent-ARS control program is deployed in the rear-wheel active steering controller (RWS-ECU), and its output is physically connected to the rear axle steering actuator via hardware drive circuitry.

[0052] The Agent-AFS and Agent-ARS communicate with each other via a chassis bus communication line or a dual-core communication (IPC) control line within the domain controller chip, enabling high-frequency, millisecond-level asynchronous interaction and transmission of control actions and status observation messages. This allows for the non-cooperative game-theoretic Nash equilibrium demodulation of the underlying 4WS control resources within a distributed hardware topology.

[0053] S3-2: Definition of Independent Actor Conflict Reward Function To simulate the conflict and cooperation among the executors in a non-cooperative game, independent optimization objectives for both are designed: The Agent-AFS output action is to add a front wheel steering angle. Its specialized reward function It tends to ensure the accuracy of driver macroscopic trajectory tracking, and is subject to weighting. Positive incentives: ; - These four coefficients are the specific performance weighted penalty factors of Agent-AFS, used to internally reconcile the trade-off between tracking accuracy and maneuverability.

[0054] Agent-ARS outputs the active rear wheel steering angle. Its specialized reward function Defined as completely eliminating body roll under extreme conditions, damping term Dynamic gain: ; - These four coefficients are specific performance weighted penalty factors for Agent-ARS, used to internally reconcile the trade-off between tracking accuracy and maneuverability.

[0055] In actual vehicle control systems, and It is directly computed by the forward propagation of the Actor Network of Agent-AFS and Agent-ARS.

[0056] Specifically, the front axle active steering control circuit (carrying agent-AFS) and the rear axle active steering control circuit (carrying agent-ARS) sample the current state at high frequency. Through matrix multiplication and activation functions (such as the Tanh function) of their respective Actor networks, continuous scalar values ​​of these two angles are directly mapped to the chassis, and their dynamic landing formulas on the chassis are as follows: The actual physical total steering angle of the front wheels : ; in, The nominal steering angle of the wheels is calculated by the driver based on the steering wheel angle and the gear ratio. Rear wheel actual physical total steering angle : .

[0057] S3-3: Online Nash Equilibrium Solution and Policy Iteration Formula Based on MAPPO Algorithm The two agents are trained on the architecture based on the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm.

[0058] During the offline centralized training phase, a globally centralized critic network is introduced, with input including joint actions between two agents. The global state value function is Its advantage function The calculation uses the generalized advantage estimation (GAE) method: ; In the formula, , As a discount factor, The smoothing coefficient for GAE; represent The globally expanded guidance state space vector unique to the time-based chassis multi-agent system during the offline training phase; Represents intelligent agents At the current control sampling time The generalized advantage value; It represents the time span step (time domain recursion exponent) in the GAE (Generalized Advantage Estimation) algorithm for making multi-step look-ahead predictions into the future time domain. Indicates time Temporal difference residuals; Represents intelligent agents At the current control sampling time The immediate, specialized Pareto reward value obtained.

[0059] Parameters of the policy network (Actor) Update using the pruned near-end objective function: ; In the formula, the importance sampling ratio is , Trim the boundaries for the strategy. Represents intelligent agents ( The strategy optimizes the network parameters and loss value; This indicates the agent's performance during the current iteration cycle. The new strategy network parameter set (parameters to be updated); for The underlying layer shifts to intelligent agents at all times The wheel-end control action response, specifically for the front axle agent, refers to the additional front wheel steering angle. Specifically referring to the active rear wheel steering angle of the rear axle intelligent agent. ; For the local dynamic physical observation vector captured by the agent at high frequency; conditional probability symbol " "Used to isolate and establish the nonlinear mapping boundary from "local input observation features" to "continuous action output probability distribution" in the actor strategy network; This represents the time-step empirical expectation operator calculated on the current training experience replay buffer data sample set; The Importance Sampling Ratio is used to quantify the ratio of the probability change between the "new strategy" and the "old strategy" in the same state of outputting the same action. Indicates time intelligent agent The generalized advantage estimate (GAE) (i.e., the physical quantity calculated in the previous step by the Critic network, representing how much better the current steering action is than the average performance); The pruning function sets the importance sampling ratio. The value of is rigidly limited to the range [ Within; The value is usually 0.1 or 0.2, which represents the maximum percentage (e.g., 20%) that the new strategy is allowed to deviate from the old strategy.

[0060] When deploying inference online, the centralized commentator network is offloaded, retaining only the independent policy executor network (ActorNetwork). Each wheel control motor execution unit operates independently and in a distributed manner, performing millisecond-level asynchronous forward inference directly through the forward matrix multiplication of the policy network. ; ; in, This represents the calculated optimal additional front wheel steering angle command; This represents the calculated optimal active rear wheel steering angle command; Indicates the optimal target value (Optimal Reference Command); and : These represent the front-wheel AFS actor policy network and the rear-wheel ARS actor policy network, respectively, which have been trained and converged offline and embedded in the vehicle chassis domain controller chip; and These represent the physical observation feature streams that are collected and input in real time by the front and rear steering actuators during the current control cycle.

[0061] Within each control cycle (2ms), the optimal turning angle sequence satisfying the equilibrium solution of the non-cooperative game is directly demodulated. This avoids the mutual friction and cancellation of steering control signals in the nonlinear region.

[0062] like Figure 3The diagram shows the evolution topology of the multi-agent reinforcement learning (SAC-MAPPO) interactive network in steps S2 and S3, which involves "upper-level cooperative game - lower-level non-cooperative Nash equilibrium game".

[0063] S4: Physical Limit Constraint Interception and Anti-Saturation Control Allocation Execution Although step S3 has already inferred the optimal turning angle sequence satisfying Nash equilibrium online based on multi-agent reinforcement learning. However, due to the data-driven nature of reinforcement learning algorithms, in the absence of prior physical boundaries, their output commands may excessively exceed the actual adhesion circle of the tire under extreme conditions, causing the actuator to saturate and fail due to hardware limitations. To ensure the "physical feasibility" and "hardware safety" of the entire method, step S4 establishes a physical hard boundary constraint interception and control allocation layer at the final gate of the system output. This step, through the following steps S4-1 to S4-4, introduces the inverse solution of the nonlinear tire brush model and anti-saturation quadratic programming to safely and faithfully map the theoretical control commands to the underlying current signal of the hardware motor. The topology of the physical friction circle hard interception and clipping and the underlying torque control allocation closed-loop structure considering the nonlinear anti-windup constraint described in steps S4-1 and S4-3 is shown in the figure below. Figure 4 As shown, combined with Figure 4 Step S4 will be described in detail: S4-1: Dynamic trimming and interception of physical friction circular boundary and inverse solution of nonlinear brush model Before the lower-level intelligent agent network outputs actions to the physical hardware, a physical hard constraint boundary interception layer is set up. Based on the wheel-end vertical force estimation model, the current vertical load of the four wheels is calculated in real time. : ; ; In the formula, Real-time vertical load (normal force, unit: N) on the left / right wheels of the front axle; This represents the real-time vertical load (normal force, unit: N) on the left / right wheels of the rear axle.

[0064] For longitudinal / lateral acceleration; The height of the center of gravity; This refers to the front and rear wheel track. Based on the road surface adhesion coefficient... Construct the physical adhesion limit friction circle for each tire: ; Indicates the first One wheel ( The wheel end is generated by the real-time longitudinal driving force or braking force actually generated by the hub motor; Combining a nonlinear single-wheel tire brush model, the maximum slip angle limit that a tire can withstand is determined. The physical mechanism derivation formula is as follows: ; In the formula, Let be the lateral stiffness of each individual wheel.

[0065] Wheel slip angle With steering angle The dynamic geometric relationship formula is as follows: ; ; and These are the actual physical steering angles of the front and rear wheels, respectively. and These are the real-time slip angles of the front and rear tires, respectively. and These are the longitudinal vehicle speed and lateral vehicle speed in the vehicle's center of gravity coordinate system, respectively. This refers to the real-time yaw rate of the actual vehicle. and These are the geometric horizontal distances from the vehicle's center of gravity to the front and rear axles, respectively.

[0066] Based on the vehicle's motion state, the limit mapping of the four-wheel independent steering angle is as follows: ; ; If the sideslip angle corresponding to the steering angle deduced in step S3 exceeds the physical limit, the interceptor will initiate a cutoff correction: ; ; in, It is a saturation truncation function.

[0067] S4-2: Solution of Macroscopic Expected Additional Yaw Moment Combining the weight factors output by the upper-level cooperative game agent Define the total desired additional yaw moment that the current control system needs to compensate for. To overcome the rigid distortion of traditional physical models in the extreme nonlinear region, a hybrid bias correction feedback control law is established: ; in, To ultimately issue the desired additional yaw moment target command to the underlying torque distribution actuator; The weight allocation factor is output in real time by the upper-level intelligent agent; and These are the real-time yaw rate of the actual vehicle and the ideal reference yaw rate, respectively. and These are the real-time yaw acceleration of the actual vehicle and the ideal reference yaw acceleration, respectively. and These are the proportional gain coefficient and differential gain coefficient, respectively, preset in the feedback correction control law. This formula introduces a differential term. Establish an advanced damping mechanism for the transient yaw acceleration of the vehicle body, in conjunction with... The dynamic gain multiplier enables millisecond-level smooth switching of control resources between "uninterrupted steering in the linear region" and "powerful differential braking correction in the nonlinear limit region," ensuring the accuracy of vehicle trajectory tracking while comprehensively strengthening the chassis's dynamic instability defense.

[0068] S4-3: Low-level torque control distribution based on nonlinear Anti-Windup correction and Lagrange multiplier KKT demodulation Establish the control distribution matrix for the chassis overdrive system. Assign the desired additional yaw moment. Total driving force with driver's longitudinal expectation As a control target allocation vector The longitudinal force output vector of a four-wheel independent hub motor is defined as follows: Construct a linear matrix relationship for control allocation: ; Among them, the effective control matrix The formal definition is: ; in, These are the front track and the rear track, respectively. These are the actual steering angles of the front wheels and the actual steering angles of the rear wheels, respectively.

[0069] To avoid system oscillations caused by individual hub motors rapidly falling into electromagnetic saturation or slip saturation under extreme operating conditions, a nonlinear anti-saturation correction mechanism is introduced in the control allocation optimization process. The control allocation problem is transformed into a quadratic dynamic optimization solution with feedback compensation: ; ; in, The actuator output vector representing the previous control cycle To find the optimal independent variable, so that the total algebraic loss value of the quadratic programming is... Minimization is a standard mathematical optimization problem; This represents the tracking error weighting diagonal matrix; This is the reference control input vector updated from the previous control time step; This represents the diagonal matrix of control energy adjustment weights; and These represent the lower and upper bound vectors of the dynamic constraints on the independent longitudinal forces of the four wheels, respectively. The wheel-end components of these upper and lower bound vectors are not fixed constants. Instead, they are determined in each control cycle by real-time traversal of the maximum electromagnetic / regenerative braking external characteristic curve limit at the current speed of the hub motor, and by forcibly coupling the real-time road surface adhesion coefficient, the dynamic vertical normal load of the wheel, and the remaining adhesion space of the tire's physical friction circle for intersection interception, thus achieving real-time closed-loop calculation. This ensures that the torque command output by the secondary planner fully exploits the chassis's overdrive capability without penetrating the physical limits of the actuator hardware or the ground adhesion safety barrier.

[0070] For the actuator electromagnetic output saturation residual vector: , This indicates the nominal control command of the hub motor at the previous moment. This indicates the actual physical force feedback from the wheel end at the previous moment.

[0071] To achieve high real-time performance in the onboard drive-by-wire chassis microcontroller, a Lagrange multiplier constructor is constructed to explicitly solve the above quadratic programming problem using KKT conditions: ; The calculated saturation error residual As a negative feedback correction term, it is injected into the SAC reinforcement learning environment observation of S2 in real time, thereby realizing bidirectional closed-loop collaboration from the control layer to the allocation layer.

[0072] The saturated residual vector The specific method for real-time injection into the upper-level reinforcement learning environment observation is as follows: the accumulated saturation residual from the previous control cycle is used... As additional state variables, they are dynamically spliced ​​into the guiding state space generated in step S1-3. In, update it to This forces the upper-level actor network to reduce or adjust the synergy factor in the next moment, thereby compelling it to do so. To avoid hardware saturation.

[0073] S4-4: Low-level hardware driver execution Within each control sampling period, the optimal longitudinal force vectors of each single wheel, calculated online by the embedded quadratic programming (QP) solver in step S4-3, are used. The optimal longitudinal force vector Specifically, this refers to the low-level interior-point quadratic programming solver's application to the saturation-resistant constrained objective function constructed in step S4-3. The output is the global optimal control vector solution after real-time convergence. The braking / driving external characteristic curves of each wheel-end motor are rigidly converted into electromagnetic command torque for the four independent hub motors. (in the formula) (This refers to the effective rolling radius of the wheel). Simultaneously, in high-frequency extraction step S4-1, the trimmed steering angle signal, generated by the saturated truncation function based on real-time road surface adhesion and dynamic axle load transfer interception, is extracted and is located within the safe adhesion friction circle. The command torque With the turning angle signal after trimming Along with the vehicle controller bus messages, they are sent in parallel through the chassis FlexRay / CAN-FD high-speed hard-wired control circuit to the controllers of the chassis brake-by-wire drive motors (MCUs) and steering-by-wire actuators, driving the physical deflection and torque output of each heterogeneous chassis mechanism, and completing the millisecond-level precise intelligent closed-loop control of the vehicle's lateral stability.

[0074] Trimmed steering angle signal Step S4-1 (physical limit constraint interception) is performed using a saturation truncation function. It is obtained through calculation. It is the safe result after offsetting the ideal game theory perspective with the hard physical constraints attached to the underlying ground.

[0075] The method described in this invention introduces a control architecture combining deep reinforcement learning and a two-layer game (upper-layer cooperative game allocation, lower-layer non-cooperative Nash equilibrium execution) into the distributed drive electric vehicle control, and deeply integrates meta-learning driving style representation and physical friction circle hard constraint interception mechanism, breaking the limitations of traditional centralized control and strong dependence on linearized tire physical models. The significant technical advantages and inventive effects of the method described in this invention are specifically reflected in the following aspects: 1. The method described in this invention significantly reduces coupling conflicts under extreme operating conditions of multi-drive-by-wire actuators, balancing high-precision path tracking with inherent stability. Traditional controllers, under extreme and abrupt operating conditions (such as low-attachment double lane change and high-speed obstacle avoidance), often experience control signal cancellation and loss of ground adhesion due to the strong physical coupling between the steering (4WS) and braking (DYC) actuators, leading to chassis vibration. The method described in this invention utilizes the global optimization capability of maximum entropy reinforcement learning (SAC) in its upper-layer cooperative game-theoretic agent, through a Pareto joint reward function and... Phase plane hazard index The precipitous penalty achieves a smooth and seamless shift in control priority between the high-smoothness steering zone and the braking limit recovery zone. Simultaneously, the lower layer uses Multi-Agent Reinforcement Learning (MARL) to solve non-cooperative Nash equilibrium solutions in real time between Active Front Steering (AFS) and Active Rear Steering (ARS), ensuring that each actuator performs its function effectively and eliminating underlying steering physical interference. This results in a synergistic improvement in the vehicle's inherent stability and trajectory following accuracy in the nonlinear limit zone.

[0076] 2. The method described in this invention completely eliminates the reliance of traditional control on high-precision tire mechanics physical models, resulting in stronger robustness under extreme conditions. Traditional model predictive control (MPC) or sliding mode control (SMC) heavily rely on simplified linear or magic tire models. When the vehicle enters the nonlinear limit saturation region with large centroid sideslip angles and high lateral accelerations, the complex flow field mechanics and tire adhesion degrade drastically, often distorting the mathematical model and causing controller failure. The method described in this invention employs a multi-agent reinforcement learning approach combining data-driven and physical feature-guided methods. Through the powerful nonlinear fitting capability of the policy network, it implicitly and adaptively learns the multidimensional mapping between tire sideslip force, slip ratio, and vertical load. Simultaneously, a tire friction circle interception layer based on physical limits is established at the action output point, enabling the controller to maintain extremely high-fidelity execution control and system robustness even in extremely complex nonlinear regions.

[0077] 3. The method described in this invention utilizes offline training and online forward inference mechanisms to overcome the technical bottleneck of real-time performance in traditional game theory and optimization algorithms. Traditional chassis coordination strategies based on game theory Nash equilibrium solutions or swarm optimization (such as KDPC-PSO) require large-scale matrix inversion or swarm evolution iterations online within each vehicle control cycle. The computational load increases exponentially with the control dimension, failing to meet the millisecond-level response requirements of vehicle controllers. The method described in this invention "offline" the complex nonlinear two-layer game computation, using MAPPO and SAC algorithms to complete high-intensity policy network training in a supercomputer or simulation environment. When the algorithm is deployed online on the vehicle microcontroller (MCU / MDC), the actuator network eliminates the centralized commentator and performs inference only through forward matrix multiplication of the lightweight Actor policy network, directly demodulating the optimal control command within 2ms, perfectly meeting the engineering implementation requirements of extremely high real-time performance for drive-by-wire chassis.

[0078] 4. The method described in this invention incorporates adaptive representations of driving habits based on meta-learning, achieving deep decoupling of "human-machine-vehicle" and personalized safety control. Existing collaborative controllers often ignore the differences in driving styles among different drivers (such as aggressive, conservative, and novice drivers), leading to overly aggressive or sluggish intervention during human-machine co-driving. The method described in this invention innovatively introduces a driving intention recognizer based on a meta-learning feature network, using a time-series rolling manipulation dataset as a meta-task to spontaneously extract task context feature vectors representing driving styles online. This feature vector is deeply involved in the dynamic weight allocation of the reinforcement learning state guidance and reward function. Under the premise of ensuring extreme safety, when the system detects that the driver has conservative habits, it can spontaneously reduce the driver's operational load through AFS (Adaptive Front-lighting) to ensure passenger comfort. When it detects aggressive or extremely erroneous operations, under the premise of ensuring extreme safety, the algorithm can achieve a smooth and seamless shift of control between the active safety allocation of human-machine co-driving and the chassis multi-actuator limit boundary rescue, thus alleviating the driver's operational burden and spontaneously constructing an intrinsic safety defense line at the physical limit edge. This achieves a higher-order adaptive human-machine co-driving experience that is more in line with human engineering principles.

[0079] 5. The method described in this invention introduces nonlinear anti-windup constraints and explicit KKT solutions, greatly improving the physical safety of the hub motor drive system. Under extreme conditions, the hub motors of distributed drive vehicles are prone to electromagnetic output or slip saturation failure due to excessive pursuit of yaw torque. If not suppressed, this can lead to integral saturation oscillation and system divergence. The method described in this invention introduces a nonlinear anti-saturation correction mechanism in the lower-level control allocation (CA), adjusting the actual electromagnetic execution residual at the wheel end. As a negative feedback correction term, it feeds back to the upper-level reinforcement learning network in real time. At the same time, by constructing a Lagrange multiplier constructor to perform KKT conditional explicit resolution on the quadratic programming problem, the torque load of the wheel that experiences physical overload or slippage saturation can be instantly and smoothly transferred to the other unsaturated wheels in a dynamic and safe manner within 1ms. This not only fully pushes the ground adhesion limit of multi-wheel independently driven vehicles, but also establishes a solid defense for the active safety of distributed drive vehicles from the physical hardware level.

Claims

1. A two-layer game-theoretic multi-agent cooperative control method for distributed-drive electric vehicles, characterized in that, The method includes the following steps: S1. State Perception and Representation: Constructing the state equation and observation equation of the nonlinear system based on tire slip ratio and physical adhesion limit, and assessing the road adhesion coefficient. Estimate; define the real-time boundary instability risk index. A meta-learning feature network is introduced to extract driver features, resulting in the final guidance state space. ; S2. Dynamic Allocation of Upper-Level Collaborative Weights: Constructing a cooperative game agent at the vehicle's upper-level control layer, and allocating weights through factors... The torque distribution between the lower-level active four-wheel steering control system and the direct yaw torque control system is regulated; a Pareto multi-objective joint reward function, which includes physical boundary penalties and kinetic energy losses, is designed to guide the motion control of the upper-level cooperative game agent. S3, Lower-level active four-wheel steering cooperative execution: Construct non-cooperative game-theoretic agents at the vehicle's bottom execution layer, including an additional front-wheel steering agent responsible for coordinating front axle steering and an active rear-wheel steering agent responsible for controlling rear axle steering; Design specialized reward functions for the two agents respectively to guide their motion control. S4. Physical Limit Constraint Interception and Anti-Saturation Control Allocation Execution: Before the lower-level intelligent agent network outputs actions to the physical hardware, a physical hard constraint boundary interception layer is set up; the macroscopic expectation additional yaw moment is solved and the underlying torque control allocation is performed; based on the underlying torque control allocation, the underlying hardware drive execution is performed.

2. The method for two-layer game-theoretic multi-agent cooperative control of a distributed drive electric vehicle according to claim 1, characterized in that, The state equation of the nonlinear system is: ; The observation equations for the nonlinear system are: ; in, and Represents the discrete time step. Represents the system state vector. Represents the observed input vector. Represents the state transition function of a nonlinear system. Represents the observation mapping function of a nonlinear system. This refers to the physical quantity directly measured and input by the vehicle's onboard sensors. and These are the process noise covariance matrix and the measurement noise covariance matrix, respectively. road surface adhesion coefficient The estimation process involves generating a symmetric Sigma point set with weighted coefficients using an unscented transformation within each control cycle. This point set is then substituted into the nonlinear system state equation and the nonlinear system observation equation for nonlinear recursive propagation. The filter gain is updated in real time using the observed residuals, thereby enabling the estimation of the current road surface adhesion coefficient. Online estimates.

3. The distributed drive electric vehicle two-layer game multi-agent cooperative control method according to claim 2, characterized in that, Define the real-time boundary instability risk index Specifically: Based on real-time estimated road adhesion coefficient With longitudinal speed The critical centroid sideslip angle boundary at which tire nonlinear slip occurs is calculated using a two-degree-of-freedom vehicle dynamics model. Boundary of angular velocity with center of mass deflection : ; ; For gravitational acceleration; establish a double asymptote nonlinear phase plane safety boundary, and vehicle state points. The algebraic transformation formula for the physical distance to the boundary is: ; Define the real-time boundary instability risk index The following segmentation mapping is satisfied: ; in, This characterizes the linear dynamic zone in which the vehicle is in an absolutely safe position. This indicates that the vehicle has fully reached or exceeded the nonlinear instability limit boundary.

4. The distributed drive electric vehicle two-layer game multi-agent cooperative control method according to claim 3, characterized in that, Guiding state space pass Obtain, among which, The task-representation context vector is obtained as follows: A meta-learning feature network is introduced to extract driver features, and a latent intent extraction network containing an encoder layer and a feature mapping layer is constructed. With sliding time window The serialized steering wheel input and path error within the dataset are used to construct the meta-task dataset. Task representation context vectors in the latent space are extracted through online scrolling using grid gradient optimization. : ; in, This indicates the current sampling time, representing the absolute time reference point for the vehicle controller to execute the current control loop. express The value range is 0 to , Indicates the steering wheel angle. Indicates the angular velocity of the steering wheel rotation. This represents the lateral distance tracking error between the vehicle's center of gravity and the desired planned path. This represents the yaw angle error between the vehicle's longitudinal centerline and the tangent direction of the desired path.

5. The distributed drive electric vehicle two-layer game multi-agent cooperative control method according to claim 4, characterized in that, The upper-level cooperative game agent employs a flexible actor-commentator algorithm, in which the actor network guides the state space. Input, output motion space Limited to a continuous interval: ; Weighting factor As a time-varying resource allocation parameter, the allocation ratio between the lower-level active four-wheel steering control system and the direct yaw moment control system is defined when counteracting the total yaw moment demand of the vehicle.

6. The distributed drive electric vehicle two-layer game multi-agent cooperative control method according to claim 5, characterized in that, The Pareto multi-objective joint reward function is: ; in, The tracking error is the lateral distance between the vehicle's center of gravity and the desired planned path. The sideslip angle is the angle between the vehicle's center of gravity and its body. The yaw rate is angular velocity. The nominal weight is a positive coefficient. For the ideal reference centroid sideslip angle, For the ideal reference yaw rate, This is a penalty term for kinetic energy loss. This is a penalty function for exponentially dangerous cliff-like constraints based on the physical boundary of the phase plane. This is a negative feedback term for manipulation burden based on meta-intention.

7. The distributed drive electric vehicle two-layer game multi-agent cooperative control method according to claim 6, characterized in that, The additional front-wheel steering agent and active rear-wheel steering agent will allocate cooperative factors from the upper layer. And the local observation states, pieced together from the local vehicle motion states, serve as the inputs for their respective network decisions. The additional front wheel steering agent outputs an action of additional front wheel steering angle. The active rear-wheel steering agent outputs the active rear-wheel steering angle. The additional front-wheel steering agent and the active rear-wheel steering agent adopt the Actor-Critic neural network structure corresponding to the multi-agent proximal strategy optimization algorithm.

8. The method for two-layer game-theoretic multi-agent cooperative control of a distributed drive electric vehicle according to claim 7, characterized in that, Specialized reward function for the front wheel steering agent for: ; in, , , and These are the specialization performance weighted penalty factors for the additional front wheel steering agent; Specialized reward function for active rear-wheel steering agent for: ; in, , , and These are the specialization performance weighted penalty factors for the active rear-wheel steering agent.

9. The distributed drive electric vehicle two-layer game multi-agent cooperative control method according to claim 8, characterized in that, Based on the wheel-end vertical force estimation model, the current vertical load of the four wheels is calculated in real time, according to the road adhesion coefficient. The physical adhesion limit friction circle of each tire is constructed. Combined with the nonlinear single-wheel tire brush model, the maximum side slip angle limit that each tire can withstand is derived. Combined with the vehicle motion state, the four-wheel independent steering angle limit is obtained. When the additional front wheel steering angle and the active rear wheel steering angle exceed the obtained corresponding four-wheel independent steering angle limit, the interceptor starts the cut-off correction.

10. A distributed drive electric vehicle two-layer game multi-agent cooperative control method according to claim 9, characterized in that, Macroeconomic expectations plus yaw moment pass: get, in, To ultimately issue the desired additional yaw moment target command to the underlying torque distribution actuator; and These are the real-time yaw rate of the actual vehicle and the ideal reference yaw rate, respectively. and These are the real-time yaw acceleration of the actual vehicle and the ideal reference yaw acceleration, respectively. and These are the proportional gain coefficient and the differential gain coefficient, respectively. The underlying torque control allocation is as follows: Establish the control distribution matrix for the chassis overdrive system, and apply the desired additional yaw moment. Total driving force with driver's longitudinal expectation As a control target allocation vector The longitudinal force output vector of the four-wheel independent hub motor is defined as Construct a linear matrix relationship for control allocation: ; Among them, the effective control matrix Defined as: in, These are the front track and the rear track, respectively. These are the actual steering angles of the front wheels and the actual steering angles of the rear wheels, respectively.