Distributed electric drive heavy-duty vehicle multi-domain fusion variable degree of freedom control system and method and vehicle

CN122808764APending Publication Date: 2026-09-25JIANGSU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611202510.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-10
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

然而,现有控制方法多围绕特定工况展开,往往采用固化控制架构,难以根据工况特征和执行器能力等因素动态调整可控自由度,易造成执行器冗余自由度之间的控制冲突与控制计算负担增加,使重载车辆在极限工况下缺乏有效的协同控制能力,引发动态失稳,危及行车安全

Benefits of technology

[0051](1)本发明构建了分布式电驱动重载车辆可变自由度控制范式,将车辆各自由度的冻结、保持、增强与降级策略编码为可学习潜变量码本,实现了车辆有效自由度的自适应组织与控制维度的动态压缩,提升了复杂工况下的控制效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122808764A_ABST
    Figure CN122808764A_ABST
Patent Text Reader

Abstract

The application discloses a distributed electric drive heavy-duty vehicle multi-domain fusion variable degree of freedom control system and method and a vehicle, constructs a variable degree of freedom control paradigm of a distributed electric drive heavy-duty vehicle, encodes freezing, maintaining, enhancing and degrading strategies of each degree of freedom of the vehicle into a learnable latent variable codebook, realizes adaptive organization of effective degrees of freedom of the vehicle and dynamic compression of a control dimension, and improves control efficiency under complex working conditions. A mechanism-embedded large language model variable degree of freedom decision method is proposed, multi-domain information such as scene, vehicle state, dynamics constraint and safety margin is fused, an interpretable degree of freedom latent code strategy is generated, and the autonomous adaptive ability of the vehicle to working condition changes, load disturbance and actuator capacity degradation is enhanced. A distributed electric drive chassis cooperative control method based on dynamics constraint reinforcement learning is designed, virtual generalized control quantities are generated based on the degree of freedom latent code, constraint control distribution and safety protection mechanisms are combined, and safe and efficient operation of the vehicle is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent vehicle dynamics control, specifically to a multi-domain fusion variable degree-of-freedom control system, method, and vehicle for distributed electric drive heavy-duty vehicles. Background Technology

[0002] Heavy-duty vehicles are critical transport equipment supporting the mobile deployment of national defense equipment, engineering construction, mining and port transportation, emergency rescue, and the circulation of bulk materials. Their operational safety and mobility directly affect reliability and transportation efficiency in complex scenarios. Due to their large mass, long wheelbase, high center of gravity, and wide load variation range, heavy-duty vehicles exhibit highly coupled degrees of freedom of motion in all directions under conditions such as high-speed lane changes, emergency obstacle avoidance, low-adhesion road surfaces, and complex road maneuvers. Distributed electric drive technology, by independently adjusting the drive, braking, steering, and suspension subsystems of each wheel, expands the vehicle's controllable degrees of freedom, enhances tire force distribution capabilities, yaw moment control capabilities, and maneuverability under complex conditions, and has become an important development direction for intelligent chassis of heavy-duty vehicles.

[0003] Currently, existing methods for coordinated control of the chassis of distributed electric-drive heavy-duty vehicles mainly include feedforward control based on vehicle dynamics models, feedback control based on sliding mode control, model predictive control, and active disturbance rejection control, as well as data-driven control based on reinforcement learning and deep learning. In engineering implementation, strategies such as rule switching, gain scheduling, and direct yaw moment control are also commonly used. However, existing control methods are mostly focused on specific operating conditions and often adopt a fixed control architecture. It is difficult to dynamically adjust the controllable degrees of freedom according to the characteristics of the operating conditions and the capabilities of the actuators. This can easily lead to control conflicts between redundant degrees of freedom of the actuators and an increased control computation burden. As a result, heavy-duty vehicles lack effective coordinated control capabilities under extreme operating conditions, leading to dynamic instability and endangering driving safety.

[0004] Therefore, there is an urgent need to construct a new paradigm that can dynamically adjust the chassis control architecture according to the characteristics of the working conditions. By dynamically locking the secondary degrees of freedom and focusing on the dominant control degrees of freedom, the high-dimensional coupled distributed chassis collaborative control problem can be transformed into a hierarchical subsystem optimization problem, thereby reducing the dimensionality and decoupling of the complex control system and realizing the safe and efficient operation of heavy-duty vehicles. Summary of the Invention

[0005] The purpose of this invention is to provide a multi-domain fusion variable degree-of-freedom control system, method, and vehicle for distributed electric drive heavy-duty vehicles. For distributed electric drive heavy-duty vehicles with multi-actuator coordination capabilities in drive, braking, steering, and suspension, a variable degree-of-freedom control paradigm oriented towards all operating conditions is established. Based on real-time task requirements, environmental complexity, and the vehicle's own state, the effective control degrees of freedom of the chassis are dynamically increased or decreased, resolving the contradiction between control accuracy and computational efficiency in a fixed control architecture. Under this paradigm, the variable degree-of-freedom strategy is implicitly encoded into a learnable latent variable codebook, establishing a multi-domain information fusion and unified state coding mechanism. A decision-making method based on a mechanism-embedded large language model is proposed, and a collaborative control method with fusion code conditional reinforcement learning and safety protection mechanisms is designed to solve the problem of vehicle stability degradation caused by multi-actuator constraint conflicts, enabling the distributed electric drive heavy-duty vehicle to operate safely and efficiently under complex conditions.

[0006] To achieve the above objectives, the control method of the present invention adopts the following technical solution:

[0007] Step S1: Collect vehicle motion, load, road environment and chassis actuator status. After time synchronization, time delay compensation and rolling time domain estimation, design a mechanism constraint cross attention mechanism, integrate multi-domain information and generate a unified state code.

[0008] Step S2: Design a mechanism-embedded large language model, evaluate vehicle stability, economy and passability based on unified state coding, construct a degree-of-freedom latent code by combining actuator health and safety margin, and determine the target latent code from the candidate set that satisfies dynamics and actuator constraints.

[0009] Step S3 proposes a dynamically constrained latent code conditional dual-delay deterministic policy gradient algorithm, which takes the unified state code, target latent code and historical control quantity as input, modulates the effective control degrees of freedom, and generates continuous virtual generalized control quantity.

[0010] Step S4: Establish the incremental mapping between the virtual generalized control quantity and the driving, braking, steering and suspension commands, reconstruct the allocation model based on the latent code and actuator health, and solve the target command under the constraints of attachment, amplitude, speed and thermal state.

[0011] Step S5: Construct a safety set of vehicle stability, tire adhesion, and actuator state, and use a control barrier function for quadratic programming to correct the target command; when the constraint is not feasible or the actuator fails, implement degradation or backup takeover through the dwell-hysteresis mechanism.

[0012] In step S1, the vehicle motion domain observation vector, chassis execution domain observation vector, and road environment domain observation vector are constructed into a multi-domain original observation vector:

[0013]

[0014] In the formula, For the first Multi-domain original observation vectors at each sampling time , and These are the observation vectors for the vehicle motion domain, chassis execution domain, and road environment domain, respectively.

[0015] A rolling time-domain estimator is constructed based on a vehicle model that includes the vehicle's longitudinal, lateral, yaw, roll, tire, and actuator dynamics. Within a sliding time window, unmeasurable vehicle state and time-varying mechanism parameters are jointly estimated, while satisfying the vehicle state range, road surface adhesion coefficient range, actuator effectiveness range, and tire friction ellipse constraints.

[0016] Furthermore, in step S1, state flow features are used as query information, geometric flow, semantic flow, state flow, and mechanistic flow features are used as keys and values, and the cross-attention weights are jointly modulated using state estimation uncertainty, sensor or actuator health, and mechanistic violation.

[0017]

[0018] In the formula, For the first Attention weights for each information tag, The confidence coefficient is determined by estimating the covariance and health status. For the degree of violation of mechanism, and These are the query vector and the key vector, respectively. For feature dimension, The mechanism penalty coefficient, The total number of information tags. The unified multi-domain state coding is represented as:

[0019]

[0020] In the formula, To unify multi-domain state coding, State flow characteristics, For the first A vector of values The projection matrix of the state flow is... This is a layer normalization operation.

[0021] In step S2, the first step is calculated based on the unified multi-domain state code, the comprehensive evaluation vector, the actuator health vector, and the safety margin vector. The required degree of control freedom:

[0022]

[0023] In the formula, For degrees of freedom, This is a comprehensive evaluation vector composed of stability, economy, and passability indicators. For the executor health vector, For the safety margin vector, and These are the demand mapping parameters and the bias, respectively. This is the activation function.

[0024] Based on the required degrees of freedom, each control degree of freedom is encoded into a state of suppression, degradation, maintenance, or enhancement, and the state is constructed from the degree of freedom modulation coefficient vector and the control priority weight vector. One candidate potential code :

[0025]

[0026] In the formula, The vector of modulation coefficients with degrees of freedom. To control the priority weight vector, The number of candidate control degrees of freedom.

[0027] Furthermore, in step S2, the decision score for the candidate latent codes simultaneously incorporates penalties for violations of dynamic constraints and latent code switching, and the target latent code is selected according to the following formula:

[0028]

[0029] In the formula, To determine the feasible set of latent codes that satisfy vehicle safety constraints and actuator capability constraints, For the mechanism of embedded large language model, The variable-degree-of-freedom latent code selected and actually used at the previous sampling time. For the latent code from the previous sampling time The degree-of-freedom modulation coefficient vector extracted from it. For Hamming distance, Modulation coefficient vector used to measure the current candidate latent code Modulation coefficient vector from the previous time step The differences between them are intended to suppress frequent switching of latent codes. For the degree of violation of dynamic constraints, and These are the mechanism violation penalty coefficient and the latent code switching penalty coefficient, respectively. Index the target latent code.

[0030] In step S3, the actor network generates a normalized control vector and uses the modulation matrix of the degrees of freedom corresponding to the target latent code to form a virtual generalized control quantity:

[0031]

[0032] In the formula, For virtual generalized control variables, For the target code, The modulation matrix of the degrees of freedom corresponding to the latent code. To control the upper limit matrix of the amplitude, For parameters The actor network, The strategy state consists of a unified state code, a target latent code, and historical control variables.

[0033] Furthermore, in step S3, state-action value estimation is performed using two parameter-independent critic networks. The target value is constructed using the smaller of the two target critic estimates, and dynamic consistency loss and control smoothing loss are added to the actor network loss.

[0034]

[0035] In the formula, For the actor's online losses, For the first commentator network, For training batch size, For vehicle dynamic consistency loss, This represents the smoothing loss corresponding to the changes in control quantity at adjacent time points. and These are the corresponding weighting coefficients.

[0036] In step S4, a mapping is established between the actuator instruction increment and the actual generalized control quantity near the current operating point:

[0037]

[0038] In the formula, The actual generalized control quantity formed after the actuator acts. This is the generalized control variable corresponding to the previous operating point. For executor instruction increment, The control allocation matrix is ​​jointly determined by the vehicle status, actuator health, and actuator participation matrix corresponding to the latent code.

[0039] Furthermore, in step S4, the actuator target instruction is obtained by solving a weighted constraint optimization problem with slack variables:

[0040]

[0041] In the formula, For the desired virtual generalized control quantity, The generalized control quantity tracking weight matrix is ​​determined by the latent code. The actuator instruction rate of change weight matrix, Use a cost-weight matrix for the executor. To track slack variables, The slack variable penalty coefficient; the optimization problem satisfies the actuator command upper and lower limits, command change rate, tire friction ellipse, actuator thermal state, and suspension travel constraints.

[0042] In step S5, safety functions are constructed for the sideslip angle, yaw rate, load transfer rate, tire adhesion utilization rate, suspension dynamic deflection, and actuator thermal state, respectively. The vehicle's comprehensive safety domain is the intersection of all safety sets, and the target command output in step S4 is used as the nominal control command. The safety correction command is obtained through quadratic programming of the control barrier function as follows:

[0043]

[0044] And satisfy:

[0045]

[0046] In the formula, For nominal executor instructions, To safely modify the instructions, To correct the weight matrix, To constrain slack variables for safety, The penalty coefficient for slack variables, For the first The safety function is estimated in the current vehicle state. The value at that location, For the current vehicle state estimate and security correction instructions Under the action, the first Each safety function predicts the value of the vehicle state at the next sampling time, where, To predict the vehicle state at the next sampling time. For the first The convergence coefficients of a security function, For the first The relaxation amount of a safety constraint.

[0047] Further, in step S5, a comprehensive risk index is constructed based on the safety function value, the slack amount optimized by the control barrier function, the control allocation residual, and the actuator health. Recovery threshold, degradation threshold, emergency takeover threshold, and minimum dwell time are set. In degradation mode, the failed actuator is frozen, and step S4 is called to reallocate control to the remaining actuators. In emergency takeover mode, the backup controller is called. The final output instruction is:

[0048]

[0049] In the formula, To safely modify the instructions, For degrade or emergency backup control commands. To take over the fusion coefficient, This is the final control command sent to the drive, braking, steering, and suspension actuators.

[0050] The beneficial effects of this invention are:

[0051] (1) This invention constructs a distributed electric drive heavy-duty vehicle variable degree of freedom control paradigm, and encodes the freezing, maintaining, enhancing and degrading strategies of each degree of freedom of the vehicle into a learnable latent variable codebook, realizing the adaptive organization of the effective degree of freedom of the vehicle and the dynamic compression of the control dimension, thereby improving the control efficiency under complex working conditions.

[0052] (2) This invention proposes a mechanism-embedded large language model variable degree of freedom decision method, which integrates multi-domain information such as scenario, vehicle state, dynamic constraints and safety margin, guides the large language model to generate interpretable degree of freedom latent code strategy, and enhances the vehicle's autonomous adaptability to changes in working conditions, load disturbances and actuator capability degradation.

[0053] (3) The present invention designs a distributed electric drive chassis cooperative control method based on dynamic constraint reinforcement learning, generates virtual generalized control quantities with the latent code of the degree of freedom as a condition, and combines constraint control allocation and safety protection mechanism to realize the safe and efficient operation of heavy-duty vehicles. Attached Figure Description

[0054] Figure 1 This is a flowchart of the multi-domain fusion variable degree-of-freedom control system for distributed electric drive heavy-duty vehicles in this invention.

[0055] Figure 2 This is a flowchart of multi-source information fusion and state representation in this invention.

[0056] Figure 3 This is a flowchart of the variable degree of freedom latent code decision-making process in this invention.

[0057] Figure 4 This is a schematic diagram of the variable degree-of-freedom latent variable codebook and degree-of-freedom modulation strategy in this invention.

[0058] Figure 5 This is a flowchart of the multi-actuator collaborative control in this invention. Detailed Implementation

[0059] This invention proposes a multi-domain fusion variable degree-of-freedom control system for distributed electric-driven heavy-duty vehicles, combined with... Figure 1-5Further explanation of the invention:

[0060] Figure 1 This is a flowchart of the multi-domain fusion variable degree-of-freedom control method or system for distributed electric drive heavy-duty vehicles in this invention. It includes the following steps:

[0061] S1: Multi-domain state information fusion and unified coding;

[0062] S2: Construction and decision-making of latent codes with variable degrees of freedom;

[0063] S3: Generation of latent code conditional virtual generalized control quantity;

[0064] S4: Multi-actuator constraint control allocation;

[0065] S5: Security boundary modification and downgrade takeover.

[0066] Step S1: Multi-domain state information fusion and unified coding. Details are as follows:

[0067] like Figure 2 As shown, step S1 is used to perform time synchronization, state reconstruction and mechanism constraint encoding on the vehicle motion domain, chassis execution domain and road environment domain information of the distributed electric drive heavy-duty vehicle, generate a unified state representation with fixed dimensions and consistent physical meaning, and output the unified state representation to step S2 as the input for decision-making.

[0068] At sampling time The system acquires vehicle attitude, vehicle speed, and wheel speed information, as well as the operating status information of the wheel-end motors, braking system, steering system, and suspension system. It also acquires vehicle load, road boundary, road curvature, road slope, road surface adhesion status, obstacle status, scene semantics, route intent, and task constraint information. This information is divided into vehicle motion domain observations, chassis execution domain observations, and road environment domain observations, forming a multi-domain raw observation vector.

[0069]

[0070] In the formula, Indicates the first Multi-domain original observation vectors at each sampling time; The vehicle motion domain observation vector includes longitudinal velocity, lateral velocity, longitudinal acceleration, lateral acceleration, vertical acceleration, yaw rate, roll angle, pitch angle, and wheel speed of each wheel; The chassis execution domain observation vector includes at least the drive torque of each wheel, braking pressure, steering angle, suspension displacement, suspension speed, suspension control force, and actuator health status. The road environment domain observation vector includes at least vehicle load, road boundary, obstacles, road slope, road curvature, road surface adhesion status, route intent, scene category, and task constraints; superscript Represents the transpose of a matrix or vector.

[0071] To address the issues of different sampling periods, communication delays, and timestamp asynchronization among various sensors and chassis controllers, a unified clock from the vehicle's central controller is used. Based on this, time alignment is performed on each information source. For continuously changing [information sources], [the following steps are taken]. For observables, first-order prediction compensation based on time delay estimation is used:

[0072]

[0073] In the formula, Indicates compensation until a unified time. The Observational measurement; Indicates the first The measured value of the information source at its actual sampling time; This represents the first-order rate of change of the corresponding observation with respect to time. Indicates the first The total delay of transmission and sampling of information sources relative to a unified clock; This represents the information source category index. For discrete information such as scene categories and task instructions, time alignment is achieved by preserving the most recent valid value, and a mask variable representing the validity of the information is generated simultaneously to prevent invalid information from participating in subsequent fusion.

[0074] After time synchronization is completed, a rolling time-domain estimator is constructed based on a discrete vehicle dynamics model that includes the vehicle's longitudinal, lateral, yaw, and roll dynamics, as well as the dynamics of the tires and actuators. This estimator has a length of... The unmeasurable state and time-varying mechanism parameters of the vehicle are jointly estimated within a sliding time window. The rolling time-domain estimation problem is expressed as:

[0075]

[0076] And satisfy:

[0077]

[0078] In the formula, Indicates the use of up to the The information at the sampling time is obtained for the first time. State estimate at each time step; This represents the estimated values ​​of the mechanism parameters; Indicates the length of the rolling estimation window; This represents the index of the sampling time within the rolling estimation window, and its range is determined by the upper and lower limits of the corresponding summation term; Indicates the start sampling time of the window The vehicle condition to be estimated This represents the prior state at that moment; This represents the observation vector after time synchronization; The vehicle state vector includes at least the longitudinal velocity, lateral velocity, center of gravity sideslip angle, yaw rate, roll angle, roll rate, longitudinal slip ratio of each wheel, and vertical load. The vector of mechanistic parameters to be estimated includes at least the road adhesion coefficient, vehicle mass, center of gravity height, tire lateral stiffness, and actuator effectiveness. This indicates known chassis input quantities; This represents the state transition function established based on a discrete vehicle dynamics model that includes the vehicle's longitudinal, lateral, yaw, roll, tire, and actuator dynamics. It is used to adjust the state of the vehicle based on its current state. Known chassis input quantities and mechanism parameters Obtain the nominal predicted state at the next sampling time; This represents an observation function established based on the observational relationship between vehicle state, mechanistic parameters, and sensor measurements, used to... and Obtain the synchronous observation vector Corresponding predictive observations; Indicates model perturbation; , and These represent the positive definite weight matrices for prior state error, measurement error, and model perturbation, respectively. Non-negative weighting coefficients representing the inhibition mechanism parameters abruptly; Represented by matrix The weighted quadratic form is defined as follows: ,For example, ; and These represent the permissible sets of vehicle state and mechanism parameters, respectively. Indicates the first A physical constraint function, This is an index of physical constraints. The physical constraints include at least the non-negativity constraint of the wheel normal load, the upper and lower bound constraints of the road adhesion coefficient, the actuator effectiveness range constraint, and the tire friction ellipse constraint.

[0079]

[0080] In the formula, , and They represent the first The longitudinal tire force, lateral tire force, and vertical load of each wheel; Indicates the first The coefficient of adhesion of the road surface where each wheel is located; Indicates the wheel index.

[0081] The rolling time-domain estimation described above yields states that are either not directly measurable or have low measurement reliability, such as the centroid sideslip angle, road surface adhesion coefficient, tire force, load transfer, and actuator equivalent effectiveness. Simultaneously, the covariance matrix of each estimate is output to characterize the degree of uncertainty of the corresponding state information.

[0082] Further, based on the vehicle dynamics equilibrium relationship, construct the mechanism residuals:

[0083]

[0084] In the formula, Indicates the first Vehicle dynamics residuals at each sampling time; Indicates the current estimated mass of the vehicle; and These represent the longitudinal velocity and lateral velocity of the vehicle's center of gravity, respectively. Indicates yaw rate; Indicates the vehicle body roll angle; and These represent the moments of inertia of the vehicle about its vertical and longitudinal axes, respectively. Indicates the first The yaw moment generated by each wheel or actuator; Indicates the first The roll moment generated by each wheel or suspension actuator; the upper dot represents the first derivative of the variable with respect to time, and the upper double dots represent the second derivative of the variable with respect to time. The mechanistic residual is used to measure the degree of consistency between multi-source observation information and vehicle dynamics laws.

[0085] After completing the state estimation, the synchronous observations and estimated states are constructed into geometric flow, semantic flow, and state flow according to information attributes. The geometric flow consists of road boundary points, obstacle points, road slope, and road curvature, and uses a point feature mapping function with shared parameters and a symmetric aggregation function to eliminate the influence of the input point order. Its geometric encoding is represented as follows:

[0086]

[0087] In the formula, Indicates geometric flow encoding; Indicates the first The sampling time of the first sampling moment The feature vector of a geometric point includes at least the relative longitudinal position, the relative lateral position, the boundary or obstacle category, the relative velocity, the local slope, and the local curvature. Indicates the number of geometric points; This represents a point feature mapping function with shared parameters. Its inputs include the relative longitudinal position, relative lateral position, boundary or obstacle category, relative velocity, local slope, and local curvature of a geometric point. Its output is the point-level feature vector corresponding to that geometric point. Each geometric point uses the same mapping parameters, and the mapping results are then aggregated by maximum value according to the feature dimension to obtain the geometric stream code. ; This indicates aggregation of maximum values ​​based on the feature dimension.

[0088] Semantic Stream transforms route intent, scene category, task constraints, and traffic state into a sequence of semantic tags, and obtains semantic features through a semantic encoder. The state flow normalizes and temporally encodes the vehicle state, mechanism parameters, actuator state, actuator health, and safety margin obtained from the rolling time-domain estimation to obtain state features. Simultaneously, the dynamic residual, tire adhesion utilization rate, load transfer rate, and actuator residual capacity shown in equation (6) are encoded as mechanistic characteristics. This results in a set of multimodal tags to be fused:

[0089]

[0090] In the formula, Indicates the first A set of multimodal markers for each sampling time; superscript , , and These represent geometric flow, semantic flow, state flow, and mechanism flow, respectively.

[0091] To avoid giving excessive weight to low-reliability sensor information or information that violates dynamic laws during the fusion process, the cross-attention weights are modulated jointly by estimation uncertainty, actuator health, and the degree of mechanism violation. The credibility coefficient of each tag is:

[0092]

[0093] In the formula, Indicates the first The credibility coefficient of each tag; This indicates the observation corresponding to the marker; Represents the trace of a matrix; Represents a scaling factor that is greater than zero; This indicates the health status of the corresponding sensor or actuator; This represents an exponential function.

[0094] No. The degree of mechanistic violation of a single tag is defined as:

[0095]

[0096] In the formula, Indicates the first Mechanism violation of each marker; This represents the dynamic residual selection and normalization matrix associated with the label; Denotes the norm 1; Indicates the number of physical constraints; Indicates according to the first The vehicle status, mechanism parameters, and the information associated with each information tag The constraint function is constructed from the physical constraint relationships, based on the current state estimate. and mechanism parameter estimates Obtain the corresponding scalar constraint function value when When this physical constraint is satisfied, When >0, the positive part This indicates the amount of violation of the physical constraint; Indicates the first A positive normalization metric for each constraint; and These represent the state estimate and mechanism parameter estimate at the current moment, respectively.

[0097] Using state flow features as query information, and geometric flow, semantic flow, state flow, and mechanistic flow features as keys and values, we calculate the cross-attention weights with mechanistic constraints:

[0098]

[0099]

[0100] In the formula, Represents a set The first in One tag; , and These represent the query vector, key vector, and value vector, respectively. , and These represent the trainable query projection matrix, key projection matrix, and value projection matrix, respectively. The feature dimensions represent the query vector and key vector; Indicates the penalty coefficient for mechanism constraints; Indicates the first Normalized attention weights for each label; Indicates the total number of tags; This represents the attention normalization summation index. Equation (12) allows information with high estimation uncertainty, low health, or high degree of mechanism violation to receive smaller fusion weights.

[0101] The cross-attention fusion result is combined with the state flow features through residual connection and normalization to obtain a unified state code.

[0102]

[0103] In the formula, This indicates the output to the unified multi-domain state code in step S2; Represents the state flow residual projection matrix; The representation layer normalization operation. The unified state coding. It has a preset fixed dimension that does not change with the number of sensors, road points, and actuators, and simultaneously represents the vehicle's operating status, road geometry, scene semantics, task constraints, dynamic mechanism status, actuator health status, and the reliability of various types of information.

[0104] Step S2, latent code construction and decision-making for variable degrees of freedom. Details are as follows:

[0105] like Figure 3 and Figure 4 As shown, step S2 receives the unified multi-domain state code output from step S1. Based on vehicle stability, economy, and passability requirements, and combined with actuator health status and safety margin, it constructs discrete latent codes characterizing the frozen, degraded, maintained, and enhanced states of chassis control degrees of freedom, and determines the variable degree of freedom control mode corresponding to the current sampling time. These latent codes serve as conditional inputs to the virtual generalized control quantity generation model in step S3, transforming the degree of freedom organization from a fixed-rule switching problem into a latent code selection problem constrained by vehicle dynamics mechanisms.

[0106] S2.1, Construction of Latent Codes with Variable Degrees of Freedom

[0107] Based on the unified multi-domain state coding The vehicle's stability, economy, and passability indices are calculated to form a comprehensive evaluation vector.

[0108]

[0109] In the formula, Indicates the first The comprehensive evaluation vector at each sampling time; , and These represent stability, economy, and passability indicators, respectively. Indicates the category of evaluation indicators; Indicates the first The first of the class of evaluation indicators One state quantity; Indicates the corresponding reference value; Indicates the normalization scale; Indicates non-negative evaluation weights; Indicates the first The number of state variables included in the evaluation index. Stability indexes include at least the center of gravity sideslip angle, yaw rate error, tire adhesion utilization rate, and load transfer rate; economic indexes include at least the driving energy consumption, braking energy loss, and actuator action amount; and passability indexes include at least the slope adaptation margin, road boundary margin, vertical impact, and obstacle distance.

[0110] The generalized degrees of freedom of vehicle controllability are represented as:

[0111]

[0112] In the formula, Indicates the longitudinal resultant force degree of freedom; Indicates the degree of freedom of the lateral resultant force; Indicates the degree of freedom of yaw moment; Indicates the degree of freedom of the roll moment; This indicates the degree of freedom for adjusting the vertical load. In other embodiments, independent wheel-end steering degrees of freedom can be added depending on the vehicle chassis configuration.

[0113] Regarding the first Each control degree of freedom is calculated based on multi-domain states, comprehensive evaluation indicators, actuator health, and safety margins to determine the required degree of freedom.

[0114]

[0115] In the formula, Indicates the first The requirement for each degree of freedom to control; Represents the health vector of the drive, braking, steering, and suspension actuators; This represents the safety margin vector corresponding to lateral stability, roll stability, tire adhesion, and actuator residual capacity. and They represent the first The parameter vector and biases for mapping the requirements of each degree of freedom; This represents an activation function that restricts the output range to zero to one. This indicates the index of the degrees of freedom.

[0116] Each degree of control freedom is encoded into a discrete modulation state based on the demand:

[0117]

[0118] In the formula, Indicates the first Modulation coefficients with one degree of freedom; This indicates that the degree of freedom is suppressed; Indicates degraded control; This indicates that routine control measures will be maintained. This indicates enhanced control; , and These are the state determination thresholds that increase sequentially.

[0119] The discrete latent code is composed of modulation coefficients for each degree of freedom and control priority weights:

[0120]

[0121] In the formula, Indicates the first in the codebook One degree of freedom latent code; A vector representing the discrete modulation coefficients for each degree of control freedom; This represents the corresponding control priority weight vector; Indicates the first The first of the hidden codes Control weights for each degree of freedom; Indicates the total number of candidate control degrees of freedom; This represents the latent code index. Multiple latent codes constitute the latent codebook. ,in This indicates the number of latent codes. The latent codebook is obtained through training on typical operating condition data, and combinations of degrees of freedom that cannot be achieved by the current chassis actuators or violate vehicle dynamics constraints are removed.

[0122] S2.2, Variable Degrees of Freedom Latent Code Decision

[0123] Unify multi-domain state coding Comprehensive evaluation vector Actuator health vector Safety margin vector and the latent code at the previous sampling time The data is converted into structured decision labels and input into a mechanism-embedded large language model. This model scores each candidate latent code within the codebook and corrects the latent code scores using mechanism violation penalties and mode switching penalties.

[0124]

[0125] In the formula, Indicates the first The candidate potential code in the first A comprehensive score for each sampling time point; The parameter is The embedded big language model of the mechanism comprehensively scores candidate latent codes based on structured decision labels such as unified multi-domain state coding, comprehensive evaluation vector, actuator health vector, safety margin vector, and latent code at the previous sampling time. This represents the variable-degree-of-freedom latent code selected and actually used at the previous sampling time. It includes the activation or suppression state of each control degree of freedom, control strength, and control priority, and is used as historical decision information for the current candidate latent code scoring. Indicates the use of the first The degree of violation of dynamic constraints at a given latent code includes at least tire adhesion constraints, roll stability constraints, and actuator capability constraints. Indicates the penalty coefficient for violating the mechanism; The Hamming distance between two degree-of-freedom modulation vectors is used to suppress frequent latent code switching. Indicates the penalty coefficient for latent code switching; Indicates the latent code from the previous sampling time. The extracted degree-of-freedom modulation coefficient vector, each component of which represents the state of the corresponding control degree of freedom being suppressed, degraded, maintained, or enhanced.

[0126] From the set of feasible latent codes that satisfy actuator health constraints and vehicle safety constraints Select the potential code with the highest overall score:

[0127]

[0128] In the formula, Represents the set of feasible latent codes in the current state; Indicates the index of the selected latent code; This represents the variable degree of freedom latent code output in step S2; Indicates the current set of feasible latent codes In the middle, compare the candidate latent codes. In the Comprehensive score at each sampling time point It returns the index of the candidate latent key with the highest overall score. ; This represents the optimal index of the selected latent code. In the formula... This indicates that the index in the latent codebook is The candidate latent code is the latent code with the highest comprehensive score within the feasible latent code set; this latent code is determined as the target variable degree of freedom latent code output in step S2. This includes the activation or suppression state, control strength, and control priority of each control degree of freedom. The latent code also includes the activation or suppression state, control strength, and control priority of each control degree of freedom, and serves as a condition variable in step S3, enabling the virtual generalized control quantity to be generated within the effective control dimension specified by the latent code.

[0129] Step S3: Generation of latent code conditional virtual generalized control variables. Details are as follows:

[0130] Step S3 receives the unified multi-domain state code output from step S1. and the variable degree of freedom latent code determined in step S2 A latent code-conditional dual-delay deep deterministic policy gradient algorithm incorporating vehicle dynamics constraints is employed to generate continuous virtual generalized control quantities at the vehicle dynamics level. This method limits the effective dimension, modulation intensity, and control priority of the control degrees of freedom through latent codes, utilizes a dual-critic network to suppress overestimation of the value function, and uses vehicle dynamics residuals and stability boundaries as policy training constraints, ensuring that the generated control quantities possess adaptability, physical consistency, and continuity.

[0131] The unified multi-domain state coding, the variable degree-of-freedom latent code, and the virtual generalized control quantity from the previous sampling time are combined to form the reinforcement learning state:

[0132]

[0133] In the formula, Indicates the first The strategy state at each sampling time; This represents the unified multi-domain state code generated in step S1; This represents the variable degree of freedom latent code generated in step S2; Represents the virtual generalized control quantity at the previous sampling time; superscript This indicates transpose.

[0134] The actor network generates a normalized control vector based on the policy state, and forms the actual virtual generalized control quantity through the degree-of-freedom modulation matrix corresponding to the latent code:

[0135]

[0136]

[0137] In the formula, The parameter is The actor network; This refers to the normalized motion output by the actor via the network. This represents the degree-of-freedom modulation matrix obtained by latent code decoding; Indicates the first Freeze, downgrade, maintain, or enhance coefficients for each degree of freedom; A diagonal matrix representing the upper limit of the amplitude of each virtual generalized control variable; Represents the hyperbolic tangent function; and These represent the desired longitudinal resultant force and the desired lateral resultant force, respectively. Indicates the desired yaw moment; Indicates the desired roll moment; This indicates the desired vertical load adjustment amount; This indicates the number of candidate control degrees of freedom. When the modulation coefficient of a certain degree of freedom is zero, the corresponding virtual generalized control quantity is suppressed; when the modulation coefficient is greater than one, the corresponding degree of freedom gains enhanced control authority.

[0138] Two critic networks with identical structures and independent parameters are used. and Estimate the state-action value separately, and construct the target value using the dual-delay deep deterministic policy gradient method:

[0139]

[0140]

[0141] In the formula, This indicates the target value of the critics' network; Indicates an immediate reward; Indicates the discount factor; Indicates the end of the round; Indicates the first A network of target commentators; Indicates the target actor network; This indicates that the target policy has been smoothed for noise after amplitude truncation. Indicates the first Mean squared error loss of a network of critics; Indicates the batch size of the experience playback; and Let represent the parameters of the online critic network and the target critic network, respectively. The target value is calculated using the smaller of the two critic estimates to mitigate the aggressive control output caused by overestimation of the control value.

[0142] To ensure that the actor network output conforms to the dynamics of heavy-duty vehicles, dynamic consistency penalties and degree-of-freedom smoothing penalties are added to the actor optimization objective:

[0143]

[0144]

[0145]

[0146] In the formula, This indicates the actor's online losses; This indicates the loss of dynamic consistency. This indicates the control of smoothing loss; and These represent the dynamic constraint weights and the control smoothing weights, respectively. and These represent the estimated vehicle state values ​​at the current and next sampling times, respectively. This represents a vehicle state prediction model composed of longitudinal, lateral, yaw, roll, and vertical dynamics. It represents online estimated parameters such as vehicle mass, center of gravity location, tire parameters, and road adhesion coefficient; and These represent the positive definite weight matrices for the dynamic residual and the control rate of change, respectively. Indicates the first One dynamic or stability constraint function; Indicates the number of constraints; The constraints include at least the center of gravity sideslip angle boundary, the yaw rate boundary, the load transfer rate boundary, and the tire adhesion utilization rate boundary.

[0147] Real-time rewards comprehensively reflect trajectory tracking, vehicle stability, energy consumption, ride comfort, and control continuity:

[0148]

[0149] In the formula, Indicates speed or trajectory tracking error; This represents the stability error comprised of the centroid sideslip angle, yaw rate error, and load transfer rate. and These represent the weight matrices for tracking performance and stability performance, respectively. This indicates the equivalent energy consumption corresponding to the driving, braking, and suspension actions. and These represent the weights of the energy consumption term and the control rate of change term, respectively.

[0150] The actor network is updated according to a preset delay period, and the target network is updated using a soft update method. After training, only the actor network is retained to perform online inference, continuously outputting based on the real-time unified state code and the degree-of-freedom latent code. This is then passed to step S4. This structure separates the discrete decision-making of the latent code from the continuous generation of the virtual generalized control quantity. Through latent code modulation, dual critic conservative estimation, and joint optimization of dynamic constraints, it reduces the risk of purely data-driven strategies generating physically unrealizable control quantities.

[0151] Step S4, multi-actuator constraint control allocation. Details are as follows:

[0152] like Figure 5 As shown, step S4 receives the virtual generalized control quantity output in step S3 and the variable degree of freedom latent code determined in step S2, establishes local control mappings from the drive, braking, steering and active suspension actuators to the vehicle's longitudinal force, lateral force, yaw moment, roll moment and vertical load, and completes multi-actuator collaborative optimization allocation under constraints of tire adhesion, actuator amplitude, rate of change, thermal load and health status, and outputs the target commands of each actuator.

[0153] The virtual generalized control quantity output in step S3 is represented as:

[0154]

[0155] In the formula, Indicates the first The expected virtual generalized control quantity at each sampling time; and These represent the desired longitudinal resultant force and the desired lateral resultant force, respectively. and These represent the desired yaw moment and the desired roll moment, respectively. This indicates the desired vertical load adjustment.

[0156] chassis actuator control vector Represented as:

[0157]

[0158] In the formula, This represents the target torque vector of each wheel-end motor; This represents the braking pressure vector of each wheel; This represents the target steering angle vector of the controllable steering wheel; This indicates that each active suspension target is a dynamic vector.

[0159] Considering the nonlinear characteristics of tire force and steering action relative to actuator commands, an incremental control allocation model is established near the current operating point:

[0160]

[0161] In the formula, This represents the actual generalized control quantity formed after the actuator operates; This represents the generalized control quantity corresponding to the operating point at the previous sampling time. Indicates the increment of the executor instruction; Represented by the vehicle state vector and chassis actuator control command vector A defined, nonlinear mapping function from actuator commands to generalized forces and generalized torques, wherein, This includes vehicle longitudinal speed, lateral speed, center of gravity sideslip angle, yaw rate, roll angle, roll rate, longitudinal slip ratio of each wheel, and vertical load. Including the target torque of each wheel-end motor, the braking pressure of each wheel, the target steering angle of the controllable steering wheel, and the target power of each active suspension, the current operating point in equation (32) is taken as the estimated value of the vehicle state. The executor instruction at the previous sampling time ; This represents the current control allocation matrix; Represents the actuator health matrix; Indicates the first Health status of each actuator; Indicates the number of actuator channels; Indicates a latent code with variable degrees of freedom A defined actuator participation matrix is ​​used to freeze unavailable channels or adjust the participation of various actuators in the current degree of freedom.

[0162] Construct a generalized control quantity tracking weight matrix based on the degree-of-freedom priority weights in the latent code. :

[0163]

[0164] In the formula, , , , and These represent the control priority weights for longitudinal, lateral, yaw, tilt, and vertical degrees of freedom, respectively. These weights change with the latent code, ensuring that control allocation prioritizes meeting the corresponding critical stability control requirements under conditions such as low-adhesion steering, long downhill braking, load transfer, or actuator degradation.

[0165] The actuator instructions are obtained by solving the following weighted constraint optimization problem with slack variables:

[0166]

[0167] In the formula, This indicates that the generalized control variable tracks the slack variable, which is used to keep the optimization problem solvable when there are multiple conflicting constraints; This represents the weight matrix indicating the rate of change of actuator instructions. This indicates that the actuator uses a cost weight matrix, which can be set based on motor efficiency, braking energy loss, and suspension energy consumption. This represents the penalty coefficient for slack variables; This represents the L2 norm.

[0168] Equation (34) satisfies actuator capability and vehicle dynamics constraints:

[0169]

[0170] In the formula, and These represent the health status of the actuator. and hidden code The lower and upper limits of instructions after participating in state correction; and These represent the lower and upper limits of the rate of change of the actuator instructions, respectively; , and They represent the first The longitudinal force, lateral force, and vertical load of each wheel; This represents the estimated road surface adhesion coefficient for the corresponding wheel; and They represent the first Current thermal state and permissible upper limit of each drive or brake actuator; , and These represent the suspension dynamic deflection and its upper and lower limits, respectively, with min representing the lower limit and max representing the upper limit. The maximum steering angle and angular rate of the steering actuator, the motor torque and torque change rate, the braking pressure and pressure increase rate, the suspension action force, and the suspension travel are all included in the above amplitude and rate of change constraints.

[0171] Find the optimal solution Then, the target instructions for each executor are generated:

[0172]

[0173] And calculate the control allocation residuals:

[0174]

[0175] In the formula, This indicates the target commands output to the wheel-end drive, braking, steering, and active suspension controllers; This indicates the control allocation residual. Step S4 synchronously outputs the target instruction, control allocation residual, slack variables, and currently active constraints to step S5 for use in safety boundary correction and degrade takeover.

[0176] Step S5, security boundary correction and downgrade takeover. Details are as follows:

[0177] Step S5 receives the target commands, control allocation residuals, slack variables, and actuator health status output from step S4. Based on vehicle lateral stability, roll stability, tire adhesion, and actuator availability, a multi-safety set is constructed. The target commands are modified using quadratic programming of the multi-safety set control barrier function with minimal intrusion. When safety constraints continue to deteriorate, actuators fail, or safety correction is not feasible, the control commands are switched to degraded control or backup controller through residence time and hysteresis monitoring mechanisms to achieve smooth takeover.

[0178] S5.1, Correction of Safety Boundary Monitoring and Control Commands

[0179] Based on the vehicle state estimate obtained in step S1, construct a discrete vehicle prediction model:

[0180]

[0181] In the formula, and They represent the first The and the first Vehicle state estimates at each sampling time point; Indicates the actuator control instruction to be corrected; Represents the discrete vehicle dynamics term without control input; The control action matrix represents the change in vehicle state from actuator commands.

[0182] Safety functions are constructed for the sideslip angle, yaw rate, load transfer rate, tire adhesion utilization, suspension dynamic deflection, and actuator thermal state, respectively. and will the The set of security is represented as:

[0183]

[0184] In the formula, Indicates the first A safe set; This represents the vehicle's overall safety domain, and ∩ represents the intersection. Indicates the first One security function; This indicates the number of safety functions. Each safety function is constructed according to the principle that a function value not less than zero indicates that the corresponding safety constraint is satisfied. The centroid sideslip angle safety function is expressed as follows: ,in, Indicates the current centroid sideslip angle. This represents the upper limit of the absolute value of the allowable sideslip angle. The yaw rate safety function is expressed as... ,in, Indicates the current yaw rate. This represents the upper limit of the absolute value of the permissible yaw rate. The load transfer rate safety function is expressed as... ,in Indicates the load transfer rate. This indicates the upper limit of the allowed load transfer rate. The tire adhesion safety function for each wheel is expressed as follows: ,in, Indicates the first The first wheel in the The estimated value of the adhesion coefficient of the road surface at each sampling time. This indicates the vertical load on the wheel at that moment. and These represent the longitudinal tire force and the lateral tire force of the wheel at that moment, respectively. The dynamic deflection safety function of a suspension is expressed as follows: ,in, Indicates the current suspension dynamic deflection. and These represent the lower and upper limits of the suspension dynamic deflection, respectively. The thermal state safety function of each actuator is expressed as: ,in, Indicates the current thermal state of the actuator. This indicates the upper limit of the permissible thermal state. The safety functions corresponding to each wheel, suspension, and actuator are all included in the total number of safety functions. The target instruction output in step S4. As nominal control instructions, safety correction instructions are obtained through quadratic programming of multi-safety-set control barrier functions:

[0185]

[0186] In the formula, This represents the target instructions for each executor output in step S4; This indicates the control command after safety modifications; This represents the control correction weight matrix; Represents the slack variables of safety constraints; This represents the penalty coefficient for slack variables; Indicates the first The safety function is estimated in the current vehicle state. The value at that point represents the current safety margin; This represents the estimated value of the current vehicle state. and security correction instructions Under the action, the first Each safety function predicts the value of the vehicle state at the next sampling time, i.e., predicts the safety margin, where... The predicted vehicle state for the next sampling time; Indicates the first The convergence coefficients of the security functions; Indicates the first The slack of a safety constraint; and These represent the lower and upper limits of instructions allowed by the executor, respectively. and represents the lower and upper limits of the rate of change of the actuator command, respectively. Equation (40) maintains the original control strategy when not approaching the safety boundary by minimizing the deviation between the safety command and the nominal command, and only corrects the necessary control channels when the predicted state may exceed the limit.

[0187] S5.2, Fault Diagnosis and Degradation Takeover

[0188] The health status of each actuator is assessed using the residual between the actual and predicted responses:

[0189]

[0190] In the formula, Indicates the first The response residuals of each actuator; Indicates the actual response of the actuator; This represents the predicted response obtained based on the actuator model; This represents the residual covariance matrix under normal operating conditions. This indicates the health status of the actuator; the lower the health status, the higher the degree of actuator degradation or failure. This indicates the executor channel index.

[0191] Risk indicators are constructed by combining safety functions, quadratic programming slack, control allocation residuals, and actuator health.

[0192]

[0193] In the formula, This represents a comprehensive risk indicator; ; Indicates a safe function index. This represents the degree of normalization violation across all security functions. The maximum value is taken from the middle value to represent the severity of the most serious violation of the current security constraint; Indicates the first Normalized scale of a security function; The optimal slack variable is expressed in equation (40); Indicates the allowable relaxation amount; This represents the control allocation residual generated in step S4; This indicates that residual allocation is allowed; Indicates the actuator channel index. This indicates that the health of all actuator channels is insufficient. The maximum value is taken from the middle value to represent the actuator with the lowest current health and the most severe degradation. Indicates the first The actuator in the first Health status at each sampling time; Indicates the minimum allowable health level of the actuator; The infinite norm of a vector is the maximum absolute value among all the components of that vector.

[0194] The settings meet the requirements. recovery threshold Downgrade threshold and emergency takeover threshold And set a minimum stay time Only when risk indicators Continue to exceed When the minimum dwell time is reached, the system transitions from normal mode to degraded mode; only when the risk indicators... Persistently below When the minimum dwell time is reached, the system returns to normal mode; when When equation (40) is not feasible, the slack exceeds the allowable upper limit, or a critical actuator fails, the system directly enters the emergency takeover mode. The dual threshold hysteresis and dwell time constraints are used to avoid frequent switching of control modes when the vehicle state fluctuates near the safety boundary.

[0195] In degraded mode, the failed actuator channels are frozen according to the health matrix, the control weights of high-risk degrees of freedom are reduced, and step S4 is called to reallocate control to the remaining available actuators; in emergency takeover mode, backup commands are generated by calling the backup controller based on vehicle longitudinal deceleration, yaw stabilization, and roll suppression. To minimize instruction abrupt changes during the takeover process, the final output instructions are presented in a continuous fusion format:

[0196]

[0197] In the formula, The security correction instruction obtained from expression (40); Indicates the output of the degraded or emergency backup controller; This represents the takeover integration coefficient, which is 0 in normal mode, gradually increases according to a preset rate of change during downgrade or emergency takeover, and is 1 in complete takeover. This indicates the safety control commands ultimately sent to the drive, braking, steering, and suspension actuators.

[0198] Based on the above control method, embodiments of the present invention further include a control system, comprising:

[0199] The unified state code generation module is capable of executing the content of step S1 above;

[0200] The target latent code generation module is capable of performing the content of step S2 above;

[0201] The continuous virtual generalized control quantity generation module is able to execute the contents of step S3 above;

[0202] The executor target instruction generation module is capable of executing the contents of step S4 above;

[0203] The hierarchical control module is capable of executing the contents of step S5 above.

[0204] Based on the above control method, this embodiment of the invention also includes a distributed electric drive heavy-duty vehicle, which is controlled according to the above steps S1-S5.

[0205] The detailed descriptions listed above are merely specific descriptions of feasible embodiments of the present invention, and are not intended to limit the scope of protection of the present invention. All equivalent methods or modifications that do not depart from the technology of the present invention should be included within the scope of protection of the present invention.

Claims

1. A multi-domain fusion variable degree-of-freedom control method for distributed electric-driven heavy-duty vehicles, characterized in that, include: Step S1: Collect vehicle motion, load, road environment and chassis actuator status. After time synchronization, time delay compensation and rolling time domain estimation, multi-domain information is fused through the mechanism constraint cross attention mechanism to generate a unified state code. Step S2: Using the mechanism-embedded large language model, the vehicle's stability, economy, and passability are evaluated based on the unified state coding. The degree-of-freedom latent code is constructed by combining the actuator health and safety margin. The target latent code is determined from the candidate set that satisfies the dynamics and actuator constraints. Step S3: The dynamic constraint latent code conditional double-delay deterministic strategy gradient algorithm is adopted. The unified state code, target latent code and historical control quantity are used as inputs to modulate the effective control degrees of freedom and generate continuous virtual generalized control quantity. Step S4: Establish the mapping from actuator instruction increment to actual generalized control quantity, construct the allocation model based on the latent code and actuator health, and solve the actuator target instruction under attachment, amplitude, rate and thermal state constraints. Step S5: Construct a safety set of vehicle stability, tire adhesion, and actuator state, and use a control barrier function for quadratic programming to correct the target command; when the constraint is not feasible or the actuator fails, implement degradation or backup takeover through the dwell-hysteresis mechanism to achieve smooth takeover of the control command.

2. The multi-domain fusion variable degree-of-freedom control method for distributed electric-driven heavy-duty vehicles according to claim 1, characterized in that, In step S1, the collected vehicle motion domain observation vector, chassis execution domain observation vector, and road environment domain observation vector are constructed into a multi-domain original observation vector: In the formula, For the first Multi-domain original observation vectors at each sampling time , and These are the observation vectors for the vehicle motion domain, chassis execution domain, and road environment domain, respectively. A rolling time-domain estimator is constructed based on a vehicle model that includes the vehicle's longitudinal, lateral, yaw, roll, tire and actuator dynamics. Within a sliding time window, unmeasurable vehicle state and time-varying mechanism parameters are jointly estimated, while satisfying the vehicle state range, road surface adhesion coefficient range, actuator effectiveness range and tire friction ellipse constraints. State flow features are used as query information, and geometric flow, semantic flow, state flow, and mechanistic flow features are used as keys and values. Specifically: geometric flow consists of road boundary points, obstacle points, road slope, and road curvature; semantic flow converts route intent, scene category, task constraints, and traffic state into semantic tag sequences and obtains semantic features through a semantic encoder; state flow is obtained by normalizing and temporally encoding vehicle state, mechanistic parameters, actuator state, actuator health, and safety margin obtained from rolling temporal estimation; and mechanistic flow is obtained by encoding dynamic residuals, tire adhesion utilization, load transfer rate, and actuator remaining capacity. Cross-attention weights are jointly modulated using state estimation uncertainty, sensor or actuator health, and mechanistic violation. In the formula, For the first Attention weights for each information tag, The confidence coefficient is determined by estimating the covariance and health status. For the degree of violation of mechanism, and These are the query vector and the key vector, respectively. For feature dimension, The mechanism penalty coefficient, The total number of information tags; The cross-attention fusion result is combined with the state flow features through residual connection and normalization to generate a unified state encoding representation: In the formula, To unify state coding, State flow characteristics, For the first A vector of values The projection matrix of the state flow is... This is a layer normalization operation.

3. The multi-domain fusion variable degree-of-freedom control method for distributed electric-driven heavy-duty vehicles according to claim 2, characterized in that, The implementation of step S2 includes: S2.1, Constructing a latent code with variable degrees of freedom According to the unified state code Calculate vehicle stability, economy, and passability indicators to obtain a comprehensive evaluation vector: In the formula, Indicates the first The comprehensive evaluation vector at each sampling time; , and These represent stability, economy, and passability indicators, respectively. Indicates the category of evaluation indicators; Indicates the first The first of the class of evaluation indicators One state quantity; Indicates the corresponding reference value; Indicates the normalization scale; Indicates non-negative evaluation weights; Indicates the first The number of state variables included in the evaluation index; stability indexes include at least the center of gravity sideslip angle, yaw rate error, tire adhesion utilization rate and load transfer rate; economic indexes include at least the driving energy consumption, braking energy loss and actuator action amount; passability indexes include at least the slope adaptation margin, road boundary margin, vertical impact and obstacle distance; The generalized degrees of freedom for vehicle controllability are represented as: In the formula, Indicates the longitudinal resultant force degree of freedom; Indicates the degree of freedom of the lateral resultant force; Indicates the degree of freedom of yaw moment; Indicates the degree of freedom of the roll moment; Indicates the degree of freedom for adjusting the vertical load; Regarding the first Each control degree of freedom is calculated based on multi-domain states, comprehensive evaluation indicators, actuator health, and safety margins to determine the required degree of freedom. In the formula, Indicates the first The requirement for each degree of freedom to control; Represents the health vector of the drive, braking, steering, and suspension actuators; This represents the safety margin vector corresponding to lateral stability, roll stability, tire adhesion, and actuator residual capacity. and They represent the first The parameter vector and biases for mapping the requirements of each degree of freedom; This represents an activation function that restricts the output range to zero to one. Indicates the index of control degrees of freedom; Each degree of control freedom is encoded into a discrete modulation state based on the demand: In the formula, Indicates the first Modulation coefficients with one degree of freedom; This indicates that the degree of freedom is suppressed; Indicates degraded control; This indicates that routine control measures will be maintained. This indicates enhanced control; , and The state determination thresholds are incremented sequentially. The discrete latent code is composed of modulation coefficients for each degree of freedom and control priority weights: In the formula, Indicates the first in the codebook One degree of freedom latent code; A vector representing the discrete modulation coefficients for each degree of control freedom; This represents the corresponding control priority weight vector; Indicates the first The first of the hidden codes Control weights for each degree of freedom; Indicates the total number of candidate control degrees of freedom; Represents the latent code index; multiple latent codes constitute the latent codebook. ,in Indicates the number of latent codes; S2.2, Variable Degrees of Freedom Latent Code Decision Unified state coding Comprehensive evaluation vector Actuator health vector Safety margin vector and the latent code at the previous sampling time The data is converted into structured decision labels and input into a mechanism-embedded large language model. The model scores each candidate latent code in the codebook and corrects the latent code scores using mechanism violation penalties and mode switching penalties. In the formula, Indicates the first The candidate potential code in the first A comprehensive score for each sampling time point; The parameter is The mechanism of embedded decision-making model; This represents the variable-degree-of-freedom latent code selected and actually used at the previous sampling time. Indicates the use of the first The degree of violation of dynamic constraints at a given latent code includes at least tire adhesion constraints, roll stability constraints, and actuator capability constraints. Indicates the penalty coefficient for violating the mechanism; This represents the Hamming distance between two degree-of-freedom modulation vectors; Indicates the penalty coefficient for latent code switching; This represents the degree-of-freedom modulation vector used at the previous sampling time. From the set of feasible latent codes that satisfy actuator health constraints and vehicle safety constraints Select the potential code with the highest overall score: In the formula, Represents the set of feasible latent codes in the current state; Indicates the index of the selected latent code; This represents the target latent code with variable degrees of freedom in the output.

4. The multi-domain fusion variable degree-of-freedom control method for distributed electric-driven heavy-duty vehicles according to claim 3, characterized in that, In step S3, the unified state code, target latent code, and historical control quantity are combined into a reinforcement learning state: In the formula, Indicates the first The strategy state at each sampling time; This represents the unified state code generated in step S1; This represents the target latent code in step S2; Represents the virtual generalized control quantity at the previous sampling time; superscript Indicates transpose; The actor network generates a normalized control vector based on the policy state, and a virtual generalized control quantity is formed using the modulation matrix of the degree of freedom corresponding to the latent code. In the formula, The parameter is The actor network; This refers to the normalized motion output by the actor via the network. This represents the degree-of-freedom modulation matrix obtained by latent code decoding; Indicates the first Freeze, downgrade, maintain, or enhance coefficients for each degree of freedom; A diagonal matrix representing the upper limit of the amplitude of each virtual generalized control variable; Represents the hyperbolic tangent function; and These represent the desired longitudinal resultant force and the desired lateral resultant force, respectively. Indicates the desired yaw moment; Indicates the desired roll moment; This indicates the desired vertical load adjustment amount; Indicates the number of candidate control degrees of freedom; Two parameter-independent critic networks are used to estimate the state-action value, and the target value is constructed according to the dual-delay deep deterministic policy gradient method. In the formula, This indicates the target value of the critics' network; Indicates an immediate reward; Indicates the discount factor; Indicates the end of the round; Indicates the first A network of target commentators; Indicates the target actor network; This indicates that the target policy has been smoothed for noise after amplitude truncation. Indicates the first Mean squared error loss of a network of critics; Indicates the batch size of the experience playback; and These represent the parameters of the online critic network and the target critic network, respectively. The target value is constructed using the smaller of the two target critic estimates, and dynamic consistency loss and control smoothing loss are added to the actor network loss: In the formula, For the actor's online losses, For the first commentator network, For training batch size, For vehicle dynamic consistency loss, This represents the smoothing loss corresponding to the changes in control quantity at adjacent time points. and These are the corresponding weighting coefficients.

5. The multi-domain fusion variable degree-of-freedom control method for distributed electric-driven heavy-duty vehicles according to claim 4, characterized in that, In step 4, a mapping is established from the actuator instruction increment to the actual generalized control quantity: In the formula, The actual generalized control quantity formed after the actuator acts. This is the generalized control variable corresponding to the previous operating point. For executor instruction increment, The control allocation matrix is ​​jointly determined by the vehicle status, actuator health, and actuator participation matrix corresponding to the latent code. Establish an incremental control allocation model: In the formula, The nonlinear mapping function representing the actuator command to the generalized force and generalized torque. This is the vehicle state vector. This refers to the control command vector for the chassis actuators. This represents the current control allocation matrix; This represents the current estimated vehicle status. Represents the actuator health matrix; Indicates the first Health status of each actuator; Indicates the number of actuator channels; Indicates a latent code with variable degrees of freedom A defined actuator participation matrix is ​​used to freeze unavailable channels or adjust the participation of various actuators in the current degree of freedom. Construct a generalized control quantity tracking weight matrix based on the priority weights of the degrees of freedom in the latent code: In the formula, , , , and These represent the control priority weights for the longitudinal, lateral, yaw, tilt, and vertical degrees of freedom, respectively. The target instruction for the actuator is obtained by solving a weighted constraint optimization problem with slack variables: In the formula, For the desired virtual generalized control quantity, The generalized control quantity tracking weight matrix is ​​determined by the latent code. The actuator instruction rate of change weight matrix, Use a cost-weight matrix for the executor. To track slack variables, The slack variable penalty coefficient; the optimization problem satisfies the actuator command upper and lower limits, command change rate, tire friction ellipse, actuator thermal state, and suspension travel constraints.

6. The multi-domain fusion variable degree-of-freedom control method for distributed electric-driven heavy-duty vehicles according to claim 5, characterized in that, The implementation of step S5 includes: Safety functions are constructed for the sideslip angle, yaw rate, load transfer rate, tire adhesion utilization, suspension dynamic deflection, and actuator thermal state, respectively. The vehicle's comprehensive safety domain is the intersection of all safety sets, the first... The set of security is represented as: In the formula, Indicates the first A safe set; Indicates the vehicle's overall safety domain; Indicates the first One security function; Indicates the number of security functions; The target instruction output in step S4 The nominal control instruction is used to obtain the modified target instruction through quadratic programming of the control barrier function as follows: And satisfy: In the formula, For safety-corrected control commands, To correct the weight matrix, To constrain slack variables for safety, The penalty coefficient for slack variables, For the first The safety function is estimated in the current vehicle state. The value at that location, For the current vehicle state estimate and security correction instructions Under the action, the first Each safety function predicts the value of the vehicle state at the next sampling time, where, To predict the vehicle state at the next sampling time. For the first The convergence coefficients of a security function, For the first The relaxation amount of a safety constraint.

7. The multi-domain fusion variable degree-of-freedom control method for distributed electric-driven heavy-duty vehicles according to claim 6, characterized in that, Step S5 further includes: fault diagnosis and downgrade takeover. The health status of each actuator is assessed using the residual between the actual and predicted responses: In the formula, Indicates the first The response residuals of each actuator; Indicates the actual response of the actuator; This represents the predicted response obtained based on the actuator model; This represents the residual covariance matrix under normal operating conditions. This indicates the health status of the actuator; the lower the health status, the higher the degree of actuator degradation or failure. Indicates the actuator channel index; Risk indicators are constructed by combining safety functions, quadratic programming slack, control allocation residuals, and actuator health. In the formula, This represents a comprehensive risk indicator; ; Indicates a safe function index. This represents the degree of normalization violation across all security functions. The maximum value is taken from the middle value to represent the severity of the most serious violation of the current security constraint; Indicates the first Normalized scale of a security function; The optimal slack variable is expressed in equation (40); Indicates the allowable relaxation amount; Indicates control allocation residuals; This indicates that residual allocation is allowed; Indicates the actuator channel index. This indicates that the health of all actuator channels is insufficient. The maximum value is taken from the middle value to represent the actuator with the lowest current health and the most severe degradation. Indicates the first The actuator in the first Health status at each sampling time; Indicates the minimum allowable health level of the actuator; Only when risk indicators Continuously exceeding the downgrade threshold Minimum stay time At that time, it will switch from normal mode to downgraded mode; Only when risk indicators Persistently below the recovery threshold Minimum stay time At that time, return to normal mode; When risk indicators Greater than or equal to the emergency takeover threshold When the slack exceeds the allowable limit or a critical actuator fails, the system enters emergency takeover mode. In degraded mode, the failed actuator is frozen and step S4 is called to reallocate control to the remaining actuators; in emergency takeover mode, the backup control command is called, and the final output command is: In the formula, To safely modify the instructions, For degrade or emergency backup control commands. To take over the fusion coefficient, This is the final control command sent to the drive, braking, steering, and suspension actuators.

8. The multi-domain fusion variable degree-of-freedom control method for distributed electric-driven heavy-duty vehicles according to claim 7, characterized in that, Recovery threshold Downgrade threshold and emergency takeover threshold satisfy: .

9. The system of the multi-domain fusion variable degree-of-freedom control method for distributed electric drive heavy-duty vehicles according to any one of claims 1-8, characterized in that, include: The unified state code generation module is capable of executing the content of step S1; The target latent code generation module is capable of executing the content of step S2; The continuous virtual generalized control quantity generation module is able to execute the content of step S3; The executor target instruction generation module is capable of executing the content of step S4; The hierarchical control module is capable of executing the content of step S5.

10. The distributed electric drive heavy-duty vehicle according to the multi-domain fusion variable degree-of-freedom control method for distributed electric drive heavy-duty vehicles according to any one of claims 1-8, characterized in that, The vehicle is controlled according to steps S1-S5.