Data-driven chassis safety filtering control method and end-to-end motion control architecture with safety guarantee

By using a differentiable safety protection layer and an end-to-end joint optimization framework, this approach solves the problems of overly conservative safety strategies in traditional methods and performance exploration in data-driven methods while ensuring safety. It achieves synergistic optimization of safety and performance and is suitable for complex multi-actuator control scenarios.

CN121763735APending Publication Date: 2026-03-31TONGJI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-18
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

In existing technologies, traditional methods are overly conservative in safety-critical scenarios, making it difficult to fully leverage the performance potential of data-driven methods while ensuring safety. Furthermore, training data for multi-actuator drive-by-wire chassis is scarce and the quality of examples varies, making it difficult to coordinate the optimization of safety and performance.

Method used

By employing a differentiable safety protection layer and an end-to-end joint optimization framework, and by constructing control barrier functions for actuator action, maneuver stability, and tire adhesion constraints, combined with multi-objective optimization and a two-stage training method, we achieve synergistic optimization of safety and performance.

Benefits of technology

While ensuring safety, the system effectively explores performance boundaries, improves motion control performance, and solves the problem of scarce training data in multi-actuator drive-by-wire chassis, achieving synergistic optimization of safety and performance and engineering feasibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121763735A_ABST
    Figure CN121763735A_ABST
Patent Text Reader

Abstract

The invention relates to a data-driven chassis safety filtering control method and an end-to-end motion control framework with safety assurance. The method comprises the following steps: establishing a vehicle kinetic equation, constructing an actuator action constraint, a manipulation stability constraint and a tire attachment constraint, and integrating into a micro safety protection layer through micro quadratic programming; constructing a strategy generation network considering a vehicle kinetic equation, combining the strategy generation network with the microdistinguishable safety protection layer to form an end-to-end joint optimization framework, and forming an inner layer-outer layer optimization mechanism inside; the expert data acquisition method based on multi-objective optimization construction and parameter adjustment is used for constructing an end-to-end joint optimization framework training data set, and comprises two stages of training: the first stage is imitation learning, and the second stage is online fine adjustment. Compared with the prior art, the method has the advantages that the performance is enhanced on the premise of ensuring the safety and stability, the safety and the performance are optimized cooperatively, and the consistency and the high quality of expert examples on the multi-target dimension are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data-driven motion control, and in particular to a data-driven chassis safety filtering control method and an end-to-end motion control architecture with safety assurance. Background Technology

[0002] With the rapid development of intelligent vehicles, autonomous driving systems, and service robots, the reliability and performance of motion control algorithms in safety-critical scenarios have become core issues in research and engineering applications. Traditional mechanism-based safety protection strategies rely on deterministic rules or conservative controllers, switching to more conservative control strategies to ensure safety when risks arise. However, this "safety-first" approach often suppresses the performance optimization potential of data-driven methods, preventing vehicles or robots from fully utilizing their kinematic and dynamic performance under normal operating conditions, thus affecting performance indicators such as responsiveness and trajectory tracking accuracy.

[0003] Currently, data-driven motion control methods have made some progress in improving control performance. However, when facing safety-critical scenarios, the mainstream approach still decouples policy generation from safety protection functions: on the one hand, high-performance policies are obtained using learning or optimization methods, and on the other hand, the output is hard-constrained or taken over when risks occur through independent rules / monitors. This type of decoupled architecture has several shortcomings: (1) the high-priority switching of the safety module will cause the policy to be frequently shielded in most actual operating scenarios, resulting in excessive conservatism; (2) the separation of policy and safety logic hinders the system from co-learning and optimizing the performance-safety trade-off under complex operating conditions; (3) existing calibration methods generally rely on manual experience or single-objective parameter tuning, making it difficult to automatically and interpretably find a suitable compromise solution among multiple objectives.

[0004] Furthermore, end-to-end optimization technology has demonstrated its advantages in end-to-end learning and joint optimization in the fields of autonomous driving and robotics. However, how to effectively apply end-to-end optimization capabilities to motion control architecture to achieve the optimal trade-off between safety and performance, while ensuring that safety constraints are not violated, remains a critical issue that urgently needs to be addressed. Therefore, a new control architecture and training / calibration method are urgently needed, enabling the system to maximize and improve motion control performance through data-driven methods under strict safety assurance, while maintaining interpretability and engineering feasibility.

[0005] Chinese invention patent CN120850747A discloses a large-model-driven automotive modeling method based on the Modelica language. The method includes: firstly, employing a domain-knowledge-enhanced BERT-CRF multi-task model to parse natural language requirements into structured triples; secondly, using a graph neural network-driven multi-objective optimization algorithm to match the optimal combination of automotive components, combining symbolic mathematical derivation and machine learning to collaboratively predict cross-disciplinary parameter feasible regions; thirdly, using graph attention networks to optimize the topology connection matrix, fusing template engines and syntax tree analysis techniques to achieve intelligent generation and dynamic verification of Modelica simulation code; and finally, constructing a multi-objective activation function through reinforcement learning to optimize the control strategy and establishing a closed-loop knowledge iteration mechanism. This invention, through a deep integration of large models and Modelica, significantly reduces the technical difficulty of automotive system modeling while maintaining the rigor of physical modeling, making it particularly suitable for complex scenarios such as new energy vehicle development and intelligent driving system integration. However, it still suffers from problems such as overly conservative safety strategies in traditional methods, difficulty in exploring performance boundaries while ensuring safety using data-driven methods, difficulty in co-optimizing safety and performance, and scarcity of training data and inconsistent example quality in multi-actuator drive-by-wire chassis.

[0006] In summary, there is currently a lack of a data-driven chassis safety filtering control method and an end-to-end motion control architecture with safety guarantees to solve or partially solve the above problems. Summary of the Invention

[0007] The purpose of this invention is to overcome the shortcomings of the existing technology by providing a data-driven chassis safety filtering control method and an end-to-end motion control architecture with safety assurance. This addresses or partially addresses the problems of overly conservative safety strategies and the difficulty in exploring performance boundaries while ensuring safety using data-driven methods in traditional methods, the difficulty in co-optimizing safety and performance, and the scarcity and inconsistent quality of training data in multi-actuator drive-by-wire chassis.

[0008] The objective of this invention can be achieved through the following technical solutions: According to one aspect of the present invention, a data-driven chassis safety filtering control method is provided, the method specifically comprising: S1. Establish vehicle dynamics equations, and based on the vehicle dynamics equations, construct actuator action constraints, handling stability constraints, and tire adhesion constraints. Transform the vehicle dynamics equations into control barrier function form, and integrate the constructed constraints into a differentiable safety protection layer through differentiable quadratic programming. S2. Construct a strategy generation network that considers the vehicle dynamics equations, and combine the strategy generation network with the differentiable safety protection layer to form an end-to-end joint optimization framework, wherein the end-to-end joint optimization framework forms an inner-outer optimization mechanism internally; S3. Construct an expert data acquisition method based on multi-objective optimization for constructing and tuning the training dataset of the end-to-end joint optimization framework. The expert data acquisition method includes two-stage training: the first stage adopts imitation learning, and the second stage adopts online fine-tuning.

[0009] As a preferred technical solution, the vehicle dynamics equations Expressed as: In the formula, It is an n-dimensional state vector; Represents the state transition function; Represents the control function. To control variables, For control constraints.

[0010] As a preferred technical solution, the specific form of the higher-order control barrier functions for each constraint obtained from the vehicle dynamics model is as follows: In the formula, Constraints on four-wheel steering and drive / braking torque, Stability constraints include constraints on the sideslip angle and yaw rate. The tire is attached to an elliptical constraint. The relative degrees of the original actuator constraints, stability constraints, and tire adhesion constraints with respect to the vehicle dynamics model. Let the state transition function vector be... For the original constraints of Lie derivative, This is the control coefficient matrix. Original tire adhesion constraint and The mixed Lie derivative, and They are all Li Dao's number operators. Learnable parameters Polynomials with derivatives For learnable parameters, For indexing, To control variables, For control constraints.

[0011] As a preferred technical solution, the constraint high-order control barrier function is integrated into a differentiable security protection layer through differentiable quadratic programming in the following specific form: In the formula, The weight matrix is ​​a positive definite quadratic term. The linear term weight matrix, As a high-dimensional hidden feature, For differentiable quadratic programming network parameters, To optimize the obtained safety control parameters, To control variables, Indicates matrix transpose. For a moment, To solve for the step size using a differentiable quadratic programming problem, At the initial moment, For time indexing, A matrix consisting of the minimum values ​​of the control variables. A matrix consisting of the maximum values ​​of the control variables.

[0012] As a preferred technical solution, the gradient of the control architecture loss function with respect to the inner optimization parameters... Specifically, this manifests as follows: In the formula, For the loss function on the control vector gradient, This represents the optimal control vector currently being solved. For optimal Lagrange multipliers, For duality sensitivity, It is a diagonal matrix. This indicates the matrix transpose.

[0013] As a preferred technical solution, the policy generation network includes a dynamic feature encoder and a multi-branch fusion decoder; The dynamic feature encoder includes a reference trajectory processor, a dynamic guided attention module, and a fusion coding module for dynamic coding and stability enhancement; The multi-branch fusion decoder includes a reference control decoder, a model uncertainty decoder, and a security constraint parameter decoder; The reference trajectory processor constructs a multi-layer perceptron network to learn the three-point curvature of the reference trajectory, expanding the feature dimension of the reference trajectory to [number missing]. : In the formula, It is a non-linear activation function. This represents a multilayer perceptron network. The x-position of the previous reference point in the reference trajectory. This represents the y-direction position of the previous reference point in the reference trajectory. Let x be the current position in the x-direction. This represents the x-direction position of the next reference point in the reference trajectory. This refers to the y-direction position of the next reference point in the reference trajectory; Encoding trajectory features containing curvature information to extract richer representations. : In the formula, For trajectory encoder, This represents the tiling operator. Let be the feature tensor composed of positions in the x-direction. Let be the feature tensor composed of positions along the y-direction. The feature tensor is the sum of all reference points on the current reference trajectory. For position encoding; The dynamics-guided attention module selects key stability features, constructs an attention gate, and learns the dynamic characteristics along the future reference trajectory in the current state as follows: In the formula, These are the possible dynamic weights for motion along the reference trajectory, derived from reasoning based on the current state. It is a Sigmoid non-linear activation function. This represents a key feature encoding network. The feature tensor is composed of longitudinal vehicle speeds. The feature tensor is composed of lateral vehicle speeds. The characteristic tensor composed of yaw angular velocities, The characteristic tensor is composed of lateral acceleration. Let the characteristic tensor be composed of the adhesion coefficients of each wheel. Represents the dot product operator. Trajectory features for embedding dynamic information; The stability-enhanced fusion coding module incorporates key stability features as prior information into the policy generation network. These key features include centroid sideslip angle, tire characteristics, and road adhesion, with specific indices as follows: In the formula, The calculated stability index, The weight of the centroid sideslip angle index, The index of the centroid sideslip angle. The weights of tire characteristic indicators, These are tire performance indicators. As the weight of the road surface adhesion index, For road surface adhesion index, The sideslip angle is the angle of the center of mass. The stability threshold for the centroid sideslip angle. For longitudinal acceleration, It is lateral acceleration. The current road surface adhesion coefficient, It is the acceleration due to gravity. The adhesion coefficient threshold for the tire entering the nonlinear region; Constructing stability gates The stability gate is then subjected to dot product attention with the dynamic prediction along the future trajectory, and the original features are encoded. The fusion is performed to obtain the final fusion feature tensor. : In the formula, For the weight of the fusion gate, To fuse feature tensors, To fuse feature weights, For fusion gate variables, ReLU is a non-linear activation function. The mathematical expression for the multi-branch fusion encoder is: In the formula, These are the reference control decoder network, the safety constraint parameter decoder network, and the model uncertainty decoder network, respectively. These are the output reference control variables, safety control barrier parameters, and model uncertainty tensors, respectively.

[0014] As a preferred technical solution, the vehicle dynamics model, after collaborative optimization by the policy generation network and the differentiable safety protection layer, is expressed as follows: In the formula, For vehicle dynamics model equations, To account for the uncertainty of the model, the state transition matrix, This is the control coefficient matrix. For the control matrix, Here is the state transition matrix. This is the uncertainty matrix of the model.

[0015] As a preferred technical solution, the expert data acquisition method based on multi-objective optimization construction and parameter tuning is specifically as follows: S3.1. Based on the motion control task and the dynamic characteristics of the controlled object, construct the system dynamic equations. : In the formula, These are the state vector and the control vector, respectively. These are the system state transition matrix and the control matrix, respectively. S3.2. Based on the requirements of the motion control task, construct the objective function: In the formula, This represents minimizing the objective function, where the objective function is defined by the control variables. The variables can be direct or indirect, and the goal is to find the optimal control variables by minimizing the combined objective function. , Indicates the first Project target weight, For the target item, The total number of targets to be considered; S3.3. A weight calibration method based on multi-objective optimization and Pareto front, and using Knee point automatic detection, is adopted to determine the multi-objective weight vector, including: sampling several candidate weight vectors in the weight space; running the control strategy for each candidate in a simulation or real environment and recording all objective values; obtaining the Pareto front through non-dominated sorting; automatically identifying the compromise point for the Pareto front using Knee point detection; and performing local fine-tuning on the selected points while considering safety lower bound constraints and weight smoothing regularization.

[0016] As a preferred technical solution, the two-stage training specifically includes: The input to the imitation learning is the expert dataset, the maximum number of training rounds, and the number of training batches. The loss function for imitation learning is constructed as follows. for: In the formula, For For the expectation operator of variables, obey probability distribution Represents the loss function. This is the strategy model for the differentiable security protection layer. For expert controller model, For regularization weights, For parameter regularization terms, This refers to the set of learnable parameters in the differentiable security protection layer; The online fine-tuning performs collaborative optimization of weights and network parameters. The inputs are pre-training parameters, the maximum online training period, and the set of operating conditions. The online training loss is a hybrid performance-safety loss. Represented as: In the formula, For expectation operator, For performance measurement, For the input state variables, For safety reasons, For regularization terms, This is a weighting factor.

[0017] According to another aspect of the present invention, an end-to-end motion control architecture with safety guarantees is provided, the architecture for implementing the above-described method, including a policy generation network and a differentiable safety protection layer, wherein the policy generation network includes a dynamic feature encoder and a multi-branch fusion decoder; The dynamic feature encoder includes a reference trajectory processor, a dynamic guided attention module, and a fusion coding module for dynamic coding and stability enhancement; The multi-branch fusion decoder includes a reference control decoder, a model uncertainty decoder, and a security constraint parameter decoder; The observed data is processed by the strategy generation network to generate output reference control variables and safety control barrier parameters. The generated data is then input into the differentiable safety protection layer to output the optimal control variables, while simultaneously generating constraint gradients for structural optimization.

[0018] Compared with the prior art, the present invention has at least one of the following beneficial effects: (1) The present invention adopts a differentiable safety protection layer, which integrates actuator constraints, road surface adhesion constraints and stability constraints in a differentiable form through differentiable quadratic programming. It embeds the end-to-end control flow in a differentiable form, which solves the problems of overly conservative safety strategies and difficulty in exploring performance boundaries using data-driven methods while ensuring safety in traditional methods. It makes the safety constraints differentiable to upstream network parameters, realizes the real-time adjustment of the safety boundary through gradient backpropagation according to the task conditions, and supports controlled learning and exploration of higher performance strategies while ensuring safety and stability.

[0019] (2) This invention constructs an end-to-end control architecture with inner and outer layers coupled: the inner layer is responsible for solving and ensuring the safety and executability of the output control commands under the current safety constraints, while the outer layer is responsible for optimizing and exploring the performance of the control strategy under the condition of satisfying the safety constraints. There is both explicit parameter interaction between the inner and outer layers and implicit transmission of the safety target to the strategy parameters through the safety-performance coupling gradient. This solves the problem of the difficulty in co-optimizing safety and performance, realizes the approximation of the strategy to the upper bound of the safety-allowed performance, and promotes the technical effect of the system approaching the Pareto efficiency boundary.

[0020] (3) This invention proposes an expert data acquisition process based on multi-objective optimization construction and parameter tuning. Offline imitation learning pre-trains the policy network through expert examples, enabling the end-to-end framework to quickly converge to a feasible solution near the safety-performance Pareto front. Progressive online fine-tuning adjusts the network parameters and multi-objective weights in small steps through task-driven sample updates and online optimization in co-simulation or controlled vehicle environments. This solves the problem of scarce training data and inconsistent example quality in multi-actuator drive-by-wire chassis, ensuring the consistency and high quality of expert examples in multi-objective dimensions. It achieves the technical effect of balancing sample efficiency, robustness and engineering deployability, and is applicable to practical applications in complex multi-actuator control scenarios. Attached Figure Description

[0021] Figure 1 A comparative diagram of different motion control architectures; Figure 2 A schematic diagram illustrating the optimization mechanism of the inner and outer layers of an end-to-end motion control architecture that balances safety and performance; Figure 3 A schematic diagram of an end-to-end vehicle motion control architecture that considers the balance between safety and performance; Figure 4 A schematic diagram of the network architecture generated by the strategy considering vehicle stability; Figure 5 Optimize offline pre-trained algorithm graphs end-to-end for safety and performance; Figure 6 For safety and performance, end-to-end optimization of the online fine-tuning algorithm diagram. Detailed Implementation

[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0023] To address the aforementioned problems in the prior art, this embodiment provides a data-driven chassis safety filtering control method and an end-to-end motion control frame with safety guarantees. Compared with existing common motion control architectures, such as... Figure 1 As shown, the specific differences of the present invention are as follows: Common motion control architectures can be categorized into four types: traditional control schemes based on safety filters, black-box data-driven control schemes, hybrid architectures based on safety learning, and the end-to-end architecture that balances safety and performance proposed in this invention.

[0024] (1) Traditional control schemes based on safety filters typically consist of a model-based controller (e.g., model predictive control, sliding mode control, etc.) and a safety filter. The model-based controller generates a reference control quantity by establishing a system dynamics model and solving a constrained optimal control problem. However, due to parameter disturbances and model uncertainties, it may lose recursive feasibility under certain operating conditions, thus failing to obtain a feasible solution that satisfies the requirement that the future state still lies within the safety set. The safety filter then aims to minimize the deviation from the reference control quantity and resolves the feasible control quantity within the safety set; therefore, the obtained solution is usually a compromise solution that sacrifices some performance to ensure safety. In this framework, both modules are mainly based on mechanism design, and the information flow is unidirectional (the reference control quantity is transmitted from the model-based controller to the safety filter). The gradient information of safety constraints does not propagate upstream, meaning that the model-based controller cannot directly perceive the sensitivity of its output to the degree of violation of safety constraints and can only rely on experience and manual parameter tuning.

[0025] (2) Black-box data-driven control schemes directly generate control strategies using end-to-end neural networks (common network forms include recurrent neural networks, fully connected networks, Transformers, etc.), and the training paradigm is mostly imitation learning or reinforcement learning. The control performance of this type of method is highly dependent on data quality, simulator fidelity, and the design of loss / reward functions; since no mechanistic information is embedded internally, its output is difficult to guarantee to meet safety constraints.

[0026] (3) The control architecture based on security learning adds a security assurance module to the pure data-driven network. This module is similar in form to a traditional security filter, but allows joint training with the policy network. Typically, the security filter is expressed in the form of a quadratic form (or a quadratic term plus a linear term): In the formula, To optimize the obtained safety control parameters, To control variables, For reference control quantity, These are the quadratic and linear terms of the security filter, respectively, determined by the observed state at time t. The decision is made and is subject to the policy generation network weight parameters. Influence, To control the range of constraints on variables, For the set of safety variables formed by differentiable safety layers, This represents the matrix transpose. Only... The gradient update strategy generates network parameters. Constraints in the safety filter are designed through a mechanism, using fixed parameters with the highest priority. Similar to traditional security filter-based methods, it is difficult to fully leverage the advantages of data-driven approaches.

[0027] The generated safety control variables are influenced by the current observation state and the policy network parameters. However, in common implementations, joint training only affects the policy network by updating the objective function parameters (e.g., the weights of quadratic terms), while the safety constraints themselves still have the highest priority in a fixed form designed by the mechanism. This prevents the full utilization of the advantages of data-driven methods in constraint structure adaptation and task-oriented optimization.

[0028] (4) The end-to-end safety-performance balance architecture proposed in this invention achieves the following characteristics by transforming the safety filter into a differentiable module: the policy generation network simultaneously outputs learnable parameters (e.g., constraint "tightening" parameters) of the reference control quantity and safety constraints, and updates the optimization structure of the safety filter in the form of real-time inner layer optimization. Specifically, the inner layer optimization is a learnable task-oriented subproblem (which can be regarded as a constraint optimization problem with learnable parameters), while the outer layer optimization updates the policy network parameters and safety constraint parameters simultaneously by jointly optimizing the loss function. Under the Karush–Kuhn–Tucker (KKT) conditional derivation based on this inner layer optimization problem, the explicit forms of four types of sensitivity / gradient between the inner and outer layers can be obtained. These terms typically include: the gradient of the loss function with respect to the inner layer optimization parameters. The gradient of the loss function with respect to the control input This includes the diagonal matrix and dual sensitivity matrix related to the constraints. It is evident that the safety constraints and performance terms are coupled at the gradient level, allowing the gradient of the safety constraints to propagate back along the optimization path and directly guide the parameter updates of the policy generation network. This enables a more effective exploration of performance limits while maintaining safety. In the formula, For the loss function on the control vector gradient, This represents the optimal control vector currently being solved. For optimal Lagrange multipliers, For duality sensitivity, It is a diagonal matrix. This indicates the matrix transpose.

[0029] The aforementioned data-driven chassis safety filtering control method specifically includes: S1. Establish vehicle dynamics equations. Based on the vehicle dynamics equations, construct actuator action constraints, handling stability constraints, and tire adhesion constraints. Transform the vehicle dynamics equations into control barrier function form. Integrate the constructed constraints into a differentiable safety protection layer through differentiable quadratic programming. S2. Construct a strategy generation network that considers vehicle dynamics equations, and combine the strategy generation network with a differentiable safety protection layer to form an end-to-end joint optimization framework. The end-to-end joint optimization framework forms an inner-outer optimization mechanism internally. S3. Construct an expert data acquisition method based on multi-objective optimization for constructing and tuning training datasets for end-to-end joint optimization framework. The expert data acquisition method includes two-stage training: the first stage adopts imitation learning, and the second stage adopts online fine-tuning.

[0030] The differentiable safety protection layer is a neural network layer that allows backpropagation of safety constraint gradients, enabling upstream modules to determine the safety of the current control strategy and guide them in parameter updates. Furthermore, the differentiable safety protection layer possesses forward invariance, ensuring recursive feasibility under safety constraints. Step S1 specifically includes: S11. For vehicle motion control tasks, construct vehicle dynamics equations; construct a vehicle model considering vehicle handling and stability characteristics and tire nonlinear dynamics, specifically expressed as: In the formula, It is an n-dimensional state vector; Represents the state transition function; Represents the control function. To control variables, For control constraints.

[0031] S12. Based on the established vehicle dynamics equations, actuator motion constraints, handling stability constraints, and tire adhesion constraints are constructed and transformed into control barrier function forms. Finally, the constructed constraints are integrated using differentiable quadratic programming techniques.

[0032] Taking into account the vehicle's actuator constraints, handling stability constraints, and tire adhesion constraints, a safety control barrier function is constructed: The formula is for four-wheel distributed independent drive and steering vehicles. Constraints on four-wheel steering and drive / braking torque, Stability constraints include constraints on the sideslip angle and yaw rate. Apply elliptical constraints to the tire.

[0033] The vehicle dynamics model yields the following specific forms of the higher-order control barrier functions for each constraint: In the formula, Constraints on four-wheel steering and drive / braking torque, Stability constraints include constraints on the sideslip angle and yaw rate. The tire is attached to an elliptical constraint. The relative degrees of the original actuator constraints, stability constraints, and tire adhesion constraints to the vehicle dynamics model. Let the state transition function vector be... For the original constraints of Lie derivative, This is the control coefficient matrix. Original tire adhesion constraint and The mixed Lie derivative, and They are all Li Dao's number operators. Learnable parameters Polynomials with derivatives For learnable parameters, For indexing, To control variables, For control constraints.

[0034] The constructed models are integrated into a differentiable security protection layer using differentiable quadratic programming techniques, specifically in the following form: In the formula, The weight matrix is ​​a positive definite quadratic term. The linear term weight matrix, As a high-dimensional hidden feature, For differentiable quadratic programming network parameters, To optimize the obtained safety control parameters, To control variables, Indicates matrix transpose. For a moment, To solve for the step size using a differentiable quadratic programming problem, At the initial moment, For time indexing, A matrix consisting of the minimum values ​​of the control variables. A matrix consisting of the maximum values ​​of the control variables.

[0035] The policy generation network and the differentiable security protection layer are combined to form an end-to-end joint optimization framework. This framework internally forms an inner-outer layer optimization mechanism, the details of which are as follows: Figure 2 , 3 As shown, its basic principle and function are as follows: Under extreme conditions, common reasons make it difficult for the safety set and the performance Pareto front to coincide or intersect: on the one hand, due to factors such as model uncertainty, actuator saturation, and environmental disturbances, the optimal control solution often lies outside the safety set; on the other hand, if the safety set is defined too conservatively, it will severely compress the feasible performance space, preventing the Pareto front from entering the safety set.

[0036] To ensure the safety of real-time control, this invention employs inner-layer optimization as a "safety guarantor." The policy generation network, with its data-driven structure, allows gradient backpropagation. The reference control variables it generates are typically distributed near the Pareto front due to differences in training distribution and network uncertainties. Compared to model-based MPC, the latter's solutions are usually located at the boundary of the constraint set but cannot guarantee recursive feasibility. Inner-layer optimization, constructed based on the Control Barrier Function (CBF) and its forward invariance theory, is used to solve for control variables that satisfy higher-order safety sets (determined by learnable constraint parameters p, which control the tightness of the constraints) on an instantaneous basis, thus ensuring that the control variables solved online always reside within the safety set.

[0037] The outer optimization layer performs performance exploration within the safety framework: on one hand, it backpropagates the sensitivity of safety constraints to control and network parameters (given by the inner KKT conditions) to the policy generation network, ensuring the policy network updates in a direction that improves performance without violating safety constraints; on the other hand, it directly optimizes network parameters using a performance-related objective function. This end-to-end joint optimization achieves the coupling of safety and performance gradients. The gradient information of safety constraints and the performance objective jointly constrain parameter updates, forming a task-driven, differentiable safety protection mechanism. As training progresses, the Pareto front's position in the parameter space gradually moves towards the higher-order safety set until it overlaps with it. At this point, the solved control quantity represents the control policy that achieves optimal performance under safety constraints.

[0038] The specific construction steps of the data-driven chassis safety filtering control method provided by this invention are as follows: S21. Construct a policy generation network that considers vehicle dynamics. The policy generation network structure is as follows: Figure 4 As shown, specifically: (1) Construct a dynamic feature encoder, which mainly includes a reference trajectory processor, a dynamic guidance attention module, and a fusion coding module for dynamic coding and stability enhancement.

[0039] The curvature of the reference trajectory affects the difficulty of vehicle motion control; a high curvature reference trajectory increases the risk of vehicle instability. To improve the model's ability to anticipate reference trajectories with different curvatures, a data-driven approach is used to help the model understand the difficulty of the motion control task. First, a multi-layer perceptron network is constructed to learn the three-point curvature of the reference trajectory, expanding the feature dimension of the reference trajectory to [missing information]. . In the formula, It is a non-linear activation function. This represents a multilayer perceptron network. The x-position of the previous reference point in the reference trajectory. This represents the y-direction position of the previous reference point in the reference trajectory. Let x be the current position in the x-direction. This represents the x-direction position of the next reference point in the reference trajectory. This refers to the y-direction position of the next reference point in the reference trajectory; Encoding trajectory features containing curvature information to extract richer representations. : In the formula, For trajectory encoder, This represents the tiling operator. Let be the feature tensor composed of positions in the x-direction. Let be the feature tensor composed of positions along the y-direction. The feature tensor is the sum of all reference points on the current reference trajectory. For location encoding, since the distance relationship is lost after the trajectory features are tiled, location encoding is added to ensure that distance information is not lost; The dynamics attention module selects longitudinal vehicle speed, lateral vehicle speed, yaw rate, lateral acceleration, and road adhesion coefficient as key features affecting vehicle handling characteristics through feature optimization from motion observations. It then constructs an attention gate and learns the dynamic characteristics of the vehicle along a future reference trajectory in the current state. In the formula, These are the possible dynamic weights for motion along the reference trajectory, derived from reasoning based on the current state. It is a Sigmoid non-linear activation function. This represents a key feature encoding network. The feature tensor is composed of longitudinal vehicle speeds. The feature tensor is composed of lateral vehicle speeds. The characteristic tensor composed of yaw angular velocities, The characteristic tensor is composed of lateral acceleration. Let the characteristic tensor be composed of the adhesion coefficients of each wheel. Represents the dot product operator. Trajectory features for embedding dynamic information; The stability-enhanced fusion coding module incorporates key stability features as prior information into the network, guiding it to focus on vehicle motion stability. These key features include center-of-gravity sideslip angle, tire characteristics, and road surface adhesion, with specific indices as follows: In the formula, The calculated stability index, The weight of the centroid sideslip angle index, The index of the centroid sideslip angle. The weights of tire characteristic indicators, These are tire performance indicators. As the weight of the road surface adhesion index, For road surface adhesion index, The sideslip angle is the angle of the center of mass. The stability threshold for the centroid sideslip angle. For longitudinal acceleration, It is lateral acceleration. The current road surface adhesion coefficient, It is the acceleration due to gravity. The adhesion coefficient threshold for the tire entering the nonlinear region; Constructing stability gates This is combined with a dot product attention function on the dynamic predictions along the future trajectory, and the original features are encoded. The fusion is performed to obtain the final fusion feature tensor. : In the formula, For the weight of the fusion gate, To fuse feature tensors, To fuse feature weights, For fusion gate variables, ReLU is a non-linear activation function. (2) Construct a multi-branch encoder, which mainly includes a reference control decoder, a model uncertainty decoder, and a safety constraint parameter decoder.

[0040] S22. Connect the policy generation network and the end-to-end differentiable security protection layer to form an end-to-end joint optimization architecture, specifically: In the formula, These consist of a reference control decoder network, a safety constraint parameter decoder network, and a model uncertainty decoder network, respectively. The parameters of each decoder are not shared. These are the output reference control variables, safety control barrier parameters, and model uncertainty tensor, respectively. Learnable parameters for various safety constraints constructed for vehicle dynamics A set of.

[0041] The vehicle dynamics model, after co-optimization by the policy generation network and the differentiable safety protection layer, is expressed as: In the formula, For vehicle dynamics model equations, To account for the uncertainty of the model, the state transition matrix, This is the control coefficient matrix. For the control matrix, Here is the state transition matrix. This is the uncertainty matrix of the model.

[0042] The cost function of a differentiable optimization problem is to minimize the distance between the optimization control variable and the reference control, i.e.: The expert data collection and training methods are as follows: S31. Design an expert data acquisition method based on multi-objective optimization and parameter tuning for constructing a training dataset for a safety-performance balanced end-to-end motion control architecture, specifically including: First, based on the motion control task and the dynamic characteristics of the controlled object, the system dynamic equations are constructed: In the formula, These are the state vector and the control vector, respectively. These are the system state transition matrix and the control matrix, respectively; for complex control systems, It exhibits time-varying and nonlinear characteristics and is a function matrix.

[0043] Then, based on the requirements of the motion control task, the objective function is constructed: In the formula, This represents minimizing the objective function, where the objective function is defined by the control variables. The variables can be direct or indirect, and the goal is to find the optimal control variables by minimizing the combined objective function. , Indicates the first Project target weight, For the target item, This represents the total number of targets that need to be considered.

[0044] Secondly, a weight calibration method based on multi-objective optimization and the Pareto front, and employing automatic Knee point detection, is used to determine the multi-objective weight vector. The basic idea is as follows: sample several candidate weight vectors in the weight space; run the control strategy for each candidate in a simulation or real environment and record all objective values; obtain the Pareto front through non-dominated sorting; automatically identify the compromise point on the Pareto front using Knee point detection (e.g., based on maximum curvature or maximum vertical distance between connection endpoints); perform local fine-tuning on the selected points (using Bayesian optimization or gradient approximation) and consider safety lower bound constraints and weight smoothing regularization, finally outputting the weight vector for use by the actual controller. The specific operation steps are as follows: 1) Represent weights and constraints in a parametric form (e.g., softmax). It can also set lower bounds for weights of security-related objectives; 2) Generate candidate vectors in the weight space using grid, Latin hypercube, random, or adaptive sampling strategies; 3) Evaluate the candidate weights on the simulated or real controller, record the multi-objective values, and perform non-dominated sorting to construct the Pareto front; 4) Perform Knee point detection on the Pareto front, including but not limited to (a) detection of the point of maximum curvature; (b) detection of the maximum vertical distance (maximum convex-back distance) between the endpoints; and (c) detection based on the second derivative or inflection point. 5) Locally fine-tune the weights of candidate Knee points (e.g., Bayesian optimization, finite difference gradient, or local search) and add regularization terms and smoothing mechanisms. If the safety lower bound is violated, back off or project. 6) The final weights are saved and can be updated in a low-frequency adaptive manner based on environmental characteristics during the online phase, or the environment can be mapped to the weights through a small policy network.

[0045] S32. Design a two-stage training algorithm. The first stage uses imitation learning as a guarantee of basic motion control capabilities, and the second stage uses online fine-tuning, so that the proposed architecture can explore higher performance under the premise of ensuring safety.

[0046] Phase one involves offline imitation pre-training, such as... Figure 5 As shown, the goal of this training phase is to establish a robust initial policy using expert examples. This ensures basic control capabilities and provides a reasonable initialization for online fine-tuning. The algorithm input is an expert dataset. Maximum number of training rounds and training batches. The algorithm constructs the offline imitation loss as: In the formula, For For the expectation operator of variables, obey probability distribution The loss function is defined according to the task. A strategy model for the formation of differentiable security protection layers. For expert controller model, For regularization weights, For parameter regularization terms, It is the set of learnable parameters in a differentiable safety protection layer.

[0047] Phase two involves online fine-tuning, such as... Figure 6 As shown. The goal of this training phase is to improve performance while meeting safety lower bounds by fine-tuning strategies under more realistic and diverse operating conditions in a controlled simulation / real vehicle environment. Online fine-tuning allows for collaborative optimization of weights and network parameters. The algorithm input consists of pre-trained parameters. Maximum online training cycle and the set of operating conditions. Define the online training loss as a performance-safety hybrid loss: In the formula, For expectation operator, For performance measurement, For the input state variables, For safety reasons, For regularization terms, This is a weighting factor.

[0048] Online fine-tuning includes inner and outer loops. The outer loop iterates through each working condition in the set of working conditions, while the inner loop runs a co-simulation under each working condition, collects trajectory data and target values, calculates them in each simulation step / batch, and updates parameters based on an optimizer with high sample efficiency.

[0049] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A data-driven chassis safety filtering control method, characterized in that, The method specifically includes: S1. Establish vehicle dynamics equations, and based on the vehicle dynamics equations, construct actuator action constraints, handling stability constraints, and tire adhesion constraints. Transform the vehicle dynamics equations into control barrier function form, and integrate the constructed constraints into a differentiable safety protection layer through differentiable quadratic programming. S2. Construct a strategy generation network that considers the vehicle dynamics equations, and combine the strategy generation network with the differentiable safety protection layer to form an end-to-end joint optimization framework, wherein the end-to-end joint optimization framework forms an inner-outer optimization mechanism internally; S3. Construct an expert data acquisition method based on multi-objective optimization for constructing and tuning the training dataset of the end-to-end joint optimization framework. The expert data acquisition method includes two-stage training: the first stage adopts imitation learning, and the second stage adopts online fine-tuning.

2. The data-driven chassis safety filtering control method according to claim 1, characterized in that, The vehicle dynamics equations Expressed as: In the formula, It is an n-dimensional state vector; Represents the state transition function; Represents the control function. To control variables, For control constraints.

3. The data-driven chassis safety filtering control method according to claim 1, characterized in that, The specific forms of the higher-order control barrier functions for each constraint obtained from the vehicle dynamics model are as follows: In the formula, Constraints on four-wheel steering and drive / braking torque, Stability constraints include constraints on the sideslip angle and yaw rate. The tire is attached to an elliptical constraint. The relative degrees of the original actuator constraints, stability constraints, and tire adhesion constraints with respect to the vehicle dynamics model. Let the state transition function vector be... For the original constraints of Lie derivative, This is the control coefficient matrix. Original tire adhesion constraint and The mixed Lie derivative, and They are all Li Dao's number operators. Learnable parameters Polynomials with derivatives For learnable parameters, For indexing, To control variables, For control constraints.

4. The data-driven chassis safety filtering control method according to claim 3, characterized in that, The aforementioned constraint high-order control barrier function, integrated into a differentiable safety protection layer through differentiable quadratic programming, takes the following specific form: In the formula, The weight matrix is ​​a positive definite quadratic term. The linear term weight matrix, As a high-dimensional hidden feature, For differentiable quadratic programming network parameters, To optimize the obtained safety control parameters, To control variables, Indicates matrix transpose. For a moment, To solve for the step size using a differentiable quadratic programming problem, At the initial moment, For time indexing, A matrix consisting of the minimum values ​​of the control variables. A matrix consisting of the maximum values ​​of the control variables.

5. The data-driven chassis safety filtering control method according to claim 1, characterized in that, The gradient of the control architecture loss function with respect to the inner optimization parameters Specifically, this manifests as follows: In the formula, For the loss function on the control vector gradient, This represents the optimal control vector currently being solved. For optimal Lagrange multipliers, For duality sensitivity, It is a diagonal matrix. This indicates the matrix transpose.

6. The data-driven chassis safety filtering control method according to claim 1, characterized in that, The policy generation network includes a dynamic feature encoder and a multi-branch fusion decoder; The dynamic feature encoder includes a reference trajectory processor, a dynamic guided attention module, and a fusion coding module for dynamic coding and stability enhancement; The multi-branch fusion decoder includes a reference control decoder, a model uncertainty decoder, and a security constraint parameter decoder; The reference trajectory processor constructs a multi-layer perceptron network to learn the three-point curvature of the reference trajectory, expanding the feature dimension of the reference trajectory to [number missing]. : In the formula, It is a non-linear activation function. This represents a multilayer perceptron network. The x-position of the previous reference point in the reference trajectory. This represents the y-direction position of the previous reference point in the reference trajectory. Let x be the current position in the x direction. This represents the x-direction position of the next reference point in the reference trajectory. This refers to the y-direction position of the next reference point in the reference trajectory; Encoding trajectory features containing curvature information to extract richer representations. : In the formula, For trajectory encoder, This represents the tiling operator. Let be the feature tensor composed of positions in the x-direction. Let be the feature tensor composed of positions in the y-direction. The feature tensor is the sum of all reference points on the current reference trajectory. For position encoding; The dynamics-guided attention module selects key stability features, constructs an attention gate, and learns the dynamic characteristics along the future reference trajectory in the current state as follows: In the formula, These are the possible dynamic weights for motion along the reference trajectory, derived from reasoning based on the current state. It is a Sigmoid non-linear activation function. This represents a key feature encoding network. The feature tensor is composed of longitudinal vehicle speeds. Let be the feature tensor composed of lateral vehicle speeds. The characteristic tensor composed of yaw angular velocity, The characteristic tensor is composed of lateral acceleration. Let the characteristic tensor be composed of the adhesion coefficients of each wheel. Represents the dot product operator. Trajectory features for embedding dynamic information; The stability-enhanced fusion coding module incorporates key stability features as prior information into the policy generation network. These key features include centroid sideslip angle, tire characteristics, and road adhesion, with specific indices as follows: In the formula, The calculated stability index, The weight of the centroid sideslip angle index, The index of the centroid sideslip angle. The weights of tire characteristic indicators, These are tire performance indicators. As the weight of the road surface adhesion index, For road surface adhesion index, The sideslip angle is the angle of the centroid. The stability threshold for the centroid sideslip angle. For longitudinal acceleration, It is lateral acceleration. The current road surface adhesion coefficient, It is the acceleration due to gravity. The adhesion coefficient threshold for the tire entering the nonlinear region; Constructing stability gates The stability gate is then subjected to dot product attention with the dynamic prediction along the future trajectory, and the original features are encoded. The fusion is performed to obtain the final fusion feature tensor. : In the formula, For the weight of the fusion gate, To fuse feature tensors, To fuse feature weights, For fusion gate variables, ReLU is a non-linear activation function. The mathematical expression for the multi-branch fusion encoder is: In the formula, These are the reference control decoder network, the safety constraint parameter decoder network, and the model uncertainty decoder network, respectively. These are the output reference control variables, safety control barrier parameters, and model uncertainty tensors, respectively.

7. The data-driven chassis safety filtering control method according to claim 1, characterized in that, The vehicle dynamics model, after being co-optimized by the policy generation network and the differentiable safety protection layer, is expressed as follows: In the formula, For vehicle dynamics model equations, To account for the uncertainty of the model, the state transition matrix, This is the control coefficient matrix. For the control matrix, Here is the state transition matrix. This is the uncertainty matrix of the model.

8. The data-driven chassis safety filtering control method according to claim 1, characterized in that, The expert data acquisition method based on multi-objective optimization construction and parameter tuning is as follows: S3.

1. Based on the motion control task and the dynamic characteristics of the controlled object, construct the system dynamic equations. : In the formula, These are the state vector and the control vector, respectively. It is all single variables The set of states, and similarly the state vectors. These are the system state transition matrix and the control matrix, respectively. S3.

2. Based on the requirements of the motion control task, construct the objective function: In the formula, This represents minimizing the objective function, where the objective function is defined by the control variables. The variables can be direct or indirect, and the goal is to find the optimal control variables by minimizing the combined objective function. , Indicates the first Project target weight, For the target item, The total number of targets to be considered; S3.

3. A weight calibration method based on multi-objective optimization and Pareto front, and using Knee point automatic detection, is adopted to determine the multi-objective weight vector, including: sampling several candidate weight vectors in the weight space; running the control strategy for each candidate in a simulation or real environment and recording all objective values; The Pareto front is obtained through non-dominated sorting; Knee point detection is used to automatically identify the compromise point on the Pareto front; local fine-tuning is performed on the selected points, and safety lower bound constraints and weight smoothing regularization are considered.

9. The data-driven chassis safety filtering control method according to claim 1, characterized in that, The two-stage training is as follows: The input to the imitation learning is the expert dataset, the maximum number of training rounds, and the number of training batches. The loss function for imitation learning is constructed as follows. for: In the formula, For For the expectation operator of variables, obey probability distribution Represents the loss function. This is the strategy model for the differentiable security protection layer. For expert controller model, For regularization weights, For parameter regularization terms, This refers to the set of learnable parameters in the differentiable security protection layer; The online fine-tuning performs collaborative optimization of weights and network parameters. The inputs are pre-training parameters, the maximum online training period, and the set of operating conditions. The online training loss is a hybrid performance-safety loss. Represented as: In the formula, For expectation operator, For performance measurement, For the input state variables, For safety reasons, For regularization terms, This is a weighting factor.

10. An end-to-end motion control architecture with safety guarantees, characterized in that, The architecture is used to implement the method of any one of claims 1-9, including a policy generation network and a differentiable security protection layer, wherein the policy generation network includes a dynamic feature encoder and a multi-branch fusion decoder; The dynamic feature encoder includes a reference trajectory processor, a dynamic guided attention module, and a fusion coding module for dynamic coding and stability enhancement; The multi-branch fusion decoder includes a reference control decoder, a model uncertainty decoder, and a security constraint parameter decoder; The observed data is processed by the strategy generation network to generate output reference control variables and safety control barrier parameters. The generated data is then input into the differentiable safety protection layer to output the optimal control variables, while simultaneously generating constraint gradients for structural optimization.

Citation Information

Patent Citations

  • Modelica language-based large model driven automobile model modeling method

    CN120850747A