Active suspension system, control method and storage medium

By combining heterogeneous master-slave controllers and lidar pre-aiming modules, the suspension damping is dynamically adjusted, solving the problems of single-point failure and unstable switching between homogeneous dual controllers in the active suspension system, and achieving safety, reliability and stability under complex working conditions.

CN120863271APending Publication Date: 2025-10-31SHANGHAI YUANYUQING NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511194229.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

The existing active suspension system control architecture has the problem of single point of failure risk and difficulty in reliably switching between homogeneous dual controllers when there is targeted interference or systemic risk, which can lead to inaccurate or uncontrollable vehicle handling and affect driving safety.

Method used

It employs heterogeneous master and slave controllers, combined with lidar pre-aiming modules and status sensors, and dynamically adjusts suspension damping through edge computing controllers. It configures different control algorithms and outputs control commands through independent paths, and achieves seamless switching by combining command switching units to avoid hardware and software failures.

Benefits of technology

It improves the safety performance of the active suspension system, avoids the problems of single point of failure and simultaneous failure of the same dual controllers at the software logic level, ensures seamless switching at the moment of failure, and enhances the safety and stability of the vehicle under complex working conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120863271A_ABST
    Figure CN120863271A_ABST
Patent Text Reader

Abstract

The invention provides an active suspension system, a control method and a storage medium, the active suspension system is provided with an edge computing controller, the edge computing controller comprises a master controller, a slave controller, a monitoring unit and an instruction switching unit, and the master controller and the slave controller are heterogeneous on the hardware level, so that the risk of common hardware defects is avoided, and the reliability of the active suspension system is improved. Besides, the master controller and the slave controller are further configured to run different control algorithms and output control instructions at the same time, so that on one hand, faults triggered by same software vulnerabilities can be avoided, and on the other hand, the control instructions output by the controller without faults can be seamlessly switched to through the instruction switching unit when faults occur. Therefore, the problems that a single-point fault risk is faced by a single-controller architecture, architecture schemes of isomorphic double controllers possibly fail at the same time on the software logic level, and reliable switching is difficult to complete at the moment that a fault occurs can be solved, and the safety is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the field of automotive suspension technology, and more specifically, to an active suspension system, control method, and storage medium. Background Technology

[0002] As the automotive industry rapidly develops towards intelligence and connectivity, the handling stability, ride comfort, and driving safety during vehicle operation have become core indicators for measuring vehicle performance. As a key component for adjusting vehicle dynamics, the performance of the active suspension system directly determines the vehicle's overall performance under complex road conditions. Therefore, extremely high requirements are placed on the control reliability and safety redundancy of the active suspension system.

[0003] Currently, most mainstream active suspension systems employ a single controller or a homogeneous dual controller design. The single controller architecture carries a significant risk of single-point failure. When the controller experiences hardware failure or software malfunction, the active suspension system loses control and cannot adjust suspension stiffness and damping according to road conditions. This can lead to minor issues like vehicle bumps and inaccurate handling, or even complete loss of vehicle control, posing a serious threat to the lives of passengers.

[0004] While the homogeneous dual-controller architecture improves redundancy to some extent through dual hardware backup, the two controllers are susceptible to the same failure sources (such as common hardware defects or failures triggered by the same software vulnerabilities) because their hardware structures and software logic are completely identical. When faced with targeted interference or systemic risks, both controllers may fail simultaneously, failing to achieve true safety redundancy. Furthermore, the existing switching mechanisms of homogeneous dual controllers mostly rely on periodic fault detection, which has a detection delay and makes it difficult to complete a reliable switchover at the moment a fault occurs, still resulting in brief control interruptions.

[0005] Therefore, the industry urgently needs an active suspension system solution that can achieve highly reliable redundant control to meet the safe operation requirements of vehicles under complex working conditions. Summary of the Invention

[0006] The following is an overview of the subject matter described in detail herein. This overview is not intended to limit the scope of the claims.

[0007] This invention aims to at least solve one of the technical problems existing in the prior art. To this end, this invention proposes an active suspension system, control method, and storage medium, which can solve the single-point failure risk faced by a single controller architecture and the problem that in a homogeneous dual-controller architecture, when encountering targeted interference or systemic risks, the two controllers may fail simultaneously at the software logic level, and it is difficult to reliably switch over at the moment of failure. At the same time, it combines a LiDAR-based pre-aiming function to further improve the safety performance of the active suspension.

[0008] In a first aspect, embodiments of the present invention provide an active suspension system, comprising:

[0009] The lidar pre-aiming module is used to obtain the elevation of the road surface in front of the vehicle in order to perceive the road conditions ahead of the vehicle;

[0010] The status sensor module includes a six-dimensional force sensor, a displacement sensor, and an inertial sensor. The status sensor module is used to sense the status of the vehicle body.

[0011] Dynamic electromagnetic actuators are used to dynamically adjust the damping characteristics of the suspension or apply active force.

[0012] The edge computing controller is electrically connected to the lidar aiming module, the status sensor module, and the dynamic electromagnetic actuator.

[0013] The edge computing controller includes a master controller, slave controllers, a monitoring unit, and an instruction switching unit.

[0014] The master controller and slave controller are heterogeneous, and the master controller and slave controller are configured to synchronously receive data from the lidar pre-aiming module and the six-dimensional force sensor and run different control algorithms to output control commands simultaneously;

[0015] The monitoring unit is connected to the main controller, slave controller, and command switching unit, which are equipped with independent interfaces, through different hardware pins. The monitoring unit is used to verify the control commands output by the main controller and, according to preset conditions, the control command switching unit selects the control commands output by the main controller or the slave controller so that the dynamic electromagnetic actuator adjusts the damping characteristics of the suspension or applies active force according to the selected control commands.

[0016] In one embodiment, the dynamic electromagnetic actuator includes an electromagnetic clutch, and a linear motor and a magnetorheological damper connected in parallel via the electromagnetic clutch, the linear motor and the magnetorheological damper being connected together to the same suspension link of the vehicle to provide dynamic main force and auxiliary damping, respectively.

[0017] The linear motor has a magnetic shielding layer on its outer wall, and the linear motor and the magnetorheological damper are configured for time-sharing control to eliminate the superposition of their magnetic fields.

[0018] The active suspension system according to the first aspect of the present invention has at least the following beneficial effects:

[0019] The active suspension system of the first aspect of this invention includes an edge computing controller for dynamically adjusting suspension damping characteristics by controlling a dynamic electromagnetic actuator based on a lidar pre-aiming module and a state sensor module. The edge computing controller includes a master controller, a slave controller, a monitoring unit, and a command switching unit. The master and slave controllers are heterogeneous at the hardware level, thus avoiding the risk of common hardware defects. Furthermore, the master and slave controllers are configured to run different control algorithms and simultaneously output control commands. Therefore, on the one hand, it avoids failures triggered by the same software vulnerability; on the other hand, it allows seamless switching to the control commands output by the unfailed controller via the command switching unit in the event of a failure. The command switching unit is connected to the master and slave controllers via independent physical paths. Therefore, the active suspension system of this invention can solve the single-point-of-failure risk faced by a single-controller architecture and the problem that in a homogeneous dual-controller architecture, both controllers may fail simultaneously at the software logic level when encountering targeted interference or systemic risks, and it is difficult to reliably switch over instantly upon failure. Combined with the lidar-based pre-aiming function, it further enhances the safety performance of the active suspension.

[0020] In a first aspect, embodiments of the present invention provide an active suspension control method, applied to an active suspension system as described in the first aspect of the present invention, comprising:

[0021] Initial parameters are obtained and standardized preprocessing is performed to obtain state parameters. The state parameters are used to characterize the road conditions and vehicle body status in front of the vehicle as perceived by the lidar aiming module and the sensor module.

[0022] The state parameters are input to the preset first control model and second control model, and the first control instruction and second control instruction with the same format are obtained respectively. The first control model is configured on the master controller and the second control model is configured on the slave controller. The architectures of the first control model and the second control model are different.

[0023] The first and second control commands are verified according to the preset gating conditions, and the first or second control command is selected as the output control command to control the dynamic electromagnetic actuator to adjust the damping characteristics of the suspension or apply the active force.

[0024] In one embodiment, the first control model includes:

[0025] The main network group consists of a single main decision network and two main evaluation networks. The main evaluation networks are used to guide the main decision network in optimization.

[0026] The target network group, corresponding to the network type and number in the main network group, includes a single target decision network and two target evaluation networks. The target evaluation network is used to guide the target decision network in optimization.

[0027] The main network group is configured to update in real time for real-time decision-making, while the target network group is configured to synchronize slowly with the main network via soft updates to serve as a benchmark during the main network update process and prevent the main network from oscillating during the update process.

[0028] In one embodiment, the first control model further includes a playback buffer for subsequent batch sampling training. State parameters are input to a preset first control model and a second control model to obtain first control instructions and second control instructions of the same format, including:

[0029] The state parameters are input into the main decision network, and Gaussian noise is added to the output of the main decision network to obtain the initial action command;

[0030] The initial action command is verified according to the preset safety constraints, and if the initial action command passes the verification, the initial action command is used as the first control command to control the dynamic electromagnetic actuator and obtain the corresponding feedback signal.

[0031] The initial action command actually applied to the dynamic electromagnetic actuator, along with the corresponding feedback signal and state parameters, are stored as empirical data in the playback buffer.

[0032] Randomly sample empirical data pairs from the playback buffer, input the empirical data pairs into the target decision network, and add clipped Gaussian noise to the output of the target decision network to obtain the target action command;

[0033] Update the main evaluation network and the main decision-making network based on empirical data and target action instructions;

[0034] Based on a preset soft update method, the target evaluation network and target decision network are updated according to the main evaluation network and the main decision network.

[0035] In one embodiment, the state space of the state parameters is set to have multiple dimensions. Each dimension of the state space of the state parameters is used to characterize the vertical acceleration of the vehicle body, the vertical velocity of the vehicle body, the rate of change of the vehicle body acceleration, the suspension dynamic travel, the relative displacement between the tire and the road surface, the relative velocity between the vehicle body and the wheel, the constraint on the output force of the dynamic electromagnetic actuator, the pre-aimed road surface elevation, and the rate of change of the pre-aimed road surface elevation.

[0036] The feedback signal includes a reward function, which is constructed based on the state parameters, and the state space of the reward function corresponds to the state space of the state parameters.

[0037] The feedback signal also includes the state parameters that are reacquired at the next moment after the initial action command is actually applied to the electromagnetic actuator.

[0038] In one embodiment, updating the main evaluation network and the main decision network based on empirical data and target action instructions includes:

[0039] The empirical data pair and the target action command are input into two target evaluation networks. The minimum value of the output of the two target evaluation networks is selected and combined with the empirical data pair to obtain the target value estimate.

[0040] The empirical data pairs and the target action instructions are input into two main evaluation networks to obtain two main value estimates.

[0041] An evaluation loss function is constructed based on the mean squared error of the two principal value estimates and the target value estimate, and the principal evaluation network is updated using the gradient descent method based on the evaluation loss function.

[0042] Select the principal value estimate output by a fixed principal evaluation network from the two principal evaluation networks as the benchmark, construct a negative loss function based on the principal value estimate used as the benchmark, and update the principal decision network based on the negative loss function using the gradient descent method.

[0043] In one embodiment, updating the target evaluation network and the target decision network based on a preset soft update method according to the main evaluation network and the main decision network includes:

[0044] When the number of updates to the main network group reaches the preset update threshold, the corresponding target evaluation network is updated according to the preset soft update coefficient and the updated main evaluation network based on the evaluation loss function.

[0045] The target decision network is updated based on the soft update coefficients and the updated master decision network based on the negative loss function.

[0046] In one embodiment, verifying the first control instruction and the second control instruction according to preset gating conditions, and selecting the first control instruction or the second control instruction as the output control instruction, includes:

[0047] The deviation value is obtained based on the first control command and the second control command;

[0048] If the deviation value is less than the preset deviation limit, the first control command is selected as the output control command.

[0049] If the deviation value is greater than the deviation limit, the second control command is selected as the output control command until more than a preset number of deviation values ​​are less than the deviation limit, then the first control command is selected as the output control command again.

[0050] When the second control command is selected as the output control command for the first time, the second control model is initialized according to the state parameters corresponding to the current first control command in order to avoid control abrupt changes.

[0051] The active suspension control method according to a second aspect of the present invention has at least the following beneficial effects:

[0052] The active suspension control method of the second aspect of the present invention is provided with a first control model and a second control model with different architectures. The first control model and the second control model are respectively configured on the master controller and the slave controller to simultaneously output the first control command and the second control command, and select a single command as the final actual output after verification to control the dynamic electromagnetic actuator to adjust the suspension damping characteristics or apply active force. Since the master controller and the slave controller are configured to run different control algorithms and output control commands simultaneously, it can avoid failures triggered by the same software vulnerability on the one hand, and seamlessly switch to the control command output by the unfailed controller when a failure occurs on the other hand. The first control command and the second control command are obtained based on the same state parameters and have the same output command format, which further reduces the delay during the switching process and prevents control abrupt changes. Therefore, the active suspension control method of the second aspect of the present invention can solve the problem that traditional homogeneous dual controllers may fail simultaneously at the software logic level when encountering targeted interference or systemic risks, and it is difficult to complete a reliable switch at the moment of failure, thereby improving the safety performance of active suspension.

[0053] Thirdly, embodiments of the present invention also provide a computer-readable storage medium storing computer-executable instructions for causing a computer to perform the active suspension control method as described in the second aspect of the embodiments above.

[0054] The third aspect of the present invention has the same improvements as the second aspect of the present invention, and therefore possesses all the beneficial effects of the second aspect of the present invention, which will not be repeated here.

[0055] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of what is particularly pointed out in the description, claims and drawings. Attached Figure Description

[0056] The accompanying drawings are provided to further understand the technical solutions of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the technical solutions of the present invention, and do not constitute a limitation on the technical solutions of the present invention.

[0057] Figure 1This is a schematic diagram of the architecture of an active suspension system provided in one embodiment of the present invention;

[0058] Figure 2 yes Figure 1 A schematic diagram of the architecture of the edge computing controller;

[0059] Figure 3 This is a flowchart of an active suspension control method provided in one embodiment of the present invention;

[0060] Figure 4 yes Figure 3 The detailed flowchart of step S200;

[0061] Figure 5 yes Figure 4 The detailed flowchart of step S250;

[0062] Figure 6 yes Figure 5 The detailed flowchart of step S260;

[0063] Figure 7 yes Figure 3 The detailed flowchart of step S300. Detailed Implementation

[0064] The technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. These embodiments are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0065] It should be noted that in the description of this disclosure, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" and "second" may explicitly or implicitly include one or more features.

[0066] Furthermore, although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0067] It should also be noted that if the embodiments of this disclosure involve directional indicators, such as up, down, left, right, front, back, etc., then the directional indicators are only used to explain the relative positional relationship and movement of the components in a certain specific posture (e.g., as shown in the figure). If the specific posture changes, the directional indicators should also change accordingly.

[0068] Furthermore, unless otherwise explicitly specified and limited, the term "connection / linkage" should be interpreted broadly, for example, it can be a fixed connection or a movable connection, a detachable connection or a non-detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection or a connection that allows communication between the two components; it can be a direct connection or an indirect connection through an intermediate medium; it can be a connection within two components or an interaction between two components.

[0069] Finally, in the description of this disclosure, references to terms such as "one embodiment / implementation," "another embodiment / implementation," or "some embodiments / implementations," etc., indicate that a specific feature, structure, material, or characteristic described in connection with an embodiment or example is included in at least two embodiments or implementations of this disclosure. In this disclosure, illustrative expressions of the above terms do not necessarily refer to the same embodiment or implementation. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or implementations.

[0070] Currently, most mainstream active suspension systems employ a single controller or a homogeneous dual controller design. The single controller architecture carries a significant risk of single-point failure. When the controller experiences hardware failure or software malfunction, the active suspension system loses control and cannot adjust suspension stiffness and damping according to road conditions. This can lead to minor issues like vehicle bumps and inaccurate handling, or even complete loss of vehicle control, posing a serious threat to the lives of passengers.

[0071] While the homogeneous dual-controller architecture improves redundancy to some extent through dual hardware backup, the two controllers are susceptible to the same failure sources (such as common hardware defects or failures triggered by the same software vulnerabilities) because their hardware structures and software logic are completely identical. When faced with targeted interference or systemic risks, both controllers may fail simultaneously, failing to achieve true safety redundancy. Furthermore, the existing switching mechanisms of homogeneous dual controllers mostly rely on periodic fault detection, which has a detection delay and makes it difficult to complete a reliable switchover at the moment a fault occurs, still resulting in brief control interruptions.

[0072] Focusing solely on the control algorithm level, the core design of existing active control algorithms is concentrated on improving control accuracy and response speed (such as algorithm optimization based on PID and model predictive control), with insufficient consideration for safety. On the one hand, the algorithms lack real-time fault diagnosis and fault tolerance capabilities. When the data collected by sensors (such as inertial sensors and displacement sensors) is abnormal or the actuators (such as electromagnetic actuators) have a response delay, the algorithm cannot identify and adjust the control strategy in time, easily outputting erroneous control commands, resulting in a mismatch between suspension actions and actual road conditions. On the other hand, existing algorithms have not established a sound safety boundary control mechanism. Under extreme conditions (such as high-speed cornering and emergency braking), they may exceed the mechanical tolerance limits of the suspension in pursuit of control effects, causing damage to suspension components and further exacerbating safety risks.

[0073] The shortcomings of existing active suspension systems in terms of control architecture redundancy and algorithm security have become key bottlenecks restricting the intelligent upgrade of vehicles. At the same time, the pre-aiming function realized by combining LiDAR can enable the suspension to respond to road conditions in advance. Therefore, it is necessary to combine the pre-aiming function at the algorithm level to further improve safety, and at the same time, to further improve the safety boundary.

[0074] At the hardware level, this invention includes an edge computing controller for dynamically adjusting suspension damping characteristics by controlling a dynamic electromagnetic actuator based on a lidar pre-aiming module and a status sensor. The edge computing controller comprises a master controller, a slave controller, a monitoring unit, and a command switching unit. The master and slave controllers are heterogeneous at the hardware level, thus avoiding the risk of common hardware defects. Furthermore, the master and slave controllers are configured to run different control algorithms and simultaneously output control commands. This avoids faults triggered by the same software vulnerability and allows seamless switching to the control commands output by the unaffected controller in the event of a fault, via the command switching unit. The command switching unit is connected to the master and slave controllers via independent physical paths. Therefore, this active suspension system addresses the single-point-of-failure risk of a single-controller architecture and the problem of simultaneous software logic failure and reliable switching between the two controllers in a homogeneous dual-controller architecture when encountering targeted interference or systemic risks. Combined with lidar-based pre-aiming functionality, this further enhances the safety performance of the active suspension.

[0075] At the software level, a first control model and a second control model with different architectures are set up. The first control model and the second control model are respectively configured on the master controller and the slave controller to simultaneously output the first control command and the second control command. After verification, a single command is selected as the final actual output to control the dynamic electromagnetic actuator to adjust the suspension damping characteristics or apply active force. Since the master controller and the slave controller are configured to run different control algorithms and output control commands simultaneously, on the one hand, it can avoid failures triggered by the same software vulnerability, and on the other hand, it can seamlessly switch to the control command output by the unfailed controller when a failure occurs. The first control command and the second control command are obtained based on the same state parameters and have the same output command format, which further reduces the delay during the switching process and prevents control abrupt changes. Therefore, the active suspension control method of the second aspect embodiment of the present invention can solve the problem that when encountering targeted interference or systemic risks, the traditional homogeneous dual controllers may fail simultaneously at the software logic level and it is difficult to complete a reliable switch at the moment of failure, thereby improving the safety performance of the active suspension.

[0076] Furthermore, the state space corresponding to the first control model in this invention expands the state dimension beyond the core of suspension dynamic characteristics to include the dimensions of the pre-aimed road surface elevation and the rate of change of the pre-aimed road surface elevation. It also expands the dimension related to safety constraints as a penalty term and combines it with additionally set physical safety constraint verification to further improve the safety performance of the control model, thereby further improving the safety performance of the active suspension.

[0077] The embodiments of the present invention will be further described below with reference to the accompanying drawings.

[0078] like Figure 1 As shown, Figure 1 This is a schematic diagram of the architecture of an active suspension system 100 provided in one embodiment of the present invention. Figure 1 In the examples, the active suspension system 100 of this embodiment includes, but is not limited to:

[0079] The lidar pre-aiming module 110 is used to obtain the elevation of the road surface in front of the vehicle in order to perceive the road conditions in front of the vehicle.

[0080] Status sensor module 120 is used to sense the status of the vehicle body;

[0081] The dynamic electromagnetic actuator 140 is used to dynamically adjust the damping characteristics of the suspension or apply active force.

[0082] The edge computing controller 130 is electrically connected to the lidar aiming module 110, the status sensor module 120, and the dynamic electromagnetic actuator 140.

[0083] Specifically, the state sensor module 120 includes a six-dimensional force sensor, a displacement sensor, and an inertial sensor.

[0084] Understandably, the state parameters constructed based on data collected by the lidar pre-aiming module 110, the six-dimensional force sensor, the displacement sensor, and the inertial sensor can simultaneously characterize the vehicle's vertical acceleration, vehicle's vertical velocity, the rate of change of vehicle acceleration, suspension dynamic travel, relative displacement between the tires and the road surface, relative velocity between the vehicle and the wheels, constraints on the output force of the dynamic electromagnetic actuator 140, pre-aimed road surface elevation, and the rate of change of pre-aimed road surface elevation. Therefore, the state parameters contain the core perception dimension of the suspension's dynamic characteristics, the perception dimension that ensures the safe operation of the suspension and tires, and future environmental information introduced by the pre-aimed road surface elevation and its rate of change.

[0085] Figure 2 yes Figure 1 A schematic diagram of the architecture of the mid-edge computing controller 130. Figure 2 In the example, the edge computing controller 130 includes a main controller 132, a slave controller 133, a monitoring unit 131, and an instruction switching unit 134. The monitoring unit 131 is connected to the main controller 132, the slave controller 133, and the instruction switching unit 134, which are equipped with independent interfaces, through different hardware pins. The monitoring unit 131 is used to verify the control commands output by the main controller 132, and controls the instruction switching unit 134 to select the control commands output by the main controller 132 or the slave controller 133 according to preset conditions, so that the dynamic electromagnetic actuator 140 adjusts the damping characteristics of the suspension or applies active force according to the selected control command.

[0086] It is understood that the master controller 132 and slave controller 133 are heterogeneous at the hardware level, thus avoiding the risk of common hardware defects. Furthermore, the master controller 132 and slave controller 133 are configured to run different control algorithms and simultaneously output control commands. Therefore, on the one hand, this avoids failures triggered by the same software vulnerability; on the other hand, it allows for seamless switching to the control commands output by the unfailed controller via the command switching unit 134 in the event of a failure. The command switching unit 134 is connected to the master controller 132 and slave controller 133 via independent physical paths. Therefore, the active suspension system 100 of this embodiment can solve the single-point-of-failure risk faced by a single-controller architecture and the problem that in a homogeneous dual-controller architecture, both controllers may fail simultaneously at the software logic level when encountering targeted interference or systemic risks, and it is difficult to reliably switch over instantly upon failure. Furthermore, the combination of lidar-based pre-aiming functionality further enhances the safety performance of the active suspension.

[0087] Understandably, apart from the independent interface, the power supply module and storage module of the main controller 132 and the slave controller 133 are set independently to ensure redundancy and safety.

[0088] To facilitate a further understanding of the possible configuration methods and coordination and switching logic of the master controller 132 and slave controller 133, the present invention provides the following specific examples:

[0089] Example 1:

[0090] The main controller 132 adopts a GPU unit + FPGA unit architecture, where the GPU unit uses NVIDIA's Jetson AGX Orin or other equivalent embedded GPUs, which support FP16 / INT8 quantization inference. The main controller 132 is configured to run a Twin Delayed Deep Deterministic Policy Gradient (TD3) model or a model based on it.

[0091] The GPU unit is responsible for running the Actor-Critic neural network of the TD3 model to generate control instructions.

[0092] The FPGA unit uses Xilinx's Kintex-7 or equivalent FPGA, and completes LiDAR point cloud preprocessing (elevation feature extraction), multi-sensor timing alignment, and CAN bus data transmission and reception through hardware-accelerated logic.

[0093] The controller 133 adopts an architecture of MCU unit + DSP unit, in which the MCU unit is NXP Semiconductors S32K344. The controller 133 is configured to run a linear quadratic regulator (LQR), which is responsible for the state feedback logic of the LQR algorithm (calculating the control quantity based on the preset feedback matrix K).

[0094] The DSP unit selected is a TI C2000 series DSP (such as TMS320F28379D) from Texas Instruments. It is responsible for fast matrix operations of the LQR algorithm to work with the MCU unit to generate control commands.

[0095] In addition, the controller 133 can also be responsible for system status monitoring (receiving CAN data from the main controller 132 to prepare for switching and to determine and calculate latency).

[0096] Example 2:

[0097] Based on the configuration of the master controller 132 and slave controller 133 shown in Example 1, the master controller 132 and slave controller 133 work together in the following three modes:

[0098] 1. Normal operating mode (main controller 132 dominant):

[0099] The main controller 132 runs TD3 or a model based on it, outputting control commands, and simultaneously:

[0100] A heartbeat signal is sent to the hardware watchdog, i.e., monitoring unit 131, every 10ms.

[0101] The current status is sent to the slave controller 133 every 5ms via the internal CAN bus.

[0102] It synchronously receives sensor data from controller 133, runs the LQR model based on the same state space as the main controller 132, and outputs control commands.

[0103] The monitoring unit 131 or the slave controller 133 compares the deviation of the control commands output by the master controller 132 and the slave controller 133 in real time. If the deviation exceeds the limit, the fault switching mode is triggered.

[0104] 2. Fault switching mode (takeover from controller 133):

[0105] The following scenarios trigger the fault switching mode through monitoring unit 131:

[0106] Scenario 1: The heartbeat signal of the main controller 132 is lost for more than 20ms;

[0107] Scenario 2: The deviation between the control commands output by the main controller 132 and the slave controller 133 exceeds the limit;

[0108] Scenario 3: The main power output corresponding to the control command output by the main controller 132 exceeds the safety limit for 3ms.

[0109] Scenario 4: The computational latency of the main controller 132 is greater than 50ms.

[0110] Switching process:

[0111] When the detection unit outputs a low-level switching signal, the instruction switching unit 134 immediately disconnects the main controller 132 path and connects the slave controller 133 path.

[0112] After receiving the switching signal from controller 133, the current output command of controller 133 is selected and output. At the same time, the state of the LQR model is initialized based on the state of the master controller 132 at the last moment to avoid control abrupt changes.

[0113] After taking over from the main controller 132, a heartbeat signal is sent to the monitoring unit 131 every 10ms until the main controller 132 completes the reset.

[0114] 3. Fault Recovery Mode (Main Controller 132 Reset Process)

[0115] After the main controller 132 has been troubleshooted (such as after a restart or when the algorithm has converged again), it sends a "recovery ready" signal to the monitoring unit 131 via the internal CAN bus.

[0116] After the monitoring unit 131 verifies the rationality of the commands of the main controller 132 for 5 consecutive cycles (deviation ≤ 0.3kN), it sends a "switch back to main" signal to the command switching unit 134.

[0117] The instruction switching unit 134 reselects the control instructions output by the main controller 132.

[0118] In one embodiment, the instruction switching unit 134 switches using a soft transition method. For example, within 0.5 seconds, the instruction weight coefficient of the master controller 132 increases linearly from 0 to 1, while the slave controller 133 decreases from 1 to 0, in order to avoid control abrupt changes (switching shocks).

[0119] In one embodiment, in addition to setting up an independent monitoring unit 131, the monitoring unit 131 can also be implemented based on the slave controller 133. For example, the master controller 132 sends its own status (calculation delay, instruction value) to the slave controller 133 every 10ms. The slave controller 133 assists in judging whether the master controller 132 is abnormal by comparing the deviation between its own LQR instruction and the master controller 132 instruction (normal ≤0.5kN).

[0120] Understandably, despite the different hardware architectures, the output instruction formats of the master controller 132 and the slave controller 133 are completely identical. Therefore, when switching from the master controller 132 to the slave controller 133, the instruction switching unit 134 can be directly selected, thereby reducing the switching delay and avoiding suspension impact.

[0121] In one embodiment, the dynamic electromagnetic actuator 140 includes an electromagnetic clutch, and a linear motor and a magnetorheological damper connected in parallel via the electromagnetic clutch. The linear motor and the magnetorheological damper are used to be connected together to the same suspension link of the vehicle to provide dynamic main force and auxiliary damping, respectively.

[0122] The linear motor has a magnetic shielding layer on its outer wall, and the linear motor and the magnetorheological damper are configured for time-sharing control to eliminate the superposition of their magnetic fields.

[0123] It is understandable that, in addition to providing active power, linear motors can also act as generators to recover energy from suspension bumps, and the state space of the state parameters can also include dimensions used to characterize the energy recovery capability.

[0124] It is understandable that by setting up an electromagnetic clutch to connect the linear motor and the magnetorheological damper in parallel, different working modes can be combined to meet different needs;

[0125] To facilitate a further understanding of the different operating modes achieved by combining a linear motor and a magnetorheological damper in parallel using an electromagnetic clutch, the present invention provides the following specific examples:

[0126] Example 3:

[0127] 1. Active control mode: The electromagnetic clutch engages the linear motor and the magnetorheological damper to work together, and the linear motor outputs active force (such as suppressing body roll when cornering at high speed), while the magnetorheological damper provides auxiliary damping through current regulation;

[0128] 2. Semi-active unpowered mode: The electromagnetic clutch separates the linear motor and the magnetorheological damper, allowing the magnetorheological damper to work independently and actively, thereby adjusting the damping separately with low energy consumption.

[0129] 3. Semi-active energy-recharged mode: The electromagnetic clutch connects the linear motor and the magnetorheological damper to make them work together, and the linear motor acts as a generator to recover energy, while the magnetorheological damper provides auxiliary damping through current regulation;

[0130] 4. Passive energy feeding mode: The electromagnetic clutch connects the linear motor and the magnetorheological damper. The linear motor acts as a generator to recover energy, and the magnetorheological damper does not actively apply current, but only generates passive damping to reduce energy consumption.

[0131] It will be understood by those skilled in the art that Figure 1 or Figure 2 The active suspension system shown does not constitute a limitation on the embodiments of the present invention and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0132] Based on the above active suspension system, various embodiments of the active suspension control method of the present invention are presented below.

[0133] like Figure 3 As shown, Figure 3 This is a flowchart of an active suspension control method provided in an embodiment of the present invention. Figure 3 In the examples, the active suspension control methods include:

[0134] Step S100: Obtain initial parameters and perform standardized preprocessing to obtain state parameters. The state parameters are used to characterize the road conditions and vehicle body status in front of the vehicle perceived by the lidar aiming module and the sensor module.

[0135] Step S200: Input the state parameters into the preset first control model and second control model to obtain the first control instruction and the second control instruction with the same format, wherein the first control model is configured on the master controller and the second control model is configured on the slave controller, and the architecture of the first control model and the second control model are different;

[0136] Step S300: Verify the first control command and the second control command according to the preset gating conditions, and select the first control command or the second control command as the output control command to control the dynamic electromagnetic actuator to adjust the damping characteristics of the suspension or apply the active force.

[0137] Specifically, the standardization preprocessing adopts normalization processing.

[0138] It is understood that the active suspension control method has a first control model and a second control model with different architectures. The first control model and the second control model are respectively configured in the master controller and the slave controller to simultaneously output the first control command and the second control command, and select a single command as the final actual output after verification to control the dynamic electromagnetic actuator to adjust the suspension damping characteristics or apply active force. Since the master controller and the slave controller are configured to run different control algorithms and output control commands simultaneously, it can avoid failures triggered by the same software vulnerability on the one hand, and seamlessly switch to the control command output by the unfailed controller when a failure occurs on the other hand. The first control command and the second control command are obtained based on the same state parameters and have the same output command format, which further reduces the delay in the switching process and prevents control abrupt changes. Therefore, the active suspension control method of the second aspect embodiment of the present invention can solve the problem that when encountering targeted interference or systemic risks, the traditional homogeneous dual controllers may fail simultaneously at the software logic level and it is difficult to complete a reliable switch at the moment of failure, thereby improving the safety performance of the active suspension.

[0139] In one embodiment, the first control model includes:

[0140] The main network group consists of a single main decision network and two main evaluation networks. The main evaluation networks are used to guide the main decision network in optimization.

[0141] The target network group, corresponding to the network type and number in the main network group, includes a single target decision network and two target evaluation networks. The target evaluation network is used to guide the target decision network in optimization.

[0142] The main network group is configured to update in real time for real-time decision-making, while the target network group is configured to synchronize slowly with the main network via soft updates to serve as a benchmark during the main network update process and prevent the main network from oscillating during the update process.

[0143] Specifically, the first control model can be an improved TD3 model, and the second control model can be an LQR model, wherein the main decision network and the target decision network are both Actor networks, and the main evaluation network and the target evaluation network are both Critic networks.

[0144] It should be noted that the TD3 model used in this embodiment of the invention has been improved. Its state space has been expanded to support the core perception dimension of suspension dynamic characteristics, the perception dimension to ensure the safe operation of the suspension and tires, and future environmental information introduced by anticipating road surface elevation and rate of change. Simultaneously, verification is performed based on additional safety constraints. This expands the state space dimensions to improve safety and support anticipation functionality, while further ensuring safety performance through additional verification. The addition of anticipation functionality further enhances the strategy's foresight and adaptability by expanding future road surface information, optimizing the temporal network structure, and designing anticipation-guided rewards. The core value lies in expanding current state control to spatiotemporal sequence control, enabling the active suspension to balance comfort, stability, and safety under complex road conditions.

[0145] However, the existing TD3 model has several problems when applied to active suspension: its state space only realizes the comprehensive perception of the suspension system through multi-dimensional dynamic parameters, only realizes the current state control, and has safety issues.

[0146] It should be noted that when switching from selecting the first control instruction to selecting the second control instruction, the state of the second control model is initialized based on the state space of the first control model at the last moment to avoid control abrupt changes.

[0147] In one embodiment, when switching from selecting the second control command to selecting the first control command, the switching process is carried out in a soft transition manner. For example, within 0.5s, the weight coefficient of the first control command increases linearly from 0 to 1, and the weight coefficient of the second control command decreases from 1 to 0, so as to avoid control abrupt changes (switching shock).

[0148] It should be noted that the first and second control commands have the same format, so they can be directly selected during switching, thereby reducing switching delay and avoiding suspension impact.

[0149] In one embodiment, the first control model further includes a replay buffer for subsequent batch sampling training. The replay buffer supports subsequent training by randomly sampling from the data pairs stored therein to break sample correlations.

[0150] Figure 4 yes Figure 3 The detailed flowchart of step S200 includes:

[0151] Step S210: Input the state parameters into the main decision network and add Gaussian noise to the output of the main decision network to obtain the initial action command;

[0152] Step S220: Verify the initial action command according to the preset safety constraints, and if the initial action command passes the verification, use the initial action command as the first control command to control the dynamic electromagnetic actuator and obtain the corresponding feedback signal.

[0153] Step S230: Store the initial action command actually applied to the dynamic electromagnetic actuator, along with the corresponding feedback signal and state parameters, as empirical data in the playback buffer.

[0154] Step S240: Randomly sample empirical data pairs from the playback buffer, input the empirical data pairs into the target decision network, and add clipped Gaussian noise to the output of the target decision network to obtain the target action command;

[0155] Step S250: Update the main evaluation network and the main decision network based on empirical data and target action instructions;

[0156] Step S260: Update the target evaluation network and target decision network according to the main evaluation network and the main decision network based on the preset soft update method.

[0157] It is understandable that the preset safety constraint can be the actual active force |F applied by the electromagnetic actuator. u If |F| is an interval, then |F| is an interval. u If the value is too large, the initial action command is deemed to have failed the verification.

[0158] Specifically, the state space of the state parameters is set to have nine dimensions. Each dimension of the state space of the state parameters is used to characterize the vertical acceleration of the vehicle body, the vertical velocity of the vehicle body, the rate of change of the vehicle body acceleration, the suspension dynamic travel, the relative displacement between the tire and the road surface, the relative velocity between the vehicle body and the wheel, the constraint on the output force of the dynamic electromagnetic actuator, the pre-aimed road surface elevation, and the rate of change of the pre-aimed road surface elevation.

[0159] Specifically, the feedback signal includes a reward function, which is constructed based on state parameters, and the state space of the reward function corresponds to the state space of the state parameters.

[0160] Specifically, the feedback signal also includes the state parameters that are reacquired at the next moment after the initial action command is actually applied to the electromagnetic actuator.

[0161] To facilitate a further understanding of the state space and reward function design corresponding to the improved TD3 model provided in the embodiments of the present invention, the present invention provides the following specific examples:

[0162] Example 4:

[0163] The state space contains nine dimensions, namely: vehicle vertical acceleration. (Comfort-related), vehicle vertical speed (Comfort-related), rate of change of vehicle acceleration (Suppressing the severity of impact response), suspension dynamic travel z s -z w (To prevent exceeding safety design limits), relative displacement qz between tires and road surface w (Affecting driving stability), relative speed between the vehicle body and wheels Constraints on the output force of dynamic electromagnetic actuators (Physical safety constraints in state space), projected road surface elevation q(t+i) (in conjunction with projecting, t+i represents the elevation from the i-th future time step), and the rate of change of the projected road surface elevation. (Combined with pre-aiming)

[0164] The constructed reward function consists of three terms, namely the core state term r. core Energy consumption item r actuator And the aiming information item r preview ,in:

[0165]

[0166] r actuator =-K7·|F u | 2 , where |F u |This is the main power actually applied to the electromagnetic actuator;

[0167]

[0168] Therefore, the reward function is ultimately constructed as follows:

[0169] r total =r core +r actuator +r preview .

[0170] At each training time step t, the main Actor network (i.e., the main decision network) interacts with the environment, generates actions, and calculates the optimized total reward r. total .

[0171] It is understandable that the preset safety constraint can be |F uIf |F| is an interval, then |F| is an interval. u If the value is too large, the initial action command is deemed to have failed the verification.

[0172] It is understood that the first control model in this embodiment of the invention achieves coordination through a closed-loop process of interactive sampling, value assessment, strategy update, and target synchronization, and incorporates the optimized reward function throughout the process; by combining the pre-aiming item and the core item, the strategy simultaneously optimizes the current vibration suppression and future risk avoidance, thereby improving adaptability under complex road conditions.

[0173] Understandably, by setting the above nine dimensions of the state space, it supports the core perception dimensions of the suspension dynamic characteristics, the perception dimensions that ensure the safe operation of the suspension and tires, and the future environmental information introduced by anticipating the road surface elevation and rate of change. At the same time, it performs verification based on additional safety constraints. Thus, while expanding the dimensions of the state space to improve safety and support the anticipation function, it further ensures safety performance based on additional verification.

[0174] Figure 5 yes Figure 4 The detailed flowchart of step S250 includes:

[0175] Step S251: Input the empirical data pair and the target action command into two target evaluation networks, and select the minimum value of the output of the two target evaluation networks and combine it with the empirical data pair to obtain the target value estimate;

[0176] Step S252: Input the empirical data pairs and the target action instructions into two main evaluation networks to obtain two main value estimates respectively;

[0177] Step S253: Construct an evaluation loss function based on the mean squared error of the two main value estimates and the target value estimate, and update the main evaluation network based on the evaluation loss function using the gradient descent method;

[0178] Step S254: Select a fixed master value estimate output by one of the two master evaluation networks as the benchmark, construct a negative loss function based on the benchmark master value estimate, and update the master decision network based on the negative loss function using the gradient descent method.

[0179] Understandably, the optimization objective of the Actor network is to maximize the value evaluated by the Critic network (i.e., to generate actions with high Q values). Therefore, a negative loss function is constructed based on the output Q1 of the main Critic network.

[0180] Specifically, the negative loss function ensures that the main power output of the electromagnetic actuator corresponding to the control command generated by the Actor network can simultaneously optimize core state suppression, energy consumption control and future risk avoidance in the extended state (including pre-aiming information) (because the Q value of the main Critic network has been incorporated into the optimized reward).

[0181] Understandably, the core function of the Critic network (i.e., the evaluation network) is to assess the value of the active force output by the electromagnetic actuator associated with the current state, that is, the value estimate.

[0182] Figure 6 yes Figure 5 The detailed flowchart of step S260 includes:

[0183] Step S261: When the number of updates of the main network group reaches a preset update threshold, the target evaluation network is updated according to the preset soft update coefficient and the updated main evaluation network based on the evaluation loss function.

[0184] Step S262: Update the corresponding target decision network based on the soft update coefficients and the updated master decision network based on the negative loss function.

[0185] Specifically, to avoid training divergence caused by drastic fluctuations in the target value estimate and the target action, this embodiment of the invention adopts a soft update mechanism to periodically and slowly synchronize the parameters of the main network to the target network.

[0186] Specifically, the parameters φ′ of the target Actor network are softly updated using the formula: φ′←τ·φ+(1-τ)·φ′

[0187] Where τ is the soft update coefficient and φ is the principal decision network parameter, ensuring that the target network parameters iterate slowly to provide a stable learning objective for the Critic and Actor networks.

[0188] Figure 7 yes Figure 3 The detailed flowchart of step S300 includes:

[0189] Step S310: Obtain the deviation value according to the first control command and the second control command;

[0190] Step S320: If the deviation value is less than the preset deviation limit, select the first control command as the output control command.

[0191] Step S330: When the deviation value is greater than the deviation limit, the second control command is selected as the output control command until more than a preset number of deviation values ​​are less than the deviation limit in the subsequent consecutive cases, then the first control command is selected as the output control command again.

[0192] Step S340: When the second control instruction is selected as the output control instruction for the first time, the second control model is initialized according to the state parameters corresponding to the current first control instruction to avoid control abrupt changes.

[0193] It should be noted that when switching from selecting the first control instruction to selecting the second control instruction, the state of the second control model is initialized based on the state space of the first control model at the last moment to avoid control abrupt changes.

[0194] In one embodiment, when switching from selecting the second control command to selecting the first control command, the switching process is carried out in a soft transition manner. For example, within 0.5s, the weight coefficient of the first control command increases linearly from 0 to 1, and the weight coefficient of the second control command decreases from 1 to 0, so as to avoid control abrupt changes (switching shock).

[0195] Furthermore, one embodiment of the present invention provides a computer-readable storage medium storing computer-executable instructions that are executed by a processor or controller, for example, by a processor in the above-described apparatus or method embodiments, causing the processor to perform the active suspension method in the above embodiments, for example, performing the above-described... Figures 2 to 7 The methods and steps in the text.

[0196] It will be understood by those skilled in the art that all or some of the steps and systems in the methods disclosed above can be implemented as software, firmware, hardware, and suitable combinations thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, as is known to those skilled in the art, communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0197] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.

Claims

1. An active suspension system, characterized in that, include: The lidar pre-aiming module is used to obtain the elevation of the road surface in front of the vehicle in order to perceive the road conditions ahead of the vehicle; The state sensor module includes a six-dimensional force sensor, a displacement sensor, and an inertial sensor, and the state sensor module is used to sense the state of the vehicle body. Dynamic electromagnetic actuators are used to dynamically adjust the damping characteristics of the suspension or apply active force. The edge computing controller is electrically connected to the lidar pre-aiming module, the status sensor module, and the dynamic electromagnetic actuator. The edge computing controller includes a master controller, a slave controller, a monitoring unit, and an instruction switching unit. The master controller and the slave controller are heterogeneous, and the master controller and the slave controller are configured to synchronously receive data from the lidar pre-aiming module and the six-dimensional force sensor and run different control algorithms to output control commands simultaneously; The monitoring unit is connected to the main controller, the slave controller, and the instruction switching unit, which are equipped with independent interfaces, through different hardware pins. The monitoring unit is used to verify the control instructions output by the main controller and, according to preset conditions, controls the instruction switching unit to select the control instructions output by the main controller or the slave controller, so that the dynamic electromagnetic actuator adjusts the damping characteristics of the suspension or applies active force according to the selected control instructions.

2. The active suspension system according to claim 1, characterized in that, The dynamic electromagnetic actuator includes an electromagnetic clutch, and a linear motor and a magnetorheological damper connected in parallel via the electromagnetic clutch. The linear motor and the magnetorheological damper are used to be connected together to the same suspension link of the vehicle to provide dynamic main power and auxiliary damping, respectively. The linear motor has a magnetic shielding layer on its outer wall, and the linear motor and the magnetorheological damper are configured for time-sharing control to eliminate the superposition of their magnetic fields.

3. An active suspension control method, applied to the active suspension system as described in any one of claims 1 to 2, characterized in that, include: Initial parameters are obtained and standardized preprocessing is performed to obtain state parameters, which are used to characterize the road conditions and vehicle body status in front of the vehicle as perceived by the lidar aiming module and the sensor module. The state parameters are input to a preset first control model and a second control model to obtain a first control instruction and a second control instruction with the same format, wherein the first control model is configured on the master controller and the second control model is configured on the slave controller, and the architectures of the first control model and the second control model are different. The first control command and the second control command are verified according to the preset selection conditions, and the first control command or the second control command is selected as the output control command to control the dynamic electromagnetic actuator to adjust the damping characteristics of the suspension or apply the active force.

4. The active suspension control method according to claim 3, characterized in that, The first control model includes: The main network group includes a single main decision network and two main evaluation networks, wherein the main evaluation networks are used to guide the main decision network in optimization. The target network group, corresponding to the network type and number in the main network group, is configured with a single target decision network and two target evaluation networks. The target evaluation network is used to guide the target decision network to perform optimization. The main network group is configured to update in real time for making real-time decisions, while the target network group is configured to slowly synchronize with the main network via soft updates, serving as a benchmark during the main network update process to prevent the main network from oscillating during the update process.

5. The active suspension control method according to claim 4, characterized in that, The first control model further includes a playback buffer for subsequent batch sampling training. The step of inputting the state parameters into a preset first control model and a second control model to obtain first and second control instructions with the same format includes: The state parameters are input into the main decision network, and Gaussian noise is added to the output of the main decision network to obtain the initial action command; The initial action command is verified according to the preset safety constraints, and if the initial action command passes the verification, the initial action command is used as the first control command to control the dynamic electromagnetic actuator and obtain the corresponding feedback signal. The initial action command actually applied to the dynamic electromagnetic actuator, along with the corresponding feedback signal and state parameters, are stored as empirical data in the playback buffer. The empirical data pairs are randomly sampled from the playback buffer and input into the target decision network. Clipped Gaussian noise is added to the output of the target decision network to obtain the target action command. The main evaluation network and the main decision network are updated based on the empirical data and the target action instructions; The target evaluation network and the target decision network are updated based on a preset soft update method according to the main evaluation network and the main decision network.

6. The active suspension control method according to claim 5, characterized in that, The state space of the state parameters is set to have multiple dimensions. Each dimension of the state space of the state parameters is used to characterize the vertical acceleration of the vehicle body, the vertical velocity of the vehicle body, the rate of change of the vehicle body acceleration, the suspension dynamic travel, the relative displacement between the tire and the road surface, the relative velocity between the vehicle body and the wheel, the constraint on the output force of the dynamic electromagnetic actuator, the pre-aimed road surface elevation, and the rate of change of the pre-aimed road surface elevation. The feedback signal includes a reward function, which is constructed based on the state parameters, and the state space of the reward function corresponds to the state space of the state parameters. The feedback signal also includes the state parameters reacquired at the next moment after the initial action command is actually applied to the electromagnetic actuator.

7. The active suspension control method according to claim 6, characterized in that, The step of updating the main evaluation network and the main decision network based on the empirical data and the target action instruction includes: The empirical data pair and the target action command are input into two target evaluation networks, and the minimum value of the output of the two target evaluation networks is selected and combined with the empirical data pair to obtain the target value estimate. The empirical data pair and the target action instruction are input into the two main evaluation networks to obtain two main value estimates respectively. An evaluation loss function is constructed based on the mean squared error of the two main value estimates and the target value estimate, and the main evaluation network is updated based on the evaluation loss function using the gradient descent method. Select a fixed master value estimate from the output of one of the two master evaluation networks as a benchmark, construct a negative loss function based on the benchmark master value estimate, and update the master decision network based on the negative loss function using the gradient descent method.

8. The active suspension control method according to claim 7, characterized in that, The preset soft update method updates the target evaluation network and the target decision network according to the main evaluation network and the main decision network, including: When the number of updates to the main network group reaches a preset update threshold, the target evaluation network is updated according to the preset soft update coefficient and the updated main evaluation network based on the evaluation loss function. The target decision network is updated based on the soft update coefficient and the updated main decision network based on the negative loss function.

9. The active suspension control method according to claim 5, characterized in that, The step of verifying the first control instruction and the second control instruction according to preset gating conditions, and selecting the first control instruction or the second control instruction as the output control instruction, includes: The deviation value is obtained based on the first control command and the second control command; If the deviation value is less than the preset deviation limit, the first control command is selected as the output control command. If the deviation value is greater than the deviation limit, the second control command is selected as the output control command until more than a preset number of deviation values ​​are less than the deviation limit, then the first control command is selected as the output control command again. When the second control instruction is selected as the output control instruction for the first time, the second control model is initialized according to the state parameters corresponding to the current first control instruction to avoid control abrupt changes.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to perform the active suspension control method as described in any one of claims 3 to 9.