Dynamic cognitive driving variable transmission ratio control method for four-wheel independent steering system

By constructing a system state cognition model and introducing deep reinforcement learning, the problem of transmission ratio adjustment in a four-wheel independent steering system under dynamic and complex working conditions was solved, realizing real-time self-learning and strategy evolution, and improving the system's responsiveness and stability.

CN121799494APending Publication Date: 2026-04-07NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202610135550.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-30
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing four-wheel independent steering systems lack real-time cognitive and strategy self-evolution capabilities under dynamic and complex working conditions, resulting in rigid transmission ratio adjustment, inflexible strategies, and insufficient adjustment precision, which affects the steering sensitivity and stability of the entire vehicle.

Method used

By collecting vehicle state parameters, a system state cognition model is constructed, a set of operating condition state labels and their associated weight matrices are generated, a control performance objective function is established, and a deep reinforcement learning method is introduced to construct a deep neural network for the transmission ratio control strategy, realizing real-time self-learning and strategy evolution, and activating the strategy correction module for anomaly correction.

Benefits of technology

It achieves comprehensive modeling and precise classification of vehicle operating status, improves the responsiveness and accuracy of transmission ratio adjustment, maintains the responsiveness and stability of the steering system under complex dynamic conditions, and enhances the control performance of the system under extreme conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121799494A_ABST
    Figure CN121799494A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic cognitive driving variable transmission ratio control method for a four-wheel independent steering system. The method comprises the following steps that vehicle state parameters and steering wheel related parameters are collected; constructing a system state cognition model according to the obtained perception vector, and generating a working condition state label set and a corresponding working condition association weight matrix; according to the generated working condition state label set and the association weight matrix, constructing a control performance objective function and forming a working condition-objective-strategy value mapping model; outputting a target transmission ratio adjustment interval and a deep reinforcement learning control strategy initial parameter under the current working condition; and adjusting the interval and the initial strategy parameter according to the output target transmission ratio. The problems that in the prior art, due to the lack of a system-level state perception and strategy self-evolution mechanism, the steering system transmission ratio adjusting rigidity is high, the strategy is rigid, the adjusting precision is insufficient under the dynamic complex working condition, and then the steering sensitivity and stability of the whole vehicle are affected are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of automobile steering, and more particularly to a dynamic cognitive drive variable transmission ratio control method for a four-wheel independent steering system. Background Technology

[0002] With the development of intelligent driving technology, four-wheel independent steering systems, as an advanced chassis actuator with high flexibility and controllability, have been gradually applied in scenarios such as intelligent electric vehicles and heavy-duty logistics transportation equipment. This type of system eliminates the traditional steering tie rod, achieving completely independent control of the steering angle of each wheel. It features high maneuverability, redundancy tolerance, and a high degree of freedom in path planning, making it particularly suitable for operating scenarios with complex structures and rapidly changing environments. However, the multi-degree-of-freedom control redundancy also places higher demands on the system's real-time control performance and adaptive capabilities.

[0003] During dynamic operation, influenced by factors such as road adhesion, vehicle speed changes, and uncertainties in driving intentions, the transmission ratio control strategy of a four-wheel independent steering system urgently needs to possess adaptive capabilities and multi-objective coordination. Traditional rule-driven adjustment methods suffer from response lag and insufficient adjustment rigidity when facing complex operating conditions, making it difficult to fully realize the potential of the four-wheel independent structure in terms of steering performance optimization, energy consumption control, and vehicle stability maintenance.

[0004] Existing research on variable transmission ratio control of steering systems includes, for example: Chinese Invention Publication No. CN120408845A, entitled "Design Method and System for Variable Transmission Ratio of Steerable-by-Wire Vehicle Based on Fuzzy Neural Network," which uses a quadratic cost function of the vehicle's dynamic state to design a multi-objective evaluation method, and learns the globally optimal control scheme based on a nonlinear control model established by a fuzzy RBF network; CN120337759A, entitled "A Design Method and System for Variable Transmission Ratio of Steerable-by-Wire System," which uses an improved gray wolf optimization algorithm to iteratively optimize the steady-state yaw rate gain and center ratio gradient coefficient to establish a three-dimensional transmission ratio model; and CN109726516A, entitled "A Variable Transmission Ratio Optimization Design Method and Dedicated System for Multi-Mode Steerable-by-Wire Power Steering System," which establishes a multi-objective optimization model with steering feel and steering sensitivity as optimization objectives to optimize the transmission ratio design.

[0005] However, existing control methods still have significant shortcomings in the following two aspects: First, in terms of operating condition perception, most methods rely on rule setting or single operating condition variables (such as vehicle speed or adhesion coefficient) for adjustment, lacking the ability to jointly perceive and dynamically identify complex environmental factors, making it difficult to achieve real-time cognition and predictive response of system status. Second, in terms of control strategy design, traditional mapping models or fuzzy rule methods cannot adjust the control strategy based on real-time feedback, lack the ability to evolve the strategy, and are difficult to cope with the optimal adjustment requirements under dynamic and changing operating conditions. Especially under extreme operating conditions or multivariate disturbance conditions, problems such as steering response lag, over-adjustment, or decreased vehicle stability are likely to occur.

[0006] Therefore, how to construct a transmission ratio control method with real-time cognition, strategy optimization and multi-condition collaborative adjustment capabilities, give full play to the advantages of the four-wheel independent steering system such as high degree of freedom, fast response and strong reconfigurability, and develop intelligent control strategies with dynamic evolution capabilities has become a key link in promoting the evolution of the four-wheel independent steering system towards higher-level intelligent driving. Summary of the Invention

[0007] The purpose of this section is to outline some aspects of embodiments of the present invention and to briefly describe some preferred embodiments. Simplifications or omissions may be made in this section, as well as in the abstract and title of this application, to avoid obscuring the purpose of these documents; however, such simplifications or omissions should not be construed as limiting the scope of the invention.

[0008] In view of the problems existing in the dynamic cognitive drive variable transmission ratio control method of the above-mentioned four-wheel independent steering system, the present invention is proposed.

[0009] Therefore, the purpose of this invention is to provide a dynamic cognitive drive variable transmission ratio control method for a four-wheel independent steering system, which solves the problem in the prior art that the lack of system-level state perception and strategy self-evolution mechanism leads to strong rigidity, rigid strategy, and insufficient adjustment accuracy in the steering system transmission ratio adjustment under dynamic and complex working conditions, thereby affecting the steering sensitivity and stability of the whole vehicle.

[0010] To solve the above-mentioned technical problems, the present invention provides the following technical solution: a dynamic cognitive drive variable transmission ratio control method for a four-wheel independent steering system, comprising the following steps: 1) Collect vehicle status parameters and steering wheel-related parameters to form a working condition perception vector; 2) Based on the perception vectors obtained above, construct a system state cognition model and generate a set of working condition state labels and their corresponding working condition association weight matrices; 3) Based on the generated set of operating condition labels and the associated weight matrix, construct the control performance objective function and form an operating condition-objective-strategy value mapping model; output the target transmission ratio adjustment range and the initial parameters of the deep reinforcement learning control strategy under the current operating condition; 4) Based on the target transmission ratio adjustment range and initial strategy parameters output above, establish a deep neural network structure for the transmission ratio control strategy; construct a dataset and output a sequence of control commands; 5) Iteratively update the network parameters based on the learned policy parameters; 6) Based on the real-time control command sequence output above and the updated policy network parameters, when the conditions are met, activate the policy correction module; and correct the abnormal reporting areas.

[0011] As a preferred embodiment of the dynamic cognitive drive variable transmission ratio control method for the four-wheel independent steering system described in this invention, step 1) specifically includes: 11) Vehicle state parameters and steering wheel related parameters include: longitudinal velocity, yaw rate, lateral acceleration, four-wheel steering angle, road adhesion coefficient, steering wheel angle, steering wheel angular velocity, and steering wheel torque characteristic parameters; 12) The working condition perception vector is constructed as follows: (1) In the formula, v x It is longitudinal velocity. It is the yaw rate. It is lateral acceleration. , , , These are the steering angles of the left front wheel, right front wheel, left rear wheel, and right rear wheel, respectively. It is the road adhesion coefficient. It's the steering wheel angle. It is the angular velocity of the steering wheel. It is the sampling time. T sw It refers to steering wheel torque.

[0012] As a preferred embodiment of the dynamic cognitive drive variable transmission ratio control method for the four-wheel independent steering system described in this invention, step 2) specifically includes: 21) Based on the perceptual vector constructed in step 1), sliding window reconstruction, temporal feature extraction, and multivariate cross-analysis are performed. The integrated model input vector is as follows: (2) In the formula, It is the mean. It is the standard deviation. It is the maximum value. It is the minimum value. It's a steepness. It is the first-order difference mean. It is the autocorrelation coefficient. It is the Pearson correlation coefficient. It is the ratio of yaw rate to lateral acceleration. It is the transmission ratio between the steering wheel angle and the front wheel angle; Construct a system state cognition model based on the principles of random forest: 22) Based on cluster analysis and dynamic feature attribution, a set of operating condition labels reflecting driving behavior, vehicle dynamics, and environmental disturbances is generated. Static similarity and dynamic transition probability are weighted and fused to obtain the final association weight. as follows: (3) In the formula, These are the weighting coefficients. It is cosine similarity. It's the transition probability; if "smooth straight driving" frequently transitions to "slight turning," then... If "emergency braking" hardly transitions into "high-speed acceleration," then... Small; Based on the final association weight Generate a working condition association weight matrix corresponding to the set of working condition status labels, which serves as the input basis for the subsequent control strategy generation.

[0013] As a preferred embodiment of the dynamic cognitive drive variable transmission ratio control method for the four-wheel independent steering system described in this invention, step 3) specifically includes: 31) Based on the set of working condition labels and the associated weight matrix output in step 2), set the yaw response gain, lateral acceleration response gain, and steering wheel feel stability as multi-objective constraints, and apply them to each working condition. For each of the three control objectives, a "performance deviation index" is defined, using the "squared relative error" method. and The difference between the actual yaw response gain and the ideal yaw response gain, and the difference between the actual lateral acceleration response gain and the ideal lateral acceleration response gain, are quantified using a "penalty term for exceeding the upper limit". The actual steering wheel torque fluctuation and the ideal steering wheel torque fluctuation are quantified, and the three single-target indicators are weighted and fused according to "control priority" to obtain a single-condition comprehensive indicator. : (4) In the formula, It is the target weight, which satisfies .

[0014] Weight matrix associated with working conditions By weighted summing of all single-condition comprehensive indicators, a global multi-objective control performance objective function is obtained. : (5) 32) Introduce the policy value function from deep reinforcement learning and design a weighted penalty reward function: (6) In the formula, This is the base reward; the reward is positive when there is no deviation. These are the weighting coefficients for the control objective, with a penalty term added. To prevent the strategy from deviating from the safe range; Value assessment of different transmission ratio adjustment strategies is conducted to form a value mapping model of working condition-target-strategy; 33) Historical strategies are filtered and optimized through a policy gradient mechanism, and the optimized strategies are then used. For optimal parameters, in the current state downsampling Each action yields an action set. Based on vehicle hardware limitations, actions exceeding the mechanical range are eliminated, and the 5th and 95th percentiles of the filtered actions are used as the adjustment range. , The target transmission ratio adjustment range is Output the target transmission ratio adjustment range and the initial parameters of the deep reinforcement learning control strategy under the current operating conditions, as the basis for online updating of the real-time control strategy.

[0015] As a preferred embodiment of the dynamic cognitive drive variable transmission ratio control method for the four-wheel independent steering system described in this invention, step 4) specifically includes: 41) Based on the target transmission ratio adjustment range and initial strategy parameters output in step 3), a deep neural network structure for the transmission ratio control strategy is established with the working condition state vector as input. 42) Using the real-time feedback of vehicle yaw rate, lateral acceleration, and steering wheel torque as the reward signal in reinforcement learning, define the state vector: (7) In the formula, It is the longitudinal speed of the vehicle. It is the yaw rate. It is lateral acceleration. It is the torque required to turn the steering wheel. It is an estimated value of the road surface adhesion coefficient. It is the deviation in yaw rate. , It's the front wheel steering angle. It's the wheelbase. It is a stability factor. It is lateral acceleration deviation. , It's a deviation in steering wheel torque.

[0016] Construct a reward signal based on real-time feedback, weighted sum of single-objective rewards, and add a safety constraint penalty term: (8) In the formula, ; Construct a state-action-reward triplet dataset to enable online training and dynamic updating of the policy network. 43) Based on the current transmission ratio With the target transmission ratio Generate the sequence using linear interpolation: (9) In the formula, This is the current transmission ratio. It is the target transmission ratio; Output the real-time transmission ratio control command sequence corresponding to the independent target steering angles of the left and right front wheels, and iteratively update the learned strategy network parameters to improve the adaptive and generalization capabilities of the control strategy under different operating conditions.

[0017] As a preferred embodiment of the dynamic cognitive drive variable transmission ratio control method for the four-wheel independent steering system described in this invention, step 5) specifically includes: pre-training the network with basic working condition data, gradually fine-tuning the upper output layer, verifying it with extreme working conditions that were not involved in the training, and stopping the current iteration after the average reward reaches the target.

[0018] As a preferred embodiment of the dynamic cognitive drive variable transmission ratio control method for the four-wheel independent steering system described in this invention, step 6) specifically includes: 61) Based on the real-time control command sequence and updated strategy network parameters output above, when disturbance characteristics such as sudden changes in road adhesion conditions, sudden changes in steering wheel input, system response lag, or deterioration of yaw control performance are detected, if the abnormal threshold of any disturbance characteristic is met and it lasts for 2 control cycles, the strategy correction module is activated, the current working condition is marked as "disturbance working condition", and the strategy correction module is activated. 62) Introducing policy fine-tuning and experience replay mechanisms from deep reinforcement learning, priority sampling is performed on regions with abnormal rewards, and the policy network is retrained. To enhance optimization of low-reward samples, a reward-weighted term is added to the loss function, with samples having higher weights for lower rewards. (10) In the formula, It is PPO-clip loss For the dominant function, , It is the return-weighted loss. These are weighting coefficients; The output corrected control strategy incremental parameters are then weighted and fused together with the categorized compensation terms, superimposed onto the basic command, and the output adjustment signal compensation term is generated. (11) In the formula, To compensate for the weight, It is transmission ratio gain compensation. It is a compensation term for instruction smoothing. It is an advance compensation item. It is a yaw deviation compensation; to enhance the control robustness and response stability of the system in unstructured environments.

[0019] As a preferred embodiment of the dynamic cognitive drive variable transmission ratio control method for the four-wheel independent steering system described in this invention, the four-wheel independent steering control system includes: a hydraulic cylinder linear displacement sensor 1, a hydraulic cylinder 2, a proportional valve 3, a valve core displacement sensor 4, a steering wheel drive module 5, a steering wheel angle sensor 6, a vehicle speed sensor 7, a PC 8, a CAN controller 9, a CAN transceiver 10, CAN HIGH 11, CAN LOW 12, a main control module 13, and a CAN bus 14. The hydraulic cylinder linear displacement sensor 1 is installed on the hydraulic cylinder 2 to detect the linear displacement of the hydraulic cylinder 2; the hydraulic cylinder 2 serves as a steering actuator to drive the steering wheel to rotate; the proportional valve 3 controls the oil flow and direction of the hydraulic cylinder 2; the valve core displacement sensor 4 detects the displacement of the valve core of the proportional valve 3; the steering wheel drive module 5 drives the steering wheel to achieve the steering action. The steering wheel angle sensor 6 detects the rotation angle of the steering wheel; the vehicle speed sensor 7 detects the vehicle speed; the PC 8 serves as the host computer for system monitoring or parameter setting; and the main control module 13 serves as the core control unit of the system, processing sensor signals and outputting control commands. The CAN controller 9 processes the CAN bus 14 communication protocol; the CAN transceiver 10 realizes level conversion and signal transmission and reception between the CAN controller 9 and the CAN bus 14; the CAN HIGH 11 and CAN LOW 12 together constitute the twisted pair signal line of the CAN bus 14; the CAN bus 14 connects each control unit to realize data communication and coordinated control of the vehicle steering system; The control system transmits the detection signals from the steering wheel angle sensor 6, vehicle speed sensor 7, hydraulic cylinder linear displacement sensor 1, and valve core displacement sensor 4 to the main control module 13 via the CAN bus 14. The main control module 13 generates control commands according to the control strategy and sends them to the steering wheel drive module 5 via the CAN bus 14 to drive the proportional valve 3 and hydraulic cylinder 2 to achieve four-wheel independent steering function.

[0020] The beneficial effects of this invention are: 1. This invention not only introduces a multi-source operating condition perception mechanism in the transmission ratio adjustment process to achieve comprehensive modeling and accurate classification of vehicle operating status, but also combines the operating condition labels and state weights output by the dynamic cognitive model to complete targeted information extraction and state feature reconstruction before the transmission ratio strategy is generated, providing reliable data support and state prior information for subsequent strategy control.

[0021] 2. This invention introduces a deep reinforcement learning method into the variable transmission ratio control process to achieve self-learning, self-updating, and strategy evolution of the control strategy under different operating conditions. It does not rely on fixed mapping rules or experience-based adjustment models, and has online adaptability and historical experience-driven strategy optimization capabilities, thereby improving the steering system's response capability and adjustment accuracy to complex dynamic operating conditions.

[0022] 3. When abnormal states such as operating condition disturbances, attachment changes, or control delays occur, the present invention activates a deep reinforcement learning strategy correction mechanism, and corrects the controller output by means of experience playback and strategy fine-tuning, so as to achieve continuous optimization of the control strategy and dynamic robustness improvement; it can maintain the sensitivity and stability of steering response during vehicle driving, and significantly enhance the control performance of the system under extreme conditions. Attached Figure Description

[0023] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein: Figure 1 This is a schematic diagram illustrating the principle and structure of the dynamic cognitive drive variable transmission ratio control method for the four-wheel independent steering system of the present invention.

[0024] Figure 2 This is a schematic diagram of the system structure of the dynamic cognitive drive variable transmission ratio control method for the four-wheel independent steering system of the present invention. Detailed Implementation

[0025] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0026] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0027] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0028] Secondly, the present invention is described in detail with reference to the schematic diagrams. When detailing the embodiments of the present invention, for ease of explanation, the cross-sectional views illustrating the device structure may be partially enlarged, not according to the usual scale. Furthermore, the schematic diagrams are merely examples and should not limit the scope of protection of the present invention. In addition, actual fabrication should include three-dimensional spatial dimensions of length, width, and depth.

[0029] Reference Figures 1-2 A dynamic cognitive drive variable transmission ratio control method for a four-wheel independent steering system is provided, comprising the following steps: 1) Collect vehicle status parameters and steering wheel-related parameters to form a working condition perception vector; Specifically, step 1) includes: 11) Vehicle state parameters and steering wheel related parameters include: longitudinal velocity, yaw rate, lateral acceleration, four-wheel steering angle, road adhesion coefficient, steering wheel angle, steering wheel angular velocity, and steering wheel torque characteristic parameters; 12) The working condition perception vector is constructed as follows: (1) In the formula, v x It is longitudinal velocity. It is the yaw rate. It is lateral acceleration. , , , These are the steering angles of the left front wheel, right front wheel, left rear wheel, and right rear wheel, respectively. It is the road adhesion coefficient. It's the steering wheel angle. It is the angular velocity of the steering wheel. It is the sampling time. T sw It refers to steering wheel torque; 2) Based on the perception vectors obtained above, construct a system state cognition model and generate a set of working condition state labels and their corresponding working condition association weight matrices; Specifically, step 2) includes: 21) Based on the perceptual vector constructed in step 1), sliding window reconstruction, temporal feature extraction, and multivariate cross-analysis are performed. The integrated model input vector is as follows: (2) In the formula, It is the mean. It is the standard deviation. It is the maximum value. It is the minimum value. It's a steepness. It is the first-order difference mean. It is the autocorrelation coefficient. It is the Pearson correlation coefficient. It is the ratio of yaw rate to lateral acceleration. It is the transmission ratio between the steering wheel angle and the front wheel angle; Construct a system state cognition model based on the principles of random forest: 22) Based on cluster analysis and dynamic feature attribution, a set of operating condition labels reflecting driving behavior, vehicle dynamics, and environmental disturbances is generated. Static similarity and dynamic transition probability are weighted and fused to obtain the final association weight. as follows: (3) In the formula, These are the weighting coefficients. It is cosine similarity. It's the transition probability; if "smooth straight driving" frequently transitions to "slight turning," then... If "emergency braking" hardly transitions into "high-speed acceleration," then... Small; Based on the final association weight Generate a working condition association weight matrix corresponding to the set of working condition status labels, which serves as the input basis for the subsequent control strategy generation; 3) Based on the generated set of operating condition labels and the associated weight matrix, construct the control performance objective function and form an operating condition-objective-strategy value mapping model; output the target transmission ratio adjustment range and the initial parameters of the deep reinforcement learning control strategy under the current operating condition; Specifically, step 3) includes: 31) Based on the set of working condition labels and the associated weight matrix output in step 2), set the yaw response gain, lateral acceleration response gain, and steering wheel feel stability as multi-objective constraints, and apply them to each working condition. For each of the three control objectives, a "performance deviation index" is defined, using the "squared relative error" method. and The difference between the actual yaw response gain and the ideal yaw response gain, and the difference between the actual lateral acceleration response gain and the ideal lateral acceleration response gain, are quantified using a "penalty term for exceeding the upper limit". The actual steering wheel torque fluctuation and the ideal steering wheel torque fluctuation are quantified, and the three single-target indicators are weighted and fused according to "control priority" to obtain a single-condition comprehensive indicator. : (4) In the formula, It is the target weight, which satisfies .

[0030] Weight matrix associated with working conditions By weighted summing of all single-condition comprehensive indicators, a global multi-objective control performance objective function is obtained. : (5) 32) Introduce the policy value function from deep reinforcement learning and design a weighted penalty reward function: (6) In the formula, This is the base reward; the reward is positive when there is no deviation. These are the weighting coefficients for the control objective, with a penalty term added. To prevent the strategy from deviating from the safe range; Value assessment of different transmission ratio adjustment strategies is conducted to form a value mapping model of working condition-target-strategy; 33) Historical strategies are filtered and optimized through a policy gradient mechanism, and the optimized strategies are then used. For optimal parameters, in the current state downsampling Each action yields an action set. Based on vehicle hardware limitations, actions exceeding the mechanical range are eliminated, and the 5th and 95th percentiles of the filtered actions are used as the adjustment range. , The target transmission ratio adjustment range is Output the target transmission ratio adjustment range and the initial parameters of the deep reinforcement learning control strategy under the current operating conditions, as the basis for online updating of the real-time control strategy; 4) Based on the target transmission ratio adjustment range and initial strategy parameters output above, establish a deep neural network structure for the transmission ratio control strategy; construct a dataset and output a sequence of control commands; Specifically, step 4) includes: 41) Based on the target transmission ratio adjustment range and initial strategy parameters output in step 3), a deep neural network structure for the transmission ratio control strategy is established with the working condition state vector as input. 42) Using the real-time feedback of vehicle yaw rate, lateral acceleration, and steering wheel torque as the reward signal in reinforcement learning, define the state vector: (7) In the formula, It is the longitudinal speed of the vehicle. , It is lateral acceleration. It is the torque required to turn the steering wheel. It is an estimated value of the road surface adhesion coefficient. It is the deviation in yaw rate. , It's the front wheel steering angle. It's the wheelbase. It is a stability factor. It is lateral acceleration deviation. , It's a deviation in steering wheel torque.

[0031] Construct a reward signal based on real-time feedback, weighted sum of single-objective rewards, and add a safety constraint penalty term: (8) In the formula, ; Construct a state-action-reward triplet dataset to enable online training and dynamic updating of the policy network. 43) Based on the current transmission ratio With the target transmission ratio Generate the sequence using linear interpolation: (9) In the formula, This is the current transmission ratio. It is the target transmission ratio; Output the real-time transmission ratio control command sequence corresponding to the independent target steering angles of the left and right front wheels, and iteratively update the learned strategy network parameters to improve the adaptive and generalization capabilities of the control strategy under different working conditions. 5) Iteratively update the network parameters based on the learned policy parameters; Specifically, step 5) includes: pre-training the network with basic working condition data, gradually fine-tuning the upper output layer, and verifying it with extreme working conditions that were not used in the training. The current iteration is stopped after the average reward reaches the target. 6) Based on the real-time control command sequence output above and the updated policy network parameters, when the conditions are met, activate the policy correction module; and correct the abnormal reporting areas. Specifically, step 6) includes: 61) Based on the real-time control command sequence and updated strategy network parameters output above, when disturbance characteristics such as sudden changes in road adhesion conditions, sudden changes in steering wheel input, system response lag, or deterioration of yaw control performance are detected, if the abnormal threshold of any disturbance characteristic is met and it lasts for 2 control cycles, the strategy correction module is activated, the current working condition is marked as "disturbance working condition", and the strategy correction module is activated. 62) Introducing policy fine-tuning and experience replay mechanisms from deep reinforcement learning, priority sampling is performed on regions with abnormal rewards, and the policy network is retrained. To enhance optimization of low-reward samples, a reward-weighted term is added to the loss function, with samples having higher weights for lower rewards. (10) In the formula, It is PPO-clip loss For the dominant function, , It is the return-weighted loss. These are weighting coefficients; The output corrected control strategy incremental parameters are then weighted and fused together with the categorized compensation terms, superimposed onto the basic command, and the output adjustment signal compensation term is generated. (11) In the formula, To compensate for the weight, It is transmission ratio gain compensation. It is a compensation term for instruction smoothing. It is an advance compensation item. It is yaw rate compensation; to enhance the control robustness and response stability of the system in unstructured environments; The four-wheel independent steering control system includes: hydraulic cylinder linear displacement sensor 1, hydraulic cylinder 2, proportional valve 3, valve core displacement sensor 4, steering wheel drive module 5, steering wheel angle sensor 6, vehicle speed sensor 7, PC 8, CAN controller 9, CAN transceiver 10, CAN HIGH 11, CAN LOW 12, main control module 13, and CAN bus 14. The hydraulic cylinder linear displacement sensor 1 is installed on the hydraulic cylinder 2 to detect the linear displacement of the hydraulic cylinder 2; the hydraulic cylinder 2 acts as a steering actuator to drive the steering wheel to rotate; the proportional valve 3 controls the oil flow and direction of the hydraulic cylinder 2; the valve core displacement sensor 4 detects the displacement of the valve core of the proportional valve 3; the steering wheel drive module 5 drives the steering wheel to achieve the steering action. Steering wheel angle sensor 6 detects the steering wheel rotation angle; vehicle speed sensor 7 detects the vehicle speed; PC 8 acts as a host computer for system monitoring or parameter setting; main control module 13 acts as the core control unit of the system, processing sensor signals and outputting control commands. CAN controller 9 processes the CAN bus 14 communication protocol; CAN transceiver 10 realizes level conversion and signal transmission and reception between CAN controller 9 and CAN bus 14; CAN HIGH 11 and CAN LOW 12 together constitute the twisted pair signal line of CAN bus 14; CAN bus 14 connects various control units to realize data communication and coordinated control of the vehicle steering system; The control system transmits the detection signals from the steering wheel angle sensor 6, vehicle speed sensor 7, hydraulic cylinder linear displacement sensor 1, and valve core displacement sensor 4 to the main control module 13 via the CAN bus 14. The main control module 13 generates control commands according to the control strategy and sends them to the steering wheel drive module 5 via the CAN bus 14 to drive the proportional valve 3 and hydraulic cylinder 2 to achieve four-wheel independent steering function.

[0032] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A dynamic cognitive drive variable transmission ratio control method for a four-wheel independent steering system, characterized in that, Includes the following steps: 1) Collect vehicle status parameters and steering wheel-related parameters to form a working condition perception vector; 2) Based on the perception vectors obtained above, construct a system state cognition model and generate a set of working condition state labels and their corresponding working condition association weight matrices; 3) Based on the generated set of operating condition labels and the associated weight matrix, construct the control performance objective function and form an operating condition-objective-strategy value mapping model; output the target transmission ratio adjustment range and the initial parameters of the deep reinforcement learning control strategy under the current operating condition; 4) Based on the target transmission ratio adjustment range and initial strategy parameters output above, establish a deep neural network structure for the transmission ratio control strategy; construct a dataset and output a sequence of control commands; 5) Iteratively update the network parameters based on the learned policy parameters; 6) Based on the real-time control command sequence output above and the updated policy network parameters, activate the policy correction module when the conditions are met; And correct any abnormal reporting areas.

2. The dynamic cognitive drive variable transmission ratio control method for a four-wheel independent steering system according to claim 1, characterized in that: Step 1) specifically includes: 11) Vehicle status parameters and steering wheel related parameters include: longitudinal velocity, yaw rate, lateral acceleration, four-wheel steering angle, road adhesion coefficient, steering wheel angle, steering wheel angular velocity, and steering wheel torque characteristic parameters; 12) The working condition perception vector is constructed as follows: (1) In the formula, v x It is longitudinal velocity. It is the yaw rate. It is lateral acceleration. , , , These are the steering angles of the left front wheel, right front wheel, left rear wheel, and right rear wheel, respectively. It is the road adhesion coefficient. It's the steering wheel angle. It is the angular velocity of the steering wheel. It is the sampling time. T sw It refers to steering wheel torque.

3. The dynamic cognitive drive variable transmission ratio control method for a four-wheel independent steering system according to claim 2, characterized in that: Step 2) specifically includes: 21) Based on the perceptual vector constructed in step 1), sliding window reconstruction, temporal feature extraction, and multivariate cross-analysis are performed. The integrated model input vector is as follows: (2) In the formula, It is the mean. It is the standard deviation. It is the maximum value. It is the minimum value. It's a steepness. It is the first-order difference mean. It is the autocorrelation coefficient. It is the Pearson correlation coefficient. It is the ratio of yaw rate to lateral acceleration. It is the transmission ratio between the steering wheel angle and the front wheel angle; Construct a system state cognition model based on the principles of random forest: 22) Based on cluster analysis and dynamic feature attribution, a set of operating condition labels reflecting driving behavior, vehicle dynamics, and environmental disturbances is generated. Static similarity and dynamic transition probability are weighted and fused to obtain the final association weight. as follows: (3) In the formula, These are the weighting coefficients. It is cosine similarity. It's the transition probability; if "smooth straight driving" frequently transitions to "slight turning," then... If "emergency braking" hardly transitions into "high-speed acceleration," then... Small; Based on the final association weight Generate a working condition association weight matrix corresponding to the set of working condition status labels, which serves as the input basis for the subsequent control strategy generation.

4. The dynamic cognitive drive variable transmission ratio control method for a four-wheel independent steering system according to claim 3, characterized in that: Step 3) specifically includes: 31) Based on the set of working condition labels and the associated weight matrix output in step 2), set the yaw response gain, lateral acceleration response gain, and steering wheel feel stability as multi-objective constraints, and apply them to each working condition. For each of the three control objectives, a "performance deviation index" is defined, using the "squared relative error" method. and The difference between the actual yaw response gain and the ideal yaw response gain, and the difference between the actual lateral acceleration response gain and the ideal lateral acceleration response gain, are quantified using a "penalty term for exceeding the upper limit". The actual steering wheel torque fluctuation and the ideal steering wheel torque fluctuation are quantified, and the three single-target indicators are weighted and fused according to "control priority" to obtain a single-condition comprehensive indicator. : (4) In the formula, It is the target weight, which satisfies ; Weight matrix associated with working conditions By weighted summing of all single-condition comprehensive indicators, a global multi-objective control performance objective function is obtained. : (5) 32) Introduce the policy value function from deep reinforcement learning and design a weighted penalty reward function: (6) In the formula, This is the base reward; the reward is positive when there is no deviation. These are the weighting coefficients for the control objective, with a penalty term added. To prevent the strategy from deviating from the safe range; Value assessment of different transmission ratio adjustment strategies is conducted to form a value mapping model of working condition-target-strategy; 33) Historical strategies are filtered and optimized through a policy gradient mechanism, and the optimized strategies are then used. For optimal parameters, in the current state downsampling Each action yields an action set. Based on vehicle hardware limitations, actions exceeding the mechanical range are eliminated, and the 5th and 95th percentiles of the filtered actions are used as the adjustment range. , The target transmission ratio adjustment range is Output the target transmission ratio adjustment range and the initial parameters of the deep reinforcement learning control strategy under the current operating conditions, as the basis for online updating of the real-time control strategy.

5. The dynamic cognitive drive variable transmission ratio control method for a four-wheel independent steering system according to claim 1, characterized in that: Step 4) specifically includes: 41) Based on the target transmission ratio adjustment range and initial strategy parameters output in step 3), a deep neural network structure for the transmission ratio control strategy is established with the working condition state vector as input. 42) Using the real-time feedback of vehicle yaw rate, lateral acceleration, and steering wheel torque as the reward signal in reinforcement learning, define the state vector: (7) In the formula, It is the longitudinal speed of the vehicle. It is the yaw rate. It is lateral acceleration. It is the torque required to turn the steering wheel. It is an estimated value of the road surface adhesion coefficient. It is the deviation in yaw rate. , It's the front wheel steering angle. It's the wheelbase. It is a stability factor. It is lateral acceleration deviation. , It's a steering wheel torque deviation; Construct a reward signal based on real-time feedback, weighted sum of single-objective rewards, and add a safety constraint penalty term: (8) In the formula, ; Construct a state-action-reward triplet dataset to enable online training and dynamic updating of the policy network. 43) Based on the current transmission ratio With the target transmission ratio Generate the sequence using linear interpolation: (9) In the formula, This is the current transmission ratio. It is the target transmission ratio; Output the real-time transmission ratio control command sequence corresponding to the independent target steering angles of the left and right front wheels, and iteratively update the learned strategy network parameters to improve the adaptive and generalization capabilities of the control strategy under different operating conditions.

6. The dynamic cognitive drive variable transmission ratio control method for a four-wheel independent steering system according to claim 5, characterized in that: Step 5) specifically includes: pre-training the network with basic working condition data, gradually fine-tuning the upper output layer, and verifying it with extreme working conditions that were not involved in the training. The current iteration is stopped after the average reward reaches the target.

7. The dynamic cognitive drive variable transmission ratio control method for a four-wheel independent steering system according to claim 1, characterized in that: Step 6) specifically includes: 61) Based on the real-time control command sequence and updated strategy network parameters output above, when disturbance characteristics such as sudden changes in road adhesion conditions, sudden changes in steering wheel input, system response lag, or deterioration of yaw control performance are detected, if the abnormal threshold of any disturbance characteristic is met and it lasts for 2 control cycles, the strategy correction module is activated, the current working condition is marked as "disturbance working condition", and the strategy correction module is activated. 62) Introducing policy fine-tuning and experience replay mechanisms from deep reinforcement learning, priority sampling is performed on regions with abnormal rewards, and the policy network is retrained. To enhance optimization of low-reward samples, a reward-weighted term is added to the loss function, with samples having higher weights for lower rewards. (10) In the formula, It is PPO-clip loss For the dominant function, , It is the return-weighted loss. These are weighting coefficients; The output corrected control strategy incremental parameters are then weighted and fused together with the categorized compensation terms, superimposed onto the basic command, and the output adjustment signal compensation term is generated. (11) In the formula, To compensate for the weight, It is transmission ratio gain compensation. It is a compensation term for instruction smoothing. It is an advance compensation item. It is a yaw deviation compensation; to enhance the control robustness and response stability of the system in unstructured environments.

8. The dynamic cognitive drive variable transmission ratio control method for a four-wheel independent steering system according to claim 1, characterized in that: The four-wheel independent steering control system includes: hydraulic cylinder linear displacement sensor (1), hydraulic cylinder (2), proportional valve (3), valve core displacement sensor (4), steering wheel drive module (5), steering wheel angle sensor (6), vehicle speed sensor (7), PC (8), CAN controller (9), CAN transceiver (10), CAN HIGH (11), CAN LOW (12), main control module (13), and CAN bus (14). The hydraulic cylinder linear displacement sensor (1) is installed on the hydraulic cylinder (2) to detect the linear displacement of the hydraulic cylinder (2); the hydraulic cylinder (2) serves as a steering actuator to drive the steering wheel to rotate; the proportional valve (3) controls the oil flow and direction of the hydraulic cylinder (2); the valve core displacement sensor (4) detects the displacement of the valve core of the proportional valve (3); the steering wheel drive module (5) drives the steering wheel to achieve steering action; The steering wheel angle sensor (6) detects the rotation angle of the steering wheel; the vehicle speed sensor (7) detects the vehicle speed; the PC (8) serves as the host computer for system monitoring or parameter setting; the main control module (13) serves as the core control unit of the system, processing sensor signals and outputting control commands. The CAN controller (9) processes the CAN bus (14) communication protocol; the CAN transceiver (10) realizes the level conversion and signal transmission and reception between the CAN controller (9) and the CAN bus (14); the CAN HIGH (11) and CAN LOW (12) together constitute the twisted pair signal line of the CAN bus (14); the CAN bus (14) connects each control unit to realize the data communication and coordinated control of the vehicle steering system; The control system transmits the detection signals of the steering wheel angle sensor (6), vehicle speed sensor (7), hydraulic cylinder linear displacement sensor (1), and valve core displacement sensor (4) to the main control module (13) via the CAN bus (14). The main control module (13) generates control commands according to the control strategy and sends them to the steering wheel drive module (5) via the CAN bus (14) to drive the proportional valve (3) and hydraulic cylinder (2) to realize the four-wheel independent steering function.

Citation Information

Patent Citations

  • A variable transmission ratio optimization design method of a multi-mode drive-by-wire power steering system and a special system thereof

    CN109726516A

  • Variable transmission ratio design method and system for steer-by-wire system

    CN120337759A

  • Steer-by-wire vehicle variable transmission ratio design method and system based on fuzzy neural network

    CN120408845A