Vehicle slip rate double-layer control method based on visual identification and deep reinforcement learning

By introducing a double-layer control method of visual recognition and deep reinforcement learning in the vehicle control system, the accuracy problem of vehicle slip rate recognition and control in the prior art is solved, and efficient and precise braking under different road surface conditions is achieved.

CN120207289APending Publication Date: 2025-06-27WUHAN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510465104.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art has problems of low accuracy and sensitivity to external environmental interference in identifying and controlling vehicle slip rate, especially under different types of road surfaces and environmental changes.

Method used

A two-layer control method based on visual recognition and deep reinforcement learning is adopted to identify the road state through machine vision technology, and the vehicle's best slip rate is obtained by combining the deep Q-learning neural network. The lower layer uses an integrated slip mode controller to accurately track the slip rate.

Benefits of technology

It realizes high-precision identification of different road surface types and quickly adjusts the optimal slip rate. The lower control module can accurately track the optimal slip rate, verifying the accuracy of braking time and braking distance under different road surface conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120207289A_ABST
    Figure CN120207289A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of vehicle control, and particularly relates to a vehicle slip rate double-layer control method based on visual identification and deep reinforcement learning. A double-layer control structure is designed, the upper layer adopts a machine vision technology to recognize road surface state information, and tracking of the optimal slip rate under different road surface states is achieved in combination with a deep reinforcement learning algorithm. By using the control method of the invention, different road surface types can be identified with high precision, and the corresponding optimal slip rate can be rapidly and efficiently adjusted and tracked. And the lower layer adopts an integral sliding mode controller of an exponential approaching rate to accurately control the optimal slip rate. And the accuracy of the braking time and the braking distance under the optimal slip rate determined by different methods is verified. Through the structure, the accuracy of the braking time and the braking distance under the optimal slip rate determined by different methods is verified, and the excellent performance of the system under various road surface conditions is shown.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of vehicle control, and particularly relates to a dual-layer control method for vehicle slip rate based on visual recognition and deep reinforcement learning. Background Art

[0002] With the rapid development of intelligent connected vehicles, the public's expectations for the safety and reliability of automobiles have been continuously increasing. As an important active safety technology, the Anti-lock Braking System (ABS) has received extensive attention in the automotive industry in recent years. As the development trend of future braking systems, the brake decision-making structure and the actuator are the core technologies of the electronic brake-by-wire system, which have an important impact on the braking effect. In response to the above problems, many scholars have conducted a large amount of research in this area.

[0003] In braking decision-making, it is necessary to control the vehicle to stay near the optimal slip rate. Different types of road surfaces and environmental changes will directly affect the optimal slip rate of the vehicle. There are many methods for road surface condition recognition, which are mainly divided into two categories. One is to use sensors for recognition, and the second is to recognize based on the response of vehicle dynamics parameters, mainly using vehicle and tire dynamics models to achieve. The first type of method has high requirements for sensor accuracy, is more sensitive to external environmental interference, and is easily restricted; the second type is based on the slip rate curves obtained from different tire models, and the slip rate curves of different models often vary, and a large number of tests are required to obtain them, which has many uncertain effects on the system parameters and affects the accuracy of identification. Therefore, the present invention proposes a method combining deep reinforcement learning and vision to directly obtain the road surface slip rate. Summary of the Invention

[0004] In view of the problems existing in the prior art, the present invention provides a dual-layer control method for vehicle slip rate based on visual recognition and deep reinforcement learning.

[0005] The present invention is realized through the following technical solutions:

[0006] A dual-layer control method for vehicle slip rate based on visual recognition and deep reinforcement learning, comprising the following steps:

[0007] S1. First, based on the Electronic Hydraulic Brake (EHB), a model is established, and upper and lower layer control modules are designed;

[0008] S2. The upper layer control module collects road surface images through machine vision, and uses an adaptively improved YOLOv8 algorithm for road surface type classification;

[0009] S3. Based on the classification results, obtain the optimal slip ratio of the vehicle through a deep Q - learning neural network (DQN), specifically including the design of the environment and reward function, the construction of the agent, parameter setting, and training;

[0010] S4. Input the optimal slip ratio into the lower - layer control module, and track the slip ratio of the quarter - vehicle model based on the integral sliding - mode controller;

[0011] S5. Use Matlab / Simulink for simulation verification.

[0012] Furthermore, in step S1, the process of modeling the electro - hydraulic brake model is as follows:

[0013] When the braking system works, the ECU receives signals to control the motor, enabling the ball screw to push the hydraulic piston to move, generating hydraulic braking force and realizing the braking process. The motor model is

[0014]

[0015] where: \(u\) is the PWM duty cycle; \(U\) m is the voltage; \(L\) m is the inductance; \(i\) m is the current magnitude; \(R\) m is the electric drive resistance; \(K\) m is the back - electromotive - force coefficient; \(\omega\) a is the angular velocity of the motor rotor; \(J\) c is the equivalent rotor moment of inertia; is the rotor angular acceleration; \(K\) c is the torque coefficient; \(T\) d is the load torque;

[0016] During braking, it is divided into two parts: eliminating mechanical clearance and motor stalling. The time for eliminating clearance is very short, so only the motor - stalling part is considered, that is, \(\omega\) a = 0, The relationship between torque and current is as shown in the formula

[0017] \(T\) d = \(K\) c \(i\) m

[0018] The braking torque generated on the brake disc is as shown in the formula

[0019] \(T\) f = 2\(\mu_1F\) f \(r\) m

[0020] where the generated hydraulic braking force \(F\) f can be expressed by the formula as

[0021]

[0022] Where: μ1 is the friction coefficient of the brake disc; r m is the effective radius of the brake disc; η2 is the transmission efficiency of the ball screw; P d is the lead of the screw; A1 is the cross-sectional area of the master cylinder; A2 is the cross-sectional area of the brake cylinder;

[0023] When the motor is stalled, it needs to overcome the maximum static friction torque. Rearranging the above formula gives the output torque as

[0024]

[0025] Where: is expressed as the output torque of the wheel brake; T s is the maximum static friction torque of the brake.

[0026] Furthermore, in step S2, the road surface types include dry asphalt, wet asphalt, cement road, snow road, ice road surface, and dry cobblestones.

[0027] Furthermore, the specific operations of step S2 are as follows:

[0028] (1) Obtain 100 images of each of the six road surfaces from the public dataset roboflow, for a total of 600 images as the training set, and number them; the six road surfaces include dry asphalt, wet asphalt, cement road, snow road, ice road surface, and dry cobblestones;

[0029] (2) Then select 20 images of each road surface, for a total of 100 images as the validation set. This part is different from the training set images and is numbered;

[0030] (3) Define the corresponding sets and the labels of the six road surfaces respectively, and arrange the six types of road surfaces and the label number sequences in order;

[0031] (4) Use the yolov8n-cls model and set the number of training iterations to start training; save the training results and reserve the road surface types for training.

[0032] (5) The training results are automatically saved in the training directory, and then the trained road surface types are applied to the subsequent algorithms.

[0033] Furthermore, in step S3, the specific operations of the environment design are as follows:

[0034] Based on the optimal slip ratio data provided by the semi-empirical tire model and the optimal slip ratio identifier, construct a 7×7 grid, and convert the optimal slip ratios of the six basic road surfaces into the final positions in the grid matrix, and select the positions near the optimal slip ratio as the starting positions of the mass points according to different road surfaces;

[0035] λ = S(x, y)

[0036] λ * = S fin (x, y)

[0037] Where: λ is the percentage of slip rate, S(x, y) represents the corresponding horizontal and vertical coordinates in the grid matrix, and S fin (x, y) represents the optimal slip rate λ * The corresponding coordinate value;

[0038] During braking, the vehicle slip rate tracking is transformed into a quadruple {S, A, γ, R}. S represents the state set, and the lateral position deviation e of the mass point in the grid matrix is defined x and the longitudinal position deviation e y , A is the action set, defined as the motion unit u, and is in four directions: up, down, left, and right. γ is the discount factor, and R is defined as the value return function that evaluates the mass point tracking composed of errors and actions

[0039]

[0040] R in the formula t+k represents the cumulative return obtained k times after time t.

[0041] Furthermore, in step S3, the specific operation of designing the reward return function R is as follows: The reward return function R is defined as

[0042]

[0043] Where: represents the unit value in the motion direction, and e x is the lateral position deviation and e y is the longitudinal position deviation; P s is any state value in the reward value set {-10, -1, 5, 10}. The value -10 represents the edge line reward value and cannot be crossed. 5 represents the reward value in the area adjacent to the end point, 10 represents the reward value in the end point area, and -1 represents the reward value in the remaining areas, which can be represented by the matrix grid P. The formula is as follows:

[0044]

[0045] Design different area reward values according to the different position coordinates of the optimal slip rate. The mass point continuously approaches the position of the optimal slip rate according to the action state u(t) and is transformed into the corresponding slip rate of the wheel by formula (13).

[0046] Further, in step S3, the specific operations for the agent construction are as follows: Utilize the established virtual environment, then design a neural network to approximate the Q-value function, and create a target network with the same structure to stabilize the training process; then initialize an experience replay buffer to store the experiences of the interaction between the agent and the environment, and break data correlation through random sampling; the agent selects actions using the ε-greedy policy, interacts with the environment during the training process and stores the experiences in the buffer, and at the same time randomly samples a batch of experiences from the buffer to calculate the target Q-value, uses the mean squared error as the loss function to update the Q-network parameters, and synchronizes the parameters of the target network regularly.

[0047] Further, the specific operations of step S4 are as follows: Use sliding mode control to control the optimal slip ratio values on different road surfaces trained by the deep reinforcement learning algorithm. Realize the braking of the vehicle by building a quarter-vehicle model and combining it with the motor. Design a sliding mode surface by introducing an integral term, and in order to weaken the chattering phenomenon, replace the sgn(s) function in the controller with the saturation function sat(s).

[0048] Compared with the prior art, the present invention has the following technical effects:

[0049] The present invention designs a double-layer control structure, in which the upper layer uses machine vision technology to identify road surface state information and combines the deep reinforcement learning algorithm to achieve the tracking of the optimal slip ratio under different road surface states.

[0050] Using the control method of the present invention can accurately identify different road surface types, and quickly and efficiently adjust and track the corresponding optimal slip ratio. The lower layer uses an integral sliding mode controller with an exponential approach rate to precisely control the optimal slip ratio. And verify the accuracy of the braking time and braking distance under the optimal slip ratio determined by different methods. Through this structure, the present invention verifies the accuracy of the braking time and braking distance under the optimal slip ratio determined by different methods, demonstrating the excellent performance of the system under various road surface conditions. Description of the Drawings

[0051] Figure 1 It is a schematic diagram of the force on a single-wheel brake;

[0052] Figure 2 It is the control structure of the EHB system;

[0053] Figure 3 It is the slip ratio identification process;

[0054] Figure 4 It is a schematic diagram of the yolov8 structure;

[0055] Figure 5 It is the training result diagram of yolov8n;

[0056] Figure 6 It is the structural diagram of the DQN algorithm;

[0057] Figure 7 It is the network structure diagram;

[0058] Figure 8 It is the convergence graph of the cumulative reward of the agent training under different road surfaces;

[0059] Figure 9 It is the upper and lower layer control flow chart;

[0060] Figure 10 It is the simulation result graph of the dry asphalt road surface;

[0061] Figure 11 The simulation result graph of the wet asphalt road surface;

[0062] Figure 12 The simulation result graph of the snow road surface;

[0063] Figure 13 The simulation result graph of the road surface changing from snow to dry cement. Specific implementation manners

[0064] The following further describes the present invention in conjunction with embodiments and the accompanying drawings.

[0065] The present invention is modeled based on electronic hydraulic braking (EHB), designs an upper and lower layer control structure. The upper layer uses machine vision to identify the road surface and combines deep reinforcement learning to obtain the optimal slip ratio of the vehicle; the lower layer builds a quarter-vehicle model and uses an integral sliding mode controller (ISMC) to track the optimal slip ratio of the vehicle. Finally, Matlab / Simulink is used for simulation to verify the feasibility of this method.

[0066] 1. Modeling of the electro-hydraulic braking system

[0067] 1.1. 1 / 4 vehicle tire model

[0068] Ignoring the unevenness of the road surface and not considering the influence of the braking intensity on the normal forces of the front and rear wheels. Accordingly, a 1 / 4 vehicle model can be established to simplify the representation of the overall vehicle motion characteristics. The force analysis of the single-wheel braking motion is as Figure 1 shown.

[0069] During the wheel movement, the mechanical equation can be expressed as

[0070]

[0071] F f =μ c F z (2)

[0072] F z= N = mg(3)

[0073] The tire dynamics equation can be expressed as

[0074]

[0075] Where: m is the vehicle mass; v a is the vehicle speed; F f is the frictional force between the wheel and the ground; μ c is the longitudinal utilization adhesion coefficient of the tire; F z is the normal force of the ground on the tire; N is the wheel load; g is the acceleration due to gravity; J is the moment of inertia of the wheel; ω a is the angular velocity of the wheel; R is the wheel radius; T b is the braking torque.

[0076] The slip ratio λ can represent the sliding component during wheel movement, as shown in the formula:

[0077]

[0078] 1.2. Tire Model

[0079] The tire steady-state model is generally divided into theoretical models, empirical models, and semi-empirical models. The peak adhesion coefficient and the optimal slip ratio μ-s corresponding to different models are different. In this embodiment, the research is carried out based on the known empirical model, the Burckhardt model, and the optimal slip ratio estimated by the Kiencke_RLS optimal slip ratio identifier. The μ-s relationship corresponding to the two models is as follows:

[0080]

[0081] Using the above formula, the optimal slip ratio can be calculated as

[0082]

[0083] Where c1, c2, and c3 are parameters related to the road surface.

[0084] The optimal slip ratio λ c and the optimal slip ratio λ estimated by the identifier m can be shown in Table 1

[0085] Table 1 Comparison of Optimal Slip Ratios Obtained by Different Methods

[0086] Road surface <![CDATA[λ c > <![CDATA[λ m > Dry asphalt 0.1700 0.1512 Wet asphalt 0.1300 0.1314 Concrete road 0.1600 0.1546 Snowy road 0.0600 0.0732 Ice road surface 0.0315 0.0396 Dry cobblestones 0.4000 0.3842

[0087] 1.3. Electro-Hydraulic Brake Model

[0088] The hydraulic brake has the advantages of fast response and precise braking force control. The schematic diagram of its braking system is asFigure 2 as shown

[0089] When the system is working, the ECU receives signals to control the motor, making the ball screw drive the hydraulic piston to move, generating hydraulic braking force and realizing the braking process. Motor model:

[0090]

[0091] In the formula: u is the PWM duty cycle; U m is the voltage; L m is the inductance; i m is the current magnitude; R m is the electric drive resistance; K m is the back electromotive force coefficient; ω a is the angular velocity of the motor rotor; J c is the equivalent rotor moment of inertia; is the rotor angular acceleration; K c is the torque coefficient; T d is the load torque;

[0092] During braking, it is divided into two parts: eliminating mechanical clearance and motor stall. The time for eliminating clearance is very short, so only the motor stall part is considered, that is, ω a = 0, The relationship between torque and current is as shown in the formula

[0093] T d = K c i m (9)

[0094] The braking torque generated on the brake disc is as shown in the formula:

[0095] T f = 2μ1F f r m (10)

[0096] Among them, the generated hydraulic braking force F b can be expressed by the formula as

[0097]

[0098] In the formula: μ1 is the brake disc friction coefficient; r m is the effective radius of the brake disc; η2 is the transmission efficiency of the ball screw; P d is the lead of the screw; A1 is the cross-sectional area of the master cylinder; A2 is the cross-sectional area of the brake cylinder.

[0099] When the motor is stalled, it needs to overcome the maximum static friction torque. Rearranging the above formula, the output torque can be obtained as

[0100]

[0101] In the formula: represents the output torque of the wheel brake; T s is the maximum static friction torque of the brake.

[0102] 2. Optimal Slip Ratio Identification Design Based on Deep Reinforcement Learning

[0103] This embodiment proposes a new identification model that classifies road surfaces through the YOLOv8 algorithm and combines the DQN deep reinforcement learning algorithm to track the optimal slip ratio of various road surfaces. The identification process is as follows Figure 3 shown.

[0104] 2.1 Road Surface Classification Based on YOLOv8

[0105] YOLOv8 is a high-precision object detection algorithm, and its architecture consists of an input end (Input), a backbone network (Backbone), a neck network (Neck), and a prediction head (Head) (see Figure 4 ). First, the input end receives the road surface image, and the backbone network (Backbone) is responsible for extracting the low-level features of the image and generating partial feature maps. Next, the neck network (Neck) fuses these low-level features with other deep features to generate high-dimensional feature maps. Finally, the prediction head (Head) decodes these high-dimensional features to generate the final prediction result map. In the image classification task, YOLOv8 uses a Softmax classifier for classification determination.

[0106] Road Surface Classification Steps

[0107] (1) Obtain 100 images of each of the six road surfaces from the public dataset roboflow, for a total of 600 images as the training set, and number them.

[0108] (2) Then select 20 images of each road surface type, for a total of 100 images as the validation set. This part is different from the training library images and is also numbered.

[0109] (3) Define the corresponding libraries and the labels of the six road surfaces respectively, and arrange the six types of road surfaces in order corresponding to the label number sequence.

[0110] (4) Open the YOLOv8 algorithm in Visual Studio Code, create a new training window, call the yolov8n-cls model through the Python language, and set the training iteration times to start training.

[0111] (5) The training results are automatically saved in the training directory, and then the trained road surface types are applied to the subsequent algorithms.

[0112] The training results are as follows Figure 5As shown in the figure, "loss" in the figure represents the loss values of the training set and the validation set, indicating the error change during the training process. "metrics" is used as an evaluation metric. "accuracy_top1" represents the accuracy rate that the first candidate category predicted by the model is the correct category. As the number of iterations (epoch) increases, the accuracy becomes higher. "Accuracy_top5" represents the proportion that the correct category appears in the top five predicted categories, and the accuracy rate is 100%, indicating that the model can always include the correct category in the top five predicted categories.

[0113] After training for 500 rounds, at the 3rd round of training, the loss values of the training set and the validation set have started to converge and remain stable. The average recognition accuracy is as high as 99%. The following table shows the output road surface type numbers.

[0114] Table 2 corresponds to the road surface numbers

[0115] Road surface type Number Dry asphalt 1 Wet asphalt 2 Dry cement 3 Dry cobblestones 4 Snowy ground 5 Ice road surface 6

[0116] 2.2. Optimal slip ratio design based on DQN reinforcement learning

[0117] Reinforcement learning is a machine learning method that learns how to take actions to maximize by the interaction between the agent and the environment. The neural network is used to evaluate the environmental state and convert it into numerical inputs to calculate the Q-values of each action. These Q-values represent the cumulative rewards that can be obtained by selecting this action in the current state. The agent will execute the obtained optimal action, accept the new state, and update the network parameters through feedback data, continuously training and exploring to learn the optimal strategy (see Figure 6 ).

[0118] 2.2.1. Construction of the DQN reinforcement learning slip ratio environment

[0119] In this embodiment, based on the slip ratios provided by both the semi-empirical tire model and the optimal slip ratio identifier, a virtual grid matrix environment is designed. The appropriate optimal slip ratio is selected from Table 1 and substituted. Since the optimal slip ratio of common road surfaces will not be higher than 50%, a 7×7 grid is designed, and the optimal slip ratios of six basic road surfaces are converted into the final positions in the grid matrix by the formula, and the positions near the optimal slip ratio are selected as the starting positions of the mass points according to different road surfaces.

[0120] λ = S(x,y)(13)

[0121] λ * = S fin (x,y)(14)

[0122] In the formula, λ is the slip ratio percentage value, S(x,y) represents the corresponding horizontal and vertical coordinates in the grid matrix, and S fin (x,y) represents the optimal slip ratio λ *The corresponding coordinate values.

[0123] During braking, the vehicle slip ratio tracking can be transformed into a quadruple {S, A, γ, R}, as shown in the formula. S represents the state set, and the lateral position deviation e of the particle in the grid matrix is defined. x and the longitudinal position deviation e y , A is the action set, defined as the motion unit u, and it is in four directions: up, down, left, and right. γ is the discount factor, and R is defined as the value return function that evaluates the particle tracking composed of errors and actions.

[0124]

[0125] R in the formula t+k represents the cumulative return obtained k times after starting from time t. The reward return function R can be defined by the following formula

[0126]

[0127] In the formula represents the unit value in the motion direction, and P s is any state value in the reward value set {-10, -1, 5, 10}. The value -10 represents the edge line reward value and cannot be crossed. 5 represents the reward value in the area adjacent to the end point. 10 represents the reward value in the end point area. -1 represents the reward value in the remaining areas, which can be represented by the matrix grid P. The formula is as follows

[0128]

[0129] Design different area reward values according to different optimal slip ratio position coordinates. The particle continuously approaches the position of the optimal slip ratio according to the action state u(t) and is transformed into the corresponding slip ratio of the wheel by formula (13).

[0130] 2.2.2. Agent construction and training result analysis

[0131] Use a deep neural network to approximate the function Q(s t , a t ,), and adopt the value approximation function Q(s t , a t , θ - ) to replace the Q table. The corresponding formula is as follows:

[0132]

[0133] Q(s t , a t , θ - ) ≈ Q(s t | a t )(19)

[0134] Where: Q(s t |a t ) is the value of taking action a t in state s t , r t is the immediate reward obtained after taking the action, α ∈ (0, 1) is the learning rate, γ ∈ (0, 1) is the discount factor, representing the weight of future rewards, and θ - are the neural network parameters.

[0135] The DQN also needs to define a loss function L(θ) to measure the difference between the current Q value and the target Q value, and uses the gradient descent method to approach the target value. As shown in Equation (20), and to improve the training stability and data utilization efficiency, the "Experience Replay" method is usually adopted. Each time the learned data (S t , A t , R t , S t+1 ) is stored in the experience replay pool M. The experience data of length N is taken from the pool, and the average loss function value is calculated, as shown in Equation (22). The update of parameter θ is as shown in Equation (23):

[0136]

[0137] Where θ are the parameters of the behavior network (current Q value), and θ - are the parameters of the target (target Q value) network.

[0138] In network training, a neural network is also needed to approximate the state-action value function Q(s t |a t ), see Figure 7 . Among them, the input layer corresponds to the state s(t) in reinforcement learning, with two neurons, corresponding to the lateral position deviation e x and the longitudinal position deviation e y respectively. Then there are three fully connected layers. The first two layers are set to 100 and 200 neuron parameters respectively. The last layer needs to reduce the number of output features to be close to the target and is set to 30 neuron parameters. The activation function uses the Tanh function shown in Equation (24) to enable the network to fit complex functions. The last layer is the output layer corresponding to different actions.

[0139]

[0140] After building the agent and the environment model, three road conditions of dry asphalt, dry cobblestone, and snow are selected for training. The control algorithm is built in Matlab software, and the grid environments of the optimal slip ratios for the three road surfaces are established respectively. An agent with a learning rate of 0.001 is selected for training. The specific training parameters are shown in Table 3

[0141] Table 3 Training Parameters Corresponding to Different Road Surfaces

[0142]

[0143]

[0144] Construct an agent for training according to the parameters in the above table, and its convergence effect is as Figure 8 shown. Agents in three different road surface states can converge relatively quickly. When the number of training times reaches 160 times for the dry asphalt road surface, the return value of a single iteration converges stably to about 7, and the error range is within (-2, 0); when the number of training times reaches 188 times for the dry cobblestone road surface, the return value of a single iteration converges stably to about 8 and remains within the error range; when the number of training times reaches 52 times for the snow road surface, the return value of a single iteration converges stably to about 2 and converges to the predetermined range.

[0145] 3. Design of Slip Ratio SMC Controller

[0146] Sliding mode control (SMC) shows strong robustness to model uncertainties and external disturbances, has a fast response speed, is insensitive to parameter changes and has anti-disturbance ability, and can handle nonlinear systems. In the vehicle braking system, the friction characteristics of different road surfaces are highly nonlinear, and the vehicle is affected by continuously changing system parameters such as load and tire wear. Therefore, combining the advantages of sliding mode control, the sliding mode control theory is selected to design the EHB controller.

[0147] According to the output results of the reinforcement learning algorithm, select the vehicle slip ratio around λ * and the controller tracking error is e = λ * -λ(25)

[0148] To improve the system stability, design the integral sliding mode surface function as:

[0149]

[0150] where: C is a constant; t is time

[0151] Take the derivative of the above formula to get

[0152]

[0153] Then substitute the formula into the above formula to get

[0154]

[0155] Select the exponential reaching law function:

[0156]

[0157] Where: ε is the gain, used to weaken interference and reduce chattering; k is the feedback gain, both greater than 0; sgn(s) is the sign function. By controlling the current i m is the input of the EHB controller, denoted by u a and can be derived from the above formula as

[0158]

[0159] Define the Lyapunov function as

[0160] Substitute into the above formula to satisfy which can represent the system stability.

[0161] To weaken the chattering phenomenon, replace the sgn(s) function in the controller with the saturation function sat(s), and the function is expressed as:

[0162]

[0163] Where: k = 1 / Δ, is the saturation bandwidth; Δ is the boundary layer.

[0164] 4. Simulation verification

[0165] 4.1. Simulation analysis

[0166] To verify the performance of the proposed strategy, use Matlab / Simulink software for simulation and result analysis. By establishing a two-layer structure, as Figure 9 shown, the upper layer is the road surface recognition layer, which uses vision and deep reinforcement learning methods to identify the slip ratio, and the lower layer is the brake actuator layer, which builds a quarter-vehicle model, an integral sliding mode controller model, and an actuator model.

[0167] In this embodiment, the optimal slip ratio λ provided by the semi-empirical tire model c and the slip ratio λ provided by the slip ratio identifier m are used to verify the performance of the controller. Two types of road surfaces are simulated respectively: braking tests on a single road surface type and a variable road surface type. In the single-road surface braking scenario, research is carried out on three road surface types: dry asphalt, wet asphalt, and snow. In the variable road surface scenario, a gradually changing road surface from snow to dry cement is set, and the road surface is switched after 1 s. Braking is carried out when the vehicle speed v0 is 20 m / s. As Figures 10 - 13 are the simulation results, including speed change, braking torque, current magnitude, braking distance change, and slip ratio error curve. The simulation parameters are shown in Table 4.

[0168] Table 4 Simulation parameters

[0169]

[0170] 4.1. Result Analysis

[0171] It can be seen from Figure 10 and Figure 11 that under the same road surface type, the braking distance at the optimal slip ratio is shorter. Taking the dry asphalt road surface as an example, when the optimal slip ratios are 0.17 and 0.1512 respectively, the braking time and distance of the vehicle are (1.89 s, 18.9 m) and (1.98 s, 19.9 m) respectively. On the wet asphalt road surface, when the optimal slip ratios are 0.13 and 0.1314 respectively, the braking distance and time are (2.75 s, 27.6 m) and (2.97 s, 29.8 m). Since the peak adhesion coefficient of the wet asphalt road surface is smaller than that of the dry asphalt road surface, the braking distance is longer and the braking time will also increase at the optimal slip ratio. As Figure 12 shown, when braking on the snow road surface, when the optimal slip ratios are 0.06 and 0.0732 respectively, the braking distance and time are (11.6 s, 115.9 m) and (13.42 s, 134.2 m). The asphalt road surface is a high-adhesion road surface with an adhesion coefficient greater than 0.8. Compared with the low-adhesion road surface, the adhesion coefficient of the snow is less than 0.2. Therefore, the braking distance and time are longer than those of the asphalt road surface at the optimal slip ratio. As Figure 13 shown, when the road surface changes from the snow road surface to the dry cobblestone road surface, the slip ratios change from 0.06 to 0.16 and from 0.0732 to 0.1546 respectively, and the braking time and distance are (2.15 s, 22.6 m) and (2.21 s, 23.8 m) respectively, indicating that when the ISMC controller adopts the optimal slip ratio control under the current road surface conditions, it can also achieve the optimal effect.

[0172] The present invention designs a double-layer control structure. The upper layer uses machine vision method to identify road surface state information and combines with the deep reinforcement learning algorithm to track the optimal slip ratio of different road surface states. The lower layer uses an integral sliding mode controller with an exponential approach rate to control the optimal slip ratio, and the following conclusions are verified through simulation experiments:

[0173] (1) Classify and verify different road surfaces through the machine vision method, and the average accuracy can reach about 99%.

[0174] (2) Use the deep reinforcement learning algorithm to track and verify the optimal slip ratio under different road surface states, and the agents in different states can converge and stabilize quickly.

[0175] (3) In the optimal slip ratio control experiment, this control method can accurately control different slip ratios, verifying the accuracy of the braking time and braking distance at the optimal slip ratio.

Claims

1. A dual-layer control method for vehicle slip rate based on visual recognition and deep reinforcement learning, characterized in that: The steps include: S1. First, model the electronic hydraulic brake EHB as the basis and design the upper and lower control modules; S2, the upper control module collects road images through machine vision and uses the improved YOLOv8 algorithm to classify road types; S3. Based on the classification results, the optimal slip rate of the vehicle is obtained through a deep Q-learning neural network (DQN), which includes environment and reward function design, agent construction, parameter setting and training; S4, inputting the optimal slip ratio into the lower control module, and tracking the slip ratio of the quarter vehicle model based on the integral sliding mode controller; S5. Use Matlab / Simulink for simulation verification.

2. The vehicle slip rate dual-layer control method based on visual recognition and deep reinforcement learning according to claim 1 is characterized in that: In step S1, the process of modeling the electronic hydraulic brake EHB is as follows: When the brake system is working, the ECU receives signals to control the motor, causing the ball screw to push the hydraulic piston to move, generating hydraulic braking force and realizing the braking process. The motor model is Where: u is the PWM duty cycle; U m is voltage; L m is the inductance; i m is the current size; R m is the electric drive resistance; K m is the back electromotive force coefficient; ω a is the motor rotor angular velocity; J c is the equivalent rotor moment of inertia; is the rotor angular acceleration; K c is the torque coefficient; T d is the load torque; The braking process is divided into two parts: eliminating mechanical clearance and motor stalling. The time to eliminate clearance is very short, so only the motor stalling part is considered, that is, ω a =0, The relationship between torque and current is as follows: T d =K c i m The braking torque generated on the brake disc is as follows: T f =2μ1F f r m Among them, the hydraulic pressure braking force F f It can be expressed as Where μ1 is the friction coefficient of the brake disc; r m is the effective radius of the brake disc; η2 is the transmission efficiency of the ball screw; P d is the lead of the screw rod; A1 is the cross-sectional area of ​​the master cylinder; A2 is the cross-sectional area of ​​the brake cylinder; The motor needs to overcome the maximum static friction torque when it is stalled. The output torque can be obtained by rearranging the above formula: Where: Expressed as the output torque of the wheel brake; T s is the maximum static friction torque of the brake.

3. The vehicle slip rate dual-layer control method based on visual recognition and deep reinforcement learning according to claim 1 is characterized in that: In step S2, the road surface types include dry asphalt, wet asphalt, cement road, snow road, icy road, and dry cobblestone.

4. The vehicle slip rate dual-layer control method based on visual recognition and deep reinforcement learning according to claim 1 is characterized in that: The specific operations of step S2 are as follows: (1) 100 images of each of the six types of road surfaces were obtained from the public data set roboflow, a total of 600 images as the training set, and they were numbered; the six types of road surfaces include dry asphalt, wet asphalt, cement road, snow road, ice road, and dry cobblestone; (2) Select 20 images of each road surface, a total of 100 images, as the validation set. These images are different from the training library images and are numbered. (3) Define the corresponding sets and labels of the six types of road surfaces, and arrange the six types of road surfaces in order corresponding to the label number sequences; (4) Use the yolov8n-cls model and set the number of training iterations to start training; Save the training results and keep the trained road type for later use. (5) The training results are automatically saved in the training directory, and then the trained road type is used in the subsequent algorithms.

5. The vehicle slip rate dual-layer control method based on visual recognition and deep reinforcement learning according to claim 2 is characterized in that: In step S3, the specific operations of environment design are as follows: Based on the optimal slip rate data provided by the semi-empirical tire model and the optimal slip rate identifier, a 7×7 grid is constructed, and the optimal slip rates of the six basic road surfaces are converted into the final positions in the grid matrix. According to different road surfaces, the position near the optimal slip rate is selected as the starting position of the particle. λ=S(x,y)(13) l * =S fin (x,y)(14) Where: λ is the percentage of slip rate, S(x,y) represents the corresponding horizontal and vertical coordinates converted into the grid matrix, S fin (x,y) represents the optimal slip ratio λ * The corresponding coordinate values; The vehicle slip rate tracking during braking is converted into a quaternion {S, A, γ, R}, where S represents the state set and defines the lateral position deviation e of the particle in the grid matrix. x and longitudinal position deviation e y , A is the action set, defined as the motion unit u, and in four directions, γ is the discount factor, and R is defined as the value return function of the evaluation particle tracking composed of error and action R in the formula t+k It represents the cumulative return after k times starting from time t.

6. The vehicle slip rate dual-layer control method based on visual recognition and deep reinforcement learning according to claim 2 is characterized in that: In step S3, the specific operation of the reward return function R is as follows: The reward function R is defined as Where: Indicates the unit value in the direction of motion, e x is the lateral position deviation and e y is the longitudinal position deviation; P s is any state value in the reward value set {-10,-1,,5,10}. The value -10 represents the edge line reward value and cannot be crossed. 5 represents the reward value of the area adjacent to the end point. 10 represents the reward value of the end point area. -1 represents the reward value of the remaining area. It can be represented by the matrix grid P. The formula is as follows: The reward values ​​of different areas are designed according to the coordinates of the optimal slip rate position. The particle continuously approaches the position of the optimal slip rate according to the action state u(t), and is converted into the slip rate corresponding to the wheel by formula (13).

7. The vehicle slip rate dual-layer control method based on visual recognition and deep reinforcement learning according to claim 2 is characterized in that: In step S3, the specific operations of the agent construction are as follows: using the built virtual environment, a neural network is designed to approximate the Q-value function, and a target network with the same structure is created to stabilize the training process; then an experience replay buffer is initialized to store the experience of the agent's interaction with the environment, and data correlation is broken by random sampling; the agent adopts the ε-greedy strategy to select actions, interacts with the environment during training and stores the experience in the buffer, and randomly samples a batch of experience from the buffer to calculate the target Q value, uses the mean square error as the loss function to update the Q network parameters, and periodically synchronizes the parameters of the target network.

8. The vehicle slip rate dual-layer control method based on visual recognition and deep reinforcement learning according to claim 2 is characterized in that: The specific operation of step S4 is as follows: using sliding mode control to control the optimal slip rate values ​​of different road surfaces trained by the deep reinforcement learning algorithm, building a quarter vehicle model and combining it with a motor to achieve vehicle braking, designing the sliding mode surface by introducing an integral term, and in order to weaken the vibration phenomenon, using the saturation function sat(s) to replace the sgn(s) function in the controller.