A braking system and method
The braking system uses a Q-network to dynamically adjust brake pressure based on machine learning, addressing the inefficiencies of ABS by minimizing wheel lock-up and ensuring safer, quicker stops with improved steering control.
Patent Information
- Application Number
- PCT/TR2024/051117
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-10-01
- Publication Date
- 2025-11-27
AI Technical Summary
Existing anti-lock braking systems (ABS) increase braking distance on snowy and gravel surfaces, and there is a need for a system that minimizes wheel lock-up to achieve safer and quicker stops.
A braking system utilizing a Q-network based on machine learning to adjust brake pressure dynamically, learning from external conditions to operate at the global maximum point before partial lock-up, ensuring optimal brake pressure without wheel lock-up, integrated with a brake control unit and hydraulic control system.
The system achieves shorter braking distances and maintains 100% steering control by operating in the full roll zone, reducing the risk of wheel lock-up and enhancing safety and performance on various road conditions.
Smart Images

Figure TR2024051117_27112025_PF_FP_ABST
Abstract
Description
[0001] A BRAKING SYSTEM AND METHOD
[0002] Technical Field
[0003] The present invention relates to an anti-lock braking system and braking method for ensuring a safer and quicker stop for all types of wheeled motorized land and air vehicles.
[0004] Prior Art
[0005] Motorized air and land vehicles have wheels used to move the vehicle forward, and braking systems applied to the wheels are used to stop or slow down the vehicle movement. In the aforementioned braking systems, the calipers close with the pressure of the oil in the hydraulic mechanism, and with the closing of the calipers, the brake pads and the discs in the wheels stick to each other, thus slowing down or stopping the vehicle.
[0006] In the prior art, full locking of the wheel occurs when the vehicle is under braking, reducing vehicle maneuverability. In the current technique, ABS (Anti-lock Braking System) is the solution to the technical problem of locking of the vehicle wheel. The said ABS (Anti-lock Braking System) is a braking system that prevents the wheels from locking during sudden braking of vehicles in compulsory situations at various speeds under all kinds of load conditions and all road conditions. ABS is a system that controls the change in the speed of each wheel in braking situations through a control unit. In the event of a sudden decrease in the number of revolutions (e.g. when braking on slippery surfaces) and locking of the wheel, the control unit automatically reduces the brake pressure. When the wheel accelerates again, the wheel is braked by increasing the brake pressure again. This phase takes place many times per second. The wheel can be in 3 states during braking. The first state is when the wheel is completely, i.e. 100% locked. The second state is partial lock-up. The third state is the zone where the wheel is in full roll during braking. ABS is activated in the event of 100% or partial locking of any one of the wheels, adjusting the brake pressure to lock the wheel by 20%, i.e. it operates in the partial lock-up zone. The ABS braking system is mainly used to maintain the steering control of the vehicle when the wheels are locked during braking and also provides a reduction in braking distance compared to conventional braking. However, although the ABS braking system improves steerability on gravel or thick snow, the braking distance is greater than conventional braking, and the reason for this is unknown. The ABS braking system increases the braking distance on snowy and gravel surfaces rather than decreasing it compared to conventional braking. Low braking distances for vehicles are very important for vehicle and passenger safety.
[0007] Patent application US6622077B2 in the prior art relates to an ABS braking system for preventing locking of vehicle wheels under braking, in which brake pressure can be adjusted using wheel speed sensor data.
[0008] Brief Description of the Invention
[0009] The present invention aims to realize a braking system for all types of wheeled motorized land and air vehicles to ensure a safer and quicker stop.
[0010] This invention aims at achieving a braking system that provides a minimum or zero wheel lock-up rate, thereby enabling a quicker stop.
[0011] This invention aims at achieving a braking system that allows the optimum brake pressure to be adjusted without lock-up ratio data in the full rolling zone where partial or full lock-up does not occur.
[0012] This invention aims to develop a braking system that learns the best braking performance value through a learning method aided by machine-learning on the basis of the external conditions that may affect the braking distance, and that can be applied in vehicles.
[0013] Detailed Description of Invention
[0014] Figures relating to the braking system realized to achieve the purpose of the present invention:
[0015] Figure 1 is a representative view of the working algorithm of the brake module to be used in the braking system which is subject of the invention.
[0016] Figure 2 is a representative view of the working algorithm of the braking system which is subject of the invention.
[0017] Figure 3 is a graph of the vehicle deceleration rate, lock-up rate and the pressure applied to the calipers of the braking system which is subject of the invention as well as the full roll, partial lock-up, and full lock-up phases of the wheel.
[0018] Figure 4 is a representative view of the working algorithm of the braking system in one embodiment of the invention.
[0019] The parts in the figures are individually numbered, and the parts shown by these numbers are given below.
[0020] 1. Braking system
[0021] 2. Brake module
[0022] 21. Experience replay memory
[0023] 22. Environment simulator
[0024] 23. Training data
[0025] 24. Target neural network
[0026] 3. Q network
[0027] 33. LSTM neural network
[0028] 34. Cell state 35. Hidden state
[0029] 36. Sensor data (It)
[0030] 4. Brake control unit
[0031] 5. Hydraulic control unit
[0032] A braking system (1) that is used in vehicle braking systems to ensure a safer and quicker stop / braking for all types of wheeled motorized land and air vehicles, and that is characterized with the following:
[0033] - Used in the vehicle and adapted to work as integrated to the algorithm in the vehicle brake control unit (4) processor,
[0034] - Adapted to receive sensor data in response to changing conditions acting on the vehicle and wheel, and to adjust the pressure value so that the wheel operates at its global maximum point prior to partial lock-up, and
[0035] - At least one Q network (3) consisting of an artificial neural network that learns how to provide optimum brake pressure under changing conditions and a brake module (2) used for training this network.
[0036] The Q network (3) of the invention is trained in the brake module (2), and after the training is completed, it is integrated into the software in the braking system (1) and used as a decision algorithm. The inventive brake module (2) is used so that the Q network (3) can learn to calculate the optimum brake pressure of the vehicle wheel under different conditions without full or partial lock-up of the wheel, and operates within the processor in the brake control unit (4) in the braking system used in vehicles for braking. The said brake module (2) is adapted so that the Q network (3) can learn to calculate the variation of the global maximum point of the wheel before lock-up for different conditions of the wheel and the optimum pressure required for a safe and quick stop of the vehicle. It is necessary to train the Q network (3) to be included in the brake control unit by performing tests with the data received from the sensors under different conditions such as snowy, icy, wet, dry ground, different road slopes, different vehicle speeds. The more different conditions the Q network (3) algorithm is trained under, the better it will work. While the ABS braking system tries to keep the wheel lock-up rate at 20% as shown in Figure 3, the Q network (3) of the invention operates at the global maximum point of the wheel in the full roll zone, before 1 % lock-up, where better deceleration and 100% steering control will be achieved, as shown in Figure 3. The ABS braking system operates in the partial lock-up zone where the longitudinal friction acting on the wheel is both static and kinetic frictions. The braking system (1) operates in the full roll zone where there is no longitudinal kinetic friction and only static friction acts longitudinally on the wheel. The brake pressure value in this zone, corresponding to the global maximum point where the deceleration is maximum, is constantly changing in response to changing conditions (such as speed, brake pressure, vehicle weight, tire tread depth and pressure). The curve in Figure 3 remains generally the same in shape but changes its shape and position on the graph to some extent in response to changing conditions.
[0037] The invention comprises a brake control unit (4) used in vehicle braking systems and integrated into the vehicle to ensure safer and quicker stopping / braking in all types of wheeled motorized land and air vehicles, a brake module (2) trained in vehicle braking system tests, and at least one Q network (3) adapted to receive sensor data based on changing conditions affecting the wheel, tracking the change in the brake pressure value at the global maximum point before partial locking of the wheel, and learning to apply the optimum brake pressure under changing conditions.
[0038] The said brake module (2) has at least one Q network (3) consisting of an artificial neural network to learn and use the optimum pressure under changing conditions. Neural networks have connection weights as parameters. These weights are selected randomly at the beginning of training. These weights are gradually updated during the training period through reinforcement learning.
[0039] The Q network (3) receives data from sensors on the vehicle, on the basis of changing conditions acting on the wheel. It adjusts the brake pressure by the sensor data it receives. In one embodiment of the invention, the brake module (2) is adapted to train the Q network (3) through machine learning tests to enable it to learn the variation of the global maximum point prior to wheel lock-up in response to changing conditions.
[0040] In one embodiment of the invention, the Q-network (3) operates in a state-actionstate cycle to perform braking.
[0041] In one embodiment of the invention, the brake Q-network (3) performs the braking operation with the “Reinforcement learning” artificial intelligence technique featuring the “Partially Observable Markov Decision Process” (POMDP) in which the past is taken into account. Thanks to POMDP, it works even when there is a missing state value, i.e. when there are not enough sensor data about the operating environment it is in. For this, either the Q network (3) should be provided with past sensor values or an artificial neural network such as RNN or LSTM that keeps statistics of the past must be integrated into the decision network (3).
[0042] In one embodiment of the invention, the Q network (3) has the optimization property as it is an artificial neural network. In this way, it can work at the global maximum point in this system, aiming for maximum deceleration.
[0043] In one embodiment of the invention, the brake module (2) is used in training tests of the Q network (3). In one embodiment of the invention, the said Q-network (3) may be a fully connected artificial neural network or a combination of different artificial neural networks.
[0044] In one embodiment of the invention, the brake module (2) braking method (200) includes the following method steps:
[0045] 201. Start of the process by applying the brake,
[0046] 202. Transmission of data from the sensors in the tested vehicle to the Q network, which is an artificial neural network, 203. Measuring the brake fluid pressure delivered to the brake when approximately 3% wheel lock-up occurs,
[0047] 204. Calculating the reward with the new sensor data,
[0048] 205. Updating the Q network weights with the reward value,
[0049] Repeating the process steps 202, 203, 204, 205 continuously until the Q network is ready,
[0050] 206. That the Q network is in the learned state and ready to be integrated as Q network in the vehicle control unit.
[0051] In this way, the artificial neural network learns the optimum use of pressure under changing conditions and can be integrated as a Q network into the braking systems of vehicles of the same model to be produced.
[0052] In one embodiment of the invention, the brake module (2) is in data communication with at least one or more or all of the wheel speed sensor, vehicle acceleration sensor, vehicle tilt sensor, steering angle sensor and / or tire pressure sensor in the vehicle under test, and the data from these sensors are provided as input to the brake module (2). The rewards to be used in the brake module (2) are calculated from wheel longitudinal force sensor data, vehicle acceleration sensor data, and wheel lock-up rate. The wheel brake fluid pressure value is stored and updated when 3% wheel lock-up occurs. For time steps after the update, this brake fluid pressure value is taken as the z value until another 3% lock-up occurs and used as the state value. Then, in the later time steps, new sensor data and this new state information are fed into the Q network (3) algorithm. This is a duty cycle of the algorithm, followed by a new cycle. In today's commonly used hydraulic control units, approximately 15- 20 cycles occur in 1 second during braking. In practice, measuring 1% lock-up is misleading, because every time the brakes are applied, a small amount of partial wheel lock-up occurs. The same is true for acceleration. Each time the vehicle accelerates, a small amount of partial skidding occurs. Therefore, with the tests to be performed, the beginning of the lock-up limit in practice can be accepted as a value of approximately 3%. The brake fluid pressure value is stored and updated when 3% wheel lock-up occurs again in the subsequent time steps. For time steps after the update, this brake fluid pressure value is taken as the z value until another 3% lock-up occurs and used as the state value.
[0053] In one embodiment of the invention, the wheel longitudinal force sensor is conveniently positioned between the suspension system and the chassis and measures the wheel longitudinal force exerted by the wheel on the chassis during braking. It is sufficient to use it as a reward in the brake module (2) only during the training of the algorithm. There is no need to aim to measure this force perfectly when placing the sensor in the construction. In order for the algorithm to be able to make a comparison, it is sufficient for the sensor to present to the algorithm of the brake module (2) how the force changes relative to the off-brake state and during braking. 1 unit must be used for each vehicle wheel and must measure the force in the direction of the wheel for both the front and rear wheels.
[0054] The steering angle sensor value “Y” is the absolute angle value and only used in the algorithm for the front wheels. The tilt sensor measures the tilt of the vehicle, i.e. the slope of the road. When the brake module (2) and the decision network (3) use a reinforcement learning algorithm that complies with the Partially Observable Markov Decision Process, they can more or less perform their tasks without one or more of the state values such as tire pressure and temperature, brake fluid pressure, which have a minor effect.
[0055] In one embodiment of the invention, the output of the algorithm, i.e. its actions, depends on the mechanical structure of the braking system. The actions of the algorithm can be continuous or discrete depending on this structure. In other words, the brake module (2) and Q network (3) algorithms can be adapted to braking system hardware with discrete output (on-off actuators), to hardware with continuous output, and to actuators with both continuous and discrete outputs. In one embodiment of the invention, the brake module (2) comprises an experience replay memory (21), training data (23) (mini batch), a loss, a target network (24), and a Q network, also used as a decision network. When the Q network, which is a decision network, completes the necessary training, it is integrated into the vehicle braking system and can adjust the required brake pressure not only for the situations it encountered during the training but also for the situations it has not experienced and can transmit commands to apply the pressure.
[0056] In one embodiment of the invention, the brake module (2) algorithm is as follows: The brake module (2) presented here as an example of reinforcement learning algorithms that can be used in the aforementioned brake module (2) is a Deep Q Network (DQN) artificial intelligence technique with POMDP feature.
[0057] Algorithm (Passenger car):
[0058] The states, or inputs, of the algorithm are:
[0059] » vehicle speed (to be calculated from wheel speed sensors and acceleration) V » 4 wheel lock-up rate (to be calculated) C = { Oi, O2, O3, O4 }
[0060] » 4 wheel speed sensors Hi = { Hi, H2, H3, H4 }
[0061] » 4 brake fluid pressure sensors Bi = { Bi, B2, B3, B4 }
[0062] » vehicle acceleration sensor AC
[0063] » vehicle tilt sensor E
[0064] » steering angle sensor Y (only used in the front wheel algorithm)
[0065] » 4 tire pressure sensors
[0066] Li = { Li, L2, L3, L4 }
[0067] » 4 tire temperature sensors
[0068] Si ={Si, S2, S3, S4}
[0069] >>brake fluid pressure value at moment of 3% partial lock-up Zi = { zi, Z2, Z3, Z4 }
[0070] The rewards to be used in the algorithm are:
[0071] » 4 wheel longitudinal force sensors Fi = { Fi, F2, F3, F4 }
[0072] » 4 wheel lock-up rate (to be calculated) Oi » 1 vehicle acceleration sensor AC
[0073] Vehicle speed and wheel lock-up rates are calculated with the data from the wheel speed sensors. In the brake fluid line between the hydraulic control unit (5) and the brake calipers, it would be useful to have one brake fluid static or total pressure sensor for each wheel for the algorithm to work better. This pressure sensor is placed between the hydraulic control unit (5) and the brake caliper for each wheel in the brake fluid line.
[0074] The wheel longitudinal force sensor is conveniently placed between the suspension system and the chassis and measures the “wheel longitudinal force” exerted by the wheel on the chassis during braking; and it is sufficient to use it as a reward only during the training of the algorithm. There is no need to aim to measure this force perfectly when placing the sensor in the construction. In order for the algorithm to be able to make a comparison, it is sufficient for the sensor to present to the algorithm how the force changes relative to the off-brake state and during braking. 1 unit must be used for each wheel and must measure the force in the direction of the wheel for both the front and rear wheels.
[0075] In practice, measuring 1% lock-up is misleading, because every time the brakes are applied, a small amount of wheel lock-up occurs. Same is true for acceleration. Each time vehicle accelerates, a small amount of skidding occurs. Therefore, it is possible to detect beginning of lock-up limit in practice with tests to be performed, and this rate is accepted as 3% approximately. The “z ” values are stored and updated when 3% wheel lock-up occurs. For time steps after the update, this z value is taken as the z value until another 3% lock-up occurs and used as the state value.
[0076] When the brake module (2) and the decision network (3) use a reinforcement learning algorithm that complies with the Partially Observable Markov Decision Process, they can more or less perform their tasks without one or more of the state values such as tire pressure and temperature, brake fluid pressure, which have a minor effect. Current state: current cycle state values (St)
[0077] Next state: next cycle state values (St+i)
[0078] Reward: reward (R)
[0079] Action: action (Ak) - A1: action to be applied to wheel 1 t: current cycle t-1: previous cycle t+1: next cycle
[0080] • i index: index of the relevant wheel
[0081] Where S = { V, Oi,.. Oi ,Hi,.. Hi ,Bi,.. Bi , AC , E , Y ,Ei,.. ,Si,.. Si ,zi,..Zi} It:
[0082] It = { At-n-i , { St-n } , • • • , AM , { St } } with “n” being an arbitrary number, values going n time steps back are given to the algorithm, thus enabling the algorithm to make comparisons against the past. The past state values that are not yet possible for the It set elements are the residual values from the previous braking job when the Q network first started running. As the time step occurs and values are available, they are replaced by their measured and calculated values respectively.
[0083] Arbitrary Parameters: y (Discount Factor, a value between 0 and 1), a (braking coefficient, a value between 0 and 1), K, M, P, c, m, x, g
[0084] For training the Q network, i.e. updating its weights, the backpropagation algorithm
[0011] and the “Stochastic Gradient Descent”
[0010] optimization algorithm can be used.
[0085] The “Mean Squared Error” method can be used to calculate the loss value.
[0086] Reward: R = Ri+ R2+R3+R4
[0087] Under difficult conditions, in both acceleration and deceleration, exactly 0% relative speed between the tire and the ground is not possible. Therefore, the onset of partial lock-up must in practice be taken at around 3%. R1 : When the first partial or full lock-up of the wheel occurs, the hydraulic pressure at the moment of 3% lock-up is recorded, and this value is called z. The z value is used throughout the episode, i.e. until the algorithm is deactivated. When the algorithm is activated again, a new z value is assigned.
[0088] Ri = c*(B-z) / B for O < 3%
[0089] R2 = - O*m for 3% < O < 100%
[0090] R3= F * x for 0 < O < co
[0091] R4= AC * g for 0 < O < co
[0092] R = Ri +R2 +R3+R4
[0093] Example algorithm for the 1st wheel (k=l) of a 4- wheeled vehicle (w=l, 2, 3, 4) using a hydraulic control unit (5) with discrete output.
[0094] In one embodiment of the invention, a method of training a brake module (2) for passenger cars;
[0095] Providing the target network (24) and the Q network (3) with the same weights and providing the existing experience replay memory (21),
[0096] 101. Reading data coming from sensors upon sudden braking,
[0097] 102. Saving the values It (including At-i), At and Rtand It+iin an experience replay memory (21),
[0098] 103. Transmission of It set values to the Q network (3) and operation of the Q network,
[0099] 104. Generating 3 values with a sum of 1 in the softmax function in the Q network and selecting exclusively the output with the highest value from the increase, decrease, and hold outputs as the action,
[0100] 105. Execution of braking by transmitting the selected action or an arbitrary action according to epsilon-greedy as a command to the hydraulic control unit (5),
[0101] 106. Saving the resulting It+i, At+i, Rt+i, and It+2 as experience data sets in the memory,
[0102] 107. Repeating steps 102, 103, 104, 105 and 106 for an arbitrary number of time steps “K”, - 108.
[0103] 109. Saving in experience replay memory (21),
[0104] 110. Transfer of as many data sets as the arbitrary number of “M” from the training data (23) in the experience replay memory to the training batch,
[0105] 111. Taking one of these sets from the training batch and finding current states, next cycle states, current action, and reward in the set,
[0106] 112. Using the current states It from the selected set as input for the Q network (3), and receiving outputs from the Q network (however, they are not given to the hydraulic control unit),
[0107] 113. Giving, out of the resulting outputs, the QA value of the previously applied action (At) in the selected set to Loss,
[0108] 114. Giving the next cycle states (It+i) from the selected set as input to the target network (24) and receiving outputs from the target network,
[0109] 115. Assigning the QB value with the highest value among these outputs to Loss,
[0110] 116. Giving the reward (Rt) to Loss,
[0111] Repeating operations from 110 to 116 for an arbitrary number “M”,
[0112] 117. Training the Q network with gradient backpropagation, using the resulting Loss value,
[0113] 118. Repeating the training of the Q network for an arbitrary number “P”,
[0114] 119. Then, copying the resulting new Q network weights to the target network (24) so that the two networks are identical again,
[0115] Repetition of operations from 102 to 119 until the Q network (3) receives a sufficiently high reward.
[0116] In this braking algorithm, the weights of the Q network and the target network (24) are initially randomized to be the same. Initially, the target network (24) is a copy of the Q network (3). The said environment simulator (22) is the external world. The experience replay memory (21) is used to store past experiences. An anti-lock braking system (1) for ensuring a safer and quicker stop for all types of wheeled motorized land and air vehicles includes at least one Q network (3) that is;
[0117] Used as integrated with a brake control unit (4), and
[0118] Trained in a brake module (2),
[0119] That received the sensor data, that is adapted to determine the brake pressure value that must be applied to operate at the global maximum point before 1% lock-up and to transmit a command to a hydraulic control unit (5) to apply the optimum pressure that must be applied to the vehicle wheel, and that is an artificial neural network.
[0120] In the invented braking system (1), the aforementioned Q network (3) compares the vehicle data (sensor data, etc.) it receives with the previously learned data and determines the (optimum) brake pressure value based on the previously used data and transmits a command to the hydraulic control unit (5) to apply this pressure. The braking system (1) of the invention immediately reduces the brake pressure when a full or partial lock-up occurs during braking and applies the optimum braking pressure that can be applied so that not even the slightest partial lock-up of the wheel occurs. That is, it applies a braking pressure that maximizes the deceleration of the vehicle by operating in the full roll zone. With the invented product, shorter braking distances and 100% steering control will be provided to automobile, motorcycle, and airplane users compared to the ABS braking system.
[0121] In one embodiment of the invention, the Q network (3) receives data from sensors on the vehicle that also measure data affecting the brake performance. In one embodiment of the invention, the Q network (3) is in data communication with at least one or more or all of the wheel speed sensor, vehicle acceleration sensor, vehicle tilt sensor, steering angle sensor and / or tire pressure sensor in the vehicle under test, and the data from these sensors are provided as input to the Q network (3). In one embodiment of the invention, the Q network (3) referred to in the braking system is a decision network trained in the aforementioned brake module (2).
[0122] The algorithm starts to work for any wheel when a lock-up occurs at any wheel at any rate, i.e. it starts to interfere with the brake fluid pressure. The pressure does not exceed the pressure exerted by the driver but decreases and increases continuously below this value. A separate algorithm for each wheel runs individually. In each algorithm, a pre-trained Q network is used to decide the action to be applied according to the data coming from the sensors. The first action is to switch to decrease mode for the wheel in question, before the algorithm is activated. The algorithm then receives the state values. The values and the action (decrease action in the first step) are given as input to the decision network. This is the artificial neural network Q network in the training algorithm. The action coming out of the network, i.e. one of the increase, decrease, and hold actions, is given to the hydraulic control unit (5) in order to be applied. Then, in the next time step, the state values and the action selected in the previous step are given to the Q network again, and the network again selects an action. This cycle ends when the algorithm produces increase output 3 times in a row. In other words, from this moment on, the brake module (2) has no influence on the brake pressure for the wheel in question. In run mode, rewards, loss, experience replay memory (21), and target network (24) are not used.
[0123] The Q network (3) shown in Figure 4 is a combination of one LSTM neural network and one fully connected neural network following it. In this embodiment, everything remains the same as in the first embodiment, except that there is an additional LSTM layer. And the past sensor values are not given to the fully connected neural network here. Instead, the hidden state vector produced by the LSTM, which is a statistic of the past, is given. This embodiment, like the first one, is also for discrete action. The LSTM neural network uses two vectors of decimal numbers. The first one is the cell state vector and the second one is the hidden state vector. The numbers in these vectors are statistical numbers, i.e. parameters.
[0124] Where S = { V, Oi,.. Oi ,Hi,.. Hi ,Bi,.. Bi , AC , E , Y ,Li,.. Li ,Si,.. Si ,zi,..Zi} It: It = { AM , {St} }
[0125] Apart from that, the values in this embodiment are the same as in the first embodiment, except that this embodiment also uses the LSTM parameters.
[0126] When the Q network first starts running, the cell state and hidden state values of LSTM are the residual values from the previous braking job. They can be accepted as zero at the beginning of the training.
[0127] Operation mode:
[0128] When the vehicle brakes suddenly, in case of partial or full lock-up, the algorithm starts to intervene in the brake pressure. It values read from the sensors are given to the LSTM layer. The output received from the LSTM layer is a vector identical to the hidden state and is given to the fully connected neural network. At the same time, the hidden state and cell state vectors are given back to the LSTM layer to be used in the next time step. As in the first example, the fully connected neural network generates one of the increase, decrease, or hold options. The result obtained is given to the hydraulic control unit for implementation. In the next time step, the It values are again given to the LSTM layer. This cycle ends when the algorithm produces increase output 3 times in a row. In other words, from this moment on, the Q network (3) has no influence on the brake pressure for the wheel in question. Data from the sensors continue to be read. If a lock-up occurs at any wheel at any rate, the process goes back to the step of giving the It values to the LSTM layer. Otherwise, the algorithm terminates when the braking job is finished. Training mode:
[0129] The fully connected neural network in the Q network is trained as in the first embodiment. The LSTM layer is trained using the “B ackpropagation Through Time” method
[0013] with the values obtained optionally for approximately 10 time steps. For this purpose, in the training of the fully connected neural network, the weight and bias values of the LSTM layer are updated using the gradient values (optionally for 10 time steps) generated at the input of the fully connected neural network by applying backpropagation backwards from the output.
[0130] Operation steps of the braking system (1) for passenger cars in one embodiment of the invention;
[0131] 301. Braking of the vehicle.
[0132] 302. Reading data coming from sensors.
[0133] 303. During braking, the algorithm starts to work for any wheel when a lock-up occurs at any wheel at any rate, i.e. it starts to interfere with the brake fluid pressure.
[0134] 304. As the first action, the Q network (3) sends the decrease command to the hydraulic control unit (5) for the wheel in question, before the algorithm is activated.
[0135] 305. It values obtained in response to the sensor data changing with the application of action are given as input to the Q network (3).
[0136] 306. The action coming out of the network, i.e. one of the increase, decrease, and hold actions, is given again to the hydraulic control unit (5) in order to be applied.
[0137] 307. Steps 304 and 305 continue to be executed in a state-action- state cycle.
[0138] 308. This cycle ends when the algorithm produces increase output 3 times in a row. In other words, from this moment on, the brake module (2) has no influence on the brake pressure for the wheel in question.
[0139] 309. Data from the sensors continue to be read.
[0140] 310. If a lock-up occurs at any wheel at any rate, the process goes back to the step 303. 311. Termination of the algorithm when braking is finished.
[0141] In one embodiment of the invention, a technique using RNN, LSTM, or GRU is used, in which a hidden state and statistics of the past state and action values are kept instead of the method of giving the past state values. When past states are given to the algorithm, or statistics of the past are given to the algorithm using RNN or its derivatives, the algorithm can operate in the full rolling zone. An example is the “Action- specific Deep Recurrent Q network” [1] technique.
[0142] In one embodiment of the invention, the braking system and / or brake module (2) comprises multiple equivalent Q networks (3) with different activation functions and initial weights, different numbers of layers and neurons, both in training mode and in operation mode, to avoid erroneous output. The Q network (3), which will be used as a decision maker, may produce erroneous output from time to time. This is because neural networks, due to their generalization capabilities, miscalculate about one out of every 20 results they produce. Multiple networks are used to solve this problem. The networks are trained by giving each the same state values and rewards. The outputs are subjected to a validation assessment before they are used. Whichever mode (increase, decrease, hold) is the most common output among the 5 result values, that mode is given to the hydraulic control unit (5) for implementation. Artificial neural networks do not produce exact results due to their convergence property. As long as they do not make mistakes, the outputs of the networks will be close to each other. If a control valve is used, i.e., the multiple parallel Q networks of the example used have both discrete and continuous outputs, the incorrect output(s) can be detected automatically with the Unsupervised Learning - Proximity based outlier detection algorithms, to calculate the correct continuous output. The PyOD [9] code library in the Python software language can be used for this. From the moment the training module is activated, all continuous outputs of the 5 networks are added together in the form of an unbounded data stream, and erroneous outputs (outliers) are detected and not included in the calculation. Out of the 5 outputs, the correct ones are averaged and the resulting one number (continuous value) is given to the hydraulic control unit (5), i.e. the control valve, to be applied. This number can be a voltage value depending on the type of the control valve.
[0143] In one embodiment of the invention, a two-way solenoid control valve, at least one for each vehicle wheel, is positioned in the brake fluid line between the hydraulic control unit and the caliper (before the hydraulic pressure sensor). The ABS braking system tries to maintain the wheel partial lock-up rate at the local maximum point of 20%; however, because the hydraulic control unit is not mechanically sensitive enough, the lock-up rate oscillates between 10% and 80%. The same problem occurs with the brake module (2) when it is running at global maximum. To overcome this, a two-way solenoid control valve was conceived. This method is not essential for the operation of the brake module (2). With this valve, the pressure value for each wheel can be adjusted in decrease and increase modes. For example, when the Q network selects the increase mode for the time step in question, it also selects a control valve voltage value for this increase mode. In other words, the pressure does not rise at the same rate in each increase mode. Likewise, the pressure does not decrease at the same rate in each decrease mode. There is no need to apply any voltage value for the hold mode. In addition to the increase, decrease, and hold outputs, the Q network also has 1 control valve voltage output for each wheel. Unlike others, this output is continuous, not discrete. P-DQN (Parametrized Deep Q-Networks Learning) reinforcement learning method can be used for training a Q network with both discrete and continuous outputs.
[0144] If the hydraulic control unit (5) is used, the Q network intervenes in the brake pressure. However, even in case of using a brake -by-wire braking system, the Q network under patent can adjust the brake pressure. The invention can also perform braking jointly with an electric motor that performs regenerative braking.
[0145] In one embodiment of the invention, there is a separate Q network for each wheel or one Q network for centralized control of each wheel. The wheels can be under different conditions during braking. For example, there may be icy ground on the right side of the road and dry asphalt on the left. In such a situation, the right wheels may be more prone to locking while the left wheels may not be locked. In the presented algorithm, each wheel is controlled individually and a separate Q network is used for each wheel; however, of course, the state information of the wheels is also provided to each other’s control system as state information. In another embodiment of the invention, multi-agent reinforced braking algorithms [8] can be used to control all wheels with a common artificial intelligence algorithm.
[0146] Braking system for passenger cars with continuous action output (1), Q network (3), and brake module (2):
[0147] There may be situations where the conventional hydraulic control unit (5), which produces discrete action in the braking system, is not used. The Q network (3) may only need to generate continuous action output, especially in brake-by-wire systems, when regenerative braking is used as the actuator system, or when a servomotor is used to do the braking job. In such a case, an actor network can be used as Q network (3). A method can be used where this network is trained with the [reference 12] Deep Deterministic Policy Gradient reinforcement learning method. Here again, a separate Q network (3) for each wheel is activated in the event of a partial or full lock-up of the wheel in question and starts to precisely adjust the brake pressure.
[0148] In this example, It and Rtvalues are the same as in example 1.
[0149] Operating status:
[0150] If full or partial lock-up occurs during sudden braking, the Q network (3) starts to intervene in the brake pressure. As in the previous example, the It values are given to the final pre-trained actor network, and the network is run. The continuous action value(s) (such as voltage, amperage, etc.) produced by the network are fed to the actuator system for implementation. In this example, we can call these values “At”. At = xtIn this example, At does not contain any discrete value; it is a continuous value or set of values. If the brake pressure continuous action value increases in 3 consecutive time steps, the Q network (3) is deactivated and stops intervening in the brake pressure. If a lock-up occurs afterward, it is activated again. Here, as in the previous example, the actor network must be trained in advance under different braking conditions. For training, the actor network is used together with a critic network to criticize the decisions made by the actor network and 1 target network each to update these two networks. Each one of these 4 networks is fully connected neural network.
[0151] The number of inputs of the actor network and the actor's target networks is equal to the number of It set elements. The number of outputs is equal to the number of At set elements.
[0152] The number of inputs of the critic network and the target networks of the critic network is equal to (It + At). Their output numbers are 1.
[0153] The weights and biases of these neural networks are random at the beginning of training. During training, there is also a replay buffer which holds an arbitrary number of sets consisting of It, Rt, At, and It+i values. In this example, the It and Rt values are obtained in the same way as in the previous example.
[0154] Training of the Network:
[0155] For training, the actor network is first given Itvalues, and the network produces continuous output. These outputs are given to the actuator system for implementation. This process is repeated for an arbitrary number of times in an action- state- action cycle. The collected It, Rt, At, and It+ivalues are saved in the replay buffer.
[0156] For the training of the actor network, Itvalues are given from the replay buffer to the network, and the action output is obtained at the output of the network. With this output, the same It values are given to the critic network this time. In order to maximize the output of the critic network, the deterministic gradient descent method is applied to this network, and the weights and biases of the network are updated.
[0157] For the training of the critic network, It+ivalues are being fed as input to the actor's target network. After running the network, the At+ivalue or values are obtained at its output. This value or values together with It+ivalues are again fed as input to the target network of the critic network. The output of this network is Qt+i. In order to minimize the difference between the output of the critic network, Qt, and (Rt+(discount factor*Qt+i)), the weights and biases of the critic network are updated by stochastic gradient descent method. Target networks are updated with a “soft update”.
[0158] In one embodiment of the invention, a braking system (1) for airplanes with both continuous and discrete action outputs, a Q network (3), and a brake module (2) algorithm:
[0159] In this example, a brake module (2) for the landing gear of an airplane, the braking system algorithm, and state, action and reward information are given. In this algorithm, the state values and actions of the past time step are given to the algorithm as well as the total arbitrary value “n”. Examples of reinforcement learning braking algorithms that could be used in this example are the brake module presented here, the artificial intelligence technique in [5].
[0160] Additionally, this algorithm uses a solenoid-controlled two-way control valve for more precise operation. This valve, one for each caliper, is placed in the brake fluid lines between the hydraulic control unit (5) and the hydraulic pressure sensor. For this, the reinforced braking algorithm must be able to produce both discrete and continuous outputs. In this algorithm, a reinforced brake module with parameterized output as in reference [5] can be used. The symbol for control valves is D.
[0161] When the airplane is landing on the runway, the angle of the ground spoilers (YS) during braking is given to the algorithm as a state value. If there are 2 ground spoilers, YSi and YS2 are given to the algorithm as state values. If control surfaces other than ground spoilers are used for braking when landing on the runway, or if these control surfaces are used instead of ground spoilers, their angle values can also be used as state values. Furthermore, the algorithm is given the motor power value -m being the number of motors- (EDi, ED2 ,..., EDm) and, if applicable, the motor reverse start power value (ETi, ET2 ,..., ETm).
[0162] On most airplanes today, the load on the front wheels is low, so they do not contribute much to braking when landing on the runway. The aft landing gear is equipped with a vertical force sensor, which measures the vertical load on the landing gear. The values from these sensors are given to the algorithm as state values KTi and KT2 if, for example, there are two aft landing gears. In addition, a total of 3 wheel longitudinal force sensors, one for the front landing gear and one for each of the two aft landing gears, are only used for reward calculation during training. This sensor measures the force exerted on the landing gear in the direction of the airplane. Both force sensors can be placed on the struts of the landing gears. If all 3 landing gears can be steered to ensure safe landing of the airplane in a crosswind, all 3 “Y” angle sensors are used as state values.
[0163] One ground speed “V” is also given to the algorithm as state information.
[0164] Likewise, when there are 2 aft landing gears and each of these gears has 2 brake assemblies; the number of state information is 4 for each of O, H, B, L, S, D.
[0165] Where S={V, Oi, Hi, Bi, AC, Yi, Y2, Y3, E, Li, Si, EDi, ED2,.., EDm, ETi, ET2,.., ETm, Di, KTi, KT2, YSI, YS2}; all states of the algorithm, including past states and actions are as follows:
[0166] It = { At-n-i , {St-n} , • • •, AM , {St} }. With “n” being an arbitrary number, values going n time steps back are given to the algorithm, thus enabling the algorithm to make comparisons against the past. The past state values that are not yet possible for the It set elements are the values left over from the previous braking job when the Q network (3) first started running. As the time step occurs and values are available, they are replaced by their measured and calculated values respectively. The front landing gear of the airplane does not contribute much to braking as it is not heavily loaded and thus is not taken into account during braking. The aft landing gear may have one or more than one wheel sets comprising more than one wheel together. Each of these sets is connected to a single brake fluid line (if there is 1 independent caliper in the set), and each set is controlled by one algorithm. An algorithm generates one discrete value and one continuous value based on this discrete value and these two values are fed to the hydraulic control unit for implementation. The number of different algorithms to be used depends on the number of wheel sets. If desired, the entire system can be controlled through multiagent reinforced braking algorithms that can be used to control all wheels with a common artificial intelligence algorithm. Actions (At), on the other hand;
[0167] For a single wheel set, 1 discrete action (increase, decrease, or hold) is selected, and a continuous control valve setting value is selected accordingly. In this way, the hydraulic control unit (5) determines the direction of flow of the hydraulics and thus the rate of increase and decrease of the pressure. In other words, how fast the pressure will increase, if increase is selected, and how fast the pressure will decrease, if decrease is selected, is adjusted with the control valve. If hold is selected, the control valve setting value from the previous time step continues to be applied. At= {kt,Xkt}. The discrete value is ktand the corresponding continuous value is Xkt. Xkt value is selected from the xk set.
[0168] The rewards to be used in the algorithm are:
[0169] » 3 wheel longitudinal force sensors. Fi, F2, F3
[0170] » Wheel lock-up rate. Ok (all wheels in a wheel set have the same lock-up rate.)
[0171] » 1 vehicle acceleration sensor AC
[0172] Rewards: k: For the wheel in question controlled by the module
[0173] Ri = c*(B-z) / B for Ok < 3%
[0174] R2 = - Ok*m for 3% < Ok < 100%
[0175] R3 = Fk * x for 0 < Ok < 00
[0176] R4 = AC * g for 0 < O < co
[0177] Rt = Ri +R2 +R3 + R4
[0178] In this method, two fully connected neural networks, a policy network and a value network, are used together as decision network (3). The policy network generates a set of continuous action values Xk, i.e. one throttle valve current (or voltage, if used) value each for the increase, decrease and hold options. The network generates these values directly, because the network is deterministic. However, the network is trained with the stochastic gradient descent method. The network takes It values as input and therefore the number of inputs is equal to the number of elements of the It set. The output is 3 continuous numbers, i.e. the values of the Xk set.
[0179] The value network determines the discrete value “kt” and the corresponding Xkt. For this, the network generates 1 Q value for each of the options increase, decrease, and hold. The discrete action with the highest Q value (kt) and its subordinate, i.e. the continuous action value (xkt) generated by the policy network for this discrete action, are the pair of actions to be applied together. The network is stochastic, and it is trained using the stochastic gradient descent method. The network takes as input the It values and the xk values obtained from the policy network. Therefore, the number of inputs is equal to the number of elements of the It set plus the number of elements of Xk set. The output number is 3.
[0180] Operating status:
[0181] If full or partial lock-up occurs during sudden braking, the Q network (3) starts to intervene in the brake pressure. The It values are given as input to the final pretrained policy network. The network generates the continuous action values (xk). These values are also fed to the pre-trained final value network along with the It values. At the output of the network, a Q value each is obtained for the increase, decrease, and hold options. Of the options, the one with the highest Q is the discrete action option and the pair of Xkt values subject to this option is the action to be applied. In this way, the cycle is repeated over and over again. If the increase command is given in 3 consecutive time steps, the decision network (3) is deactivated and stops intervening in the brake pressure. If a lock-up occurs afterward, it is activated again.
[0182] During the training process, Alpha (a) and beta (p) are used as learning coefficients for the value network and the policy network, respectively. Gamma (y) is being used as a discount factor to calculate the target value. Alpha, beta, and gamma values are arbitrary parameters.
[0183] Additionally, experience replay is the memory that holds the as many experiences of the {It, At, Rt, It+i } sets as the arbitrary number “D” during the training process. The replay buffer (or mini-batch) is a small memory to which as many random sample sets as the arbitrary number “B” are retrieved from the experience replay. Training of the Network:
[0184] In order to generate the sets to generate the experience replay, first the It values are given to the policy network, and the network is run. The Xk values coming out of the network and again the It values are given to the value network this time and the network is run. Whichever option has a higher Q value at the output of the network is discrete action and its subordinate Xkt is continuous action. According to the epsilon-greedy method, either this determined pair of actions or a random pair of actions is chosen as the action to be applied to the environment. The resulting { It, At, Rt, It+i } set is saved. As many state-action-state-action cycles as the number "D" are executed, and the sets formed are saved in the experience replay.
[0185] Then, as many random sets as the number “B” are imported into the replay memory from the experience replay memory. The (lb, Ab, Rb, and Ib+i) set is retrieved from the replay memory. Ib is fed into the value network. The Q value generated by the network for Ab is the “predicted Q”. In order to obtain the target value (ytargetb), Ib+i is given to the policy network. The Xb values obtained from the network and again the Ib+i values are given to the value network this time. The “Maximum Q value” is obtained at the output of the network.
[0186] The “ytargetb” value is obtained by multiplying the “maximum Q value” by gamma and summing the result with Rb. This process is repeated for each set in the replay memory. The loss value of the value network for all sets in the replay memory is calculated using the least squares loss function with the resulting “ytargetb” and “predicted Q” values. Likewise, for the time step b, the Q(Ib, xk(Ib)) values of the increase, decrease, and hold options of the value network and the loss value of the policy network are calculated.
[0187] Both networks are updated using the stochastic gradient descent method with the loss values and alpha and beta learning rates of both networks. Then, random sets are imported into the replay buffer from the experience replay again, and the training cycle is repeated as many times as desired to complete the training of the policy and value networks.
[0188] Using artificial intelligence in braking system and brake module of the invention:
[0189] It is not possible to apply conventional mathematical optimization methods to search for the global maximum during braking. Therefore, we use a data- based method.
[0190] It can operate in the pre-partial lock-up zone thanks to the algorithm, that is, by comparing with the past. Otherwise, when operating in this zone, it is not possible to tell how much further the brake pressure can be increased before partial lock-up starts; there is no sensor to measure this.
[0191] In one embodiment of the invention, the braking system comprises multiple equivalent Q networks (3) with different activation functions and initial weights, different numbers of layers and neurons, to avoid an erroneous output.
[0192] When past states are given to the algorithm, or statistics of the past are given to the algorithm using RNN or its derivatives, the Q network gains POMDP property. This allows the algorithm to work in the full roll zone. It is difficult to work in this zone as there is no lock-up rate information for this zone.
[0193] References:
[0194] 1. [Pengfei Zhu, Xin Li, Pascal Poupart, Guanghui Miao] On Improving Deep Reinforcement Learning for POMDPs
[0195] 2. [Matthew Hausknecht and Peter Stone] Deep Recurrent Q-Learning for Partially Observable MDPs
[0196] 3. [Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusul, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, loannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg & Demis Hassabis] Human-level control through deep reinforcement learning
[0197] 4. [Zhou Fan, Rui Su, Weinan Zhang and Yong Yu] Hybrid Actor-Critic Reinforcement Learning in Parameterized Action Space
[0198] 5. [Jiechao Xiong, Qing Wang, Zhuoran Yang, Peng Sun, Lei Han, Yang Zheng, Haobo Fu, Tong Zhang, Ji Liu and Han Liu] Parametrized Deep Q- Networks Learning: Reinforcement Learning with Discrete-Continuous Hybrid Action Space
[0199] 6. [Hochreiter and Schmidhuber, 1997] Sepp Hochreiter and J’ urgen Schmidhuber. Long short-term memory. Neural Computation, 9(8): 1735— 1780, 1997.
[0200] 7. [Paul J Werbos.] B ackpropagation through time: what it does and how to do it. Proceedings of the IEEE, 78(10): 1550-1560, 1990. 8. [Haotian Fu, Hongyao Tang, Jianye Hao, Zihan Lei, Yingfeng Chen, Changjie Fan] Deep Multi-Agent Reinforcement Learning with Discrete- Continuous Hybrid Action Spaces
[0201] 9. Python Outlier Detection (PyOD) https: / / ithub.com / yzhao062 / pyod
[0202] 10. Ruder, S. (2016). An overview of gradient descent optimization algorithms. ArXiv Preprint ArXiv: 1609.04747.
[0203] I L S. Linnainmaa. The representation of the cumulative rounding error of an algorithm as a Taylor expansion of the local rounding errors. Master's Thesis (in Finnish), Univ. Helsinki, 1970.
[0204] 12. [Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, Daan Wierstra] Continuous control with deep reinforcement learning
[0205] 13. [Paul j. Werbos] B ackpropagation through time: what it does and how to do it.
[0206] Dictionary:
[0207] HCU: Hydraulic Control Unit, Hydraulic Modulator
[0208] Current state: Current state (St)
[0209] Next state: Next state (St+i)
[0210] Reward: Reward (R)
[0211] Action: Action (A)
[0212] Discount Factor: Discount Factor
[0213] Discrete: Discrete
[0214] Continuous: Continuous Increase: Increase pressure
[0215] Decrease: Decrease pressure
[0216] Hold: Hold the pressure constant
[0217] Experience replay: Experience replay tuple y (Discount Factor, a value between 0 and 1): Discount factor a (braking coefficient): Learning rate, a.k.a “step size”
[0218] Control valve: (Flow control valve / Speed control valve). Control valve
[0219] Episode: Episode
[0220] Artificial Neural Network: Artificial Neural Network, Fully Connected Layer
[0221] (FCL)
[0222] Ground Speed: Ground Speed
[0223] Ground Spoiler: Ground Spoiler
[0224] Control Surface: Control Surface (flap, spoiler, aileron, horizontal stabilizer, etc.)
[0225] Strut: Landing Gear Strut, main element of landing gear construction
[0226] Reinforcement Learning: Reinforcement Learning
[0227] ADRQN: Action Deep Recurrent Q Network
[0228] Exclusively: Exclusively (The output with the highest value, i.e. one of the increase, decrease, and hold outputs, is selected as action exclusively)
[0229] Unsupervised Learning: Unsupervised Learning
[0230] Proximity Based Outlier Detection: Proximity Based Outlier Detection
[0231] Unbounded Data Stream: Unbounded Data Stream
[0232] Wheel Longitudinal Force Sensor: Wheel Longitudinal Force Sensor
[0233] Cell State: Long-term memory
[0234] Hidden State: Short-term memory
[0235] LSTM: Long-Short Term Memory Neural Network
Claims
CLAIMS1. A braking system (1) that is used in vehicle braking systems to ensure a safer and quicker stop / braking for all types of wheeled motorized land and air vehicles, and characterized by comprising:- a brake control unit (4) integrated into the vehicle,- a brake module (2) used for training in vehicle braking system tests, and- at least one decision network (3) adapted to receive sensor data based on changing conditions affecting the wheel, tracking the change in the brake pressure value at the global maximum point before partial locking of the wheel, and learning to apply the optimum brake pressure under changing conditions.
2. The braking system (1) as in claim 1, characterized with a brake module (2) adapted to be trained in machine learning tests and to track the change of the global maximum point before partial locking of the wheel due to changing conditions.
3. The braking system (1) as in claim 1 or 2, characterized by at least one decision network (3) comprising an artificial neural network adapted to learn how to generate an output for a brake fluid pressure value to be transmitted to the brake based on data from the sensors.
4. The braking system (1) as in any one of the abovementioned claims, characterized by a brake module (2) that performs braking with reinforcement learning algorithms operating in a state-action- state cycle.
5. The braking system (1) as in any one of the abovementioned claims, characterized by the decision network (3), which is an artificial neural network that includes gradient descent and thus has an optimization property.
6. The braking system (1) as in any one of the abovementioned claims, characterized by a decision network (3) (Q network), which is an artificial neural network whose weights in vehicle tests are initially random and which updates itsweights by learning the brake pressure required according to changing vehicle driving conditions and adapted to produce a closer comparison between the desired output and the actual output in the next cycle.
7. The braking system (1) as in any one of the abovementioned claims, characterized by a reward unit that takes the brake pressure value generated by the decision network (3) as data output and is adapted to compare it with the expected output.
8. The braking system (1) as in any one of the abovementioned claims, which is in communication with at least one or more or all of the wheel speed sensor, vehicle acceleration sensor, vehicle tilt sensor, steering angle sensor and / or tire pressure sensor in the vehicle under test.
9. The braking system (1) as in any one of the abovementioned claims, which takes the wheel longitudinal force sensor data as input.
10. The braking system (1) as in any one of the abovementioned claims, which comprises an experience replay memory (21), training data (mini batch) (23), a loss, a target network (24), and a Q network (3) also used as decision network.
11. The braking system (1) that comprises multiple equivalent Q networks (3) with different activation functions and initial weights, different numbers of layers and neurons, to avoid an erroneous output.
12. The braking system (1) as in any one of the abovementioned claims, characterized by the decision network acquiring the POMDP property when past states are given to the algorithm or when the statistics of the past are given to the algorithm using RNN or its derivatives.
13. The braking system (1) as in any one of the abovementioned claims, which includes a separate decision network for each wheel or one decision network for centralized control of each wheel.
14. A braking method (200) and characterized by comprising the following steps:- Starting the process by pressing the brake (201),- Transmission of data from the sensors in the tested vehicle to the decision network, which is an artificial neural network (202),- Measuring the brake fluid pressure delivered to the brake when approximately 3% wheel lock-up occurs (203),- Evaluation of the pressure value state realized with sensor data by a reward unit (204),- Updating the weights of the decision network, which is an artificial neural network, as a result of the evaluation (205),- Repeating the process steps 201, 202, 203, 204, 205 continuously until the decision network becomes ready for use (206),- That the decision network is in the learned state and ready to be integrated as decision network in the vehicle control unit (207).
15. The braking system (1) as in claims 1 to 11, comprising a decision network (3) trained according to the braking method (200) of Claim 14.
Citation Information
Patent Citations
Automobile anti-lock braking control method and device realized based on neural network algorithm, vehicle and storage medium
CN112918448A
Vehicle speed control method and device based on reinforcement learning, equipment and medium
CN116552474A
System for preserving value of virtual Items and virtual assets
KR1020230160480A