Multi-data fusion and reinforcement learning hybrid vehicle fleet safety and energy saving method and system
By using multi-data fusion and reinforcement learning methods, high-precision slope perception and energy management of hybrid commercial vehicle fleets under complex road conditions were achieved, solving the problems of low slope perception accuracy and unstable energy distribution in traditional technologies, and improving the stability and energy efficiency of the fleet.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHANGAN UNIV
- Filing Date
- 2026-05-26
- Publication Date
- 2026-07-10
AI Technical Summary
Traditional hybrid commercial vehicles struggle to achieve high-precision slope perception, unstable energy distribution, and coordinated fleet control under complex road conditions, making it difficult to track speed and maintain a safe distance. This is especially true in mountainous areas or urban elevated highways where slopes change frequently, where existing technologies cannot meet the requirements for real-time performance and accuracy.
Employing multi-data fusion and reinforcement learning methods, slope estimation is achieved through multi-source data synchronization, spatial registration, and adaptive Kalman filters. Combined with an improved vehicle longitudinal dynamics model and a dual-network deep reinforcement learning model, coordinated control of the engine and motor is realized, and energy management is achieved by switching between multiple drive modes.
It improves the accuracy of gradient estimation and fleet stability, reduces fuel consumption by 8.5%, reduces vehicle distance fluctuation by 35%, maintains fleet stability and safety in complex road conditions, and improves energy efficiency, with a braking energy recovery contribution rate of 15%.
Smart Images

Figure CN122354474A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of commercial vehicle control, specifically to a method and system for the safety and energy saving of hybrid vehicle fleets using multi-data fusion and reinforcement learning. Background Technology
[0002] In recent years, with the implementation of energy conservation and emission reduction policies and the rapid development of intelligent connected vehicle technology, hybrid commercial vehicles have gradually become the mainstream in the transportation industry. Hybrid vehicles can significantly reduce fuel consumption and emissions through the coordinated work of the engine and motor. However, under complex road conditions, especially in mountainous areas or urban elevated roads with frequent changes in slope, vehicle load fluctuations and power demands change drastically, making it difficult for traditional safety and energy-saving strategies to meet the requirements in terms of real-time performance and accuracy.
[0003] Existing research often employs strategies to minimize equivalent fuel consumption or model predictive control based on slope feedforward compensation to achieve energy allocation, optimizing engine and motor power through static models. In terms of vehicle longitudinal control, traditional car-following controllers such as PID (Proportional-Integral-Derivative Controller) or fixed-gain MPC (Model Predictive Control) are widely used to achieve speed tracking and maintaining a safe distance. In recent years, deep reinforcement learning (DRL) has provided new ideas for vehicle safety and energy conservation. Through interaction between the agent and the environment, optimal decision-making can be achieved without relying on a model.
[0004] However, the above solutions still have the following drawbacks: 1. The perception of road slope, vehicle status and environmental disturbances relies on a single sensor, such as GPS (Global Positioning System) or accelerometer, resulting in low perception accuracy and strong lag; 2. Traditional car-following controllers are difficult to achieve accurate speed tracking and safe distance maintenance when facing complex slopes and wind resistance disturbances. Moreover, as the fleet size increases, single-vehicle control is difficult to achieve optimal global energy consumption, resulting in a decrease in overall energy allocation efficiency; 3. Existing deep reinforcement learning is mostly focused on single-vehicle control and lacks a unified framework for multi-vehicle collaboration and multi-source perception fusion. At the same time, slope and obstacle perception are still limited to a single sensor and cannot cope with complex working conditions such as rain, fog and changes in lighting. Summary of the Invention
[0005] To address the problems mentioned in the prior art, this invention proposes a method and system for safe and energy-saving hybrid vehicle fleets based on multi-data fusion and reinforcement learning. This method achieves high-precision perception of gradient and environment, adaptive energy allocation, and coordinated safe control of the vehicle fleet, overcoming the problems of low perception accuracy, unstable energy allocation, and insufficient vehicle control coordination in the prior art.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: This invention proposes a multi-data fusion and reinforcement learning-based method for the safety and energy conservation of hybrid vehicle fleets, comprising the following steps: Collect multi-source data from each vehicle in the fleet; Multi-source data is synchronized in time to obtain synchronized data; the synchronized data is then spatially registered to obtain a fused dataset. The fused dataset is input into an adaptive multi-source extended Kalman filter to estimate the road slope and generate the slope estimation results. Based on the slope estimation results and the pre-set desired speed trajectory, an improved vehicle longitudinal dynamics model is constructed. Based on the improved vehicle longitudinal dynamics model, the equivalent energy required for speed tracking, the equivalent energy required for spacing tracking, and the energy lost due to control acceleration are used as optimization objectives to solve the problem and obtain the desired acceleration and target distance control signals. The desired acceleration is converted into the vehicle's required power, which is then input into a pre-built dual-network deep reinforcement learning model for decision processing, and outputs power control commands. Based on the power control command and the target distance control signal, the engine, motor and braking system of each vehicle in the fleet are controlled to work together, and braking energy is recovered through multiple mode switching to achieve energy saving for each vehicle in the fleet.
[0007] As a further improvement of the present invention, the step of time synchronization of multi-source data to obtain synchronized data includes: The HSC-DDC time synchronization unit is used to synchronize the time of multi-source data. The specific process is as follows: It provides a nanosecond-level reference clock, and based on the nanosecond-level reference clock, it combines the vehicle's inertial measurement unit to correct multi-source data to obtain clock-synchronized data; Dynamically predict the data after clock synchronization to generate delay compensation. Based on the delay compensation amount, the clock-synchronized data is time-aligned to obtain preliminary synchronized data; the predicted delay value is calculated based on the delay compensation amount. When the predicted delay value exceeds the dynamic threshold, the neighboring frame interpolation-confidence scoring mechanism is triggered based on the preliminary synchronization data to process the predicted delay value and obtain a confidence score. The confidence score is compared with a preset threshold, and data with a confidence score lower than the preset threshold is removed to obtain the synchronized data.
[0008] As a further improvement of the present invention, spatial registration is performed on the synchronized data to obtain a fused dataset, including: Semantic augmentation-based fast point cloud-visual registration is used to perform time synchronization on the synchronized data. The specific process is as follows: The image and point cloud data in the synchronized data are processed as follows: The image is input into a lightweight convolutional neural network for semantic segmentation processing to obtain the semantic segmentation result; Based on the semantic segmentation results, the core road regions of the point cloud data are filtered to obtain the filtered point cloud dataset. Dynamic distortion correction is performed on the filtered point cloud dataset to obtain the corrected point cloud data; The corrected point cloud data is coarsely registered with the image to obtain the coarse registration result; The coarse registration result is refined to obtain the accurate registration result; The precise matching result is fused with the semantic segmentation result to generate a fused dataset containing road surface geometric features, slope information, and semantic labels.
[0009] As a further improvement of this invention, the fused dataset is input into an adaptive multi-source extended Kalman filter for road slope estimation, generating slope estimation results, including: First, the state variables in the adaptive multi-source extended Kalman filter are extended to obtain an extended state vector that includes vehicle longitudinal velocity, acceleration, road slope angle, radar ranging value, surface normal angle, slope change rate, and radar ranging error. The observation residual is calculated based on the extended state vector. The process noise covariance matrix and measurement noise covariance matrix in the adaptive multi-source extended Kalman filter are dynamically weighted based on the observation residual to obtain the updated process noise covariance matrix and the updated measurement noise covariance matrix. The slope estimation result is generated by iteratively updating the updated process noise covariance matrix and the updated measurement noise covariance matrix.
[0010] As a further improvement to the present invention, based on the slope estimation results and the pre-set desired speed trajectory, an improved vehicle longitudinal dynamics model is constructed, including: The slope disturbance acceleration is calculated based on the slope estimation results. By incorporating the slope disturbance acceleration into the state equation of the vehicle's longitudinal dynamics model, the corrected state equation is obtained. The modified state equations are calibrated based on the pre-set desired velocity trajectory to obtain an improved vehicle longitudinal dynamics model.
[0011] As a further improvement of this invention, based on the improved vehicle longitudinal dynamics model, the equivalent energy required for speed tracking, the equivalent energy required for distance tracking, and the energy lost due to control acceleration are used as optimization objectives to solve for the desired acceleration and target vehicle distance control signal, including: A multi-objective optimization problem is constructed by defining a prediction time domain and a control time domain. Using the equivalent energy required for velocity tracking, the equivalent energy required for spacing tracking, and the energy lost due to control acceleration as cost terms for the objective problem, a cost function is constructed, the expression of which is as follows:
[0012] In the formula: The cost function is expressed in joules. This represents the equivalent energy required for velocity tracking, measured in joules. Weighting coefficients representing the equivalent energy required for velocity tracking; For vehicle mass; N represents the length of the prediction time domain; Indicates the index of the current time step; express The actual speed at any given moment; for Reference speed at any given moment; This represents the equivalent energy required for spacing tracking, measured in joules. Weighting coefficients for the equivalent energy required for spacing tracking; This is the equivalent stiffness coefficient; for Actual distance between vehicles at any given time; To maintain a safe following distance; This represents the energy lost due to controlled acceleration, measured in joules. Weighting coefficients to control the energy loss due to acceleration; It is a time constant; To control acceleration; The weight coefficients in the cost function are dynamically adjusted based on the slope estimation results to obtain the adjusted cost function. By combining the improved vehicle longitudinal dynamics model with the adjusted cost function, a multi-performance objective optimization problem is obtained. The multi-performance objective optimization problem is solved to obtain the desired acceleration and target vehicle distance control signals.
[0013] As a further improvement of the present invention, the desired acceleration is converted into the vehicle's required power, and the vehicle's required power is input into a pre-constructed dual-network deep reinforcement learning model for decision processing, outputting power control commands, including: Construct the state space, action space, and reward function of a dual-network deep reinforcement learning model; The state space includes vehicle power demand, battery state of charge, and engine efficiency. The action space is the engine power distribution ratio; The expression for the reward function is as follows:
[0014] In the formula: Represents the reward function; Sampling period The energy consumed by fuel internally, measured in joules; Sampling period The electrical energy consumed internally, measured in joules; Sampling period The energy lost due to internal power fluctuations, measured in joules; This is the weighting coefficient for fuel energy consumption; This is a weighting coefficient for electricity consumption; This is the weighting coefficient for energy loss due to power fluctuations; For time integration variables; Fuel consumption rate; This refers to the calorific value of fuel oil. This is the battery open-circuit voltage; This refers to the battery current. This represents the change in power. The dual-network deep reinforcement learning model is trained to obtain a well-trained dual-network deep reinforcement learning model. The vehicle's required power, battery state of charge, and engine efficiency are input into a trained dual-network deep reinforcement learning model, which then outputs power control commands.
[0015] As a further improvement of this invention, based on power control commands and target distance control signals, the engines, motors, and braking systems of each vehicle in the convoy are controlled to work in coordination, and braking energy is recovered through multiple mode switching to achieve energy saving for each vehicle in the convoy, including: When the vehicle's power demand is not greater than a preset percentage of the motor's peak power and the battery's state of charge is not less than a threshold, it enters pure electric drive mode, shuts off the engine, and is driven solely by the motor. When the vehicle's power demand is not less than a preset percentage of the motor's peak power, it enters the range-extended drive mode, which allows the engine to operate in the optimal efficiency range to generate electricity. The electricity is first supplied to the motor drive, and the remaining electricity is used to charge the battery. When the vehicle's power demand is no greater than a preset percentage of the motor's peak power and the battery's state of charge is no greater than a threshold, the range extender charging mode is entered, allowing the engine to operate and generate electricity within its optimal efficiency range. The electricity is prioritized to meet the vehicle's driving needs, and the remaining electricity is used to charge the battery.
[0016] This invention proposes a multi-data fusion and reinforcement learning hybrid vehicle fleet safety and energy-saving system, comprising: The data acquisition module is used to collect multi-source data from each vehicle in the fleet; The decision module is used to synchronize multi-source data in time to obtain synchronized data; spatially register the synchronized data to obtain a fused dataset; input the fused dataset into an adaptive multi-source extended Kalman filter for road slope estimation to generate slope estimation results; based on the slope estimation results and a pre-set desired speed trajectory, construct an improved vehicle longitudinal dynamics model; based on the improved vehicle longitudinal dynamics model, solve for the equivalent energy required for speed tracking, the equivalent energy required for distance tracking, and the energy lost due to control acceleration as optimization objectives to obtain the desired acceleration and target distance control signals; convert the desired acceleration into vehicle power demand, input the vehicle power demand into a pre-constructed dual-network deep reinforcement learning model for decision processing, and output power control commands; The execution module controls the engines, motors, and braking systems of each vehicle in the convoy to work together according to the power control commands and target distance control signals. It recovers braking energy through multiple mode switching to achieve energy saving for each vehicle in the convoy.
[0017] This invention proposes a multi-data fusion and reinforcement learning hybrid vehicle safety and energy-saving device, comprising a processor and a memory, wherein the processor executes a computer program stored in the memory to implement the multi-data fusion and reinforcement learning hybrid vehicle safety and energy-saving method as described above.
[0018] Compared with the prior art, the present invention achieves the following technical effects: This invention addresses the technical problems of low sensing accuracy and strong lag in traditional single-sensor systems by collecting multi-source data. Through sequential time synchronization and spatial registration of the multi-source data, alignment in both time and space is achieved, providing an accurate and reliable foundation for subsequent slope estimation. Furthermore, employing an adaptive multi-source extended Kalman filter for road slope estimation reduces the mean square error of the estimation results from 0.3858 to 0.1247, a reduction of over 67%, providing high-precision slope estimation results for subsequent processing. By constructing an improved vehicle longitudinal dynamics model, the slope estimation results can be used as the control optimization objective, achieving predictive compensation for slope disturbances. In complex slope scenarios with ±5°, vehicle speed tracking error is reduced by 40%, and platoon spacing fluctuation is reduced by 35%, effectively solving the problems of spacing fluctuation and speed instability caused by slope changes in traditional control. Even in complex road conditions with alternating continuous slopes and obstacles, the stability and safety of the platoon can still be maintained.
[0019] The dual-network deep reinforcement learning model of this invention uses vehicle power demand, battery state of charge, and engine efficiency as the state space, and engine power allocation ratio as the action space, to achieve adaptive optimization of engine and motor power allocation. Compared with existing methods, fuel consumption is reduced by approximately 8.5%, and battery SOC fluctuations are strictly controlled within ±0.5%, effectively extending the lifespan of the power battery. Simultaneously, through switching between multiple modes such as pure electric drive, range-extended drive, and range-extended charging, and through regenerative braking, the fleet can achieve a regenerative braking contribution rate of over 15% in urban congestion conditions, thereby improving overall energy efficiency. Attached Figure Description
[0020] Figure 1 This is a schematic diagram of the overall process of the present invention; Figure 2 This is a schematic diagram of the HSC-DDC time synchronization unit processing of the present invention; Figure 3 This is a schematic diagram of the semantic enhancement fast point cloud-visual matching process of the present invention; Figure 4 This is a schematic diagram of the multi-objective optimization of the improved vehicle longitudinal dynamics model of the present invention. Detailed Implementation
[0021] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.
[0022] like Figure 1As shown, this embodiment proposes a multi-data fusion and reinforcement learning-based method for the safety and energy conservation of hybrid vehicle fleets, including the following steps: Collect multi-source data from each vehicle in the fleet; Multi-source data is synchronized in time to obtain synchronized data; the synchronized data is then spatially registered to obtain a fused dataset. The fused dataset is input into an adaptive multi-source extended Kalman filter to estimate the road slope and generate the slope estimation results. Based on the slope estimation results and the pre-set desired speed trajectory, an improved vehicle longitudinal dynamics model is constructed. Based on the improved vehicle longitudinal dynamics model, the equivalent energy required for speed tracking, the equivalent energy required for spacing tracking, and the energy lost due to control acceleration are used as optimization objectives to solve the problem and obtain the desired acceleration and target distance control signals. The desired acceleration is converted into the vehicle's required power, which is then input into a pre-built dual-network deep reinforcement learning model for decision processing, and outputs power control commands. Based on the power control command and the target distance control signal, the engine, motor and braking system of each vehicle in the fleet are controlled to work together, and braking energy is recovered through multiple mode switching to achieve energy saving for each vehicle in the fleet.
[0023] The present invention will be further explained below with reference to the accompanying drawings and specific embodiments: Step 1: This embodiment requires acquiring multi-source data from each vehicle in the fleet. Specifically, multi-source data is collected by various sensors installed at the front of each vehicle, including millimeter-wave radar, lidar, camera, GPS, and IMU (Inertial Measurement Unit).
[0024] Millimeter-wave radar acquires distance information of obstacles ahead at a fixed sampling frequency, and has high ranging accuracy and is less affected by rain and fog. LiDAR is used to acquire point cloud data that can describe the geometry of the road surface. Cameras are responsible for capturing images of the surrounding road, and the camera frame rate and exposure time can be automatically adjusted according to the current vehicle speed to ensure clear images at different driving speeds. GPS is used to provide vehicle location information. IMU is used to measure vehicle acceleration information, and each vehicle has its own IMU. Through the coordinated work of the above-mentioned multiple sensors, various information about the vehicle's own status and the surrounding environment can be comprehensively captured.
[0025] Step 2: Since the sensors in Step 1 differ in hardware structure, sampling frequency, and data transmission delay, directly using unsynchronized data for fusion will result in time misalignment, which will affect the accuracy of subsequent slope estimation results. Therefore, this embodiment introduces an HSC-DDC (Heterogeneous Clock Synchronization-Dynamic Delay Compensation) time synchronization unit to perform time synchronization processing on multi-source data.
[0026] like Figure 2 As shown, firstly, on the hardware side, a GPS dual-mode timing module provides a nanosecond-level reference clock, and combined with an IMU to correct clock drift caused by sensor bumps in real time, thereby achieving clock alignment for multiple sensors. Secondly, on the software side, this embodiment uses an LSTM (Long Short-Term Memory) delay prediction model to dynamically predict the clock-synchronized data and generate a delay compensation amount; based on the delay compensation amount, timestamp interpolation calibration is performed on the clock-synchronized data to achieve timing alignment and obtain preliminary synchronized data.
[0027] Subsequently, the actual predicted delay value is calculated based on the delay compensation amount. When the predicted delay value exceeds the dynamic threshold (the dynamic threshold can be adjusted according to road conditions, with a maximum of ≤1μs), the neighboring frame interpolation-confidence scoring mechanism is triggered based on the preliminary synchronization data. The confidence of this part of the data is evaluated and a confidence score is obtained. The confidence score is compared with the preset threshold (0.6 points), and low-confidence data with scores lower than the preset threshold are removed. Finally, the synchronized data is obtained, ensuring that the frame-level alignment accuracy of the synchronized data is ≤200ns.
[0028] After time synchronization is completed, the synchronized data under different coordinates need to be unified into the same spatial coordinate system. Therefore, this embodiment uses Semantic Enhanced-FastPoint Cloud-Visual Registration (SE-FPV) to achieve spatial registration.
[0029] like Figure 3 As shown, the images captured by the camera are first input into a lightweight convolutional neural network for semantic segmentation processing, resulting in semantic segmentation results containing semantic labels such as roads, vehicles, and obstacles. Then, based on the semantic segmentation results, the point cloud data is filtered to remove the core road regions, eliminating redundant and noisy point clouds, effectively reducing the data volume, and obtaining the filtered point cloud dataset.
[0030] Dynamic distortion correction is performed on the filtered point cloud dataset based on IMU to compensate for point cloud distortion caused by vehicle movement, resulting in corrected point cloud data. The corrected point cloud data and the image are then coarsely registered using NDT (Normal Distribution Transform) to roughly align the corrected point cloud data and the image globally. Subsequently, the ICP (Iterative Closest Point) algorithm is used for fine registration to achieve pixel-level high-precision matching, resulting in accurate matching results. Finally, the accurate matching results are fused with the semantic segmentation results to generate a fused dataset containing road surface geometric features, slope information, and semantic labels.
[0031] Step 3: In this embodiment, the fused dataset obtained in Step 2 is input into an adaptive multi-source extended Kalman filter for road slope estimation. Traditional Kalman filters often limit the state vector to a finite dimension, making it difficult to handle complex conditions such as rapid slope changes. Therefore, this embodiment makes the following improvements to the Kalman filter: First, the state vector is extended, defined as an extended state vector that includes the vehicle's longitudinal velocity, acceleration, road slope angle, radar ranging value, surface normal angle, slope change rate, and radar ranging error. The extended state vector can be expressed as:
[0032] In the formula: X is the extended state vector; The longitudinal speed of the vehicle; For acceleration; The road slope angle; This is the radar ranging value; The surface normal angle; This represents the rate of change of slope. denoted as radar ranging error; T represents transpose.
[0033] Extended state vectors comprehensively capture dynamic parameters related to road slope from two dimensions: vehicle operation and road features. This enables better tracking of real-time changes in road slope and improves the ability to estimate dynamic slope.
[0034] Based on the extended state vector, the state transition equation and observation equation are constructed, and their expressions are as follows:
[0035]
[0036] In the formula: For the first The extended state vector at time step; This is the state transition function; For the first The control input vector at each time step; For the first Time-matter noise; For the first The extended state vector at time -1; For the first The observation vector at time; These represent the observed values of vehicle speed, altitude obtained from GPS, and vehicle acceleration, respectively; T represents the transpose. , , These are the vehicle's longitudinal speed, altitude obtained from GPS, and acceleration, respectively. For observation functions; For measuring noise.
[0037] To improve the adaptability of the adaptive multi-source extended Kalman filter under complex operating conditions, this embodiment introduces an adaptive weight adjustment mechanism. Specifically, the process noise covariance matrix and the measurement noise covariance matrix are dynamically adjusted based on the observation residuals, thereby enabling the adaptive multi-source extended Kalman filter to maintain optimal estimation performance under different operating conditions.
[0038] The formula for calculating the observation residual is as follows:
[0039] In the formula: For the first Observe the residuals at all times; For the first Previous estimates of the state variables at time step.
[0040] The update formulas for the process noise covariance matrix and the measurement noise covariance matrix are as follows:
[0041]
[0042] In the formula: For the first Time-matrix noise covariance matrix; For the first Time-matter noise; For adjustment coefficients; This is represented as covariance operation; For the first The noise covariance matrix is measured at any given time. For the first The noise covariance matrix is measured at any given time. Forgetting factor; For the first Observe the transpose of the residuals at all times.
[0043] The updated process noise covariance matrix and the updated measurement noise covariance matrix are obtained through the above calculations.
[0044] In practical applications, when the slope estimation error is large (greater than 0.1°), it indicates that the weights of the road slope angle and slope change rate in the process noise covariance matrix are improperly set, and the weights of the slope angle and slope change rate should be increased. When the radar ranging fluctuation is large (greater than 0.3m), it indicates that the weight of the radar ranging value in the measurement noise covariance matrix is too low, and the weight of the radar ranging value should be increased. Through the above dynamic adjustments, the adaptive multi-source extended Kalman filter can always maintain the optimal estimation performance under different working conditions.
[0045] After multiple iterations based on the updated process noise covariance matrix and the updated measurement noise covariance matrix, the slope estimation result is finally generated.
[0046] In a preferred embodiment, after obtaining the slope estimation result, the slope estimation result can be further smoothed by using a moving average filter to eliminate high-frequency fluctuations and make the slope estimation result more stable.
[0047] Step 4: Traditional vehicle longitudinal dynamics models usually do not consider the influence of slope, resulting in poor control performance when driving on slopes. Therefore, this embodiment improves the traditional vehicle longitudinal dynamics model to obtain an improved vehicle longitudinal dynamics model.
[0048] like Figure 4 The process of this step is shown. The slope disturbance acceleration is introduced as a feedforward compensation term into the state equation of the vehicle longitudinal dynamics model to construct an improved vehicle longitudinal dynamics model. Then, by combining the reference trajectory input (actual speed, actual distance, and control acceleration) and state feedback information, a secondary programming solution is performed with the equivalent energy required for speed tracking, the equivalent energy required for distance tracking, and the energy lost due to control acceleration as optimization objectives. Finally, the desired acceleration and target distance control signals are output.
[0049] First, based on the road slope angle obtained from the slope estimation results in step three, the slope disturbance acceleration is calculated in conjunction with the gravitational acceleration, as shown in the following expression:
[0050] In the formula: It is the acceleration due to gravity; For slope disturbance acceleration; This refers to the road slope angle.
[0051] By incorporating the gradient disturbance acceleration into the state equation of the vehicle's longitudinal dynamics model, the corrected state equation is obtained, and its calculation formula is as follows:
[0052] In the formula: To find the derivative with respect to the extended state vector; For extended state vectors; For the desired acceleration; To extend the dynamic relationship between state vectors; This is to assess the current impact on its own speed and vehicle spacing. The effect of disturbance on acceleration state; It is the identity matrix; This is the output vector.
[0053] The modified state equations are calibrated based on the preset desired velocity trajectory, resulting in an improved vehicle longitudinal dynamics model.
[0054] The embodiment constructs a multi-objective optimization problem based on an improved vehicle longitudinal dynamics model. Specifically, it requires setting a prediction time domain and a control time domain. The prediction time domain determines the length of time the controller predicts forward, while the control time domain determines the length of the control sequence output by the controller. In this embodiment, the prediction time domain is set to 10, corresponding to the vehicle's motion state within the next 1.5 seconds; the control time domain is set to 3, used to balance control response speed and computational complexity.
[0055] Based on this, a cost function is constructed with the equivalent energy required for velocity tracking, the equivalent energy required for spacing tracking, and the energy lost due to control acceleration as the objective problem. Its expression is as follows:
[0056] In the formula: The cost function is expressed in joules. This represents the equivalent energy required for velocity tracking, measured in joules. Weighting coefficients representing the equivalent energy required for velocity tracking; For vehicle mass; N represents the length of the prediction time domain; Indicates the index of the current time step; express The actual speed at any given moment; for Reference speed at any given moment; This represents the equivalent energy required for spacing tracking, measured in joules. Weighting coefficients for the equivalent energy required for spacing tracking; This is the equivalent stiffness coefficient, with a value ranging from 0.2 to 0.8. for Actual distance between vehicles at any given time; To maintain a safe following distance; This represents the energy lost due to controlled acceleration, measured in joules. Weighting coefficients to control the energy loss due to acceleration; The time constant is 0.1 to 0.5 seconds. To control acceleration.
[0057] It should be noted that the weighting coefficients in this embodiment are not fixed values, but are dynamically adjusted based on the slope estimation results. This dynamic adjustment strategy allows the controller to adapt to the control requirements under different operating conditions. Specifically, in uphill conditions, the vehicle needs to overcome additional slope resistance, resulting in decreased acceleration performance. In this case, the weighting coefficient of the equivalent energy required for distance tracking is appropriately increased. Prioritize maintaining a safe following distance to avoid following too closely due to insufficient power; when driving downhill, vehicles are prone to speed overshoot due to the gravitational component, so increase the weighting coefficient for controlling acceleration energy loss in this situation. Strengthen braking intervention to suppress abnormal speed increases; under flat road conditions, use the weighting coefficient of the equivalent energy required for speed tracking. Prioritize ensuring speed tracking accuracy to achieve smooth following.
[0058] In this embodiment, the velocity deviation weighting coefficient, spacing deviation weighting coefficient, and control acceleration weighting coefficient in the cost function are determined through offline calibration, as detailed below: Under typical driving conditions, Bayesian optimization is used to achieve the optimal weight combination with the goal of minimizing the cost function. The initial calibration results are: 2.0, 5.0, and 1.0. In practical applications, the above coefficients can be fine-tuned according to driver style or driving condition preferences, with the adjustment range not exceeding ±30% of the calibration value.
[0059] To meet the real-time solution requirements of automotive embedded controllers, the state equations need to be discretized. This embodiment uses the first-order Euler discretization method. After discretization, the state equations are substituted into the cost function to obtain a multi-performance objective optimization problem, which is then transformed into a standard quadratic programming problem.
[0060] In the formula: For control quantity sequence The transpose of the matrix; for The transpose of the matrix; for The transpose of the matrix; The error weighted diagonal matrix for the state variables; The weight matrix is used to control the increment; This is the mapping coefficient matrix from control variables to state terms; It is a direct weighted diagonal matrix of control quantities; for The transpose of the matrix; The weighted diagonal matrix is used to control the increment; This is the relaxation factor coefficient matrix; To predict the sequence of control variables in the time domain; To optimize the constant term in the objective function.
[0061] Solving the multi-performance objective optimization problem yields the desired acceleration and target vehicle distance control signals.
[0062] Step 5: After obtaining the desired acceleration, it is necessary to convert the desired acceleration into specific control commands. Therefore, this embodiment uses a dual-network deep reinforcement learning model to output power control commands.
[0063] First, we define the state space, action space, and reward function. Specifically, the state space includes the vehicle's power demand, battery state of charge, and engine efficiency. These parameters directly reflect the vehicle's energy demand and energy storage status. The action space is defined as the engine power distribution ratio. In this embodiment, the engine power command is divided into multiple continuous power intervals, corresponding to the engine's full operating range from idle to full load. This division method avoids power distribution jumps and improves the smoothness of power output.
[0064] The expression for the reward function is as follows:
[0065] In the formula: Represents the reward function; Sampling period The energy consumed by fuel internally, measured in joules; Sampling period The electrical energy consumed internally, measured in joules; Sampling period The energy lost due to internal power fluctuations, measured in joules; The weighting factor for fuel energy consumption is set to 0.7. The weighting factor for electricity consumption is set to 0.5; The weighting factor for energy loss due to power fluctuations is set to 0.2; For time integration variables; Fuel consumption rate; This refers to the calorific value of fuel oil. This is the battery open-circuit voltage; This refers to the battery current. This represents the change in power.
[0066] During the training phase of the dual-network deep reinforcement learning model, state data of the vehicle during driving is collected, including the vehicle's power demand, battery state of charge, and engine efficiency. The model selects actions based on the current policy, receives a reward after execution, and stores the experience in an experience replay buffer with a capacity of 10. 5 During training, a priority experience replay mechanism is adopted to prioritize the learning of experience samples with high reward error, thereby accelerating the convergence speed. After training, the real-time vehicle power demand, battery state of charge, and engine efficiency are input into the trained dual-network deep reinforcement learning model, which outputs power control commands that include the engine power allocation ratio.
[0067] Step Six: After obtaining the power control command and target distance control signal, convert them into specific execution actions. In this embodiment, the power control command and target distance control signal are received through the vehicle CAN (Controller Area Network) bus, and different drive modes are switched according to the vehicle power requirements and battery status.
[0068] When the vehicle's power demand is no greater than a preset percentage (30%) of the motor's peak power and the battery's state of charge is no less than a threshold (60%), it enters pure electric drive mode, shuts off the engine, and is driven solely by the motor. This mode is suitable for low-power demand scenarios such as short-distance commuting in the city and following other vehicles in traffic jams.
[0069] When the vehicle's power demand is not less than a preset percentage (30%) of the motor's peak power, it enters the range-extended drive mode, which allows the engine to operate in the optimal efficiency range to generate electricity. The electric energy is prioritized to drive the motor, and the remaining electricity is used to replenish the battery. This ensures power output while avoiding the engine operating in an inefficient range.
[0070] When the vehicle's power demand is not greater than a preset percentage (30%) of the motor's peak power but the battery's state of charge is not greater than the threshold (60%), the range-extending charging mode is entered. The engine continues to operate in the optimal efficiency range to generate electricity, and the generated electricity is used to meet the vehicle's driving needs first. The remaining part is used to charge the battery. After the battery's state of charge recovers to the threshold (60%), the mode is switched back to pure electric drive mode.
[0071] Based on the same inventive concept, this invention also provides a multi-data fusion and reinforcement learning hybrid vehicle safety and energy-saving system. Since the principle of this multi-data fusion and reinforcement learning hybrid vehicle safety and energy-saving system is similar to that of the aforementioned multi-data fusion and reinforcement learning hybrid vehicle safety and energy-saving method, the implementation of this multi-data fusion and reinforcement learning hybrid vehicle safety and energy-saving system can refer to the implementation of the multi-data fusion and reinforcement learning hybrid vehicle safety and energy-saving method, and the repeated parts will not be described again.
[0072] In specific implementation, the multi-data fusion and reinforcement learning hybrid vehicle safety and energy-saving system provided in this embodiment of the invention specifically includes: The data acquisition module is used to collect multi-source data from each vehicle in the fleet; The decision module is used to synchronize multi-source data in time to obtain synchronized data; spatially register the synchronized data to obtain a fused dataset; input the fused dataset into an adaptive multi-source extended Kalman filter for road slope estimation to generate slope estimation results; based on the slope estimation results and a pre-set desired speed trajectory, construct an improved vehicle longitudinal dynamics model; based on the improved vehicle longitudinal dynamics model, solve for the equivalent energy required for speed tracking, the equivalent energy required for distance tracking, and the energy lost due to control acceleration as optimization objectives to obtain the desired acceleration and target distance control signals; convert the desired acceleration into vehicle power demand, input the vehicle power demand into a pre-constructed dual-network deep reinforcement learning model for decision processing, and output power control commands; The execution module controls the engines, motors, and braking systems of each vehicle in the convoy to work together according to the power control commands and target distance control signals. It recovers braking energy through multiple mode switching to achieve energy saving for each vehicle in the convoy.
[0073] Accordingly, embodiments of the present invention also provide a multi-data fusion and reinforcement learning hybrid vehicle safety and energy-saving device, including a processor and a memory, wherein the processor executes a computer program stored in the memory to implement the multi-data fusion and reinforcement learning hybrid vehicle safety and energy-saving method provided in embodiments of the present invention.
[0074] For more detailed information on the above methods, please refer to the relevant content disclosed in the foregoing embodiments, which will not be repeated here.
[0075] Accordingly, embodiments of the present invention also provide a computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the above-described multi-data fusion and reinforcement learning hybrid vehicle safety and energy-saving method provided in embodiments of the present invention.
[0076] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems, devices, and storage media disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and relevant details can be found in the method section.
[0077] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0078] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0079] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0080] The above provides a detailed description of the multi-data fusion and reinforcement learning hybrid vehicle safety and energy-saving method, system, equipment, and storage medium provided by this invention. Specific examples have been used to illustrate the principles and implementation methods of this invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.
Claims
1. A method for safe and energy-saving hybrid vehicle fleets using multi-data fusion and reinforcement learning, characterized in that, Includes the following steps: Collect multi-source data from each vehicle in the fleet; Synchronize data from multiple sources to obtain synchronized data; Spatial registration is performed on the synchronized data to obtain a fused dataset; The fused dataset is input into an adaptive multi-source extended Kalman filter to estimate the road slope and generate the slope estimation results. Based on the slope estimation results and the pre-set desired speed trajectory, an improved vehicle longitudinal dynamics model is constructed. Based on the improved vehicle longitudinal dynamics model, the equivalent energy required for speed tracking, the equivalent energy required for spacing tracking, and the energy lost due to control acceleration are used as optimization objectives to solve the problem and obtain the desired acceleration and target distance control signals. The desired acceleration is converted into the vehicle's required power, which is then input into a pre-built dual-network deep reinforcement learning model for decision processing, and outputs power control commands. Based on the power control command and the target distance control signal, the engine, motor and braking system of each vehicle in the fleet are controlled to work together, and braking energy is recovered through multiple mode switching to achieve energy saving for each vehicle in the fleet.
2. The multi-data fusion and reinforcement learning hybrid vehicle safety and energy-saving method according to claim 1, characterized in that, The process of synchronizing multi-source data in time to obtain synchronized data includes: The HSC-DDC time synchronization unit is used to synchronize the time of multi-source data. The specific process is as follows: It provides a nanosecond-level reference clock, and based on the nanosecond-level reference clock, it combines the vehicle's inertial measurement unit to correct multi-source data to obtain clock-synchronized data; Dynamically predict the data after clock synchronization to generate delay compensation. Based on the delay compensation amount, the clock-synchronized data is time-aligned to obtain preliminary synchronized data; the predicted delay value is calculated based on the delay compensation amount. When the predicted delay value exceeds the dynamic threshold, the neighboring frame interpolation-confidence scoring mechanism is triggered based on the preliminary synchronization data to process the predicted delay value and obtain a confidence score. The confidence score is compared with a preset threshold, and data with a confidence score lower than the preset threshold is removed to obtain the synchronized data.
3. The multi-data fusion and reinforcement learning hybrid vehicle safety and energy-saving method according to claim 1, characterized in that, Spatial registration is performed on the synchronized data to obtain a fused dataset, including: Semantic augmentation-based fast point cloud-visual registration is used to perform time synchronization on the synchronized data. The specific process is as follows: The image and point cloud data in the synchronized data are processed as follows: The image is input into a lightweight convolutional neural network for semantic segmentation processing to obtain the semantic segmentation result; Based on the semantic segmentation results, the core road regions of the point cloud data are filtered to obtain the filtered point cloud dataset. Dynamic distortion correction is performed on the filtered point cloud dataset to obtain the corrected point cloud data; The corrected point cloud data is coarsely registered with the image to obtain the coarse registration result; The coarse registration result is refined to obtain the accurate registration result; The precise matching result is fused with the semantic segmentation result to generate a fused dataset containing road surface geometric features, slope information, and semantic labels.
4. The multi-data fusion and reinforcement learning hybrid vehicle safety and energy-saving method according to claim 1, characterized in that, The fused dataset is input into an adaptive multi-source extended Kalman filter for road slope estimation, generating slope estimation results, including: First, the state variables in the adaptive multi-source extended Kalman filter are extended to obtain an extended state vector that includes vehicle longitudinal velocity, acceleration, road slope angle, radar ranging value, surface normal angle, slope change rate, and radar ranging error. The observation residual is calculated based on the extended state vector. The process noise covariance matrix and measurement noise covariance matrix in the adaptive multi-source extended Kalman filter are dynamically weighted based on the observation residual to obtain the updated process noise covariance matrix and the updated measurement noise covariance matrix. The slope estimation result is generated by iteratively updating the updated process noise covariance matrix and the updated measurement noise covariance matrix.
5. The multi-data fusion and reinforcement learning hybrid vehicle safety and energy-saving method according to claim 1, characterized in that, Based on the slope estimation results and the pre-defined desired speed trajectory, an improved vehicle longitudinal dynamics model is constructed, including: The slope disturbance acceleration is calculated based on the slope estimation results. By incorporating the slope disturbance acceleration into the state equation of the vehicle's longitudinal dynamics model, the corrected state equation is obtained. The modified state equations are calibrated based on the pre-set desired velocity trajectory to obtain an improved vehicle longitudinal dynamics model.
6. The multi-data fusion and reinforcement learning hybrid vehicle safety and energy-saving method according to claim 1, characterized in that, Based on the improved vehicle longitudinal dynamics model, the equivalent energy required for speed tracking, the equivalent energy required for distance tracking, and the energy lost due to control acceleration are used as optimization objectives to obtain the desired acceleration and target distance control signals, including: A multi-objective optimization problem is constructed by defining a prediction time domain and a control time domain. Using the equivalent energy required for velocity tracking, the equivalent energy required for spacing tracking, and the energy lost due to control acceleration as cost terms for the objective problem, a cost function is constructed, the expression of which is as follows: In the formula: The cost function is expressed in joules. This represents the equivalent energy required for velocity tracking, measured in joules. Weighting coefficients representing the equivalent energy required for velocity tracking; For vehicle mass; N represents the length of the prediction time domain; Indicates the index of the current time step; express The actual speed at any given moment; for Reference speed at any given moment; This represents the equivalent energy required for spacing tracking, measured in joules. Weighting coefficients for the equivalent energy required for spacing tracking; This is the equivalent stiffness coefficient; for Actual distance between vehicles at any given time; To maintain a safe following distance; This represents the energy lost due to controlled acceleration, measured in joules. Weighting coefficients to control the energy loss due to acceleration; It is a time constant; To control acceleration; The weight coefficients in the cost function are dynamically adjusted based on the slope estimation results to obtain the adjusted cost function. By combining the improved vehicle longitudinal dynamics model with the adjusted cost function, a multi-performance objective optimization problem is obtained. The multi-performance objective optimization problem is solved to obtain the desired acceleration and target vehicle distance control signals.
7. The multi-data fusion and reinforcement learning hybrid vehicle safety and energy-saving method according to claim 1, characterized in that, The desired acceleration is converted into the vehicle's required power, which is then input into a pre-built dual-network deep reinforcement learning model for decision processing. The resulting power control commands include: Construct the state space, action space, and reward function of a dual-network deep reinforcement learning model; The state space includes vehicle power demand, battery state of charge, and engine efficiency. The action space is the engine power distribution ratio; The expression for the reward function is as follows: In the formula: Represents the reward function; Sampling period The energy consumed by fuel internally, measured in joules; Sampling period The electrical energy consumed internally, measured in joules; Sampling period The energy lost due to internal power fluctuations, measured in joules; This is a weighting coefficient for fuel energy consumption; This is a weighting coefficient for electricity consumption; This is the weighting coefficient for energy loss due to power fluctuations; For time integration variables; Fuel consumption rate; This refers to the calorific value of fuel oil. This is the battery open-circuit voltage; This refers to the battery current. This represents the change in power. The dual-network deep reinforcement learning model is trained to obtain a well-trained dual-network deep reinforcement learning model. The vehicle's required power, battery state of charge, and engine efficiency are input into a trained dual-network deep reinforcement learning model, which then outputs power control commands.
8. The multi-data fusion and reinforcement learning hybrid vehicle safety and energy-saving method according to claim 1, characterized in that, Based on power control commands and target distance control signals, the engine, motor, and braking system of each vehicle in the convoy are controlled to work in coordination. Through multiple mode switching, braking energy recovery is achieved to realize energy conservation for each vehicle in the convoy, including: When the vehicle's power demand is not greater than a preset percentage of the motor's peak power and the battery's state of charge is not less than a threshold, it enters pure electric drive mode, shuts off the engine, and is driven solely by the motor. When the vehicle's power demand is not less than a preset percentage of the motor's peak power, it enters the range-extended drive mode, which allows the engine to operate in the optimal efficiency range to generate electricity. The electricity is first supplied to the motor drive, and the remaining electricity is used to charge the battery. When the vehicle's power demand is no greater than a preset percentage of the motor's peak power and the battery's state of charge is no greater than a threshold, the range extender charging mode is entered, allowing the engine to operate and generate electricity within its optimal efficiency range. The electricity is prioritized to meet the vehicle's driving needs, and the remaining electricity is used to charge the battery.
9. A multi-data fusion and reinforcement learning hybrid vehicle fleet safety and energy-saving system, characterized in that, include: The data acquisition module is used to collect multi-source data from each vehicle in the fleet; The decision module is used to synchronize multi-source data over time to obtain synchronized data. Spatial registration is performed on the synchronized data to obtain a fused dataset; the fused dataset is then input into an adaptive multi-source extended Kalman filter for road slope estimation to generate slope estimation results; Based on the slope estimation results and the pre-set desired speed trajectory, an improved vehicle longitudinal dynamics model is constructed. Based on the improved vehicle longitudinal dynamics model, the equivalent energy required for speed tracking, the equivalent energy required for spacing tracking, and the energy loss of control acceleration are used as optimization objectives to solve for the desired acceleration and target distance control signals. The desired acceleration is converted into the vehicle's required power, and the vehicle's required power is input into a pre-constructed dual-network deep reinforcement learning model for decision processing, outputting power control commands. The execution module controls the engines, motors, and braking systems of each vehicle in the convoy to work together according to the power control commands and target distance control signals. It recovers braking energy through multiple mode switching to achieve energy saving for each vehicle in the convoy.
10. A multi-data fusion and reinforcement learning hybrid vehicle fleet safety and energy-saving device, characterized in that, It includes a processor and a memory, wherein the processor executes a computer program stored in the memory to implement the multi-data fusion and reinforcement learning hybrid vehicle safety and energy-saving method as described in any one of claims 1 to 8.