Deep learning-based energy consumption optimization control method and system for pure electric wide-body vehicle
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-27
- Publication Date
- 2026-08-11
AI Technical Summary
然而,现有技术方案大多隐含地假设驾驶行为是固定的,或在单一驾驶风格下进行训练和验证
1.通过第一深度学习模型对驾驶行为特征进行实时识别,输出驾驶风格类别及其置信度,将模糊、难以量化的驾驶员操作习惯转化为可供决策系统直接利用的语义标签。同时结合第二深度学习模型对工况特征的识别,构建融合车辆状态、驾驶风格与工况类别的综合状态向量,使深度强化学习策略能够清晰感知驾驶员操作特性、车辆动力学状态及道路环境信息的完整上下文信息。这一机制使得控制策略能够针对激进型驾驶风格提供更敏捷的动力响应,针对平稳型驾驶风格强化能量回收与平顺性,实现“千人千面”的自适应优化,降低因驾驶员差异带来的能耗波动。
Smart Images

Figure CN122540172A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of vehicle control technology, specifically a method and system for optimizing energy consumption control of pure electric wide-body vehicles based on deep learning. Background Technology
[0002] The statements in this section merely refer to the background art related to this invention and do not necessarily constitute prior art.
[0003] As mining operations move towards greener and more efficient models, pure electric wide-body vehicles are gradually becoming the mainstay of mining transportation due to their advantages such as zero emissions and low operating costs. Mining vehicle operations have distinct scenario characteristics: relatively fixed work routes form a highly repetitive work cycle of "loading-heavy-loaded transport-empty return-unloading"; however, the working environment is extremely harsh, with large load variations (from empty to fully loaded), significant changes in road gradient (continuous long uphill / downhill slopes), and complex road conditions. These characteristics place extremely high demands on the energy efficiency of the entire vehicle.
[0004] However, in actual mine operations, a multi-shift system of "changing drivers but not vehicles" is commonly adopted, meaning that the same vehicle is driven in shifts by drivers with different driving styles (such as aggressive and moderate drivers). This results in significant fluctuations in vehicle energy consumption even when the vehicle is traveling on the same physical road surface and under the same load conditions, due to significant differences in the accelerator and brake pedal operation habits of different drivers.
[0005] To address this issue, existing technologies have proposed energy consumption optimization control methods for pure electric engineering vehicles, mainly including the following technical solutions: Rule-based energy management strategies rely on pre-defined logical thresholds or MAPs (such as lookup tables) to control power output, regenerative braking intensity, and thermal management system operating points. Their advantages include simple logic and high reliability, but their disadvantage is that the rules are static and non-adaptive. When faced with varying driving behaviors and complex combinations of operating conditions, fixed rules cannot be dynamically adjusted, often sacrificing fuel economy for power performance, or impacting driving experience and operational efficiency in different driver scenarios to ensure fuel economy.
[0006] Model-based optimization control methods, represented by Model Predictive Control (MPC), establish vehicle longitudinal dynamics models, battery models, etc., and combine them with operating condition predictions to solve for the optimal control sequence online. Theoretically, their control performance is superior to rule-based methods. However, their performance is highly dependent on the accuracy of the model and the accuracy of the operating condition predictions. In mining environments, key parameters such as load, gradient, and road rolling resistance have significant uncertainties and are difficult to obtain accurately in real time, leading to model mismatch and a significant reduction in control effectiveness. Furthermore, MPC has a heavy online computational burden and lacks the ability to learn and iteratively evolve from historical operating data.
[0007] Machine learning / reinforcement learning methods in single scenarios: In recent years, some research has attempted to apply machine learning, especially deep reinforcement learning (DRL), to energy management. These methods improve adaptability by learning policies instead of fixed rules. However, most existing solutions implicitly assume that driving behavior is fixed or are trained and validated under a single driving style. When directly applied to multi-driver scenarios where "the driver changes but the car remains the same," the policy network cannot perceive and distinguish the current driver's control characteristics, resulting in severely insufficient generalization ability and potentially even "incompatibility," leading to increased energy consumption fluctuations or safety risks.
[0008] It is evident that existing technologies cannot enable vehicle energy consumption control strategies to effectively perceive and adapt to the differences in operating habits of different drivers under the "change drivers but not vehicles" operation model. This leads to a "mismatch" between the strategy and driver behavior, resulting in significant fluctuations in energy consumption and limitations in optimization effects. Summary of the Invention
[0009] This invention provides a deep learning-based energy consumption optimization control method and system for pure electric wide-body vehicles, constructing a closed-loop system of "perception-decision-evolution". Through deep learning, driver style and operating conditions are perceived in real time, transforming fuzzy behaviors into quantifiable labels; then, deep reinforcement learning dynamically generates multi-dimensional control parameters to collaboratively optimize subsystems such as power and regeneration; finally, actual energy consumption comparison generates reward feedback, driving the model to self-evolve online, and using federated learning to achieve cross-vehicle knowledge sharing, continuously adapting to optimal energy consumption under the "driver change without vehicle change" scenario.
[0010] The first aspect of this invention discloses a deep learning-based energy consumption optimization control method for pure electric wide-body vehicles, comprising the following steps: Vehicle operation data, environmental and road segment data are collected, and after preprocessing and feature extraction, driving behavior characteristics and operating condition characteristics are obtained; The driving behavior features are processed by the first deep learning model, which outputs the driving style category and its confidence level to represent the current driver's operating habits. The operating condition features are processed by a second deep learning model, which outputs the operating condition category and its confidence level to characterize the current operating scenario of the vehicle. By integrating real-time vehicle status information, driving style category and confidence level, and operating condition category and confidence level, a comprehensive state vector representing the coupling relationship between "human-vehicle-road" is constructed. The combined state vector is input into a pre-trained deep reinforcement learning model, and the output is a control parameter vector for co-optimizing multiple subsystems. The obtained control parameter vector is mapped into executable control commands, which are then sent to the dynamic response subsystem, regenerative braking distribution subsystem, thermal management subsystem, and auxiliary load subsystem for execution.
[0011] Furthermore, during the execution of control commands, the actual energy consumption and performance indicators within the current control cycle are calculated, a reward signal is generated by comparing it with the benchmark energy consumption, and the parameters of the deep reinforcement learning model are updated based on the reward signal.
[0012] Furthermore, the driving behavior characteristics are constructed using a sliding time window and include at least one of the following types of characteristics: Operation sequence characteristics include at least the timing segments of accelerator pedal opening, brake pedal pressure, vehicle speed, acceleration, and torque request; Statistical characteristics, including at least one of the following: mean, variance, extreme values, kurtosis, frequency of rapid acceleration events, and frequency of rapid deceleration events; Coupling characteristics include at least one of pedal-vehicle speed response relationship and braking-deceleration efficiency.
[0013] Furthermore, the first deep learning model is implemented using any of the following architectures: Network architecture combining 1D-CNN with BiLSTM / GRU attention mechanism; Transformer encoder network structure combined with pooling layers; A lightweight network structure combining a multilayer perceptron with a Softmax output layer; The output of the first deep learning model includes the probability distribution of driving style categories, driving style labels, confidence scores, and style embedding vectors.
[0014] Furthermore, the operating condition characteristics include at least one of the following types of information: Vehicle longitudinal status information, including at least one of vehicle speed, acceleration, motor torque, power, SOC, battery temperature, and motor temperature; Road gradient and resistance information, including at least one of real-time gradient, gradient change rate, and road rolling resistance estimation; Load estimation information is obtained by combining suspension pressure signals or longitudinal dynamic back-propagation with operational phase correction. Forward prediction information includes forward slope sequences or forward road segment types from maps, high-precision positioning, or V2X.
[0015] Furthermore, the working condition categories are defined using a hierarchical combination method, including combinations of load state and slope state; The second deep learning model takes continuous slope values as input, combines the slope change rate and filtering smoothing mechanism to handle continuous slope changes, and outputs the probability distribution of working condition categories, working condition labels, confidence scores, and working condition embedding vectors.
[0016] Furthermore, by integrating real-time vehicle status information, driving style category and confidence level, and operating condition category and confidence level, a comprehensive state vector is constructed, including: The driving style and operating condition are adaptively weighted based on the confidence level. When the confidence level is lower than the preset threshold, the corresponding category is set to an unknown state. Extract the style embedding vector of the first deep learning model and the working condition embedding vector of the second deep learning model, and generate a fused representation by concatenation, gated fusion or feature crossing. The vehicle's real-time status information, fused representation, driving style probability distribution, and operating condition probability distribution are combined to construct a comprehensive state vector.
[0017] Furthermore, the control parameter vector includes at least one or more of the following parameters: Torque response coefficient, used to correct the mapping relationship between pedal opening and motor requested torque; The regenerative braking intensity coefficient is used to determine the proportion allocated to regenerative braking during braking. Thermal management operating point factor, used to adjust the operating point of cooling pumps, fans, or heating strategies; Auxiliary load power factor, used to limit the upper limit of power for air conditioning or hydraulic systems; Predictive control activation level coefficient is used to adjust energy management gain in advance based on information about the road ahead.
[0018] Furthermore, the reward signal is calculated using the following formula: r=w1·(-Eunit)+w2·ΔE+w3·Rrecu-λ1·Psafe-λ2·Pcomfort-λ3·Pwear; Wherein, Eunit is the energy consumption per unit mileage or per unit operation cycle, ΔE is the improvement in energy consumption relative to the baseline, Rrecu is the energy recovery rate, Psafe, Pcomfort, and Pwear are the safety boundary penalty, comfort penalty, and wear penalty, respectively; the baseline energy consumption is derived from the rule-based strategy energy consumption, the moving average of the previous stable version strategy energy consumption, or the historical average energy consumption of the same road segment; the weights w1, w2, w3, λ1, λ2, and λ3 are determined through engineering constraint priority calibration, data-driven offline optimization, or adaptive online adjustment.
[0019] A second aspect of the present invention discloses a deep learning-based energy consumption optimization control system for pure electric wide-body vehicles, comprising: The data acquisition and feature extraction module is configured to: collect vehicle operation data, environmental and road segment data, and obtain driving behavior features and operating condition features through preprocessing and feature extraction; The driving style recognition module is configured to: pass driving behavior features through a first deep learning model to output a driving style category and its confidence level that characterizes the current driver's operating habits; The operating condition identification module is configured to: use the operating condition features through a second deep learning model to output the operating condition category and its confidence level, which characterize the current operating scenario of the vehicle; The state construction module is configured to: integrate real-time vehicle state information, driving style category and confidence level, and operating condition category and confidence level to construct a comprehensive state vector representing the coupling relationship between "human-vehicle-road"; The policy generation module is configured to: take the state vector as input to a pre-trained deep reinforcement learning model and output a control parameter vector for co-optimizing multiple subsystems; The control execution module is configured to map the obtained control parameter vector into executable control commands, which are then sent to the power response subsystem, regenerative braking distribution subsystem, thermal management subsystem, and auxiliary load subsystem for execution.
[0020] Compared with existing technologies, one or more of the above technical solutions have the following beneficial effects: 1. A first deep learning model is used to identify driving behavior characteristics in real time, outputting driving style categories and their confidence levels. This transforms vague and difficult-to-quantify driver operating habits into semantic labels that can be directly used by the decision-making system. Simultaneously, a second deep learning model is used to identify operating condition characteristics, constructing a comprehensive state vector that integrates vehicle state, driving style, and operating condition category. This allows the deep reinforcement learning strategy to clearly perceive the complete contextual information of driver operating characteristics, vehicle dynamics, and road environment. This mechanism enables the control strategy to provide a more agile power response for aggressive driving styles and enhance energy recovery and smoothness for stable driving styles, achieving personalized adaptive optimization and reducing energy consumption fluctuations caused by driver differences.
[0021] 2. The output control parameter vector comprises a multi-dimensional set of parameters, including torque response coefficient, regenerative braking intensity coefficient, thermal management operating point coefficient, auxiliary load power coefficient, and predictive control activation level. This can be applied to multiple subsystems simultaneously, achieving vehicle-level energy consumption optimization. Based on this, by comparing actual operating energy consumption with baseline energy consumption and combining multiple objective constraints such as safety, comfort, and wear to construct a reward signal, a "optimization-feedback-correction" online closed-loop mechanism is formed. The deep reinforcement learning model continuously updates parameters based on the reward signal, enabling the control strategy to evolve continuously with vehicle aging, environmental changes, and driving behavior drift, reducing reliance on manual calibration and maintaining optimal energy consumption in the long term.
[0022] 3. When constructing the comprehensive state vector, not only are the recognition results of driving style and working condition category directly utilized, but their respective confidence levels are also introduced for adaptive weight allocation. When the confidence level of a certain recognition result is low, it is automatically downweighted or set to an unknown state to avoid misidentification and misleading decisions. Simultaneously, the style embedding vector from the first deep learning model and the working condition embedding vector from the second deep learning model are extracted, and a fused representation is generated using concatenation, gated fusion, or feature cross-validation methods, fully transmitting high-dimensional semantic information to the policy network. This mechanism enables the deep reinforcement learning model to more fully understand the coupling relationship between driving behavior and working condition scenarios, maintaining the accuracy and stability of decision-making in complex and ever-changing mining environments, while providing a confidence basis for safety fallback and version rollback mechanisms. Attached Figure Description
[0023] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0024] Figure 1 A schematic diagram of the energy consumption optimization control system architecture for a pure electric wide-body vehicle provided in one or more embodiments of the present invention; Figure 2 A schematic diagram of an online control process provided for one or more embodiments of the present invention; Figure 3 This is a schematic diagram of an offline training process provided for one or more embodiments of the present invention. Detailed Implementation
[0025] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0026] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0027] Terminology Explanation: Pure electric wide-body vehicles: Pure electric wide-body dump trucks or transport vehicles used in mining / engineering scenarios, equipped with power systems, battery systems, braking systems, thermal management systems, and on-board controllers (such as VCUs).
[0028] Driving style (driving behavior category): A behavior category formed by the driver's operating habits such as pedaling, braking, and speed changes, such as aggressive, steady, and economical driving; it is obtained by classifying driving behavior characteristics through a model.
[0029] Operating condition (vehicle operating condition category): The category of the vehicle's operating situation, such as heavy-load uphill, heavy-load downhill, unloaded uphill, unloaded downhill, etc., or the situational state formed by the combination of load, gradient, road conditions, and operation cycle stage.
[0030] Control parameter vector (action vector): A set of multi-dimensional control quantities output by the strategy model, used to adjust power response (pedal-torque mapping / torque change rate, etc.), regenerative braking distribution, thermal management operating point, auxiliary load power limit, etc.
[0031] Deep reinforcement learning (DRL): a learning method that models a state-action-reward relationship and optimizes a policy through interaction with the environment. This invention is used to generate energy consumption optimization control policies.
[0032] Reward: An evaluation signal used to assess the quality of a strategy and drive learning updates; it can consist of energy consumption reduction and constraints and penalties such as safety, comfort, and wear.
[0033] Federated learning (FL): A training mechanism in which multiple vehicles upload model parameters / gradients without uploading the original data, and the server aggregates them to form a global model and distributes updated models.
[0034] Offline training: The model is trained / iterated on the vehicle or an offline computing platform based on historical driving data. After training, the model can be deployed to the vehicle for execution.
[0035] Energy consumption indicators: Indicators used to measure vehicle energy consumption, such as energy consumption per unit mileage, energy consumption per unit operating cycle, and energy recovery rate.
[0036] Safety fallback / rollback mechanism: When the policy output is abnormal or the safety boundary (temperature, current, braking stability, etc.) is triggered, the protection mechanism switches to rule control or the previous stable version model.
[0037] As described in the background section, existing technologies cannot effectively perceive and adapt to the differences in the operating habits of different drivers under the "changing drivers but not changing vehicles" operation model. This leads to a "mismatch" between the strategy and driver behavior, resulting in large fluctuations in energy consumption and limitations in optimization effects.
[0038] In existing technologies, strategy optimization and updates typically rely on offline data collection and manual calibration, failing to compare "optimized actual energy consumption" with "historical energy consumption" in real time and to feed this comparison result directly back to the strategy model as a reward signal to drive it to self-correct online or periodically, making it difficult to achieve continuous performance evolution.
[0039] Meanwhile, to improve the model's generalization ability, the ideal approach is to aggregate operational data from multiple vehicles and drivers for centralized training. However, this would incur significant data transmission costs and data compliance risks (such as data confidentiality requirements in mining areas). Relying solely on single-vehicle data for training, on the other hand, would result in underfitting and weak generalization ability due to limited data volume and coverage of specific scenarios.
[0040] In the scenario of mining operation cycle (loading-transportation-unloading), there are characteristics such as large range of working conditions (significant changes in load, slope, road conditions, and cycle time), strong repetitiveness of operation cycle, and "changing people but not changing vehicles", which leads to significant differences in driving behavior.
[0041] Therefore, this solution provides a deep learning-based energy consumption optimization control method and system for pure electric wide-body vehicles; it constructs a closed-loop intelligent system of "perception-decision-evolution". Through deep learning, it perceives driver style and operating conditions in real time, transforming fuzzy behaviors into quantifiable labels; then, deep reinforcement learning dynamically generates multi-dimensional control parameters to collaboratively optimize subsystems such as power and regeneration; finally, it uses actual energy consumption comparisons to generate reward feedback, driving the model to self-evolve online, and leverages federated learning to achieve cross-vehicle knowledge sharing, continuously adapting to optimal energy consumption under the "driver change without vehicle change" model.
[0042] The architecture of a deep learning-based energy consumption optimization control system for pure electric wide-body vehicles is as follows: Figure 1 As shown in Table 1.
[0043] Table 1 System Architecture
[0044] The aforementioned modules can be implemented by the processor and memory in the vehicle control unit (VCU / domain controller) or deployed in a distributed manner in the vehicle network.
[0045] Online closed-loop energy consumption optimization control methods, such as Figure 2 As shown, it includes the following steps: S1 Data Acquisition: Collects vehicle operation data and environmental / road segment data.
[0046] S2 Preprocessing and Feature Extraction: Time synchronization, filtering, and windowing statistics are performed on the data to obtain driving behavior features fd and operating condition features fc.
[0047] S3 Driving Style Recognition: Input fd into the driving style recognition model, and output driving style category c. d and confidence level.
[0048] S4 Working Condition Recognition: Input fc into the working condition recognition model and output the working condition category cc and confidence level.
[0049] S5 State Construction: Merging Vehicle State xvehicle and c d, cc constructs the state vector s=[xvehicle,c d c c ].
[0050] S6 Strategy Generation: Input the state s into the energy consumption optimization strategy model and output the control parameter vector a.
[0051] S7 control execution: Generates executable control commands based on a and applies them to subsystems such as power response, regenerative braking distribution, thermal management, and auxiliary loads.
[0052] S8 Energy Consumption Comparison and Reward Calculation: Calculate energy consumption per unit mile / cycle and compare it with the baseline energy consumption to obtain ΔE; combine safety / comfort / wear indicators to form the reward r.
[0053] S9 Local Update: Performs online or periodic incremental updates to the policy model parameters based on the reward r, and records the model version.
[0054] S10 Federated / Offline Branch: When the network connection and trigger conditions are met, local parameters are uploaded to participate in federated aggregation and receive the global model; when the network connection conditions are not met, offline training is performed based on local data and updates are deployed.
[0055] In step S3, driving style recognition, based on driving operation signals from onboard data sources (CAN / sensors / positioning, etc.), a sliding time window segmentation is used (e.g., 3–10 s, step size 0.5–1 s). After preprocessing the data within each time window (filtering, outlier removal, time alignment), a driving behavior feature fd is constructed. This feature may include: Operation sequence characteristics: timing segments such as accelerator pedal opening / rate of change, brake pedal or brake pressure, vehicle speed, acceleration, jerk, torque request / actual torque, etc. Statistical characteristics: mean, variance, extreme values, kurtosis, number of rising edges, frequency of rapid acceleration / deceleration events, pedal variation, etc. Coupling characteristics: pedal-vehicle speed / torque response relationship, braking-deceleration efficiency, car-following stability, etc.
[0056] The above features can be input into the model in either the form of "statistical vectors" or "multi-channel time series tensors".
[0057] The driving style recognition model in step S3 can be implemented using any of the following structures: 1D-CNN + BiLSTM / GRU + Attention + Fully Connected Softmax; TransformerEncoder + Pooling + Fully Connected Softmax; Statistical feature vectors + multilayer perceptron (MLP) + Softmax (lightweight deployment version).
[0058] The above structure is used to learn style distinction representations from driving behavior characteristics.
[0059] In step S3, the training method for driving style recognition can be supervised learning or semi-supervised learning: Supervised learning: Driving style categories labeled manually or weakly labeled by rules are used as labels, and cross-entropy loss is used for training; Semi-supervised / two-stage: First, cluster the feature embeddings to form a "style prototype", and then fine-tune the classifier with a small amount of labeled data.
[0060] Stability can be improved by using methods such as category balancing and time window voting / smoothing.
[0061] The manual annotation criteria in this embodiment are as follows: manual annotation is based on sliding time window segments (e.g., 5-second window, 1-second step), combined with statistical indicators of driving behavior within the segment for judgment. The annotation criteria adopt three types of indicators: "event intensity + frequency + smoothness". Rapid acceleration event: Appears in the window The duration or frequency of (e.g., 1.0–1.5 m / s²); Rapid deceleration event: Appears in the window The duration or frequency of (e.g., -1.2 to -1.8 m / s²); Jerk indicator: , inside the window or ; Pedal fluctuation: RMS or peak value, number of frequent pedal vibrations; Braking intensity: peak / average braking pressure, braking trigger frequency; Speed stability: speed variance, speed fluctuation amplitude.
[0062] If rule-based weak labeling is used, since the essence of rule-based weak labeling is to coarsely classify samples into "aggressive / stable / economical" categories using measurable indicators to form trainable initial labels, the rules can be divided into rules based on thresholds and event counts, and rules based on comprehensive scores.
[0063] The rules based on thresholds and event counts are defined in units of time windows: Rapid acceleration count ; Rapid deceleration counting ; jerk strength or ; pedal fluctuation .
[0064] Examples of weak annotation rules: like and → Marked as "radical"; like and and → Marked as "stable type"; like Furthermore, if the throttle opening is relatively small and the recovery ratio is high (or energy consumption is low), it is labeled as "Economy".
[0065] Based on the rules of comprehensive scoring, a style score is constructed: ; in It can be the braking intensity (average / peak pressure or braking frequency).
[0066] Weak labeling based on score segments: radical; Economy / Stable; the middle section is stable (or "neutral").
[0067] Weak labeling can provide initial labels that can be generated on a large scale; subsequent corrections / fine-tuning can be done using a small number of high-quality manual labels.
[0068] In step S3, the recommended output format of the driving style recognition model should include: Style category probability distribution (Softmax output, K class probabilities); Style tags ; Confidence level, such as Or distribution entropy, used to determine the credibility of subsequent fusion and control strategies; You can also select style embedding vectors (From the second to last layer), used for "style-condition" fusion and strategy input.
[0069] When the confidence level is insufficient, "unknown / neutral" can be output or majority voting can be used to smooth the style and avoid style jitter.
[0070] The style tags in this scheme are divided into: Aggressive type: strong and frequent acceleration / braking operations, rapid speed changes, and large jerk; Stable type: moderate acceleration / braking operation, small speed fluctuation, and small jerk; Economy model: The pedal opening is smaller, the operation is smoother, and it tends to accelerate smoothly, with higher energy recovery and lower energy consumption per unit distance.
[0071] The boundaries between different styles can be defined using either a "threshold type" or a "quantile type" method. The criteria for threshold-based partitioning are as follows: radical: and ; smooth: and ; Economy: On a stable basis, the average pedal value is smaller, energy consumption is lower, and the recycling rate is higher.
[0072] When using quantile-based partitioning, take historical data... The 70% / 30% quantile is used as It automatically adapts to different vehicle models and operating conditions. For example, based on 5-second window statistics, let's assume... , , , (Example value, actual calibrable); Segment A: Within 5 seconds, there are 3 instances of rapid acceleration and 2 instances of rapid deceleration. → Radical type; Segment B: The number of rapid accelerations / decelerations is 0–1 times. Smooth pedal movement → Stable pedal type; Fragment C: Similarly However, the average pedal energy consumption is lower, the recovery rate is higher, and the energy consumption per unit mileage is significantly lower than the average → Economy type.
[0073] In step S4, the input feature fc of the working condition identification model consists of "vehicle longitudinal state + road slope / resistance + load estimation + optional road segment information", as follows: Vehicle longitudinal status: vehicle speed acceleration Torque / power, braking signal, SOC, current / voltage, temperature, etc.; Slope / Resistance Information: Slope (IMU or map elevation calculation), slope change rate Rolling resistance / adhesion estimation, etc.; Load estimation It can be calculated from suspension pressure / hydraulic signals, or by reverse calculation through longitudinal dynamics and correction in conjunction with loading / unloading operation stages; Optional prediction information: Forward slope sequence from map / high-precision positioning / V2X Or the type of road segment ahead (used to predict operating conditions).
[0074] Similarly, a sliding window approach is used to construct statistical / time series features to ensure adaptability to continuous changes.
[0075] In this embodiment, load estimation The process involves longitudinal dynamics back-calculation combined with corrections during the loading / unloading phase, specifically: 1) The basis of dynamic back-calculation; Simplified longitudinal dynamics of the vehicle: ; in It can be converted from the output torque of the wheel end / motor (e.g.) ), The slope angle is denoted by .
[0076] Instantaneous load estimation is obtained: ; in, To prevent small amounts of the denominator from approaching zero, in actual implementation only... It is activated when a certain threshold is reached or when observable conditions are met (to avoid instability in back-calculation when stationary / uniform).
[0077] 2) Correction in conjunction with the loading / unloading phase; The loading / unloading phase can be identified by work signals or location / geofencing (loading point / unloading point) and triggered as a "quality status update". The correction methods are as follows (choose one of the two): Preferred correction method: event-triggered exponential smoothing correction.
[0078] When a "loading complete" event is detected (or the user leaves after entering a loading point), the mass estimate is shifted to the "heavy load nominal mass". "convergence: ; When the "unloading complete" event is detected, the quality estimate is shifted to the "no-load nominal quality". "convergence: ; in The weights can be calibrated to correct them; this allows for a rapid return to a reasonable quality range when back-calculation noise is high.
[0079] Optional correction method: Bayesian / Kalman filter correction.
[0080] Treating quality as a slowly changing state: ; Observation as : ; When a load / unload event occurs, a "prior mean" is given. or It also reduces prior covariance, enabling rapid correction.
[0081] In step S4 of the working condition identification model, the working condition categories can be defined using a hierarchical combination, including: Load condition: No load / Half load / Heavy load; Slope condition: Uphill / Flat / Downhill; Optional operation stages: loading area / transportation section / unloading area / waiting, etc.
[0082] The final operating condition category is a combination of the above dimensions, such as "heavy load uphill" or "unloaded downhill".
[0083] In step S4, the fusion of slope, load, and vehicle speed in the working condition identification model can be achieved through early fusion or late fusion. Early integration: Directly concatenate the input to MLP / LSTM / Transformer; Late fusion: The embedding / confidence of the load branch and the slope branch are obtained separately, and then gating fusion or attention fusion is used to form a unified working condition representation. .
[0084] The preferred implementation method is "early integration + time-series smoothing", which makes deployment simpler and more robust.
[0085] In step S4 of the working condition identification model, the thresholds for "heavy load" and "uphill" are determined using either "calibrable rules" or "model learning": Example of rule labeling: Uphill: And continue ;downhill: And continue ; Overload: ,in The threshold can be divided according to the ratio of no-load / full-load calibration values (for example, taking 0.6–0.8 of the full-load range), and the specific threshold can be determined by vehicle calibration.
[0086] Model learning: Employing multi-task learning to simultaneously regress continuous variables Output discrete class probabilities The threshold is formed by the training data to create the decision boundary, and there is no need to fix a dead threshold.
[0087] In step S4, the slope is identified in the working condition model. As a continuous input, it is directly fed into the model, while also incorporating the slope change rate. Furthermore, filtering / Kalman filtering or time window majority voting / HMM is used to smooth the operating condition categories, avoiding frequent switching of operating conditions caused by slope boundary jitter.
[0088] Optionally, step S4, the working condition recognition model incorporates information about the road segment ahead (predicted working conditions), obtained through maps / high-precision positioning / V2X. Slope sequence, speed limit, or road type form the predictive working condition characteristics. It is used to adjust control parameters in advance to achieve "predictive operating condition adaptation" and further improve energy consumption optimization and safety.
[0089] Offline training process as follows Figure 3 As shown, the process includes historical data collection, cleaning and labeling, dataset partitioning, training of the driving style classification model, training of the driving condition recognition model, initial reinforcement learning strategy training, and performance evaluation. Once the performance meets the requirements, the model is deployed to the vehicle. If the performance does not meet the requirements, sample supplementation, relabeling, hyperparameter tuning, and retraining iterations are performed.
[0090] Key data structures, action vectors, and reward design.
[0091] (1) State vector s: contains at least the following elements (expandable): Vehicle operating status: vehicle speed v, acceleration a, pedal opening, brake signal, motor torque / power, SOC, battery / motor temperature, etc. Road and working environment: gradient g (from map / IMU), road segment type / operation stage, ambient temperature, etc.; Load / resistance estimation: load mass estimation (m), road resistance rating, etc.; Discrete semantic information: driving style category cd, operating condition category cc (e.g., heavy / unloaded × uphill / downhill).
[0092] (2) Control parameter vector a (action vector): To facilitate on-board execution and interpretability, a multi-dimensional parameter vector form is preferred, as shown in Table 2.
[0093] Table 2 Multidimensional parameter vectors
[0094] (3) Reward r: The reward function for multi-objective constraint optimization is formed by taking energy consumption reduction as the main term and superimposing penalty terms for safety, comfort and wear constraints, as shown in the following formula: r=w1·(-Eunit)+w2·ΔE+w3·Rrecu-λ1·Psafe-λ2·Pcomfort-λ3·Pwear; Wherein, Eunit is the energy consumption per unit mileage or per unit operating cycle, ΔE is the improvement in energy consumption relative to the baseline, Rrecu is the recovery rate, Psafe, Pcomfort, and Pwear are the penalty terms for safety boundary, comfort (such as jerk), and wear (such as the proportion of friction braking), respectively, and w and λ are the weights.
[0095] 1) The process of determining safety boundary penalties (hard constraints take precedence) is as follows: safety-related quantities (battery temperature, motor temperature, bus current / power, SOC lower limit, braking stability index, etc.) are defined as constraints. Over-limit penalties can be applied using ReLU / squared penalties: ; Optional, add a lower bound constraint (such as SOC): ; The larger safety penalty weight is equivalent to a hard constraint; or a "safety projection / shield" can be used to directly project the action into the feasible domain.
[0096] 2) The process for determining the comfort jerk penalty is as follows: jerk definition: ; The jerk metric within the window can be expressed as RMS or peak value: , ; Example of a penalty: ; Or, when the threshold is exceeded Punishment: .
[0097] 3) Wear penalty: Friction braking ratio (can be calculated from braking distribution); The total braking torque (or braking power) consists of regenerative braking and friction braking: ; Friction braking ratio can be defined as energy ratio: ; in It can be estimated from the friction braking torque and wheel speed; or in discrete form, summation can be used instead of integration.
[0098] Wear and tear penalty items: ; Additional penalties can be applied to "high-intensity friction braking events" (such as the duration / number of times the braking pressure exceeds the threshold) to reflect the impact of brake pad thermal decay on lifespan.
[0099] Federated learning aggregation and model distribution. When a vehicle meets the network connectivity and upload trigger conditions (e.g., parking / charging, network availability, accumulated data volume reaching a threshold, or energy consumption improvement exceeding a threshold), the onboard communication unit uploads local model parameters or gradients to the server-side aggregation unit. The server uses aggregation strategies such as FedAvg / FedProx and can combine anomaly detection, gradient pruning, compression, and encryption mechanisms to form a global model. Subsequently, the global model is distributed to the vehicle for updates according to vehicle type / scenario / version strategy.
[0100] To ensure vehicle safety and controllability, a safety fallback mechanism is preferred: when the confidence level of the strategy output is lower than the threshold or when safety boundaries such as temperature / current / braking stability are triggered, the control unit switches to the preset rule control strategy or the previous stable version model; at the same time, version number management and gray-scale updates are performed on the vehicle model to ensure that the update can be rolled back if it fails.
[0101] In the energy consumption optimization strategy model of step S6, the state It consists of vehicle state, environmental state, style / operating condition output, and their uncertainties. Unlike conventional RL control, this scheme incorporates the probability distribution and confidence level of style / operating condition into the state. For example: .
[0102] action The strategy output is a "control parameter vector" rather than a single control variable. Actions may include power distribution, regenerative braking distribution, thermal management / auxiliary load control parameters, and optional prediction weight parameters.
[0103] award It consists of a combination of energy consumption, comfort, safety, wear and tear, and constraint penalties.
[0104] To ensure security and deployability, action boundaries and security projection / fallback mechanisms are introduced: Each dimension of the action has upper and lower limits (determined by system capabilities, regulations, and safety boundaries); if the strategy output may cause constraints such as current, temperature, and braking stability to be violated, the action is corrected to the feasible region through "Action Projection / Safety Shield". When the confidence level is insufficient or anomaly detection is triggered, the system falls back to the rule-based policy / calibration policy to ensure functional safety.
[0105] The energy consumption optimization strategy model is an Actor-Critic structure (MLP or Actor with LSTM handles the observability and hysteresis); the output layer uses Tanh / Clamp to map actions to the legal range; the algorithm is for continuous actions, and SAC / TD3 / PPO, etc. are optional; the initial strategy for training in the offline simulation environment is preferred (vehicle longitudinal dynamics + battery / thermal model + braking recovery model), and small-step online updates or periodic updates are performed after going online.
[0106] The combination of "explicit style / condition input + safety projection + fallback" enables this solution to be more than a simple conventional RL control.
[0107] The "style-condition" approach can employ a "two-level fusion" to simultaneously handle semantics and uncertainty, inputting style probabilities. Operating condition probability Each has its own confidence level; the style and working condition weights are adaptively allocated based on the confidence level, and the side with low confidence level is downweighted or set to "unknown" to avoid misidentification directly affecting the strategy.
[0108] Extracting embedding vectors from style recognition and work condition recognition networks (Second to last layer); The fusion method can be splicing. Gating fusion Or add cross items Expressing coupling relationships; obtaining fusion characterization And construct the final state vector For use by policy networks.
[0109] The above mechanism is a preferred implementation method and is not limited to a single fusion method.
[0110] The weights in the reward function employ a combination of "normalization and multiple determination methods." Each reward item is first normalized in terms of dimensions: energy consumption is based on baseline energy consumption. Normalization: Penalty terms such as current, temperature, and jerk are normalized according to their respective thresholds / upper limits to ensure that the weights are comparable.
[0111] The weights are determined as follows: Method A: Prioritize engineering constraints. Set the weights of safety-related penalties (over-temperature, over-current, braking stability, etc.) to be sufficiently large, making them equivalent to "hard constraints," and then optimize energy consumption and comfort within the feasible domain that satisfies the safety constraints.
[0112] Method B: Data-driven calibration. Based on offline data / simulation, perform grid search or Bayesian optimization on the weights to select the weight combination that satisfies the constraints and is energy-optimal / approximately Pareto optimal.
[0113] Method C: Adaptive Weights. Part of the weights are treated as Lagrange multipliers and adaptively adjusted online based on the degree of constraint violation, achieving constrained reinforcement learning (RL) and avoiding the mismatch of fixed weights under different operating conditions.
[0114] Baseline energy consumption Optional sources include: energy consumption of rule-based / calibrated strategies, moving average of energy consumption of the previous stable version of the strategy, and historical average energy consumption of the same road segment (grouped by load / slope), which are used to calculate energy consumption gain and keep training stable.
[0115] This solution uses deep learning to classify driving behavior by style, and explicitly incorporates the driving differences caused by "changing drivers but not changing vehicles" into the control decision input, thereby reducing energy consumption fluctuations caused by the mismatch between a unified strategy and the behavior of multiple drivers.
[0116] This solution uses operating condition recognition and style-operating condition fusion modeling to enable the strategy to adaptively adjust the power response and recovery intensity for different scenarios such as heavy-load uphill, downhill recovery, and empty transport, thereby achieving a more stable energy consumption optimization effect.
[0117] This solution employs a closed-loop mechanism of "energy consumption comparison before and after optimization - reward feedback - strategy correction" to directly feed the actual operating results back to the strategy update stage, reducing reliance on manual calibration and improving long-term adaptive and generalization capabilities.
[0118] This solution uses a federated learning aggregation and distribution mechanism to achieve knowledge sharing across vehicles / drivers without centralized uploading of raw data, thereby improving model convergence and generalization. It also supports offline training and deployment to meet the needs of continuous updates under weak network / network outage conditions.
[0119] This solution's strategy outputs parameters in the form of a vector and applies them to multiple subsystems, including power, recovery, thermal management, and auxiliary loads. While reducing energy consumption, it also takes into account comprehensive technical effects such as safety boundaries, comfort, and brake wear.
[0120] Online closed-loop energy consumption optimization control (corresponding to) Figure 1 , Figure 2 During vehicle operation, CAN signals and environmental / road segment information are collected at a preset sampling frequency; fd and fc are obtained after preprocessing and feature extraction; driving style category cd and operating condition category cc are output respectively; after fusing to form state s, the deep reinforcement learning strategy model outputs control parameter vector a; a is mapped to power response, regenerative braking distribution, thermal management and auxiliary load control commands and executed; energy consumption and performance indicators after execution are collected, reward r is calculated and the local model is updated; when the network connection and trigger conditions are met, the data is uploaded and participates in federated aggregation, and the global model is received and distributed to complete the vehicle-side update.
[0121] Offline training and deployment (corresponding) Figure 3Collect historical driving data from multiple drivers and under multiple working conditions, clean and label the data, and divide it into training and testing sets; train driving style classification models and working condition recognition models; train an initial reinforcement learning strategy based on the training environment and evaluate its performance; once the performance meets the target, generate a deployment model and distribute it to vehicles; if the performance does not meet the target, perform sample supplementation / relabeling / parameter tuning and iterative training until the target is met.
[0122] Safety fallback and rollback (optional). Set safety boundary thresholds and policy confidence thresholds; when the output is abnormal or the boundary is triggered, immediately roll back to the rule policy or the previous stable model version; at the same time, record the trigger log for offline analysis and parameter correction to improve the reliability of subsequent models.
[0123] The above are merely preferred embodiments of this solution and are not intended to limit the solution. Various modifications and variations can be made to this solution by those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this solution should be included within the scope of protection of this solution.
Claims
1. A deep learning-based energy consumption optimization control method for a pure electric wide-body vehicle, characterized in that, Includes the following steps: Vehicle operation data, environmental and road segment data are collected, and after preprocessing and feature extraction, driving behavior characteristics and operating condition characteristics are obtained; The driving behavior features are processed by the first deep learning model, which outputs the driving style category and its confidence level to represent the current driver's operating habits. The operating condition features are processed by a second deep learning model, which outputs the operating condition category and its confidence level to characterize the current operating scenario of the vehicle. By integrating real-time vehicle status information, driving style category and confidence level, and operating condition category and confidence level, a comprehensive state vector representing the coupling relationship between "human-vehicle-road" is constructed. The combined state vector is input into a pre-trained deep reinforcement learning model, and the output is a control parameter vector for co-optimizing multiple subsystems. The obtained control parameter vector is mapped into executable control commands, which are then sent to the dynamic response subsystem, regenerative braking distribution subsystem, thermal management subsystem, and auxiliary load subsystem for execution.
2. The deep learning-based energy consumption optimization control method for a pure electric wide-body vehicle according to claim 1, characterized in that, During the execution of control commands, the actual energy consumption and performance indicators within the current control cycle are calculated. A reward signal is generated by comparing the energy consumption with the baseline energy consumption, and the parameters of the deep reinforcement learning model are updated based on the reward signal.
3. The deep learning-based energy consumption optimization control method for a pure electric wide-body vehicle according to claim 2, characterized in that, The reward signal is calculated using the following formula: r=w1·(-Eunit)+w2·ΔE+w3·Rrecu-λ1·Psafe-λ2·Pcomfort-λ3·Pwear; Wherein, Eunit is the energy consumption per unit mileage or per unit operation cycle, ΔE is the improvement in energy consumption relative to the baseline, Rrecu is the energy recovery rate, Psafe, Pcomfort, and Pwear are the safety boundary penalty, comfort penalty, and wear penalty, respectively; the baseline energy consumption is derived from the rule-based strategy energy consumption, the moving average of the previous stable version strategy energy consumption, or the historical average energy consumption of the same road segment; the weights w1, w2, w3, λ1, λ2, and λ3 are determined through engineering constraint priority calibration, data-driven offline optimization, or adaptive online adjustment.
4. The deep learning-based energy consumption optimization control method for a pure electric wide-body vehicle according to claim 1, wherein, Driving behavior characteristics are constructed using a sliding time window and include at least one of the following categories of characteristics: Operation sequence characteristics include at least the timing segments of accelerator pedal opening, brake pedal pressure, vehicle speed, acceleration, and torque request; Statistical characteristics, including at least one of the following: mean, variance, extreme values, kurtosis, frequency of rapid acceleration events, and frequency of rapid deceleration events; Coupling characteristics include at least one of pedal-vehicle speed response relationship and braking-deceleration efficiency.
5. The energy consumption optimization control method for pure electric wide-body vehicles based on deep learning as described in claim 1, characterized in that, The first deep learning model is implemented using any of the following architectures: Network architecture combining 1D-CNN with BiLSTM / GRU attention mechanism; Transformer encoder network structure combined with pooling layers; A lightweight network structure combining a multilayer perceptron with a Softmax output layer; The output of the first deep learning model includes the probability distribution of driving style categories, driving style labels, confidence scores, and style embedding vectors.
6. The deep learning-based energy consumption optimization control method for pure electric wide-body vehicles according to claim 1, wherein, Operating condition characteristics include at least one of the following types of information: Vehicle longitudinal status information, including at least one of vehicle speed, acceleration, motor torque, power, SOC, battery temperature, and motor temperature; Road gradient and resistance information, including at least one of real-time gradient, gradient change rate, and road rolling resistance estimation; Load estimation information is obtained by combining suspension pressure signals or longitudinal dynamic back-propagation with operational phase correction. Forward prediction information includes forward slope sequences or forward road segment types from maps, high-precision positioning, or V2X.
7. The deep learning-based energy consumption optimization control method for pure electric wide-body vehicles according to claim 1, wherein, The working condition categories are defined using a hierarchical combination method, including combinations of load conditions and slope conditions; The second deep learning model takes continuous slope values as input, combines the slope change rate and filtering smoothing mechanism to handle continuous slope changes, and outputs the probability distribution of working condition categories, working condition labels, confidence scores, and working condition embedding vectors.
8. The deep learning-based energy consumption optimization control method for pure electric wide-body vehicles according to claim 1, wherein, By integrating real-time vehicle status information, driving style category and confidence level, and operating condition category and confidence level, a comprehensive state vector is constructed, including: The driving style and operating condition are adaptively weighted according to the confidence level. When the confidence level is lower than the preset threshold, the corresponding category is set to an unknown state. Extract the style embedding vector of the first deep learning model and the working condition embedding vector of the second deep learning model, and generate a fused representation by concatenation, gated fusion or feature crossing. The vehicle's real-time status information, fused representation, driving style probability distribution, and operating condition probability distribution are combined to construct a comprehensive state vector.
9. The deep learning-based energy consumption optimization control method for pure electric wide-body vehicles according to claim 1, wherein, The control parameter vector shall include at least one or more of the following parameters: Torque response coefficient, used to correct the mapping relationship between pedal opening and motor requested torque; The regenerative braking intensity coefficient is used to determine the proportion allocated to regenerative braking during braking. Thermal management operating point factor, used to adjust the operating point of cooling pumps, fans, or heating strategies; Auxiliary load power factor, used to limit the upper limit of power for air conditioning or hydraulic systems; Predictive control activation level coefficient is used to adjust energy management gain in advance based on information about the road ahead.
10. The energy consumption optimization control system based on deep learning for pure electric wide-body vehicle, for realizing the energy consumption optimization control method based on deep learning for pure electric wide-body vehicle as claimed in claim 1, characterized in that, include; The data acquisition and feature extraction module is configured to: collect vehicle operation data, environmental and road segment data, and obtain driving behavior features and operating condition features through preprocessing and feature extraction; The driving style recognition module is configured to: pass driving behavior features through a first deep learning model to output a driving style category and its confidence level that characterizes the current driver's operating habits; The operating condition identification module is configured to: use the operating condition features through a second deep learning model to output the operating condition category and its confidence level, which characterize the current operating scenario of the vehicle; The state construction module is configured to: integrate real-time vehicle state information, driving style category and confidence level, and operating condition category and confidence level to construct a comprehensive state vector representing the coupling relationship between "human-vehicle-road"; The policy generation module is configured to: take the state vector as input to a pre-trained deep reinforcement learning model and output a control parameter vector for co-optimizing multiple subsystems; The control execution module is configured to map the obtained control parameter vector into executable control commands, which are then sent to the power response subsystem, regenerative braking distribution subsystem, thermal management subsystem, and auxiliary load subsystem for execution.