New energy vehicle soc variation estimation method integrated with multi-time window model

CN122815201APending Publication Date: 2026-09-25CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610980160.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-02
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0006]本发明旨在解决现有方法中整数SOC标签失真、多尺度片段预测精度波动、短时序输入无法预测及工况适应性差的问题,提供一种集成多时间窗口模型的新能源汽车SOC变化量估计方法

Benefits of technology

本发明通过基于能耗与平均速度的K-Means工况聚类和分层XGBoost/LightGBM残差级联建模,实现了对不同驾驶工况的差异化准确预测,克服了单一全局模型难以适应高速巡航、持续爬坡、城市拥堵等差异化工况的问题,有效减小了因驾驶工况差异和电池老化等因素带来的预测偏差。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122815201A_ABST
    Figure CN122815201A_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of new energy automobile battery management, and specifically relates to a new energy automobile SOC change amount estimation method integrated with a multi-time window model. In view of the problems of integer SOC label distortion, multi-scale segment prediction accuracy fluctuation, short time sequence input unable to be predicted and poor working condition adaptability in the existing method, the following scheme is provided: performing SOC fine processing on real vehicle operation data and constructing 10-minute and 60-minute standard segment data sets; performing working condition clustering based on energy consumption and average speed, and constructing a cascade structure of an XGBoost basic predictor and a LightGBM residual correction model; adapting an arbitrary length segment to the standard model prediction through four adaptive strategies of direct, copy, disassemble and expand; and extrapolating complete segment global statistical characteristics from the previous 10-minute time sequence data by a bidirectional LSTM network, inputting the cascade model and outputting the SOC change amount. The present application is used for new energy automobile cloud SOC change amount estimation, and realizes high-precision prediction of driving segments of arbitrary length from several minutes to more than ten hours.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of new energy vehicle battery management technology, specifically relating to a method for estimating the change in SOC of new energy vehicles by integrating a multi-time window model. Background Technology

[0002] In the actual operation of new energy vehicles, the change in battery state of charge (SOC) is a crucial parameter for vehicle energy consumption assessment, range prediction, trip planning, and cloud-based battery status analysis. Especially in cloud-based analysis scenarios based on vehicle-to-everything (V2X) and real-vehicle operation data, accurately estimating the SOC change within any driving segment can help improve the real-time performance and reliability of range prediction, reducing users' range anxiety.

[0003] Existing technologies utilize historical vehicle network operating data, vehicle operating condition information, environmental information, and machine learning models to predict the energy consumption or remaining driving range of electric vehicles. For example, some solutions collect historical vehicle operating data and combine it with operating condition prediction models and machine learning energy consumption prediction models to predict future energy consumption; others extract vehicle time-series features using models such as LSTM, XGBoost, TCN, or BiLSTM to achieve short-term energy consumption or remaining range prediction. These solutions improve the automation level of electric vehicle energy consumption prediction to some extent.

[0004] However, in the application of real-world vehicle operation data, existing solutions still have the following shortcomings: First, the SOC data uploaded by the vehicle is usually recorded in the form of integer percentages. Within a short driving segment, the actual SOC change may be less than one percentage point, resulting in the recorded results showing zero change or a step change, which in turn causes distortion of the training labels. Second, existing models are usually built for fixed time windows, fixed travel segments, or fixed prediction scales, making it difficult to adapt to multi-scale driving segments ranging from a few minutes to more than ten hours at the same time. The prediction accuracy is prone to significant fluctuations with the length of the segment. Third, most existing models rely on the statistical characteristics of the complete driving segment as input. When only short-term data of the first part of the segment is obtained, it is difficult to predict the SOC change of the entire journey in advance. Fourth, a single global model is difficult to fully adapt to differentiated driving conditions such as high-speed cruising, urban congestion, and continuous uphill climbing, resulting in insufficient stability of cross-condition prediction.

[0005] Therefore, there is an urgent need in this field to provide a method for estimating the SOC change of new energy vehicles that can be applied to real vehicle cloud data, adapt to different driving conditions and time scales, and estimate the SOC change of a complete segment when only short-term initial data is acquired, so as to improve the accuracy, generalization ability and real-time application value of SOC change estimation under driving segments of arbitrary length. Summary of the Invention

[0006] This invention aims to solve the problems of integer SOC label distortion, multi-scale segment prediction accuracy fluctuation, unpredictability of short time series inputs, and poor adaptability to operating conditions in existing methods, and provides a method for estimating the SOC change of new energy vehicles by integrating a multi-time window model.

[0007] To achieve the above objectives, the present invention adopts the following technical solution: A method for estimating the state of charge (SOC) change of new energy vehicles by integrating a multi-time window model includes the following steps: We acquire real-vehicle operation data of new energy vehicles, preprocess and refine the acquired real-vehicle operation data using SOC, and construct two standard scale segment datasets containing 10-minute segments and 60-minute segments. Feature extraction is performed on each segment in the standard-scale segment dataset to obtain multidimensional statistical features; Based on the aforementioned multidimensional statistical characteristics, a working condition clustering and hierarchical residual integrated modeling strategy is adopted to construct a prediction model for the SOC change of each working condition category. The hierarchical residual ensemble model includes a base predictor and a residual correction model; During prediction, the working condition category of the segment to be predicted is first determined, the basic predictor of the corresponding working condition category is called to output the initial prediction value, and then the initial prediction value is corrected by the residual correction model to obtain the predicted value of SOC change under the working condition. Based on the actual duration of the segment to be predicted, a multi-time window adaptive fusion strategy is adopted to adapt the segment to be predicted of any duration to the standard scale model corresponding to the standard scale segment dataset through expansion or decomposition. The prediction results of each segment are then fused to obtain the final SOC change.

[0008] Furthermore, the SOC refinement process includes: Select a charging segment to calibrate the actual total battery capacity; When the SOC value jumps to an integer value, the last integer value data point before the jump is used as the reference point. Based on the SOC integer value, cumulative discharge capacity, and actual total capacity of the battery at the reference point, each data point in the discharge segment is refined point by point to obtain a continuous SOC value.

[0009] Furthermore, the construction of the standard-scale fragment dataset includes: The original vehicle driving data is divided into basic time windows of fixed duration. Extract the dynamic features of each basic time window, perform dimensionality reduction on the dynamic features, and divide the high-dynamic class and low-dynamic class based on the first principal component score; When generating target scale segments using a sliding window, the sliding step size is adaptively adjusted based on the proportion of high dynamic windows within the candidate segments: overlapping movement is used when the proportion of high dynamic windows is not less than a preset threshold, otherwise non-overlapping movement is used. The SOC change of each segment is calculated as the supervised training label.

[0010] Furthermore, the dynamic characteristics of the basic time window include velocity variance, current variance, acceleration variance, and velocity amplitude; the dimensionality reduction process uses principal component analysis, with the mean minus one standard deviation as the classification threshold between the high-dynamic class and the low-dynamic class.

[0011] Furthermore, the integrated modeling of working condition clustering and hierarchical residuals includes: Based on the two dimensions of energy consumption per unit mileage and average speed, K-Means clustering was used to divide all segment samples into several typical driving condition categories, including high-speed cruising, continuous uphill climbing, and urban congestion. XGBoost base predictors are trained independently for each type of work condition, and the multidimensional statistical features are used as input to output the initial predicted value of SOC change. The LightGBM model is used to learn the residual between the initial predicted value and the true value, and the residual correction value is output. The initial predicted value is added to the residual correction value to obtain the final predicted value of SOC change under this operating condition.

[0012] Furthermore, for 10-minute scale segments, no working condition stratification is introduced; instead, a unified single model is used to process the prediction of all 10-minute scale segments.

[0013] Furthermore, the multi-time-window adaptive integration strategy includes: The duration of the segment to be predicted is determined. If the duration of the segment to be predicted is within the preset standard duration range, the corresponding standard scale model is directly input to obtain the predicted value of SOC change. If the duration of the segment to be predicted is greater than the first threshold and less than the second threshold, the segment to be predicted is split into multiple sub-segments, each sub-segment is input into the corresponding standard scale model, and the prediction results are summed. If the duration of the segment to be predicted is within the third threshold range, the segment to be predicted will be extended to the standard duration through temporal self-copying, and then scaled after being input into the corresponding standard scale model. If the duration of the segment to be predicted is within the fourth threshold range, data is extracted from the beginning of the segment to be predicted and spliced ​​to the end to form a standard duration input, wherein the voltage data is regenerated according to the linear decay formula; If the duration of the segment to be predicted exceeds the fifth threshold, the standard duration portion is cut out, and the remaining portion is returned to the adaptive integration strategy for recursive processing.

[0014] Furthermore, it also includes a short-time global feature extrapolation step: Using multidimensional time-series data with a preset duration before the segment as input, the temporal features are extracted through a bidirectional long short-term memory network (BiLSTM), and mapped to the estimated values ​​of the multidimensional statistical features of the complete segment through a fully connected layer. The estimated values ​​are then input into the hierarchical residual ensemble model, and the predicted value of SOC change is output.

[0015] Furthermore, for segments with large changes exceeding a preset duration threshold, when the initially predicted SOC change exceeds the preset threshold, the complete segment is divided into several sub-segments smaller than the preset duration threshold. For each sub-segment, the feature extrapolation of the bidirectional long short-term memory network and the cascaded prediction of the hierarchical residual ensemble model are performed, and the prediction results of each sub-segment are accumulated.

[0016] Furthermore, the method is deployed on a cloud server or an in-vehicle terminal; the in-vehicle terminal collects real-vehicle operating data and executes the method via the in-vehicle CAN bus.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention achieves differentiated and accurate prediction of different driving conditions by using K-Means clustering based on energy consumption and average speed and hierarchical XGBoost / LightGBM residual cascade modeling. It overcomes the problem that a single global model is difficult to adapt to different driving conditions such as high-speed cruising, continuous uphill climbing, and urban congestion, and effectively reduces prediction bias caused by differences in driving conditions and battery aging.

[0018] This invention designs four adaptive segment processing strategies—direct, copy, disassemble, and expand—and combines them with recursion and boundary fusion mechanisms to enable driving segments of any length, from a few minutes to more than ten hours, to be intelligently adapted to a standard-scale sub-model for accurate estimation, thereby improving multi-timescale compatibility.

[0019] This invention fills the technical gap of short-term input and long-term prediction by constructing a bidirectional LSTM network to extrapolate the global statistical features of the complete segment from the first 10 minutes of time-series data. This enables the system to estimate the SOC change of the complete journey in advance under the condition of only acquiring short-term time-series data of the first segment, thus improving the real-time performance and practicality of the application. Attached Figure Description

[0020] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will now be described in detail with reference to the accompanying drawings, wherein... Figure 1 A flowchart of the method for estimating the SOC change of new energy vehicles using an integrated multi-time window model provided in an embodiment of the present invention; Figure 2 This is a comparison chart of the effects of SOC fine-tuning in the embodiments of the present invention; Figure 3 Flowchart for constructing the 10-minute segment dataset and the 60-minute segment dataset provided in this embodiment of the invention; Figure 4 This is a flowchart of the working condition clustering and hierarchical modeling process in an embodiment of the present invention; Figure 5 This is a schematic diagram of the adaptive integration strategy in an embodiment of the present invention; Figure 6 This is a diagram of the BiLSTM network structure for short-time global feature extrapolation in an embodiment of the present invention; Figure 7 This is a prediction result diagram of a coupled model in an embodiment of the present invention that does not use the multi-time window concept; Figure 8 This is a graph showing the SOC change estimation results using an integrated multi-time window model in an embodiment of the present invention. Detailed Implementation

[0021] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0022] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual pictures, and should not be construed as limiting the invention. To better illustrate the embodiments of the invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.

[0023] Example 1 This embodiment provides a method for estimating the state of charge (SOC) change of new energy vehicles by integrating a multi-time window model. See also... Figure 1The method includes the following steps: data acquisition and preprocessing, dataset construction, feature extraction, operating condition clustering and hierarchical residual ensemble modeling, multi-time window adaptive ensemble prediction, and short-time-series global feature extrapolation and prediction. This method can be deployed on a cloud server, which communicates with the in-vehicle terminal to acquire real-world vehicle operation data uploaded by the in-vehicle terminal.

[0024] S1. Data acquisition and SOC fine-tuning.

[0025] Real-world operational data of new energy vehicles is acquired through a vehicle-to-everything (V2X) platform. This data is collected based on the national standard GB / T32960 and includes fields such as vehicle status, charging status, vehicle speed, total current, total voltage, state of charge (SOC), and cumulative mileage. The real-world operational data covers the actual operation of multiple pure electric passenger vehicles, encompassing driving and charging statuses.

[0026] Data preprocessing includes the following operations: First, the raw data is sorted chronologically to ensure that the data collected from the same vehicle are correctly ordered in chronological order; second, outliers (such as jump values ​​caused by sensor failure or communication interruption) and alarm data are filtered and removed; third, the data multipliers caused by storage compression or transmission truncation in the raw data are restored, that is, the scaled values ​​are restored to the original dimensions to ensure the consistency of the values ​​of each physical quantity.

[0027] Because the SOC values ​​uploaded from the vehicle to the cloud are integer percentages, the actual SOC change within short time segments is often less than one percentage point. After rounding or down-rounding, the change is recorded as zero, leading to distorted model training labels and making the data unsuitable for direct data-driven fine-grained analysis. To address this issue, this step first performs SOC fine-grained processing, specifically including: S1-1, Actual total battery capacity calibration.

[0028] Select a charging segment of the vehicle, choosing a charging process with a State of Charge (SOC) greater than 60% to ensure data stability. Utilize the accumulated charge and SOC changes during the charging process to calculate the actual total battery capacity of the vehicle in its current state. .

[0029]

[0030] in, To accumulate power for charging, This represents the change in SOC.

[0031] S1-2, Refinement of discharge segments point by point.

[0032] When the SOC value undergoes an integer jump, such as from 70% to 69%, the last data point before the jump, which still holds the integer value before the jump, is considered relatively accurate. This point is recorded as the reference point int, and its corresponding cumulative discharge capacity is also recorded. and SOC integer value Based on the above benchmarks, the refined continuous SOC value of any point x within the segment is calculated using the following formula:

[0033] in, The refined SOC value for point x; The cumulative discharge capacity at point x is obtained by integrating the discharge current prior to that point using ampere-hours. The cumulative discharge capacity at the reference point int; This refers to the actual total capacity of the battery as specified above.

[0034] See Figure 2 The image shows a comparison of the SOC refinement process before and after this step. The refined SOC value exhibits a continuous change characteristic, eliminating the step-by-step change problem of the original integer SOC.

[0035] S2, Construction of standard-scale fragment dataset.

[0036] Based on refined SOC labels, two standard-scale segment datasets were constructed: a 10-minute segment dataset and a 60-minute segment dataset. Analysis of real-vehicle data revealed a highly uneven distribution of original driving segment lengths: a large number of segments were concentrated in the short duration range of around 10 minutes, while samples of medium-to-long-duration segments (such as long-distance driving exceeding 1 hour or even 12 hours) were scarce, resulting in a severely skewed distribution of segment lengths in the training samples. To address this imbalanced data distribution, this step employed a sliding window sampling and adaptive augmentation strategy to balance the number of samples across each scale. See [link to relevant documentation]. Figure 3 This step specifically includes: S2-1, Basic Window Division.

[0037] The raw vehicle driving data is divided into basic time windows of fixed duration. For a 60-minute scale, each 60-minute segment needs to be composed of 12 consecutive 5-minute basic windows, i.e., 12 segments × 5 minutes / segment = 60 minutes; for a 10-minute scale, each 10-minute segment needs to be composed of 5 consecutive 2-minute basic windows, i.e., 5 segments × 2 minutes / segment = 10 minutes.

[0038] S2-2, Basic Window Dynamic Feature Extraction.

[0039] For each basic window, the following four statistical features are extracted to measure driving dynamics: velocity variance (V) std The standard deviation of vehicle speed within a window reflects the degree of speed fluctuation. The larger the speed variance, the more drastic the speed changes within that window.

[0040] Current variance (I) std ): The standard deviation of the total current within the window reflects the degree of dynamic change in the motor load.

[0041] Acceleration variance (Acc) std ): The standard deviation of acceleration within the window reflects the severity of acceleration / deceleration behavior.

[0042] Velocity amplitude (V) abs ): The absolute value of the difference between the maximum and minimum speeds within the window, i.e. This reflects the range of speed changes.

[0043] The formula for calculating the integral of acceleration is as follows:

[0044] Used to characterize the cumulative magnitude of acceleration within a window; The formula for calculating velocity variance is:

[0045] Used to characterize the degree of fluctuation of each variable (velocity, current, acceleration) within the window.

[0046] S2-3, PCA dimensionality reduction and dynamic characteristic classification.

[0047] Principal component analysis (PCA) is used to analyze the high-dimensional space formed by the above four dynamic features. Dimensionality reduction was performed. The score of the first principal component was used as a comprehensive quantitative indicator of the dynamic characteristics of the basic window. Based on the PCA score distribution of the entire sample, the dimensionality was reduced by one standard deviation from the mean. The basic window is divided into the following categories, with the classification threshold being: High dynamic range: This corresponds to a window where driving behavior is intense and speed current fluctuates frequently; Low dynamic class: This corresponds to a window where driving behavior is smooth and the speed and current are stable.

[0048] The distribution of PCA scores exhibits a bimodal characteristic, indicating that the high-dynamic and low-dynamic windows are statistically distinct.

[0049] S2-4, Adaptive sliding window expansion.

[0050] When generating dataset segments at the target scale (10 minutes or 60 minutes) using a sliding window, the proportion of highly dynamic basic windows within the candidate segments is first calculated, and the sliding step size is adaptively adjusted based on this proportion. When the proportion of high dynamic window is ≥0.5, overlapping movement is adopted, that is, the sliding step is smaller than the window size, to achieve dense expansion and increase the number of high dynamic training samples under this segment type; When the proportion of high dynamic windows is less than 0.5, non-overlapping movement is adopted, that is, the sliding step size is equal to the window size, to achieve sparse expansion and avoid excessive redundancy of low dynamic samples.

[0051] By employing the aforementioned adaptive sliding strategy, the distribution ratio of samples with varying dynamic characteristics in the training set is balanced, thereby improving the model's ability to generalize to different dynamic characteristics.

[0052] S2-5, Fragment tag generation.

[0053] For each generated segment, the change in SOC ΔSOC is calculated based on the refined SOC value and used as a supervised training label. At the same time, the number of high dynamic windows and the number of large change windows in the segment are counted and recorded for segment characteristic judgment in subsequent steps.

[0054] S3. Feature extraction and filtering.

[0055] Statistical features covering four dimensions—current, voltage, temperature, and driving behavior—were extracted from each segment, initially constructing a full feature set containing 43 features. In the feature selection stage, a triple-screening mechanism was employed, combining correlation analysis, model feature weight evaluation, and manual identification, ultimately retaining 16 core features from the 43 initial selections. This feature selection process fully considered various influencing factors of real-vehicle SOC, including driving conditions and battery pack inconsistencies, and based on multiple screening principles, ensured that the selected features had sufficient explanatory power and predictive contribution to SOC changes. These 16 core features include: Driving behavior characteristics: mileage, number of rapid accelerations, proportion of high-speed time, average speed, and rate of change of speed; Battery state characteristics: discharge current time integral, net current time integral, average current; voltage-related: voltage standard deviation, minimum voltage; power-related: net power integral, average power; energy-related: discharge energy; temperature-related: average temperature, average temperature difference. Battery cell characteristics: number of cells in a single cell.

[0056] S4. Working condition clustering and hierarchical residual integrated modeling.

[0057] To overcome the challenge of adapting to diverse driving conditions with a single global model, this step, after comparing and analyzing various modeling approaches (including a globally unified model, independent models for different driving conditions, ensemble learning models, and cascaded residual correction models), ultimately adopts a driving condition clustering and hierarchical residual ensemble modeling strategy based on energy consumption and average speed to establish a mapping relationship between the 16-dimensional features of a complete segment and the change in SOC. This scheme comprehensively considers the impact of driving condition differences and battery aging and degradation on the estimation of SOC changes, improving prediction robustness through a dual mechanism of driving condition hierarchies and residual cascades. See also Figure 4 Specifically, it includes the following sub-steps: S4-1, Working Condition Clustering.

[0058] On a segment basis, two dimensions of features are extracted: energy consumption per unit distance (kWh / km) and average speed (km / h). Taking a 60-minute time window as an example, based on this two-dimensional feature space, the K-Means clustering algorithm is used to divide all segment samples into three typical driving conditions with significant differences. Their distribution characteristics in the energy consumption-average speed two-dimensional feature space are as follows: (a) High-speed cruising condition: The average speed is relatively high and the energy consumption is moderate, which is suitable for smooth driving scenarios such as highways; (b) Continuous climbing condition: energy consumption is significantly higher, average speed is moderate or lower, corresponding to mountainous or long slope driving scenarios; (c) Urban congestion conditions: lower average speed and lower energy consumption, corresponding to driving scenarios with frequent starts and stops in urban areas.

[0059] The three types of operating conditions exhibit clear clustering and separation characteristics in the two-dimensional feature space, supporting the effectiveness of subsequent hierarchical modeling. For the 10-minute scale segment, due to the small time scale, small SOC variation, and no obvious regular differences between segments, operating condition stratification can be omitted, and a unified model can be used for processing.

[0060] S4-2, Hierarchical XGBoost Basic Prediction.

[0061] After stratifying by work condition, the XGBoost base predictor is trained independently for the training samples under each work condition. The XGBoost model for each work condition takes the 16-dimensional statistical features selected by S3 as input and outputs the initial predicted value ΔSOC of the change in SOC under that work condition. pseudo .

[0062] S4-3, LightGBM residual correction.

[0063] The initial predicted value ΔSOC output by the XGBoost base predictor pseudo Compare the error with the true value ΔSOCtrue and calculate the residual Error = ΔSOC. true ΔSOC pseudo The LightGBM model is used to learn the mapping relationship between the residual error and the 16-dimensional input features, and the residual correction value ΔSOC is output. residual .

[0064] S4-4, Cascaded Prediction.

[0065] During online prediction, the segment to be predicted is first classified into its operating condition category. The XGBoost model and LightGBM model for the corresponding operating condition category are then called, and their outputs are summed to obtain the predicted final SOC change value for that operating condition: ΔSOC. final =ΔSOC pseudo +ΔSOC residual Among them, ΔSOC pseudo The initial prediction value, ΔSOC, is the output of the XGBoost base predictor. residual This refers to the residual correction value output by the LightGBM residual correction model.

[0066] Through the cascaded structure of XGBoost and LightGBM described above, the base predictor is responsible for capturing the main mapping relationships within the driving conditions, while the residual correction model is responsible for compensating for the systematic biases of the base predictor. Together, they achieve a more accurate estimate of SOC changes. In addition, this step effectively reduces prediction biases caused by differences in driving conditions and battery aging by using driving condition stratification (modeling data separately under different driving conditions). The independent model for each driving condition only needs to learn the data distribution characteristics within that condition, avoiding mutual interference between data from different driving conditions in the global model and performance degradation caused by battery state decay over time.

[0067] S4-5, 10-minute model processing.

[0068] For 10-minute segments, due to their small time scale (only 10 minutes), the change in SOC within the segment is usually small and the difference in SOC change patterns under different working conditions is not significant. Therefore, the 10-minute model does not introduce working condition stratification and uses a unified single XGBoost model to process the prediction of all 10-minute segment segments in order to reduce model complexity and avoid overfitting.

[0069] In another implementation, the base predictor and residual correction model can also employ other gradient boosting algorithms or ensemble learning methods, such as CatBoost, gradient boosting decision trees, etc., to replace XGBoost and LightGBM, as long as the cascaded structure of base prediction and residual correction can be achieved. Furthermore, for special operating conditions, such as dedicated driving scenarios for new energy buses or logistics vehicles, the number of cluster categories can be increased accordingly to adapt to finer-grained operating condition classifications.

[0070] S5, multi-time-window adaptive ensemble prediction.

[0071] To overcome the limitations of single models in estimating SOC changes in time segments of arbitrary length, which only perform well around 10 and 60 minutes, this step develops a multi-window model ensemble algorithm based on an adaptive strategy. See [link to relevant documentation]. Figure 5 This algorithm recursively calls hierarchical residual ensemble models trained on S4 at two scales: 10 minutes and 60 minutes. It intelligently expands or decomposes segments of arbitrary length into a series of standard segments of approximately 10 minutes and 60 minutes, performs accurate estimations on each segment, and then fuses the results to calculate the final SOC change of the target segment. The specific strategy is as follows: (a) Direct model: For segments with a duration of 10 minutes or 55-65 minutes, since their length matches the standard scale model, no additional processing is required. The corresponding 10-minute model or 60-minute model can be directly input to obtain the predicted value of SOC change.

[0072] (b) Decomposition model: For segments with a duration of 10 minutes to 27.5 minutes (excluding the 10-minute and 55-65-minute intervals), they are divided into multiple segments with no overlap according to a 10-minute sliding window. Each segment is fed into the 10-minute model in sequence, and the SOC increment of each segment is obtained and summed to obtain the SOC change of the original segment.

[0073] (c) Replication Model: For segments with a duration of 27.5-32.5 minutes, the segments are extended to about 60 minutes by time-domain self-replication. For example, the segment data is repeatedly spliced ​​on the time axis, and after being input into the 60-minute model, it is scaled by 0.5 times as the predicted value of the SOC change of the original segment.

[0074] (d) Augmentation Model: For segments with a duration of 30-60 minutes (excluding the 55-65 minute interval), the required augmentation data is extracted from the beginning of the segment and spliced ​​to the end to form a complete 60-minute input. During the augmentation process, the voltage data is regenerated according to the linear decay formula, which is:

[0075]

[0076] in, To expand the regenerated voltage value corresponding to time t; , These are the voltage values ​​at the end and beginning of the original segment, respectively; This is the voltage linear attenuation coefficient calculated based on the voltage changes at the beginning and end of the original segment; To expand any point in the segment, This is the end time of the original segment; The duration of the original segment.

[0077] This simulates the natural monotonic decrease in voltage during battery discharge. If the expanded segment still does not meet the conditions of the direct model, the adaptive strategy is returned to continue recursive processing.

[0078] (e) Recursive processing of ultra-long segments: For ultra-long segments with a duration of more than 65 minutes, first cut out 60 minutes of the segment and process them using the above strategy. The remaining segments are returned to the adaptive strategy for recursive processing until all segments meet the above strategy conditions.

[0079] S6. Short-time global feature extrapolation and prediction.

[0080] To address the issue in real-world travel scenarios where only the first 10 minutes of time-series data are available, making it impossible to directly calculate the statistical features of the complete segment, this step constructs a Bidirectional Long Short-Term Memory (BiLSTM) network to extrapolate from local time-series data to global features, and cascades it with a base prediction model to complete long-term ΔSOC prediction. The BiLSTM network consists of two Bidirectional LSTM layers. The first layer has 128 hidden units and is used to extract local dynamic change features from the first 10 minutes of time-series data. The second layer has 64 hidden units and is used to form a global representation of the time-series data, which is then mapped to the global statistical features of the complete segment via a fully connected layer. See also... Figure 6 Specifically, it includes the following sub-steps: S6-1, BiLSTM feature extrapolation.

[0081] Using the first 10 minutes of multidimensional time-series data as input, a bidirectional LSTM network is constructed to extract time-series features. This multidimensional time-series data includes raw time-series signals such as vehicle speed, total current, total voltage, and acceleration. A fully connected layer is set at the end of the BiLSTM to map the time-series representation to estimated values ​​of 16-dimensional statistical features of the complete segment. This feature set is completely consistent with the features used in the basic modeling stage described in S3.

[0082] S6-2, Cascaded Prediction.

[0083] The 16-dimensional extrapolated features output from S6-1 are input into the working condition hierarchical residual ensemble model trained by S4. First, the working condition category is determined, and then the corresponding XGBoost and LightGBM cascaded sub-models are called to output the final predicted value of the SOC change of the segment.

[0084] S6-3, Multiple windows are used to evenly divide segments with large changes.

[0085] For segments with large changes exceeding 90 minutes in length, when the initial prediction value output by S6-2 exceeds the 10% SOC preset threshold, the complete segment is divided into several sub-segments of less than 90 minutes. For each sub-segment, feature extrapolation of S6-1 and cascaded prediction of S6-2 are performed separately. The prediction results of each sub-segment are summed to obtain the final SOC change of the complete segment, thereby reducing the cumulative error of long-term extrapolation.

[0086] In another embodiment of the present invention, the above method can be run on an in-vehicle terminal instead of a cloud server, utilizing the local computing resources of the in-vehicle terminal to estimate real-time SOC changes, thereby reducing communication latency and dependence on network bandwidth. The in-vehicle terminal executes steps S1 to S6 as described above, wherein the data acquisition in S1 is replaced by directly collecting real vehicle operating data through the in-vehicle CAN bus.

[0087] In another embodiment of the present invention, the three core modules mentioned above—work condition clustering hierarchical modeling, multi-window adaptive integration, and short-time extrapolation—can be implemented individually or in any combination. For example, with sufficient prior information about the route, such as the known work condition distribution of the complete route segment, only S4 and S5 can be implemented, omitting the short-time extrapolation step of S6, and the actual statistical characteristics of the complete segment can be used directly for prediction.

[0088] In another embodiment of the invention, the adaptive strategy in S5 can be further extended to include more standard-scale models, such as adding dedicated models at 30-minute or 120-minute scales, to further refine the granularity of multi-window integration and prediction accuracy. The corresponding dataset construction and condition clustering steps need to be adapted to include model training at this scale.

[0089] To verify the effectiveness of the method of this invention, actual operating data from 40 pure electric passenger vehicles over two consecutive months was used, covering driving and charging states. Key vehicle parameters are as follows: battery capacity of 39.0 kWh, vehicle weight of 2300 kg. Data collection fields were based on the requirements of the national standard GB / T 32960, generated by different vehicles at different times, covering diverse driving conditions and environmental conditions.

[0090] Two control groups were set up: Experimental group (method of this invention): adopts the complete process of steps S2 to S6, including SOC fine-tuning, working condition clustering hierarchical residual ensemble modeling, multi-time window adaptive ensemble prediction, and BiLSTM short-time global feature extrapolation.

[0091] Control group (coupled model without multi-time window concept): Steps S1 to S4 are the same as the method of this invention, but the multi-time window adaptive integration in step S5 is omitted, and the prediction result of a single 60-minute scale model is directly used as the SOC change estimate for any length segment.

[0092] The training set for both experiments used all driving data from 40 pure electric passenger vehicles over two months, while the test set used the first 10 minutes of time-series data from all driving data of each vehicle. This simulated the condition of only realizing part of the travel data in the actual scenario, and the root mean square error (RMSE) and mean absolute error (MAE) were used as evaluation metrics.

[0093] See test results Figure 7 and Figure 8 .

[0094] in, Figure 7 The scatter plot shows the prediction results of the coupled model without the multi-time window approach. The horizontal axis represents the change in actual SOC, and the vertical axis represents the change in predicted SOC. Observations reveal the following significant problems with this control group method: Prediction error is relatively small and accuracy is acceptable only when the segment length is close to 60 minutes, i.e., in the 55-65 minute range; when the segment length deviates from 60 minutes, prediction accuracy drops sharply, exhibiting a clear "V-shaped" performance degradation curve; for short segments (<30 minutes) and long segments (>90 minutes), the prediction error amplifies to unacceptable levels. The root cause of this phenomenon is that the training data for the single 60-minute model only covers segments close to 60 minutes in length, and this model lacks generalization ability for segments of other lengths that deviate from its training data distribution.

[0095] in, Figure 8 The image shows a scatter plot of the SOC change estimation results using the integrated multi-time window model (10-minute and 60-minute dual-scale model) of this invention. The figure shows that: across the entire segment length range from several minutes to over 200 minutes, the prediction error distribution is uniform and concentrated, with no obvious performance degradation intervals; the boundary effect of segment length is effectively eliminated, and the prediction accuracy for short, medium, and long segments remains at a high level; the multi-time window adaptive strategy allows each segment length to be adapted to its matching standard scale model.

[0096] The above results demonstrate that the present invention improves the accuracy and robustness of SOC change estimation for electric vehicles through a collaborative strategy of hierarchical modeling based on operating conditions and adaptive integration of multiple time windows, and exhibits excellent generalization ability, especially when facing complex driving segments of different lengths and operating conditions.

[0097] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to specific embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for estimating the change in SOC of new energy vehicles by integrating a multi-time window model, characterized in that, Includes the following steps: We acquire real-vehicle operation data of new energy vehicles, preprocess and refine the acquired real-vehicle operation data using SOC, and construct two standard scale segment datasets containing 10-minute segments and 60-minute segments. Feature extraction is performed on each segment in the standard-scale segment dataset to obtain multidimensional statistical features; Based on the aforementioned multidimensional statistical characteristics, a working condition clustering and hierarchical residual integrated modeling strategy is adopted to construct a prediction model for the SOC change of each working condition category. The hierarchical residual ensemble model includes a base predictor and a residual correction model; During prediction, the working condition category of the segment to be predicted is first determined, the basic predictor of the corresponding working condition category is called to output the initial prediction value, and then the initial prediction value is corrected by the residual correction model to obtain the predicted value of SOC change under the working condition. Based on the actual duration of the segment to be predicted, a multi-time window adaptive fusion strategy is adopted to adapt the segment to be predicted of any duration to the standard scale model corresponding to the standard scale segment dataset through expansion or decomposition. The prediction results of each segment are then fused to obtain the final SOC change.

2. The method for estimating the SOC change of new energy vehicles using an integrated multi-time window model according to claim 1, characterized in that, The SOC fine-tuning process includes: Select a charging segment to calibrate the actual total battery capacity; When the SOC value jumps to an integer value, the last integer value data point before the jump is used as the reference point. Based on the SOC integer value, cumulative discharge capacity, and actual total capacity of the battery at the reference point, each data point in the discharge segment is refined point by point to obtain a continuous SOC value.

3. The method for estimating the SOC change of new energy vehicles using an integrated multi-time window model according to claim 1, characterized in that, The construction of the standard-scale fragment dataset includes: The original vehicle driving data is divided into basic time windows of fixed duration. Extract the dynamic features of each basic time window, perform dimensionality reduction on the dynamic features, and divide the high-dynamic class and low-dynamic class based on the first principal component score; When generating target scale segments using a sliding window, the sliding step size is adaptively adjusted based on the proportion of high dynamic windows within the candidate segments: overlapping movement is used when the proportion of high dynamic windows is not less than a preset threshold, otherwise non-overlapping movement is used. The SOC change of each segment is calculated as the supervised training label.

4. The method for estimating the SOC change of new energy vehicles using an integrated multi-time window model according to claim 3, characterized in that, The dynamic characteristics of the basic time window include velocity variance, current variance, acceleration variance, and velocity amplitude; the dimensionality reduction process uses principal component analysis, with the mean minus one standard deviation as the classification threshold between the high-dynamic class and the low-dynamic class.

5. The method for estimating the SOC change of new energy vehicles using an integrated multi-time window model according to claim 1, characterized in that, The integrated modeling of working condition clustering and hierarchical residuals includes: Based on the two dimensions of energy consumption per unit mileage and average speed, K-Means clustering was used to divide all segment samples into several typical driving condition categories, including high-speed cruising, continuous uphill climbing, and urban congestion. XGBoost base predictors are trained independently for each type of work condition, and the multidimensional statistical features are used as input to output the initial predicted value of SOC change. The LightGBM model is used to learn the residual between the initial predicted value and the true value, and the residual correction value is output. The initial predicted value is added to the residual correction value to obtain the final predicted value of SOC change under this operating condition.

6. The method for estimating the SOC change of new energy vehicles using an integrated multi-time window model according to claim 5, characterized in that, For 10-minute scale segments, no load condition stratification is introduced; instead, a unified single model is used to process the prediction of all 10-minute scale segments.

7. The method for estimating the SOC change of new energy vehicles using an integrated multi-time window model according to claim 1, characterized in that, The multi-time-window adaptive integration strategy includes: The duration of the segment to be predicted is determined. If the duration of the segment to be predicted is within the preset standard duration range, the corresponding standard scale model is directly input to obtain the predicted value of SOC change. If the duration of the segment to be predicted is greater than the first threshold and less than the second threshold, the segment to be predicted is split into multiple sub-segments, each sub-segment is input into the corresponding standard scale model, and the prediction results are summed. If the duration of the segment to be predicted is within the third threshold range, the segment to be predicted will be extended to the standard duration through temporal self-copying, and then scaled after being input into the corresponding standard scale model. If the duration of the segment to be predicted is within the fourth threshold range, data is extracted from the beginning of the segment to be predicted and spliced ​​to the end to form a standard duration input, wherein the voltage data is regenerated according to the linear decay formula; If the duration of the segment to be predicted exceeds the fifth threshold, the standard duration portion is cut out, and the remaining portion is returned to the adaptive integration strategy for recursive processing.

8. The method for estimating the SOC change of new energy vehicles using an integrated multi-time window model according to claim 1, characterized in that, It also includes a short-time global feature extrapolation step: Using multidimensional time-series data with a preset duration before the segment as input, the temporal features are extracted through a bidirectional long short-term memory network (BiLSTM), and mapped to the estimated values ​​of the multidimensional statistical features of the complete segment through a fully connected layer. The estimated values ​​are then input into the hierarchical residual ensemble model, and the predicted value of SOC change is output.

9. The method for estimating the SOC change of new energy vehicles using an integrated multi-time window model according to claim 8, characterized in that, For segments with large changes exceeding a preset duration threshold, when the initially predicted SOC change exceeds the preset threshold, the complete segment is divided into several sub-segments smaller than the preset duration threshold. For each sub-segment, the feature extrapolation of the bidirectional long short-term memory network and the cascaded prediction of the hierarchical residual ensemble model are performed, and the prediction results of each sub-segment are accumulated.

10. The method for estimating the SOC change of new energy vehicles using an integrated multi-time window model according to claim 1, characterized in that, The method is deployed on a cloud server or in-vehicle terminal; the in-vehicle terminal collects real vehicle operation data and executes the method via the in-vehicle CAN bus.