Battery large model construction method and system based on multi-source heterogeneity and physical constraints

By constructing a large-scale battery model with multi-source heterogeneity and physical constraints, the problem of multi-source data integration is solved, and the stability and generalization performance of the battery state prediction model are improved, making it adaptable to complex operating conditions and suitable for predicting battery health status and remaining service life.

CN122451802APending Publication Date: 2026-07-24MIRATTERY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
MIRATTERY CO LTD
Filing Date
2026-05-08
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing technologies cannot effectively integrate multi-source heterogeneous battery data, resulting in limited stability and generalization performance of battery state prediction models. Traditional supervised learning algorithms are highly dependent on labeled data, have insufficient generalization ability for laboratory data, and cannot adapt to real vehicle environments.

Method used

By constructing a large battery model based on multi-source heterogeneity and physical constraints, multi-dimensional profile labels are built using cloud and laboratory data, feature dimension alignment is performed, a large pre-trained model of physically constrained batteries is constructed, and deep fusion of multi-source data is achieved through hierarchical learning rate decay and weight decay for coordinated fine-tuning.

Benefits of technology

It improves the practicality, stability, and cross-platform portability of the battery state prediction model, effectively utilizes massive amounts of unlabeled operational data, adapts to complex operating conditions, and enhances the model's generalization performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122451802A_ABST
    Figure CN122451802A_ABST
Patent Text Reader

Abstract

The application provides a battery large model construction method and system based on multi-source heterogeneity and physical constraints, relates to the technical field of battery model construction, and comprises the following steps: acquiring cloud battery data to construct a multi-dimensional portrait label, and obtaining a first training data set; acquiring laboratory battery data for standardization processing, and obtaining a second training data set; performing feature dimension alignment according to the first training data set and the second training data set, and obtaining a multi-source heterogeneous data set; constructing a physical constraint battery pre-training large model, training the physical constraint battery pre-training large model through the multi-source heterogeneous data set, and obtaining an initial physical constraint battery large model; and performing hierarchical learning rate decay and weight decay collaborative fine-tuning on the initial physical constraint battery large model based on a target task label, and obtaining a target physical constraint battery large model. The application utilizes multi-source heterogeneous data to learn the general representation of battery aging, and constructs physical constraints, thereby significantly improving the practicability and stability of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of battery model construction technology, specifically to a method and system for constructing a large battery model based on multi-source heterogeneity and physical constraints. Background Technology

[0002] With the rapid development of the new energy industry, the widespread application of electric vehicles and energy storage systems has driven the exponential growth of Battery Management System (BMS) data. While current cloud-stored real-vehicle operation data and laboratory test data have significant research value, the long-standing phenomenon of data silos has made it difficult to effectively integrate and extract value from multi-source heterogeneous data.

[0003] In constructing prediction models for battery state of health (SOH) and remaining useful life (RUL), existing methods face a dual dilemma: first, traditional supervised learning algorithms are heavily reliant on labeled data, making it difficult to effectively utilize massive amounts of unlabeled operational data; second, models trained solely on laboratory data often suffer from insufficient generalization ability, failing to adapt to the complex operating conditions of real-world vehicle environments. This contradiction directly leads to significant limitations in the stability and generalization performance of the prediction models. Therefore, achieving deep fusion of multi-source heterogeneous data has become a key issue in improving the performance of battery state prediction models. Summary of the Invention

[0004] In view of the above-mentioned shortcomings of the prior art, this application provides a method and system for constructing a large battery model based on multi-source heterogeneity and physical constraints, which effectively solves the problem that the current modeling method cannot achieve deep fusion of multi-source heterogeneous data, resulting in poor performance of the battery state prediction model.

[0005] Firstly, this application provides a method for constructing a large-scale battery model based on multi-source heterogeneity and physical constraints, the method comprising: Obtain cloud-based battery data, construct multi-dimensional profile labels based on the cloud-based battery data, and obtain the first training dataset; Acquire laboratory battery data, and perform standardization processing based on the laboratory battery data to obtain a second training dataset; By aligning the feature dimensions of the first training dataset and the second training dataset, a multi-source heterogeneous dataset is obtained. A pre-trained large model of a physically constrained battery is constructed, and the pre-trained large model of the physically constrained battery is trained using the multi-source heterogeneous dataset to obtain an initial physical constrained battery large model. Based on the target task label, the initial physical constraint battery model is finely adjusted by hierarchical learning rate decay and weight decay to obtain the target physical constraint battery model.

[0006] In an optional implementation, the physically constrained battery pre-trained large model includes at least a data encoding layer, a physically constrained masking layer, an encoder layer, a hidden space layer, and a decoder layer. The objective function of the physically constrained battery pre-trained large model includes at least a mask reconstruction loss term and a physical consistency regularization term. The expression of the objective function is as follows:

[0007] In the above formula, Represents the objective function value. This represents the mask mean square error constraint term. This represents the first-order difference continuity constraint term. This represents the second-order difference smoothing constraint term. The monotonically increasing constraint terms represent the remaining battery power and event capacity. This indicates the battery parameter sorting constraint. , , and These represent the weight coefficients of the first-order difference continuity constraint, the second-order difference smoothing constraint, the monotonically increasing constraint of remaining power and event capacity, and the battery parameter sorting constraint, respectively.

[0008] In an optional implementation, training the pre-trained large model of the physically constrained battery using the multi-source heterogeneous dataset to obtain the initial large model of the physically constrained battery includes: The multi-source heterogeneous dataset is input into the data encoding layer for training of temporal data encoding, local block unit encoding, and positional encoding to obtain structured representation data; The structured representation data is input into the physical constraint mask layer to train the physical constraint partial mask, thereby obtaining masked data segments and unmasked data segments; The unmasked data segment is input into the encoder layer for deep feature extraction training to obtain the encoded output features; The encoded output features are input into the hidden space layer for latent space representation processing and training to obtain a latent space vector. The masked data fragments are input into the decoder layer based on the latent spatial vector and mask label to perform data reconstruction training, thereby obtaining the initial physical constraint battery large model.

[0009] In an optional implementation, the step of performing hierarchical learning rate decay and weight decay coordinated fine-tuning on the initial physically constrained battery large model based on the target task label to obtain the target physically constrained battery large model includes: Based on the target task label, the encoder of the initial physical constraint battery large model is transferred to the target task, and a loss function is constructed; The encoder's hierarchical learning rate decay parameter is adjusted; The objective function weight coefficients of the initial physical constraint battery model are weighted and decayed based on the loss function to obtain the target physical constraint battery model.

[0010] In an optional implementation, the step of constructing multi-dimensional profile labels based on the cloud battery data to obtain a first training dataset includes: The cloud-based battery data is used to construct a multi-dimensional profile tag vector based on battery operating conditions. The multi-dimensional profile tag vector includes at least one or more of the following: regional information tag, battery mileage range tag, charging number range tag, charging ratio tag, charge / discharge depth tag, battery model tag, battery chemical system tag, and operating temperature range tag. The cloud battery data is divided into multiple data hierarchical units based on the multidimensional profile label vector, and the sampling weight of each data hierarchical unit is defined to obtain the first training dataset.

[0011] In an optional implementation, the standardization process based on the laboratory battery data to obtain a second training dataset includes: The laboratory battery data was sliced ​​to obtain multiple charge and discharge event datasets; Each of the charging and discharging event datasets is preprocessed to obtain multiple preprocessed datasets. The data preprocessing includes at least data cleaning, data resampling, and time alignment. The degradation stage is identified for each of the preprocessed datasets to obtain the second training dataset.

[0012] In an optional implementation, the step of aligning the feature dimensions based on the first training dataset and the second training dataset to obtain a multi-source heterogeneous dataset includes: Construct a unified feature space, which includes at least excitation features, response features, mask features, and encoding features; Perform dimension alignment mapping on the second training dataset to generate a unified representation dataset; Based on the unified feature space, feature alignment is performed on the unified representation dataset and the first training dataset to obtain the multi-source heterogeneous dataset.

[0013] Secondly, this application provides a battery large-scale model construction system based on multi-source heterogeneity and physical constraints, the system comprising: The first data processing module is used to acquire cloud battery data, construct multi-dimensional profile labels based on the cloud battery data, and obtain the first training dataset. The second data processing module is used to acquire laboratory battery data and perform standardization processing on the laboratory battery data to obtain a second training dataset. The feature dimension alignment module is used to perform feature dimension alignment based on the first training dataset and the second training dataset to obtain a multi-source heterogeneous dataset. The model supervised training module is used to construct a pre-trained large model of a physically constrained battery. The pre-trained large model of the physically constrained battery is trained using the multi-source heterogeneous dataset to obtain an initial physical constrained battery large model. The model parameter fine-tuning module is used to perform hierarchical learning rate decay and weight decay coordinated fine-tuning on the initial physical constraint battery large model based on the target task label, so as to obtain the target physical constraint battery large model.

[0014] Thirdly, this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the battery large model construction method based on multi-source heterogeneity and physical constraints as described in the first aspect of this application.

[0015] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the battery large model construction method based on multi-source heterogeneity and physical constraints as described in the first aspect of this application.

[0016] This application provides a method and system for constructing a large-scale battery model based on multi-source heterogeneity and physical constraints. It utilizes multi-dimensional profiling and hierarchical sampling based on massive cloud data. Through dimensional alignment and unified representation, cloud-based battery operating data and laboratory battery data can be used collaboratively within the same model, thus combining advantages in scale and standardization. Simultaneously, based on a physical constraint masking mechanism, the pre-training task is transformed into a response reconstruction task subject to physical prior constraints. The pre-trained large-scale model can fully leverage multi-source heterogeneous data to learn general characteristics of battery aging, reducing dependence on labels such as prediction parameters. In the application stage of the large-scale battery model, the hierarchical adjustment learning strategy retains the general aging knowledge obtained from pre-training while supporting rapid adaptation to new battery types and operating environments under limited labeling conditions, significantly improving the practicality, stability, and cross-platform transferability of the large-scale battery model. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of the battery large model construction method based on multi-source heterogeneity and physical constraints provided in the embodiments of this application; Figure 2 This is a schematic diagram of the standard charging data of a certain battery cell after data cleaning in an embodiment of this application; Figure 3 This is a schematic diagram of the cell aging trajectory under standard operating conditions in the embodiments of this application; Figure 4 This is a schematic diagram of heterogeneous multi-source data from the battery cell to the cloud in an embodiment of this application; Figure 5 This is a schematic diagram of the encoder's hierarchical learning rate decay in an embodiment of this application; Figure 6 This is a schematic diagram illustrating the fine-tuning error of the target physical constraint battery large model in this application embodiment under a small sample size of 5%; Figure 7 This is a schematic diagram of the loss curve during the fine-tuning training process of the target physical constraint battery large model in the embodiments of this application; Figure 8 This is a schematic diagram of the battery large model construction system based on multi-source heterogeneity and physical constraints provided in the embodiments of this application; Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0019] Explanation of key component symbols: 200. Battery large model construction system based on multi-source heterogeneity and physical constraints; 210. First data processing module; 220. Second data processing module; 230. Feature dimension alignment module; 240. Model supervised training module; 250. Model parameter fine-tuning module; 300. Electronic device; 310. Processor; 320. Communication interface; 330. Memory; 340. Communication bus. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be further described clearly and completely below with reference to the accompanying drawings of the embodiments. It should be noted that the described embodiments are merely some embodiments of this application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0021] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of this application.

[0023] In constructing prediction models for battery state of health (SOH) and remaining useful life (RUL), traditional supervised learning algorithms are heavily reliant on labeled data, making it difficult to effectively utilize massive amounts of unlabeled operational data. Models trained solely on laboratory data often suffer from insufficient generalization ability, failing to adapt to the complex operating conditions of real-world vehicle environments. Existing methods face a dual dilemma in constructing these prediction models: firstly, traditional supervised learning algorithms are heavily reliant on labeled data, making it difficult to effectively utilize massive amounts of unlabeled operational data; secondly, models trained solely on laboratory data often suffer from insufficient generalization ability, failing to adapt to the complex operating conditions of real-world vehicle environments. Currently, traditional methods based on manual feature engineering, deep learning methods based on single-source data, battery modeling methods based on transfer learning or pre-training, and small-sample modeling methods based on laboratory data all fail to achieve deep fusion of multi-source heterogeneous data, directly leading to significantly limited stability and generalization performance of the prediction models.

[0024] Example 1 This application provides a method for constructing a large battery model based on multi-source heterogeneity and physical constraints, which effectively solves the problem that current modeling methods cannot achieve deep fusion of multi-source heterogeneous data, resulting in poor performance of battery state prediction models. Figure 1 This is a schematic diagram of the battery large model construction method based on multi-source heterogeneity and physical constraints provided in the embodiments of this application, such as... Figure 1 As shown, the method includes the following steps: S100. Obtain cloud battery data, construct multi-dimensional profile labels based on cloud battery data, and obtain the first training dataset.

[0025] In this embodiment, massive amounts of battery operation data can be collected from the cloud to obtain cloud-based battery data. The cloud-based battery data is then used to construct a multi-dimensional profile tag vector based on battery operating conditions. For example, a multi-dimensional profile tag vector can be constructed for each charging event, charge / discharge segment, or cycle sample. This multi-dimensional profile tag vector includes, but is not limited to, one or more of the following: geographic information tags, battery mileage range tags, charging cycle range tags, charging ratio tags, charge / discharge depth tags, battery model tags, battery chemical system tags, and operating temperature range tags.

[0026] In one implementation, the geographic information tag includes key geographic information. To protect user privacy, the geographic information dimension is only accurate to seven major regions: Central China, East China, North China, South China, Northeast China, Northwest China, and Southwest China. The battery mileage range tag is used to mark the battery's cumulative mileage; the charge cycle range tag is used to mark the cumulative number of charges; the charge ratio tag is used to mark the fast charging ratio or the regular charging ratio; the charge / discharge depth tag is used to mark the battery's charge / discharge depth; the battery model tag is used to mark different battery model information, such as models A, B, C, D, and E; the battery chemistry system tag is used to mark the battery's chemistry system, such as nickel-cobalt-manganese ternary lithium (NCM) and lithium iron phosphate (LFP) systems; and the operating temperature range tag is used to mark the battery's operating temperature range.

[0027] For example, the first i The multidimensional profile label vectors for each charge / discharge event sample are as follows:

[0028] in, i Indicates the first i The sequence number of the next charge / discharge event. Indicates the first i A multidimensional profile label vector for each charge / discharge event sample. Indicates a regional information tag. Indicates the battery range label. Labels indicating the range of charging cycles. A label indicating the charging percentage. Label indicating charge / discharge depth The label indicates the battery model. Indicates the battery chemistry system label, Labels indicating the operating temperature range.

[0029] The cloud battery data is divided into multiple data hierarchical units based on the multidimensional profile label vector, and the sampling weight of each data hierarchical unit is defined to obtain the first training dataset.

[0030] Based on the aforementioned multi-dimensional profile tag vectors, the cloud battery data is divided into several data hierarchical units. Define sampling weights for each data hierarchical unit. ,in This indicates the sample size allocated to this data stratification unit. The aging information index for this data stratification unit consists of charge / discharge depth and coverage. Discharge depth is obtained by the difference in State of Charge (SOC) before and after charging. For example, if charging starts at SOC1 and ends at SOC2, the charge / discharge depth of this charging event is SOC2 - SOC1. The charge / discharge depth is obtained by statistically analyzing the charge / discharge depth of all charging events. Coverage is obtained by acquiring the number of charging cycles, including fast charging cycles C1 and slow charging cycles C2. Fast charging coverage is obtained by calculating C1 / (C1+C2), and slow charging coverage is obtained by calculating C2 / (C1+C2). The sample quality indicators for this data stratification unit consist of missing rate, anomaly rate, and signal stability. The sample quality indicators are used to evaluate the data quality of this charging event and whether the data is reliable. It is necessary to exclude some unreliable and abnormal data. For example, the missing rate is the proportion of missing frames in the entire charging event to the total number of frames; the anomaly rate is the proportion of abnormal charging frames (voltage, current, or temperature exceeding the range or abnormal jump judgment) in the entire charging event to the total number of frames; and the signal stability is the proportion of the number of dropped frames in the entire charging event to the total number of frames. , and This indicates that hyperparameters can be adjusted and set according to the actual situation.

[0031] Understandably, based on the sampling weights of each data stratification unit and the total sampling plan quota, the sampling quota for each data stratification unit can be calculated as follows:

[0032] In the above formula, Indicates the first k The sampling quota for each data stratification unit, where Q represents the total sampling plan quota. J This indicates the total number of data layer units. j Indicates the sequence number of the data layer unit.

[0033] Based on this, the first training dataset obtained through the above stratified sampling method simultaneously covers different regions, usage intensities, aging stages, and battery types, thus providing a high-quality, representative, and diverse real-world dataset for subsequent large-scale model training. Furthermore, stratified sampling and quality screening significantly improve the coverage, balance, and aging information density of the training data, reducing the waste of training resources caused by redundant samples.

[0034] S200. Obtain laboratory battery data, perform standardization processing based on the laboratory battery data, and obtain the second training dataset.

[0035] In this embodiment, the laboratory battery data primarily comes from cell charge-discharge cycle tests under standard operating conditions, including battery data such as constant current / constant voltage charging, constant current discharging, and capacity decay processes. The standardization process based on the laboratory battery data is as follows: First, the laboratory battery data is sliced ​​to obtain multiple charge / discharge event datasets. Cyclic slicing or event slicing can be performed to encode individual charge or discharge events and record unique identifiers, such as using the event's start time as a unique identifier.

[0036] Then, data preprocessing is performed on each charge and discharge event dataset to obtain multiple preprocessed datasets. This data preprocessing includes, but is not limited to, data cleaning, data resampling, and time alignment.

[0037] For example, due to the influence of sensors, devices and environment, it is necessary to align the noise and sampling frequency of the charging and discharging event dataset. For example, it can be uniformly normalized to 5s~10s for resampling, and information such as voltage, current, temperature and capacity increment can be collected. For events with maximum values, minimum values, noise or sampling abnormalities, the corresponding data of the event can be discarded. Figure 2 This is a schematic diagram of standard charging data for a certain battery cell after data cleaning, as shown in the embodiments of this application. Figure 2 As shown, the horizontal axis represents incremental capacity, the vertical axis represents voltage, and each curve represents the charging curve under each SOH. As the SOH ages, the cell charging curve shows significant changes.

[0038] Finally, degradation stage labels are applied to each preprocessed dataset to obtain the second training dataset.

[0039] Understandably, laboratory battery data can obtain standard capacity calibration data, have accurate capacity labels, can mark the capacity of events, and obtain the state of energy (SOH) at different cycle and aging stages. Figure 3 This is a schematic diagram of the cell aging trajectory under standard operating conditions in the embodiments of this application, such as... Figure 3 As shown, under laboratory conditions, the state of oxygen (SOH) of multiple cells changes with the number of cycles. By analyzing the change trajectory data, the charging curves of the cells at different SOH stages can be accurately obtained.

[0040] Based on this, the second training dataset obtained from laboratory battery data can provide standard degradation paths and high-quality sampling data. It has fewer outliers and can provide high-quality standard aging trajectories, which can serve as an important source of supplementary cloud battery data mechanisms.

[0041] S300. Align the feature dimensions based on the first and second training datasets to obtain a multi-source heterogeneous dataset.

[0042] In this embodiment, the first training dataset obtained based on cloud-based battery data and the second training dataset obtained based on laboratory battery data belong to different modal signals. Therefore, aligning their feature dimensions is crucial for extracting rich aging mechanism features. Feature dimension alignment specifically includes the following steps: First, a unified feature space is constructed, which includes, but is not limited to, stimulus features, response features, mask features, and encoding features.

[0043] For example, in order to unify the two types of training set data, the embodiments of this application can construct a unified feature space as follows:

[0044] In the above formula, Representing a unified feature space, The excitation characteristics can include information such as current and state of charge (SOC). The response characteristics can include information such as voltage statistics, temperature statistics, and capacity statistics. This represents a mask feature used to characterize modal availability mask information, i.e., information that needs to be aligned with the battery pack, such as... V min and V max These signals are not available in the laboratory data, so they need to be compensated or transformed. This mask feature solves the feature alignment problem caused by the inconsistency in the frequency and dimension of cross-scene data acquisition, making the model robust in processing non-full observation data. It represents the coding characteristics used to characterize source domain coding and type coding. For example, the information source coding indicates whether the data is from the laboratory or the cloud.

[0045] Then, dimension alignment mapping is performed on the second training dataset to generate a unified representation dataset.

[0046] Understandably, the first training dataset obtained from cloud-based battery data typically includes multi-dimensional statistics such as current, remaining state of charge (SOC) of the battery pack, minimum voltage, average voltage, maximum voltage, minimum temperature, maximum temperature, and incremental capacity. However, the second training dataset obtained from laboratory battery data often only contains a limited set of signals, such as cell voltage, current, and capacity. Therefore, dimensionality alignment mapping is performed on the second training dataset to obtain a unified representation dataset. For example, for the voltage of a single laboratory cell... The unified representation is generated using the dimension alignment mapping function A() as follows:

[0047] In the above formula, This represents the minimum voltage of a single battery cell after dimension alignment mapping. This represents the average voltage of a single battery cell after dimension-aligned mapping. This represents the maximum voltage of a single battery cell after dimension alignment mapping. It represents the prior statistical parameters used to characterize the discrete distribution of voltage, the statistical regularity of series and parallel structures, or the empirical boundary in cloud data of the same battery system.

[0048] The mapped unified expression parameter values ​​can be determined using either the equivalence alignment method or the prior discrete compensation method. The equivalence alignment method is determined as follows:

[0049] The prior discrete compensation method is determined as follows:

[0050] In the above formula, This represents the compensation voltage obtained through learning.

[0051] Finally, feature alignment is performed on the unified representation dataset and the first training dataset based on the unified feature space to obtain a multi-source heterogeneous dataset.

[0052] By aligning the feature dimensions of the unified representation dataset and the first training dataset through a unified feature space, a unified and standardized multi-source heterogeneous dataset can be obtained. This enables large-scale real knowledge in the cloud and high-quality aging knowledge in the laboratory to learn together in the same model and obtain rich feature information. Figure 4 This is a schematic diagram of heterogeneous multi-source data from the battery cell to the cloud in an embodiment of this application, such as... Figure 4 As shown, the solid line represents the individual voltage curve of a laboratory cell. The laboratory only has the voltage of a single cell, unlike the cloud-based battery pack which contains multiple cells. V min , V avg and V max The dashed lines represent the unified representation of dimensions in the aligned cloud battery data, including aligned voltage 1 and aligned voltage 2.

[0053] Based on this, by dimensional alignment and unified representation of the first and second training datasets, rich mechanistic features of diverse battery aging trajectories can be obtained, enabling cloud-based operational data and laboratory cell data to be used collaboratively in the same model, thus combining the advantages of scale and standardization.

[0054] S400. Construct a pre-trained large model of a physically constrained battery. Train the pre-trained large model of a physically constrained battery using a multi-source heterogeneous dataset to obtain the initial physical constrained battery large model.

[0055] In the embodiments of this application, the physical constraint battery pre-trained large model includes at least a data encoding layer, a physical constraint mask layer, an encoder layer, a hidden space layer, and a decoder layer.

[0056] As an optional implementation of this application, the data encoding layer may include a time-series data encoding module, a local block unit encoding module, and a position encoding module, wherein: The time-series data encoding module, given a cleaned data input sequence, aligns and quantizes it to convert it into time-series initialization units. For example, the time-series initialization unit corresponding to a single charging segment in a multi-source heterogeneous dataset after processing by the time-series data encoding module is... ,in, T D represents the sequence length, and D represents the feature dimension. Indicates the first t Multidimensional observations at each moment.

[0057] Furthermore, to learn shared battery aging characteristics across types and improve generalization capabilities across battery systems and models, multi-dimensional feature encoding is introduced into the input of the time-series data encoding module. This includes, but is not limited to, one or more of the following: battery type encoding, chemical system encoding, battery platform encoding, data source encoding, temperature zone encoding, and operating condition encoding. Let the output of the shared encoder be... Where z represents the aforementioned multidimensional feature encoding, and X represents the initial temporal unit, this forces the physically constrained battery pre-trained large model to simultaneously observe real-world complex operating conditions in the cloud, standard laboratory aging trajectories, different chemical systems such as nickel-cobalt-manganese ternary lithium (NCM) and lithium iron phosphate (LFP), as well as different vehicle models and battery pack structures during the pre-training phase. Therefore, the physically constrained battery pre-trained large model learns aging representations that are more vertically focused, rather than short-term general fitting patterns on a single dataset.

[0058] The local block unit encoding module divides the long-term sequence into several local continuous segments, transforming the model from time-series modeling to local temporal pattern modeling, thus improving the ability to extract local evolution patterns of charging and discharging. Assuming the local block length is... P The sequence is then divided into N=T / P The local block, the first i The local blocks are represented as follows:

[0059] In the above formula, Indicates the first i Feature representation of a local block Indicates the firsti Temporal characteristics of local blocks.

[0060] The first i The flattened local blocks projected onto the embedding space are represented as follows:

[0061] In the above formula, Indicates the first i The original features of each local block are used to perform block-based feature representation encoding. Where d represents the feature dimension, Represents the flattening function. Indicates the coding weight of a local block. , This represents the encoding bias of a local block.

[0062] Obtained by performing convolution operations using a convolutional neural network:

[0063] In the above formula, This represents the features of a local block after convolutional block encoding. Encoding of the original features This represents the convolution kernel of a convolutional neural network. This represents the stride of the convolutional neural network.

[0064] The local block can be obtained in the end. patch The sequence is as follows:

[0065] By using local block processing, each data unit can simultaneously contain dynamic information within a local time window, rather than instantaneous information at a single moment, making it more suitable for characterizing local slope changes, plateau changes, and short-term fluctuation patterns during battery charging.

[0066] In this embodiment, since the Transformer structure itself does not explicitly contain sequence order information, a position encoding module is introduced to preserve the sequential relationship of each local block on the original timeline. Let the position encoding matrix be... Then the final representation of the input encoder is ,in The position encoding is processed using learnable parameters as input to the encoder.

[0067] Battery aging-related characteristics have obvious time evolution attributes. The same voltage or temperature value appears at different stages and has different physical meanings. Introducing a position encoding module can enhance the model's ability to identify the charging process and state evolution sequence.

[0068] In this embodiment, the charging segment sequence in the multi-source heterogeneous dataset is input into the data encoding layer and sequentially sent to the time-series data encoding module, the local block unit encoding module, and the position encoding module to obtain local blocks that combine local multivariate coupling information and time-series position information, thereby obtaining structured representation data.

[0069] In this embodiment, the physical constraint masking layer employs a masking strategy with physical constraints to mask certain local blocks in the structured representation data, constructing an incomplete input while maintaining temporal continuity and physical correlation. Notably, the physical constraint masking layer uses a temporal asymmetric conditional masking mechanism oriented towards the electrochemical excitation-response mechanism, learning the physical mapping from excitation to response, rather than a simple data completion relationship.

[0070] Specifically, an asymmetric conditional mask matrix is ​​constructed for the battery charging timing data. Based on the physical interactions during battery charging, the input features are divided into a set of excitation features. and response class feature set Among them, excitation features such as current and SOC characterize external input conditions, while response features such as voltage, temperature, and capacity increments characterize the system state output. For any feature dimension... d Its mask probability is defined as follows:

[0071] In the above formula , Representing feature dimension d The mask probability, Representing feature dimension d The probability of the incentive-type feature mask. Representing feature dimension d The response class feature mask probability, where .

[0072] The input after the mask is And it satisfies time series constraints That is, the masking operation does not change the original time order, where express t Time-of-flight feature dimension d Input features after masking express t Time-of-time feature dimension d The original characteristics, express t Time-of-flight feature dimension d The mask matrix. The model maximizes the conditional probability. Learning in Motivational Information With locally visible response information Response to shading under certain conditions The basic loss function for reconstruction is expressed as follows:

[0073] In the above formula, The model mask loss function value is further superimposed with voltage sequence constraint loss and time smoothing constraint loss to enhance the consistency between the reconstruction result and the physical laws of the battery.

[0074] Understandably, by keeping stimulus features unmasked or with a low-proportion mask to fully preserve driving information, and applying a high-proportion random mask to response features (ideally 75%), the model's ability to reconstruct key state variables is enhanced. Simultaneously, the masking operation only changes feature visibility, not the temporal order, thus maintaining the dynamic continuity of the charging trajectory. Based on this mechanism, the model is constrained during pre-training to infer the masked response quantity based on stimulus information and locally visible response information, thereby learning a deep representation that conforms to the battery stimulus response pattern, improving generalization ability across battery systems and the few-sample fine-tuning effect for downstream tasks.

[0075] The encoder layer is used to extract deep features from unmasked local blocks to obtain global charge and discharge features, so as to learn high-order temporal dependencies and implicit aging characteristics during battery charging.

[0076] For example, if the encoder input is ,use L The Transformer Encoder layers are stacked, and the recursive form is as follows:

[0077] In the above formula, express l The recursive representation of the layer Transformer Encoder, Indicates the first l The layer's encoded output.

[0078] A single layer of self-attention can be represented as follows:

[0079] In the above formula, Q Represents the query vector. K Represents the key vector. , V Represents a value vector. Represents the dimension of the key vector. H Indicates the encoded output. , and These represent the weight coefficients of the query vector, key vector, and value vector, respectively.

[0080] The final encoder layer output is The encoder layer captures long-distance temporal dependencies through a multi-head self-attention mechanism, enabling it to learn the mutual influence between different time segments and thus extract a deep representation that reflects the battery's health status and aging trend.

[0081] The hidden space layer is used to map the encoder output into a more compact and transferable latent space vector to carry core information about battery aging mechanisms and operating states. Let the encoder output be... ,in h i Indicates the first i The latent space vector can be obtained by aggregation, such as mean pooling, using the following latent space representation:

[0082] In the above formula, Z represents the latent space vector after mean pooling.

[0083] Based on this, the latent space vector can provide a more abstract and compressed representation of the aging state, dynamic characteristics, and health mode of the current charging segment. This representation can be used for decoding and reconstruction, as well as as a shared basic feature for downstream health status, remaining lifetime, or classification tasks.

[0084] The decoder layer recovers the local blocks of the masked response class by combining the latent space vector with the mask label, thus achieving the goal of self-supervised pre-training.

[0085] For example, the latent space vector and mask labels are combined to restore the complete sequence length. The `Restore()` function reassembles the visible data blocks and mask markers into a complete sequence based on their original local block positions. The complete sequence is then fed into... L The TransformerDecoder layer yields the following recursive form:

[0086] In the above formula, express l The recursive representation of the layer Transformer Decoder, Indicates the first l The layer's decoding output.

[0087] Finally, the results are output at the local block level. The decoder layer is tasked with recovering the occluded response information from the latent representation, forcing the encoder to learn the core representations related to response prediction rather than just memorizing local visible values, thereby improving the effectiveness and transfer value of pre-training.

[0088] Finally, a reconstruction consistency loss function is constructed based on the masked region, and a physically constrained model objective function is constructed by combining it with a physical consistency constraint term. This allows the learned feature representation to better characterize battery charging behavior and its aging evolution. The objective function of the physically constrained battery pre-trained large model includes at least a mask reconstruction loss term and a physical consistency regularization term. The mask reconstruction loss term only calculates the deviation between the predicted and true values ​​on the masked segments to drive the model to recover missing temporal information from fragmented charging and discharging segments that are only partially visible. The physical consistency regularization term further constrains the reconstruction results in terms of temporal continuity, local smoothness, monotonic evolution of state variables, and multivariate coupling relationships. This ensures that the feature representation learned by the model not only has numerical reconstruction capabilities but also satisfies the physical rationality and evolutionary consistency of the battery charging process. The expression of the objective function is as follows:

[0089] In the above formula, Represents the objective function value. This represents the mask mean square error constraint term. This represents the first-order difference continuity constraint term. This represents the second-order difference smoothing constraint term. The monotonically increasing constraint terms represent the remaining battery power and event capacity. This indicates the battery parameter sorting constraint. , , and These represent the weight coefficients of the first-order difference continuity constraint, the second-order difference smoothing constraint, the monotonically increasing constraint of remaining power and event capacity, and the battery parameter sorting constraint, respectively, which can be dynamically adjusted according to actual conditions.

[0090] In this embodiment, the mask mean square error constraint term is the mask reconstruction loss term, used to reconstruct the charge-discharge curve from finitely visible fragmented charge-discharge segments. The first-order difference continuity constraint term is used to ensure that the changing trend of the reconstructed signal at adjacent time points remains consistent with the actual charging process, thereby avoiding unreasonable abrupt changes at the reconstruction boundary. The second-order difference smoothing constraint term, by penalizing the second-order difference of the reconstructed sequence, suppresses high-frequency jitter and non-physical oscillations, making the reconstruction result more consistent with the gradual dynamic evolution characteristics of the battery charging process. The monotonically increasing constraint terms for remaining capacity and event capacity are used to constrain the changing direction of state variables with definite physical evolution directions in the time dimension, thereby enhancing the model's ability to preserve the charging stage patterns and state evolution trends. The battery parameter ranking constraint term constrains the order priority of battery parameters such as maximum voltage, minimum voltage, average voltage, highest temperature, and lowest temperature, enabling the model to learn not only numerical fitting but also abstracting the boundary relationships and evolutionary laws in the battery response sequence.

[0091] Based on the model structure of the pre-trained large-scale physical constraint battery model described above, the initial large-scale physical constraint battery model can be obtained by training with multi-source heterogeneous datasets. The specific training steps include the following: S410. Input the multi-source heterogeneous dataset into the data encoding layer for training of temporal data encoding, local block unit encoding and positional encoding to obtain structured representation data.

[0092] S420. Input the structured representation data into the physical constraint mask layer to train the physical constraint partial mask, and obtain the masked data segment and the unmasked data segment.

[0093] S430. Input the unmasked data segment into the encoder layer for deep feature extraction training to obtain the encoded output features.

[0094] S440. Input the encoded output features into the hidden space layer for latent space representation processing and training to obtain the latent space vector.

[0095] S450. Based on the latent space vector and mask label, the masked data fragments are input to the decoder layer for data reconstruction training to obtain the initial physical constraint battery large model.

[0096] Based on this, the structured representation of the original charging segment is completed through the data encoding layer. The physical constraint masking mechanism transforms the pre-training task into a response reconstruction task subject to physical prior constraints by masking the stimulus features with low masking and the response features with high masking while maintaining time order. The encoder is responsible for learning the high-order latent representation of the battery's dynamic behavior, and the decoder is responsible for recovering the masked response information. Finally, the model optimization is driven by the combination of reconstruction error and physical consistency constraints to obtain the initial physical constraint battery model. This can significantly enhance the model's ability to represent the battery aging mechanism and improve the model's generalization performance for different battery types and downstream tasks with few samples.

[0097] S500 performs hierarchical learning rate decay and weight decay coordinated fine-tuning on the initial physical constraint battery model based on the target task label to obtain the target physical constraint battery model.

[0098] In this embodiment, the initial physically constrained battery model, based on physical constraints, has already learned general structural knowledge of the battery aging process during the pre-training stage. Downstream tasks only require a small number of labels to complete task mapping. In actual fine-tuning, only a small number of target task labels are needed to enable the model to quickly adapt to new battery types, new operating condition distributions, or new platforms. Therefore, to simultaneously maintain pre-training knowledge and quickly adapt to new tasks in scenarios with a small number of sample labels, a few-sample fine-tuning approach combined with layer-wise learning rate decay (LLRD) and weight decay (WD) is used to collaboratively optimize the initial physically constrained battery model, obtaining the target physically constrained battery model. The collaborative optimization specifically includes the following steps: First, based on the target task label, the encoder of the initial physical constraint battery large model is transferred to the target task, and a loss function is constructed.

[0099] For example, in the downstream target task, the encoder of the initial physically constrained large battery model is transferred to a regression task of battery health state or battery remaining life, representing a given encoder output data block. H The global representation is obtained by using the pooling operator. Then, by using a simple regression probe, accurate predictions of battery health status or remaining battery life can be achieved, obtaining predicted values. , The hidden layer output of the encoder in the initial physical constraint battery large model.

[0100] Optionally, the loss function can be constructed using mean squared error or smoothed L1 loss, for example... , Indicates the predicted value. Indicates the tag value. N This indicates the total number of values.

[0101] Then, the encoder's hierarchical learning rate decay parameter is adjusted. Let the encoder have a total of... L Layer, number l The learning rate of a layer is defined as ,in Indicates the parameter to be tuned. The preferred value is between 0.6 and 0.85. This represents the learning rate for the top layer, which is the largest, and the learning rate decreases as you move closer to the bottom layer. For task headers, a larger learning rate can be set. ,in The preferred values ​​are 2 to 5.

[0102] Figure 5 This is a schematic diagram of the encoder's hierarchical learning rate decay in an embodiment of this application, as shown below. Figure 5 As shown, in the encoder, the last 4 layers are unfrozen, and different learning rates are applied to the last 4 layers. The learning rate is larger and more adjustments are made closer to the top, while the learning rate is smaller and less adjustments are made closer to the front. This helps to enhance the stability of the pre-trained model as much as possible.

[0103] Finally, based on the loss function, the weight coefficients of the objective function of the initial physically constrained battery model are weighted attenuated to obtain the target physically constrained battery model. For example, regularization is applied only to the weight parameters, while no weight attenuation is applied to the bias term, normalization parameters, and position encoding. The loss function is fine-tuned as follows:

[0104] In the above formula, Indicates fine-tuning loss. This represents the loss between the predicted and actual values. For the decay weight parameter, This indicates that the parameters are being adjusted.

[0105] Based on this, the underlying general features such as feature extraction and cleaning are slightly updated and fine-tuned, making them less susceptible to damage. The high-level aging features and task heads adapt faster, and the fine-tuning is more stable. With a few iterations, fine-tuning can achieve convergence, which greatly improves the training speed and generalization performance of the model. The features perform more stably when transferring between small samples and cross-types.

[0106] To verify the effectiveness of the target physical constraint battery large model provided in the embodiments of this application, fine-tuning training was performed with 5% of the labeled samples. Figure 6 This is a schematic diagram illustrating the fine-tuning error of the target physical constraint battery large model in this application embodiment under a small sample size of 5%, as shown below. Figure 6As shown, with 5% of the labeled samples, the model still has good error performance on the test set, with RMSE < 1%, indicating that the target physical constraint battery model converges faster and has better and more stable performance. Figure 7 This is a schematic diagram of the loss curve during the fine-tuning training process of the target physical constraint battery large model in this application embodiment, as shown in the figure. Figure 7 As shown, during the fine-tuning training process, the model loss changes with the number of iterations. The model stabilizes after about 20 iterations, demonstrating the rapid adaptability of fine-tuning.

[0107] Therefore, the target physical constraint battery model obtained in this application, after fine-tuning and training in downstream tasks, can be used to predict battery health status or remaining lifespan. Furthermore, with a small number of labels, it can quickly converge and iterate to a stable performance, demonstrating the powerful advantages of fine-tuning under the target physical constraint battery model. It can stably and quickly adapt to the prediction of battery health status or remaining lifespan for different battery types. Simultaneously, the target physical constraint battery model can better abstract the aging characteristics at different stages, achieving more stable transfer between different battery types, different chemical systems, and different operating conditions. Compared with traditional supervised methods, it significantly reduces the dependence on the target domain label size, and improves training efficiency and RMSE by at least 30%, effectively saving costs.

[0108] The battery large-scale model construction method based on multi-source heterogeneity and physical constraints provided in this application embodiment utilizes multi-dimensional profiling and hierarchical sampling based on massive cloud data. Through dimensional alignment and unified expression, cloud-based battery operation data and laboratory battery data can be used collaboratively in the same model, thus combining the advantages of scale and standardization. Simultaneously, based on a physical constraint masking mechanism, the pre-training task is transformed into a response reconstruction task subject to physical prior constraints. The pre-trained large-scale model can fully leverage multi-source heterogeneous data to learn general characteristics of battery aging, reducing dependence on labels such as prediction parameters. Even with only a few labels, it can still achieve superior performance in fewer training rounds, demonstrating greater engineering deployment value.

[0109] Example 2 Based on the same technical concept as Embodiment 1 above, this application provides a battery large model construction system based on multi-source heterogeneity and physical constraints. Figure 8 This is a schematic diagram of the battery large model construction system based on multi-source heterogeneity and physical constraints provided in the embodiments of this application, as shown below. Figure 8 As shown, the battery large model construction system 200 based on multi-source heterogeneity and physical constraints includes: The first data processing module 210 is used to acquire cloud battery data, construct multi-dimensional profile labels based on the cloud battery data, and obtain the first training dataset.

[0110] The second data processing module 220 is used to acquire laboratory battery data, perform standardization processing on the laboratory battery data, and obtain a second training dataset.

[0111] The feature dimension alignment module 230 is used to align the feature dimensions based on the first training dataset and the second training dataset to obtain a multi-source heterogeneous dataset.

[0112] The model supervised training module 240 is used to construct a pre-trained large model of a physically constrained battery. The pre-trained large model of the physically constrained battery is trained using a multi-source heterogeneous dataset to obtain an initial large model of the physically constrained battery.

[0113] The model parameter fine-tuning module 250 is used to perform hierarchical learning rate decay and weight decay collaborative fine-tuning on the initial physical constraint battery large model based on the target task label, so as to obtain the target physical constraint battery large model.

[0114] The battery large-scale model construction system based on multi-source heterogeneity and physical constraints provided in this application embodiment encompasses a comprehensive and innovative combination of data, tasks, training mechanisms, and transfer mechanisms. This includes high-quality data construction, unification of multi-source heterogeneity, pre-training based on physical constraints, cross-type and platform-specific knowledge generalization, and finally, efficient fine-tuning methods with few samples. In practical engineering applications, it can quickly adapt to fine-tuning of new batteries and platforms, achieving stable convergence speed and accuracy with a small number of labels, effectively improving development efficiency and cost cycle time, and providing stable battery state prediction capabilities.

[0115] It is understood that the implementation method of the battery large model construction method based on multi-source heterogeneity and physical constraints in the above embodiment 1 is also applicable to this embodiment and can achieve the same technical effect, so it will not be described again here.

[0116] Example 3 Based on the same concept, this application also provides an electronic device. Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application, such as... Figure 9 As shown, the electronic device 300 may include a processor 310, a communication interface 320, a memory 330, and a communication bus 340, wherein the processor 310, the communication interface 320, and the memory 330 communicate with each other via the communication bus 340. The processor 310 can call logical instructions in the memory 330 to execute the steps of the battery large model construction method based on multi-source heterogeneity and physical constraints as described in the above embodiments. For example, it includes: S100. Obtain cloud battery data, construct multi-dimensional profile labels based on cloud battery data, and obtain the first training dataset; S200. Obtain laboratory battery data, perform standardization processing based on the laboratory battery data, and obtain the second training dataset. S300. Align the feature dimensions based on the first and second training datasets to obtain a multi-source heterogeneous dataset. S400. Construct a pre-trained large model of a physically constrained battery. Train the pre-trained large model of a physically constrained battery using a multi-source heterogeneous dataset to obtain the initial physical constrained battery large model. S500 performs hierarchical learning rate decay and weight decay coordinated fine-tuning on the initial physical constraint battery model based on the target task label to obtain the target physical constraint battery model.

[0117] The processor 310 can be a central processing unit (CPU). The processor can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or combinations of the above types of chips.

[0118] Furthermore, the logical instructions in the aforementioned memory 330 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0119] The memory 330 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created by the processor, etc. Furthermore, the memory may include high-speed random access memory and non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory may optionally include memory remotely located relative to the processor, which can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0120] Example 4 Based on the same concept, embodiments of this application also provide a computer-readable storage medium storing a computer program containing at least one piece of code executable by a master control device to control the master control device to implement the steps of the battery large model construction method based on multi-source heterogeneity and physical constraints as described in the above embodiments. For example, it includes: S100. Obtain cloud battery data, construct multi-dimensional profile labels based on cloud battery data, and obtain the first training dataset; S200. Obtain laboratory battery data, perform standardization processing based on the laboratory battery data, and obtain the second training dataset. S300. Align the feature dimensions based on the first and second training datasets to obtain a multi-source heterogeneous dataset. S400. Construct a pre-trained large model of a physically constrained battery. Train the pre-trained large model of a physically constrained battery using a multi-source heterogeneous dataset to obtain the initial physical constrained battery large model. S500 performs hierarchical learning rate decay and weight decay coordinated fine-tuning on the initial physical constraint battery model based on the target task label to obtain the target physical constraint battery model.

[0121] Based on the same technical concept, this application also provides a computer program, which, when executed by a main control device, is used to implement the above-described method embodiments.

[0122] The computer program may be stored, in whole or in part, on a computer-readable storage medium packaged with the processor, or in part or in whole on a memory not packaged with the processor.

[0123] Based on the same technical concept, this application also provides a processor for implementing the above-described method embodiments. The processor can be a chip.

[0124] In summary, the battery large-scale model construction method and system based on multi-source heterogeneity and physical constraints provided in this application utilizes multi-dimensional profiling and hierarchical sampling based on massive cloud data. Through dimensional alignment and unified representation, cloud-based battery operation data and laboratory battery data can be used collaboratively in the same model, thus combining the advantages of scale and standardization. Simultaneously, based on a physical constraint masking mechanism, the pre-training task is transformed into a response reconstruction task subject to physical prior constraints. The pre-trained large-scale model can fully leverage multi-source heterogeneous data to learn general characteristics of battery aging, reducing dependence on labels such as prediction parameters. In the application stage of the battery large-scale model, the hierarchical adjustment learning strategy not only retains the general aging knowledge obtained from pre-training but also supports rapid adaptation to new battery types and operating environments under limited labeling conditions, significantly improving the practicality, stability, and cross-platform transferability of the battery large-scale model.

[0125] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0126] The embodiments described above are merely examples of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these modifications and improvements all fall within the protection scope of this application.

[0127] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for constructing a large-scale battery model based on multi-source heterogeneity and physical constraints, characterized in that, The method includes: Obtain cloud-based battery data, construct multi-dimensional profile labels based on the cloud-based battery data, and obtain the first training dataset; Acquire laboratory battery data, and perform standardization processing based on the laboratory battery data to obtain a second training dataset; By aligning the feature dimensions of the first training dataset and the second training dataset, a multi-source heterogeneous dataset is obtained. A pre-trained large model of a physically constrained battery is constructed, and the pre-trained large model of the physically constrained battery is trained using the multi-source heterogeneous dataset to obtain an initial physical constrained battery large model. Based on the target task label, the initial physical constraint battery model is finely adjusted by hierarchical learning rate decay and weight decay to obtain the target physical constraint battery model.

2. The method for constructing a large-scale battery model based on multi-source heterogeneity and physical constraints according to claim 1, characterized in that, The physically constrained battery pre-trained large model includes at least a data encoding layer, a physically constrained masking layer, an encoder layer, a hidden space layer, and a decoder layer. The objective function of the physically constrained battery pre-trained large model includes at least a mask reconstruction loss term and a physical consistency regularization term. The expression of the objective function is as follows: In the above formula, Represents the objective function value. This represents the mask mean square error constraint term. This represents the first-order difference continuity constraint term. This represents the second-order difference smoothing constraint term. The monotonically increasing constraint terms represent the remaining battery power and event capacity. This indicates the battery parameter sorting constraint. , , and These represent the weight coefficients of the first-order difference continuity constraint, the second-order difference smoothing constraint, the monotonically increasing constraint of remaining power and event capacity, and the battery parameter sorting constraint, respectively.

3. The method for constructing a large-scale battery model based on multi-source heterogeneity and physical constraints according to claim 2, characterized in that, The step of training the pre-trained large model of the physically constrained battery using the multi-source heterogeneous dataset to obtain the initial large model of the physically constrained battery includes: The multi-source heterogeneous dataset is input into the data encoding layer for training of temporal data encoding, local block unit encoding, and positional encoding to obtain structured representation data; The structured representation data is input into the physical constraint mask layer to train the physical constraint partial mask, thereby obtaining masked data segments and unmasked data segments; The unmasked data segment is input into the encoder layer for deep feature extraction training to obtain the encoded output features; The encoded output features are input into the hidden space layer for latent space representation processing and training to obtain a latent space vector. The masked data fragments are input into the decoder layer based on the latent spatial vector and mask label to perform data reconstruction training, thereby obtaining the initial physical constraint battery large model.

4. The method for constructing a large-scale battery model based on multi-source heterogeneity and physical constraints according to claim 1, characterized in that, The step of performing hierarchical learning rate decay and weight decay coordinated fine-tuning on the initial physically constrained battery large model based on the target task label to obtain the target physically constrained battery large model includes: Based on the target task label, the encoder of the initial physical constraint battery large model is transferred to the target task, and a loss function is constructed; The encoder's hierarchical learning rate decay parameter is adjusted; The objective function weight coefficients of the initial physical constraint battery model are weighted and decayed based on the loss function to obtain the target physical constraint battery model.

5. The method for constructing a large-scale battery model based on multi-source heterogeneity and physical constraints according to claim 1, characterized in that, The process of constructing multi-dimensional profile labels based on the cloud battery data to obtain the first training dataset includes: The cloud-based battery data is used to construct a multi-dimensional profile tag vector based on battery operating conditions. The multi-dimensional profile tag vector includes at least one or more of the following: regional information tag, battery mileage range tag, charging number range tag, charging ratio tag, charge / discharge depth tag, battery model tag, battery chemical system tag, and operating temperature range tag. The cloud battery data is divided into multiple data hierarchical units based on the multidimensional profile label vector, and the sampling weight of each data hierarchical unit is defined to obtain the first training dataset.

6. The method for constructing a large-scale battery model based on multi-source heterogeneity and physical constraints according to claim 1, characterized in that, The standardization process based on the laboratory battery data to obtain the second training dataset includes: The laboratory battery data was sliced ​​to obtain multiple charge and discharge event datasets; Each of the charging and discharging event datasets is preprocessed to obtain multiple preprocessed datasets. The data preprocessing includes at least data cleaning, data resampling, and time alignment. The degradation stage is identified for each of the preprocessed datasets to obtain the second training dataset.

7. The method for constructing a large-scale battery model based on multi-source heterogeneity and physical constraints according to claim 1, characterized in that, The step of aligning feature dimensions based on the first training dataset and the second training dataset to obtain a multi-source heterogeneous dataset includes: Construct a unified feature space, which includes at least excitation features, response features, mask features, and encoding features; Perform dimension alignment mapping on the second training dataset to generate a unified representation dataset; Based on the unified feature space, feature alignment is performed on the unified representation dataset and the first training dataset to obtain the multi-source heterogeneous dataset.

8. A battery large-scale model construction system based on multi-source heterogeneity and physical constraints, characterized in that, The system includes: The first data processing module is used to acquire cloud battery data, construct multi-dimensional profile labels based on the cloud battery data, and obtain the first training dataset. The second data processing module is used to acquire laboratory battery data and perform standardization processing on the laboratory battery data to obtain a second training dataset. The feature dimension alignment module is used to perform feature dimension alignment based on the first training dataset and the second training dataset to obtain a multi-source heterogeneous dataset. The model supervised training module is used to construct a pre-trained large model of a physically constrained battery. The pre-trained large model of the physically constrained battery is trained using the multi-source heterogeneous dataset to obtain an initial physical constrained battery large model. The model parameter fine-tuning module is used to perform hierarchical learning rate decay and weight decay coordinated fine-tuning on the initial physical constraint battery large model based on the target task label, so as to obtain the target physical constraint battery large model.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the battery large model construction method based on multi-source heterogeneity and physical constraints as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the battery large model construction method based on multi-source heterogeneity and physical constraints as described in any one of claims 1-7.