Energy storage battery health status prediction method and device based on federated learning

By performing two clustering and gradient algorithm collaborative training on the energy storage battery operation data based on a federated learning method, the problem of insufficient energy storage battery data in new energy power stations was solved, and efficient and accurate health status prediction and data privacy protection were achieved.

CN119988984BActive Publication Date: 2025-09-12BEIJING YINHU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510458068.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-09-12
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

There is little adverse health data on energy storage batteries in new energy power stations, which affects the accuracy of evaluation model training. In addition, data is difficult to share between power stations, leading to privacy protection and corporate competition issues.

Method used

A federated learning-based method is used to cluster the energy storage battery operation data twice, automatically screen negative and positive samples, and use the cloud server aggregate gradient algorithm to collaboratively train the local model to ensure a balanced ratio of positive and negative samples and protect data privacy.

Benefits of technology

The accuracy and training efficiency of the energy storage battery health prediction model are improved, the cost of data sharing is reduced, and data privacy and model effectiveness are guaranteed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988984B_ABST
    Figure CN119988984B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for predicting the health status of energy storage batteries based on federated learning, which belongs to the field of battery management. The method includes: for each power zone in the current new energy power station, executing: clustering the energy storage battery operation data samples of the power zone twice to respectively screen out the negative samples and positive samples of the current new energy power station data set; based on the federated learning framework, each new energy power station uses its own data set to train a local model, and the cloud server uses the aggregated gradient algorithm to aggregate and update the training results and gradient information of the local models of each new energy power station, and sends them to each new energy power station for local model update until a target prediction model that meets expectations is obtained. This solution can automatically screen out positive and negative samples with a balanced ratio, ensure the number of negative samples, and improve the effectiveness of each local model. In addition, it can protect the data privacy of each new energy power station while improving the accuracy and training efficiency of the final prediction model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of battery management technology, and in particular to a method and device for predicting the health status of an energy storage battery based on federated learning. Background Art

[0002] New energy is an important pillar of the future sustainable energy system and will gradually replace traditional petrochemical energy and occupy a dominant position.

[0003] Because acquiring adverse health data for energy storage batteries in new energy power plants is costly, and current data on energy storage batteries primarily comes from batteries in new energy electric vehicles, which can vary widely from battery to battery. Consequently, there is limited adverse health data for energy storage batteries in new energy power plants, impacting the accuracy of evaluation model training. Furthermore, for privacy reasons and to mitigate corporate competition, adverse health data from various new energy power plants is difficult to share.

[0004] Therefore, there is an urgent need for a method to predict the health status of energy storage batteries based on federated learning. Summary of the Invention

[0005] This invention provides a method and device for predicting the health status of energy storage batteries based on federated learning, which can solve the problem that there is a lack of adverse health data for energy storage batteries in new energy power plants, which affects the accuracy of evaluation model training. The technical solution is as follows:

[0006] In one aspect, a method for predicting the health status of an energy storage battery based on federated learning is provided, the method comprising:

[0007] For each power zone in the current new energy power station, the following steps are performed: cluster the energy storage battery operation data samples in the power zone, record the samples in the cluster categories other than the cluster category with the largest number of samples as negative samples, and add them to the data set of the current new energy power station;

[0008] Aggregate the samples in the cluster category with the largest number of samples in each power zone for secondary clustering, and filter the positive samples to the dataset of the current new energy power station based on the preset negative sample ratio, the total number of negative samples in the dataset at this time, and the secondary clustering results;

[0009] Based on the federated learning framework, each new energy power station uses its own data set to train a local model. The cloud server uses the aggregated gradient algorithm to aggregate and update the training results and gradient information of the local models of each new energy power station, and sends it to each new energy power station for local model updates until a target prediction model that meets the expectations is obtained.

[0010] On the other hand, a device for predicting the health status of an energy storage battery based on federated learning is provided, for implementing the method described in any embodiment of the specification, the device comprising:

[0011] The primary clustering unit is used to perform the following operations for each power zone in the current new energy power station: clustering the energy storage battery operation data samples in the power zone, recording the samples in the cluster categories other than the cluster category with the largest number of samples as negative samples, and adding them to the data set of the current new energy power station;

[0012] The secondary clustering unit is used to aggregate the samples in the cluster category with the largest number of samples in each power zone for secondary clustering, so as to filter the positive samples to the data set of the current new energy power station based on the preset negative sample ratio, the total number of negative samples in the data set at this time, and the secondary clustering results;

[0013] The training unit is used to train local models based on the federated learning framework. Each new energy power station uses its own data set to train a local model. The cloud server uses the aggregated gradient algorithm to aggregate and update the training results and gradient information of the local models of each new energy power station, and sends them to each new energy power station for local model updates until a target prediction model that meets the expectations is obtained.

[0014] On the other hand, a computer device is provided, comprising a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to implement the steps of the above-mentioned method.

[0015] On the other hand, a computer-readable storage medium is provided, wherein a computer program is stored in the storage medium, and when the computer program is executed by a processor, the steps of the above-mentioned method are implemented.

[0016] On the other hand, a computer program product is provided, comprising a computer program, wherein the computer program implements the steps of the above method when executed by a processor.

[0017] The technical solution provided by the present invention can at least bring the following beneficial effects:

[0018] Since most of the energy storage battery operation data are positive samples, and the model training needs to balance the ratio of positive and negative samples, when the ratio of positive and negative samples is unbalanced, the trained model will lose its predictive significance, greatly reducing the model performance. The screening of positive and negative samples of the operation data often requires manual labeling and selection, which involves a large amount of data and is time-consuming and labor-intensive. This solution clusters the energy storage battery operation data samples twice, and automatically screens out negative samples and positive samples based on the results of the two clusterings, which is highly efficient. In addition, the samples in the cluster categories other than the cluster category with the largest number of samples in the first clustering result are recorded as negative samples, which can maximize the guarantee of a sufficient number of negative samples. Based on the pre-set negative sample ratio, the total number of negative samples in the data set at this time, and the secondary clustering results, the positive samples are screened to ensure a balanced ratio of positive and negative samples and improve the effectiveness of each local model.

[0019] In addition, the federated learning framework allows each new energy power station to collaboratively train the prediction model by sharing the underlying model parameters. While protecting the data privacy and sensitivity of each new energy power station, it can save the cost of training the model and improve the accuracy and training efficiency of the final prediction model. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0021] Figure 1 This is a flow chart of a method for predicting the health status of an energy storage battery based on federated learning provided by one embodiment of the present invention;

[0022] Figure 2 This is a structural diagram of a device for predicting the health status of an energy storage battery based on federated learning provided by one embodiment of the present invention;

[0023] Figure 3 This is a hardware architecture diagram of a computer device provided by one embodiment of the present invention. DETAILED DESCRIPTION

[0024] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0025] The specific implementation of the above concept is described below.

[0026] Please refer to Figure 1 An embodiment of the present invention provides a method for predicting the health status of an energy storage battery based on federated learning, the method comprising:

[0027] Step 100: For each power zone in the current new energy power station, perform the following steps: cluster the energy storage battery operation data samples in the power zone, record the samples in the cluster categories other than the cluster category with the largest number of samples as negative samples, and add them to the data set of the current new energy power station;

[0028] Step 102: Aggregate the samples in the cluster category with the largest number of samples in each power zone and perform secondary clustering to filter the positive samples to the dataset of the current new energy power station based on the preset negative sample ratio, the total number of negative samples in the dataset at this time, and the secondary clustering results;

[0029] In step 104, based on the federated learning framework, each new energy power station uses its own data set to train a local model. The cloud server uses the aggregated gradient algorithm to aggregate and update the training results and gradient information of the local models of each new energy power station, and sends them to each new energy power station for local model update until a target prediction model that meets the expectations is obtained.

[0030] In the embodiment of the present invention, since most of the energy storage battery operation data are positive samples, and the training of the model needs to balance the ratio of positive and negative samples, when the ratio of positive and negative samples is unbalanced, the trained model will lose its predictive significance, greatly reducing the model performance. The screening of positive and negative samples of the operation data often requires manual labeling and selection, which involves a large amount of data and is time-consuming and labor-intensive. This solution clusters the energy storage battery operation data samples twice, and automatically screens out negative samples and positive samples based on the results of the two clusterings, which is highly efficient. In addition, the samples in the cluster categories other than the cluster category with the largest number of samples in the first clustering result are recorded as negative samples, which can maximize the guarantee of a sufficient number of negative samples. Based on the pre-set negative sample ratio, the total number of negative samples in the data set at this time, and the secondary clustering results, the positive samples are screened, which can ensure a balanced ratio of positive and negative samples and improve the effectiveness of each local model.

[0031] In addition, the federated learning framework allows each new energy power station to collaboratively train the prediction model by sharing the underlying model parameters. While protecting the data privacy and sensitivity of each new energy power station, it can save the cost of training the model and improve the accuracy and training efficiency of the final prediction model.

[0032] Described below Figure 1 How to perform the steps shown.

[0033] For step 100:

[0034] In step 100, the energy storage battery operation data samples for each power zone are clustered separately. Assuming there are three new energy power stations, each with two power zones, the negative sample generation process for each new energy power station's data set is as follows: first, cluster the first power zone of the new energy power station, record the samples in the cluster categories other than the cluster category with the largest number of samples as negative samples, and add them to the data set of the current new energy power station; then, cluster the second power zone of the new energy power station, record the samples in the cluster categories other than the cluster category with the largest number of samples as negative samples, and also add them to the data set of the current new energy power station, to obtain the negative samples of the current new energy power station. Since all samples in the cluster categories other than the cluster category with the largest number of samples are recorded as negative samples, the number of negative samples can be guaranteed.

[0035] Regarding step 102:

[0036] In some embodiments, step 102 includes:

[0037] Aggregate the samples in the cluster category with the largest number of samples in each electrical area for secondary clustering. The number of categories in the secondary clustering result is recorded as S;

[0038] Determine the required number of positive samples based on the preset negative sample ratio and the total number of negative samples in the dataset at this time;

[0039] Based on the required number of positive samples and the number of categories in the secondary clustering results, the number of samples selected from each category is determined to obtain the positive samples of the current new energy power station.

[0040] In this embodiment, the positive sample generation process in the data set of each new energy power station is: first, the samples in the cluster category with the largest number of samples in the first clustering results of the two power areas of the new energy power station are aggregated together, and then secondary clustering is performed to group samples with similar operating data change trends. Since the negative samples have been determined in step 100, in order to ensure a balanced ratio of positive and negative samples, it is necessary to determine the required number of positive samples based on the pre-set negative sample ratio and the total number of negative samples in the data set of the new energy power station at this time. The negative sample ratio can be set to 1 / 2 or 2 / 5 to avoid an imbalance in the ratio. Then, based on the required number of positive samples and the number of categories in the secondary clustering results, the number of samples to be selected from each category can be determined to obtain the positive samples of the current new energy power station.

[0041] In some embodiments, the number of samples selected from each category is equal, and the number of samples selected from each category is determined by the following formula:

[0042]

[0043] Where, The number of samples selected for each category, is the total number of negative samples in the dataset, is the preset negative sample ratio, is the number of categories in the secondary clustering result.

[0044] In this embodiment, in order to improve the generalization ability of the model, samples of various changing trends must be covered. They cannot be randomly selected from the secondary clustering samples according to the required number of positive samples. They can be evenly selected from each category according to the required number of positive samples. This can ensure that the number of samples for each changing trend is the same or similar, and the generated local model is more effective.

[0045] Regarding step 104:

[0046] In some implementations, step 104 may include:

[0047] Each renewable energy power station uses its own data set to train a local model, calculate the gradient information of the tree nodes, and encrypt the gradient information and send it to the cloud server;

[0048] The cloud server uses the aggregated gradient algorithm to map the sample features of the current node to the buckets one by one, aggregate the gradient information statistics, and encrypt the aggregated results and send them to each new energy power station;

[0049] Each new energy power station decrypts the aggregated results, determines the sample features based on the optimal segmentation algorithm, and sends them to the cloud server;

[0050] The cloud server determines the threshold of sample features based on the sample features, splits the tree nodes, and sends the split results to each new energy power station;

[0051] Each new energy power station splits its nodes based on the splitting results of the cloud server, and sends the information gain of the split child nodes to the cloud server for updating until the splitting is completed and a target prediction model that meets the expectations is obtained.

[0052] In this example, a gradient boosting decision tree is used as the target prediction model, and training is performed by iteratively optimizing the loss of the true label and the predicted label. In each iteration, a new tree is added to minimize the loss without increasing the complexity of the model. Therefore, the objective function for the tth iteration is set as:

[0053]

[0054] in, Refers to the prediction result of sample i after the tth round of iteration, Represents the prediction results of the first t-1 trees, Represents the model of the t-th tree. Loss function By real value y i and predicted values express:

[0055]

[0056] Where n is the number of samples, is the true value and is the predicted value.

[0057] The goal of gradient boosting decision trees is to iteratively optimize the loss between true and predicted values. In each iteration, a new tree is added to minimize the loss without increasing the model's complexity. To achieve high model prediction accuracy, the model's coefficient of variance should be small. Therefore, adding a regularization term to the training objective function can reduce overfitting. Therefore, the objective function L is set to consist of two parts: the training loss and the regularization term:

[0058]

[0059] The loss function L can be obtained by using the second-order Taylor expansion. Therefore, the objective function L of the t-th iteration is set as follows:

[0060]

[0061] in, is the loss function, and are the first-order gradient and second-order gradient of the loss function, respectively. Represents the tree model of the t-th round algorithm iteration, Represents the complexity of the new tree, where T represents the number of leaf nodes, represents the weight of the leaf node, represents a hyperparameter.

[0062] In this embodiment, the cloud server performs gradient aggregation in the following manner:

[0063]

[0064] Where, is the gradient aggregation result, is the number of new energy power stations, The gradient information of the current node sent by the a-th new energy power station, is the average loss of the local model of the new energy power station, is the loss value of the local model of the new energy power station in this round of update, is the loss value of the last update of the local model of the new energy power station, is a function, When it is negative, Get zero, When non-negative, equal .

[0065] In this embodiment, if the loss value of the current update is greater than the loss value of the previous round, the local model is invalidated, its gradient weight is set to 0, and the contribution of the local model to the global model is reduced. If the loss value of the current update is less than the loss value of the previous round, the weight is calculated based on the average loss, thereby improving the effectiveness of the model and training efficiency.

[0066] In this embodiment, the information gain is calculated using the following formula:

[0067]

[0068] Where, is the sample gradient information, U is the sample space, is the number of samples in the sample space of the current node, and are the number of samples of the left child node and the number of samples of the right child node, respectively. is the important sample set retained for the left child node, It is the important sample set retained for the right child node.

[0069] In this embodiment, information gain calculation is performed using gradient information containing weights, which can ensure the effectiveness and training efficiency of the model.

[0070] Please refer to Figure 2 The embodiment of the present invention provides a device for predicting the health status of an energy storage battery based on federated learning, which is used to implement the method of any embodiment of the specification. The device includes:

[0071] The primary clustering unit 201 is configured to perform, for each power zone in the current new energy power station, the following operations: clustering the energy storage battery operation data samples in the power zone, recording the samples in the cluster categories other than the cluster category with the largest number of samples as negative samples, and adding them to the data set of the current new energy power station;

[0072] The secondary clustering unit 202 is used to aggregate the samples in the cluster category with the largest number of samples in each power zone for secondary clustering, so as to filter the positive samples into the data set of the current new energy power station based on the preset negative sample ratio, the total number of negative samples in the data set at this time, and the secondary clustering results;

[0073] The training unit 203 is used to train the local model of each new energy power station using its own data set based on the federated learning framework. The cloud server uses the aggregated gradient algorithm to aggregate and update the training results and gradient information of the local models of each new energy power station, and sends them to each new energy power station for local model update until a target prediction model that meets the expectations is obtained.

[0074] In one embodiment of the present invention, the secondary clustering unit 202 is configured to perform:

[0075] Aggregate the samples in the cluster category with the largest number of samples in each electrical area for secondary clustering. The number of categories in the secondary clustering result is recorded as S;

[0076] Determine the required number of positive samples based on the preset negative sample ratio and the total number of negative samples in the dataset at this time;

[0077] Based on the required number of positive samples and the number of categories in the secondary clustering results, the number of samples selected from each category is determined to obtain the positive samples of the current new energy power station.

[0078] In one embodiment of the present invention, the number of samples selected from each category in the secondary clustering unit 202 is equal, and the number of samples selected from each category is determined by the following formula:

[0079]

[0080] Where, The number of samples selected for each category, is the total number of negative samples in the dataset, is the preset negative sample ratio, is the number of categories in the secondary clustering result.

[0081] In one embodiment of the present invention, the training unit 203 is configured to perform:

[0082] Each renewable energy power station uses its own data set to train a local model, calculate the gradient information of the tree nodes, and encrypt the gradient information and send it to the cloud server;

[0083] The cloud server uses the aggregated gradient algorithm to map the sample features of the current node to the buckets one by one, aggregate the gradient information statistics, and encrypt the aggregated results and send them to each new energy power station;

[0084] Each new energy power station decrypts the aggregated results, determines the sample features based on the optimal segmentation algorithm, and sends them to the cloud server;

[0085] The cloud server determines the threshold of sample features based on the sample features, splits the tree nodes, and sends the split results to each new energy power station;

[0086] Each new energy power station splits its nodes based on the splitting results of the cloud server, and sends the information gain of the split child nodes to the cloud server for updating until the splitting is completed and a target prediction model that meets the expectations is obtained.

[0087] In one embodiment of the present invention, the cloud server in the training unit 203 performs gradient aggregation in the following manner:

[0088]

[0089] Where, is the gradient aggregation result, is the number of new energy power stations, The gradient information of the current node sent by the a-th new energy power station, is the average loss of the local model of the new energy power station, is the loss value of the local model of the new energy power station in this round of update, is the loss value of the last update of the local model of the new energy power station, is a function, When it is negative, Get zero, When non-negative, equal .

[0090] In one embodiment of the present invention, the information gain in the training unit 203 is calculated using the following formula:

[0091]

[0092] Where, is the sample gradient information, U is the sample space, is the number of samples in the sample space of the current node, and are the number of samples of the left child node and the number of samples of the right child node, respectively. is the important sample set retained for the left child node, It is the important sample set retained for the right child node.

[0093] It should be noted that the above-mentioned embodiment of the energy storage battery health prediction device based on federated learning is only illustrated by the division of the above-mentioned functional units. In actual applications, the above-mentioned functions can be assigned to different functional units as needed, that is, the internal structure of the device can be divided into different functional units to complete all or part of the functions described above. In addition, the above-mentioned device embodiment and the method embodiment are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0094] The embodiment of the present application also provides a computer device, please refer to Figure 3 The computer device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set, or instruction set, and the at least one instruction, at least one program, code set, or instruction set is loaded and executed by the processor to implement the energy storage battery health status prediction method based on federated learning provided by the above-mentioned method embodiments.

[0095] An embodiment of the present application also provides a computer-readable storage medium, on which is stored at least one instruction, at least one program, code set, or instruction set, and the at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor to implement the energy storage battery health status prediction method based on federated learning provided in the above-mentioned method embodiments.

[0096] An embodiment of the present application also provides a computer program product, which includes a computer program. A processor of a computer device reads the computer program from a computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the energy storage battery health status prediction method based on federated learning described in any of the above embodiments.

[0097] For the convenience of description, the above systems or devices are described as being divided into various modules or units according to their functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0098] Through the description of the above embodiments, it can be seen that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application or certain parts of the embodiments.

[0099] Finally, it should be noted that, in this document, relational terms such as first, second, third, and fourth are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.

[0100] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A method for predicting the health status of energy storage batteries based on federated learning, characterized in that: include: For each power zone in the current new energy power station, the following steps are performed: cluster the energy storage battery operation data samples in the power zone, record the samples in the cluster categories other than the cluster category with the largest number of samples as negative samples, and add them to the data set of the current new energy power station; Aggregate the samples in the cluster category with the largest number of samples in each power zone for secondary clustering, and filter the positive samples to the dataset of the current new energy power station based on the preset negative sample ratio, the total number of negative samples in the dataset at this time, and the secondary clustering results; Based on the federated learning framework, each renewable energy power station uses its own data set to train a local model. The cloud server uses the aggregated gradient algorithm to aggregate and update the training results and gradient information of each renewable energy power station's local model, and sends it to each renewable energy power station for local model update until a target prediction model that meets the expectations is obtained. The samples in the cluster category with the largest number of samples in each power zone are aggregated for secondary clustering, so as to filter positive samples to the data set of the current new energy power station based on the preset negative sample ratio, the total number of negative samples in the data set at this time, and the secondary clustering results, including: Aggregate the samples in the cluster category with the largest number of samples in each electrical area for secondary clustering. The number of categories in the secondary clustering result is recorded as S; Determine the required number of positive samples based on the preset negative sample ratio and the total number of negative samples in the dataset at this time; Based on the required number of positive samples and the number of categories in the secondary clustering results, the number of samples to be selected from each category is determined to obtain the positive samples of the current new energy power station; The number of samples selected in each category is equal, and the number of samples selected in each category is determined by the following formula: Where, The number of samples selected for each category, is the total number of negative samples in the dataset, is the preset negative sample ratio, is the number of categories in the secondary clustering result; Based on the federated learning framework, each new energy power station uses its own data set to train a local model. The cloud server uses the aggregated gradient algorithm to aggregate and update the training results and gradient information of the local models of each new energy power station, and sends it to each new energy power station for local model update until a target prediction model that meets the expectations is obtained, including: Each new energy power station uses its own data set to train a local model, calculate the gradient information of the tree nodes, and encrypt the gradient information and send it to the cloud server; The cloud server uses the aggregated gradient algorithm to map the sample features of the current node to the buckets one by one, aggregates the gradient information statistics, and encrypts the aggregated results and sends them to each new energy power station; Each new energy power station decrypts the aggregation results, determines the sample features based on the optimal segmentation algorithm, and sends them to the cloud server; The cloud server determines the threshold of the sample features based on the sample features, splits the tree nodes, and sends the split results to each new energy power station; Each new energy power station performs node splitting based on the splitting result of the cloud server, and sends the information gain of the split child node to the cloud server for updating until the splitting is completed, and obtains the expected target prediction model.

2. The method according to claim 1, wherein The cloud server performs gradient aggregation in the following way: Where, is the gradient aggregation result, is the number of new energy power stations, The gradient information of the current node sent by the a-th new energy power station, is the average loss of the local model of the new energy power station, is the loss value of the local model of the new energy power station in this round of update, is the loss value of the last update of the local model of the new energy power station, is a function, When it is negative, Get zero, When non-negative, equal .

3. The method according to claim 1, wherein Information gain is calculated using the following formula: Where, is the sample gradient information, U is the sample space, is the number of samples in the sample space of the current node, and are the number of samples of the left child node and the number of samples of the right child node, respectively. is the important sample set retained for the left child node, is the important sample set retained for the right child node.

4. A device for predicting the health status of an energy storage battery based on federated learning, used to implement the method according to any one of claims 1 to 3, characterized in that: The device comprises: The primary clustering unit is used to perform the following operations for each power zone in the current new energy power station: clustering the energy storage battery operation data samples in the power zone, recording the samples in the cluster categories other than the cluster category with the largest number of samples as negative samples, and adding them to the data set of the current new energy power station; The secondary clustering unit is used to aggregate the samples in the cluster category with the largest number of samples in each power zone for secondary clustering, so as to filter the positive samples to the data set of the current new energy power station based on the preset negative sample ratio, the total number of negative samples in the data set at this time, and the secondary clustering results; The training unit is used to train local models based on the federated learning framework. Each new energy power station uses its own data set to train a local model. The cloud server uses the aggregated gradient algorithm to aggregate and update the training results and gradient information of the local models of each new energy power station, and sends them to each new energy power station for local model updates until a target prediction model that meets the expectations is obtained.

5. A computer device, characterized in that: The computer device includes a memory and a processor, the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to implement the steps of any one of the methods described in claims 1-3.

6. A computer-readable storage medium, characterized in that The storage medium stores a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 3.

7. A computer program product, characterized in that The method comprises a computer program, wherein when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 3 are implemented.

Citation Information

Patent Citations

  • Clustering-based picture classification method and system, device and storage medium

    CN111507380A

  • New energy power station operation data intelligent analysis method and analysis system

    CN119134311A