A Cross-Operating Condition Degradation Trend Hierarchical Federal Prediction Method Based on Ring Distillation

Through the ring distillation method, the prediction accuracy reduction caused by working conditions is solved in the federal system, and the cross-work conditions product degradation trend is achieved, and the flexibility and expansion capabilities of the federal prediction system are improved.

CN118839799BActive Publication Date: 2025-07-18BEIHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311377103.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-23
Publication Date
2025-07-18
Estimated Expiration
2043-10-23

AI Technical Summary

Technical Problem

The prior art is difficult to effectively combine the differences in the use conditions of individual products and common degradation laws in the federal system, resulting in a decrease in prediction accuracy and low learning efficiency.

Method used

A hierarchical federal prediction method for cross-condition degradation trend based on annular distillation is adopted, through the common law mining stage and the individual characteristic adaptation stage, and using knowledge distillation to share and personalize the adaptation model parameters under different operating conditions, a model collaboration mechanism is established that shares common degradation laws and individual characteristic adaptation across operating conditions and parallelizes the common degradation laws of cross-condition products.

Benefits of technology

It improves the flexibility and expansion capabilities of the federal prediction system, realizes coordinated prediction from similar degradation laws in single working conditions to complex degradation laws in multiple working conditions, and improves the accuracy and generalization capabilities of the prediction model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118839799B_ABST
    Figure CN118839799B_ABST
Patent Text Reader

Abstract

The present invention discloses a cross-condition degradation trend hierarchical federated prediction method based on annular distillation, including: in the common law mining stage, client clusters under different conditions participate in each communication round in turn, and clients under the same condition share model parameters. After reaching the maximum communication round of this stage, the knowledge information suitable for the deep degradation laws of different conditions is retained in the model of each condition cluster, which can be used to supervise the update of subsequent personalized models. In the individual characteristic adaptation stage, each condition cluster generates a personalized prediction model for different clients through a super network, and iteratively optimizes the model through knowledge distillation, ultimately achieving accurate prediction of the degradation trend of individual products.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of federated prediction, and particularly to a cross-condition degradation trend hierarchical federated prediction method based on circular distillation. Background Art

[0002] The core of the federated prediction of the degradation trend of population products is to design an efficient knowledge transfer system to help each individual in the federated system screen out applicable degradation knowledge from other rich types of degradation trends, thereby improving the prediction accuracy of the degradation trend of individual products. Therefore, it is considered to further improve the prediction performance of the federated system by means of cross-condition product degradation data. However, in addition to the differences in degradation laws caused by the characteristics of individual products themselves, the distribution shift of degradation laws caused by different product usage conditions is also a realistic problem that urgently needs to be solved in the federated prediction of the degradation trend of population products. Although personalized federated learning methods can help clients obtain the required knowledge from similar models, a large number of clients lacking reference value will instead reduce the learning efficiency of the target client and affect the prediction accuracy of individual products. On the other hand, the model obtained by completely personalized training may fall into the optimization goal of the degradation data of a single product, and its generalization ability to other individuals is greatly reduced, which is not an ideal state for the federated system. Summary of the Invention

[0003] The present invention provides a cross-condition degradation trend hierarchical federated prediction method based on circular distillation, so as to solve the technical problem of how to obtain a common model with strong generalization ability and how to perform personalized adaptation of the model according to the data characteristics of each client, which will help improve the ability of the prediction model and the scalability of the federated architecture.

[0004] An embodiment of the present invention provides a cross-condition degradation trend hierarchical federated prediction method based on circular distillation, including:

[0005] Before each round of iteration in the common law mining stage, the server sequentially selects a working condition cluster in a cyclic manner, and uses the cluster sampling method to select multiple clients participating in the current round from all the clients in the working condition cluster;

[0006] Each client participating in the current round uses the working condition cluster prediction model of the previous round as the teacher model for knowledge distillation, and uses the working condition cluster prediction model of the current round as the student model for knowledge distillation, and locally trains the working condition cluster prediction model of the current round to obtain the parameters of the working condition cluster prediction model after the current round of training;

[0007] The server uses the federated average algorithm to perform federated aggregation processing on the parameters of the working condition cluster prediction model after the current round of training sent by the multiple clients to obtain the parameters of the working condition cluster prediction model after federated aggregation processing;

[0008] Repeat the steps in the common law mining stage until, after determining the maximum number of rounds, all clients under each working condition cluster obtain the working condition cluster prediction model parameters after federated aggregation processing by loading the maximum number of rounds sent by the server, and obtain a working condition cluster prediction model with common knowledge information in different working condition degradation data;

[0009] Before each iteration of each working condition cluster in the individual characteristic adaptation stage, the server uses the cluster sampling method to select multiple clients participating in the current round from all clients in each working condition cluster, and sends the personalized prediction model parameters of the current round of the client generated by using the hypernetwork corresponding to each client to the corresponding client;

[0010] Each client participating in the current round uses the working condition cluster prediction model of the previous working condition cluster as the teacher model for knowledge distillation, and the personalized prediction model of the current round as the student model for knowledge distillation, locally trains the personalized prediction model of the current round, obtains the change amount of the personalized prediction model parameters of the current round, and sends the change amount of the personalized prediction model parameters of the current round to the server;

[0011] The server uses the change amount of the personalized prediction model parameters of each client in the current round to update the personalized prediction model parameters and the hypernetwork, and obtains the updated personalized prediction model parameters and hypernetwork;

[0012] Repeat the steps in the above individual characteristic adaptation stage until, after the server determines the maximum number of rounds, it sends the personalized prediction model parameters of each client obtained in the maximum number of rounds to the client, so that each client obtains a prediction model applicable to predicting the degradation trend of the individual device by loading the personalized prediction model parameters obtained in the maximum number of rounds.

[0013] Preferably, clients with the same working condition form a working condition cluster, and in the common law mining stage, the working condition cluster prediction models and working condition cluster prediction model parameters of all clients in the working condition cluster are the same. In the individual characteristic adaptation stage, the personalized prediction models and personalized prediction model parameters of all clients in the working condition cluster are different.

[0014] Preferably, each client participating in the current round uses the working condition cluster prediction model of the previous round as the teacher model for knowledge distillation, and the working condition cluster prediction model of the current round as the student model for knowledge distillation, and locally trains the working condition cluster prediction model of the current round. The obtained working condition cluster prediction model parameters after the current round of training include:

[0015] Each client participating in the current round receives the working condition cluster prediction model sent by the server in the previous round, uses the working condition cluster prediction model in the previous round as the teacher model for knowledge distillation, and uses the working condition cluster prediction model in the current round as the student model for knowledge distillation;

[0016] The client calculates the validation loss value of the student model using the local training set data;

[0017] When the validation loss value of the student model is less than the set first validation loss threshold, the client uses the first constraint to locally train the student model, and at the same time adds the constraint of the teacher model on the extracted degradation features during the training process, and uses the parameters of the trained student model as the parameters of the working condition cluster prediction model after the current round of training;

[0018] When the validation loss value of the student model is not less than the set first validation loss threshold, the client uses the second constraint to locally train the teacher model, and uses the parameters of the trained teacher model as the parameters of the working condition cluster prediction model after the current round of training.

[0019] Preferably, the client uses the first constraint to locally train the student model, and at the same time adding the constraint of the teacher model on the extracted degradation features during the training process includes:

[0020] L1 = L MSE + λ0L KD

[0021]

[0022] where L1 represents the training loss of the first constraint; L MSE represents the mean square error loss; λ0 represents a constant coefficient used to balance the attention of the model to the knowledge of the current data and the common knowledge from the teacher model; L KD represents the knowledge distillation loss; f tea / stu (x) represents the degradation features extracted by the teacher / student model encoder when taking x as the input.

[0023] Preferably, the client uses the second constraint to locally train the teacher model includes:

[0024] L2 = L M.SE

[0025] where L2 represents the training loss of the second constraint.

[0026] Preferably, each client participating in the current round locally trains the personalized prediction model of the current round by using the working condition cluster prediction model of the previous working condition cluster as the teacher model for knowledge distillation and the personalized prediction model of the current round as the student model for knowledge distillation. The obtained change amount of the personalized prediction model parameters of the current round includes:

[0027] Each client participating in the current round uses the pre-stored working condition cluster prediction model of the previous working condition cluster sent by the server as the teacher model for knowledge distillation and the personalized prediction model of the current round sent by the server as the student model for knowledge distillation;

[0028] The client respectively calculates the validation loss value of the student model and the validation loss value of the teacher model by using the local training set data;

[0029] If the validation loss value of the teacher model is less than the set second validation loss threshold or the validation loss value of the student model, the client locally trains the student model using the third constraint and adds the constraint of the teacher model on the extracted degraded features during the training process to obtain the change amount of the personalized prediction model parameters of the current round;

[0030] If the validation loss value of the teacher model is not less than the set second validation loss threshold and the validation loss value of the student model, the client locally trains the student model using the second constraint to obtain the change amount of the personalized prediction model parameters of the current round.

[0031] Preferably, the client locally trains the student model using the third constraint and adding the constraint of the teacher model on the extracted degraded features during the training process includes:

[0032] L3 = L MSE + λ1L KD

[0033]

[0034]

[0035] Among them, L3 represents the training loss of the third constraint; λ1 represents the distillation loss coefficient in the third constraint; represents the validation loss value of the student model; represents the validation loss value of the teacher model.

[0036] Preferably, the server sends the personalized prediction model parameters of the current round of the client generated by using the hypernetwork corresponding to each client to the corresponding client, including:

[0037] The server is equipped with a hypernetwork for each client to generate aggregated weights for each network layer of the personalized prediction model for the client, and generates a client aggregated weight matrix by using the hypernetwork corresponding to each client and the client embedding vector;

[0038] The server uses the client aggregated weight matrix to generate the personalized prediction model parameters for the current round of the client, and sends the generated personalized prediction model parameters for the current round of the client to the client.

[0039] Preferably, the server uses the client aggregated weight matrix to generate the personalized prediction model parameters for the current round of the client, including:

[0040]

[0041]

[0042] Among them, represents the personalized prediction model parameters for the current round of client i, represents the parameters of the l-th layer in the personalized prediction model; represents the parameters of the l-th layer in the personalized prediction model of client k; represents the weight of the l-th layer network in the personalized prediction model of the k-th client applied to the corresponding layer in client i; L represents the number of network layers in the personalized prediction model; K represents the number of clients in the federated system.

[0043] Preferably, the server updates the personalized prediction model parameters and the hypernetwork by using the change amount of the personalized prediction model parameters for each client in the current round, and the updated personalized prediction model parameters and hypernetwork obtained include:

[0044]

[0045]

[0046]

[0047]

[0048]

[0049] Among them, Δω k represents the change amount of the personalized prediction model parameters of client k in the current round; represents the personalized prediction model parameters of client k in the current round; Denote the personalized prediction model parameters of client k after the update in the current round; Ω represents the set of personalized prediction model parameters of all clients before the start of the current round; respectively denote the embedding vector e in the hypernetwork of client k k , the model parameters ψ in the hypernetwork k gradient; Δe k Denote the change in the embedding vector in the hypernetwork of client k in the current round; Denote the embedding vector in the hypernetwork of client k after the update in the current round; Denote the embedding vector in the hypernetwork of client k in the current round; Δψ k Denote the change in the personalized model parameters in the hypernetwork of client k in the current round; Denote the model parameters in the hypernetwork of client k after the update in the current round; Denote the model parameters in the hypernetwork of client k in the current round.

[0050] The beneficial effect of the present invention is that based on the degradation knowledge circular distillation in the common law mining stage and the individual characteristic adaptation stage, a model cooperation mechanism for sharing the common degradation laws of cross-condition products and parallel adaptation of individual characteristics can be established, effectively improving the flexibility and expansion ability of the federated prediction system, and realizing the evolution from collaborative prediction using single-condition similar degradation laws to collaborative prediction using multi-condition complex degradation laws. Brief Description of the Drawings

[0051] Figure 1 is a flowchart of a cross-condition degradation trend hierarchical federated prediction method based on circular distillation provided by the present invention;

[0052] Figure 2 is a structural diagram of a cross-condition degradation trend hierarchical federated prediction system based on circular distillation provided by the present invention;

[0053] Figure 3 is a detailed flowchart of a cross-condition degradation trend hierarchical federated prediction method based on circular distillation provided by the present invention;

[0054] Figure 4 is a schematic diagram of the battery capacity retention rate curves at three test temperatures provided by the present invention;

[0055] Figure 5 is a schematic diagram of the test error of the battery degradation trend personalized federated prediction comparison test provided by the present invention;

[0056] Figure 6 is a schematic diagram of the test error of the cross-condition battery degradation trend hierarchical federated prediction ablation test provided by the present invention;

[0057] Figure 7It is a schematic diagram of the test error of the cross-condition battery degradation trend hierarchical federated prediction comparison test provided by the present invention. Detailed implementation manners

[0058] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In the subsequent descriptions, suffixes such as "module", "component" or "unit" used to represent elements are only for the convenience of describing the present invention, and they have no specific meaning by themselves. Therefore, "module", "component" or "unit" can be used interchangeably.

[0059] Figure 1 It is a flowchart of a cross-condition degradation trend hierarchical federated prediction method based on annular distillation provided by the present invention, as Figure 1 shown, including:

[0060] Step S101: Before each round of iteration in the common law mining stage, the server sequentially selects a working condition cluster in a cyclic manner, and uses the cluster sampling method to select multiple clients participating in the current round from all the clients in the working condition cluster;

[0061] Step S102: Each client participating in the current round uses the working condition cluster prediction model of the previous round as the teacher model for knowledge distillation, and uses the working condition cluster prediction model of the current round as the student model for knowledge distillation, and locally trains the working condition cluster prediction model of the current round to obtain the parameters of the working condition cluster prediction model after the current round of training;

[0062] Step S103: The server uses the federated averaging algorithm to perform federated aggregation processing on the parameters of the working condition cluster prediction model after the current round of training sent by the multiple clients to obtain the parameters of the working condition cluster prediction model after federated aggregation processing;

[0063] Step S104: Repeat the steps in the above-mentioned common law mining stage until it is judged that the maximum number of rounds is reached. Then, all the clients under each working condition cluster obtain the working condition cluster prediction model with the general knowledge information in different working condition degradation data by loading the parameters of the working condition cluster prediction model after federated aggregation processing obtained in the maximum number of rounds sent by the server;

[0064] Step S105: Before each round of iteration of each working condition cluster in the individual characteristic adaptation stage, the server uses the cluster sampling method to select multiple clients participating in the current round from all the clients in each working condition cluster, and distributes the personalized prediction model parameters of the current round of the client generated by using the super network corresponding to each client to the corresponding client;

[0065] Step S106: Each client participating in the current round locally trains the personalized prediction model of the current round by using the working condition cluster prediction model of the previous working condition cluster as the teacher model for knowledge distillation and the personalized prediction model of the current round as the student model for knowledge distillation, obtains the change amount of the personalized prediction model parameters of the current round, and sends the change amount of the personalized prediction model parameters of the current round to the server side;

[0066] Step S107: The server side updates the personalized prediction model parameters and the hypernetwork by using the change amount of the personalized prediction model parameters of each client in the current round, and obtains the updated personalized prediction model parameters and hypernetwork;

[0067] Step S108: Repeat the steps of the above individual characteristic adaptation stage until after the server side determines that the maximum number of rounds is reached, the server side distributes the personalized prediction model parameters of each client obtained in the maximum number of rounds to the client, so that each client obtains a prediction model applicable to predicting the degradation trend of the individual device by loading the personalized prediction model parameters obtained in the maximum number of rounds.

[0068] Among them, clients with the same working conditions form a working condition cluster, and in the common law mining stage, the working condition cluster prediction models and working condition cluster prediction model parameters of all clients in the working condition cluster are the same. In the individual characteristic adaptation stage, the personalized prediction models and personalized prediction model parameters of all clients in the working condition cluster are different.

[0069] Further, each client participating in the current round locally trains the working condition cluster prediction model of the current round by using the working condition cluster prediction model of the previous round as the teacher model for knowledge distillation and the working condition cluster prediction model of the current round as the student model for knowledge distillation. The working condition cluster prediction model parameters after the current round of training obtained include: each client participating in the current round receives the working condition cluster prediction model of the previous round sent by the server side, and uses the working condition cluster prediction model of the previous round as the teacher model for knowledge distillation and the working condition cluster prediction model of the current round as the student model for knowledge distillation; the client calculates the validation loss value of the student model by using the local training set data; when the validation loss value of the student model is less than the set first validation loss threshold, the client locally trains the student model by using the first constraint, and at the same time adds the constraint of the teacher model on the extracted degradation features during the training process, and takes the student model parameters after training as the working condition cluster prediction model parameters after the current round of training; when the validation loss value of the student model is not less than the set first validation loss threshold, the client locally trains the teacher model by using the second constraint, and takes the teacher model parameters after training as the working condition cluster prediction model parameters after the current round of training.

[0070] Specifically, the client uses the first constraint to locally train the student model, and the constraints of the teacher model on the extracted degenerate features added during the training process include:

[0071] L1 = L MSE + λ0L KD

[0072]

[0073] where L1 represents the training loss of the first constraint; L MSE represents the mean square error loss; λ0 represents a constant coefficient used to balance the attention of the model to the current data knowledge and the common knowledge from the teacher model; L KD represents the knowledge distillation loss; f tea / stu (x) represents the degenerate features extracted by the teacher / student model encoder when taking x as the input.

[0074] Among them, the client uses the second constraint to locally train the teacher model, including:

[0075] L2 = L M.SE

[0076] where L2 represents the training loss of the second constraint.

[0077] Furthermore, each client participating in the current round locally trains the personalized prediction model of the current round by using the working condition cluster prediction model of the previous working condition cluster as the teacher model for knowledge distillation and the personalized prediction model of the current round as the student model for knowledge distillation, and the change amount of the parameters of the personalized prediction model of the current round obtained includes: each client participating in the current round uses the pre-stored working condition cluster prediction model of the previous working condition cluster sent by the server as the teacher model for knowledge distillation and the personalized prediction model of the current round sent by the server as the student model for knowledge distillation; the client respectively calculates the validation loss value of the student model and the validation loss value of the teacher model by using the local training set data; if the validation loss value of the teacher model is less than the set second validation loss threshold or the validation loss value of the student model, the client uses the third constraint to locally train the student model, and at the same time adds the constraints of the teacher model on the extracted degenerate features during the training process to obtain the change amount of the parameters of the personalized prediction model of the current round; if the validation loss value of the teacher model is not less than the set second validation loss threshold and the validation loss value of the student model, the client uses the second constraint to locally train the student model to obtain the change amount of the parameters of the personalized prediction model of the current round.

[0078] Specifically, the client uses the third constraint to locally train the student model, and the constraints imposed by the teacher model on the extracted degraded features during the training process include:

[0079] L3 = L MSE + λ1L KD

[0080]

[0081]

[0082] Among them, L3 represents the training loss of the third constraint; λ1 represents the distillation loss coefficient in the third constraint; represents the validation loss value of the student model; represents the validation loss value of the teacher model.

[0083] Furthermore, the server side distributes the personalized prediction model parameters of the current round of the client generated by the hypernetwork corresponding to each client to the corresponding client, including: The server side equips each client with a hypernetwork for generating aggregation weights for each network layer of the client's personalized prediction model, and generates a client aggregation weight matrix using the hypernetwork corresponding to each client and the client embedding vector; The server side uses the client aggregation weight matrix to generate the personalized prediction model parameters of the current round of the client, and distributes the generated personalized prediction model parameters of the current round of the client to the client.

[0084] Specifically, the server side uses the client aggregation weight matrix to generate the personalized prediction model parameters of the current round of the client, including:

[0085]

[0086]

[0087] Among them, represents the personalized prediction model parameters of the current round of client i, represents the parameters of the l-th layer in the personalized prediction model; represents the parameters of the l-th layer in the personalized prediction model of client k; represents the weight of the l-th layer network in the personalized prediction model of the k-th client applied to the corresponding layer in client i; L represents the number of layers of the network in the personalized prediction model; K represents the number of clients in the federated system.

[0088] Specifically, the server side updates the personalized prediction model parameters and the hypernetwork by using the change amount of the personalized prediction model parameters of each client in the current round, and obtaining the updated personalized prediction model parameters and hypernetwork includes:

[0089]

[0090]

[0091]

[0092]

[0093]

[0094] where, Δω k represents the change amount of the personalized prediction model parameters of client k in the current round; represents the personalized prediction model parameters of client k in the current round; represents the personalized prediction model parameters of client k after update in the current round; Ω represents the set of personalized prediction model parameters of all clients before the start of the current round; respectively represent the embedding vector e in the hypernetwork of client k k , the model parameters ψ in the hypernetwork k 's gradient; Δe k represents the change amount of the embedding vector in the hypernetwork of client k in the current round; represents the embedding vector in the hypernetwork of client k after update in the current round; represents the embedding vector in the hypernetwork of client k in the current round; Δψ k represents the change amount of the personalized model parameters in the hypernetwork of client k in the current round; represents the model parameters in the hypernetwork of client k after update in the current round; represents the model parameters in the hypernetwork of client k in the current round.

[0095] The embodiment of the present invention proposes a CyclicDistillation based Hierarchical Federated Learning (CDHFL) algorithm. Individuals under the same working condition form a client cluster, and clusters under different working conditions further form a cyclic federated system. The prediction models with common degradation knowledge are transferred between cyclic federations through knowledge distillation, while accurate prediction models applicable to different individuals are generated inside the working condition clusters, such as Figures 2 - 3As shown. Specifically, CDHFL includes two main stages, the common law mining stage and the individual characteristic adaptation stage. In the first stage, individuals under the same working conditions use the same cluster prediction model. In each communication round, a working condition cluster is sequentially selected for model aggregation. At the same time, the prediction model formed by the previous working condition cluster will be used as the teacher model for the next working condition cluster. By means of knowledge distillation, the common degradation law information in the teacher model is efficiently utilized, and the general knowledge in the degradation data of different working conditions is solidified into the prediction model through multiple rounds of iteration, improving the generalization ability of the prediction models for each working condition cluster. In the second stage, within each working condition cluster, a personalized federated learning method based on model hierarchical aggregation is used. Based on the cluster prediction model obtained in the first stage, a personalized prediction model is established for each client. At the same time, the teacher model is still used to constrain the training of the local personalized model to prevent the loss of common degradation knowledge. Based on the circular distillation of degradation knowledge in the above two stages, a model collaboration mechanism for sharing common product degradation laws across working conditions and adapting individual characteristics in parallel can be established, effectively improving the flexibility and scalability of the federated prediction system, and realizing the evolution from collaborative prediction using similar degradation laws of a single working condition to collaborative prediction using complex degradation laws of multiple working conditions.

[0096] CDHFL aims to accumulate common product degradation laws among multiple client clusters divided according to working conditions through knowledge distillation and further adapt the individual degradation data of each client. Therefore, hierarchical federated circular distillation contains two stages: the common law mining stage and the individual characteristic adaptation stage.

[0097] In the common law mining stage, each working condition cluster is trained sequentially, and the cluster model trained in the previous step will be used as the teacher for the next cluster. After the common law mining of the client degradation data under different working condition clusters is completed, it enters the individual characteristic adaptation stage. The individual characteristic adaptation stage is carried out for each working condition, and the clients under the same working condition are trained. On this basis, in order not to lose the common degradation knowledge, the constraint of the cluster model obtained in the common law mining stage will be added during the model training process.

[0098] Different from traditional output-oriented knowledge distillation that learns a simplified model by imitating the output of the teacher model on the same data, both of the above two training stages adopt feature-oriented knowledge distillation, guiding the training of the student model by measuring the difference in the features extracted by the teacher model and the student model. During the calculation of the model training loss, in addition to considering the basic mean square error loss, the knowledge distillation loss is added.

[0099] (1) Common law mining stage

[0100] The main purpose of the common pattern mining stage is to deeply explore the underlying logic of degradation time series data from seemingly different degradation manifestations across working conditions and individuals, screen the patterns of the degradation process in the time dimension, retain the deeper degradation patterns suitable for describing group products for modeling, so as to improve the expression ability of the prediction model for the common degradation patterns of the group, rather than losing the generalization ability for group products after easily fitting the surface degradation patterns of specific objects.

[0101] In the common pattern mining stage, clients under the same working condition use the same set of prediction model parameters. That is to say, the model parameters of all clients within the working condition cluster are the same as those of the cluster model. Different working condition clusters are trained sequentially in a cyclic manner, and a batch of clients are selected from the current working condition cluster by cluster sampling to participate in this communication round. The local training process of the client transfers the degradation patterns learned in the previous working condition cluster to the current working condition cluster through knowledge distillation, retains the useful knowledge, and discards the useless knowledge. After cyclic training, the knowledge applicable to the degradation trends of different individuals in each working condition, that is, the common degradation patterns, will be retained.

[0102] In the traditional local training process of the client, only local data is used to train the received model. However, in the local training process of the client in the common pattern mining stage, it is a knowledge distillation process. The model of the current working condition cluster is received as the student model, and the model of the previous working condition cluster is received as the teacher model, and the performance of the student model is verified. When is better than the set threshold t1, the client will use the student model for training, and the constraint of the teacher model on the extracted degradation features will be added during the training process, that is:

[0103] L = L MSE + λ0L KD (1)

[0104]

[0105] Among them, L MSE represents the mean square error loss; λ0 represents a constant coefficient used to balance the attention of the model to the knowledge of the current data and the common knowledge from the teacher model; L KD represents the knowledge distillation loss; f tea / stu (x) represents the degradation features extracted by the teacher / student model encoder when taking x as the input.

[0106] When it is the case, only the teacher model is used for local training:

[0107] L = L MSE (3)

[0108] (2) Individual Characteristic Adaptation Stage

[0109] The main purpose of the individual characteristic adaptation stage is to perform personalized adaptation for the unique degradation laws of different clients on the basis of fully exploring the cluster model of the common degradation laws of group products, and to divide and train the cluster prediction model with good generalization performance into a batch of accurate prediction models for different client individuals.

[0110] In the individual characteristic adaptation stage, the training of each working condition cluster is independent of each other. The personalized federated learning algorithm of model hierarchical aggregation is used to generate and iterate the client models within the working condition cluster. At the same time, to prevent the loss of common knowledge, the client local update process also adopts the method of knowledge distillation. The client receives the previous working condition cluster model obtained in the common law mining stage as the teacher model, and receives the client model obtained by model hierarchical aggregation as the student model, and verifies their performance respectively to obtain and For the set threshold t2, when or it means that the teacher model is still applicable to the current client degradation data. The client model is updated using the same knowledge distillation strategy as in the common law mining stage, and the coefficient λ1 of the distillation loss is dynamically adjusted according to the applicability of the teacher model:

[0111]

[0112] When and only the student model is used for local training.

[0113] Algorithm Implementation

[0114] The overall process of CDHF is as Figure 3 shown. In the common law mining stage, the client clusters under different working conditions participate in each communication round in turn. The client models with the same working conditions share model parameters. After reaching the maximum communication round of this stage, the knowledge information suitable for the deep degradation laws of different working conditions is retained in the working condition cluster models, which can be used to supervise the update of subsequent personalized models. In the individual characteristic adaptation stage, each working condition cluster generates personalized prediction models for different clients through a hypernetwork, and iterates the models through the knowledge distillation method, and finally realizes the accurate prediction of the degradation trend of individual products.

[0115] The implementation of the CDHFL algorithm is as shown in Algorithm 7.

[0116]

[0117]

[0118] Case Verification

[0119] Data Introduction

[0120] This chapter uses the cycle life test data of pouch lithium-ion batteries from a new energy company to verify the effectiveness of the above-mentioned federated prediction method for the degradation trend of population products. The cycle life test requires long-term testing of batteries with different formulations under pre-designed test conditions until their capacity reaches a pre-set failure threshold, which is an important part of the design and development of lithium-ion power batteries.

[0121] The battery capacity degradation dataset was obtained from the cycle life tests of 100 pouch lithium-ion batteries in eight groups (A1 to H1) under three temperature stresses of 25 °C, 45 °C, and 60 °C. Their cathode and separator materials are exactly the same, while the anode materials and electrolyte solutions vary from group to group. The specific formulation differences and the number of batteries in each group are shown in Table 1.

[0122] Table 1: Formulations and Quantities of Lithium-Ion Power Batteries at Different Test Temperatures

[0123]

[0124] Select one battery from each of the eight groups at each test temperature, and the curve obtained after maximum-minimum normalization of the capacity retention rate data before the battery capacity decays to 80% (failure threshold) is as Figure 4 shown. The horizontal axis is the number of sampling points, and the vertical axis is the capacity retention rate. It can be seen that the degradation trends of different battery individuals vary greatly, and as the temperature increases, the battery life is significantly shortened.

[0125] In the subsequent case, a single battery serves as a client, and the client's local dataset is sliced from the normalized capacity retention rate data of this battery and divided into a training dataset and a test dataset according to a certain proportion.

[0126] Cross-Operating Condition Degradation Trend Hierarchical Federated Prediction Case

[0127] (1) Case Setup

[0128] The cross-operating condition degradation trend hierarchical federated prediction verification case divides the degradation data of lithium-ion batteries at three temperatures into three operating condition clusters, as Figure 5 shown. Each cluster contains all battery individuals under one operating condition. Federated learning based on ring distillation can be self-organized by the three operating condition clusters or coordinated by the server.

[0129] (2) Ablation Experiment and Model Setup

[0130] The proposed CDHFL introduces a hierarchical structure based on LLApFL and uses circular knowledge distillation to enable the flow of degradation law information between layers. Therefore, the necessity of the proposed method is verified by ablation of the circular distillation method and the model hierarchical aggregation method.

[0131] 1) CluSamLLApFL: The personalized federated learning method is directly used to train the cross-condition battery degradation trend prediction model, and the supernetwork needs to screen effective knowledge from all prediction models.

[0132] 2) CluSamFedAvg: Neither the circular distillation training strategy nor the personalized adaptation for battery individuals is adopted. Only the cluster sampling method is used to federally aggregate the client prediction models.

[0133] The model structures in each algorithm are kept consistent, and the detailed model settings are shown in Table 2.

[0134] Table 2: Model settings for the hierarchical federated prediction case of cross-condition degradation trends

[0135]

[0136] 3) Analysis of ablation test results

[0137] The hierarchical federated prediction results of cross-condition battery degradation trends are shown in Table 3. By comparing the prediction indicators of CDHFL and CluSamLLApFL, it can be seen that the proposed hierarchical circular distillation method helps each client efficiently obtain more useful degradation information by mining the common degradation laws across conditions, avoiding the impact of the introduction of a large amount of useless knowledge on the accurate aggregation of the supernetwork, and significantly improving the performance of the personalized model. On the other hand, although CluSamFedAvg achieves the best prediction effect for batteries with more training data at 25°C, the reason is that the battery degradation trends at these two temperatures are relatively slow and the data volume is large, and the global prediction model completely tends to the part of the clients with the largest amount of training data. Its performance drops sharply for the battery individuals at 60°C with a small data volume and the clients with only 50% of the degradation data as training data, and the comprehensive prediction accuracy is much lower than that of CluSamLLApFL and CDHFL.

[0138] Table 3: Comparative test results of hierarchical federated prediction of cross-condition battery degradation trends

[0139]

[0140] Through Figure 6From the average prediction error of the ablation experiment of the hierarchical federated prediction of cross-condition battery degradation trends, it can be seen that although the convergence process of CluSamFedAvg is relatively smooth, the performance is the worst after the model converges. The convergence speed of CluSamLLApFL is much slower than that in the single-condition scenario. The supernetwork needs more time to learn the network layers required by the target client from more heterogeneous clients, and the final prediction accuracy also decreases. The proposed CDHFL quickly learns the basic degradation laws of the battery from different conditions in the common law mining stage, obtains a teacher model with better generalization ability, and lays a good foundation for the personalized model in the individual characteristic adaptation stage, thus achieving the most accurate prediction of battery degradation trends.

[0141] 4) Comparative experiments and model settings

[0142] To further verify the superiority of the proposed CDHFL, it is also compared with methods such as FedFomo, pFedHN, LG-FedAvg, and LocalOnly.

[0143] 5) Analysis of comparative experiment results

[0144] The results of the comparative experiment on the hierarchical federated prediction of cross-condition battery degradation trends are shown in Table 4. Similar to the single-condition battery degradation trend prediction experiment, the methods of LocalOnly and LG-FedAvg have the worst performance. Due to the increase in the number of clients, the training of the supernetwork in pFedHN is more difficult, and it cannot provide effective personalized model parameters for each client, resulting in a significant reduction in performance. FedFomo can select a more suitable model for the target client from more degradation models, and its performance is better than that in the single-condition scenario with fewer clients. CDHFL further reduces the mean squared error and mean absolute error to about 1 / 2 on the basis of FedFomo, and the prediction accuracy is significantly improved.

[0145] Table 4: Results of the comparative experiment on the hierarchical federated prediction of cross-condition battery degradation trends

[0146]

[0147] Further analyze the average prediction error of each method in the comparative experiment on the hierarchical federated prediction of cross-condition battery degradation trends. As Figure 7 shown, the three methods of Localonly, LG-FedAvg, and pFedHN cannot enable the prediction model to learn battery degradation knowledge, and CDHFL leads FedFomo in terms of the early convergence speed and the late prediction accuracy.

[0148] The preferred embodiments of the present invention have been described above with reference to the accompanying drawings, and thus do not limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention shall fall within the scope of the present invention.

Claims

1. A cross-condition degradation trend hierarchical federated prediction method based on annular distillation, characterized in that Including: Before each round of iteration in the common law mining stage, the server sequentially selects a working condition cluster in a cyclic manner, and uses the cluster sampling method to select multiple clients participating in the current round from all the clients in the working condition cluster. Each client is a single battery. Each working condition cluster is generated by dividing the battery degradation data at multiple temperatures according to the temperature. Each working condition cluster contains all the battery individuals at one temperature; Each battery participating in the current round uses the working condition cluster prediction model of the previous round as the teacher model for knowledge distillation, and the working condition cluster prediction model of the current round as the student model for knowledge distillation, and uses the local training set data to locally train the working condition cluster prediction model of the current round to obtain the parameters of the working condition cluster prediction model after the current round of training. The local training set data is generated by splitting the capacity retention rate data after normalizing the battery; The server uses the federated averaging algorithm to perform federated aggregation processing on the parameters of the working condition cluster prediction model after the current round of training sent by the multiple batteries to obtain the parameters of the working condition cluster prediction model after federated aggregation processing; Repeat the steps of the above common law mining stage until, after determining that the maximum number of rounds is reached, all the batteries under each working condition cluster obtain the parameters of the working condition cluster prediction model after federated aggregation processing obtained in the maximum number of rounds sent by the server, and obtain a working condition cluster prediction model with general knowledge information in different working condition degradation data; Before each round of iteration of each working condition cluster in the individual characteristic adaptation stage, the server uses the cluster sampling method to select multiple batteries participating in the current round from all the batteries in each working condition cluster, and sends the personalized prediction model parameters of the current round of the battery generated by using the supernetwork corresponding to each battery to the corresponding battery; Each battery participating in the current round uses the working condition cluster prediction model of the previous working condition cluster as the teacher model for knowledge distillation, and the personalized prediction model of the current round as the student model for knowledge distillation, locally trains the personalized prediction model of the current round to obtain the change amount of the personalized prediction model parameters of the current round, and sends the change amount of the personalized prediction model parameters of the current round to the server; The server uses the change amount of the personalized prediction model parameters of each battery in the current round to update the personalized prediction model parameters and the supernetwork to obtain the updated personalized prediction model parameters and supernetwork; Repeat the steps of the above individual characteristic adaptation stage until the server, after determining that the maximum number of rounds is reached, sends the personalized prediction model parameters of each battery obtained in the maximum number of rounds to the battery, so that each battery obtains a cross-working condition battery degradation trend prediction model applicable to the battery by loading the personalized prediction model parameters obtained in the maximum number of rounds.

2. The method according to claim 1, wherein Batteries with the same working condition form a working condition cluster. In the common law mining stage, the working condition cluster prediction models and the working condition cluster prediction model parameters of all the batteries in the working condition cluster are the same. In the individual characteristic adaptation stage, the personalized prediction models and the personalized prediction model parameters of all the batteries in the working condition cluster are different.

3. The method according to claim 2, wherein Each battery participating in the current round uses the working condition cluster prediction model of the previous round as the teacher model for knowledge distillation, and the working condition cluster prediction model of the current round as the student model for knowledge distillation, and locally trains the working condition cluster prediction model of the current round using the local training set data, and the obtained parameters of the working condition cluster prediction model after the current round of training include: Each battery participating in the current round receives the working condition cluster prediction model of the previous round sent by the server side, and uses the working condition cluster prediction model of the previous round as the teacher model for knowledge distillation, and the working condition cluster prediction model of the current round as the student model for knowledge distillation; The battery calculates the validation loss value of the student model using the local training set data; When the validation loss value of the student model is less than the set first validation loss threshold, the battery locally trains the student model using the first constraint, and at the same time adds the constraint of the teacher model on the extracted degradation features during the training process, and takes the parameters of the trained student model as the parameters of the working condition cluster prediction model after the current round of training; When the validation loss value of the student model is not less than the set first validation loss threshold, the battery locally trains the teacher model using the second constraint, and takes the parameters of the trained teacher model as the parameters of the working condition cluster prediction model after the current round of training.

4. The method according to claim 3, characterized in that, The battery locally trains the student model using the first constraint, and at the same time adding the constraint of the teacher model on the extracted degradation features during the training process includes: L1 = L MSE + λ0L KD Among them, L1 represents the training loss of the first constraint; L MSE represents the mean squared error loss; λ0 represents a constant coefficient used to balance the attention of the model to the current data knowledge and the common knowledge from the teacher model; L KD represents the knowledge distillation loss; f tea / stu (x) represents the degraded features extracted by the teacher / student model encoder when x is input.

5. The method according to claim 4, characterized in that The battery locally trains the teacher model using the second constraint includes: L2 = L MSE Among them, L2 represents the training loss of the second constraint.

6. The method according to claim 5, wherein Each battery participating in the current round uses the working condition cluster prediction model of the previous working condition cluster as the teacher model for knowledge distillation, and the personalized prediction model of the current round as the student model for knowledge distillation, and locally trains the personalized prediction model of the current round, and the obtained change in the parameters of the personalized prediction model of the current round includes: Each battery participating in the current round uses the pre-stored working condition cluster prediction model of the previous working condition cluster sent by the server side as the teacher model for knowledge distillation, and the personalized prediction model of the current round sent by the server side as the student model for knowledge distillation; The battery calculates the validation loss value of the student model and the validation loss value of the teacher model respectively using the local training set data; If the validation loss value of the teacher model is less than the set second validation loss threshold or the validation loss value of the student model, the battery locally trains the student model using the third constraint, and at the same time adds the constraint of the teacher model on the extracted degradation features during the training process, and obtains the change in the parameters of the personalized prediction model of the current round; If the validation loss value of the teacher model is not less than the set second validation loss threshold and the validation loss value of the student model, the battery locally trains the student model using the second constraint, and obtains the change in the parameters of the personalized prediction model of the current round.

7. The method according to claim 6, wherein The battery uses the third constraint to locally train the student model, and the constraint of the teacher model on the extracted degraded features added during the training process includes: L3 = L MSE + λ1L KD Among them, L3 represents the training loss of the third constraint; λ1 represents the distillation loss coefficient in the third constraint; represents the validation loss value of the student model; represents the validation loss value of the teacher model.

8. The method according to claim 1, wherein The server side will send the personalized prediction model parameters of the current round of each battery generated by the hypernetwork corresponding to each battery to the corresponding battery, including: The server side equips each battery with a hypernetwork for generating aggregation weights for each network layer of the battery's personalized prediction model, and uses the hypernetwork corresponding to each battery and the battery embedding vector to generate a battery aggregation weight matrix; The server side uses the battery aggregation weight matrix to generate the personalized prediction model parameters of the current round of the battery, and sends the generated personalized prediction model parameters of the current round of the battery to the battery.

9. The method according to claim 8, wherein The server side uses the battery aggregation weight matrix to generate the personalized prediction model parameters of the current round of the battery, including: Among them, represents the personalized prediction model parameters of battery i in the current round, represents the parameters of the l-th layer in the personalized prediction model; represents the parameters of the l-th layer in the personalized prediction model of battery k; represents the weight of the l-th layer network in the personalized prediction model of the k-th battery applied to the corresponding layer in battery i; L represents the number of layers of the network in the personalized prediction model; K represents the number of batteries in the federated system.

10. The method according to claim 9, wherein The server side updates the personalized prediction model parameters and the hypernetwork using the change amount of the personalized prediction model parameters of each battery in the current round, and obtains the updated personalized prediction model parameters and hypernetwork, including: Among them, Δω k represents the change amount of the personalized prediction model parameters of battery k in the current round; represents the personalized prediction model parameters of battery k in the current round; represents the personalized prediction model parameters of battery k after the update in the current round; Ω represents the set of all battery personalized prediction model parameters before the start of the current round; respectively represent the gradients of the embedding vector e k in the hypernetwork of battery k, the model parameters ψ k ; Δe k represents the change amount of the embedding vector in the hypernetwork of battery k in the current round; represents the embedding vector in the hypernetwork of battery k after the update in the current round; represents the embedding vector in the hypernetwork of battery k in the current round; Δψ k represents the change amount of the personalized prediction model parameters in the hypernetwork of battery k in the current round; represents the model parameters in the hypernetwork of battery k after the update in the current round; represents the model parameters in the hypernetwork of battery k in the current round.

Citation Information

Patent Citations

  • Lithium battery long-term degradation trend prediction method

    CN112257348A

  • Target detection method and apparatus based on federated learning, and device and storage medium

    WO2021189906A1