Server-side adaptive parameter aggregation method

Through the server-side adaptive parameter aggregation method, the improved joint mean algorithm and adaptive aggregation algorithm are used to solve the problems of data islands, privacy security and hardware performance upper limits in the training of traditional transmission line defect detection models, and achieve higher detection accuracy and model stability.

CN119989024APending Publication Date: 2025-05-13STATE GRID SICHUAN ELECTRIC POWER CORP ELECTRIC POWER RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510064298.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-14
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

There are problems of data silos, privacy and security and hardware performance limits in the training of traditional transmission line defect detection models, resulting in insufficient data diversity and generalization capabilities.

Method used

The server-side adaptive parameter aggregation method is adopted, and the advantages of different algorithms are fused through the improved joint mean algorithm and adaptive aggregation algorithm, parameter aggregation and model training are carried out to generate a more stable and generalized global model.

Benefits of technology

It significantly improves the detection accuracy and model stability of transmission line defect detection, and solves the problems of data imbalance, training efficiency and model generalization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989024A_ABST
    Figure CN119989024A_ABST
Patent Text Reader

Abstract

The invention discloses a server-side adaptive parameter aggregation method, and belongs to the technical field of power grid transmission line defect detection. Comprising the following steps: carrying out joint training on K clients, loading a local data set by each client, and entering a state of waiting for initializing a global model; a global model is issued to each client, and the client loads the global model and takes the global model as an initial training local model parameter of the client; training is carried out according to a specified training round, if the specified round is reached, the joint training task is ended, otherwise, each client uses a local data set to carry out training until all the K clients are trained; and taking the calculated hierarchical model difference as a pseudo gradient of an adaptive aggregation algorithm, and generating and storing a two-stage global model. According to the method, the joint mean value algorithm and the adaptive aggregation algorithm are fused, advantage complementation is achieved, defect detection tasks are better processed, false detection and missing detection are reduced, and the problems of data imbalance, training efficiency and model generalization ability are solved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of a patent application entitled "A server-side adaptive parameter aggregation method and system". The application date of the original application is September 14, 2024, and the application number is 202411289022.4. Technical Field

[0002] The present invention relates to the technical field of power grid transmission line defect detection, and in particular to a service-side adaptive parameter aggregation method. Background Art

[0003] As a key link in power transmission, the safe and stable operation of transmission lines is directly related to the reliability of power supply. Transmission lines are spread all over the country, covering a vast geographical area. Therefore, efficient and accurate detection of their defects has become a core challenge in power maintenance.

[0004] Traditional transmission line defect detection mainly relies on manual inspection and single-node computer vision technology. Circuit defect data is scattered in devices or systems in different regions, resulting in the following problems in the training of traditional transmission line defect detection models: (1) There is a lack of effective sharing and connection between data, forming data islands, making it impossible to fully integrate and analyze data collected in different environments, thereby affecting the data diversity and generalization ability of the transmission line defect detection model. (2) Data privacy and security issues are prominent. (3) There is an upper limit to computer hardware performance, which makes it impossible to train large-scale data on a single node.

[0005] Although federated learning can solve the above problems well, existing federated learning algorithms still have some problems: (1) Model parameters of different training nodes or devices are different, and a single parameter aggregation algorithm is difficult to guarantee the convergence and stability of the global model. (2) The computing power and data quality of different training nodes or devices are different, and a single parameter aggregation algorithm is difficult to guarantee that all nodes or devices can contribute to the global model. (3) Only using a single parameter aggregation algorithm for parameter aggregation makes it difficult to fully utilize the advantages of different algorithms to improve the performance of the model. Summary of the invention

[0006] In order to overcome the shortcomings of the prior art, the purpose of the present invention is to provide a server-side adaptive parameter aggregation method, which solves the problems of data imbalance, training efficiency and model generalization ability, especially in the application of transmission line defect detection, significantly improves the detection accuracy and model stability.

[0007] To achieve the above object, the present invention provides the following solutions:

[0008] A server-side adaptive parameter aggregation method comprises the following steps:

[0009] S1, the server selects K clients for joint training, and each client loads the local data set and enters the state of waiting for initialization of the global model;

[0010] S2. In the initial stage, after the server initializes the global model, it sends the global model parameters to each client. The client loads the initialized global model parameters and uses them as the initial training local model parameters of the client. In the iterative stage, as the joint learning proceeds, the server generates a second-stage global model after each iteration and sends the second-stage global model to the client for training.

[0011] S3, determine whether the client has reached the preset training rounds, if so, end the joint training task and the process, otherwise, each client uses the improved joint mean algorithm to train the local model using the local data set and enter S4;

[0012] S4. Determine whether all K clients have completed training. If so, the server uses the improved joint mean algorithm to aggregate the local model parameters uploaded by the client based on the training results to obtain a first-stage global model and enter S5. Otherwise, return to S3 until all clients have completed training.

[0013] S5. Calculate the hierarchical model difference based on the current one-stage global model and the global model of the previous round of training;

[0014] S6, using the hierarchical model difference as the pseudo gradient of the adaptive aggregation algorithm, calculating the first-order moment estimation and the second-order moment estimation involved in the adaptive aggregation algorithm, and using the adaptive aggregation algorithm for adaptive aggregation, generating a two-stage global model, and storing the parameters of the two-stage global model of this round, and dynamically adjusting the learning rate through the first-order moment estimation and the second-order moment estimation, and sending the two-stage global model parameters to the client, and returning to the iteration stage in S2;

[0015] In the initial stage, after the server initializes the global model, it sends the global model parameters to each client, and the client loads the initialized global model parameters and uses them as the initial training local model parameters of the client, which are specifically:

[0016] A1. At the beginning, after the server initializes the global model, it sends the global model parameters to each client.

[0017] A2. The client loads the initialized global model parameters and uses them as initial parameters. The global model parameters are used on the local data set to perform local model training to generate local model parameters.

[0018] A3. After the training is completed, the client sends the updated local model parameters back to the server.

[0019] A4. The server collects all updated local model parameters of the clients and aggregates them to generate a new global model.

[0020] A5. The server sends the new global model parameters to the client, and the client uses the new global model parameters as new initial parameters to train the local model.

[0021] The expressions of the local model parameters are as follows:

[0022]

[0023] in, represents the local model parameters of the kth client in the t+1th training round, represents the local model parameters of the kth client in the tth training round, η represents the learning rate, g k represents the gradient calculated by the local model of the kth client when training on the local dataset;

[0024] The expression of the one-stage global model in S4 is as follows:

[0025]

[0026] Among them, w t+1 represents the one-stage global model obtained after the aggregation of the t+1th training round, k represents the client index, K represents the total number of clients, n represents the total number of client samples, and n k represents the number of samples of the kth client;

[0027] The expression of the hierarchical model difference in S5 is as follows:

[0028] Δ t =[w avg,1 -w t,1 ,w avg,2 -w t,2 ,...,w avg,i -w t,i ]

[0029]

[0030] Among them, Δ t represents the hierarchical model difference, i represents the level of the global model, and w avg,i represents the global model whose i-th layer parameters are aggregated using the joint mean algorithm, w t,i represents the i-th layer parameter of the previous round of global model, w avg represents the global model aggregated by the joint mean algorithm in the first stage, θ represents the proportional parameter between the number of client samples and the loss, and Lk represents the loss ratio of the kth client among all clients, l k and l i″′ Both represent the local model loss value of the k-th client trained on the local dataset, and i″′ represents the index number of the client;

[0031] The expression of the first-order moment estimate is as follows:

[0032] m t =β1m t-1 +(1-β1)Δ t

[0033] Among them, m t represents the first-order moment estimate of the t-th training round, t represents the training round, β1 represents the decay rate of the first-order moment estimate, m t-1 represents the first-order moment estimate of the t-1th training round, Δ t represents pseudo gradient;

[0034] The expression of the second-order moment estimate is as follows:

[0035]

[0036] Among them, v t represents the second-order moment estimate of the t-th training round, β2 represents the decay rate of the second-order moment estimate, and v t-1 represents the second-order moment estimate of the t-1th training round;

[0037] The updating formula of the global model parameters is as follows:

[0038]

[0039] Among them, w t represents the global model before updating, σ represents the learning rate, ε represents a constant, and λ represents the weight decay term;

[0040] The expression of the global optimal objective function of the server is as follows:

[0041]

[0042] Among them, min means taking a small value, w means the global model parameter, F(w) means the global optimal objective function of the server, and F k (w) represents the local objective function of the kth client, n k represents the data volume of the kth client, i" represents the i"th sample, d k represents the local dataset of the kth client, f i″(w) represents the loss value of the local client defect detection model SSD on sample i″, N represents the number of default boxes corresponding to the matched real boxes, L conf represents the classification loss, x represents the matching indicator variable, c represents the confidence score of each category predicted, λ represents the hyperparameter for balancing the weight between classification loss and localization loss, and L loc represents the positioning loss, l represents the offset of each predicted default box relative to the original shape, g represents the parameters of the real box, Pos represents the set of positive class labels, i′ represents the i′th default box, Indicates whether the i′th default box matches the jth true bounding box of the set pos belonging to the positive class label, represents the probability that the i′th default box belongs to the set pos of positive labels, Neg represents the set of negative labels, Represents the probability that the i'th default box belongs to the set Neg of negative class labels, Indicates the confidence that the i′th default box belongs to the set pos of the positive class label, m represents different coordinate types, cx represents the horizontal coordinate of the center point of the object, cy represents the vertical coordinate of the center point of the object, and h represents the height. Indicates whether the i′th default box matches the j′th real bounding box, smooth L1 represents the smooth L1 loss function, represents the predicted coordinates, Indicates the actual coordinates.

[0043] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0044] In view of the problems existing in the traditional transmission line defect detection model training method, the present invention aims to integrate the advantages of different algorithms, utilize the complementarity of different algorithms, and integrate the improved joint mean algorithm with the adaptive aggregation algorithm. While ensuring the universality and good effect of the joint mean algorithm, combined with the adaptability and stability of the adaptive aggregation algorithm, the performance and generalization ability of the global model are further improved, and the problems of data islands, privacy security and hardware performance limits are solved, and the training efficiency and generalization ability of the global model are improved. The present invention makes full use of the complementary advantages of different algorithms, fully considers the contribution of different nodes or devices under the premise of ensuring stability and convergence, and effectively improves the performance and generalization ability of the joint learning model of transmission line defect detection.

[0045] The present invention proposes an improved model hierarchical difference calculation method, which not only considers the number of client data samples, but also considers the impact of factors such as client local data distribution and quality on the aggregation effect. The loss in the client training process is used as an indirect indicator to measure the client local data distribution and quality factors. In this way, the contribution of the client to the global model can be more accurately reflected, thereby improving the performance and generalization ability of the global model. Compared with the traditional method that only considers the number of client data samples, the present invention can better utilize the diversity and quality of client data, thereby reducing the weight distribution bias and improving accuracy.

[0046] The present invention introduces parameters to adjust the ratio between the number of client data samples and the training loss, so that the global model can better adapt to the distribution and quality differences of client data. The present invention can better handle the problem of uneven data distribution, improve the generalization ability of the global model, and reduce the performance degradation caused by data distribution differences.

[0047] The present invention introduces a weight decay term λ to control the complexity of the global model and improve the generalization performance of the global model. In this way, the overfitting problem can be avoided and the stability and generalization performance of the model can be improved.

[0048] The present invention is based on the FedAvg algorithm and the FedAdam algorithm, and integrates the two algorithms to achieve complementary advantages, further improve the effect of joint learning, and can better handle defect detection tasks, reduce false detections and missed detections, and effectively solve the problems of data imbalance, training efficiency and global model generalization ability. Especially in the application of transmission line defect detection, the detection accuracy and the stability of the global model are significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0050] Figure 1 The present invention is a flow chart of the method. DETAILED DESCRIPTION

[0051] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0052] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0053] The terms "first", "second", "third" and "fourth" in the specification and claims of the present application and the drawings are used to distinguish different objects rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a series of steps, processes, methods, etc. are not limited to the listed steps, but may optionally include steps that are not listed, or may optionally include other step elements inherent to these processes, methods, products or devices.

[0054] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0055] like Figure 1 As shown, the present invention provides a server-side adaptive parameter aggregation method, comprising the following steps:

[0056] S1, the server selects K clients for joint training, and each client loads the local data set and enters the state of waiting for initialization of the global model;

[0057] S2. In the initial stage, after the server initializes the global model, it sends the global model parameters to each client. The client loads the initialized global model parameters and uses them as the initial training local model parameters of the client. In the iterative stage, as the joint learning proceeds, the server generates a second-stage global model after each iteration and sends the second-stage global model to the client for training.

[0058] In the initial stage, after the server initializes the global model, it sends the global model parameters to each client, and the client loads the initialized global model parameters and uses them as the initial training local model parameters of the client, which are specifically:

[0059] A1. At the beginning, after the server initializes the global model, it sends the global model parameters to each client.

[0060] A2. The client loads the initialized global model parameters and uses them as initial parameters. The global model parameters are used on the local data set to perform local model training to generate local model parameters.

[0061] A3. After the training is completed, the client sends the updated local model parameters back to the server.

[0062] A4. The server collects all updated local model parameters of the clients and aggregates them to generate a new global model.

[0063] A5. The server sends the new global model parameters to the client, and the client uses the new global model parameters as new initial parameters to train the local model.

[0064] S3, determine whether the client has reached the preset training rounds, if so, end the joint training task and the process, otherwise, each client uses the improved joint mean algorithm to train the local model using the local data set and enter S4;

[0065] S4. Determine whether all K clients have completed training. If so, the server uses the improved joint mean algorithm to aggregate the local model parameters uploaded by the client based on the training results to obtain a first-stage global model and enter S5. Otherwise, return to S3 until all clients have completed training.

[0066] S5. Calculate the hierarchical model difference based on the current one-stage global model and the global model of the previous round of training;

[0067] S6, using the hierarchical model difference as the pseudo gradient of the adaptive aggregation algorithm, calculating the first-order moment estimation and the second-order moment estimation involved in the adaptive aggregation algorithm, and using the adaptive aggregation algorithm for adaptive aggregation, generating a two-stage global model, and storing the parameters of the two-stage global model of this round, and dynamically adjusting the learning rate through the first-order moment estimation and the second-order moment estimation, and sending the two-stage global model parameters to the client, and returning to the iteration stage in S2;

[0068] In this embodiment, the present invention integrates the federated mean algorithm (FedAvg) and the adaptive aggregation algorithm (FedAdam), fully utilizes the complementary advantages of different algorithms, iteratively trains to find the global optimal parameter w, and improves the global model performance.

[0069] In this embodiment, a power transmission line defect detection system based on the server-side adaptive parameter aggregation algorithm (SAPAA-MMF) with multi-method fusion is constructed. First, the server initializes the global model parameters and distributes them to each client. Each client performs local training based on the local data set, updates the global model parameters using the stochastic gradient descent method, and then uploads the updated global model parameters to the server. The update formula of the global model parameters is as follows:

[0070]

[0071] Among them, w t represents the global model before updating, σ represents the learning rate, ε represents a constant, and λ represents the weight decay term;

[0072] In this embodiment, after receiving the local model parameters uploaded by the client, the server first uses the improved joint mean algorithm to aggregate the parameters and generate a global model. Then, the hierarchical difference between the global model and the previous round of global model is calculated and used as the pseudo gradient of the adaptive aggregation algorithm FedAdam algorithm. During the calculation process, the aggregation weight of the global model parameters is adjusted according to the number of samples and training loss of the client to ensure the balance of contributions from different clients.

[0073] In this embodiment, the joint mean algorithm is mainly used to solve the problem of data distribution differences and has good effect and adaptability. In the joint mean algorithm, the server is used to aggregate the local model parameters uploaded by the client, and the client performs local model training on the local data set. After local training, the local model parameters are generated.

[0074]

[0075] in, represents the local model parameters of the kth client in the t+1th training round, represents the local model parameters of the kth client in the tth training round, η represents the learning rate, g k represents the gradient calculated by the local model of the kth client when training on the local dataset;

[0076] In this embodiment, the server is used to aggregate the local model parameters uploaded by the client. The server aggregates the local model parameters uploaded by each client according to the following formula:

[0077]

[0078] Among them, w t+1represents the one-stage global model obtained after the aggregation of the t+1th training round, k represents the client index, K represents the total number of clients, n represents the total number of client samples, and n k represents the number of samples of the kth client;

[0079] In this embodiment, the server generates a global model after aggregation and sends it to each client for subsequent rounds of training until the global model converges or reaches the end condition.

[0080] In this embodiment, the adaptive aggregation algorithm is used to solve the convergence speed and stability problems of the global model. The adaptive aggregation algorithm dynamically adjusts the learning rate through the first-order moment estimation and the second-order moment estimation to accelerate the convergence and stability of the global model. The expression of the first-order moment estimation in the adaptive aggregation algorithm is as follows:

[0081] m t =β1m t-1 +(1-β1)△ t

[0082] Among them, m t represents the first-order moment estimate of the t-th training round, t represents the training round, β1 represents the decay rate of the first-order moment estimate, m t-1 represents the first-order moment estimate of the t-1th training round, Δ t represents pseudo gradient;

[0083] By choosing different values ​​of β1, the first-order moment estimation is used to simulate the influence of momentum, which helps to improve the stability of the global model parameter update.

[0084] The expression of the second-order moment estimation in this embodiment is as follows:

[0085]

[0086] Among them, v t represents the second-order moment estimate of the t-th training round, β2 represents the decay rate of the second-order moment estimate, and v t-1 represents the second-order moment estimate of the t-1th training round;

[0087] By considering the second-order moment estimation, the adaptive aggregation algorithm can more flexibly respond to gradient changes in the parameter space. In some cases, the gradient will change dramatically, and the second-order moment estimation can capture the trend of this change and help adjust the learning rate to make the algorithm more robust. Finally, the global model parameter update formula is:

[0088]

[0089] Among them, w trepresents the global model before update, σ represents the learning rate, ε represents a constant, a small constant added for numerical stability to prevent division by zero, and λ represents the weight decay term;

[0090] In this embodiment, the pseudo gradient is based on the adaptive aggregation algorithm, which is integrated with the weighting of the client data samples by the joint mean algorithm, and the hierarchical model difference is used as the pseudo gradient of the adaptive aggregation algorithm:

[0091] △ t =[w avg,1 -w t,1 , w avg,2 -w t,2 ,...,w avg,i -w t,i ]

[0092] In this embodiment, the global model w is aggregated using the joint mean algorithm in the first stage. avg The optimized expression is as follows:

[0093]

[0094] The expression for the loss ratio of the kth client to all clients is as follows:

[0095]

[0096] Among them, Δ t represents the hierarchical model difference, i represents the level of the global model, and w avg,i represents the global model whose i-th layer parameters are aggregated using the joint mean algorithm, w t,i represents the i-th layer parameter of the previous round of global model, w avg represents the global model aggregated by the joint mean algorithm in the first stage, θ represents the proportional parameter between the number of client samples and the loss, and L k represents the loss ratio of the kth client among all clients, l k and l i″′ Both represent the local model loss value of the k-th client trained on the local dataset, and i″′ represents the index number of the client;

[0097] This formula converts a set of losses into a probability distribution so that its value range falls between (0,1). Input l k Taking a negative number is to make the smaller l k After exponential operation, it approaches 1, while the larger l k It approaches 0, corresponding to a larger loss, the global model learning effect is poor, and the loss ratio should be small, while a smaller loss, the global model learning effect is good, and the loss ratio should be large.

[0098] In this embodiment, the adaptive aggregation algorithm introduces a weight decay penalty in its update to control the complexity of the global model to avoid overfitting. The optimized global model is updated as follows:

[0099]

[0100] λ represents the weight decay term, which is a non-negative real number and is usually set to 0.0001 during training.

[0101] In this embodiment, the expression of the global optimal objective function of the server is as follows:

[0102]

[0103] Among them, min means taking a small value, w means the global model parameter, F(w) means the global optimal objective function of the server, and F k (w) represents the local objective function of the kth client, n k represents the data volume of the kth client, i″ represents the i″th sample, and d k represents the local dataset of the kth client, f i″ (w) represents the loss value of the local client defect detection model SSD on sample i″, N represents the number of default boxes corresponding to the matched real boxes, L conf represents the classification loss, x represents the matching indicator variable, c represents the confidence score of each category predicted, λ represents the hyperparameter for balancing the weight between classification loss and localization loss, and L loc represents the positioning loss, l represents the offset of each predicted default box relative to the original shape, g represents the parameters of the real box, Pos represents the set of positive class labels, i′ represents the i′th default box, Indicates whether the i′th default box matches the jth true bounding box of the set pos belonging to the positive class label, represents the probability that the i′th default box belongs to the set pos of positive labels, Neg represents the set of negative labels, Represents the probability that the i'th default box belongs to the set Neg of negative class labels, Indicates the confidence that the i′th default box belongs to the set pos of the positive class label, m represents different coordinate types, cx represents the horizontal coordinate of the center point of the object, cy represents the vertical coordinate of the center point of the object, and h represents the height. Indicates whether the i′th default box matches the j′th real bounding box, smooth L1 represents the smooth L1 loss function, represents the predicted coordinates, Indicates the actual coordinates.

[0104] In this embodiment, five professionally labeled transmission line defect data sets are selected for algorithm performance evaluation through experimental verification. The results show that the detection accuracy of the present invention is higher than that of the improved joint mean algorithm FedAvgM, the federated YOGI algorithm FedYogi, the truncated mean algorithm TrimmedAvg, and other algorithms in typical defect detection tasks such as cement pole breakage and anti-vibration hammer slippage, and the false detection and missed detection are significantly reduced.

[0105] In this embodiment, visualization analysis and Friedman statistical test are used to further confirm the significant advantages of the present invention in the evaluation indicators of the server and client, and prove the effectiveness of multi-method fusion and the superiority of the overall performance of the algorithm.

[0106] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0107] The principles and implementation methods of the present invention are described in this article using specific examples. The description of the above embodiments is only used to help understand the method and core idea of ​​the present invention. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. A server-side adaptive parameter aggregation method, characterized in that: The following steps are involved: S1, the server selects K clients for joint training, and each client loads the local data set and enters the state of waiting for initialization of the global model; S2. In the initial stage, after the server initializes the global model, it sends the global model parameters to each client. The client loads the initialized global model parameters and uses them as the initial training local model parameters of the client. In the iterative stage, as the joint learning proceeds, the server generates a second-stage global model after each iteration and sends the second-stage global model to the client for training. S3, determine whether the client has reached the preset training rounds, if so, end the joint training task and the process, otherwise, each client uses the improved joint mean algorithm to train the local model using the local data set and enter S4; S4. Determine whether all K clients have completed training. If so, the server uses the improved joint mean algorithm to aggregate the local model parameters uploaded by the client based on the training results to obtain a first-stage global model and enter S5. Otherwise, return to S3 until all clients have completed training. S5. Calculate the hierarchical model difference based on the current one-stage global model and the global model of the previous round of training; S6, using the hierarchical model difference as the pseudo gradient of the adaptive aggregation algorithm, calculating the first-order moment estimation and the second-order moment estimation involved in the adaptive aggregation algorithm, and using the adaptive aggregation algorithm for adaptive aggregation, generating a two-stage global model, and storing the parameters of the two-stage global model of this round, and dynamically adjusting the learning rate through the first-order moment estimation and the second-order moment estimation, and sending the two-stage global model parameters to the client, and returning to the iteration stage in S2; In the initial stage, after the server initializes the global model, it sends the global model parameters to each client, and the client loads the initialized global model parameters and uses them as the initial training local model parameters of the client, which are specifically: A1. At the beginning, after the server initializes the global model, it sends the global model parameters to each client. A2. The client loads the initialized global model parameters and uses them as initial parameters. The global model parameters are used on the local data set to perform local model training to generate local model parameters. A3. After the training is completed, the client sends the updated local model parameters back to the server. A4. The server collects all updated local model parameters of the clients and aggregates them to generate a new global model. A5. The server sends the new global model parameters to the client, and the client uses the new global model parameters as new initial parameters to train the local model. The expressions of the local model parameters are as follows: in, represents the local model parameters of the kth client in the t+1th training round, represents the local model parameters of the kth client in the tth training round, η represents the learning rate, g k represents the gradient calculated by the local model of the kth client when training on the local dataset; The expression of the one-stage global model in S4 is as follows: Among them, w t+1 represents the one-stage global model obtained after the aggregation of the t+1th training round, k represents the client index, K represents the total number of clients, n represents the total number of client samples, and n k represents the number of samples of the kth client; The expression of the hierarchical model difference in S5 is as follows: Δ t =[w avg,1 -w t,1 ,w avg,2 -w t,2 ,...,w avg,i -w t,i ] Among them, Δ t represents the hierarchical model difference, i represents the level of the global model, and w avg,i represents the global model whose i-th layer parameters are aggregated using the joint mean algorithm, w t,i represents the i-th layer parameter of the previous round of global model, w avg represents the global model aggregated by the joint mean algorithm in the first stage, θ represents the proportional parameter between the number of client samples and the loss, and L k represents the loss ratio of the kth client among all clients, l k and l i″′ Both represent the local model loss value of the k-th client trained on the local dataset, and i″′ represents the index number of the client; The expression of the first-order moment estimate is as follows: m t =β1m t-1 +(1-β1)Δ t Among them, m t represents the first-order moment estimate of the t-th training round, t represents the training round, β1 represents the decay rate of the first-order moment estimate, m t-1 represents the first-order moment estimate of the t-1th training round, Δ t represents pseudo gradient; The expression of the second-order moment estimate is as follows: Among them, v t represents the second-order moment estimate of the t-th training round, β2 represents the decay rate of the second-order moment estimate, and v t-1 represents the second-order moment estimate of the t-1th training round; The updating formula of the global model parameters is as follows: Among them, w t represents the global model before updating, σ represents the learning rate, ε represents a constant, and λ represents the weight decay term; The expression of the global optimal objective function of the server is as follows: Among them, min means taking a small value, w means the global model parameter, F(w) means the global optimal objective function of the server, and F k (w) represents the local objective function of the kth client, n k represents the data volume of the kth client, i″ represents the i″th sample, and d k represents the local dataset of the kth client, f i″ (w) represents the loss value of the local client defect detection model SSD on sample i″, N represents the number of default boxes corresponding to the matched real boxes, L conf represents the classification loss, x represents the matching indicator variable, c represents the confidence score of each category predicted, λ represents the hyperparameter for balancing the weight between classification loss and localization loss, and L loc represents the positioning loss, l represents the offset of each predicted default box relative to the original shape, g represents the parameters of the real box, Pos represents the set of positive class labels, i′ represents the i′th default box, Indicates whether the i′th default box matches the jth true bounding box of the set pos belonging to the positive class label, represents the probability that the i′th default box belongs to the set pos of positive labels, Neg represents the set of negative labels, Represents the probability that the i'th default box belongs to the set Neg of negative class labels, Indicates the confidence that the i′th default box belongs to the set pos of the positive class label, m represents different coordinate types, cx represents the horizontal coordinate of the center point of the object, cy represents the vertical coordinate of the center point of the object, and h represents the height. Indicates whether the i′th default box matches the j′th real bounding box, smooth L1 represents the smooth L1 loss function, represents the predicted coordinates, Indicates the actual coordinates.