Method for predicting heat transfer coefficient after drying based on federal learning

By collaboratively training BP neural networks among distributed nodes through the federated learning framework and the FedProx algorithm, combined with Optuna hyperparameter optimization, the computational resource dependence and data applicability issues of post-drying heat transfer coefficient prediction were resolved, achieving high-precision heat transfer coefficient prediction.

CN120671930APending Publication Date: 2025-09-19SHANGHAI JIAOTONG UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510962974.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

The existing methods for predicting the heat transfer coefficient after drying rely on large computing resources, have poor applicability, are difficult to cope with different fluid working conditions, and have data privacy protection issues.

Method used

A federated learning mechanism is adopted to build a local BP neural network model in the offline stage through multiple distributed nodes, and collaborative training is performed without sharing the original data. The Flower framework and FedProx algorithm are used to aggregate model parameters, combined with Optuna hyperparameter optimization to build a global prediction model.

Benefits of technology

Without sharing data, the model's generalization ability and prediction accuracy are significantly improved, the mean absolute percentage error (MAPE) is reduced to no more than 15.88%, and the coefficient of determination (R²) is increased to 0.9753, adapting to multi-source heterogeneous data scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671930A_ABST
    Figure CN120671930A_ABST
Patent Text Reader

Abstract

A method for predicting a heat transfer coefficient after drying based on federated learning comprises the following steps: in an offline stage, constructing a plurality of distributed processing nodes, after each node collects original data in the offline stage and constructs a local BP neural network model, while performing offline training and hyper-parameter optimization locally, not sharing the original data and only uploading model parameters; the federated learning server adopts an aggregation algorithm to aggregate the received model parameters so as to eliminate the problem of inconsistent data distribution among the nodes; and in the online stage, the federal learning server predicts the heat transfer coefficient in real time after drying through the aggregated model parameters. According to the method, a plurality of distributed nodes are utilized, on the premise that original data are not shared, a federal learning mechanism is adopted to carry out cooperative training on a BP neural network model, local model parameters of all the nodes are aggregated to construct a global prediction model, and therefore prediction precision and generalization ability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technology in the field of heat transfer, specifically a method for predicting heat transfer coefficient after drying based on federated learning. Background Art

[0002] Current prediction methods for the heat transfer coefficient after drying mostly rely on empirical correlations or semi-theoretical models, which have problems such as high dependence on computing resources, poor applicability, and difficulty in coping with different fluid conditions. Summary of the Invention

[0003] In response to the problems of existing methods for predicting the heat transfer coefficient after drying up, such as high dependence on computing resources and poor adaptability, the present invention proposes a method for predicting the heat transfer coefficient after drying up based on federated learning. By utilizing multiple distributed nodes and adopting a federated learning mechanism to collaboratively train the BP neural network model without sharing the original data, a global prediction model is constructed by aggregating the local model parameters of each node. This method has the advantages of improving the model generalization ability, protecting data privacy, and adapting to multi-source heterogeneous data.

[0004] The present invention is achieved through the following technical solutions:

[0005] The present invention relates to a method for predicting the heat transfer coefficient after drying up based on federated learning. In the offline stage, multiple distributed processing nodes are constructed. Each node collects raw data in the offline stage and builds a local BP neural network model. Then, offline training and hyperparameter optimization are performed locally without sharing the raw data and only uploading the model parameters. The federated learning server uses an aggregation algorithm to aggregate the received model parameters to eliminate the problem of inconsistent data distribution among the nodes. In the online stage, the federated learning server uses the aggregated model parameters to perform real-time prediction of the heat transfer coefficient after drying up.

[0006] The offline phase specifically includes:

[0007] 1) The federated learning server initializes the model parameters of the BP neural network as the global model and sends them to each distributed processing node .

[0008] The federated learning server is preferably built based on the open source federated learning framework Flower. The server is used to coordinate the initialization, reception, aggregation and synchronization of model parameters of each node, and has functions such as task scheduling, communication management and global parameter maintenance.

[0009] The Flower framework is implemented using, but not limited to, the technology described in "FLOWER: A FRIENDLY FEDERATED LEARNING FRAMEWORK" (Daniel J. Beutel et al., 2020).

[0010] 2) Each distributed processing node receives the model parameters and performs offline training locally, specifically: ,in: The learning rate is generally in the range of 10 -5 ~10 -2 ; is the local loss function of the i-th node, represents the model parameters of the i-th node in the t-th round; are the model parameters of the previous round.

[0011] The offline training described above uses the hyperparameter optimization library Optuna to optimize the loss function of the distributed processing node validation set using the Bayesian optimization algorithm. The optimization goal is to minimize the loss function of the validation set. An early stopping mechanism is used. If any experiment performs poorly in the validation phase, the system will terminate the experiment early to save computing resources.

[0012] The loss function is preferably , where n is the number of samples, is the true value of the i-th sample, is the predicted value of the i-th sample.

[0013] The early stopping mechanism is as follows: when the loss of the validation set decreases for five consecutive training runs or the loss of the current validation set is less than the optimal loss of the previous experiment, early stopping is performed.

[0014] The local data of the distributed processing node include: Weber number (Wev), Reynolds number (R eTP ), Froude number (Fr L ), Prandtl number (Prw), ( ), ( ), ( ), ( ), ( ), equilibrium gas content (x e ), ( ), boiling number (Bo), ( ).

[0015] 3) Each distributed processing node sends the model parameters back to the federated learning server for aggregation, and the federated learning server sends the aggregated model parameters to each distributed processing node again.

[0016] The aggregation preferably adopts the FedProx algorithm, specifically: , where: K is the total number of distributed processing nodes participating in the aggregation; is the number of local data samples of the kth distributed processing node; The total amount of data across all devices; No. The loss function of the distributed processing nodes; λ is the regularization coefficient, which is a positive real number that controls the weight of the regularization term and is generally between 0.001 and 1.00; is a regularization term used to measure the difference between the global model and the local model; is the global model parameter, For the The model parameters of each organization.

[0017] 4) Loop steps 2) and 3) until the stopping condition is met.

[0018] The stopping conditions include, but are not limited to, reaching a preset maximum number of polymerization rounds or the total loss being lower than a preset value.

[0019] The present invention relates to a post-drying heat transfer coefficient prediction system for implementing the above-mentioned method, comprising: a local data processing module and a local training module of the distributed processing node deployed on a distributed processing node, and a parameter initialization and model optimization module and a server-side aggregation and synchronization module deployed on the server side, wherein: the local data processing module receives original experimental data, cleans, extracts features, and performs maximum and minimum normalization processing on the data, and divides the data into a training set and a test set in proportion, and outputs standardized training samples; the parameter initialization and model optimization module performs global model initialization (using Kaiming initialization) and coordinates the Optuna library to implement hyperparameter search, determining structural parameters including initial learning rate, number of hidden layers, number of neurons in each layer, etc.; the local training module of the distributed processing node receives server-side model parameters, performs local model training in combination with data processed by the local data processing module, optimizes BP neural network parameters using the gradient descent method, and calculates a local loss function; the server-side aggregation and synchronization module receives local model parameters uploaded by each distributed processing node, updates global parameters based on the FedProx aggregation algorithm, and sends the updated parameters back to each distributed processing node for use in the next round of training.

[0020] Technical Effects

[0021] The present invention integrates the hyperparameter optimization method based on BP neural network (using Optuna Bayesian search mechanism) and FedProx aggregation strategy into the federated learning architecture to solve the problem of heat transfer coefficient prediction after drying. Without sharing the original data, a global model with a unified structure is constructed, and a regularization term aggregation mechanism is introduced to address the problem of non-independent and identically distributed (Non-IID) data, thereby improving the stability and consistency of cross-node model training. Compared with the prior art, the present invention achieves a significant reduction in the average MAPE from the traditional method (up to 70.55%) to no more than 15.88%, and the highest R² is increased to 0.9753, indicating that this method has stronger prediction accuracy and generalization ability under complex working conditions. Especially in scenarios where the distribution of non-IID data is severely unbalanced, the FedProx aggregation strategy exhibits better convergence performance, effectively alleviates the problem of model offset, and provides a deployable solution for high-privacy thermal computing modeling. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 Flowchart of the present invention;

[0023] Figure 2 This is a diagram of the hyperparameter optimization structure of the present invention;

[0024] Figure 3 This is a hyperparameter optimization flowchart of the present invention;

[0025] Figure 4 This is a schematic diagram of MAPE on the server side under the FedProx algorithm;

[0026] Figure 5 The server side of the FedProx algorithm Schematic diagram;

[0027] Figure 6 、 Figure 7 、 Figure 8 These are the frequency distribution histograms of the Prandtl numbers of the three node data sets participating in federated learning. DETAILED DESCRIPTION

[0028] like Figure 1As shown, this embodiment involves a method for predicting the heat transfer coefficient after drying based on federated learning. Each node pre-processes the raw data locally, including data cleaning and normalization, and then performs hyperparameter search and automatically tunes the BP neural network structure (such as learning rate, number of layers, number of neurons, etc.); after obtaining the optimal structure, the local BP neural network training process is started, and the network parameters are iteratively optimized using local data; after each node completes training locally, the model parameters are uploaded to the central server; the central server performs model aggregation operations (such as the FedProx algorithm) and generates new global model parameters; the new parameters are sent to each node, and the next round of training continues until the aggregation round is met or the loss converges. The entire process achieves collaborative modeling while ensuring that the data does not leave the local area, and is suitable for predicting the heat transfer coefficient after drying in multi-node and privacy-sensitive scenarios.

[0029] This embodiment specifically includes the following steps:

[0030] Step 1) Select multiple nodes (for example, three) to jointly train a shared BP neural network. During the offline phase, these nodes each collect raw data and build their own BP neural network framework. During the data processing phase, the three nodes locally divide the data into training and test sets, uniformly using 80% for training and 20% for testing. To account for the impact of varying dimensional ranges across features, the data is normalized using minimum and maximum normalization. , is the normalized value, For the sample, and are the sample maximum and sample minimum values, respectively.

[0031] Step 2) BP neural network parameter initialization. To improve the training stability and optimization efficiency of the model, the federated learning server uses the Kaiming initialization method to initialize the bias to zero and sends the initialized parameters to the three nodes through communication technology to ensure that the three nodes have the same initialization parameters.

[0032] The Kaming initialization method is to set the weight initialization of the current layer by calculating the variance of the activation value of the previous layer, specifically: the weight matrix , where: N is the normal distribution, n in is the number of input units of the current layer.

[0033] like Figure 2As shown, the BP neural network performs hyperparameter optimization on the federated learning server through the Optuna hyperparameter optimization platform before training. The Optuna hyperparameter optimization platform includes: a federated learning algorithm module that can select the FedAvg or FedProx algorithms for model parameter aggregation; a hyperparameter scheduling module that dynamically adjusts neural network structural parameters (such as learning rate and number of layers) through the Optuna hyperparameter optimization platform; an RPC communication interface that distributes and transmits parameters between the server and participating nodes via the remote call protocol (RPC); and participating nodes that deploy the Flower federated learning distributed processing node and run the BP neural network training process, using the underlying PyTorch framework. This structure enables centralized hyperparameter optimization and collaborative model updates across heterogeneous nodes, providing an efficient training strategy for the post-drying heat transfer coefficient prediction task.

[0034] like Figure 3 As shown, the hyperparameter optimization specifically includes:

[0035] Step a: Initialize the hyperparameter search space: Set the value range of the neural network hyperparameters in the Optuna hyperparameter optimization platform;

[0036] Step b: Generate a set of candidate parameters: Optuna uses a Bayesian optimization strategy to generate a set of candidate parameters.

[0037] Step c: Send parameters to the server: The server uses the current candidate parameters to build the BP neural network model structure;

[0038] Step d: Parameter broadcast to each node: Synchronize the model initial parameters to participating nodes through the federated learning framework;

[0039] Step e: Local model training: Each node performs several rounds of training using local data.

[0040] Step f: Aggregate and calculate validation error: The server collects model parameters and aggregates them, calculating the loss value on the validation set.

[0041] Step g updates the Bayesian model: based on the validation loss feedback, updates the posterior distribution of the hyperparameters;

[0042] Step h: Iterative search: Repeat the above process until the preset number of rounds is reached or the loss converges.

[0043] Step 3) The three nodes receive the parameters sent by the server and use the local training set to train the BP neural network for 1 to 2 rounds to avoid the model being overly dependent on the data of any one node, which would cause the model to perform well on the data of that node but poorly on the data of other nodes.

[0044] Step 4) Each node extracts the model weights and bias parameters obtained through local training and sends them to the server.

[0045] Step 5) After receiving the parameters uploaded by the three nodes, the server processes the parameters using an aggregation algorithm to obtain new BP neural network parameters.

[0046] like Figure 4 As shown, this embodiment uses the FedProx algorithm for aggregation, and records the change trend of the mean absolute percentage error (MAPE) indicator of the global model on the test set after each round of aggregation on the server side. It can be observed that in the initial training rounds (such as the first 20 rounds), the MAPE fluctuates greatly, which shows that the model has not yet converged. As the number of training rounds increases, the MAPE shows an overall downward trend and tends to stabilize after about 80 rounds, and finally maintains at about 10% in multiple training rounds, indicating that the federated prediction framework proposed in the present invention still maintains a high prediction accuracy and convergence performance with the participation of different distributed nodes. .

[0047] like Figure 5 As shown in the figure, under the FedProx algorithm aggregation mechanism, the variation of the coefficient of determination R² during the federated learning training process with each training round is recorded. As can be seen from the figure, the R² value increases overall and gradually approaches 1.0, indicating that the prediction model has strong fitting and explanatory power for the post-drying heat transfer coefficient. Although there is a certain degree of fluctuation in the first 20 rounds of training, the overall trend is steadily improving, eventually reaching and maintaining above 0.95, reflecting that the federated BP neural network model constructed by this invention has good generalization performance under heterogeneous data conditions. , where: y, and are the target value, predicted value and average target value respectively, and n is the number of samples.

[0048] Step 6) The server sends the new parameters to the three nodes, and the three nodes use the local data to train the BP neural network for 1-2 rounds to obtain their own BP neural network parameters.

[0049] Step 7) Repeat steps 3 to 6 until the preset aggregation rounds or other stopping criteria are reached, and the training ends.

[0050] After specific actual experiments, the system was used in Windows 11 system, i5-13400F processor, NVIDIA RTX 4070 Super GPU, and 32GB memory environment. Python 3.8 was used for model training, PyTorch 1.13.1 and Flower 1.7.0 frameworks were used to implement federated learning, and Optuna 3.0.3 was used for hyperparameter search.

[0051] The post-drying heat transfer coefficient was predicted using the following steps: During the data processing phase, the three nodes performed minimum-maximum normalization on the input features of their local data (BEC, BIS, and SWE, respectively). 80% of the data was divided into a training set and 20% into a test set. Before training, the BP neural network hyperparameters were tuned using the Bayesian optimization algorithm within the federated learning Flower framework. The hyperparameters that required tuning included the initial learning rate of the Adam optimizer, the number of hidden layers in the BP neural network, the number of neurons per layer, the number of training rounds per node for the three nodes, and the number of server aggregations. After hyperparameter optimization, each node determined its own local BP neural network structure. The final optimized hyperparameters were: initial learning rate: 0.000624; number of hidden layers: 7; number of neurons per layer: 274, 118, 143, 191, 232, 170, 176; number of training rounds per node: 2; and number of server aggregations: 144, as shown in Table 1.

[0052] Table 1

[0053] The Flower server selects the FedProx aggregation algorithm and uses the Kaiming initialization method to initialize the bias to zero. These initialization parameters are then sent to the three nodes via communication technology, ensuring that all three nodes have the same initialization parameters. After receiving the initial parameters, the three nodes deploy them to their local BP neural networks and perform one or two rounds of training using local data. After training, the weights and biases are extracted from the local neural network and sent back to the Flower server via communication. The server then uses the FedProx algorithm to perform a weighted average of the received parameters and biases to obtain the new round of parameters. These new parameters are then sent to the three nodes via communication. This process repeats until the pre-defined number of aggregation rounds is reached.

[0054] After the model is trained, the generalization ability of the model is tested, using MAPE and R 2 As shown in Table 2, the MAPEs on the three node test sets are 9.28%, 11.24%, and 15.88%. The MAPEs on the three nodes are 0.9514, 0.9417, and 0.8966, respectively. If the server uses the general FedAvg aggregation algorithm, the MAPEs on the three nodes are 17.15%, 37.8%, and 29.22%, respectively. They are 0.9585, 0.9662 and 0.82 respectively. It can be seen that FedProx performs well on all three nodes, while the FedAvg algorithm has a higher MAPE on nodes 2 and 3, and is Reaching,the lowest value.,The FedProx algorithm has a good coverage for the heterogeneous problem of,data federated learning.

[0055] Table 2

[0056] like Figures 6 to 8 As shown in Figure 2, the frequency distribution of the key thermal parameter, Prandtl number (Pr), in the local data of the three participating nodes (BEC, BIS, and SWE). Figure 6 The Prandtl number in the BEC data set is concentrated around 1.0, showing a clear unimodal skewed distribution, indicating that the data at this node is relatively concentrated and the physical properties of the samples are highly consistent. Figure 7 The BIS dataset in

[15] presents a multimodal discrete distribution with a large data span and drastic frequency changes, reflecting the typical non-independent and identically distributed (non-IID) characteristics of the node data, which poses a challenge to federated training. Figure 8 The SWE data shown covers a wider range and has a more uniform distribution of Prandtl numbers, reflecting that the data comes from a variety of complex experimental conditions. Figures 6 to 8 As can be seen, data distribution varies significantly between nodes, and direct merging and modeling would introduce significant bias. This paper introduces the FedProx aggregation strategy based on local training to effectively control parameter drift between different models. It also incorporates hyperparameter optimization methods to improve the model's stability and adaptability to heterogeneous data, thereby enhancing predictive performance.

[0057] Compared with existing technologies, the present invention adopts new technologies in the following key links and brings significant improvements: 1. In terms of model training structure, the present invention introduces the federated learning framework Flower, which enables multiple nodes to collaboratively train BP neural networks without sharing original data, solving the problems of privacy and heterogeneity of cross-node experimental data. 2. In terms of parameter optimization, the Optuna hyperparameter optimization algorithm is combined with federated learning for the first time to automatically search for the structure and learning rate parameters of the neural network, thereby improving the generalization ability and stability of the model. 3. In terms of model aggregation strategy, the present invention adopts the FedProx aggregation algorithm, introduces regularization constraints on the global model and the local model, and effectively alleviates the interference of non-independent and identically distributed (Non-IID) data on model convergence. In terms of actual performance, compared with traditional empirical correlation methods (such as Miropol'ski and Groeneveld, as shown in Table 3), the present invention reduces the MAPE from 40.40% to 70.55% to 8.34% to 15.88% on the three node test sets, and the R² is significantly improved. Compared with the standard federation strategy FedAvg, the FedProx solution used in the present invention reduces the MAPE by approximately 29% and 13% on nodes 2 and 3, respectively, and significantly improves the R² index, verifying the adaptability of the present invention to data heterogeneity problems and the prediction accuracy, as shown in Table 2.

[0058] Table 3

[0059] This paper combines the Optuna Bayesian optimization algorithm, BP neural network, Kaiming initialization method, FedProx federated aggregation strategy, and the Flower framework to construct a post-drying heat transfer coefficient prediction method suitable for heterogeneous data scenarios. This combination is not only the first systematic integration of post-drying modeling tasks, but also demonstrates excellent stability and prediction accuracy under non-independent and identically distributed (Non-IID) data conditions, resolving the issues of traditional methods prone to divergence and difficulty in parameter adjustment in cross-node modeling.

[0060] In summary, compared with existing technologies, this method effectively mitigates issues such as model non-convergence and overfitting caused by sample distribution differences and uneven data scale by introducing the FedProx regularization term at the model aggregation layer and incorporating a hyperparameter adaptive optimization mechanism at the local model training layer. Experimental results show that on three heterogeneous data sources, the prediction error (MAPE) significantly outperforms both the traditional empirical formula and the standard federated strategy (FedAvg), while maintaining a high R² value. This demonstrates the potential for this method to be widely adopted in engineering scenarios and its non-obvious technical contribution.

[0061] The above-mentioned specific implementation can be partially adjusted in different ways by those skilled in the art without departing from the principles and purpose of the present invention. The scope of protection of the present invention shall be based on the claims and shall not be limited by the above-mentioned specific implementation. All implementation schemes within its scope shall be subject to the constraints of the present invention.

Claims

1. A method for predicting heat transfer coefficient after drying based on federated learning, characterized in that: In the offline stage, multiple distributed processing nodes are constructed. Each node collects raw data and builds a local BP neural network model. Then, offline training and hyperparameter optimization are performed locally without sharing the raw data and only uploading the model parameters. The federated learning server uses an aggregation algorithm to aggregate the received model parameters to eliminate the problem of inconsistent data distribution between nodes. In the online stage, the federated learning server uses the aggregated model parameters to perform real-time prediction of the heat transfer coefficient after drying.

2. The method for predicting heat transfer coefficient after drying based on federated learning according to claim 1, characterized in that: The offline phase specifically includes: 1) The federated learning server initializes the model parameters of the BP neural network as the global model and sends them to each distributed processing node ; 2) Each distributed processing node receives the model parameters and performs offline training locally, specifically: ,in: Take the value of learning rate; is the local loss function of the i-th node, represents the model parameters of the i-th node in the t-th round; is the model parameter of the previous round; 3) Each distributed processing node sends the model parameters back to the federated learning server for aggregation, and the federated learning server sends the aggregated model parameters to each distributed processing node again; 4) Loop steps 2) and 3) until the stopping condition is met.

3. The method for predicting heat transfer coefficient after drying based on federated learning according to claim 2, wherein: The offline training described above uses the hyperparameter optimization library Optuna to optimize the loss function of the distributed processing node validation set using the Bayesian optimization algorithm. The optimization goal is to minimize the loss function of the validation set and adopt an early stopping mechanism, that is, to terminate the experiment early when the performance in the validation phase is poor to save computing resources.

4. The method for predicting heat transfer coefficient after drying based on federated learning according to claim 2, wherein: The loss function is , where n is the number of samples, is the true value of the i-th sample, is the predicted value of the i-th sample.

5. The method for predicting heat transfer coefficient after drying based on federated learning according to claim 2, wherein: The early stopping mechanism is as follows: when the loss of the validation set decreases for five consecutive training runs or the loss of the current validation set is less than the optimal loss of the previous experiment, early stopping is performed.

6. The method for predicting heat transfer coefficient after drying based on federated learning according to claim 2, wherein: The local data of the distributed processing node include: Weber number (Wev), Reynolds number (R eTP ), Froude number (Fr L ), Prandtl number (Prw), ( ), ( ), ( ), ( ), ( ), equilibrium gas content (x e ), ( ), boiling number (Bo), ( ).

7. The method for predicting heat transfer coefficient after drying based on federated learning according to claim 2, wherein: The aggregation adopts the FedProx algorithm, specifically: ,in: The number of participating devices, No. The amount of data per device, Total data volume across all devices , No. The loss function of distributed processing nodes; w is the global model parameter to be optimized on the server side; K is the total number of distributed processing nodes participating in the aggregation; is the number of local data samples of the kth distributed processing node; The total amount of data across all devices; No. The loss function of distributed processing nodes; is a regularization term used to measure the difference between the global model and the local model. The global model parameter Model parameters of the nodes The regularization term between , Regularization parameter, used to control the extent to which the local model deviates from the current parameters of the global model.

8. The method for predicting heat transfer coefficient after drying based on federated learning according to claim 2, wherein: The stopping conditions include reaching a preset maximum number of polymerization rounds or the total loss being lower than a preset value.

9. A system for predicting heat transfer coefficient after drying for implementing the method according to any one of claims 1 to 8, characterized in that: include: The local data processing module and the local training module of the distributed processing node are deployed on the distributed processing node, and the parameter initialization and model optimization module and the server-side aggregation and synchronization module are deployed on the server side. Among them: the local data processing module receives the original experimental data, cleans, extracts features, and performs maximum and minimum normalization processing on the data, and divides it into training sets and test sets in proportion, and outputs standardized training samples; the parameter initialization and model optimization module performs global model initialization, coordinates the Optuna library to implement hyperparameter search, and determines the structural parameters; the local training module of the distributed processing node receives the server-side model parameters, combines the data processed by the local data processing module to perform local model training, uses the gradient descent method to optimize the BP neural network parameters, and calculates the local loss function; the server-side aggregation and synchronization module receives the local model parameters uploaded by each distributed processing node, updates the global parameters based on the FedProx aggregation algorithm, and sends the updated parameters back to each distributed processing node for use in the next round of training.