Distributed model training method and device and electronic equipment
Through the federated averaging algorithm and security aggregation algorithm, the problems of high data transmission cost, large latency and difficult data privacy protection in distributed model training are solved, and efficient model training and data privacy protection are achieved.
Patent Information
- Application Number
- CN202510165140.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-27
AI Technical Summary
The existing distributed model training methods have high data transmission costs, large communication delays, low model training efficiency, and are difficult to meet the privacy protection needs of distributed computing nodes for local data.
The federated averaging algorithm is used to weight the model update information of each distributed computing node, generate global model parameters, and encrypt it through a secure aggregation algorithm to avoid data leakage.
It reduces the data transmission cost and communication delay of the model training process, improves the model training efficiency, and at the same time realizes data privacy protection of distributed computing nodes, avoiding data security risks.
Smart Images

Figure CN120218184A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and particularly to a distributed model training method, apparatus, and electronic device. Background Art
[0002] With the rapid development of artificial intelligence (AI) technology, the training and deployment of AI models have become an important support for the digital transformation of various industries. At present, with the wide application of edge computing devices and Internet of Things (IoT) devices, the training and deployment of AI models need to be carried out in a distributed and multi-node environment.
[0003] Most of the existing distributed model training methods adopt centralized training. Centralized training usually requires aggregating local data collected by a large number of distributed computing nodes to a central server, training the model by the central server according to the aggregated local data, and then the central server distributes the trained model to each distributed computing node, and each distributed computing node can deploy the received model to run locally.
[0004] Since a large amount of local data collected by each distributed computing node needs to be sent to the central server, this will increase the data transmission cost and communication delay in the model training process, resulting in low model training efficiency; at the same time, since the local data of each distributed computing node can be directly obtained by the central server, the data privacy of the distributed computing nodes is leaked, and there are data security risks. Especially in scenarios involving sensitive data such as medical and financial data, the centralized training method cannot meet the privacy protection requirements of distributed computing nodes for local data.
[0005] Therefore, the existing distributed model training methods have high data transmission costs, large communication delays, and low model training efficiency, and it is difficult to meet the privacy protection requirements of distributed computing nodes for local data. Summary of the Invention
[0006] The present invention provides a distributed model training method, apparatus, and electronic device to solve the defects in the existing distributed model training methods, such as high data transmission cost, large communication delay, low model training efficiency, and difficulty in meeting the privacy protection requirements of distributed computing nodes for local data.
[0007] The present invention provides a distributed model training method, which is applied to a central server. The central server communicates with multiple distributed computing nodes. The method includes: receiving model update information of each distributed computing node; a model update information is obtained by training based on local data collected by a distributed computing node; based on the federated averaging algorithm, performing weighted averaging on each model update information to generate global model parameters; generating a global model based on the global model parameters; and sending the global model to each distributed computing node respectively, so that each distributed computing node deploys the global model.
[0008] According to the distributed model training method provided by the present invention, each model update information is encrypted based on the secure aggregation algorithm; based on the federated averaging algorithm, performing weighted averaging on each model update information to generate global model parameters, including: decrypting each model update information respectively to obtain a plurality of target model update information; and based on the federated averaging algorithm, performing weighted averaging on each target model update information to generate global model parameters.
[0009] According to the distributed model training method provided by the present invention, generating a global model based on the global model parameters includes: generating an initial global model based on the global model parameters; obtaining the computing resource information of each distributed computing node, the task load information of each distributed computing node, and the network resource information of each distributed computing node; and optimizing the initial global model based on the computing resource information of each distributed computing node, the task load information of each distributed computing node, and the network resource information of each distributed computing node to generate a global model.
[0010] According to the distributed model training method provided by the present invention, before sending the global model to each distributed computing node respectively, it further includes: obtaining the hardware configuration information of each distributed computing node and the network resource information of each distributed computing node; and performing compression processing and quantization processing on the global model based on the hardware configuration information of each distributed computing node and the network resource information of each distributed computing node.
[0011] According to the distributed model training method provided by the present invention, before receiving the model update information of each distributed computing node, it further includes: generating initial model parameters and a training task; and sending the initial model parameters and the training task to each distributed computing node respectively, so that each distributed computing node collects local data based on the training task, and trains a local model based on the local data and the initial model parameters to obtain the model update information of the local model.
[0012] According to the distributed model training method provided by the present invention, the local model is obtained by performing quantization-aware training based on local data, initial model parameters, and noise generated based on the differential privacy mechanism.
[0013] A distributed model training method provided by the present invention, after sending the global model to each distributed computing node respectively so that each distributed computing node deploys the global model, further includes: receiving model feedback data of each distributed computing node; based on the model feedback data of each distributed computing node, iteratively optimizing the global model to obtain an optimized global model; sending the optimized global model to each distributed computing node respectively so that each distributed computing node updates and deploys the optimized global model.
[0014] A distributed model training method provided by the present invention, when a distributed computing node runs the global model, generates interpretive analysis information of the global model through the SHAP algorithm.
[0015] The present invention also provides a distributed model training device, including: a receiving module, configured to receive model update information of each distributed computing node; a model update information is obtained by training based on local data collected by a distributed computing node; a federated learning module, configured to perform weighted averaging on each model update information based on the federated averaging algorithm to generate global model parameters; a generating module, configured to generate a global model based on the global model parameters; a sending module, configured to send the global model to each distributed computing node respectively so that each distributed computing node deploys the global model.
[0016] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, it implements any one of the above-mentioned distributed model training methods.
[0017] For the distributed model training method, device, and electronic device provided by the present invention, after each distributed computing node obtains the corresponding model update information by training based on the collected local data, it sends the model update information to the central server, and the central server uses the federated averaging algorithm to perform weighted averaging on the model update information corresponding to each distributed computing node to generate global model parameters, and then generates a global model, and then distributes the trained global model to each distributed computing node. In this way, each distributed computing node does not need to send a large amount of local data to the central server, but only transmits the trained model update information, reducing the data transmission cost and communication delay in the model training process, effectively improving the model training efficiency. At the same time, since the central server cannot directly obtain the local data of each distributed computing node, the possibility of data privacy leakage of the distributed computing node is avoided, reducing the data security risk, and can meet the privacy protection requirements of the distributed computing node for local data. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0019] Figure 1 It is a schematic flowchart of the distributed model training method provided by the present invention.
[0020] Figure 2 It is a schematic system architecture diagram of the distributed AI model training and inference integrated platform provided by the present invention.
[0021] Figure 3 It is a schematic working flowchart of the distributed AI model training and inference integrated platform provided by the present invention.
[0022] Figure 4 It is a schematic structural diagram of the distributed model training device provided by the present invention.
[0023] Figure 5 It is a schematic structural diagram of the electronic device provided by the present invention. Specific embodiments
[0024] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0025] Please refer to Figures 1 to 3 , Figure 1 It is a schematic flowchart of the distributed model training method provided by the present invention, Figure 2 It is a schematic system architecture diagram of the distributed AI model training and inference integrated platform provided by the present invention, Figure 3 It is a schematic working flowchart of the distributed AI model training and inference integrated platform provided by the present invention. In this embodiment, the distributed model training method is applied to the central server of the distributed AI model training and inference integrated platform. The central server communicates with multiple distributed computing nodes of the distributed AI model training and inference integrated platform. The distributed model training method includes steps S110 to S140, and the specific steps are as follows: S110: Receive the model update information of each distributed computing node.
[0026] A model update information is obtained by training based on the local data collected by a distributed computing node.
[0027] As Figure 2 shown, the distributed AI model training and inference integration platform consists of multiple distributed computing nodes and a central server. Each distributed computing node and the central server communicate and cooperate through a secure network connection. The system architecture of the distributed AI model training and inference integration platform mainly includes multiple distributed computing nodes, a central server, an integrated training and inference module, a security and interpretability module, and an intelligent scheduling and management module.
[0028] Among them, the functions and working processes of each module are as follows: (1) Distributed computing nodes: Each distributed computing node is configured with a local data management module, a local model training module, a federated learning client module, and a secure aggregation client module.
[0029] The local data management module is used to manage and store the private data (i.e., local data) of the corresponding distributed computing node, and can perform data preprocessing, feature extraction, and data privacy protection on the local data, such as data encryption and differential privacy protection.
[0030] The local model training module is used to train the local AI model according to the local data using the federated learning algorithm. The model training uses the federated learning algorithm, which can ensure that each distributed computing node can cooperate to train the global model without uploading the local data. The local model training module integrates a lightweight deep learning framework and supports efficient operation in resource-constrained environments.
[0031] The federated learning client module is used to communicate with the central server, and can regularly upload the model update information (such as model gradients or weights) of the local model to the central server, and receive the global model aggregated by the central server for local deployment and update.
[0032] To ensure the security of the uploaded model, the secure aggregation client module can use the secure aggregation algorithm to encrypt the model update information to be uploaded, ensuring that the model update information is not stolen or tampered with during the transmission process.
[0033] (2) Central server: The central server is configured with a federated learning server module, a global model optimization module, and a resource scheduling and load balancing module.
[0034] The federated learning server module is the core module of federated learning. The federated learning server module can receive the encrypted model update information uploaded by each distributed computing node, decrypt and aggregate the encrypted model update information using the Secure Aggregation algorithm, then perform weighted averaging on each model update information based on the federated averaging algorithm to generate global model parameters, and generate a global model based on the global model parameters. The generated global model can then be distributed to each distributed computing node via the federated learning client module.
[0035] The central server can further optimize and adjust the received model update information through the global model optimization module: this module can combine the adaptive federated learning strategy, and dynamically adjust the weight allocation and parameter update of the global model according to the computing resources, task load, data quality, and network resources of each distributed computing node to optimize the global model and ensure the efficiency and accuracy of the global model.
[0036] The resource scheduling and load balancing module optimizes the global resource allocation and task scheduling by real-time monitoring the resource usage and network load of each distributed computing node, and adopts deep learning prediction algorithms and reinforcement learning techniques to ensure the stable operation of the platform and the maximization of resource utilization.
[0037] (3) Integrated training and inference module: The integrated training and inference module can be divided into a model packaging and pushing module and an edge computing support module.
[0038] The model packaging and pushing module is deployed on the central server. After the central server completes the optimization of the global model, the model packaging and pushing module can package the optimized global model and automatically push it to each distributed computing node. The model packaging and pushing module can perform adaptive compression processing and quantization processing on the global model according to the hardware configuration information and network resource information of each distributed computing node to ensure that the global model can adapt to the hardware configuration and network conditions of different distributed computing nodes.
[0039] Each distributed computing node is configured with an edge computing support module, which is used to optimize the model running performance on the edge distributed computing node. It adopts lightweight model compression and quantization techniques, such as quantization aware training (QAT) technology, to reduce the computing burden of the model on the edge distributed computing node device and ensure the real-time performance and response speed of the model.
[0040] (4) Security and interpretability module: Each distributed computing node is configured with a security and interpretability module, which can be divided into a differential privacy module and an interpretable AI module.
[0041] To further enhance data privacy protection, the differential privacy module can introduce noise during the training process of the local model to ensure the security of training data and prevent malicious nodes from obtaining local data through reverse inference.
[0042] The interpretable AI module enables users to monitor the behavior of the globally deployed model locally in real time and provides an explanatory analysis of the model's decision-making process. When the distributed computing nodes run the global model, this module can generate explanatory analysis information of the global model through the SHAP (SHapley Additive exPlanations) algorithm to help users understand the basis for the model output results and improve the transparency and credibility of the model.
[0043] (5) Intelligent Scheduling and Management Module: The intelligent scheduling and management module is deployed on the central server and can be divided into a task management module and a monitoring and feedback module.
[0044] The task management module is used to schedule and manage the training tasks of the central server. By dynamically allocating training tasks, it ensures the load balance of each distributed computing node and realizes the efficient operation of the platform.
[0045] The monitoring and feedback module can monitor the running status of each distributed computing node in the platform in real time, collect key performance indicators, and adjust the model parameters according to the feedback information to maintain the stability and performance optimization of the platform.
[0046] It should be noted that the local model training module, federated learning client module, secure aggregation client module, differential privacy module, and federated learning server module in this embodiment can jointly form a federated learning module. The federated learning module is the core component of the present invention, aiming to achieve the collaborative training and deployment of distributed AI models. Especially in a multi-node distributed environment, it ensures data privacy and training efficiency. During the application process of federated learning, each distributed computing node first uses the local model training module to independently perform local model training using local data. This process includes data preprocessing, model initialization, and iterative training to ensure that the distributed computing node can update the model parameters using its private data. After the training is completed, each distributed computing node does not upload the original local data, but uploads the model update information to the central server through the federated learning client module. To ensure the security of this information during transmission, the secure aggregation client module will first encrypt the model update information to prevent it from being intercepted or tampered with during network transmission. In addition, the differential privacy module will introduce noise during the local model training process to further protect data privacy, ensuring that even if the data is obtained, the original training data cannot be inferred. The central server receives the encrypted model update information uploaded by each distributed computing node through the federated learning server module and uses the federated averaging algorithm to perform weighted averaging on this information to generate a global model. During the process of generating the global model, the central server will also dynamically adjust the optimization strategy of the global model in combination with the computing power, data quality, and network conditions of each distributed computing node to ensure the adaptability and accuracy of the model on different distributed computing nodes. The optimized global model is then pushed to each distributed computing node. After each distributed computing node receives these parameters, it will update the local model and thus enter the next round of training iteration.
[0047] The innovation of the federated learning module lies in its integration of multiple technologies, including adaptive federated learning strategies, secure aggregation technology, differential privacy technology, and lightweight edge computing technology. The adaptive federated learning strategy improves the optimization efficiency and adaptability of the global model by dynamically adjusting the weights and participation frequencies of each distributed computing node in model updates, especially suitable for situations where the resources of each distributed computing node are uneven or the data heterogeneity is relatively high. The combination of secure aggregation and differential privacy technologies provides multi-level data privacy protection for federated learning, preventing data from being leaked or reverse-inferred by malicious nodes in a distributed environment. In addition, the lightweight edge computing technology enables edge distributed computing nodes to participate in the federated learning process and complete the efficient training and deployment of the model without significantly increasing the computing burden. These innovations enable the federated learning module not only to effectively protect data privacy but also to achieve efficient AI model training and deployment in multi-scenario and multi-node environments, expanding the scope and effect of federated learning in practical applications.
[0048] Through the above system architecture design, the distributed AI model training and deployment integration platform can efficiently and securely train and deploy AI models in a multi-node distributed environment, ensuring the high performance of the models, the security of the data, and the optimal utilization of system resources.
[0049] Understandably, as Figure 3 shown, based on the above system architecture, when model training is required, the central server can generate initial model parameters and corresponding training tasks, and send the initial model parameters and training tasks to each distributed computing node respectively.
[0050] After each distributed computing node receives the initial model parameters and training tasks, it can perform data collection and data preprocessing according to the training tasks: each distributed computing node can collect local data from its affiliated local environment based on the training tasks, where the collected local data may include images, texts, sensor data, etc., which may vary according to different specific application scenarios; after the local data collection is completed, the local data management module of each distributed computing node can perform data preprocessing on the collected local data, and the steps of data preprocessing can include data cleaning, feature extraction, data normalization, etc., to facilitate subsequent model training; at the same time, the local data management module will apply the Differential Privacy technology to generate noise and introduce the noise during the data preprocessing process to protect data privacy, and each distributed computing node can utilize local computing resources to perform local model training based on the collected local data, the initial model parameters sent by the central server, and the noise generated based on the differential privacy mechanism.
[0051] During the local model training process, conventional deep learning training algorithms (such as gradient descent) can be used to iteratively update the local model parameters, and the Quantization-Aware Training (QAT) method can be adopted to reduce the computational burden of the local model and improve the training efficiency. After one round or multiple rounds of training are completed, each distributed computing node can obtain the trained local model and the model update information of the local model.
[0052] Furthermore, each distributed computing node can upload the model update information (such as model gradients or weights) generated by the training to the central server through the corresponding federated learning client module.
[0053] Preferably, in order to protect the data security of the model update information, each distributed computing node can encrypt the model update information through a secure aggregation algorithm to ensure that the model update information is not stolen or tampered with during the transmission process.
[0054] Furthermore, the central server can receive the model update information uploaded by each distributed computing node to achieve global model aggregation.
[0055] S120: Based on the federated averaging algorithm, perform weighted averaging on each model update information to generate global model parameters.
[0056] Specifically, the federated learning server module of the central server can adopt the Federated Averaging (FedAvg) algorithm to perform weighted averaging on each aggregated model update information to generate global model parameters.
[0057] S130: Generate a global model based on the global model parameters.
[0058] The central server can generate an initial global model based on the global model parameters, and further optimize the initial global model through the global model optimization module. The optimization process needs to consider the data quality, computing power, and network conditions of each distributed computing node, and dynamically adjust the weights of each distributed computing node in the update of the initial global model to ensure the optimal performance of the global model.
[0059] Specifically, the central server can optimize the initial global model according to the computing resource information of each distributed computing node, the task load information of each distributed computing node, and the network resource information of each distributed computing node, adjust the weights of each distributed computing node in the update of the initial global model, and generate a global model.
[0060] S140: Send the global model to each distributed computing node respectively, so that each distributed computing node deploys the global model.
[0061] Preferably, after the optimization of the global model is completed, the model packaging and pushing module of the central server can perform adaptive processing such as compression and quantization on the global model according to the hardware configuration information of each distributed computing node and the network resource information of each distributed computing node to ensure that the model can run efficiently in the edge device or the node environment with limited computing power; after the adaptive processing is completed, the central server can send the global model to each distributed computing node respectively, so that each distributed computing node deploys the global model and completes the rapid online of the global model on each distributed computing node.
[0062] In the distributed model training method provided in this embodiment, after each distributed computing node trains to obtain corresponding model update information based on the collected local data, it sends the model update information to the central server. The central server uses the federated averaging algorithm to perform weighted averaging on the model update information corresponding to each distributed computing node to generate global model parameters, and then generates a global model, and then distributes the trained global model to each distributed computing node. In the above manner, each distributed computing node does not need to send a large amount of local data to the central server, but only transmits the trained model update information, reducing the data transmission cost and communication delay in the model training process, effectively improving the model training efficiency. At the same time, since the central server cannot directly obtain the local data of each distributed computing node, the possibility of data privacy leakage of the distributed computing node is avoided, and the data security risk is reduced, which can meet the need of the distributed computing node for privacy protection of local data.
[0063] In some embodiments, each model update information is encrypted based on the secure aggregation algorithm; based on the federated averaging algorithm, performing weighted averaging on each model update information to generate global model parameters includes: decrypting each model update information respectively to obtain a plurality of target model update information; based on the federated averaging algorithm, performing weighted averaging on each target model update information to generate global model parameters.
[0064] When model training is required, the central server can generate initial model parameters and corresponding training tasks, and send the initial model parameters and training tasks to each distributed computing node respectively.
[0065] After each distributed computing node receives the initial model parameters and training tasks, it can collect local data from its affiliated local environment according to the training tasks, and perform local model training based on the collected local data, the initial model parameters sent by the central server, and the noise generated based on the differential privacy mechanism, to obtain the trained local model and the model update information of the local model.
[0066] After obtaining the model update information, each distributed computing node encrypts the model update information through the secure aggregation algorithm and uploads the encrypted model update information to the central server.
[0067] After the central server receives the encrypted model update information uploaded by each distributed computing node, it needs to decrypt each model update information respectively to obtain a plurality of target model update information, and the target model update information is the model update information restored by decryption.
[0068] Furthermore, the central server can perform weighted averaging on each target model update information based on the federated averaging algorithm to generate global model parameters.
[0069] In some embodiments, generating a global model based on global model parameters includes: generating an initial global model based on the global model parameters; obtaining the computing resource information of each distributed computing node, the task load information of each distributed computing node, and the network resource information of each distributed computing node; and optimizing the initial global model based on the computing resource information of each distributed computing node, the task load information of each distributed computing node, and the network resource information of each distributed computing node to generate a global model.
[0070] After obtaining the global model parameters, the central server can generate an initial global model based on the global model parameters and further optimize the initial global model through a global model optimization module. The optimization process needs to consider the data quality, computing power, and network conditions of each distributed computing node, and dynamically adjust the weights of each distributed computing node in the update of the initial global model to ensure the optimal performance of the global model.
[0071] Specifically, during the operation of the platform, the intelligent scheduling and management module can monitor the computing resource usage, task load, and network status of each distributed computing node in real time to obtain the computing resource information, task load information, and network resource information of each distributed computing node; the central server can optimize the initial global model according to the computing resource information of each distributed computing node, the task load information of each distributed computing node, and the network resource information of each distributed computing node, adjust the weights of each distributed computing node in the update of the initial global model, and generate a global model.
[0072] In some embodiments, before sending the global model to each distributed computing node respectively, it further includes: obtaining the hardware configuration information of each distributed computing node and the network resource information of each distributed computing node; and performing compression processing and quantization processing on the global model based on the hardware configuration information of each distributed computing node and the network resource information of each distributed computing node.
[0073] After completing the optimization of the global model, the model packaging and pushing module of the central server can perform adaptive processing such as compression and quantization on the global model according to the hardware configuration information of each distributed computing node and the network resource information of each distributed computing node to ensure that the model can run efficiently in an edge device or a node environment with limited computing power; after completing the adaptive processing, the central server can send the global model to each distributed computing node respectively, so that each distributed computing node deploys the global model and completes the rapid online deployment of the global model on each distributed computing node.
[0074] In some embodiments, before receiving the model update information of each distributed computing node, it further includes: generating initial model parameters and a training task; sending the initial model parameters and the training task to each distributed computing node respectively, so that each distributed computing node collects local data based on the training task, and trains a local model based on the local data and the initial model parameters to obtain the model update information of the local model.
[0075] When model training is required, the central server can generate initial model parameters and corresponding training tasks, and send the initial model parameters and the training tasks to each distributed computing node respectively.
[0076] Optionally, during the operation of the platform, the intelligent scheduling and management module can monitor the computing resource usage, task load, and network status of each distributed computing node in real time, and according to the monitored information, adopt deep learning prediction algorithms and reinforcement learning (such as Deep Q-Learning) technologies to dynamically adjust the allocation of training tasks and the scheduling of resources, ensuring that the platform can still maintain stability and high efficiency under high load conditions, and realizing intelligent scheduling and load balancing of resources.
[0077] After each distributed computing node receives the initial model parameters and the training task, it can perform data collection and data preprocessing according to the training task: each distributed computing node can collect local data from its local environment based on the training task, where the collected local data may include images, texts, sensor data, etc., which may vary according to different specific application scenarios; after the local data collection is completed, the local data management module of each distributed computing node can perform data preprocessing on the collected local data, and the steps of data preprocessing can include data cleaning, feature extraction, data standardization, etc., for the convenience of subsequent model training; at the same time, the local data management module will apply differential privacy technology to generate noise and introduce noise during the data preprocessing process to protect data privacy, and each distributed computing node can use local computing resources to perform local model training based on the collected local data, the initial model parameters sent by the central server, and the noise generated based on the differential privacy mechanism.
[0078] During the local model training process, conventional deep learning training algorithms (such as gradient descent) can be used to iteratively update the local model parameters, and the quantization aware training (QAT) method can be adopted to reduce the computational burden of the local model and improve the training efficiency. After one round or multiple rounds of training are completed, each distributed computing node can obtain the trained local model and the model update information of the local model.
[0079] Furthermore, each distributed computing node can upload the model update information (such as model gradients or weights) generated during training to the central server through the corresponding federated learning client module.
[0080] Preferably, to protect the data security of the model update information, each distributed computing node can encrypt the model update information through a secure aggregation algorithm to ensure that the model update information is not stolen or tampered with during transmission.
[0081] Furthermore, the central server can receive the model update information uploaded by each distributed computing node to achieve global model aggregation.
[0082] In some embodiments, the local model is obtained through quantization-aware training based on local data, initial model parameters, and noise generated based on the differential privacy mechanism.
[0083] In some embodiments, after the global model is sent to each distributed computing node respectively so that each distributed computing node deploys the global model, it further includes: receiving the model feedback data of each distributed computing node; iteratively optimizing the global model based on the model feedback data of each distributed computing node to obtain an optimized global model; sending the optimized global model to each distributed computing node respectively so that each distributed computing node updates and deploys the optimized global model.
[0084] Specifically, after each distributed computing node deploys the global model in the local environment, the user can run the global model on the local client of the distributed computing node. During the running process of the global model, the central server can continuously collect the model feedback data of each distributed computing node. The model feedback data includes model performance metrics, resource usage, and user experience, etc.
[0085] Furthermore, the central server can iteratively optimize the global model based on the model feedback data of each distributed computing node to obtain an optimized global model; the central server will regularly release the optimized global model and distribute the optimized global model to each distributed computing node respectively so that each distributed computing node updates and deploys the optimized global model to form a continuous closed-loop of training, optimization, and deployment.
[0086] In some embodiments, when the distributed computing node runs the global model, it generates interpretive analysis information of the global model through the SHAP algorithm.
[0087] Specifically, when the distributed computing node runs the global model, the explainable AI module can use the SHAP algorithm to generate interpretive analysis information of the global model to help users understand the decision-making process of the model and make adjustments when necessary.
[0088] The distributed model training method provided in this embodiment improves the training efficiency, security, flexibility, and adaptability of the distributed AI model training and inference integration platform by combining multiple technologies. It can achieve efficient training and deployment of AI models in complex environments with multiple nodes and multiple scenarios. Compared with the prior art, the distributed model training method provided in this embodiment has at least the following technical advantages: (1)Significantly enhanced data privacy protection: By combining the federated learning algorithm, differential privacy technology, and secure aggregation technology, the present invention achieves high-level data privacy protection in a multi-node distributed environment. Each distributed computing node does not need to upload the original data, but only needs to upload the encrypted model update information to participate in the training of the global model, avoiding the risk of data leakage. In addition, differential privacy technology introduces noise during the training process, which can effectively avoid the possibility of malicious nodes obtaining the original data through reverse inference, further improving the data security.
[0089] (2)Improved model training and deployment efficiency: The federated learning module proposed in the present invention uses an adaptive federated learning strategy to dynamically adjust the weight distribution and participation frequency during model training according to the computing power, data quality, and network conditions of each distributed computing node. This adaptive optimization strategy not only improves the training efficiency of the global model but also ensures the adaptability and performance of the model in different distributed computing node environments. Through the integrated training and inference process, the global model can be quickly pushed to each distributed computing node to achieve efficient model deployment.
[0090] (3)Intelligent resource scheduling and load balancing: Through the intelligent resource scheduling and load balancing mechanism, the resource usage, task load, and network conditions of each distributed computing node are monitored in real time, and the distribution of resources and task scheduling are dynamically adjusted using deep learning and reinforcement learning algorithms. This intelligent scheduling method not only improves the resource utilization rate of the platform but also ensures the efficiency of training task execution and system stability in a multi-node environment.
[0091] (4)Edge computing support and model lightweighting: For edge computing scenarios, the present invention supports lightweight model training and deployment on edge devices with limited computing resources by introducing technologies such as quantization-aware training (QAT). This edge computing support enables the distributed system to be applied in more scenarios, especially in the Internet of Things and edge computing environments, enhancing the practical application value and flexibility of the model.
[0092] (5) Enhanced system security and reliability: Through multi-level security measures, the present invention significantly enhances the overall security of the system. The secure aggregation technology ensures that during the federated learning process, the communication and data transmission between nodes are encrypted, preventing data from being intercepted or tampered with during transmission. In addition, the system also supports the asynchronous federated learning mode, which can effectively address the issues of communication latency and uneven computing resources in a distributed environment, improving the robustness and reliability of the system.
[0093] (6) Improved model interpretability: During the model deployment and operation process, the present invention integrates interpretability AI technology, providing real-time explanations of the model decision-making process through the SHAP (SHapley Additive exPlanations) algorithm. Users can intuitively understand the basis for the model output results, which not only improves the transparency and user trust of the model but also provides a basis for further optimization of the model.
[0094] (7) Wide application adaptability: The system architecture and workflow of the present invention are highly flexible and can adapt to a variety of application scenarios, including but not limited to intelligent monitoring, healthcare, financial risk control, and smart city construction. The modular design and adaptive optimization capabilities of the system enable it to be customized and extended according to different application requirements, with broad market application prospects.
[0095] The present invention also provides a distributed model training device. Please refer to Figure 4 , Figure 4 is a schematic structural diagram of the distributed model training device provided by the present invention. In this embodiment, the distributed model training device includes a receiving module 410, a federated learning module 420, a generating module 430, and a sending module 440.
[0096] The receiving module 410 is used to receive the model update information of each distributed computing node.
[0097] A model update information is trained based on the local data collected by a distributed computing node.
[0098] The federated learning module 420 is used to perform weighted averaging on each model update information based on the federated average algorithm to generate global model parameters.
[0099] The generating module 430 is used to generate a global model based on the global model parameters.
[0100] The sending module 440 is used to send the global model to each distributed computing node respectively, so that each distributed computing node deploys the global model.
[0101] In some embodiments, each model update information is encrypted based on the secure aggregation algorithm.
[0102] The federated learning module 420 is configured to decrypt each model update information respectively to obtain a plurality of target model update information; and generate global model parameters by performing weighted average on each target model update information based on the federated averaging algorithm.
[0103] In some embodiments, the generation module 430 is configured to generate an initial global model based on the global model parameters; obtain the computing resource information of each distributed computing node, the task load information of each distributed computing node, and the network resource information of each distributed computing node; and optimize the initial global model based on the computing resource information of each distributed computing node, the task load information of each distributed computing node, and the network resource information of each distributed computing node to generate a global model.
[0104] In some embodiments, the sending module 440 is configured to obtain the hardware configuration information of each distributed computing node and the network resource information of each distributed computing node; and perform compression processing and quantization processing on the global model based on the hardware configuration information of each distributed computing node and the network resource information of each distributed computing node.
[0105] In some embodiments, the sending module 440 is configured to generate initial model parameters and training tasks; and send the initial model parameters and training tasks to each distributed computing node respectively, so that each distributed computing node collects local data based on the training tasks, and trains a local model based on the local data and the initial model parameters to obtain model update information of the local model.
[0106] In some embodiments, the local model is obtained by performing quantization-aware training based on local data, initial model parameters, and noise generated based on the differential privacy mechanism.
[0107] In some embodiments, the distributed model training device further includes an optimization module.
[0108] The optimization module is configured to receive model feedback data of each distributed computing node; perform iterative optimization on the global model based on the model feedback data of each distributed computing node to obtain an optimized global model; and send the optimized global model to each distributed computing node respectively, so that each distributed computing node updates and deploys the optimized global model.
[0109] In some embodiments, when the distributed computing node runs the global model, interpretive analysis information of the global model is generated through the SHAP algorithm.
[0110] The present invention also provides an electronic device. Figure 5 is a schematic structural diagram of the electronic device provided by the present invention, as Figure 5As shown, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communications interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 can call the logical instructions in the memory 530 to execute the distributed model training method.
[0111] In addition, when the logical instructions in the above-mentioned memory 530 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0112] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.
[0113] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disks, optical discs, etc., and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments.
[0114] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A distributed model training method, characterized in that: Applied to a central server, the central server communicating with a plurality of distributed computing nodes, the method comprising: Receiving model update information of each of the distributed computing nodes; one of the model update information is obtained based on training of local data collected by one of the distributed computing nodes; Based on the federated average algorithm, weighted average is performed on each of the model update information to generate global model parameters; Based on the global model parameters, generating a global model; The global model is sent to each of the distributed computing nodes respectively, so that each of the distributed computing nodes deploys the global model.
2. The distributed model training method according to claim 1, characterized in that: Each of the model update information is encrypted based on a secure aggregation algorithm; The method of performing weighted averaging on each model update information based on the federated averaging algorithm to generate global model parameters includes: Decrypting each of the model update information respectively to obtain multiple target model update information; Based on the federated averaging algorithm, weighted averaging is performed on each target model update information to generate global model parameters.
3. The distributed model training method according to claim 1, characterized in that: The generating of the global model based on the global model parameters comprises: Based on the global model parameters, generating an initial global model; Acquire computing resource information of each of the distributed computing nodes, task load information of each of the distributed computing nodes, and network resource information of each of the distributed computing nodes; Based on the computing resource information of each of the distributed computing nodes, the task load information of each of the distributed computing nodes, and the network resource information of each of the distributed computing nodes, the initial global model is optimized to generate the global model.
4. The distributed model training method according to claim 1, characterized in that: Before sending the global model to each of the distributed computing nodes respectively, the method further includes: Acquire hardware configuration information of each of the distributed computing nodes and network resource information of each of the distributed computing nodes; Based on the hardware configuration information of each of the distributed computing nodes and the network resource information of each of the distributed computing nodes, the global model is compressed and quantized.
5. The distributed model training method according to claim 1, characterized in that: Before receiving the model update information of each distributed computing node, the method further includes: Generate initial model parameters and training tasks; The initial model parameters and the training tasks are sent to each of the distributed computing nodes respectively, so that each of the distributed computing nodes collects the local data based on the training tasks, and trains the local model based on the local data and the initial model parameters to obtain model update information of the local model.
6. The distributed model training method according to claim 5, characterized in that: The local model is obtained by performing quantization-aware training based on the local data, the initial model parameters and noise generated based on a differential privacy mechanism.
7. The distributed model training method according to claim 1, characterized in that: After sending the global model to each of the distributed computing nodes respectively so that each of the distributed computing nodes deploys the global model, the method further includes: Receiving model feedback data from each of the distributed computing nodes; Iteratively optimizing the global model based on the model feedback data of each of the distributed computing nodes to obtain an optimized global model; The optimized global model is sent to each of the distributed computing nodes respectively, so that each of the distributed computing nodes updates and deploys the optimized global model.
8. The distributed model training method according to claim 1, characterized in that: When the distributed computing node runs the global model, explanatory analysis information of the global model is generated through the SHAP algorithm.
9. A distributed model training device, characterized in that: include: A receiving module, used for receiving model update information of each distributed computing node; One of the model update information is obtained based on local data training collected by one of the distributed computing nodes; A federated learning module, used for performing weighted averaging on each of the model update information based on a federated averaging algorithm to generate global model parameters; A generating module, used for generating a global model based on the global model parameters; The sending module is used to send the global model to each of the distributed computing nodes respectively, so that each of the distributed computing nodes deploys the global model.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, it implements the distributed model training method as described in any one of claims 1 to 8.