Multi-user MEC federated learning heterogeneous data optimization method and system based on dynamic sampling

By collecting and classifying data in a multi-user MEC system in real time, adjusting the number of iterations dynamically, and using dynamic gradient calculation and federated averaging algorithm, the problems of data heterogeneity and device heterogeneity are solved, training efficiency and accuracy are improved, and the adaptability and transparency of the system are enhanced.

CN120387066APending Publication Date: 2025-07-29INST OF WAR STUDIES ACAD OF MILITARY SCI OF THE CHINESE PEOPLES LIBERATION ARMY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510395186.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

There are problems in the multi-user MEC system with low training efficiency and poor model training results due to data heterogeneity and device heterogeneity.

Method used

The terminal device collects the local sample dataset size, calculation rate and network bandwidth parameters in real time, uses a power law distribution classifier for classification, dynamically adjusts the number of iterations, and uses dynamic gradient calculation and federated averaging algorithm for model training to monitor system parameters changes in real time.

Benefits of technology

It improves data processing efficiency, ensures the accuracy and stability of model training, enhances the adaptability and transparency of the system, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387066A_ABST
    Figure CN120387066A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-user MEC federated learning heterogeneous data optimization method and system based on dynamic sampling, and relates to the technical field of edge computing, and the method comprises the following steps: collecting equipment data set size, computing rate, communication time and bandwidth in real time through a terminal sensor, carrying out the preprocessing, and then carrying out the power law distribution classification; large sample equipment adopts dynamic gradient calculation, small sample equipment adopts full data training, the number of iterations is set and dynamically adjusted according to the size of a data set, training data is received through a communication interface, and system parameters are dynamically displayed. According to the method, the efficiency and accuracy of multi-user MEC federation learning in heterogeneous data processing are remarkably improved, resource allocation is optimized through dynamic sampling and power law distribution classification, efficient training of large sample equipment and accurate modeling of small sample equipment are ensured, meanwhile, the number of iterations and real-time system parameter display are dynamically adjusted, and the efficiency and accuracy of data processing are improved. And the flexibility and the monitorability of the system are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of edge computing, and specifically to a multi-user MEC federated learning heterogeneous data optimization method and system based on dynamic sampling. Background Art

[0002] Multi-user MEC, namely multi-access mobile edge computing, is an edge computing technology aiming to transfer computing resources and capabilities from a centralized data center to the edge of the network, making it closer to data sources and end users. In the MEC environment, multi-users usually refer to multiple end users or devices, which can be uniformly connected to the MEC platform through different access networks and receive unified management and control of the network. The core advantage of MEC lies in its ability to provide low-latency, high-bandwidth, and real-time processing applications and services for end users. By completing computing tasks at edge nodes, MEC reduces signal transmission latency, improves data processing speed, and reduces the demand for transmission bandwidth, thereby reducing the risk of network congestion. In addition, MEC also supports flexible application deployment and management, and can dynamically adjust computing resources and services according to actual project needs. Multi-user MEC is widely used in fields such as intelligent transportation, industrial automation, virtual reality, smart cities, and Internet of Things devices, providing strong computing service support for the intelligent transformation of these fields and Internet of Things applications.

[0003] To solve the problems of data heterogeneity and device heterogeneity existing in multi-user MEC, the prior art adopts a centralized data processing method, uploading the data of all terminal devices to the central server for unified processing. However, there will still be situations such as high data transmission latency, large bandwidth occupation, and low training efficiency caused by device performance differences, which will further lead to problems of limited overall system performance and poor model training effects. In view of these challenges, a multi-user MEC federated learning heterogeneous data optimization method and system based on dynamic sampling are proposed. Summary of the Invention

[0004] The purpose of the present invention is to provide a multi-user MEC federated learning heterogeneous data optimization method and system based on dynamic sampling to solve the problems mentioned in the above background art.

[0005] To solve the above technical problems, the technical solutions adopted by the present invention are as follows: In the first aspect, a multi-user MEC federated learning heterogeneous data optimization method based on dynamic sampling includes the following steps:

[0006] S1. Real-time collect the local sample dataset size, computing rate, communication time, and network bandwidth parameters of each device through local sensors of the terminal device;

[0007] S2. Preprocess the collected local sample dataset size, computing rate, communication time, and network bandwidth parameters;

[0008] S3. Classify the size of the collected local sample data set by threshold through a power-law distribution classifier;

[0009] S4. Through a dynamic sampling optimization algorithm, perform dynamic gradient calculation on large-sample devices, and directly use all data for local model training on small-sample devices through the federated averaging algorithm;

[0010] S5. Set the initial number of iterations according to the size of the local sample data set of each terminal device, and dynamically adjust the number of iterations;

[0011] S6. Through communication interface technology, receive the training loss function values, accuracy rates, and training times of each device, and dynamically display the change curves of system parameters.

[0012] A further improvement of the technical solution of the present invention lies in: in S1, the process of collecting the size of the local sample data set, calculation rate, communication time, and network bandwidth parameters of each device through the local sensors of the terminal device includes:

[0013] In a multi-user collaborative mobile edge computing system, deploy terminal devices in the same network environment and connect them to the MEC server. The terminal devices include smart phones, tablets, and sensor nodes. The terminal devices obtain the total amount of data stored locally through the file system interface and record its size to form the local sample data set of the device , estimate the calculation rate by monitoring the CPU utilization rate and the number of currently executed tasks. Each time communicating with the MEC server, record the total time from sending a request to receiving a response, monitor the current upload and download rates through the network interface, and take their average value as the network bandwidth.

[0014] A further improvement of the technical solution of the present invention lies in: in S2, the process of preprocessing the size of the collected local sample data set, calculation rate, communication time, and network bandwidth parameters includes:

[0015] Adopt the 3σ principle to remove the noise data in the size of the local sample data set, calculation rate, communication time, and network bandwidth parameters, use linear interpolation to fill in the missing values in the local sample data set, record the calculation rate, communication time, and network bandwidth of each terminal device, check whether they are within a reasonable range. If there are data points deviating from the expected values, use statistical methods to identify and remove the outliers, and fill in the corresponding missing data with the average values of the calculation rate, communication time, and network bandwidth of the terminal device.

[0016] A further improvement of the technical solution of the present invention lies in: in S3, the process of classifying the size of the collected local sample data set by threshold through a power-law distribution classifier includes:

[0017] According to the size of the standardized local sample data set, use the power-law distribution classifier and the cumulative distribution function to classify the data, and set as the 80th percentile of the sample data set size. If , then the device belongs to a large sample. If , then the device belongs to a small sample.

[0018] A further improvement of the technical solution of the present invention lies in: in S4, through the dynamic sampling optimization algorithm, the process of performing dynamic gradient calculation on large-sample devices includes:

[0019] Calculate the variance of the local sample data set of the terminal device. For large-sample devices, perform dynamic subset screening according to the sample variance, and select the first 50% of the data with the smallest variance as the subset as the training sample. For each iteration, use the selected subset to calculate the gradient on this subset, and adjust the sample size of the next iteration according to the sample variance estimation obtained during the batch gradient calculation process. Each device sends the updated model parameters to the MEC server, and the MEC server aggregates and calculates the average value of the global model parameters.

[0020] A further improvement of the technical solution of the present invention lies in: in S4, through the federated averaging algorithm, the process of directly using the full amount of data for local model training on small-sample devices includes:

[0021] For each iteration, the small-sample device calculates the gradient on the entire data set and updates the local model parameters. After each terminal device completes training locally, it sends the updated model parameters to the MEC server. The MEC server aggregates the updates of the terminal devices, calculates the average value of the global model parameters, and sends the new global model parameters back to each terminal device as the initial parameters for the next round of training.

[0022] A further improvement of the technical solution of the present invention lies in: in S5, set the initial number of iterations according to the size of the local sample data set of each terminal device, and the process of dynamically adjusting the number of iterations includes:

[0023] Monitor the change rate of the loss function in each round of training to adjust the remaining number of iterations of each terminal device, set the global maximum number of iterations, and set the initial number of iterations according to the size of the local sample data set and the average value of the sample data volume of each device , after each round of training, calculate the change rate of the loss function , if , then reduce the remaining number of iterations , if , then increase the remaining number of iterations , update the remaining number of iterations, set the loss function change threshold, and stop training when the change in the loss function is less than the loss function change threshold within 5 consecutive rounds. Among them, represents the average value of the sample dataset size of the acquisition terminal device, and is the adjustment coefficient.

[0024] A further improvement of the technical solution of the present invention lies in: in S6, through the communication interface technology, receiving the training loss function value, accuracy rate, and training time of each device, and the process of dynamically displaying the change curve of the system parameters includes:

[0025] The terminal device sends the training loss function value, accuracy rate, and training time to the MEC server in real time through the communication interface. The MEC server aggregates the data and calculates the global loss function value, global accuracy rate, and system average training time, and dynamically displays the comparison chart of the system average training time, the convergence trend chart of the loss function, and the change curve chart of the accuracy rate under different threshold strategies.

[0026] In the second aspect, a multi-user MEC federated learning heterogeneous data optimization system based on dynamic sampling is used to implement the above-mentioned multi-user MEC federated learning heterogeneous data optimization method based on dynamic sampling, including an MEC server, and the MEC server is communicatively connected to a dynamic sampling data acquisition module, a dynamic gradient training module, and a real-time performance monitoring module, wherein, the electrical signals are connected between the modules;

[0027] The dynamic sampling data acquisition module collects the local data of each device in real time through the local sensors of the terminal device, preprocesses the collected local data, and classifies the size of the collected local sample dataset according to the threshold through the power-law distribution classifier;

[0028] The dynamic gradient training module performs dynamic gradient calculation on large-sample devices through the dynamic sampling optimization algorithm, directly uses the full amount of data for local model training on small-sample devices through the federated averaging algorithm, sets the initial number of iterations according to the size of the local sample dataset of each terminal device, and dynamically adjusts the number of iterations;

[0029] The real-time performance monitoring module receives the training loss function value, accuracy rate, and training time of each device through the communication interface technology, and dynamically displays the change curve of the system parameters.

[0030] Due to the adoption of the above technical solution, the technical progress achieved by the present invention compared with the prior art is:

[0031] 1. The present invention provides a multi-user MEC federated learning heterogeneous data optimization method and system based on dynamic sampling, which can significantly improve the data processing efficiency. By collecting and analyzing key parameters such as the size of the local sample data set and the computing rate of terminal devices in real time, the present invention realizes the accurate evaluation of device performance. Furthermore, through the dynamic sampling optimization algorithm and the federated averaging algorithm, it can flexibly handle data sets of different sizes, avoiding resource waste and low training efficiency, thereby improving the overall data processing efficiency.

[0032] 2. The present invention provides a multi-user MEC federated learning heterogeneous data optimization method and system based on dynamic sampling, which effectively solves the problem of data heterogeneity. By classifying the size of the sample data set through a power-law distribution classifier, the present invention can adopt different processing strategies for data sets of different scales, ensuring the accuracy and stability of model training. This feature enables the present invention to perform well in dealing with complex and changing heterogeneous data, improving the adaptability and robustness of the system.

[0033] 3. The present invention provides a multi-user MEC federated learning heterogeneous data optimization method and system based on dynamic sampling, which realizes the dynamic monitoring and display of system parameters. By receiving key indicators such as the training loss function value, accuracy, and training time of each device in real time through communication interface technology, the present invention can dynamically display the change curve of system parameters, helping users intuitively understand the system operation status, timely adjust the training strategy, and ensure the efficiency and accuracy of model training. This feature enhances the transparency and controllability of the system and improves the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0035] Figure 1 is the flowchart of the present invention;

[0036] Figure 2 is the block diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0038] Example 1, as Figure 1 shown, the present invention provides a multi - user MEC federated learning heterogeneous data optimization method based on dynamic sampling, including the following steps:

[0039] S1. Through the local sensors of the terminal devices, the sizes of the local sample data sets, computing rates, communication times, and network bandwidth parameters of each device are collected in real - time. In a multi - user collaborative mobile edge computing system, the terminal devices are deployed in the same network environment and connected to the MEC server. The terminal devices include smart phones, tablets, and sensor nodes. The terminal devices obtain the total amount of data stored locally through the file system interface and record its size to form the local sample data set of the device , estimate the computing rate by monitoring the CPU utilization rate and the number of currently executing tasks. Each time communicating with the MEC server, record the total time from sending the request to receiving the response. Monitor the current upload and download rates through the network interface and take their average value as the network bandwidth;

[0040] S2. Pre - process the collected sizes of the local sample data sets, computing rates, communication times, and network bandwidth parameters. Use the 3σ principle to remove the noise data in the sizes of the local sample data sets, computing rates, communication times, and network bandwidth parameters. Use the linear interpolation method to fill in the missing values in the local sample data set. Record the computing rate, communication time, and network bandwidth of each terminal device, and check whether they are within a reasonable range. If there are data points deviating from the expected values, use statistical methods to identify and remove the outliers, and fill in the corresponding missing data with the mean values of the computing rate, communication time, and network bandwidth of the terminal device;

[0041] S3. Through the power - law distribution classifier, classify the sizes of the collected local sample data sets according to a threshold. According to the standardized sizes of the local sample data sets, use the power - law distribution classifier and the cumulative distribution function to classify the data. Set as the 80th percentile of the sample data set size. If , then the device belongs to the large - sample type. If , then the device belongs to the small - sample type;

[0042] S4. Through the dynamic sampling optimization algorithm, perform dynamic gradient calculation on large-sample devices. Through the federated averaging algorithm, directly use the full amount of data for local model training on small-sample devices. Calculate the variance of the local sample data set of the terminal device. For large-sample devices, perform dynamic subset screening according to the sample variance, and select the first 50% of the data with the smallest variance as the subset as the training sample. For each iteration, use the selected subset to calculate the gradient on this subset, and adjust the sample size of the next iteration according to the sample variance estimation obtained during the batch gradient calculation process. Each device sends the updated model parameters to the MEC server, and the MEC server aggregates and calculates the average value of the global model parameters. For each iteration, small-sample devices calculate the gradient on the entire data set and update the local model parameters. After each terminal device completes training locally, it sends the updated model parameters to the MEC server. The MEC server aggregates the updates of the collected terminal devices, calculates the average value of the global model parameters, and sends the new global model parameters back to each terminal device as the initial parameters for the next round of training;

[0043] S5. Set the initial number of iterations according to the size of the local sample data set of each terminal device, dynamically adjust the number of iterations, monitor the change rate of the loss function in each round of training to adjust the remaining number of iterations of each terminal device, set the global maximum number of iterations, and set the initial number of iterations according to the size of the local sample data set of each device and the average value of the sample data volume , after each round of training, calculate the change rate of the loss function , if , then reduce the remaining number of iterations , if , then increase the remaining number of iterations , update the remaining number of iterations, set the loss function change threshold, and stop training when the change of the loss function is less than this loss function change threshold for 5 consecutive rounds. Among them, represents the average value of the sample data set sizes of the collected terminal devices, and are adjustment coefficients;

[0044] S6. Through the communication interface technology, receive the training loss function values, accuracies, and training times of each device, and dynamically display the change curves of the system parameters. The terminal device sends the training loss function value, accuracy, and training time to the MEC server in real time through the communication interface. The MEC server aggregates the data and calculates the global loss function value, global accuracy, and system average training time, and dynamically displays the comparison chart of the system average training time, the convergence trend chart of the loss function, and the change curve chart of the accuracy under different threshold strategies.

[0045] Example 2, as Figure 2As shown, based on Embodiment 1, the present invention provides a technical solution: a multi-user MEC federated learning heterogeneous data optimization system based on dynamic sampling, which is used to implement a multi-user MEC federated learning heterogeneous data optimization method based on dynamic sampling. It includes an MEC server, and the MEC server is communicatively connected to a dynamic sampling data acquisition module, a dynamic gradient training module, and a real-time performance monitoring module. Among them, the electrical signals are connected between the modules;

[0046] The dynamic sampling data acquisition module collects the local data of each device in real time through the local sensors of the terminal devices, preprocesses the collected local data, and classifies the size of the collected local sample data set according to a threshold through a power-law distribution classifier;

[0047] The dynamic gradient training module performs dynamic gradient calculation on large-sample devices through a dynamic sampling optimization algorithm, directly uses the full amount of data for local model training on small-sample devices through the federated averaging algorithm, sets the initial number of iterations according to the size of the local sample data set of each terminal device, and dynamically adjusts the number of iterations;

[0048] The real-time performance monitoring module receives the training loss function values, accuracies, and training times of each device through communication interface technology, and dynamically displays the change curves of system parameters.

[0049] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A multi-user MEC federated learning heterogeneous data optimization method based on dynamic sampling, characterized by: It includes the following steps: S1. Through the local sensors of the terminal device, the sizes of the local sample data sets, calculation rates, communication times, and network bandwidth parameters of each device are collected in real time; S2. The collected sizes of the local sample data sets, calculation rates, communication times, and network bandwidth parameters are preprocessed; S3. Through the power-law distribution classifier, the sizes of the collected local sample data sets are classified according to a threshold; S4. Through the dynamic sampling optimization algorithm, dynamic gradient calculation is performed on large-sample devices, and through the federated averaging algorithm, local model training is directly performed on small-sample devices using the full amount of data; S5. The initial number of iterations is set according to the size of the local sample data set of each terminal device, and the number of iterations is dynamically adjusted; S6. Through the communication interface technology, the training loss function values, accuracies, and training times of each device are received, and the change curves of the system parameters are dynamically displayed.

2. The multi-user MEC federated learning heterogeneous data optimization method based on dynamic sampling according to claim 1, wherein: In the above S1, the process of collecting the sizes of the local sample data sets, calculation rates, communication times, and network bandwidth parameters of each device through the local sensors of the terminal device includes: In a multi-user collaborative mobile edge computing system, terminal devices are deployed in the same network environment and connected to the MEC server. The terminal devices include smartphones, tablets, and sensor nodes. The terminal devices obtain the total amount of data stored locally through the file system interface, record its size, and form the local sample data set of the device. , estimate the computing rate by monitoring the CPU utilization rate and the number of tasks currently being executed. Each time communicating with the MEC server, record the total time from sending a request to receiving a response, monitor the current upload and download rates through the network interface, and take their average value as the network bandwidth.

3. The multi - user MEC federated learning heterogeneous data optimization method based on dynamic sampling according to claim 2, characterized in that: In the above S2, the process of preprocessing the collected sizes of the local sample data sets, calculation rates, communication times, and network bandwidth parameters includes: The 3σ principle is used to remove the noise data in the sizes of the local sample data sets, calculation rates, communication times, and network bandwidth parameters. The linear interpolation method is used to fill in the missing values in the local sample data sets. The calculation rates, communication times, and network bandwidths of each terminal device are recorded, and it is checked whether they are within a reasonable range. If there are data points deviating from the expected values, statistical methods are used to identify and remove the outliers, and the means of the calculation rates, communication times, and network bandwidths of the terminal device are used to fill in the corresponding missing data.

4. The method for optimizing heterogeneous data using multi-user MEC federated learning based on dynamic sampling according to claim 3 is characterized by: In the above S3, the process of classifying the sizes of the collected local sample data sets according to a threshold through the power-law distribution classifier includes: According to the size of the standardized local sample data set, use the power-law distribution classifier and the cumulative distribution function to classify the data, and set as the 80th percentile of the sample data set size. If , then the device belongs to a large sample. If , then the device belongs to a small sample.

5. The multi - user MEC federated learning heterogeneous data optimization method based on dynamic sampling according to claim 4, characterized in that: In the above S4, the process of performing dynamic gradient calculation on large-sample devices through the dynamic sampling optimization algorithm includes: Calculate the variance of the local sample data set of the terminal device. For large-sample devices, dynamic subset screening is performed according to the sample variance, and the first 50% of the data with the smallest variance is selected as the subset as the training sample. For each iteration, the selected subset is used to calculate the gradient on this subset, and the sample size for the next iteration is adjusted according to the sample variance estimate obtained during the batch gradient calculation process. Each device sends the updated model parameters to the MEC server, and the MEC server aggregates and calculates the average value of the global model parameters.

6. The multi - user MEC federated learning heterogeneous data optimization method based on dynamic sampling according to claim 5, characterized in that: In the above S4, the process of directly performing local model training on small-sample devices using the full amount of data through the federated averaging algorithm includes: For each iteration, the small-sample device calculates the gradient on the entire data set and updates the local model parameters. After each terminal device completes training locally, it sends the updated model parameters to the MEC server. The MEC server aggregates the updates of the collected terminal devices, calculates the average value of the global model parameters, and sends the new global model parameters back to each terminal device as the initial parameters for the next round of training.

7. The multi-user MEC federated learning heterogeneous data optimization method based on dynamic sampling according to claim 6, wherein: In S5, the process of setting the initial number of iterations according to the size of the local sample data set of each terminal device and dynamically adjusting the number of iterations includes: Monitor the change rate of the loss function in each round of training to adjust the remaining number of iterations for each terminal device, set the global maximum number of iterations, and set the initial number of iterations according to the size of the local sample dataset of each device and the average value of the sample data volume , calculate the change rate of the loss function after each round of training , if , then reduce the remaining number of iterations , if , then increase the remaining number of iterations , update the remaining number of iterations, set the loss function change threshold, and stop training when the change of the loss function is less than this loss function change threshold for 5 consecutive rounds, where represents the average value of the sample dataset sizes of the acquisition terminal devices, and are adjustment coefficients 8. The multi - user MEC federated learning heterogeneous data optimization method based on dynamic sampling according to claim 7, characterized in that: In S6, the process of receiving the training loss function value, accuracy rate, and training time of each device through the communication interface technology and dynamically displaying the system parameter change curve includes: The terminal device sends the training loss function value, accuracy and training time to the MEC server in real time through the communication interface. The MEC server summarizes the data and calculates the global loss function value, global accuracy and system average training time, and dynamically displays the system average training time comparison chart under different threshold strategies, the loss function convergence trend chart and the accuracy change curve chart.

9. A multi-user MEC federated learning heterogeneous data optimization system based on dynamic sampling, used to implement the multi-user MEC federated learning heterogeneous data optimization method based on dynamic sampling according to any one of claims 1 to 8, comprising a MEC server, characterized in that: The MEC server is communicatively connected to a dynamic sampling data acquisition module, a dynamic gradient training module, and a real-time performance monitoring module, wherein electrical signals are connected between the modules.

10. The multi - user MEC federated learning heterogeneous data optimization system based on dynamic sampling according to claim 9, characterized in that: The dynamic sampling data acquisition module collects local data of each device in real time through the local sensor of the terminal device, pre-processes the collected local data, and classifies the collected local sample data set size according to the threshold value through the power law distribution classifier; The dynamic gradient training module performs dynamic gradient calculations on large sample devices using a dynamic sampling optimization algorithm, and directly uses the full data for local model training on small sample devices using a federated averaging algorithm. The initial number of iterations is set based on the size of the local sample dataset of each terminal device, and the number of iterations is dynamically adjusted. The real-time performance monitoring module receives the training loss function value, accuracy and training time of each device through communication interface technology, and dynamically displays the system parameter change curve.