Balancing method for data consistency and data volume uniformity and storage medium

Through data consistency processing, reweighted data allocation and real-time feedback adaptive adjustment, the problems of data inconsistency and allocation imbalance in federated learning are solved, and the stability, accuracy and efficiency of model training are improved, and global and local consistency are balanced.

CN120429637APending Publication Date: 2025-08-05WUHAN YANGTZE COMM IND GRP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510480217.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

In large-scale federated learning, client data inconsistency, uneven data allocation, dynamic adjustment problems during training, and contradictions between global and local consistency are difficult to solve, resulting in low training efficiency and insufficient model accuracy.

Method used

Through data consistency processing, reweighted data allocation, real-time feedback adaptive adjustment and global and local consistency joint optimization, we ensure data consistency and uniformity, including data standardization, denoising processing, format conversion, weight calculation, dynamic data volume adjustment and adaptive adjustment based on training feedback, and the introduction of joint optimization objective functions.

Benefits of technology

It improves the stability and accuracy of model training, avoids waste of computing resources, improves training efficiency and accuracy, and achieves a balance between global and local consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429637A_ABST
    Figure CN120429637A_ABST
Patent Text Reader

Abstract

The invention provides a data consistency and data volume uniformity balancing method, which comprises the steps of data consistency processing, data volume uniformity adjustment, self-adaptive adjustment based on real-time feedback, self-adaptive adjustment based on training feedback, global and local consistency joint optimization and the like. The problems that in the prior art, data standardization and treatment efficiency is low, and the whole process from data treatment to intelligent analysis cannot be comprehensively covered are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine learning technology, and in particular to a method for balancing data consistency and data volume uniformity. Background Art

[0002] In large-scale federated learning training, how to effectively solve the following technical problems is the core goal of this invention: 1) Client data inconsistency: Due to the different data sources of each client, the data may be noisy, scale-inconsistent, or format-inconsistent, affecting the stability and accuracy of the training model.

[0003] 2) Unbalanced data distribution: In distributed training, if the amount of data is unevenly distributed among the clients, it may cause the computing load of some clients to be too large or too small, thereby affecting the overall training efficiency and even causing computing bottlenecks. 3) Dynamic adjustment problems during training: As training progresses, the amount of data and training progress of each client will change. How to continuously adjust the data distribution strategy in a dynamic environment to maintain the efficiency of training and the high accuracy of the model has become an urgent problem to be solved. 4) The balance problem between global and local consistency: In distributed training, how to deal with the contradiction between global consistency and local consistency and avoid local optimization affecting the global training effect is the key to improving training quality.

[0004] Currently, data consistency and data distribution strategies in federated learning mostly adopt static methods or dynamic adjustment strategies based on simple rules. The main problems of existing technologies include:

[0005] 1) Static data allocation: In traditional methods, most systems rely on predefined data allocation strategies. This strategy cannot adapt to dynamic changes in the training process and can easily lead to some clients having too much or too little data, thus affecting training efficiency.

[0006] 2) Lack of effective data consistency guarantee: Existing methods mostly rely on the consistency assumption of data sources, ignoring the possible differences between different clients in data preprocessing, data format conversion, etc., which leads to inconsistencies in the training process and affects the accuracy of training results.

[0007] 3) Insufficient optimization of global and local consistency: Existing methods generally only focus on the consistency of local data or the consistency of the global model, and lack an effective joint optimization mechanism for global and local consistency. As a result, the data allocation and optimization during the training process cannot be optimized globally.

[0008] 4) Poor dynamic adaptability during training: Traditional methods often find it difficult to dynamically adjust data allocation strategies based on training progress, error feedback, etc., resulting in wasted computing resources or unbalanced training progress. Summary of the Invention

[0009] This invention proposes a method to balance data consistency and data volume uniformity, solving the problems of low efficiency of data standardization and management and the inability to fully cover the entire process from data management to intelligent analysis in existing technologies. The technical solution of this invention is achieved as follows:

[0010] A method for balancing data consistency and data volume uniformity includes the following steps:

[0011] Step S1: Data consistency processing, which performs consistency preprocessing on the data of each client to ensure that the data of all clients meets the unified standard, thereby ensuring the efficiency and consistency of training;

[0012] Step S2: Data volume uniformity adjustment, based on a re-weighted data distribution strategy to dynamically balance the data volume between clients, ensuring that the computing load of each client tends to be balanced;

[0013] Step S3: Adaptive adjustment based on real-time feedback: adjust the data volume and computing load of each node in real time according to the training progress of each client and the changes in the client's data;

[0014] Step S4: Adaptive adjustment based on training feedback, dynamically adjusting data allocation by collecting performance data of each client in real time;

[0015] Step S5: Joint optimization of global and local consistency, introducing a joint optimization objective function to ensure a balance between global consistency and local consistency.

[0016] As a preferred technical solution, the data consistency processing in step S1 specifically includes: data standardization, denoising, and data format conversion. The data standardization adjusts the data to the same scale to prevent certain data from dominating the model training process due to being too large or too small, thereby ensuring the stability of the training. The data standardization formula is as follows:

[0017]

[0018] in, represents the standardized data, x i represents the original data of client i, μ i and σ i are the mean and standard deviation of client i’s data respectively;

[0019] The denoising method includes median filtering or mean filtering;

[0020] The data format conversion converts the data into a standard format to ensure the compatibility and interoperability of the data.

[0021] As a preferred technical solution, the data volume uniformity adjustment in step S2 includes:

[0022] S21: Weight calculation and data volume allocation: Assign a weight to each client. This weight is proportional to the amount of data the client has. The weight calculation formula for each client is as follows:

[0023]

[0024] Among them, D i represents the amount of data held by client i, D total Indicates the total amount of global data, w i represents the weight of client i;

[0025] S22: Dynamic data volume adjustment: During the training process, the data volume distribution is dynamically adjusted according to the weight of each client and real-time feedback. Specifically, the system controls the data input amount of each client to balance the data load of all clients, avoiding situations where some clients are overloaded or underloaded.

[0026] As a preferred technical solution, the adaptive adjustment based on real-time feedback in step S3 is based on the following real-time feedback adjustment mechanism: according to the training progress of each client and the data changes of the client, the data volume and computing load of each node are adjusted in real time. Specifically, the system adjusts the weight and data volume distribution of the client according to the following formula:

[0027] D i (t) = D i (t-1)×Decay Factor(t);

[0028] Among them, Decay Factor (t) is a decay factor that gradually decreases as the training progresses to ensure that the adjustment of the data volume tends to be stable;

[0029] The specific adjustment strategy is: within each training step, the system will fine-tune the data allocation of each client based on real-time feedback; if a client has a large error or slow calculation progress, the system will automatically increase the amount of data for that client to improve its training efficiency; otherwise, the amount of data will be reduced.

[0030] As a preferred technical solution, the adaptive adjustment based on training feedback in step S4 automatically optimizes data allocation by collecting the performance loss of each client, ensuring that each client obtains an appropriate amount of data to achieve the optimization of the global training effect.

[0031] The specific adjustment formula is as follows:

[0032] ΔD i(t) = α × Performance Loss (t)

[0033] Where ΔD i (t) is the data volume adjustment of client i at time step t, α is the adjustment coefficient,

[0034] Performance Loss(t) is the performance loss value of client i at time step t;

[0035] The specific adjustment strategy is: clients with poor performance will receive more data to improve their training efficiency; while clients with better performance will receive less data to prevent excessive data from affecting their calculation progress.

[0036] As a preferred technical solution, step S5 of joint optimization of global and local consistency includes the following steps: dynamically adjusting the client's data allocation through the training effect of the global model to improve the overall accuracy and stability of the model; ensuring the data quality and training stability of each client through local consistency optimization; introducing a joint optimization objective function to ensure a balance between global consistency and local consistency; the optimization function is as follows:

[0037] L total =λ1*L global +λ2*L local

[0038] Among them, L global and L local They represent the loss of global consistency and local consistency respectively, and λ1 and λ2 are the corresponding weights. By adjusting λ1 and λ2 to dynamically optimize data distribution, we can ensure the stability of the global model and the data consistency within each client during the training process.

[0039] A non-temporary storage medium is used to store a program for executing the above-mentioned method for balancing data consistency and data amount uniformity.

[0040] Compared with the existing technology, this solution has the following beneficial effects:

[0041] (1) Data consistency assurance: Through consistency preprocessing and local data optimization, data consistency is ensured during the training process, thereby improving the stability and accuracy of model training.

[0042] (2) Balanced data distribution: Through the reweighted data distribution strategy, the data volume is dynamically adjusted to avoid the waste of computing resources and improve the overall computing efficiency of the system.

[0043] (3) Real-time dynamic adjustment: Real-time adjustment based on training feedback can automatically optimize data allocation during training, thereby improving training efficiency and accuracy.

[0044] (4) Global and local consistency optimization: The introduction of global consistency and local consistency joint optimization can balance the global goals and local training effects during the training process to ensure the best training results. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0046] Figure 1 The figure is a flow chart of a method for balancing data consistency and data volume uniformity according to the present invention. DETAILED DESCRIPTION

[0047] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments of the present invention. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0048] Reference Figure 1 The present invention proposes a method for balancing client data consistency and data volume uniformity in a dynamic process. Its core technical solution is to ensure data consistency and balance in a federated learning environment through data consistency preprocessing, reweighted data distribution, dynamic adjustment mechanism, training feedback adaptive adjustment and joint optimization of global and local consistency.

[0049] The following is a detailed description of the technical solution.

[0050] 1. Data consistency processing

[0051] In distributed training, data inconsistencies can easily occur due to differences in data sources across different clients, impacting the stability and accuracy of model training. Therefore, the first step of this invention is to perform consistency preprocessing on each client's data to ensure that all client data meets a unified standard, thereby ensuring efficient and consistent training.

[0052] 1.1 Data Standardization:

[0053] To eliminate the impact of inconsistent data scales, we normalize the data of each client. By adjusting the data to the same scale, we can prevent some data from dominating the model training process due to being too large or too small, ensuring the stability of training. The data normalization formula is as follows:

[0054]

[0055] in, represents the standardized data, x i represents the original data of client i, μ i and σ i are the mean and standard deviation of the client i data respectively.

[0056] 1.2 Denoising

[0057] Data may contain noise information, and denoising is used to improve the quality and reliability of the data. Common denoising methods include:

[0058] 1.2.1 Median filtering: remove outliers by sorting the data and taking the median.

[0059] 1.2.2 Mean filtering: Smooth the data, eliminate high-frequency noise, and retain the main features of the data.

[0060] 1.3 Data format conversion:

[0061] Because different clients may use different data formats, we convert data into standard formats (such as JSON and CSV) to ensure data compatibility and interoperability. This step ensures that data can be smoothly transferred and shared between clients, avoiding errors or incompatibilities caused by format differences.

[0062] After these preprocessing steps, all client data has been standardized, denoised, and converted into a unified format, providing a stable foundation for subsequent model training.

[0063] 2. Data volume uniformity adjustment

[0064] In distributed training, data imbalances can cause some clients to be overloaded with computational loads while others are underloaded, impacting overall training efficiency and the utilization of computing resources. Therefore, this paper dynamically balances the data load across clients using a reweighted data allocation strategy, ensuring that the computational load on each client is balanced.

[0065] 2.1 Weight calculation and data volume allocation:

[0066] We assign a weight to each client, which is proportional to the amount of data the client has. The weight of each client is calculated as follows:

[0067]

[0068] Among them, D i represents the amount of data held by client i, D total Indicates the total amount of global data, w i The weight of client i is calculated by this formula. The system can adjust the proportion of data received by each client based on the client's data volume and computing power, thus avoiding extremely uneven data distribution.

[0069] 2.2 Dynamic data volume adjustment:

[0070] During training, the system dynamically adjusts data allocation based on each client's weight and real-time feedback. Specifically, the system controls the amount of data input for each client to balance the data load across all clients, preventing situations where some clients are overloaded or underloaded. This adjustment mechanism ensures efficient use of training resources and avoids excessive load on some clients and computational bottlenecks.

[0071] 3. Dynamic adjustment mechanism

[0072] As training progresses, factors such as the client's training progress and data volume will constantly change. To ensure efficient system operation, the present invention has designed a dynamic adjustment mechanism. This mechanism adaptively adjusts data allocation through real-time feedback to ensure optimal results during training.

[0073] 3.1 Real-time feedback adjustment mechanism:

[0074] The system will adjust the data volume and computing load of each node in real time based on the training progress of each client (such as convergence speed, training error, etc.) and the client's data changes. Specifically, the system adjusts the client's weight and data volume distribution according to the following formula:

[0075] D i (t) = D i (t-1)×Decay Factor(t)

[0076] Decay Factor (t) is a decay factor that decreases gradually as training progresses to ensure that the data volume adjustment tends to be stable. This formula makes data volume allocation more flexible and can be adaptively optimized as training progresses.

[0077] 3.2 Adjustment strategy:

[0078] Within each training timestep, the system fine-tunes the data allocation for each client based on real-time feedback (such as convergence speed and training error). If a client's error is large or its computational progress is slow, the system automatically increases the amount of data allocated to that client to improve training efficiency; otherwise, the system reduces the amount of data allocated to that client. This mechanism not only ensures balanced data distribution but also improves training efficiency.

[0079] 4. Adaptive adjustment based on training feedback

[0080] This paper further proposes an adaptive adjustment mechanism based on training feedback, which aims to dynamically adjust data allocation by collecting performance data from each client in real time. By collecting performance losses (such as training error and computation time) from each client, the system can automatically optimize data allocation, ensuring that each client receives the appropriate amount of data to achieve optimal global training results.

[0081] 4.1 Performance loss feedback:

[0082] Each client generates a performance loss value based on its current training performance (such as training error, computation time, etc.). This value reflects the efficiency and progress of the current client training and serves as the basis for data allocation adjustment. The specific adjustment formula is as follows:

[0083] ΔD i (t) = α × Performance Loss (t)

[0084] Where ΔD i (t) is the data volume adjustment amount of client i at time step t, α is the adjustment coefficient, and Performance Loss(t) is the performance loss value of client i at time step t.

[0085] 4.2 Adaptive Data Allocation:

[0086] Based on training feedback, the system dynamically adjusts the amount of training data for each client during training. Clients with poor performance receive more data to improve their training efficiency, while clients with better performance receive less data to prevent excessive data from impacting their computational progress. This adaptive adjustment mechanism significantly improves training efficiency and avoids computational bottlenecks and wasted resources.

[0087] 5. Joint optimization of global and local consistency

[0088] In distributed training, balancing global and local consistency is a crucial issue. Global consistency ensures the effectiveness of overall model training, while local consistency ensures data consistency within each client. To optimize this balance, this paper proposes a joint optimization strategy for global and local consistency.

[0089] 5.1 Global consistency optimization:

[0090] Global consistency focuses on the overall effectiveness of model training. It evaluates consistency by measuring the loss function of the global model. The optimization goal is to dynamically adjust client data allocation based on the training effectiveness of the global model to improve the overall accuracy and stability of the model.

[0091] 5.2 Local consistency optimization:

[0092] Local consistency focuses on data consistency within each client. In a distributed environment, each client may hold different datasets, so local consistency optimization is necessary to ensure data quality and training stability for each client.

[0093] 5.3 Joint Optimization Strategy:

[0094] The present invention ensures a balance between global consistency and local consistency by introducing a joint optimization objective function. The optimization function is as follows:

[0095] L total =λ1*L global +λ2*L local

[0096] Here, and represent the loss of global and local consistency, respectively, and λ1 and λ2 are the corresponding weights. By adjusting λ1 and λ2, the system can dynamically optimize data allocation, ensuring both global model stability and data consistency within each client during training.

[0097] Compared with the prior art, this application has the following beneficial effects:

[0098] (1) Data consistency assurance: Through consistency preprocessing and local data optimization, data consistency is ensured during the training process, thereby improving the stability and accuracy of model training.

[0099] Balanced data distribution: Through the re-weighted data distribution strategy, the data volume is dynamically adjusted to avoid the waste of computing resources and improve the overall computing efficiency of the system.

[0100] (2) Real-time dynamic adjustment: Real-time adjustment is performed in combination with training feedback, which can automatically optimize data allocation during training and improve training efficiency and accuracy.

[0101] (3) Global and local consistency optimization: The introduction of global consistency and local consistency joint optimization can balance the global goals and local training effects during the training process to ensure the best training results.

[0102] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for balancing data consistency and data volume uniformity, characterized in that: The steps include: Step S1: Data consistency processing, which performs consistency preprocessing on the data of each client to ensure that the data of all clients meets the unified standard, thereby ensuring the efficiency and consistency of training; Step S2: Data volume uniformity adjustment, based on a re-weighted data distribution strategy to dynamically balance the data volume between clients, ensuring that the computing load of each client tends to be balanced; Step S3: Adaptive adjustment based on real-time feedback: adjust the data volume and computing load of each node in real time according to the training progress of each client and the changes in the client's data; Step S4: Adaptive adjustment based on training feedback, dynamically adjusting data allocation by collecting performance data of each client in real time; Step S5: Joint optimization of global and local consistency, introducing a joint optimization objective function to ensure a balance between global consistency and local consistency.

2. A method for balancing data consistency and data volume uniformity according to claim 1, characterized in that: The data consistency processing in step S1 specifically includes: data standardization, denoising, and data format conversion. Data standardization adjusts the data to the same scale to prevent certain data from dominating the model training process due to being too large or too small, thereby ensuring the stability of the training. The data standardization formula is as follows: in, represents the standardized data, x i represents the original data of client i, μ i and σ i are the mean and standard deviation of client i’s data respectively; The denoising method includes median filtering or mean filtering; The data format conversion converts the data into a standard format to ensure the compatibility and interoperability of the data.

3. The method for balancing data consistency and data volume uniformity according to claim 1, wherein: The data volume uniformity adjustment in step S2 includes: S21: Weight calculation and data volume allocation: Assign a weight to each client. This weight is proportional to the amount of data the client has. The weight calculation formula for each client is as follows: Among them, D i represents the amount of data held by client i, D total Indicates the total amount of global data, w i represents the weight of client i; S22: Dynamic data volume adjustment: During the training process, the data volume distribution is dynamically adjusted according to the weight of each client and real-time feedback. Specifically, the system controls the data input amount of each client to balance the data load of all clients, avoiding situations where some clients are overloaded or underloaded.

4. The method for balancing data consistency and data volume uniformity according to claim 1, wherein: The adaptive adjustment based on real-time feedback in step S3 is based on the following real-time feedback adjustment mechanism: the data volume and computing load of each node are adjusted in real time according to the training progress of each client and the data changes of the client. Specifically, the system adjusts the weight and data volume distribution of the client according to the following formula: D i (t)=D i (t-1)×Decay Factor(t); Among them, DecayFactor(t) is a decay factor that gradually decreases as the training progresses to ensure that the adjustment of the data volume tends to be stable; The specific adjustment strategy is: within each training step, the system will fine-tune the data allocation of each client based on real-time feedback; if a client has a large error or slow calculation progress, the system will automatically increase the amount of data for that client to improve its training efficiency; otherwise, the amount of data will be reduced.

5. The method for balancing data consistency and data volume uniformity according to claim 1, wherein: The adaptive adjustment based on training feedback in step S4 automatically optimizes data allocation by collecting performance losses of each client, ensuring that each client obtains an appropriate amount of data to achieve the optimization of the global training effect. The specific adjustment formula is as follows: ΔD i (t)=α×Performance Loss(t) Where ΔD i (t) is the data volume adjustment of client i at time step t, α is the adjustment coefficient, Performance Loss(t) is the performance loss value of client i at time step t; The specific adjustment strategy is: clients with poor performance will receive more data to improve their training efficiency; while clients with better performance will receive less data to prevent excessive data from affecting their calculation progress.

6. A method for balancing data consistency and data volume uniformity according to claim 1, characterized in that: The step S5 of global and local consistency joint optimization includes the following steps: dynamically adjusting the client data allocation based on the training effect of the global model to improve the overall accuracy and stability of the model; ensuring the data quality and training stability of each client through local consistency optimization; introducing a joint optimization objective function to ensure a balance between global consistency and local consistency; the optimization function is as follows: L total =λ1*L global +λ2*L local Among them, L global and L local They represent the loss of global consistency and local consistency respectively, and λ1 and λ2 are the corresponding weights. By adjusting λ1 and λ2 to dynamically optimize data distribution, we can ensure the stability of the global model and the data consistency within each client during the training process.

7. A non-temporary storage medium, characterized in that: It is used to store a program for executing a method for balancing data consistency and data amount uniformity as described in any one of claims 1 to 6 above.