Data processing method based on network awareness, electronic equipment and storage medium

By adopting network-aware data processing methods in distributed machine learning, using multiple rounds of gradient transmission to obtain network state parameters, and adaptively adjust the gradient compression rate and compression strategy, the problem that the gradient compression method in the existing technology cannot be adjusted in real time, and more efficient network transmission and model training performance is achieved.

CN119940450APending Publication Date: 2025-05-06HONG KONG UNIV OF SCI & TECH (GUANGZHOU)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411997147.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing gradient compression method cannot be adjusted in real time according to dynamic changes in the network environment, resulting in network congestion when bandwidth is tight, and network resources cannot be fully utilized when bandwidth is abundant, and communication and model training performance cannot be balanced.

Method used

A network-aware data processing method is proposed, which obtains the minimum transmission round-trip time and bottleneck bandwidth through multiple rounds of gradient transmission, calculates the bandwidth time delay product based on these parameters, senses the network state, and adaptively updates the gradient compression rate based on the bandwidth time delay product, and selects a gradient compression strategy for adaptive adjustment.

Benefits of technology

It realizes adjusting the gradient compression rate according to dynamic changes in the network environment, reducing bandwidth and resource waste, accelerating the development of large-scale distributed machine learning models, and improving network transmission efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940450A_ABST
    Figure CN119940450A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method based on network awareness, electronic equipment and a storage medium. According to the method, multiple rounds of gradient transmission are performed in a starting stage, rapid increase of a gradient compression ratio is completed, minimum transmission round-trip time and a bottleneck bandwidth are calculated, after a network sensing stage is entered, a bandwidth delay product is calculated according to the minimum transmission round-trip time and the bottleneck bandwidth, the gradient compression ratio is updated according to the bandwidth delay product, and before training reaches a preset condition, the gradient compression ratio is updated according to the bandwidth delay product. And repeatedly selecting the gradient compression strategy, and for gradient compression and transmission, calculating the bandwidth delay product of each round of transmission to update the gradient compression rate until the training reaches a preset condition. The gradient compression strategy comprises quantization, pruning and rarefaction operation. According to the method, the parameters of the network state are obtained through multiple rounds of gradient transmission for network state sensing, self-adaptive adjustment of the gradient compression ratio is achieved, the network transmission efficiency is improved, the network bandwidth is fully utilized, the gradient transmission quantity is improved, and therefore the model convergence speed and the model accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning, and in particular to a data processing method, electronic device and storage medium based on network perception. Background Art

[0002] In recent years, with the rapid expansion of the scale of deep learning models, distributed machine learning has become the main means to deal with computing and storage bottlenecks. In distributed data parallel training, each node needs to regularly synchronize the calculated model parameter gradients with other nodes. This large-scale gradient transmission has a very high demand for network bandwidth. In order to reduce the impact of large-scale gradients on communication and model training performance, congestion control algorithms and gradient compression are applied in distributed machine learning, but while alleviating network congestion and improving training efficiency, there are also some defects, such as new challenges in convergence speed, model accuracy and network utilization. In addition, existing gradient compression methods such as quantization, pruning and sparsification techniques can effectively reduce the amount of gradient transmission, thereby alleviating network pressure, but these methods usually use a fixed compression ratio and cannot be adjusted in real time according to the dynamic changes of the network environment. Therefore, the relevant technology may still cause network congestion when the bandwidth is tight, and fail to fully utilize network resources when the bandwidth is sufficient, and cannot achieve a good balance between communication and model training performance. Summary of the invention

[0003] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, the present invention proposes a data processing method based on network perception, which achieves a balance between communication and model training performance, reduces bandwidth and resource waste, and accelerates the development of large-scale distributed machine learning models.

[0004] The present invention also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements a data processing method based on network perception when executing the computer program.

[0005] The present invention also provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are used to execute a data processing method based on network perception.

[0006] In a first aspect, a data processing method based on network perception proposed in an embodiment of the present invention includes:

[0007] Performing model training on the machine learning model of the target node using the image data or the text data to obtain a first model parameter gradient;

[0008] In the startup phase, at least one round of gradient transmission is performed according to the gradient compression rate. In each round of gradient transmission: the first model parameter gradient is compressed according to the gradient compression strategy and the gradient compression rate, the compressed gradient is sent to each node in the node network, the second model parameter gradient returned by each node according to the compressed gradient is received, and the machine learning model is trained according to the second model parameter gradient of each node to update the first model parameter gradient;

[0009] Record the round trip time and amount of gradient transmission for each round of gradient transmission;

[0010] After each round of gradient transmission in the startup phase, when the gradient compression rate is greater than 1, or when the gradient transmission round trip time of the current transmission round is greater than an end threshold, the startup phase is ended and the network perception phase is entered; otherwise, the gradient compression rate is adjusted according to the first gradient growth adjustment parameter and the step of performing at least one round of gradient transmission according to the gradient compression rate is performed;

[0011] In the network perception stage, the gradient compression rate is updated according to the gradient transmission round-trip time, the gradient transmission amount, the gradient reduction parameter and the second gradient growth adjustment parameter, and the step of performing at least one round of gradient transmission according to the gradient compression rate is performed;

[0012] After each round of gradient transmission in the network perception phase, when a preset condition is met, it is determined that the model training is completed.

[0013] According to some embodiments of the present invention, adjusting the gradient compression rate according to the first gradient growth adjustment parameter includes:

[0014] When the sum of the gradient compression rate and the first gradient growth adjustment parameter is less than 1, updating the gradient compression rate to the sum of the gradient compression rate and the first gradient growth adjustment parameter;

[0015] When the sum of the gradient compression rate and the first gradient growth adjustment parameter is greater than or equal to 1, the gradient compression rate is updated to 1.

[0016] According to some embodiments of the present invention, compressing the first model parameter gradient according to the gradient compression strategy and the gradient compression rate includes:

[0017] When the gradient compression rate is less than the gradient quantization threshold, calculating the L2 norm of the first model parameter gradient;

[0018] When the L2 norm is greater than the gradient density threshold, a quantization operation is performed on the value of the first model parameter gradient, and twice the gradient compression rate is used to perform pruning and sparse operations on the first model parameter gradient;

[0019] When the gradient compression rate is greater than or equal to the gradient quantization threshold, or when the gradient compression rate is less than the gradient quantization threshold and the L2 norm is less than or equal to the gradient density threshold, pruning and sparsification operations are performed on the first model parameter gradient.

[0020] According to some embodiments of the present invention, the step of pruning and sparsifying the first model parameter gradient includes:

[0021] Calculate the pruning ratio according to the gradient compression rate;

[0022] Retain the first model parameter gradient of the weight before the pruning ratio;

[0023] The absolute values ​​of the first model parameter gradients obtained after the model training are sorted from large to small, and the first model parameter gradients corresponding to the first model parameters sorted in front are gradient transmitted, and the remaining first model parameter gradients that have not been transmitted are locally accumulated and accumulated as residuals in the next gradient update.

[0024] According to some embodiments of the present invention, the calculation formula for calculating the L2 norm of the first model parameter gradient is as follows:

[0025]

[0026] Among them, L 2_norm is the L2 norm, g is the first model parameter gradient, n is the total number of first model parameter gradients, g i is the gradient of the i-th first model parameter.

[0027] According to some embodiments of the present invention, updating the gradient compression rate according to the gradient transmission round-trip time, the gradient transmission amount, the gradient reduction parameter, and the second gradient growth adjustment parameter includes:

[0028] Update the estimated bandwidth of the corresponding gradient transmission round according to the gradient transmission round-trip time and the gradient transmission amount;

[0029] Update the minimum gradient transmission round-trip time and bottleneck bandwidth according to the gradient transmission round-trip time of each round of gradient transmission and the estimated bandwidth;

[0030] Calculate the bandwidth-delay product of each round of gradient transmission according to the minimum transmission round-trip time and the bottleneck bandwidth;

[0031] The gradient compression rate is updated according to the bandwidth-delay product, the gradient reduction parameter and the second gradient growth adjustment parameter of each round of gradient transmission.

[0032] According to some embodiments of the present invention, updating the estimated bandwidth of the corresponding gradient transmission round according to the gradient transmission round-trip time and the gradient transmission amount includes:

[0033] The estimated bandwidth of each round of transmission is calculated according to the gradient transmission round trip time and the gradient transmission amount of each round of gradient transmission. The calculation formula is as follows:

[0034]

[0035] Among them, EBB is the estimated bandwidth of each round of gradient transmission, data_size is the gradient transmission amount of each round of gradient transmission, and RTT is the round-trip time of gradient transmission.

[0036] According to some embodiments of the present invention, updating the gradient compression rate according to the bandwidth-delay product, the gradient reduction parameter, and the second gradient growth adjustment parameter of each round of gradient transmission includes:

[0037] When the gradient transmission amount is less than or equal to N times the bandwidth-delay product, updating the gradient compression rate of the current transmission round according to the second gradient growth adjustment parameter; wherein N is a number greater than 0 and less than 1;

[0038] When the gradient transmission amount is greater than N times the bandwidth-delay product, updating the gradient compression rate of the current transmission round according to the gradient reduction parameter;

[0039] The calculation formula is as follows:

[0040]

[0041] Among them, ratio is the gradient compression ratio, β2 is the second gradient growth adjustment parameter, α is the gradient reduction parameter, data_size is the amount of gradient transmission in each round of transmission, and BDP is the bandwidth-delay product.

[0042] In a second aspect, an electronic device proposed according to an embodiment of the present invention includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the network-aware data processing method described in the first aspect is implemented.

[0043] In a third aspect, a computer-readable storage medium is proposed according to an embodiment of the present invention, wherein the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the network-aware data processing method as described in the first aspect above.

[0044] The network-aware data processing method according to an embodiment of the present invention has at least the following beneficial effects:

[0045] Through multiple rounds of gradient transmission, the minimum transmission round-trip time and bottleneck bandwidth are obtained. The bandwidth-delay product is calculated based on the minimum transmission round-trip time and bottleneck bandwidth to evaluate the network capacity, that is, the network state is perceived, and the gradient compression rate is adaptively updated according to the bandwidth-delay product, forming a calculation formula for adjusting the gradient compression rate based on network capacity. At the same time, the gradient compression strategy is selected to adaptively adjust the gradient compression rate to minimize the trade-off between transmission efficiency, convergence speed and model accuracy in distributed machine learning. Through multiple rounds of gradient transmission, the parameters of the network state are obtained for network perception, and the gradient compression rate is adaptively adjusted according to the network state to improve network transmission efficiency.

[0046] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The present invention will be further described below with reference to the accompanying drawings and embodiments, wherein:

[0048] Figure 1 A flow chart of a data processing method based on network perception provided by an embodiment of the present invention;

[0049] Figure 2 A detailed flow chart of step S500 in the data processing method based on network perception provided by an embodiment of the present invention;

[0050] Figure 3 A flowchart of a gradient compression strategy provided for an embodiment of the present invention;

[0051] Figure 4 A flowchart of the startup phase and the network perception phase in the model training provided by an embodiment of the present invention;

[0052] Figure 5 A module diagram of a data processing method based on network perception provided by an embodiment of the present invention;

[0053] Figure 6 A schematic diagram of gradient transmission under distributed machine learning provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0054] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.

[0055] In the description of the present invention, it should be understood that descriptions involving orientations, such as up, down, front, back, left, right, etc., and orientations or positional relationships indicated are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present invention.

[0056] In the description of the present invention, "several" means more than one, "many" means more than two, "greater than", "less than", "exceed", etc. are understood to exclude the number itself, and "above", "below", "within", etc. are understood to include the number itself. If there is a description of "first" or "second", it is only used for the purpose of distinguishing the technical features, and cannot be understood as indicating or implying the relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features.

[0057] In the description of the present invention, unless otherwise clearly defined, terms such as setting, installing, connecting, etc. should be understood in a broad sense, and technicians in the relevant technical field can reasonably determine the specific meanings of the above terms in the present invention based on the specific content of the technical solution.

[0058] In the description of the present invention, the description with reference to the terms "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0059] The present invention provides a data processing method based on network perception, which obtains the minimum transmission round-trip time and bottleneck bandwidth through multiple rounds of gradient transmission, and calculates the bandwidth-delay product according to the minimum transmission round-trip time and bottleneck bandwidth to evaluate the network capacity, that is, perceives the network state, and adaptively updates the gradient compression rate according to the bandwidth-delay product, forms a calculation formula for adjusting the gradient compression rate based on the network capacity, and selects the gradient compression strategy to adaptively adjust the gradient compression rate, thereby minimizing the trade-off between transmission efficiency, convergence speed and model accuracy in distributed machine learning. The parameters of the network state are obtained through multiple rounds of gradient transmission for network perception, and the gradient compression rate is adaptively adjusted according to the network state, thereby improving the network transmission efficiency.

[0060] The method of the embodiment of the present invention is further described below based on the accompanying drawings.

[0061] like Figure 1 As shown, the network-aware data processing method of the present invention includes:

[0062] S100: Perform model training on a machine learning model of a target node using image data or text data to obtain a first model parameter gradient.

[0063] A machine learning model is a computational model that can learn from data and generalize to solve specific problems or tasks, such as image classification models and large language models. In the field of machine learning, gradients are mainly used in optimization algorithms to find the minimum value of the loss function. The loss function is a function that measures the difference between the model's predicted value and the true value. The label is the supervisory information in the model training process, which provides the correct output corresponding to each input data. In the training of the image classification model, the label includes the category label of the image classification. In the training of the large language model, the label includes the sentiment label in the text classification. In the model training process, the image data or text data is input into the machine learning model, the loss value is calculated according to the model output and the label, and the first model parameter gradient is calculated by the back propagation algorithm. In addition, after the first model parameter gradient is obtained, the first model parameter can be updated according to the first model parameter gradient, thereby updating the machine learning model. The first model parameter refers to the parameter of the machine learning model.

[0064] S200: In the startup phase, at least one round of gradient transmission is performed according to the gradient compression rate. In each round of gradient transmission: the first model parameter gradient is compressed according to the gradient compression strategy and the gradient compression rate, the compressed gradient is sent to each node in the node network, the second model parameter gradient returned by each node according to the compressed gradient is received, and the machine learning model is trained according to the second model parameter gradient of each node to update the first model parameter gradient.

[0065] It should be noted that, during the startup phase, a value ranging from 0 to 0.001 is set as the initial value of the gradient compression rate. A smaller initial value of the gradient compression rate helps to accurately detect network capacity and avoid network congestion. Rapidly increasing the compression rate during the startup phase helps to quickly detect network capacity and ensure convergence speed.

[0066] The network-aware data processing method provided by the present invention is executed by a target node among multiple nodes in a node network. A node network refers to a network structure composed of multiple nodes communicating with each other. A node can be a computer, a server, etc. In one example, referring to Figure 6, the node network includes 3 nodes, denoted as J1, J2 and J3. If J1 is the target node, J1 compresses the first model parameter gradient according to the gradient compression strategy and the gradient compression rate, sends the compressed gradient to J2 and J3, receives the second model parameter gradient returned by J2 and J3 according to the compressed gradient, and performs model training on the machine learning model according to the second model parameter gradient of J2 and J3 to update the first model parameter gradient, and then performs the next round of gradient transmission according to the first model parameter gradient. Among them, J2 can be used as a target node, so in addition to receiving the compressed gradient sent by J1, J2 can also receive the compressed gradient sent by J3. Similarly, J3 can be used as a target node, and in addition to receiving the compressed gradient sent by J1, J3 can also receive the compressed gradient sent by J2. After receiving the compressed gradients sent by other nodes, J2 or J3 can update the gradients of its own machine learning model together with its own accumulated gradients, and then use its own image data or text data to train the machine learning model to obtain new model parameter gradients. After compressing them, the second model parameter gradients can be obtained, and then the second model parameter gradients can be returned to J1 for global aggregation update.

[0067] S300: Recording the gradient transmission round trip time and gradient transmission amount of each round of gradient transmission.

[0068] Among them, the round-trip time of gradient transmission refers to the time when the target node starts to send the first model parameter gradient to other nodes in each round of gradient transmission, and the time difference between the time when the target node receives the second model parameter gradient from all other nodes. The round-trip time of gradient transmission does not include the time for waiting for model training, gradient update, and compression of other nodes. In one example, nodes J1, J2, and J3 cooperate to perform distributed machine learning tasks, and perform model update and training based on their own gradients and the aggregated results of gradients from other nodes in the same round, and generate gradients for the next round; for the target node J1, it starts to send the first model gradient to nodes J2 and J3 after generating the gradient, and starts to receive the second model gradient from J2 and J3 at the same time. At this time, the round-trip time of gradient transmission in this round of gradient transmission is the time from when J1 starts to send the gradient to when it receives the gradients sent by other nodes.

[0069] The transmission amount refers to the size of the gradient data after the target node compresses the first model parameter gradient in each round of gradient transmission, that is, the size of the data actually sent by the target node.

[0070] It should be noted that the gradient transmission round trip time and the gradient transmission amount are used to update the estimated bandwidth of the corresponding gradient transmission round.

[0071] S400: After each round of gradient transmission in the startup phase, when the gradient compression rate is greater than 1, or when the round-trip time of the gradient transmission of the current transmission round is greater than the end threshold, the startup phase is ended and the network perception phase is entered; otherwise, the gradient compression rate is adjusted according to the first gradient growth adjustment parameter and the step of performing at least one round of gradient transmission according to the gradient compression rate is executed.

[0072] In one embodiment, the step of adjusting the gradient compression rate according to the first gradient growth adjustment parameter includes: when the sum of the gradient compression rate and the first gradient growth adjustment parameter is less than 1, updating the gradient compression rate to the sum of the gradient compression rate and the first gradient growth adjustment parameter; when the sum of the gradient compression rate and the first gradient growth adjustment parameter is greater than or equal to 1, updating the gradient compression rate to 1, and the calculation formula is as follows:

[0073] ratio=min(1,ratio+β1)

[0074] ratio refers to the gradient compression ratio, and β1 refers to the first gradient growth adjustment parameter.

[0075] Reference Figure 4 In the startup phase, before entering the network perception phase, the gradient compression rate is only updated according to the first gradient growth adjustment parameter. After entering the network perception phase, it will continue to be in the network perception phase, and the gradient compression rate will no longer be updated according to the first gradient growth adjustment parameter but according to the parameters of the network perception phase.

[0076] S500: In the network perception stage, the gradient compression rate is updated according to the gradient transmission round-trip time, the gradient transmission amount, the gradient reduction parameter and the second gradient growth adjustment parameter, and the step of performing at least one round of gradient transmission according to the gradient compression rate in step S200 is executed.

[0077] It should be noted that the minimum gradient transmission round-trip time and bottleneck bandwidth are calculated based on the gradient transmission round-trip time and gradient transmission amount of each round, and the network capacity can be estimated based on the minimum gradient transmission round-trip time and bottleneck bandwidth, and the gradient compression rate is adjusted based on the estimated network capacity, the gradient reduction parameter and the second gradient growth adjustment parameter.

[0078] S600: After each round of gradient transmission in the network perception phase, when a preset condition is met, it is determined that the model training is finished.

[0079] Reference Figure 4The preset conditions include the set rounds of gradient transmission or the set accuracy of model training. When the preset conditions are not met, the gradient compression and transmission will continue, that is, the corresponding gradient compression strategy will continue to be selected for gradient compression and transmission, and the bandwidth-delay product of each round of gradient transmission will be obtained. The gradient compression rate is updated according to the bandwidth-delay product of each round of gradient transmission until the preset conditions are met and the training is terminated.

[0080] In addition, in one embodiment, referring to Figure 2 ,exist Figure 1 Step S500 of the illustrated embodiment also includes but is not limited to the following steps:

[0081] S501: updating the estimated bandwidth of the corresponding gradient transmission round according to the gradient transmission round-trip time and the gradient transmission amount;

[0082] The estimated bandwidth calculation formula is as follows:

[0083]

[0084] Among them, EBB is the estimated bandwidth of each round of gradient transmission, data_size is the gradient transmission amount of each round of gradient transmission, and RTT is the round-trip time of gradient transmission. By updating the estimated bandwidth of each round of gradient transmission, calculating the bottleneck bandwidth of the subsequent steps, dynamically adjusting the bandwidth-delay product of the subsequent steps, and further dynamically adjusting the gradient compression rate, adaptive gradient compression is achieved.

[0085] S502: updating the minimum gradient transmission round-trip time and bottleneck bandwidth according to the gradient transmission round-trip time and estimated bandwidth of each round of gradient transmission;

[0086] It should be noted that when the round-trip time of the gradient transmission of the current transmission round is less than the minimum round-trip time of the gradient transmission, the minimum round-trip time of the gradient transmission is updated to the round-trip time of the gradient transmission of the current transmission round. When the estimated bandwidth of the current transmission round is greater than the bottleneck bandwidth, the bottleneck bandwidth is updated to the estimated bandwidth of the current transmission round. The calculation formula is as follows:

[0087] RTprop=min(RTT,RTprop)

[0088] BtlBw=max(EBB,BtlBw)

[0089] Among them, RTT is the round-trip time of gradient transmission, RTprop is the minimum round-trip time of gradient transmission; EBB is the estimated bandwidth of each round of gradient transmission, and BtlBw is the bottleneck bandwidth.

[0090] S503: Calculate the bandwidth-delay product of each round of gradient transmission according to the minimum transmission round-trip time and the bottleneck bandwidth. The calculation formula is as follows:

[0091] BDP=BtlBw×RTprop

[0092] Among them, BDP is the bandwidth-delay product, RTprop is the minimum round-trip time, and BtlBw is the bottleneck bandwidth.

[0093] It should be noted that in this embodiment, the bandwidth-delay product represents the estimated value of the network capacity. The minimum transmission round-trip time is selected to be multiplied by the bottleneck bandwidth in order to more accurately estimate the network capacity, thereby optimizing the network transmission strategy and improving the transmission efficiency. In contrast, the product of the transmission round-trip time and the estimated bandwidth of each round only reflects the gradient transmission volume within a transmission cycle, and cannot accurately reflect the actual capacity of the network. At the same time, the gradient compression rate is dynamically adjusted based on the bandwidth-delay product, so that the gradient transmission volume of the transmission can adapt to the network capacity and maximize the bandwidth utilization.

[0094] S504: updating the gradient compression rate according to the bandwidth-delay product, the gradient reduction parameter and the second gradient growth adjustment parameter of each round of gradient transmission.

[0095] It should be noted that the steps of updating the gradient compression rate corresponding to the current transmission round include:

[0096] When the gradient transmission amount is less than or equal to N times the bandwidth-delay product, the gradient compression rate of the current transmission round is updated according to the second gradient growth adjustment parameter;

[0097] When the amount of gradient transmission is greater than N times the bandwidth-delay product, the gradient compression rate of the current transmission round is updated according to the gradient reduction parameter; where N is a number greater than 0 and less than 1;

[0098] The calculation formula is as follows:

[0099]

[0100] Wherein, BDP is the bandwidth delay product, date_size is the gradient transmission amount, ratio is the gradient compression ratio, α is the gradient reduction parameter, and β2 is the second gradient growth adjustment parameter. In this embodiment, N is 0.9.

[0101] It should be noted that according to the data processing method based on network perception in an embodiment of the present invention. The method performs multiple rounds of gradient transmission in the startup phase, and simultaneously calculates the minimum transmission round-trip time and bottleneck bandwidth, thereby completing the rapid growth of the gradient compression rate. After entering the network perception phase, the bandwidth-delay product is calculated based on the minimum transmission round-trip time and bottleneck bandwidth, and the gradient compression rate is updated based on the bandwidth-delay product. Before the training reaches the preset conditions, the selection of the gradient compression strategy, the gradient compression and transmission are repeated, and the bandwidth-delay product of each round of transmission is calculated to update the gradient compression rate until the training reaches the preset conditions. The gradient compression strategy includes quantization, pruning, and sparse operations. The method disclosed in the present invention obtains the parameters of the network state through multiple rounds of gradient transmission for network perception, realizes adaptive adjustment of the gradient compression rate according to the network state, and improves network transmission efficiency.

[0102] The gradient compression strategy mentioned in the above method is described in detail below.

[0103] like Figure 5 As shown in FIG. 1 , the gradient compression strategy includes quantization, pruning, and thinning operations. When compressing the gradient in the startup phase and the network state perception phase, the gradient compression strategy will be used to compress the gradient that needs to be transmitted.

[0104] Furthermore, in some embodiments of the present invention, Figure 3 As shown in Figure 1, the gradient compression strategy includes quantization, pruning, and thinning operations, including:

[0105] When the gradient compression rate is less than the gradient quantization threshold, calculating the L2 norm of the first model parameter gradient;

[0106] When the L2 norm is greater than the gradient density threshold, a quantization operation is performed on the value of the first model parameter gradient, and twice the gradient compression rate is used to perform pruning and sparse operations on the first model parameter gradient;

[0107] When the gradient compression rate is greater than or equal to the gradient quantization threshold, or when the gradient compression rate is less than the gradient quantization threshold and the L2 norm is less than or equal to the gradient density threshold, pruning and sparse operations are performed on the first model parameter gradient.

[0108] It should be noted that in the gradient compression strategy, when the gradient compression rate is less than the gradient quantization threshold and the L2 norm of the first model parameter gradient is greater than the gradient density threshold, the gradient compression rate is temporarily updated to twice the original value for calculating the pruning ratio. After all compression calculations are completed, the overall gradient compression rate is the original value.

[0109] Furthermore, in some embodiments of the present invention, the pruning and thinning operations include:

[0110] Calculate the pruning ratio based on the gradient compression rate;

[0111] Keep the first model parameter gradient before the weight pruning ratio;

[0112] The absolute values ​​of the first model parameter gradients obtained after model training are sorted from large to small, and the first model parameter gradients corresponding to the first model parameters sorted in front are gradient transmitted, and the remaining first model parameter gradients that have not been transmitted are locally accumulated and accumulated as residuals in the next gradient update.

[0113] It should be noted that the quantization operation on the value of gradient data includes reducing the gradient from 32-bit floating point to 16 bits, so that the amount of data occupied by a single gradient is lower, thereby increasing the amount of gradient transmission after pruning and sparsification, thereby improving the accuracy of the model under the condition of a certain amount of data transmission.

[0114] It should be noted that the first model parameter gradient before the pruning ratio is retained, and unimportant parameters are set to 0 to reduce data redundancy and the computational complexity of the model.

[0115] It should be noted that the remaining untransmitted first model parameter gradients are saved locally and accumulated as residuals in the next round of gradient updates to improve training accuracy. Based on the residual accumulation mechanism, it ensures that the discarded gradient information will not be lost for a long time, thereby improving the convergence of the model in the case of sparseness.

[0116] It should be noted that the calculation formulas for calculating the L2 norm of the first model parameter gradient and the pruning ratio are as follows:

[0117]

[0118] prunning_rate=0.5×(1-ratio)

[0119] Among them, L 2_norm is the L2 norm, g is the first model parameter gradient, n is the total number of first model parameter gradients, g i is the gradient of the first model parameter of the i-th model, and prunning_rate is the pruning ratio. Using the L2 norm can intuitively reflect the overall size of the gradient and keep the direction of the original gradient unchanged.

[0120] It should be noted that existing gradient compression methods such as quantization, pruning and sparsification techniques can effectively reduce the amount of gradient transmission, thereby alleviating network pressure, but these methods usually use a fixed compression ratio and cannot be adjusted in real time according to dynamic changes in the network environment. In this embodiment, by judging whether the gradient compression rate is less than the gradient quantization threshold and judging whether the L2 norm of the gradient is greater than the gradient density threshold, the first model parameter gradient is quantized, pruned and sparsified, so as to achieve real-time adjustment of the first model parameter gradient according to the dynamic changes of the network environment.

[0121] In a second aspect, an embodiment of the present application further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and wherein the processor implements the network-aware data processing method of the first aspect when executing the computer program.

[0122] In a third aspect, an embodiment of the present application further provides a storage medium, which is a computer-readable storage medium and stores a computer program. When the computer program is executed by a processor, the network-aware data processing method based on the first aspect described above is implemented.

[0123] As a non-transient computer-readable storage medium, the memory can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof. The device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and are implemented to be located in one place, or may also be distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment.

[0124] It will be appreciated by those skilled in the art that all or some of the steps and systems in the disclosed method above may be implemented as software, firmware, hardware and appropriate combinations thereof. Some physical components or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor or a microprocessor, or may be implemented as hardware, or may be implemented as an integrated circuit, such as an application specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or a non-transitory medium) and a communication medium (or a temporary medium). As known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, disk storage or other magnetic storage devices, or any other medium that may be used to store desired information and may be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically include computer readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0125] The embodiments of the present invention are described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the above embodiments. Various changes can be made within the knowledge of ordinary technicians in the relevant technical field without departing from the purpose of the present invention. In addition, the embodiments of the present invention and the features in the embodiments can be combined with each other without conflict.

Claims

1. A data processing method based on network perception, characterized in that: Applied to a target node among a plurality of nodes in a node network, the method comprises: Performing model training on the machine learning model of the target node using the image data or the text data to obtain a first model parameter gradient; In the startup phase, at least one round of gradient transmission is performed according to the gradient compression rate. In each round of gradient transmission: the first model parameter gradient is compressed according to the gradient compression strategy and the gradient compression rate, the compressed gradient is sent to each node in the node network, the second model parameter gradient returned by each node according to the compressed gradient is received, and the machine learning model is trained according to the second model parameter gradient of each node to update the first model parameter gradient; Record the round trip time and amount of gradient transmission for each round of gradient transmission; After each round of gradient transmission in the startup phase, when the gradient compression rate is greater than 1, or when the gradient transmission round trip time of the current transmission round is greater than an end threshold, the startup phase is ended and the network perception phase is entered; otherwise, the gradient compression rate is adjusted according to the first gradient growth adjustment parameter and the step of performing at least one round of gradient transmission according to the gradient compression rate is performed; In the network perception stage, the gradient compression rate is updated according to the gradient transmission round-trip time, the gradient transmission amount, the gradient reduction parameter and the second gradient growth adjustment parameter, and the step of performing at least one round of gradient transmission according to the gradient compression rate is performed; After each round of gradient transmission in the network perception phase, when a preset condition is met, it is determined that the model training is completed.

2. The data processing method based on network perception according to claim 1 is characterized in that: The step of adjusting the gradient compression rate according to the first gradient growth adjustment parameter includes: When the sum of the gradient compression rate and the first gradient growth adjustment parameter is less than 1, updating the gradient compression rate to the sum of the gradient compression rate and the first gradient growth adjustment parameter; When the sum of the gradient compression rate and the first gradient growth adjustment parameter is greater than or equal to 1, the gradient compression rate is updated to 1.

3. The data processing method based on network perception according to claim 1 is characterized in that: The compressing the first model parameter gradient according to the gradient compression strategy and the gradient compression rate includes: When the gradient compression rate is less than the gradient quantization threshold, calculating the L2 norm of the first model parameter gradient; When the L2 norm is greater than the gradient density threshold, a quantization operation is performed on the value of the first model parameter gradient, and twice the gradient compression rate is used to perform pruning and sparse operations on the first model parameter gradient; When the gradient compression rate is greater than or equal to the gradient quantization threshold, or when the gradient compression rate is less than the gradient quantization threshold and the L2 norm is less than or equal to the gradient density threshold, pruning and sparsification operations are performed on the first model parameter gradient.

4. The data processing method based on network perception according to claim 3 is characterized in that: The step of pruning and sparsifying the gradient of the first model parameters includes: Calculate the pruning ratio according to the gradient compression rate; Retain the first model parameter gradient of the weight before the pruning ratio; The absolute values ​​of the first model parameter gradients obtained after the model training are sorted from large to small, and the first model parameter gradients corresponding to the first model parameters sorted in front are gradient transmitted, and the remaining first model parameter gradients that have not been transmitted are locally accumulated and accumulated as residuals in the next gradient update.

5. The data processing method based on network perception according to claim 3 is characterized in that: The calculation formula for calculating the L2 norm of the first model parameter gradient is as follows: Among them, L 2_norm is the L2 norm, g is the first model parameter gradient, n is the total number of first model parameter gradients, G i is the gradient of the i-th first model parameter.

6. The data processing method based on network perception according to claim 1 is characterized in that: The updating of the gradient compression rate according to the gradient transmission round trip time, the gradient transmission amount, the gradient reduction parameter and the second gradient growth adjustment parameter comprises: Update the estimated bandwidth of the corresponding gradient transmission round according to the gradient transmission round-trip time and the gradient transmission amount; Update the minimum gradient transmission round-trip time and bottleneck bandwidth according to the gradient transmission round-trip time of each round of gradient transmission and the estimated bandwidth; Calculate the bandwidth-delay product of each round of gradient transmission according to the minimum transmission round-trip time and the bottleneck bandwidth; The gradient compression rate is updated according to the bandwidth-delay product, the gradient reduction parameter and the second gradient growth adjustment parameter of each round of gradient transmission.

7. The data processing method based on network perception according to claim 6 is characterized in that: The updating of the estimated bandwidth of the corresponding gradient transmission round according to the gradient transmission round-trip time and the gradient transmission amount includes: The estimated bandwidth of each round of transmission is calculated according to the gradient transmission round trip time and the gradient transmission amount of each round of gradient transmission. The calculation formula is as follows: Among them, EBB is the estimated bandwidth of each round of gradient transmission, data_size is the gradient transmission amount of each round of gradient transmission, and RTT is the round-trip time of gradient transmission.

8. The data processing method based on network perception according to claim 6 is characterized in that: The updating of the gradient compression rate according to the bandwidth-delay product, the gradient reduction parameter and the second gradient growth adjustment parameter of each round of gradient transmission includes: When the gradient transmission amount is less than or equal to N times the bandwidth-delay product, updating the gradient compression rate of the current transmission round according to the second gradient growth adjustment parameter; wherein N is a number greater than 0 and less than 1; When the gradient transmission amount is greater than N times the bandwidth-delay product, updating the gradient compression rate of the current transmission round according to the gradient reduction parameter; The calculation formula is as follows: Among them, ratio is the gradient compression ratio, β2 is the second gradient growth adjustment parameter, α is the gradient reduction parameter, data_size is the amount of gradient transmission in each round of transmission, and BDP is the bandwidth-delay product.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the method according to any one of claims 1 to 8.