Model Training Method, Device, Equipment and Storage Medium

The self-adaptive gradient clipping technique stabilizes large model training by addressing gradient anomalies, improving stability and efficiency through a method that utilizes both current and historical gradient values for clipping.

CN117610677BActive Publication Date: 2025-07-15BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311426412.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-31
Publication Date
2025-07-15
Estimated Expiration
2043-10-31

AI Technical Summary

Technical Problem

There are problems of instability in training convergence during large model training, such as training divergence and loss spikes, resulting in low training stability.

Method used

Adaptive gradient cutting method is adopted to obtain the values of this round of gradient and historical gradient, and combine the cutting methods of different particle sizes to cut abnormal gradients in real time to suppress training divergence and loss spikes.

Benefits of technology

It effectively improves the stability of model training, reduces training time, and improves the efficiency and accuracy of model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117610677B_ABST
    Figure CN117610677B_ABST
Patent Text Reader

Abstract

The present disclosure provides a model training method, apparatus, device, and storage medium. The present disclosure relates to the field of artificial intelligence technologies, specifically to technical fields such as natural language processing, deep learning, computer vision, and image processing. The specific solution is as follows: Deploy the model to be trained on a computing device; train the model to be trained deployed on the computing device to obtain a target model; wherein, training the model to be trained deployed on the computing device includes: obtaining a value representing the gradient of the current round and a value representing the historical gradient; based on the value representing the historical gradient and the value representing the gradient of the current round, obtaining a reference value corresponding to the target clipping method; when the reference value reaches the clipping criterion corresponding to the target clipping method, clip the gradient of the current round using the target clipping method. According to the solution of the present disclosure, it is possible to promptly discover abnormal gradients during model training, effectively suppress the occurrence of training divergence or loss spikes, and thereby effectively improve the stability of model training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to technical fields such as natural language processing, deep learning, computer vision, and image processing. Background Art

[0002] In the field of artificial intelligence, large model training is a current hot topic. However, during the training process of large models, problems such as unstable training convergence often occur. For example, problems such as training divergence or loss spikes occur during the training of large models, resulting in low stability of large model training. Therefore, how to effectively improve the stability of large model training is an urgent problem to be solved. Summary of the Invention

[0003] The present disclosure provides a model training method, apparatus, device, and storage medium.

[0004] According to a first aspect of the present disclosure, there is provided a model training method, including:

[0005] Deploy a model to be trained on a computing device;

[0006] Train the model to be trained deployed on the computing device to obtain a target model;

[0007] Wherein, training the model to be trained deployed on the computing device includes:

[0008] Obtain a value representing the current round of gradient and a value representing the historical gradient;

[0009] Based on the value representing the historical gradient and the value representing the current round of gradient, obtain a reference value corresponding to the target clipping method;

[0010] When the reference value reaches the clipping standard corresponding to the target clipping method, clip the current round of gradient using the target clipping method.

[0011] According to a second aspect of the present disclosure, there is provided an image processing method, including:

[0012] Input an image to be processed into the trained target model, where the trained target model is obtained by training according to the training method of the first aspect;

[0013] According to the trained target model, perform at least one image processing operation including image classification, image recognition, and image segmentation on the image to be processed to obtain an image processing result.

[0014] According to a third aspect of the present disclosure, there is provided a natural language processing method, including:

[0015] Input the first type of data to be processed into the trained target model, where the trained target model is obtained by training according to the training method of the first aspect;

[0016] According to the trained target model, perform at least one natural language processing including information extraction, text classification, text recognition, speech recognition, question answering on the first type of data to be processed, and obtain the natural language processing result.

[0017] According to the fourth aspect of the present disclosure, a computer vision processing method is provided, including:

[0018] Input the second type of data to be processed into the trained target model, where the trained target model is obtained by training according to the training method of the first aspect;

[0019] According to the trained target model, perform at least one computer vision processing including image recognition, object detection, semantic segmentation, video understanding, and image generation on the second type of data to be processed, and obtain the computer vision processing result.

[0020] According to the fifth aspect of the present disclosure, a model training device is provided, including:

[0021] A deployment module for deploying the model to be trained on a computing device;

[0022] A training module for training the model to be trained deployed on the computing device to obtain a target model;

[0023] Among them, the training module includes:

[0024] A first acquisition sub-module for acquiring the value representing the current round of gradient and the value representing the historical gradient;

[0025] A second acquisition sub-module for obtaining a reference value corresponding to the target clipping method based on the value representing the historical gradient and the value representing the current round of gradient;

[0026] A clipping sub-module for clipping the current round of gradient by using the target clipping method when the reference value reaches the clipping standard corresponding to the target clipping method.

[0027] According to the sixth aspect of the present disclosure, an image processing device is provided, including:

[0028] A first input module for inputting the image to be processed into the trained target model, where the trained target model is obtained by training according to the training method of the first aspect;

[0029] An image processing module for performing at least one image processing including image classification, image recognition, and image segmentation on the image to be processed according to the trained target model, and obtaining the image processing result.

[0030] According to a seventh aspect of the present disclosure, there is provided a natural language processing apparatus, including:

[0031] A second input module, configured to input first type of data to be processed into a trained target model, where the trained target model is obtained by training according to the training method of the first aspect;

[0032] A natural language processing module, configured to perform at least one natural language processing including information extraction, text classification, text recognition, speech recognition, question answering on the first type of data to be processed according to the trained target model, so as to obtain a natural language processing result.

[0033] According to an eighth aspect of the present disclosure, there is provided a computer vision processing apparatus, including:

[0034] A third input module, configured to input second type of data to be processed into a trained target model, where the trained target model is obtained by training according to the training method of the first aspect;

[0035] A computer vision processing module, configured to perform at least one computer vision processing including image recognition, object detection, semantic segmentation, video understanding, and image generation on the second type of data to be processed according to the trained target model, so as to obtain a computer vision processing result.

[0036] According to a ninth aspect of the present disclosure, there is provided an electronic device, including: at least one processor; a memory communicatively connected to the at least one processor; the memory stores instructions that can be executed by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the model training method provided in the first aspect and / or the image processing method provided in the second aspect and / or the natural language processing method provided in the third aspect and / or the computer vision processing method provided in the fourth aspect.

[0037] According to a tenth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause a computer to execute the model training method provided in the first aspect and / or the image processing method provided in the second aspect and / or the natural language processing method provided in the third aspect and / or the computer vision processing method provided in the fourth aspect.

[0038] According to an eleventh aspect of the present disclosure, there is provided a computer program product, including a computer program stored on a storage medium, where the computer program, when executed by a processor, implements the model training method provided in the first aspect and / or the image processing method provided in the second aspect and / or the natural language processing method provided in the third aspect and / or the computer vision processing method provided in the fourth aspect.

[0039] According to the solution of the present disclosure, it is possible to timely detect gradient anomalies during model training, effectively suppress the occurrence of training divergence or loss spikes, and thus effectively improve the stability of model training.

[0040] The above summary is only for the purpose of the specification and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features of the present application will become apparent by reference to the drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In the drawings, unless otherwise specified, the same reference numerals throughout the several views refer to the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments disclosed in the present application and should not be regarded as limiting the scope of the present application.

[0042] Figure 1 is a schematic diagram of the loss curve during model training;

[0043] Figure 2 is a schematic flowchart of a model training method according to an embodiment of the present disclosure;

[0044] Figure 3 is a schematic diagram of the global level and local level divided according to granularity according to an embodiment of the present disclosure;

[0045] Figure 4 is a schematic architecture diagram of model training according to an embodiment of the present disclosure;

[0046] Figure 5 is a schematic diagram comparing the loss curves before and after optimization during model training according to an embodiment of the present disclosure;

[0047] Figure 6 is a schematic flowchart of an image processing method according to an embodiment of the present disclosure;

[0048] Figure 7 is a schematic flowchart of a natural language processing method according to an embodiment of the present disclosure;

[0049] Figure 8 is a schematic flowchart of a computer vision processing method according to an embodiment of the present disclosure;

[0050] Figure 9 is a schematic structural diagram of a model training device according to an embodiment of the present disclosure;

[0051] Figure 10 is a schematic structural diagram of an image processing device according to an embodiment of the present disclosure;

[0052] Figure 11It is a schematic structural diagram of a natural language processing device according to an embodiment of the present disclosure;

[0053] Figure 12 It is a schematic structural diagram of a computer vision processing device according to an embodiment of the present disclosure;

[0054] Figure 13 It is a schematic scenario diagram of a model training method according to an embodiment of the present disclosure;

[0055] Figure 14 It is a schematic scenario diagram of an image processing method according to an embodiment of the present disclosure;

[0056] Figure 15 It is a schematic scenario diagram of a natural language processing method according to an embodiment of the present disclosure;

[0057] Figure 16 It is a schematic scenario diagram of a computer vision processing method according to an embodiment of the present disclosure;

[0058] Figure 17 It is a schematic structural diagram of an electronic device for implementing the model training method and / or image processing method and / or natural language processing method and / or computer vision processing method according to an embodiment of the present disclosure. Detailed implementation manners

[0059] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted below.

[0060] In the embodiments of the specification, claims, and the above-mentioned drawings of the present disclosure, terms such as "first", "second", and "third" are used to distinguish similar objects and do not necessarily describe a specific order or sequence. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.

[0061] In the related art, problems such as unstable training convergence often occur during model training, such as training divergence or loss spikes during the training process. Figure 1 Shows a schematic diagram of the loss curve during the model training process, such as Figure 1As shown, during the training process of the large model, the loss curve shows loss spikes, indicating that the model experiences unstable convergence during training.

[0062] In related technologies, the following methods are usually used to solve the problem of unstable model training:

[0063] (1) Reduce the learning rate;

[0064] (2) Gradient scaling;

[0065] (3) Gradient clipping;

[0066] (4) Improve through the network model, such as the Post Layer Norm model;

[0067] (5) Add an auxiliary loss, such as the Max-z loss;

[0068] (6) Replace the data;

[0069] (7) Roll back to a recent checkpoint with good status and continue training.

[0070] The model training method proposed in this disclosure belongs to the gradient clipping method. Table 1 shows the existing gradient clipping methods, as shown in Table 1:

[0071]

[0072]

[0073] Table 1

[0074] Among them, GC directly compares with the preset constant 1.0. If the L2-Norm of the current gradient is larger than the preset constant 1.0, the gradient of this round is clipped. While AGC, LAMB, and Clippy introduce weights. Both AGC and LAMB use the L2-Norm of the weights, while Clippy uses the L1-Norm. Among them, GC belongs to the fixed gradient clipping method, while AGC, LAMB, and Clippy belong to the adaptive gradient clipping methods. Here, the expansions of L1-Norm and L2-Norm are the Euclidean norms.

[0075] According to Table 1, the current adaptive gradient clipping methods include AGC, LAMB, and Clippy. The clipping conditions of the above clipping methods all introduce weights additionally. Although a certain degree of adaptability is achieved, the historical gradient changes are not considered.

[0076] To at least partially address one or more of the above problems and other potential problems, the present disclosure proposes a model training method with adaptive gradient clipping. By mixing different granularities together to address gradient anomalies, while considering the historical gradient changes, it can promptly detect gradient anomalies during model training and promptly clip the current round of gradients, effectively suppressing the occurrence of training divergence or loss spikes, thereby effectively improving the stability of model training and helping to save the overall training time.

[0077] An embodiment of the present disclosure provides a model training method. Figure 2 It is a flowchart of the model training method according to an embodiment of the present disclosure. This model training method can be applied to a model training device. The model training device is located on an electronic device. The electronic device includes but is not limited to fixed devices and / or mobile devices. For example, fixed devices include but are not limited to servers, and the server can be a cloud server or a general server. For example, mobile devices include but are not limited to mobile phones, tablets, etc. In some possible implementation manners, this model training method can also be implemented by a processor invoking computer-readable instructions stored in a memory. As Figure 2 shown, this model training method includes:

[0078] S201: Deploy the model to be trained on a computing device;

[0079] S202: Train the model to be trained deployed on the computing device to obtain a target model;

[0080] Among them, S202 includes:

[0081] S202a: Obtain the value representing the current round of gradients and the value representing the historical gradients;

[0082] S202b: Based on the value representing the historical gradients and the value representing the current round of gradients, obtain a reference value corresponding to the target clipping method;

[0083] S202c: When the reference value reaches the clipping criterion corresponding to the target clipping method, use the target clipping method to clip the current round of gradients.

[0084] In an embodiment of the present disclosure, the computing device can be a server device or a terminal device. For example, if the computing device can be an image processing device, the target model is an image processing model; for another example, if the computing device can be a natural language processing device, the target model is a natural language processing model; for still another example, if the computing device can be multiple computer vision processing devices, the target model is a computer vision processing model. The above is only an exemplary illustration and does not limit all possible types of the computing device, and only not all are enumerated here.

[0085] In the embodiments of the present disclosure, the model to be trained may be a Transformer model, a Convolutional Neural Network (CNN), or a Recurrent Neural Network (RNN). The above is only an exemplary illustration and does not limit all possible structures of the model to be trained. Here, an exhaustive list is not provided.

[0086] In the embodiments of the present disclosure, the target model may be an image processing model, which is used to perform at least one of image classification, image recognition, and image segmentation on the image to be processed to obtain an image processing result. The target model may also be a natural language processing model, which is used to perform at least one of information extraction, text classification, text recognition, speech recognition, and question answering on the first type of data to be processed to obtain a natural language processing result. The target model may also be a computer vision processing model, which is used to perform at least one of image recognition, object detection, semantic segmentation, video understanding, and image generation on the second type of data to be processed to obtain a computer vision processing result.

[0087] In the embodiments of the present disclosure, the value representing the current round of gradient may be the norm value of the current round of gradient, such as the L1 norm (denoted as L1-Norm). The value representing the current round of gradient can be solved by formula (1):

[0088]

[0089] where x represents each element in the current round of gradient. Here, the value of the current round of gradient is calculated in real time through the above formula (1). Specifically, when it is detected that a preset condition is satisfied, the value representing the current round of gradient is obtained. Here, the preset condition may be set by the system by default or set according to user requirements. For example, the preset condition may be to obtain the value representing the current round of gradient every time a step is updated. Another example is that the preset condition may be to obtain the value representing the current round of gradient every two steps. Another example is that the preset condition may be to obtain the value representing the current round of gradient every t seconds. The above is only an exemplary illustration and does not limit all possible situations of the preset condition. Here, an exhaustive list is not provided.

[0090] It should be noted that the value representing the historical gradient can also be calculated by the above formula (1).

[0091] In the embodiments of the present disclosure, the value representing the historical gradient may be the norm value of the historical gradient, such as the L2 norm (denoted as L2-Norm). The value representing the historical gradient can be solved by formula (2):

[0092]

[0093] Expanding formula (2) can obtain formula (3):

[0094]

[0095] Among them, x represents each element in the historical gradient. Here, the value representing the historical gradient can be calculated through the above formula (2) or formula (3). Specifically, when it is detected that a preset condition is satisfied, the value representing the historical gradient is obtained; or the value representing the historical gradient is obtained from the data source of the computing device. Here, the preset condition can be set by the system default or according to user requirements. For example, the preset condition can be to obtain the value representing the current round of gradient every time a step is updated. Another example is that the preset condition can also be to obtain the value representing the current round of gradient every two updates. Another example is that the preset condition can also be to obtain the value representing the current round of gradient every t seconds. The above is only an exemplary illustration and does not limit all possible situations of the preset condition, but exhaustive listing is not done here.

[0096] It should be noted that the value representing the current round of gradient can also be calculated in real time through the above formula (2) or formula (3).

[0097] Figure 3 Shows a schematic diagram of the global level and the local level according to the granularity division, such as Figure 3 As shown, according to the granularity, the target cropping method can be divided into two levels: the global level and the local level. Among them, the local level can be divided into the tensor level, the split tensor level, and the element-by-element level, etc. Here, for the global level, the tensor level, the split tensor level, and the element-by-element level, the granularity becomes finer and finer in turn.

[0098] In the embodiments of the present disclosure, the reference value is calculated based on the value representing the current round of gradient and the value representing the historical gradient. Specifically, the reference value can be equal to the ratio of the value representing the historical gradient to the value representing the current round of gradient; the reference value can also be equal to the ratio of the value representing the current round of gradient to the value representing the historical gradient.

[0099] In the embodiments of the present disclosure, the clipping criteria corresponding to different clipping methods may be different. Exemplarily, the clipping criteria may be specific numerical values or thresholds. The clipping criteria corresponding to each clipping method may be different. For example, for the global-level clipping method, when the ratio of the value representing the current-round gradient to the value representing the historical gradient is greater than λ1, it meets the clipping criteria; for the tensor-level clipping method, when the ratio of the value representing the current-round gradient to the value representing the historical gradient is greater than λ2, it meets the clipping criteria; for the split-tensor-level clipping method, when the ratio of the value representing the current-round gradient to the value representing the historical gradient is greater than λ3, it meets the clipping criteria; for the element-wise-level clipping method, when the ratio of the value representing the current-round gradient to the value representing the historical gradient is greater than λ4, it meets the clipping criteria; where the values of λ1, λ2, λ3, and λ4 may be the same or different. Another example is that for the global-level clipping method, when the ratio of the value representing the historical gradient to the value representing the current-round gradient is less than λ5, it meets the clipping criteria; for the tensor-level clipping method, when the ratio of the value representing the historical gradient to the value representing the current-round gradient is less than λ6, it meets the clipping criteria; for the split-tensor-level clipping method, when the ratio of the value representing the historical gradient to the value representing the current-round gradient is less than λ7, it meets the clipping criteria; for the element-wise-level clipping method, when the ratio of the value representing the historical gradient to the value representing the current-round gradient is less than λ8, it meets the clipping criteria; where the values of λ5, λ6, λ7, and λ8 may be the same or different.

[0100] It should be noted that for different clipping methods, even if their clipping criteria are the same, since the reference values calculated by different clipping methods may be different, there may be a situation where some clipping methods meet the clipping criteria while some do not.

[0101] Figure 4 shows a schematic diagram of the architecture of model training, such as Figure 4As shown, deploy the model to be trained on a computing device; train the model to be trained deployed on the computing device to obtain a target model; training the model to be trained deployed on the computing device includes: obtaining a value representing the current round of gradients and a value representing the historical gradients; based on the value representing the historical gradients and the value representing the current round of gradients, obtaining a reference value corresponding to a target clipping method. The target clipping methods include clipping method 1, clipping method 2, and clipping method n; for example, if the reference value meets the clipping criteria corresponding to clipping method 1 and clipping method 2, and the reference value does not meet the clipping criteria of clipping method n, then use clipping method 1 and clipping method 2 to clip the current round of gradients, and do not use clipping method n to clip the current round of gradients. If the reference value only meets the clipping criteria of clipping method 1, then only use clipping method 1 to clip the current round of gradients, and do not use other clipping methods to clip the current round of gradients. If the reference value only meets the clipping criteria of clipping method 2, then only use clipping method 2 to clip the current round of gradients, and do not use other clipping methods to clip the current round of gradients. Here, among the clipping method 1, clipping method 2, and clipping method n, one clipping method can be a global-level clipping method, and the remaining clipping methods can be local-level clipping methods.

[0102] In the technical solution of the embodiment of the present disclosure, deploy the model to be trained on a computing device; train the model to be trained deployed on the computing device to obtain a target model; wherein, training the model to be trained deployed on the computing device includes: obtaining a value representing the current round of gradients and a value representing the historical gradients; based on the value representing the historical gradients and the value representing the current round of gradients, obtaining a reference value corresponding to a target clipping method; in the case where the reference value reaches the clipping criteria corresponding to the target clipping method, use the target clipping method to clip the current round of gradients. In this way, both the historical gradient changes and different granularities are considered, and adaptive gradient clipping can be achieved with a small computational cost; during the training of large models, gradient anomalies that occur during the model training process can be detected in a timely manner, thereby effectively suppressing the occurrence of training divergence or loss spikes, and then effectively improving the stability of model training.

[0103] In the embodiment of the present disclosure, the model training method may further include: obtaining granularity indication information, where the granularity indication information is used to indicate the granularity used in the current training; determining a target clipping method based on the granularity indication information.

[0104] In the embodiment of the present disclosure, the granularity used in the model training may include global granularity and local granularity. Among them, the global granularity may include the global level, and the local granularity may include the tensor level, the split tensor level, and the element-by-element level.

[0105] In the disclosed embodiment, the granularity indication information can be determined based on the information input by the user on the control panel. Specifically, the control panel displays a granularity parameter that can be selected by the user; the granularity indication information is determined based on the selection operation of the granularity parameter. The granularity indication information is used to indicate the granularity used in this training; the granularity indicated by the granularity indication information can be one or more granularities used in this training; the granularity indicated by the granularity indication information can be both a global granularity and one or more local granularities.

[0106] The technical solution of the embodiment of the present disclosure obtains granularity indication information; and determines the target cropping method based on the granularity indication information. In this way, the computing device can flexibly set the granularity of model training, provide training application scenarios with different cropping granularities, and improve the flexibility of model training, so that target models with different accuracies can be obtained based on the same training samples.

[0107] In the disclosed embodiment, clipping the current round gradient includes: determining the current round clipping factor; when it is determined that the current round gradient needs to be clipped, replacing the current round gradient with the product of the current round clipping factor and the current round gradient.

[0108] In the disclosed embodiment, the adaptive gradient clipping method can be simplified to obtain formula (4):

[0109] g=σ*g (4)

[0110] Among them, σ is the clipping factor, and g is the current round gradient. The present disclosure proposes the design of "σ"; the adaptive gradient clipping method of the present disclosure not only takes into account different granularities, but also introduces the consideration of historical gradient changes. In this way, gradient anomalies can be discovered in time, and the abnormal current round gradients can be clipped, thereby suppressing the generation of training divergence or loss spikes.

[0111] The technical solution of the embodiment of the present disclosure determines the clipping factor of the current round; when it is determined that the gradient of the current round needs to be clipped, the product of the clipping factor of the current round and the gradient of the current round is used to replace the gradient of the current round. In this way, by designing the clipping factor, the gradient of the current round can be clipped in scenarios of different granularities, which helps to suppress the occurrence of training divergence or loss spikes without affecting the learning speed, so that the model can be trained more stably and the stability of model training is improved.

[0112] In some embodiments, the adopting target clipping mode to clip the current round of gradients includes: adopting the target clipping mode to clip the current round of gradients, including: when the target clipping mode is a global level clipping mode, determining the clipping factor of the current round according to a first formula, wherein the first formula is: Among them, ||g ema ||2=β||g ema||2 + (1 - β)||g||2, where λ is a preset value, g ema is the historical gradient, g is the current round's gradient, β represents a preset adjustment factor, and σ represents the clipping factor for this round.

[0113] Here, the value of λ can be set or adjusted according to user requirements such as precision requirements or speed requirements, etc.

[0114] In the embodiments of the present disclosure, this global-level clipping includes: squaring the elements in the tensor corresponding to each gradient, then globally summing, and then taking the square root to obtain the global-level ||g||2; recording the historical information of the gradient through exponential moving average, and determining the global-level ||g ema ||2 according to the historical information; according to the global-level ||g||2 and the global-level ||g ema ||2, determining whether the global-level clipping condition is met. Here, when the global-level clipping condition is met, clip the current round's gradient; the clipping condition is The clipping condition means that: if the global ||g||2 at the current step is greater than λ times the global ||g ema ||2, clip the current round's gradient; if the global ||g||2 at the current step is less than or equal to λ times the global ||g ema ||2 or the global ||g||2 at the current step is equal to λ times the global ||g ema ||2, do not clip the current round's gradient. Exemplarily, if tensor 1 is A = [a11, a12, a13, a21, a22, a23], where a11 = 1, a12 = 2, a13 = 3, a21 = 4, a22 = 5, a23 = 6, specifically, tensor 1 is represented as A = [1, 2, 3][4, 5, 6], and tensor 2 is B = [b11, b12, b13], where b11 = 7, b12 = 8, b13 = 9, specifically, tensor 2 is represented as B = [7, 8, 9]; then Similarly, use the exponential moving average method to record g ema , and calculate ||g||2 based on a similar way of obtaining ||g||2 as above. ema Based on the above ||g||2 and ||g ema ||2, perform the following operations on tensor 1 and tensor 2 respectively: If do not clip the current round's gradient (such as tensor 1 and tensor 2); if calculate σ according to the first formula, and clip the current round's gradient (such as tensor 1 and tensor 2) based on this σ and the above formula (4).

[0115] Here, if λ = 1.2, it means that the fluctuation range between the current round's gradient and the historical gradient does not exceed 20%.

[0116] Exemplarily, if ||g||2 of the gradient g at the current step is 0.5, and the historical gradient g ema ||g ema ||2 is 1.0, and λ = 1.2, no clipping is performed. Specifically, then σ = 1.0, ||g||2 = σ * ||g||2 = 1.0 * ||g||2 = 1×0.5 = 0.5, which is equivalent to not performing clipping.

[0117] Exemplarily, if ||g||2 of the gradient g at the current step is 5.0, and the historical gradient g ema ||g ema ||2 is 1.0, and λ = 1.2, then clipping is performed. Specifically, then σ = 0.24, ||g||2 = σ * ||g||2 = 0.24 * ||g||2 = 0.24×5 = 1.2, and ||g||2 of the gradient in this round is clipped from 5.0 to 1.2. Using the existing clipping method, if ||g||2 of the gradient g in the previous step is 5.0 and λ = 0.1, then σ = 0.02, ||g||2 = σ * ||g||2 = 0.02 * ||g||2 = 0.02×5 = 0.1, and ||g||2 of the gradient in this round is clipped from 5.0 to 0.1. Obviously, the clipping multiple of the existing clipping method is much larger than that of the present disclosure. The clipping method of the present disclosure is milder and more reasonable, can quickly and timely identify abnormal gradients, and timely clip abnormal gradients, thereby controlling the gradient of each step within a certain range, and thus can stabilize the training of the model.

[0118] The technical solution of the embodiment of the present disclosure uses a global-level clipping method to clip the gradient in this round, which can consider from a global perspective, judge whether to clip the gradient in this round from a global perspective, and decide to perform clipping on all tensors based on the judgment result. All tensors are clipped using the unified judgment result, which helps to improve the clipping efficiency and thus helps to improve the speed of model training.

[0119] In some embodiments, using a target clipping method to clip the gradient in this round includes: when the target clipping method is a tensor-level clipping method, determining the clipping factor in this round according to the first formula, where the first formula is: where ||g ema ||2 = β||g ema ||2 + (1 - β)||g||2, λ is a preset value, g ema is the historical gradient, g is the gradient in this round, β represents a preset adjustment factor, and σ represents the clipping factor in this round.

[0120] Here, λ can be set or adjusted according to user requirements such as precision requirements or speed requirements, etc.

[0121] In the embodiments of the present disclosure, the tensor level is also called the layerwise level. Tensor-level pruning includes: for each non-split parameter, during gradient pruning, for each complete tensor, calculate its own ||g ema ||2 and ||g||2, and then each complete tensor performs gradient pruning separately. Here, when the pruning condition at the tensor level is reached, the gradient of this round is pruned; the pruning condition is This pruning condition means that if the ||g||2 of the current step is larger than λ times of ||g ema ||2, the gradient of this round is pruned; if the ||g||2 of the current step is smaller than λ times of ||g ema ||2 or the ||g||2 of the current step is equal to λ times of ||g ema ||2, the gradient of this round is not pruned. Exemplarily, if tensor 1 is A = [a11, a12, a13, a21, a22, a23], where a11 = 1, a12 = 2, a13 = 3, a21 = 4, a22 = 5, a23 = 6, specifically, tensor 1 is represented as A = [1, 2, 3][4, 5, 6], and tensor 2 is B = [b11, b12, b13], where b11 = 7, b12 = 8, b13 = 9, specifically, tensor 2 is represented as B = [7, 8, 9]; then the Similarly, the sliding average method is used to record g ema , and ||g ema ||2 of tensor 1 is calculated based on a method similar to the above for obtaining ||g||2; according to ||g||2 and ||g ema ||2 of tensor 1, it is determined whether to prune the gradient of this round for tensor 1; for Similarly, the sliding average method is used to record g ema , and ||g ema ||2 of tensor 2 is calculated based on a method similar to the above for obtaining ||g||2; according to ||g||2 and ||g ema ||2 of tensor 2, it is determined whether to prune the gradient of this round for tensor 2. Among them, the pruning of tensor 1 and tensor 2 does not interfere with each other. ||g ema ||2 of tensor 1 and ||g ema ||2 of tensor 2 are independent. Based on the above, the following operations are performed on tensor 1:

[0122] If Do not prune the gradient of this round; if Calculate σ according to the first formula, and clip the current round of gradient based on this σ and the above formula (4).

[0123] Exemplarily, if ||g||2 of the current step gradient g of tensor 1 is 0.2, and ||g ema ||2 of the historical gradient g ema is 1.0, and λ = 1.2, no clipping is performed. Specifically,

[0124]

[0125] then σ = 1.0, ||g||2 = σ * ||g||2 = 1 * ||g||2 = 1 × 0.2 = 0.2, which is equivalent to no clipping.

[0126] Exemplarily, if ||g||2 of the current step gradient g of tensor 2 is 10, and ||g ema ||2 of the historical gradient g ema is 1.0, and λ = 1.2, clipping is performed. Specifically, then σ = 0.12, ||g||2 = σ * ||g||2 = 0.12 * ||g||2 = 0.12 × 10 = 1.2, and ||g||2 of the current round gradient g of this step is clipped from 10 to 1.2.

[0127] The technical solution of the embodiment of the present disclosure clips the current round of gradient by adopting a tensor-level clipping method, determines whether to clip the current round of gradient from the tensor-level perspective, and each tensor decides whether to perform clipping on itself according to the judgment result of its own tensor, which helps to improve the diversity of clipping and further helps to improve the flexibility of model training. Clipping the current round of gradient by adopting a tensor-level clipping method helps to improve the clipping effect. Compared with the global-level clipping, the tensor-level clipping is more delicate and the clipping effect is better.

[0128] In some embodiments, clipping the current round of gradient by adopting a target clipping method includes: when the target clipping method is a split tensor-level clipping method, determining the clipping factor of the current round according to the first formula, where the first formula is: where ||g ema ||2 = β||g ema ||2 + (1 - β)||g||2, λ is a preset value, g ema is the historical gradient, g is the current round gradient, β represents a preset adjustment factor, and σ represents the clipping factor of the current round.

[0129] Here, λ can be set or adjusted according to user requirements such as accuracy requirements or speed requirements, etc.

[0130] In the embodiments of the present disclosure, the split tensor, i.e., the distributed split parameter, for example, the split parameter of tensor parallelism. The split tensor level clipping includes: during gradient clipping, for each tensor in each split state, calculate their respective ||g ema ||2 and ||g||2 independently, and perform gradient clipping on each tensor in each split state. Here, when the clipping condition at the split tensor level is reached, clip the current round of gradients; the clipping condition is This clipping condition means that if the global ||g||2 at the current step is larger than λ times of ||g ema ||2, clip the current round of gradients; if the global ||g||2 at the current step is smaller than λ times of ||g ema ||2 or the global ||g||2 at the current step is equal to λ times of ||g ema ||2, do not clip the current round of gradients. Exemplarily, if tensor 1 is A = [a11, a12, a13, a21, a22, a23], where a11 = 1, a12 = 2, a13 = 3, a21 = 4, a22 = 5, a23 = 6, specifically, tensor 1 is represented as A = [1, 2, 3][4, 5, 6], and tensor 3 is C = [c11, c12, c13, c21, c22, c23], where c11 = 7, c12 = 8, c13 = 9, c21 = 1, c22 = 2, c23 = 3, specifically, tensor 3 is represented as B = [7, 8, 9][1, 2, 3]; if tensor 1 is split into tensor 1-1 and tensor 1-2, that is, for tensor 1-1 For tensor 1-2 Split tensor 3 into tensor 3-1 and tensor 3-2, that is, for tensor 3-1 For tensor 3-2 Among them, perform the following operations on tensor 1-1, tensor 1-2, tensor 3-2, and tensor 3-2 respectively: if Do not clip the current round of gradients; if Calculate σ according to the first formula, and clip the current round of gradients based on this σ and the above formula (4).

[0131] Tensor 1-1 decides whether to clip the current round of gradients of tensor 1-1 according to the above calculation results; tensor 1-2 decides whether to clip the current round of gradients of tensor 1-2 according to the above calculation results; tensor 3-1 decides whether to clip the current round of gradients of tensor 3-1 according to the above calculation results; tensor 3-2 decides whether to clip the current round of gradients of tensor 3-2 according to the above calculation results; tensor 1-1, tensor 1-2, tensor 3-1, and tensor 3-2 calculate their respective corresponding ||g ema||2 and ||g||2, and based on their own calculation results, decide whether to clip themselves, without interfering with each other. The ||g ema ||2 of Tensor 1-1, Tensor 1-2, Tensor 3-1, and Tensor 3-2 are independent of each other. Record the ||g ema ||2 of Tensor 1-1, Tensor 1-2, Tensor 3-1, and Tensor 3-2 respectively through exponential moving average.

[0132] Exemplarily, if ||g||2 of the current-step gradient g of Tensor 1-1 is 0.2, and ||g ema of the historical gradient g ema ||2 is 1.0, and λ = 1.2, no clipping is performed. If ||g||2 of the current-step gradient g of Tensor 1-2 is 10, and ||g ema of the historical gradient g ema || 2 = 1.0, and λ = 1.2, clipping is performed. Specifically, then σ = 0.12, ||g||2 = σ * ||g||2 = 0.12 * ||g||2 = 0.12 × 10 = 1.2, and clip ||g||2 = 10 of the current-round gradient g to ||g||2 = 1.2.

[0133] The technical solution of the embodiments of the present disclosure uses a sliced-tensor-level clipping method to clip the current-round gradient, determines whether to clip the current-round gradient from the perspective of the sliced-tensor level, and each sliced tensor decides whether to perform clipping on itself according to the judgment result of the sliced tensor after slicing, which helps to improve the diversity of clipping, helps to improve the flexibility of model training, and helps to improve the speed of model training. Using the sliced-tensor-level clipping method to clip the current-round gradient helps to improve the clipping effect. Compared with the global-level and tensor-level clippings, the sliced-tensor-level clipping is more delicate, helps to save the computing resources of the devices where each sliced tensor is located, and helps to improve the training efficiency and stability of the model.

[0134] In some embodiments, using a target clipping method to clip the current-round gradient includes: when the target clipping method is an element-wise level clipping method, determining the clipping factor of the current round according to the second formula, where the second formula is: where, λ is a preset value, v i is the historical gradient, g i is the current-round gradient, β represents a preset adjustment factor, and σ represents the clipping factor of the current round.

[0135] In an embodiment of the present disclosure, this per-element level is a strategy with the finest granularity compared to the global level, tensor level, and sliced tensor level. This per-element level clipping includes: when performing gradient clipping, for each independent element, calculating its respective |g i | and v i ; if the |g i | of the current step is larger than λ times v i , clip the gradient of this round; if the |g i | of the current step is smaller than λ times v i or the |g i | of the current step is equal to λ times v i , do not clip the gradient of this round. Here, by using the information of the second-order gradient variance in Adaptive Moment Estimation (Adam) and Adam With Weight Decay (Adam W), additional storage overhead can be avoided.

[0136] In an embodiment of the present disclosure, in this per-element level clipping, L1-Norm = L2-Norm.

[0137] Exemplarily, if tensor 1 is A = [a11, a12, a13, a21, a22, a23], where a11 = 1, a12 = 2, a13 = 3, a21 = 4, a22 = 5, a23 = 6, specifically, tensor 1 is represented as A = [1, 2, 3][4, 5, 6]; in tensor 1: x1 = 1; x2 = 2; x3 = 3; x4 = 4; x5 = 5; x6 = 6; then for the six elements x1, x2, x3, x4, x5, x6, calculate their respective ||g ema ||2 and ||g||2, and based on their respective calculation results, decide whether to clip, but the six elements x1, x2, x3, x4, x5, and x6 do not interfere with each other. The ||g ema ||2 of the six elements x1, x2, x3, x4, x5, and x6 are recorded separately through exponential moving average.

[0138] The technical solution of the embodiments of the present disclosure determines whether to clip the current round of gradients from the per-element level perspective. Each element included in each tensor decides whether to perform clipping on itself based on its own judgment result, which helps to improve the diversity of clipping, improve the flexibility of model training, and help to improve the speed of model training. Using the per-element level clipping method to clip the current round of gradients helps to improve the clipping effect. Compared with the global level, tensor level, and split tensor level clipping, the per-element level clipping is more delicate, which helps to save the computing resources of the computing device and helps to improve the training efficiency and stability of the model.

[0139] In some embodiments, the target clipping method includes the global level clipping method. When the reference value reaches the clipping standard corresponding to the target clipping method, the current round of gradients is clipped using the target clipping method, including: when the reference value reaches the clipping standard corresponding to the global level clipping method, the current round of gradients is clipped using the global level clipping method.

[0140] In some embodiments, the global level clipping includes: squaring the elements in the tensor corresponding to each gradient, then globally summing, and then taking the square root to obtain the global level ||g||2; recording the historical information of the gradients through exponential moving average, and determining the global level ||g ema ||2 according to the historical information; judging whether the global level clipping condition is reached based on the global level ||g||2 and the global level ||g ema ||2. Here, when the global level clipping condition is reached, the current round of gradients is clipped.

[0141] In some embodiments, a model to be trained is deployed on a computing device. The training data of the model to be trained includes pure image data and annotated image data; the model to be trained deployed on the computing device is trained to obtain a target model, and the target model is an image recognition model; wherein, training the model to be trained deployed on the computing device includes: obtaining a value representing the current round of gradients and a value representing the historical gradients; obtaining a reference value corresponding to the target clipping method based on the value representing the historical gradients and the value representing the current round of gradients; when the reference value reaches the clipping standard corresponding to the global level, the current round of gradients is clipped using the global level.

[0142] The technical solution of the embodiments of the present disclosure clips the current round of gradients using the global level clipping method when the reference value reaches the clipping standard corresponding to the global level clipping method. In this way, the global level clipping granularity can be adopted according to actual needs, which helps to improve the flexibility of adaptive gradient clipping. At the same time, adopting the global level clipping granularity helps to improve the training efficiency of the model.

[0143] In some embodiments, the target clipping method includes a local-level clipping method. When the reference value reaches the clipping criterion corresponding to the target clipping method, the target clipping method is used to clip the current-round gradient, including: when the reference value reaches the clipping criterion corresponding to the local-level clipping method, the local-level clipping method is used to clip the current-round gradient.

[0144] Here, the local-level clipping method may include at least one of a tensor-level clipping method, a sliced-tensor-level clipping method, and an element-wise-level clipping method.

[0145] In some embodiments, a model to be trained is deployed on a computing device. The training data of the model to be trained includes audio data and text data; the model to be trained deployed on the computing device is trained to obtain a target model, and the target model is a speech recognition model; wherein, training the model to be trained deployed on the computing device includes: obtaining a value representing the current-round gradient and a value representing the historical gradient; based on the value representing the historical gradient and the value representing the current-round gradient, obtaining a reference value corresponding to the target clipping method; when the reference value reaches the clipping criterion corresponding to the local level, the local level is used to clip the current-round gradient.

[0146] The technical solution of the embodiments of the present disclosure uses the local-level clipping method to clip the current-round gradient when the reference value reaches the clipping criterion corresponding to the local-level clipping method. In this way, the local-level clipping granularity can be adopted according to actual needs, which helps to improve the flexibility of adaptive gradient clipping. At the same time, adopting a finer local-level clipping granularity helps to improve the effect of model training.

[0147] In some embodiments, when the reference value reaches the clipping criterion corresponding to the local-level clipping method, using the local-level clipping method to clip the current-round gradient includes at least one of the following: when the reference value reaches the clipping criterion corresponding to the first type of local-level clipping method, using the first type of local-level clipping method to clip the current-round gradient; when the reference value reaches the clipping criterion corresponding to the second type of local-level clipping method, using the second type of local-level clipping method to clip the current-round gradient; when the reference value reaches the clipping criterion corresponding to the third type of local-level clipping method, using the third type of local-level clipping method to clip the current-round gradient; wherein, the first type of local-level clipping method, the second type of local-level clipping method, and the third type of local-level clipping method are local-level clipping methods with different granularities.

[0148] In embodiments of the present disclosure, the local level may include a first type of local pruning level, a second type of local pruning level, and a third type of local pruning level. If the pruning granularity of the first type of local pruning level is at the tensor level, the pruning granularity of the second type of local pruning level is at the split tensor level, and the pruning granularity of the third type of local pruning level is at the element-by-element level. Here, the pruning granularity decreases step by step. When the pruning granularity is smaller, the pruning degree is more delicate.

[0149] In embodiments of the present disclosure, the local level pruning scheme may include: only adopting the first type of local pruning level; or only adopting the second type of local pruning level; or only adopting the third type of local pruning level; or adopting the first type of local pruning level and the second type of local pruning level; or adopting the first type of local pruning level and the third type of local pruning level; or adopting the second type of local pruning level and the third type of local pruning level; or simultaneously adopting the first type of local pruning level, the second type of local pruning level, and the third type of local pruning level.

[0150] The technical solution of the embodiments of the present disclosure divides the local level pruning into the tensor level, the split tensor level, and the element-by-element level according to different granularities. In this way, mixing and matching the tensor level, the split tensor level, and the element-by-element level of the local level pruning can give application scenarios with different pruning granularities, which helps to improve the flexibility of adaptive gradient pruning, thereby improving the effect of model training.

[0151] In some embodiments, the target pruning method includes a global level pruning method and at least one local level pruning method. When the reference value reaches the pruning standard corresponding to the target pruning method, the target pruning method is used to prune the current round of gradients, including: when the reference value reaches the pruning standard corresponding to the global level pruning method, the global level pruning method is used to prune the current round of gradients; when the reference value does not reach the pruning standard corresponding to the global level pruning method, at least one local level pruning method is used to prune the current round of gradients.

[0152] That is to say, when the pruning standard corresponding to the global level pruning method is reached, the global level pruning method is used to prune the current round of gradients; only when the pruning standard corresponding to the global level pruning method is not reached, the local level pruning method is used to prune the current round of gradients.

[0153] In the embodiments of the present disclosure, the solution of mixing the global level and the local level may include: when the reference value reaches the clipping criterion corresponding to the global level clipping method, clipping the current round of gradients by using the global level clipping method; when the reference value does not reach the clipping criterion corresponding to the global level clipping method, only clipping the current round of gradients by using the first type of local clipping level; or when the reference value does not reach the clipping criterion corresponding to the global level clipping method, only clipping the current round of gradients by using the second type of local clipping level; or when the reference value does not reach the clipping criterion corresponding to the global level clipping method, only clipping the current round of gradients by using the third type of local clipping level; or when the reference value does not reach the clipping criterion corresponding to the global level clipping method, clipping the current round of gradients by using the first type of local clipping level and the second type of local level clipping; or when the reference value does not reach the clipping criterion corresponding to the global level clipping method, clipping the current round of gradients by using the first type of local clipping level and the third type of local level clipping; or when the reference value does not reach the clipping criterion corresponding to the global level clipping method, clipping the current round of gradients by using the second type of local clipping level and the third type of local level clipping; or clipping the current round of gradients by using the first type of local level clipping, the second type of local clipping level, and the third type of local level clipping at the same time.

[0154] For the technical solution of the embodiments of the present disclosure, when the global level clipping effective condition is reached, global level clipping is adopted; when the global level clipping effective condition is not reached, local level clipping is adopted. In this way, global level clipping and local level clipping can be flexibly called, which helps to improve the flexibility of adaptive gradient clipping, and thus helps to improve the stability of model training.

[0155] In some embodiments, the target clipping method includes a global level clipping method and at least one local level clipping method. When the reference value reaches the clipping criterion corresponding to the target clipping method, clipping the current round of gradients by using the target clipping method includes: when the reference value reaches the clipping criterion corresponding to the global level clipping method, clipping the current round of gradients by using the global level clipping method; when the reference value reaches the clipping criterion corresponding to at least one local level clipping method, clipping the current round of gradients by using the at least one local level clipping method.

[0156] In some embodiments, when the reference value meets the clipping criteria corresponding to the global-level clipping method and when the reference value meets the clipping criteria corresponding to the local-level clipping method, the solution of simultaneously adopting global-level clipping and local-level clipping may include: simultaneously adopting the global-level clipping method and the first type of local-level clipping; or simultaneously adopting the global-level clipping method and the second type of local-level clipping; or simultaneously adopting the global-level clipping method and the third type of local-level clipping; or simultaneously adopting the global-level clipping method, the first type of local-level clipping, and the second type of local-level clipping; or simultaneously adopting the global-level clipping method, the first type of local-level clipping, and the third type of local-level clipping; or simultaneously adopting the global-level clipping method, the second type of local-level clipping, and the third type of local-level clipping; or simultaneously adopting the global-level clipping method, the first type of local-level clipping, the second type of local-level clipping method, and the third type of local-level clipping.

[0157] Here, the global-level clipping method and the local-level clipping method are carried out synchronously.

[0158] The technical solution of the embodiments of the present disclosure, when the reference value meets the clipping criteria corresponding to the global-level clipping method and when the clipping criteria corresponding to the local-level clipping method, adopts both the global-level clipping method and the local-level clipping method. In this way, the intensity of adaptive gradient clipping can be improved, thereby improving the clipping effect of each step and the stability of model training. At the same time, it can adaptively adjust the clipping conditions with relatively small computational overhead and give application scenarios with different clipping granularities.

[0159] In some embodiments, the target clipping method includes a global-level clipping method and at least one local-level clipping method. When the reference value meets the clipping criteria corresponding to the target clipping method, clipping the current-round gradient using the target clipping method includes: when the reference value meets the clipping criteria corresponding to the global-level clipping method, clipping the current-round gradient using the global-level clipping method; after clipping the current-round gradient using the global-level clipping method, if the new reference value determined based on the clipped current-round gradient meets the clipping criteria corresponding to at least one local-level clipping method, then continue to clip the current-round gradient using the at least one local-level clipping method.

[0160] In some embodiments, when the reference value meets the clipping criteria corresponding to the global-level clipping method and the local-level clipping method, the solution of simultaneously adopting global-level clipping and local-level clipping may include: Solution (1) simultaneously adopting the global-level clipping method and the first type of local-level clipping; or Solution (2) simultaneously adopting the global-level clipping method and the second type of local-level clipping; or Solution (3) simultaneously adopting the global-level clipping method and the third type of local-level clipping; or Solution (4) simultaneously adopting the global-level clipping method, the first type of local-level clipping, and the second type of local-level clipping; or Solution (5) simultaneously adopting the global-level clipping, the first type of local-level clipping, and the third type of local-level clipping; or Solution (6) simultaneously adopting the global-level clipping method, the second type of local-level clipping, and the third type of local-level clipping; or Solution (7) simultaneously adopting the global-level clipping method, the first type of local-level clipping, the second type of local-level clipping method, and the third type of local-level clipping.

[0161] When the new reference value determined based on the clipped current-round gradient meets the clipping criteria corresponding to at least one local-level clipping method, the solution may include: If it meets the clipping criteria of the first type of local-level clipping, use the first type of local-level clipping method to continue clipping the clipped current-round gradient; or if it meets the clipping criteria of the second type of local-level clipping, use the second type of local-level clipping method to continue clipping the clipped current-round gradient; or if it meets the clipping criteria of the third type of local-level clipping, use the third type of local-level clipping method to continue clipping the clipped current-round gradient.

[0162] In the embodiments of the present disclosure, if the new reference value determined based on the clipped current-round gradient meets the clipping criteria corresponding to at least one local-level clipping method, at least one local-level clipping method is used to continue clipping the current-round gradient, including: If it meets the clipping criteria of the first type of local-level clipping, use the first type of local-level clipping method to continue clipping the clipped current-round gradient; after using the first type of local-level clipping method for clipping, if the new reference value determined based on the clipped current-round gradient meets the clipping criteria corresponding to the second type of local clipping, use the second type of local-level clipping method to continue clipping the current-round gradient after the first type of local clipping; after using the second type of local-level clipping method for clipping, if the new reference value determined based on the clipped current-round gradient meets the clipping criteria corresponding to the third type of local clipping, use the third type of local-level clipping method to continue clipping the current-round gradient after the second type of local clipping.

[0163] In the technical solution of the embodiment of the present disclosure, the global-level clipping method is first used to clip the current-round gradient; when the new reference value determined for the current-round gradient after clipping the current-round gradient reaches the effective state of the local-level clipping method, the local-level clipping method is then used to clip the current-round gradient. In this way, each round of gradient can be effectively clipped, which helps to improve the effect of adaptive gradient clipping and further improve the stability of model training.

[0164] Figure 5 is a schematic diagram of the comparison of the loss curves before and after optimization in the model training process according to the embodiment of the present disclosure. As Figure 5 shown, there are multiple loss spikes in the loss curve before optimization; the loss spikes in the loss curve after optimization are effectively eliminated, which can help the large model training to converge more stably.

[0165] The embodiment of the present disclosure provides an image processing method. Figure 6 is a schematic flowchart of the image processing method according to the embodiment of the present disclosure. This image processing method can be applied to an image processing device, and this image processing device can be applied to an electronic device. The electronic device includes but is not limited to a fixed device and / or a mobile device. For example, the fixed device includes but is not limited to a server, and the server can be a cloud server or a general server. For example, the mobile device includes but is not limited to a mobile phone, a tablet computer, etc. In some possible implementation manners, this image processing method can also be implemented by a processor calling computer-readable instructions stored in a memory. As Figure 6 shown, this image processing method includes:

[0166] S601: Input the image to be processed into the trained target model, and the trained target model is obtained by training according to the training method described above;

[0167] S602: According to the trained target model, perform at least one image processing including image classification, image recognition, and image segmentation on the image to be processed to obtain an image processing result.

[0168] In some embodiments, the image to be processed can be obtained from an image data source; it can also be crawled from a web page to obtain the image to be processed; it can also be multiple frames of images intercepted from a video to obtain the image to be processed.

[0169] In the embodiments of the present disclosure, when the target model is used for image recognition, the target model is an image recognition model. Specifically, a plurality of images to be recognized are obtained, and the images to be recognized are input into the image recognition model to obtain the recognition results of the images to be recognized output by the image recognition model. Exemplarily, a plurality of images to be recognized are obtained; the plurality of images to be recognized are respectively an image including "Philodendron", an image including "Epipremnum aureum", and an image including "Syngonium podophyllum". The plurality of images to be recognized are input into the image recognition model, and the recognition results are "Philodendron", "Epipremnum aureum", and "Syngonium podophyllum".

[0170] The technical solution of the embodiments of the present disclosure uses the target model trained by adaptive gradient clipping for image processing, which can provide a target model with high stability for the field of image processing, thereby improving the stability and accuracy of image processing.

[0171] The embodiments of the present disclosure provide a natural language processing method. Figure 7 is a schematic flowchart of the natural language processing method according to the embodiments of the present disclosure. The natural language processing method can be applied to a natural language processing device. The natural language processing device can be applied to an electronic device. The electronic device includes but is not limited to a fixed device and / or a mobile device. For example, the fixed device includes but is not limited to a server, and the server can be a cloud server or a general server. For example, the mobile device includes but is not limited to a mobile phone, a tablet computer, etc. In some possible implementation manners, the natural language processing method can also be implemented by a processor calling computer-readable instructions stored in a memory. As Figure 7 shown, the natural language processing method includes:

[0172] S701: Input the first type of data to be processed into the trained target model, and the trained target model is obtained by training according to the training method described above;

[0173] S702: According to the trained target model, perform at least one natural language processing including information extraction, text classification, text recognition, speech recognition, question answering on the first type of data to be processed to obtain a natural language processing result.

[0174] In some embodiments, the first type of data to be processed can be obtained from a data source; it can also be crawled from a web page to obtain the first type of data to be processed.

[0175] In some embodiments, the first type of data to be processed can be multiple text data or long text data, or multiple audio data or long audio data.

[0176] In an embodiment of the present disclosure, when the target model is used for text recognition, the target model is a text recognition model. Specifically, a plurality of data to be recognized are obtained; the data to be recognized are input into the text recognition model, and the recognition result of the data to be recognized output by the text recognition model is obtained. Exemplarily, a plurality of data to be recognized are obtained; the data to be recognized include "thorny roses", "thorny Chinese roses", and "thorny roses"; the data to be recognized are input into the text recognition model, and the text recognition result "thorny roses; thorny Chinese roses; thorny roses" is obtained.

[0177] The technical solution of the embodiment of the present disclosure uses the target model trained by adaptive gradient clipping for natural language processing, which can provide a target model with high stability for the field of natural language processing, thereby improving the stability and accuracy of natural language processing results.

[0178] The embodiment of the present disclosure provides a computer vision processing method. Figure 8 FIG. is a schematic flowchart of the computer vision processing method according to the embodiment of the present disclosure. The computer vision processing method can be applied to a computer vision processing device. The computer vision processing device can be applied to an electronic device. The electronic device includes but is not limited to a fixed device and / or a mobile device. For example, the fixed device includes but is not limited to a server, and the server can be a cloud server or a general server. For example, the mobile device includes but is not limited to a mobile phone, a tablet computer, etc. In some possible implementation manners, the computer vision processing method can also be implemented by a processor calling computer-readable instructions stored in a memory. As Figure 8 shown, the computer vision processing method includes:

[0179] S801: Input the second type of data to be processed into the trained target model, and the trained target model is obtained by training according to the training method described above;

[0180] S802: According to the trained target model, perform at least one computer vision processing including image recognition, target detection, semantic segmentation, video understanding, and image generation on the second type of data to be processed, and obtain a computer vision processing result.

[0181] In some embodiments, the second type of data to be processed can be multiple text data or long text data, or multiple audio data or long audio data, or multiple image data.

[0182] In an embodiment of the present disclosure, when the target model is used for target detection, the target model is a target detection model. Specifically, a plurality of images to be detected are obtained; the images to be detected are input into an image detection model, and detection results of the images to be detected output by the image detection model are obtained. Exemplarily, an image of a bird's nest and a big tree is obtained as the image to be detected, and the image to be detected is input into the target detection model, and the output target detection result is the bird's nest and the big tree.

[0183] The technical solution of the embodiment of the present disclosure uses the target model obtained by adaptive gradient clipping training for image processing, which can provide a target model with high stability for the field of computer vision processing, thereby improving the stability and accuracy of computer vision processing.

[0184] It should be understood that Figure 1 、 Figure 3 、 Figure 4 and Figure 5 The schematic diagrams shown are merely exemplary and not restrictive, and they are extensible. Those skilled in the art can make various obvious changes and / or substitutions based on the examples of Figure 1 、 Figure 3 、 Figure 4 and Figure 5 The obtained technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.

[0185] The embodiment of the present disclosure provides a model training device, as shown in Figure 9 . The model training device may include: a deployment module 901, configured to deploy a model to be trained on a computing device; a training module 902, configured to train the model to be trained deployed on the computing device to obtain a target model; wherein, the training module 902 includes: a first obtaining sub-module, configured to obtain a value representing the gradient of this round and a value representing the historical gradient; a second obtaining sub-module, configured to obtain a reference value corresponding to the target clipping method based on the value representing the historical gradient and the value representing the gradient of this round; a clipping sub-module, configured to clip the gradient of this round using the target clipping method when the reference value reaches the clipping standard corresponding to the target clipping method.

[0186] In some embodiments, the model training device further includes: an obtaining module (not shown in Figure 9 ) for obtaining granularity indication information, where the granularity indication information is used to indicate the granularity used in this training; a determining module (not shown in Figure 9 ) for determining a target clipping method based on the granularity indication information.

[0187] In some embodiments, the clipping sub-module is configured to: determine a clipping factor for this round; when it is determined that the gradient of this round needs to be clipped, replace the gradient of this round with the product of the clipping factor of this round and the gradient of this round.

[0188] In some embodiments, the clipping sub-module is configured to: when the target clipping mode is the global-level clipping mode, determine the clipping factor of this round according to the first formula, where the first formula is: where ||g ema ||2 = β||g ema ||2 + (1 - β)||g||2, λ is a preset value, g ema is the historical gradient, g is the gradient of this round, β represents a preset adjustment factor, and σ represents the clipping factor of this round.

[0189] In some embodiments, the clipping sub-module is configured to: when the target clipping mode is the tensor-level clipping mode, determine the clipping factor of this round according to the first formula, where the first formula is: where ||g ema ||2 = β||g ema ||2 + (1 - β)||g||2, λ is a preset value, g ema is the historical gradient, g is the gradient of this round, β represents a preset adjustment factor, and σ represents the clipping factor of this round.

[0190] In some embodiments, the clipping sub-module is configured to: when the target clipping mode is the sliced-tensor-level clipping mode, determine the clipping factor of this round according to the first formula, where the first formula is: where ||g ema ||2 = β||g ema ||2 + (1 - β)||g||2, λ is a preset value, g ema is the historical gradient, g is the gradient of this round, β represents a preset adjustment factor, and σ represents the clipping factor of this round.

[0191] In some embodiments, the clipping sub-module is configured to: when the target clipping mode is the element-level clipping mode, determine the clipping factor of this round according to the second formula, where the second formula is: where λ is a preset value, v i is the historical gradient, g i is the gradient of this round, β represents a preset adjustment factor, and σ represents the clipping factor of this round.

[0192] In some embodiments, the target clipping mode includes the global-level clipping mode. The clipping sub-module is configured to: when the reference value reaches the clipping standard corresponding to the global-level clipping mode, clip the gradient of this round using the global-level clipping mode.

[0193] In some embodiments, the target clipping method includes a local-level clipping method. The clipping sub-module is configured to: when the reference value reaches the clipping criterion corresponding to the local-level clipping method, clip the current-round gradient using the local-level clipping method.

[0194] In some embodiments, when the reference value reaches the clipping criterion corresponding to the local-level clipping method, clipping the current-round gradient using the local-level clipping method includes at least one of the following: when the reference value reaches the clipping criterion corresponding to the first type of local-level clipping method, clip the current-round gradient using the first type of local-level clipping method; when the reference value reaches the clipping criterion corresponding to the second type of local-level clipping method, clip the current-round gradient using the second type of local-level clipping method; when the reference value reaches the clipping criterion corresponding to the third type of local-level clipping method, clip the current-round gradient using the third type of local-level clipping method; wherein, the first type of local-level clipping method, the second type of local-level clipping method, and the third type of local-level clipping method are local-level clipping methods with different granularities.

[0195] In some embodiments, the target clipping method includes a global-level clipping method and at least one local-level clipping method. The clipping sub-module is configured to: when the reference value reaches the clipping criterion corresponding to the global-level clipping method, clip the current-round gradient using the global-level clipping method; when the reference value does not reach the clipping criterion corresponding to the global-level clipping method, clip the current-round gradient using at least one local-level clipping method.

[0196] In some embodiments, the target clipping method includes a global-level clipping method and at least one local-level clipping method. The clipping sub-module is configured to: when the reference value reaches the clipping criterion corresponding to the global-level clipping method, clip the current-round gradient using the global-level clipping method; when the reference value reaches the clipping criterion corresponding to at least one local-level clipping method, clip the current-round gradient using at least one local-level clipping method.

[0197] In some embodiments, the target clipping method includes a global-level clipping method and at least one local-level clipping method. The clipping sub-module is configured to: when the reference value reaches the clipping criterion corresponding to the global-level clipping method, clip the current-round gradient using the global-level clipping method; after clipping the current-round gradient using the global-level clipping method, if the new reference value determined based on the clipped current-round gradient reaches the clipping criterion corresponding to at least one local-level clipping method, continue to clip the current-round gradient using at least one local-level clipping method.

[0198] Those skilled in the art should understand that the functions of the processing modules in the model training device according to the embodiments of the present disclosure can be understood with reference to the relevant descriptions of the foregoing model training method. Each processing module in the model training device according to the embodiments of the present disclosure can be implemented by an analog circuit that implements the functions of the embodiments of the present disclosure, or can also be implemented by the operation of software that executes the functions of the embodiments of the present disclosure on an electronic device.

[0199] The model training device according to the embodiments of the present disclosure can effectively suppress the occurrence of training divergence and loss spikes during the model training process, and does not affect the learning speed, thereby effectively improving the stability of model training.

[0200] The embodiments of the present disclosure provide an image processing device, as Figure 10 shown. The image processing device includes: a first input module 1001, configured to input an image to be processed into a trained target model, where the trained target model is obtained by training according to the training method described above; an image processing module 1002, configured to perform at least one of image classification, image recognition, and image segmentation on the image to be processed according to the trained target model, to obtain an image processing result.

[0201] Those skilled in the art should understand that the functions of the processing modules in the image processing device according to the embodiments of the present disclosure can be understood with reference to the relevant descriptions of the foregoing image processing method. Each processing module in the image processing device according to the embodiments of the present disclosure can be implemented by an analog circuit that implements the functions described in the embodiments of the present disclosure, or can also be implemented by the operation of software that executes the functions described in the embodiments of the present disclosure on an electronic device.

[0202] The image processing device according to the embodiments of the present disclosure performs image processing using a target model obtained by adaptive gradient clipping training, and can improve the stability and accuracy of image processing.

[0203] The embodiments of the present disclosure provide a natural language processing device, as Figure 11 shown. The natural language processing device includes: a second input module 1101, configured to input a first type of data to be processed into a trained target model, where the trained target model is obtained by training according to the training method described above; a natural language processing module 1102, configured to perform at least one of information extraction, text classification, text recognition, speech recognition, and question answering on the first type of data to be processed according to the trained target model, to obtain a natural language processing result.

[0204] Those skilled in the art should understand that the functions of the various processing modules in the natural language processing device according to the embodiments of the present disclosure can be understood with reference to the relevant descriptions of the aforementioned natural language processing methods. The various processing modules in the natural language processing device according to the embodiments of the present disclosure can be implemented by an analog circuit that implements the functions described in the embodiments of the present disclosure, or can also be implemented by the operation of software that executes the functions described in the embodiments of the present disclosure on an electronic device.

[0205] The natural language processing device according to the embodiments of the present disclosure performs natural language processing using a target model obtained by adaptive gradient clipping training, which can improve the stability and accuracy of natural language processing.

[0206] The embodiments of the present disclosure provide a computer vision processing device, such as Figure 12 shown. The computer vision processing device includes: a third input module 1201 for inputting second type of data to be processed into a trained target model, where the trained target model is obtained by training according to the training method described above; a computer vision processing module 1202 for performing at least one computer vision process including image recognition, object detection, semantic segmentation, video understanding, and image generation on the second type of data to be processed according to the trained target model, to obtain a computer vision processing result.

[0207] Those skilled in the art should understand that the functions of the various processing modules in the computer vision processing device according to the embodiments of the present disclosure can be understood with reference to the relevant descriptions of the aforementioned computer vision processing methods. The various processing modules in the computer vision processing device according to the embodiments of the present disclosure can be implemented by an analog circuit that implements the functions described in the embodiments of the present disclosure, or can also be implemented by the operation of software that executes the functions described in the embodiments of the present disclosure on an electronic device.

[0208] The computer vision processing device according to the embodiments of the present disclosure performs computer vision processing using a target model obtained by adaptive gradient clipping training, which can improve the stability and accuracy of computer vision processing.

[0209] The embodiments of the present disclosure provide a schematic diagram of a model training scenario, such as Figure 13 shown.

[0210] As described above, the model training method provided by the embodiments of the present disclosure is applied to an electronic device. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices.

[0211] Specifically, the electronic device may specifically perform the following operations:

[0212] Deploy the model to be trained on a computing device;

[0213] Train the model to be trained deployed on the computing device to obtain a target model;

[0214] Among them, training the model to be trained deployed on the computing device includes:

[0215] Obtain the value representing the current round of gradient and the value representing the historical gradient;

[0216] Based on the value representing the historical gradient and the value representing the current round of gradient, obtain a reference value corresponding to the target pruning method;

[0217] When the reference value reaches the pruning standard corresponding to the target pruning method, use the target pruning method to prune the current round of gradient.

[0218] Among them, the value representing the current round of gradient and the value representing the historical gradient can be obtained from a data source. The data source can be various forms of data storage devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The data source can also represent various forms of mobile devices, such as, personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. In addition, the data source and the user terminal can be the same device.

[0219] It should be understood that Figure 13 The scene diagram shown is merely illustrative and not restrictive. Those skilled in the art can make various obvious changes and / or substitutions based on Figure 13 the examples, and the obtained technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.

[0220] Embodiments of the present disclosure provide a schematic diagram of a scene for image processing, as Figure 14 shown.

[0221] As mentioned above, the image processing method provided by the embodiments of the present disclosure is applied to an electronic device. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices.

[0222] Specifically, the electronic device may specifically perform the following operations:

[0223] Input the image to be processed into the trained target model, which is obtained by training according to the training method described above.

[0224] According to the trained target model, perform at least one image processing operation including image classification, image recognition, and image segmentation on the image to be processed, and obtain an image processing result.

[0225] Among them, the image to be processed can be obtained from an image data source. The image data source can be various forms of data storage devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The image data source can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. In addition, the image data source and the user terminal can be the same device.

[0226] It should be understood that Figure 14 the scene diagram shown is merely illustrative and not restrictive. Those skilled in the art can make various obvious changes and / or substitutions based on Figure 14 the examples, and the obtained technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.

[0227] The embodiments of the present disclosure provide a schematic diagram of a scene for natural language processing, as Figure 15 shown.

[0228] As mentioned above, the natural language processing method provided by the embodiments of the present disclosure is applied to an electronic device. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices.

[0229] Specifically, the electronic device can perform the following operations:

[0230] Input the first type of data to be processed into the trained target model, which is obtained by training according to the training method described above.

[0231] According to the trained target model, perform at least one natural language processing operation including information extraction, text classification, text recognition, speech recognition, and question answering on the first type of data to be processed, and obtain a natural language processing result.

[0232] Among them, the first type of data to be processed can be obtained from a data source. The data source can be various forms of data storage devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The data source can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. In addition, the data source and the user terminal can be the same device.

[0233] It should be understood that Figure 15 the illustrated scenario diagram is merely illustrative rather than restrictive, and those skilled in the art can make various obvious changes and / or substitutions based on Figure 15 the examples, and the obtained technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.

[0234] The embodiments of the present disclosure provide a schematic diagram of a scenario for computer vision processing, as Figure 16 shown.

[0235] As mentioned above, the computer vision processing method provided by the embodiments of the present disclosure is applied to an electronic device. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices.

[0236] Specifically, the electronic device can specifically perform the following operations:

[0237] Input the second type of data to be processed into the trained target model, and the trained target model is obtained according to the training method described above;

[0238] According to the trained target model, perform at least one computer vision processing including picture recognition, object detection, semantic segmentation, video understanding, and picture generation on the second type of data to be processed, and obtain a computer vision processing result.

[0239] Among them, the second type of data to be processed can be obtained from a data source. The data source can be various forms of data storage devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The data source can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. In addition, the data source and the user terminal can be the same device.

[0240] It should be understood that Figure 16The scene diagram shown is merely illustrative and not restrictive. Those skilled in the art can make various obvious changes and / or substitutions based on Figure 16 the examples, and the resulting technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.

[0241] In the technical solutions of the present disclosure, the acquisition, storage, and application of user personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0242] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0243] Figure 17 A schematic block diagram of an exemplary electronic device 1700 that can be used to implement the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital assistant, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0244] As Figure 17 shown, the device 1700 includes a computing unit 1701, which can execute various appropriate actions and processes according to the computer program stored in a read-only memory (ROM) 1702 or the computer program loaded from a storage unit 1708 into a random access memory (RAM) 1703. In the RAM 1703, various programs and data required for the operation of the device 1700 can also be stored. The computing unit 1701, the ROM 1702, and the RAM 1703 are connected to each other via a bus 1704. An input / output (I / O) interface 1705 is also connected to the bus 1704.

[0245] Multiple components in the device 1700 are connected to the I / O interface 1705, including: an input unit 1706, such as a keyboard, a mouse, etc.; an output unit 1707, such as various types of displays, speakers, etc.; a storage unit 1708, such as a magnetic disk, an optical disc, etc.; and a communication unit 1709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1709 allows the device 1700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0246] The computing unit 1701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1701 executes the various methods and processes described above, such as model training methods and / or image processing methods and / or natural language processing methods and / or computer vision processing methods. For example, in some embodiments, the model training method and / or image processing method and / or natural language processing method and / or computer vision processing method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1700 via the ROM 1702 and / or the communication unit 1709. When the computer program is loaded into the RAM 1703 and executed by the computing unit 1701, one or more steps of the model training method and / or image processing method and / or natural language processing method and / or computer vision processing method described above can be executed. Alternatively, in other embodiments, the computing unit 1701 can be configured to execute the model training method and / or image processing method and / or natural language processing method and / or computer vision processing method in any other suitable way (e.g., by means of firmware).

[0247] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0248] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0249] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0250] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0251] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0252] A computer system may include a client and a server. The client and the server are generally far away from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0253] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.

[0254] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A model training method, comprising: Deploying a model to be trained on a computing device; The model to be trained is at least used for image processing; Training the model to be trained deployed on the computing device to obtain a target model; Wherein, the training of the model to be trained deployed on the computing device includes: Obtaining a value representing the current round of gradient and a value representing the historical gradient; Based on the value representing the historical gradient and the value representing the current round of gradient, obtaining a reference value corresponding to the target pruning method; When the reference value reaches the pruning standard corresponding to the target pruning method, pruning the current round of gradient using the target pruning method; Wherein, the pruning of the current round of gradient includes: Determining the pruning factor of the current round; When it is determined that the current round of gradient needs to be pruned, replacing the current round of gradient with the product of the pruning factor of the current round and the current round of gradient; Wherein, the pruning of the current round of gradient using the target pruning method includes: When the target pruning method is a global-level pruning method, determining the pruning factor of the current round according to the first formula, where the first formula is: , where, , is a preset value, is the historical gradient, g is the gradient of this round, represents the preset adjustment factor, represents the clipping factor of this round.

2. The method according to claim 1, further comprising: Obtaining granularity indication information, where the granularity indication information is used to indicate the granularity adopted in this training; Determining the target pruning method based on the granularity indication information.

3. The method according to claim 1, wherein The pruning of the current round of gradient using the target pruning method includes: When the target pruning method is a tensor-level pruning method, determining the pruning factor of the current round according to the first formula, where the first formula is: , where , is a preset value, is the historical gradient, g is the gradient of this round, represents the preset adjustment factor, represents the clipping factor of this round.

4. The method according to claim 1, wherein The pruning of the current round of gradient using the target pruning method includes: When the target pruning method is a sliced tensor-level pruning method, determining the pruning factor of the current round according to the first formula, where the first formula is: , where , is a preset value, is the historical gradient, g is the gradient of this round, represents the preset adjustment factor, represents the clipping factor of this round.

5. The method according to claim 1, wherein The pruning of the current round of gradient using the target pruning method includes: When the target pruning method is an element-wise pruning method, determining the pruning factor of the current round according to the second formula, where the second formula is: , where , is a preset value, is the historical gradient, is the current round gradient, represents the preset adjustment factor, represents the clipping factor for this round.

6. The method according to claim 1, wherein The target pruning method includes a global-level pruning method, and when the reference value reaches the pruning standard corresponding to the target pruning method, pruning the current round of gradient using the target pruning method includes: When the reference value reaches the pruning standard corresponding to the global-level pruning method, pruning the current round of gradient using the global-level pruning method.

7. The method according to claim 1, wherein The target pruning method includes a local-level pruning method, and when the reference value reaches the pruning standard corresponding to the target pruning method, pruning the current round of gradient using the target pruning method includes: When the reference value reaches the pruning standard corresponding to the local-level pruning method, pruning the current round of gradient using the local-level pruning method.

8. The method according to claim 7, wherein When the reference value reaches the pruning standard corresponding to the local-level pruning method, pruning the current round of gradient using the local-level pruning method includes at least one of the following: When the reference value reaches the clipping criterion corresponding to the first type of local-level clipping method, use the first type of local-level clipping method to clip the current-round gradient; When the reference value reaches the clipping criterion corresponding to the second type of local-level clipping method, use the second type of local-level clipping method to clip the current-round gradient; When the reference value reaches the clipping criterion corresponding to the third type of local-level clipping method, use the third type of local-level clipping method to clip the current-round gradient; Among them, the first type of local-level clipping method, the second type of local-level clipping method, and the third type of local-level clipping method are local-level clipping methods with different granularities.

9. The method according to claim 1, wherein The target clipping method includes a global-level clipping method and at least one local-level clipping method. When the reference value reaches the clipping criterion corresponding to the target clipping method, using the target clipping method to clip the current-round gradient includes: When the reference value reaches the clipping criterion corresponding to the global-level clipping method, use the global-level clipping method to clip the current-round gradient; When the reference value does not reach the clipping criterion corresponding to the global-level clipping method, use the at least one local-level clipping method to clip the current-round gradient.

10. The method according to claim 1, wherein, The target clipping method includes a global-level clipping method and at least one local-level clipping method. When the reference value reaches the clipping criterion corresponding to the target clipping method, using the target clipping method to clip the current-round gradient includes: When the reference value reaches the clipping criterion corresponding to the global-level clipping method, use the global-level clipping method to clip the current-round gradient; When the reference value reaches the clipping criterion corresponding to the at least one local-level clipping method, use the at least one local-level clipping method to clip the current-round gradient.

11. The method according to claim 1, wherein, The target clipping method includes a global-level clipping method and at least one local-level clipping method. When the reference value reaches the clipping criterion corresponding to the target clipping method, using the target clipping method to clip the current-round gradient includes: When the reference value reaches the clipping criterion corresponding to the global-level clipping method, use the global-level clipping method to clip the current-round gradient; After using the global-level clipping method to clip the current-round gradient, if the new reference value determined based on the clipped current-round gradient reaches the clipping criterion corresponding to the at least one local-level clipping method, then use the at least one local-level clipping method to continue to clip the current-round gradient.

12. The method according to claim 1, wherein, The model to be trained is also used for natural language processing.

13. The method according to claim 1, wherein, The model to be trained is also used for computer vision processing.

14. A model training device, comprising: A deployment module, configured to deploy a model to be trained on a computing device; The model to be trained is at least used for image processing; A training module, configured to train the model to be trained deployed on the computing device to obtain a target model; Among them, the training module includes: The first acquisition sub-module is configured to acquire the value representing the current-round gradient and the value representing the historical gradient; The second acquisition sub-module is configured to obtain a reference value corresponding to the target clipping method based on the value representing the historical gradient and the value representing the current-round gradient; The clipping sub-module is configured to clip the current-round gradient using the target clipping method when the reference value meets the clipping criterion corresponding to the target clipping method; Wherein, the clipping sub-module is configured to: Determine the clipping factor for the current round; When it is determined that the current-round gradient needs to be clipped, replace the current-round gradient with the product of the clipping factor for the current round and the current-round gradient; Wherein, the clipping sub-module is configured to: When the target clipping method is a global-level clipping method, determine the clipping factor for the current round according to the first formula, where the first formula is: , where , is a preset value, is the historical gradient, g is the gradient of this round, represents the preset adjustment factor, represents the clipping factor of this round.

15. The apparatus according to claim 14, further comprising: An acquisition module, configured to acquire granularity indication information, where the granularity indication information is used to indicate the granularity adopted in the current training; A determination module, configured to determine the target clipping method based on the granularity indication information.

16. The device according to claim 14, wherein, The clipping sub-module is configured to: When the target clipping method is a tensor-level clipping method, determine the clipping factor for the current round according to the first formula, where the first formula is: , where , is a preset value, is the historical gradient, g is the gradient of this round, represents the preset adjustment factor, represents the clipping factor of this round.

17. The apparatus according to claim 14, wherein, The clipping sub-module is configured to: When the target clipping method is a split-tensor-level clipping method, determine the clipping factor for the current round according to the first formula, where the first formula is: , where , is a preset value, is the historical gradient, g is the gradient of this round, represents the preset adjustment factor, represents the clipping factor of this round.

18. The device according to claim 14, wherein, The clipping sub-module is configured to: When the target clipping method is an element-wise-level clipping method, determine the clipping factor for the current round according to the second formula, where the second formula is: , where, , is a preset value, is the historical gradient, is the current round gradient, represents the preset adjustment factor, represents the clipping factor for this round.

19. The apparatus according to claim 14, wherein, The target clipping method includes a global-level clipping method, and the clipping sub-module is configured to: When the reference value meets the clipping criterion corresponding to the global-level clipping method, clip the current-round gradient using the global-level clipping method.

20. The apparatus according to claim 14, wherein The target clipping method includes a local-level clipping method, and the clipping sub-module is configured to: When the reference value meets the clipping criterion corresponding to the local-level clipping method, clip the current-round gradient using the local-level clipping method.

21. The apparatus according to claim 20, wherein, The step of clipping the current-round gradient using the local-level clipping method when the reference value meets the clipping criterion corresponding to the local-level clipping method includes at least one of the following: When the reference value meets the clipping criterion corresponding to the first type of local-level clipping method, clip the current-round gradient using the first type of local-level clipping method; When the reference value meets the clipping criterion corresponding to the second type of local-level clipping method, clip the current-round gradient using the second type of local-level clipping method; When the reference value meets the clipping criterion corresponding to the third type of local-level clipping method, clip the current-round gradient using the third type of local-level clipping method; Wherein, the first type of local-level clipping method, the second type of local-level clipping method, and the third type of local-level clipping method are local-level clipping methods with different granularities.

22. The apparatus according to claim 14, wherein, The target clipping method includes a global-level clipping method and at least one local-level clipping method. The clipping sub-module is configured to: When the reference value reaches the clipping criterion corresponding to the global-level clipping method, clip the current round of gradients using the global-level clipping method; When the reference value does not reach the clipping criterion corresponding to the global-level clipping method, clip the current round of gradients using the at least one local-level clipping method.

23. The apparatus according to claim 14, wherein The target clipping method includes a global-level clipping method and at least one local-level clipping method. The clipping sub-module is configured to: When the reference value reaches the clipping criterion corresponding to the global-level clipping method, clip the current round of gradients using the global-level clipping method; When the reference value reaches the clipping criterion corresponding to the at least one local-level clipping method, clip the current round of gradients using the at least one local-level clipping method.

24. The device according to claim 14, wherein, The target clipping method includes a global-level clipping method and at least one local-level clipping method. The clipping sub-module is configured to: When the reference value reaches the clipping criterion corresponding to the global-level clipping method, clip the current round of gradients using the global-level clipping method; After clipping the current round of gradients using the global-level clipping method, if the new reference value determined based on the clipped current round of gradients reaches the clipping criterion corresponding to the at least one local-level clipping method, continue to clip the current round of gradients using the at least one local-level clipping method.

25. The apparatus according to claim 14, wherein, The to-be-trained model is also used for natural language processing.

26. The apparatus according to claim 14, wherein The to-be-trained model is also used for computer vision processing.

27. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-13.

28. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-13.

29. A computer program product, comprising a computer program stored on a storage medium, where the computer program, when executed by a processor, implements the method according to any one of claims 1-13.

Citation Information

Patent Citations

  • Federal learning model training privacy protection method and system based on hybrid strategy

    CN116167084A