Training Method, Device, Electronic Device and Storage Medium of Neural Network

By generating fixed-directional gradient noise in neural network training and performing gradient fusion, the problem of inefficient neural network training is solved, and more efficient training process and cost compression are achieved.

CN119761449BActive Publication Date: 2025-07-11ZHEJIANG LAB
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510275403.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-07-11
Estimated Expiration
2045-03-10

AI Technical Summary

Technical Problem

The existing neural network training methods are inefficient and have high training costs, and are prone to falling into local minimum values, resulting in gradient disappearance or explosion, making it difficult to quickly and accurately implement functions.

Method used

By generating directional gradient noise in a fixed direction throughout the training process, and fusing the gradient noise of the connection weight that meets the gradient threshold requirements with the first gradient, a fusion gradient is obtained, which is used to update the target connection weight, avoid local minimum values, and improve training efficiency.

Benefits of technology

It improves the efficiency of neural network training, reduces training time and cost, realizes effective search for larger parameter space, and avoids the problem of gradient disappearance or explosion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119761449B_ABST
    Figure CN119761449B_ABST
Patent Text Reader

Abstract

The present application discloses a training method for a neural network, a training device for a neural network, an electronic device, and a computer-readable storage medium. The training method for the neural network is at least applied to a visual processing scenario and includes: inputting a training sample into the neural network to be trained to obtain an output result of the neural network under each current connection weight; the training sample at least includes visual data; determining a first gradient of each connection weight based on the output result and the annotation information of the training sample; generating directional gradient noise for each connection weight respectively; for a target connection weight that satisfies the relationship requirement between the first gradient and a set gradient threshold, fusing the directional gradient noise corresponding to the target connection weight and the first gradient to obtain a fused gradient, and updating the target connection weight according to the fused gradient; repeating the training until the obtained neural network meets the set requirements. It can improve the training efficiency of the neural network, reduce the training time, and compress the training cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of neural network training, and particularly to a neural network training method, a neural network training device, an electronic device, and a computer-readable storage medium. Background Art

[0002] With the development of the field of artificial intelligence and the emergence of large models, the scale of neural networks has increased exponentially, and the learning data has grown rapidly. The huge amount of computation has led to a rapid increase in the training cost of neural networks. The training cost of current mainstream large models often reaches tens of millions or even hundreds of millions of US dollars, resulting in huge economic cost consumption. How to develop an effective neural network training method that can quickly and accurately implement functions, improve training efficiency, and reduce training costs is an urgent problem to be solved in the current field of artificial intelligence.

[0003] The training method of neural networks is the key to determining the efficiency and performance of neural networks. However, different from the rapid development of the scale of neural networks, the development of neural network training methods is relatively slow. The current mainstream neural network training methods are still the traditional backpropagation algorithm and gradient descent algorithm. The current neural network training method aims to minimize the loss function between the output result of the neural network and the annotation of the training samples. Through the backpropagation algorithm, the gradient of the loss function is calculated, and the gradient descent algorithm is used to iteratively search along the opposite direction of the corresponding gradient of the neural network parameters, so that the neural network evolves to the local minimum of the loss function to realize complex functions such as network prediction, classification, and generation.

[0004] However, the existing gradient descent algorithm can only search and update in a relatively small effective parameter space, resulting in low efficiency of neural network parameter update and prone to problems such as saddle points of the loss function, leading to gradient disappearance or gradient explosion, resulting in low neural network training efficiency, long training time, and high cost. Summary of the Invention

[0005] In view of the above technical problems, this application provides a neural network training method, a neural network training device, an electronic device, and a computer-readable storage medium. The technical solutions are as follows:

[0006] According to the first aspect of this application, a neural network training method is provided, which is at least applied to a visual processing scenario. The method includes:

[0007] Input a training sample into the neural network to be trained to obtain the output result of the neural network under each current connection weight; the training sample includes at least visual data;

[0008] Based on the output result and the annotation information of the training sample, determine the first gradient of each connection weight;

[0009] Generate directional gradient noise for each of the connection weights; wherein the initial direction of the directional gradient noise is randomly generated and remains fixed during the training process;

[0010] For a target connection weight that satisfies the relationship requirement between the first gradient and a set gradient threshold, fuse the directional gradient noise corresponding to the target connection weight and the first gradient to obtain a fused gradient, and update the target connection weight according to the fused gradient;

[0011] Repeat the training until the obtained neural network meets the set requirements.

[0012] According to a second aspect of the present application, there is provided a training device for a neural network, which is at least applied to a visual processing scenario. The device includes:

[0013] An input unit, configured to input a training sample into a neural network to be trained, and obtain an output result of the neural network under each current connection weight; the training sample includes at least visual data;

[0014] A determination unit, configured to determine the first gradient of each connection weight based on the output result and the annotation information of the training sample;

[0015] A generation unit, configured to generate directional gradient noise for each connection weight respectively; wherein the initial direction of the directional gradient noise is randomly generated and remains fixed during the training process;

[0016] An update unit, configured to, for a target connection weight that satisfies the relationship requirement between the first gradient and a set gradient threshold, fuse the directional gradient noise corresponding to the target connection weight and the first gradient to obtain a fused gradient, and update the target connection weight according to the fused gradient;

[0017] A training unit, configured to repeat the training until the obtained neural network meets the set requirements.

[0018] According to a third aspect of the present application, there is provided an electronic device, which includes:

[0019] A processor;

[0020] A memory for storing instructions executable by the processor;

[0021] Wherein, the processor is configured to implement the method as described in the first aspect.

[0022] According to a fourth aspect of the present application, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the method as described in the first aspect are implemented.

[0023] Taking the visual processing scenario as an example, the technical solution provided by this application uses the training samples and the annotation information of the training samples as a training sample set. The training samples may include visual data. The training samples are input into the neural network to be trained, and the output results of the neural network under each current connection weight are obtained. Based on the output results and the annotation information of the training samples input into the neural network, the first gradient of each connection weight of the neural network is determined according to the backpropagation algorithm. Directional gradient noise is generated for each connection weight of the neural network. Among them, the initial direction of any directional gradient noise is randomly generated and remains fixed in any subsequent round of training and will not change. For the target connection weights that meet the relationship requirements between the first gradient and the set gradient threshold, the directional gradient noise corresponding to the target connection weights and the first gradient are fused to obtain a fused gradient, and the fused gradient is used as the new gradient of the target connection weights. And according to the fused gradient, the gradient descent algorithm is used to update the target connection weights, and the above training samples and the annotation information of the training samples are used to repeat the training until the neural network meets the set requirements.

[0024] It can be seen that when training a neural network in a gradient descent manner, by generating directional gradient noise with fixed positive and negative signs (i.e., positive and negative directions) throughout the training process for each connection weight of the neural network, and fusing the directional gradient noise corresponding to the target connection weights that meet the relationship requirements between the first gradient and the set gradient threshold and the first gradient to obtain a fused gradient, and using the fused gradient as the new gradient to update the target connection weights. The directivity of the gradient noise, that is, only generating the initial direction of the gradient noise once and reusing it in subsequent training, enables the training process to perform a more effective search in the parameter space, continuously search in a larger effective parameter space, and at the same time avoid the neural network falling into local minima, thereby improving the training efficiency of the neural network, reducing the training time, and compressing the training cost.

[0025] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this application. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments recorded in this application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0027] Figure 1 It is a schematic diagram of a neural network training scenario in related technologies;

[0028] Figure 2 It is a schematic flowchart of a method for training a neural network according to an embodiment of the present application;

[0029] Figure 3 It is a schematic diagram of a neural network training scenario according to an embodiment of the present application;

[0030] Figure 4 It is a schematic structural diagram of a training device for a neural network according to an embodiment of the present application;

[0031] Figure 5 It is a schematic structural diagram of an electronic device according to an embodiment of the present application. Detailed implementation manners

[0032] In order to enable those skilled in the art to better understand the technical solutions in the present application, the following will describe the technical solutions in the embodiments of the present application in detail with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art shall fall within the protection scope of the present application.

[0033] With the development of the field of artificial intelligence and the emergence of large models, the scale of neural networks has increased exponentially, and the learning data has grown rapidly. The huge amount of computation has led to a rapid increase in the training cost of neural networks. The training cost of current mainstream large models often reaches tens of millions or even hundreds of millions of US dollars, resulting in huge economic cost consumption. How to develop an effective neural network training method that can quickly and accurately achieve functions, improve training efficiency, and reduce training costs is an urgent problem to be solved in the current field of artificial intelligence.

[0034] In related technologies, the training method of neural networks is the key to determining the efficiency and performance of neural networks. However, different from the rapidly developing scale of neural networks, the development of neural network training methods in related technologies is relatively slow. The current mainstream neural network training methods are still traditional backpropagation algorithms and gradient descent algorithms. For example, Figure 1 In the neural network training scenario shown, the current neural network training process can be generally summarized as: initializing network parameters → inputting training data → forward propagation → calculating the loss function → backpropagation → (using the gradient descent method) updating network parameters → repeating training → training completed. It can be seen that the current neural network training method aims to minimize the loss function between the output result of the neural network and the annotation of the training samples. Through the backpropagation algorithm, the gradient of the loss function is calculated, and using the gradient descent algorithm, the parameters of the neural network are iteratively searched along the opposite direction of the corresponding gradient, so that the neural network evolves to the local minimum of the loss function to achieve complex functions such as network prediction, classification, and generation.

[0035] However, the current gradient descent algorithm can only search and update in a relatively small effective parameter space, resulting in low efficiency of parameter update in the neural network and being prone to problems such as saddle points of the loss function, which may lead to gradient vanishing or gradient explosion, thus making the neural network training inefficient, with long training time and high cost.

[0036] It should be noted that the above introduction to the neural network training scenario in the related art is only an exemplary display. In actual applications, there may be other neural network training scenarios, which are not specifically limited herein.

[0037] To address the above problems, the present application provides a training method for a neural network, which is at least applied to a visual processing scenario and can improve the training efficiency of the neural network, reduce the training time, and compress the training cost. As Figure 2 shown, the method may include the following steps:

[0038] S201. Input the training samples into the neural network to be trained, and obtain the output results of the neural network under each current connection weight.

[0039] The training samples include at least visual data.

[0040] S202. Based on the output results and the annotation information of the training samples, determine the first gradient of each connection weight.

[0041] S203. Generate directional gradient noise for each connection weight respectively.

[0042] Among them, the initial direction of the directional gradient noise is randomly generated and remains fixed during the training process.

[0043] S204. For the target connection weights that meet the relationship requirements between the first gradient and the set gradient threshold, fuse the directional gradient noise corresponding to the target connection weights and the first gradient to obtain a fused gradient, and update the target connection weights according to the fused gradient.

[0044] S205. Repeat the training until the obtained neural network meets the set requirements.

[0045] When training a neural network in a gradient descent manner, the technical solution provided by the embodiments of the present application generates, for each connection weight of the neural network, a directional gradient noise with fixed positive and negative signs (i.e., positive and negative directions) throughout the training process, and fuses the directional gradient noise corresponding to the target connection weight that satisfies the relationship requirement between the first gradient and the set gradient threshold and the first gradient to obtain a fused gradient, and uses the fused gradient as the new gradient to update the target connection weight. The directionality of the gradient noise, that is, only generating the initial direction of the gradient noise once and reusing it in subsequent training, enables the training process to perform a more effective search in the parameter space, continuously search in a larger effective parameter space, and at the same time avoid the neural network falling into local minima, thereby improving the training efficiency of the neural network, reducing the training time, and compressing the training cost.

[0046] It can be understood that the order of S202 and S203 is not limited.

[0047] There can be various specific implementations for the execution subject of the neural network training method provided by the embodiments of the present application. As an example, the execution subject can be a server or an electronic device such as a desktop computer, a laptop computer, or a mobile terminal. For the sake of description, in any of the following embodiments, the server is used as the execution subject as an example to illustrate the neural network training method provided by the embodiments of the present application. Among them, the server can be a single device or include multiple devices. For example, a distributed server, and the present application does not specifically limit this.

[0048] There can be various specific implementations for the training samples input to the neural network. As an example, taking the visual processing scenario as an example, the training samples can include visual data. As another example, the visual data can be image data or other data. In other scenarios, the training samples can also be other types of data, such as text data in the text processing scenario. Therefore, the specific implementation of the training samples is not limited.

[0049] There can be various specific implementations for the above-mentioned set requirements that the neural network needs to meet. As an example, the neural network meeting the set requirements can mean that the neural network can complete a specified task. As another example, the neural network meeting the set requirements can mean that the number of repeated training times of the neural network has reached the number threshold. Therefore, the specific implementation of the above-mentioned set requirements that the neural network needs to meet is not limited.

[0050] As an example, when the above setting requirement means that the neural network can complete a specified task, the server can obtain the training samples related to the requirements of the specified task that the neural network needs to complete and the annotation information of the training samples. Taking the type of the specified task as any machine learning task applying supervised learning or self-supervised learning training process as an example, the specified task can be any task such as visual scene processing, including classification tasks, recognition tasks, generation tasks, etc. Therefore, the type of the specified task is not specifically limited.

[0051] For ease of description, the above specified task will be described exemplarily as an image classification task below. As an example, when the above specified task is an image classification task, the training samples input into the neural network to be trained can be images to be classified. As another example, each image to be classified can be preprocessed, and the preprocessing methods can include cropping, scaling, and / or rotation, etc. After the image to be classified is preprocessed, it is converted into a one-dimensional vector form and then input into the neural network to be trained. As another example, the annotation information of the training samples can be the category information of the images to be classified, where each image to be classified can correspond to one category. Under the training data set composed of the images to be classified and the annotation information of the images to be classified, after the server completes the training of the neural network, the recognition and classification of images can be realized. Taking a specific training scenario as an example below, the server can perform image classification on the surrounding environment images of the unmanned device to determine whether there are obstacles in the environment images. The corresponding training samples can be the environment images collected by the image acquisition device historically, and the annotation information of the training samples can be the information on whether the environment images contain obstacles, that is, the annotation information is two image categories: there are obstacles and there are no obstacles.

[0052] As an example, before inputting the training samples into the neural network to be trained, the neural network can be initialized. Specifically, the connection weights of the neural network can be initially set. That is, before the training samples are input into the neural network, the initially set connection weights are the current connection weights of the neural network.

[0053] As an example, the method for initializing the neural network can be determined based on the architecture of the neural network. Among them, the architecture of the neural network can be any network architecture applying supervised learning or self-supervised learning training process. For example, the network architecture can be a deep neural network, a convolutional neural network, or a Transformer architecture, or other types of architectures. Therefore, the type of the neural network architecture is not specifically limited.

[0054] For ease of description, the following takes the architecture of a neural network as a multi-layer feedforward neural network as an example to give an exemplary introduction to a way of initializing a neural network. As an example, the server can build a multi-layer feedforward neural network containing L layers of neurons. As another example, the first layer of the multi-layer feedforward neural network containing L layers of neurons can be the input layer, the L-th layer is the output layer, and the other layers can be hidden layers. The number of neurons in the input layer can be the dimension of the input data , and the number of neurons in the output layer can be the number of image categories , and the number of neurons in the hidden layer, that is, the width of the neural network, can be N. Neurons between adjacent layers are fully connected. Specifically, the connection weight from the l-th layer to the l+1-th layer is defined as , where are the number of neurons in the l-th layer and the l+1-th layer respectively. It should be noted that the server can set any connection weight initialization that meets the network architecture requirements, and the embodiments of the present application do not limit the initialization method of the neural network

[0055] As an example, the trainability of the network can be used as a reference standard to specifically introduce an initialization method with adjustable trainability based on a multi-layer feedforward neural network (adjustable trainability means that the trainability of the neural network can be adjusted through initialization settings). Of course, this method does not constitute an improper limitation to the embodiments of the present application. As another example, the server can adjust the trainability of the neural network initialization by constructing a connection weight perturbation term and initializing the connection weight of the neural network as the sum of the identity matrix and this connection weight perturbation term. As another example, the server can select two forms of Bernoulli perturbation term and Gaussian perturbation term as the connection weight perturbation term. The connection weights under the two connection weight perturbation terms can be respectively expressed as:

[0056] (1)

[0057] (2)

[0058] where is the Bernoulli perturbation term, and its diagonal element , and the non-diagonal elements follow a Bernoulli distribution with probability p , that is, the probability that the connection weight term takes 1 is p, and the probability that it takes 0 is 1-p is the Gaussian perturbation term, which follows a Gaussian distribution with a mean of 0 and a standard deviation of , and N is the width of the neural network. The server can adjust the trainability of the neural network initialization by adjusting the perturbation intensity parameters, that is, the Bernoulli distribution probability p and the Gaussian distribution standard deviation parameter . The larger the perturbation intensity parameter, the weaker the trainability of the neural network initialization .

[0059] As an example, the output result of the neural network under each current connection weight can be compared with the annotation information of the training samples input to the neural network, and the first gradient of each connection weight can be determined according to the backpropagation algorithm. It can be understood that the first gradient of the connection weight can refer to the gradient of the connection weight when the directional gradient noise is not fused in this round of training. If the first gradient fuses the directional gradient noise, the gradient after the two are fused is the new gradient of the connection weight.

[0060] There can be various specific ways to determine the above first gradient. As an example, the server can convert the training samples with the above annotation information input to the neural network into target vectors, and convert the output results of the neural network under each current connection weight into vector forms, and based on the target vectors and the output results converted into vector forms, determine the loss function of the neural network; based on this loss function, according to the backpropagation algorithm, determine the above first gradient. As another example, taking the above specified task as an image classification task, the format of the target vector is a one-hot vector indicating the image category, that is, the category item corresponding to the training image is 1, and the remaining category items are 0. As another example, the server can determine the loss function by comparing the target vector with the above output result to describe the difference between the target vector and the output result. The server can flexibly set the form of the loss function. For example, a loss function including higher-order terms or other regularization terms is not specifically limited in this regard.

[0061] As another example, the loss function can be expressed as follows:

[0062] (3)

[0063] Among them, represents the parameters of the neural network. Specifically, can be the connection weights and biases of the neural network, is the loss function under the parameter The server can calculate the loss function by comparing the output of the neural network under the parameter with the target vector under the input of the neural network.

[0064] To complete the training task of the neural network, the server needs to minimize the loss function as much as possible. The general gradient descent algorithm can search iteratively along the opposite direction of the gradient of the corresponding loss function for the parameters of the neural network, so that the neural network evolves to the local minimum of the loss function to obtain a better task effect. As an example, the embodiments of the present application can optimize the training method based on the general gradient descent algorithm, so it is also necessary to calculate the gradient of the loss function in the parameter space. As another example, since the loss function is a function of all the parameters of the neural network, the server can determine the first gradient corresponding to each connection weight of the neural network and the absolute value of the first gradient according to the backpropagation algorithm.

[0065] The following takes a multi-layer feedforward neural network as an example for exemplary introduction:

[0066] (4)

[0067] Among them, the connection weight term from the l-th layer neurons to the l+1-th layer neurons of the neural network The corresponding gradient Is the partial derivative of the loss function With respect to it. Among them, Are the parameters of the neural network and can be regarded as the set of all .

[0068] As another example, in order to measure the training situation of the neural network at multiple scales, when the neural network includes multiple layers, the server can also calculate the mean value of the absolute values of the first gradients of each connection weight within each layer in the neural network , and the mean value of the absolute values of the first gradients of each connection weight in the neural network (that is, all connection weight terms of the neural network) :

[0069] (5)

[0070] (6)

[0071] Among them, Represents the mean value of the absolute values of the first gradients corresponding to all connection weights from the l-th layer to the l+1-th layer of the neural network, Represents the mean value of the absolute values of the first gradients of each connection weight in the neural network (that is, all connection weight terms of the neural network). The above three absolute values related to the first gradient (4), (5), (6): the absolute value of the first gradient corresponding to each connection weight of the neural network , the mean value of the absolute values of the first gradients of each connection weight within each layer in the neural network , the mean of the absolute values of the first gradients of the connection weights in the neural network (i.e., all connection weight terms in the entire neural network) , which can respectively correspond to the three different gradient judgment and fusion methods provided in the embodiments of the present application.

[0072] To increase the search range of the parameter space and avoid the ineffective search of the traditional gradient descent algorithm in multi-parameter dimensions, as an example, the server can generate directional gradient noise with fixed positive and negative signs to correct the first gradient calculated by the backpropagation algorithm in S202. The server can randomly generate directional gradient noise for each connection weight according to a specified probability. In subsequent training, the sign (i.e., the initial direction) of the generated directional gradient noise is fixed, while the intensity can be flexibly adjusted. Randomly generate the initial direction of any directional gradient noise, that is, randomly generate its positive or negative sign. Once the positive and negative signs are generated, they are fixed and used in subsequent training without change.

[0073] The following is an exemplary introduction to a specific process of generating directional gradient noise:

[0074] As an example, to generate the initial direction of the directional gradient noise, the server will randomly generate the direction of the directional gradient noise for each connection weight according to a specified probability.

[0075] The above-specified probability can have various specific implementations. As an example, the specified probability can be implemented as follows:

[0076] (7)

[0077] (8)

[0078] Among them, represents the gradient noise corresponding to the connection weight term , sgn( ) represents its positive and negative signs (1, -1), that is, 50% probability for the positive direction and 50% probability for the negative direction. Of course, the specific value of the specified probability is not limited to this.

[0079] As another example, to generate the intensity of the directional gradient noise, the server can generate the intensity of the directional gradient noise assigned as positive according to a specified method.

[0080] The above-specified method can have multiple specific implementations. As an example, the specified method can be a randomly generated method. As another example, the randomly generated method can refer to randomly generating according to a Gaussian distribution, where the mean of the Gaussian distribution is a positive value and its standard deviation is much smaller than its mean; it can also be randomly generated according to other types of distributions. As another example, the specified method can also be a method based on a preset value. Therefore, the specific implementation of the above-specified method is not limited.

[0081] The frequency of generating the intensity of the directional gradient noise can have multiple specific implementations. As an example, the intensity of the directional gradient noise can also be fixed during the training process after generation, just like the initial direction of the directional gradient noise, without further change. As another example, during multiple rounds of training, the intensity of the directional gradient noise can also be generated multiple times. For example, the intensity of the directional gradient noise can be regenerated once in each round of training. Therefore, the specific implementation of the frequency of generating the intensity of the directional gradient noise is not limited.

[0082] In addition, the intensity of generating the directional gradient noise can be flexibly adjusted in multiple ways during training. As an example, the intensity of the directional noise can be adjusted according to the training stage. As another example, the intensity of the directional noise can also be adjusted according to the magnitude of the loss function value. Therefore, the specific implementation of the implementation method of adjusting the intensity of generating the directional gradient noise is not limited.

[0083] The following takes the above-specified method of randomly generating according to a Gaussian distribution as an example for an exemplary introduction: Taking a Gaussian distribution with a positive mean and a standard deviation much smaller than the mean as an example, the server can randomly generate the magnitude of the intensity of the directional gradient noise in the following way:

[0084] (9)

[0085] Among them, represents the intensity of the directional gradient noise, , are the mean and standard deviation of the Gaussian distribution respectively. Specifically, , needs to satisfy and to ensure that the intensity of the directional gradient noise generated by the server is positive. When , the probability that the server generates a negative value is extremely low and can be ignored. Of course, the generation method of the intensity of the directional gradient noise is not limited to this.

[0086] As another example, the directional gradient noise can be obtained by multiplying the initial direction of the directional gradient noise by the intensity of the directional gradient noise:

[0087] (10)

[0088] Among them, is the connection weight term corresponding to the directional gradient noise of the initial direction, indicating the connection weight term corresponding to the directional gradient noise of the intensity.

[0089] It should be noted that the above introduction to the generation method of the intensity of the directional gradient noise is only an exemplary display. In actual applications, there may be other generation methods. The server can generate the intensity of the directional gradient noise with a positive value in any other way, and can also flexibly select the generation frequency of the intensity of the directional gradient noise or flexibly adjust the intensity of the directional gradient noise according to factors such as the progress of training, such as the progress of training or the size of the loss function. Therefore, the specific generation and update methods of the intensity of the directional gradient noise are not specifically limited. However, the initial direction of the directional gradient noise can only be generated once and needs to be repeatedly used in the subsequent training with the same direction to ensure the effectiveness of the parameter space search.

[0090] As an example, the server can fuse the first gradient calculated by backpropagation in S202 and the directional gradient noise generated in S203. To prevent the gradient noise from affecting the effective gradient information transmission and hindering the gradient descent process, the server can set a gradient threshold. When the relevant value of the first gradient calculated by backpropagation is less than or equal to the set gradient threshold, that is, when the calculated gradient information is weak, the directional gradient noise is fused to correct the gradient calculated by backpropagation, specifically expanding the parameter space search range.

[0091] As an example, the server can set three gradient thresholds, such as the weight gradient threshold , the layer gradient threshold , and the network gradient threshold . That is, the set gradient threshold can include at least one of the weight gradient threshold, the layer gradient threshold, and the network gradient threshold. The specific size of the set gradient threshold in the embodiments of the present application is not limited, and the server can set a reasonable gradient threshold according to the specific training scenario.

[0092] The above three set gradient thresholds can be used to determine whether the relevant values of three types related to the first gradient meet the requirements. Taking a multi-layer feedforward neural network as an example, when the above set gradient threshold includes the weight gradient threshold, the above relationship requirements include the absolute value of the first gradient and the weight gradient threshold The size relationship, i.e., the first relationship requirement; when the neural network includes multiple layers, setting the gradient threshold can include the above-mentioned layer gradient thresholds; the above relationship requirement can include the mean value of the absolute values of the first gradients of each connection weight within the layer and the layer gradient threshold The size relationship, i.e., the second relationship requirement. At this time, the target connection weights include all connection weight terms within the target layer that meet the second relationship requirement; when the set gradient threshold includes the network gradient threshold, the above relationship requirement can include the mean value of the absolute values of the first gradients of all connection weight terms in the neural network and the network gradient threshold The third relationship requirement, at this time, the target connection weights include all connection weight terms of the neural network that meet the third relationship requirement.

[0093] There are various ways to fuse the directional gradient noise corresponding to the target connection weights and the first gradient. As an example, the server can sum the directional gradient noise corresponding to the target connection weights and the first gradient to obtain the fused gradient of the target connection weights. According to the determination method of the relationship between the value related to the first gradient and the set gradient threshold, the above three methods for determining the target connection weights can be provided (the first relationship requirement, the second relationship requirement, and the third relationship requirement corresponding to the weight gradient threshold, the layer gradient threshold, and the network gradient threshold described above), corresponding to three optional noise fusion methods:

[0094] As the first example, the absolute value of the first gradient determines the gradient fusion. The server determines whether the absolute value of the first gradient of any connection weight is less than or equal to the weight gradient threshold, that is, whether it meets the first relationship requirement. If it meets, then this connection weight is the target connection weight, and the sum of the first gradient of this connection weight and the directional noise gradient (i.e., the fused gradient) is used as the new gradient of this connection weight, and the gradients corresponding to the other connection weight terms remain unchanged (still the first gradient). For example, the expression of gradient fusion can be as follows:

[0095] (11)

[0096] As the second example, the mean value of the absolute values of the first gradients of each connection weight within the layer determines the gradient fusion. The server determines whether the mean value of the absolute values of the first gradients of each connection weight within the layer is less than or equal to the layer gradient threshold, that is, whether it meets the second relationship requirement. If it meets, then this layer is the target layer, and the sum of the first gradient of each connection weight within the target layer and the directional noise gradient (i.e., the fused gradient) is used as the new gradient of each connection weight, and the gradients corresponding to the connection weights within the other layers outside the target layer remain unchanged (still the first gradient). For example, the expression of gradient fusion can be as follows:

[0097] (12)

[0098] As a third example, the mean of the absolute values of the first gradients of all connection weight terms in the neural network determines gradient fusion. The server determines whether the mean of the absolute values of the first gradients of all connection weight terms in the neural network is less than or equal to the network gradient threshold, that is, whether the third relational requirement is satisfied. If it is satisfied, the sum of the first gradient of all connection weight terms in the neural network and the directional noise gradient (i.e., the fused gradient) is used as the new gradient of each connection weight. For example, the expression of gradient fusion can be as follows:

[0099] (13)

[0100] Wherein, and are respectively the first gradient and the directional gradient noise corresponding to the connection weight term , is the Iverson bracket. If the condition x is satisfied, = 1. If the condition x is not satisfied, = 0.

[0101] The embodiments of the present application do not limit the above three gradient fusion methods. The server can flexibly select any one of the three fusion methods for gradient fusion, or flexibly combine the three gradient fusion methods in different training stages.

[0102] Considering that due to the phenomenon of gradient disappearance often existing in the neural network to be trained, the gradient cannot be backpropagated from the output side to the input side. Specifically, the gradient magnitude and information intensity on the output side are often much larger than those on the input side. Therefore, the mean of the absolute values of the first gradients of all connection weight terms in the neural network is often larger than the mean of the absolute values of the first gradients of each connection weight within the layer. To address this problem, as an example, for the training process of the same neural network, the network gradient threshold needs to be greater than the weight gradient threshold and the layer gradient threshold , otherwise it is difficult to perform effective gradient fusion.

[0103] As an example, the server can use the fused gradient calculated in S204 to implement the network weight update of the above target connection weight. The specific update method can refer to the general gradient descent algorithm:

[0104] (14)

[0105] Wherein, is the learning rate, which is updated with the gradient descent weight. The neural network can evolve in the reverse direction of the fused gradient at a constant or optimized step size.

[0106] It can be understood that in a certain round of training, if the target connection weights determined in the neural network are updated with the fused gradient, then in this round of training, the other connection weights in the neural network except the target connection weights are updated with the unfused gradient, that is, the first gradient.

[0107] Repeat the training until the obtained neural network meets the set requirements. There can be various specific implementations for this set requirement. As an example, the neural network meeting the set requirements can mean that the neural network can complete a specified task. As another example, the neural network meeting the set requirements can mean that the number of repeated training times of the neural network has reached the number threshold. Therefore, the specific implementation of the above set requirements that the neural network needs to meet is not limited. As an example, the repeated training in S205 can refer to repeating the process of S201 - S204 to train the neural network until the neural network meets the set requirements.

[0108] Considering that in the case where the set gradient threshold is the layer gradient threshold or the network gradient threshold, and the relationship requirement is the second relationship requirement or the third relationship requirement, since the average operation performed by the server on the absolute value of the first gradient will lose the specific gradient intensity information, resulting in a judgment error and losing part of the loss function gradient information. To address this problem, when the set gradient threshold is the layer gradient threshold and the relationship requirement is the second relationship requirement, or when the set gradient threshold is the network gradient threshold and the relationship requirement is the third relationship requirement, the way to update each connection weight of the neural network alternates between the first way and the second way in each round of training; the first way includes: updating the target connection weights with the fused gradient and updating the connection weights in the neural network except the target connection weights with the first gradient; the second way includes: updating all connection weight terms in the neural network with the first gradient. By alternately performing the first way and the second way in multiple rounds of training to update each connection weight of the neural network, the effective transmission of the loss function information can be ensured. As another example, when the set gradient threshold is the weight gradient threshold and the relationship requirement is the first relationship requirement, there is no restriction on alternating different update ways in multiple rounds of training. When the server performs repeated training, each connection weight of the neural network can be updated in the above first way in each round of training.

[0109] It should be noted that the main principle of the neural network training method described in any embodiment of the present application above lies in the directivity of gradient noise. Therefore, for the generation of the initial direction of the directional gradient noise, only the gradient noise with a fixed direction can perform the most effective search for the parameter space. Therefore, the server only generates the initial direction of the gradient noise once and reuses it in subsequent training. If the server repeatedly generates and uses different gradient noise directions during the training process, an effective search for the parameter space cannot be completed. In addition, due to the existence of many local minima in the over-parameterized neural network, when the server generates the initial direction of the directional gradient noise, the specific selection of the initial direction may not be restricted.

[0110] Figure 3 The following is a schematic diagram of a specific training scenario of a neural network provided by an embodiment of the present application:

[0111] Among them, after inputting the training data into the network to be trained, the server can obtain the network output of the neural network under this connection weight. By comparing the network output with the training sample annotation, the server can calculate the gradient according to the backpropagation algorithm, determine the absolute value of the gradient corresponding to each connection of the network, the average value of the absolute value of the gradient of each layer and the average value of the absolute value of the gradient of the entire network. This process can be regarded as the calculation part of the gradient descent training process based on the fusion of directional gradient noise; subsequently, the server can generate directional gradient noise with a fixed sign for each connection weight, and by setting a gradient threshold, confirm the relationship between the absolute value of each gradient and the gradient threshold, fuse the directional gradient noise and the backpropagation calculated gradient (i.e., the first gradient), and finally use the fused gradient as the new gradient value to update the weight of each connection of the network. This process can be regarded as the training part of the gradient descent training process based on the fusion of directional gradient noise. The server will repeat the calculation part and the training part until the network can meet the set requirements, such as correctly identifying the image category of the training sample in the visual processing scenario.

[0112] As an example, the neural network training method described in any embodiment above can be applied to machine learning tasks that apply supervised learning or self-supervised learning during the training process. The machine learning tasks include classification tasks, recognition tasks or generation tasks, and / or, the neural network training method described in any embodiment above can be applied to network architectures that apply supervised learning or self-supervised learning during the training process. The network architectures include deep neural networks, convolutional neural networks, and Transformer architectures, and their initialization modes can depend on the network architecture.

[0113] Corresponding to the above method embodiment, an embodiment of the present application also provides a neural network training device, which is at least applied to the visual processing scenario. See Figure 4 As shown, the device may include:

[0114] An input unit 401 for inputting training samples into a neural network to be trained, and obtaining output results of the neural network under each current connection weight; the training samples at least include visual data;

[0115] A determination unit 402 for determining a first gradient of each connection weight based on the output results and the annotation information of the training samples;

[0116] A generation unit 403 for generating directional gradient noise for each connection weight respectively; wherein, the initial direction of the directional gradient noise is randomly generated and remains fixed during the training process;

[0117] An update unit 404 for, for a target connection weight that meets the relationship requirement between the first gradient and a set gradient threshold, fusing the directional gradient noise corresponding to the target connection weight and the first gradient to obtain a fused gradient, and updating the target connection weight according to the fused gradient;

[0118] A training unit 405 for repeating the training until the obtained neural network meets the set requirements.

[0119] As an example, the determination unit 402 is specifically configured to convert the training samples with the annotation information into a target vector, and convert the output results into a vector form, determine a loss function of the neural network based on the target vector and the output results converted into a vector form; and determine the first gradient based on the loss function.

[0120] As an example, the directional gradient noise is obtained by multiplying the initial direction of the directional gradient noise by the intensity of the directional gradient noise; wherein, the initial direction of the directional gradient noise is randomly generated according to a specified probability; the intensity of the directional random gradient noise is generated in a specified manner and is assigned a positive value.

[0121] As an example, the set gradient threshold includes at least one of a weight gradient threshold, a layer gradient threshold, and a network gradient threshold; wherein, the network gradient threshold is greater than the weight gradient threshold and the layer gradient threshold; when the set gradient threshold includes the weight gradient threshold, the relationship requirement includes a first relationship requirement between the absolute value of the first gradient and the weight gradient threshold; when the neural network includes multiple layers, the set gradient threshold includes the layer gradient threshold; the relationship requirement includes a second relationship requirement between the average value of the absolute values of the first gradients of the connection weights within the layer and the layer gradient threshold, and the target connection weights include all connection weight items within the target layer that satisfy the second relationship requirement; when the set gradient threshold includes the network gradient threshold, the relationship requirement includes a third relationship requirement between the average value of the absolute values of the first gradients of all connection weight items in the neural network and the network gradient threshold, and the target connection weights include all connection weight items of the neural network that satisfy the third relationship requirement.

[0122] As an example, satisfying the first relationship requirement includes: the absolute value of the first gradient is less than or equal to the weight gradient threshold; satisfying the second relationship requirement includes: the average value of the absolute values of the first gradients of the connection weights within the layer is less than or equal to the layer gradient threshold; satisfying the third relationship requirement includes: the average value of the absolute values of the first gradients of all connection weight items in the neural network is less than or equal to the network gradient threshold.

[0123] As an example, the update unit 404 is specifically configured to sum the directional gradient noise corresponding to the target connection weights and the first gradient to obtain the fused gradient.

[0124] As an example, when the set gradient threshold is the layer gradient threshold and the relationship requirement is the second relationship requirement, or when the set gradient threshold is the network gradient threshold and the relationship requirement is the third relationship requirement, the update unit 404 is further configured to update each connection weight of the neural network in an alternating manner between a first manner and a second manner in each round of training; the first manner includes: updating the target connection weights with the fused gradient and updating the connection weights other than the target connection weights in the neural network with the first gradient; the second manner includes: updating all connection weight items in the neural network with the first gradient.

[0125] The present application also provides a computer program product, including a computer program, which when executed by a processor implements the training method of the neural network described in any one of the above embodiments.

[0126] The present application also provides an electronic device, as Figure 5 shown, the electronic device includes:

[0127] Processor 501;

[0128] Memory 502 for storing instructions executable by the processor;

[0129] Wherein, the processor 501 is configured to implement the neural network training method described in any of the above embodiments.

[0130] This application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the neural network training method described in any of the above embodiments.

[0131] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial descriptions of the method embodiments. The device embodiments described above are only illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this application. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0132] The above embodiments can be applied to one or more computer devices. The computer device is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions. The hardware of the computer device includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0133] The computer device can be any electronic product that can perform human-computer interaction with users. For example, personal computers, tablet computers, smart phones, personal digital assistants (PDAs), game consoles, Internet Protocol Televisions (IPTVs), smart wearable devices, etc.

[0134] The computer device may further include a network device and / or a user device. Among them, the network device includes, but is not limited to, a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of hosts or network servers based on cloud computing (Cloud Computing).

[0135] The network where the computer device is located includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, virtual private network (Virtual Private Network, VPN), etc.

[0136] The step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, it is within the protection scope of this patent; adding insignificant modifications or introducing insignificant designs to the algorithm or process, but not changing the core design of its algorithm and process, are all within the protection scope of this application.

[0137] Among them, the description of "specific examples" or "some examples", etc. means that the specific features, structures, materials or characteristics described in connection with the embodiments or examples are included in at least one embodiment or example of this application. In this application, the schematic expression of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0138] Those skilled in the art will readily think of other implementation manners of this application after considering the specification and practicing the invention herein. This application is intended to cover any variations, uses or adaptations of this application. These variations, uses or adaptations follow the general principles of this application and include the common general knowledge or conventional technical means in the technical field not claimed in this application. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of this application are pointed out by the following claims.

[0139] It should be understood that this application is not limited to the exact structure already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is only limited by the appended claims.

[0140] The above is only the specific implementation manner of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this application.

Claims

1. A training method for a neural network, applied to an image classification scenario, characterized in that, including: Inputting training samples into a neural network to be trained to obtain the output results of the neural network under each current connection weight; The training samples include images to be classified; Based on the output results and the annotation information of the training samples, determining the first gradient of each connection weight; Generating directional gradient noise for each connection weight respectively; wherein, the initial direction of the directional gradient noise is randomly generated and remains fixed during the training process; For a target connection weight that satisfies the relationship requirement between the first gradient and a set gradient threshold, fusing the directional gradient noise corresponding to the target connection weight and the first gradient to obtain a fused gradient, and updating the target connection weight according to the fused gradient; Repeating the training until the obtained neural network can complete the image classification task; Wherein, the set gradient threshold includes at least one of a weight gradient threshold, a layer gradient threshold, and a network gradient threshold; the network gradient threshold is greater than the weight gradient threshold and the layer gradient threshold; When the set gradient threshold includes the weight gradient threshold, the relationship requirement includes a first relationship requirement between the absolute value of the first gradient and the weight gradient threshold; When the neural network includes multiple layers, the set gradient threshold includes the layer gradient threshold; the relationship requirement includes a second relationship requirement between the average value of the absolute values of the first gradients of the connection weights within the layer and the layer gradient threshold, and the target connection weights include all connection weight items within the target layer that satisfy the second relationship requirement; When the set gradient threshold includes the network gradient threshold, the relationship requirement includes a third relationship requirement between the average value of the absolute values of the first gradients of all connection weight items in the neural network and the network gradient threshold, and the target connection weights include all connection weight items of the neural network that satisfy the third relationship requirement.

2. The method according to claim 1, wherein The determining the first gradient of each connection weight based on the output results and the annotation information of the training samples includes: Converting the training samples with the annotation information into target vectors, and converting the output results into vector forms, and determining the loss function of the neural network based on the target vectors and the output results converted into vector forms; Based on the loss function, determining the first gradient.

3. The method according to claim 1, wherein The directional gradient noise is obtained by multiplying the initial direction of the directional gradient noise by the intensity of the directional gradient noise; Wherein, the initial direction of the directional gradient noise is randomly generated according to a specified probability; The intensity of the directional random gradient noise is generated in a specified manner and is assigned a positive value.

4. The method according to claim 1, wherein Satisfying the first relationship requirement includes: the absolute value of the first gradient is less than or equal to the weight gradient threshold; Satisfying the second relationship requirement includes: the average value of the absolute values of the first gradients of the connection weights within the layer is less than or equal to the layer gradient threshold; Satisfying the third relationship requirement includes: the average value of the absolute values of the first gradients of all connection weight items in the neural network is less than or equal to the network gradient threshold.

5. The method according to claim 1, wherein Fusing the directional gradient noise corresponding to the target connection weight and the first gradient includes: Summing the directional gradient noise corresponding to the target connection weight and the first gradient to obtain the fused gradient.

6. The method according to claim 1, wherein The method further includes: When the set gradient threshold is the layer gradient threshold and the relationship requirement is the second relationship requirement, or when the set gradient threshold is the network gradient threshold and the relationship requirement is the third relationship requirement, the manner of updating each connection weight of the neural network alternates between the first manner and the second manner in each round of training; The first manner includes: Updating the target connection weight with the fused gradient and updating the connection weights other than the target connection weight in the neural network with the first gradient; The second manner includes: Updating all connection weight terms in the neural network with the first gradient.

7. A training device for a neural network, applied to an image classification scenario, characterized in that It includes: An input unit for inputting training samples into a neural network to be trained to obtain the output result of the neural network under each current connection weight; The training samples include images to be classified; A determination unit for determining the first gradient of each connection weight based on the output result and the annotation information of the training samples; A generation unit for generating directional gradient noise for each connection weight respectively; wherein, the initial direction of the directional gradient noise is randomly generated and remains fixed during training; An update unit for, for a target connection weight that satisfies the relationship requirement between the first gradient and the set gradient threshold, fusing the directional gradient noise corresponding to the target connection weight and the first gradient to obtain a fused gradient, and updating the target connection weight according to the fused gradient; A training unit for repeatedly training until the obtained neural network can complete the image classification task; Wherein, the set gradient threshold includes at least one of a weight gradient threshold, a layer gradient threshold, and a network gradient threshold; the network gradient threshold is greater than the weight gradient threshold and the layer gradient threshold; When the set gradient threshold includes the weight gradient threshold, the relationship requirement includes the first relationship requirement between the absolute value of the first gradient and the weight gradient threshold; When the neural network includes multiple layers, the set gradient threshold includes the layer gradient threshold; the relationship requirement includes the second relationship requirement between the mean value of the absolute values of the first gradients of the connection weights within the layer and the layer gradient threshold, and the target connection weights include all connection weight terms within the target layer that satisfy the second relationship requirement; When the set gradient threshold includes the network gradient threshold, the relationship requirement includes the third relationship requirement between the mean value of the absolute values of the first gradients of all connection weight terms in the neural network and the network gradient threshold, and the target connection weights include all connection weight terms of the neural network that satisfy the third relationship requirement.

8. An electronic device, characterized in that, It includes: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to implement the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps in the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Privacy-sensitive neural network training

    US20250077871A1