A Gradient-Based Method and Device for Measuring the Learning Difficulty of Samples
By combining gradient modulus mean and variance with local anomaly factor and logistic regression model, the problem of inaccurate sample learning difficulty measurement in the existing technology is solved, efficient and accurate measurement in the stochastic gradient descent learning problem is achieved, and the model performance is improved.
Patent Information
- Application Number
- CN202111658842.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-30
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-12-30
AI Technical Summary
The existing sample learning difficulty measurement method has a loss function value that tends to 0 in the later stage of training and the predicted value is affected by a variety of factors, resulting in low accuracy and lack of unified standards for most application scenarios.
By recording the gradient mode mean and variance of each sample during the training process, combining local anomaly factors and logistic regression models, a sample learning difficulty measurement method is constructed, and hyperparameters are adjusted to distinguish samples with high learning difficulty.
Without increasing training time and calculation complexity, the sample learning difficulty is accurately measured, which is suitable for stochastic gradient descent learning problems and improve model performance.
Smart Images

Figure CN114492736B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer application technologies, and in particular, to a method and device for measuring the learning difficulty of samples based on gradients. Background Art
[0002] In many fields such as image recognition and object detection, different learning strategies need to be applied to samples of different difficulties. And the division of the learning difficulty of samples is the core technology among them.
[0003] In recent years, many algorithms have respectively proposed to distinguish the learning difficulty of samples based on the magnitude of the loss function value of the samples during the training process or the prediction after Softmax, and use different learning strategies to improve the performance or accuracy of the model. However, according to existing research, the loss function value of the training samples tends to 0 in the later stage of training, and the prediction is affected by various factors, and the accuracy of reflecting the learning difficulty of samples has a certain degree of fluctuation. The above two measurement criteria are not applicable to some application scenarios with high requirements for the measurement accuracy of the learning difficulty of samples.
[0004] Currently, although some methods different from the above two measurements have been proposed, most of their applicable scenarios are limited to individual learning problems, and there is no unified standard, a relatively reasonable standard applicable to most application scenarios. Summary of the Invention
[0005] To solve the deficiencies of the existing technology, the present invention is based on the gradients of all stages of a sample in a complete training process. The average contribution of the sample to the model optimization in the entire learning process is reflected by the mean of the norms of the gradients, and the degree of dispersion of the contribution of the sample to the optimization in the entire learning process is reflected by the variance of the norms of the gradients. The greater the contribution of the sample to the model optimization, the greater its relative learning difficulty. The present invention adopts the following technical solutions:
[0006] A method for measuring the learning difficulty of samples based on gradients, comprising the following steps:
[0007] S1, collect image data as training samples, construct a deep neural network, and use stochastic gradient descent for training;
[0008] S2, use the deep neural network to perform a complete training on the image samples; record the gradient vectors of the layer before the Softmax layer of each image sample during K iterations in a training process;
[0009] S3, calculate the norms of the gradients of each image sample for K iterations;
[0010] S4, based on the norms of the gradients of each image sample for K iterations, obtain the learning difficulty measurement of each image sample;
[0011] S5. Given the number of samples and the weighted parameter within the neighborhood of a given sample, using the learning difficulty gap between two image samples as the distance metric, calculate the local anomaly factor for each image sample, and determine the anomaly points based on the anomaly factor.
[0012] S6. For the anomaly points, construct a linear regression model, find the parameter that maximizes the local anomaly factor of the anomaly points, and adjust it so that the image samples with higher learning difficulty, as the anomaly points, can be better distinguished from other image samples.
[0013] Further, the K - fold iteration in S2 means that in a complete training process of each image sample, with as the interval, there are a total of K - fold iterations of the gradient vector in the layer before the Softmax layer.
[0014] Further, the gradient vector in S2 is expressed as i represents the i - th training sample, and N represents the number of features describing the sample.
[0015] Further, the norm in S3 is the 1 - norm and / or 2 - norm in one - dimensional space, that is, the sum of the absolute values of the components in each direction. The 1 - norm and 2 - norm in one - dimensional space are equivalent, so the used norm can be selected arbitrarily.
[0016] Further, the norm in S3 is expressed as |·| represents the norm operation.
[0017] Further, S4 is based on the norm of the gradient of each sample for K - fold iterations to obtain the mean value of the gradient and the variance According to the mean value and the variance to obtain the learning - difficulty metric for each sample, including the following steps:
[0018] S41. Mean value:
[0019] S42. Variance:
[0020] S43. Learning - difficulty metric of the image sample: μ represents a hyper - parameter. By introducing the hyper - parameter μ, a weighted sum of the mean value of the norm of the gradients of each sample and the variance of the norm of the gradients is calculated. By adjusting this hyper - parameter, the emphasis direction of the learning - difficulty metric of the samples can be changed.
[0021] At the same time, paying attention to the variance of the norm of the gradient and incorporating the degree of dispersion of the learning - difficulty change of the image sample during the entire training process into the learning - difficulty metric of each image sample can more accurately locate the sample difficulty.
[0022] Furthermore, S5 includes the following steps:
[0023] S51. Based on the learning difficulty gap between two image samples x i and x j being and the number of samples k within a given sample neighborhood and the custom weighted parameter μ, calculate the distance of the sample x i from the k-th nearest sample;
[0024] S52. Define the k-distance neighborhood N k (x i ), that is, all samples within the k-distance of the sample x i ;
[0025] S53. Give the reach-distance k (x i , x j ), that is, if the sample x j is within N k (x i ), then the reach-distance is the k-distance of the sample x i , otherwise it is the true distance between the two samples;
[0026] S54. Calculate the local reachability density:
[0027]
[0028] The smaller the local reachability density, the greater the possibility that it is an outlier;
[0029] S55. Calculate the local dispersion factor:
[0030]
[0031] When LoF k > 1, the image sample corresponding to the data is an outlier, that is, the difficulty gap between this sample and other samples is relatively large, and it is considered that the learning difficulty of the corresponding sample is relatively high.
[0032] Furthermore, the linear regression model in S6 is a logistic linear regression model.
[0033] A gradient-based sample learning difficulty measurement device includes one or more processors for implementing the described gradient-based sample learning difficulty measurement method.
[0034] The advantages and beneficial effects of the present invention are as follows:
[0035] 1. The mean of the magnitude of the gradient only reflects the average performance of the sample during one training process. However, during this training process, the learning difficulty of the sample will change and fluctuate to a certain extent as the model is gradually optimized. Since the model starts learning from relatively easy samples first and optimizes in the gradient direction of these simple samples. As the model parameters are updated and the model becomes more optimized, the gradients of the simple samples hardly change, while at this time the model will focus on optimizing in the gradient direction of relatively difficult samples. The degree of fluctuation of relatively difficult samples is greater than that of relatively easy samples. Therefore, only focusing on the mean of the gradients of each sample will result in a measured learning difficulty that is lower than the actual learning difficulty of the sample. By also considering the variance of the magnitude of the gradient and incorporating the degree of dispersion of the change in the learning difficulty of the sample throughout the training process into the measurement of the learning difficulty of each sample, the sample difficulty can be more accurately located;
[0036] 2. Without changing the time cost of training, the present invention can obtain relatively accurate sample learning difficulty while training once, without increasing the computational complexity of training. And this measurement method is mainly based on the optimization method of stochastic gradient descent, so it is applicable to all learning problems using stochastic gradient descent;
[0037] 3. Introduce a hyperparameter μ to perform a weighted sum of the mean of the magnitude of the gradient of each sample and the variance of the magnitude of the gradient. By adjusting this hyperparameter, the focus direction of the measurement of the learning difficulty of the sample can be changed. Introduce a local outlier factor to analyze the mean and variance, and adjust the hyperparameter μ through a logistic regression model to better distinguish samples with higher learning difficulty as outliers from other samples. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 is the flowchart of the method of the present invention.
[0039] Figure 2 is the guiding diagram of the deep neural network structure and the layer where the used gradient is located in the present invention.
[0040] Figure 3 is the schematic diagram of the device structure of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] The following will describe the specific embodiments of the present invention in detail with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for explaining and illustrating the present invention, and are not used to limit the present invention.
[0042] Taking an IC image data set as an example, it consists of ten categories, with 5000 samples in each category.
[0043] The execution environment of the present invention uses a server with a 3.0 GHz central processing unit, an Nvida 3090Ti GPU processor, and 16 gigabytes of memory, and a learning program is compiled in the Pytorch environment using the Python language, realizing image recognition based on a deep neural network. Other execution environments can also be used, which will not be elaborated here.
[0044] As Figure 1 shown, a gradient-based method for measuring the learning difficulty of samples measures the learning difficulty based on the gradient vector of image samples during the training process. The steps are as follows:
[0045] S1: First, according to 50,000 given training samples, construct a deep neural network ResNet-101 and train it using stochastic gradient descent for a total of 200 epochs. The initial learning rate is set to 0.1, and the learning rate is multiplied by 0.1 every 40 epochs;
[0046] The optimization method defined in the present invention is stochastic gradient descent, which is the most common way to optimize deep neural networks in machine learning and can be understood and reproduced by workers with relevant work experience in deep learning;
[0047] The sample set, deep neural network model, and specific parameters used in the present invention are not restrictive and can be specifically changed according to requirements;
[0048] S2: Use the deep neural network constructed in step 1 to perform a complete training on the given training samples; record the gradient vectors of the layer before the Softmax layer for K iterations (the value of K in the present invention is taken as 20 and can be selected according to requirements) during one training process for each sample i represents the i-th training sample, and N represents the number of features describing the sample.
[0049] The present invention records, for each sample during a complete training process, at intervals, a total of K iterations of the gradient vectors of the layer before the Softmax layer.
[0050] The deep network structure in this part and the layer where the recorded gradient is located are as Figure 2 shown.
[0051] S3: Calculate the norm of the gradients of each sample for K iterations |·| represents the norm operation.
[0052] The norm operation used in the present invention is the 1-norm in one-dimensional space, that is, the sum of the absolute values of the components in each direction. The 1-norm and 2-norm in one-dimensional space are equivalent, so the used norm can be selected by oneself.
[0053] S4: Based on the norms of the gradients for each sample over K iterations, calculate the mean and variance of the gradients, and obtain the learning difficulty metric for each sample, specifically:
[0054] 1) Mean
[0055] 2) Variance
[0056] 3) Learning difficulty metric for the sample
[0057] S5: Denote the learning difficulty gap between two samples x i and x j as Under the given k value (the number of samples within the given sample neighborhood) and μ value (a user-defined weighting parameter) (in this invention, k is selected as 5000 and μ as 0.5, which can be chosen according to specific circumstances), using the difficulty gap between the two samples as the distance metric, calculate the Local Outlier Factor (LoF k ) for each sample; This invention selects to use the Local Outlier Factor to analyze the array composed of the mean and variance of the norms of the gradients of each sample, as follows:
[0058] Based on the difficulty gap d(x i , x j ), that is, the distance metric, and the given k value, the k-distance of sample x i can be calculated, which is the distance to the k-th nearest sample to sample x i ;
[0059] Define the k-distance neighborhood N k (x i ), that is, all samples within the k-distance of sample x i ;
[0060] Give the reach-distance k (x i , x j ), that is, if a certain sample x j is within N k (x i ), then the reach-distance is the k-distance of sample x i , otherwise it is the true distance between the two samples;
[0061] Calculate the local reachability density:
[0062]
[0063] The smaller the local reachability density, the greater the likelihood that it is an outlier;
[0064] Calculation method of local discrete factor:
[0065]
[0066] According to the general understanding of the local outlier factor, it is considered that when LoF k > 1, the point corresponding to the array is an outlier, that is, the difficulty gap between this sample and other samples is relatively large, and it is considered that the learning difficulty of the corresponding sample is relatively high.
[0067] The analysis method of the array composed of the mean and variance selected in the present invention is not unique, and can be selected and set according to requirements and computing power.
[0068] S6: For the outlier, construct a logistic regression model; use the constructed logistic regression model to find the parameter μ that maximizes the local outlier factor (LoF k ) of the outlier.
[0069] The present invention sets the outlier as the sample corresponding to the array with LoF k > 1, and can be adjusted accordingly according to the actual learning problem;
[0070] The present invention selects to construct a logistic regression model and uses this regression model to adjust the hyperparameter μ in the sample learning difficulty metric to find the μ value that maximizes the local outlier factor of the outlier, that is, makes the outlier more deviate from the average level.
[0071] Corresponding to the foregoing embodiment of a gradient-based sample learning difficulty metric method, the present invention also provides an embodiment of a gradient-based sample learning difficulty metric device.
[0072] See Figure 3 , an embodiment of a gradient-based sample learning difficulty metric device provided by an embodiment of the present invention includes one or more processors for implementing a gradient-based sample learning difficulty metric method in the foregoing embodiment.
[0073] An embodiment of a gradient-based sample learning difficulty metric device of the present invention can be applied to any device with data processing capabilities, and the any device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. From the hardware level, as Figure 3 shown, it is a hardware structure diagram of any device with data processing capabilities where a gradient-based sample learning difficulty metric device of the present invention is located. Except forFigure 3 In addition to the processor, memory, network interface, and non-volatile memory shown, any device with data processing capabilities where the device in the embodiment is located may generally include other hardware according to the actual functions of the device with data processing capabilities, which will not be elaborated here.
[0074] For the specific implementation process of the functions and roles of each unit in the above device, please refer to the implementation process of the corresponding steps in the above method, which will not be elaborated here.
[0075] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0076] The embodiment of the present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements a method for measuring the learning difficulty of samples based on gradients in the above embodiment.
[0077] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device of any device with data processing capabilities, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by the device with data processing capabilities, and may also be used to temporarily store the data that has been output or will be output.
[0078] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A gradient-based method for measuring the learning difficulty of samples, characterized in that It includes the following steps: S1. Collect image data as training samples, construct a deep neural network, and use stochastic gradient descent for training; S2. Use the deep neural network to conduct a complete training on the image samples; record the gradient vectors of the layer before the Softmax layer for K iterations during one training process of each image sample; S3. Calculate the norm of the gradients of each image sample for K iterations; the norm is denoted as |·| represents the norm operation; S4. Based on the norms of the gradients of each image sample for K iterations, obtain the learning difficulty metric of each image sample; Magnitude of the gradient for K iterations per sample Obtain the mean of the gradient and variance Based on the mean and variance Obtain the learning difficulty measure for each sample, including the following steps: S41, Mean: S42, Variance: S43, Learning difficulty metric of image samples: μ represents a hyperparameter; S5. Under the given number of samples in the sample neighborhood and the weighting parameter, use the learning difficulty gap between two image samples as the distance metric to obtain the local outlier factor of each image sample, and determine the outlier points according to the outlier factor; specifically, it includes the following steps: S51, based on two image samples x i and x j The learning difficulty gap between them is and the number of samples k within the given sample neighborhood, and the custom weighted parameter μ, calculate the distance of sample x i from the k-th nearest sample; S52, define the k-distance neighborhood N k (x i ), that is, all samples within the k-distance of sample x i ; S53, give the reach - distance k (x i ,x j ), that is, if the sample x j is within N k (x i ), then the reach - distance is the k - th distance of the sample x i , otherwise it is the true distance between the two samples; S54. Calculate the local reachability density: The smaller the local reachability density, the greater the possibility that it is an outlier point; S55. Calculate the local dispersion factor: When LoF k > 1, the image sample corresponding to the data is an outlier; S6. For the outlier points, construct a linear regression model, find the parameters that maximize the local outlier factor of the outlier points, and adjust them to distinguish the image samples with higher learning difficulty as outlier points from other image samples.
2. The method for measuring the learning difficulty of samples based on gradients according to claim 1, wherein The K iterations in S2 refer to the gradient vectors of the previous layer of the Softmax layer for a total of K iterations at intervals of for each image sample in a complete training process.
3. A gradient-based method for measuring the learning difficulty of samples according to claim 1, characterized in that The gradient vector in S2 is expressed as where i represents the i-th training sample and N represents the number of features describing the sample.
4. A gradient-based method for measuring the learning difficulty of samples according to claim 1, characterized in that The norm in S3 is the 1-norm and / or 2-norm in one-dimensional space, that is, the sum of the absolute values of the components in each direction.
5. A method for measuring the learning difficulty of samples based on gradients according to claim 1, characterized in that The linear regression model in S6 is a logistic linear regression model.
6. A gradient-based sample learning difficulty measurement device, characterized in that, It includes one or more processors for implementing a gradient-based sample learning difficulty metric method according to any one of claims 1-5.
Citation Information
Patent Citations
A depth measurement learning method and device for self-adaptive sample synthesis
CN109902805A
Deep neural network regression model based on density screening
CN112861989A