Convolutional neural network training method for label noise image processing, electronic equipment and storage medium
By introducing time-average neural network and momentum update strategies into the convolutional neural network training method, the selection deviation and correction deviation problems in label noise data processing are solved, and more stable noise robustness and efficient training effects are achieved.
Patent Information
- Application Number
- CN202510077022.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art has selection bias and correction bias when processing label noise data, resulting in insufficient robustness of deep neural networks in noise environments.
A convolutional neural network training method is proposed, using the sample weight weighted training loss generated by the time-average neural network, and combining the momentum update strategy to dynamically correct the label of noise samples to optimize the label confidence.
The robustness of convolutional neural networks to image label noise is enhanced, especially in high noise-rate environments, without the need for complex hyperparameter tuning and sample selection processes.
Smart Images

Figure CN120012855A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of label noise image processing, and in particular relates to a convolutional neural network training method, electronic equipment and storage medium for label noise image processing. Background Art
[0002] Deep neural networks have made significant progress in recent years in tasks such as image classification and object detection. This achievement is mainly attributed to their supervised training on large-scale labeled datasets. However, building these high-quality large-scale datasets is costly and time-consuming. Therefore, researchers have explored more cost-effective annotation methods. Although these alternative methods reduce costs, they inevitably introduce label noise, that is, the phenomenon that data is incorrectly labeled. The existence of the label noise problem seriously impairs the generalization ability of deep neural networks, because deep neural networks have strong memory capabilities and are prone to overfitting noisy labels, resulting in a significant drop in the performance of the model on the test set. Therefore, how to effectively train deep neural networks on noisy label data has become one of the important challenges in the current field of machine learning.
[0003] At present, researchers have proposed a variety of learning methods for image label noisy datasets to cope with the challenges brought by label noise. Existing noise learning methods can be roughly divided into two categories: data selection and label correction.
[0004] The first category is data selection methods. The core idea of this type of method is to select relatively clean data sets from noisy data sets and use these clean data to train the model, thereby improving the robustness of the model. Specifically, since convolutional neural networks will give priority to learning simple patterns in the data at the beginning of training, this means that they will first remember the data with clean labels. Therefore, many studies (such as Co-teaching, etc.) have proposed methods to identify clean data by selecting small loss data.
[0005] The second type of method belongs to the label correction method. The core idea of this type of method is to increase the scale of reliable training data, thereby enhancing the robustness of the model in a noisy environment. Specifically, by correcting the wrong labels in the noisy data set, the data can more accurately reflect its true category. For example, some studies achieve label correction by estimating the noise conversion matrix; in addition, there are also methods that use the prediction results of the convolutional neural network itself to correct the label.
[0006] However, when screening clean samples in previous label correction methods, the selection bias problem of noise filtering was ignored, that is, there are high-confidence label noise samples in the clean dataset and hard samples (i.e., samples that are difficult to learn but are still clean) in the label noise dataset. This type of data selection bias problem will lead to error accumulation during the training of deep neural networks. In addition, when using the prediction results of the deep neural network itself for label correction, the model usually assigns higher confidence to certain easy-to-classify categories. This correction bias may cause the label noise correction part to mistakenly correct some difficult-to-classify samples, thereby affecting the label noise learning performance, especially in class-imbalanced datasets. Summary of the invention
[0007] The problem to be solved by the present invention is to improve the label noise filtering performance and the stability of label noise correction, and propose a convolutional neural network training method, electronic device and storage medium for label noise image processing.
[0008] To achieve the above object, the present invention is implemented through the following technical solutions:
[0009] A convolutional neural network training method for label noisy image processing comprises the following steps:
[0010] S1. Obtain a labeled noisy image dataset, initialize the convolutional neural network parameters before training, use the initial parameters of the convolutional neural network to initialize the parameters of the time average neural network, and use uniform distribution to initialize the corrected labels of the labeled noisy image dataset;
[0011] S2. Input the labeled noisy image dataset obtained in step S1 into the convolutional neural network, calculate the confidence of the labeled noisy image dataset according to the output of the convolutional neural network, and divide the dataset into a clean dataset and a noisy dataset based on the noise label of the labeled noisy image dataset;
[0012] S3. Update the corrected label obtained in step S1 using the time average neural network, and calculate the sample weight of the label noise image data set based on the time average neural network;
[0013] S4. Calculate the cross entropy loss between the clean data set and the original label, and calculate the cross entropy loss between the noisy data set and the corrected label. Combine the sample weights to weight the two types of cross entropy losses, and use the weighted loss to update the convolutional neural network.
[0014] S5. Using the convolutional neural network parameters updated in step S4, the time-averaged neural network parameters are updated using a momentum update strategy;
[0015] S6. Repeat steps S2 to S5 until the preset maximum number of iterations is reached to complete the convolutional neural network training under the label noisy image.
[0016] Furthermore, the specific implementation method of step S1 includes the following steps:
[0017] S1.1. Set the initial parameters of the convolutional neural network f(·,θ) to θ0, and set the parameters of the convolutional neural network including the multi-layer weight matrix and bias term, where W (l) represents the weight matrix of the lth layer, b (l) represents the bias term of the lth layer;
[0018] Use normal distribution to initialize the weight matrix, the expression is:
[0019]
[0020] Among them, σ is the reciprocal square root of the number of input units, that is, n in is the number of input units in the lth layer, and the bias term is initialized to 0;
[0021] S1.2. Setting the time average neural network uses the same architecture as the convolutional neural network. The parameters of the time average neural network are Including the weight matrix of the time-averaged neural network and the bias term of the time-averaged neural network The initialization method is to set it to the same as the initial parameters of the convolutional neural network;
[0022] S1.3. Corrected labels for noisy image datasets using uniformly distributed initialization The labeled noisy image dataset is recorded as:
[0023]
[0024] Among them, x i represents the i-th training data, is a noisy label, N is the total number of data in the dataset, and C is the total number of categories.
[0025] Furthermore, the specific implementation method of step S2 includes the following steps:
[0026] S2.1. Input the labeled noisy image dataset obtained in step S1 into the convolutional neural network, and use the convolutional neural network f(·,θ) to obtain the predicted output y of the labeled noisy image dataset i =f(x i ,θ), for y i =f(x i ,θ) is normalized to obtain the confidence vector pi , the expression is:
[0027]
[0028] Confidence vector p i Characterizes the convolutional neural network for training data x i Estimation of the probability distribution of belonging to each category;
[0029] S2.2. For the c-th data subset, calculate the average confidence value of the data subset of this category υ c , the expression is:
[0030]
[0031] Among them, N c is the total number of c-th subsets, For training data x k Confidence of belonging to class c;
[0032] S2.3. Combine the labels of the labeled noisy image dataset to divide the training data into a clean dataset and a noisy dataset. If the data x in category c k Confidence Greater than the average confidence value of the data in category c c , then the training data x i The data is divided into clean data set, otherwise it is divided into noisy data set.
[0033] Furthermore, the specific implementation method of step S3 includes the following steps:
[0034] S3.1. Using the time-averaged neural network to obtain the output f(x) of the labeled noise image dataset i ,θ * ), the momentum update strategy is used to update the corrected labels of the label noisy image dataset, and the expression is:
[0035]
[0036] in, represents the corrected label of the noise dataset at training time t, t represents the training time, φ represents the weight coefficient of momentum update, φ∈(0,1), m c represents the one-hot encoding of the cth class;
[0037] S3.2. The prediction output matrix Y of the convolutional neural network f(·,θ) for the labeled noisy image dataset i,j , i∈N,j∈C is transformed into the corresponding one-hot matrix M i,j ,i∈N,j∈C, where
[0038] The stable prediction y of the time-averaged neural network for the labeled noisy image dataset * =f(x i ,θ * ) is normalized, and the expression is:
[0039]
[0040] Among them, p(y i * ∣x i ,θ * ) represents a more stable probability distribution estimate of the training data xi belonging to each category by the time-averaged neural network;
[0041] Combined with M i,j , i∈N, j∈C, calculate the sample weight W of the labeled noisy image dataset i , the expression is:
[0042]
[0043] Among them, P i,j Represents the time average neural network for the training data x i The probability estimate of belonging to class j.
[0044] Furthermore, the specific implementation method of step S4 includes the following steps:
[0045] S4.1. Calculate the cross entropy loss between the clean dataset and the original label. The cross entropy loss between the clean dataset and the original label weighted by the clean sample weight is as follows:
[0046]
[0047] in, is the total number of clean data sets, represents the sample weight of clean data;
[0048] S4.2. Calculate the cross entropy loss of the noisy dataset and the corrected labels. The cross entropy loss of the noisy dataset and the corrected labels weighted by the noise sample weight is as follows:
[0049]
[0050] in, is the total number of noise data sets, Represents the sample weight of the noise data;
[0051] S4.3. Calculate the total loss value as follows:
[0052] Loss=Loss clean +λLossnoise
[0053] Among them, λ is the weight coefficient of the total loss, λ∈(0,1);
[0054] S4.4. Update the convolutional neural network parameters θ using the total loss.
[0055] Furthermore, the specific implementation method of step S5 includes the following steps:
[0056] The momentum update strategy is used to update the time-averaged neural network parameters, and the expression is:
[0057]
[0058] Among them, α represents the weight coefficient of momentum update, α∈(0,1).
[0059] An electronic device includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of a convolutional neural network training method for label noise image processing when executing the computer program.
[0060] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements a convolutional neural network training method for label noisy image processing.
[0061] Beneficial effects of the present invention:
[0062] The convolutional neural network training method for label noisy image processing described in the present invention takes into account the selection bias problem in the label noise filtering process, and uses the sample weights generated by the time-averaged neural network to weight the training loss, thereby enhancing the robustness of the convolutional neural network to image label noise, especially showing more stable performance in a high noise rate environment.
[0063] The convolutional neural network training method for label noise image processing described in the present invention designs a time-averaged neural network based on momentum update, which can dynamically correct the labels of noise samples and continuously optimize the label confidence during the training process. This method not only reduces the dependence on additional resources and complex algorithms, but also avoids problems such as high computational cost and complex parameter tuning, so that the model can maintain accuracy while greatly improving training efficiency.
[0064] The convolutional neural network training method for label noise image processing described in the present invention does not require complex hyperparameter tuning, nor does it rely on error-prone sample selection processes, and is applicable to various practical scenarios. Whether on a small-scale image label noise dataset or a large-scale image label noise dataset, the present invention can easily achieve efficient and robust training effects, and has strong practicality and operability. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 A flowchart of a convolutional neural network training method for label noise image processing according to the present invention;
[0066] Figure 2 The present invention provides an overall framework for a convolutional neural network training method that is robust to image label noise. DETAILED DESCRIPTION
[0067] In order to make the purpose, technical solution and advantages of the present invention more clear, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention, that is, the specific embodiments described are only part of the embodiments of the present invention, rather than all of the specific embodiments. The components of the specific embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations, and the present invention can also have other embodiments.
[0068] Therefore, the following detailed description of the specific embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents the selected specific embodiments of the present invention. Based on the specific embodiments of the present invention, all other specific embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present invention.
[0069] In order to further understand the content, features and effects of the present invention, the following specific implementation methods are given as examples, and the attached Figure 1 and attached Figure 2 The detailed instructions are as follows:
[0070] Embodiment 1:
[0071] A convolutional neural network training method for label noisy image processing comprises the following steps:
[0072] S1. Obtain a labeled noisy image dataset, initialize the convolutional neural network parameters before training, use the initial parameters of the convolutional neural network to initialize the parameters of the time average neural network, and use uniform distribution to initialize the corrected labels of the labeled noisy image dataset;
[0073] Furthermore, the specific implementation method of step S1 includes the following steps:
[0074] S1.1. Set the initial parameters of the convolutional neural network f(·,θ) to θ0, and set the parameters of the convolutional neural network including the multi-layer weight matrix and bias term, where W (l) represents the weight matrix of the lth layer, b (l) represents the bias term of the lth layer;
[0075] Use normal distribution to initialize the weight matrix, the expression is:
[0076]
[0077] Among them, σ is the reciprocal square root of the number of input units, that is, n in is the number of input units in the lth layer, and the bias term is initialized to 0;
[0078] S1.2. Setting the time average neural network uses the same architecture as the convolutional neural network. The parameters of the time average neural network are Including the weight matrix of the time-averaged neural network and the bias term of the time-averaged neural network The initialization method is to set it to the same as the initial parameters of the convolutional neural network;
[0079] S1.3. Corrected labels for noisy image datasets using uniformly distributed initialization The labeled noisy image dataset is recorded as:
[0080]
[0081] Among them, x i represents the i-th training data, is a noisy label, N is the total number of data in the dataset, and C is the total number of categories.
[0082] S2. Input the labeled noisy image dataset obtained in step S1 into the convolutional neural network, calculate the confidence of the labeled noisy image dataset according to the output of the convolutional neural network, and divide the dataset into a clean dataset and a noisy dataset based on the noise label of the labeled noisy image dataset;
[0083] Furthermore, the specific implementation method of step S2 includes the following steps:
[0084] S2.1. Input the labeled noisy image dataset obtained in step S1 into the convolutional neural network, and use the convolutional neural network f(·,θ) to obtain the predicted output y of the labeled noisy image dataset i =f(x i,θ), for y i =f(x i ,θ) is normalized to obtain the confidence vector p i , the expression is:
[0085]
[0086] Confidence vector p i Characterizes the convolutional neural network for training data x i Estimation of the probability distribution of belonging to each category;
[0087] S2.2. For the c-th data subset, calculate the average confidence value of the data subset of this category υ c , the expression is:
[0088]
[0089] Among them, N c is the total number of c-th subsets, For training data x k Confidence of belonging to class c;
[0090] S2.3. Combine the labels of the labeled noisy image dataset to divide the training data into a clean dataset and a noisy dataset. If the data x in category c k Confidence Greater than the average confidence value of the data in category c c , then the training data x i The data is divided into clean data set, otherwise it is divided into noisy data set.
[0091] S3. Update the corrected label obtained in step S1 using the time average neural network, and calculate the sample weight of the label noise image data set based on the time average neural network;
[0092] Furthermore, the specific implementation method of step S3 includes the following steps:
[0093] S3.1. Using the time-averaged neural network to obtain the output f(x) of the labeled noise image dataset i ,θ * ), the momentum update strategy is used to update the corrected labels of the label noisy image dataset, and the expression is:
[0094]
[0095] in, represents the corrected label of the noise dataset at training time t, t represents the training time, φ represents the weight coefficient of momentum update, φ∈(0,1), m c represents the one-hot encoding of the cth class;
[0096] S3.2. The prediction output matrix Y of the convolutional neural network f(·,θ) for the labeled noisy image dataset i,j , i∈N,j∈C is transformed into the corresponding one-hot matrix M i,j ,i∈N,j∈C, where
[0097] The stable prediction y of the time-averaged neural network for the labeled noisy image dataset * =f(x i ,θ * ) is normalized, and the expression is:
[0098]
[0099] Among them, p(y i * ∣x i ,θ * ) represents a more stable probability distribution estimate of the training data xi belonging to each category by the time-averaged neural network;
[0100] Combined with M i,j , i∈N, j∈C, calculate the sample weight W of the labeled noisy image dataset i , the expression is:
[0101]
[0102] Among them, P i,j Represents the time average neural network for the training data x i The probability estimate of belonging to class j.
[0103] S4. Calculate the cross entropy loss between the clean data set and the original label, and calculate the cross entropy loss between the noisy data set and the corrected label. Combine the sample weights to weight the two types of cross entropy losses, and use the weighted loss to update the convolutional neural network.
[0104] Furthermore, the specific implementation method of step S4 includes the following steps:
[0105] S4.1. Calculate the cross entropy loss between the clean dataset and the original label. The cross entropy loss between the clean dataset and the original label weighted by the clean sample weight is as follows:
[0106]
[0107] in, is the total number of clean data sets, represents the sample weight of clean data;
[0108] S4.2. Calculate the cross entropy loss of the noisy dataset and the corrected labels. The cross entropy loss of the noisy dataset and the corrected labels weighted by the noise sample weight is as follows:
[0109]
[0110] in, is the total number of noise data sets, Represents the sample weight of the noise data;
[0111] S4.3. Calculate the total loss value as follows:
[0112] Loss=Loss clean +λLoss noise
[0113] Among them, λ is the weight coefficient of the total loss, λ∈(0,1);
[0114] S4.4. Update the convolutional neural network parameters θ using the total loss.
[0115] S5. Using the convolutional neural network parameters updated in step S4, the time-averaged neural network parameters are updated using a momentum update strategy;
[0116] Furthermore, the specific implementation method of step S5 is to use the momentum update strategy to update the time average neural network parameters, and the expression is:
[0117]
[0118] Among them, α represents the weight coefficient of momentum update, α∈(0,1).
[0119] S6. Repeat steps S2 to S5 until the preset maximum number of iterations is reached to complete the convolutional neural network training under the label noisy image.
[0120] Embodiment 2:
[0121] A convolutional neural network training method for label noisy image processing includes the following steps:
[0122] S1. Obtain a labeled noisy image dataset. Before training, initialize the parameters of the convolutional neural network ResNet18, use the initial parameters of the convolutional neural network to initialize the parameters of the time-averaged neural network, and use uniform distribution to initialize the corrected labels of the labeled noisy image dataset. The specific steps are as follows:
[0123] First, before training begins, the parameters θ0 of the ResNet18 neural network need to be initialized; assuming that the ResNet18 neural network consists of multiple layers, the parameters of each layer include the weight matrix W (l)and the bias term b (l) , where l represents the lth layer;
[0124] Initialize the weight matrix W using a normal distribution (l) :
[0125]
[0126] Where σ is the reciprocal square root of the input dimension of the layer, i.e. n in is the number of input units of this layer, and the bias term b (l) Usually initialized to b (l) =0;
[0127] The time average neural network uses the same architecture as the ResNet18 neural network, and its parameters Including weight matrix and the bias term Its initialization method is to set it to the same as the initial parameters of the ResNet18 neural network; specifically, for each layer l, the initialization time average neural network parameters are:
[0128]
[0129] In order to enhance the robustness of training, the initial correction label y pseudo,c is the uniform distribution of the category set C:
[0130]
[0131] Where |C| represents the number of categories.
[0132] S2. Input the labeled noisy image dataset obtained in step S1 into the convolutional neural network ResNet18, calculate the confidence of the labeled noisy image dataset according to the output of the convolutional neural network ResNet18, and divide the dataset into a clean dataset and a noisy dataset based on the noise label of the labeled noisy image dataset. The specific steps are as follows:
[0133] The noisy label image dataset is recorded as:
[0134]
[0135] Among them, x i represents the i-th training data, is a label that may be noisy, N is the total number of data in the dataset, and C is the total number of categories;
[0136] During the training cycle, the ResNet18 neural network f(·,θ) is used to obtain the predicted output y of the dataset i =f(xi ,θ); for y i =f(x i ,θ) is normalized to get the confidence:
[0137]
[0138] p(y i ∣x i ,θ) represents the probability distribution estimation of the training data xi belonging to each category by the convolutional neural network ResNet18; for the c-th data subset, the average confidence υ of the data subset of this category is calculated c , the expression is:
[0139]
[0140] Among them, N c is the total number of c-th subsets, For training data x k The confidence that it belongs to category c; Combined with the labels of the labeled noisy image dataset, the training data is divided into a clean dataset and a noisy dataset. If the data x in category c k Confidence Greater than the average confidence value of the data in category c c , then the training data x i Divide it into a clean data set, otherwise divide it into a noisy data set, and divide the data set into:
[0141] Clean dataset D clean : Contains all samples with high confidence;
[0142] Noisy dataset D noise : Include all samples with lower confidence.
[0143] S3. Use the time average neural network to update the corrected label obtained in step S1, and calculate the sample weight of the label noise image data set based on the time average neural network. The specific steps are as follows:
[0144] First, the time average neural network is used to obtain the output f(x i ,θ * ), the momentum update strategy is used to update the corrected labels of the label noisy image dataset, and the expression is:
[0145]
[0146] in, represents the corrected label of the noise dataset at training time t, t represents the training time, φ represents the weight coefficient of momentum update, φ∈(0,1), m cRepresents the unique hot encoding of the cth category. Through this correction strategy, the label of the noise sample can be dynamically adjusted to gradually approach the correct category;
[0147] Secondly, calculate the confidence weight W of the samples in the data set i ,i∈N, is used to measure the credibility of the sample during the training process. The specific steps are as follows:
[0148] The prediction output matrix Y of the ResNet18 neural network f(·,θ) for the data set i,j , i∈N,j∈C is transformed into the corresponding one-hot matrix M i,j ,i∈N,j∈C:
[0149]
[0150] The stable prediction y of the time-averaged neural network for the data set * =f(x i ,θ * ) for normalization:
[0151]
[0152] Combined with M i,j ,i∈N,j∈C, get the confidence weight of the data set: W i ,i∈N:
[0153]
[0154] S4. Calculate the cross entropy loss between the clean data set and the original label, and calculate the cross entropy loss between the noisy data set and the corrected label. Combine the sample weights to weight the two types of cross entropy losses, and use the weighted loss to update the convolutional neural network. The specific steps are as follows:
[0155] Calculate the cross entropy loss between the clean dataset and the original label. The cross entropy loss between the clean dataset and the original label weighted by the clean sample weight is as follows:
[0156]
[0157] in, is the total number of clean data sets, represents the sample weight of clean data;
[0158] Calculate the cross entropy loss of the noisy dataset and the corrected labels. The cross entropy loss of the noisy dataset and the corrected labels weighted by the noise sample weight is as follows:
[0159]
[0160] in, is the total number of noise data sets, Represents the sample weight of the noise data;
[0161] Calculate the total loss value: Loss = Loss clean +λLoss noise ; Among them, λ∈(0,1) is the weight coefficient; SGD is used to update the ResNet8 neural network parameter θ using the total loss.
[0162] S5. Using the convolutional neural network parameters updated in step S4, the time-averaged neural network parameters are updated using a momentum update strategy. The specific steps are as follows:
[0163] During the training cycle, the time-averaged neural network is updated using the momentum update strategy:
[0164]
[0165] Among them, α∈(0,1) represents the weight coefficient of momentum update, and t and t-1 represent the time.
[0166] S6. Repeat steps S2 to S5 until the preset maximum number of iterations is reached to complete the convolutional neural network training under the label noisy image.
[0167] The method proposed in this embodiment is experimentally analyzed:
[0168] In the experiment, ResNet18 was selected as the deep neural network model, and all hyperparameters followed the settings of the original paper. The experiment was conducted on the CIFAR10 and CIFAR100 image classification datasets, and two types of noise were introduced: symmetric noise and asymmetric noise. Symmetric noise means that each sample is independently and randomly assigned to a non-real label, the noise ratio is set to 20%, 50%, and 80%, and the noise probability is uniformly distributed. Asymmetric noise means that all samples in a category are only assigned to specific other categories, and the probability of incorrect label assignment is 40%, which is consistent with previous studies. It is worth noting that asymmetric noise is only defined on the CIFAR10 dataset.
[0169] On all datasets, the performance of deep neural networks trained by different methods on the test set was evaluated, and the proposed method was compared with the well-known noise detection algorithms Co-Teaching+, M-correction, and MOIT. The experimental results are detailed in Tables 1 and 2, which record the best accuracy achieved on the test set, where Asym represents asymmetric noise and Sym represents symmetric noise. The results show that the proposed method outperforms the existing noise detection algorithms in all test groups.
[0170] Table 1 Test accuracy of noisy learning on CIFAR10
[0171]
[0172] Table 2 Test accuracy of noisy learning on CIFAR100
[0173]
[0174] Embodiment 3:
[0175] An electronic device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of a convolutional neural network training method for label noise image processing described in Example 1 or Example 2 are implemented.
[0176] The computer device of the present invention may be a device including a processor and a memory, such as a single chip microcomputer including a central processing unit. Furthermore, the processor is used to implement the steps of the above-mentioned convolutional neural network training method for label noise image processing when executing the computer program stored in the memory.
[0177] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0178] The memory may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), etc. In addition, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0179] Embodiment 4:
[0180] A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, a convolutional neural network training method for label noise image processing described in Example 1 or Example 2 is implemented.
[0181] The computer-readable storage medium of the present invention can be any form of storage medium that can be read by a processor of a computer device, including but not limited to non-volatile memory, volatile memory, ferroelectric memory, etc. A computer program is stored on the computer-readable storage medium. When the processor of the computer device reads and executes the computer program stored in the memory, the steps of the above-mentioned convolutional neural network training method for label noise image processing can be implemented.
[0182] The computer program includes computer program code, which may be in source code form, object code form, executable file or some intermediate form, etc. The computer readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer readable media do not include electric carrier signals and telecommunication signals.
[0183] It should be noted that relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0184] Although the present application has been described above with reference to specific embodiments, various modifications may be made thereto and parts thereof may be replaced with equivalents without departing from the scope of the present application. In particular, as long as there is no structural conflict, the various features in the specific embodiments disclosed in the present application may be used in combination with each other in any manner, and the fact that these combinations are not exhaustively described in this specification is only for the sake of omitting space and saving resources. Therefore, the present application is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.
Claims
1. A convolutional neural network training method for label noisy image processing, characterized in that: The steps include: S1. Obtain a labeled noisy image dataset, initialize the convolutional neural network parameters before training, use the initial parameters of the convolutional neural network to initialize the parameters of the time average neural network, and use uniform distribution to initialize the corrected labels of the labeled noisy image dataset; S2. Input the labeled noisy image dataset obtained in step S1 into the convolutional neural network, calculate the confidence of the labeled noisy image dataset according to the output of the convolutional neural network, and divide the dataset into a clean dataset and a noisy dataset based on the noise label of the labeled noisy image dataset; S3. Update the corrected label obtained in step S1 using the time average neural network, and calculate the sample weight of the label noise image data set based on the time average neural network; S4. Calculate the cross entropy loss between the clean data set and the original label, and calculate the cross entropy loss between the noisy data set and the corrected label. Combine the sample weights to weight the two types of cross entropy losses, and use the weighted loss to update the convolutional neural network. S5. Using the convolutional neural network parameters updated in step S4, the time-averaged neural network parameters are updated using a momentum update strategy; S6. Repeat steps S2 to S5 until the preset maximum number of iterations is reached to complete the convolutional neural network training under the label noisy image.
2. The convolutional neural network training method for label noise image processing according to claim 1, characterized in that: The specific implementation method of step S1 includes the following steps: S1.
1. Set the initial parameters of the convolutional neural network f(·,θ) to θ0, and set the parameters of the convolutional neural network including the multi-layer weight matrix and bias term, where W (l) represents the weight matrix of the lth layer, b (l) represents the bias term of the lth layer; Use normal distribution to initialize the weight matrix, the expression is: Among them, σ is the reciprocal square root of the number of input units, that is, n in is the number of input units in the lth layer, and the bias term is initialized to 0; S1.
2. Setting the time average neural network uses the same architecture as the convolutional neural network. The parameters of the time average neural network are Including the weight matrix of the time-averaged neural network and the bias term of the time-averaged neural network The initialization method is to set it to the same as the initial parameters of the convolutional neural network; S1.
3. Corrected labels for noisy image datasets using uniformly distributed initialization The labeled noisy image dataset is recorded as: Among them, x i represents the i-th training data, is a noisy label, N is the total number of data in the dataset, and C is the total number of categories.
3. The convolutional neural network training method for label noise image processing according to claim 2, characterized in that: The specific implementation method of step S2 includes the following steps: S2.
1. Input the labeled noisy image dataset obtained in step S1 into the convolutional neural network, and use the convolutional neural network f(·,θ) to obtain the predicted output y of the labeled noisy image dataset i =f(x i ,θ), for y i =f(x i ,θ) is normalized to obtain the confidence vector p i , the expression is: Confidence vector p i Characterizes the convolutional neural network for training data x i Estimation of the probability distribution of belonging to each category; S2.
2. For the c-th data subset, calculate the average confidence value of the data subset of this category υ c , the expression is: Among them, N c is the total number of c-th subsets, For training data x k Confidence of belonging to class c; S2.
3. Combine the labels of the labeled noisy image dataset to divide the training data into a clean dataset and a noisy dataset. If the data x in category c k Confidence Greater than the average confidence value of the data in category c c , then the training data x i The data is divided into clean data set, otherwise it is divided into noisy data set.
4. The convolutional neural network training method for label noise image processing according to claim 3, characterized in that: The specific implementation method of step S3 includes the following steps: S3.
1. Using the time-averaged neural network to obtain the output f(x) of the labeled noise image dataset i ,θ * ), the momentum update strategy is used to update the corrected labels of the label noisy image dataset, and the expression is: in, represents the corrected label of the noise dataset at training time t, t represents the training time, φ represents the weight coefficient of momentum update, φ∈(0,1), m c represents the one-hot encoding of the cth class; S3.
2. The prediction output matrix Y of the convolutional neural network f(·,θ) for the labeled noisy image dataset i,j , i∈N,j∈C is transformed into the corresponding one-hot matrix M i,j ,i∈N,j∈C, where The stable prediction y of the time-averaged neural network for the labeled noisy image dataset * =f(x i ,θ * ) is normalized, and the expression is: Among them, p(y i * ∣x i ,θ * ) represents the time average neural network for the training data x i More stable probability distribution estimates belonging to each class; Combined with M i,j , i∈N, j∈C, calculate the sample weight W of the labeled noisy image dataset i , the expression is: Among them, P i,j Represents the time average neural network for the training data x i The probability estimate of belonging to class j.
5. The convolutional neural network training method for label noise image processing according to claim 4, characterized in that: The specific implementation method of step S4 includes the following steps: S4.
1. Calculate the cross entropy loss between the clean dataset and the original label. The cross entropy loss between the clean dataset and the original label weighted by the clean sample weight is as follows: in, is the total number of clean data sets, represents the sample weight of clean data; S4.
2. Calculate the cross entropy loss of the noisy dataset and the corrected labels. The cross entropy loss of the noisy dataset and the corrected labels weighted by the noise sample weight is as follows: in, is the total number of noise data sets, Represents the sample weight of the noise data; S4.
3. Calculate the total loss value as follows: Loss=Loss clean +λLoss noise Among them, λ is the weight coefficient of the total loss, λ∈(0,1); S4.
4. Update the convolutional neural network parameters θ using the total loss.
6. The convolutional neural network training method for label noise image processing according to claim 5, characterized in that: The specific implementation method of step S5 is to use the momentum update strategy to update the time average neural network parameters, and the expression is: Among them, α represents the weight coefficient of momentum update, α∈(0,1).
7. An electronic device, characterized in that: The invention comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of a convolutional neural network training method for label noise image processing as described in any one of claims 1 to 6 when executing the computer program.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the convolutional neural network training method for label noisy image processing described in any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Storage and calculation integration-oriented neural network training method and device and storage medium
CN122088624A
Noise label detection and correction method based on deep learning
CN122313065A