Image noise recognition model training method and device, and image noise recognition method and device

By designing a tag fusion gain expression method for noise data, the characteristics of noise data are directly extracted and the noise training data set is automatically generated, which solves the limitations of the existing noise tag recognition method and achieves high accuracy and robust noise recognition effect.

WO2025107505A1PCT designated stage expired Publication Date: 2025-05-30GUANGZHOU MARITIME INST
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/089547
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-23
Filing Date
2024-04-24
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing noise label identification methods have limitations, which are easy to ignore the inherent characteristics of noise labels, and artificially designed models and strategies may have limitations, resulting in identification errors and missed detection.

Method used

Design a data feature expression method for label fusion gain expression for noise data, directly extract the characteristics of noise data, and construct a noise training data set through the label fusion gain expression vector, automatically generate noise training data without manual annotation.

Benefits of technology

It realizes accurate mining of the inherent characteristics of noise data, accurately judging noise data, improves the dimension and information amount of data feature expression, avoids overfitting in network training, and improves the robustness and accuracy of noise recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024089547_30052025_PF_FP_ABST
    Figure CN2024089547_30052025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of image processing, and discloses an image noise recognition model training method and device, and an image noise recognition method and device. The image noise recognition model training method comprises: obtaining a sample image and an image label thereof; using the sample image and the image label to calculate a label fusion gain expression vector corresponding to the sample image; constructing a noise training data set on the basis of the image label and the label fusion gain expression vector; and training an initial image noise recognition model on the basis of the noise training data set to obtain a target image noise recognition model, the target image noise recognition model being used for image noise recognition. According to the present application, by designing a data feature expression method for label fusion gain expression of noise data, features of the noise data are directly extracted, inherent features of the noise data are accurately mined, and accurate determination is realized. The dimension of data feature expression is improved, overfitting during network training is avoided, and manual annotation of noise data is not required.
Need to check novelty before this filing date? Find Prior Art

Description

Image noise recognition model training method, image noise recognition method and device

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on November 23, 2023, with application number 2023115798761 and invention name “Image noise recognition model training method, image noise recognition method and device”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of image processing technology, and in particular to an image noise recognition model training method, an image noise recognition method and a device. Background Art

[0003] In recent years, with the rapid development of artificial intelligence (AI) technology, large-scale data has become increasingly important in AI research. Data labeling is crucial for the application of large-scale data, and its quality directly impacts the quality of algorithms. Data labeling is the process of assigning labels to the attributes or features of data. Noisy data is a major factor affecting data labeling quality. Noise refers to data with incorrect labeling—that is, data whose labels are inconsistent with their true labels. Labels for noisy data are called noisy labels. Noise is an inevitable occurrence in large-scale data. Regardless of the data collection method or data labeling approach, no noise can be completely eliminated. Noisy data can easily lead to overfitting in deep neural networks and negatively impact both training and generalization performance, leading to reduced learning effectiveness and training quality.

[0004] In the related art, research on noise label detection mostly relies on indirect identification methods, which tend to overlook the inherent characteristics of the noise labels themselves. Furthermore, these methods assume that the distribution of noise data conforms to certain physical or mathematical laws, and artificially design mathematical models or model training strategies to infer noise data. However, these artificially designed models and strategies may have limitations. Therefore, indirect identification methods may lead to recognition errors and missed detections.

[0005] Summary of the Invention

[0006] In view of this, the present application provides an image noise recognition model training method, an image noise recognition method and an apparatus to solve the problem of limitations in existing noise label recognition.

[0007] In a first aspect, the present application provides a method for training an image noise recognition model, the method comprising:

[0008] Get sample images and their image labels;

[0009] Using the sample image and image label, calculate the label fusion gain expression vector corresponding to the sample image;

[0010] Based on the image label and label fusion gain expression vector, a noise training dataset is constructed;

[0011] Based on the noise training data set, the initial image noise recognition model is trained to obtain a target image noise recognition model, which is used to perform image noise recognition.

[0012] In this application, by designing a data feature expression method for label fusion gain expression for noise data, the features of the noise data are directly extracted and the "noise" is directly judged. This can accurately mine the inherent characteristics of the noise data and achieve accurate judgment. The data feature expression method using label fusion gain expression is used to increase the dimension of data feature expression, increase the amount of information, and avoid overfitting during network training. Based on existing data, the noise training data set is automatically generated without the need for manual labeling of the noise data.

[0013] In an optional embodiment, using the sample image and the image label, calculating the label fusion gain expression vector corresponding to the sample image includes:

[0014] Perform feature recognition on the sample image to obtain the image category probability corresponding to the sample image;

[0015] Calculate the difference between the image category probability and the image label to obtain the label difference corresponding to the sample image;

[0016] Based on the category probability and label difference, the gain network is trained to obtain a trained gain network;

[0017] Based on the trained gain network, the label difference, image category probability and image label are gained to obtain the label fusion gain expression vector.

[0018] In this method, the label difference, image category probability and image label are gained through the trained gain network, the dimensions of vectors such as label difference, category probability and image label are increased, and the amount of information of noise features is expanded. In the process of directly identifying noise features, overfitting of network training can be effectively avoided, and more prior information can be brought to the noise learning network.

[0019] In an optional embodiment, based on the trained gain network, the label difference, image category probability and image label are gained to obtain a label fusion gain expression vector, including:

[0020] Input the label difference into the trained gain network to obtain the gain label difference;

[0021] Input the image category probability and image label into the trained gain network to obtain the category probability gain and image label gain;

[0022] Calculate the difference between the category probability gain and the image label gain to obtain the label gain difference;

[0023] The gain label difference is combined with the label gain difference to obtain the label fusion gain expression vector.

[0024] In this way, the label fusion gain expression vector is obtained by fusing the label difference, image category probability and image label, which further enhances the information content of the noise feature expression vector and avoids overfitting of the noise learning network.

[0025] In an optional embodiment, a noise training dataset is constructed based on the image label and the label fusion gain expression vector, including:

[0026] The image label and label fusion gain expression vector are used as negative sample data of the noise training dataset;

[0027] Based on the preset ratio, the pseudo label corresponding to the image label is calculated, and the label fusion gain expression vector corresponding to the pseudo label is calculated using the pseudo label and the sample image. The pseudo label and the label fusion gain expression vector corresponding to the pseudo label are used as the positive sample data of the noise training data set;

[0028] The negative sample data is combined with the positive sample data to obtain a noisy training data set.

[0029] In this method, by modifying the labels of the data set and artificially creating noise data, a training sample is provided for the image noise recognition model, which facilitates the image noise recognition model to directly identify the noise feature data, further improving the robustness and accuracy of noise recognition.

[0030] In an optional embodiment, calculating a pseudo label corresponding to an image label based on a preset ratio includes:

[0031] Based on a preset ratio, the image labels are divided into a first data set, a second data set, and a third data set;

[0032] The image labels in the first data set, the second data set, and the third data set are modified to obtain pseudo labels corresponding to the image labels.

[0033] In this method, noise is artificially created by dividing the noise into three categories, ensuring that the noise data is relatively evenly distributed within different confidence ranges, further improving the accuracy of the trained image noise recognition model.

[0034] In a second aspect, the present application provides an image noise recognition method, the method comprising:

[0035] Get the image to be identified and its image label;

[0036] Based on the image to be identified and its image label, the label fusion gain expression vector corresponding to the image to be identified is calculated;

[0037] The label fusion gain expression vector corresponding to the image to be identified is input into the image noise recognition model to identify the noise probability that the image to be identified is noise. Based on the noise probability, it is determined whether the image to be identified is noise, wherein the image noise recognition model is trained using the image noise recognition model training method of any one of the first aspects.

[0038] In this application, by using a trained noise recognition model, the inherent characteristics of noise data can be accurately mined, and it can be directly determined whether the image to be identified is noise data and whether the label of the image to be identified is correct.

[0039] In a third aspect, the present application provides an image noise recognition model training device, the device comprising:

[0040] A first image acquisition module, configured to acquire a sample image and its image label;

[0041] The first fusion gain calculation module is used to calculate the label fusion gain expression vector corresponding to the sample image using the sample image and the image label;

[0042] A noise training dataset construction module is used to construct a noise training dataset based on image labels and label fusion gain expression vectors;

[0043] The model training module is used to train the initial image noise recognition model based on the noise training data set to obtain the target image noise recognition model, which is used to perform image noise recognition.

[0044] In a fourth aspect, the present application provides an image noise recognition device, comprising:

[0045] A second image acquisition module is used to acquire the image to be identified and its image label;

[0046] The second fusion gain calculation module is used to calculate the label fusion gain expression vector corresponding to the image to be identified based on the image to be identified and its image label;

[0047] The noise recognition module is used to input the label fusion gain expression vector corresponding to the image to be recognized into the image noise recognition model, identify the noise probability that the image to be recognized is noise, and based on the noise probability, determine whether the image to be recognized is noise, wherein the image noise recognition model is trained using the image noise recognition model training device of the third aspect.

[0048] In a fifth aspect, the present application provides a computer device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, computer instructions stored in the memory, and the processor executing the computer instructions to thereby execute the image noise recognition model training method of the above-mentioned first aspect or any corresponding embodiment thereof or the image noise recognition method of the second aspect.

[0049] In a sixth aspect, the present application provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the image noise recognition model training method of the above-mentioned first aspect or any corresponding embodiment thereof, or to execute the image noise recognition method of the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the specific implementation methods of the present application or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the specific implementation methods or the description of the prior art. Obviously, the drawings described below are some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0051] FIG1 is a flow chart of an image noise recognition model training method according to an embodiment of the present application.

[0052] FIG2 is a flowchart of image noise recognition model training and image noise recognition according to an embodiment of the present application.

[0053] FIG3 is a flow chart of another method for training an image noise recognition model according to an embodiment of the present application.

[0054] FIG4 is a schematic diagram of calculating a label fusion gain expression vector according to an embodiment of the present application.

[0055] FIG5 is a flowchart of another image noise recognition model training method according to an embodiment of the present application.

[0056] FIG6 is a flow chart of an image noise recognition method according to an embodiment of the present application.

[0057] FIG7 is a structural block diagram of an image noise recognition model training device according to an embodiment of the present application.

[0058] FIG8 is a structural block diagram of an image noise recognition device according to an embodiment of the present application.

[0059] FIG9 is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present application. DETAILED DESCRIPTION

[0060] To make the purpose, technical solutions, and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of this application.

[0061] In the related art, research on noise label detection mostly relies on indirect identification methods, which tend to overlook the inherent characteristics of the noise labels themselves. Furthermore, these methods assume that the distribution of noise data conforms to certain physical or mathematical laws, and artificially design mathematical models or model training strategies to infer noise data. However, these artificially designed models and strategies may have limitations. Therefore, indirect identification methods may lead to recognition errors and missed detections.

[0062] To solve the above problems, an image noise recognition model training method is provided in an embodiment of the present application for use in a computer device. It should be noted that its execution subject can be an image noise recognition model training device, which can be implemented as part or all of a computer device through software, hardware, or a combination of software and hardware. The computer device can be a terminal, a client, or a server. The server can be a single server or a server cluster composed of multiple servers. The terminal in the embodiment of the present application can be a smart phone, a personal computer, a tablet computer, or other intelligent hardware devices. In the following method embodiments, the execution subject is a computer device as an example for explanation.

[0063] The computer device in this embodiment is suitable for use scenarios in which image noise labels are detected and identified. This application provides an image noise recognition model training method. By designing a data feature expression method for label fusion gain expression for noise data, the features of the noise data are directly extracted, and "whether it is noise" is directly judged. It can accurately mine the inherent features of the noise data and achieve accurate judgment. By utilizing the data feature expression method for label fusion gain expression, the dimension of data feature expression is improved, the amount of information is increased, and overfitting is avoided during network training. Based on existing data, the automatic generation of noise training data sets is realized, and there is no need for manual labeling of noise data.

[0064] According to an embodiment of the present application, an embodiment of a method for training an image noise recognition model is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0065] In this embodiment, a method for training an image noise recognition model is provided, which can be used in the above-mentioned computer device. FIG1 is a flow chart of the method for training an image noise recognition model according to an embodiment of the present application. As shown in FIG1 , the process includes the following steps:

[0066] Step S101: Obtain a sample image and its image label.

[0067] In one example, all images in the dataset D = {x, y} are obtained, where x is the sample image data and y is the image label in the form of one-hot encoding. The form of the image label is not limited in this application. Let y argmaxy is the actual category that the image label refers to.

[0068] Step S102 : Using the sample image and the image label, calculate and obtain the label fusion gain expression vector corresponding to the sample image.

[0069] In one example, by inputting the above data into the image feature extraction network f i Train to realize image feature extraction network f i The ResNet network is used as the image feature extraction network. In this application, there is no restriction on the type of image feature extraction network. Image feature extraction network f i The vector obtained by the Softmax function is recorded as s, which is called the image category probability. Its length is equal to the number of categories n in the data set. By designing the gain network f g The network is trained to obtain a trained gain network. The trained gain network is used to calculate the label fusion gain expression vector corresponding to the sample image.

[0070] Step S103: constructing a noise training dataset based on the image label and the label fusion gain expression vector.

[0071] In one example, to use a binary classification network to directly judge whether the data is "noise", it is necessary to have a data set corresponding to the problem that includes positive and negative samples of "whether it is noise" (referred to as the noise training data set, Noise Training Set, NTS). The noise training data set is constructed using the original data set D by constructing pseudo labels.

[0072] Step S104: training the initial image noise recognition model based on the noise training data set to obtain a target image noise recognition model.

[0073] In an embodiment of the present application, a target image noise recognition model is used to perform image noise recognition.

[0074] In one example, an image noise recognition model uses NTS as training samples and the fusion gain expression vector u as input. The model consists of five fully connected layers, with the first layer having a length of 512, the third and fourth layers having a length of 1024, and the activation function being ReLU. The final fully connected layer, with a length of 1, serves as the network's output layer and uses a Sigmoid activation function. The model is trained with batches of 1024 data samples until all data are traversed 50 times.

[0075] For the input label fusion gain expression vector u of the image noise recognition model, its output is the probability f that the image belongs to noise f (u), the loss function in the image noise recognition model training process is the binary cross entropy loss function, as shown in the following formula:

[0076] Where v is the image label of the image sample in the noise training dataset NTS, v = 1 means that the image sample is a positive sample (i.e., noise data), v = 0 means that the image sample is a negative sample (i.e., non-noise data), log is the number operation, u is the label fusion gain expression vector, f f (u) The probability that the image belongs to noise, L f is the loss function value of the noise identification model.

[0077] In one implementation scenario, FIG2 is a flow chart of image noise recognition model training and image noise recognition according to an embodiment of the present application. As shown in FIG2, the original image x is input into the image feature extraction network f i , get the category probability s corresponding to the image. In the process of image noise recognition model training, a pseudo label is generated based on the image label y. Construct a noise training data set NTS containing positive samples and negative samples, combined with image labels y and pseudo labels With the category probability s, generate the label fusion gain expression vector u of the data in the noise training dataset NTS, and use the noise training dataset to train the image noise recognition model f f The image noise recognition model obtained through learning and training can identify the noise probability corresponding to the original image.

[0078] The image noise recognition model training method provided in this embodiment directly extracts the features of the noise data and directly determines whether it is noise by designing a data feature expression method based on label fusion gain expression for noise data. This method can accurately mine the inherent characteristics of the noise data and achieve precise judgment. The data feature expression method based on label fusion gain expression is used to increase the dimension of data feature expression, increase the amount of information, and avoid overfitting during network training. Based on existing data, the noise training data set is automatically generated without the need for manual labeling of the noise data.

[0079] In this embodiment, a method for training an image noise recognition model is provided, which can be used in the above-mentioned computer device. FIG3 is a flow chart of another method for training an image noise recognition model according to an embodiment of the present application. As shown in FIG3 , the flow chart includes the following steps:

[0080] Step S301: Obtain a sample image and its image label. Please refer to step S101 of the embodiment shown in FIG1 for details, which will not be repeated here.

[0081] Step S302 : Using the sample image and the image label, calculate and obtain the label fusion gain expression vector corresponding to the sample image.

[0082] Specifically, the above step S302 includes:

[0083] Step S3021: perform feature recognition on the sample image to obtain the image category probability corresponding to the sample image.

[0084] In one example, let y argmaxy is the actual category that the image label refers to. By inputting the above data into the image feature extraction network f i Train to realize image feature extraction network f i The weights of are calculated. ResNet network is used as the image feature extraction network. Image feature extraction network f i The calculated vector obtained by processing the Softmax function is recorded as s, where s is the image category probability and its length is equal to the number of categories n in the dataset.

[0085] Step S3022: Calculate the difference between the image category probability and the image label to obtain the label difference corresponding to the sample image.

[0086] In one example, the label difference u l is the difference between the image category probability s and the image label y, and the label difference u is calculated l The formula is as follows: l =sy

[0087] Step S3023: Training the gain network based on the category probability and the label difference to obtain a trained gain network.

[0088] In one example, the gain network consists of a decoding network and an encoding network. The input of the decoding network is a vector of length n. The decoding network f g-de It consists of 6 fully connected layers, where the length of the first and second layers is 128, the length of the third and fourth layers is 256, and the length of the fifth and sixth layers is 512. The activation function is ReLU. The input of the encoding network is the decoding network f g-de The output of the encoding network f g-en It consists of 5 fully connected layers, where the lengths of the first and second layers are 256, the lengths of the third and fourth layers are 128, the length of the fifth layer is n, and the activation function is ReLU.

[0089] The training of the gain network includes: the training data of the gain network consists of the image category probability s and image label y of all images in the dataset, and the network is trained with 512 images per batch until all the data are traversed 50 times. For the input x of the gain network, its output is f g (x), then the loss function in the gain network training process is the sum of the Euclidean distances between the original input vector and the vector obtained after decoding-encoding processing. The formula of the loss function is as follows:

[0090] Among them, x i is the i-th input vector of the gain network, f g (x i ) is the output result of the gain network for the i-th input vector, and d(a,b) is the Euclidean distance between vector a and vector b.

[0091] For the trained gain network f g , only using the decoding network f g-de Part of it is used to increase the gain of the image category probability vector s, so that the amount of information in the input vector can be increased to avoid overfitting of information.

[0092] Step S3024: Based on the trained gain network, the label difference, image category probability and image label are gained to obtain a label fusion gain expression vector.

[0093] In some optional implementations, step S3024 includes:

[0094] Step a1: input the label difference into the trained gain network to obtain the gain label difference.

[0095] In one example, the label difference u l Input to the decoder network f of the trained gain network g-de, obtain the gain label difference u d , the specific formula is as follows: d =f g-de (u l )

[0096] In step a2, the image category probability and image label are input into the trained gain network to obtain the category probability gain and image label gain.

[0097] In one example, the class probability s and the image label y are input into the decoder network f of the trained gain network. g-de In the equation, we get the category probability gain f g-de (s) and image label gain f g-de (y).

[0098] Step a3: Calculate the difference between the category probability gain and the image label gain to obtain the label gain difference.

[0099] In one example, the label gain difference u of the image to be identified is calculated by the following formula: e : u e =f g-de (s)-f g-de (y)

[0100] Step a4: Combine the gain label difference with the label gain difference to obtain a label fusion gain expression vector.

[0101] In one example, the gain label difference u d Difference u from label gain e Splice in sequence to obtain the label fusion gain expression vector u: u=[u d u e ]

[0102] In one implementation scenario, FIG4 is a schematic diagram of calculating a label fusion gain expression vector according to an embodiment of the present application. As shown in FIG4 , the image feature extraction network f i The calculated image category probability s after Softmax function processing, the difference label difference u between the image category probability s and the image label y is calculated l , the label difference u l Input to the decoder network f of the trained gain network g-de , obtain the gain label difference u d , respectively input the category probability s and image label y into the decoder network f of the trained gain network g-de In the equation, we get the category probability gain f g-de (s) and image label gain f g-de (y), calculate the label gain difference u of the image to be identified e, the gain label difference u d Difference u from label gain e Perform splicing in sequence to obtain the label fusion gain expression vector u.

[0103] In this way, the label fusion gain expression vector is obtained by fusing the label difference, image category probability and image label, which further enhances the information content of the noise feature expression vector and avoids overfitting of the noise learning network.

[0104] Step S303: construct a noise training dataset based on the image label and the label fusion gain expression vector. For details, please refer to step S103 of the embodiment shown in FIG1 , which will not be described in detail here.

[0105] Step S304: Based on the noise training data set, the initial image noise recognition model is trained to obtain a target image noise recognition model. For details, please refer to step S103 of the embodiment shown in FIG1 , which will not be described in detail here.

[0106] The image noise recognition model training method provided in this embodiment uses a trained gain network to gain label differences, image category probabilities, and image labels. This increases the dimensionality of these vectors, expanding the information content of noise features. This effectively avoids overfitting of network training during direct noise feature recognition and provides more prior information to the noise learning network. By fusing label differences, image category probabilities, and image labels to generate a label fusion gain expression vector, the information content of the noise feature expression vector is further enhanced, preventing overfitting of the noise learning network.

[0107] In this embodiment, a method for training an image noise recognition model is provided, which can be used in the above-mentioned computer device. FIG5 is a flowchart of another method for training an image noise recognition model according to an embodiment of the present application. As shown in FIG5 , the process includes the following steps:

[0108] Step S501: Obtain a sample image and its image label. Please refer to step S301 of the embodiment shown in FIG3 for details, which will not be repeated here.

[0109] Step S502: Using the sample image and the image label, calculate and obtain the label fusion gain expression vector corresponding to the sample image. For details, please refer to step S302 of the embodiment shown in FIG3 , which will not be described in detail here.

[0110] Step S503: construct a noise training dataset based on the image label and the label fusion gain expression vector.

[0111] Specifically, the above step S503 includes:

[0112] Step S5031: Use the image label and the label fusion gain expression vector as negative sample data of the noise training data set.

[0113] In one example, the data in the noise training dataset NTS consists of two parts: NTS = {u, v}, where u is the label fusion gain expression vector; v is the noise label, with 1 indicating noise and 0 indicating non-noise. For all image data x and original labels y in the original dataset D, the label fusion gain expression vector u is calculated. The noise label v for this part of the data is set to 0, indicating non-noise data, which constitutes the negative samples of NTS.

[0114] In step S5032, based on a preset ratio, a pseudo label corresponding to the image label is calculated, and a label fusion gain expression vector corresponding to the pseudo label is calculated using the pseudo label and the sample image. The pseudo label and the label fusion gain expression vector corresponding to the pseudo label are used as positive sample data of the noise training data set.

[0115] In some optional implementations, the above step S5032 includes:

[0116] Step b1: Divide the image labels into a first data set, a second data set, and a third data set based on a preset ratio.

[0117] Step b2: modify the image labels in the first data set, the second data set, and the third data set to obtain pseudo labels corresponding to the image labels.

[0118] In one example, randomly select 1 / 3 of the data in the original data set D as the first data set, and let the actual category y in the first data set argmaxy = 0, for each image label y in the first dataset, divide the image label y by y argmaxy Randomly select a component y from the other components k1 , modify y k1 The value is a random value in the range of (0.99,1]; argmaxy with y k1 Randomly select a component y from the other components k2 , modify its value to 1-y k1 .

[0119] (2) Randomly select half of the data from the remaining 2 / 3 of the original data set D as the second data set, and let the actual category y in the second data set argmaxy = 0, for each image label y in the second dataset, divide the image label y by y argmaxy Randomly select a component y from the other components k1 , modify y k1The value is a random value in the range of (0.95,0.99]; argmaxy with y k1 Randomly select a component y from the other components k2 , modify its value to (0,1-y k1 ] a random value in the range; in the image label y divided by y argmaxy 、y k1 with y k2 Randomly select a component y from the other components k3 , modify its value to 1-y k1 -y k2 .

[0120] (3) The remaining data that are not selected are used as the third data set. For each image label y in the third data set, ensure that each component y i The sum of is 1, and the subscript of the largest component is not the actual category y of the original label argmaxy Under the premise, randomly assign values ​​to each component in y.

[0121] (4) The label of each image obtained in the above steps (1)-(3) is called a pseudo label Using the image data x and pseudo labels of each image in the original dataset D Calculate its label fusion gain expression vector u, and let this part of the data v=1, which represents noise data. The above data constitutes the positive sample of NTS.

[0122] Step S5033: Combine the negative sample data with the positive sample data to obtain a noise training data set.

[0123] Step S504: Based on the noise training data set, the initial image noise recognition model is trained to obtain a target image noise recognition model. For details, please refer to step S304 of the embodiment shown in FIG3 , which will not be described in detail here.

[0124] The image noise recognition model training method provided in this embodiment modifies dataset labels to artificially generate noise data, providing training samples for the image noise recognition model. This facilitates the model's direct identification of noise signature data, further improving the robustness and accuracy of noise recognition. By classifying artificial noise generation into three categories, the generated noise data is ensured to be relatively evenly distributed within different confidence ranges, further improving the accuracy of the trained image noise recognition model.

[0125] In this embodiment, a method for identifying image noise is provided, which can be used in the above-mentioned computer device. FIG6 is a flow chart of the method for identifying image noise according to an embodiment of the present application. As shown in FIG6 , the flow chart includes the following steps:

[0126] Step S601: Obtain an image to be identified and its image label.

[0127] Step S602 : Based on the image to be identified and its image label, a label fusion gain expression vector corresponding to the image to be identified is calculated.

[0128] Step S603: input the label fusion gain expression vector of the image to be identified into the image noise recognition model to identify the noise probability of the image to be identified as noise, and determine whether the image to be identified is noise based on the noise probability.

[0129] In an embodiment of the present application, the image noise recognition model is trained using the image noise recognition model training method of any one of the above embodiments.

[0130] In one example, for a trained image feature extraction network f i , gain network f g and noise learning network f f (), for the image data I = {x, y} to be judged as noise data, follow the following steps to judge whether its label is a noise label:

[0131] (1) Input the image data x into the image feature extraction network f i , get its category probability s.

[0132] (2) Using the category probability s and data label y, according to the trained gain network f g As in step 3, the label fusion gain expression vector u is obtained.

[0133] (3) Input the label fusion gain expression vector u into the trained noise learning network f f (), the output f of the network is obtained f (u), the output can be considered as the probability that the image belongs to noise.

[0134] (4) Such as f f If (u)>0.5, the data is considered to be noise data and its label is a noise label; otherwise, the data is considered to be non-noise data and its label is a non-noise label.

[0135] The image noise recognition method provided in this embodiment can directly determine whether the image to be recognized is noise data and whether the label of the image to be recognized is correct by using a trained noise recognition model. It can accurately mine the inherent characteristics of noise data and achieve accurate judgment.

[0136] In this embodiment, an image noise recognition model training device is also provided. The device is used to implement the above-mentioned embodiments and optional implementation methods. The details that have been described will not be repeated here. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0137] This embodiment provides an image noise recognition model training device, as shown in FIG7 , including:

[0138] The first image acquisition module 701 is used to acquire a sample image and its image label. Please refer to step S101 of the embodiment shown in FIG1 for details, which will not be repeated here.

[0139] The first fusion gain calculation module 702 is used to calculate the label fusion gain expression vector corresponding to the sample image using the sample image and the image label. Please refer to step S102 of the embodiment shown in Figure 1 for details, which will not be repeated here.

[0140] The noise training data set construction module 703 is used to construct a noise training data set based on the image label and the label fusion gain expression vector. For details, please refer to step S103 of the embodiment shown in Figure 1, which will not be repeated here.

[0141] The model training module is used to train the initial image noise recognition model based on the noise training data set to obtain a target image noise recognition model, which is used to perform image noise recognition. For details, please refer to step S104 of the embodiment shown in Figure 1, which will not be repeated here.

[0142] In some optional implementations, the first fusion gain calculation module 702 includes:

[0143] The feature recognition unit is used to perform feature recognition on the sample image and obtain the image category probability corresponding to the sample image.

[0144] The label difference calculation unit is used to calculate the difference between the image category probability and the image label to obtain the label difference corresponding to the sample image.

[0145] The gain network training unit is used to train the gain network based on the category probability and the label difference to obtain a trained gain network.

[0146] The vector gain unit is used to gain the label difference, image category probability and image label based on the trained gain network to obtain the label fusion gain expression vector.

[0147] In some optional embodiments, the vector gain unit includes:

[0148] The gain label difference calculation subunit is used to input the label difference into the trained gain network to obtain the gain label difference.

[0149] The category sample probability gain calculation subunit is used to input the image category probability and image label into the trained gain network to obtain the category probability gain and image label gain.

[0150] The label gain difference calculation subunit is used to calculate the difference between the category probability gain and the image label gain to obtain the label gain difference.

[0151] The vector gain subunit is used to combine the gain label difference with the label gain difference to obtain a label fusion gain expression vector.

[0152] In some optional implementations, the noise training data set construction module 703 includes:

[0153] The negative sample data construction unit is used to use the image label, the sample image and the label fusion gain expression vector as the negative sample data of the noise training data set.

[0154] The positive sample data construction unit is used to calculate the corresponding pseudo label in the image label based on a preset ratio, use the pseudo label and the sample image to calculate the label fusion gain expression vector corresponding to the pseudo label, and use the pseudo label and the label fusion gain expression vector corresponding to the pseudo label as the positive sample data of the noise training data set.

[0155] The training set construction unit is used to combine negative sample data with positive sample data to obtain a noise training data set.

[0156] In some optional implementations, the positive sample data construction unit includes:

[0157] The data set division subunit is used to divide the image labels into a first data set, a second data set and a third data set based on a preset ratio.

[0158] The pseudo label acquisition subunit is used to modify the image labels in the first data set, the second data set, and the third data set to obtain pseudo labels corresponding to the image labels.

[0159] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0160] The image noise recognition model training device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0161] This embodiment also provides an image noise training device, which is used to implement the above-mentioned embodiments and optional implementations. Details that have already been described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0162] This embodiment provides an image noise recognition device, as shown in FIG8 , including:

[0163] The second image acquisition module 801 is used to acquire the image to be identified and its image label. Please refer to step S601 of the embodiment shown in FIG6 for details, which will not be repeated here.

[0164] The second fusion gain calculation module 802 is used to calculate the label fusion gain expression vector corresponding to the image to be identified based on the image to be identified and its image label. Please refer to step S602 of the embodiment shown in Figure 6 for details, which will not be repeated here.

[0165] Noise identification module 803 is configured to input the label fusion gain expression vector corresponding to the image to be identified into the image noise identification model, determine the noise probability that the image to be identified is noise, and determine whether the image to be identified is noise based on the noise probability. For details, please refer to step S603 of the embodiment shown in Figure 6 and will not be repeated here.

[0166] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0167] The image noise recognition device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0168] An embodiment of the present application further provides a computer device having the image noise recognition model training device shown in FIG. 7 and the image noise recognition device shown in FIG. 8 .

[0169] Please refer to Figure 9, which is a schematic diagram of the structure of a computer device provided in an optional embodiment of the present application. As shown in Figure 9, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are connected to each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in or on the memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 9 takes a processor 10 as an example.

[0170] The processor 10 may be a central processing unit, a network processor, or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.

[0171] The memory 20 stores instructions that can be executed by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.

[0172] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created based on the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0173] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0174] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or a communication network.

[0175] The embodiments of the present application also provide a computer-readable storage medium. The above-mentioned method according to the embodiment of the present application can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.

[0176] Although the embodiments of the present application have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present application, and such modifications and variations shall fall within the scope defined by the appended claims.

Claims

1. A method for training an image noise recognition model, characterized in that: The method comprises: Get sample images and their image labels; Using the sample image and the image label, calculate and obtain a label fusion gain expression vector corresponding to the sample image; Based on the image label and the label fusion gain expression vector, a noise training data set is constructed; Based on the noise training data set, an initial image noise recognition model is trained to obtain a target image noise recognition model, and the target image noise recognition model is used to perform image noise recognition.

2. The method according to claim 1, characterized in that The step of calculating a label fusion gain expression vector corresponding to the sample image by using the sample image and the image label includes: Performing feature recognition on the sample image to obtain an image category probability corresponding to the sample image; Calculating the difference between the image category probability and the image label to obtain the label difference corresponding to the sample image; Based on the category probability and the label difference, the gain network is trained to obtain a trained gain network; Based on the trained gain network, the label difference, the image category probability and the image label are gained to obtain the label fusion gain expression vector.

3. The method according to claim 2, characterized in that The step of performing gain on the label difference, the image category probability and the image label based on the trained gain network to obtain the label fusion gain expression vector includes: Inputting the label difference into the trained gain network to obtain the gain label difference; Inputting the image category probability and the image label into the trained gain network to obtain category probability gain and image label gain; Calculate the difference between the category probability gain and the image label gain to obtain the label gain difference; The gain label difference is combined with the label gain difference to obtain the label fusion gain expression vector.

4. The method according to claim 1, characterized in that: The step of constructing a noise training data set based on the image label and the label fusion gain expression vector includes: Using the image label and the label fusion gain expression vector as negative sample data of the noise training data set; Based on a preset ratio, a pseudo label corresponding to the image label is calculated, and a label fusion gain expression vector corresponding to the pseudo label is calculated using the pseudo label and the sample image, and the pseudo label and the label fusion gain expression vector corresponding to the pseudo label are used as positive sample data of the noise training data set; The negative sample data and the positive sample data are combined to obtain the noise training data set.

5. The method according to claim 4, characterized in that The step of calculating a pseudo label corresponding to the image label based on a preset ratio includes: Based on the preset ratio, the image tags are divided into a first data set, a second data set and a third data set; The image labels in the first data set, the second data set and the third data set are modified to obtain the image labels The pseudo-label corresponding to the signature.

6. A method for identifying image noise, characterized in that: The method comprises: Get the image to be identified and its image label; Based on the image to be identified and its image label, a label fusion gain expression vector corresponding to the image to be identified is calculated; The label fusion gain expression vector corresponding to the image to be identified is input into the image noise recognition model to identify the noise probability that the image to be identified belongs to noise, and based on the noise probability, it is determined whether the image to be identified belongs to noise, wherein the image noise recognition model is trained using the image noise recognition model training method according to any one of claims 1 to 5.

7. An image noise recognition model training device, characterized in that: The device comprises: A first image acquisition module, used to acquire a sample image and its image label; A first fusion gain calculation module, used to calculate a label fusion gain expression vector corresponding to the sample image using the sample image and the image label; A noise training data set construction module, used to construct a noise training data set based on the image label and the label fusion gain expression vector; The model training module is used to train the initial image noise recognition model based on the noise training data set to obtain a target image noise recognition model, and the target image noise recognition model is used to perform image noise recognition.

8. An image noise recognition device, characterized in that: The device comprises: The second image acquisition module is used to acquire the image to be identified and its image label; A second fusion gain calculation module, used for calculating a label fusion gain expression vector corresponding to the image to be identified based on the image to be identified and its image label; A noise recognition module is used to input the label fusion gain expression vector corresponding to the image to be recognized into an image noise recognition model, identify the noise probability that the image to be recognized belongs to noise, and based on the noise probability, determine whether the image to be recognized belongs to noise, wherein the image noise recognition model is trained using the image noise recognition model training device described in claim 7.

9. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the image noise recognition model training method described in any one of claims 1 to 5 or the image noise recognition method described in claim 6 by executing the computer instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the image noise recognition model training method described in any one of claims 1 to 5 or the image noise recognition method described in claim 6.

Citation Information

Patent Citations

  • Unsupervised pedestrian re-identification method based on sample filtering and pseudo label refining

    CN114332517A

  • Noise sample identification method and device

    CN116229196A

  • Image noise recognition model training method, image noise recognition method and device

    CN117710763A

  • Method and apparatus with label noise processing

    US20230252771A1