Methods, apparatus, equipment and storage media for identifying noise in image sample sets

By measuring the model's fit state and loss value during image sample training, and identifying and filtering out noisy samples, the problem of long processing time in existing technologies is solved, and rapid identification and improved model robustness are achieved.

CN112233102BActive Publication Date: 2025-10-28TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011157403.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-26
Publication Date
2025-10-28
Estimated Expiration
2040-10-26

AI Technical Summary

Technical Problem

Existing technologies are time-consuming to identify noise in image sample sets and are difficult to respond quickly to urgent needs, such as in the face of rapidly spreading infectious diseases, resulting in poor robustness of the trained models.

Method used

By training on an image sample set, using a pre-defined standard model to measure the model's fit, and recording the loss value at each fit state, noise samples can be identified and filtered out. Noise samples can be determined with only one complete training session.

Benefits of technology

It improves the rate of noisy sample identification, reduces the time spent on sample screening, enhances the cleanliness of the training set, strengthens the robustness of the model, and enables rapid response to sudden demands.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112233102B_ABST
    Figure CN112233102B_ABST
Patent Text Reader

Abstract

This invention relates to the field of machine learning technology, specifically to a method, apparatus, device, and storage medium for identifying noise in an image sample set. The method includes: acquiring an image sample set; training a model in a first fitting state based on the image sample set until the model enters a second fitting state, wherein the fitting state of the model characterizes the degree of fit between the model and the image sample set; the first fitting state and the second fitting state include zero or at least one intermediate fitting state; the fitting state of the model is determined based on a preset standard model; acquiring the loss value corresponding to each image sample in each fitting state; calculating the loss statistics of each image sample based on the loss values ​​corresponding to each image sample in each fitting state; and identifying noisy image samples in the image sample set based on the loss statistics corresponding to each image sample.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of machine learning technology, and in particular to a method, apparatus, device, and storage medium for identifying noise in a set of image samples. Background Technology

[0002] With the research and advancement of artificial intelligence (AI) technology, AI technology has been researched and applied in many fields, such as finance, healthcare, and the gaming industry.

[0003] In the field of intelligent healthcare, machine learning techniques can be used to process relevant medical images to identify disease attributes. Specifically, a machine learning model can be trained using several manually labeled training sample images. This trained model can then be used to identify disease regions and disease attributes within relevant medical images. However, training this model requires collecting multiple medical images of the disease as training samples. These samples may contain noisy data, such as poorly sampled images (e.g., containing metal artifacts) or images showing missing target regions due to patient displacement, leading to poor robustness of the trained model. Therefore, it is necessary to quickly remove this noisy data to improve the "cleanliness" of the training samples, thereby enhancing the model's robustness.

[0004] In existing technologies, noise samples can generally be identified to some extent by recording the loss values ​​of samples at different stages of training and then statistically analyzing them (e.g., calculating the mean or variance of the samples). However, the typical training process tends to transition from underfitting to overfitting. Directly adopting this training process presents two problems: firstly, if a noise sample is fitted, the loss value decreases rapidly; secondly, it is difficult to determine when a noise sample is fitted. These two issues lead to poor reliability of the statistical results. Therefore, existing technologies have proposed the concept of iterative training, which involves adjusting the learning rate so that it linearly decreases from its original value and then returns to its initial value, repeating this process to allow the network to alternate between underfitting and overfitting to identify noise samples. However, this approach requires repeated training of the network, making the sample selection process very time-consuming and computationally demanding. This makes it unacceptable to spend a significant amount of time selecting training data in the face of rapidly spreading infectious diseases. Summary of the Invention

[0005] To address the aforementioned problems in the prior art, the present invention aims to provide a method, apparatus, device, and storage medium for identifying noise in an image sample set, which can improve the rate of identifying noisy samples in an image sample set and significantly reduce the time spent on sample screening.

[0006] To address the above problems, this invention provides a method for identifying noise in an image sample set, comprising:

[0007] Obtain an image sample set;

[0008] The model in the first fitting state is trained based on the image sample set until the model enters the second fitting state. The fitting state of the model represents the degree of fit between the model and the image sample set. There are zero or at least one intermediate fitting state between the first fitting state and the second fitting state. The fitting state of the model is determined based on a preset standard model.

[0009] Obtain the loss value of each image sample in the image sample set for each fitting state;

[0010] Based on the loss value corresponding to each image sample in each fitting state, calculate the loss statistics of each image sample;

[0011] Based on the loss statistics corresponding to each image sample in the image sample set, the noisy image samples in the image sample set are identified.

[0012] Another aspect of the present invention provides a device for identifying noise in an image sample set, comprising:

[0013] Image sample set acquisition module, used to acquire image sample sets;

[0014] The first model training module is used to train the model in the first fitting state based on the image sample set until the model enters the second fitting state. The fitting state of the model represents the degree of fit between the model and the image sample set. There are zero or at least one intermediate fitting state between the first fitting state and the second fitting state. The fitting state of the model is determined based on a preset standard model.

[0015] The loss value acquisition module is used to acquire the loss value of each image sample in the image sample set for each fitting state.

[0016] The loss statistics determination module is used to calculate the loss statistics of each image sample based on the loss value corresponding to each fitting state of each image sample.

[0017] The noise sample determination module is used to identify noisy image samples in the image sample set based on the loss statistics corresponding to each image sample in the image sample set.

[0018] In another aspect, the present invention provides an electronic device, including a processor and a memory, wherein the memory stores at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the above-described method for identifying noise in a set of image samples.

[0019] In another aspect, the present invention provides a computer-readable storage medium storing at least one instruction or at least one program, wherein the at least one instruction or the at least one program is loaded and executed by a processor to implement the method for identifying noise in a set of image samples as described above.

[0020] In another aspect, the present invention provides a computer program product or computer program comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the aforementioned method for identifying noise in an image sample set.

[0021] Due to the above technical solution, the present invention has the following beneficial effects:

[0022] The method for identifying noise in an image sample set of the present invention trains a model in a first fitting state based on an image sample set containing noise to be identified until the model enters a second fitting state. During the model training process, comparative learning is performed based on a preset standard model to determine multiple fitting states of the model during the training process. Noise samples are determined according to the loss value corresponding to each image sample in the image sample set in each fitting state. Noise image samples can be determined with only one complete training, which can improve the speed of identifying noise samples in the image sample set, significantly reduce the time spent on sample screening, and improve the cleanliness of the training set. Attached Figure Description

[0023] To more clearly illustrate the technical solutions of the present invention, the accompanying drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without any creative effort.

[0024] Figure 1 This is a schematic diagram of the implementation environment provided in one embodiment of the present invention;

[0025] Figure 2 This is a flowchart of a method for identifying noise in an image sample set according to an embodiment of the present invention;

[0026] Figure 3This is a schematic diagram of the structure for identifying noise in an image sample set according to an embodiment of the present invention;

[0027] Figure 4 This is a schematic diagram illustrating the network representation similarity between two 10-layer models provided in an embodiment of the present invention;

[0028] Figure 5 This is a schematic diagram illustrating the network representation similarity between a 14-layer model and a 32-layer model provided in an embodiment of the present invention;

[0029] Figure 6 This is a flowchart of a method for identifying noise in an image sample set according to another embodiment of the present invention;

[0030] Figure 7 This is a schematic diagram of the structure of a noise recognition device for image sample sets provided in an embodiment of the present invention;

[0031] Figure 8 This is a schematic diagram of the structure of a noise recognition device for image sample sets provided in another embodiment of the present invention;

[0032] Figure 9 This is a schematic diagram of the structure of a server provided in one embodiment of the present invention. Detailed Implementation

[0033] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to have perception, reasoning, and decision-making capabilities. AI technology is a comprehensive discipline involving a wide range of fields, encompassing both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0034] The solutions provided in this invention relate to the field of machine learning in artificial intelligence. Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, and other disciplines. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance.

[0035] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0036] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, apparatus, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0037] First, the relevant terms involved in the embodiments of this invention are explained as follows:

[0038] Convolutional Neural Networks (CNNs) are the foundational network framework for deep learning. They utilize operations such as convolution and pooling to extract image features and perform tasks such as image classification and segmentation.

[0039] Feature Map: A feature map obtained by convolving an image with a filter. A new feature map can be generated by convolving a feature map with a filter.

[0040] Centered Kernel Alignment (CKA): A metric function for measuring the similarity of network representations, proposed by Hinton in 2019.

[0041] Hilbert-Schmidt Independence Criterion (HSIC): A statistical measure designed to assess whether two sets are independent of each other.

[0042] Reference manual attached Figure 1 This illustration shows a schematic diagram of an implementation environment provided by an embodiment of the present invention. This implementation environment may include a terminal 110 and a server 120. The terminal 110 and the server 120 can be directly or indirectly connected via wired or wireless communication, which is not limited herein. Specifically, the server 120 can access the data provided by the terminal 110.

[0043] The terminal 110 may include physical devices such as smartphones, tablets, laptops, desktop computers, digital assistants, smart speakers, smart wearable devices, in-vehicle terminals, and servers, and may also include software running on the physical device, such as applications, but is not limited thereto. The operating system running on the terminal 110 in this embodiment of the invention may include, but is not limited to, Android, iOS, Linux, Windows, etc.

[0044] The server 120 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0045] In practical applications, the terminal 110 can acquire multiple images through an image acquisition device and send the acquired images to the server 120. The server 120 can use the acquired images for machine learning training to obtain models such as image classification models and image recognition models. Since image quality may be substandard due to changes in lighting or movement of the subject, the server 120 can also use the method provided in this embodiment to quickly filter out noisy images, improving the "cleanliness" of the training set and thus enhancing the robustness of the trained model.

[0046] The image sample noise recognition method provided in this invention can be applied to various scenarios requiring image sample training, and therefore can be widely used in public security, banking, customs, airports, intelligent video surveillance, intelligent healthcare, and other fields. For example, in the field of intelligent healthcare, a disease attribute recognition model can be trained using relevant medical images. During model training, multiple medical images of the disease need to be collected as training sample data, which is then used for model training. For instance, multiple computed tomography (CT) images from multiple hospitals can be collected as training samples, including images of COVID-19, community-acquired pneumonia, and normal images. The training samples may contain noisy images with substandard sampling quality (e.g., containing metal artifacts) or missing target areas due to patient displacement. The method provided in this invention can be used to quickly filter out these noisy images, and the training samples after filtering out noisy images can be used for model training to obtain a machine learning model that can recognize COVID-19 and community-acquired pneumonia.

[0047] It should be noted that, Figure 1 This is just one example.

[0048] Reference manual attached Figure 2 This illustrates the flowchart of a method for identifying noise in an image sample set according to an embodiment of the present invention. This method can be applied to... Figure 1 The server side in the middle, specifically such as Figure 2 As shown, the method may include the following steps:

[0049] S210: Obtain the image sample set.

[0050] In this embodiment of the invention, the image sample set includes multiple images within the target domain. The source of the image sample set varies depending on the application scenario. For example, in a disease attribute recognition scenario, the image sample set originates from medical images from various hospitals, such as CT scans, raw pathological images obtained directly through electron microscopy, or slices of pathological images, etc., combining multiple medical images into an image sample set. As another example, in a facial expression recognition scenario, the image sample set originates from facial images of various end users, such as facial photos taken while a user is using an application, combining multiple facial photos into an image sample set.

[0051] S220: The model in the first fitting state is trained based on the image sample set until the model enters the second fitting state. The fitting state of the model represents the degree of fitting between the model and the image sample set. There are zero or at least one intermediate fitting state between the first fitting state and the second fitting state. The fitting state of the model is determined based on a preset standard model.

[0052] In this embodiment of the invention, the model can be a deep learning model, which may include a convolutional neural network, such as a residual neural network (ResNet). The deep learning model can also be configured according to actual needs, and this embodiment of the invention does not limit this. For example, the deep learning model may include four convolutional layers and one fully connected layer. The architecture of the preset standard model is the same as the model architecture of the model mentioned above. The preset standard model can be a pre-trained model, which is obtained by pre-training using multiple natural image samples. The pre-trained model can be an ImageNet pre-trained model, i.e., a model obtained by pre-training using the ImageNet dataset.

[0053] In this embodiment of the invention, the first fitting state can be an underfitting state, the second fitting state can be an overfitting state, and the intermediate fitting states can include fitting states. There can be zero or at least one intermediate fitting state, and this embodiment of the invention does not limit the number of such states. The training process gradually transitions from underfitting to fitting, and then gradually to overfitting. In the early stages of training, the model converges relatively quickly because the model tends to learn simple samples (i.e., samples that are easy to train) first. In the later stages of training, the model tends to learn difficult samples (i.e., samples that are difficult to train), and noisy image samples are also included as part of the difficult samples and are trained in the later stages.

[0054] In one possible embodiment, training the model in the first fitting state based on the image sample set until the model enters the second fitting state may include:

[0055] Calculate the network representation similarity between the current model and the preset standard model during the training process;

[0056] The fitting state of the model is determined based on the network representation similarity between the current model and the preset standard model;

[0057] If the model's fit state belongs to the second fit state, then training is stopped.

[0058] In this embodiment of the invention, a preset standard model can be introduced, and the fitting state of the model can be determined by the network expression similarity between the current model and the preset standard model during the training process. Specifically, since the training process generally includes multiple training epochs, each epoch refers to training the model once using all image samples in the image sample set. After each training epoch, the network expression similarity between the current model and the preset standard model can be calculated, and the fitting state of the model can be determined based on the network expression similarity. For example, if the calculated network expression similarity is 20%, it indicates that the model's fitting state is underfitting; if the calculated network expression similarity is 50%, it indicates that the model's fitting state is well-fitting; and if the calculated network expression similarity is 80%, it indicates that the model's fitting state is overfitting.

[0059] In this embodiment of the invention, the second fitting state can be an overfitting state. If the fitting state of the model is an overfitting state, then training is stopped.

[0060] In one possible embodiment, the number of epochs can be preset. The deep learning model is trained on all image samples in the image sample set for the preset number of epochs, so that after training, the model's fitting state belongs to the second fitting state. During the training process, after each training epoch, the network representation similarity between the current model and the preset standard model can be calculated, and the model's fitting state can be determined based on the network representation similarity. For example, 100 epochs can be considered as a complete training process. It should be noted that the specific number of epochs is not limited in this embodiment of the invention. For example, 80 epochs can also be considered as a complete training process.

[0061] In this embodiment of the invention, due to the differences in content between natural images and domain-specific images (e.g., medical images, such as chest CT slices of patients), the high-level semantic information of the network in models trained from the two types of images often differs. This makes it difficult to use the high-level information of the pre-trained model to supervise models trained using domain-specific images. Therefore, only shallow k-layer network features can be used as a supervisory network, allowing the model trained on the image sample set to perform comparative learning based on the provided supervisory network, thus determining the model's fit state. Here, the number of layers k in the supervisory network is a hyperparameter, which can be adjusted to balance the relationship between computational cost and accuracy during the training phase.

[0062] Specifically, refer to the appendix of the instruction manual. Figure 3In this embodiment of the invention, the shallow features of the pre-trained model are used as "probes" for the fitting state of the model. Thus, with only one complete network training, multiple fitting states of the model during the training process can be determined based on the network expression similarity between the model and the pre-trained model. The loss value of the image sample in each fitting state is recorded, thereby filtering out noisy image samples in the image sample set.

[0063] In this embodiment of the invention, the network expression similarity between the current model and the supervised network can be characterized by a metric function such as the Centered Kernel Alignment (CKA) index. CKA can be used to reveal the relationship between different convolutional kernels of convolutional neural networks trained based on different random initializations. For example, based on a preset independence criterion (such as the Hilbert-Schmidt Independence Criterion (HSIC)), CKA can be used to estimate the network expression similarity between the convolutional kernel K in the current model and the convolutional kernel L in the supervised network after each training epoch. The expression is as follows:

[0064]

[0065] That is, optionally, in one possible embodiment, the network representation similarity between the current model and the preset standard model during the computational training process may include:

[0066] Obtain a shallow network with a preset number of layers from the preset standard model, and use it as a supervision network;

[0067] Extract multiple convolutional kernels corresponding to the current model to obtain multiple first convolutional kernels;

[0068] Extract multiple convolutional kernels corresponding to the supervision network to obtain multiple second convolutional kernels;

[0069] Calculate the similarity between each first convolutional kernel and each second convolutional kernel separately to obtain multiple similarity scores;

[0070] The average of the multiple similarities is used as the network representation similarity between the current model and the preset standard model.

[0071] Specifically, in convolutional neural networks, weight parameters are a very important concept. Convolutional layers typically use multiple different convolutional kernels, each with corresponding weight parameters. These weight parameters utilize the local correlation of images to extract and enhance image features. It can be understood that the similarity between the first and second convolutional kernels can be reflected by the similarity between the weight parameters corresponding to the first and second convolutional kernels. That is, optionally, in some possible embodiments, calculating the similarity between each first convolutional kernel and each second convolutional kernel separately to obtain multiple similarities may specifically include:

[0072] Extract the weight information corresponding to each first convolution kernel to obtain multiple first weight information, and extract the weight information corresponding to each second convolution kernel to obtain multiple second weight information;

[0073] The similarity between each first weight information and each second weight information is calculated separately to obtain the calculation results.

[0074] In one example, refer to the appendix of the instruction manual. Figure 4 and Figure 5 , Figure 4 This is a schematic diagram illustrating the network representation similarity between two 10-layer models. Figure 5 This diagram illustrates the network representation similarity between a 14-layer model and a 32-layer model. The horizontal and vertical axes represent the number of layers in the corresponding model structures, and the brightness of the convolutional kernels represents the similarity between the two models; brighter areas indicate higher representational similarity between the two kernels.

[0075] In one possible embodiment, the Earth Mover's Distance (EMD) and other distance metrics can also be used to express the network representation similarity between the current model and the preset standard model.

[0076] S230: Obtain the loss value corresponding to each image sample in the image sample set for each fitting state.

[0077] In this embodiment of the invention, during model training, the model's fitting state gradually transitions from a first fitting state to a second fitting state. The first fitting state can be an underfitting state, and the second fitting state can be an overfitting state. The first fitting state and the second fitting state may not include intermediate fitting states, or they may include one or more intermediate fitting states. The second fitting state and the intermediate fitting states can be determined based on the network representation similarity between the current model and the preset standard model during training. It should be noted that the number of intermediate fitting states can be set according to actual needs, and this embodiment of the invention does not limit this.

[0078] In this embodiment of the invention, during model training, in the early stages (when the model is underfitting), the model can fit clean image samples very well, so the loss of noisy image samples will be larger than the loss of clean image samples, and the difference in loss is significant. In the later stages (when the model is overfitting), the model gradually fits the noisy image samples, and the loss value of the noisy image samples will gradually decrease, making the difference in loss between the two less significant. Therefore, we can identify noisy image samples by statistically analyzing the loss value of each image sample at each fitting state. Specifically, during model training, the loss value of each image sample can be recorded after each training period, and the loss value can be calculated according to a pre-defined loss function.

[0079] In this embodiment of the invention, after determining the various fitting states of the model, for each image sample, one of the multiple loss values ​​corresponding to each fitting state can be selected as the loss value corresponding to the image sample in the current fitting state. For example, assuming the model's fitting states include three states: underfitting (network expression similarity 0%-30%), fitting (network expression similarity 30%-70%), and overfitting (network expression similarity 70%-100%), then the loss values ​​of each image sample in the image sample set when the network expression similarity is 20%, 50%, and 80% can be obtained as the loss values ​​corresponding to the image sample in the underfitting, fitting, and overfitting states, respectively.

[0080] S240: Calculate the loss statistics of each image sample based on the loss value corresponding to each fitting state.

[0081] In this embodiment of the invention, since the loss value of noisy image samples gradually decreases as the training process changes from underfitting to overfitting, while the loss value of clean image samples changes less throughout the training process, noisy image samples can be determined based on the changes in the loss values ​​of each image sample in the image sample set.

[0082] In one possible embodiment, calculating the loss statistics of each image sample based on the loss value corresponding to each fitting state may include:

[0083] The mean and variance of the loss values ​​corresponding to each image sample in each fitting state are calculated and used as the loss statistics for each image sample.

[0084] S250: Identify noisy image samples in the image sample set based on the loss statistics corresponding to each image sample in the image sample set.

[0085] In this embodiment of the invention, during the early stages of training (when the model is underfitting), the model can fit clean image samples well, so the loss of noisy image samples will be greater than the loss of clean image samples, and the difference in loss is quite significant. During the later stages of training (when the model is overfitting), the model gradually fits the noisy image samples, and the loss value of the noisy image samples will gradually decrease, making the difference in loss between the two less significant. Therefore, noisy image samples can be determined based on the mean and variance of each image sample.

[0086] In one possible embodiment, identifying noisy image samples in the image sample set based on the loss statistics corresponding to each image sample in the image sample set may include:

[0087] The image samples in the image sample set are sorted according to the mean and the variance to obtain the target sorting result;

[0088] A predetermined number of image samples that rank highly in the target sorting results are identified as noise image samples.

[0089] Specifically, sorting the image samples in the image sample set according to the mean and the variance to obtain the target sorting result may include:

[0090] The image samples in the image sample set are sorted according to the order of their mean values ​​from largest to smallest to obtain a first sorting result;

[0091] The image samples in the image sample set are sorted according to the order of variance from largest to smallest to obtain a second sorting result;

[0092] The target sorting result is determined based on the first sorting result and the second sorting result.

[0093] Specifically, a comprehensive ranking can be calculated based on the ranking number of each image sample in the first ranking result and the ranking number in the second ranking result. Then, each image sample in the image sample set is sorted from smallest to largest according to the comprehensive ranking to obtain the target ranking result. The top N image samples in the target ranking result are identified as noise image samples.

[0094] In one example, suppose the sequence number of a certain image sample in the first sorting result is L. M The sorting number in the second sorting result is L. S Then, based on these two permutation numbers, the comprehensive sorting of the image samples L = αL can be calculated. M +βL S Wherein, α and β can be set according to actual needs. This embodiment of the invention does not limit this. For example, α and β can both be set to 0.5.

[0095] In summary, the noise identification method in the image sample set of the present invention trains a model in a first fitting state based on the image sample set containing the noise to be identified until the model enters a second fitting state. During the model training process, comparative learning is performed based on a preset standard model to determine multiple fitting states of the model during training. Noise samples are determined based on the loss value corresponding to each image sample in the image sample set in each fitting state. Noise image samples can be determined with only one complete training, which can improve the speed of identifying noise samples in the image sample set, significantly reduce the time spent on sample screening, not only improve the cleanliness of the training set, but also enable rapid response to sudden demands.

[0096] In one possible embodiment, referencing the appendix to the specification... Figure 6 The method may further include the following steps:

[0097] S260: Remove the noisy image samples from the image sample set to obtain a clean image sample set;

[0098] S270: The model is trained based on the clean image sample set to obtain the trained target model.

[0099] In this embodiment of the invention, after obtaining a clean image sample set, the clean image sample set can be used as training samples to train the model. During the training process, the parameters of the model are adjusted until the model converges, thus obtaining the trained target model.

[0100] In summary, the noise identification method in the image sample set of the present invention obtains a clean image sample set by identifying and removing noisy samples in the image sample set, and uses the clean image sample set to train a machine learning model to obtain a target model, which can enhance the robustness of the trained model.

[0101] Reference manual attached Figure 7 This illustrates the structure of a noise identification device 700 for image sample sets provided in one embodiment of the present invention. For example... Figure 7 As shown, the device 700 may include:

[0102] Image sample set acquisition module 710, used to acquire image sample sets;

[0103] The first model training module 720 is used to train the model in the first fitting state based on the image sample set until the model enters the second fitting state. The fitting state of the model represents the degree of fitting between the model and the image sample set. There are zero or at least one intermediate fitting state between the first fitting state and the second fitting state. The fitting state of the model is determined based on a preset standard model.

[0104] The loss value acquisition module 730 is used to acquire the loss value of each image sample in the image sample set for each fitting state.

[0105] The loss statistics calculation module 740 is used to calculate the loss statistics of each image sample based on the loss value corresponding to each fitting state of each image sample.

[0106] The noise sample determination module 750 is used to identify noisy image samples in the image sample set based on the loss statistics corresponding to each image sample in the image sample set.

[0107] In one possible embodiment, such as Figure 8 As shown, the device 700 may further include:

[0108] The noise sample removal module 760 is used to remove the noisy image samples from the image sample set to obtain a clean image sample set.

[0109] The second model training module 770 is used to train the model based on the clean image sample set to obtain the trained target model.

[0110] It should be noted that the apparatus provided in the above embodiments is only illustrated by the division of the above functional modules when implementing its functions. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus provided in the above embodiments and the corresponding method embodiments belong to the same concept, and the specific implementation process can be found in the corresponding method embodiments, which will not be repeated here.

[0111] An embodiment of the present invention also provides an electronic device, which includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the method for identifying noise in an image sample set as provided in the above method embodiments.

[0112] Memory can be used to store software programs and modules. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory. Memory can primarily include a program storage area and a data storage area. The program storage area can store the operating system, application programs required for the functions, etc.; the data storage area can store data created based on the use of the device, etc. Furthermore, memory can include high-speed random access memory, and can also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, memory can also include a memory controller to provide the processor with access to the memory.

[0113] The method embodiments provided in this invention can be executed in a terminal, server, or similar computing device; that is, the aforementioned electronic device may include a terminal, server, or similar computing device. Taking running on a server as an example, such as... Figure 9The diagram illustrates the structural schematic of a server for a noise identification method in a set of running image samples provided in an embodiment of the present invention. The server 900 can vary considerably depending on its configuration or performance, and may include one or more central processing units (CPUs) 910 (e.g., one or more processors) and memory 930, and one or more storage media 920 (e.g., one or more mass storage devices) for storing application programs 923 or data 922. The memory 930 and storage media 920 may be temporary or persistent storage. The program stored in the storage media 920 may include one or more modules, each module including a series of instruction operations on the server. Furthermore, the CPU 910 may be configured to communicate with the storage media 920 and execute the series of instruction operations in the storage media 920 on the server 900. Server 900 may also include one or more power supplies 960, one or more wired or wireless network interfaces 950, one or more input / output interfaces 940, and / or one or more operating systems 921, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0114] The input / output interface 940 can be used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of server 900. In one example, the input / output interface 940 includes a network interface controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In one example, the input / output interface 940 can be a radio frequency (RF) module for wireless communication with the Internet. This wireless communication can use any communication standard or protocol, including but not limited to Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, and Short Messaging Service (SMS).

[0115] Those skilled in the art will understand that Figure 9 The structure shown is for illustrative purposes only; the server 900 may also include more advanced components. Figure 9 More or fewer components than shown, or with Figure 9 Different configurations shown.

[0116] An embodiment of the present invention also provides a computer-readable storage medium, which can be disposed in an electronic device to store at least one instruction or at least one program related to implementing a method for identifying noise in a set of image samples. The at least one instruction or the at least one program is loaded and executed by the processor to implement the method for identifying noise in a set of image samples provided in the above-described method embodiment.

[0117] Optionally, in embodiments of the present invention, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0118] One embodiment of the present invention also provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the noise recognition method for the image sample set provided in the various optional implementations described above.

[0119] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, specific embodiments have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0120] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0121] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.

[0122] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for identifying noise in an image sample set, characterized in that, include: Obtain an image sample set; The model in the first fitting state is trained based on the image sample set until the model enters the second fitting state. The fitting state of the model represents the degree of fit between the model and the image sample set. There are zero or at least one intermediate fitting state between the first fitting state and the second fitting state. The fitting state of the model is determined based on a preset standard model. Obtain the loss value of each image sample in the image sample set for each fitting state; Based on the loss value corresponding to each image sample in each fitting state, calculate the loss statistics of each image sample; Based on the loss statistics corresponding to each image sample in the image sample set, the noisy image samples in the image sample set are identified.

2. The method according to claim 1, characterized in that, The step of training the model in the first fitting state based on the image sample set until the model enters the second fitting state includes: Calculate the network representation similarity between the current model and the preset standard model during the training process; The fitting state of the model is determined based on the network representation similarity between the current model and the preset standard model; If the model's fit state belongs to the second fit state, then training is stopped.

3. The method according to claim 1 or 2, characterized in that, The model is a deep learning model; the preset standard model is a pre-trained model, which is obtained by pre-training multiple natural image samples.

4. The method according to claim 2, characterized in that, The network representation similarity between the current model and the preset standard model during the computational training process includes: Obtain a shallow network with a preset number of layers from the preset standard model, and use it as a supervision network; Extract multiple convolutional kernels corresponding to the current model to obtain multiple first convolutional kernels; Extract multiple convolutional kernels corresponding to the supervision network to obtain multiple second convolutional kernels; Calculate the similarity between each first convolutional kernel and each second convolutional kernel separately to obtain multiple similarity scores; The average of the multiple similarities is used as the network representation similarity between the current model and the preset standard model.

5. The method according to claim 1 or 2, characterized in that, The step of calculating the loss statistics of each image sample based on the loss value corresponding to each fitting state includes: The mean and variance of the loss values ​​corresponding to each image sample in each fitting state are calculated and used as the loss statistics of each image sample. The step of identifying noisy image samples in the image sample set based on the loss statistics corresponding to each image sample in the image sample set includes: The image samples in the image sample set are sorted according to the mean and the variance to obtain the target sorting result; A predetermined number of image samples that rank highly in the target sorting results are identified as noise image samples.

6. The method according to claim 5, characterized in that, The step of sorting each image sample in the image sample set according to the mean and the variance to obtain the target sorting result includes: The image samples in the image sample set are sorted according to the order of their mean values ​​from largest to smallest to obtain a first sorting result; The image samples in the image sample set are sorted according to the order of variance from largest to smallest to obtain a second sorting result; The target sorting result is determined based on the first sorting result and the second sorting result.

7. The method according to claim 1 or 2, characterized in that, The method further includes: The noisy image samples are removed from the image sample set to obtain a clean image sample set; The model is trained based on the clean image sample set to obtain the trained target model.

8. A device for identifying noise in an image sample set, characterized in that, include: Image sample set acquisition module, used to acquire image sample sets; The first model training module is used to train the model in the first fitting state based on the image sample set until the model enters the second fitting state. The fitting state of the model represents the degree of fit between the model and the image sample set. There are zero or at least one intermediate fitting state between the first fitting state and the second fitting state. The fitting state of the model is determined based on a preset standard model. The loss value acquisition module is used to acquire the loss value of each image sample in the image sample set for each fitting state. The loss statistics determination module is used to calculate the loss statistics of each image sample based on the loss value corresponding to each fitting state of each image sample. The noise sample determination module is used to identify noisy image samples in the image sample set based on the loss statistics corresponding to each image sample in the image sample set.

9. An electronic device, characterized in that, The electronic device includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the method for identifying noise in an image sample set as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one instruction or at least one program, which is loaded and executed by a processor to implement the method for identifying noise in an image sample set as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Video denoising method based on cascaded deep residual network

    CN110930327A

  • Training method and device of image classification model

    CN111507419A