Image classification method and device, model training method and device and electronic equipment
By using target noise images and target losses in the image classification model, the problem of low accuracy in traditional models when distinguishing background and semantic information is solved, and higher image classification accuracy is achieved.
Patent Information
- Application Number
- CN202311820757.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-27
- Publication Date
- 2025-06-27
AI Technical Summary
When traditional deep neural networks process images, it is difficult to accurately distinguish background information and semantic information, resulting in a decrease in image classification accuracy.
By obtaining sample images and target noise images, input them into the target classification model for classification, calculate and determine the target loss, to constrain the size relationship between the category confidence scores, and then train the target classification model.
It effectively improves the accuracy of image classification, avoids the model incorrectly fits the background into a key feature, and improves the ability to recognize semantic region features.
Smart Images

Figure CN120219784A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and particularly to an image classification method, a model training method, an apparatus, and an electronic device. Background Art
[0002] With the development of technology, the application of image classification has become more and more extensive. Currently, deep neural networks are generally used to classify images. However, when the background and semantic information in the images are complex and similar, traditional deep neural networks often misfit the background information as semantic information, thereby reducing the accuracy of image classification. Summary of the Invention
[0003] The following is an overview of the subject matter described in detail in the present application. This overview is not intended to limit the scope of protection of the claims.
[0004] Embodiments of the present application provide an image classification method, a model training method, an apparatus, and an electronic device, which can improve the accuracy of image classification.
[0005] On the one hand, an embodiment of the present application provides an image classification method, including:
[0006] Obtaining a sample image and at least one type of target noise image, where the target noise image is obtained by adding noise to the sample image, and the noise addition positions of different types of target noise images are different, and the noise addition positions include the background area or the semantic area of the sample image;
[0007] Inputting the sample image and the target noise image into a target classification model for classification respectively to obtain a first category confidence score corresponding to the sample image and a second category confidence score corresponding to the target noise image;
[0008] Determining a target loss according to the first category confidence score and the second category confidence score, and training the target classification model according to the target loss, where the target loss is used to constrain the magnitude relationship between the first category confidence score and the second category confidence score, and the magnitude relationship is determined according to the noise addition position;
[0009] Obtaining an image to be classified, and inputting the image to be classified into the trained target classification model for classification.
[0010] On the other hand, an embodiment of the present application further provides a model training method, including:
[0011] Obtain a sample image and at least one type of target noise image, where the target noise image is obtained by adding noise to the sample image, and the noise addition positions of different types of target noise images are different, and the noise addition positions include the background area or the semantic area of the sample image;
[0012] Input the sample image and the target noise image into a target classification model for classification respectively to obtain a first category confidence score corresponding to the sample image and a second category confidence score corresponding to the target noise image;
[0013] Determine a target loss according to the first category confidence score and the second category confidence score, and train the target classification model according to the target loss, where the target loss is used to constrain the magnitude relationship between the first category confidence score and the second category confidence score, and the magnitude relationship is determined according to the noise addition position.
[0014] On the other hand, an embodiment of the present application also provides an image classification device, including:
[0015] A first image acquisition module, configured to obtain a sample image and at least one type of target noise image, where the target noise image is obtained by adding noise to the sample image, and the noise addition positions of different types of target noise images are different, and the noise addition positions include the background area or the semantic area of the sample image;
[0016] A first score acquisition module, configured to input the sample image and the target noise image into a target classification model for classification respectively to obtain a first category confidence score corresponding to the sample image and a second category confidence score corresponding to the target noise image;
[0017] A first training module, configured to determine a target loss according to the first category confidence score and the second category confidence score, and train the target classification model according to the target loss, where the target loss is used to constrain the magnitude relationship between the first category confidence score and the second category confidence score, and the magnitude relationship is determined according to the noise addition position;
[0018] An inference module, configured to obtain an image to be classified, and input the image to be classified into the trained target classification model for classification.
[0019] Further, at least one type of the target noise image includes a first target noise image obtained by adding noise to the background area and a second target noise image obtained by adding noise to the semantic area, and the first training module is further configured to:
[0020] Determine a first loss according to the first category confidence score and the second category confidence score corresponding to the first target noise image, where the first loss is used to constrain the first category confidence score to be equal to the second category confidence score corresponding to the first target noise image;
[0021] Determine a second loss according to the first category confidence score and the second category confidence score corresponding to the second target noise image, where the second loss is used to constrain the first category confidence score to be greater than the second category confidence score corresponding to the second target noise image;
[0022] Determine a target loss according to the first loss and the second loss.
[0023] Further, the background region and the semantic region are determined by semantic segmentation, and the first training module is further configured to:
[0024] Obtain a first segmentation confidence of the background region and a second segmentation confidence of the semantic region;
[0025] Normalize according to the first segmentation confidence and the second segmentation confidence to obtain a first weight of the background region and a second weight of the semantic region;
[0026] Use the second weight as the weight of the first loss, use the first weight as the weight of the second loss, and weight the first loss and the second loss to obtain a target loss.
[0027] Further, the first image acquisition module is further configured to:
[0028] Obtain a sample image, input the sample image into a semantic segmentation model for semantic segmentation to obtain an original background image corresponding to the background region and an original semantic image corresponding to the semantic region;
[0029] Add noise to the original background image and the original semantic image respectively to obtain a noise background image corresponding to the original background image and a noise semantic image corresponding to the original semantic image;
[0030] Merge the noise background image and the original semantic image to obtain one type of target noise image, and merge the noise semantic image and the original background image to obtain the other type of target noise image.
[0031] Further, the first image acquisition module is further configured to:
[0032] Input the sample image into the semantic segmentation model, and perform multiple scale-downs on the sample image in sequence to obtain the sample image at multiple scales;
[0033] Encode the corresponding sample images based on the encoders corresponding to each scale in the order from high to low in scale to obtain target encoded images for each scale, where the input of the encoder at the current scale is obtained by merging the sample image at the current scale and the target encoded image at the previous scale;
[0034] Decode the corresponding target encoded images based on the decoders corresponding to each scale in the order from low to high in scale to obtain decoded images for each scale, where the input of the decoder at the current scale is obtained by merging the target encoded image at the current scale and the decoded image at the previous scale;
[0035] Merge the decoded image at the current scale with the upsampled image at the previous scale to obtain a merged image, where the upsampled image is obtained by upsampling the decoded image at the previous scale;
[0036] Perform semantic segmentation on the merged image corresponding to the original scale to obtain a background image corresponding to the background region and a semantic image corresponding to the semantic region.
[0037] Further, the first image acquisition module is further configured to:
[0038] Encode the corresponding sample images based on the encoders corresponding to each scale to obtain original encoded images for each scale;
[0039] Determine the spatial attention matrix and the position attention matrix corresponding to the sample image corresponding to each scale respectively, and merge the spatial attention matrix and the position attention matrix corresponding to each scale to obtain a merged attention matrix;
[0040] Perform max pooling on the merged attention matrices corresponding to multiple scales to obtain a target attention matrix;
[0041] Adjust the original encoded images corresponding to each scale based on the target attention matrix to obtain target encoded images for each scale.
[0042] Further, the first image acquisition module is further configured to:
[0043] Determine a first noise intensity corresponding to the original background image and a second noise intensity corresponding to the original semantic image;
[0044] Generate a first noise image having the same size as the original background image according to the first noise intensity, and merge the first noise image with the original background image to obtain a noise background image corresponding to the original background image;
[0045] Generate a second noise image with the same size as the original semantic image according to the second noise intensity, and merge the second noise image with the original semantic image to obtain a noise semantic image corresponding to the original semantic image.
[0046] Further, the first image acquisition module is further configured to:
[0047] Input the original background image and the original semantic image into a noise intensity prediction model, extract a first image feature of the original background image and a second image feature of the original semantic image, splice the first image feature and the second image feature to obtain a spliced image feature, and perform linear regression on the spliced image feature to obtain a noise intensity prediction result, where the noise intensity prediction result includes a first noise standard deviation corresponding to the original background image and a second noise standard deviation corresponding to the original semantic image;
[0048] Determine a first noise intensity corresponding to the original background image according to the first noise standard deviation, and determine a second noise intensity corresponding to the original semantic image according to the second noise standard deviation.
[0049] Further, the number of the noise intensity prediction results is multiple, the noise intensity prediction results are output by the same output layer of the noise intensity prediction model, the output layer also outputs intensity weights corresponding to the respective noise intensity prediction results, and the first image acquisition module is further configured to:
[0050] Weight the multiple first noise standard deviations according to the intensity weights to obtain a first weighted standard deviation;
[0051] Weight the multiple second noise standard deviations according to the intensity weights to obtain a second weighted standard deviation;
[0052] Determine a first noise intensity corresponding to the original background image according to the first weighted standard deviation, and determine a second noise intensity corresponding to the original semantic image according to the second weighted standard deviation.
[0053] On the other hand, an embodiment of the present application further provides a model training device, including:
[0054] A second image acquisition module, configured to acquire a sample image and at least one type of target noise image, where the target noise image is obtained by adding noise to the sample image, and the noise addition positions of different types of target noise images are different, and the noise addition positions include a background area or a semantic area of the sample image;
[0055] A second score acquisition module, configured to input the sample image and the target noise image into a target classification model respectively for classification, so as to obtain a first category confidence score corresponding to the sample image and a second category confidence score corresponding to the target noise image;
[0056] A second training module, configured to determine a target loss according to the first category confidence score and the second category confidence score, and train the target classification model according to the target loss, wherein the target loss is used to constrain the magnitude relationship between the first category confidence score and the second category confidence score, and the magnitude relationship is determined according to the noise addition position.
[0057] On the other hand, an embodiment of the present application further provides an electronic device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the above image classification method or model training method is implemented.
[0058] On the other hand, an embodiment of the present application further provides a computer-readable storage medium, where the storage medium stores a computer program, and the computer program is executed by a processor to implement the above image classification method or model training method.
[0059] On the other hand, an embodiment of the present application further provides a computer program product, which includes a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the above image classification method or model training method.
[0060] The embodiments of the present application at least include the following beneficial effects: By obtaining a sample image and at least one type of target noise image, inputting the sample image and the target noise image into a target classification model respectively for classification to obtain a first category confidence score corresponding to the sample image and a second category confidence score corresponding to the target noise image, determining a target loss according to the first category confidence score and the second category confidence score, and training the target classification model according to the target loss. Since the target loss is used to constrain the magnitude relationship between the first category confidence score and the second category confidence score, and the magnitude relationship is determined according to the noise addition position, therefore, even if noise is added to at least one of the background area or the semantic area of the sample image, the category confidence score output by the target classification model can be constrained based on the target loss, which helps to prevent the target classification model from wrongly fitting the background as the key feature for classification. When a to-be-classified image is obtained later and input into the trained target classification model for classification, the accuracy of image classification can be effectively improved.
[0061] Other features and advantages of the present application will be described in the subsequent specification, and in part will become apparent from the specification, or will be understood by implementing the present application. Description of the Drawings
[0062] The drawings are used to provide a further understanding of the technical solution of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the technical solution of the present application and do not constitute a limitation to the technical solution of the present application.
[0063] Figure 1 Schematic diagram of an optional implementation environment provided for an embodiment of the present application;
[0064] Figure 2 Schematic diagram of an optional process flow of an image classification method provided for an embodiment of the present application;
[0065] Figure 3 Schematic diagram of a target noise image provided for an embodiment of the present application;
[0066] Figure 4 Schematic diagram of the process of determining a target loss provided for an embodiment of the present application;
[0067] Figure 5 Schematic diagram of the process of determining a target loss provided for another embodiment of the present application;
[0068] Figure 6 Schematic diagram of the process of determining a target loss provided for another embodiment of the present application;
[0069] Figure 7 Schematic diagram of an image classification process provided for an embodiment of the present application;
[0070] Figure 8 Schematic diagram of the magnitude relationship between the first category confidence score and the second category confidence score provided for an embodiment of the present application;
[0071] Figure 9 Schematic diagram of the process of determining a target loss provided for another embodiment of the present application;
[0072] Figure 10 Schematic diagram of the effect of a semantic segmentation model provided for an embodiment of the present application;
[0073] Figure 11 Schematic diagram of the process of generating a target noise image provided for an embodiment of the present application;
[0074] Figure 12 Schematic diagram of the process of determining noise intensity provided for an implementation of the present application;
[0075] Figure 13 Schematic diagram of the architecture of a semantic segmentation model provided for an embodiment of the present application;
[0076] Figure 14 Schematic diagram of the process for generating a target encoded image provided by an embodiment of the present application;
[0077] Figure 15 An optional flowchart of the model training method provided by an embodiment of the present application;
[0078] Figure 16 An optional overall flowchart of the image classification method provided by an embodiment of the present application;
[0079] Figure 17 An optional overall flowchart of the model training method provided by an embodiment of the present application;
[0080] Figure 18 An optional structural diagram of the image classification device provided by an embodiment of the present application;
[0081] Figure 19 An optional structural diagram of the model training device provided by an embodiment of the present application;
[0082] Figure 20 Partial structural block diagram of the terminal provided by an embodiment of the present application;
[0083] Figure 21 Partial structural block diagram of the server provided by an embodiment of the present application. Detailed implementation manners
[0084] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0085] It should be noted that in each specific implementation manner of the present application, when it comes to performing relevant processing based on data related to the characteristics of the target object, such as target object attribute information or attribute information set, etc., the permission or consent of the target object will be obtained first. Moreover, the collection, use and processing of these data will comply with relevant laws, regulations and standards. Among them, the target object may be a user. In addition, when an embodiment of the present application needs to obtain target object attribute information, it will obtain the separate permission or separate consent of the target object through pop-up windows or by jumping to a confirmation page, etc. After clearly obtaining the separate permission or separate consent of the target object, the necessary data related to the target object for the normal operation of the embodiment of the present application will be obtained.
[0086] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of that module or unit.
[0087] To facilitate the understanding of the technical solutions provided in the embodiments of the present application, some key terms used in the embodiments of the present application are explained here first:
[0088] Digital Image Processing refers to the process of processing digital image data using computer algorithms and techniques. Digital image processing can be applied to various fields, including computer vision, medical imaging, remote sensing images, image synthesis, image enhancement, image segmentation, and object detection, etc. Digital image processing techniques use mathematics and algorithms to change the characteristics, quality, or information content of an image in order to extract useful information or improve the visualization effect of the image. These techniques can be used to remove noise in the image, adjust the contrast and brightness of the image, enhance the details of the image, change the color space of the image, rotate and scale the image, edge detection, and object extraction, etc. The basic steps of digital image processing techniques include image acquisition, image preprocessing, image enhancement, feature extraction, image segmentation, and object recognition, etc. In the processing process, commonly used algorithms and techniques include filtering, transformation, edge detection, morphological operations, feature extraction, and machine learning and deep learning, etc.
[0089] Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems that can perceive the environment, acquire knowledge, and use knowledge to achieve optimal results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, pre-trained models, also known as large models or foundation models, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. Artificial intelligence software technology mainly includes several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0090] Computer Vision (CV) is a science that studies how to enable machines to "see". More specifically, it refers to using cameras and computers to replace human eyes for tasks such as target recognition and measurement in machine vision, and further performing image processing to make the images processed by the computer more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc. technologies, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0091] To solve the above problems, the embodiments of the present application provide an image classification method, a model training method, a device, and an electronic device, which can improve the accuracy of image classification.
[0092] The methods provided by the embodiments of the present application can be applied to different scenarios, including but not limited to scenarios such as artificial intelligence, medical detection, and intelligent transportation.
[0093] Refer to Figure 1 , Figure 1 which is a schematic diagram of an optional implementation environment provided by the embodiments of the present application. This implementation environment includes a terminal 101 and a server 102, where the terminal 101 and the server 102 are connected through a communication network.
[0094] The terminal 101 can be a mobile phone, a computer, an intelligent voice interaction device, an intelligent household appliance, a vehicle-mounted terminal, etc., but is not limited thereto. The terminal 101 and the server 102 can be directly or indirectly connected through wired or wireless communication methods, and the embodiments of the present application do not limit this here. Optionally, the terminal 101 can pre-store sample images, at least one type of target noise image, and images to be classified, or download the sample images, target noise images, and images to be classified from an open-source information source. The terminal 101 can send the sample images, target noise images, and images to be classified to the server 102.
[0095] The server 102 can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Additionally, the server 102 can also be a node server in a blockchain network. Optionally, the server 102 can pre-store a target classification model. When obtaining the sample images and at least one type of target noise image, the sample images and the target noise images can be input into the target classification model for classification to obtain the first category confidence score corresponding to the sample images and the second category confidence score corresponding to the facial noise images; determine the target loss according to the first category confidence score and the second category confidence score, and train the target classification model according to the target loss; then, after obtaining the images to be classified, the images to be classified can be input into the trained target classification model for classification.
[0096] Exemplarily, the terminal 101 may pre-store sample images and at least one type of target noise images, while the server 102 may pre-store images to be classified; the terminal 101 may send the stored sample images and at least one type of target noise images to the server 102. After the server 102 receives these image data, the sample images and the target noise images may be respectively input into a target classification model for classification to obtain a first category confidence score corresponding to the sample images and a second category confidence score corresponding to the target noise images; then, the server 102 may determine a target loss according to the first category confidence score and the second category confidence score, and train the target classification model according to the target loss, thereby completing the process of training the target classification model. If the terminal 101 also has an image classification task, that is, the terminal 101 also stores images to be classified, or downloads images to be classified from an open source information source, the terminal 101 may continue to send the images to be classified to the server 102, so that the server 102 may input the images to be classified into the trained target classification image for classification, and return the classification result to the terminal 101, thereby completing the image classification task. Since the target loss is used to constrain the magnitude relationship between the first category confidence score and the second category confidence score, and the magnitude relationship is determined according to the noise addition position, even if noise is added to at least one of the background area or the semantic area of the sample image, the category confidence score output by the target classification model can be constrained based on the target loss, which helps to prevent the target classification model from erroneously fitting the background as the key feature for classification. When subsequently obtaining an image to be classified and inputting the image to be classified into the trained target classification model for classification, the accuracy of image classification can be effectively improved.
[0097] Refer to Figure 2 , Figure 2 FIG. is an optional flowchart of an image classification method provided by an embodiment of the present application. The image classification method may be executed by a terminal, or may also be executed by a server, or may also be executed in cooperation by a terminal and a server. In the embodiment of the present application, taking the image classification method being executed by the server as an example for illustration, the image classification method includes but is not limited to the following steps 201 to step 204.
[0098] Step 201, obtain a sample image and at least one type of target noise image.
[0099] In a possible implementation manner, the sample image may be an image with a classification label for indicating a classification type. The classification labels of these sample images may be obtained by manually annotating the images, or may be generated by pre-identifying and annotating the images using a machine learning model.
[0100] Refer to Figure 3 , Figure 3Schematic diagram of the target noise image provided by the embodiment of the present application. In a possible implementation, the target noise image is obtained by adding noise to the sample image. The noise addition positions of different types of target noise images are different. The noise addition positions include the background area or the semantic area of the sample image. Therefore, the target noise image includes at least two types. One type of target noise image is the image obtained by adding noise to the background area of the sample image, that is, the target noise image with a perturbed background (such as Figure 3 shown in the first target noise schematic diagram); Another type of target noise image is the image obtained by adding noise to the semantic area of the sample image, that is, the target noise image with a perturbed semantics (such as Figure 3 shown in the second target noise schematic diagram). Among them, when the sample image carries a classification label, the target noise image can also carry the corresponding classification label. In addition, these sample images and target noise images can be selected from different application fields. For example, they can be natural images, medical images, biological images, etc. Specifically, for a medical image classification task, these sample images can include images of different lesions or organs. At this time, the semantic area of the image refers to the organ or lesion area of interest in the medical image. The semantic area can be the specific name of the organ or the type of medical lesion. For example, in a lung image classification task, the semantics can be lungs, nodules, tumors, etc.; The background area usually refers to the area part of the image other than the area of interest, that is, the part of the medical image other than the organ or area of interest. At this time, the background area can include the surrounding tissues of the organ, other organs, air, or noise generated by the scanning device, etc.
[0101] It should be noted that the semantic area of an image usually refers to the area containing the key features related to the target object in the image. These features can usually be used to identify and classify the target object. For example, for an image with a target object of a car, the image contains a car, then the wheels, body shape, headlights, doors, etc. of the car can be used as the features of the semantic area of the image. The background area of an image usually refers to the area irrelevant to the target object. The features of the background area usually do not contain information closely related to the semantic target object classification. For example, for an image with a target object of a car, the features of the background area can be background texture, color, illumination, etc.
[0102] In a possible implementation, the semantic regions and background regions of the divided images in the target noise image can be achieved through manual annotation. Specifically, it can rely on manual selection of the target objects in the image as semantic regions according to prior knowledge by humans, and define the regions unrelated to the target objects as background regions; or, the pixels in the image can be classified into different categories through a semantic segmentation algorithm (such as a semantic segmentation network model), that is, the pixels used to form the target object are marked as the pixels of the semantic region, while the pixels unrelated to the target object (such as the background) are marked as the pixels of the background region; or, it can be achieved through a target detection algorithm (such as object bounding box detection), locate and annotate the target objects in the image, and then the regions occupied by the target objects can be defined as semantic regions, and the regions other than the target objects are defined as background regions.
[0103] It should be noted that the sample images and the target noise images can be pre-stored in the server, or can be pre-stored in the terminal. Or, the sample images and the target noise images can be downloaded from open-source information sources. In addition, after obtaining the sample images, the terminal or the server can add noise to the corresponding regions of the sample images to generate the target noise images.
[0104] In a possible implementation, after obtaining the sample images and the target noise images, preprocessing before classification can be performed on the sample images and the target noise images. For example, first, the sizes of all sample images and target noise images are unified. Based on the pixel size, the sample images and the target noise images are adjusted to square images of the same size. Then, each image is cut to emphasize the region of interest in the image. Subsequently, the square region corresponding to the region of interest is extracted from the original image to form a new sample image or a new target noise image. This preprocessing method initially cuts the image through image segmentation technology to find the location of the region of interest, and then performs multiple iterative cuts on the region of interest, which can enhance the segmentation effect, avoid the limitations of human tissues and image sources, be applicable to the segmentation of multiple images, and improve the data versatility of the sample images and the target noise images.
[0105] Step 202: Input the sample image and the target noise image into the target classification model for classification respectively, to obtain the first category confidence score corresponding to the sample image and the second category confidence score corresponding to the target noise image.
[0106] In a possible implementation, the target classification model can be an open-source pre-trained model, that is, a model pre-trained on a large-scale dataset. For example, for the natural image classification task, the target classification model can be pre-trained on a natural image atlas and has good feature extraction ability for natural images; for the medical image classification task, the target classification model can be trained on a medical image atlas and has good medical image feature extraction ability. Additionally, the target classification model can also be a model with initialized weights, that is, a model that has not undergone any training process, and is trained using sample images and target noise images to meet the corresponding image classification requirements.
[0107] In a possible implementation, the class confidence score is a measure of the certainty of the target classification model for each prediction result. Specifically, the class confidence score can be the confidence of the target classification model in predicting a certain class, that is, the class confidence score can be the maximum value among the probability values of each class output by the target classification model.
[0108] In a possible implementation, the sample image can be first input into the target classification model for training and classification prediction to obtain the classification result of the sample image and the corresponding first class confidence score. At the same time, the target classification model completes preliminary training. Then, the target noise image is input into the target classification model for classification prediction to obtain the classification result of the target noise image and the corresponding second class confidence score. Alternatively, the sample image and the target noise image can be spliced, and then the spliced image is input into the target classification model for classification prediction to obtain the first class confidence score corresponding to the sample image and the second class confidence score corresponding to the target noise image. For some image classification tasks with high sample data acquisition costs, such as medical image classification tasks, due to the limitations of medical image imaging conditions, medical images are prone to problems such as noise, artifacts, and poor resolution, and medical images usually contain complex tissue structures, pathological phenomena, physiological changes, etc. that affect the classification difficulty. Therefore, the accuracy of medical image classification needs to be based on a large amount of sample data. However, the acquisition cost of medical image data is relatively high, resulting in a scarcity of medical image sample data and making it difficult to classify medical images accurately and at low cost. The image classification method provided in the embodiments of the present application expands the corresponding noise images on the basis of the sample images as training sample data, increases the data volume of the training sample data, helps to improve the training effect of the target classification model, and at the same time can reduce the data dependence on the sample images.
[0109] In a possible implementation, the feature learned by the target classification model can be determined to be related to noise or to have an obvious response to noise by comparing the first category confidence score and the second category confidence score. If the difference between the first category confidence score and the second category confidence score is large, it indicates that the target classification model may be affected by noise, and subsequent processing needs to be adjusted to improve the classification effect.
[0110] Step 203: Determine the target loss according to the first category confidence score and the second category confidence score, and train the target classification model according to the target loss.
[0111] In a possible implementation, the target loss is used to constrain the magnitude relationship between the first category confidence score and the second category confidence score, and the magnitude relationship is determined according to the noise addition position. Since the category confidence score output by the target classification model for the input image directly affects the learning degree of the target classification model for the input image, by adjusting the magnitude relationship between the first category confidence score and the second category confidence score for different noise addition positions respectively, the learning ability of the target classification model is different when classifying and predicting target noise images with different noise addition positions. Under the guidance of the target loss, the response of the target classification model to the features of the image background region can be suppressed, the situation of fitting the background as the key feature can be effectively reduced, and the recognition ability of the model for the features of the semantic region can be improved.
[0112] Refer to Figure 4 , Figure 4 is a schematic diagram of the process for determining the target loss provided by an embodiment of the present application. As Figure 4 shown, when the target noise image input to the target classification model is a noise image with perturbed semantics (target noise image A), that is, generated by adding noise to the semantic region of the sample image, and the second category confidence score corresponding to the target noise image A is the second category confidence score A, therefore, the target loss can be determined according to the first category confidence score corresponding to the sample image and the second category confidence score A corresponding to the target noise image A. At this time, the target loss can be used to constrain the first category confidence score to be greater than the second category confidence score. For example, the first category confidence score can be constrained to remain unchanged, while the second category confidence score is decreased, that is, when the same sample image and target noise image are input subsequently, the obtained first category confidence score remains unchanged, and the obtained second category confidence score is lower than the second category confidence score obtained before training.
[0113] Refer to Figure 5 , Figure 5 is a schematic diagram of the process for determining the target loss provided by another embodiment of the present application. As Figure 5As shown, when the target noise image input to the target score model is a noise image with a perturbed background (target noise image B), that is, generated by adding noise to the background region of the sample image, and the second-class confidence score corresponding to the target noise image B is the second-class confidence score B. Therefore, the target loss can be determined according to the first-class confidence score corresponding to the sample image and the second-class confidence score B corresponding to the target noise image B. At this time, the target loss can be used to constrain the first-class confidence score to be equal to the second-class confidence score.
[0114] Referring to Figure 6 , Figure 6 is a schematic diagram of the process for determining the target loss provided by another embodiment of the present application. As Figure 6 shown, when the target noise image input to the target classification model has two types, that is, a noise image with a perturbed semantics (target noise image A) and a noise image with a perturbed background (target noise image B) are input simultaneously, then the target loss can be determined by the first-class confidence score, the second-class confidence score A, and the second-class confidence score B. At this time, the target loss can be used to constrain the first-class confidence score to be greater than the second-class confidence score corresponding to the noise image with a perturbed semantics (i.e., the second-class confidence score A), and at the same time constrain the second-class confidence score corresponding to the noise image with a perturbed background (i.e., the second-class confidence score B) to be equal to the first-class confidence score. In addition, in addition to being able to constrain the first-class confidence score to be equal to the second-class confidence score B, the target loss can also constrain the first-class confidence score to be less than the second-class confidence score B, that is, increase the second-class confidence score B, so as to improve the learning ability of the subsequent target classification model when facing the noise image with a perturbed background. Thus, calibration of the confidence estimation of the target classification model is achieved. Furthermore, the target classification model can be trained by the target loss, so that the target classification model can maintain the class confidence score (confidence level) of the classification prediction unchanged when the background region of the image is perturbed by noise, and reduce the class confidence score (confidence level) of the classification prediction when facing the semantic region of the image perturbed by noise. Therefore, the target loss can help the target classification model identify and focus on the features of the semantic region, reduce the response to the image background, and help avoid the target classification model wrongly fitting the background region as the key feature of the classification, thereby improving the classification performance.
[0115] In a possible implementation, the target loss can be determined according to the input situation of the target noise image. Specifically, the degree of constraint of the target loss on the magnitude relationship between the first-class confidence score and the second-class confidence score can depend on the type of the target noise image input to the target classification model. For example, as more target noise images with perturbed semantics are input to the target classification model, the degree of constraint of the target loss on the second-class confidence score corresponding to the target noise image with perturbed semantics is higher, that is, the second-class confidence score corresponding to the target noise image with perturbed semantics is lower. Since the second-class confidence score obtained by inputting the target noise image with perturbed semantics to the trained target classification model for classification prediction has been constrained and decreased, and then the first-class confidence and the decreased second-class confidence score are adjusted by the target loss again, so that the target loss constrains the magnitude relationship between the two again, enhancing the constraint effect again. When subsequent target noise images with perturbed semantics are input to the target classification model, the obtained second-class confidence score is lower, thereby being able to reduce the response of the target classification model to the image background and helping to improve the classification performance of the model.
[0116] Step 204: Obtain the image to be classified, and input the image to be classified into the trained target classification model for classification.
[0117] In a possible implementation, the image to be classified can belong to the same application field as the sample image. For example, both the image to be classified and the sample image can belong to natural images or medical images. Additionally, the application field to which the image to be classified belongs can be different from the application field to which the sample image belongs. For example, the image to be classified belongs to medical images while the sample image belongs to natural images.
[0118] In a possible implementation, the image to be classified can be pre-stored in the server. When the server performs image classification on the internally stored image data, it can directly call the internally stored image data as the image to be classified and input it into the trained target classification model for classification, and can also obtain the corresponding classification result and sort the corresponding image data based on the classification result. Additionally, the image to be classified can be obtained by being sent from a terminal. For example, the image to be classified can be obtained by being captured by the imaging device of the terminal, and then the terminal sends the image to be classified to the server. Thus, the server can input the image to be classified into the target classification model for classification prediction and return the classification result of the image to be classified to the terminal.
[0119] Refer to Figure 7 , Figure 7 which is a schematic diagram of the image classification process provided by the embodiments of this application. As Figure 7As shown, first obtain a sample image and at least one type of target noise image, input the sample image and the target noise image into a target classification model for classification respectively, obtain the first category confidence score corresponding to the sample image and the second category confidence score corresponding to the target noise image. By supplementing the target noise image on the basis of the sample image, the data volume of the training samples is increased, which helps to improve the classification performance of the target classification model. Then, determine the target loss through the first category confidence score and the second category confidence score, and train the target classification model according to the target loss. The target loss is used to constrain the magnitude relationship between the first category confidence score and the second category confidence score, and the magnitude relationship is determined based on the noise addition position. If there are multiple types of target noise images input to the target classification model, there will be corresponding multiple second category confidence scores. For example, as Figure 7 shown, the target noise images input to the target classification model include a target noise image with noise added to the background area (target noise image A) and a target noise image with noise added to the semantic area (target noise image B). Therefore, the second category confidence score (second category confidence score A) of the target noise image with noise added to the background area (target noise image A) and the second category confidence score (second category confidence score B) of the target noise image with noise added to the semantic area (target noise image B) can be obtained. Thus, in the process of determining the target loss, the target loss can be comprehensively determined by combining the first category confidence score and the two second category confidence scores (i.e., second category confidence score A and second category confidence score B), and the magnitude relationship between the first category confidence score and the two second category confidence scores is constrained by the target loss. Specifically, the target loss constrains that the second category confidence score corresponding to the target noise image of the same category as target noise image A is equal to the first category confidence score corresponding to the sample image, and constrains that the second category confidence score corresponding to the target noise image of the same category as target noise image B is less than the first category confidence score. Therefore, the target classification model can better understand the characteristics and meanings of both the semantic area and the background area in the sample image. Even if noise is added to at least one of the background area or the semantic area of the sample image, the category confidence score output by the target classification model can be constrained based on the target loss, which helps to prevent the target classification model from wrongly fitting the background as the key feature of classification and can effectively improve the accuracy of image classification. Therefore, after the target classification model is trained, the obtained image to be classified is input into the trained target classification model for classification, and a more accurate image classification result can be obtained.
[0120] In a possible implementation manner, during the process of determining the target loss, the first loss may be determined first according to the first category confidence score and the second category confidence score corresponding to the first target noise image; then, the second loss may be determined according to the first category confidence score and the second category confidence score corresponding to the second target noise image; and then, the target loss is determined according to the first loss and the second loss.
[0121] In a possible implementation manner, as Figure 3 shown, at least one type of target noise image includes a first target noise image obtained by adding noise to the background region and a second target noise image obtained by adding noise to the semantic region. Referring to Figure 8 , Figure 8 is a schematic diagram of the magnitude relationship between the first category confidence score and the second category confidence score provided by an embodiment of the present application. As Figure 8 shown, the first target noise image is the target noise image with the background region disturbed, the second category confidence score corresponding to the first target noise image is the second category confidence score A, and the first loss is used to constrain the first category confidence score to be equal to the second category confidence score A corresponding to the first target noise image. The second target noise image is the target noise image with the semantic region disturbed, the second category confidence score corresponding to the second target noise image is the second category confidence score B, and the second loss is used to constrain the first category confidence score to be greater than the second category confidence score B corresponding to the second target noise image.
[0122] In a possible implementation, since the focus of the image classification task is to identify and understand the target objects in the image, and the semantic regions usually contain key features related to the target objects in the image, by focusing on the features of the semantic regions, the target classification model can more accurately extract and utilize the information related to the classification of the target objects, and can more accurately improve the classification performance. The background region usually contains a large amount of irrelevant information or clutter signals, and the features of the background region may interfere with the classification judgment of the target classification model. For example, factors such as texture, color, and illumination in the background region of the image may cause misjudgment of the classification model in the European and American versions. Therefore, it is necessary to guide the target classification model to ignore the features of the background region, which can reduce the influence of irrelevant information and enable the target classification model to focus more on the feature extraction and classification judgment of the semantic region (i.e., the target object) of the image. In addition, in the same application field or the same category of image datasets, the features of the background region vary greatly between different images, while the features of the semantic region usually tend to be stable or consistent. Therefore, guiding the model to ignore the features of the background region can help improve the robustness of the target classification model and the reliability of the model's classification judgment under different backgrounds. Noise is added to the background region of the first target noise image, while the features of the semantic region are not disturbed, that is, it can be considered that the key features related to the target object in the first target noise image are not affected by the noise. Therefore, by constraining the first-class confidence score to be equal to the second-class confidence score corresponding to the first target noise image through the first loss, it is possible to avoid the features of the background region from having an excessive interference on the classification result of the target classification model, and enable the target classification model to focus on the features of the target object, that is, the features of the semantic region, during the classification inference process. In contrast, since the semantic region in the second target noise image is disturbed by noise, that is, the key features related to the target object are affected by the noise, it is difficult for the target classification model to perform accurate classification inference on the second target noise image. Therefore, by using the second loss to constrain the decrease of the second-class confidence score corresponding to the second target noise image, so that the first-class confidence score is greater than the second-class confidence score corresponding to the second target noise image, it is possible to avoid the target classification model from focusing on the features of the background region or the features of the semantic region disturbed by noise, and prevent the target classification model from misfitting the background region as the key features, thereby realizing the estimation and calibration of the class confidence score of the target classification model and improving the classification performance of the target classification model.
[0123] Refer to Figure 9 , Figure 9 FIG. is a schematic diagram of the process for determining the target loss provided by another embodiment of the present application. As Figure 9 shown, the first loss is based on the first-class confidence score Conf(x) corresponding to the sample image x and the second-class confidence score Conf(x' s corresponding to the first target noise image x' s)It is determined that specifically, the functional formula of the first loss L1 can be shown as the following formula (1):
[0124] L1 = max[0, Conf(x’ s ) - Conf(x)] (1).
[0125] It can be seen that in the process of minimizing the loss, the difference between the confidence score Conf(x) of the first category and the confidence score Conf(x’ s ) of the second category is restricted by using the maximum value return function and the limit value 0, so as to realize the regularization of the first loss L1. In order to make the first loss L1 reach the minimum value of 0, the target classification model needs to continuously reduce the confidence score Conf(x’ s ) of the second category until the confidence score Conf(x) of the first category is greater than the confidence score Conf(x’ s ) of the second category. In addition, the magnitude relationship between the confidence score Conf(x) of the first category and the confidence score Conf(x’ s ) of the second category can be constrained by adjusting the value of the limit.
[0126] The second loss is determined according to the confidence score Conf(x) of the first category corresponding to the sample image x and the confidence score Conf(x’ b ) of the second category corresponding to the second target noise image x’ b ). Specifically, the functional formula of the second loss L1 can be shown as the following formula (2):
[0127] L2 = [Conf(x’ b ) - Conf(x)] 2 (2).
[0128] It can be seen that in the process of minimizing the loss, the difference between the confidence score Conf(x) of the first category and the confidence score Conf(x’ b ) of the second category is restricted by using the sum of squares formula, so as to realize the regularization of the second loss L2. In order to make the second loss L2 reach the minimum value of 0, it is necessary to guide the target classification model to continuously adjust the confidence score Conf(x’ b ) of the second category so that the confidence score Conf(x) of the first category is equal to the confidence score Conf(x’ b ) of the second category. In addition, the magnitude relationship between the confidence score Conf(x) of the first category and the confidence score Conf(x’ b ) of the second category can be constrained by using the maximum value return function and the limit value.
[0129] When the first loss L1 and the second loss L2 are obtained, the target loss L can be determined by synthesizing the first loss L1 and the second loss L2. For example, the first loss L1 and the second loss L2 can be added together to determine the target loss L. Specifically, the functional formula of the target loss L can be shown as the following formula (3):
[0130] L = max[0, Conf(x’ s ) - Conf(x)] + [Conf(x’ b ) - Conf(x)] 2 = L1 + L2 (3).
[0131] It can be seen that in the process of minimizing the loss, it is necessary to guide the target classification model to continuously lower the confidence score Conf(x’ s ) of the second category corresponding to the first target noise image, so that the confidence score Conf(x) of the first category is greater than the confidence score Conf(x’ s ) of the second category. At the same time, it is also necessary to guide the target classification model to continuously adjust the confidence score Conf(x’ b ) of the second category corresponding to the second target noise image, so that the confidence score Conf(x) of the first category is equal to the confidence score Conf(x’ b ) of the second category. Therefore, the target loss L can be used to constrain the confidence of the category of the sample image with the background area perturbed to remain unchanged, while the confidence of the category of the sample image with the semantic area perturbed decreases, so as to realize the calibration of the confidence estimation of the target classification model, effectively avoid the target classification model focusing on the features of the background area, reduce the situation that the target classification model misfits the features of the background area as the key features of image classification, and improve the classification performance of the target classification model.
[0132] In a possible implementation, the first loss can also be used to constrain the confidence score of the second category corresponding to the first target noise image to be increased to be equal to the confidence score of the first category, or to constrain the confidence score of the second category corresponding to the first target noise image to be decreased to be equal to the confidence score of the first category; and the second loss can also be used to constrain the confidence score of the second category corresponding to the second target noise image to decrease, including the case where the confidence score of the second category is less than the confidence score of the first category, that is, when the confidence score of the second category corresponding to the second target noise image is less than the confidence score of the first category, the confidence score of the second category is still decreased, further suppressing the target classification model from focusing on the features of the background area.
[0133] In a possible implementation, the first loss can also be used to constrain the confidence score of the first category to be less than the confidence score of the second category corresponding to the first target noise image, that is, the first loss can be used to increase the confidence score of the second category corresponding to the first target noise image.
[0134] In a possible implementation, the background region and the semantic region are determined through semantic segmentation. In the process of determining the target loss through the first loss and the second loss, the first segmentation confidence of the background region and the second segmentation confidence of the semantic region can be obtained first; then, normalization is performed according to the first segmentation confidence and the second segmentation confidence to obtain the first weight of the background region and the second weight of the semantic region; then, the second weight is used as the weight of the first loss, and the first weight is used as the weight of the second loss, and the first loss and the second loss are weighted to obtain the target loss.
[0135] In a possible implementation, the first segmentation confidence of the background region may refer to the confidence of the background region determined in the process of semantic segmentation of the sample image. The first segmentation confidence can be obtained by calculating the average value of the confidences corresponding to each pixel point in the background region of the sample image, or by taking the maximum or minimum value of the confidences corresponding to each pixel point in the background region of the sample image. Correspondingly, the second segmentation confidence of the semantic region may refer to the confidence of the semantic region determined in the process of semantic segmentation of the sample image. The second segmentation confidence can be obtained by calculating the average value of the confidences corresponding to each pixel point in the semantic region of the sample image, or by taking the maximum or minimum value of the confidences corresponding to each pixel point in the semantic region of the sample image. It should be noted that the semantic segmentation of the sample image can be realized by manual division or by a semantic segmentation algorithm.
[0136] In a possible implementation, the first weight obtained by normalizing the first segmentation confidence is used as the weight of the second loss, and the second weight obtained by normalizing the second segmentation confidence is used as the weight of the first loss to cross-adjust the first loss and the second loss to obtain the target loss. The specific weighting formula can refer to the following formula (4):
[0137] L = pL1 + qL2 (4).
[0138] Among them, L represents the target loss, L1 represents the first loss, L2 represents the second loss, p represents the first weight obtained by normalization of the second confidence, and q represents the second weight obtained by normalization of the first confidence. It is equivalent to that when the probability of being classified as the background area is higher, that is, the higher the value of the first confidence, it can be considered that the characteristics of the background area can be ignored, and the target classification model can be prevented from fitting the characteristics of the background area as the key features of the target object classification as much as possible. Therefore, using the high value of the first confidence as the weight of the second loss can effectively help the target classification model focus on the second loss and guide the target classification model to focus on the characteristics of the semantic area. When the probability of being classified as the background area is lower, it can be considered that the characteristics of the background area are more and more variable, and some irrelevant information may be classified into the semantic area, or some features related to the target object are classified into the background area. Therefore, the low value of the first confidence can be used as the weight of the second loss to balance the attention of the target classification model to the semantic area and the background area, avoid misjudgment due to excessive attention to a single data with low reliability, and can help improve the classification performance.
[0139] Similarly, when the probability of being divided into a semantic area is higher, that is, the value of the second confidence is higher, it can be considered that the second category confidence score corresponding to the second target noise image whose semantic area is not disturbed by noise is more accurate, and the second loss determined is also more accurate. Therefore, the high value of the second confidence can be used as the weight of the first loss to increase the attention of the target classification model to the first loss, so that the target classification model accelerates the convergence speed of the first loss. When the probability of being divided into a semantic area is lower, that is, the value of the second confidence is lower, it means that some irrelevant information may be divided into the semantic area, or some features related to the target object are divided into the background area. Therefore, the low value of the first confidence can be used as the weight of the second loss to balance the attention of the target classification model to the semantic area and the background area, avoid misjudgment due to excessive attention to a single data with low reliability, and help improve classification performance.
[0140] Therefore, by using the segmentation confidence of the background area and the semantic area as weights respectively and cross-configuring the losses determined based on the perturbed images of the corresponding areas, the accuracy of the sample images in semantic segmentation can be fully considered, which helps to improve the classification performance of the target classification model.
[0141] In one possible implementation, after obtaining the first weight of the background area and the second weight of the semantic area, the first weight can be used as the weight of the first loss, and the second weight can be used as the weight of the second loss. The first loss and the second loss are weighted to obtain the target loss.
[0142] In a possible implementation, after obtaining the first weight of the background region and the second weight of the semantic region, the first weight and the second weight can be calculated to obtain the weight of the first loss and the weight of the second loss respectively, and the first loss and the second loss are weighted to obtain the target loss. Specifically, the ratio of the first weight to the second weight can be used as the weight of the first loss, and the ratio of the second weight to the first weight can be used as the weight of the second loss. Alternatively, different weight coefficients can be assigned to the first weight and the second weight respectively, and the first weight and the second weight are weighted and calculated using the corresponding weight coefficients. Among them, when calculating the weight of the first loss, a weight coefficient with a larger value can be assigned to the first weight, and when calculating the weight of the second loss, a weight coefficient with a larger value can be assigned to the second weight.
[0143] In a possible implementation, after obtaining the sample image, noise addition processing can be performed on the sample image to generate different types of target noise images. Specifically, the sample image can be obtained first, and the sample image is input into the semantic segmentation model for semantic segmentation to obtain the original background image corresponding to the background region and the original semantic image corresponding to the semantic region. Then, noise is added to the original background image and the original semantic image respectively to obtain the noise background image corresponding to the original background image and the noise semantic image corresponding to the original semantic image. Then, the noise background image and the original semantic image are merged to obtain one type of target noise image, and the noise semantic image and the original background image are merged to obtain another type of target noise image.
[0144] In a possible implementation, the semantic segmentation model can be a computer vision algorithm that assigns each pixel in the image to a specific semantic category, where the semantic segmentation model can be an open-source pre-trained model. Specifically, refer to Figure 10 , Figure 10 which is the effect schematic diagram of the semantic segmentation model provided by the embodiments of the present application. As Figure 10 shown, when the sample image is input into the semantic segmentation model for semantic segmentation, the semantic segmentation model can distinguish each pixel in the image into a semantic category and a background category, integrate all the pixel points of the background category to form the original background image corresponding to the background region, and integrate all the pixel points of the semantic category to form the original semantic image corresponding to the semantic region. By dividing the semantic region and the background region of the sample image through the semantic segmentation model, it can provide a basis for subsequent realization of the target classification model to align the category confidence scores.
[0145] In a possible implementation, during the process of adding noise to the original background image and the original semantic image respectively, the first noise intensity corresponding to the original background image and the second noise intensity corresponding to the original semantic image can be determined first. Then, a first noise image with the same size as the original background image is generated according to the first noise intensity, and the first noise image is merged with the original background image to obtain a noise background image corresponding to the original background image. Next, a second noise image with the same size as the original semantic image is generated according to the second noise intensity, and the second noise image is merged with the original semantic image to obtain a noise semantic image corresponding to the original semantic image.
[0146] Refer to Figure 11 , Figure 11 which is a schematic diagram of the process of generating a target noise image provided by an embodiment of the present application. As Figure 11 shown, a background noise matrix with the same size as the original background image can be randomly generated, and the background noise matrix is added to the original background image pixel by pixel, so that a background interference image with the background area perturbed by noise can be obtained. At the same time, a semantic noise matrix with the same size as the original semantic image can also be randomly generated, and the semantic noise matrix is added to the original semantic image pixel by pixel to obtain a semantic interference image with the semantic area perturbed by noise. Then, for each sample image, the background interference image is merged with the original semantic image to generate a first target noise image with the background area perturbed by noise; the semantic interference image is merged with the original background image to generate a second target noise image with the semantic area perturbed by noise, so that two sets of training sample data expanded based on the sample image can be obtained. It should be noted that the noise intensity of adding noise to the semantic area and the background area can be controlled by adjusting the standard deviation of their respective noise matrices.
[0147] In a possible implementation, during the process of determining the first noise intensity corresponding to the original background image and the second noise intensity corresponding to the original semantic image, the original background image and the original semantic image can be input into a noise intensity prediction model to extract the first image feature of the original background image and the second image feature of the original semantic image. The first image feature and the second image feature are spliced to obtain a spliced image feature, and linear regression is performed on the spliced image feature to obtain a noise intensity prediction result. The noise intensity prediction result includes the first noise standard deviation corresponding to the original background image and the second noise standard deviation corresponding to the original semantic image, and the noise intensity prediction result can reflect the linear relationship between the spliced image feature and the noise intensity prediction result. Then, the first noise intensity corresponding to the original background image is determined according to the first noise standard deviation, and the second noise intensity corresponding to the original semantic image is determined according to the second noise standard deviation.
[0148] In a possible implementation, the first image feature and the second image feature can be subjected to scale unification processing, such as cropping, normalization, etc., to ensure that the scales of the first image feature and the second image feature are equal in a certain dimension, so as to realize the splicing of the first image feature and the second image feature. By splicing the first image feature of the original background image and the second image feature of the original semantic image, a more comprehensive spliced image feature can be formed, which helps the noise intensity prediction model to discover the relationship between the original semantic image and the original background image in terms of noise intensity, and improves the accuracy of noise intensity prediction. In addition, for different image classification tasks, the frequencies and intensities of the noise added to the semantic region and the noise added to the background region are different. By splicing the first image feature and the second image feature, it helps to improve the discrimination ability and prediction ability of the noise intensity prediction model for different noise types and intensities, and improves the generalization ability of the noise intensity prediction model.
[0149] In a possible implementation, after obtaining the noise intensity prediction result, that is, after obtaining the first noise standard deviation and the second noise standard deviation, the first noise intensity corresponding to the original background image and the second noise intensity corresponding to the original semantic image can be determined by using a pre-set regression model or a relationship look-up table between the standard deviation and the noise intensity, and by using the first mapping relationship between the first noise standard deviation and the first noise intensity of the original background image, and the second mapping relationship between the second noise standard deviation and the second noise intensity of the original semantic image. Among them, the first mapping relationship and the second mapping relationship can be established in advance through relevant data samples of background noise and semantic noise, and the corresponding noise standard deviations.
[0150] In a possible implementation, the number of noise intensity prediction results can be multiple. The noise intensity prediction results are output by the same output layer of the noise intensity prediction model, and the output layer also outputs the intensity weights corresponding to each noise intensity prediction result. Therefore, the multiple first noise standard deviations can be weighted according to the intensity weights to obtain the first weighted standard deviation; at the same time, the multiple second noise standard deviations are weighted according to the intensity weights to obtain the second weighted standard deviation; then, the first noise intensity corresponding to the original background image is determined according to the first weighted standard deviation, and the second noise intensity corresponding to the original semantic image is determined according to the second weighted standard deviation.
[0151] In a possible implementation, the intensity weights may include a first intensity weight corresponding one-to-one to the first noise standard deviation and a second intensity weight corresponding one-to-one to the second noise standard deviation. Therefore, the first weighted standard deviation can be obtained by performing weighted summation based on each first intensity weight and the corresponding first noise standard deviation. At the same time, the second weighted standard deviation can be obtained by performing weighted summation based on each second intensity weight and the corresponding second noise standard deviation. Then, the first noise intensity corresponding to the original background image can be determined according to the first weighted standard deviation, and the second noise intensity corresponding to the original semantic image can be determined according to the second weighted standard deviation.
[0152] Referring to Figure 12 , Figure 12 As shown in the schematic diagram of the process for determining the noise intensity provided by the implementation of this application, it can be seen that by inputting the original semantic image and the original background image into the noise intensity prediction model for noise prediction, multiple noise intensity prediction results and the intensity weights corresponding to the multiple noise intensity prediction results can be obtained respectively. Among them, each noise intensity prediction result includes a first noise standard deviation and a second noise standard deviation respectively. Therefore, taking the noise intensity prediction result as the reference basis, weighted summation is performed using the corresponding intensity weight and each first noise standard deviation to obtain the first weighted standard deviation. At the same time, weighted summation is continued using the corresponding intensity weight and each second noise standard deviation to obtain the second weighted standard deviation. Furthermore, the first noise intensity and the second noise intensity can be determined respectively using the first weighted standard deviation and the second weighted standard deviation.
[0153] In a possible implementation, the sample image can be first input into the semantic segmentation model, and the sample image is successively subjected to multiple scale-down operations to obtain sample images at multiple scales. Then, in the order from high to low in terms of scale, each corresponding sample image is encoded based on the encoder corresponding to each scale to obtain the target encoded image at each scale. Among them, the input of the encoder at the current scale is obtained by merging the sample image at the current scale and the target encoded image at the previous scale. Then, in the order from low to high in terms of scale, each corresponding target encoded image is decoded based on the decoder corresponding to each scale to obtain the decoded image at each scale. Among them, the input of the decoder at the current scale is obtained by merging the target encoded image at the current scale and the decoded image at the previous scale. The decoded image at the current scale is merged with the upsampled image at the previous scale to obtain a merged image, where the upsampled image is obtained by upsampling the decoded image at the previous scale. Semantic segmentation is performed on the merged image corresponding to the original scale to obtain the background image corresponding to the background region and the semantic image corresponding to the semantic region.
[0154] Referring to Figure 13 , Figure 13Schematic diagram of the architecture of the semantic segmentation model provided by the embodiments of the present application. The semantic segmentation model includes a first encoder E1, a second encoder E2, a third encoder E3, a fourth encoder E4, and first decoders D1, D2, D3, D4 corresponding to the same scales one by one. Among them, the first scale corresponding to the first encoder E1 and the first decoder D1 is the largest, the second scale corresponding to the second encoder E2 and the second decoder D2 is larger than the third scale corresponding to the third encoder E3 and the third decoder D3, and the fourth scale corresponding to the fourth encoder E4 and the fourth decoder D4 is the smallest. After the sample image is input into the semantic segmentation model, multiple scale descents can be performed successively according to the scales corresponding to each encoder and decoder to obtain sample images at multiple scales, including the sample image TL1 at the first scale, the sample image LT2 at the second scale, the sample image LT3 at the third scale, and the sample image LT4 at the fourth scale. Specifically, the semantic segmentation model includes multiple cascaded convolutional filters and max pooling layers, so as to be able to compress and reduce the scale of the input image data, reduce the amount of data processing, and be able to combine context information, that is, combine the image information of the front and back scales, which can help the semantic segmentation model learn image features; and the semantic segmentation model also includes corresponding cascaded convolutional filters and upsampling layers to restore the scale of the input image data until it is restored to an image equal to the original scale of the sample image.
[0155] The sample images at each scale need to be merged with the target encoded image output by the encoder at the previous scale and then input into the encoder at the current scale for encoding. Since the first encoder E1 is the encoder with the largest scale and also the first encoder in the semantic segmentation model, therefore, two sample images LT1 at the first scale can be merged, and the merged image can be input into the first encoder E1 for encoding to obtain the first target encoded image EP1. Among them, the first scale is the same as the original scale of the sample image, and two sample images LT1 at the first scale can be merged and used as the input of the first encoder E1.
[0156] Next, the first target encoded image EP1 is downscaled through a max pooling layer to form the first target encoded image EP1 at the second scale, then merged with the sample image LT2 at the second scale, and the merged image is input into the second encoder E2 for encoding to obtain the second target encoded image EP2; and so on. After downscaling the target encoding scale corresponding to the previous scale to the current scale, it is merged with the sample image at the current scale, and the merged image is used as the input to the encoder at the current scale until the fourth encoder E4 outputs the fourth target encoded image EP4. Since the fourth encoder E4 is the encoder with the smallest scale and also the last encoder in the semantic segmentation model, the fourth target encoded image EP4 is directly input into the fourth decoder D4 for decoding to obtain the fourth decoded image DP4 at the fourth scale. Then, the fourth decoded image DP4 at the fourth scale is upsampled to obtain the fourth decoded image DP4 at the third scale, which is then merged with the third target encoded image EP3 at the third scale, and the merged image is input into the decoder D3 at the third scale for decoding to obtain the third decoded image DP3 at the third scale; and so on. After upsampling the decoded image at the previous scale to the current scale, it is merged with the target encoded image corresponding to the current scale, and the merged image is used as the input to the encoder at the current scale for decoding until the current scale is the first scale. Specifically, after obtaining the second decoded image DP2 at the second scale, the second decoded image DP2 at the second scale is upsampled to obtain the second decoded image DP2 at the first scale, which is then merged with the target encoded image EP1 at the first scale, and the merged image is input into the decoder D1 at the first scale for decoding to obtain the first decoded image DP1 at the first scale.
[0157] Next, the fourth decoded image DP4 at the fourth scale is upsampled to the third scale and merged with the third decoded image DP3 at the third scale to obtain the merged image H3 at the third scale; the merged image H3 at the third scale is upsampled to the second scale and merged with the second encoded image DP2 at the second scale to obtain the merged image H2 at the second scale; then, the merged image H2 at the second scale is continuously upsampled to obtain the upsampled image at the first scale, which is then merged with the first decoded image DP1 at the first scale to obtain the merged image H1 at the first scale, which is the merged image H1 corresponding to the original scale. Thus, semantic segmentation of the merged image H1 can obtain a more accurate background image corresponding to the background region and a semantic image corresponding to the semantic region.
[0158] Due to the complexity and similarity of the background information and semantic information in some images, traditional image classification models or semantic segmentation models are difficult to accurately divide the background area and the semantic area, resulting in the easy fitting of background information as semantic information, poor calibration of class confidence estimation, and affecting the accuracy and reliability of image classification. Especially for the classification task of medical images, since medical images usually contain multiple tissue structures, such as muscles, bones, fats, blood vessels, etc., these tissue structures have similarities in appearance or texture, which are likely to interfere with the model and make accurate classification difficult. In addition, artifacts, image blurring, and low contrast are likely to occur in medical images, resulting in unclear boundaries between the background information and semantic information in medical images, increasing the difficulty of the model for image classification. Therefore, through multi-stage iterative encoding, decoding, and merging of sample images, target encoded images, and decoded images at multiple different scales, the data accuracy of the images can be merged, and the semantic segmentation model can perform semantic segmentation on the merged images to generate a background image corresponding to the background area and a semantic image corresponding to the semantic area, enabling the semantic segmentation model to output more accurate background images and semantic images, thereby helping to improve the inference accuracy of subsequent classification of sample images.
[0159] In a possible implementation, during the encoding of sample images at each scale, the corresponding sample images can be encoded respectively based on the encoders corresponding to each scale to obtain the original encoded images at each scale; then, the spatial attention matrix and the position attention matrix corresponding to the sample images at each scale are determined respectively, and the spatial attention matrix and the position attention matrix corresponding to each scale are merged to obtain a merged attention matrix; then, the merged attention matrices corresponding to multiple scales are subjected to max pooling to obtain a target attention matrix; and the original encoded images at each scale are adjusted respectively based on the target attention matrix to obtain the target encoded images at each scale.
[0160] In a possible implementation, the attention matrix is a matrix used to represent the correlation degree between different elements (each pixel point in the image). By calculating the attention weights, the degree of attention to different pixel points in the original encoded image can be achieved, so that the image can be adjusted using the attention matrix, enabling the model to focus on the regions of interest related to the semantic segmentation task. Specifically, using the attention matrix can make the model ignore the background or noise in the image, focus on the boundaries of the target objects, and enhance the key features in the image, thereby improving the accuracy and precision of semantic segmentation. Among them, the attention matrix can include a spatial attention matrix, a position attention matrix, a channel attention matrix, a multi-head attention matrix, a cross attention matrix, etc.
[0161] In a possible implementation, the spatial attention matrix can be used to capture the relationships between different regions in the original encoded image, so as to facilitate the distinction between the semantic regions and background regions in the original encoded image. The spatial attention matrix focuses on calculating the relative positions and distances of each pixel point in the original encoded image in space to obtain the importance weights between each pixel, so that the boundary shape of the target object in the original encoded image can be captured through the spatial attention matrix. The position attention matrix can be used to calculate the dependence relationships between each pixel point in different positions in the original encoded image to capture the correlation between each pixel point. Therefore, the positions of any two similar features can contribute to each other for improvement (i.e., improving the weights), regardless of the distance between them. Specifically, this can be achieved by introducing position encoding in the self-attention mechanism.
[0162] Referring to Figure 14 , Figure 14 is a schematic diagram of the process for generating the target encoded image provided by the embodiments of the present application. As Figure 14 shown, for the medical image classification task, each scale's encoder encodes the corresponding medical sample image, and the original encoded images of each scale can be obtained. It should be noted that when the non-first encoder encodes the medical sample image, the medical sample image of the corresponding scale needs to be merged with the target encoded image of the previous scale and then the merged image is encoded. Then, after each scale's medical sample image is input into the convolutional neural network to extract features, a medical sample feature map can be obtained. The medical sample feature map is respectively input into the spatial attention module and the position attention module, so that the spatial attention matrix and the position attention matrix corresponding to each scale's medical sample image can be extracted. Specifically, the spatial attention module can use the convolutional layer to perform convolution extraction on the medical sample feature map to obtain a convolutional feature map, multiply the convolutional feature map by its transposed feature map, and then use the softmax function to process to obtain the spatial attention matrix. The position attention module can also use the convolutional layer to perform convolution features on the medical sample feature image to obtain a convolutional feature map, multiply the convolutional feature map by its transposed feature map, and then use the softmax function to process to obtain the spatial attention feature map. Then, multiply the transposed spatial attention feature map by the convolutional feature map, and after the multiplication result is restored to the same scale as the medical sample feature map, add it to the medical sample feature map to obtain the position attention matrix.
[0163] After obtaining the spatial attention matrix and the position attention matrix at each scale, the spatial attention matrix and the position attention matrix can be superimposed and merged based on each scale to generate multiple merged attention matrices, and then the multiple merged attention matrices are respectively subjected to max pooling to obtain the target attention matrix. Furthermore, the original encoded images at the corresponding scales are adjusted by using the target attention matrices at each scale, and thus the target encoded images at each scale can be obtained. Since both the spatial attention matrix and the position attention matrix focus on exploring the relationships between each pixel point in the image in terms of spatial position, for the high-difficulty medical image classification task, medical images usually contain key features such as complex tissue structures and lesions. However, the randomness of the positions and spatial distributions of these key features is strong, and at the same time, medical images also contain a large amount of noise. Therefore, by comprehensively adjusting the original encoded images using the spatial attention matrix and the position attention matrix, it can assist the model to better understand and distinguish the semantic information and background information in the original encoded images, improve the model's detection and segmentation capabilities for target objects and regions of interest, reduce the model's attention to noise or irrelevant regions, and thus improve the accuracy and robustness of the image task.
[0164] Referring to Figure 15 , Figure 15 FIG. is an optional flowchart of the model training method provided by an embodiment of the present application. This model training method can be executed by a terminal, or can also be executed by a server, or can also be executed in cooperation between the terminal and the server. In the embodiment of the present application, taking the model training method being executed by the server as an example for illustration, this model training method includes but is not limited to the following steps 1501 to step 1503.
[0165] Step 1501, obtain a sample image and at least one type of target noise image, where the target noise image is obtained by adding noise to the sample image.
[0166] In a possible implementation manner, the noise addition positions of different types of target noise images are different, and the noise addition positions include the background region or the semantic region of the sample image.
[0167] In a possible implementation, after obtaining the sample image, noise addition processing can be performed on the sample image to generate different types of target noise images. Specifically, the sample image can be obtained first, and the sample image is input into a semantic segmentation model for semantic segmentation to obtain the original background image corresponding to the background region and the original semantic image corresponding to the semantic region. Then, noise is added to the original background image and the original semantic image respectively to obtain the noise background image corresponding to the original background image and the noise semantic image corresponding to the original semantic image. Then, the noise background image and the original semantic image are merged to obtain one type of target noise image, and the noise semantic image and the original background image are merged to obtain another type of target noise image.
[0168] In a possible implementation, in the process of adding noise to the original background image and the original semantic image respectively, the first noise intensity corresponding to the original background image and the second noise intensity corresponding to the original semantic image can be determined first. Then, a first noise image with the same size as the original background image is generated according to the first noise intensity, and the first noise image is merged with the original background image to obtain the noise background image corresponding to the original background image. Then, a second noise image with the same size as the original semantic image is generated according to the second noise intensity, and the second noise image is merged with the original semantic image to obtain the noise semantic image corresponding to the original semantic image.
[0169] In a possible implementation, in the process of determining the first noise intensity corresponding to the original background image and the second noise intensity corresponding to the original semantic image, the original background image and the original semantic image can be input into a noise intensity prediction model to extract the first image feature of the original background image and the second image feature of the original semantic image. The first image feature and the second image feature are spliced to obtain a spliced image feature, and linear regression is performed on the spliced image feature to obtain a noise intensity prediction result. Among them, the noise intensity prediction result includes the first noise standard deviation corresponding to the original background image and the second noise standard deviation corresponding to the original semantic image. The noise intensity prediction result can reflect the linear relationship between the spliced image feature and the noise intensity prediction result. Then, the first noise intensity corresponding to the original background image is determined according to the first noise standard deviation, and the second noise intensity corresponding to the original semantic image is determined according to the second noise standard deviation.
[0170] In a possible implementation, the number of noise intensity prediction results can be multiple. The noise intensity prediction results are output by the same output layer of the noise intensity prediction model, and the output layer also outputs the intensity weights corresponding to each noise intensity prediction result. Therefore, the multiple first noise standard deviations can be weighted according to the intensity weights to obtain a first weighted standard deviation. At the same time, the multiple second noise standard deviations are weighted according to the intensity weights to obtain a second weighted standard deviation. Then, the first noise intensity corresponding to the original background image is determined according to the first weighted standard deviation, and the second noise intensity corresponding to the original semantic image is determined according to the second weighted standard deviation.
[0171] In a possible implementation, the intensity weights can include a first intensity weight corresponding one-to-one to the first noise standard deviation and a second intensity weight corresponding one-to-one to the second noise standard deviation. Therefore, the first weighted standard deviation can be obtained by weighted summation of each first intensity weight and the corresponding first noise standard deviation. At the same time, the second weighted standard deviation is obtained by weighted summation of each second intensity weight and the corresponding second noise standard deviation. Then, the first noise intensity corresponding to the original background image can be determined according to the first weighted standard deviation, and the second noise intensity corresponding to the original semantic image can be determined according to the second weighted standard deviation.
[0172] Step 1502: Input the sample image and the target noise image into the target classification model for classification respectively, to obtain a first category confidence score corresponding to the sample image and a second category confidence score corresponding to the target noise image.
[0173] Step 1503: Determine the target loss according to the first category confidence score and the second category confidence score, and train the target classification model according to the target loss.
[0174] In a possible implementation, the target loss is used to constrain the magnitude relationship between the first category confidence score and the second category confidence score. The magnitude relationship is determined according to the noise addition position. Since the category confidence score output by the target classification model for the input image directly affects the learning degree of the target classification model for the input image, by adjusting the magnitude relationship between the first category confidence score and the second category confidence score for different noise addition positions respectively, the learning ability of the target classification model when classifying and predicting target noise images with different noise addition positions is different. Under the guidance of the target loss, the response of the target classification model to the features of the image background area can be suppressed, effectively reducing the situation of fitting the background as key features, and improving the recognition ability of the model for the features of the semantic area.
[0175] In a possible implementation, during the process of determining the target loss, the first loss may be determined first according to the first category confidence score and the second category confidence score corresponding to the first target noise image; then, the second loss may be determined according to the first category confidence score and the second category confidence score corresponding to the second target noise image; then, the target loss is determined according to the first loss and the second loss.
[0176] In a possible implementation, the background region and the semantic region are determined by semantic segmentation. During the process of determining the target loss by the first loss and the second loss, the first segmentation confidence of the background region and the second segmentation confidence of the semantic region may be obtained first; then, normalization is performed according to the first segmentation confidence and the second segmentation confidence to obtain the first weight of the background region and the second weight of the semantic region; then, the second weight is used as the weight of the first loss, and the first weight is used as the weight of the second loss, and the first loss and the second loss are weighted to obtain the target loss.
[0177] Through the image classification method and the model training method provided by the above embodiments, since the target loss is used to constrain the magnitude relationship between the first category confidence score and the second category confidence score, and the magnitude relationship is determined according to the noise addition position, therefore, even if noise is added to at least one of the background region or the semantic region of the sample image, the category confidence score output by the target classification model can be constrained based on the target loss, and it is possible to align the category confidence scores of the target classification model for both the background region and the semantic region, which helps to prevent the target classification model from erroneously fitting the background as the key feature for classification and improve the classification performance of the model. At the same time, the embodiments of the present application can expand two sets of noise interference data as training sample data based on the sample image, increase the data volume of the training sample data, help to improve the training effect of the target classification model while reducing the data dependence on the sample image, so as to be able to identify and judge more accurately when facing complex classification tasks, thereby improving the completion quality of the image classification task.
[0178] The image classification method provided by the embodiments of the present application will be described in detail below.
[0179] Refer to Figure 16 , Figure 16 which is an optional overall flowchart of the image classification method provided by the embodiments of the present application. Among them, the image classification method includes but is not limited to the following steps 1601 to step 1612:
[0180] Step 1601, obtain a sample image.
[0181] Step 1602, input the sample image into a semantic segmentation model for semantic segmentation to obtain the original background image corresponding to the background region and the original semantic image corresponding to the semantic region.
[0182] In this step, the sample image can be input into the semantic segmentation model, and the sample image is sequentially subjected to multiple scale-down operations to obtain sample images at multiple scales; in the order from high to low scales, the corresponding sample images are respectively encoded based on the encoders corresponding to each scale to obtain target encoded images at each scale, where the input of the encoder at the current scale is obtained by merging the sample image at the current scale and the target encoded image at the previous scale; then, in the order from low to high scales, the corresponding target encoded images are respectively decoded based on the decoders corresponding to each scale to obtain decoded images at each scale, where the input of the decoder at the current scale is obtained by merging the target encoded image at the current scale and the decoded image at the previous scale; the decoded image at the current scale is merged with the upsampled image at the previous scale to obtain a merged image, where the upsampled image is obtained by upsampling the decoded image at the previous scale; then, semantic segmentation is performed on the merged image corresponding to the original scale to obtain a background image corresponding to the background region and a semantic image corresponding to the semantic region.
[0183] Step 1603, determine the first noise intensity corresponding to the original background image and the second noise intensity corresponding to the original semantic image.
[0184] In this step, the original background image and the original semantic image can be input into the noise intensity prediction model to extract the first image feature of the original background image and the second image feature of the original semantic image, splice the first image feature and the second image feature to obtain a spliced image feature, perform linear regression on the spliced image feature to obtain a noise intensity prediction result, where the noise intensity prediction result includes the first noise standard deviation corresponding to the original background image and the second noise standard deviation corresponding to the original semantic image; then, determine the first noise intensity corresponding to the original background image according to the first noise standard deviation, and determine the second noise intensity corresponding to the original semantic image according to the second noise standard deviation.
[0185] Step 1604, generate a first noise image with the same size as the original background image according to the first noise intensity, and merge the first noise image with the original background image to obtain a noise background image corresponding to the original background image.
[0186] Step 1605, generate a second noise image with the same size as the original semantic image according to the second noise intensity, and merge the second noise image with the original semantic image to obtain a noise semantic image corresponding to the original semantic image.
[0187] Step 1606, merge the noise background image and the original semantic image to obtain one type of target noise image, and merge the noise semantic image and the original background image to obtain another type of target noise image.
[0188] In this step, the noise addition positions of different types of target noise images are different, and the noise addition positions include the background area or the semantic area of the sample image; specifically, at least one type of target noise image includes a first target noise image obtained by adding noise to the background area and a second target noise image obtained by adding noise to the semantic area.
[0189] Step 1607: Input the sample image and the target noise image into the target classification model for classification respectively, to obtain the first category confidence score corresponding to the sample image and the second category confidence score corresponding to the target noise image.
[0190] Step 1608: Determine the first loss according to the first category confidence score and the second category confidence score corresponding to the first target noise image.
[0191] In this step, the first loss is used to constrain the first category confidence score to be equal to the second category confidence score corresponding to the first target noise image.
[0192] Step 1609: Determine the second loss according to the first category confidence score and the second category confidence score corresponding to the second target noise image.
[0193] In this step, the second loss is used to constrain the first category confidence score to be greater than the second category confidence score corresponding to the second target noise image.
[0194] Step 1610: Determine the target loss according to the first loss and the second loss.
[0195] In this step, the target loss is used to constrain the magnitude relationship between the first category confidence score and the second category confidence score, and the magnitude relationship is determined according to the noise addition position. Specifically, the first segmentation confidence of the background area and the second segmentation confidence of the semantic area can be obtained, and then normalized according to the first segmentation confidence and the second segmentation confidence to obtain the first weight of the background area and the second weight of the semantic area. Then, the second weight is used as the weight of the first loss, and the first weight is used as the weight of the second loss, and the first loss and the second loss are weighted to obtain the target loss.
[0196] Step 1611: Train the target classification model according to the target loss.
[0197] Step 1612: Obtain the image to be classified, and input the image to be classified into the trained target classification model for classification.
[0198] Through the image classification method from step 1601 to step 1612 above, a sample image and at least one type of target noise image are obtained. The sample image and the target noise image are respectively input into the target classification model for classification, obtaining the first category confidence score corresponding to the sample image and the second category confidence score corresponding to the target noise image. The target loss is determined according to the first category confidence score and the second category confidence score, and the target classification model is trained according to the target loss. Since the target loss is used to constrain the magnitude relationship between the first category confidence score and the second category confidence score, and the magnitude relationship is determined according to the noise addition position, therefore, even if noise is added to at least one of the background region or the semantic region of the sample image, the category confidence score output by the target classification model can be constrained based on the target loss, which helps to avoid the target classification model wrongly fitting the background as the key feature for classification. When obtaining an image to be classified later and inputting the image to be classified into the trained target classification model for classification, the accuracy of image classification can be effectively improved.
[0199] The model training method provided by the embodiments of the present application will be described in detail below.
[0200] Refer to Figure 17 , Figure 17 which is an optional overall flow schematic diagram of the model training method provided by the embodiments of the present application. Among them, the model training method includes but is not limited to the following steps 1701 to step 1713:
[0201] Step 1701, obtain a sample image.
[0202] Step 1702, input the sample image into a semantic segmentation model for semantic segmentation, obtaining the original background image corresponding to the background region and the original semantic image corresponding to the semantic region.
[0203] In this step, the sample image can be input into the semantic segmentation model, and the sample image is successively subjected to multiple scale-down operations to obtain sample images at multiple scales; in the order from high to low scales, the corresponding sample images at each scale are respectively encoded based on the encoders corresponding to each scale to obtain target encoded images at each scale, where the input of the encoder at the current scale is obtained by merging the sample image at the current scale and the target encoded image at the previous scale; then, in the order from low to high scales, the corresponding target encoded images at each scale are respectively decoded based on the decoders corresponding to each scale to obtain decoded images at each scale, where the input of the decoder at the current scale is obtained by merging the target encoded image at the current scale and the decoded image at the previous scale; the decoded image at the current scale is merged with the upsampled image at the previous scale to obtain a merged image, where the upsampled image is obtained by upsampling the decoded image at the previous scale; then, semantic segmentation is performed on the merged image corresponding to the original scale to obtain a background image corresponding to the background region and a semantic image corresponding to the semantic region.
[0204] Step 1703: Determine the first noise intensity corresponding to the original background image and the second noise intensity corresponding to the original semantic image.
[0205] In this step, the original background image and the original semantic image can be input into the noise intensity prediction model to extract the first image feature of the original background image and the second image feature of the original semantic image, splice the first image feature and the second image feature to obtain a spliced image feature, perform linear regression on the spliced image feature to obtain a noise intensity prediction result, where the noise intensity prediction result includes the first noise standard deviation corresponding to the original background image and the second noise standard deviation corresponding to the original semantic image; then, determine the first noise intensity corresponding to the original background image according to the first noise standard deviation, and determine the second noise intensity corresponding to the original semantic image according to the second noise standard deviation.
[0206] Step 1704: Generate a first noise image with the same size as the original background image according to the first noise intensity, and merge the first noise image with the original background image to obtain a noise background image corresponding to the original background image.
[0207] Step 1705: Generate a second noise image with the same size as the original semantic image according to the second noise intensity, and merge the second noise image with the original semantic image to obtain a noise semantic image corresponding to the original semantic image.
[0208] Step 1706: Merge the noise background image and the original semantic image to obtain one type of target noise image, and merge the noise semantic image and the original background image to obtain another type of target noise image.
[0209] In this step, the noise addition positions of different types of target noise images are different. The noise addition positions include the background area or the semantic area of the sample image. Specifically, at least one type of target noise image includes a first target noise image obtained by adding noise to the background area and a second target noise image obtained by adding noise to the semantic area.
[0210] Step 1707: Input the sample image and the target noise image into the target classification model for classification to obtain the first category confidence score corresponding to the sample image and the second category confidence score corresponding to the target noise image.
[0211] Step 1708: Determine the first loss according to the first category confidence score and the second category confidence score corresponding to the first target noise image.
[0212] In this step, the first loss is used to constrain the first category confidence score to be equal to the second category confidence score corresponding to the first target noise image.
[0213] Step 1709: Determine the second loss according to the first category confidence score and the second category confidence score corresponding to the second target noise image.
[0214] In this step, the second loss is used to constrain the first category confidence score to be greater than the second category confidence score corresponding to the second target noise image.
[0215] Step 1710: Obtain the first segmentation confidence of the background area and the second segmentation confidence of the semantic area.
[0216] Step 1711: Normalize according to the first segmentation confidence and the second segmentation confidence to obtain the first weight of the background area and the second weight of the semantic area.
[0217] Step 1712: Use the second weight as the weight of the first loss and the first weight as the weight of the second loss, and weight the first loss and the second loss to obtain the target loss.
[0218] In this step, the target loss is used to constrain the magnitude relationship between the first category confidence score and the second category confidence score, and the magnitude relationship is determined according to the noise addition position.
[0219] Step 1713: Train the target classification model according to the target loss.
[0220] Through the model training method from step 1701 to step 1713 as described above, first obtain a sample image and at least one type of target noise image. Input the sample image and the target noise image into the target classification model for classification respectively to obtain the first category confidence score corresponding to the sample image and the second category confidence score corresponding to the target noise image. Determine the target loss according to the first category confidence score and the second category confidence score, and train the target classification model according to the target loss. Since the target loss is used to constrain the magnitude relationship between the first category confidence score and the second category confidence score, and the magnitude relationship is determined according to the noise addition position, therefore, even if noise is added to at least one of the background area or the semantic area of the sample image, the category confidence score output by the target classification model can be constrained based on the target loss, which helps to prevent the target classification model from wrongly fitting the background as the key feature for classification. Subsequently, when obtaining an image to be classified and inputting the image to be classified into the trained target classification model for classification, the accuracy of image classification can be effectively improved.
[0221] It can be understood that although the steps in the above respective flowcharts are shown sequentially according to the indication of the arrows, these steps are not necessarily executed sequentially according to the order indicated by the arrows. Unless there is a clear indication in this embodiment, the execution of these steps has no strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above flowcharts may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0222] Refer to Figure 18 , Figure 18 FIG. is an optional structural schematic diagram of an image classification device provided by an embodiment of the present application. The image classification device 1800 includes:
[0223] A first image acquisition module 1801, configured to acquire a sample image and at least one type of target noise image, where the target noise image is obtained by adding noise to the sample image, and the noise addition positions of different types of target noise images are different, and the noise addition positions include the background area or the semantic area of the sample image;
[0224] A first score acquisition module 1802, configured to input the sample image and the target noise image into the target classification model for classification respectively to obtain the first category confidence score corresponding to the sample image and the second category confidence score corresponding to the target noise image;
[0225] The first training module 1803 is configured to determine a target loss according to a first-class confidence score and a second-class confidence score, and train a target classification model according to the target loss, where the target loss is used to constrain the magnitude relationship between the first-class confidence score and the second-class confidence score, and the magnitude relationship is determined according to the noise addition position;
[0226] The inference module 1804 is configured to obtain an image to be classified, and input the image to be classified into the trained target classification model for classification.
[0227] In a possible implementation, at least one type of target noise image includes a first target noise image obtained by adding noise to a background region and a second target noise image obtained by adding noise to a semantic region. The first training module 1803 is further configured to:
[0228] Determine a first loss according to the first-class confidence score and the second-class confidence score corresponding to the first target noise image, where the first loss is used to constrain the first-class confidence score to be equal to the second-class confidence score corresponding to the first target noise image;
[0229] Determine a second loss according to the first-class confidence score and the second-class confidence score corresponding to the second target noise image, where the second loss is used to constrain the first-class confidence score to be greater than the second-class confidence score corresponding to the second target noise image;
[0230] Determine the target loss according to the first loss and the second loss.
[0231] In a possible implementation, the background region and the semantic region are determined by semantic segmentation. The first training module 1803 is further configured to:
[0232] Obtain a first segmentation confidence of the background region and a second segmentation confidence of the semantic region;
[0233] Normalize according to the first segmentation confidence and the second segmentation confidence to obtain a first weight of the background region and a second weight of the semantic region;
[0234] Use the second weight as the weight of the first loss, use the first weight as the weight of the second loss, and weight the first loss and the second loss to obtain the target loss.
[0235] In a possible implementation, the first image acquisition module 1801 is further configured to:
[0236] Obtain a sample image, input the sample image into a semantic segmentation model for semantic segmentation, and obtain an original background image corresponding to the background region and an original semantic image corresponding to the semantic region;
[0237] Add noise to the original background image and the original semantic image respectively to obtain a noisy background image corresponding to the original background image and a noisy semantic image corresponding to the original semantic image;
[0238] Merge the noisy background image and the original semantic image to obtain one type of target noisy image, and merge the noisy semantic image and the original background image to obtain another type of target noisy image.
[0239] In a possible implementation, the first image acquisition module 1801 is further configured to:
[0240] Input the sample image into the semantic segmentation model, and perform multiple scale-down operations on the sample image in sequence to obtain sample images at multiple scales;
[0241] In the order from high to low in terms of scale, encode the corresponding sample images respectively based on the encoders corresponding to each scale to obtain target encoded images at each scale, where the input of the encoder at the current scale is obtained by merging the sample image at the current scale and the target encoded image at the previous scale;
[0242] In the order from low to high in terms of scale, decode the corresponding target encoded images respectively based on the decoders corresponding to each scale to obtain decoded images at each scale, where the input of the decoder at the current scale is obtained by merging the target encoded image at the current scale and the decoded image at the previous scale;
[0243] Merge the decoded image at the current scale with the upsampled image at the previous scale to obtain a merged image, where the upsampled image is obtained by upsampling the decoded image at the previous scale;
[0244] Perform semantic segmentation on the merged image corresponding to the original scale to obtain a background image corresponding to the background region and a semantic image corresponding to the semantic region.
[0245] In a possible implementation, the first image acquisition module 1801 is further configured to:
[0246] Encode the corresponding sample images respectively based on the encoders corresponding to each scale to obtain original encoded images at each scale;
[0247] Determine the spatial attention matrix and the position attention matrix corresponding to the sample image at each scale respectively, and merge the spatial attention matrix and the position attention matrix corresponding to each scale to obtain a merged attention matrix;
[0248] Perform max pooling on the merged attention matrices corresponding to multiple scales to obtain a target attention matrix;
[0249] Adjust the original encoded images corresponding to each scale based on the target attention matrix to obtain the target encoded images for each scale.
[0250] In a possible implementation, the first image acquisition module 1801 is further configured to:
[0251] Determine the first noise intensity corresponding to the original background image and the second noise intensity corresponding to the original semantic image;
[0252] Generate a first noise image with the same size as the original background image according to the first noise intensity, and merge the first noise image with the original background image to obtain a noise background image corresponding to the original background image;
[0253] Generate a second noise image with the same size as the original semantic image according to the second noise intensity, and merge the second noise image with the original semantic image to obtain a noise semantic image corresponding to the original semantic image.
[0254] In a possible implementation, the first image acquisition module 1801 is further configured to:
[0255] Input the original background image and the original semantic image into a noise intensity prediction model, extract the first image feature of the original background image and the second image feature of the original semantic image, splice the first image feature and the second image feature to obtain a spliced image feature, perform linear regression on the spliced image feature to obtain a noise intensity prediction result, where the noise intensity prediction result includes the first noise standard deviation corresponding to the original background image and the second noise standard deviation corresponding to the original semantic image;
[0256] Determine the first noise intensity corresponding to the original background image according to the first noise standard deviation, and determine the second noise intensity corresponding to the original semantic image according to the second noise standard deviation.
[0257] In a possible implementation, the number of noise intensity prediction results is multiple, the noise intensity prediction results are output by the same output layer of the noise intensity prediction model, and the output layer also outputs the intensity weights corresponding to each noise intensity prediction result. The first image acquisition module 1801 is further configured to:
[0258] Weight the multiple first noise standard deviations according to the intensity weights to obtain a first weighted standard deviation;
[0259] Weight the multiple second noise standard deviations according to the intensity weights to obtain a second weighted standard deviation;
[0260] Determine the first noise intensity corresponding to the original background image according to the first weighted standard deviation, and determine the second noise intensity corresponding to the original semantic image according to the second weighted standard deviation.
[0261] Refer toFigure 19 , Figure 19 FIG. 3 is an alternative structural schematic diagram of the model training apparatus provided by the embodiment of the present application. The model training apparatus 1900 includes:
[0262] A second image acquisition module 1901, configured to acquire a sample image and at least one type of target noise image, where the target noise image is obtained by adding noise to the sample image, and the noise addition positions of different types of target noise images are different. The noise addition positions include the background area or the semantic area of the sample image;
[0263] A second score acquisition module 1902, configured to input the sample image and the target noise image into a target classification model for classification respectively, to obtain a first category confidence score corresponding to the sample image and a second category confidence score corresponding to the target noise image;
[0264] A second training module 1903, configured to determine a target loss according to the first category confidence score and the second category confidence score, and train the target classification model according to the target loss, where the target loss is used to constrain the magnitude relationship between the first category confidence score and the second category confidence score, and the magnitude relationship is determined according to the noise addition position.
[0265] In a possible implementation manner, at least one type of target noise image includes a first target noise image obtained by adding noise to the background area and a second target noise image obtained by adding noise to the semantic area. The second training module 1903 is further configured to:
[0266] Determine a first loss according to the first category confidence score and the second category confidence score corresponding to the first target noise image, where the first loss is used to constrain the first category confidence score to be equal to the second category confidence score corresponding to the first target noise image;
[0267] Determine a second loss according to the first category confidence score and the second category confidence score corresponding to the second target noise image, where the second loss is used to constrain the first category confidence score to be greater than the second category confidence score corresponding to the second target noise image;
[0268] Determine the target loss according to the first loss and the second loss.
[0269] In a possible implementation manner, the background area and the semantic area are determined by semantic segmentation. The second training module 1903 is further configured to:
[0270] Obtain a first segmentation confidence of the background area and a second segmentation confidence of the semantic area;
[0271] Normalize according to the first segmentation confidence and the second segmentation confidence to obtain the first weight of the background region and the second weight of the semantic region;
[0272] Use the second weight as the weight of the first loss and the first weight as the weight of the second loss, and weight the first loss and the second loss to obtain the target loss.
[0273] In a possible implementation, the second image acquisition module 1901 is further configured to:
[0274] Acquire a sample image, input the sample image into a semantic segmentation model for semantic segmentation to obtain an original background image corresponding to the background region and an original semantic image corresponding to the semantic region;
[0275] Add noise to the original background image and the original semantic image respectively to obtain a noise background image corresponding to the original background image and a noise semantic image corresponding to the original semantic image;
[0276] Merge the noise background image and the original semantic image to obtain one type of target noise image, and merge the noise semantic image and the original background image to obtain another type of target noise image.
[0277] In a possible implementation, the second image acquisition module 1901 is further configured to:
[0278] Input the sample image into a semantic segmentation model, and perform multiple scale-down operations on the sample image in sequence to obtain sample images at multiple scales;
[0279] In the order from high to low in scale, encode the corresponding sample images based on the encoders corresponding to each scale to obtain target encoded images at each scale, where the input of the encoder at the current scale is obtained by merging the sample image at the current scale and the target encoded image at the previous scale;
[0280] In the order from low to high in scale, decode the corresponding target encoded images based on the decoders corresponding to each scale to obtain decoded images at each scale, where the input of the decoder at the current scale is obtained by merging the target encoded image at the current scale and the decoded image at the previous scale;
[0281] Merge the decoded image at the current scale with the upsampled image at the previous scale to obtain a merged image, where the upsampled image is obtained by upsampling the decoded image at the previous scale;
[0282] Perform semantic segmentation on the merged image corresponding to the original scale to obtain a background image corresponding to the background region and a semantic image corresponding to the semantic region.
[0283] In a possible implementation, the second image acquisition module 1901 is further configured to:
[0284] Encode the corresponding sample images respectively based on the encoders corresponding to each scale to obtain the original encoded images of each scale;
[0285] Determine the spatial attention matrix and the position attention matrix corresponding to the sample images of each scale respectively, and merge the spatial attention matrix and the position attention matrix corresponding to each scale to obtain a merged attention matrix;
[0286] Perform max pooling on the merged attention matrices corresponding to multiple scales to obtain a target attention matrix;
[0287] Adjust the original encoded images of each scale respectively based on the target attention matrix to obtain the target encoded images of each scale.
[0288] In a possible implementation, the second image acquisition module 1901 is further configured to:
[0289] Determine the first noise intensity corresponding to the original background image and the second noise intensity corresponding to the original semantic image;
[0290] Generate a first noise image with the same size as the original background image according to the first noise intensity, and merge the first noise image with the original background image to obtain a noise background image corresponding to the original background image;
[0291] Generate a second noise image with the same size as the original semantic image according to the second noise intensity, and merge the second noise image with the original semantic image to obtain a noise semantic image corresponding to the original semantic image.
[0292] In a possible implementation, the second image acquisition module 1901 is further configured to:
[0293] Input the original background image and the original semantic image into a noise intensity prediction model, extract the first image feature of the original background image and the second image feature of the original semantic image, splice the first image feature and the second image feature to obtain a spliced image feature, and perform linear regression on the spliced image feature to obtain a noise intensity prediction result, where the noise intensity prediction result includes the first noise standard deviation corresponding to the original background image and the second noise standard deviation corresponding to the original semantic image;
[0294] Determine the first noise intensity corresponding to the original background image according to the first noise standard deviation, and determine the second noise intensity corresponding to the original semantic image according to the second noise standard deviation.
[0295] In a possible implementation, the number of noise intensity prediction results is multiple, and the noise intensity prediction results are output by the same output layer of the noise intensity prediction model. The output layer also outputs the intensity weights corresponding to the respective noise intensity prediction results. The second image acquisition module 1901 is further configured to:
[0296] Weight the multiple first noise standard deviations according to the intensity weights to obtain a first weighted standard deviation;
[0297] Weight the multiple second noise standard deviations according to the intensity weights to obtain a second weighted standard deviation;
[0298] Determine the first noise intensity corresponding to the original background image according to the first weighted standard deviation, and determine the second noise intensity corresponding to the original semantic image according to the second weighted standard deviation.
[0299] The electronic device provided by the embodiments of the present application for executing the above image classification method or model training method may be a terminal. Refer to Figure 20 , Figure 20 which is a partial structural block diagram of the terminal provided by the embodiments of the present application. The terminal includes: a camera component 2010, a first memory 2020, an input unit 2030, a display unit 2040, a sensor 2050, an audio circuit 2060, a wireless fidelity (WiFi) module 2070, a first processor 2080, and a power supply 2090 and other components. Those skilled in the art can understand that Figure 20 the terminal structure shown in
[0300] does not constitute a limitation on the terminal, and may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0301] The camera component 2010 can be used to collect images or videos. Optionally, the camera component 2010 includes a front camera and a rear camera. Usually, the front camera is disposed on the front panel of the terminal, and the rear camera is disposed on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth camera, a wide-angle camera, and a telephoto camera, to implement functions such as background blurring by fusing the main camera and the depth camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (Virtual Reality) shooting function or other fusion shooting functions.
[0302] The input unit 2030 can be used to receive input digital or character information and generate key signal inputs related to the settings and function control of the terminal. Specifically, the input unit 2030 can include a touch panel 2031 and other input devices 2032.
[0303] The display unit 2040 can be used to display the input information or provided information and various menus of the terminal. The display unit 2040 can include a display panel 2041.
[0304] The audio circuit 2060, speaker 2061, and microphone 2062 can provide an audio interface.
[0305] The power supply 2090 can be alternating current, direct current, a disposable battery, or a rechargeable battery.
[0306] The number of sensors 2050 can be one or more. The one or more sensors 2050 include but are not limited to: acceleration sensors, gyroscope sensors, pressure sensors, optical sensors, etc. Among them:
[0307] The acceleration sensor can detect the magnitudes of accelerations on the three coordinate axes of the coordinate system established by the terminal. For example, the acceleration sensor can be used to detect the components of the gravitational acceleration on the three coordinate axes. The first processor 2080 can control the display unit 2040 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor. The acceleration sensor can also be used for the collection of game or user movement data.
[0308] The gyroscope sensor can detect the body direction and rotation angle of the terminal. The gyroscope sensor can cooperate with the acceleration sensor to collect the 3D actions of the user on the terminal. According to the data collected by the gyroscope sensor, the first processor 2080 can implement the following functions: motion sensing (such as changing the UI according to the user's tilting operation), image stabilization during shooting, game control, and inertial navigation.
[0309] The pressure sensor can be arranged on the side frame of the terminal and / or the lower layer of the display unit 2040. When the pressure sensor is arranged on the side frame of the terminal, it can detect the holding signal of the user on the terminal, and the first processor 2080 can perform left / right hand recognition or quick operation according to the holding signal collected by the pressure sensor. When the pressure sensor is arranged on the lower layer of the display unit 2040, the first processor 2080 can control the operable controls on the UI interface according to the pressure operation of the user on the display unit 2040. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0310] The optical sensor is used to collect the ambient light intensity. In one embodiment, the first processor 2080 may control the display brightness of the display unit 2040 according to the ambient light intensity collected by the optical sensor. Specifically, when the ambient light intensity is high, the display brightness of the display unit 2040 is increased; when the ambient light intensity is low, the display brightness of the display unit 2040 is decreased. In another embodiment, the first processor 2080 may also dynamically adjust the shooting parameters of the camera assembly 2010 according to the ambient light intensity collected by the optical sensor.
[0311] In this embodiment, the first processor 2080 included in the terminal may execute the image classification method or the model training method of the previous embodiment.
[0312] The electronic device provided in the embodiment of the present application for executing the above image classification method or model training method may also be a server. Refer to Figure 21 , Figure 21 which is a partial structural block diagram of the server provided in the embodiment of the present application. The server 2100 may vary greatly due to configuration or performance differences, and may include one or more central processing units (Central Processing Units, abbreviated as CPU), that is, the second processor 2122 as shown in Figure 21 (for example, one or more processors) and the second memory 2132, and one or more storage media 2130 (for example, one or more mass storage devices) for storing the application program 2142 or data 2144. Among them, the second memory 2132 and the storage media 2130 may be transient storage or persistent storage. The program stored in the storage media 2130 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server 2100. Further, the second processor 2122 may be configured to communicate with the storage media 2130 and execute a series of instruction operations in the storage media 2130 on the server 2100.
[0313] The server 2100 may further include one or more power supplies 2126, one or more wired or wireless network interfaces 2150, one or more input / output interfaces 2158, and / or one or more operating systems 2141, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0314] The processor in the server 2100 may be used to execute the image classification method or the model training method.
[0315] The embodiments of the present application also provide a computer-readable storage medium for storing a computer program, which is used to execute the image classification method or the model training method in the foregoing respective embodiments.
[0316] The embodiments of the present application also provide a computer program product, which includes a computer program stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the above-mentioned image classification method or model training method.
[0317] Terms such as "first", "second", "third", "fourth", etc. (if any) in the description of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances to describe the embodiments of the present application. For example, it can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0318] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0319] It should be understood that in the description of the embodiments of the present application, the meaning of "a plurality (or multiple items)" is more than two. Understandings such as greater than, less than, exceeding, etc. do not include the present number, and understandings such as above, below, within, etc. include the present number.
[0320] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.
[0321] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0322] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0323] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0324] It should also be understood that the various embodiments provided by the embodiments of the present application can be combined arbitrarily to achieve different technical effects.
[0325] The above is a specific description of the preferred embodiment of the present application. However, the present application is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions under the condition of not violating the spirit of the present application. These equivalent deformations or substitutions are all included within the scope defined by the claims of the present application.
Claims
1. An image classification method, characterized in that, Including: Obtaining a sample image and at least one type of target noise image, where the target noise image is obtained by adding noise to the sample image, and the noise addition positions of different types of target noise images are different, and the noise addition positions include the background region or the semantic region of the sample image; Inputting the sample image and the target noise image into a target classification model for classification respectively to obtain a first category confidence score corresponding to the sample image and a second category confidence score corresponding to the target noise image; Determining a target loss according to the first category confidence score and the second category confidence score, and training the target classification model according to the target loss, where the target loss is used to constrain the magnitude relationship between the first category confidence score and the second category confidence score, and the magnitude relationship is determined according to the noise addition position; Obtaining an image to be classified, and inputting the image to be classified into the trained target classification model for classification.
2. The image classification method according to claim 1, wherein At least one type of the target noise image includes a first target noise image obtained by adding noise to the background region and a second target noise image obtained by adding noise to the semantic region. The determining the target loss according to the first category confidence score and the second category confidence score includes: Determining a first loss according to the first category confidence score and the second category confidence score corresponding to the first target noise image, where the first loss is used to constrain the first category confidence score to be equal to the second category confidence score corresponding to the first target noise image; Determining a second loss according to the first category confidence score and the second category confidence score corresponding to the second target noise image, where the second loss is used to constrain the first category confidence score to be greater than the second category confidence score corresponding to the second target noise image; Determining the target loss according to the first loss and the second loss.
3. The image classification method according to claim 2, wherein The background region and the semantic region are determined by semantic segmentation. The determining the target loss according to the first loss and the second loss includes: Obtaining a first segmentation confidence of the background region and a second segmentation confidence of the semantic region; Normalizing according to the first segmentation confidence and the second segmentation confidence to obtain a first weight of the background region and a second weight of the semantic region; Using the second weight as the weight of the first loss and the first weight as the weight of the second loss, and weighting the first loss and the second loss to obtain the target loss.
4. The image classification method according to claim 1, wherein The obtaining the sample image and at least one type of target noise image includes: Obtaining a sample image, inputting the sample image into a semantic segmentation model for semantic segmentation to obtain an original background image corresponding to the background region and an original semantic image corresponding to the semantic region; Adding noise to the original background image and the original semantic image respectively to obtain a noise background image corresponding to the original background image and a noise semantic image corresponding to the original semantic image; Merge the noise background image and the original semantic image to obtain one type of target noise image, and merge the noise semantic image and the original background image to obtain the other type of the target noise image.
5. The image classification method according to claim 4, wherein Input the sample image into a semantic segmentation model for semantic segmentation to obtain the background image corresponding to the background region and the semantic image corresponding to the semantic region, including: Input the sample image into the semantic segmentation model, and perform multiple scale-down operations on the sample image in sequence to obtain the sample images at multiple scales; In the order of decreasing scale, encode the corresponding sample images respectively based on the encoders corresponding to each scale to obtain the target encoded images at each scale. Among them, the input of the encoder at the current scale is obtained by merging the sample image at the current scale and the target encoded image at the previous scale; In the order of increasing scale, decode the corresponding target encoded images respectively based on the decoders corresponding to each scale to obtain the decoded images at each scale. Among them, the input of the decoder at the current scale is obtained by merging the target encoded image at the current scale and the decoded image at the previous scale; Merge the decoded image at the current scale with the upsampled image at the previous scale to obtain a merged image, where the upsampled image is obtained by upsampling the decoded image at the previous scale; Perform semantic segmentation on the merged image corresponding to the original scale to obtain the background image corresponding to the background region and the semantic image corresponding to the semantic region.
6. The image classification method according to claim 5, wherein The encoding the corresponding sample images respectively based on the encoders corresponding to each scale to obtain the target encoded images at each scale includes: Encode the corresponding sample images respectively based on the encoders corresponding to each scale to obtain the original encoded images at each scale; Determine the spatial attention matrix and the position attention matrix corresponding to the sample images at each scale respectively, and merge the spatial attention matrix and the position attention matrix corresponding to each scale to obtain a merged attention matrix; Perform max pooling on the merged attention matrices corresponding to multiple scales to obtain a target attention matrix; Adjust the original encoded images at each scale respectively based on the target attention matrix to obtain the target encoded images at each scale.
7. The image classification method according to claim 4, wherein Adding noise to the original background image and the original semantic image respectively to obtain the noise background image corresponding to the original background image and the noise semantic image corresponding to the original semantic image includes: Determine the first noise intensity corresponding to the original background image and the second noise intensity corresponding to the original semantic image; Generate a first noise image with the same size as the original background image according to the first noise intensity, and merge the first noise image with the original background image to obtain the noise background image corresponding to the original background image; Generate a second noise image with the same size as the original semantic image according to the second noise intensity, and merge the second noise image with the original semantic image to obtain a noise semantic image corresponding to the original semantic image.
8. The image classification method according to claim 7, characterized in that The determining the first noise intensity corresponding to the original background image and the second noise intensity corresponding to the original semantic image includes: Input the original background image and the original semantic image into a noise intensity prediction model, extract the first image feature of the original background image and the second image feature of the original semantic image, splice the first image feature and the second image feature to obtain a spliced image feature, and perform linear regression on the spliced image feature to obtain a noise intensity prediction result, where the noise intensity prediction result includes the first noise standard deviation corresponding to the original background image and the second noise standard deviation corresponding to the original semantic image; Determine the first noise intensity corresponding to the original background image according to the first noise standard deviation, and determine the second noise intensity corresponding to the original semantic image according to the second noise standard deviation.
9. The image classification method according to claim 8, wherein The number of the noise intensity prediction results is multiple, the noise intensity prediction results are output by the same output layer of the noise intensity prediction model, and the output layer also outputs the intensity weights corresponding to the respective noise intensity prediction results. The determining the first noise intensity corresponding to the original background image according to the first noise standard deviation and the second noise intensity corresponding to the original semantic image according to the second noise standard deviation includes: Weight the multiple first noise standard deviations according to the intensity weights to obtain a first weighted standard deviation; Weight the multiple second noise standard deviations according to the intensity weights to obtain a second weighted standard deviation; Determine the first noise intensity corresponding to the original background image according to the first weighted standard deviation, and determine the second noise intensity corresponding to the original semantic image according to the second weighted standard deviation.
10. A model training method, characterized in that, Includes: Obtain a sample image and at least one type of target noise image, where the target noise image is obtained by adding noise to the sample image, and the noise addition positions of different types of target noise images are different, and the noise addition positions include the background area or the semantic area of the sample image; Input the sample image and the target noise image into a target classification model for classification to obtain a first category confidence score corresponding to the sample image and a second category confidence score corresponding to the target noise image; Determine a target loss according to the first category confidence score and the second category confidence score, and train the target classification model according to the target loss, where the target loss is used to constrain the magnitude relationship between the first category confidence score and the second category confidence score, and the magnitude relationship is determined according to the noise addition position.
11. An image classification device, characterized in that, Includes: A first image acquisition module, configured to acquire a sample image and at least one type of target noise image, wherein the target noise image is obtained by adding noise to the sample image, and the noise addition positions of different types of target noise images are different, and the noise addition positions include the background region or the semantic region of the sample image; A first score acquisition module, configured to input the sample image and the target noise image into a target classification model for classification respectively, to obtain a first category confidence score corresponding to the sample image and a second category confidence score corresponding to the target noise image; A first training module, configured to determine a target loss according to the first category confidence score and the second category confidence score, and train the target classification model according to the target loss, wherein the target loss is used to constrain the magnitude relationship between the first category confidence score and the second category confidence score, and the magnitude relationship is determined according to the noise addition position; An inference module, configured to acquire an image to be classified, and input the image to be classified into the trained target classification model for classification.
12. A model training device, characterized in that, including: A second image acquisition module, configured to acquire a sample image and at least one type of target noise image, wherein the target noise image is obtained by adding noise to the sample image, and the noise addition positions of different types of target noise images are different, and the noise addition positions include the background region or the semantic region of the sample image; A second score acquisition module, configured to input the sample image and the target noise image into a target classification model for classification respectively, to obtain a first category confidence score corresponding to the sample image and a second category confidence score corresponding to the target noise image; A second training module, configured to determine a target loss according to the first category confidence score and the second category confidence score, and train the target classification model according to the target loss, wherein the target loss is used to constrain the magnitude relationship between the first category confidence score and the second category confidence score, and the magnitude relationship is determined according to the noise addition position.
13. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the image classification method according to any one of claims 1 to 9, or implements the model training method according to claim 10.
14. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the image classification method according to any one of claims 1 to 9, or implements the model training method according to claim 10.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the image classification method according to any one of claims 1 to 9, or implements the model training method according to claim 10.
Citation Information
Cited By
Image classification model training method and device, image classification method and device and computer equipment
CN120894631A