Training Method of Image Classification Model, Semantic Segmentation Method and Related Devices
Through the methods of feature extraction, reconstruction and coverage density screening, the core pixel set of the target domain image is filtered out, solving the problem of high pixel-level annotation cost, and achieving efficient image classification model training and pixel-level classification of the target domain image.
Patent Information
- Application Number
- CN202311510724.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-10
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-11-10
AI Technical Summary
When migrating the image classification model from the source domain to the target domain, the pixel-level annotation work of the target domain image is high, resulting in high annotation cost.
Features of the target domain image are extracted through the feature extraction network, feature reconstruction is performed, pixel coverage density is determined, core pixel sets are filtered, and image classification model is fine-tuned and trained based on the annotation information of the core pixel set.
The pixel annotation amount and annotation cost are reduced, and the fine-tuning training effect of the image classification model is ensured, so that the model can effectively classify the images in the target domain.
Smart Images

Figure CN117315404B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and more particularly, to a method for training an image classification model, a semantic segmentation method, and related devices. Background Art
[0002] In order to transfer an image classification model for implementing a semantic segmentation task from a source domain to a target domain, it is necessary to perform pixel-level annotation on the target domain image, and then use the annotation information of the pixels in the target domain image and the target domain image to train the image classification model. However, the annotation workload for performing pixel-level annotation on all pixels in the target domain image is large, resulting in a high annotation cost. Summary of the Invention
[0003] In view of the above problems, embodiments of the present application propose a method for training an image classification model, a semantic segmentation method, and related devices to solve the problems of large annotation workload and high annotation cost for performing pixel-level annotation on all pixels in the target domain image in the related art.
[0004] According to one aspect of the embodiments of the present application, there is provided a method for training an image classification model, including: determining a feature map of a target domain image based on features extracted from the target domain image by a feature extraction network in the image classification model; the image classification model is pre-trained by a source domain image and annotation information of each pixel in the source domain image; performing feature reconstruction based on the feature map to obtain a reconstructed feature map, where the reconstructed feature of a pixel in the reconstructed feature map is reconstructed according to the features of the neighboring pixels of the pixel; determining the coverage density of each pixel in the target domain image according to the feature map and the reconstructed feature map; screening a core pixel set to be pixel-annotated in the target domain image according to the coverage density of each pixel in the target domain image and the features of each pixel in the feature map; the coverage density is used to scale the feature distance between two pixels; fine-tuning and training the image classification model according to the annotation information of the pixels in the core pixel set and the target domain image.
[0005] According to one aspect of the embodiments of the present application, there is provided a semantic segmentation method, including: obtaining an image to be processed; classifying the image to be processed by an image classification model based on the image to be processed, and outputting a classification result of the image to be processed, where the image classification model is trained according to the method as above; the classification result indicates the category to which each pixel in the image to be processed belongs; determining a semantic segmentation result of the image to be processed according to the category to which each pixel in the image to be processed belongs.
[0006] According to one aspect of the embodiments of the present application, there is provided a training device for an image classification model, including: a feature map acquisition module, configured to determine a feature map of a target domain image based on features extracted from the target domain image by a feature extraction network in the image classification model; the image classification model is pre-trained through source domain images and annotation information of each pixel in the source domain images; a feature reconstruction module, configured to perform feature reconstruction based on the feature map to obtain a reconstructed feature map, and the reconstructed features of pixels in the reconstructed feature map are reconstructed according to the features of neighboring pixels of the pixel; a coverage density determination module, configured to determine the coverage density of each pixel in the target domain image according to the feature map and the reconstructed feature map; a screening module, configured to screen a core pixel set to be pixel-labeled in the target domain image according to the coverage density of each pixel in the target domain image and the features of each pixel in the feature map; the coverage density is used to scale the feature distance between two pixels; a first fine-tuning training module, configured to perform fine-tuning training on the image classification model according to the annotation information of pixels in the core pixel set and the target domain image.
[0007] Based on the above solution, the screening module includes: a first determination unit, configured to determine a first pixel set, and the first pixels in the first pixel set are pixels in the target domain image; a second determination unit, configured to, for each second pixel in a second pixel set, determine a first target pixel with the smallest feature distance from the second pixel in the first pixel set; the feature distance is determined according to the features of two pixels; the second pixels in the second pixel set are pixels in the target domain image, and the intersection of the second pixel set and the first pixel set is empty; a reference distance determination unit, configured to scale the feature distance between the second pixel and the corresponding first target pixel according to the coverage density of the first target pixel corresponding to the second pixel, to obtain a reference distance between the second pixel and the first pixel set, and the reference distance has a negative correlation with the coverage density; a second target pixel determination unit, configured to determine a second target pixel with the largest reference distance from the first pixel set in the second pixel set; an adding unit, configured to add the second target pixel to the first pixel set; a third determination unit, configured to determine the core pixel set based on the updated first pixel set.
[0008] Based on the above solution, the third determination unit is further configured to: count the number of pixels in the updated first pixel set; if it is determined that the annotation requirement is met according to the number of pixels, use the updated first pixel set as the core pixel set; The training device of the image classification model further includes: a removal module, configured to: if it is determined that the annotation requirement is not met according to the number of pixels, remove the second target pixel from the second pixel set to obtain a new second pixel set, and return to execute for each second pixel in the second pixel set, determine the first target pixel with the smallest feature distance from this second pixel in the first pixel set.
[0009] Based on the above solution, the image classification model further includes a first classifier; the training device of the image classification model further includes: a first classification module, configured to classify the result of feature extraction of the target domain image by the first classifier based on the feature extraction network to obtain a first classification result, and the first classification result indicates the probability that each pixel in the target domain image belongs to each of multiple categories; a target probability determination module, configured to determine the largest N target probabilities for each pixel among the probabilities that the pixel belongs to multiple categories, where N is a positive integer; a labeling gain determination module, configured to calculate the labeling gain for each pixel according to the N target probabilities determined for each pixel; a second pixel set determination module, configured to screen out the pixels with a labeling gain greater than the labeling gain threshold from the target domain image, and construct the second pixel set according to the pixels with a labeling gain greater than the labeling gain threshold.
[0010] Based on the above solution, N is 2, and the N target probabilities determined for a pixel include the maximum probability and the second largest probability among the probabilities that the pixel belongs to multiple categories; the labeling gain determination module is configured to: for each pixel, calculate the sum of the maximum probability and the second largest probability corresponding to this pixel to obtain a reference result for this pixel; for each pixel, subtract the reference result corresponding to this pixel from 1 to obtain the labeling gain for this pixel.
[0011] Based on the above solution, the feature reconstruction module includes: a dynamic mask convolution processing unit, configured to perform dynamic mask convolution processing on the feature map to obtain a spatial modulation map, where the feature value of the central pixel of the convolution window for each convolution in the spatial modulation map is 0; a fusion unit, configured to fuse the spatial modulation map and the feature map to obtain the reconstructed feature map.
[0012] Based on the above solution, the dynamic mask convolution processing unit is configured to: perform dynamic convolution on the feature map to obtain an intermediate feature map; superimpose the intermediate feature map with the mask tensor corresponding to the convolution window to obtain a reference feature map; perform normalization processing on the features in the reference feature map to obtain the spatial modulation map.
[0013] Based on the above solution, the coverage density determination module includes: an acquisition unit, configured to acquire, for each pixel, the feature of the pixel from the feature map and the reconstructed feature of the pixel from the reconstructed feature map; a feature difference determination unit, configured to calculate the feature difference between the feature of the pixel and the reconstructed feature of the pixel; and a coverage density determination unit, configured to determine the coverage density of the pixel according to the feature difference of the pixel.
[0014] Based on the above solution, the image classification model further includes a first classifier; the reconstructed feature map is obtained by the dynamic mask convolution network reconstructing based on the feature map; the first fine-tuning training module includes: a first acquisition module, configured to acquire a first classification result obtained by the first classifier classifying the features extracted from the target domain image by the feature extraction network; a first loss determination module, configured to determine a first loss according to the first classification result and the annotation information of the pixels in the core pixel set; a feature reconstruction loss determination module, configured to determine a feature reconstruction loss according to the feature map and the reconstructed feature map; and an adjustment module, configured to adjust the parameters of at least one of the image classification model and the dynamic mask convolution network according to the first loss and the feature reconstruction loss.
[0015] Based on the above solution, the feature map is obtained by the first convolutional network performing dimensionality reduction processing on the first feature map, and the first feature map is the feature extracted from the target domain image by the feature extraction network; the training device of the image classification model further includes: a second classification module, configured to classify by the second classifier based on the feature map to obtain a second classification result; a second loss determination module, configured to determine a second loss according to the second classification result and the annotation information of the pixels in the core pixel set; correspondingly, the adjustment module is configured to: if the number of iterations is less than the first number threshold, determine a target loss according to the first loss, the second loss, and the feature reconstruction loss; and adjust the parameters of the image classification model, the dynamic mask convolution network, the first convolutional network, and the second classifier according to the target loss.
[0016] Based on the above solution, the adjustment module further includes: if the number of iterations is not less than the first number threshold, adjust the parameters of the image classification model according to the first loss; and adjust the parameters of the dynamic mask convolution network, the first convolutional network, and the second classifier according to the weighted result of the second loss and the feature reconstruction loss.
[0017] Based on the above solution, the training device of the image classification model further includes: a second fine-tuning training module, configured to perform fine-tuning training on the image classification model according to the first source domain image and the label information of the first source domain image.
[0018] Based on the above solution, the second fine-tuning training module is used to: extract features from the first source domain image by the feature extraction network to obtain a first feature map of the first source domain image; classify the first feature map of the first source domain image by a first classifier to obtain a third classification result; determine a first loss for the first source domain image according to the annotation information of the first source domain image and the third classification result; determine a feature map of the first source domain image based on the first feature map of the first source domain image; perform feature reconstruction on the feature map of the first source domain image through the dynamic mask convolution network to obtain a reconstructed feature map of the first source domain image; determine a feature reconstruction loss for the first source domain image according to the feature map of the first source domain image and the reconstructed feature map of the first source domain image; determine a third loss for the first source domain image according to the probabilities that each pixel in the third classification result belongs to the actual class and belongs to other classes; and adjust the parameters of at least one of the image classification model and the dynamic mask convolution network according to the first loss, the feature reconstruction loss, and the third loss for the first source domain image.
[0019] According to one aspect of the embodiments of the present application, a semantic segmentation device is provided, including: a to-be-processed image acquisition module, configured to acquire a to-be-processed image; a third classification module, configured to classify the to-be-processed image by an image classification model to output a classification result of the to-be-processed image, where the image classification model is trained according to the training method of the above image classification model; the classification result indicates the class to which each pixel in the to-be-processed image belongs; and a semantic segmentation result determination module, configured to determine a semantic segmentation result of the to-be-processed image according to the class to which each pixel in the to-be-processed image belongs.
[0020] According to one aspect of the embodiments of the present application, an electronic device is provided, including: a processor; a memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the training method of the above image classification model is implemented or the above semantic segmentation method is implemented.
[0021] According to one aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by a processor, the training method of the above image classification model is implemented or the above semantic segmentation method is implemented.
[0022] According to one aspect of the embodiments of the present application, a computer program product is provided, including computer instructions, and when the computer instructions are executed by a processor, the training method of the above image classification model is implemented or the above semantic segmentation method is implemented.
[0023] In this application, after obtaining the feature map of the target domain image, the feature of each pixel is reconstructed based on the features of the neighboring pixels of the feature image pixel to obtain a reconstructed feature map. Then, based on the feature map and the reconstructed feature map, the coverage density of each pixel in the target domain image is determined. This coverage density can reflect the correlation between a pixel and its neighboring pixels in the feature space. Based on this, the feature distance between two pixels is scaled according to the coverage density of each pixel in the target domain image, and based on this, a core pixel set to be pixel-labeled is screened in the target domain image. Since the feature distance between two pixels is scaled by the coverage density, that is, if the coverage density of a pixel is large, the feature distance after scaling based on the coverage density decreases; conversely, if the coverage density of a pixel is small, the feature distance after scaling according to the coverage density increases. In this way, it is possible to avoid selecting the neighboring pixels of the pixels with high correlation with a pixel in the neighborhood of the pixel in the target domain image into the core pixel set, which can ensure the training effect of the image classification model using the pixels in the core pixel set and avoid waste of the annotation budget. After determining the core pixel set for the target domain image, the pixels in the core pixel set are labeled, rather than labeling all the pixels in the target domain image. This can greatly reduce the amount of pixel labeling, reduce the cost of pixel labeling, and ensure the fine-tuning training effect of the image classification model, enabling the image classification model to effectively perform pixel-level classification on the images in the target domain. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0025] Figure 1 is a schematic diagram of the application scenario of the present application shown according to an embodiment of the present application.
[0026] Figure 2 shows a schematic diagram of the relationship between the estimated coverage density and the average radial distance observed in the experiment.
[0027] Figure 3 is a flowchart of a method for training an image classification model shown according to an embodiment of the present application.
[0028] Figure 4 is a flowchart of performing feature reconstruction shown according to an embodiment of the present application.
[0029] Figure 5A shows a schematic diagram of determining the core pixel set according to the greedy algorithm.
[0030] Figure 5B A schematic diagram showing the determination of the core pixel set by combining the greedy algorithm and the coverage density of pixels in the target domain image is shown.
[0031] Figure 6 Shown is Figure 3 A flowchart of step 340 in a corresponding embodiment in one embodiment.
[0032] Figure 7 Shown is Figure 6 A flowchart of the steps before step 620 in a corresponding embodiment.
[0033] Figure 8 Shown is Figure 3 A flowchart of step 350 in a corresponding embodiment in one embodiment.
[0034] Figure 9 It is a flowchart showing the fine-tuning training of an image classification model using a first source domain image according to an embodiment of the present application.
[0035] Figure 10 It is a flowchart showing the fine-tuning training of an image classification model according to an embodiment of the present application.
[0036] Figure 11 It is a block diagram of a training device for an image classification model according to an embodiment of the present application.
[0037] Figure 12 It is a block diagram of a semantic segmentation device according to an embodiment of the present application.
[0038] Figure 13 A schematic diagram of the structure of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown. Detailed implementation manners
[0039] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.
[0040] In addition, the described features, structures or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application.
[0041] The flowcharts shown in the accompanying drawings are only illustrative and not necessarily include all contents and operations / steps, nor are they necessarily executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined. Therefore, the actual execution order may change according to the actual situation. In the following description, the terms "first / second", etc. only distinguish similar objects and do not represent a specific order for the objects. Understandably, "first / second" can be interchanged in a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0042] As used herein, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0043] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.
[0044] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0045] Computer Vision Technology (CV) Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to using electronic devices such as cameras and computers to replace the human eye for tasks such as object recognition and measurement in machine vision, and further performing graphic processing to make the computer-processed images more suitable for human eye observation or transmission to instrument detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. The large model technology has brought important changes to the development of computer vision technology. Pretrained models in the visual field such as visual models can be quickly and widely applied to downstream specific tasks after fine-tuning. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, optical character recognition, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0046] The solution provided in the embodiments of this application involves technologies such as computer vision in artificial intelligence and transfer learning based on active domain adaptation, and will be specifically described through the following embodiments.
[0047] Figure 1 It is a schematic diagram of the application scenario of this application shown according to an embodiment of this application, as Figure 1 shown, this application scenario includes a server 120, and the server 120 can be a physical server or a cloud server, and specific limitations are not made here.
[0048] The server 120 can train an image classification model according to the training method of the image classification model provided in this application, and its training process includes: Step S1, extracting features of the target domain image through the feature extraction network in the image classification model; Step S2, determining the feature map of the target domain image; Step S3, performing feature reconstruction based on the feature map to obtain a reconstructed feature map; Step S4, determining the coverage density of each pixel in the target domain image; Step S5, screening a core pixel set to be pixel-labeled in the target domain image; Step S6, performing fine-tuning training on the image classification model according to the annotation information of the pixels in the core pixel set and the target domain image. The implementation details in the training process are described below.
[0049] Based on the above training process, the image classification model can be migrated to the target domain, and the to-be-processed image from the target domain can be classified to determine the semantic segmentation result of the to-be-processed image.
[0050] Based on the image classification model after fine-tuning training, the server 120 can provide semantic segmentation services for terminals. In this case, the application scenario further includes the terminal 110, which can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a vehicle-mounted terminal, a smart TV, a smart home device, etc., and is not specifically limited herein.
[0051] The terminal 110 is communicatively connected to the server 120 through a wired or wireless network. The server 120 can receive the image to be processed sent by the terminal 110, classify the image to be processed through the image classification model after fine-tuning training, output the semantic segmentation result of the image to be processed, and send the semantic segmentation result to the terminal.
[0052] The training method of the image classification model and the method of providing semantic segmentation services for terminals based on the image classification model after fine-tuning training can be executed by the same electronic device, for example, both can be executed by the server 120, or can be executed by different electronic devices, and is not specifically limited herein.
[0053] In some embodiments, the image classification model trained according to the method of the present application can be applied to semantic segmentation of road images. For example, it can perform semantic segmentation on the road images sent by the vehicle-mounted terminal in real time to provide information on free space on the road and detect lane markings, traffic signs, etc., so as to realize automatic environmental perception. For example, it can be applied to perceive the road environment where the vehicle is located during the process of autonomous driving.
[0054] In some embodiments, the image classification model trained according to the method of the present application can be applied to semantic segmentation of human images to determine the clothing area in the human image, which is convenient for subsequent virtual clothing based on the clothing area. The above are only exemplary scenarios of image semantic segmentation, and the image classification model of the present application is not limited to being applied to the above scenarios.
[0055] In the related art, in order to transfer the model from the source domain to the target domain, the method of domain adaptation is usually adopted for transfer learning. During the process of domain adaptation, in order to reduce the annotation cost of samples in the target domain, the method of the core set is usually adopted to screen the target domain samples that need to be annotated from the target domain. Among them, the main idea of the core set method is to use a small core set (Core-set) to approximate the entire training set (i.e., the entire training set in the target domain), and tend to select diverse target domain samples. In other words, the core set method is to select a fixed-size subset from the entire training set, called the core set, with the aim of covering the entire training set with the smallest radius based on the samples in the core set to enhance the diversity of the selected samples.
[0056] However, directly applying the core-set method to the semantic segmentation task cannot solve the problem that samples (i.e., pixels) at different positions in the image feature space have different representative capabilities for other pixels within their coverage range. That is to say, the correlation between the selected samples in the target domain and their local context (such as the selected pixel in the image and its neighboring pixels) in the feature space is ignored (such as the probability distribution of other samples within the local range of the sample), resulting in a waste of the annotation budget. That is, in the scenario of pixel-level semantic segmentation tasks, if pixels with too high feature similarity are selected for annotation in the target domain, it will lead to a waste of the annotation cost.
[0057] To solve the above problems, the solution of this application is proposed. First, introduce the optimization objective of the Core-set method. Let χ represent the data space, and γ = {1, …, C} represent the annotation space, where C is the number of semantic segmentation categories. The training set consists of independently and identically distributed samples (samples are pixels in the semantic segmentation scenario) sampled from Z = χ × γ, which can be expressed as {x t , y t} t∈[n] ~p Z , where y t is the annotation information of sample x t , n is the size of the training set, and [n] is the set composed of subscripts {1, 2, …, n}. The goal of the Core-set method is to select a small set of annotated samples to minimize the Core-set loss, that is:
[0058]
[0059] Among them, is the Core-set loss, b is the annotation budget, A s represents the algorithm for learning to fit η c (x) = p(y = c|x), and l(·, ·, A s ) represents a non-negative and bounded loss function; the purpose of formula (1) is to find a set of annotated samples s such that the performance of the model on the set of annotated samples s is close to its performance on the entire training set, and the total annotation cost of the set of annotated samples s is equal to the annotation budget b, that is, |s| = b; for images, the annotation budget b can be measured by the proportion of annotated pixels in the image or the number of annotated pixels in the image.
[0060] The following Theorem 1 shows that the Core-set loss is bounded.
[0061] Theorem 1. Upper bound of the classical Core-set loss: Given the selected set s, if l(·, y, A s ) is λl -Lipsc(itz continuous, with absolute value less than L, η c (x) is λ η -Lipschitz continuous, and l(x k ,y k ,A s ) = 0, Then there exists a radius such that the following holds with probability not less than 1 - γ:
[0062]
[0063] where λ l is the Lipschitz constant corresponding to the loss function l(·, y, A s ), that is, the slope of the loss function l(·, y, A s ) is less than λ l ; similarly, λ η is also a Lipschitz constant.
[0064] According to Formula 2, minimizing the upper bound of the Core - set loss requires finding a subset s that can cover all training samples with a minimum radius δ. The radius δ can be rewritten as the following Formula 3:
[0065]
[0066] Ν s (k) is the coverage range of the sample x k , k ∈ s; arg min(.) represents the value of the variable when (.) takes the minimum value.
[0067] As shown in the above Formula (4), the upper bound of the classical Core - set loss only considers the farthest points covered by each sample in s, making it a relatively loose upper bound.
[0068] The inventors of this application found through research that the Core - set loss can be bounded by a tighter upper bound. Based on the assumptions in Theorem 1 and the coverage ranges defined by Formulas 3 and 4, the following Formula 5 holds with probability not less than 1 - γ:
[0069]
[0070] As shown in the above Formulas 5 and 6, the upper bound of the Core - set loss depends on the maximum value of δ s (k), that is, the expected distance between x k and other samples in its coverage range. Denote δ s (k) as the sample x kThe average radial distance. It can be seen that minimizing the maximum average radial distance is very important for reducing the Core-set loss.
[0071] Sample x k The average radial distance δ s (k) can be expressed as:
[0072] δ s (k) = ∫ x |x - x k |p(x|π(x) = k)dx; (Formula 7)
[0073] It can be seen from Formula 7 that δ s (k) depends on the coverage range Ν k of x s (k) of the conditional probability distribution p(x|π(x) = k), and this distribution is also called the coverage sample distribution of sample x k . When the samples are uniformly distributed in the feature space, the coverage sample distributions of all samples are the same. However, the samples are often non-uniformly distributed in the feature space. At this time, the more dispersed the coverage sample distribution of a sample is (that is, the smaller the coverage density), the larger its average radial distance.
[0074] Figure 2 shows a schematic diagram of the relationship between the estimated coverage density and the average radial distance observed in the experiment, as Figure 2 shown, the larger the estimated coverage density, the smaller the average radial distance. This rule also applies in the case of taking pixels as samples. For an image, due to the spatial continuity of the image, pixels adjacent in space are often adjacent in the semantic space. The features of the neighboring pixels of a pixel can be regarded as sampled from the coverage sample distribution p(x|π(x) = k) of the central pixel. Therefore, calculating the distance between the features of the central pixel and its neighboring pixels and averaging is equivalent to using the Monte Carlo method to approximate the average radial distance shown in Formula 7 above. However, this simple estimate may have a large deviation due to insufficient pixels in the neighborhood window. Therefore, in order to improve the quality of the estimate with a limited number of neighboring pixels, in the present application, the features of the neighboring pixels are aggregated to reconstruct the features of the central pixel. Thus, the coverage density of the pixel can be determined based on the difference between the features of the pixel and the reconstructed features of the pixel. Since the coverage density and the average radial distance are negatively correlated, the average radial distance of the labeled samples with a low coverage density is large, and as the coverage density increases, the distribution of the average radial distance also approaches zero.
[0075] Therefore, in the solution of the present application, the feature distance between two pixels is scaled by the coverage density of the pixel, so that a larger δ sSamples with a smaller coverage area are assigned, and the maximum average radial distance is thus reduced. As a result, the Core-set loss corresponding to the core set (i.e., the core pixel set to be labeled) selected based on the coverage density is also relatively small. That is, the performance of the model trained based on the core set selected by the coverage density is better. In other words, the new upper bound derived from the solution of the present application is related to the average radial distance between each sample in the training set (e.g., the pixels to be selected in the image) and the selected sample closest to the sample (e.g., the pixels that have been determined to need labeling in the image). Therefore, more labeling budgets should be allocated to those samples that are sparser in the neighborhood. That is to say, in the solution of the present application, during the process of determining the core pixel set, the local distribution of the samples is considered, thereby making the effect of training the model with the selected samples (pixels) better.
[0076] Figure 3 is a flowchart of a method for training an image classification model according to an embodiment of the present application. This method can be executed by an electronic device with processing capabilities, such as a server. Refer to Figure 3 As shown, this method at least includes steps 310 to 340, which are introduced in detail as follows:
[0077] Step 310, based on the features extracted from the target domain image by the feature extraction network in the image classification model, determine the feature map of the target domain image; the image classification model is pre-trained through source domain images and the annotation information of each pixel in the source domain images.
[0078] The present application is proposed to achieve the migration of the image classification model from the source domain to the target domain. The source domain refers to the domain where the existing knowledge belongs in transfer learning. In the present application, the source domain can also be understood as the domain where the image classification model is applied before migration. The target domain in transfer learning refers to the domain of new knowledge to be learned, and it can also be understood that the target domain is the domain to which the image classification model is to be migrated.
[0079] The target domain image refers to an image from the target domain, and the source domain image refers to an image from the source domain. For example, the source domain image can be an image collected in a scene with normal imaging conditions (such as lighting, shooting position, shooting angle), and the target domain image can be an image collected in a scene with abnormal imaging conditions (such as strong or weak lighting, a relatively biased shooting angle, etc.). Another example is that the source domain image is a road image collected in area A, and the target domain image is a road image collected in area B. In this way, even if the geographical environments of area A and area B cause large differences in the collected road images, the image classification model trained according to the method of the present application can accurately classify the road images to be classified collected in area B.
[0080] Before migration, the image classification model first uses the source domain images and the annotation information of each pixel in the source domain images. The pre-trained image classification model is used as a pre-training model (Pre-training model), also known as the foundation model or large model, which refers to a deep neural network (Deep neural network, DNN) with a large number of parameters. It is trained on a large amount of data, and the function approximation ability of the large-parameter DNN is used to enable the PTM to extract common features from the data. Through techniques such as fine-tuning, parameter-efficient fine-tuning (Parameter-Efficient Fine-Tuning, PEFT), and prompt-tuning, the pre-training model can be migrated to other scenarios to be applicable to downstream tasks.
[0081] In this application, the classification task of the image classification model can perform pixel-level classification on the image to achieve semantic segmentation of the image. That is, the classification task of the image classification model is the semantic segmentation task, which is to identify the objects presented in the image and all the pixels belonging to each object. Therefore, during the pre-training process, pixel-level annotation is performed on the source domain images to obtain the annotation information of each pixel in the source domain images. This annotation information is used to indicate the category to which the pixel belongs. After that, the image classification model is pre-trained using the source domain images and the annotation information of the pixels in the source domain images. After pre-training the image classification model with a large number of source domain images and the corresponding annotation information, the image classification model can accurately perform pixel-level classification on the images in the source domain, that is, accurately perform semantic segmentation.
[0082] The pre-training process of the image classification model can include: the image classification model extracts features from the source domain images and then performs classification to obtain the classification results of the source domain images. This classification result indicates the classification to which each pixel in the source domain image belongs, or indicates the probability that each pixel in the source domain image belongs to each category. After that, according to the classification results of the source domain images and the annotation information of the pixels in the source domain images, the prediction loss is calculated, and the parameters of the image classification model are adjusted according to the prediction loss until the first training end condition is reached. The first training end condition can be that the number of iterations reaches the first number threshold, the loss function converges, etc.
[0083] The image classification model can be constructed through one or more neural networks, such as convolutional neural networks, fully connected neural networks, etc., which will not be specifically limited herein. The image classification model includes a feature extraction network and a classifier. The feature extraction network is used to extract features from the input image. For the convenience of distinction, the classifier in the image classification model is called the first classifier, and this first classifier is used to predict the probability that each pixel belongs to each category for the result output by the feature extraction network, so as to output the classification result of each pixel. In some embodiments, the feature extraction network can be a Resnet network (such as Resnet101), or a Swin Transformer, which will not be specifically limited herein.
[0084] In some embodiments, the feature extraction network can extract features from the target domain image, and use the feature extraction result output by the feature extraction network as the feature map of the target domain image.
[0085] In other embodiments, after the feature extraction network extracts features from the target domain image and obtains the feature extraction result, the result obtained by further processing (such as dimensionality reduction processing, etc.) of the feature extraction result can also be used as the feature map of the target domain image.
[0086] Step 320: Based on the feature map, perform feature reconstruction to obtain a reconstructed feature map. The reconstructed feature of a pixel in the reconstructed feature map is reconstructed according to the features of the neighboring pixels of this pixel.
[0087] In this application, the feature represented in the reconstructed feature map is called the reconstructed feature. In step 320, for the pixel in the target domain image, the feature of this pixel is reconstructed through the features of the neighboring pixels of this pixel in the feature map, and used as the reconstructed feature of this pixel in the reconstructed feature map.
[0088] In some embodiments, considering that in an image, the pixel information of a pixel has a high similarity with the pixel information of the neighboring pixels of this pixel, therefore, the features of the neighboring pixels of the pixel can be obtained from the feature map, and the features of the neighboring pixels are weighted and fused as the reconstructed feature of the pixel. The neighboring pixels of a pixel refer to other pixels close to this pixel in the target domain image. For example, for pixel A, the pixels adjacent to pixel A in the target domain image and located above, below, left, and right of pixel A respectively can be used as the neighboring pixels of pixel A; another example is that the pixels centered on pixel A and with a distance from pixel A not exceeding the first distance threshold in the target domain image can be used as the neighboring pixels of pixel A.
[0089] In some embodiments, step 320 includes the following steps A1 to A2:
[0090] Step A1: Perform dynamic masked convolution on the feature map to obtain a spatial modulation map, where the eigenvalue of the central pixel of the convolution window in each convolution of the spatial modulation map is 0.
[0091] The dynamic masked convolution process at least includes dynamic convolution processing and masking processing. Dynamic convolution processing refers to adjusting convolution parameters (such as the size of the convolution kernel) adaptively according to the image for convolution processing. That is, during the dynamic convolution process, for different images, the convolution parameters adapted to the image can be used to process the image. Masking processing refers to setting the features of some positions in the image to a specified value. For example, in Step A1, the value of the feature of the central pixel in the convolution window of each convolution can be set to a specified value (such as -∞) to ensure that the eigenvalue of the central pixel of the convolution window in each convolution of the spatial modulation map is 0.
[0092] In some embodiments, Step A1 includes: performing dynamic convolution on the feature map to obtain an intermediate feature map; superimposing the intermediate feature map with the mask tensor corresponding to the convolution window to obtain a reference feature map; and normalizing the features in the reference feature map to obtain a spatial modulation map.
[0093] It is worth mentioning that since the sizes of the convolution kernels used during the dynamic convolution process may be different, during the dynamic convolution process, the sizes of the convolution windows in each convolution may be different. For example, if the size of the convolution kernel in one convolution is 3*3, correspondingly, the size of the convolution window in this convolution is 3*3. The mask tensor corresponding to the convolution window corresponds to the size of the convolution window, that is, if the size of the convolution window is a*a, the mask tensor corresponding to this convolution window is a matrix of a*a. In the mask tensor, the value corresponding to the central pixel is specified, such as -∞, and the values of other positions are 0. For example, a mask tensor corresponding to a 3*3 convolution window can be represented as: Based on the mask tensors corresponding to each convolution window, after superimposing the intermediate feature map with the mask tensor corresponding to the convolution window, the feature representing the central pixel in the convolution window in the reference feature map can be set to -∞, and the features of other pixels that are not the central pixel of the convolution window are the same as their values in the intermediate feature map.
[0094] In a specific embodiment, the features in the reference feature map can be normalized based on the softmax function. Since the feature representing the central pixel in the convolution window in the reference feature map is set to -∞, correspondingly, the feature representing the central pixel in the convolution window in the normalized spatial modulation map is set to zero.
[0095] Step A2: Fuse the spatial modulation map and the feature map to obtain a reconstructed feature map.
[0096] In some embodiments, the spatial modulation map and the feature map can be multiplied bit by bit (i.e., element-wise product or Schur product), and the resulting result can be used as the reconstructed feature map.
[0097] In some embodiments, to ensure the feature representation ability of the obtained reconstructed feature map, the spatial modulation map and the feature map can also be multiplied bit by bit to obtain a bit-by-bit multiplication result, and the bit-by-bit multiplication result can be convolved through a 1×1 convolution kernel, and the result of the convolution can be used as the reconstructed feature map.
[0098] In a specific embodiment, the process of step 320 above can be implemented by a dynamic mask convolution network. The dynamic mask convolution network includes a second convolution network, a non-linear layer, and a third convolution network. Among them, the second convolution network is used for dynamic convolution, and the learnable parameters of the dynamic convolution of the second convolution network can be determined during the training process. The non-linear layer can be set with a softmax function to normalize the reference feature map; a 1×1 convolution kernel can be set in the third convolution network. The process of feature reconstruction based on this dynamic mask convolution network can be as Figure 4 shown. The feature map F (W×H×D, where W is the width, H is the height, and D is the number of channels) is processed by the spatial modulation map generator φ(·) to obtain the spatial modulation map M. Among them, the spatial modulation map generator φ(·) is composed of the second convolution network h(·) and the non-linear layer softmaxσ(·). The spatial modulation map can be expressed as:
[0099]
[0100] Among them, M -∞ represents the mask tensor. In this mask tensor, the th channel is set to -∞, and the values of other channels are zero to exclude the influence of the central pixel; K is the convolution kernel size of the dynamic convolution. In Formula 1, h(F) is the intermediate feature map M1 obtained by performing dynamic convolution. The intermediate feature map M1 and the mask tensor M -∞ are superimposed to obtain the reference feature map M2. Then, the reference feature map M2 is normalized through the non-linear layer softmaxσ(·) to obtain the spatial modulation map
[0101] Among them, the spatial modulation map is fused with the feature map, and the result of the fusion is used as the intermediate reconstructed feature map. The intermediate reconstructed feature map can be expressed as:
[0102]
[0103] Among them, Represents the intermediate reconstructed feature map The feature of the pixel at position (i, j) in the output channel o, where d represents the channels in F, d ∈ D; ΔK represents (u, v) represents the position of the parameter in the parameter matrix corresponding to the convolution kernel during the dynamic convolution process Represents the spatial modulation map The feature obtained after the pixel (i, j) is convolved through the parameter matrix corresponding to the convolution kernel Are the learnable parameters of the dynamic convolution
[0104] Subsequently, the intermediate reconstructed feature map is improved through the third convolution network ψ(·) To obtain the reconstructed feature map
[0105] Step 330: Determine the coverage density of each pixel in the target domain image according to the feature map and the reconstructed feature map
[0106] The coverage density of a pixel refers to a statistic for estimating the coverage sample distribution. The coverage density of a pixel is used to reflect the distribution of the features of the pixel in the image, that is, the higher the coverage density of the pixel, the denser the distribution of the features of the pixel in the image. Since the image is continuous in space, pixels adjacent in space are often adjacent in the semantic space. Therefore, the features of neighboring pixels can be regarded as sampled from the coverage sample distribution (i.e., the conditional probability distribution) p(x|π(x) = k) of the central pixel x k However, this simple estimation may have a large deviation due to insufficient pixels in the neighborhood window. To improve the quality of the estimation with a limited number of neighboring pixels, in this application, the features of neighboring pixels are aggregated through dynamic masked convolution to reconstruct the features of the central pixel
[0107] Specifically, step 330 includes: for each pixel, obtain the feature of the pixel from the feature map and the reconstructed feature of the pixel from the reconstructed feature map; calculate the feature difference between the feature of the pixel and the reconstructed feature of the pixel; determine the coverage density of the pixel according to the feature difference of the pixel
[0108] The above process can be represented by the following formula 3. The coverage density D of the pixel (i, j) i,j Is:
[0109]
[0110] Where β and τ are hyperparameters and can be preset Represents the reconstructed feature of the pixel (i, j), and F :,i,j Represents the feature of the pixel (i, j); Denote the reconstruction error at the pixel (i, j); Denote the square of the L2 norm. Can be regarded as an estimate of the average radial distance. Pixel x k The average radial distance δ s (k) refers to pixel x k And this pixel x k The expected distance between this pixel x
[0111] Step 340, according to the coverage density of each pixel in the target domain image and the features of each pixel in the feature map, screen the core pixel set to be pixel-labeled in the target domain image; the coverage density is used to scale the feature distance between two pixels.
[0112] The purpose of determining the core pixel set is: after labeling the pixels in the core pixel set, the training effect of training the image classification model using the labeling information of the pixels in the core pixel set is close to the training effect of training the image classification model by labeling all the pixels in the target domain image, or in other words, the difference in the training effect is less than a preset threshold. In this way, since the training effects of the two image classification models are close, only the pixels in the core pixel set need to be labeled, and there is no need to label the other pixels in the target domain image except the core pixel set. In this way, on the basis of ensuring the training effect of the image classification model, the amount of pixel labeling is reduced, and the labeling cost is reduced.
[0113] The core pixel set to be pixel-labeled can be screened by combining the greedy algorithm and the coverage density of the pixels in the target domain image. First, combine Figure 5A To illustrate the principle of determining the core set according to the greedy algorithm. Suppose Figure 5A Pixels K1-K3 in are the pixels to be determined whether to be labeled (for easy distinction, the set of pixels to be determined whether to be labeled is called the second pixel set, and correspondingly, the pixels to be determined whether to be labeled are called the second pixels), and the other pixels except pixels K1-K3 are the pixels that have been determined to need to be labeled (for easy description, the set of pixels that have been determined to need to be labeled is called the first pixel set, and the pixels that have been determined to need to be labeled are called the first pixels).
[0114] Such as Figure 5AAs shown, for each second pixel, the feature distance between the second pixel and each first pixel in the first pixel set is calculated (this feature distance is calculated based on the features of the first pixel and the second pixel). After that, the second pixel is connected to the first pixel in the first pixel set with the shortest feature distance to the second pixel (for easy distinction, the determined first pixel is called the first target pixel), and the feature distance between the second pixel and the first target pixel determined for this second pixel is called the feature distance between the second pixel and the first pixel set. Then, based on the feature distances between each second pixel and the first pixel set, sample selection is performed, that is, the second pixel with the largest feature distance from the first pixel set (i.e., Figure 5A the pixel M1 in) is added to the first pixel set, and this process is repeated, and the finally output first pixel set is used as the core pixel set.
[0115] The process of screening the core pixel set to be pixel-labeled by combining the greedy algorithm and the coverage density of pixels in the target domain image can be as Figure 5B shown. Compared with Figure 5A the process shown, Figure 5B the process shown adds a new step between the step of distance calculation and the step of sample selection, that is, density-aware distance weighting, that is, using the coverage density to scale the feature distance between two pixels. Specifically, after determining the first target pixel for each second pixel, based on the coverage density of the first target pixel, the feature distance between the second pixel and the first target pixel determined for this second pixel is scaled as the feature distance between the second pixel and the first pixel set. As Figure 5B shown, the coverage density of the first target pixel determined for the pixel K1 as the second pixel is small, and the coverage densities of the pixels K3 and K2 as the first target pixels are large. After scaling based on the coverage density, the feature distance between the second pixel connected to the pixel K3 and the first pixel set is reduced compared to Figure 5A the feature distance shown. In Figure 5B this case, after scaling based on the coverage density, the feature distance between the pixel M2 and the first pixel set is the largest. Therefore, the pixel M2 is added to the first pixel set instead of the pixel M1 being added to the first pixel set.
[0116] The method of screening the core pixel set to be pixel-labeled by combining the greedy algorithm and the coverage density of pixels in the target domain image is equivalent to scaling the distance in the coverage range of the first target pixel by the coverage density of the first target pixel. Through this step, when the coverage density of the pixel is highly correlated with the average radial distance, through this step, the selected pixel can be converted from the farthest edge to the edge with the largest δ s Therefore, with a larger δ sSamples are assigned a smaller coverage area, and thus the maximum average radial distance is minimized.
[0117] In some embodiments, as Figure 6 shown, step 340 includes at least the following steps 610 - 660, which are introduced in detail as follows:
[0118] Step 610, determine a first pixel set, where the first pixels in the first pixel set are pixels in the target domain image.
[0119] The first pixel set can be a set of pixels that need to be labeled determined during initialization. The pixels in this first pixel set can be any specified number of pixels from the target domain image. After that, based on the pixels in the first pixel set, a core pixel set is determined. For ease of distinction, the pixels in the first pixel set are called first pixels.
[0120] Step 620, for each second pixel in the second pixel set, determine a first target pixel in the first pixel set with the smallest feature distance to this second pixel; the feature distance is determined according to the features of the two pixels; the second pixels in the second pixel set are pixels in the target domain image, and the intersection of the second pixel set and the first pixel set is empty.
[0121] The second pixel set is a set of pixels for which it is to be determined whether pixel labeling is required, that is, to screen the pixels to be pixel-labeled from the second pixel set. This second pixel set can be a set of all other pixels in the target domain image except the first pixel set, or a set of some pixels selected from the other pixels in the target domain image except the first pixel set. For ease of distinction, the pixels in the second pixel set are called second pixels.
[0122] In step 620, for each second pixel, calculate the feature distance between the second pixel and each first pixel according to the features of the second pixel and the first pixel. Then, take the first pixel in the first pixel set with the smallest feature distance to the second pixel as the first target pixel corresponding to this second pixel. Through this process, the corresponding first target pixel can be determined for each second pixel respectively.
[0123] In some embodiments, the L2 norm result between the feature of the second pixel and the feature of the first pixel can be used as the feature distance between the two. In other embodiments, the Euclidean distance or cosine distance between the feature of the second pixel and the feature of the first pixel can also be used to determine the feature distance between the two.
[0124] Step 630: Scale the feature distance between the second pixel and the corresponding first target pixel according to the coverage density of the first target pixel corresponding to the second pixel, to obtain the reference distance between the second pixel and the first pixel set, where the reference distance is negatively correlated with the coverage density.
[0125] That is, for each second pixel, use the result obtained by scaling the feature distance between the second pixel and the corresponding first target pixel by the coverage density of the first target pixel corresponding to the second pixel as the reference distance between the second pixel and the first pixel set.
[0126] In some embodiments, the feature distance between the second pixel and the corresponding first target pixel can be divided by the coverage density of the first target pixel corresponding to the second pixel, and the obtained result is used as the reference distance between the second pixel and the first pixel set.
[0127] Step 640: Determine the second target pixel in the second pixel set with the largest reference distance from the first pixel set.
[0128] The second target pixel refers to the second pixel in the second pixel set with the largest reference distance from the first pixel set.
[0129] Step 650: Add the second target pixel to the first pixel set.
[0130] Step 660: Determine the core pixel set based on the updated first pixel set.
[0131] The updated first pixel set is at least a subset of the core pixel set. In some implementations, the updated first pixel set can be used as the core pixel set. In some embodiments, Step 660 includes the following Steps 661 - 663, which are introduced in detail as follows:
[0132] Step 661: Count the number of pixels in the updated first pixel set.
[0133] Step 662: Determine whether the annotation requirement is met according to the number of pixels; if yes, execute Step 663; if no, execute Step 670.
[0134] The annotation requirement can be a pixel annotation quantity threshold. If the number of pixels reaches the pixel annotation quantity threshold, it is determined that the annotation requirement is met. The annotation requirement can also be an annotation pixel ratio. In this case, the ratio of the number of pixels in the updated first pixel set to the total number of pixels in the target domain image can be calculated. If this ratio reaches the annotation pixel ratio, it is determined that the annotation requirement is met. This annotation requirement is used to limit the annotation cost. The more pixels need to be annotated, the higher the annotation cost. Therefore, based on the annotation requirement, the annotation cost can be restricted.
[0135] Step 663: Use the updated first pixel set as the core pixel set.
[0136] Correspondingly, the method further includes step 670: Remove the second target pixel from the second pixel set to obtain a new second pixel set. Then return to execute step 620.
[0137] The process shown in steps 610 to 670 above can be described in Table 1 below:
[0138] Table 1
[0139]
[0140]
[0141] In Table 1 above, is the minimum feature distance between the first pixel set s 0 and the second pixel; d k is the coverage density of the first target pixel x t corresponding to the second pixel x k ; r t is the reference distance between the second pixel x t and the first pixel set s 0 . u is the second target pixel with the largest reference distance between the second pixel set and the first pixel set. The finally output pixel set s is the core pixel set.
[0142] Step 350: Fine-tune and train the image classification model according to the annotation information of the pixels in the core pixel set and the target domain image.
[0143] Among them, the annotation information of the pixels in the core pixel set can be obtained by an annotator annotating the pixels in the core pixel level, and this annotation information indicates the category to which the pixel belongs. For each target domain image, the core pixel level to be annotated for this target domain image can be determined according to the process shown in steps 310 - 340 above.
[0144] During the process of fine-tuning and training the image classification model, the target domain image can be feature-extracted through the feature extraction network in the image classification model, and then pixel classification is performed to output a pixel classification result. This pixel classification result can include the predicted categories of the pixels in the core pixel set. Subsequently, according to the predicted categories of the pixels in the core pixel set and the annotation information of the pixels in the core pixel set, the model loss is calculated, and the parameters of the image classification model are adjusted backward according to the model loss until the training end condition is reached.
[0145] In some embodiments, since determining the core pixel set requires using the feature extraction network to extract features from the target domain image, and then the dynamic mask convolution network is needed to reconstruct the feature map to estimate the coverage density of pixels, and the accuracy of the coverage density directly affects the training effect of the image classification model. Therefore, to balance the accuracy of the coverage density and the training effect of the image classification model, the dynamic mask convolution network can be jointly trained during the fine-tuning training process of the image classification model, that is, the parameters of the image classification model and the dynamic mask convolution network are updated by combining the determined loss.
[0146] In this application, the process of determining the core pixel level is the process of actively selecting samples. Since the process of actively selecting samples needs to perform feature extraction based on the feature extraction network of the image classification model, the active selection of samples and the fine-tuning training of the image classification model can be alternated. For example, after initially determining the core pixel set for some target domain images and performing annotation, the image classification model can be fine-tuned for a period of time (such as iterating N times) based on the annotated core pixel level and the corresponding target domain image; then, the image classification model iterated N times during the fine-tuning training is used to actively select samples (i.e., determine the core pixel set) for other target domain images; then iterate, actively select samples, and repeat the above process. For example, the active selection of samples can be performed when the image classification model iterates 10000, 12000, 14000, 16000, and 18000 steps during the fine-tuning training. Since the more iterations, the better the performance of the image classification model and the higher the accuracy of feature extraction by the feature extraction network in the corresponding image classification model, by alternately performing the active selection of samples in this way, it is ensured that the features extracted later are more accurate. In this way, the more accurate the coverage density determined for the pixels, the more representative the determined core pixel set, and thus the training effect of the image classification model is further optimized.
[0147] In this application, after obtaining the feature map of the target domain image, the feature of each pixel is reconstructed based on the features of the neighboring pixels of the feature image pixel to obtain a reconstructed feature map. Then, based on the feature map and the reconstructed feature map, the coverage density of each pixel in the target domain image is determined. This coverage density can reflect the correlation between the pixel and its neighboring pixels in the feature space. Based on this, the feature distance between two pixels is scaled according to the coverage density of each pixel in the target domain image, and based on this, a core pixel set to be pixel-labeled is screened in the target domain image. Since the feature distance between two pixels is scaled by the coverage density, that is, if the coverage density of a pixel is large, the feature distance after scaling based on the coverage density decreases; conversely, if the coverage density of a pixel is small, the feature distance after scaling according to the coverage density increases. In this way, it is possible to avoid selecting the neighboring pixels of the pixels in the neighborhood of a pixel in the target domain image that have a high correlation with the pixel into the core pixel set, which can ensure the training effect of the image classification model using the pixels in the core pixel level and avoid waste of annotation budget. After determining the core pixel set for the target domain image, only the pixels in the core pixel set need to be labeled, rather than all the pixels in the target domain image. This can greatly reduce the amount of pixel labeling, reduce the cost of pixel labeling, and ensure the fine-tuning training effect of the image classification model, enabling the image classification model to effectively perform pixel-level classification on the images in the target domain.
[0148] In some embodiments, the image classification model further includes a first classifier; as Figure 7 shown, before step 620, the method further includes the following steps 710 - step 740, which are introduced in detail as follows:
[0149] Step 710, the first classifier classifies based on the result of feature extraction of the target domain image by the feature extraction network to obtain a first classification result, and the first classification result indicates the probability that each pixel in the target image belongs to each of multiple categories.
[0150] For the sake of distinction, the classification result output by the first classifier for the target domain image is referred to as the first classification result. It can be understood that if the result of feature extraction of the target domain image by the feature extraction network is used as the feature map in the above steps, in step 710, the first classifier classifies this feature map to obtain the first classification result.
[0151] Step 720, for each pixel, determine the largest N target probabilities for the pixel among the probabilities that the pixel belongs to multiple categories, where N is a positive integer.
[0152] For example, if N is 1, the maximum probability of a pixel is used as the target probability; if N is 2, the maximum probability and the second largest probability among the probabilities that a pixel belongs to multiple categories are used as the target probabilities. Among them, N is less than the total number of categories.
[0153] Step 730: Calculate the annotation gain of each pixel according to the N target probabilities determined for each pixel.
[0154] The annotation gain of a pixel refers to the gain brought to the training of the image classification model by annotating this pixel. For an image classification model, if the target probability corresponding to a pixel is larger, it indicates that the image classification model has a stronger recognition ability for this pixel. Thus, if this pixel is subsequently used for fine-tuning the training of the image classification model, the actual improvement in the performance of the image classification model is not significant. Based on this principle, the annotation gain of a pixel is calculated according to the N target probabilities determined for the pixel. The annotation gain of a pixel is negatively correlated with the target probability corresponding to this pixel.
[0155] In some embodiments, for each pixel, the sum of the N target probabilities determined for this pixel can be calculated, and then the difference between 1 and the sum of the probabilities is used as the annotation gain of the pixel.
[0156] In some embodiments, if N is 2, the N target probabilities determined for a pixel include the maximum probability and the second-largest probability among the probabilities that the pixel belongs to multiple categories; Step 730 includes: for each pixel, calculate the sum of the maximum probability and the second-largest probability corresponding to this pixel to obtain a reference result for this pixel; for each pixel, subtract the reference result corresponding to this pixel from 1 to obtain the annotation gain of this pixel.
[0157] The above process can be described by the following formula 12:
[0158] I i,j = 1 - (P 1*,i,j + P 2*,i,j ); (Formula 12)
[0159] where I i,j is the annotation gain of pixel (i, j), P 1*,i,j is the maximum probability corresponding to pixel (i, j); P 2*,i,j is the second-largest probability corresponding to pixel (i, j).
[0160] Step 740: Screen out the pixels in the target image whose annotation gain is greater than the annotation gain threshold, and construct a second pixel set according to the pixels whose annotation gain is greater than the annotation gain threshold.
[0161] Among them, the annotation gain threshold can be set according to actual needs. It is worth mentioning that the number of pixels filtered based on the set annotation gain threshold is not less than the pixel annotation quantity threshold defined by the annotation requirements. Of course, the pixels in the target domain image can also be sorted in descending order of annotation gain, and then the first set number of pixels can be selected and added to the second pixel set. The set number is greater than the maximum number of pixels required in the core pixel set.
[0162] In some embodiments, the annotation gain threshold can be determined in combination with the annotation requirements. For example, based on the annotation requirements, the maximum number of pixels required in the core pixel set can be determined, and it can be magnified based on this maximum number of pixels. For example, if the maximum number of pixels is K, the maximum number of pixels in the second pixel set can be set to αK, where α is greater than 1.
[0163] In the above embodiments, based on the annotation gain of each pixel in the target domain image, the pixels with annotation gain greater than the annotation gain threshold are filtered and added to the second pixel set. In this way, subsequently, only the pixels that need to be annotated are selected from the second pixel set, rather than performing the Figure 6 process shown for all other pixels in the target domain image except the first pixel set. In this way, the time for determining the core pixel set can be shortened, and the efficiency of actively selecting samples can be improved.
[0164] In some embodiments, the image classification model further includes a first classifier; the reconstructed feature map is obtained by the dynamic mask convolutional network reconstructing based on the feature map; the feature map is output by the first convolutional network processing the first feature map, and the first feature map is obtained by the feature extraction network extracting features from the target domain image; that is to say, in this embodiment, the feature map is obtained by the first convolutional network further processing the first feature map. For example, the first convolutional network can perform dimensionality reduction processing on the first feature map. For example, the first convolutional network can be a convolutional network with a bottleneck structure. In this way, after dimensionality reduction processing, the processing amount for subsequent feature reconstruction can be reduced, and the efficiency of subsequent determining the coverage density can be improved. As Figure 8 shown, step 350 includes the following steps:
[0165] Step 810, obtain the first classification result obtained by the first classifier classifying based on the features extracted from the target domain image by the feature extraction network. As described above, the first classification result indicates the probability that each pixel in the target domain image belongs to each of multiple categories. Of course, it can be determined according to.
[0166] Step 820, determine the first loss according to the first classification result and the annotation information of the pixels in the core pixel set.
[0167] Based on mean squared error loss function, cross-entropy loss function, absolute value loss function, etc., the loss can be calculated according to the first classification result and the annotation information of the pixels in the core pixel set, and the calculated loss is called the first loss. The first classification result reflects the category to which each pixel in the predicted target domain image belongs, and the first classification result also reflects the category to which each pixel in the predicted core pixel set belongs. The first loss determined based on the first classification result and the annotation information of the pixels in the core pixel set can reflect the difference between the actual category to which the pixels in the core pixel set belong and the predicted category.
[0168] Step 830: Determine the feature reconstruction loss according to the feature map and the reconstructed feature map.
[0169] This feature reconstruction loss is used for the reconstruction error between the reconstructed feature map obtained by the dynamic mask convolutional network and the output feature map. The purpose of training based on the feature reconstruction loss is to minimize the reconstruction error. In some embodiments, the mean square error (MSE) between the feature map and the reconstructed feature map can be calculated as the feature reconstruction loss.
[0170] As shown in Equation 11 above, the coverage density is negatively correlated with the reconstruction error of pixel (i, j). Therefore, minimizing the reconstruction error is equivalent to maximizing the coverage density, that is, the training objective of the dynamic mask convolutional network can be expressed as: where g represents the neural network that processes the first feature map to obtain the feature map. For example, this neural network can be a convolutional network with a bottleneck structure to achieve the purpose of dimensionality reduction; φ represents the above spatial modulation map generator, and ψ represents the above third convolutional network. where g represents the neural network that processes the first feature map to obtain the feature map. For example, this neural network can be a convolutional network with a bottleneck structure to achieve the purpose of dimensionality reduction; φ represents the above spatial modulation map generator, and ψ represents the above third convolutional network.
[0171] Step 840: Adjust the parameters of at least one of the image classification model and the dynamic mask convolutional network according to the first loss and the feature reconstruction loss.
[0172] In some embodiments, in each iteration process, the target loss can be determined according to the first loss, the second loss, and the feature reconstruction loss, and the parameters of the image classification model and the dynamic mask convolutional network can be adjusted according to the target loss.
[0173] In some embodiments, the feature map is obtained by reducing the dimensionality of the first feature map through the first convolutional network, and the first feature map is the feature extracted by the feature extraction network for the target domain image. The method further includes: Step B1: Classify based on the feature map by the second classifier to obtain the second classification result; Step B2: Determine the second loss according to the second classification result and the annotation information of the pixels in the core pixel set.
[0174] In step B1, the feature map is used as the input of the second classifier. The second classifier performs classification based on the feature map, and the classification result output by the second classifier is called the second classification result. The second classification result also correspondingly indicates the probabilities of each pixel in the target domain image predicted based on the feature map belonging to each of multiple categories. The second classifier can be a classifier different from the first classifier or the same classifier. It is worth mentioning that when the first classifier and the second classifier are the same classifier, due to different inputs (i.e., the first feature map is input once and the feature map is input another time, and the feature map is obtained by dimensionality reduction processing of the first feature map), there may be differences in the classification results obtained from the two classifications.
[0175] Similarly, in step B2, based on mean squared error loss function, cross-entropy loss function, absolute value loss function, etc., the loss can be calculated according to the second classification result and the annotation information of the pixels in the core pixel set, and the calculated loss is called the second loss. The second classification result reflects the category to which each pixel in the predicted target domain image belongs, and the second classification result also reflects the category to which each pixel in the predicted core pixel set belongs. The second loss determined based on the second classification result and the annotation information of the pixels in the core pixel set can reflect the difference between the actual category to which the pixels in the core pixel set belong and the predicted category.
[0176] Correspondingly, step 840 may include: if the number of iterations is less than the first number threshold, determine the target loss according to the first loss, the second loss, and the feature reconstruction loss; adjust the parameters of the image classification model, the dynamic mask convolutional network, the first convolutional network, and the second classifier according to the target loss. If the number of iterations is not less than the first number threshold, adjust the parameters of the image classification model according to the first loss; and adjust the parameters of the dynamic mask convolutional network, the first convolutional network, and the second classifier according to the weighted result of the second loss and the feature reconstruction loss.
[0177] That is to say, after the number of iterations in the fine-tuning training process reaches the first number threshold, the gradient from the dynamic mask convolutional network will not be passed back to the feature extraction network, thereby avoiding the gradient of the dynamic mask convolutional network from affecting the convergence of the image classification model (especially the feature extraction network).
[0178] In the above embodiments, since the first feature map output by the feature extraction network is not directly used for feature reconstruction, but the feature map obtained by performing dimensionality reduction on the first feature map is used as the basis for feature reconstruction. To avoid the dynamic mask convolutional network simply erasing the differences between different pixels during the fine-tuning training process due to the feature reconstruction loss, resulting in learning the reconstructed features of different pixels into the same features, a second classifier as an auxiliary is introduced, and the second loss calculated based on the second classification result output by the second classifier based on the feature map is also used to supervise the fine-tuning training process of the image classification model. Thus, introducing the second loss can effectively supervise the accuracy of feature reconstruction and enable the dynamic mask convolutional network to retain the class information of different pixels during the feature reconstruction process.
[0179] In some other embodiments, if the features extracted by the feature extraction network for the target domain image are used as the feature map and no first convolutional network for dimensionality reduction is introduced between the feature extraction network and the dynamic mask convolutional network, in this case, the first loss can be used to avoid the dynamic mask convolutional network simply erasing the differences between different pixels during the fine-tuning training process due to the feature reconstruction loss. Therefore, in this case, an auxiliary second classifier does not need to be introduced. Correspondingly, based on the first loss and the feature reconstruction loss, the parameters of at least one of the image classification model and the dynamic mask convolutional network can be adjusted; for example, if the number of iterations is less than the first number threshold, the target loss is determined according to the first loss and the feature reconstruction loss; according to the target loss, the parameters of the image classification model and the dynamic mask convolutional network are adjusted; if the number of iterations is not less than the first number threshold, the parameters of the image classification model are adjusted according to the first loss; and the parameters of the dynamic mask convolutional network are adjusted according to the weighted result of the first loss and the feature reconstruction loss. The weighting coefficients of the first loss and the feature reconstruction loss can be the same (for example, both are 1) or different, and are not specifically limited herein.
[0180] In some embodiments, the method further includes: fine-tuning the image classification model according to the first source domain image and the label information of the first source domain image.
[0181] Wherein, the first source domain image refers to the source domain image used for fine-tuning the image classification model. The first source domain image can be a part of the source domain images selected from the source domain images used in the pre-training process. Since the source domain images have been pre-annotated, there is no need to screen the pixels that need to be annotated for the first source domain image.
[0182] In some embodiments, as Figure 9 shown, the steps of fine-tuning the image classification model according to the first source domain image and the label information of the first source domain image include the following steps 910 - step 917:
[0183] Step 910: The feature extraction network extracts features from the first source domain image to obtain the first feature map of the first source domain image.
[0184] Step 911: The first classifier classifies based on the first feature map of the first source domain image to obtain the third classification result.
[0185] Step 912: According to the annotation information and the third classification result of the first source domain image, determine the first loss for the first source domain image.
[0186] Step 913: Determine the feature map of the first source domain image based on the first feature map of the first source domain image.
[0187] In some embodiments, the first feature map of the first source domain image can be used as the feature map of the first source domain image. In other embodiments, the first feature map of the first source domain image can be dimensionally reduced by the first convolutional network to obtain the feature map of the first source domain image.
[0188] Step 914: The dynamic mask convolutional network performs feature reconstruction on the feature map of the first source domain image to obtain the reconstructed feature map of the first source domain image.
[0189] Step 915: According to the feature map and the reconstructed feature map of the first source domain image, determine the feature reconstruction loss for the first source domain image.
[0190] The implementation details of the above steps 910 - 915 are similar to the process of training based on the target domain image, and will not be elaborated here.
[0191] Step 916: According to the probabilities that each pixel in the third classification result belongs to the actual class and belongs to other classes, determine the third loss for the first source domain image.
[0192] In some embodiments, the third loss for the first source domain image can be calculated according to the following formula 13:
[0193] L ML =∑ i ∑ j ∑ c≠y [m - P y,i,j +P c,i,j + ;(Formula 13)
[0194] where, L ML represents the third loss; [x] + represents taking the maximum value between 0 and x; P y,i,j is the probability that the pixel (i, j) determined by classification belongs to the actual class corresponding to this pixel, and y can also be understood as representing the channel of the actual class corresponding to the pixel (i, j); Pc,i,j is the probability that the pixel (i, j) belongs to other classes except the actual class corresponding to the pixel. c can also be understood as representing the channels of other classes except the actual class corresponding to the pixel. Where m is a preset constant. For example, m can be 1, 1.1, 1.5, etc., and specific limitations are not provided here. It can be understood that m - P y,i,j +P c,i,j can reflect the discrimination accuracy of pixels of different classes. For a pixel, if P y,i,j is higher, correspondingly, m - P y,i,j +P c,i,j is smaller, indicating that the recognition accuracy of the class represented by channel y is higher.
[0195] Step 917: Adjust the parameters of at least one of the image classification model and the dynamic mask convolution network according to the first loss, the feature reconstruction loss, and the third loss for the first source domain image.
[0196] In some embodiments, during each iteration process, the first target loss can be determined according to the first loss, the feature reconstruction loss, and the third loss for the first source domain image, and the parameters of the image classification model and the dynamic mask convolution network can be adjusted according to the first target loss.
[0197] In some embodiments, if the feature map for the first source domain image is obtained by reducing the dimension of the first feature map of the first source domain image through the first convolution network, and the first feature map is the feature extracted by the feature extraction network for the first source domain image; the method further includes the following steps C1 and C2: Step C1, the second classifier classifies based on the feature map of the first source domain image to obtain the fourth classification result. Step C2, according to the fourth classification result and the label information of the first source domain image, determine the second loss for the first source domain image.
[0198] Correspondingly, step 917 can include: If the number of iterations is less than the first number threshold, determine the target loss according to the first loss, the second loss, the feature reconstruction loss, and the third loss for the first source domain image; according to the target loss, adjust the parameters of the image classification model, the dynamic mask convolution network, the first convolution network, and the second classifier. If the number of iterations is not less than the first number threshold, adjust the parameters of the image classification model according to the weighted result of the first loss and the third loss for the first source domain image; and adjust the parameters of the dynamic mask convolution network, the first convolution network, and the second classifier according to the weighted result of the second loss and the feature reconstruction loss for the first source domain image.
[0199] In some other embodiments, if the features extracted by the feature extraction network for the first source domain image are used as the feature map, and no first convolutional network for dimensionality reduction processing is introduced between the feature extraction network and the dynamic mask convolutional network, in this case, an auxiliary second classifier may not be introduced. Correspondingly, step 917 may include: if the number of iterations is less than the first number threshold, determining the target loss for the first source domain image according to the first loss, the feature reconstruction loss, and the third loss for the first source domain image; adjusting the parameters of the image classification model, the dynamic mask convolutional network, and the first convolutional network according to the target loss of the first source domain image; if the number of iterations is not less than the first number threshold, adjusting the parameters of the image classification model according to the first loss and the third loss for the first source domain image; and adjusting the parameters of the dynamic mask convolutional network according to the weighted result of the first loss and the feature reconstruction loss for the first source domain image. The weighting coefficients of the first loss and the feature reconstruction loss may be the same (for example, both are 1) or different, and no specific limitation is made here.
[0200] Figure 10 is a schematic diagram of training an image classification model according to an embodiment of the present application. In Figure 10 it, the double-dashed arrow indicates the processing flow for the target domain image, the dashed arrow indicates the processing flow for the first source domain image, and the solid arrow indicates the processing flow shared by the target domain image and the first source domain image. As Figure 10 shown, for the target domain image, first, feature extraction is performed through the feature extraction network to obtain the first feature map of the target domain image; thereafter, the first feature map is input into the first classifier for classification (i.e., target domain prediction), and pixel screening can be performed based on the first classification result to determine the second pixel set (the specific implementation process can be referred to the corresponding embodiment above Figure 7 shown). In addition, the first feature map of the target domain image can also be input into the bottleneck network (i.e., the first convolutional network) for dimensionality reduction processing to obtain the feature map of the target domain image; thereafter, on the one hand, the feature map is reconstructed through the dynamic mask convolutional network to output the reconstructed feature map, and then the coverage density of each pixel is determined according to the reconstructed feature map and the feature map, and based on the coverage density of each pixel and the feature map of the target domain image, a density-aware greedy algorithm is used to determine the core pixel set. After the core pixel set is labeled, fine-tuning training is performed on the image classification model.
[0201] After labeling the pixels in the core pixel set, the annotation information of each pixel in the core pixel set and the target domain image can be used to fine-tune the image classification model. The specific process is as follows: First, the target domain image is feature-extracted through a feature extraction network to obtain the first feature map of the target domain image; then, the first feature map is input into the first classifier for classification (i.e., target domain prediction), and the first loss can be calculated based on the first classification result and the annotation information of each pixel in the core pixel set; in addition, the first feature map of the target domain image is input into the bottleneck network (i.e., the first convolutional network) for dimensionality reduction processing to obtain the feature map of the target domain image; then, on the one hand, the feature map is reconstructed through the dynamic mask convolutional network, and the reconstructed feature map is output, and the feature reconstruction loss (i.e., Figure 10 the LMSE in
[0202] is calculated based on the feature map and the reconstructed feature map; on the other hand, the feature map of the target domain image is input into the second classifier, and the second loss for the target domain image is calculated based on the second classification result output by the second classifier and the annotation information of each pixel in the core pixel set; then, according to the first loss, the feature reconstruction loss, and the second loss, the parameters of the image classification model (including the feature reconstruction network and the first classifier), the bottleneck network, the dynamic mask convolutional network, and the second classifier are adjusted. Figure 10 The process of fine-tuning the image classification model using the first source domain image and the annotation information of each pixel in the first source domain image is as follows: First, the first source domain image is feature-extracted through the feature extraction network to obtain the first feature map of the first source domain image; then, the first feature map is input into the first classifier for classification (i.e., source domain prediction), and the first loss for the first source domain image can be calculated based on the output third classification result and the annotation information of each pixel in the first source domain image; and the third loss for the first source domain image is calculated; in addition, the first feature map of the first source domain image is input into the bottleneck network (i.e., the first convolutional network) for dimensionality reduction processing to obtain the feature map of the first source domain image; then, on the one hand, the feature map of the first source domain image is reconstructed through the dynamic mask convolutional network, and the reconstructed feature map of the first source domain image is output, and the feature reconstruction loss (i.e.,
[0203] In some embodiments, considering that there may be differences in the resolutions of different source domain images in the source domain dataset and there may be differences in the resolutions of different target domain images in the target domain dataset, the resolutions of different source domain images in the source domain dataset can be unified to the same resolution (e.g., the first resolution) in advance. Similarly, the resolutions of different target domain images in the target domain dataset are unified to the same resolution (e.g., the second resolution), which is convenient for subsequent processing. Herein, the first resolution and the second resolution may be the same or different. For example, the first resolution may be 1280×720 and the second resolution may be 1280×640.
[0204] By comparing the method of training an image classification model according to the present application with an unsupervised domain adaptation method, experiments show that the method according to the present application can achieve a model performance improvement of more than 10 Mean Intersection over Union (MIoU, which refers to the ratio of the intersection and union of two sets of true values and predicted values, i.e., Mean Intersection over Union, MIoU); further, by comparing the method of training an image classification model according to the present application with other active domain methods, it is found that the present solution can achieve a model performance improvement of 1 to 2 MIoU. In addition, the performance of the image classification model trained by the present solution is close to the performance of the model trained with all annotations in the target domain, and it can achieve 98% of the performance of the full supervision scheme with 2.2% to 5% of the annotations. It can be seen that training an image classification model in the manner according to the present application can effectively ensure the training effect on the basis of a small number of annotations.
[0205] The present application also provides a semantic segmentation method, which can be executed by an electronic device. The electronic device can be a server, a terminal, etc., and is not specifically limited herein. The semantic segmentation method includes: obtaining an image to be processed; classifying the image to be processed by an image classification model based on the image to be processed, and outputting a classification result of the image to be processed. The image classification model is trained according to the training method of the image classification model as described above; the classification result indicates the category to which each pixel in the image to be processed belongs; and according to the category to which each pixel in the image to be processed belongs, determining the semantic segmentation result of the image to be processed.
[0206] Among them, the image to be processed can be an image from the target domain. After fine-tuning and training the image classification model according to the method of this application, the image classification model can accurately perform pixel-level classification on the images in the target domain, thereby realizing semantic segmentation. For example, the solution of this application can be applied to the industrial quality inspection scenario. The target domain can be a field where imaging conditions such as illumination, shooting position, and shooting angle change. After fine-tuning and training the image classification model according to the method of this application, the image classification model can accurately adapt to the images collected after changes in imaging conditions such as illumination, shooting position, and shooting angle, and accurately perform semantic segmentation on the images in the target domain. Thus, the problem that the performance of the pre-trained image classification model (such as the reduction of semantic segmentation accuracy) deteriorates due to changes in imaging conditions such as illumination, shooting position, and shooting angle in the industrial quality inspection scenario can be solved.
[0207] Since the classification result of the image to be processed indicates the categories to which the pixels in the image to be processed belong, therefore, based on the classification result of the image to be processed, the pixels belonging to the same category in the image to be processed can be determined, that is, the semantic segmentation result of the image to be processed can be determined. Further, the pixels belonging to the same category in the image to be processed can be shown in the form of a mask, and different masks are used for different categories of pixels. In this way, it is convenient for users to distinguish different categories of pixels.
[0208] The following introduces the device embodiments of this application, which can be used to execute the methods in the above embodiments of this application. For the details not disclosed in the device embodiments of this application, please refer to the above method embodiments of this application.
[0209] Figure 11 is a block diagram of a training device for an image classification model shown according to an embodiment of this application. The training device for the image classification model can be configured in an electronic device and is used to implement the training method for the image classification model provided by this application. As Figure 11As shown in the figure, the training device of the image classification model includes: a feature map acquisition module 1110, configured to determine a feature map of a target domain image based on features extracted from the target domain image by a feature extraction network in the image classification model; the image classification model is pre-trained by source domain images and annotation information of each pixel in the source domain images; a feature reconstruction module 1120, configured to perform feature reconstruction based on the feature map to obtain a reconstructed feature map, where the reconstructed features of pixels in the reconstructed feature map are reconstructed according to the features of neighboring pixels of the pixel; a coverage density determination module 1130, configured to determine the coverage density of each pixel in the target domain image according to the feature map and the reconstructed feature map; a screening module 1140, configured to screen a core pixel set to be pixel-annotated in the target domain image according to the coverage density of each pixel in the target domain image and the features of each pixel in the feature map; the coverage density is used to scale the feature distance between two pixels; a first fine-tuning training module 1150, configured to perform fine-tuning training on the image classification model according to the annotation information of pixels in the core pixel set and the target domain image.
[0210] In some embodiments, the screening module 1140 includes: a first determination unit, configured to determine a first pixel set, where the first pixels in the first pixel set are pixels in the target domain image; a second determination unit, configured to, for each second pixel in a second pixel set, determine a first target pixel in the first pixel set with the smallest feature distance from the second pixel; the feature distance is determined according to the features of two pixels; the second pixels in the second pixel set are pixels in the target domain image, and the intersection of the second pixel set and the first pixel set is empty; a reference distance determination unit, configured to scale the feature distance between the second pixel and the corresponding first target pixel according to the coverage density of the first target pixel corresponding to the second pixel to obtain a reference distance between the second pixel and the first pixel set, and the reference distance is negatively correlated with the coverage density; a second target pixel determination unit, configured to determine a second target pixel in the second pixel set with the largest reference distance from the first pixel set; an addition unit, configured to add the second target pixel to the first pixel set; a third determination unit, configured to determine a core pixel set based on the updated first pixel set.
[0211] In some embodiments, the third determination unit is further configured to: count the number of pixels in the updated first pixel set; if it is determined that the annotation requirement is met according to the number of pixels, use the updated first pixel set as the core pixel set; the training device of the image classification model further includes: a removal module, configured to: if it is determined that the annotation requirement is not met according to the number of pixels, remove the second target pixel from the second pixel set to obtain a new second pixel set, and return to execute determining, for each second pixel in the second pixel set, a first target pixel in the first pixel set with the smallest feature distance from the second pixel.
[0212] In some embodiments, the image classification model further includes a first classifier; the training device of the image classification model further includes: a first classification module, configured to classify, by the first classifier, based on the result of feature extraction of the target domain image by the feature extraction network, to obtain a first classification result, where the first classification result indicates the probability that each pixel in the target domain image belongs to each of multiple categories; a target probability determination module, configured to determine, for each pixel, the largest N target probabilities among the probabilities that the pixel belongs to multiple categories, where N is a positive integer; a labeling gain determination module, configured to calculate the labeling gain of each pixel according to the N target probabilities determined for each pixel; a second pixel set determination module, configured to screen out the pixels in the target domain image whose labeling gain is greater than the labeling gain threshold, and construct a second pixel set according to the pixels whose labeling gain is greater than the labeling gain threshold.
[0213] In some embodiments, N is 2, and the N target probabilities determined for a pixel include the maximum probability and the second largest probability among the probabilities that the pixel belongs to multiple categories; the labeling gain determination module is configured to: for each pixel, calculate the sum of the maximum probability and the second largest probability corresponding to the pixel to obtain a reference result for the pixel; for each pixel, subtract the reference result corresponding to the pixel from 1 to obtain the labeling gain of the pixel.
[0214] In some embodiments, the feature reconstruction module 1120 includes: a dynamic masked convolution processing unit, configured to perform dynamic masked convolution processing on the feature map to obtain a spatial modulation map, where the feature value of the central pixel of the convolution window in each convolution in the spatial modulation map is 0; a fusion unit, configured to fuse the spatial modulation map and the feature map to obtain a reconstructed feature map.
[0215] In some embodiments, the dynamic masked convolution processing unit is configured to: perform dynamic convolution on the feature map to obtain an intermediate feature map; superimpose the intermediate feature map and the mask tensor corresponding to the convolution window to obtain a reference feature map; normalize the features in the reference feature map to obtain a spatial modulation map.
[0216] In some embodiments, the coverage density determination module 1130 includes: an acquisition unit, configured to, for each pixel, acquire the feature of the pixel from the feature map and the reconstructed feature of the pixel from the reconstructed feature map; a feature difference determination unit, configured to calculate the feature difference between the feature of the pixel and the reconstructed feature of the pixel; a coverage density determination unit, configured to determine the coverage density of the pixel according to the feature difference of the pixel.
[0217] In some embodiments, the image classification model further includes a first classifier; the reconstructed feature map is obtained by the dynamic mask convolutional network based on the feature map for reconstruction; correspondingly, the first fine-tuning training module 1150 includes: a first acquisition module, configured to acquire a first classification result obtained by the first classifier for classifying the features extracted from the target domain image by the feature extraction network; a first loss determination module, configured to determine a first loss according to the first classification result and the annotation information of the pixels in the core pixel set; a feature reconstruction loss determination module, configured to determine a feature reconstruction loss according to the feature map and the reconstructed feature map; and an adjustment module, configured to adjust the parameters of at least one of the image classification model and the dynamic mask convolutional network according to the first loss and the feature reconstruction loss.
[0218] In some embodiments, the feature map is obtained by the first convolutional network reducing the dimension of the first feature map, and the first feature map is the feature extracted by the feature extraction network for the target domain image; the training device of the image classification model further includes: a second classification module, configured to classify by the second classifier based on the feature map to obtain a second classification result; a second loss determination module, configured to determine a second loss according to the second classification result and the annotation information of the pixels in the core pixel set; correspondingly, the adjustment module is configured to: if the number of iterations is less than the first number threshold, determine a target loss according to the first loss, the second loss, and the feature reconstruction loss; and adjust the parameters of the image classification model, the dynamic mask convolutional network, the first convolutional network, and the second classifier according to the target loss.
[0219] In some embodiments, the adjustment module further includes: if the number of iterations is not less than the first number threshold, adjust the parameters of the image classification model according to the first loss; and adjust the parameters of the dynamic mask convolutional network, the first convolutional network, and the second classifier according to the weighted result of the second loss and the feature reconstruction loss.
[0220] In some embodiments, the training device of the image classification model further includes: a second fine-tuning training module, configured to perform fine-tuning training on the image classification model according to the first source domain image and the label information of the first source domain image.
[0221] In some embodiments, the second fine-tuning training module is configured to: extract features from the first source domain image by the feature extraction network to obtain a first feature map of the first source domain image; classify the first feature map of the first source domain image by the first classifier to obtain a third classification result; determine a first loss for the first source domain image according to the annotation information and the third classification result of the first source domain image; determine a feature map of the first source domain image based on the first feature map of the first source domain image; reconstruct the features of the feature map of the first source domain image through the dynamic mask convolution network to obtain a reconstructed feature map of the first source domain image; determine a feature reconstruction loss for the first source domain image according to the feature map and the reconstructed feature map of the first source domain image; determine a third loss for the first source domain image according to the probabilities that each pixel in the third classification result belongs to the actual class and other classes; and adjust the parameters of at least one of the image classification model and the dynamic mask convolution network according to the first loss, the feature reconstruction loss, and the third loss for the first source domain image.
[0222] Figure 12 is a block diagram of a semantic segmentation device shown according to an embodiment of the present application. The semantic segmentation device can be configured in an electronic device to implement the semantic segmentation method provided by the present application. As Figure 12 shown, the semantic segmentation device includes: a to-be-processed image acquisition module 1210, configured to acquire a to-be-processed image; a third classification module 1220, configured to classify the to-be-processed image by the image classification model and output a classification result of the to-be-processed image, where the image classification model is trained according to the training method of the image classification model as described above; the classification result indicates the class to which each pixel in the to-be-processed image belongs; and a semantic segmentation result determination module 1230, configured to determine a semantic segmentation result of the to-be-processed image according to the class to which each pixel in the to-be-processed image belongs.
[0223] Figure 13 shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application. It should be noted that Figure 13 the computer system 1300 of the electronic device shown is only an example, and should not bring any limitation to the functions and usage scopes of the embodiments of the present application. The electronic device can be used to execute the training method of the image classification model as described above or execute the semantic segmentation method as described above.
[0224] As Figure 13As shown, computer system 1300 includes a Central Processing Unit (CPU) 1301, which can perform various appropriate actions and processes according to the program stored in the Read-Only Memory (ROM) 1302 or the program loaded from the storage section 1308 into the Random Access Memory (RAM) 1303, such as executing the methods in the above embodiments. In the RAM 1303, various programs and data required for system operation are also stored. The CPU 1301, ROM 1302, and RAM 1303 are connected to each other via a bus 1304. An Input / Output (I / O) interface 1305 is also connected to the bus 1304.
[0225] The following components are connected to the I / O interface 1305: an input section 1306 including a keyboard, a mouse, etc.; an output section 1307 including, for example, a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), etc., and a speaker, etc.; a storage section 1308 including a hard disk, etc.; and a communication section 1309 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1309 performs communication processing via a network such as the Internet. A drive 1310 is also connected to the I / O interface 1305 as needed. A removable medium 1311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1310 as needed so that the computer program read from it can be installed into the storage section 1308 as needed.
[0226] Specifically, according to the embodiments of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments of the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 1309, and / or installed from the removable medium 1311. When the computer program is executed by the Central Processing Unit (CPU) 1301, various functions defined in the system of the present application are executed.
[0227] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0228] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0229] The units involved in the embodiments described in this application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not, in some cases, constitute a limitation on the units themselves.
[0230] As another aspect, this application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or may exist separately without being assembled into the electronic device. The above computer-readable storage medium carries computer-readable instructions, and when the computer-readable storage instructions are executed by a processor, the methods in any of the above embodiments are implemented.
[0231] According to one aspect of the embodiments of this application, there is provided a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the methods in any of the above embodiments.
[0232] In the embodiments of this application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be an overall module or a part of the unit of its function.
[0233] It should be noted that although several modules or units of the devices for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0234] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solution according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.
[0235] After considering the specification and practicing the disclosed embodiments herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the present application.
[0236] It should be understood that the present application is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A training method for an image classification model, characterized in that, Including: Determine a feature map of the target domain image based on the features extracted from the target domain image by a feature extraction network in an image classification model; The image classification model is pre-trained by source domain images and annotation information of each pixel in the source domain images; Perform feature reconstruction based on the feature map to obtain a reconstructed feature map, where the reconstructed features of the pixels in the reconstructed feature map are reconstructed according to the features of the neighboring pixels of the pixel; Determine the coverage density of each pixel in the target domain image according to the feature map and the reconstructed feature map; the higher the coverage density of a pixel, the denser the distribution of the features of the pixel in the target domain image; According to the coverage density of each pixel in the target domain image and the features of each pixel in the feature map, select a core pixel set to be pixel-labeled in the target domain image according to the greedy algorithm; the coverage density is used to scale the feature distance between two pixels; Fine-tune and train the image classification model according to the annotation information of the pixels in the core pixel set and the target domain image.
2. The method according to claim 1, characterized in that, The step of selecting a core pixel set to be pixel-labeled in the target domain image according to the coverage density of each pixel in the target domain image and the features of each pixel in the feature map according to the greedy algorithm includes: Determine a first pixel set, where the first pixels in the first pixel set are pixels in the target domain image; For each second pixel in a second pixel set, determine a first target pixel with the smallest feature distance from the second pixel in the first pixel set; the feature distance is determined according to the features of the two pixels; the second pixels in the second pixel set are pixels in the target domain image, and the intersection of the second pixel set and the first pixel set is empty; Scale the feature distance between the second pixel and the corresponding first target pixel according to the coverage density of the first target pixel corresponding to the second pixel to obtain a reference distance between the second pixel and the first pixel set, and the reference distance is negatively correlated with the coverage density; Determine a second target pixel with the largest reference distance from the first pixel set in the second pixel set; Add the second target pixel to the first pixel set; Based on the updated first pixel set, determine the core pixel set.
3. The method according to claim 2, wherein The step of determining the core pixel set based on the updated first pixel set includes: Count the number of pixels in the updated first pixel set; If it is determined that the annotation requirement is met according to the number of pixels, use the updated first pixel set as the core pixel set; If it is determined that the annotation requirement is not met according to the number of pixels, remove the second target pixel from the second pixel set to obtain a new second pixel set, and return to execute the step of determining a first target pixel with the smallest feature distance from the second pixel in the first pixel set for each second pixel in the second pixel set.
4. The method according to claim 2, wherein The image classification model further includes a first classifier; Before determining a first target pixel with the smallest feature distance from each second pixel in the second pixel set in the first pixel set, the method further includes: The first classifier classifies based on the result of feature extraction of the target domain image by the feature extraction network to obtain a first classification result, and the first classification result indicates the probability that each pixel in the target domain image belongs to each of multiple categories; For each pixel, determine the largest N target probabilities for the pixel among the probabilities that the pixel belongs to multiple categories, where N is a positive integer; Calculate the annotation gain of each pixel according to the N target probabilities determined for each pixel; Screen pixels in the target domain image whose annotation gain is greater than the annotation gain threshold, and construct the second pixel set according to the pixels whose annotation gain is greater than the annotation gain threshold.
5. The method according to claim 4, wherein N is 2, and the N target probabilities determined for a pixel include the maximum probability and the second largest probability among the probabilities that the pixel belongs to multiple categories; The calculating the annotation gain of each pixel according to the N target probabilities determined for each pixel includes: For each pixel, calculate the sum of the maximum probability and the second largest probability corresponding to the pixel to obtain the reference result of the pixel; For each pixel, subtract the reference result corresponding to the pixel from 1 to obtain the annotation gain of the pixel.
6. The method according to claim 1, characterized in that, The performing feature reconstruction based on the feature map to obtain a reconstructed feature map includes: Perform dynamic masked convolution processing on the feature map to obtain a spatial modulation map, and the feature value of the central pixel of the convolution window for each convolution in the spatial modulation map is 0; Fuse the spatial modulation map and the feature map to obtain the reconstructed feature map.
7. The method according to claim 6, characterized in that, The performing dynamic masked convolution processing on the feature map to obtain a spatial modulation map includes: Perform dynamic convolution on the feature map to obtain an intermediate feature map; Overlay the intermediate feature map with the mask tensor corresponding to the convolution window to obtain a reference feature map; Normalize the features in the reference feature map to obtain the spatial modulation map.
8. The method according to claim 1, wherein The determining the coverage density of each pixel in the target domain image according to the feature map and the reconstructed feature map includes: For each pixel, obtain the feature of the pixel from the feature map and the reconstructed feature of the pixel from the reconstructed feature map; Calculate the feature difference between the feature of the pixel and the reconstructed feature of the pixel; Determine the coverage density of the pixel according to the feature difference of the pixel.
9. The method according to claim 1, wherein The image classification model further includes a first classifier; the reconstructed feature map is reconstructed by a dynamic masked convolution network based on the feature map; The fine-tuning and training of the image classification model according to the annotation information of each pixel in the core pixel set and the target domain image includes: Obtain the first classification result obtained by the first classifier classifying the features extracted from the target domain image by the feature extraction network; Determine a first loss according to the first classification result and the annotation information of the pixels in the core pixel set; Determine a feature reconstruction loss according to the feature map and the reconstructed feature map; Adjust the parameters of at least one of the image classification model and the dynamic masked convolution network according to the first loss and the feature reconstruction loss.
10. The method according to claim 9, characterized in that The feature map is obtained by reducing the dimension of the first feature map through a first convolutional network, and the first feature map is the feature extracted by the feature extraction network for the target domain image; The method further includes: classifying by a second classifier based on the feature map to obtain a second classification result; determining a second loss according to the second classification result and the annotation information of the pixels in the core pixel set; The adjusting the parameters of at least one of the image classification model and the dynamic mask convolutional network according to the first loss and the feature reconstruction loss includes: if the number of iterations is less than the first threshold, determining a target loss according to the first loss, the second loss and the feature reconstruction loss; adjusting the parameters of the image classification model, the dynamic mask convolutional network, the first convolutional network and the second classifier according to the target loss.
11. The method according to claim 10, wherein The adjusting the parameters of at least one of the image classification model and the dynamic mask convolutional network according to the first loss and the feature reconstruction loss further includes: if the number of iterations is not less than the first threshold, adjusting the parameters of the image classification model according to the first loss; and adjusting the parameters of the dynamic mask convolutional network, the first convolutional network and the second classifier according to the weighted result of the second loss and the feature reconstruction loss.
12. The method according to any one of claims 9 to 11, characterized in that The method further includes: fine-tuning and training the image classification model according to the first source domain image and the label information of the first source domain image.
13. The method according to claim 12, wherein The fine-tuning and training the image classification model according to the first source domain image and the label information of the first source domain image includes: extracting features of the first source domain image by the feature extraction network to obtain a first feature map of the first source domain image; classifying by a first classifier based on the first feature map of the first source domain image to obtain a third classification result; determining a first loss for the first source domain image according to the annotation information and the third classification result of the first source domain image; determining the feature map of the first source domain image based on the first feature map of the first source domain image; performing feature reconstruction on the feature map of the first source domain image through the dynamic mask convolutional network to obtain a reconstructed feature map of the first source domain image; determining a feature reconstruction loss for the first source domain image according to the feature map and the reconstructed feature map of the first source domain image; determining a third loss for the first source domain image according to the probabilities that each pixel in the third classification result belongs to the actual class and belongs to other classes; adjusting the parameters of at least one of the image classification model and the dynamic mask convolutional network according to the first loss, the feature reconstruction loss and the third loss for the first source domain image.
14. A semantic segmentation method, characterized in that, including: obtaining an image to be processed; classifying by an image classification model based on the image to be processed, and outputting a classification result of the image to be processed, where the image classification model is trained according to the method described in any one of claims 1 to 13; the classification result indicates the classes to which the pixels in the image to be processed belong; Determine the semantic segmentation result of the image to be processed according to the categories to which the pixels in the image to be processed belong.
15. A training device for an image classification model, characterized in that, Including: A feature map acquisition module, configured to determine the feature map of the target domain image based on the features extracted from the target domain image by the feature extraction network in the image classification model; The image classification model is pre-trained by using the source domain image and the annotation information of each pixel in the source domain image; A feature reconstruction module, configured to perform feature reconstruction based on the feature map to obtain a reconstructed feature map, where the reconstructed feature of a pixel in the reconstructed feature map is reconstructed according to the features of the neighboring pixels of the pixel; A coverage density determination module, configured to determine the coverage density of each pixel in the target domain image according to the feature map and the reconstructed feature map; the higher the coverage density of a pixel, the denser the distribution of the features of the pixel in the target domain image; A screening module, configured to screen a core pixel set to be pixel-annotated in the target domain image according to the coverage density of each pixel in the target domain image and the features of each pixel in the feature map according to the greedy algorithm; the coverage density is used to scale the feature distance between two pixels; A first fine-tuning training module, configured to perform fine-tuning training on the image classification model according to the annotation information of the pixels in the core pixel set and the target domain image.
16. The device according to claim 15, characterized in that, The screening module includes: A first determination unit, configured to determine a first pixel set, where the first pixels in the first pixel set are pixels in the target domain image; A second determination unit, configured to, for each second pixel in the second pixel set, determine a first target pixel with the smallest feature distance from the second pixel in the first pixel set; the feature distance is determined according to the features of the two pixels; the second pixels in the second pixel set are pixels in the target domain image, and the intersection of the second pixel set and the first pixel set is empty; A reference distance determination unit, configured to scale the feature distance between the second pixel and the corresponding first target pixel according to the coverage density of the first target pixel corresponding to the second pixel to obtain the reference distance between the second pixel and the first pixel set, and the reference distance has a negative correlation with the coverage density; A second target pixel determination unit, configured to determine a second target pixel with the largest reference distance from the first pixel set in the second pixel set; An adding unit, configured to add the second target pixel to the first pixel set; A third determination unit, configured to determine the core pixel set based on the updated first pixel set.
17. The device according to claim 16, characterized in that, The third determination unit is further configured to: Count the number of pixels in the updated first pixel set; if it is determined that the annotation requirement is met according to the number of pixels, use the updated first pixel set as the core pixel set; The training device of the image classification model further includes a removal module, configured to: if it is determined that the annotation requirement is not met according to the number of pixels, remove the second target pixel from the second pixel set to obtain a new second pixel set, and return to execute determining a first target pixel with the smallest feature distance from the second pixel in the first pixel set for each second pixel in the second pixel set.
18. The device according to claim 16, characterized in that, The image classification model further includes a first classifier; the training device of the image classification model further includes: A first classification module, configured to classify, by using the first classifier, according to the result of feature extraction performed on a target domain image by the feature extraction network, to obtain a first classification result, where the first classification result indicates the probability that each pixel in the target domain image belongs to each of multiple categories; A target probability determination module, configured to determine, for each pixel, the maximum N target probabilities among the probabilities that the pixel belongs to multiple categories, where N is a positive integer; A labeling gain determination module, configured to calculate the labeling gain of each pixel according to the N target probabilities determined for each pixel; A second pixel set determination module, configured to screen out pixels in the target domain image whose labeling gain is greater than a labeling gain threshold, and construct the second pixel set according to the pixels whose labeling gain is greater than the labeling gain threshold.
19. The device according to claim 18, characterized in that, N is 2, and the N target probabilities determined for a pixel include the maximum probability and the second maximum probability among the probabilities that the pixel belongs to multiple categories; The labeling gain determination module is configured to: for each pixel, calculate the sum of the maximum probability and the second maximum probability corresponding to the pixel, to obtain a reference result of the pixel; For each pixel, subtract the reference result corresponding to the pixel from 1 to obtain the labeling gain of the pixel.
20. The device according to claim 15, characterized in that, The feature reconstruction module includes: A dynamic mask convolution processing unit, configured to perform dynamic mask convolution processing on the feature map to obtain a spatial modulation map, where the feature value of the central pixel of the convolution window in each convolution in the spatial modulation map is 0; A fusion unit, configured to fuse the spatial modulation map and the feature map to obtain the reconstructed feature map.
21. The device according to claim 20, wherein The dynamic mask convolution processing unit is configured to: perform dynamic convolution on the feature map to obtain an intermediate feature map; superimpose the intermediate feature map and a mask tensor corresponding to the convolution window to obtain a reference feature map; Normalize the features in the reference feature map to obtain the spatial modulation map.
22. The device according to claim 15, characterized in that, The coverage density determination module includes: An acquisition unit, configured to, for each pixel, acquire the feature of the pixel from the feature map and the reconstructed feature of the pixel from the reconstructed feature map; A feature difference determination unit, configured to calculate the feature difference between the feature of the pixel and the reconstructed feature of the pixel; A coverage density determination unit, configured to determine the coverage density of the pixel according to the feature difference of the pixel.
23. The device according to claim 15, characterized in that, The image classification model further includes a first classifier; the reconstructed feature map is reconstructed by a dynamic mask convolution network based on the feature map; the first fine-tuning training module includes: A first acquisition module, configured to acquire the first classification result obtained by classifying, by using the first classifier, the features extracted from a target domain image by the feature extraction network; A first loss determination module, configured to determine a first loss according to the first classification result and the labeling information of the pixels in the core pixel set; A feature reconstruction loss determination module, configured to determine a feature reconstruction loss according to the feature map and the reconstructed feature map. An adjustment module, configured to adjust parameters of at least one of the image classification model and the dynamic mask convolutional network according to the first loss and the feature reconstruction loss.
24. The device according to claim 23, wherein, The feature map is obtained by reducing the dimension of a first feature map through a first convolutional network, and the first feature map is the feature extracted by the feature extraction network for a target domain image. The training device of the image classification model further includes: A second classification module, configured to perform classification based on the feature map by a second classifier to obtain a second classification result. A second loss determination module, configured to determine a second loss according to the second classification result and the annotation information of pixels in the core pixel set. Correspondingly, the adjustment module is configured to: if the number of iterations is less than a first number threshold, determine a target loss according to the first loss, the second loss, and the feature reconstruction loss; and adjust parameters of the image classification model, the dynamic mask convolutional network, the first convolutional network, and the second classifier according to the target loss.
25. The device according to claim 24, wherein The adjustment module is further configured to: if the number of iterations is not less than the first number threshold, adjust parameters of the image classification model according to the first loss; and adjust parameters of the dynamic mask convolutional network, the first convolutional network, and the second classifier according to a weighted result of the second loss and the feature reconstruction loss.
26. The device according to any one of claims 23 to 25, characterized in that, The training device of the image classification model further includes: A second fine-tuning training module, configured to perform fine-tuning training on the image classification model according to a first source domain image and label information of the first source domain image.
27. The device according to claim 26, wherein, The second fine-tuning training module is configured to: Extract features of the first source domain image by the feature extraction network to obtain a first feature map of the first source domain image. Perform classification on the first feature map of the first source domain image by a first classifier to obtain a third classification result. Determine a first loss for the first source domain image according to the annotation information of the first source domain image and the third classification result. Determine a feature map of the first source domain image based on the first feature map of the first source domain image. Perform feature reconstruction on the feature map of the first source domain image through the dynamic mask convolutional network to obtain a reconstructed feature map of the first source domain image. Determine a feature reconstruction loss for the first source domain image according to the feature map of the first source domain image and the reconstructed feature map of the first source domain image. Determine a third loss for the first source domain image according to probabilities that each pixel in the third classification result belongs to an actual class and belongs to other classes. Adjust parameters of at least one of the image classification model and the dynamic mask convolutional network according to the first loss, the feature reconstruction loss, and the third loss for the first source domain image.
28. A semantic segmentation device, characterized in that, Includes: A to-be-processed image acquisition module, configured to acquire a to-be-processed image. A third classification module, configured to perform classification on the to-be-processed image by an image classification model to output a classification result of the to-be-processed image, where the image classification model is trained according to the method described in any one of claims 1 to 13; and the classification result indicates classes to which each pixel in the to-be-processed image belongs. A semantic segmentation result determination module, configured to determine the semantic segmentation result of the to-be-processed image according to the categories to which the pixels in the to-be-processed image belong.
29. An electronic device, characterized in that, Including: A processor; A memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the method described in any one of claims 1-14 is implemented.
30. A computer-readable storage medium having computer-readable instructions stored thereon, characterized in that, When the computer-readable instructions are executed by the processor, the method described in any one of claims 1-14 is implemented.
31. A computer program product, comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, the method described in any one of claims 1-14 is implemented.
Citation Information
Patent Citations
Semantic segmentation method and device, computer equipment and computer readable storage medium
CN116630630A
Unsupervised domain adaptation method, device, system and storage medium of semantic segmentation based on uniform clustering
US20220383052A1