Image semi-supervised semantic segmentation method and system based on pixel-level correction
By adopting teacher-student network model and pixel-level correction technology in semi-supervised semantic segmentation of images, the problem of pseudo-labels not reliable enough and neglecting inter-pixel correlation in the existing methods is solved, and image semantic segmentation with higher accuracy and stronger generalization capabilities is achieved.
Patent Information
- Application Number
- CN202411801544.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-09
- Publication Date
- 2025-05-16
AI Technical Summary
The existing semantic segmentation method based on semi-supervised learning uses a fixed confidence threshold when selecting pseudo-labels, resulting in unreliable pseudo-labels and ignore the correlation between pixels, resulting in poor model performance.
The semi-supervised semantic segmentation method based on pixel-level correction is adopted to semantic segmentation of scene images using the teacher-student network model. By constructing agent-related losses, unsupervised consistency losses and supervised losses, the teacher model's knowledge is used to guide the training of the student model, and enhance the generalization ability and robustness of the model.
It significantly improves the accuracy of image semantic segmentation, reduces dependence on a large amount of labeled data, enhances the generalization ability and robustness of the model, and enables the student model to achieve excellent segmentation results when labeled data is scarce.
Smart Images

Figure CN120014256A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and semantic segmentation, and in particular to a method and system for semi-supervised semantic segmentation of images based on pixel-level correction. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] With the continuous deepening of computer vision research, the rapid development of the Internet and the large-scale application of computer technology, computer vision technology has not only liberated human labor to a certain extent, but also greatly promoted the improvement of productivity. In this context, the accuracy of computer vision technology has become particularly important. Semantic segmentation is one of the important tasks in the field of computer vision. Its goal is to assign a semantic label to each pixel in the image, thereby dividing the image into regions with specific meanings. Semantic segmentation technology has a wide range of applications in autonomous driving, medical image analysis, remote sensing image processing and other fields.
[0004] Traditional semantic segmentation methods usually rely on a large amount of labeled data to train deep learning models. This fully supervised learning method requires accurate pixel-by-pixel labeling of massive images, which is time-consuming and labor-intensive, and requires a lot of manpower and resources, resulting in high costs. In addition, human errors may occur during the labeling process, which will further affect the effectiveness of the model. In some fields such as medicine and remote sensing, the labeling of medical images and remote sensing images requires the knowledge and experience of professionals, which further increases the difficulty and cost of labeling.
[0005] Considering that unlabeled data exists in large quantities in reality and is easy to obtain, these unlabeled data often have a wide coverage and diversity, and can provide rich learning information for the model. Therefore, semi-supervised learning methods are introduced in existing semantic segmentation. This semi-supervised learning method can effectively utilize a large amount of unlabeled data and combine it with a small amount of labeled data to improve the performance of the model. The core idea is to build a system that can self-learn and optimize, so that the model can use unlabeled data to improve its own performance while learning labeled data.
[0006] However, existing semantic segmentation methods based on semi-supervised learning still have certain problems: on the one hand, existing commonly used unsupervised methods usually use fixed confidence thresholds to select pseudo labels, and these confidence scores are not reliable enough to be associated with the accuracy of pseudo labels for unlabeled images. Moreover, for class-imbalanced class distributions, high confidence scores always tend to favor classes with dominant distributions, resulting in poor model accuracy after unsupervised training. On the other hand, existing methods usually directly apply contrastive learning to semi-supervised semantic segmentation by establishing a set of positive / negative sampling points in the feature representation space to establish an independent set of pixel consistency regularization. However, this method ignores the correlation between pixels. The above problems will lead to poor performance of the model obtained by the final semi-supervised training and poor semantic segmentation accuracy. Summary of the invention
[0007] To address the deficiencies of the above-mentioned prior art, the present invention provides a method and system for semi-supervised semantic segmentation of images based on pixel-level correction, which utilizes a teacher-student network model to perform semantic segmentation on objects in scene images, utilizes an image enhancement method to generate different instances for the same image, and inputs the different instances into a teacher-student network for image segmentation. In addition, by constructing a proxy correlation loss that considers the correlation and consistency between feature pixels, an unsupervised consistency loss across images, and a supervised loss, the knowledge of the teacher model is used to guide the training and learning of the student model. The trained student model can achieve higher-precision image semantic segmentation.
[0008] In a first aspect, the present invention provides an image semi-supervised semantic segmentation method based on pixel-level correction.
[0009] A semi-supervised semantic segmentation method for an image based on pixel-level correction, comprising:
[0010] Obtaining a scene image to be segmented;
[0011] The scene image to be segmented is input into the trained student model, and the image semantic segmentation result is output; wherein the training process of the student model includes:
[0012] Collecting a number of scene images to construct a training data set; the training data set includes labeled images and unlabeled images, wherein the labeled images are images annotated with pixel-level labels;
[0013] By performing random diversity image enhancement on the dataset images, weakly enhanced labeled images, weakly enhanced unlabeled images, and strongly enhanced unlabeled images are obtained;
[0014] Construct a teacher-student model with the same structural composition but different parameter initialization;
[0015] The teacher-student model is trained using the training dataset, and a total loss function based on the supervised loss function, the cross-image unsupervised loss function, and the agent-related loss function is constructed. The network parameters of the teacher-student model are optimized and updated through the total loss function until the total loss function is minimized, thus obtaining the trained teacher-student model.
[0016] In a second aspect, the present invention provides an image semi-supervised semantic segmentation system based on pixel-level correction.
[0017] A semi-supervised semantic segmentation system for images based on pixel-level correction, comprising:
[0018] An image acquisition module, used to acquire a scene image to be segmented;
[0019] The image semantic segmentation module is used to input the scene image to be segmented into the trained student model and output the image semantic segmentation result; wherein the training process of the student model includes:
[0020] Collecting a number of scene images to construct a training data set; the training data set includes labeled images and unlabeled images, wherein the labeled images are images annotated with pixel-level labels;
[0021] By performing random diversity image enhancement on the dataset images, weakly enhanced labeled images, weakly enhanced unlabeled images, and strongly enhanced unlabeled images are obtained;
[0022] Construct a teacher-student model with the same structural composition but different parameter initialization;
[0023] The teacher-student model is trained using the training dataset, and a total loss function based on the supervised loss function, the cross-image unsupervised loss function, and the agent-related loss function is constructed. The network parameters of the teacher-student model are optimized and updated through the total loss function until the total loss function is minimized, thus obtaining the trained teacher-student model.
[0024] In a third aspect, the present invention further provides an electronic device comprising a memory and a processor, and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the steps of the method described in the first aspect are completed.
[0025] In a fourth aspect, the present invention further provides a computer-readable storage medium for storing computer instructions, which, when executed by a processor, complete the steps of the method described in the first aspect.
[0026] One or more of the above technical solutions have the following beneficial effects:
[0027] 1. The present invention provides a method and system for semi-supervised semantic segmentation of images based on pixel-level correction, which uses a teacher-student network model to perform semantic segmentation on objects in scene images, uses an image enhancement method to generate different instances for the same image, inputs the different instances into the teacher-student network, and guides the student model to effectively learn on a large amount of unlabeled data through pseudo labels generated by the teacher model, thereby significantly improving the segmentation accuracy and reducing the dependence on a large amount of labeled data; on this basis, a proxy correlation loss, cross-image unsupervised consistency loss and supervised loss that consider the correlation and consistency between feature pixels are constructed, so that the teacher model and the student model can learn from each other and gradually improve, and the teacher model knowledge is used to guide the training and learning of the student model, thereby enhancing the generalization ability and robustness of the model, so that the student model can better adapt to complex and changeable scenes, and can still achieve excellent segmentation effects when labeled data is scarce, thereby achieving higher-precision image semantic segmentation.
[0028] 2. In the semantic segmentation method proposed in the present invention, taking into account the problem that the selected fixed confidence threshold of the existing unsupervised methods is unreliable and affects the segmentation accuracy of the model, the present invention proposes a method of using reliable labeled images to correct pseudo-labels. By performing pixel-level similarity calculations on unlabeled images and labeled images, the labeled images can guide the pseudo-labels and perform reliable pixel-level correction. Unsupervised training is performed based on the cross-image unsupervised consistency loss built in this way to ensure the effectiveness and accuracy of the semantic segmentation of the final model.
[0029] 3. Taking into account the problem that the existing methods ignore the correlation between pixels and affect the model segmentation accuracy, the present invention introduces a dense prediction task. The dense prediction task carries sufficient inter-pixel information that exceeds the consistency of a single basic pixel. This information can reveal the correlation between pixels and consistency regularization. By setting a set of representative reference points as agents, the agent-level correlation is calculated and obtained, and the agent-level correlation is converted into an agent ranking probability distribution. A proxy-related loss function based on the consistency of the agent ranking probability distribution between the teacher and student models is constructed, so that the model can be guided by more effective supervisory signals, thereby improving the model performance and segmentation accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0031] Figure 1 This is an overall flow chart of the image semi-supervised semantic segmentation method based on pixel-level correction according to an embodiment of the present invention;
[0032] Figure 2Schematic diagram of a weak image enhancement processing method in an embodiment of the present invention;
[0033] Figure 3 Schematic diagram of a model structure based on an encoder-decoder in an embodiment of the present invention;
[0034] Figure 4 Schematic diagram of image semantic segmentation results in an embodiment of the present invention. DETAILED DESCRIPTION
[0035] It should be noted that the following detailed descriptions are exemplary only, are intended to describe specific embodiments, are intended to provide further explanation of the present invention, and are not intended to limit exemplary embodiments according to the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those of ordinary skill in the art to which the present invention belongs. In addition, it should also be understood that when the terms "comprising" and / or "including" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.
[0036] Embodiment 1
[0037] This embodiment provides an image semi-supervised semantic segmentation method based on pixel-level correction, which makes full use of unlabeled image data, combines with a small amount of labeled image data, and adopts a semi-supervised learning training method, so that the model can not only effectively learn rich feature information, but also maintain or even improve the segmentation accuracy when the labeled data is insufficient, effectively improve the generalization ability and overall performance of the model, and reduce the dependence on a large amount of labeled data. The model obtained through training can better adapt to complex and changeable scenes and achieve more accurate semantic segmentation, thereby showing stronger robustness and practicality in practical applications.
[0038] The semi-supervised semantic segmentation method of the image based on pixel-level correction proposed in this embodiment is as follows: Figure 1 As shown, the specific steps include:
[0039] Step S1: construct a data set.
[0040] Several scene images are collected through image sensors such as cameras. For supervised images, a small number of images are collected and manually annotated according to the semantic information of the pixel content to obtain pixel-level labels. The images annotated with pixel-level labels are called labeled images. For unsupervised images, a large number of images are collected and initialized by denoising, cropping, etc. to facilitate subsequent use as training samples. The images without label annotations are called unlabeled images. In the above way, a training data set including labeled images and unlabeled images is constructed.
[0041] Step S2: random diversity image enhancement.
[0042] Before inputting the images in the training data set into the model, data augmentation (i.e., image augmentation) is required. Through image augmentation, the segmentation results of the same image with different degrees of enhancement remain consistent, so as to improve the robustness and generalization ability of the model. Among them, weak enhancement operations are performed on labeled images and images input to the teacher model, and strong enhancement operations are performed on images input to the student model. By performing random diversity image augmentation on the images in the data set, weakly enhanced labeled images, weakly enhanced unlabeled images, and strongly enhanced unlabeled images are obtained.
[0043] Furthermore, different from the traditional enhancement method which is limited to the image level enhancement, the present embodiment further expands the enhancement space and constructs a random diverse enhancement pool including a variety of different strong enhancement methods and weak enhancement methods, wherein the enhancement pool includes image level, feature level and distortion level-based enhancement methods.
[0044] Specifically, as shown in Table 1 below, weakly enhanced images are images enhanced using weak enhancement methods at the image space level, including random scaling, random inversion, and random cropping; and strongly enhanced images are images enhanced using at most k strong enhancement methods randomly selected from the enhancement pool (k is the preset maximum enhancement limit), including CutMix, automatic contrast, equalization, Gaussian filtering, contrast, sharpening, color balance, brightness, hue jitter, hue separation, exposure, color inversion, etc. The above image enhancement methods help simulate different scene changes, so that the model can better adapt to the diverse inputs in practical applications.
[0045] Table 1 Different image enhancement methods in the random diversity enhancement pool
[0046]
[0047] Step S3: construct a teacher-student model. The teacher model and the student model have the same composition / structure, and both use an encoder-decoder network architecture, but the parameters are initialized differently. The encoder extracts high-level semantic features from the input image, and then the decoder restores it to the same scale as the input image, and outputs the pixel-by-pixel classification prediction result.
[0048] Furthermore, the weakly enhanced unlabeled image is input into the teacher model. The encoder in the teacher model is composed of a ResNet network and a regularization module, which is used to extract high-level semantic features (also called high-dimensional features) of the input image. The result of the encoder is then input into the decoder. The decoder is composed of a DeepLabV3 network, which gradually restores the high-level semantic features extracted by the encoder to the spatial resolution of the original input image through upsampling and convolution operations, restores spatial details, and obtains the pixel-by-pixel classification result as the pseudo label y t .
[0049] The weakly enhanced labeled image and the strongly enhanced unlabeled image are input into the student model, where the weakly enhanced labeled image is input into the student model to obtain the pixel-by-pixel classification prediction result, and the cross entropy loss function is constructed with the result and the actual label of the image, which is used as the supervised loss function of the student model; the strongly enhanced unlabeled image is input into the student model. In fact, the unlabeled image is the same image as the unlabeled image input into the teacher model. The difference is that after strong enhancement operations such as random image level, feature level feature loss and strong distortion, an instance different from the input teacher model can be generated. The strongly enhanced unlabeled image is input into the student model, and the predicted classification result p is obtained through the encoder-decoder. s And output.
[0050] The encoder-decoder structure is as follows Figure 3 As shown in the figure, the encoder consists of a ResNet50 network, including sequentially connected convolutional layers, batch normalization layers, ReLU activation layers, four-stage residual blocks, global average pooling layers and fully connected layers; the decoder adopts the decoder part of the DeepLab V3 network, including multi-layer upsampling layers.
[0051] (1) The encoder consists of a ResNet50 network. In the ResNet50 network, it first passes through a convolution layer consisting of a 7×7 convolution kernel, and then performs batch normalization and ReLU activation; then it passes through four stages of residual blocks, among which the first stage contains 3 residual blocks, each of which is composed of 1×1, 3×3, and 1×1 convolution kernels, and the number of channels is 64, 64, and 256 respectively; the second stage contains 4 residual blocks, the residual block convolution kernel is the same as the first stage, and the number of output channels is 128, 128, and 512 respectively; the third stage contains 4 residual blocks, the residual block convolution kernel is the same as the first stage, and the number of output channels is 256, 256, and 1024 respectively; the fourth stage contains 4 residual blocks, the residual block convolution kernel is the same as the first stage, and the number of output channels is 512, 512, and 1028 respectively; in addition, each residual block passes the input directly to the output through a skip connection, thereby forming a residual link.
[0052] (2) The decoder uses the decoder part of the DeepLab V3 network. In the DeepLab V3 network, after multiple upsampling layers, the feature map is restored to the same size as the input image, thereby generating pixel-level predicted segmentation results.
[0053] Step S4: Use the training data set to train the teacher-student model, construct a total loss function based on the supervised loss function, the cross-image unsupervised loss function, and the agent-related loss function, and optimize and update the network parameters of the teacher-student model through the total loss function until the total loss function is minimized to obtain a trained teacher-student model.
[0054] Step S4.1, constructing a cross-image unsupervised loss function. Based on the pixel-level pseudo labels and classification prediction results of the unlabeled images output by the teacher model and the student model, a cross-image similarity method is used to perform pixel-level correction to construct a cross-image unsupervised loss function.
[0055] Considering that the existing commonly used unsupervised methods usually use a fixed confidence threshold to select pseudo-labels, ignoring the labeled images with accurate annotations, and these confidence scores are not reliable enough to be associated with the accuracy of the pseudo-labels of the unlabeled images, and for class imbalanced class distribution, high confidence always tends to be biased towards the class with the main distribution. To this end, this embodiment proposes a method for correcting pseudo-labels using reliable labeled images. First, under the guidance of the pseudo-label of the unlabeled image, a labeled image of the same class as the unlabeled image is queried; then, the pixel-level similarity of the unlabeled image and the labeled image is calculated, and the similarity calculation result can guide the reliable pixel-level correction of the pseudo-label; finally, a method based on this cross-image similarity is used to segment reliable and unreliable unlabeled images for training.
[0056] Specifically, the pixel-level pseudo labels and classification prediction results of the unlabeled images output by the teacher model and the student model are introduced by taking the corresponding pseudo labels of the unlabeled images output by the teacher model as an example.
[0057] First, randomly select a class (k classes) from the labels to form a pseudo mask Represents the positions of all pixels predicted as class k in the pseudo-label; then, a labeled image containing the k class is selected from all labeled images, and the actual class mask of the k class is obtained based on the actual pixel-level label of the labeled image In order to achieve high-dimensional consistency, high-level semantic feature maps, namely high-dimensional features F, are extracted for labeled images and unlabeled images respectively. l ,F u , F l With class mask Perform Hadamard product to extract supporting features of labeled images The supporting feature can provide the visual features of the class to be identified by the supporting image (i.e., the labeled image), i.e., the features of the kth class. By using the supporting feature, the unsupervised image can be better generalized and similar objects in the query image (i.e., the unlabeled image) can be segmented, thereby further improving the image segmentation effect.
[0058] After obtaining the supporting features, the pixel-level cosine similarity between the supporting features and the high-dimensional features of the unlabeled image is calculated as:
[0059]
[0060] In the above formula, m ′ u represents the region in the unlabeled image that is most likely to belong to class k, which can be called the CIC (Cross-ImageClass) graph. u Represents the high-dimensional features of unlabeled images, represents the supporting features of the labeled image, The transposed matrix representing the high-dimensional features. The above features are all represented in the form of feature graphs.
[0061] By calculating the similarity between unlabeled images and labeled images, more accurate and reliable information can be provided, thereby making more effective pixel-level corrections to pseudo labels. Therefore, the unsupervised loss can be optimized as:
[0062]
[0063] In the above formula, m u represents a pseudo mask, m ′ u represents the CIC graph, |B u |Total number of unlabeled images, represents the pseudo label of the jth pixel of the i-th image by the teacher model, Represents the prediction result of the student model for the jth pixel of the i-th image.
[0064] Step S4.2, constructing an agent-related loss function. Based on the high-level semantic features of the unlabeled images extracted by the teacher model and the student model, the agent-level correlation is calculated, the agent-level correlation is converted into an agent ranking probability distribution, and an agent-related loss function based on the consistency of the agent ranking probability distribution between the teacher and student models is constructed.
[0065] The dense prediction task analyzes rich inter-pixel information beyond the consistency of a single pixel, thereby revealing the possibility of closer collaboration between inter-pixel correlation and consistency regularization. To this end, this embodiment introduces this technology by setting a set of representative reference points (called agents), adding agent-level correlation, and modeling the prior inertia between pixels. Specifically, first, for each pixel in the weakly enhanced or strongly enhanced image, the agent-level correlation is obtained by comparing it with the agent; then, the KL divergence (KL divergence is a measure of information loss between two probability distributions) is used to model the relationship between agents, and ranking-aware consistency is designed instead of treating each agent equally and independently; the agent ranking is regarded as a random event rather than a deterministic arrangement. On this basis, for a given pixel, each possible ranking arrangement of the agent is considered, and the agent-level correlation is converted into an agent ranking probability distribution; finally, by constructing a loss to constrain the consistency of the agent ranking probability distribution between the teacher and student networks, the model can be guided by more effective supervisory signals.
[0066] Specifically, to better model the consistency regularization of inter-pixel correlations, a proxy correlation is constructed by comparing each pixel with a set of representative reference points.
[0067] Considering that the proxy should have a wide range of semantic contrasts with various semantic clues from the original pixels, this embodiment designs an orthogonal selection strategy to select the most representative proxy (or proxy point) from the image, that is, for the high-level semantic feature map of the unlabeled image extracted by the teacher model and the student model, a feature point is randomly selected in the feature map as the proxy point, and then a group of feature points with the maximum orthogonality to the selected proxy point are incrementally selected to form / compose a group of proxy points. On this basis, according to the group of proxy points and the feature map, the proxy correlation c=softmax(fA T ), where A is a set of proxy points and f is the pixel feature.
[0068] Secondly, in order to further utilize the correlation between agents to obtain more effective supervision signals, a perceptual ranking consistency regularization is designed. The sorting of agent correlation (also known as agent point correlation) is regarded as a random event rather than a deterministic arrangement, that is, each sorting method of the selected multiple agent points exists with a certain probability, not just from the largest to the smallest. On this basis, the agent correlation c calculated by different sorting methods is different, so all the arrangement methods are modeled. Given the agent correlation c, the probability of the arrangement under the agent correlation c is calculated. The calculation formula can be defined as:
[0069]
[0070] By the above method, all permutations are calculated The probability of the agent correlation is converted into the agent ranking probability, thereby realizing the modeling of the relationship between agents. To improve the calculation efficiency, this embodiment selects the arrangement between the top four agents in probability ranking.
[0071] Based on the proxy ranking probability of the high-level semantic feature graph extracted by the teacher model and the student model, the proxy-related loss function based on the consistency of the proxy ranking probability distribution is constructed as:
[0072]
[0073] In the above formula, |B u | represents the total number of unlabeled images, l kl represents the KL divergence, represents the proxy correlation of the teacher model, Represents the proxy correlation of the student model.
[0074] Step S4.3: Construct a supervised loss function. A supervised loss function is constructed based on the pixel-level classification prediction results of the labeled image output by the student model and the actual pixel-level labels.
[0075] Specifically, the supervised loss function is defined as:
[0076]
[0077] Among them, |B1| represents the number of supervised images, represents the true label of the jth pixel in the i-th image, Represents the prediction result of the student model for the jth pixel of the i-th image, and the pixel resolution of the input image is initialized to H×W.
[0078] Step S4.4, construct a total loss function, and use the total loss function to optimize and update the network parameters of the teacher-student model until the total loss function is minimized to obtain a trained teacher-student model.
[0079] In this embodiment, the network model is trained in an end-to-end manner, and the final loss function of the model is composed of the supervised loss L sup , cross-image unsupervised loss L usup and the agent-related loss L rank It consists of three parts and can be expressed as:
[0080] L guide =L sup +αL usup +βL rank
[0081] Among them, α and β are used to control the contribution of unsupervised loss and proxy-related loss in the training process, and supervised loss and unsupervised loss adopt a semi-supervised general method.
[0082] Through the above method, the weakly enhanced unlabeled image is input into the teacher model, and the pseudo label of the image is obtained through the encoder-decoder; the weakly enhanced labeled image and the strongly enhanced unlabeled image are input into the student model, the labeled image is trained with the label, and the unlabeled image is back-propagated and updated with the pseudo label generated by the teacher model. The teacher model does not use back-propagation update, but obtains the network parameters of the student model through EMA (exponential moving average) for gradient update.
[0083] Finally, the scene image to be segmented is obtained, and the scene image to be segmented is input into the trained student model to output a more accurate image semantic segmentation result.
[0084] like Figure 4 As shown, for the original image, its true label is obtained, and the original image is segmented using the method proposed in this embodiment and other existing methods. According to the comparison between the segmentation results and the true label, the segmentation accuracy and segmentation effect of the solution proposed in this embodiment are obviously better.
[0085] Embodiment 2
[0086] This embodiment provides an image semi-supervised semantic segmentation system based on pixel-level correction, including:
[0087] An image acquisition module, used to acquire a scene image to be segmented;
[0088] The image semantic segmentation module is used to input the scene image to be segmented into the trained student model and output the image semantic segmentation result; wherein the training process of the student model includes:
[0089] Collecting a number of scene images to construct a training data set; the training data set includes labeled images and unlabeled images, wherein the labeled images are images annotated with pixel-level labels;
[0090] By performing random diversity image enhancement on the dataset images, weakly enhanced labeled images, weakly enhanced unlabeled images, and strongly enhanced unlabeled images are obtained;
[0091] Construct a teacher-student model with the same structural composition but different parameter initialization;
[0092] The teacher-student model is trained using the training dataset, and a total loss function based on the supervised loss function, the cross-image unsupervised loss function, and the agent-related loss function is constructed. The network parameters of the teacher-student model are optimized and updated through the total loss function until the total loss function is minimized, thus obtaining the trained teacher-student model.
[0093] Embodiment 3
[0094] This embodiment provides an electronic device, including a memory and a processor, and computer instructions stored in the memory and executed on the processor. When the computer instructions are executed by the processor, the steps in the image semi-supervised semantic segmentation method based on pixel-level correction as described above are completed.
[0095] Embodiment 4
[0096] This embodiment also provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the steps in the image semi-supervised semantic segmentation method based on pixel-level correction as described above are completed.
[0097] The steps involved in the above embodiments 2 to 4 correspond to the method embodiment 1, and the specific implementation methods can refer to the relevant description part of embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.
[0098] Those skilled in the art should understand that the modules or steps of the present invention described above can be implemented by a general-purpose computer device, or alternatively, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.
[0099] The above description is only a preferred embodiment of the present invention. Although the specific implementation mode of the present invention is described in conjunction with the accompanying drawings, it is not a limitation of the protection scope of the present invention. Those skilled in the art should understand that on the basis of the technical solution of the present invention, various modifications or variations that can be made by those skilled in the art without creative work are still within the protection scope of the present invention.
Claims
1. A semi-supervised semantic segmentation method for images based on pixel-level correction, characterized in that: include: Obtaining a scene image to be segmented; The scene image to be segmented is input into the trained student model, and the image semantic segmentation result is output; wherein the training process of the student model includes: Collecting a number of scene images to construct a training data set; the training data set includes labeled images and unlabeled images, wherein the labeled images are images annotated with pixel-level labels; By performing random diversity image enhancement on the dataset images, weakly enhanced labeled images, weakly enhanced unlabeled images, and strongly enhanced unlabeled images are obtained; Construct a teacher-student model with the same structural composition but different parameter initialization; The teacher-student model is trained using the training dataset, and a total loss function based on the supervised loss function, the cross-image unsupervised loss function, and the agent-related loss function is constructed. The network parameters of the teacher-student model are optimized and updated through the total loss function until the total loss function is minimized, thus obtaining the trained teacher-student model.
2. The method for semi-supervised semantic segmentation of an image based on pixel-level correction according to claim 1, characterized in that: Construct a random diverse enhancement pool including a variety of strong enhancement methods and weak enhancement methods; Among them, the weakly enhanced image is an image enhanced by using weak enhancement methods at the image space level, and the weak enhancement methods include random scaling, random inversion and random cropping; the strongly enhanced image is an image enhanced by at most k strong enhancement methods randomly selected from the enhancement pool, and the strong enhancement methods include CutMix, automatic contrast, equalization, Gaussian filtering, contrast, sharpening, color balance, brightness, hue jitter, hue separation, exposure, and color inversion.
3. The method for semi-supervised semantic segmentation of an image based on pixel-level correction as claimed in claim 1, characterized in that: The teacher model and the student model have the same composition and both use an encoder-decoder network architecture. After the encoder extracts high-level semantic features from the input image, the decoder restores it to the same scale as the input image and outputs the pixel-by-pixel classification prediction result. The encoder is composed of a ResNet50 network, including a convolutional layer, a batch normalization layer, a ReLU activation layer, and a four-stage residual block connected in sequence; The decoder adopts the decoder of DeepLab V3 network, including multiple upsampling layers.
4. The method for semi-supervised semantic segmentation of an image based on pixel-level correction according to claim 1, characterized in that: The weakly enhanced unlabeled image is input into the teacher model, and the output pixel-by-pixel classification prediction result is used as the pseudo label; the weakly enhanced labeled image and the strongly enhanced unlabeled image are input into the student model, and the pixel-by-pixel classification prediction result is output; Based on the pixel-level classification prediction results of the labeled images output by the student model and the actual pixel-level labels, a supervised loss function is constructed; Based on the pixel-level pseudo labels and classification prediction results of the unlabeled images output by the teacher model and the student model, the cross-image similarity method is used for pixel-level correction to construct a cross-image unsupervised loss function; Based on the high-level semantic features of unlabeled images extracted by the teacher model and the student model, the agent-level correlation is calculated and converted into agent ranking probability distribution, and an agent correlation loss function based on the consistency of agent ranking probability distribution between the teacher and student models is constructed.
5. The method for semi-supervised semantic segmentation of an image based on pixel-level correction as claimed in claim 4, characterized in that: The construction of the cross-image unsupervised loss function includes: For the pixel-level pseudo-labels and classification prediction results of the unlabeled images output by the teacher model and the student model, a class is randomly selected from the pseudo-labels and classification prediction results to form a pseudo-mask; the pseudo-mask represents the positions of all pixels predicted to be of this class in the pseudo-labels or classification prediction results; Select a labeled image containing the class and get the class mask of the class; In the encoder of the model, high-level semantic features of unlabeled and labeled images are extracted, and the high-level semantic features of labeled images are multiplied by the class mask to extract the supporting features of the labeled images, so as to calculate the pixel-level cosine similarity between the high-level semantic features of the unlabeled images and the supporting features; Based on the pixel-level cosine similarity between the unlabeled images and the labeled images output by the teacher model and the student model, pixel-level correction is performed to construct a cross-image unsupervised loss function.
6. The method for semi-supervised semantic segmentation of an image based on pixel-level correction as claimed in claim 4, characterized in that: For the high-level semantic feature map of the unlabeled image extracted by the teacher model and the student model, a feature point is randomly selected as a proxy point in the feature map, and a group of feature points with the maximum orthogonality to the selected proxy point are incrementally selected to form a proxy. Then, based on the proxy and the high-level semantic feature map, the proxy correlation is calculated. According to the proxy correlation under different proxy ranking methods, the probability of all proxy ranking methods is calculated, so as to convert the proxy correlation into proxy ranking probability; Based on the proxy ranking probabilities of the high-level semantic feature maps extracted by the teacher model and the student model, an proxy-related loss function based on the consistency of the proxy ranking probability distribution is constructed.
7. A semi-supervised semantic segmentation system for images based on pixel-level correction, characterized in that: include: An image acquisition module, used to acquire a scene image to be segmented; The image semantic segmentation module is used to input the scene image to be segmented into the trained student model and output the image semantic segmentation result; wherein the training process of the student model includes: Collecting a number of scene images to construct a training data set; the training data set includes labeled images and unlabeled images, wherein the labeled images are images annotated with pixel-level labels; By performing random diversity image enhancement on the dataset images, weakly enhanced labeled images, weakly enhanced unlabeled images, and strongly enhanced unlabeled images are obtained; Construct a teacher-student model with the same structural composition but different parameter initialization; The teacher-student model is trained using the training dataset, and a total loss function based on the supervised loss function, the cross-image unsupervised loss function, and the agent-related loss function is constructed. The network parameters of the teacher-student model are optimized and updated through the total loss function until the total loss function is minimized, thus obtaining the trained teacher-student model.
8. The image semi-supervised semantic segmentation system based on pixel-level correction as claimed in claim 7, characterized in that: The weakly enhanced unlabeled image is input into the teacher model, and the output pixel-by-pixel classification prediction result is used as the pseudo label; the weakly enhanced labeled image and the strongly enhanced unlabeled image are input into the student model, and the pixel-by-pixel classification prediction result is output; Based on the pixel-level classification prediction results of the labeled images output by the student model and the actual pixel-level labels, a supervised loss function is constructed; Based on the pixel-level pseudo labels and classification prediction results of the unlabeled images output by the teacher model and the student model, the cross-image similarity method is used for pixel-level correction to construct a cross-image unsupervised loss function; Based on the high-level semantic features of unlabeled images extracted by the teacher model and the student model, the agent-level correlation is calculated and converted into agent ranking probability distribution, and an agent correlation loss function based on the consistency of agent ranking probability distribution between the teacher and student models is constructed.
9. An electronic device, characterized in that: The invention comprises a memory and a processor and computer instructions stored in the memory and executed on the processor. When the computer instructions are executed by the processor, the steps of a method for semi-supervised semantic segmentation of an image based on pixel-level correction as claimed in any one of claims 1 to 6 are completed.
10. A computer-readable storage medium, characterized in that: Used to store computer instructions, which, when executed by a processor, complete the steps of a method for semi-supervised semantic segmentation of an image based on pixel-level correction as described in any one of claims 1-6.
Citation Information
Cited By
Medical image classification method and related equipment
CN120236150A
Semantic alignment and deformation enhancement multi-parameter MRI (Magnetic Resonance Imaging) image model construction method and system
CN120563660A
Semi-supervised image segmentation method based on adaptive pixel subdivision
CN120635465A
Method and device for training polypropylene film defect detection model and polypropylene film defect detection method
CN120783149A
Retinal vessel segmentation system based on pixel-level screening strategy
CN120997231A