A semantic segmentation map acquisition method, device, equipment and readable storage medium

By constructing a dual-branch deep semantic segmentation network and utilizing a deterministic loss function for superpixel label prediction, the resolution loss problem in deep semantic segmentation networks is solved, improving the accuracy of image semantic segmentation and saving annotation workload.

CN116563547BActive Publication Date: 2026-04-24ZHEJIANG UNIV CITY COLLEGE
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG UNIV CITY COLLEGE
Filing Date
2023-05-12
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing deep semantic segmentation networks suffer from resolution loss due to multiple downsampling in the encoder, making it difficult to recover the semantic boundary error of the image segmentation result and affecting the accuracy of image semantic segmentation.

Method used

A dual-branch deep semantic segmentation network is constructed, consisting of a student network and a teacher network. The deterministic loss function of superpixel label prediction is used, and the teacher network provides superpixel semantic consistency knowledge to compensate for resolution loss.

Benefits of technology

It improves the generalization accuracy of image semantic segmentation, saves annotation workload, and effectively compensates for the resolution loss caused by multiple downsampling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116563547B_ABST
    Figure CN116563547B_ABST
Patent Text Reader

Abstract

The application provides a semantic segmentation map acquisition method, device and equipment and a readable storage medium. The method comprises: acquiring a sample image set; constructing an initial semantic segmentation network structure, the initial semantic segmentation network structure comprising image input, a teacher network branch, a student network branch and a total loss function; training the initial semantic segmentation network structure using the sample image set to obtain a semantic segmentation network model; and performing semantic segmentation on a to-be-processed image using the semantic segmentation network model to obtain a semantic segmentation map corresponding to the to-be-processed image. The application constructs a double-branch deep semantic segmentation network, which can be trained end-to-end. The teacher network branch can provide superpixel semantic consistency knowledge of each image patch to the student network branch to make up for resolution loss caused by multiple down-sampling, thereby improving the generalization accuracy of image semantic segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image and remote sensing image depth semantic segmentation technology, and more specifically, to a method, apparatus, device, and readable storage medium for acquiring semantic segmentation maps. Background Technology

[0002] Semantic segmentation refers to pixel-level classification of images, that is, predicting a class label for each pixel. Recent deep semantic segmentation networks such as UNet, Residual UNet, and DeepLabv3+ have achieved significant success in image semantic segmentation. However, because these networks inevitably introduce multiple downsampling operations in the encoder, they cause varying degrees of resolution loss, resulting in semantic boundary errors in the image segmentation results. Even interpolation methods are difficult to recover from these errors during the decoding stage. Therefore, compensating for the resolution loss caused by downsampling is a key issue in improving the accuracy of image semantic segmentation. Summary of the Invention

[0003] The purpose of this invention is to provide a method, apparatus, device, and readable storage medium for obtaining semantic segmentation maps, so as to improve the above-mentioned problems.

[0004] To achieve the above objectives, the embodiments of this application provide the following technical solutions:

[0005] On one hand, embodiments of this application provide a method for obtaining a semantic segmentation map, the method comprising:

[0006] Obtain the sample image set;

[0007] Construct an initial semantic segmentation network structure, which includes an image input, a teacher network branch, a student network branch, and a total loss function;

[0008] The initial semantic segmentation network structure is trained using the sample image set to obtain a semantic segmentation network model;

[0009] The semantic segmentation network model is used to perform semantic segmentation on the image to be processed, thereby obtaining the semantic segmentation map corresponding to the image to be processed.

[0010] Secondly, embodiments of this application provide a semantic segmentation map acquisition device, the device including an acquisition module, a construction module, a training module and a segmentation module.

[0011] The acquisition module is used to acquire a set of sample images;

[0012] A construction module is used to construct an initial semantic segmentation network structure, which includes an image input, a teacher network branch, a student network branch, and a total loss function.

[0013] The training module is used to train the initial semantic segmentation network structure using the sample image set to obtain a semantic segmentation network model.

[0014] The segmentation module is used to perform semantic segmentation on the image to be processed using the semantic segmentation network model, so as to obtain the semantic segmentation map corresponding to the image to be processed.

[0015] Thirdly, embodiments of this application provide a semantic segmentation map acquisition device, the device including a memory and a processor. The memory is used to store a computer program; the processor is used to execute the computer program to implement the steps of the semantic segmentation map acquisition method described above.

[0016] Fourthly, embodiments of this application provide a readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the semantic segmentation map acquisition method described above.

[0017] The beneficial effects of this invention are as follows:

[0018] This invention leverages the knowledge that pixels within a superpixel have a high probability of semantic consistency to establish a novel loss function term, called the deterministic loss for superpixel predicted labels. Based on this loss function, a dual-branch deep semantic segmentation network is constructed. This network can be trained end-to-end. The teacher branch can provide the superpixel semantic consistency knowledge of each image patch to the student network to compensate for the resolution loss caused by multiple downsampling, thereby improving the generalization accuracy of image semantic segmentation. Furthermore, both the student and teacher networks can use currently mainstream segmentation networks, requiring only the addition of a single loss term.

[0019] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing embodiments of the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description

[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a schematic diagram of the semantic segmentation map acquisition method described in this embodiment of the invention;

[0022] Figure 2This is a schematic diagram of the semantic segmentation map acquisition device described in this embodiment of the invention;

[0023] Figure 3 This is a schematic diagram of the structure of the semantic segmentation map acquisition device described in this embodiment of the invention;

[0024] Figure 4 This is a structural diagram of the student network branch described in the embodiments of the present invention;

[0025] Figure 5 This is the manually annotated building result diagram described in the embodiments of the present invention;

[0026] Figure 6 This is the prediction result image obtained by inputting an image containing buildings into the first model, as described in this embodiment of the invention.

[0027] Figure 7 This is the prediction result image of inputting an image containing buildings into a semantic segmentation network model, as described in this embodiment of the invention. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0029] It should be noted that similar reference numerals or letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this invention, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0030] Example 1

[0031] like Figure 1 As shown, this embodiment provides a method for obtaining a semantic segmentation map, which includes steps S1, S2, S3 and S4.

[0032] Step S1: Obtain the sample image set;

[0033] In this step, 9000 high-resolution remote sensing image patches with a size of 256*256 are used as the sample image set;

[0034] Step S2: Construct an initial semantic segmentation network structure, which includes an image input, a teacher network branch, a student network branch, and a total loss function;

[0035] In this step, constructing the initial semantic segmentation network structure specifically includes step S21;

[0036] Step S21: Construct an initial semantic segmentation network structure. Among them, the student network branch includes a fully convolutional deep semantic segmentation network, which includes an encoder and a decoder. The encoder extracts hierarchical features of the image, including deep features and shallow features. The decoder is responsible for upsampling the extracted features and localizing them to each pixel, and through a Softmax classifier, outputs a semantic segmentation probability map; the teacher network branch includes a pre-trained superpixel downsampling model; the total loss function includes a first loss function and a second loss function. The first loss function is a semantic segmentation loss function, and the second loss function is a deterministic loss function for superpixel-level predicted labels.

[0037] In this step:

[0038] (1) The fully convolutional deep semantic segmentation network in the student network branch can be a fully convolutional network such as Unet, DeepLabV3+; among them, the structure of the student network branch is as Figure 4 shown. The encoder consists of 5 modules, namely D0, D1, D2, D3, D4; the D0 module consists of two consecutive 3*3 convolutional layers and a non-linear activation function; the Di (1 <= i <= 4) accepts the output of Di-1, and after passing through pooling, two consecutive 3*3 convolutional layers, and a non-linear activation function in sequence, outputs;

[0039] The decoder consists of 4 modules, namely U1, U2, U3, U4; the downsampling plugin consists of a 1*1 convolutional layer and a Softmax layer, accepts the output of D3, generates a feature association map with 9 channels, and performs downsampling on the output of D3 based on the feature association map; U1 uses the feature association map to upsample D4, and Ui (2 <= i <= 4) uses deconvolution for upsampling, then connects the output of D4-i, and after two consecutive 3*3 convolutional layers and a non-linear activation function, outputs;

[0040] Finally, the F(P) dimensionality reduction module accepts the output of U4 and uses a 1*1 convolutional layer to reduce its dimensionality to K channels. The output is a semantic segmentation map Pp, which is a B*K*W*H tensor, where K represents the number of classes, B represents the number of images per batch, W represents the image width, and H represents the image height. In this step, the convolution operation uses zero-padding, extending the boundary by one pixel before performing the convolution operation to ensure that the output resolution is the same as the input resolution.

[0041] In the student network branch, the output channels of each convolutional layer are fixed at 9. This convolutional layer is a plugin that can learn end-to-end and is based on the idea of ​​superpixel irregular sampling. It can effectively preserve information during downsampling and avoid information loss caused by using max pooling operations.

[0042] (2) The superpixel downsampling model in the teacher network branch can be obtained by training the superpixel sampling network using publicly available superpixel sampling networks such as SSN and FCN. During the training process, an unsupervised training mode can be adopted. The loss function of the superpixel sampling network is as follows:

[0043]

[0044]

[0045]

[0046] Where A represents the superpixel-pixel correlation matrix, p represents the pixel position; q s (p) represents the probability that a pixel belongs to a superpixel seed s, u s Represents the superpixel center feature, l s p represents the superpixel center position; p' represents the reconstructed pixel position; f(p) represents the feature of the pixel at position p, which can be RGB features or CIELAB color features, etc.; f'(p) represents the reconstructed pixel feature, dist uses the L2 norm, d represents the downsampling scale, and m represents the weight, used to balance the loss of feature and distance for superpixel partitioning; N p Let A represent the superpixel grid, and L(A) represent the loss function of the superpixel sampling network;

[0047] For an input image Sp, the trained superpixel downsampling model outputs a pixel-superpixel correlation matrix A, which can be a tensor of shape B*N*W*H, where N represents the number of superpixels, a hyperparameter that needs to be preset. The number of superpixels N is limited to: (W*H) / d 2This is equal to the number of regular grids with a height and width of d pixels, where d can also be understood as a downsampling factor; in the teacher network branch, the superpixel-level class probability distribution is also calculated based on the probability map and the correlation matrix. The calculated superpixel-level class probability distribution Ps is a B*K*N*N tensor, where each vector of length K represents the probability distribution of the predicted label of a certain superpixel.

[0048] In addition to training the superpixel sampling network to obtain the teacher network branch, the teacher network branch in this embodiment can also adopt the student network branch. However, in the specific implementation, K is not taken as the number of categories, but is limited to K=9. This means that each pixel falling in the small black rectangle can only be assigned to the surrounding 9 grids, and each regular small grid represents a superpixel seed.

[0049] (3) The total loss function L = a*L1 + (1-a)L2, where L1(Pp,T) is the first loss function, which is a commonly used semantic segmentation loss function, and T represents the real label map; the first loss function L1 is obtained by calculating the cross-entropy or cross-union ratio between the predicted semantic segmentation probability map and the real label; the parameter a in the total loss function can be simply set to 0.3, so that the second loss function has a higher weight;

[0050] L2(Ps) is the second loss function, a deterministic loss function for superpixel-level predicted labels; 'a' is a hyperparameter, ranging from 0 to 1, used to adjust the intensity of knowledge distillation learning; the second loss function L2 is derived by jointly considering the information entropy of pixel-level and superpixel-level predicted labels. By minimizing this joint information entropy, the determinism of superpixel-level predicted labels is improved. Specifically, L2 is represented as follows:

[0051] L2=EntropyLoss(Pp)+EntropyLoss(Ps)

[0052] Wherein, EntropyLoss(Pp) represents the average entropy of the prediction result for each pixel, and EntropyLoss(Ps) represents the average entropy of the prediction result for each superpixel.

[0053] Step S3: Train the initial semantic segmentation network structure using the sample image set to obtain the semantic segmentation network model;

[0054] In this step, the overall training process includes steps S31 and S32;

[0055] Step S31, Training process: Randomly shuffle the images in the training image set and divide them into m batches. For each batch of images, perform forward propagation through the initial semantic segmentation network structure to calculate the total loss value. At the same time, perform gradient backpropagation to update the network parameters. Here, m is a positive integer.

[0056] In this step, the Adam optimizer is first used for gradient optimization iteration to update the network weights, with the initial learning rate set to 0.01. Each batch of training inputs b images, such as b=16. The iteration round parameter e=100 is set, and in each round of training, the input samples are randomly shuffled so that the images in each batch are randomly extracted. The superpixel sampling interval d=4, and the larger d is, the greater the loss of accuracy. During the training process, every 1000 batches of image patches are trained and the corresponding loss value and training accuracy are output.

[0057] Step S32: Repeat the training process continuously until the total loss value converges to obtain the semantic segmentation network model.

[0058] In addition to the training steps mentioned above, during the training process, 1,000 high-resolution remote sensing patches can be used as a validation sample set. At regular intervals, the model can be tested using the validation samples, and the hyperparameters can be adjusted accordingly.

[0059] In addition, the specific training process in this step includes steps S33 and S34;

[0060] Step S33: During the training process, for each sample image, the sample image is preprocessed to obtain a preprocessed image; the preprocessed image is input into the teacher network branch and the student network branch respectively. In the student network branch, it passes through a fully convolutional deep semantic segmentation network and outputs a semantic segmentation probability map, which represents the category probability distribution of each pixel; in the teacher network branch, it passes through a trained superpixel downsampling model branch and outputs a superpixel downsampling correlation matrix, which describes the correlation between pixels and superpixels. Then, the superpixel-level category probability distribution is calculated based on the probability map and the correlation matrix.

[0061] In this step, after the image input module inputs the sample image, it performs preprocessing such as centering, normalization, and image enhancement on the sample image before inputting it into the encoder. Among them, the image enhancement method can be rotation, flipping, cropping and magnification, color space transformation, etc.

[0062] Step S34: Calculate the first loss function value based on the semantic segmentation probability map, calculate the second loss function value based on the semantic segmentation probability map and the superpixel-level category probability distribution, and calculate the total overall loss value based on the first loss function value, the second loss function value and their respective weights.

[0063] The semantic segmentation network model trained through the above steps has the following advantages:

[0064] (1) Precision compensation: When the number of superpixels N is large, the pixel set contained in each superpixel has strong semantic consistency, which is equivalent to the superpixel label having strong determinism and low information entropy. The fusion of this knowledge can significantly improve the accuracy of the semantic segmentation model.

[0065] (2) Saves annotation: Due to the introduction of additional knowledge, the amount of annotation work can be significantly reduced under the same model generalization accuracy requirements.

[0066] Step S4: Use the semantic segmentation network model to perform semantic segmentation on the image to be processed, and obtain the semantic segmentation map corresponding to the image to be processed.

[0067] In this step, the image containing buildings is processed. Figure 5 The displayed results show manually annotated buildings; when an image containing buildings is input into the first model (which is a model with the teacher network branches removed), the prediction results are as follows. Figure 6 The image containing the building is input into the model in this embodiment, and the prediction result is as follows: Figure 7 ;according to Figures 5-7 It is easy to see that, in the test results, the model in this embodiment can capture the location and boundaries of all buildings quite well.

[0068] Example 2

[0069] like Figure 2 As shown, this embodiment provides a semantic segmentation map acquisition device, which includes an acquisition module 701, a construction module 702, a training module 703, and a segmentation module 704.

[0070] The acquisition module 701 is used to acquire a sample image set;

[0071] The construction module 702 is used to construct an initial semantic segmentation network structure, which includes an image input, a teacher network branch, a student network branch, and a total loss function.

[0072] Training module 703 is used to train the initial semantic segmentation network structure using the sample image set to obtain a semantic segmentation network model;

[0073] The segmentation module 704 is used to perform semantic segmentation on the image to be processed using the semantic segmentation network model to obtain a semantic segmentation map corresponding to the image to be processed.

[0074] In one specific embodiment of this disclosure, the construction module 702 further includes a construction unit 7021.

[0075] The construction unit 7021 is used to construct the initial semantic segmentation network structure. The student network branch includes a fully convolutional deep semantic segmentation network, which includes an encoder and a decoder. The encoder extracts hierarchical features of the image, including deep features and shallow features. The decoder is responsible for upsampling the extracted features and locating them to each pixel. Through a Softmax classifier, a semantic segmentation probability map is output. The teacher network branch includes a pre-trained superpixel downsampling model. The total loss function includes a first loss function and a second loss function. The first loss function is a semantic segmentation loss function, and the second loss function is a deterministic loss function for superpixel-level predicted labels.

[0076] In one specific embodiment of this disclosure, the training module 703 further includes a training unit 7031 and a repetition unit 7032.

[0077] Training unit 7031 is used for the training process: randomly shuffling the images in the training image set and dividing them into m batches; for each batch of images, forward propagation is performed through the initial semantic segmentation network structure to calculate the total loss value, and gradient backpropagation is performed to update the network parameters, where m is a positive integer;

[0078] The repeating unit 7032 is used to continuously repeat the training process until the total loss value converges, thereby obtaining the semantic segmentation network model.

[0079] In one specific embodiment of this disclosure, the training module 703 further includes a processing unit 7033 and a decoding unit 7034.

[0080] The processing unit 7033 is configured to preprocess each sample image during training to obtain a preprocessed image; input the preprocessed image into a teacher network branch and a student network branch respectively; in the student network branch, the image is processed through a fully convolutional deep semantic segmentation network to output a semantic segmentation probability map, which represents the class probability distribution of each pixel; in the teacher network branch, the image is processed through a trained superpixel downsampling model branch to output a superpixel downsampling correlation matrix, which describes the relationship between pixels and superpixels; and then calculate the superpixel-level class probability distribution based on the probability map and the correlation matrix.

[0081] The decoding unit 7034 is used to calculate a first loss function value based on the semantic segmentation probability map, calculate a second loss function value based on the semantic segmentation probability map and the superpixel-level category probability distribution, and calculate a total loss value based on the first loss function value, the second loss function value and their respective weights.

[0082] It should be noted that the specific manner in which each module performs its operation in the apparatus described in the above embodiments has been described in detail in the embodiments of the method, and will not be elaborated here.

[0083] Example 3

[0084] Corresponding to the above method embodiments, this disclosure also provides a semantic segmentation map acquisition device. The semantic segmentation map acquisition device described below and the semantic segmentation map acquisition method described above can be referred to in correspondence.

[0085] Figure 3 This is a block diagram illustrating a semantic segmentation map acquisition device 800 according to an exemplary embodiment. Figure 3 As shown, the semantic segmentation map acquisition device 800 may include a processor 801 and a memory 802. The semantic segmentation map acquisition device 800 may also include one or more of a multimedia component 803, an I / O interface 804, and a communication component 805.

[0086] The processor 801 controls the overall operation of the semantic segmentation map acquisition device 800 to complete all or part of the steps in the semantic segmentation map acquisition method described above. The memory 802 stores various types of data to support the operation of the semantic segmentation map acquisition device 800. This data may include, for example, instructions for any application or method operating on the semantic segmentation map acquisition device 800, and application-related data such as contact data, sent and received messages, images, audio, video, etc. The memory 802 can be implemented using any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The multimedia component 803 may include a screen and an audio component. The screen may be, for example, a touchscreen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in the memory 802 or transmitted via the communication component 805. The audio component also includes at least one speaker for outputting audio signals. I / O interface 804 provides an interface between processor 801 and other interface modules, such as keyboards, mice, and buttons. These buttons can be virtual or physical. Communication component 805 is used for wired or wireless communication between the semantic segmentation map acquisition device 800 and other devices. Wireless communication includes Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, or 4G, or a combination thereof. Therefore, the corresponding communication component 805 may include a Wi-Fi module, a Bluetooth module, or an NFC module.

[0087] In an exemplary embodiment, the semantic segmentation map acquisition device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the semantic segmentation map acquisition method described above.

[0088] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, which, when executed by a processor, implement the steps of the semantic segmentation map acquisition method described above. For example, the computer-readable storage medium may be the memory 802 including the program instructions described above, which may be executed by the processor 801 of the semantic segmentation map acquisition device 800 to complete the semantic segmentation map acquisition method described above.

[0089] Example 4

[0090] Corresponding to the above method embodiments, this disclosure also provides a readable storage medium. The readable storage medium described below and the semantic segmentation map acquisition method described above can be referred to in correspondence.

[0091] A readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the steps of the semantic segmentation map acquisition method of the above method embodiments are implemented.

[0092] Specifically, the readable storage medium can be a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, or any other readable storage medium capable of storing program code.

[0093] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for obtaining a semantic segmentation map, characterized in that, include: Obtain the sample image set; Construct an initial semantic segmentation network structure, which includes an image input, a teacher network branch, a student network branch, and a total loss function; The initial semantic segmentation network structure is trained using the sample image set to obtain a semantic segmentation network model, including: Training process: Images in the training image set are randomly shuffled and divided into m batches. For each batch of images, forward propagation is performed through the initial semantic segmentation network structure to calculate the total loss value. At the same time, gradient backpropagation is performed to update the network parameters, where m is a positive integer. The training process is repeated until the total loss value converges to obtain the semantic segmentation network model. The semantic segmentation network model is used to perform semantic segmentation on the image to be processed, obtaining a semantic segmentation map corresponding to the image to be processed. This includes: during training, preprocessing each sample image to obtain a preprocessed image; inputting the preprocessed image into a teacher network branch and a student network branch respectively; in the student network branch, passing through a fully convolutional deep semantic segmentation network to output a semantic segmentation probability map, which represents the class probability distribution of each pixel; in the teacher network branch, passing through a trained superpixel downsampling model branch to output a superpixel downsampling association matrix, which describes the association relationship between pixels and superpixels; calculating the superpixel-level class probability distribution based on the probability map and the association matrix; calculating a first loss function value based on the semantic segmentation probability map; jointly calculating a second loss function value based on the semantic segmentation probability map and the superpixel-level class probability distribution; and calculating a total loss value based on the first loss function value, the second loss function value, and their respective weights.

2. The method for obtaining a semantic segmentation map according to claim 1, characterized in that, Constructing the initial semantic segmentation network structure includes: An initial semantic segmentation network structure is constructed, wherein the student network branch includes a fully convolutional deep semantic segmentation network, which includes an encoder and a decoder. The encoder extracts hierarchical features of the image, including deep features and shallow features, and the decoder is responsible for upsampling the extracted features and locating them to each pixel. Through a Softmax classifier, a semantic segmentation probability map is output. The teacher network branch includes a pre-trained superpixel downsampling model. The total loss function includes a first loss function and a second loss function. The first loss function is the semantic segmentation loss function, and the second loss function is the deterministic loss function for superpixel-level predicted labels.

3. A device for acquiring semantic segmentation maps, characterized in that, include: The acquisition module is used to acquire a set of sample images; A construction module is used to construct an initial semantic segmentation network structure, which includes an image input, a teacher network branch, a student network branch, and a total loss function. The training module is used to train the initial semantic segmentation network structure using the sample image set to obtain a semantic segmentation network model. The segmentation module is used to perform semantic segmentation on the image to be processed using the semantic segmentation network model, so as to obtain the semantic segmentation map corresponding to the image to be processed. The training module includes: The training unit is used for the training process: images in the training image set are randomly shuffled and divided into m batches. For each batch of images, forward propagation is performed through the initial semantic segmentation network structure to calculate the total loss value. At the same time, gradient backpropagation is performed to update the network parameters, where m is a positive integer. A repeating unit is used to continuously repeat the training process until the total loss value converges, thereby obtaining the semantic segmentation network model. The training module includes: The processing unit is used to preprocess each sample image during training to obtain a preprocessed image; input the preprocessed image into the teacher network branch and the student network branch respectively; in the student network branch, it passes through a fully convolutional deep semantic segmentation network and outputs a semantic segmentation probability map, which represents the class probability distribution of each pixel; in the teacher network branch, it passes through a trained superpixel downsampling model branch and outputs a superpixel downsampling correlation matrix, which describes the relationship between pixels and superpixels; and then calculate the superpixel-level class probability distribution based on the probability map and the correlation matrix. The decoding unit is configured to calculate a first loss function value based on the semantic segmentation probability map, jointly calculate a second loss function value based on the semantic segmentation probability map and the superpixel-level category probability distribution, and calculate a total loss value based on the first loss function value, the second loss function value and their respective weights.

4. The semantic segmentation map acquisition device according to claim 3, characterized in that, Build modules, including: The building unit is used to construct the initial semantic segmentation network structure. The student network branch includes a fully convolutional deep semantic segmentation network, which comprises an encoder and a decoder. The encoder extracts hierarchical features from the image, including deep and shallow features. The decoder upsamples the extracted features and locates them to each pixel. A semantic segmentation probability map is output through a Softmax classifier. The teacher network branch includes a pre-trained superpixel downsampling model. The total loss function includes a first loss function and a second loss function. The first loss function is a semantic segmentation loss function, and the second loss function is a deterministic loss function for superpixel-level label prediction.

5. A device for acquiring semantic segmentation maps, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the semantic segmentation graph acquisition method as described in claim 1 when executing the computer program.

6. A readable storage medium, characterized in that: The readable storage medium stores a computer program that, when executed by a processor, implements the steps of the semantic segmentation graph acquisition method as described in claim 1.

Citation Information

Patent Citations

  • Remote sensing image semantic segmentation model training method and device for contrast consistency learning

    CN114299380A

  • Semi-supervised satellite image semantic segmentation network construction method and device and electronic equipment

    CN115660069A