Semantic segmentation method, semantic segmentation model training method and electronic equipment

By introducing the first loss function and encoding-decoding network structure in the semantic segmentation model, the rich features of the decoding process are distilled into the encoding process, and the problem of insufficient texture details recognition in image semantic segmentation of deep neural networks is solved, achieving higher segmentation accuracy and detail extraction.

CN120339598AActive Publication Date: 2025-07-18HONOR DEVICE CO LTD

Patent Information

Application Number
CN202410034420.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-09
Publication Date
2025-07-18
Estimated Expiration
2044-01-09

AI Technical Summary

Technical Problem

Existing deep neural networks have the problem of poor semantic segmentation effect in image semantic segmentation tasks, especially insufficient texture details recognition, resulting in a high risk of false positives.

Method used

By introducing the first loss function in the semantic segmentation model, the richer feature maps are distilled into the encoding process during the decoding process, the model parameters are adjusted to improve the feature extraction capability, and the encoding-decoding network and prediction head structure are adopted, combined with the pyramid pooling module and the residual network to improve the details and accuracy of the segmentation results.

Benefits of technology

It improves the detail richness and accuracy of image segmentation results, reduces the risk of false positives, and enhances the feature extraction ability of semantic segmentation models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339598A_ABST
    Figure CN120339598A_ABST
Patent Text Reader

Abstract

The invention provides a semantic segmentation method, a semantic segmentation model training method and electronic equipment, and relates to the technical field of image processing. According to the method, the feature extraction capability of the semantic segmentation model in the encoding process can be improved, so that the image segmentation result has richer details, and the accuracy of the image segmentation result is improved. The method comprises the steps of obtaining a first image; inputting the first image into a semantic segmentation model to obtain a semantic segmentation result; wherein the semantic segmentation model comprises a first model parameter, the first model parameter is associated with a first loss function, the first loss function is determined according to a first feature map and a second feature map, and the first feature map is a feature map generated by a basic semantic segmentation model for generating the semantic segmentation model in an encoding process; the second feature map is a feature map generated in the decoding process of the basic semantic segmentation model, and the first feature map and the second feature map have the same resolution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of image processing, and in particular, to a semantic segmentation method, a method for training a semantic segmentation model, and an electronic device. Background Art

[0002] Image semantic segmentation is an important research content in the field of computer vision. Its goal is to segment an image into regions with different semantic information and label the corresponding semantic labels for each region.

[0003] Currently, using a deep neural network to process image semantic segmentation tasks is a relatively common solution in the industry, but this solution has the problem of poor semantic segmentation effect. Summary of the Invention

[0004] The embodiments of the present application provide a semantic segmentation method, a method for training a semantic segmentation model, and an electronic device, which can improve the feature extraction ability of the semantic segmentation model during the encoding process, thereby making the image segmentation result have richer details and reducing the risk of false positives in the image segmentation result.

[0005] To achieve the above object, the embodiments of the present application adopt the following technical solutions:

[0006] In a first aspect, a semantic segmentation method is provided. The method includes: obtaining a first image; inputting the first image into a semantic segmentation model to obtain a semantic segmentation result, where the semantic segmentation result is used to indicate the category information of each pixel in the first image. In the embodiments of the present application, the semantic segmentation model includes first model parameters, the first model parameters are associated with a first loss function, the first loss function is determined according to a first feature map and a second feature map, the first feature map is a feature map generated during the encoding process of a basic semantic segmentation model for generating the semantic segmentation model, the second feature map is a feature map generated during the decoding process of the basic semantic segmentation model, and the first feature map and the second feature map have the same resolution. It can be understood that considering that the feature map in the decoding process usually has more and richer features than the feature map in the encoding process, adjusting the first model parameters through the first loss function can distill the richer features in the decoding stage to the encoding stage to improve the feature extraction ability of the semantic segmentation model during the encoding process, thereby making the image segmentation result output by the semantic segmentation model have richer details and improving the accuracy of the image segmentation result.

[0007] In an implementation provided in the first aspect, the semantic segmentation model is obtained by training a basic semantic segmentation model based on training samples, where the training process includes: iteratively training the basic semantic segmentation model based on the training samples; if the basic semantic segmentation model does not converge, adjusting the model parameters of the basic semantic segmentation model based on the first loss function; if the basic semantic segmentation model converges, stopping the iterative training to obtain the semantic segmentation model.

[0008] In an implementation provided in the first aspect, the basic semantic segmentation model is used to perform N downsamplings and N upsamplings on the training samples. The first feature map is the feature map processed by the i-th downsampling of the basic semantic segmentation model, and the second feature map is the feature map obtained by the j-th upsampling of the basic semantic segmentation model, where i + j - 1 = N and i ≤ N. In this way, it can be ensured that the resolutions of the first feature map and the second feature map are the same.

[0009] In an implementation provided in the first aspect, the semantic segmentation model includes an encoder-decoder network and a prediction head; inputting the first image into the semantic segmentation model to obtain a semantic segmentation result, including: performing N downsamplings and N upsamplings on the first image in sequence through the encoder-decoder network to obtain a first intermediate feature map; processing the first intermediate feature map through the prediction head to obtain the semantic segmentation result.

[0010] In an implementation provided in the first aspect, the encoder-decoder network includes a convolution module, N downsampling modules, a pooling module, and N upsampling modules; performing N downsamplings and N upsamplings on the input image corresponding to the first image in sequence through the encoder-decoder network to obtain a first intermediate feature map, including: performing a convolution operation on the first image through the convolution module to obtain a second intermediate feature map; performing downsampling processing on the second intermediate feature map in sequence through N downsampling modules to obtain a third intermediate feature map; processing the third intermediate feature map through the pooling module to obtain a fourth intermediate feature map; performing upsampling processing on the fourth intermediate feature map in sequence through N upsampling modules to obtain the first intermediate feature map.

[0011] In an implementation provided in the first aspect, the N downsampling modules include a first downsampling module, a second downsampling module, and a third downsampling module. The first downsampling module includes a first residual network, the second downsampling module includes a first residual network and a second residual network, and the third downsampling module includes a first residual network, a second residual network, and a third residual network. The network structures of the first residual network, the second residual network, and the third residual network are different. It can be understood that the different network structures of multiple downsampling modules can extract features of the image at different resolutions, which is beneficial to extracting more details.

[0012] In an implementation provided in the first aspect, the pooling module includes a pyramid pooling module, and the pyramid pooling module includes 6 pooling layers. That is to say, the pyramid pooling module provided in the embodiments of the present application is extended from the traditional 4-channel to 6-channel, so that under the premise of capturing semantic features in a larger range, image details can be retained as much as possible, which is beneficial to improving the accuracy of the semantic segmentation result.

[0013] In an implementation provided in the first aspect, each upsampling module includes an initial module, and the initial module includes two convolutional layers and a splicing layer. The two convolutional layers are used to perform convolutional processing on the received input and send the processing results to the splicing layer, and the splicing layer is used to splice the processing results to obtain an output. The network structure of this initial module is simple, which can reduce the number of parameters of the semantic segmentation model.

[0014] In an implementation provided in the first aspect, the feature map processed by the i-th downsampling of the encoder-decoder network is the first feature map generated when the basic semantic segmentation model converges, and the feature map output by the j-th upsampling of the encoder-decoder network is the second feature map generated when the basic semantic segmentation model converges, where i + j - 1 = N and i ≤ N.

[0015] In an implementation provided in the first aspect, the first model parameters include all the parameters used before the i-th downsampling of the encoder-decoder network. That is to say, the first loss function is only used to adjust the parameters of the network layers necessary to generate the first feature map, so as to avoid the problem of identity mapping between the first feature map and the second feature map, and achieve the effect of effective training.

[0016] In an implementation provided in the first aspect, the first model parameters at least include the parameters of the convolutional module.

[0017] In an implementation provided in the first aspect, N = 3 and i = 1.

[0018] In an implementation provided in the first aspect, the first model parameters are also associated with a second loss function, and the second loss function is determined according to the predicted semantic segmentation result of the training image and the label of the training image, and the predicted semantic segmentation result is obtained by inputting the training image into the basic semantic segmentation model.

[0019] In an implementation provided in the first aspect, the semantic segmentation model further includes second model parameters, and the second model parameters include all the parameters of the semantic segmentation model, and the second model parameters are determined according to the second loss function.

[0020] Second aspect, an embodiment of the present application further provides a method for training a semantic segmentation model. The method includes: iteratively training a basic semantic segmentation model based on training samples; if the basic semantic segmentation model does not converge, adjusting the model parameters of the basic semantic segmentation model based on a first loss function; wherein, the first loss function is determined according to a first feature map and a second feature map, the first feature map is the feature map generated by the basic semantic segmentation model during the encoding process, the second feature map is the feature map generated by the basic semantic segmentation model during the decoding process, and the first feature map and the second feature map have the same resolution; if the basic semantic segmentation model converges, stopping the iterative training to obtain the semantic segmentation model.

[0021] In an implementation manner provided in the second aspect, the semantic segmentation model includes an encoder-decoder network and a prediction head; inputting a training image into the semantic segmentation model to obtain a predicted semantic segmentation result, including: performing N times of downsampling and N times of upsampling on the training image in sequence through the encoder-decoder network to obtain a first intermediate feature map; processing the first intermediate feature map through the prediction head to obtain the predicted semantic segmentation result.

[0022] In an implementation manner provided in the second aspect, the encoder-decoder network includes a convolutional module, N downsampling modules, a pooling module, and N upsampling modules; performing N times of downsampling and N times of upsampling on the input image corresponding to the training image in sequence through the encoder-decoder network to obtain a first intermediate feature map, including: performing a convolution operation on the training image through the convolutional module to obtain a second intermediate feature map; performing downsampling processing on the second intermediate feature map in sequence through N downsampling modules to obtain a third intermediate feature map; processing the third intermediate feature map through the pooling module to obtain a fourth intermediate feature map; performing upsampling processing on the fourth intermediate feature map in sequence through N upsampling modules to obtain the first intermediate feature map.

[0023] In an implementation manner provided in the second aspect, the N downsampling modules include a first downsampling module, a second downsampling module, and a third downsampling module. The first downsampling module includes a first residual network, the second downsampling module includes a first residual network and a second residual network, and the third downsampling module includes a first residual network, a second residual network, and a third residual network. The network structures of the first residual network, the second residual network, and the third residual network are different. It can be understood that the network structures of multiple downsampling modules are different, which can extract features of the image at different resolutions and is beneficial to extracting more details.

[0024] In an implementation provided in the second aspect, the pooling module includes a pyramid pooling module, and the pyramid pooling module includes 6 pooling layers. That is to say, the pyramid pooling module provided in the embodiments of the present application is extended from the traditional 4-channel to 6-channel, so that on the premise of capturing semantic features in a larger range, image details can be retained as much as possible, which is beneficial to improving the accuracy of the predicted semantic segmentation result.

[0025] In an implementation provided in the second aspect, each upsampling module includes an initial module, and the initial module includes two convolutional layers and a splicing layer. The two convolutional layers are used to perform convolutional processing on the received input and send the processing result to the splicing layer, and the splicing layer is used to splice the processing result to obtain an output. The network structure of this initial module is simple, which can reduce the number of parameters of the semantic segmentation model.

[0026] In a third aspect, an embodiment of the present application further provides an electronic device, which includes: a memory and one or more processors; the memory and the processor are coupled; the memory is used to store computer program code, and the computer program code includes computer instructions. When the computer instructions are executed by the electronic device, the electronic device is caused to execute the method as described in the first aspect, the second aspect and any of their implementation manners.

[0027] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, in which computer instructions are stored. When the computer instructions run on an electronic device, the electronic device is caused to execute the method as described in the first aspect, the second aspect and any of their implementation manners.

[0028] Among them, the technical effects brought by any of the design manners in the second aspect to the fourth aspect can refer to the technical effects brought by different design manners in the first aspect, which will not be elaborated here. Description of the Drawings

[0029] Figure 1 It is a schematic diagram of a scenario provided by an embodiment of the present application;

[0030] Figure 2 It is a schematic structural diagram of an AI system provided by an embodiment of the present application;

[0031] Figure 3 It is a schematic flowchart of the semantic segmentation method provided by an embodiment of the present application Figure 1 ;

[0032] Figure 4 It is a schematic flowchart of the semantic segmentation method provided by an embodiment of the present application Figure 2 ;

[0033] Figure 5Schematic diagram of the upsampling and downsampling processes for the encoding-decoding network provided by the embodiments of the present application;

[0034] Figure 6 Flow schematic of the semantic segmentation method provided by the embodiments of the present application Figure 3 ;

[0035] Figure 7 Schematic diagram of the structure of the residual network provided by the embodiments of the present application;

[0036] Figure 8 Further schematic diagram of the structure of the residual network provided by the embodiments of the present application;

[0037] Figure 9 Network structure diagram of a pyramid pooling module provided by the embodiments of the present application;

[0038] Figure 10 Network structure diagram of an initial module provided by the embodiments of the present application;

[0039] Figure 11 Network structure diagram of a semantic segmentation model provided by the embodiments of the present application;

[0040] Figure 12 Flow schematic diagram of a training method for a semantic segmentation model provided by the embodiments of the present application;

[0041] Figure 13 Schematic diagram of training a basic semantic segmentation model provided by the embodiments of the present application. Detailed implementation manners

[0042] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0043] The term "and / or" in this article is an association relationship describing associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The symbol " / " in this article represents an "or" relationship between associated objects. For example, A / B represents A or B.

[0044] In the description of the specification and claims in this document, terms such as "first" and "second" are used to distinguish different objects, rather than to describe a specific order of the objects. For example, the first feature map and the second feature map are used to distinguish different feature maps, rather than to describe a specific order of the feature maps.

[0045] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.

[0046] In the description of the embodiments of this application, unless otherwise specified, the meaning of "a plurality of" refers to two or more. For example, a plurality of processing units refers to two or more processing units; a plurality of elements refers to two or more elements, etc.

[0047] For the sake of clear and concise description of the following embodiments, a brief introduction to the related technologies is first given.

[0048] 1) Image segmentation

[0049] Image segmentation is a very important task in computer vision, and its goal is to classify each pixel point in the image, that is, a pixel-level classification task. Currently, there are mainly three types of image segmentation tasks: semantic segmentation, instance segmentation, and panoptic segmentation.

[0050] Semantic segmentation refers to identifying the category of each pixel in the image and assigning the corresponding category label to the pixel.

[0051] Instance segmentation refers to identifying each object or category instance present in the image and assigning different masks or bounding boxes with unique identifiers to them.

[0052] Panoptic segmentation refers to the combination of semantic segmentation and instance segmentation, and each pixel in the image can be assigned a semantic label (due to semantic segmentation) and a unique instance identifier.

[0053] Specifically, the solution of this application mainly involves semantic segmentation, which will be described below in combination with semantic segmentation. Since an image is composed of many pixels, semantic segmentation can be understood as grouping or segmenting pixels according to different semantic meanings expressed in the image. Through image semantic segmentation, the content in the image can be automatically segmented and recognized. When performing semantic segmentation, the number of classifications is first determined, and then an output channel is created for each category. Among them, a single channel represents the area where a specific category exists. The result after semantic segmentation can be expressed as (C, H, W) or C×H×W.

[0054] Exemplarily, as Figure 1 shown, after the original image undergoes semantic segmentation, the original image is segmented into four semantic categories: people, sky, flowers, and background. Correspondingly, 4 channels can be output. The result after semantic segmentation can be expressed as (4, H, W) or 4×H×W, where each channel classifies the pixel points into a single category as 1 or 0.

[0055] 2) Deep neural network

[0056] The above task of image semantic segmentation can be completed by a deep neural network. The deep neural network can include a convolutional neural network and a deconvolutional neural network.

[0057] The convolutional neural network is related to the extraction of image features and transforms the input image into a multi-dimensional feature matrix. The deconvolutional neural network is equivalent to a segmentation region generator, which can use the image features extracted by the convolutional neural network to perform semantic segmentation on objects.

[0058] The output of the entire deep neural network is a probability matrix diagram, which is the same size as the input image. The value of the element at each position in the matrix diagram represents the classification probability of the pixel point at the same position in the corresponding image, that is, the probability that the object to which the pixel point belongs is an object of a certain category.

[0059] It can be understood that the deep neural network can be regarded as an encoding and decoding network structure. Through the convolutional network (also called the encoding network), the image is convolved into a small matrix; through the deconvolutional network (also called the decoding network), the small matrix is restored to a large image.

[0060] 3) Feature map

[0061] A convolutional neural network includes, but is not limited to, one or more convolutional layers. Each convolutional layer may include multiple filters (or called convolutional kernels). Each filter is essentially an array, and the numbers in it are called convolutional weights or parameters. The role of the convolutional layer of the convolutional neural network is to perform a convolution operation on the input. For example, the filters in the convolutional layer can slide on the sample with a set stride. At each sliding position, the array of the filter is multiplied by the sample data and then added together to obtain a value. All the values obtained during the sliding process form a new array, which is called a feature map. Each value in the feature map is a feature of the feature map.

[0062] Among them, the feature map can be represented as a three-dimensional matrix of C×H×W, and this three-dimensional matrix can include C two-dimensional matrices of H×W. Among them, H represents the pixel height of the image to be processed, W represents the pixel width of the image to be processed, and C represents the number of channels of the image to be processed. For example, the number of channels of an RGB image is 3.

[0063] In the related art, it is a relatively common solution in the industry to use a deep neural network to process image semantic segmentation tasks. In this solution, a converged semantic segmentation model can be trained by using the supervised learning method, and then the converged semantic segmentation model is used to process the input image to obtain the final semantic segmentation result. However, the semantic segmentation model obtained based on the conventional supervised learning method cannot accurately identify the texture details of the image, resulting in poor semantic segmentation effects and even cases of semantic segmentation errors.

[0064] In view of this, the embodiments of the present application provide a semantic segmentation method, which can use a semantic segmentation model to obtain the semantic segmentation result of the first image. The first model parameters of the semantic segmentation model are associated with a first loss function, and the first loss function is determined according to the first feature map generated in the encoding process and the second feature map generated in the decoding process. In this way, the semantic segmentation model can distill richer features in the decoding stage to the encoding stage to improve the feature extraction ability of the semantic segmentation model in the encoding process, thereby making the image segmentation result have richer details and reducing the risk of false positives in the image segmentation result.

[0065] The semantic segmentation method provided by the embodiments of the present application can be applied to shooting scenarios, video call scenarios, medical scenarios, traffic scenarios, etc. For example, taking the shooting scenario as an example, the terminal can perform semantic segmentation on the preview image obtained by shooting to obtain the semantic segmentation result; then, based on the semantic segmentation result, adjust the local brightness of certain types of pixels in the preview image. For example, the brightness of the pixels with the category of "sky" in the preview image can be increased, and the brightness of the pixels with the category of "person" in the preview image can be decreased; finally, display the preview image after brightness adjustment.

[0066] As Figure 2 shown,Figure 2 The structural schematic diagram of the AI system provided by this application. As Figure 2 shown, the AI system includes a data center and multiple terminals (such as Figure 2 the shown terminal 111 and terminal 112). The data center can communicate with the terminals through a network, and the network can be the Internet or other networks. The network can include one or more network devices, such as the network device can be a router or a switch, etc.

[0067] The data center includes one or more servers, such as Figure 2 the shown server 120, for example, an application server that supports application services. The application server can provide image services, video services, game services, other AI processing services based on video or images, etc. In an alternative scenario, the server 120 refers to a server cluster deployed with multiple servers. The server cluster can have a rack, and the rack can establish communication for the multiple servers through a wired connection, such as a universal serial bus (USB) or a peripheral component interconnect express (PCIe) high-speed bus, etc.

[0068] The server 120 can also obtain data from the terminals, and after performing AI processing on the data, send the results of the AI processing to the corresponding terminals. The AI processing can refer to tasks such as object recognition, target detection, semantic segmentation, etc. on the data using an AI model, or can also refer to obtaining an AI model that meets the requirements based on the samples collected by the terminals.

[0069] In addition, Figure 2 the shown data center can also include other physical devices with AI processing functions, such as mobile phones, tablet computers, or other devices, etc.

[0070] The terminal can also be referred to as a terminal device, user equipment (UE), mobile station (MS), mobile terminal (MT), etc. The terminal can be a mobile phone (such as Figure 2 the shown terminal 111), a face recognition payment device with mobile payment function, a camera device with data (such as images or videos) collection and processing functions (such as Figure 2The terminal shown, such as the terminal 112), etc. The terminal can also be a tablet computer (Pad), a computer with wireless transceiver function, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal in transportation safety, a wireless terminal in a smart city, a wireless terminal in a smart home, and so on. The embodiments of the present application do not limit the specific technologies and specific device forms adopted by the terminal device.

[0071] It should be noted that the terminal can obtain the AI model stored in the server 120, etc. Furthermore, the terminal uses the AI model to perform various tasks on the information contained in the image. By way of example, the AI model can be a model for performing object detection, semantic segmentation, etc. on data for semantic segmentation.

[0072] For example, the AI model deployed in the terminal 112 can be used to implement functions such as object detection and semantic segmentation.

[0073] For another example, the AI model deployed in the terminal 111 is used to implement functions such as automatically adjusting the local brightness of an image.

[0074] For yet another example, the AI model deployed in the terminal 111 is used to implement functions such as face brushing payment and object classification (such as commodity classification).

[0075] It should be noted that Figure 2 This is only a schematic diagram and should not be construed as a limitation to the present application. The embodiments of the present application do not limit the application scenarios of the terminal and the server. For example, a frame of image collected by the terminal includes multiple different types of things. For example, the terminal can identify Figure 2 the shown class 1 (such as a person), class 2 (flowers), class 3 (such as the sky), etc.

[0076] To at least solve the above problems, based on the AI system shown below in Figure 2 the present embodiment provides a semantic segmentation method. As shown in Figure 3 shown, Figure 3 is the flow schematic of the semantic segmentation method provided by the embodiments of the present application Figure 1 , and this semantic segmentation method can be applied to Figure 2 the shown AI system. This semantic segmentation method can be executed by the terminal or the server. Here, the case where the terminal device executes the semantic segmentation method provided by the present embodiment is taken as an example for illustration.

[0077] As shown in Figure 3 shown, the semantic segmentation method provided by the embodiments of the present application includes steps S310 to S320.

[0078] S310, the terminal device obtains a first image.

[0079] The first image may be an image captured by the terminal device in real time, any frame of a video captured by the terminal device in real time, an image sent by another device to the terminal device, any frame of a video sent by another device to the terminal device, or an image pre-stored in the terminal device, etc., which is not specifically limited herein.

[0080] S320, the terminal device inputs the first image into a semantic segmentation model to obtain a semantic segmentation result.

[0081] The semantic segmentation result is used to indicate the category information of each pixel in the first image. In one possible design, the category information may be represented by the prediction confidence of the pixel belonging to the first category. In another possible design, the category information may further include the prediction confidence of the pixel belonging to other categories (such as the second category). In still another possible design, the category information may directly indicate the category to which the pixel belongs, for example, any one of the 4 categories shown as Figure 1 shown.

[0082] In one possible design, the semantic segmentation result may be presented in the form of an image. The terminal device may mark the category information of each pixel in the first image and use the marked first image as the semantic segmentation result. Exemplarily, the first image may be the Figure 1 original image shown, and the semantic segmentation result may be the Figure 1 semantically segmented image shown.

[0083] In the embodiments of the present application, the terminal device may mark the category information of each pixel in the image, and the marked image (i.e., the semantic segmentation result) can show the distribution of pixels of different categories in the image, so that the user can quickly view the distribution of pixels of different categories from the semantic segmentation result shown by the terminal device, improving the quality of experience (QoE).

[0084] In the embodiments of the present application, the semantic segmentation model may perform multiple downsampling processes and multiple upsampling processes. The process of multiple downsampling processes may also be referred to as an encoding process, and the process of multiple upsampling processes may also be referred to as a decoding process.

[0085] Downsampling is used to extract the detailed features of the image to obtain a feature map. By sequentially performing multiple downsampling processes on the input image, multiple feature maps with different resolutions can be obtained to represent the details and semantic information of the first image at different resolutions. In the embodiments of the present application, downsampling may be implemented through operations such as multi-layer convolution and pooling.

[0086] Upsampling is used to increase the resolution of an image. By successively performing multiple upsampling processes on the feature map output by downsampling, an image with the same size as the input image can be gradually obtained. In the embodiments of the present application, downsampling can be implemented by means such as nearest neighbor interpolation, bilinear interpolation, or bicubic interpolation.

[0087] The semantic segmentation model includes first model parameters, the first model parameters are associated with a first loss function, the first loss function is determined according to a first feature map and a second feature map, the first feature map is a feature map generated during the encoding process of the basic semantic segmentation model based on which the semantic segmentation model is generated, the second feature map is a feature map generated during the decoding process of the basic semantic segmentation model, and the first feature map and the second feature map have the same resolution. Among them, the semantic segmentation model can be understood as a converged basic semantic segmentation model.

[0088] It can be seen that in the embodiments of the present application, the first model parameters of the semantic segmentation model are associated with the first loss function, and the first loss function is determined according to two feature maps with the same resolution (such as the first feature map and the second feature map) generated during the encoding process and the decoding process of the basic semantic segmentation model. In this way, more abundant features in the decoding process can be distilled into the encoding process, enabling the semantic segmentation model to have stronger feature extraction capabilities during the encoding process, being able to extract more details, and thus making the image segmentation result more accurate and reducing the probability of false positives in the image segmentation result.

[0089] In the embodiments of the present application, the first model parameters are also associated with a second loss function. In other words, the first model parameters are determined according to the first loss function and the second loss function. Among them, the second loss function is determined according to the predicted semantic classification result of the training image and the label of the training image, the predicted semantic classification result is the semantic segmentation result obtained by inputting the training image into the basic semantic segmentation model, and the label of the training image is used to indicate the actual semantic segmentation situation of the training image.

[0090] In the embodiments of the present application, the semantic segmentation model further includes second model parameters, the second model parameters are different from the first model parameters, and the second model parameters are determined according to the second loss function.

[0091] In a possible design, the first loss function is the mean-square error (MSE) between the first feature map and the second feature map, and the second loss function is the mean-square error between the test result and the standard result. In other possible designs, the first loss function and the second loss function can also be others, which are not specifically limited herein.

[0092] For the process in which the foregoing terminal device obtains the semantic segmentation result by using the semantic segmentation model, the embodiments of the present application provide a possible implementation manner, such as Figure 4As shown Figure 4 is a schematic flowchart of the semantic segmentation method provided by an embodiment of the present application Figure 2 . Among them, the semantic segmentation model includes an encoder-decoder network and a prediction head

[0093] As Figure 4 shown, the semantic segmentation model can also preprocess the first image to obtain an input image corresponding to the first image

[0094] In the embodiment of the present application, the preprocessing includes operations such as changing the resolution and standardization

[0095] Among them, the purpose of changing the resolution is to unify the resolution of the input image, facilitating the processing by the semantic segmentation model. For example, the semantic segmentation model can change the resolution of the first image to a preset resolution, and the preset resolution is, for example, 384×384, etc

[0096] The standardization operation can adjust the pixel values in the first image so that they follow a standard normal distribution or have a specific mean and standard deviation. In a possible design, the mean and variance used by the semantic segmentation model for the standardization operation can be calculated based on the large-scale image data in the ImageNet (Image Network) dataset

[0097] In the embodiment of the present application, the semantic segmentation module can first adjust the resolution of the first image to the preset resolution, and then perform a standardization operation on the first image after the resolution adjustment to obtain an input image corresponding to the first image

[0098] It should be noted that the foregoing preprocessing can also include other graphic processing operations, such as flipping, translation, cropping, etc., which are not limited herein

[0099] In a possible design, the semantic segmentation model can also not preprocess the first image. After the terminal device obtains the input image corresponding to the first image, it directly inputs the input image corresponding to the first image into the semantic segmentation model

[0100] In the semantic segmentation method provided in this embodiment, the foregoing S320 may include S410 to S420

[0101] S410, the terminal device performs N times of downsampling and N times of upsampling on the input image corresponding to the first image through the encoder-decoder network to obtain a first intermediate feature map

[0102] Among them, as Figure 5 shown, the process of the encoder-decoder network performing N times of downsampling and N times of upsampling on the input image may include: performing a first downsampling process on the input image to obtain an intermediate feature map m d1; The result obtained from the first downsampling (i.e., the intermediate feature map m d1 ) is subjected to a second downsampling process to obtain the intermediate feature map m d2 ; and so on. The result obtained from the (N - 1)-th downsampling (i.e., the intermediate feature map m d(N-1) ) is subjected to the N-th downsampling process to obtain the intermediate feature map m dN .

[0103] Next, still as Figure 5 shown, the result of the N-th downsampling (i.e., the intermediate feature map m dN ) is subjected to a first upsampling process, and the result of the first upsampling and the result of the (N - 1)-th downsampling (i.e., the intermediate feature map m d(N-1) ) are fused to obtain the intermediate feature map m uN ; the intermediate feature map m uN is subjected to a second upsampling, and the result of the second upsampling and the result of the (N - 2)-th downsampling (i.e., the intermediate feature map m d(N-2) ) are fused to obtain the intermediate feature map m u(N - 1); and so on. The intermediate feature map m u3 is subjected to the (N - 1)-th upsampling, and the result of the (N - 1)-th upsampling and the structure of the first downsampling (i.e., the intermediate feature map m d1 ) are fused to obtain the intermediate feature map m u2 ; the intermediate feature map m u2 is subjected to the N-th upsampling, and the result of the N-th upsampling and the input image are fused to obtain the intermediate feature map m u1 (which can also be referred to as the first intermediate feature map).

[0104] Among them, the resolution of the intermediate feature map m d1 to the intermediate feature map m dN decreases in sequence, and the number of channels increases in sequence, indicating that the semantic segmentation model extracts more and more detailed features. In addition, the resolution of the intermediate feature map m uN to the intermediate feature map m u1 increases in sequence and gets closer and closer to the resolution of the input image.

[0105] The basic semantic segmentation model and the semantic segmentation model have the same network structure. The difference is that the basic semantic segmentation model is not converged, while the semantic segmentation model is converged. Therefore, in Figure 4In the provided flowchart, the above-mentioned first feature map can specifically be the feature map processed by the i-th downsampling of the basic semantic segmentation model, and the second feature map can specifically be the feature map obtained by the j-th upsampling of the basic semantic segmentation model, where i + j - 1 = N and i ≤ N. In this way, the first feature map and the second feature map can be the feature maps generated by the basic semantic segmentation model during the encoding process and the decoding process respectively, and the resolutions of the first feature map and the second feature map can be kept consistent.

[0106] Correspondingly, the feature map processed by the i-th downsampling of the semantic segmentation model is the first feature map generated when the basic semantic segmentation model converges, and the feature map output by the j-th upsampling of the semantic segmentation model is the second feature map generated when the basic semantic segmentation model converges.

[0107] In this case, the first model parameters can include the parameters used before the i-th downsampling of the semantic segmentation model, and the second model parameters can include all the parameters of the semantic segmentation model.

[0108] S420. The terminal device processes the first intermediate feature map through a prediction head to obtain a semantic segmentation result.

[0109] In the embodiment of the present application, the prediction head can convert the first intermediate feature map into a feature map with a size of C × H × W, where H × W is the resolution of the image, which is the same as the resolution of the input image, and C is the number of channels, which is the same as the number of types that the semantic segmentation model can segment. Then, the Softmax activation function is used to calculate a corresponding probability value for each category, and the type information corresponding to each pixel point is obtained through the Argmax function to complete the semantic segmentation task.

[0110] Exemplarily, the size of the input image is 3 × 384 × 384, indicating that the input image is an image with 3 color channels and a resolution of 384 × 384, and the semantic segmentation model can recognize 4 different categories; then the result of the prediction head can be 4 × 384 × 384.

[0111] Optionally, on the basis of Figure 4 shown in S410, the process of the terminal device obtaining the first intermediate feature map will be described below in combination with Figure 6 to illustrate. Figure 6 This is the flowchart of the semantic segmentation method provided by the embodiment of the present application Figure 3 . Among them, the encoding-decoding network includes a convolution module, N downsampling modules, a pooling module, and N upsampling modules.

[0112] Figure 6 Shows Figure 4 a possible implementation manner of S410 in

[0113] S411. The terminal device performs a convolution operation on the input image corresponding to the first image through a convolution module to obtain a second intermediate feature map.

[0114] Among them, the role of the convolution module is to perform a convolution operation on the input. For example, the filter in the convolutional layer can slide on the sample with a set stride. At each sliding position, the values in the filter array are multiplied by the sample data and then added to obtain a value. All the values obtained during the sliding process form a new array, and this new array is called a feature map.

[0115] The convolution operation can change the resolution and the number of channels of the image. In the embodiment of the present application, the convolution module can use the conv2d function to perform a two-dimensional convolution operation on the input image to obtain a second intermediate feature. The number of channels of the second intermediate feature is greater than that of the input image, and the resolution of the second intermediate feature is less than that of the input image.

[0116] S412. The terminal device sequentially performs downsampling processing on the second intermediate feature map through N downsampling modules to obtain a third intermediate feature map.

[0117] In the embodiment of the present application, taking the N downsampling modules including downsampling module 1, downsampling module 2,..., downsampling module N as an example, the terminal device can use downsampling module 1 to perform the first downsampling processing on the second intermediate feature map to obtain intermediate feature map m d1 ; use downsampling module 2 to perform the second downsampling processing on intermediate feature map m d1 to obtain intermediate feature map m d2 ; and so on. Use downsampling module N to perform the Nth downsampling processing on intermediate feature map m d(N-1) to obtain intermediate feature map m dN , and intermediate feature map m dN is the third intermediate feature map.

[0118] Among them, the network structures of the N downsampling modules are different to respectively extract features of the image at different resolutions.

[0119] In the embodiment of the present application, the N downsampling modules can be composed of multiple residual networks, and the network structures of the multiple residual networks are different. In the embodiment of the present application, the multiple residual networks can be obtained by permutation and combination of the first network layer, the second network layer, the third network layer, and the fourth network layer.

[0120] Taking the multiple residual networks including the first residual network, the second residual network, and the third residual network as an example, Figure 7 shows the schematic diagrams of the network structures of the first residual network, the second residual network, and the third residual network. As Figure 7As shown, the first residual network includes two first network layers, two second network layers, and a third network layer. The second residual network includes two first network layers, a second network layer, and a third network layer. The third residual network includes a first network layer, a second network layer, a third network layer, and a fourth network layer.

[0121] Among them, both the first network layer and the fourth network layer can perform convolution, normalization, and rectifying linear units (ReLU) operations on the input. The parameters used for the convolution operation in the first network layer and the convolution operation in the fourth network layer are different. The second network layer can be used to perform convolution and normalization operations on the input, and the third network layer can be used to perform ReLU operations on the input.

[0122] In the first residual network, the input x passes through two first network layers and the second network layer in sequence to obtain an intermediate quantity x1. The input x passes through the second network layer to obtain an intermediate quantity x2. Then, the intermediate quantity x1 and the intermediate quantity x2 are fused, and the fused result is input to the third network layer for processing to obtain the output y.

[0123] In the second residual network, the input x passes through two first network layers and the second network layer in sequence to obtain an intermediate quantity x1. Then, the intermediate quantity x1 and the input x are fused, and the fused result is input to the third network layer for processing to obtain the output y.

[0124] In the third residual network, the input x passes through the first network layer, the fourth network layer, and the second network layer in sequence to obtain an intermediate quantity x3. Then, the intermediate quantity x3 and the input x are fused, and the fused result is input to the third network layer for processing to obtain the output y.

[0125] Figure 8 Shows a further schematic diagram of the network structure of the first residual network, the second residual network, and the third residual network. As Figure 8 shown, the first network layer includes a first convolutional (conv) layer, a normalization layer, and a rectifying linear units (ReLU) layer. The second network layer includes a first convolutional layer and a normalization layer. The third network layer includes a ReLU layer. The fourth network layer includes a second convolutional layer, a normalization layer, and a ReLU layer.

[0126] In a possible design, the first convolutional layer can be a Conv2d layer, the second convolutional layer can be a Conv2d with Dilation layer, the normalization layer can be a BN (batch normalization) layer, and the ReLU layer can be a LeakyReLU layer.

[0127] S413. The terminal device processes the third intermediate feature map through a pooling module to obtain a fourth intermediate feature map.

[0128] Among them, the pooling module may include a pyramid pooling module, and the pyramid pooling module may perform multi-scale feature fusion using pooling of different sizes.

[0129] Exemplarily, Figure 9 A network structure diagram of a pyramid pooling module is shown. As Figure 9 shown, the pyramid pooling module includes 6 channels and a concatenation layer. Each channel includes a pooling layer, a first network layer, and an upsampling layer. The pooling layer can perform a pooling operation, the first network layer can perform a convolution operation, and the upsampling layer can perform an upsampling operation.

[0130] Figure 9 The processing flow of the shown pyramid pooling module is as follows: First, perform pooling operations of multiple different sizes on the input (for example, the third intermediate feature map) to obtain feature maps of multiple sizes. The multiple different sizes are, for example, 1×1, 2×2, 3×3, 4×4, 6×6, 12×12; then perform convolution operations on the feature maps of multiple sizes respectively to reduce the number of channels; then perform upsampling processing on the convolved feature maps respectively to obtain feature maps of multiple same sizes, and finally concatenate the feature maps of multiple same sizes on the channels to obtain the output (for example, the fourth intermediate feature map).

[0131] In a possible design, the above pooling layer may be an adaptive average pooling 2d layer, the first network layer may include a Conv2d layer, a BN layer, and a LeakyReLU layer, and the concatenation layer may include a concatenation layer.

[0132] It can be seen that the pyramid pooling module provided in the embodiments of the present application is extended from the traditional 4 channels to 6 channels, so that it can retain image details as much as possible on the premise of capturing semantic features in a larger range, which is beneficial to improving the accuracy of the semantic segmentation result.

[0133] S414. The terminal device sequentially performs upsampling processing on the fourth intermediate feature map through N upsampling modules to obtain the first intermediate feature map.

[0134] In the embodiments of the present application, taking N upsampling modules including upsampling module 1, upsampling module 2,..., upsampling module N as an example, the terminal device may use upsampling module 1 to perform the first upsampling processing on the fourth intermediate feature map, and then fuse the result of the first upsampling processing and the intermediate feature map m d(N-1) to obtain the intermediate feature map m uN ; use upsampling module 2 to perform the second upsampling processing on the intermediate feature map m uN and then fuse the result of the second upsampling processing and the intermediate feature map m d(N-2)Obtain the intermediate feature map m uN-1 ; and so on. Use the upsampling module N to upsample the intermediate feature map m u2 for the Nth upsampling process, and then fuse the fourth intermediate feature map to obtain the intermediate feature map m u1 . The intermediate feature map m u1 is the first intermediate feature map.

[0135] In a possible design, the N upsampling modules may include an inception module. Figure 10 The network structure diagram of the inception module is shown. As Figure 10 shown, the inception module includes two second network layers and a splicing layer. After the input x is processed by the two second network layers respectively, a splicing operation is performed through the splicing layer to obtain the output y. Among them, the convolutional kernels of the two second network layers are 1×9 and 9×1 respectively. The second network layer may include a Conv layer and a BN layer, and the splicing layer may include a concatenation layer.

[0136] The network structure of the inception module provided by the embodiments of the present application is simple, and the number of parameters of the semantic segmentation model can be reduced.

[0137] It should be noted that the semantic segmentation model provided by the embodiments of the present application may not only include the modules described above, but may also include other network layers or algorithm modules, which are not limited in this application.

[0138] Next, taking N = 3 as an example, the network structure of a semantic segmentation model provided by the embodiments of the present application will be described in detail. Figure 11 The network structure diagram of a semantic segmentation model provided by the embodiments of the present application is shown. As Figure 11 shown, the semantic segmentation model includes a convolutional module, a residual network module, a downsampling module 1, a downsampling module 2, a downsampling module 3, a pooling module, an upsampling module 1, an upsampling module 2, an upsampling module 3, and a prediction head. Among them, the convolutional module includes 3 network layers composed of a Conv2d layer and a BN2d layer. The residual network module includes the first residual network and the second residual network described above. The downsampling module 1 includes a MaxPool2d layer, the first residual network, and the second residual network. The downsampling module 2 and the downsampling module 3 include a MaxPool2d layer, the first residual network, the second residual network, and the third residual network. The pooling module includes a network layer composed of a Conv2d layer, a BN2d layer, and a LeakyReLU layer and a pyramid pooling module. The downsampling module 1, the downsampling module 2, and the downsampling module 3 all include an inception module; the prediction head includes a network layer composed of a Conv2d layer and a BN2d layer, a Conv2d layer, and an upsampling layer.

[0139] Among them, an input image with a size of 3×384×384 is input into the convolutional module, and a feature map FM1 with a size of 64×96×96 can be obtained. The feature map FM1 is input into the residual network module to obtain a feature map FM2 with a size of 128×96×96. The feature map FM2 is input into the downsampling module 1 to obtain a feature map FM3 with a size of 256×48×48. The feature map FM3 is input into the downsampling module 2 to obtain a feature map FM4 with a size of 512×24×24. The feature map FM4 is input into the downsampling module 3 to obtain a feature map FM5 with a size of 1024×12×12. The feature map FM5 is input into the pooling module to obtain a feature map FM6 with a size of 128×12×12. The feature map FM6 is input into the upsampling module 1, and then fused with the feature map FM7 with a size of 128×24×24 obtained by performing Conv2d and BN2d processing on the feature map FM4 and the result of the upsampling module 1 to obtain a feature map FM8 with a size of 128×24×24. The feature map FM8 is input into the upsampling module 2, and then fused with the feature map FM9 with a size of 128×48×48 obtained by performing Conv2d and BN2d processing on the feature map FM3 and the result of the upsampling module 2 to obtain a feature map FM10 with a size of 128×48×48. The feature map FM10 is input into the upsampling module 3, and then fused with the feature map FM11 with a size of 128×96×96 obtained by performing Conv2d and BN2d processing on the feature map FM2 and the result of the upsampling module 3 to obtain a feature map FM12 with a size of 128×96×96. Finally, the feature map FM12 is input into the prediction head to obtain a semantic segmentation result with a size of 4×384×384.

[0140] It can be seen that the embodiment of the present application can use residual modules with three different network structures to extract features of an image at different resolutions during the encoding process, and can more reasonably extract the details and semantic information of the image at different stages of feature encoding, capturing a larger range of semantic information. Then, during the decoding process, the pyramid pooling module and the Inception module are used to achieve the fusion of image features at different scales, and the feature map is gradually restored to the original input resolution size through the upsampling method based on bilinear interpolation. Among them, the present invention expands the classic pyramid pooling module from 4 layers to 6 layers, so as to retain the detail information of the image as much as possible on the premise of capturing a larger range of semantic features. At the same time, the simple design of the Inception module effectively reduces the number of model parameters. Finally, the model uses the Softmax activation function to calculate a corresponding probability value for each category, and obtains the predicted category corresponding to each pixel point through the Argmax function to complete the semantic segmentation task.

[0141] The embodiment of the present application also provides a training method for a semantic segmentation model, which is used to distill richer image features in the decoding stage to the encoding stage, improve the feature extraction ability of the encoder, and thus improve the performance of the semantic segmentation model.

[0142] As Figure 12 shown, it is a schematic flowchart of the training method for the semantic segmentation model provided by the embodiment of the present application. The training method of this semantic segmentation model can be executed by the above-mentioned server, or by the above-mentioned terminal device, or by other devices with computing capabilities. Here, taking the terminal device executing the training method for the semantic segmentation model provided in this embodiment as an example for illustration. As Figure 12 shown, the training method for the semantic segmentation model provided by the embodiment of the present application includes steps S710 to S740.

[0143] S710, the terminal device iteratively trains the basic semantic segmentation model based on training samples.

[0144] Among them, the training samples include training images with labels, and the labels are used to indicate the true semantic segmentation results of the corresponding training images. During the iterative training process, the terminal device can input the training images with labels into the basic semantic segmentation model to obtain the predicted semantic segmentation results of the training images.

[0145] It should be noted that the network structure of the basic semantic segmentation model is similar to the network structure of the semantic segmentation model described above, and will not be elaborated here.

[0146] S720, the terminal device determines whether the basic semantic segmentation model converges.

[0147] If the basic semantic segmentation model does not converge, then execute S730; if the basic semantic segmentation model converges, then execute S740.

[0148] For example, the terminal device can determine whether the basic semantic segmentation model converges according to the first loss function and the second loss function.

[0149] Among them, the first loss function can be determined according to the first feature map and the second feature map. The first feature map is the feature map generated by the basic semantic segmentation model during the encoding process, and the second feature map is the feature map generated by the basic semantic segmentation model during the decoding process, and the first feature map and the second feature map have the same resolution.

[0150] The terminal device can perform convolution processing on the first feature map to map the first feature map to the dimension where the second feature map is located, and then calculate the mean square error between the second feature map and the first feature map after convolution processing. This mean square error is the above-mentioned first loss function.

[0151] Understandably, the first loss function is used to reflect the difference between the first feature map and the second feature map. The smaller the first loss function is, the closer the first feature map is to the second feature map.

[0152] The second loss function can be determined according to the label of the training image and the predicted semantic segmentation result of the training image. Among them, the smaller the second loss function is, the closer the predicted semantic segmentation result is to the label. In other words, the better the effect of the basic semantic segmentation model is.

[0153] S730, the terminal device updates the model parameters of the basic semantic segmentation model based on the first loss function and the second loss function.

[0154] The basic semantic segmentation model includes first model parameters and second model parameters. Among them, the first model parameters include the model parameters used in the process of generating the first feature map, and the second model parameters include all the model parameters of the basic semantic segmentation model, that is, the second model parameters include the first model parameters. Among them, the terminal device can adjust the first model parameters based on the first loss function and the second loss function, and adjust the first model parameters based on the second loss function, so as to avoid the problem of identity mapping between the first feature map and the second feature map of the basic semantic segmentation model.

[0155] Exemplarily, as Figure 13 shown, taking the network structure of the basic semantic segmentation model as the Figure 11 shown network structure as an example, the first feature map can be the feature map FM2 output by the residual network module, and the second feature map can be the feature map FM12. Among them, the terminal device can perform a Conv2d operation on the feature map FM2, and then calculate the mean square error between the feature map FM12 and the feature map FM2 after the Conv2d operation, and use this mean square error as the above-mentioned first loss function. Among them, this first loss function can be used to adjust the parameters of the convolutional module and the residual network module, that is, the first model parameters include the parameters of the convolutional module and the residual network module.

[0156] The second loss function is determined according to the label of the input image and the predicted semantic segmentation result output by the basic semantic segmentation model. This second loss function can be used to update all the model parameters of the basic semantic segmentation model.

[0157] S740, the terminal device stops iterative training and obtains a semantic segmentation model.

[0158] It can be seen that during the process of training the basic semantic segmentation model, the first loss function and the second loss function can be obtained in one training process. Then, the first model parameters are adjusted according to the first loss function and the second loss function, and the second model parameters are adjusted according to the second loss function. When the first loss function and the second loss function converge, the determined first model parameters and second model parameters are used as the model parameters of the initial semantic segmentation model to obtain the semantic segmentation model.

[0159] It can be understood that the training method of the semantic segmentation model provided by the embodiments of the present application satisfies three conditions:

[0160] First, the basic semantic segmentation model is iteratively trained using training images with labels, that is, the basic semantic segmentation model is optimized by means of supervised learning.

[0161] Second, the local network of the basic semantic segmentation model (that is, the network between the first feature map and the second feature map) satisfies the local Markov property.

[0162] Among them, when a stochastic process, given the present state and all past states, the conditional probability distribution of its future state depends only on the current state; in other words, when the present state is given, it is conditionally independent of the past states (that is, the historical path of the process), then this stochastic process has the Markov property. The local Markov property means that given the adjacent variables of a certain variable, then this variable is conditionally independent of other variables. Then, since the first feature map is the only input feature of the local network of the basic semantic segmentation model, and the second feature map is the only output feature of the local network of the basic semantic segmentation model, the local network of the basic semantic segmentation model satisfies the local Markov property.

[0163] Third, the feature map generated during the decoding process not only obtains the detailed features extracted during the encoding process by the bridging network, but also captures rich semantic information from deeper networks. Therefore, the feature map generated during the decoding process has richer features than the feature map generated during the encoding process and can better meet the needs of label prediction during the supervised learning process.

[0164] In other words, the training method of the semantic segmentation model provided by the embodiments of the present application satisfies the condition that "in the paradigm of supervised training, the information in the output features of the local Markov process is richer than the input features".

[0165] Therefore, in the embodiments of the present application, the first feature map generated during the encoding process is used as the student feature, and the second feature map generated during the decoding process is used as the teacher feature. The features of the decoding process can be distilled into the encoding process to achieve a spiral enhancement of the feature representation in the encoding and decoding processes, thereby improving the performance of the semantic segmentation model. In addition, the experimental results show that within a certain range, the more complex the semantic segmentation model is, the more obvious the improvement in the segmentation performance brought by this training method is.

[0166] The embodiments of the present application further provide a chip system, which includes at least one processor and at least one interface circuit. The processor and the interface circuit can be interconnected by a line. For example, the interface circuit can be used to receive signals from other devices (such as the memory of an electronic device). For another example, the interface circuit can be used to send signals to other devices (such as the processor). Exemplarily, the interface circuit can read the instructions stored in the memory and send the instructions to the processor. When the instructions are executed by the processor, the electronic device can execute each step in the above embodiments. Of course, the chip system can also include other discrete devices, and the embodiments of the present application do not make specific limitations on this.

[0167] The embodiments of the present application further provide a computer storage medium, which includes computer instructions. When the computer instructions run on the above electronic device, the electronic device can execute each function or step that the mobile phone executes in the above method embodiments.

[0168] The embodiments of the present application further provide a computer program product. When the computer program product runs on a computer, the computer can execute each function or step that the mobile phone executes in the above method embodiments.

[0169] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0170] In the several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.

[0171] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may be a single physical unit or multiple physical units, that is, it may be located in one place or distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0172] In addition, each functional unit in various embodiments of the present application may be integrated in a processing unit, may exist as individual physical units, or two or more units may be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0173] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions for causing a device (which may be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: USB flash drives, mobile hard disks, read only memory (ROM), random access memory (RAM), magnetic disks, or optical discs and other various media that can store program codes.

[0174] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A semantic segmentation method, characterized in that, The method includes: Obtaining a first image; Inputting the first image into a semantic segmentation model to obtain a semantic segmentation result; Wherein, the semantic segmentation result is used to indicate the class information of each pixel in the first image, the semantic segmentation model includes first model parameters, the first model parameters are associated with a first loss function, the first loss function is determined according to a first feature map and a second feature map, the first feature map is a feature map generated during the encoding process of a basic semantic segmentation model used to generate the semantic segmentation model, the second feature map is a feature map generated during the decoding process of the basic semantic segmentation model, and the first feature map and the second feature map have the same resolution.

2. The method according to claim 1, wherein The semantic segmentation model is obtained by training the basic semantic segmentation model based on training samples, wherein the training process includes: Iteratively training the basic semantic segmentation model based on the training samples; If the basic semantic segmentation model does not converge, adjusting the model parameters of the basic semantic segmentation model based on the first loss function; If the basic semantic segmentation model converges, stopping the iterative training to obtain the semantic segmentation model.

3. The method according to claim 2, characterized in that, The basic semantic segmentation model is used to perform N times of downsampling and N times of upsampling on the training samples, the first feature map is the feature map processed by the basic semantic segmentation model at the i-th downsampling, the second feature map is the feature map obtained by the basic semantic segmentation model at the j-th upsampling, i + j - 1 = N, and i ≤ N.

4. The method according to any one of claims 1 to 3, characterized in that The semantic segmentation model includes an encoder-decoder network and a prediction head; The inputting the first image into the semantic segmentation model to obtain a semantic segmentation result includes: Successively performing N times of downsampling and N times of upsampling on the first image through the encoder-decoder network to obtain a first intermediate feature map; Processing the first intermediate feature map through the prediction head to obtain the semantic segmentation result.

5. The method according to claim 4, wherein The encoder-decoder network includes a convolution module, N downsampling modules, a pooling module, and N upsampling modules; The successively performing N times of downsampling and N times of upsampling on the input image corresponding to the first image through the encoder-decoder network to obtain a first intermediate feature map includes: Performing a convolution operation on the first image through the convolution module to obtain a second intermediate feature map; Successively performing downsampling processing on the second intermediate feature map through the N downsampling modules to obtain a third intermediate feature map; Processing the third intermediate feature map through the pooling module to obtain a fourth intermediate feature map; Successively performing upsampling processing on the fourth intermediate feature map through the N upsampling modules to obtain the first intermediate feature map.

6. The method according to claim 5, wherein The N downsampling modules include a first downsampling module, a second downsampling module, and a third downsampling module. The first downsampling module includes a first residual network, the second downsampling module includes the first residual network and a second residual network, the third downsampling module includes the first residual network, the second residual network, and a third residual network, and the network structures of the first residual network, the second residual network, and the third residual network are different.

7. The method according to claim 5, wherein The pooling module includes a pyramid pooling module, and the pyramid pooling module includes six pooling layers.

8. The method according to claim 5, characterized in that, Each upsampling module includes an initial module, and the initial module includes two convolutional layers and a splicing layer. The two convolutional layers are used to perform convolutional processing on the received input and send the processing results to the splicing layer, and the splicing layer is used to splice the processing results to obtain an output.

9. The method according to any one of claims 5-8, characterized in that, The feature map processed by the i-th downsampling of the encoding-decoding network is the first feature map generated when the basic semantic segmentation model converges, and the feature map output by the j-th upsampling of the encoding-decoding network is the second feature map generated when the basic semantic segmentation model converges, where i + j - 1 = N and i ≤ N.

10. The method according to claim 9, wherein The first model parameters include all the parameters used by the encoding-decoding network before the i-th downsampling.

11. The method according to claim 10, characterized in that, The first model parameters at least include the parameters of the convolutional module.

12. The method according to any one of claims 9-11, characterized in that, N = 3 and i = 1.

13. The method according to any one of claims 1-12, characterized in that The first model parameters are also associated with a second loss function, and the second loss function is determined according to the predicted semantic segmentation result of the training image and the label of the training image, and the predicted semantic segmentation result is obtained by inputting the training image into the basic semantic segmentation model.

14. The method according to claim 13, wherein The semantic segmentation model further includes second model parameters, and the second model parameters include all the parameters of the semantic segmentation model, and the second model parameters are determined according to the second loss function.

15. A method for training a semantic segmentation model, characterized in that, The method includes: Performing iterative training on the basic semantic segmentation model based on training samples; If the basic semantic segmentation model does not converge, adjusting the model parameters of the basic semantic segmentation model based on a first loss function; wherein, the first loss function is determined according to a first feature map and a second feature map, the first feature map is the feature map generated during the encoding process of the basic semantic segmentation model, the second feature map is the feature map generated during the decoding process of the basic semantic segmentation model, and the first feature map and the second feature map have the same resolution; If the basic semantic segmentation model converges, stopping the iterative training to obtain a semantic segmentation model.

16. An electronic device, characterized in that, The electronic device includes: a memory and one or more processors; the memory is coupled to the processor; the memory is used to store computer program code, and the computer program code includes computer instructions, and when the computer instructions are executed by the electronic device, the electronic device executes the method according to any one of claims 1-15.

17. A computer-readable storage medium, characterized in that, Computer instructions are stored in the computer-readable storage medium, and when the computer instructions run in the electronic device, the electronic device executes the method according to any one of claims 1-15.

Citation Information

Patent Citations

  • Image semantic segmentation method, system and equipment based on full-scale dense connection

    CN115601542A

  • Image coding and decoding method and device, computer equipment and storage medium

    CN115690241A

  • Network training method and device, electronic equipment and storage medium

    CN116050498A

  • Semi-supervised segmentation model construction and image analysis method, device and system

    CN116051574A

  • Image processing method based on semantic information and electronic equipment

    CN116206100A

Cited By

  • Interface intelligent preview method and system based on AI and MJPEG

    CN122526691A