Semantic segmentation method, training method, device and equipment for semantic segmentation model
By introducing a feature extraction model with a fusion attention mechanism into the semantic segmentation method, the problem of difficult to deal with special distribution images in spatial locations in the prior art is solved, and the segmentation performance of semantic segmentation is significantly improved.
Patent Information
- Application Number
- CN202210386561.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-11
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-04-11
AI Technical Summary
In the prior art, the semantic segmentation method is based only on image pixel features, and it is difficult to effectively process images with special distributions in spatial positions, resulting in the need to improve the segmentation performance.
A semantic segmentation method is adopted, combining semantic classification model and feature extraction model that integrates attention mechanism, and the structural features of the spatial position distribution of the original image are obtained through the attention mechanism, a target feature map is generated, and a semantic classification model is input to obtain the semantic segmentation results.
Through the fusion attention mechanism, the fusion of context information related to the spatial location distribution of the original image in the target feature map is strengthened, and the performance of semantic segmentation results is improved.
Smart Images

Figure CN114820633B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and in particular, to a semantic segmentation method, a training method of a semantic segmentation model, a device, and a device. Background Art
[0002] Semantic segmentation is one of the important research directions in the field of computer vision technology and is widely used in various industries, such as autonomous driving, satellite remote sensing, medical treatment, etc. Semantic segmentation aims to segment different objects in an image based on image pixels.
[0003] Generally, the semantic segmentation methods in the related art only perform segmentation based on the pixel features of an image, and for images with a special distribution in spatial positions, the segmentation performance of the semantic segmentation methods in the related art needs to be improved. Summary of the Invention
[0004] This application provides a semantic segmentation method, a training method of a semantic segmentation model, a device, and a device, which can solve the problems in the related art.
[0005] In a first aspect, a semantic segmentation method is provided. The method includes:
[0006] Obtain an original image and a semantic segmentation model. The semantic segmentation model includes a semantic classification model and a feature extraction model with a fusion attention mechanism. The attention mechanism is used to obtain the structural features of the original image based on the spatial position distribution of the original image;
[0007] Call the feature extraction model with the fusion attention mechanism to perform feature extraction on the original image to obtain a target feature map that fuses the structural features of the original image;
[0008] Input the target feature map into the semantic classification model to obtain the semantic categories of each pixel point in the original image output by the semantic classification model;
[0009] Based on the semantic categories of each pixel point in the original image, obtain the semantic segmentation result of the original image.
[0010] In a possible implementation manner, the feature extraction model with the fusion attention mechanism includes a feature extraction module, an attention module, and a fusion module;
[0011] The step of calling the feature extraction model with the fusion attention mechanism to perform feature extraction on the original image to obtain a target feature map that fuses the structural features of the original image includes:
[0012] Call the feature extraction module to extract multiple feature scale maps corresponding to the original image at different resolutions;
[0013] Call the attention module to obtain the channel attention maps corresponding to the multiple feature scale maps respectively. The channel attention maps are used to indicate the weight values of each row in the first spatial dimension of the feature scale map corresponding to each feature channel, and the first spatial dimension is determined by the spatial position distribution of the original image;
[0014] Call the fusion module to fuse the multiple feature scale maps and the channel attention maps corresponding to the multiple feature scale maps respectively to obtain the target feature map.
[0015] In a possible implementation manner, the attention module includes a pooling layer, a downsampling layer, a convolutional layer, and an upsampling layer;
[0016] The calling the attention module to obtain the channel attention maps corresponding to the multiple feature scale maps respectively includes:
[0017] For any one of the multiple feature scale maps, call the pooling layer to pool and compress the vectors in the second spatial dimension of the any one feature scale map to obtain the first intermediate feature in the first spatial dimension, and the second spatial dimension is the spatial dimension other than the first spatial dimension in the any one feature scale map;
[0018] Call the downsampling layer to perform downsampling processing on the first intermediate feature to obtain a refined second intermediate feature;
[0019] Call the convolutional layer to extract the context information between the vectors in the first spatial dimension of the second intermediate feature to obtain a third intermediate feature;
[0020] Call the upsampling layer to perform upsampling processing on the third intermediate feature to obtain the channel attention map corresponding to the any one feature scale map, and the size of the channel attention map is the same as the size of the any one feature scale map.
[0021] In a possible implementation manner, the calling the fusion module to fuse the multiple feature scale maps and the channel attention maps corresponding to the multiple feature scale maps respectively to obtain the target feature map includes:
[0022] Call the fusion module to fuse each of the multiple feature scale maps with the channel attention map corresponding to each of the multiple feature scale maps to obtain the fusion result corresponding to each of the multiple feature scale maps;
[0023] Cascade the fusion results corresponding to each of the multiple feature scale maps to obtain the target feature map.
[0024] In a possible implementation, the step of calling the fusion module to fuse each feature scale map in the multiple feature scale maps with the channel attention map corresponding to each feature scale map to obtain the fusion result corresponding to each feature scale map includes:
[0025] Call the fusion module to multiply each feature scale map in the multiple feature scale maps by the channel attention map corresponding to each feature scale map to obtain the product result corresponding to each feature scale map;
[0026] Add each feature scale map to the product result corresponding to each feature scale map to obtain the fusion result corresponding to each feature scale map.
[0027] In a second aspect, a training method for a semantic segmentation model is provided, and the method includes:
[0028] Obtain a sample image, the semantic label of the sample image, and an initial semantic segmentation model. The initial semantic segmentation model includes an initial semantic classification model and an initial feature extraction model with a fusion attention mechanism. The attention mechanism is used to obtain the structural features of the sample image based on the spatial position distribution of the sample image;
[0029] Call the initial feature extraction model with the fusion attention mechanism to perform feature extraction on the sample image to obtain a target feature map that fuses the structural features of the sample image;
[0030] Input the target feature map into the initial semantic classification model, and based on the output of the initial semantic classification model, obtain the semantic category of each pixel point in the sample image;
[0031] Train the initial semantic segmentation model based on the semantic category of each pixel point in the sample image and the semantic label of the sample image to obtain a semantic segmentation model.
[0032] In a possible implementation, the step of training the initial semantic segmentation model based on the semantic category of each pixel point in the sample image and the semantic label of the sample image includes:
[0033] Obtain a loss function value based on the semantic category of each pixel point in the sample image and the semantic label of the sample image;
[0034] Iteratively adjust the parameters of the initial semantic classification model and the initial feature extraction model with the fusion attention mechanism in the initial semantic segmentation model according to the loss function value until the convergence condition is met.
[0035] In a third aspect, a semantic segmentation device is provided, and the device includes:
[0036] A first acquisition module, configured to acquire an original image and a semantic segmentation model, where the semantic segmentation model includes a semantic classification model and a feature extraction model integrated with an attention mechanism, and the attention mechanism is configured to acquire the structural features of the original image based on the spatial position distribution of the original image;
[0037] A feature extraction module, configured to call the feature extraction model integrated with the attention mechanism to perform feature extraction on the original image, so as to obtain a target feature map integrating the structural features of the original image;
[0038] A classification module, configured to input the target feature map into the semantic classification model, so as to obtain the semantic categories of each pixel point in the original image output by the semantic classification model;
[0039] A second acquisition module, configured to acquire the semantic segmentation result of the original image based on the semantic categories of each pixel point in the original image.
[0040] In a possible implementation manner, the feature extraction model integrated with the attention mechanism includes a feature extraction module, an attention module, and a fusion module; the feature extraction module is configured to call the feature extraction module to extract a plurality of feature scale maps corresponding to the original image at different resolutions; call the attention module to acquire the channel attention maps corresponding to the plurality of feature scale maps respectively, where the channel attention map is used to indicate the weight value of each row in the first spatial dimension of the feature scale map corresponding to each feature channel, and the first spatial dimension is determined by the spatial position distribution of the original image; call the fusion module to fuse the plurality of feature scale maps and the channel attention maps corresponding to the plurality of feature scale maps respectively, so as to obtain the target feature map.
[0041] In a possible implementation manner, the attention module includes a pooling layer, a downsampling layer, a convolutional layer, and an upsampling layer; the feature extraction module is configured to, for any one of the plurality of feature scale maps, call the pooling layer to perform pooling compression on the vector in the second spatial dimension of the any one of the feature scale maps, so as to obtain a first intermediate feature in the first spatial dimension, where the second spatial dimension is the spatial dimension other than the first spatial dimension in the any one of the feature scale maps; call the downsampling layer to perform downsampling processing on the first intermediate feature, so as to obtain a refined second intermediate feature; call the convolutional layer to extract the context information between the vectors in the first spatial dimension of the second intermediate feature, so as to obtain a third intermediate feature; call the upsampling layer to perform upsampling processing on the third intermediate feature, so as to obtain the channel attention map corresponding to the any one of the feature scale maps, and the size of the channel attention map is the same as the size of the any one of the feature scale maps.
[0042] In a possible implementation, the feature extraction module is configured to call the fusion module to fuse each feature scale map in the multiple feature scale maps with the channel attention map corresponding to each feature scale map to obtain a fusion result corresponding to each feature scale map; concatenate the fusion results corresponding to each feature scale map to obtain the target feature map.
[0043] In a possible implementation, the feature extraction module is configured to call the fusion module to multiply each feature scale map in the multiple feature scale maps by the channel attention map corresponding to each feature scale map to obtain a product result corresponding to each feature scale map; add each feature scale map to the product result corresponding to each feature scale map to obtain a fusion result corresponding to each feature scale map.
[0044] In a fourth aspect, a training device for a semantic segmentation model is provided, and the device includes:
[0045] A first acquisition module, configured to acquire a sample image, a semantic label of the sample image, and an initial semantic segmentation model, where the initial semantic segmentation model includes an initial semantic classification model and an initial feature extraction model integrating an attention mechanism, and the attention mechanism is configured to obtain a structural feature of the sample image based on a spatial position distribution of the sample image;
[0046] The feature extraction module is configured to call the initial feature extraction model integrating the attention mechanism to perform feature extraction on the sample image to obtain a target feature map integrating the structural feature of the sample image;
[0047] The classification module is configured to input the target feature map into the initial semantic classification model and output semantic categories of each pixel point in the sample image based on an output of the initial semantic classification model;
[0048] The training module is configured to train the initial semantic segmentation model based on the semantic categories of each pixel point in the sample image and the semantic label of the sample image to obtain a semantic segmentation model.
[0049] In a possible implementation, the training module is configured to obtain a loss function value based on the semantic categories of each pixel point in the sample image and the semantic label of the sample image; iteratively adjust parameters of the initial semantic classification model and the initial feature extraction model integrating the attention mechanism in the initial semantic segmentation model according to the loss function value until a convergence condition is satisfied.
[0050] Fifth aspect, there is also provided a computer device, which includes a processor and a memory. At least one program code is stored in the memory, and the at least one program code is loaded and executed by the processor to enable the computer device to implement the semantic segmentation method described in any one of the above, or to enable the computer device to implement the training method of the semantic segmentation model described in any one of the above.
[0051] Sixth aspect, there is also provided a computer-readable storage medium, in which at least one program code is stored, and the at least one program code is loaded and executed by a processor to enable a computer to implement the semantic segmentation method described in any one of the above, or to enable a computer device to implement the training method of the semantic segmentation model described in any one of the above.
[0052] Seventh aspect, there is also provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to enable the computer device to execute the semantic segmentation method described in any one of the above, or to enable the computer device to execute the training method of the semantic segmentation model described in any one of the above.
[0053] The technical solution provided by this application can at least bring the following beneficial effects:
[0054] The technical solution provided by this application uses an attention mechanism to obtain a target feature map that fuses the structural features of the spatial position distribution of the original image, so that the obtained target feature map strengthens the fusion of context information related to the spatial position distribution of the original image, and further improves the performance of the semantic segmentation result obtained according to the target feature map. Description of the Drawings
[0055] To more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0056] Figure 1 is a schematic diagram of the implementation environment of a semantic segmentation method provided by an embodiment of this application;
[0057] Figure 2 is a flowchart of a semantic segmentation method provided by an embodiment of this application;
[0058] Figure 3It is a schematic diagram of an urban street view image provided by an embodiment of the present application;
[0059] Figure 4 It is a schematic diagram of a high-attention network provided by an embodiment of the present application;
[0060] Figure 5 It is a schematic diagram of a semantic segmentation model provided by an embodiment of the present application;
[0061] Figure 6 It is a flowchart of a training method for a semantic segmentation model provided by an embodiment of the present application;
[0062] Figure 7 It is a comparison diagram of segmented images provided by an embodiment of the present application;
[0063] Figure 8 It is a schematic diagram of a semantic segmentation device provided by an embodiment of the present application;
[0064] Figure 9 It is a schematic diagram of a training device for a semantic segmentation model provided by an embodiment of the present application;
[0065] Figure 10 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application;
[0066] Figure 11 It is a schematic diagram of the structure of a server provided by an embodiment of the present application. Detailed implementation manners
[0067] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0068] Semantic segmentation refers to identifying an image at the pixel level and predicting the class label of each pixel in the image, that is, labeling the object category to which each pixel in the image belongs. An embodiment of the present application provides a semantic segmentation method that can take into account the special distribution of the image in terms of spatial position and improve the segmentation performance of semantic segmentation.
[0069] See Figure 1 , Figure 1 shows a schematic diagram of the implementation environment of the semantic segmentation method provided by an embodiment of the present application. The implementation environment includes: a computer device 101. As Figure 1As shown, the computer device 101 can refer to a terminal or a server. Exemplarily, the terminal can be any electronic product that can perform human-computer interaction with the user in one or more ways such as a keyboard, touchpad, touch screen, remote control, voice interaction, or handwriting device. For example, a personal computer (PC), smartphone, personal digital assistant (PDA), wearable device, pocket PC (PPC), tablet computer, intelligent vehicle console, etc. The server can be an independent physical server or a server cluster or distributed system composed of multiple physical servers.
[0070] When the semantic segmentation method is applied to a server, the server can obtain the original image from an image database, or the server can obtain the original image from a connected terminal and perform semantic segmentation on the original image through the semantic segmentation method for further analysis operations. When the image semantic segmentation method is applied to a terminal, an image processing software can be installed in the terminal, and the semantic segmentation method can be deployed in the image processing software; the terminal can perform semantic segmentation on the captured image or the pre-stored image through the semantic segmentation method for further image processing.
[0071] Those skilled in the art should understand that the above computer device 101 is only an example, and other existing or future possible computer devices can also be applicable to this application and should also be included within the protection scope of this application and are hereby incorporated by reference.
[0072] Based on the above Figure 1 shown implementation environment, the semantic segmentation method provided by the embodiments of this application can be applied to the computer device 101. The computer device 101 can be a terminal or a server, and the embodiments of this application do not limit this. As Figure 2 shown, the semantic segmentation method provided by the embodiments of this application includes the following steps 201 to step 204.
[0073] Step 201, obtain an original image and a semantic segmentation model. The semantic segmentation model includes a semantic classification model and a feature extraction model integrating an attention mechanism. The attention mechanism is used to obtain the structural features of the original image based on the spatial position distribution of the original image.
[0074] In the embodiments of this application, the original image refers to the image that needs to be processed to obtain the semantic segmentation result. Optionally, the original image is an image with a special distribution in the spatial position. For example, the original image is an image with distribution features in the longitudinal position or the horizontal position.
[0075] Exemplarily, an image having a distribution feature in the longitudinal position may be an urban street scene image captured by a vehicle front camera. Refer to Figure 3 , Figure 3 FIG. 5 is a schematic diagram of an urban street scene image provided by an embodiment of the present application. The urban street scene image may be divided into upper, middle, and lower segments. Among them, the upper segment mainly includes objects such as the sky and street lights, the middle segment mainly includes objects such as trees and vehicles, and the lower segment mainly includes objects such as roads. Since objects such as cars and roads do not appear in the sky, the pixels in the upper segment area of the urban street scene image generally do not belong to categories such as cars and roads, that is, there are significant differences in the distribution of the urban street scene image in the longitudinal position.
[0076] In a possible implementation manner, the embodiment of the present application does not limit the acquisition method of the original image. In an exemplary embodiment, the ways for the computer device to acquire the original image include but are not limited to: the computer device acquires the original image from an image database; or, an image acquisition device having a communication connection with the computer device sends the acquired original image to the computer device; or, the computer device acquires the original image uploaded manually, etc. Optionally, the original image may refer to the acquired image of the image acquisition device, or may be an image obtained after preprocessing the acquired image of the image acquisition device. The embodiment of the present application does not limit this. Exemplarily, the ways for preprocessing the acquired image include but are not limited to at least one of cropping and data augmentation.
[0077] It should be noted that the number of original images is one or more. For the case where the number of original images is multiple, each original image can obtain a semantic segmentation result according to the manner of steps 202 to 204. The embodiment of the present application takes the number of original images as one as an example for illustration.
[0078] In the embodiment of the present application, the semantic segmentation model is used to obtain the semantic categories of each pixel point in the original image. The semantic segmentation model is obtained by training an initial semantic segmentation model. The way to train the semantic segmentation model can be referred to the embodiment shown in Figure 6 , which will not be elaborated here for the time being. The semantic segmentation model in the embodiment of the present application may refer to a semantic segmentation model trained by a real-time training method, or may refer to extracting a pre-trained and stored semantic segmentation model. The embodiment of the present application does not limit this.
[0079] Optionally, the semantic segmentation model includes a semantic classification model and a feature extraction model integrating an attention mechanism. Among them, the feature extraction model integrating an attention mechanism is used to obtain a target feature map integrating the structural features of the spatial position distribution of the original image based on the original image, and the semantic classification model is used to obtain the semantic categories of each pixel point in the original image based on the target feature map.
[0080] After obtaining the original image and the semantic segmentation model, the subsequent steps 202 to 204 are performed to obtain the semantic segmentation result of the original image.
[0081] Step 202: Invoke the feature extraction model with a fusion attention mechanism to extract features from the original image, obtaining a target feature map that fuses the structural features of the original image.
[0082] In an exemplary embodiment of the present application, the feature extraction model with a fusion attention mechanism includes a feature extraction module, an attention module, and a fusion module. Optionally, the feature extraction module is used to extract the feature scale map of the original image; the attention module is used to obtain the channel attention map corresponding to the feature scale map according to the attention mechanism, and the channel attention map is used to indicate the weight value of each row corresponding to each feature channel in the specified spatial dimension of the feature scale map, that is, the channel attention map can represent the structural features of the original image; the fusion module is used to fuse the feature scale map and the corresponding channel attention map.
[0083] It can be understood that during the process of fusing the feature scale map and the corresponding channel attention map, the activation degree of the feature channels of the feature scale map is adjusted by the weight values in the channel attention map, and the weight values in the channel attention map are obtained based on the spatial position distribution structure of the original image, and the feature scale map is obtained based on the pixel features of the original image. Therefore, the obtained target feature map fuses the structural features and pixel features of the original image. Thus, the obtained target feature map is more in line with the spatial position distribution structure of the original image, and further makes the semantic segmentation result obtained according to the target feature map more accurate.
[0084] In a possible implementation manner, invoking the feature extraction model with a fusion attention mechanism to extract features from the original image and obtaining a target feature map that fuses the structural features of the original image includes, but is not limited to, the following steps 2021 to 2023.
[0085] Step 2021: Invoke the feature extraction module to extract multiple feature scale maps corresponding to the original image at different resolutions.
[0086] The embodiments of the present application do not limit the network model adopted by the feature extraction module. For example, the network model adopted by the feature extraction module can be a network model that restores the high resolution through a symmetric low-to-high process, a network model that generates a high resolution representation by using transposed convolution layers, or a High Resolution Network (HRNET) model, etc. No matter which network model is adopted, the feature scale maps corresponding to the original image at different resolutions can be obtained.
[0087] In an exemplary embodiment, taking the HRNET model adopted by the feature extraction module as an example, HRNET is a computational model for obtaining image feature information, and can maintain a high resolution representation throughout the entire operation process. HRNET starts with a group of high resolution convolutions, then gradually adds low resolution convolution branches and connects them in parallel. This network combines the feature scale maps of different resolutions in parallel, with each resolution having its own path, and continuously exchanges information through multi-resolution fusion among the parallel operation combinations during the whole process. Therefore, calling HRNET can extract multiple feature scale maps corresponding to the original image at different resolutions.
[0088] Step 2022, call the attention module to obtain the channel attention maps corresponding to the multiple feature scale maps respectively.
[0089] In a possible implementation manner, different attention modules correspond to the original images with different spatial position distributions. Exemplarily, if the original image has a distribution feature in the vertical position, there will be significant differences in the feature information between each horizontal segment of the original image. In this case, the attention mechanism is used to extract the structural features representing the vertical position relationship between the features of each horizontal segment, that is, the weight values of the feature channels corresponding to each vertical segment are obtained through this attention module to obtain the channel attention map, so that the feature information in the vertical direction of the original image can be selectively emphasized through this channel attention map.
[0090] In another example, if the original image has a distribution feature in the horizontal position, there will be significant differences in the feature information between each vertical segment of the original image. In this case, the attention mechanism is used to extract the structural features representing the horizontal position relationship between the features of each vertical segment, that is, the weight values of the feature channels corresponding to each vertical segment are obtained through this attention module to obtain the channel attention map, so that the feature information in the horizontal direction of the original image can be selectively emphasized through this channel attention map.
[0091] Optionally, the channel attention map is used to indicate the weight values of each row of the feature scale map corresponding to each feature channel in the first spatial dimension, where the first spatial dimension is determined by the spatial position distribution of the original image, and the second spatial dimension is the spatial dimension other than the first spatial dimension in any feature scale map. For example, if the original image is an image with distribution features in the longitudinal position, the first spatial dimension is the height, and the second spatial dimension is the width; if the original image is an image with distribution features in the horizontal position, the first spatial dimension is the width, and the second spatial dimension is the height.
[0092] In a possible implementation manner, the attention module includes a pooling layer, a downsampling layer, a convolutional layer, and an upsampling layer; for any one of the multiple feature scale maps, calling the attention module to obtain the channel attention map corresponding to the any one of the feature scale maps includes, but is not limited to, the following steps 1-step 4.
[0093] Step 1, call the pooling layer to pool and compress the vectors in the second spatial dimension of any one of the feature scale maps to obtain the first intermediate feature in the first spatial dimension.
[0094] The embodiments of the present application do not limit the pooling method adopted by the pooling layer. For example, the pooling method can be average pooling, global average pooling, or maximum pooling, etc. The second spatial dimension of any one of the feature scale maps is compressed by the pooling compression method to facilitate extracting the relevant information between the vectors in the first spatial dimension.
[0095] For example, when the first spatial dimension is the height, see Figure 4 , Figure 4 which is a schematic diagram of a height-driven attention network (HANet) provided by the embodiments of the present application. As Figure 4 shown, any one of the feature scale maps is represented by X1(H*W*C), where H (height) represents the height of any one of the feature scale maps, W (width) represents the width of any one of the feature scale maps, and C (channel) represents the number of channels of any one of the feature scale maps. Call the pooling layer to pool and compress the vectors in the width of the any one of the feature scale maps X1(H*W*C) to obtain the first intermediate feature X1(H*C) in the height. Among them, the vector of the h-th row of the first intermediate feature X1(H*C) can be expressed as:
[0096]
[0097] Step 2, call the downsampling layer to perform downsampling processing on the first intermediate feature to obtain the refined second intermediate feature.
[0098] The embodiments of the present application do not limit the downsampling method used in the downsampling layer, and it is sufficient to downsample the first spatial dimension of the first intermediate feature to a hyperparameter. For example, the downsampling process can be implemented by the pooling method shown above, or by the method of dilated convolution, etc. Optionally, the value of the hyperparameter can be set according to experience or flexibly adjusted according to the application scenario. For example, the hyperparameter is set to 16. This step 2 filters out the features with little effect and information redundancy in the first intermediate feature through downsampling, retains the key features in the first intermediate feature, makes the obtained second intermediate feature more refined, and at the same time reduces the computational complexity of the network.
[0099] For example, still taking Figure 4 as an example, the first intermediate feature X1(H*C) is processed by the downsampling layer to obtain the refined second intermediate feature X1(H'*C), where H' represents the hyperparameter.
[0100] Step 3, call the convolutional layer to extract the context information between the vectors in the first spatial dimension of the second intermediate feature to obtain the third intermediate feature.
[0101] In the embodiments of the present application, the convolutional layer is used to extract the context information between the vectors in the first spatial dimension of the second intermediate feature through a convolutional operation, and the convolutional operation can obtain fine local features. The embodiments of the present application do not limit the structure of the convolutional layer. The process of extracting the context information between the vectors in the first spatial dimension of the second intermediate feature is an internal processing process of the convolutional layer, and the specific process of calling convolutional layers with different structures to extract the context information between the vectors in the first spatial dimension of the second intermediate feature may be different. In an exemplary embodiment, the convolutional layer includes one or more convolutional operations, and the structures of different convolutional operations may be the same or different.
[0102] Exemplarily, for the case where the number of convolutional operations is multiple, the multiple convolutional operations are connected in series in sequence. The process of calling the convolutional layer to process the second intermediate feature to obtain the third intermediate feature is as follows: input the second intermediate feature into the first convolutional operation layer to obtain the feature output by the first convolutional operation layer; starting from the second convolutional operation layer, input the feature output by the previous convolutional operation layer into the next convolutional operation layer to obtain the feature output by the next convolutional operation layer until the feature output by the last convolutional operation layer is obtained, and take the feature output by the last convolutional operation layer as the third intermediate feature.
[0103] In a possible implementation manner, each convolutional operation is composed of a sub-convolutional layer and an activation layer. The size of the convolutional kernel of the sub-convolutional layer is set according to experience or flexibly adjusted according to the actual application scenario. The embodiments of the present application do not limit this. For example, the size of the convolutional kernel of the convolutional layer is 1×3. The activation function used in the activation layer is set according to experience or flexibly adjusted according to the actual application scenario. The embodiments of the present application do not limit this. For example, the activation function used in the activation layer is the ReLU (Rectified Linear Unit) function or the Sigmoid (S-shaped) function. During the process of calling a convolutional operation to process features (or images), first call the sub-convolutional layer for processing, and then call the activation layer to process the features output by the convolutional layer.
[0104] Exemplarily, still taking Figure 4 as an example, the context information between the vectors of the first spatial dimension of the second intermediate feature X1 (H'*C) is extracted through the convolutional layer to obtain a third intermediate feature X1' (H'*C). The size of the third intermediate feature is the same as that of the second intermediate feature. In the exemplary embodiment, the convolutional layer includes 3 convolutional operations. The first convolutional operation includes a one-dimensional convolutional kernel, the second convolutional operation includes a one-dimensional convolutional kernel and a ReLU activation layer, and the second convolutional operation includes a one-dimensional convolutional kernel and a ReLU activation layer.
[0105] Step 4, call the upsampling layer to perform upsampling processing on the third intermediate feature to obtain a channel attention map corresponding to any feature scale map.
[0106] The embodiments of the present application do not limit the upsampling method used by the upsampling layer. It is only necessary to upsample the size of the third intermediate feature to the size of the first any feature scale map. For example, the upsampling method can be a bilinear interpolation method or a transposed convolution, etc. In this step 4, upsampling is used to make the size of the channel attention map the same as the size of the any feature scale map, which is convenient for the fusion of the channel attention map and the any feature scale map.
[0107] Exemplarily, still taking Figure 4 as an example, the third intermediate feature X1' (H'*C) is upsampled to a channel attention map X1' (H*C) through the upsampling layer. Thus, a channel attention map corresponding to any feature scale map is obtained.
[0108] Step 2023, call the fusion module to fuse multiple feature scale maps and the channel attention maps respectively corresponding to the multiple feature scale maps to obtain a target feature map.
[0109] In a possible implementation, the fusion module is called to fuse multiple feature scale maps and the channel attention maps corresponding to the multiple feature scale maps respectively to obtain a target feature map, including: calling the fusion module to fuse each feature scale map in the multiple feature scale maps with the channel attention map corresponding to each feature scale map to obtain a fusion result corresponding to each feature scale map; concatenating the fusion results corresponding to each feature scale map to obtain a target feature map.
[0110] The embodiments of the present application do not limit the method for fusing each feature scale map in the multiple feature scale maps with the channel attention map corresponding to each feature scale map. For example, directly multiplying each feature scale map in the multiple feature scale maps by the channel attention map corresponding to each feature scale map; or, performing weighted multiplication on each feature scale map in the multiple feature scale maps and the channel attention map corresponding to each feature scale map, and taking the obtained product result as the corresponding fusion result; or, multiplying each feature scale map in the multiple feature scale maps by the channel attention map corresponding to each feature scale map to obtain a product result corresponding to each feature scale map, and adding the product results corresponding to each feature scale map and the feature scale map to obtain a fusion result corresponding to each feature scale map.
[0111] Thus, through the above steps 2021 - 2023, the target feature map corresponding to the original image is extracted. Since the target feature map is obtained from the feature scale map and the channel attention map corresponding to the feature scale, and the channel attention map is determined according to the spatial position distribution of the original image, the obtained target feature map can represent the structural features of the spatial position distribution of the original image. Therefore, the feature information of the original image that the target feature map can represent is more accurate.
[0112] Step 203: Input the target feature map into a semantic classification model to obtain the semantic categories of each pixel point in the original image output by the semantic classification model.
[0113] After obtaining the target feature map, the semantic category corresponding to each pixel in the original image can be obtained based on the semantic classification model. Optionally, input the target feature map into the semantic classification model to obtain the semantic categories of each pixel point in the original image output by the semantic classification model. The embodiments of the present application do not limit the type of the semantic classification model. For example, the semantic classification model can be a Fully Convolutional Networks (FCN), a Pyramid Scene Parsing Network (PSPNet), or an Object Contextual Representations (OCR) model, etc.
[0114] Step 204: Obtain the semantic segmentation result of the original image based on the semantic categories of each pixel point in the original image.
[0115] In the embodiments of the present application, after obtaining the semantic categories of each pixel point in the original image, the areas composed of pixel points belonging to the same semantic category can be regarded as the image areas where the same object is located. Thus, the areas where each object in the original image is located, that is, the semantic segmentation result of the original image, can be obtained. That is to say, it can be known from the semantic segmentation result of the original image which areas in the original image include which objects.
[0116] The embodiments of the present application do not limit the form of the semantic segmentation result of the original image. Exemplarily, the form of the semantic segmentation result of the original image is a numerical pair. One pixel point corresponds to one numerical pair, and the numerical pair corresponding to one pixel point includes the position coordinates of the pixel point, the category to which the pixel point belongs, and the probability that the pixel point belongs to the category to which it belongs. Optionally, the numerical pairs of each pixel included in the semantic segmentation result of the original image can be visualized as an image. For example, pixels of different categories in the original image can be visualized as different semi-transparent colors to obtain the semantic segmentation map of the original image.
[0117] The semantic segmentation method provided by the embodiments of the present application uses an attention mechanism to obtain a target feature map that fuses the structural features of the spatial position distribution of the original image, so that the obtained target feature map strengthens the fusion of context information related to the spatial position distribution of the original image, and further improves the performance of the semantic segmentation result obtained according to the target feature map.
[0118] Exemplarily, taking the feature extraction model as HRNET and the semantic classification model as OCR model as an example, for the semantic segmentation of urban street view images with distribution characteristics in the vertical position. Refer to Figure 5 , Figure 5 is a schematic diagram of a semantic segmentation model provided by the embodiments of the present application. As Figure 5 shown, the semantic segmentation model includes HRNET fused with HANET and the OCR model. Among them, HRNET fused with HANET includes 4 feature extraction layers at different resolutions, and the resolution becomes lower as it goes down. Since the information between each feature extraction layer is fused with each other through operations such as convolution units (conv unit), upsampling, and downsampling, the feature scale maps of each obtained feature extraction layer maintain high-resolution representative feature information. Optionally, the model structure of this HANET is as Figure 4 shown.
[0119] As Figure 5As shown, after the original image is input into the semantic segmentation model, it first enters the HRNET integrated with HANET. The HRNET extracts feature scale maps X1, X2, X3, and X4 at different resolutions. Then, the channel attention maps X1', X2', X3', and X4' corresponding to the respective feature scale maps X1, X2, X3, and X4 are obtained through HANET respectively. Furthermore, each feature scale map X1, X2, X3, and X4 is multiplied by its corresponding channel attention map X1', X2', X3', and X4' respectively, and the results of each multiplication are added to the respective feature scale maps X1, X2, X3, and X4 to obtain the fusion results corresponding to each feature extraction layer. Finally, the fusion results corresponding to each feature extraction layer are cascaded to obtain the target feature map corresponding to the original image.
[0120] Optionally, the OCR model is a computational model used to represent the semantic categories of pixels in an image. After obtaining the target feature map of the original image through the HRNET integrated with HANET, the semantic categories of each pixel in the original image are calculated through the OCR method. Refer to Figure 5 , the OCR model includes the following parts: First, a rough semantic segmentation result, that is, soft object regions, is obtained through the intermediate layer of the backbone network; Second, K groups of vectors are calculated through the pixel representations output by the deep layer of the backbone network and the soft object regions, where K is an integer greater than 1, that is, object regions representations, and each vector corresponds to the feature representation of a semantic category; Third, the relationship matrix between the pixel representation and the object region representation is calculated, that is, the pixel-regions relation; Fourth, based on the values of the pixel representation and the object region representation of each pixel in the relation matrix, each object region representation is weighted and summed to obtain the object context representation, that is, OCR; Finally, based on OCR and the pixel features, an enhanced feature representation as context information is obtained, and this enhanced feature representation can be used to predict the semantic category of each pixel.
[0121] Based on the above Figure 1 shown implementation environment, the embodiments of the present application provide a training method for a semantic segmentation model. The training method of the semantic segmentation model can be applied to the computer device 101. The computer device 101 can be a terminal or a server, and the embodiments of the present application do not limit this. As Figure 6As shown in the figure, the training method of the semantic segmentation model provided by the embodiment of the present application includes the following steps 601 to 604.
[0122] Step 601: Obtain a sample image, a semantic label of the sample image, and an initial semantic segmentation model. The initial semantic segmentation model includes an initial semantic classification model and an initial feature extraction model with a fusion attention mechanism. The attention mechanism is used to obtain the structural features of the sample image based on the spatial position distribution of the sample image.
[0123] Exemplarily, the sample image refers to the image required for training the initial semantic segmentation model. Among them, the sample image is Figure 2 the same type of image as the original image in the illustrated embodiment. For example, the sample image and the original image are both urban street scene images to ensure the segmentation effect of the trained semantic segmentation model on the original image. It should be noted that the sample image mentioned in the embodiment of the present application refers to the sample image based on which the initial semantic segmentation model is trained once. The number of sample images can be one or multiple, and the embodiment of the present application does not limit this. Exemplarily, the number of sample images is multiple to ensure the model training effect.
[0124] In an exemplary embodiment, the way for the computer device to obtain the sample image is that the computer device uses the training images in a certain public dataset as the sample images. For example, the computer device uses the training images in the Cityscapes dataset as the sample images. The Cityscapes dataset has 5000 images of driving scenes in urban environments, among which there are 2975 training data, 1525 test data, and dense pixel category labels with 19 categories.
[0125] Optionally, the semantic label of the sample image is used to provide supervision information for the training process of the initial semantic segmentation model, and the semantic label of the sample image is used to provide information on whether each pixel point in the sample image belongs to the category of the reference object. In an exemplary embodiment, the information on whether each pixel point in the sample image belongs to the category of the reference object can be obtained through manual annotation, and the semantic label of the sample image refers to the pixel-level label of the sample image.
[0126] Step 602: Invoke the initial feature extraction model with a fusion attention mechanism to perform feature extraction on the original image, and obtain a target feature map that fuses the structural features of the sample image.
[0127] For the implementation manner of this step 602, refer to Figure 2 step 202 in the illustrated embodiment, which will not be elaborated here.
[0128] Step 603: Input the target feature map into the initial semantic classification model, and obtain the semantic categories of each pixel point in the sample image based on the output of the initial semantic classification model.
[0129] For the implementation of this step 603, refer to Figure 2 step 203 in the embodiments shown, which will not be elaborated here.
[0130] Step 604: Train the initial semantic segmentation model based on the semantic categories of each pixel point in the sample image and the semantic label of the sample image to obtain the semantic segmentation model.
[0131] After obtaining the semantic categories of each pixel point in the sample image, train the initial semantic segmentation model based on the semantic categories of each pixel point in the sample image and the semantic label of the sample image to obtain the trained semantic segmentation model.
[0132] In a possible implementation manner, training the initial semantic segmentation model based on the semantic categories of each pixel point in the sample image and the semantic label of the sample image includes: obtaining the loss function value based on the semantic categories of each pixel point in the sample image and the semantic label of the sample image; iteratively adjusting the parameters of the initial semantic classification model and the initial feature extraction model with a fusion attention mechanism in the initial semantic segmentation model according to the loss function value until the convergence condition is met.
[0133] Optionally, there is no limitation on the method of obtaining the loss function value based on the semantic categories of each pixel point in the sample image and the semantic label of the sample image. For example, use the weighted sum of the cross-entropy loss functions between the semantic categories of each pixel point in the sample image and each semantic label of the sample image as the loss function value; or use the weighted sum of the mean square error (MSE) functions between the semantic categories of each pixel point in the sample image and the semantic label of the sample image as the loss function value.
[0134] Exemplarily, in the process of calculating the weighted sum, the weights corresponding to the cross-entropy loss function or the MSE function between the semantic categories of each pixel point in the sample image and each semantic label of the sample image are set according to experience or flexibly adjusted according to the application scenario, which is not limited in the embodiments of the present application.
[0135] In a possible implementation, the parameters of the initial semantic classification model and the initial feature extraction model with a fusion attention mechanism in the initial semantic segmentation model are iteratively adjusted according to the loss function value until the convergence condition is met, and a semantic segmentation model is obtained, including: if the loss function value is greater than the target threshold, the parameters of the initial semantic classification model and the initial feature extraction model with a fusion attention mechanism in the initial semantic segmentation model are iteratively adjusted according to the loss function value until the convergence condition is met, and a semantic segmentation model is obtained; if the loss function value is less than or equal to the target threshold and the convergence condition is met, the initial semantic segmentation model under the current model parameters is used as the trained semantic segmentation model.
[0136] Optionally, the target threshold can be set according to experience or flexibly adjusted according to experience. For example, the target threshold is any value greater than or equal to 0 and less than or equal to 1.
[0137] In a possible implementation, the training method of the semantic segmentation model provided in the embodiments of the present application can be applied to a computer device with strong data processing capabilities such as a personal computer or a server. The semantic segmentation model trained by the above method can be implemented as an application program or a part of an application program and installed in the terminal, enabling the terminal to have the semantic segmentation ability of images. Alternatively, the semantic segmentation model trained by the above method can be applied to the background server of the application program, so that the server provides the semantic segmentation service of images for the application program in the terminal.
[0138] Exemplarily, taking Figure 5 the shown semantic segmentation model as an example, the Cityscapes dataset is used to train the semantic segmentation model. The comparison results of the model performance between the semantic segmentation model (HRNET+OCR integrated with HAENT) trained based on the method provided in the embodiments of the present application and the semantic segmentation model (HRNET+OCR) in the related art are shown in Table 1. Optionally, mIoU is used as the evaluation index (definition of mIoU: calculate the ratio of the intersection and union of the two sets of the true value and the predicted value).
[0139] Table 1
[0140] Semantic segmentation model Related technology (mIoU) This application (mIoU) Train 74.86 82.71 Motorcycle 66.23 68.93 Bus 90.21 92.7 Cyclist 68 69.28 Bicycle 80.23 81.1 Person 85.13 85.54 Street lamp 76.56 76.7 Road 98.27 98.36 Sky 95.49 95.58
[0141] It can be seen from Table 1 that the semantic segmentation model trained based on the method provided in the embodiments of the present application can obtain higher accuracy than the semantic segmentation model in the related art.
[0142] In a possible implementation, the computer device displays a segmentation image corresponding to the test image according to the semantic segmentation result of the test image, so as to determine the semantic segmentation performance of the semantic segmentation model according to the segmentation image, where different categories of objects are labeled in the segmentation image. Optionally, the computer device pre-assigns marker colors to each category, and then fills each pixel point with the corresponding marker color according to the category of the object to which each pixel point belongs, so as to generate the semantic segmentation image corresponding to the image.
[0143] Please refer to Figure 7 , which shows a comparison diagram of the segmentation images obtained by respectively performing semantic segmentation on a test image by a semantic segmentation model (HRNET+OCR fused with HAENT) trained by the method provided in the embodiments of the present application and a semantic segmentation model (HRNET+OCR) in the related art. As can be seen from the small box in Figure 7 , compared with the road fences and electric vehicles segmented by the semantic segmentation model (HRNET+OCR) in the related art, the road fences and electric vehicles segmented by the semantic segmentation model (HRNET+OCR fused with HAENT) trained by the method provided in the embodiments of the present application are more accurate. Therefore, the semantic segmentation performance of the semantic segmentation model (HRNET+OCR fused with HAENT) trained by the method provided in the embodiments of the present application is improved.
[0144] Based on the technical solution provided in the embodiments of the present application, a semantic segmentation model including a semantic classification model and a feature extraction model integrating an attention mechanism can be trained, which lays a foundation for achieving the target feature map that obtains the structural features integrating the spatial position distribution of the original image through the attention mechanism, and can make the obtained target feature map strengthen the context information fusion related to the spatial position distribution of the original image, and further improve the performance of the semantic segmentation result obtained according to the target feature map.
[0145] See Figure 8 , the embodiments of the present application provide a semantic segmentation device, which includes:
[0146] A first acquisition module 801, configured to acquire an original image and a semantic segmentation model, where the semantic segmentation model includes a semantic classification model and a feature extraction model integrating an attention mechanism, and the attention mechanism is used to acquire the structural features of the original image based on the spatial position distribution of the original image;
[0147] A feature extraction module 802, configured to call the feature extraction model integrating the attention mechanism to perform feature extraction on the original image, so as to obtain a target feature map integrating the structural features of the original image;
[0148] A classification module 803, configured to input a target feature map into a semantic classification model to obtain the semantic categories of each pixel point in the original image output by the semantic classification model;
[0149] A second acquisition module 804, configured to obtain a semantic segmentation result of the original image based on the semantic categories of each pixel point in the original image.
[0150] In a possible implementation manner, the feature extraction model integrating an attention mechanism includes a feature extraction module, an attention module, and a fusion module;
[0151] A feature extraction module 802, configured to call the feature extraction module to extract multiple feature scale maps corresponding to the original image at different resolutions; call the attention module to obtain channel attention maps respectively corresponding to the multiple feature scale maps, where the channel attention map is used to indicate the weight values of each row in the first spatial dimension of the feature scale map corresponding to each feature channel, and the first spatial dimension is determined by the spatial position distribution of the original image; call the fusion module to fuse the multiple feature scale maps and the channel attention maps respectively corresponding to the multiple feature scale maps to obtain a target feature map.
[0152] In a possible implementation manner, the attention module includes a pooling layer, a downsampling layer, a convolutional layer, and an upsampling layer;
[0153] A feature extraction module 802, configured to, for any one of the multiple feature scale maps, call the pooling layer to perform pooling compression on the vector in the second spatial dimension of the any one feature scale map to obtain a first intermediate feature in the first spatial dimension, where the second spatial dimension is the spatial dimension other than the first spatial dimension in the any one feature scale map; call the downsampling layer to perform downsampling processing on the first intermediate feature to obtain a refined second intermediate feature; call the convolutional layer to extract the context information between the vectors in the first spatial dimension of the second intermediate feature to obtain a third intermediate feature; call the upsampling layer to perform upsampling processing on the third intermediate feature to obtain a channel attention map corresponding to the any one feature scale map, and the size of the channel attention map is the same as the size of the any one feature scale map.
[0154] In a possible implementation manner, the feature extraction module 802 is configured to call the fusion module to fuse each of the multiple feature scale maps with the channel attention map corresponding to each feature scale map to obtain a fusion result corresponding to each feature scale map; concatenate the fusion results corresponding to each feature scale map to obtain a target feature map.
[0155] In a possible implementation, the feature extraction module 802 is configured to call the fusion module to multiply each feature scale map in multiple feature scale maps by the channel attention map corresponding to each feature scale map to obtain a product result corresponding to each feature scale map; and add each feature scale map and the product result corresponding to each feature scale map to obtain a fusion result corresponding to each feature scale map.
[0156] The semantic segmentation method provided by the embodiments of this application uses an attention mechanism to obtain a target feature map that fuses the structural features of the spatial position distribution of the original image, so that the obtained target feature map strengthens the fusion of context information related to the spatial position distribution of the original image, and further improves the performance of the semantic segmentation result obtained based on this target feature map.
[0157] See Figure 9 , the embodiments of this application provide a training device for a semantic segmentation model, and the device includes:
[0158] The first acquisition module 901 is configured to acquire a sample image, the semantic label of the sample image, and an initial semantic segmentation model. The initial semantic segmentation model includes an initial semantic classification model and an initial feature extraction model integrating an attention mechanism. The attention mechanism is used to acquire the structural features of the sample image based on the spatial position distribution of the sample image;
[0159] The feature extraction module 902 is configured to call the initial feature extraction model integrating the attention mechanism to perform feature extraction on the original image to obtain a target feature map that fuses the structural features of the sample image;
[0160] The classification module 903 is configured to input the target feature map into the initial semantic classification model and output the semantic categories of each pixel point in the sample image based on the output of the initial semantic classification model;
[0161] The training module 904 is configured to train the initial semantic segmentation model based on the semantic categories of each pixel point in the sample image and the semantic label of the sample image to obtain a semantic segmentation model.
[0162] In a possible implementation, the training module 904 is configured to obtain a loss function value based on the semantic categories of each pixel point in the sample image and the semantic label of the sample image; and iteratively adjust the parameters of the initial semantic classification model and the initial feature extraction model integrating the attention mechanism in the initial semantic segmentation model until the convergence condition is met.
[0163] The training device for the semantic segmentation model provided by the embodiments of the present application can train a semantic segmentation model including a semantic classification model and a feature extraction model with a fusion attention mechanism, laying a foundation for achieving the target feature map that obtains the fused structural features of the original image through the attention mechanism, enabling the obtained target feature map to strengthen the context information fusion related to the spatial position distribution of the original image, and further improving the performance of the semantic segmentation result obtained based on the target feature map.
[0164] It should be understood that when the device provided in the above embodiments realizes its functions, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiments and the method embodiments belong to the same concept, and the specific implementation process can be seen in the method embodiments, which will not be repeated here.
[0165] Please refer to Figure 10 , which shows a schematic structural diagram of a computer device provided by an embodiment of the present application. The computer device can be a terminal, for example, it can be: a smart phone, a tablet computer, a vehicle-mounted terminal, a notebook computer or a desktop computer. The terminal may also be called by other names such as user equipment, portable terminal, laptop terminal, desktop terminal, etc.
[0166] Generally, the terminal includes: a processor 701 and a memory 702.
[0167] The processor 701 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 701 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 701 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 701 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 701 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0168] The memory 702 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 702 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 702 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 701 to implement the semantic segmentation method provided in the method embodiments of the present application.
[0169] In some embodiments, the terminal may also optionally include: a peripheral device interface 703 and at least one peripheral device. The processor 701, the memory 702, and the peripheral device interface 703 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 703 through a bus, signal lines, or a circuit board. Specifically, the peripheral devices include at least one of a radio frequency circuit 704, a display screen 705, a camera assembly 706, an audio circuit 707, and a power supply 709.
[0170] The peripheral device interface 703 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 701 and the memory 702. In some embodiments, the processor 701, the memory 702, and the peripheral device interface 703 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 701, the memory 702, and the peripheral device interface 703 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.
[0171] The radio frequency circuit 704 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 704 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 704 converts an electrical signal into an electromagnetic signal for transmission, or converts a received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 704 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 704 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: metropolitan area network, generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area network, and / or Wireless Fidelity (WiFi) network. In some embodiments, the radio frequency circuit 704 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.
[0172] The display screen 705 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 705 is a touch display screen, the display screen 705 also has the ability to collect touch signals on or above the surface of the display screen 705. The touch signals can be input as control signals to the processor 701 for processing. At this time, the display screen 705 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there can be one display screen 705, which is set on the front panel of the terminal; in some other embodiments, there can be at least two display screens 705, which are respectively set on different surfaces of the terminal or are in a foldable design; in still some other embodiments, the display screen 705 can be a flexible display screen, which is set on the curved surface or the folding surface of the terminal. Even, the display screen 705 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 705 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0173] The camera module 706 is used to capture images or videos. Optionally, the camera module 706 includes a front camera and a rear camera. Generally, the front camera is set on the front panel of the terminal, and the rear camera is set on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera respectively, to implement functions such as the combination of the main camera and the depth-of-field camera to achieve the background blurring function, the combination of the main camera and the wide-angle camera to achieve panoramic shooting and VR (Virtual Reality) shooting functions or other combined shooting functions. In some embodiments, the camera module 706 can also include a flash. The flash can be a single-color-temperature flash or a two-color-temperature flash. The two-color-temperature flash refers to the combination of a warm-light flash and a cold-light flash, which can be used for light compensation under different color temperatures.
[0174] The audio circuit 707 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 701 for processing, or input to the radio frequency circuit 704 to enable voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 701 or the radio frequency circuit 704 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 707 may further include a headphone jack.
[0175] The power supply 709 is used to supply power to each component in the terminal. The power supply 709 may be alternating current, direct current, a disposable battery or a rechargeable battery. When the power supply 709 includes a rechargeable battery, the rechargeable battery may support wired charging or wireless charging. The rechargeable battery may also be used to support fast charging technology.
[0176] In some embodiments, the terminal further includes one or more sensors 710. The one or more sensors 710 include but are not limited to: an acceleration sensor 711, a gyroscope sensor 712, a pressure sensor 713, an optical sensor 715, and a proximity sensor 716.
[0177] The acceleration sensor 711 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established by the terminal. For example, the acceleration sensor 711 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 701 can control the display screen 705 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 711. The acceleration sensor 711 can also be used for collecting game or user movement data.
[0178] The gyroscope sensor 712 can detect the body direction and rotation angle of the terminal. The gyroscope sensor 712 can cooperate with the acceleration sensor 711 to collect the 3D actions of the user on the terminal. According to the data collected by the gyroscope sensor 712, the processor 701 can implement the following functions: motion sensing (such as changing the UI according to the user's tilting operation), image stabilization during shooting, game control, and inertial navigation.
[0179] The pressure sensor 713 can be disposed on the side frame of the terminal and / or the lower layer of the display screen 705. When the pressure sensor 713 is disposed on the side frame of the terminal, it can detect the holding signal of the user for the terminal, and the processor 701 can perform left / right hand recognition or quick operation according to the holding signal collected by the pressure sensor 713. When the pressure sensor 713 is disposed on the lower layer of the display screen 705, the processor 701 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 705. The operable controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.
[0180] The optical sensor 715 is used to collect the ambient light intensity. In one embodiment, the processor 701 can control the display brightness of the display screen 705 according to the ambient light intensity collected by the optical sensor 715. Specifically, when the ambient light intensity is high, the display brightness of the display screen 705 is increased; when the ambient light intensity is low, the display brightness of the display screen 705 is decreased. In another embodiment, the processor 701 can also dynamically adjust the shooting parameters of the camera module 706 according to the ambient light intensity collected by the optical sensor 715.
[0181] The proximity sensor 716, also known as a distance sensor, is usually disposed on the front panel of the terminal. The proximity sensor 716 is used to collect the distance between the user and the front of the terminal. In one embodiment, when the proximity sensor 716 detects that the distance between the user and the front of the terminal is gradually decreasing, the processor 701 controls the display screen 705 to switch from the lit state to the off state; when the proximity sensor 716 detects that the distance between the user and the front of the terminal is gradually increasing, the processor 701 controls the display screen 705 to switch from the off state to the lit state.
[0182] Those skilled in the art can understand that Figure 10 the structure shown in
[0183] does not constitute a limitation on the computer device, and may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component layout. Figure 11 , Figure 11It is a schematic structural diagram of a server provided by an embodiment of the present application. The server 400 may vary greatly due to different configurations or performances, and may include one or more processors 401 and one or more memories 402. Among them, at least one program instruction is stored in the one or more memories 402, and the at least one program instruction is loaded and executed by the one or more processors 401 to implement the semantic segmentation method provided by each of the above method embodiments. Of course, the server 400 may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input / output. The server 400 may also include other components for implementing the functions of the device, which will not be elaborated here.
[0184] In an exemplary embodiment, a computer device is also provided. The computer device includes a processor and a memory, and at least one program code is stored in the memory. The at least one program code is loaded and executed by one or more processors to enable the computer device to implement any of the above semantic segmentation methods.
[0185] In an exemplary embodiment, a computer-readable storage medium is also provided. At least one program code is stored in the computer-readable storage medium, and the at least one program code is loaded and executed by a processor of a computer device to enable the computer to implement any of the above semantic segmentation methods.
[0186] Optionally, the above computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, etc.
[0187] In an exemplary embodiment, a computer program product or a computer program is also provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes any of the above semantic segmentation methods.
[0188] In the description, claims, and drawings of this application, terms such as "first", "second", "third", and "fourth" are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.
[0189] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. For example, the original images and semantic segmentation models involved in this application are obtained under full authorization.
[0190] The above are only optional embodiments of this application and are not intended to limit this application. Any modifications, equivalent replacements, improvements, etc. made within the principles of this application shall be included within the protection scope of this application.
Claims
1. A semantic segmentation method, characterized in that, the method includes: obtaining an original image and a semantic segmentation model, where the semantic segmentation model includes a semantic classification model and a feature extraction model integrating an attention mechanism, and the attention mechanism is used to obtain the structural features of the original image based on the spatial position distribution of the original image; invoking the feature extraction model integrating the attention mechanism to perform feature extraction on the original image to obtain a target feature map integrating the structural features of the original image; inputting the target feature map into the semantic classification model to obtain the semantic categories of each pixel point in the original image output by the semantic classification model; obtaining the semantic segmentation result of the original image based on the semantic categories of each pixel point in the original image; the feature extraction model integrating the attention mechanism includes a feature extraction module, an attention module, and a fusion module; the step of invoking the feature extraction model integrating the attention mechanism to perform feature extraction on the original image to obtain a target feature map integrating the structural features of the original image includes: invoking the feature extraction module to extract multiple feature scale maps corresponding to the original image at different resolutions; invoking the attention module to obtain the channel attention maps corresponding to the multiple feature scale maps respectively, where the channel attention map is used to indicate the weight value of each row in the first spatial dimension of the feature scale map corresponding to each feature channel, and the first spatial dimension is determined by the spatial position distribution of the original image; invoking the fusion module to fuse the multiple feature scale maps and the channel attention maps corresponding to the multiple feature scale maps respectively to obtain the target feature map; the step of invoking the fusion module to fuse the multiple feature scale maps and the channel attention maps corresponding to the multiple feature scale maps respectively to obtain the target feature map includes: invoking the fusion module to fuse each feature scale map in the multiple feature scale maps with the channel attention map corresponding to each feature scale map to obtain the fusion result corresponding to each feature scale map; cascading the fusion results corresponding to each feature scale map to obtain the target feature map; the step of invoking the fusion module to fuse each feature scale map in the multiple feature scale maps with the channel attention map corresponding to each feature scale map to obtain the fusion result corresponding to each feature scale map includes: invoking the fusion module to multiply each feature scale map in the multiple feature scale maps by the channel attention map corresponding to each feature scale map to obtain the product result corresponding to each feature scale map; adding the each feature scale map and the product result corresponding to each feature scale map to obtain the fusion result corresponding to each feature scale map.
2. The method according to claim 1, characterized in that, the attention module includes a pooling layer, a downsampling layer, a convolutional layer, and an upsampling layer; the step of invoking the attention module to obtain the channel attention maps corresponding to the multiple feature scale maps respectively includes: For any one of the multiple feature scale maps, call the pooling layer to pool and compress the vectors in the second spatial dimension of the any one of the feature scale maps, to obtain a first intermediate feature in the first spatial dimension, where the second spatial dimension is the spatial dimension other than the first spatial dimension in the any one of the feature scale maps; Call the downsampling layer to perform downsampling processing on the first intermediate feature, to obtain a refined second intermediate feature; Call the convolutional layer to extract the context information between the vectors in the first spatial dimension of the second intermediate feature, to obtain a third intermediate feature; Call the upsampling layer to perform upsampling processing on the third intermediate feature, to obtain a channel attention map corresponding to the any one of the feature scale maps, where the size of the channel attention map is the same as the size of the any one of the feature scale maps.
3. A semantic segmentation device, characterized in that, the device includes: a first acquisition module, configured to acquire an original image and a semantic segmentation model, where the semantic segmentation model includes a semantic classification model and a feature extraction model integrating an attention mechanism, and the attention mechanism is configured to acquire the structural features of the original image based on the spatial position distribution of the original image; a feature extraction module, configured to call the feature extraction model integrating the attention mechanism to perform feature extraction on the original image, to obtain a target feature map integrating the structural features of the original image; a classification module, configured to input the target feature map into the semantic classification model, to obtain the semantic categories of each pixel point in the original image output by the semantic classification model; a second acquisition module, configured to acquire the semantic segmentation result of the original image based on the semantic categories of each pixel point in the original image; the feature extraction model integrating the attention mechanism includes a feature extraction module, an attention module, and a fusion module; the calling the feature extraction model integrating the attention mechanism to perform feature extraction on the original image, to obtain a target feature map integrating the structural features of the original image, includes: calling the feature extraction module to extract multiple feature scale maps corresponding to the original image at different resolutions; calling the attention module to acquire the channel attention maps respectively corresponding to the multiple feature scale maps, where the channel attention map is used to indicate the weight value of each row in the first spatial dimension of the feature scale map corresponding to each feature channel, and the first spatial dimension is determined by the spatial position distribution of the original image; calling the fusion module to fuse the multiple feature scale maps and the channel attention maps respectively corresponding to the multiple feature scale maps, to obtain the target feature map; the calling the fusion module to fuse the multiple feature scale maps and the channel attention maps respectively corresponding to the multiple feature scale maps, to obtain the target feature map, includes: calling the fusion module to fuse each feature scale map in the multiple feature scale maps with the channel attention map corresponding to each feature scale map, to obtain a fusion result corresponding to each feature scale map; cascading the fusion results corresponding to each feature scale map, to obtain the target feature map; Invoking the fusion module to fuse each feature scale map in the multiple feature scale maps with the channel attention map corresponding to each feature scale map to obtain a fusion result corresponding to each feature scale map includes: Invoking the fusion module to multiply each feature scale map in the multiple feature scale maps by the channel attention map corresponding to each feature scale map to obtain a product result corresponding to each feature scale map; Adding each feature scale map to the product result corresponding to each feature scale map to obtain a fusion result corresponding to each feature scale map.
4. A computer device, characterized in that, the computer device includes a processor and a memory, and at least one computer program or instruction is stored in the memory, and the at least one computer program or instruction is loaded and executed by the processor so that the computer device implements the semantic segmentation method according to any one of claims 1 to 2.
Citation Information
Patent Citations
Improved HRnet based on attention mechanism
CN112270213A