A semantic segmentation method and related devices

By employing a depth convolutional network with deformable self-attention to identify and aggregate semantically relevant features in semantic segmentation, the method addresses the high computational complexity and instability of existing self-attention mechanisms, achieving improved efficiency and accuracy.

CN112749706BActive Publication Date: 2025-07-15TENCENT TECH SHANGHAI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010557682.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-06-17
Publication Date
2025-07-15
Estimated Expiration
2040-06-17

AI Technical Summary

Technical Problem

The existing self-attention mechanism is highly complex in semantic segmentation tasks and has unstable performance, especially when applied to fully convolutional networks, resulting in limited semantic segmentation accuracy and efficiency.

Method used

By building a fully convolutional network, a feature map is generated and channel reduction is performed, M square areas related to the semantics of the target query point are determined, feature aggregation is performed, computational complexity is reduced and semantic segmentation accuracy is improved.

Benefits of technology

The computational complexity of semantic segmentation is reduced, the accuracy and speed of semantic segmentation is improved, and the most relevant context information is adaptively aggregated to generate smaller but more accurate attention maps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112749706B_ABST
    Figure CN112749706B_ABST
Patent Text Reader

Abstract

The present application provides a semantic segmentation method and related devices, which reduce the computational complexity during semantic segmentation through a deformable self-attention mechanism, reduce the waiting time, and increase the accuracy of semantic segmentation. The method includes: inputting a target image into a deep convolutional network to obtain a first feature map; performing a channel reduction operation on the first feature map to obtain a second feature map; determining M square regions and a target square region that are semantically related to a target query point; aggregating the features of the M square regions and the features of the target square region to obtain an aggregated feature; determining an output feature according to the first feature map and the aggregated feature; and performing semantic segmentation on the fused output feature to obtain a semantic segmentation map corresponding to the target image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision, and in particular, to a semantic segmentation method and related devices. Background Art

[0002] The self-attention mechanism calculates the response of each query point in the feature map as a weighted average of the features at all positions. These weights are measured by the pairwise "relationships" (correlation degrees) between all elements, that is, the more similar the features are, the greater the weight used for feature aggregation. Its time and space complexity are both O(CH 2 W 2 ), because it is necessary to describe the "relationships" of each pixel pair, where H×W is the spatial scale of the feature map and C is the number of channels of the feature map. Such a mechanism helps the semantic segmentation task to perform more accurate pixel classification and can help each pixel obtain dense context information.

[0003] Currently, two context aggregation methods based on the attention mechanism are mainly adopted. For each query point, the initial self-attention mechanism generates an attention map of H×W, while the criss-cross attention module generates a sparse attention map of only (H+W+1). After repeated criss-cross attention operations, each query point in the finally output feature map can capture long-range correlations from all pixels.

[0004] Because the time complexity of the self-attention mechanism is high, some methods have been proposed to reduce its computational complexity. The Criss-Cross NET (CCNET) proposed the Criss-Cross Attention (CCA) module, where each query point is connected to other positions in the same row and the same column to predict a sparse attention map. The attention map corresponding to each query point has only H+W-1 weights. While the weight of the attention map of each query point under the initial self-attention mechanism is H×W. In addition, in order to capture richer semantics, through the Recurrent Criss-Cross Attention (RCCA) module, which consists of two consecutive CCA modules, the number of weights of the attention map is reduced from H 2 W 2 to 2(H+W-1).

[0005] The RCCA stacks two CCA modules to obtain global spatial dependencies. However, its theoretical defects lead to unstable training. That is, although running the CCA module twice can provide dense connections and gradient paths. However, when the query point is semantically related to the key point in the second CCA module but not related to the "turning point" in the first CCA module, the correlation calculation between them is very likely to be disrupted. Therefore, when the RCCA module is applied to a fully convolutional network (FCN) with dilated convolutions, its performance is not stable. Summary of the Invention

[0006] This application provides a semantic segmentation method and related devices, which reduce the computational complexity during semantic segmentation and increase the accuracy of semantic segmentation.

[0007] The first aspect of this application provides a semantic segmentation method, including:

[0008] Input the target image into a deep convolutional network to obtain a first feature map, where the deep convolutional network is constructed in a fully convolutional manner, and the target image is the image to be semantically segmented;

[0009] Perform a channel reduction operation on the first feature map to obtain a second feature map, where the number of channels of the second feature map is less than that of the first feature map;

[0010] Determine M square regions and a target square region that are semantically related to the target query point. The target query point is any one of the query points corresponding to the second feature map. The side length of each square region in the M square regions is L, and the target square region is a square region centered on the target query point with a side length of L, where the values of M and L need to satisfy (M + 1) × L 2 < H × W, and L is an odd number, H is the height of the target image, and W is the width of the target image;

[0011] Aggregate the features of the M square regions and the features of the target square region to obtain an aggregated feature;

[0012] Determine the output feature according to the first feature map and the aggregated feature;

[0013] Perform semantic segmentation on the fused output feature to obtain a semantic segmentation map corresponding to the target image.

[0014] Optionally, the determining of the M square regions that are semantically related to the target query point includes:

[0015] Calculate the responses of the target query point to all elements in the M square regions and the target square region;

[0016] Determine the indices of the target query point to all elements in the M square regions and the target square region according to the responses of the target query point to all elements in the M square regions and the target square region;

[0017] Determine the offset of the target query point according to the indices of the target query point to all elements in the M square regions and the target square region and the side length L;

[0018] Determine the M square regions according to the offset of the target query point and the side length L.

[0019] Optionally, the aggregating the features of the M square regions and the features of the target square region to obtain an aggregated feature includes:

[0020] Calculate the similarity between the target query point and each of the M square regions and the target square region to obtain the weights of the target query point to each of the M square regions and the target square region;

[0021] Normalize the weights of each region;

[0022] Perform a weighted sum of the normalized weights of each region and the positions of each region to obtain an aggregated feature.

[0023] Optionally, the calculating the similarity between the target query point and each of the M square regions and the target square region to obtain the weights of the target query point to each of the M square regions and the target square region includes:

[0024] Determine the actual indices of the elements related to the target query point in the M square regions and the target square region;

[0025] Based on the actual indices, calculate the similarity between the target query point and each of the M square regions and the target square region to obtain the weights of the target query point to each of the M square regions and the target square region.

[0026] Optionally, the determining the output feature according to the first feature map and the aggregated feature includes:

[0027] Determine a third feature map according to the aggregated feature;

[0028] Determine the proportional parameter;

[0029] Calculate the third feature map and the first feature map through the proportional parameter to obtain the output feature.

[0030] The second aspect of the present application provides a semantic segmentation device, including:

[0031] A first determination unit for inputting a target image into a deep convolutional network to obtain a first feature map, where the deep convolutional network is constructed in a fully convolutional manner, and the target image is an image to be semantically segmented;

[0032] A channel reduction unit for performing a channel reduction operation on the first feature map to obtain a second feature map, where the number of channels of the second feature map is less than the number of channels of the first feature map;

[0033] A second determination unit for determining M square regions and a target square region that are semantically related to a target query point, where the target query point is any one of the query points corresponding to the second feature map, the side length of each square region in the M square regions is L, and the target square region is a square region with the target query point as the center and a side length of L, where the values of M and L need to satisfy (M + 1) × L 2 < H × W, and L is an odd number, H is the height of the target image, and W is the width of the target image;

[0034] An aggregation unit for aggregating the features of the M square regions and the features of the target square region to obtain an aggregated feature;

[0035] A third determination unit for determining an output feature according to the first feature map and the aggregated feature;

[0036] A segmentation unit for performing semantic segmentation on the fused output feature to obtain a semantic segmentation map corresponding to the target image.

[0037] Optionally, the second determination unit is specifically used for:

[0038] Calculate the responses of the target query point to all elements in the M square regions and the target square region;

[0039] Determine the indexes of the target query point to all elements in the M square regions and the target square region according to the responses of the target query point to all elements in the M square regions and the target square region;

[0040] Determine the offset of the target query point based on the target query point, the M square regions, the indexes of all elements in the target square region, and the side length L.

[0041] Determine the M square regions according to the offset of the target query point and the side length L.

[0042] Optionally, the aggregation unit is specifically configured to:

[0043] Calculate the similarity between the target query point and each of the M square regions and the target square region to obtain the weights of the target query point and each of the M square regions and the target square region.

[0044] Perform normalization processing on the weights of each region.

[0045] Perform weighted summation on the normalized weights of each region and the positions of each region to obtain an aggregated feature.

[0046] Optionally, when the aggregation unit calculates the similarity between the target query point and each of the M square regions and the target square region to obtain the weights of the target query point and each of the M square regions and the target square region, it includes:

[0047] Determine the actual indexes of the elements related to the target query point in the M square regions and the target square region.

[0048] Based on the actual indexes, calculate the similarity between the target query point and each of the M square regions and the target square region to obtain the weights of the target query point and each of the M square regions and the target square region.

[0049] Optionally, the third determination unit is specifically configured to:

[0050] Determine a third feature map according to the aggregated feature.

[0051] Determine a scaling parameter.

[0052] Calculate the third feature map and the first feature map through the scaling parameter to obtain the output feature.

[0053] A third aspect of the present application provides a computer device, which includes at least one connected processor, a memory, and a transceiver, wherein the memory is used to store program code, and the program code is loaded and executed by the processor to implement the steps of the semantic segmentation method described above.

[0054] The fourth aspect of the present application provides a computer-readable storage medium, which includes instructions that, when running on a computer, cause the computer to execute the steps of the semantic segmentation method described above.

[0055] In summary, it can be seen that in the embodiments provided by the present application, by determining M regions that are semantically related to the target query point, and the side length of the M regions is L, the most relevant context information at each position is adaptively aggregated. The computational complexity corresponding to the self-attention calculation mechanism provided by the present application is O(HW((M + 1)L 2 ), the values of M and L can be flexibly set, and the computational complexity can be significantly lower than that of the existing self-attention mechanism. Furthermore, the duration of semantic segmentation can be reduced, and the accuracy of semantic segmentation can be improved. Description of the Drawings

[0056] Figure 1a It is a self-attention diagram generated by the self-attention mechanism (Self-attention) provided by the embodiments of the present application;

[0057] Figure 1b It is a sparse attention diagram generated by the cross-shaped cross-attention mechanism provided by the embodiments of the present application;

[0058] Figure 2 It is an overall architecture diagram of the semantic segmentation model based on the deformable self-attention network provided by the embodiments of the present application;

[0059] Figure 3 It is a flowchart of the semantic segmentation method provided by the embodiments of the present application;

[0060] Figure 4 It is a diagram of the deformable matrix multiplication provided by the embodiments of the present application;

[0061] Figure 5 It is a virtual structure diagram of the semantic segmentation device provided by the embodiments of the present application;

[0062] Figure 6 It is a hardware structure diagram of the server provided by the embodiments of the present application. Detailed Embodiments

[0063] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments.

[0064] In the description and claims of this application and the above-mentioned drawings, terms such as "first" and "second" are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or modules does not necessarily have to be limited to those steps or modules clearly listed, but may include other steps or modules that are not clearly listed or are inherent to these processes, methods, products or devices. The division of modules in this application is only a logical division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some feature vectors can be ignored or not executed. In addition, the displayed or discussed coupling, direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection between modules can be in an electrical or other similar form, which is not limited in this application. And the modules or sub-modules described as separate components may or may not be physically separated, may or may not be physical modules, or may be distributed to multiple circuit modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this application.

[0065] The embodiments of this application relate to the fields of artificial intelligence and machine learning. The relevant content of artificial intelligence and machine learning will be described below:

[0066] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields. For example, common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0067] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence is also the study of the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning and decision-making.

[0068] Artificial intelligence technology is an interdisciplinary subject that involves a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0069] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between humans and computers using natural languages. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural languages, that is, the languages people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include technologies such as text processing, semantic understanding, machine translation, robot question answering, and knowledge graphs.

[0070] Machine Learning (ML) is an interdisciplinary subject that involves multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0071] Please refer to Figure 1a , Figure 1a which is the self-attention schematic diagram generated by the self-attention mechanism (Self-attention) provided in the embodiment of this application, Figure 1b which is the sparse attention schematic diagram generated by the cross-cross attention mechanism provided in the embodiment of this application. Refer to Figure 1a and Figure 1b for each query point, Figure 1a for query point 101 in , the initial self-attention mechanism generates an attention map of H×W, while the cross-cross attention module (Criss-Crossattention) generates a sparse attention map of only (H+W+1). After repeated cross-cross attention operations, each query point 102 in the final output feature map can capture long-range correlations from all pixels. For the sake of clear display, the remaining other connections are ignored.

[0072] Because of the high time complexity of the self-attention mechanism, some methods have been proposed to reduce its computational complexity. For example, Figure 1b As shown in Figure 1b , CCNet proposed the Criss-Cross Attention (CCA) module. Each query point is connected to other positions in the same row and column to predict a sparse attention map. The attention map corresponding to each query point has only H+W-1 weights, while the weight of the attention map of each query point under the initial self-attention mechanism is H×W. In addition, to capture richer semantics, CCNet proposed the Recurrent Criss-Cross Attention (RCCA) module, which consists of two consecutive CCA modules, thus reducing the number of weights of the attention map from H 2 W 2 to 2(H+W-1).

[0073] RCCA stacks two Criss-Cross Attention (CCA) modules to obtain global spatial dependencies. However, its theoretical defect leads to unstable training. That is, although running two CCA modules can provide dense connections and gradient paths. However, when the query point is semantically related to the Key point in the second CCA module but not related to the "turning point" in the first CCA module, the correlation calculation between them is very likely to be damaged. Therefore, when the RCCA module is applied to the FCN with dilated convolution, its performance is not stable.

[0074] In view of this, the embodiments of the present application provide a semantic segmentation method. By inputting an image into a deep convolutional network and generating a feature map Z, where the network is constructed in a fully convolutional manner, the spatial size of Z is H×W. Then, following the settings of CCNet, the last two down-convolutional layers are removed, and dilated convolution is applied in the subsequent layers, which retains more details without increasing the parameters. Therefore, both the width and height of Z are 1 / 8 of the input image. And a convolutional layer is used to reduce the number of channels of the feature map Z to obtain the feature map X, and M square regions and the target query region related to the target query point are determined. Then, the features of the M square regions and the features of the target square region are aggregated to obtain the aggregated feature E. Then, the aggregated feature E and the feature map Z are connected to obtain the final output feature H. Finally, the output feature is feature-fused, and the fused feature is semantically segmented to generate a segmentation map. In this way, compared with the initial self-attention mechanism, features that are not very relevant to the query point still contribute to the weights of feature aggregation. In the present application, the weights used for aggregating features are all calculated from semantically related regions, which makes the generated attention map smaller but more accurate, thereby improving the speed and accuracy of semantic segmentation.

[0075] Please refer to Figure 2 , Figure 2 which is the overall architecture diagram of the semantic segmentation model based on the deformable self-attention network provided by the embodiments of the present application. Taking the 201 target image as the input, a feature map Z is generated through the 202 Convolutional Neural Networks (CNN), and the feature map Z is reduced through the 203 channel reduction (Reduction) to obtain the feature map X. By inputting the feature map X into the deformable self-attention module 204, the deformable self-attention module 204 operates on it, and then the output feature H is subjected to 205 semantic segmentation (Segmentation) to obtain the segmentation map 206. Among them, the deformable self-attention module 204 uses an additional convolutional layer 2041 to generate M two-dimensional offsets 2043 for each query position (i.e., query point) in the feature map X. These offsets are used to point to M semantically related regions in the Key, and the size of each region is L×L. In addition, an additional region is added to each query point, where this position itself serves as the center of this region. Then, the M semantically related regions and the additional regions with each query point as the center and L as the side length are subjected to feature aggregation. Specifically, the attention map 2046 and the value 2042 are processed through the Deformable Matrix Multiplication 2047 to generate the aggregated feature 2048, and the aggregated feature 2048 is connected through the connection 206. The connected aggregated feature 2048 is subjected to semantic segmentation, and finally the segmentation map 207 is generated. In this way, the similarity score of each query point can be obtained through the Deformable Matrix Multiplication. Therefore, the proposed deformable attention module can effectively obtain dense context information by only aggregating similar features from semantically related regions.

[0076] As Figure 2 shown, the deformable self-attention module 204 uses three 1×1 convolutions 2045 of Wq, 2044 of Wk, and 2042 of Wv to convert the input feature map X ∈ R C×H×W into three different embeddings: Query: Key: and Value: V ∈ R C×H×W , and its process can be expressed as:

[0077] Q = WqX, K = WkX, V = VqX;

[0078] Among them, C represents the number of channels, H and W represent the height and width of the feature map, and then Q and K are reconstructed into V is reconstructed into R C×N , where N = H×W represents the number of all elements in the feature map, and an additional convolutional layer 2043 is used to obtain "offsets": O = WoX + bo ∈ R N*2M , where M is the offset learned for each query point. Each two-dimensional "offset" points to a square semantic-related region in the Key, and the side length of these regions is L. The learned offsets can make the semantic-related regions be located at any position in the feature map. Since the neighborhood of each query point contributes a large part of the correlation to feature aggregation, an additional semantic-related region is added for each query point, and this region is centered on the position of the query point itself. In other words, the Deformable Self-attention mechanism of the present application provides M + 1 regions for each query point and performs effective feature aggregation. It can be understood that in the current self-attention mechanism, features that are not very relevant to the query point still contribute to the weights of feature aggregation. In comparison, in the Deformable Self-attention mechanism provided by the present application, the weights used to aggregate features are all calculated from semantic-related regions, which makes the attention map generated by the embodiments of the present application smaller but more accurate.

[0079] The following combines Figure 3 The semantic segmentation method provided by the embodiments of the present application is described from the perspective of a semantic segmentation device. The semantic segmentation device can be a server or a service unit in the server, and specific limitations are not made.

[0080] Please refer to Figure 3 , Figure 3 , which is a schematic flowchart of the semantic segmentation method provided by the embodiments of the present application, including:

[0081] 301. Input the target image into the deep convolutional network to obtain the first feature map.

[0082] In this embodiment, the semantic segmentation device first obtains the target image, which is the image to be semantically segmented, such as Figure 2 201 in

[0083] 302. Perform a channel reduction operation on the first feature map to obtain a second feature map.

[0084] In this embodiment, after obtaining the first feature map, the semantic segmentation device can use a convolutional layer to reduce the number of channels of the first feature map to obtain a second feature map, which is represented by feature map X here.

[0085] 303. Determine M square regions and a target square region that are semantically related to the target query point.

[0086] In this embodiment, when the semantic segmentation device obtains the second feature map after channel reduction, it can determine M square regions and a target square region that are semantically related to the target query point. Among them, the target query point is any one of the query points corresponding to the second feature map. The side length of each of the M square regions is L, and the target square region is a square region centered on the target query point with a side length of L. Among them, the values of M and L need to satisfy (M + 1)×L 2 <H×W, L is an odd number. Generally, L is taken from {3, 5, 7, 9}, and M is taken from {11, 14, 17, 20}, where H is the height of the target image and W is the width of the target image.

[0087] In one embodiment, the semantic segmentation device determines M square regions and a target square region that are semantically related to the target query point, including:

[0088] Calculate the responses of the target query point to all elements in the M square regions and the target square region;

[0089] Determine the indices of the target query point and all elements in the M square regions and the target square region according to the responses of the target query point to all elements in the M square regions and the target square region;

[0090] Determine the offset of the target query point according to the indices of the target query point and all elements in the M square regions and the target square region and the side length L;

[0091] Determine the M square regions according to the offset of the target query point and the side length L.

[0092] The following is combined with Figure 4 for illustration. Please refer to Figure 4 , Figure 4 is a schematic diagram of the deformable matrix multiplication provided by the embodiment of the present application. As shown in Figure 4As shown in the figure, for the i-th spatial position in Query (Q) (i.e., the target query point), the semantic-related region is located according to the learned "offset" and the side length L of the semantic-related region. Specifically, for the i-th query point, the initial attention mechanism calculates H×W "relations" (i.e., similarities), while the deformable attention mechanism provided in this application calculates (M + 1)×L responses of all elements in the semantic-related region corresponding to this query point. 2 These relevant features can be stacked as: with a size of (where is the number of channels corresponding to the key), and using j ∈ {0, 1, 2,..., (M + 1)×L 2 - 1} to represent the index of the "relation", then the index of the "offset" for the i-th query point can be expressed as: Thus, the "offset" itself can be described as follows:

[0093]

[0094] where d x is the horizontal-axis offset of the i-th query point relative to the elements in each of the M + 1 square regions, and d y is the vertical-axis offset of the i-th query point relative to the elements in each of the M + 1 square regions. By applying r = j mod L 2 to represent the element index in each semantic-related region, where r ∈ {0, 1, 2,..., L 2 - 1}, their coordinate offsets from the center point can be expressed as:

[0095] where d′ x , The actual index of all relevant elements in Key is expressed as: j′ = i + (d y + d′ y )×W + (d x + d x ′).

[0096] The offset offset learned for each query point is usually a decimal. Therefore, the actual index j' of the semantic-related elements is not an integer either. Following the operation of deformable convolution, the values of K j′ and V j′ are obtained through bilinear interpolation, which makes the learned "offsets" differentiable during the training process. After learning the offset of the i-th feature point, based on the M offsets of this i-th feature point and the side length L of the region, the semantic-related region, i.e., M square regions, can be located.

[0097] As Figure 4 shown, according to the learned offsets and the side L of the semantically relevant region, locate the corresponding semantically relevant region for the i-th query point 4011 in Query 401 in Key 402 ( Figure 4 illustrated by taking M as 3 and L as 3 in ), then, stack the semantically relevant elements in the semantically relevant region (Stack Features) 403, and obtain a feature 404 with a size of

[0098] 304. Aggregate the features of the M square regions and the features of the target square region to obtain the aggregated features.

[0099] In this embodiment, the semantic segmentation device calculates the similarity between the target query point and each of the M square regions and the target square region to obtain the weights of the target query point and each of the M square regions and the target square region; then, normalize the weights of each region, and perform weighted summation of the normalized weights of each region and the positions of each region to determine the third feature map. Among them, when the semantic segmentation device calculates the similarity between the target query point and each of the M square regions and the target square region to obtain the weights of the target query point and each of the M square regions and the target square region, it can first determine the actual indices of the elements related to the target query point in the M square regions and the target square region; then, based on the actual indices, calculate the similarity between the target query point and each of the M square regions and the target square region to obtain the weights of the target query point and each of the M square regions and the target square region. The following is a specific description:

[0100] Obtain the similarity score in the Deformable self-attention (DSA) mechanism through deformable matrix multiplication Specifically, e ij refers to the similarity score between the i-th query point and the j'-th element in Key, and is calculated by the following formula:

[0101]

[0102] where i ∈ {0, 1, 2,..., N - 1} and j ∈ {0, 1, 2,..., (M + 1) × L 2If it is {-1}, then the calculation of the attention map in the proposed deformable self-attention module is as follows:

[0103]

[0104] Among them, all the weights for aggregating similar features are obtained by applying the softmax function along each row.

[0105] Finally, the aggregated feature E ∈ R N×C The vector of the i-th feature on is calculated by the weighted sum of all (M + 1) × L 2 positions in Value V:

[0106]

[0107] Among them, c ∈ {0, 1, 2,..., C} represents the index of the number of channels C in Value V.

[0108] 305. Determine the output feature according to the first feature map and the aggregated feature.

[0109] In this embodiment, after obtaining the aggregated feature, the semantic segmentation device can determine the output feature according to the first feature map and the aggregated feature. Specifically, the third feature map can be determined according to the aggregated feature, and the scale parameter can be determined at the same time. Then, the third feature map and the first feature map are calculated through the scale parameter to obtain the output feature. The following is a specific description: The semantic segmentation device can reconstruct the aggregated feature E into R H×W×C , and obtain the output feature H through the following consensus calculation:

[0110] H = αE + X;

[0111] Among them, α is a scale parameter.

[0112] 306. Perform semantic segmentation on the fused output feature to obtain a semantic segmentation map corresponding to the target image.

[0113] In this embodiment, after obtaining the output feature, the semantic segmentation device can fuse the output feature H, that is, use several convolutional layers for feature fusion, and perform semantic segmentation on the fused output feature to obtain a semantic segmentation map corresponding to the target object, such as Figure 2 the semantic segmentation map 207 in.

[0114] In summary, it can be seen that in the embodiment provided by this application, by determining M regions semantically related to the target query point, and the side length of the M regions is L, the most relevant context information at each position is adaptively aggregated. The computational complexity corresponding to the self-attention calculation mechanism provided by this application is O(HW((M + 1)L 2) The values of M and L can be flexibly set, and the computational complexity can be significantly lower than that of the existing self-attention mechanism. Furthermore, the duration of semantic segmentation can be reduced, and the accuracy of semantic segmentation can be improved.

[0115] The present application has been described above from the perspective of the semantic segmentation method. Next, the present application will be described from the perspective of the semantic segmentation device.

[0116] Please refer to Figure 5 , Figure 5 which is a schematic virtual structure diagram of a semantic segmentation device provided by an embodiment of the present application, including:

[0117] A first determination unit 501, configured to input a target image into a deep convolutional network to obtain a first feature map, where the deep convolutional network is constructed in a fully convolutional manner, and the target image is an image to be semantically segmented;

[0118] A channel reduction unit 502, configured to perform a channel reduction operation on the first feature map to obtain a second feature map, where the number of channels of the second feature map is less than the number of channels of the first feature map;

[0119] A second determination unit 503, configured to determine M square regions and a target square region that are semantically related to a target query point, where the target query point is any one of the query points corresponding to the second feature map, the side length of each of the M square regions is L, and the target square region is a square region centered on the target query point with a side length of L, where the values of M and L need to satisfy (M + 1) × L 2 < H × W, and L is an odd number, H is the height of the target image, and W is the width of the target image;

[0120] An aggregation unit 504, configured to aggregate the features of the M square regions and the features of the target square region to obtain an aggregated feature;

[0121] A third determination unit 505, configured to determine an output feature according to the first feature map and the aggregated feature;

[0122] A segmentation unit 506, configured to perform semantic segmentation on the fused output feature to obtain a semantic segmentation map corresponding to the target image.

[0123] Optionally, the second determination unit 503 is specifically configured to:

[0124] Calculate the responses of the target query point to all elements in the M square regions and the target square region;

[0125] Determine the indexes of the target query point with respect to all elements in the M square regions and the target square region based on the response of the target query point with respect to the M square regions and all elements in the target square region;

[0126] Determine the offset of the target query point based on the indexes of the target query point with respect to the M square regions and all elements in the target square region and the side length L;

[0127] Determine the M square regions based on the offset of the target query point and the side length L.

[0128] Optionally, the aggregation unit 504 is specifically configured to:

[0129] Calculate the similarity between the target query point and each of the M square regions and the target square region to obtain the weights of the target query point with respect to each of the M square regions and the target square region;

[0130] Perform normalization processing on the weights of each region;

[0131] Perform weighted summation on the normalized weights of each region and the positions of each region to obtain the aggregated feature.

[0132] Optionally, when the aggregation unit 504 calculates the similarity between the target query point and each of the M square regions and the target square region to obtain the weights of the target query point with respect to each of the M square regions and the target square region, it includes:

[0133] Determine the actual indexes of the elements related to the target query point in the M square regions and the target square region;

[0134] Based on the actual indexes, calculate the similarity between the target query point and each of the M square regions and the target square region to obtain the weights of the target query point with respect to each of the M square regions and the target square region.

[0135] Optionally, the third determination unit 505 is specifically configured to:

[0136] Determine the third feature map based on the aggregated feature;

[0137] Determine the scaling parameter;

[0138] Calculate the third feature map and the first feature map through the scaling parameter to obtain the output feature.

[0139] In summary, it can be seen that in the embodiments provided by the present application, by determining M regions semantically related to the target query point, and the side length of the M regions is L, the most relevant context information at each position is adaptively aggregated. The computational complexity corresponding to the self-attention calculation mechanism provided by the present application is O(HW((M + 1)L 2 ), and the values of M and L can be flexibly set. The computational complexity can be significantly lower than that of the existing self-attention mechanism, thereby reducing the duration of semantic segmentation and improving the accuracy of semantic segmentation.

[0140] Figure 6 FIG. 6 is a schematic structural diagram of a server provided by an embodiment of the present application. The server 600 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 622 (for example, one or more processors) and a memory 632, and one or more storage media 630 (for example, one or more mass storage devices) for storing application programs 642 or data 644. Among them, the memory 632 and the storage media 630 may be transient storage or persistent storage. The program stored in the storage media 630 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 622 may be configured to communicate with the storage media 630 and execute a series of instruction operations in the storage media 630 on the server 600.

[0141] The server 600 may further include one or more power supplies 626, one or more wired or wireless network interfaces 650, one or more input / output interfaces 658, and / or one or more operating systems 641, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, and so on.

[0142] The steps performed by the semantic segmentation device in the above embodiments may be based on the Figure 6 shown server structure.

[0143] The embodiment of the present application further provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, the steps of the above-mentioned semantic segmentation method are implemented.

[0144] The embodiment of the present application further provides a processor, and the processor is used to run a program, wherein when the program runs, the steps of the above-mentioned semantic segmentation method are executed.

[0145] An embodiment of the present application also provides a terminal device, which includes a processor, a memory, and a program stored on the memory and executable on the processor. The program code is loaded and executed by the processor to implement the steps of the above-mentioned semantic segmentation method.

[0146] The present application also provides a computer program product, which is suitable for executing the steps of the above-mentioned semantic segmentation method when executed on a data processing device.

[0147] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0148] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and modules can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0149] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0150] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products of the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one or more flows in the flowchart and / or one or more blocks in the block diagram.

[0151] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in one or more flows in the flowchart and / or one or more blocks in the block diagram.

[0152] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, causing a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one flow or more flows of the flowchart and / or one block or more blocks of the block diagram.

[0153] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0154] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0155] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storage of information such as computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.

[0156] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the element.

[0157] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0158] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A semantic segmentation method, characterized in that, Comprising: Inputting a target image into a deep convolutional network to obtain a first feature map, wherein the deep convolutional network is constructed in a fully convolutional manner, and the target image is an image to be semantically segmented; Performing a channel reduction operation on the first feature map to obtain a second feature map, where the number of channels of the second feature map is less than that of the first feature map; Determine M square regions and a target square region that are semantically related to the target query point. The target query point is any one of the query points corresponding to the second feature map. The side length of each square region in the M square regions is L, and the target square region is a square region centered on the target query point with a side length of L. Among them, the values of M and L need to satisfy (M + 1) × L 2 < H × W, and L is an odd number, H is the height of the target image, and W is the width of the target image; the M square regions are determined according to M two-dimensional offsets generated by the convolutional layer for the target query point in the second feature map; the M two-dimensional offsets are used to point to the M square regions in Key that are semantically related to the target query point; Aggregating the features of the M square regions and the features of the target square region to obtain an aggregated feature; Determining an output feature based on the first feature map and the aggregated feature, and fusing the output feature; Performing semantic segmentation on the fused output feature to obtain a semantic segmentation map corresponding to the target image.

2. The method according to claim 1, wherein The determining the M square regions semantically related to the target query point includes: Determining the M square regions according to the M two-dimensional offsets corresponding to the target query point and the side length L.

3. The method according to claim 1, wherein The aggregating the features of the M square regions and the features of the target square region to obtain an aggregated feature includes: Calculating the similarity between the target query point and each of the M square regions and the target square region to obtain the weights of the target query point and each of the M square regions and the target square region; Performing a normalization process on the weights of each region; Performing a weighted sum of the normalized weights of each region and the positions of each region to obtain an aggregated feature.

4. The method according to claim 3, wherein The calculating the similarity between the target query point and each of the M square regions and the target square region to obtain the weights of the target query point and each of the M square regions and the target square region includes: Determining the actual indices of the elements related to the target query point in the M square regions and the target square region; Based on the actual indices, calculating the similarity between the target query point and each of the M square regions and the target square region to obtain the weights of the target query point and each of the M square regions and the target square region.

5. The method according to claim 1, characterized in that, The determining the output feature according to the first feature map and the aggregated feature includes: Determining a third feature map according to the aggregated feature; Determining a scaling parameter; Calculating the third feature map and the first feature map through the scaling parameter to obtain the output feature.

6. A semantic segmentation device, characterized in that, Comprising: A first determining unit for inputting a target image into a deep convolutional network to obtain a first feature map, wherein the deep convolutional network is constructed in a fully convolutional manner, and the target image is an image to be semantically segmented; A channel reduction unit for performing a channel reduction operation on the first feature map to obtain a second feature map, where the number of channels of the second feature map is less than that of the first feature map; A second determination unit, configured to determine M square regions and a target square region that are semantically related to a target query point, where the target query point is any one of the query points corresponding to the second feature map, the side length of each square region in the M square regions is L, and the target square region is a square region with side length L centered on the target query point. Among them, the values of M and L need to satisfy (M + 1) × L 2 < H × W, and L is an odd number, H is the height of the target image, and W is the width of the target image; the M square regions are determined according to M two-dimensional offsets generated by the convolutional layer for the target query point in the second feature map; the M two-dimensional offsets are used to point to M square regions in Key that are semantically related to the target query point; An aggregation unit for aggregating the features of the M square regions and the features of the target square region to obtain an aggregated feature; A third determination unit, configured to determine an output feature according to the first feature map and the aggregated feature; the apparatus is further configured to fuse the output feature; A segmentation unit, configured to perform semantic segmentation on the fused output feature to obtain a semantic segmentation map corresponding to the target image.

7. The device according to claim 6, characterized in that, The second determination unit is specifically configured to: Determine the M square regions according to the M two-dimensional offsets corresponding to the target query point and the side length L.

8. The device according to claim 6, characterized in that, The aggregation unit is specifically configured to: Calculate the similarity between the target query point and each of the M square regions and the target square region to obtain the weights of the target query point and each of the M square regions and the target square region; Perform normalization processing on the weights of each region; Perform weighted summation on the weights of each region after normalization processing and the positions of each region to obtain an aggregated feature.

9. The device according to claim 8, characterized in that, The aggregation unit calculates the similarity between the target query point and each of the M square regions and the target square region to obtain the weights of the target query point and each of the M square regions and the target square region, including: Determine the actual indexes of the elements related to the target query point in the M square regions and the target square region; Based on the actual indexes, calculate the similarity between the target query point and each of the M square regions and the target square region to obtain the weights of the target query point and each of the M square regions and the target square region.

10. The device according to claim 6, wherein The third determination unit is specifically configured to: Determine a third feature map according to the aggregated feature; Determine a scale parameter; Calculate the third feature map and the first feature map through the scale parameter to obtain the output feature.

11. A terminal device, characterized in that, Comprising a processor, a memory, and a program stored on the memory and executable on the processor, the program is loaded and executed by the processor to implement the steps of the semantic segmentation method according to any one of claims 1 to 5 above.

12. A computer-readable storage medium, characterized in that, Comprising instructions, when the instructions run on a computer, causing the computer to execute the steps of the semantic segmentation method according to any one of claims 1 to 5 above.

13. A computer program product, characterized in that, Comprising a program, when executed on a data processing device, adapted to execute the steps of the semantic segmentation method according to any one of claims 1 to 5 above.

Citation Information

Patent Citations

  • Image classification model training method and image processing method and device

    CN109784424A

  • Image classification method and device, storage medium and equipment

    CN110210572A