Semantic segmentation method and related product
Through the method of combining semantic encoder and image encoder with attention map fusion layer, the problem of limited recognition ability of traditional semantic segmentation models in unknown categories is solved, and higher semantic segmentation accuracy and distinction ability of confusing categories are achieved.
Patent Information
- Application Number
- CN202510607864.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-13
AI Technical Summary
Traditional semantic segmentation models have the problem of limited recognition capabilities when dealing with unknown categories, especially in real-world application scenarios, and obtaining pixel-level tags for unknown categories can be difficult and expensive.
The semantic encoder is used to convert multiple tags into multiple text representations one by one, and the target image is converted into image representation through the image encoder. Combined with the attention map fusion layer, the label of each pixel is determined.
It improves the accuracy of semantic segmentation, can effectively deal with unknown categories, enhances the ability to distinguish easily confused categories, and thus achieves better semantic segmentation performance.
Smart Images

Figure CN120125828A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly relates to a semantic segmentation method and related products. Background Art
[0002] The semantic segmentation task refers to achieving pixel-level classification of image data, that is, given a set of semantic labels, assigning a semantic label to each pixel in the image. After being trained, traditional semantic segmentation models only have the ability to recognize and segment a limited number of known categories with labels provided during training. However, in real-world application scenarios, there are still many unknown categories without labels provided during training, and obtaining pixel-level labels for these categories may be difficult and expensive. Therefore, traditional semantic segmentation models are greatly limited in practical applications.
[0003] Chinese invention patent CN106295706B proposes an image automatic segmentation and semantic annotation method based on a shape vision knowledge base, which realizes automatic segmentation and semantic annotation of unknown shapes through the shape vision knowledge base.
[0004] Chinese invention patent application CN117710784A proposes an unknown category object detection method and system. The method mainly includes generating pseudo-label information for unknown category objects in each image sample in the training sample set according to the constructed semantic image segmentation model; training the category-related object detection head and the category-unrelated object detection head respectively; then obtaining the target image to be detected, and respectively extracting and fusing all category-related features and all category-unrelated features in the target image to be detected; inputting the fused category-related features into the trained category-related object detection head to obtain the detection result of known category objects, and inputting the fused category-unrelated features into the trained category-unrelated object detection head to obtain the detection result of foreground objects; obtaining the detection result of unknown category objects in the target image according to the detection result of foreground objects and the detection result of known category objects, and this method can effectively improve the detection accuracy of unknown category objects.
[0005] The present invention provides a new semantic segmentation method and related products. Summary of the Invention
[0006] The present invention provides a semantic segmentation method and related products to improve the accuracy of semantic segmentation.
[0007] The present invention achieves the above object through the following solutions.
[0008] The present invention provides a semantic segmentation method capable of processing unknown categories, including: Using a semantic encoder to convert multiple labels into corresponding multiple text representations, and each label represents the target category name of the pixel region segmented by semantic segmentation; Convert the target image into a corresponding image representation using an image encoder; Determine the label of each pixel in the target image based on the image representation and the multiple text representations; Among them, the image encoder includes f + 1 Transformer blocks connected in sequence, and each Transformer block contains a self-attention mechanism layer, f ≥ 2, and an attention map fusion layer is inserted into the self-attention mechanism layer of the last Transformer block; Among them, the attention map fusion layer is used for: Judge the number of the Transformer block in which the global block first appears in the attention maps calculated by the previous f Transformer blocks, and this number is denoted as g ; Fuse the attention maps calculated by at least two of the g-th Transformer block to the last Transformer block to obtain a fused attention map, where the information fusion operation must select the attention map calculated by the last Transformer block; Among them, the self-attention mechanism layer of the last Transformer block is used to calculate the image representation output by the image encoder according to its V matrix and the fused attention map.
[0009] In some embodiments, the attention maps calculated by the previous f Transformer blocks are Q-K attention maps; The attention map fusion layer uses the following function to judge whether a global block appears in the attention map: ; Among them, = 1 indicates that a global block appears in the attention map calculated by the Transformer block numbered i , = 0 indicates that no global block appears in the attention map calculated by the Transformer block numbered i , indicates the i -th attention vector of the attention map calculated by the Transformer block numbered j , is a preset constant used to prevent the product of the attention vectors from being too small and exceeding the storage range of the computer.
[0010] In some embodiments, in the step of converting a target image into a corresponding image representation by using an image encoder, image representations of multiple pixel blocks segmented from the target image are obtained; Determining a label for each pixel in the target image according to the image representation and the multiple text representations includes: Calculating the approximation degree between the image representation of each pixel block and each text representation; Using the bilinear interpolation method to obtain the approximation degree between each pixel point in each pixel block and each text representation, and further determining the label corresponding to each pixel point.
[0011] In some embodiments, the image representation output by the self-attention mechanism layer of the last Transformer block is denoted as Z, ; Proj(·) is a projection function, is a fused attention map, is the Value matrix of the self-attention mechanism layer of the last Transformer block.
[0012] In some embodiments, the attention map calculated by the last Transformer block is a Q-Q attention map.
[0013] In some embodiments, the first f Transformer blocks each include: a normalization layer, a self-attention mechanism layer, a residual layer, a normalization layer, a feed-forward neural network layer, and a residual layer, which are sequentially arranged.
[0014] In some embodiments, the fused attention map is the mean of the attention maps calculated by the g-th Transformer block to the last Transformer block; or the fused attention map is the mean of the attention maps calculated by a continuously set number of Transformer blocks starting from the g-th Transformer block and the last Transformer block.
[0015] The present invention provides an electronic device, including a memory and a processor, where the memory stores a program, and the processor runs the program to execute the above method.
[0016] The present invention provides a computer program product, which executes the above method when running on a processor.
[0017] The present invention provides a storage medium, where instructions are stored on the storage medium, and the instructions execute the above method when running.
[0018] In the attention mechanism of the last Transformer block in the visual encoder, each block in the attention map can interact with itself and adjacent blocks, thereby improving the distinguishability between block-level features and integrating image-level global features from global blocks. Global context information can provide global understanding at the image level in semantic segmentation tasks. In this pixel-level classification task of semantic segmentation, it can help distinguish some easily confused categories, and thus achieve better semantic segmentation performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is the information flow diagram of the semantic segmentation method of the present invention that can handle unknown categories.
[0020] Figure 2 is the structural block diagram of the electronic device of the present invention.
[0021] Figure 3 is the network structure of the first visual encoder in the test example of the present invention.
[0022] Figure 4 is the network structure of the last visual encoder in the test example of the present invention.
[0023] Figure 5 is the semantic segmentation result of the test example of the present invention before model transformation.
[0024] Figure 6 is the semantic segmentation result of the test example of the present invention after model transformation. DETAILED DESCRIPTION OF THE INVENTION
[0025] The present invention will be further described below in conjunction with specific embodiments, but the protection scope of the present invention is not limited thereto.
[0026] In combination with Figure 1 , the present invention provides a semantic segmentation method, including: Using a semantic encoder to convert multiple labels into corresponding multiple text representations, and each label represents the target category name of the pixel region segmented by semantic segmentation; After dividing the target image into blocks, inputting it into an image encoder to obtain the image representation corresponding to each image block of the target image; Determining the label of each pixel in the target image based on the image representation and the multiple text representations; Wherein, the image encoder includes f + 1 Transformer blocks connected in sequence, and each Transformer block includes a self-attention mechanism layer, f ≥ 2, and an attention map fusion layer is inserted into the self-attention mechanism layer of the last Transformer block; Among them, the attention map fusion layer is used for: Determine the f Transformer block number of the first appearance of the global block in the attention maps calculated by the previous g Transformer blocks, and this number is denoted as ; Fuse the attention maps calculated from at least two of the g-th to the last Transformer blocks to obtain a fused attention map, where the information fusion operation must select the attention map calculated by the last Transformer block; Among them, the self-attention mechanism layer of the last Transformer block is used to calculate the image representation output by the image encoder according to its V matrix and the fused attention map.
[0027] The semantic encoder and the image encoder are pre-trained and are aligned in the representation space.
[0028] In the attention map of the deep Transformer block, a set of special pixel blocks will appear. All pixel blocks in the attention map have a high attention score for this set of special pixel blocks, indicating that all pixel blocks have greatly integrated the information of this set of special pixel blocks. This set of special pixel blocks is called the global block, and it is considered that they have better integrated the global context information at the image level.
[0029] In the attention mechanism of the last Transformer block in the visual encoder, each pixel block in the attention map can interact with itself and adjacent pixel blocks, thereby improving the distinguishability between pixel block features and integrating the global features at the image level from the global block. Global context information can provide a global understanding of the image in the semantic segmentation task. In this pixel-level classification task of semantic segmentation, it can help distinguish some easily confused categories and thus achieve better semantic segmentation performance.
[0030] In some embodiments, the attention maps calculated by the previous f Transformer blocks are Q-K attention maps; The attention map fusion layer uses the following function to determine whether a global block appears in the attention map: ; Among them, =1 indicates that a global block appears in the attention map calculated by the Transformer block numbered i , =0 indicates that no global block appears in the attention map calculated by the Transformer block numbered i , Indicates the i -th attention vector of the attention map calculated by the Transformer block numbered j , is a preset constant used to prevent the product of attention vectors from being too small and exceeding the storage range of the computer. Here, the vector multiplication is dot product.
[0031] The Q-K attention map is the matrix obtained by multiplying the Query matrix by the transpose of the Key matrix and then processing it through the Softmax function. The specific formula is as follows: , where is the Q-K attention map, is the Query matrix, is the Key matrix, is the dimension of the Query matrix.
[0032] In some embodiments, takes values between 50 and 200. The typical value of
[0033] In some embodiments, the image representation output by the self-attention mechanism layer of the last Transformer block is denoted as Z, ; Proj(·) is a projection function, is the fused attention map, is the Value matrix of the self-attention mechanism layer of the last Transformer block.
[0034] In some embodiments, the attention map calculated by the last Transformer block is the Q-Q attention map.
[0035] The Q-Q attention map is the matrix obtained by multiplying the Query matrix by the transpose of the Query matrix and then processing it through the Softmax function. This enables each pixel block to mainly focus on itself in the attention mechanism.
[0036] That is , is the attention map calculated by the last Transformer block, is the dimension of the Query matrix, and Q is the Query matrix.
[0037] In some embodiments, the fused attention map is the mean of the attention maps calculated from the g-th Transformer block to the last Transformer block.
[0038] In other embodiments, the fused attention map is the weighted sum of the attention maps calculated from the g-th Transformer block to the last Transformer block. The weight coefficients increase sequentially according to the connection order of the Transformer blocks.
[0039] In another embodiment, the fused attention map is the mean of the attention maps calculated from the g-th Transformer block for a continuously set number of Transformer blocks and the last Transformer block.
[0040] For example, the fused attention map is the mean of the attention maps calculated from the g-th Transformer block and the last Transformer block.
[0041] For another example, the fused attention map is the mean of the attention maps calculated from the g-th Transformer block, the (g + 1)-th Transformer block, and the last Transformer block.
[0042] The present invention does not limit the operating mechanism of the semantic encoder. In some embodiments, the semantic encoder operates as follows.
[0043] Provide multiple labels, each label representing the target category name of the pixel region segmented semantically.
[0044] The target category name is, for example: any target category name that needs to be distinguished, such as person, skateboard, grassland, sofa, chair, etc.
[0045] Construct the multiple labels into corresponding target texts respectively according to a preset template.
[0046] The preset template is, for example: "a picture of *", where * represents the label. Then the target text corresponding to the label "person" is "a picture of a person", the target text corresponding to the label "skateboard" is "a picture of a skateboard", and the target text corresponding to the label "grassland" is "a picture of a grassland".
[0047] Perform a segmentation process on each target text to obtain a continuous plurality of words corresponding to each target text.
[0048] For example, the target text "a picture of a person" is segmented into 4 words: "a", "person", "of", "a picture".
[0049] The text segmentation process can adopt existing mature segmentation methods, and the present invention does not limit this.
[0050] Input the continuous plurality of words corresponding to each target text into the text encoder to obtain the encoded text representation.
[0051] The text encoder can adopt existing and mature text encoders. The present invention does not limit the type of the text encoder, nor does it limit the training method of the text encoder.
[0052] After text encoding, multiple consecutive words are converted into a text representation, which is represented as a vector.
[0053] Encoding the target text (i.e., a complete sentence) is beneficial to enriching semantic expression.
[0054] In some embodiments, the visual encoder first preprocesses the target image to make the target image conform to the set format representation, then inputs the preprocessed target image into the embedding layer, and then inputs the image patches output by the embedding layer into the first Transformer block.
[0055] In the preprocessing operation, first, the target image is scaled proportionally, that is, the aspect ratio of the target image is kept unchanged, and by scaling, the long side or the short side of the target image reaches the specified size, and then appropriate values (such as black pixels) are filled in the insufficient parts. Subsequently, pixel value normalization is performed on the target image. In the original RGB image data, the pixel values of each color component are in the range of 0 to 255. The normalization method converts the pixel values of each color into a normal distribution with a mean of 0 and a standard deviation of 1. The specific operation is to first calculate the mean and standard deviation of the pixel values of each color of the target image, and then subtract the corresponding mean from each pixel value and divide by the corresponding standard deviation. The purpose of this is to make the model training process more stable and efficient. Because the pixel value ranges and distributions of different images may be different, normalization can unify the pixel value distributions of all images into a similar range, facilitating the model to learn the features of the images and avoiding certain features from being overemphasized or ignored due to differences in pixel value ranges.
[0056] The embedding layer divides the input target image into pixel patches of a fixed size. For example, for an image of size (where is the height, is the width, and the number of channels of the RGB image is 3), it will be divided into N pixel patches of size . The side length of the P pixel patch. These pixel patches are mapped into N serialized vectors of D dimensions through linear embedding (a fully connected layer), and each serialized vector is added to the corresponding position encoding vector to form , where is the serialized vector containing position information corresponding to the i-th pixel patch, .
[0057] frontf The structure of a Transformer block can be designed according to the prior art, and the present invention does not limit this. The former f each of the Transformer blocks only contains one self-attention mechanism layer.
[0058] In some embodiments, referring to Figure 3 , the former f each of the Transformer blocks includes: a normalization layer, a self-attention mechanism layer, a residual layer, a normalization layer, a feed-forward neural network layer, and a residual layer, which are arranged in sequence.
[0059] Both normalization layers are layer normalization.
[0060] The first residual layer sums the input of the first normalization layer and the output of the self-attention mechanism layer.
[0061] The second residual layer sums the input of the second normalization layer and the output of the feed-forward neural network layer.
[0062] The feed-forward neural network layer is the first fully connected layer, an activation layer, and the second fully connected layer connected in sequence.
[0063] The following takes the first Transformer block as an example for illustration.
[0064] Input the serialized vector into the first normalization layer to obtain the normalized serialized vector.
[0065] For the input vector , its mean value and standard deviation will be calculated, and then the input vector is converted to: This step helps to stabilize the training process, reduce the problem of gradient vanishing or explosion, and make the distribution of the input data more stable, which is beneficial to subsequent operations.
[0066] Input the serialized vector output by the first normalization layer into the self-attention mechanism layer to obtain the serialized vector output by the self-attention mechanism layer.
[0067] The operation of the self-attention mechanism layer is as follows: for each input normalized serialized vector, three different linear layers are used to generate a Query (Q) matrix, a Key (K) matrix, and a Value (V) matrix. Among them, the attention map A is obtained through Query (Q) and Key (K), ; where Let \(d\) represent the dimension of the Key matrix, \(Q\) represent the Query matrix, and \(K\) represent the Key matrix. Using the attention map \(A\) as weights, the Value (V) matrix is weighted to obtain the serialized vector output by the self-attention mechanism layer.
[0068] The serialized vector input to the first normalization layer and the serialized vector output by the self-attention mechanism layer are subjected to residual connection, that is, vector addition operation, to obtain the serialized vector of the residual connection.
[0069] Residual connection helps prevent the vanishing gradient, enabling the network to be deeper, and also allowing the network to learn more complex features because it allows information to be directly passed from the input layer to subsequent layers, rather than relying solely on the transformation of the current layer (such as the self-attention mechanism layer here).
[0070] The serialized vector of the residual connection is input to the second normalization layer to obtain the normalized serialized vector of the residual connection. The operation of the second normalization layer is the same as that of the first normalization layer.
[0071] The normalized serialized vector of the residual connection is input to the feed-forward neural network layer to obtain the serialized vector output by the feed-forward neural network layer.
[0072] The feed-forward neural network layer has two fully-connected layers and one activation layer (the activation layer is located between the two fully-connected layers).
[0073] The function of the first fully-connected layer of the feed-forward neural network layer is to map the input features (i.e., the normalized serialized vector of the residual connection) to a higher-dimensional space, providing more abundant representation ability for subsequent non-linear transformations. This helps the network learn more complex feature combinations and relationships because higher dimensions can store more information, enabling the network to capture more subtle relationships between input features.
[0074] The high-dimensional vector output by the first fully-connected layer is input to the activation layer. The function of the activation layer is to perform non-linear mapping on the linearly transformed high-dimensional vector, allowing the network to learn non-linear feature relationships. Through non-linear activation, the network can better fit various complex functions.
[0075] The function of the second fully-connected layer of the feed-forward neural network layer is to map the high-dimensional features after dimension expansion and non-linear transformation back to the original dimension, or to a dimension suitable for the next processing stage. This ensures that the dimension of the output features is compatible with the subsequent operations of the network.
[0076] The vector addition of the serialized vector of the residual connection and the output of the feed-forward neural network layer is performed to obtain the output of the current Transformer block.
[0077] In some embodiments, refer toFigure 4 , the last Transformer block only includes a normalization layer and a self-attention mechanism layer connected in sequence, and an attention map fusion layer is inserted into the self-attention mechanism layer.
[0078] The following is an example to illustrate the complete operation process of the image encoder.
[0079] The preprocessed target image is divided into 9 pixel blocks in 3 rows and 3 columns. After being processed by the embedding layer, 9 serialized vectors are obtained. After these 9 serialized vectors are input into the first normalization layer, 9 normalized serialized vectors are also obtained. Then, after these 9 normalized serialized vectors are input into the self-attention mechanism layer, assuming single-head attention, 9 serialized vectors are still obtained. These 9 serialized vectors are sequentially processed by the second normalization layer and the feed-forward neural network layer to obtain 9 serialized vectors. These 9 serialized vectors are then summed one by one with the 9 serialized vectors received by the second normalization layer to obtain 9 serialized vectors.
[0080] The self-attention mechanism layer of the last Transformer block still outputs 9 serialized vectors. These 9 serialized vectors respectively correspond to the 9 pixel blocks into which the target image is divided, representing the D-dimensional image features extracted for each pixel block. The similarity is calculated for each pixel block and the text representation. Operationally, it is a matrix multiplication of the 9×D image features and the C×D text representations of C texts.
[0081] In some embodiments, in the step of using the image encoder to convert the target image into a corresponding image representation, the image representations of the multiple pixel blocks segmented from the target image are obtained; Determining the label of each pixel in the target image according to the image representation and the multiple text representations includes: Calculating the approximation degree between the image representation of each pixel block and each text representation; Using the bilinear interpolation method to obtain the approximation degree between each pixel point in each pixel block and each text representation, and further determining the label corresponding to each pixel point.
[0082] The image features of each image block are D-dimensional, that is, a 1×D vector. The 10 labels correspond to a 10×D text representation matrix. The matrix multiplication of the image features and the transpose of the text representation matrix results in a 1×10 vector, representing the similarity between the current image block and the 10 labels respectively. Each image block will have a similar similarity matrix.
[0083] Currently, the obtained similarity is at the pixel-block level, that is, the similarity of each pixel block with these 10 labels respectively. In order to obtain the similarity matrix at the pixel level, a bilinear interpolation operation is performed to obtain the similarity matrix at the pixel level. For each pixel position, the argmax function is used to take the label corresponding to the maximum similarity as the segmentation result of the current pixel.
[0084] Based on the same inventive concept, referring to Figure 2 , the present invention also provides an electronic device, including a memory and a processor, the memory stores a program, and the processor runs the program to execute the above method.
[0085] Based on the same inventive concept, the present invention also provides a computer program product, which executes the above method when running on a processor.
[0086] Based on the same inventive concept, the present invention also provides a storage medium, and instructions are stored on the storage medium, and the instructions execute the above method when running.
[0087] The processor is, for example, any known processor type such as a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), or a combination thereof.
[0088] The memory and the storage medium are, for example, various media that can store program codes such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.
[0089] A test example is provided below.
[0090] A multi-modal pre-training model is provided for Lan M, Chen C, Ke Y, et al. Clearclip: Decomposing clip representations for dense vision-language inference[C] / / European Conference on Computer Vision. Springer, Cham, 2025: 143-160. In the visual encoder of this model, the last Transformer block multiplies the self-computed attention map with the Value matrix to obtain the image representation.
[0091] In this test example, the structures of the first f Transformer blocks are all Figure 3In the structure shown, the attention map calculated by the self-attention mechanism layer is the Q-K attention map.
[0092] In this test case, the attention map calculated by the self-attention mechanism layer of the last Transformer block is the Q-Q attention map.
[0093] The structure of the last Transformer block in this test case before modification is Figure 4 different from the structure shown only in that: an attention map fusion layer is not inserted in the self-attention mechanism layer.
[0094] Figure 5 This is the test result of the multi-modal pre-training model. The local area in the chair region (red) is misidentified as a sofa (green area).
[0095] An attention map fusion layer is added to the last Transformer block of the multi-modal pre-training model. The fused attention map is the mean of the attention maps calculated by the g-th Transformer block, the (g + 1)-th Transformer block, and the last Transformer block. The model parameters of the remaining modules of the multi-modal pre-training model remain unchanged. The fused attention map is multiplied by the Value matrix to obtain the image representation.
[0096] Figure 6 This is the test result after modification. After modifying the multi-modal pre-training model, the error of confusing a chair with a sofa is avoided.
[0097] The protection scope of the present invention is not limited to the above embodiments. Obviously, those skilled in the art can make various changes and deformations to the present invention without departing from the scope and spirit of the present invention. If these changes and deformations fall within the scope of the claims of the present invention and their equivalent technologies, the intention of the present invention also includes these changes and deformations.
Claims
1. A semantic segmentation method, characterized in that: include: A semantic encoder is used to convert multiple labels into multiple text representations in one-to-one correspondence, where each label represents the target category name of the pixel area obtained by semantic segmentation; An image encoder is used to convert the target image into a corresponding image representation; Determine a label for each pixel in the target image according to the image representation and the plurality of text representations; The image encoder comprises sequentially connected f +1 Transformer block, each of which contains a self-attention mechanism layer, f ≥2, an attention graph fusion layer is inserted into the self-attention mechanism layer of the last Transformer block; The attention graph fusion layer is used to: Before judgment f The Transformer block number where the global block first appears in the attention map calculated by the Transformer blocks is recorded as g ; The first g The information fusion operation must use the attention map calculated by the last Transformer block; Among them, the self-attention mechanism layer of the last Transformer block is used to calculate the image representation output by the image encoder according to its V matrix and the fused attention map.
2. The semantic segmentation method according to claim 1, characterized in that: forward f The attention map calculated by the Transformer block is the QK attention map; The attention map fusion layer uses the following function to determine whether a global block appears in the attention map: ; in, =1 means the number is i The global block appears in the attention map calculated by the Transformer block. =0 means the number is i There is no global block in the attention map calculated by the Transformer block. Indicates the number i The attention map calculated by the Transformer block is j attention vector, It is a preset constant used to prevent the product of the attention vector from being too small and exceeding the computer's storage range.
3. The semantic segmentation method according to claim 1, characterized in that: In the step of converting the target image into a corresponding image representation using an image encoder, an image representation of each of the plurality of pixel blocks segmented from the target image is obtained; The step of determining a label of each pixel in the target image according to the image representation and the plurality of text representations comprises: Calculate the approximation between the image representation of each pixel block and each text representation; The bilinear interpolation method is used to obtain the approximation between each pixel point in each pixel block and each text representation, and then determine the label corresponding to each pixel point.
4. The semantic segmentation method according to claim 1, characterized in that: The image representation output by the self-attention mechanism layer of the last Transformer block is denoted as Z. ; Proj(·) is the projection function, is the fused attention map, is the Value matrix of the self-attention mechanism layer of the last Transformer block.
5. The semantic segmentation method according to claim 1, characterized in that: The attention map calculated by the last Transformer block is the QQ attention map.
6. The semantic segmentation method according to claim 1, characterized in that: forward f Each Transformer block includes: a normalization layer, a self-attention mechanism layer, a residual layer, a normalization layer, a feedforward neural network layer and a residual layer arranged in sequence.
7. The semantic segmentation method according to claim 1, characterized in that: The fused attention map is the average of the attention maps calculated from the g-th Transformer block to the last Transformer block; or the fused attention map is the average of the attention maps calculated from a set number of consecutive Transformer blocks starting from the g-th Transformer block and the last Transformer block.
8. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a program, and the processor runs the program to perform the method according to any one of claims 1 to 7.
9. A computer program product, characterized in that When the processor is run on a processor, the method according to any one of claims 1 to 7 is executed.
10. A storage medium, characterized in that: The storage medium stores instructions, which, when executed, execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
An Automatic Image Segmentation and Semantic Annotation Method Based on Shape Visual Knowledge Base
CN106295706B
Unknown category target detection method and system
CN117710784A
Weak supervision semantic segmentation method based on attention fusion
CN116912501A
Fine-grained image classification method based on feature fusion and semantic enhancement
CN118799646A
Method and apparatus with video object identification
US20240221356A1