High-efficiency open vocabulary-oriented panoramic segmentation method
By adopting multi-scale feature extractor, lightweight aggregator, vocabulary-aware selection module, bidirectional dynamic embedding expert and lightweight decoder in the open vocabulary panoramic segmentation method, the problems of large computing overhead and slow reasoning in the prior art are solved, and efficient and low-cost open vocabulary panoramic segmentation are achieved.
Patent Information
- Application Number
- CN202510109507.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-23
AI Technical Summary
The existing open-vocabulary panoramic segmentation method has large calculation overhead and slow inference speed. The two-stage method leads to high calculation overhead for visual feature and loss of context information. The single-stage method lacks spatial positioning capabilities, resulting in unacceptable calculation overhead and inference speed.
A panoramic segmentation method for efficient open vocabulary is adopted, including multi-scale feature extractor and lightweight aggregator for visual feature extraction and aggregation, a text encoder is used to encode any category of vocabulary, and the semantic understanding of visual aggregation features is improved based on the vocabulary perception selection module. Through bidirectional dynamic embedding experts, an instance embedding with semantic perception and spatial perception is generated, and a lightweight decoder is used for mask prediction and refinement.
It realizes low computing overhead and efficient open vocabulary panoramic segmentation, improves the inference speed, and maintains competitive performance, solving the problems of large computing overhead and slow inference speed in existing methods.
Smart Images

Figure CN120032371A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of open vocabulary segmentation of visual language models, and in particular relates to an efficient open vocabulary panoramic segmentation method. Background Art
[0002] Open vocabulary panoptic segmentation aims to segment all objects in an image based on a provided vocabulary of arbitrary categories. It mainly relies on the zero-shot capability of large-scale visual language models through cross-modal alignment pre-training. The application of open vocabulary panoptic segmentation has far-reaching implications for enhancing scene understanding in fields such as autonomous driving and robotics, and has attracted widespread research interest.
[0003] Existing open-vocabulary panoptic segmentation methods suffer from high computational overhead and slow inference speed. (1) Two-stage methods adopt a non-shared, inefficient pipeline. Class-independent masks are first generated, and then image crops obtained from these masks are processed using another visual language model backbone network to extract features for individual classification, which results in high computational overhead of visual features and loss of contextual information. (2) Single-stage methods adopt a shared, inefficient pipeline. They adopt a single shared CNN-based frozen CLIP visual encoder as the backbone to extract multi-scale features, which is suitable for image segmentation tasks of high-resolution images. However, the CNN-based CLIP backbone only gives features the ability to distinguish instances, but lacks spatial localization capabilities. This leads to them often adopting a heavyweight mask decoder composed of many transformer layers, including self-attention and cross-attention mechanisms (self-attention captures contextual information between different queries, while cross-attention focuses on specific regions of the feature map related to each query to provide perception of spatial details) to compensate for the lack of spatial perception, resulting in unacceptable computational overhead and slower inference speed. Therefore, it is of great significance to propose a new low-cost and efficient open-vocabulary panoptic segmentation method. Summary of the invention
[0004] To solve the above problems, the present invention proposes an efficient open vocabulary panoptic segmentation method.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] A method for efficient open vocabulary panoptic segmentation, comprising the following steps:
[0007] S1, visual feature extraction and aggregation based on multi-scale feature extractor and lightweight aggregator;
[0008] S2. Use the text encoder to encode any category of words to obtain text embedding Among them, N classrepresents the number of categories; D represents the number of channels of text embedding;
[0009] S3, based on the vocabulary-aware selection module, improves the semantic understanding of visual aggregation features and reduces the feature interaction burden of the mask decoder;
[0010] S4, based on bidirectional dynamic embedding experts, generates instance embeddings with semantic awareness and spatial awareness by dynamically assigning expert weights;
[0011] S5. Based on a lightweight decoder, the object kernel is used to perform mask prediction and refinement layer by layer, and the dot product of the object kernel and text embedding is used as category prediction.
[0012] Preferably, the specific process of step S1 is:
[0013] S11. In the visual feature extraction stage, the feature extractor uses the CNN-based CLIP visual encoder as the backbone network to extract the multi-scale feature map C i ,i∈{2,3,4,5};
[0014] S12, lightweight aggregator uses a modulated deformable convolutional feature pyramid structure to enhance and fuse features from different scales to obtain the intermediate FPN feature P i ,i∈{2,3,4,5}; then aggregate the features of different scales to obtain the visual aggregate feature It is used to reduce computational overhead and speed up inference. D represents the number of feature map channels, H represents the image height, and W represents the image width. The calculation formula for feature aggregation is:
[0015]
[0016] Among them, Up i represents bilinear interpolation operation; N is the number of feature layers; Conv(·) and Fuse(·) respectively represent the use of 1×1 and 3×3 convolutions to adjust the feature dimension and fuse the visual aggregation feature F. agg .
[0017] Preferably, the specific process of step S3 is:
[0018] S31, visual aggregation features F from lightweight aggregator agg and the text embedding E from the text encoder t Mapped to the same feature dimension through depth-wise separable 2D convolution and linear layers respectively;
[0019] S32, aggregate the visual features F along the channel dimension agg and text embedding E t Split into multi-head visual aggregation features and multi-head text embedding Used to process complex semantic feature relationships in different subspaces; where h represents the number of heads;
[0020] S33. Processing multi-head visual aggregation features through matrix operations and multi-head text embedding Gaining multiple attention And perform softmax function on the multi-head attention A to produce smooth logits, and select the maximum value on the vocabulary dimension to obtain the vocabulary-aware attention weight The calculation formula is:
[0021]
[0022] S34, introduce two learnable parameters, namely the scaling factor γ and the offset factor δ, to enable the network to adaptively adjust the word-aware attention weights under various semantic distributions Lexical-aware attention weights Expand along the channel dimension and then aggregate features with multi-head vision Perform element-wise multiplication to reshape the weighted features into the original aggregate feature dimension to obtain the visual semantic aggregate feature. The calculation formula of the visual semantic aggregate feature is:
[0023]
[0024] in, represents visual semantic aggregation features; Reshape(·) represents the reshaping operation.
[0025] Preferably, the specific process of step S4 is:
[0026] S41. The first block of CLIP's ViT-B visual encoder is used as a spatial perception extractor to extract spatial location information. The input image is processed by the first block of ViT-B to generate visual tags, and the visual tags are reshaped into a two-dimensional spatial dimension to obtain the visual spatial feature F. v , using two transposed convolutions to perform four-fold upsampling to obtain spatially aware features And use mask pooling to perceive the feature F from space s Extract the region of interest from
[0027] S42, introduce a weight distribution router to estimate the importance of the current expert embedding and adaptively adjust the relevant weight coefficients; generate mask embedding E through linear layers respectively m Dynamic parameters and spatially aware embedding E s Dynamic parameters Dynamic parameters are evenly split into router parameters along the channel dimension and And the fusion parameters and Where d represents the number of channels, and the calculation formula is:
[0028] (P f ,P r )=s(E m W m ),
[0029] (Q f ,Q r )=s(E s W s ),
[0030] in, W m Indicates E m The projection matrix is used to generate dynamic parameters; W s Indicates E s The projection matrix is used to generate dynamic parameters; s(·) is the splitting operation along the channel dimension;
[0031] S43, through P r and Q r The total router parameters are aggregated by the element-wise product between them; two router linear layers and sigmoid activation function are used to provide adaptive weighting capabilities for experts, so that the router adaptively assigns weights to the embedded experts. The calculation formula is:
[0032] P t =P r Q r ,
[0033] α m =σ(LN(Linear m (P t ))),
[0034] α s =σ(LN(Linear s (P t ))),
[0035] Among them, P t represents the total router parameter; α m represents the weight coefficient of mask embedding; α s Represents the weight coefficient of spatial perception embedding; Linear m (·) represents the mask embedded router linear layer; Linear s(·) denotes the linear layer of the spatially aware embedded router; LN(·) denotes the layer normalization operation; σ denotes the sigmoid activation function;
[0036] S44. Perform weighted summation of mask embedding and spatially aware embedding according to mask embedding expert weights and spatially aware embedding expert weights, and use layer normalization and GELU activation function to obtain instance embedding Used to aggregate instance and spatial information, the calculation formula is:
[0037]
[0038] Preferably, the specific process of step S5 is:
[0039] S51, the lightweight decoder consists of only three modules in the following order at each layer: initial attention module, dynamic deep attention module and late attention module; the initial attention module aggregates the visual semantic features from the vocabulary-aware selection module with the initial mask or predicted mask from the previous layer Perform dot product to extract initial attention features
[0040] S52. Using dynamic deep attention mechanism, a set of learnable object kernels With the initial attention feature F init Perform cross-dimensional feature interaction to obtain the refined object kernel The calculation formula is:
[0041]
[0042] Among them, DyDepthwiseAttention(·) represents dynamic depth-wise attention; represents the projection matrix of the object kernel K to generate the depth convolution kernel, m represents the convolution kernel size; r(·) represents the view operation that reshapes the input into N×1×m; * represents the parameter r(·) and F init 1D convolution operation for the input; the object kernel after refinement Focus on the relationship between different object cores, enrich information through multi-head self-attention and feed-forward neural network; the refined object core As a region proposal network, it provides the potential for generating mask proposals;
[0043] S53, through the three-layer perceptron, the refined object core Mapped to a mask kernel, each binary mask By aggregating features of the i-th mask kernel and visual semantics Dot product is performed to obtain; then the visual semantic aggregation features are Perform mask pooling with the predicted mask M to obtain mask embedding The lightweight decoder contains three decoding layers and repeats step S5 three times to obtain the final mask prediction and category prediction.
[0044] After adopting the above technical scheme, the present invention has the following beneficial effects: a novel single-stage, shared, efficient and spatially aware open vocabulary panoramic segmentation framework proposed by the present invention includes a lightweight aggregator and a lightweight decoder, which greatly reduces the computational overhead of the model and improves the reasoning speed; the vocabulary-aware selection module guides the visual aggregation features to select features that are more relevant to the text according to the semantic importance of the text, thereby improving the semantic understanding of the visual aggregation features and reducing the feature interaction burden on the mask decoder; in view of the obvious advantages of the CLIP backbone network based on ViT, through the introduction of the bidirectional dynamic embedding expert, a weight distribution router is used to evaluate the importance of the embedding expert and dynamically distribute the expert weight to generate a semantically aware and spatially aware instance embedding to improve the accuracy of mask recognition. Therefore, while achieving comparable performance, the method aims to reduce the model computational overhead and speed up the reasoning speed, and has significant practical value and application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 is a flow chart of the present invention;
[0046] Figure 2 It is a schematic diagram of the network structure of the present invention. DETAILED DESCRIPTION
[0047] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0048] like Figure 1 to Figure 2 As shown, a method for efficient open vocabulary panoptic segmentation comprises the following steps:
[0049] S1, visual feature extraction and aggregation based on multi-scale feature extractor and lightweight aggregator;
[0050] The specific process of step S1 is:
[0051] S11. In the visual feature extraction stage, the feature extractor uses the CNN-based CLIP visual encoder as the backbone network to extract the multi-scale feature map C i ,i∈{2,3,4,5};
[0052] S12, lightweight aggregator uses a modulated deformable convolutional feature pyramid structure to enhance and fuse features from different scales to obtain the intermediate FPN feature P i ,i∈{2,3,4,5}; then aggregate the features of different scales to obtain the visual aggregate features It is used to reduce computational overhead and speed up inference. D represents the number of feature map channels, H represents the image height, and W represents the image width. The calculation formula for feature aggregation is:
[0053]
[0054] Among them, Up i represents bilinear interpolation operation; N is the number of feature layers; Conv(·) and Fuse(·) respectively represent the use of 1×1 and 3×3 convolutions to adjust the feature dimension and fuse the visual aggregation feature F. agg ;
[0055] S2. Use the text encoder to encode any category of words to obtain text embedding Among them, N class represents the number of categories; D represents the number of channels of text embedding;
[0056] S3, based on the vocabulary-aware selection module, improves the semantic understanding of visual aggregation features and reduces the feature interaction burden of the mask decoder;
[0057] The specific process of step S3 is:
[0058] S31, visual aggregation features F from lightweight aggregator agg and the text embedding E from the text encoder t Mapped to the same feature dimension through depth-wise separable 2D convolution and linear layers respectively;
[0059] S32, aggregate the visual features F along the channel dimension agg and text embedding E t Split into multi-head visual aggregation features and multi-head text embedding Used to process complex semantic feature relationships in different subspaces; where h represents the number of heads;
[0060] S33. Processing multi-head visual aggregation features through matrix operations and multi-head text embedding Gaining multiple attention And perform softmax function on the multi-head attention A to produce smooth logits, and select the maximum value on the vocabulary dimension to obtain the vocabulary-aware attention weight The calculation formula is:
[0061]
[0062] S34, introduce two learnable parameters, namely the scaling factor γ and the offset factor δ, to enable the network to adaptively adjust the word-aware attention weights under various semantic distributions Lexical-aware attention weights Expand along the channel dimension and then aggregate features with multi-head vision Perform element-wise multiplication to reshape the weighted features into the original aggregate feature dimension to obtain the visual semantic aggregate feature. The calculation formula of the visual semantic aggregate feature is:
[0063]
[0064] in, represents visual semantic aggregation features; Reshape(·) represents the reshaping operation.
[0065] S4, based on bidirectional dynamic embedding experts, generates instance embeddings with semantic awareness and spatial awareness by dynamically assigning expert weights;
[0066] The specific process of step S4 is:
[0067] S41. The first block of CLIP's ViT-B visual encoder is used as a spatial perception extractor to extract spatial location information. The input image is processed by the first block of ViT-B to generate visual tags, and the visual tags are reshaped into a two-dimensional spatial dimension to obtain the visual spatial feature F. v , using two transposed convolutions to perform four-fold upsampling to obtain spatially aware features And use mask pooling to perceive the feature F from space s Extract the region of interest from
[0068] S42, introduce a weight distribution router to estimate the importance of the current expert embedding and adaptively adjust the relevant weight coefficients; generate mask embedding E through linear layers respectively m Dynamic parameters and spatially aware embedding E s Dynamic parameters Dynamic parameters are evenly split into router parameters along the channel dimension and And the fusion parameters and Where d represents the number of channels, and the calculation formula is:
[0069] (P f ,P r )=s(Em W m ),
[0070] (Q f ,Q r )=s(E s W s ),
[0071] in, W m Indicates E m The projection matrix is used to generate dynamic parameters; W s Indicates E s The projection matrix is used to generate dynamic parameters; s(·) is the splitting operation along the channel dimension;
[0072] S43, through P r and Q r The total router parameters are aggregated by the element-wise product between them; two router linear layers and sigmoid activation function are used to provide adaptive weighting capabilities for experts, so that the router adaptively assigns weights to the embedded experts. The calculation formula is:
[0073] P t =P r Q r ,
[0074] α m =σ(LN(Linear m (P t ))),
[0075] α s =σ(LN(Linear s (P t ))),
[0076] Among them, P t represents the total router parameter; α m represents the weight coefficient of mask embedding; α s Represents the weight coefficient of spatially aware embedding; Linear m (·) represents the mask embedded router linear layer; Linear s (·) denotes the linear layer of the spatially aware embedded router; LN(·) denotes the layer normalization operation; σ denotes the sigmoid activation function;
[0077] S44. Perform weighted summation of mask embedding and spatially aware embedding according to mask embedding expert weights and spatially aware embedding expert weights, and use layer normalization and GELU activation function to obtain instance embedding Used to aggregate instance and spatial information, the calculation formula is:
[0078]
[0079] S5, based on a lightweight decoder, uses the object kernel to perform mask prediction and refinement layer by layer, and uses the object kernel and text embedding for dot product as category prediction;
[0080] The specific process of step S5 is:
[0081] S51, the lightweight decoder consists of only three modules in the following order at each layer: initial attention module, dynamic deep attention module and late attention module; the initial attention module aggregates the visual semantic features from the vocabulary-aware selection module with the initial mask or predicted mask from the previous layer Perform dot product to extract initial attention features
[0082] S52. Using dynamic deep attention mechanism, a set of learnable object kernels With the initial attention feature F init Perform cross-dimensional feature interaction to obtain the refined object kernel The calculation formula is:
[0083]
[0084] Among them, DyDepthwiseAttention(·) represents dynamic depth-wise attention; represents the projection matrix of the object kernel K to generate the depth convolution kernel, m represents the convolution kernel size; r(·) represents the view operation that reshapes the input into N×1×m; * represents the parameter r(·) and F init 1D convolution operation for the input; the object kernel after refinement Focus on the relationship between different object cores, enrich information through multi-head self-attention and feed-forward neural network; the refined object core As a region proposal network, it provides the potential for generating mask proposals;
[0085] S53, through the three-layer perceptron, the refined object core Mapped to a mask kernel, each binary mask By aggregating features of the i-th mask kernel and visual semantics Dot product is performed to obtain; then the visual semantic aggregation features are Perform mask pooling with the predicted mask M to obtain mask embedding The lightweight decoder contains three decoding layers and repeats step S5 three times to obtain the final mask prediction and category prediction.
[0086] Performance Test:
[0087] The open vocabulary semantic, instance, and panoptic segmentation performance of the present invention is evaluated on the ADE20K dataset, and the open vocabulary semantic segmentation performance of the present invention is evaluated on the ADE20K, PASCAL Context, and PASCAL VOC datasets. During inference, the shortest side of the input image is adjusted to 640 while ensuring that the long side does not exceed 2560. For fair comparison, all experiments are performed three times and the average is taken.
[0088] Table 1: Performance test of different open vocabulary panoptic segmentation methods on the ADE20k dataset
[0089]
[0090] Table 2: Performance tests of different open vocabulary semantic segmentation methods on multiple datasets
[0091]
[0092] It can be seen from Table 1 and Table 2 that, whether in the open vocabulary panoramic segmentation method or the open vocabulary semantic segmentation method, after introducing vocabulary-aware selection and bidirectional dynamic embedding experts, the efficient open vocabulary panoramic segmentation method (EOV-Seg) adopted by the present invention has lower computational overhead and faster inference speed while maintaining competitive performance, achieving the best balance between performance and speed.
[0093] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A method for efficient open vocabulary panoptic segmentation, characterized in that: The following steps are involved: S1, visual feature extraction and aggregation based on multi-scale feature extractor and lightweight aggregator; S2. Use the text encoder to encode any category of words to obtain text embedding Among them, N class represents the number of categories; D represents the number of channels of text embedding; S3, based on the vocabulary-aware selection module, improves the semantic understanding of visual aggregation features and reduces the feature interaction burden of the mask decoder; S4, based on bidirectional dynamic embedding experts, generates instance embeddings with semantic awareness and spatial awareness by dynamically assigning expert weights; S5. Based on a lightweight decoder, the object kernel is used to perform mask prediction and refinement layer by layer, and the dot product of the object kernel and text embedding is used as category prediction.
2. The method for efficient open vocabulary panoptic segmentation according to claim 1, characterized in that: The specific process of step S1 is: S11. In the visual feature extraction stage, the feature extractor uses the CNN-based CLIP visual encoder as the backbone network to extract the multi-scale feature map C i ,i∈{2,3,4,5}; S12, lightweight aggregator uses a modulated deformable convolutional feature pyramid structure to enhance and fuse features from different scales to obtain the intermediate FPN feature P i ,i∈{2,3,4,5}; then aggregate the features of different scales to obtain the visual aggregate feature It is used to reduce computational overhead and speed up inference. D represents the number of feature map channels, H represents the image height, and W represents the image width. The calculation formula for feature aggregation is: Among them, Up i represents bilinear interpolation operation; N is the number of feature layers; Conv(·) and Fuse(·) respectively represent the use of 1×1 and 3×3 convolutions to adjust the feature dimension and fuse the visual aggregation feature F. agg .
3. The method for efficient open vocabulary panoptic segmentation according to claim 1, characterized in that: The specific process of step S3 is: S31, visual aggregation features F from lightweight aggregator agg and the text embedding E from the text encoder t Mapped to the same feature dimension through depth-wise separable 2D convolution and linear layers respectively; S32, aggregate the visual features F along the channel dimension agg and text embedding E t Split into multi-head visual aggregation features and multi-head text embedding Used to process complex semantic feature relationships in different subspaces; where h represents the number of heads; S33. Processing multi-head visual aggregation features through matrix operations and multi-head text embedding Gaining multiple attention And perform softmax function on the multi-head attention A to produce smooth logits, and select the maximum value on the vocabulary dimension to obtain the vocabulary-aware attention weight The calculation formula is: S34, introduce two learnable parameters, namely the scaling factor γ and the offset factor δ, to enable the network to adaptively adjust the word-aware attention weights under various semantic distributions Lexical-aware attention weights Expand along the channel dimension and then aggregate features with multi-head vision Perform element-wise multiplication to reshape the weighted features into the original aggregate feature dimension to obtain the visual semantic aggregate feature. The calculation formula of the visual semantic aggregate feature is: in, represents visual semantic aggregation features; Reshape(·) represents the reshaping operation.
4. The method for efficient open vocabulary panoptic segmentation according to claim 1, characterized in that: The specific process of step S4 is: S41. The first block of CLIP's ViT-B visual encoder is used as a spatial perception extractor to extract spatial location information. The input image is processed by the first block of ViT-B to generate visual tags, and the visual tags are reshaped into a two-dimensional spatial dimension to obtain the visual spatial feature F. v , using two transposed convolutions to perform four-fold upsampling to obtain spatially aware features And use mask pooling to perceive the feature F from space s Extract the region of interest from S42, introduce a weight distribution router to estimate the importance of the current expert embedding and adaptively adjust the relevant weight coefficients; generate mask embedding E through linear layers respectively m Dynamic parameters and spatially aware embedding E s Dynamic parameters Dynamic parameters are evenly split into router parameters along the channel dimension and And the fusion parameters and Where d represents the number of channels, and the calculation formula is: (P f ,P r )=s(E m W m ), (Q f ,Q r )=s(E s W s ), in, W m Indicates E m The projection matrix is used to generate dynamic parameters; W s Indicates E s The projection matrix is used to generate dynamic parameters; s(·) is the splitting operation along the channel dimension; S43, through P r and Q r The total router parameters are aggregated by the element-wise product between them; two router linear layers and sigmoid activation function are used to provide adaptive weighting capabilities for experts, so that the router adaptively assigns weights to the embedded experts. The calculation formula is: P t =P r ·Q r , a m =σ(LN(Linear m (P t ))), a s =σ(LN(Linear s (P t ))), Among them, P t represents the total router parameter; α m represents the weight coefficient of mask embedding; α s Represents the weight coefficient of spatially aware embedding; Linear m (·) represents the mask embedded router linear layer; Linear s (·) denotes the linear layer of the spatially aware embedded router; LN(·) denotes the layer normalization operation; σ denotes the sigmoid activation function; S44. Perform weighted summation of mask embedding and spatially aware embedding according to mask embedding expert weights and spatially aware embedding expert weights, and use layer normalization and GELU activation function to obtain instance embedding Used to aggregate instance and spatial information, the calculation formula is:
5. The method for efficient open vocabulary panoptic segmentation according to claim 1, characterized in that: The specific process of step S5 is: S51, the lightweight decoder consists of only three modules in the following order at each layer: initial attention module, dynamic deep attention module and late attention module; the initial attention module aggregates the visual semantic features from the vocabulary-aware selection module With the initial mask or predicted mask from the previous layer Perform dot product to extract initial attention features S52. Using dynamic deep attention mechanism, a set of learnable object kernels With the initial attention feature F init Perform cross-dimensional feature interaction to obtain the refined object kernel The calculation formula is: Among them, DyDepthwiseAttention(·) represents dynamic depth-wise attention; represents the projection matrix of the object kernel K to generate the depth convolution kernel, m represents the convolution kernel size; r(·) represents the view operation that reshapes the input into N×1×m; * represents the parameter r(·) and F init 1D convolution operation for the input; the object kernel after refinement Focus on the relationship between different object cores, enrich information through multi-head self-attention and feed-forward neural network; the refined object core As a region proposal network, it provides the potential for generating mask proposals; S53, through the three-layer perceptron, the refined object core Mapped to a mask kernel, each binary mask By aggregating features of the i-th mask kernel and visual semantics Dot product is performed to obtain; then the visual semantic aggregation features are Perform mask pooling with the predicted mask M to obtain mask embedding The lightweight decoder contains three decoding layers and repeats step S5 three times to obtain the final mask prediction and category prediction.
Citation Information
Cited By
Open vocabulary semantic segmentation method and device based on feature interaction and multi-modal data fusion
CN120655924A
Method and device for open-vocabulary semantic segmentation based on feature interaction and multi-modal data fusion
CN120655924B
Unmanned vehicle inspection small target detection method based on efficient attention mechanism
CN121074367A
Unmanned vehicle inspection small target detection method based on efficient attention mechanism
CN121074367B
Open vocabulary segmentation method based on parallel cost aggregation and expert perception
CN122154691A