Panoramic segmentation method

By activating boot query and dual-path Transformer decoder in instances in the panoramic segmentation network, the adaptability and computing resource consumption of static query methods in complex scenarios is solved, and a more efficient panoramic segmentation effect is achieved.

CN120451556APending Publication Date: 2025-08-08SHANDONG XIEHE UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510556977.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing static query methods lack adaptability in panoramic segmentation, and it is difficult to cope with complex scenarios where the number of targets changes, the computing resources are consumed and the convergence speed is slow, so it is impossible to effectively distinguish similar instances from retaining background information.

Method used

A panoramic segmentation network is adopted, including a backbone network, a pixel decoder and a dual-path Transformer decoder, and dynamic queries are generated through instance activation guide queries, combining convolutional attention mechanism and auxiliary classification heads to achieve enhanced and precise segmentation of multi-scale features.

Benefits of technology

It improves the adaptability and accuracy of panoramic segmentation, reduces computing resource consumption, can better capture global and local relationships in the image, enhances the robustness of target segmentation under different scales and occlusion situations, and breaks through the feature degradation bottleneck of single-path decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451556A_ABST
    Figure CN120451556A_ABST
Patent Text Reader

Abstract

The invention provides a panoramic segmentation method, and the method comprises the steps: inputting an initial image into a panoramic segmentation network, and obtaining a panoramic segmentation result; the method for obtaining the segmentation result comprises the following steps: sending an initial image into a backbone network, and carrying out multi-scale feature extraction to obtain a plurality of feature maps E3, E4 and E5 with different resolutions; mapping the feature maps to a multi-channel feature map, and inputting the multi-channel feature map to a pixel decoder to obtain enhanced multi-scale feature maps E3 ', E4' and E5 '; generating Na object queries from the feature map E4 by using an instance activation guide query strategy, and connecting the Na object queries with Nb background auxiliary queries to generate a total query QIA; the total query QIA and the enhanced pixel feature E3'are used as the input of a double-path Transform decoder; in the dual-path Transform decoder, a dual-path mode alternative updating strategy is adopted to update a pixel feature E3'and query Q, and a target category and a segmentation mask are predicted in each layer of the decoder; according to the invention, adaptability and segmentation precision of panoramic segmentation of different scenes can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image segmentation, and in particular to a panoramic segmentation method. Background Art

[0002] In the field of image segmentation, semantic segmentation can only achieve pixel-level classification but cannot distinguish instances of the same type. Instance segmentation can distinguish foreground objects but often ignores background information. Panoptic segmentation, by combining the advantages of both, achieves both semantic category recognition and instance-level differentiation for every pixel in an image, providing a more comprehensive representation of information for global scene analysis. Panoptic segmentation assigns each pixel both a semantic label and the attribute of belonging to a specific instance, making it a key technology for achieving human-like understanding in computer vision.

[0003] Query-based methods, with their end-to-end training and superior performance, have become the mainstream approach in the field of panoptic segmentation. These methods elegantly transform the panoptic segmentation problem into a query matching and decoding process, capturing both global semantic information and local details through deep interaction between query vectors and image features. However, static query methods, such as Mask2Former, still suffer from the following three drawbacks:

[0004] 1. Static queries use a preset fixed number of queries and lack the ability to adapt to the complexity of the scene. They are difficult to handle complex scenarios with changing target numbers and are prone to missing or confusing instances.

[0005] 2. The initialization of static queries is independent of the initial image content, which limits the model's ability to generalize to diverse scenarios.

[0006] 3. Static queries require iterative optimization of multiple layers of decoders to obtain effective representation, which consumes large computing resources and has slow convergence speed.

[0007] Therefore, there is an urgent need for a panoramic segmentation method that can solve the above defects. Summary of the Invention

[0008] The purpose of the present invention is to provide a panoptic segmentation method, aiming to address the shortcomings of the traditional static query method Mask2Former.

[0009] To achieve the above objectives, in a first aspect, the present invention provides a panoptic segmentation method, the steps of which include:

[0010] Inputting the initial image into a pre-built panoramic segmentation network to obtain a panoramic segmentation result of the initial image;

[0011] The pre-built panoptic segmentation network includes a backbone network for extracting multi-scale features from the initial image, a pixel decoder for feature enhancement of the multi-scale features, an instance activation guided query module for generating dynamic queries, and a dual-path Transformer decoder for deep interaction between queries and features and final segmentation prediction;

[0012] The steps for obtaining the panoptic segmentation result are as follows:

[0013] S1, feature extraction: Input the initial image into the backbone network to obtain multiple initial feature maps C3, C4 and C5 with different resolutions;

[0014] S2, feature enhancement: Perform a convolutional block attention module (CBAM) operation on the initial feature map to generate E3, E4, and E5, use a convolutional layer to map these feature maps to multi-channel feature maps, and input them into the pixel decoder to obtain enhanced multi-scale feature maps E3′, E4′, and E5′;

[0015] S3, query generation: Generate Na object queries from the feature graph E4 using the instance activation guided query strategy, and connect it with Nb background auxiliary queries to generate the total query Q IA ;

[0016] S4, decoding and prediction: the total query Q IA Together with the enhanced pixel feature E3′, it is used as the input of the dual-path Transformer decoder;

[0017] In the dual-path Transformer decoder, a dual-path alternating update strategy is adopted to update the pixel features E3′ and the query Q, and the target category and segmentation mask are predicted in each layer of the decoder to achieve panoramic segmentation of the image.

[0018] As a further improvement of the above scheme, the convolutional attention mechanism includes a channel attention module and a spatial attention module; and the convolutional attention mechanism can be expressed as follows:

[0019]

[0020] Among them, F represents the input feature map, Mc and Ms represent the channel attention module and spatial attention module respectively, represents the element-wise multiplication operation, F′ represents the feature map enhanced by the channel attention module, and F″ represents the final enhanced feature map.

[0021] As a further improvement of the above solution, the pixel decoder is used to enhance feature representation, and a PPM-FPN (Pyramid Pooling Module-Feature Pyramid Network) network is used to enhance features.

[0022] As a further improvement to the above scheme, the instance activation guided query strategy selects pixels that exhibit high semantics from the underlying multi-scale feature map as embeddings and directly uses them as query vectors.

[0023] As a further improvement to the above solution, the instance activation guide query strategy specifically includes:

[0024] The auxiliary classification head structure is optimized to optimize the object query generation process. It consists of two convolutional layers, as shown below:

[0025] F1=Conv3×3(E4)

[0026] P = Conv1 × 1 (F1);

[0027] Activate query selection. Based on the foreground probability, select the pixel with the local maximum foreground probability from the feature map E4 as the basis for object query. Then select the pixel with the global highest foreground probability, and finally select Na pixel embedding for object query.

[0028] Pixel matching mechanism, using category prediction and location cost function L loc The allocation cost is calculated to evaluate the matching quality.

[0029] As a further improvement to the above scheme, the steps of the method for selecting the local maximum foreground probability pixel are as follows:

[0030] Find the pixel with the highest foreground probability within the 8-neighborhood of a pixel;

[0031] The spatial 8-neighborhood index set of pixel i is defined as δ(i), if the foreground probability of pixel i p i,ki If the foreground probability of a pixel is greater than or equal to the foreground probability of all pixels in its neighborhood, the pixel is considered to be a local maximum, as shown in the following formula:

[0032]

[0033] As a further improvement to the above solution, the final object query is specifically represented as follows:

[0034] Q=TopK({E4[i]|LocalMax(i)=1},Na)

[0035] The TopK function selects the foreground with the highest probability p i,kiThe local maximum pixel.

[0036] As a further improvement to the above solution, the matching cost function in the pixel matching mechanism is shown as follows:

[0037] L(i,g)=λ cls ·L cls (i,g)+λ loc ·L loc (i,g)

[0038] Where L(i,g) represents the total matching cost between pixel i and target g, L cls (i,g) represents the category prediction loss, L loc (i,g) is the location cost function, λ cls and λ loc are the balance weights of category loss and position loss respectively.

[0039] As a further improvement of the above solution, the position cost function L loc The specific expressions are as follows:

[0040]

[0041] Among them, i represents the pixel position on the feature map, and g represents the real labeled target object.

[0042] As a further improvement to the above solution, the dual-path alternating update strategy specifically includes:

[0043] Position embedding: Set a fixed-size embedding matrix P∈R for the learnable spatial position S×S×256 , where S is determined by rounding off the square root of the number of instance-guided queries Na; in the forward propagation of the model, P adapts to feature maps of different sizes through interpolation operations;

[0044] Pixel feature update: This includes a cross-attention layer and a feedforward layer. In the cross-attention layer, the model achieves information fusion and transfer by calculating the similarity between the query and the key (pixel feature). For each pixel feature, the cross-attention layer calculates the attention weight between it and all queries, and then performs a weighted average of the queries based on the obtained attention weight to obtain the updated pixel feature. The position embedding is added to the query and key to enhance position sensitivity.

[0045] Query update: This includes a masked attention layer, a self-attention layer, and a feedforward network for feature transformation integration. The masked attention mechanism limits the attention scope of each query to the foreground region of the previous layer's prediction mask. The self-attention mechanism is introduced after the masked attention layer to enable comprehensive information exchange between queries, and the position embedding is added to the query and key of both the masked attention and self-attention layers.

[0046] Prediction Generation: At each decoder layer, two independent multilayer perceptrons (MLPs) are used to refine the instance activation guided query.

[0047] Two independent multi-layer perceptrons are used: the category prediction MLP for predicting the object category and the mask embedding MLP for generating mask embeddings.

[0048] As a further improvement to the above scheme, the steps for P to adapt to feature maps of different sizes through interpolation operations are as follows:

[0049] First, P is interpolated into a matrix of the same size as the feature map E3′ and flattened as the position embedding of the pixel feature X to capture more detailed spatial information.

[0050] Then, P is resized to the same size as the feature map E4′ for position embedding of instance activation-guided queries.

[0051] As a further improvement of the above scheme, the overall loss function of the panoptic segmentation network is shown as follows:

[0052] L=L IA +L pre

[0053] Among them L IA is the instance activation loss function of the auxiliary classification head, L IA =λ cls ·L cls +λ loc ·L loc , where L cls is the category prediction loss, L loc is the location matching cost, λ cls ,λ loc Is a hyperparameter balancing factor used to balance classification loss and mask loss;

[0054] L pre is the loss function for prediction generation, D represents the number of layers of the Transformer decoder, i=0 represents the prediction loss of the IA-guided query before entering the Transformer decoder, and L ice and L i dice denote the binary cross entropy loss and dice loss of the segmentation mask respectively; L i cls is the cross entropy loss for object classification with a “no object” weight of 0.1; ce ,λ dice and λ cls is a hyperparameter that balances the three losses.

[0055] In a second aspect, the present invention also provides a panoramic segmentation device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of a panoramic segmentation method as described in the first aspect are implemented.

[0056] In a third aspect, the present invention further provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of a panoramic segmentation method as described in the first aspect.

[0057] Since the present invention adopts the above technical solution, the beneficial effects of this application are:

[0058] The present invention provides a panoramic segmentation method that dynamically generates object queries based on the semantic information of the underlying multi-scale features through instance activation guided query (IA-guided), abandoning the limitation of traditional methods that rely on preset instance templates, thereby improving the segmentation adaptability and segmentation accuracy of different scenes; this mechanism enables the model to autonomously perceive the potential target distribution characteristics in the image, and combined with the synergistic effect of background auxiliary query, effectively covers the diverse instance forms in complex scenes, and significantly enhances the robustness of target segmentation for different scales, occlusion conditions and distribution densities; by utilizing the semantic prior of the underlying feature map (such as the foreground probability distribution predicted by the auxiliary classification head), the initial query is given higher information entropy and target directionality. This design makes the query initialization stage The dual-path Transformer decoder has a strong focusing ability on key areas, which greatly reduces the redundant computing burden of the Transformer decoder. At the same time, it realizes the deep interactive iteration of queries and pixel features through the dual-path alternating update strategy, so as to better capture the global and local relationships in the image. In addition, through the interaction between pixel features and queries, the dual-path update strategy can generate more refined feature representations; a dual-path Transformer decoder is used to decode pixel features and object queries in parallel, and a dynamic association between pixel-level details and instance-level semantics is established through a cross-path attention mechanism; it breaks through the feature degradation bottleneck in single-path decoding, while retaining multi-scale spatial details, it strengthens the representation ability of target boundaries, and improves the panoramic quality.

[0059] Through the cascade design of IA-guided query generation and dual-path decoding, a complete chain of evidence is formed from low-level feature activation to high-level semantic parsing, so that the mapping relationship between the query vector and the target instance has a visual explanation basis, providing a clear improvement direction for model optimization and error analysis; the present invention effectively solves the technical pain points of existing methods such as instance template rigidity, missed detection of small targets, and blurred mask edges, and is particularly suitable for application scenarios such as autonomous driving and remote sensing image analysis that have strict requirements on real-time performance and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying any creative work.

[0061] Figure 1 This is a schematic diagram of the structure of a panoptic segmentation network PSM-DIO disclosed in the present invention;

[0062] Figure 2 It is a structural schematic diagram of the auxiliary sorting head disclosed in the present invention;

[0063] Figure 3 It is a visualization diagram of the segmentation results of each method disclosed in the present invention on the Cityscapes dataset;

[0064] Figure 4 This is a visualization diagram of the segmentation results on the COCO dataset of each method disclosed in the present invention;

[0065] Figure 5 Schematic diagram of the change of the loss function during the training process of the panoramic segmentation network PSM-DIO disclosed in the present invention;

[0066] Figure 6 It is the PQ indicator on the validation set during the training process of the panoramic segmentation network PSM-DIO disclosed in the present invention.

[0067] The realization of the objectives, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0068] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of the present invention.

[0069] It should be noted that all directional indications (such as up, down, etc.) in the embodiments of the present invention are only used to explain the relative position relationship, movement status, etc. between the various components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.

[0070] In addition, the terms "first," "second," and so on, used in this disclosure are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, features specified as "first" or "second" may explicitly or implicitly include at least one of these features.

[0071] Moreover, the technical solutions between the various embodiments of the present invention may be combined with each other, but this must be based on the fact that ordinary technicians in this field can implement them. When the combination of technical solutions is mutually contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0072] Example 1

[0073] See also Figure 1 The present invention provides a panoptic segmentation method, the steps of which include:

[0074] Inputting the initial image into a pre-built panoramic segmentation network to obtain a panoramic segmentation result of the initial image;

[0075] The pre-built panoptic segmentation network includes a backbone network for extracting multi-scale features from the initial image, a pixel decoder for enhancing the multi-scale features, an instance activation guided query module for generating dynamic queries, and a dual-path Transformer decoder for deep interaction between queries and features and final segmentation prediction; these modules achieve the conversion from image data to accurate segmentation results in a layer-by-layer progressive manner;

[0076] The steps for obtaining the panoptic segmentation result are as follows:

[0077] S1. Feature extraction: Input the initial image into the backbone network to obtain three initial feature maps C3, C4 and C5 with different resolutions. The resolutions of these images are 1 / 8, 1 / 16 and 1 / 32 of the original initial image respectively;

[0078] S2, Feature Enhancement: A Convolutional Block Attention Module (CBAM) is applied to the three initial feature maps to generate E3, E4, and E5, which are then converted into new feature maps with 256 channels using a 1×1 convolutional layer. The resulting feature maps are then fed into a pixel decoder, which aggregates contextual information and generates and outputs a series of enhanced multi-scale feature maps E3′, E4′, and E5′.

[0079] S3, query generation: Generate Na object queries from the feature graph E4 using the instance activation guided query strategy, and concatenate them with Nb background auxiliary queries to generate the total query Q IA ;

[0080] S4, decoding and prediction: the total query Q IA Together with the enhanced pixel feature E3′, it is used as the input of the dual-path Transformer decoder;

[0081] In the dual-path Transformer decoder, as Figure 1 As shown on the right side of the dotted line, the decoder updates the pixel features E3′ and the query Q in a dual-path alternating update strategy, and can predict the target category and segmentation mask in each layer of the decoder to achieve panoramic segmentation of the image;

[0082] Compared with the traditional static query method, the present invention uses instance activation guided query (IA-guided) to dynamically generate object queries based on the semantic information of the underlying multi-scale features, abandoning the limitation of the traditional method that relies on preset instance templates, thereby improving the segmentation adaptability and segmentation accuracy of different scenes; this mechanism enables the model to autonomously perceive the potential target distribution characteristics in the image, and combined with the synergistic effect of background auxiliary query, effectively covers the diverse instance forms in complex scenes, and significantly enhances the robustness of target segmentation for different scales, occlusion conditions and distribution densities; by utilizing the semantic prior of the underlying feature map (such as the foreground probability distribution predicted by the auxiliary classification head), the initial query is given higher information entropy and target directionality. This design makes the query initialization At this stage, it has a strong focusing ability on key areas, which greatly reduces the redundant computing burden of the Transformer decoder. At the same time, it realizes the deep interactive iteration of queries and pixel features through the dual-path alternating update strategy, so as to better capture the global and local relationships in the image. In addition, through the interaction between pixel features and queries, the dual-path update strategy can generate more refined feature representations; a dual-path Transformer decoder is used to decode pixel features and object queries in parallel, and a dynamic association between pixel-level details and instance-level semantics is established through a cross-path attention mechanism; it breaks through the feature degradation bottleneck in single-path decoding, while retaining multi-scale spatial details, it strengthens the representation ability of target boundaries, and improves the panoramic quality.

[0083] As a preferred embodiment, the convolutional attention mechanism includes a channel attention module and a spatial attention module; these two modules can generate a series of attention feature maps in the channel dimension and the spatial dimension respectively; then, these feature maps are multiplied with the original input feature map to adaptively adjust the feature representation, which can more effectively capture the information of features in different channels and spatial dimensions, thereby significantly improving the prediction effect of semantic classes and instance classes, and significantly improving the accuracy of panoramic segmentation; the convolutional attention mechanism can be expressed as follows:

[0084]

[0085] Among them, F represents the input feature map, Mc and Ms represent the channel attention module and spatial attention module respectively, represents the element-wise multiplication operation, F′ represents the feature map enhanced by the channel attention module, and F″ represents the final enhanced feature map;

[0086] CBAM combines spatial attention and channel attention modules to leverage their complementary strengths, comprehensively capturing both spatial and channel information. Furthermore, the CBAM module is flexible and lightweight, making it easily embedded in any convolutional neural network architecture.

[0087] As a preferred embodiment, the pixel decoder (PixelDecoder) is a neural network component used for image processing and computer vision tasks; its main function is to downsample and fuse high-resolution feature maps to generate low-resolution features with more semantic information, providing input for subsequent decoders.

[0088] The present invention uses refined pixel features from the Transformer decoder to generate segmentation masks. This setup reduces the pixel decoder's requirement for heavy context aggregation, allowing the use of a lightweight pixel decoder module. To better balance accuracy and speed, the present invention uses a PPM-FPN (PyramidPoolingModule-FeaturePyramidNetwork) network to enhance features, enabling the utilization of the Transformer decoder's refined pixel features, rather than relying on the traditional pixel decoder in Mask2Former.

[0089] As a preferred embodiment, the instance activation guided query strategy selects pixels exhibiting high semantics from the underlying multi-scale feature map as embeddings and directly uses them as query vectors. This design can overcome the limitations of traditional static queries and learnable queries (which are independent of the input image content) and prevent instance confusion or instance loss. Instance activation guided queries utilize the semantic information in the underlying feature map to more efficiently encode target information in the image, reduce the burden on the Transformer decoder, and improve the relevance of the query to the image content.

[0090] As a further improvement to the above solution, the instance activation guide query strategy specifically includes:

[0091] Optimize the auxiliary classification head structure to optimize the object query generation process (see Figure 2 ), which includes two convolutional layers, specifically expressed as follows:

[0092] F1=Conv3×3(E4)

[0093] P = Conv1 × 1 (F1);

[0094] These two convolutional layers form a lightweight semantic segmentation network, which is used to extract semantic information from the underlying feature map. The first convolutional layer uses a 3×3 convolution kernel size to capture local features and enhance the model's spatial perception. The second convolutional layer uses a 1×1 convolution kernel size to integrate local features and generate the final category probability prediction. Using this structural design, the auxiliary classification head can effectively extract information that is helpful for classification and generate more accurate category probability predictions.

[0095] Activate query selection. Based on the foreground probability, select the pixel with the local maximum foreground probability from the feature map E4 as the basis for object query. Then select the pixel with the global highest foreground probability. Finally, select the pixel Na to embed for object query. This layer-by-layer screening method ensures that the selected pixels are representative locally and can better reflect the feature information of the target in the global scope.

[0096] Specifically, through the prospect probability p i,ki To determine the foreground probability of each pixel, where ki is the class index corresponding to the maximum activation value:

[0097]

[0098] To simplify the computational complexity, the problem is abstracted into a pair of classification tasks (i.e., foreground and background). The category with the highest activation value is selected to determine the corresponding category index. Based on these foreground probabilities, the pixel embedding with a higher foreground probability is selected from the feature map E4 as the basis for object query. Finally, Na pixel embedding is selected for object query.

[0099] The steps of the method for selecting the local maximum foreground probability pixel are as follows:

[0100] Find the pixel with the highest foreground probability within the 8-neighborhood of a pixel;

[0101] The spatial 8-neighborhood index set of pixel i is defined as δ(i), if the foreground probability of pixel i p i,ki If the foreground probability of a pixel is greater than or equal to the foreground probability of all pixels in its neighborhood, the pixel is considered to be a local maximum, as shown in the following formula:

[0102]

[0103] The final object query is specifically represented as follows:

[0104] Q=TopK({E4[i]|LocalMax(i)=1},Na)

[0105] The TopK function selects the foreground with the highest probability p i,ki The local maximum pixel;

[0106] Pixel matching mechanism, using category prediction and location cost function L loc Calculate allocation cost to assess matching quality;

[0107] Specifically, the matching cost function is shown as follows:

[0108] L(i,g)=λ cls ·L cls (i,g)+λ loc ·Lloc (i,g)

[0109] Where L(i,g) represents the total matching cost between pixel i and target g, L cls (i,g) represents the category prediction loss, L loc (i,g) is the location cost function, λ cls and λ loc are the balance weights of category loss and position loss respectively;

[0110] The location cost function L loc The specific expressions are as follows:

[0111]

[0112] Here, i represents the pixel position on the feature map, and g represents the real labeled target object; this position cost design ensures that only pixels falling inside the target object can participate in the matching process, thereby improving the accuracy of category and mask prediction;

[0113] IA-guided's prior information comes from the semantic information of the underlying feature map; using an auxiliary classification head to predict the probability of each pixel belonging to the foreground, this means that IA-guided already "knows" which areas in the image are more likely to contain objects during initialization. IA-guided has higher initial information entropy and can more effectively encode object information in the image; this enables the model to more effectively utilize limited query resources, reduce redundant computations, and improve sensitivity to objects.

[0114] The activation query selection provided by this invention prioritizes non-local maximum probability pixels because higher-scoring pixels have already been selected in their neighborhoods. This strategy further improves the accuracy of object queries. This dual screening mechanism ensures that the selected pixels are both locally representative (having the highest foreground probability in the neighborhood) and globally significant (having the highest foreground probability among all local maximum pixels).

[0115] As a preferred embodiment, the dual-path alternating update strategy specifically includes:

[0116] Position embedding: Position information is crucial for distinguishing different instances with similar semantics, especially when dealing with similar objects. This paper uses learnable position embedding to enhance the inference speed and adaptability of the model. Unlike traditional non-parametric sinusoidal position embedding, this paper uses a fixed-size learnable spatial position embedding matrix P∈R S×S×256, where S is determined by rounding off the square root of the number of instance-guided queries Na, aiming to balance computational efficiency and the fineness of position information. This embedding matrix uses learning to optimize the representation of spatial positions during training. In the forward propagation of the model, P uses interpolation operations to adapt to feature maps of different sizes: P is interpolated into a matrix of the same size as the feature map E3′, and after flattening, it is used as the position embedding of pixel features to capture more refined spatial information. Then, P is adjusted to the same size as the feature map E4′ for the position embedding of instance activation-guided queries. This interpolation strategy ensures that the model can fully utilize position information in feature representations at different levels. In order to further improve the model's ability to distinguish complex backgrounds or similar semantic objects, an additional Nb learnable position embedding is introduced. In this way, the model can better capture spatial position information and better distinguish different objects in complex scenes;

[0117] Pixel feature update: In model design, the self-attention mechanism usually brings a lot of computational and memory overhead when processing long sequences. In the pixel feature update process, the present invention introduces a method based on the cross-attention mechanism to avoid the high computational cost of the self-attention mechanism; the pixel feature update process includes a cross-attention layer and a feedforward layer. In the cross-attention layer, the model calculates the similarity between the query and the key (pixel feature) to realize information fusion and transmission, thereby efficiently aggregating global information. For each pixel feature, the cross-attention layer calculates the attention weight between it and all queries, and then performs a weighted average on the queries based on these weights to obtain the updated pixel features. In order to enhance position sensitivity, position embeddings are added to both queries and keys to ensure that the model can more accurately capture spatial position information when aggregating features, thereby better distinguishing different objects in complex scenes;

[0118] Query Update: The query update adopts an asymmetric design, centered on a unique multi-step update process consisting of a masked attention layer, a self-attention layer, and a feedforward network. Masked Attention: The masked attention mechanism limits the attention of each query to the foreground region of the mask predicted by the previous layer. This design effectively focuses the model's attention on relevant image regions, preventing queries from focusing on irrelevant areas, and improving the accuracy and efficiency of feature extraction. Self-Attention: The self-attention mechanism introduced after the masked attention layer enables comprehensive information exchange between queries, capturing the complex relationships and dependencies between queries. This allows the model to comprehensively consider information from different queries and make more comprehensive and accurate predictions. To further enhance the model's expressiveness, positional embeddings are incorporated into both the query and key in the masked attention and self-attention layers. This design considers the importance of spatial information in image segmentation tasks, enabling the model to better understand and utilize positional relationships in the image. Feedforward Network: The entire query update process uses a feedforward network for final feature transformation and integration. This step enhances the model's nonlinear capabilities and promotes deep fusion and optimization of information extracted in the previous steps. The goal of query updating is to optimize the representation of each query based on pixel features and the prediction results of the previous layer so that it better represents the potential target instance;

[0119] Prediction Generation: At each decoder layer, two independent three-layer Multilayer Perceptrons (MLPs) are used to refine instance activation-guided queries, thereby achieving more accurate instance segmentation; Category Prediction MLP: The first MLP is specifically used to predict object categories. In addition to predicting possible object categories, it also includes a special "no object" category, enabling the model to effectively distinguish foreground objects from backgrounds, thereby improving segmentation accuracy. The probability distribution of all possible categories is predicted to lay the foundation for subsequent confidence evaluation; Mask Embedding MLP: The second MLP focuses on generating mask embeddings, which encode the spatial information and shape characteristics of the object, laying the foundation for accurate pixel-level segmentation. Using independent MLPs to generate mask embeddings, the model is able to better capture the detailed features of the object without being directly affected by the category prediction task. In order to transform these refined features into actual segmentation results, a linear projection method is used. Specifically, the refined pixel features are mapped to the mask feature space using a linear transformation. Because queries and pixel features exist in different representation spaces in different decoder layers, a dynamic mask generation method is used to multiply the query embedding with the mask embedding of the pixel features to generate a unique segmentation mask for each query. The linear projection parameters of the MLP and each Transformer decoder layer are not shared, allowing each layer to independently learn its specific feature transformation. To comprehensively assess the reliability of each prediction, the class probability is multiplied by the segmentation mask score to obtain a composite confidence score.

[0120] As a preferred embodiment, the overall loss function of the panoptic segmentation network is shown as follows:

[0121] L=L IA +L pre

[0122] Among them L IA is the instance activation loss function of the auxiliary classification head, L IA =λ cls ·L cls +λ loc ·L loc , where L cls is the category prediction loss, L loc is the location matching cost, λ cls ,λ loc Is a hyperparameter balancing factor used to balance classification loss and mask loss;

[0123] L pre is the loss function for prediction generation, D represents the number of layers of the Transformer decoder, i=0 represents the prediction loss of the IA-guided query before entering the Transformer decoder, and L i ce and L i dice denote the binary cross entropy loss and dice loss of the segmentation mask respectively; L i cls is the cross entropy loss for object classification with a “no object” weight of 0.1; ce ,λ dice and λ cls is a hyperparameter that balances the three losses.

[0124] For details, see Figure 5 , showing the overall loss function decreasing over the course of training for the panoptic segmentation network proposed in this invention at 90,000 iterations. The overall loss function starts out very high, then drops sharply over the first few thousand iterations, indicating rapid initial learning. After approximately 40,000 iterations, the loss continues to decrease, but at a much slower rate. By 90,000 iterations, the loss appears to have almost fully converged and stabilized.

[0125] In order to further illustrate the effectiveness of the panoramic segmentation method provided by the present invention, a number of methods such as top-down method, bottom-up method, static query method, dynamic query method are compared with the panoramic segmentation method provided by the present invention, and the common and representative indicators PQ, PQ th and PQ st. Mask2Former is selected as a comparison example in the static query field. The panoramic segmentation network provided by the present invention incorporates CMBA, instance-guided activation strategy, dual-path Transformer decoder, and selects Res50 as the backbone network of the entire model. The parameter batch size is set to 16, the learning rate is 0.0001, the number of training iterations is 90,000, and the AdamW optimizer is used. The experiment is carried out on a single GPU. The GPU model used is NVIDIA Tesla A100, the video memory is 40GB, the deep learning framework is Pytorch 2.1.0, and the CUDA version is 12.1; the panoramic segmentation results on the Cityscapes and MSCOCO datasets are shown in Tables 1 and 2 respectively; the results show that the panoramic segmentation network based on dynamic instance query proposed by the present invention has the same backbone network configuration, which outperforms all other methods.

[0126] See Table 1, the method PSM-DIQ provided by the present invention performs well in PQ, PQ th and PQ st The indicators reached 63.9%, 56.8% and 69.3% respectively, which were 1.8, 2.0 and 2.0 percentage points higher than the baseline network Mask2Former.

[0127] Table 1 Comparison of panoramic segmentation results on the Cityscapes dataset

[0128]

[0129] See Table 2, PQ, PQ of the method PSM-DIQ provided by the present invention on the MSCOCO dataset th ,PQ st The results are 53.6%, 59.6%, and 44.5%, respectively. They are 1.7, 1.9, and 1.5 percentage points higher than the baseline network Mask2Former. PSM-DIQ also outperforms all listed top-down methods, bottom-up methods, and other query methods.

[0130] Table 2 Comparison of panoramic segmentation results on the MSCOCO dataset

[0131]

[0132] See also Figure 6Figure 3 shows the Panoptic Quality (PQ) metric on the validation set throughout training. Performance improves significantly in the early stages, reaching approximately 55–60% at 20,000 iterations. Performance continues to improve, but at a slower rate thereafter. Training progresses steadily, with no signs of overfitting. Overall, the PQ metric indicates that the model is learning effectively, and its performance on the validation set improves over time.

[0133] To further demonstrate the effectiveness of the panoptic segmentation method proposed in this paper, three complex environments were randomly selected from the Cityscapes and MSCOCO panoptic segmentation datasets for visualization. The panoptic segmentation visualization results use different colors to mark each segment. From left to right, they are the input image, the ground truth (GT), the Mask2Former prediction, and the PSMDIQ (Ours) prediction.

[0134] See also Figure 3 The red areas indicate that the method PSM-DIQ provided by the present invention can more accurately identify and segment targets such as pedestrians, traffic signs, and street lights on the road, while Mask2Former has omissions and blurring in these areas. When processing complex scenes, PSMDIQ can better capture detailed information such as trees, buildings, and traffic signs. In contrast, Mask2Former has poor segmentation effects in these areas, and is prone to blurring and missed detections. The segmentation results of PSMDIQ are also more coherent, reducing the fragmented segmentation phenomenon of Mask2Former.

[0135] See also Figure 4 The PSM-DIQ method proposed in this paper performs well in a variety of complex scenes, particularly in diverse environments such as cats, aerial perspectives, and skiing scenes. The red circled areas indicate that PSM-DIQ better preserves image details, such as skiers and small objects in the distance. In aerial images, PSM-DIQ more accurately identifies the boundaries of ground objects, reducing missegmentation.

[0136] Compared to Mask2Former, the method provided by the present invention has a clearer segmentation effect, especially when dealing with complex scenes, with significantly improved segmentation effect. In the scene, the Mask2Former method has a poor segmentation effect for complex scenes, and much detailed information is missing. The method provided by the present invention more accurately restores the true segmentation scene, with an accuracy closer to the true label.

[0137] Example 2

[0138] The present invention further provides a panoptic segmentation device, comprising a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein when the computer program instructions are executed by the processor, the device is triggered to execute some or all of the steps in embodiment 1;

[0139] A processor may include one or more processing units, for example, an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors.

[0140] The controller can be the nerve center and command center of the electronic device. The controller can generate operation control signals based on instruction operation codes and timing signals to complete the control of instruction fetching and execution.

[0141] The processor may also include a memory for storing instructions and data. In some embodiments, the memory in the processor is a cache memory. This memory can store instructions or data that the processor has just used or is reusing. If the processor needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces processor latency, and thus improves system efficiency.

[0142] Example 3

[0143] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein when the program is executed, the device where the storage medium is located is controlled to execute part or all of the steps in Example 1.

[0144] The computer-readable storage medium may include a high-speed RAM memory, and may also include a non-volatile memory (nonvolatile memory), such as at least one disk storage. It is understood that the storage medium can be a random access memory (RAM), a magnetic disk, a hard disk, a solid state disk (SSD), or a non-volatile memory, and other machine-readable media that can store program code.

[0145] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods or storage media. Therefore, embodiments of the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0146] The above are only preferred embodiments of the present invention and are not intended to limit the patent scope of the present invention. All equivalent structural transformations made based on the contents of the present invention's description and drawings, or directly or indirectly applied to other related technical fields, are within the scope of patent protection of the present invention.

Claims

1. A panoptic segmentation method, characterized in that: The steps include: Inputting the initial image into a pre-built panoramic segmentation network to obtain a panoramic segmentation result of the initial image; The pre-built panoptic segmentation network includes a backbone network for extracting multi-scale features from the initial image, a pixel decoder for feature enhancement of the multi-scale features, an instance activation guided query module for generating dynamic queries, and a dual-path Transformer decoder for deep interaction between queries and features and final segmentation prediction; The steps to obtain panoramic segmentation results are as follows: S1, feature extraction: Input the initial image into the backbone network to obtain multiple initial feature maps C3, C4 and C5 with different resolutions; S2, feature enhancement: Perform a convolutional attention mechanism operation on the initial feature map to generate E3, E4, and E5, use a convolutional layer to map these feature maps to multi-channel feature maps, and input them into the pixel decoder to obtain enhanced multi-scale feature maps E3′, E4′, and E5′; S3, query generation: Generate Na object queries from the feature graph E4 using the instance activation guided query strategy, and connect it with Nb background auxiliary queries to generate the total query Q IA ; S4, decoding and prediction: the total query Q IA Together with the enhanced pixel feature E3′, it is used as the input of the dual-path Transformer decoder; In the dual-path Transformer decoder, a dual-path alternating update strategy is adopted to update the pixel features E3′ and the query Q, and the target category and segmentation mask are predicted in each layer of the decoder to achieve panoramic segmentation of the image.

2. A panoptic segmentation method according to claim 1, characterized in that: The convolutional attention mechanism includes a channel attention module and a spatial attention module; and the convolutional attention mechanism can be expressed as follows: Among them, F represents the input feature map, Mc and Ms represent the channel attention module and spatial attention module respectively, represents the element-wise multiplication operation, F′ represents the feature map enhanced by the channel attention module, and F″ represents the final enhanced feature map.

3. A panoptic segmentation method according to claim 1 or 2, characterized in that: The instance activation guided query strategy selects pixels that exhibit high semantic meaning from the underlying multi-scale feature map as embeddings and directly uses them as query vectors.

4. A panoptic segmentation method according to claim 1 or 2, characterized in that: The instance activation guidance query strategy specifically includes: The auxiliary classification head structure is optimized to optimize the object query generation process. It consists of two convolutional layers, as shown below: F1=Conv3×3(E4) P = Conv1 × 1 (F1); Activate query selection. Based on the foreground probability, select the pixel with the local maximum foreground probability from the feature map E4 as the basis for object query. Then select the pixel with the global highest foreground probability, and finally select Na pixel embedding for object query. Pixel matching mechanism, using category prediction and location cost function L loc The allocation cost is calculated to evaluate the matching quality.

5. A panoptic segmentation method according to claim 4, characterized in that: The steps of the method for selecting the local maximum foreground probability pixel are as follows: Find the pixel with the highest foreground probability within the 8-neighborhood of a pixel; The spatial 8-neighborhood index set of pixel i is defined as δ(i), if the foreground probability of pixel i p i,ki If the foreground probability of a pixel is greater than or equal to the foreground probability of all pixels in its neighborhood, the pixel is considered to be a local maximum, as shown in the following formula:

6. A panoptic segmentation method according to claim 4, characterized in that: The final object query is specifically represented as follows: Q=TopK({E4[i]|LocalMax(i)=1},Na) The TopK function selects the foreground with the highest probability p i,ki The local maximum pixel.

7. A panoptic segmentation method according to claim 4, characterized in that: The matching cost function in the pixel matching mechanism is as follows: L(i,g)=λ cls ·L cls (i,g)+λ loc ·L loc (i,g) Where L(i,g) represents the total matching cost between pixel i and target g, L cls (i,g) represents the category prediction loss, L loc (i,g) is the location cost function, λ cls and λ loc are the balance weights of category loss and position loss respectively.

8. A panoptic segmentation method according to claim 4, characterized in that: The dual-path alternating update strategy specifically includes: Position embedding: Set an embedding matrix P∈R that can learn spatial positions S×S×256 , where S is determined by rounding off the square root of the number of instance-guided queries Na; in the forward propagation of the model, P adapts to feature maps of different sizes through interpolation operations; Pixel feature update: This includes a cross-attention layer and a feed-forward layer. In the cross-attention layer, the model achieves information fusion and transfer by calculating the similarity between the query and the key. For each pixel feature, the cross-attention layer calculates the attention weight between it and all queries, and then performs a weighted average of the queries based on the obtained attention weight to obtain the updated pixel feature. The position embedding is added to the query and key. Query update: This includes a masked attention layer, a self-attention layer, and a feedforward network for feature transformation integration. The masked attention mechanism limits the attention scope of each query to the foreground region of the previous layer's predicted mask. The self-attention mechanism is introduced after the masked attention, and the position embedding is added to the query and key of both the masked attention and self-attention layers. Prediction Generation: At each decoder layer, two independent multilayer perceptrons are used to refine the instance activation guided query; the two independent multilayer perceptrons are respectively the category prediction MLP for predicting the object category and the mask embedding MLP for generating the mask embedding.

9. A panoptic segmentation method according to claim 8, characterized in that: The steps for P to adapt to feature maps of different sizes through interpolation are as follows: First, P is interpolated into a matrix of the same size as the feature map E3′ and flattened as the position embedding of the pixel feature X to capture more detailed spatial information. Then, P is resized to the same size as the feature map E4′ for position embedding of instance activation-guided queries.

10. A panoptic segmentation method according to claim 8, characterized in that: The overall loss function of the panoptic segmentation network is shown as follows: L=L IA +L pre Among them L IA is the instance activation loss function of the auxiliary classification head, L IA =λ cls ·L cls +λ loc ·L loc , where L cls is the category prediction loss, L loc is the location matching cost, λ cls ,λ loc is a hyperparameter balancing factor; L pre is the loss function for prediction generation, D represents the number of layers of the Transformer decoder, i=0 represents the prediction loss of the IA-guided query before entering the Transformer decoder, and L i ce and L i dice denote the binary cross entropy loss and dice loss of the segmentation mask respectively; L i cls is the cross entropy loss for object classification with a “no object” weight of 0.1; ce ,λ dice and λ cls is a hyperparameter that balances the three losses.