Hyperbolic space open scene semantic segmentation method based on image language supervision
By fine-tuning large language visual models in hyperbolic space, migrating image and text encoders from image-level classification to pixel-level classification, and adjusting feature hierarchy using Mobius multiplication, the problem of semantic segmentation in the open world is solved, achieving efficient, accurate and interpretable pixel-level semantic segmentation.
Patent Information
- Application Number
- CN202510161027.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art is difficult to effectively perform semantic segmentation in the open world, especially when encountering objects or categories that have never been seen, lack pixel-level visual understanding capabilities and problems of overfitting and insufficient interpretability.
By fine-tuning large language visual models in hyperbolic space, the image and text encoder are directly migrated from image-level classification to pixel-level classification, and the hierarchy of features is adjusted using Mobius multiplication to make the model have semantic segmentation capabilities.
Efficient pixel-level semantic segmentation in open scenarios is realized, which reduces the overhead of computing resources and improves the interpretability of the model.
Smart Images

Figure CN120107582A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image semantic segmentation, and in particular relates to an open scene semantic segmentation method in a hyperbolic space based on image language supervision. Background Art
[0002] Deep neural networks have achieved remarkable results in many computer vision tasks. However, deep neural networks constructed by traditional fully supervised training can only recognize categories that have appeared in the training phase. For semantic segmentation, during the prediction phase, categories that are not included in the training data will be incorrectly classified as a category that has appeared. This greatly limits the usability of deep neural networks, making it difficult for semantic segmentation models trained based on a single dataset to cope with complex scenes in the open world.
[0003] In order to deal with objects or categories that have never been seen before, the academic community has proposed the topic of open set semantic segmentation. Open set semantic segmentation is a fully supervised training using images in a closed category set and corresponding pixel-level annotations during the training phase; in the prediction phase, it can identify the pixel categories of images in any open category set. CLIP (Contrastive Language-Image Pre-training, a pre-training model based on contrastive text-image pairs), a large visual language model, can effectively provide the ability to recognize images in an open category set. Therefore, when applying CLIP to open set semantic segmentation, the key is how to give it segmentation capabilities during the training phase.
[0004] At present, unsupervised domain adaptation methods for this problem are mainly divided into two categories, but both have shortcomings. One method is to fine-tune the parameters in the image encoder of CLIP. For example, Side Adapter Network for Open-Vocabulary Semantic Segmentation (SAN) fine-tunes the adapter network of the image encoder connected to CLIP to enable the image encoder to have segmentation capabilities. At the same time, the parameters in the text encoder of CLIP are frozen to ensure the open category recognition ability of CLIP. The open category recognition ability of this technology is for the image level rather than the pixel level, so the trained model cannot guarantee the pixel-level visual understanding ability. Another related solution is to fine-tune some parameters in the CLIP encoder, such as CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation (CAT-Seg) by fine-tuning the parameters of the multi-head attention layer in the CLIP encoder to enable CLIP to have segmentation capabilities. Although this method has pixel-level understanding capabilities, the introduction of too many parameters will cause overfitting to the closed category set, and lack the interpretability of fine-tuning.
[0005] Semantic segmentation refers to the classification of each pixel on an image, which is one of the basic tasks in the field of computer vision. Based on a large amount of labeled data, deep neural networks have achieved significant performance improvements in semantic segmentation. However, in complex scenes in the open world, it is extremely difficult to obtain labeled data of all categories. Therefore, open set semantic segmentation performs fully supervised training in a category set with labeled data, and then identifies the pixel category of an image in any open category set during the prediction phase. Especially in the field of autonomous driving, since there are many rare scenes in autonomous driving, in complex road environments, autonomous driving vehicles may encounter obstacles not covered in the training data (such as temporary road facilities, new vehicles, or sudden objects). Open set semantic segmentation algorithms can help vehicles identify these unknown objects, so that they can take appropriate avoidance measures and improve driving safety.
[0006] However, the limitations of existing methods make the realization of this goal still face many challenges. Therefore, there is an urgent need for a more efficient, accurate and interpretable open set semantic segmentation method to cope with complex and changing scenes in the open world. Summary of the invention
[0007] In order to overcome the shortcomings of the above-mentioned prior art, the present invention provides an open set semantic segmentation method in a hyperbolic space based on image-language supervision, which utilizes the property of hyperbolic space that can encode hierarchical structures. By fine-tuning a large language vision model in the hyperbolic space, the image and text encoders of the large language vision model are directly migrated from the image-level classification level to the pixel-level classification level, thereby enabling the large language vision model to have semantic segmentation capabilities while retaining the open scene capabilities.
[0008] A method for semantic segmentation of open scenes in hyperbolic space based on image language supervision, which has the following characteristics:
[0009] Construct a hyperbolic space training framework for image and language domains and perform training;
[0010] An open scene semantic segmentation model is constructed according to the training framework of the image domain and the language domain, wherein the open scene semantic segmentation model transfers the image-level classification capability of the image encoder and the text encoder to the pixel-level semantic segmentation capability by adjusting the hyperbolic radius;
[0011] The image to be segmented is input into the trained open scene semantic segmentation model to obtain the semantic segmentation result.
[0012] Further, the hyperbolic space training framework of the image domain and the language domain is constructed and trained, specifically:
[0013] -Introducing image encoder and text encoder training of large-scale visual language models:
[0014] -The parameters of the image encoder and text encoder are fixed and not updated, retaining their open scene capabilities;
[0015] -Map the features obtained by the image encoder and text encoder to the hyperbolic space to obtain the hyperbolic space features:
[0016] F 1 =p(F)
[0017] In the formula, F and F 1 denote the feature map generated by the image / text encoder and the hyperbolic space feature map, respectively, and p(·) denotes the mapping function.
[0018] - Perform Möbius multiplication on the hyperbolic space features with a learnable parameter matrix to change the hierarchical structure of the features:
[0019] F 2 =F 1 ☉W
[0020] In the formula, F 2 They represent the updated feature maps, ☉ represents the Möbius multiplication, and W represents the learnable parameter matrix.
[0021] -Map the hyperbolic space features back to the original feature space:
[0022] F 3 =q(F 2 )
[0023] In the formula, F 3 represents the mapped back feature map, and q(·) represents the mapping function.
[0024] Furthermore, the mapping function p(·) and the inverse mapping function q(·) are bidirectional projection functions between the hyperbolic space and the Euclidean space, respectively, and satisfy:
[0025] F 1 =p(F) and F 3 =q(F 2 ).
[0026] Furthermore, the training objectives of the open scene semantic segmentation model are:
[0027] - Keep the open scene recognition capabilities of the image encoder and text encoder unchanged;
[0028] -Through hierarchical adjustment in hyperbolic space, the model is transferred from image-level classification capability to pixel-level semantic segmentation capability, and only the parameters of the learnable parameter matrix W of the block are optimized.
[0029] Furthermore, each submatrix of the learnable parameter matrix W corresponds to a feature channel at a different semantic level in the hyperbolic space, and dynamic adaptation of the hyperbolic radius is achieved by adjusting the scaling coefficient of the submatrix.
[0030] Further, the adjustment of the hyperbolic radius is achieved by:
[0031] - In hyperbolic space, reduce the hyperbolic radius of pre-trained model features to enhance the fine-grained feature expression capability of pixel-level segmentation tasks;
[0032] -The reduction range of the hyperbolic radius is obtained by learning the parameters of the learnable parameter matrix W.
[0033] Compared with the prior art, the present invention has the following beneficial effects:
[0034] 1) The present invention fine-tunes a large language vision model in a hyperbolic space, and directly migrates the image and text encoders of the large language vision model from the image-level classification level to the pixel-level classification level, thereby directly endowing the large language vision model with semantic segmentation capabilities.
[0035] 2) The present invention utilizes the Möbius multiplication to directly change the semantic level of features in the hyperbolic space, thereby reducing the learnable parameters required to endow a large language vision model with semantic segmentation capabilities, thereby reducing the cost of computing resources.
[0036] 3) The present invention has good theoretical guarantee and explainability. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 Schematic diagram of the hierarchy of the text encoder before and after fine-tuning a large language vision model. The hierarchy of features required to complete pixel-level classification is different from that required to complete image-level classification.
[0038] Figure 2 It is a flow chart of the open set semantic segmentation in hyperbolic space under image language supervision according to the present invention. DETAILED DESCRIPTION
[0039] The present invention is described in detail below in conjunction with specific drawings and embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several adjustments and improvements can also be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.
[0040] This embodiment provides an open scene semantic segmentation method based on a hyperbolic space supervised by image language, which aims to adjust the hyperbolic radius of the CLIP text encoder and image encoder to align with the pixel-level semantic hierarchy required for the segmentation task. Specifically, several block diagonal scaling matrices are introduced. Using the Möbius matrix multiplication operation to multiply these matrices with the features of the text encoder and image encoder of CLIP after projecting into the hyperbolic space is equivalent to performing a scaling transformation to adjust the hyperbolic radius of the features of the text encoder and image encoder of CLIP. This can adjust their semantic structure hierarchy and ultimately improve the performance of the model. Therefore, only a small number of learnable parameters need to be introduced to give CLIP segmentation capabilities.
[0041] We evaluate our method in the context of migrating from a closed-category set to an open-category set, using COCO-Stuff as the closed-category set, ADE20K (A-847 and A-150), PASCAL-VOC (PAS-20 and PAS-20 b ) and PASCAL-Context (PC-459 and PC-59) as different open category collection scenes. As usual, we report the segmentation results on ADE20K, PASCAL-VOC, and PASCAL-Context.
[0042] Table 1 Segmentation effect of open category set scene
[0043]
[0044] As shown in Table 1, the proposed method achieves the best results based on ViT-B / 16 and ViT-L / 14 in the three open category set segmentation scenarios. b In the open scene of PC-59, the model based on ViT-B / 16 improved by 1.0 points, and the model based on ViT-L / 14 improved by 0.9 points, which shows the effectiveness of the present invention.
Claims
1. A method for semantic segmentation of open scenes in hyperbolic space based on image language supervision, characterized in that: include: Construct a hyperbolic space training framework for image and language domains and perform training; An open scene semantic segmentation model is constructed according to the training framework of the image domain and the language domain, wherein the open scene semantic segmentation model transfers the image-level classification capability of the image encoder and the text encoder to the pixel-level semantic segmentation capability by adjusting the hyperbolic radius; The image to be segmented is input into the trained open scene semantic segmentation model, and the pixel-level semantic segmentation result is output.
2. The method for semantic segmentation of open scenes in hyperbolic space based on image language supervision according to claim 1, characterized in that: Construct a hyperbolic space training framework for the image domain and language domain and perform training. Specifically: -Fixed the parameters of the image encoder and text encoder for training large visual language models, preserving their open scene capabilities; -The feature maps output by the image encoder and the text encoder are mapped to the hyperbolic space through the mapping function p(·) to generate the hyperbolic space feature map F1: F1=p(F) Where F and F1 represent the feature map and hyperbolic space feature map generated by the image / text encoder respectively; -In the hyperbolic space, the hyperbolic space feature map F1 is Mobius multiplied with a learnable parameter matrix W to adjust the semantic hierarchy of the features to generate an updated hyperbolic space feature map F2: F2=F1☉W Where ☉ represents the Möbius multiplication; - The updated hyperbolic space feature map F2 is mapped back to the original feature space through the inverse mapping function q(·) to obtain the feature map F3: F3 = q(F2).
3. The method for semantic segmentation of open scenes in hyperbolic space based on image language supervision according to claim 2, characterized in that: The mapping function p(·) and the inverse mapping function q(·) are bidirectional projection functions between the hyperbolic space and the Euclidean space, respectively, and satisfy: F1 = p(F) and F3 = q(F2).
4. The method for semantic segmentation of open scenes in hyperbolic space based on image language supervision according to claim 2, characterized in that: The training objectives of the open scene semantic segmentation model are: - Keep the open scene recognition capabilities of the image encoder and text encoder unchanged; -Through hierarchical adjustment in hyperbolic space, the model is transferred from image-level classification capability to pixel-level semantic segmentation capability, and only the parameters of the learnable parameter matrix W of the block are optimized.
5. The method for semantic segmentation of open scenes in hyperbolic space based on image language supervision according to any one of claims 1 to 4, characterized in that: Each submatrix of the learnable parameter matrix W corresponds to a feature channel at a different semantic level in the hyperbolic space, and dynamic adaptation of the hyperbolic radius is achieved by adjusting the scaling coefficient of the submatrix.
6. The method for semantic segmentation of open scenes in hyperbolic space based on image language supervision according to any one of claims 1 to 4, characterized in that: The adjustment of the hyperbolic radius is achieved by: - In hyperbolic space, reduce the hyperbolic radius of pre-trained model features to enhance the fine-grained feature expression capability of pixel-level segmentation tasks; -The reduction range of the hyperbolic radius is obtained by learning the parameters of the learnable parameter matrix W.
Citation Information
Cited By
Cloud platform auditing method based on collaborative auditing
CN121073058A
Mutual supervision optimization method for multi-modal reasoning large model and segmentation large model
CN121789001A
A multi-modal reasoning large model and segmentation large model mutual supervision optimization method
CN121789001B