A zero-shot referential image segmentation method based on hierarchical cues and directional clues

CN119049057BActive Publication Date: 2026-09-25ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411011071.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-26
Publication Date
2026-09-25
Estimated Expiration
2044-07-26

AI Technical Summary

Technical Problem

[0005]本发明的目的就是为了克服上述现有技术存在的掩码生成网络在精度和处理复杂分割场景方面表现不佳且容易产生大量冗余和琐碎的掩码以及CLIP无法有效利用指称表达式中明确的方向性线索的问题,而提供一种基于层次化提示和方向性线索的零样本指称图像分割方法

Benefits of technology

[0044]1.本发明引入了一个基于分层提示的掩码生成网络,该网络显著提高了掩码质量并减少了冗余。此外,通过一种利用文本线索中的空间方向描述来提取视觉特征的方法,解决了CLIP模型中空间灵敏度和细节识别的局限性,这将引导CLIP视觉编码器聚焦于图像中特定空间中的对象,显著的提高图像分割能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119049057B_ABST
    Figure CN119049057B_ABST
Patent Text Reader

Abstract

The application discloses a zero-shot referring image segmentation method based on hierarchical hints and directional clues. First, all object instance masks in the input image are obtained through a hierarchical hint mask generation network; comprehensive-focus visual features are obtained by extracting and fusing comprehensive visual features and focus visual features based on directional clues. Then, a pre-trained model BLIP2 is used to generate title text and negative expression text, and a CLIP text encoder is used to extract text features; for the input text T, comprehensive-focus text features are obtained by extracting and fusing comprehensive text features and focus text features. Finally, the cosine similarity between the image I and the text T is calculated by using the pre-trained model CLIP, the mask center is taken as the position information by using a spatial rectifier, and the mask with the highest matching score is selected. The application can generate accurate instance masks in occlusion and complex scenes, solves the problem that CLIP is not sensitive to spatial position information, and shows excellent performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image segmentation technology, specifically to a zero-sample reference image segmentation method based on hierarchical cues and directional clues. Background Technology

[0002] Referencing image segmentation aims to accurately segment target regions in an image as specified by a natural language description. This process requires precise alignment between the image and text, as well as a comprehensive understanding of both visual and textual elements. This presents a significant challenge in the field of visual language. Furthermore, generating accurate mask annotations and corresponding natural language descriptions is both time-consuming and expensive, and errors in hand-annotated text can affect the final results. To address these issues, researchers have explored weakly supervised methods. However, these methods perform poorly and still rely on high-quality training datasets. In contrast, zero-shot referencing image segmentation is a method that accurately identifies masks most relevant to the referencing expression without relying on pixel-level annotations. This process involves mask generation and mask-to-text matching, which is crucial for generating accurate, high-quality masks and exploring visual-textual relationships, making it more attractive and valuable for research.

[0003] Recent multimodal pre-trained models have demonstrated impressive capabilities in visual and language understanding. Of particular note is the visual-language model CLIP, which excels at capturing global similarity between text and images. Its outstanding performance in various image-level zero-shot tasks, including image retrieval, object detection, and semantic segmentation, attests to this. However, the CLIP model is trained on image-text pairs, while referential image segmentation tasks require dense pixel-level predictions. Therefore, directly applying a pre-trained CLIP model to zero-shot referential image segmentation tasks that require fine-grained region-text matching presents a challenge. Furthermore, the CLIP model is insensitive to spatial location information, which is inevitably reduced during the model's image encoding process. However, referential expressions often contain explicit directional cues (e.g., "woman on the right," "bread on the table," etc.), which are valuable but easily overlooked.

[0004] The quality of the masks generated by the mask generation network plays a crucial role in determining the upper limit of performance for this task. In Global-Local CLIP, FreeSOLO is used as the mask generation network, but it performs poorly in terms of accuracy and handling complex segmentation scenes. TAS uses SAM as the mask generation network to improve the quality of mask generation, but due to insufficient cueing, SAM tends to generate a large amount of redundant and trivial masks. Therefore, current zero-shot referential image segmentation schemes struggle to achieve good segmentation performance. Summary of the Invention

[0005] The purpose of this invention is to overcome the problems of existing technologies, such as poor performance of mask generation networks in terms of accuracy and handling of complex segmentation scenarios, the generation of a large number of redundant and trivial masks, and CLIP's inability to effectively utilize explicit directional cues in the denotation expression. In order to provide a zero-sample denotation image segmentation method based on hierarchical cues and directional cues.

[0006] To achieve the above objectives, this invention provides a zero-sample reference image segmentation method based on hierarchical cues and directional clues, comprising the following steps:

[0007] S1. Obtain the mask of all object instances in the input image through a hierarchical cue mask generation network;

[0008] S2. Extract comprehensive visual features based on the input image I and the input text T. and focal visual features Simultaneously, the two are fused to obtain the comprehensive-focus visual feature F. V ;

[0009] S3. Obtain the input text T, and use the pre-trained model BLIP2 to generate the title text and negative expression text, and use the CLIP text encoder to extract the title features. Negative text features Then, for the input text T, the CLIP text encoder is used to extract comprehensive text features. and focus text features Simultaneously, the two are fused to obtain the comprehensive-focused text feature F. T ;

[0010] S4. Use the pre-trained model CLIP to calculate the cosine similarity between the input image I and the input text T. Introduce a spatial rectifier to use the center of the mask as the position information of each mask scheme, and select the mask with the highest matching score in the corresponding direction region.

[0011] Preferably, the hierarchical cue mask generation network integrates three advanced models: Recognize Anything Plus Model (RAM++), Grounding DINO, and Segment Anything in High Quality (HQ-SAM).

[0012] Preferably, step S1 specifically includes the following steps:

[0013] S101. Based on the input image I, use RAM++ to generate category labels for all object instances as a first-level prompt;

[0014] S102. Use Grounding DINO to generate target detection boxes from the category labels as a second-layer prompt;

[0015] S103, HQ-SAM uses the first layer cue and the second layer cue to generate a segmentation mask for all object instances in the image.

[0016] Preferably, step S2 specifically includes the following steps:

[0017] S201. Capture directional descriptive cues in the input text T, and generate a directional bias matrix D∈R using the sigmoid function. W×H×C ;

[0018] S202, Set the direction bias matrix D∈R W×H×C and a mask scaled to the size of the feature map The input image I is encoded using element-wise multiplication, and then the CLIP visual encoder is used to extract comprehensive visual features from the encoded image.

[0019]

[0020] S203. Perform a masking operation on the input image I, then crop it, retaining only the portion containing the suggested masked area, and extract the focal visual features using the CLIP visual encoder.

[0021]

[0022] S204. Integrate the comprehensive visual features and the focal visual features Obtain comprehensive-focal visual features F V ;

[0023]

[0024] Where α is a hyperparameter between [0,1].

[0025] Preferably, step S3 specifically includes the following steps:

[0026] S301. Based on the input image I, generate title text and negative expression text using the pre-trained model BLIP2, and extract title features using the CLIP text encoder. and negative text features

[0027] S302. Based on the input text T, understand the overall meaning and the target object noun phrase, and use the pre-trained CLIP text encoder to extract comprehensive text features.

[0028] S303. Use spatial dependency parsing to identify and select target noun phrases containing the core nouns of the sentence, and then use these selected noun phrases to extract focus text features.

[0029] S304. Integrate the comprehensive text features. and the focus text feature Obtain comprehensive-focused text features F V

[0030]

[0031] Where β is a hyperparameter between [0,1].

[0032] Preferably, step S4 specifically includes the following steps:

[0033] S401. Use the pre-trained model CLIP to compute the comprehensive-focus visual feature F. V and the comprehensive-focused text feature F T cosine similarity S V-T

[0034] S V-T =cosine(F V ,F T );

[0035] S402. Use the pre-trained model CLIP to calculate the title features. With the aforementioned integrated-focus text feature F T cosine similarity S C-T

[0036]

[0037] S403. Use the pre-trained model CLIP to calculate the negative text features. With the aforementioned integrated-focus visual feature F V cosine similarity S N-T

[0038]

[0039] S404. Using the spatial rectifier introduced by TAS, the center of the mask is used as the position information of each mask scheme, and the mask with the highest matching score in the corresponding direction region is selected.

[0040] Preferably, RAM++ is an open image annotation model that generates category labels for each object instance in an image through multi-granular text supervision.

[0041] Preferably, the Grounding DINO is an open-set object detector that can accurately locate and detect object instances based on input category label cues, generating detailed bounding boxes for specific objects.

[0042] Preferably, the HQ-SAM is an improved version of SAM, which can generate a more detailed mask for the target area based on input points, boxes, or text prompts.

[0043] Compared with the prior art, the beneficial effects of the present invention are:

[0044] 1. This invention introduces a mask generation network based on hierarchical cues, which significantly improves mask quality and reduces redundancy. Furthermore, by employing a method that utilizes spatial orientation descriptions from textual cues to extract visual features, it overcomes the limitations of spatial sensitivity and detail recognition in the CLIP model. This guides the CLIP visual encoder to focus on objects in specific spaces within an image, significantly improving image segmentation capabilities.

[0045] 2. This invention not only solves the problem of generating a large number of invalid and redundant masks, but also generates accurate, detailed and comprehensive instance masks even in occluded and dense scenes, demonstrating excellent performance. Attached Figure Description

[0046] Figure 1 This is a schematic diagram of the method architecture of the present invention;

[0047] Figure 2 This is a schematic diagram of the hierarchical prompt mask generation network of the present invention;

[0048] Figure 3 This is a visual representation of the present invention; Detailed Implementation

[0049] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0050] This invention proposes a zero-shot reference image segmentation method based on hierarchical cues and directional clues, the specific steps of which are as follows:

[0051] S1. Generate segmentation mask: Obtain the mask of all object instances in the input image through a hierarchical cue mask generation network;

[0052] S2. Visual Feature Extraction and Fusion: Extract comprehensive visual features based on the input image I and the input text T. and focal visual features Simultaneously, the two are fused to obtain the comprehensive-focus visual feature F. V ;

[0053] S3. Obtain the input text T, and use the pre-trained model BLIP2 to generate the title text and negative expression text, and use the CLIP text encoder to extract the title features. Negative text features Then, for the input text T, the CLIP text encoder is used to extract comprehensive text features. and focus text features Simultaneously, the two are fused to obtain the comprehensive-focused text feature F. T ;

[0054] S4. Text-assisted mask selection: The pre-trained model CLIP is used to calculate the cosine similarity between the input image I and the input text T. A spatial rectifier is introduced to use the center of the mask as the position information of each mask scheme, and the mask with the highest matching score in the corresponding direction region is selected.

[0055] Specifically, step S1 includes the following steps:

[0056] (1) Specifically, given an input image I, the network initially uses RAM++ to generate category labels for all object instances as first-layer cues.

[0057] (2) These labels were then used by Grounding DINO to generate object detection boxes as a second layer of cues.

[0058] (3) Finally, HQ-SAM uses these bounding box cues to generate segmentation masks for all object instances in the image.

[0059] Specifically, step S2 includes the following steps:

[0060] (1) Given input text T, capture directional descriptive cues in the text and generate a directional bias matrix using the sigmoid function.

[0061] (2) The direction bias matrix D∈R W×H×C and a mask scaled to the size of the feature map The input image I is encoded using element-wise multiplication, and then the CLIP visual encoder is used to extract comprehensive visual features from the encoded image.

[0062]

[0063] (3) To obtain the visual features of a specific region, the input image I is masked and then cropped using Crop(), retaining only the portion containing the proposed masked region. The focus visual features are then extracted using the CLIP visual encoder.

[0064]

[0065] (4) Integrate the comprehensive visual features and the focal visual features to obtain the comprehensive-focal visual feature F. V .

[0066]

[0067] In fact, α is a hyperparameter between [0,1].

[0068] Specifically, step S3 includes the following steps:

[0069] (1) Given an input image I, use the pre-trained model BLIP2 to generate title text and negative expression text, and use CLIP text encoder to extract title features. and negative text features

[0070] (2) Given an input expression T, understanding the overall meaning and the target object noun phrase is crucial. A pre-trained CLIP text encoder is used to extract comprehensive text features.

[0071] (3) Spatial dependency parsing is used to identify and select target noun phrases containing the core nouns of the sentence, and then these selected noun phrases are used to extract focus text features.

[0072] (4) Integrate the comprehensive text features and the focus text features to obtain the comprehensive-focus text feature F. V .

[0073]

[0074] Where β is a hyperparameter between [0,1].

[0075] Specifically, step S4 includes the following steps:

[0076] (1) Use the pre-trained model CLIP to compute the comprehensive-focal visual features F. V and comprehensive-focused text features F T cosine similarity S V-T.

[0077] S V-T =cosine(F V ,F T )

[0078] (2) Use the pre-trained model CLIP to compute title features. With comprehensive-focused text features F T cosine similarity S C-T .

[0079]

[0080] (3) Use the pre-trained model CLIP to calculate negative text features. With comprehensive-focal visual features F V cosine similarity S N-T .

[0081]

[0082] (4) Using the spatial rectifier introduced by TAS, the center of the mask is used as the position information for each mask scheme, and the mask with the highest matching score in the corresponding direction region is selected. The final mask... It is selected by choosing the one with the highest matching score S.

[0083] S = S V-T +ηS C-T +θS N-V

[0084]

[0085] This invention also provides an innovative framework consisting of four parts: a mask generation network based on hierarchical cues, visual feature extraction and fusion, text feature extraction and fusion, and text cue-assisted mask selection. The mask generation network accurately generates masks based on hierarchical cues; it uses directional descriptions in the text as cues to guide the visual encoder to focus on specific regions of the image to extract comprehensive visual features and focal visual features. Furthermore, a text encoder is used to extract comprehensive text features and focal text features. Finally, with the help of text cues, the most relevant mask is selected by calculating the cosine similarity between all mask images and the reference expression, achieving accurate image segmentation.

[0086] The hierarchical cue-based mask generation network integrates three high-level models: Recognize Anything PlusModel (RAM++), Grounding DINO, and Segment Anything in High Quality (HQ-SAM). RAM++ is an open-set image annotation model that generates class labels for each object instance in an image through multi-granular text supervision. Grounding DINO, as an open-set object detector, accurately locates and detects object instances based on input class label cues, generating detailed bounding boxes for specific objects. HQ-SAM is an improved version of SAM, not only enhancing the quality of segmentation masks but also generating more detailed masks for target regions based on input points, boxes, or text cues. Furthermore, it can generate masks for all existing instances in the image without additional cues. The network initially utilizes RAM++ to generate class labels for all object instances as the first-layer cue. These labels are then used by Grounding DINO to generate object detection boxes as the second-layer cue. Finally, HQ-SAM uses these bounding box cues to generate segmentation masks for all object instances in the image. This hierarchical hint-based approach not only solves the problem of generating a large number of invalid and redundant masks, but also produces accurate, detailed and comprehensive instance masks even in occluded and dense scenes, demonstrating excellent performance.

[0087] Figure 2 A qualitative comparison between the method of the present invention and the baseline method is given, showing that the method of the present invention achieves higher accuracy in mask generation. Figure 2 b describes a hierarchical cue mask generation network that initially utilizes RAM++ to generate class labels for all object instances as the first-layer cue. These labels are then used by Grounding DINO to generate object detection bounding boxes as the second-layer cue. Finally, HQ-SAM uses these bounding box cues to generate segmentation masks for all object instances in the image. This hierarchical cue-based approach not only solves the problem of generating a large number of invalid and redundant masks, but also produces accurate, detailed, and comprehensive instance masks even in occluded and dense scenes, demonstrating superior performance.

[0088] Because CLIP is insensitive to spatial location information, this information is inevitably reduced during the image encoding process of the model. However, denotative expressions often contain explicit directional cues (e.g., "woman on the right," "bread on the table," etc.), which are valuable but easily overlooked. To emphasize this type of location information, this invention captures directional descriptive cues in the text and proposes a directional bias matrix. This matrix is ​​then used to encode the image through element-wise multiplication.

[0089] I'=I⊙D

[0090] Where I is a given image, and D∈R W×H×C Let D be the direction bias matrix, and ⊙ be the Hadamard product. If the text contains explicit directional cues (such as up, down, left, right) after parsing, then D will be the direction bias matrix generated by the sigmoid function, with bias values ​​ranging from 1 to 0 along the given direction axis. The horizontal direction axis represents left, right, west, east, etc., and the vertical direction axis represents up, down, north, south, etc. The diagonal direction axis represents directions such as southeast, northwest, etc. Lower values ​​indicate areas that should be given less attention. If no directional cues are detected, matrix D will be filled with 1s.

[0091] Visual features are extracted from encoded images using a CLIP pre-trained model. To perform more detailed analysis of the masked regions and their surrounding information, this invention optimizes the CLIP visual encoder based on the method proposed by Global-Local CLIP. In the ResNet-50 model, the global average pooling layer is replaced by an attention pooling layer utilizing a multi-head attention mechanism. In the ViT-B / 32 model, the image is segmented and processed through Transformer layers, with mask tokens applied only in the final layer. The output of the class tokens (CLS) from ViT is used as a comprehensive visual representation.

[0092] Therefore, the comprehensive visual features of the present invention can be represented as:

[0093]

[0094] in, To scale to I ′ The mask after adjusting the size of the feature map.

[0095] Focal visual features. To obtain the visual features of a specific region, the image is masked and then cropped, retaining only the portion containing the proposed masked area. The cropped and masked image is then fed into the CLIP visual encoder, whose purpose is to capture focal visual features.

[0096] F f V =∮ CLIP (Crop(I⊙M))

[0097] Crop() represents a cropping operation, the purpose of which is to focus only on the target object itself.

[0098] Once the comprehensive, focused visual features are obtained, they are fused by direct addition:

[0099]

[0100] In fact, α is a hyperparameter between [0,1].

[0101] As described in TAS, the domain gap between the natural image and the masked image affects the alignment of visual and textual features, and the presence of numerous objects in the image that are irrelevant to the denotation can be distracting. Therefore, this invention utilizes the CLIP text encoder to extract title features from the title text. Extracting negative text features from negative expressions

[0102] Synthetic Text Features. Similar to visual features, understanding the overall meaning and target object noun phrases within a given expression T is crucial. To this end, this invention uses a pre-trained CLIP text encoder to extract synthetic text features.

[0103] Focused Text Features. While the CLIP text encoder can extract text features aligned with an image, the presence of multiple clauses makes it difficult to focus attention on key nouns. To address this issue, this invention uses spatial dependency parsing to identify and select target noun phrases containing the core nouns of the sentence. These selected noun phrases are then used to extract focused text features.

[0104] Similar to visual features, the combined text features and focus text features are fused using a direct addition method:

[0105]

[0106] To identify the most relevant mask associated with the reference expression, this invention uses a pre-trained model CLIP to calculate the cosine similarity between the mask image and the reference expression. Considering the cosine similarity between the composite-focus visual text features, the title features and the composite-focus text features, and the negative text expression features and the composite-focus visual features, a matching score is calculated. The formula for calculating the matching score is as follows:

[0107] S V-T =cosine(F V ,F T )

[0108]

[0109]

[0110] T nIt is a set of noun phrases. Since all scores are cosine similarities calculated in the common CLIP feature space, the final matching score can be obtained by linearly combining the above three scores. As mentioned above, CLIP is not sensitive to spatial relationships. Therefore, this invention uses a spatial rectifier introduced by TAS, using the center of the mask as the position information of each mask scheme, and selects the mask with the highest matching score for the corresponding directional region. In this way, this invention forces CLIP to focus on specific regions when processing directional descriptions, thereby correcting erroneous predictions.

[0111] S = S V-T +ηS C-T +θS N-V

[0112]

[0113] Final mask It is selected by choosing the one with the highest matching score S.

[0114] Qualitative analysis. Figure 3 A qualitative comparison between the method of the present invention and conventional methods is shown. The results in the figure demonstrate that the method of the present invention can effectively extract detailed features from images, understand complex reference expressions, including spatial descriptors, and select the most accurate target object mask.

[0115] While the invention has been described herein with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the invention. Therefore, it should be understood that many modifications can be made to the exemplary embodiments, and other arrangements can be designed without departing from the spirit and scope of the invention as defined by the appended claims. It should be understood that different dependent claims and features described herein can be combined in ways different from those described in the original claims. It is also understood that features described in conjunction with individual embodiments can be used in other described embodiments.

Claims

1. A zero-sample reference image segmentation method based on hierarchical cues and directional clues, characterized in that, Includes the following steps: S1. Obtain the mask of all object instances in the input image through a hierarchical cue mask generation network; S2, Based on the input image Extract comprehensive visual features from input text T. and focal visual features At the same time, the two are merged to obtain a comprehensive-focus visual feature. ; S3. Obtain the input text T, and use the pre-trained model BLIP2 to generate the title text and negative expression text, and use the CLIP text encoder to extract the title features. Negative text features Then, for the input text T, the CLIP text encoder is used to extract comprehensive text features. and focus text features Simultaneously, the two are integrated to obtain comprehensive-focused text features. ; S4. Use the pre-trained model CLIP to compute the input image. The cosine similarity between the input text T and the spatial rectifier is introduced to use the center of the mask as the position information of each mask scheme, and the mask with the highest matching score in the corresponding directional region is selected. Step S1 specifically includes the following steps: S101, Based on the input image RAM++ is used to generate category labels for all object instances as the first level of hints; S102. Use Grounding DINO to generate target detection boxes from the category labels as a second-layer prompt; S103, HQ-SAM uses the first layer cue and the second layer cue to generate a segmentation mask for all object instances in the image; Step S2 specifically includes the following steps: S201. Capture directional descriptive cues in the input text T and generate a directional bias matrix using the sigmoid function. ; S202, Adjust the direction offset matrix and a mask scaled to the size of the feature map The input image is processed by element-wise multiplication. The image is then encoded, and then the CLIP visual encoder is used to extract the composite visual features from the encoded image. ; 。 2. The zero-sample reference image segmentation method based on hierarchical cues and directional clues according to claim 1, characterized in that, The hierarchical cue mask generation network integrates three advanced models: Recognize Anything Plus Model (RAM++), Grounding DINO, and Segment Anything in High Quality (HQ-SAM).

3. The zero-sample reference image segmentation method based on hierarchical cues and directional clues according to claim 1, characterized in that, Step S2 further includes the following steps: S203, the input image Perform a masking operation, then crop the image, retaining only the portion containing the proposed masked area. Use the CLIP visual encoder to extract the focal visual features. ; ; S204. Integrate the comprehensive visual features and the focal visual features To obtain comprehensive-focal visual features ; ; Where α is a hyperparameter between [0,1].

4. The zero-sample reference image segmentation method based on hierarchical cues and directional clues according to claim 3, characterized in that, Step S3 specifically includes the following steps: S301, Based on the input image The pre-trained model BLIP2 was used to generate title text and negative expression text, and the CLIP text encoder was used to extract title features. and negative text features ; S302. Based on the input text T, understand the overall meaning and the target object noun phrase, and use the pre-trained CLIP text encoder to extract comprehensive text features. ; S303. Use spatial dependency parsing to identify and select target noun phrases containing the core nouns of the sentence, and then use these selected noun phrases to extract focus text features. ; S304. Integrate the comprehensive text features. and the focus text feature To obtain comprehensive-focused text features ; ; in It is a hyperparameter between [0,1].

5. The zero-sample reference image segmentation method based on hierarchical cues and directional clues according to claim 4, characterized in that, Step S4 specifically includes the following steps: S401. Use the pre-trained model CLIP to compute the integrated-focus visual features. and the aforementioned integrated-focus text features cosine similarity ; ; S402. Use the pre-trained model CLIP to calculate the title features. With the aforementioned integrated-focus text features cosine similarity ; ; S403. Use the pre-trained model CLIP to calculate the negative text features. With the aforementioned integrated-focus visual features cosine similarity ; ; S404. Using the spatial rectifier introduced by TAS, the center of the mask is used as the position information of each mask scheme, and the mask with the highest matching score in the corresponding direction region is selected.

6. The zero-sample reference image segmentation method based on hierarchical cues and directional clues according to claim 2, characterized in that, RAM++ is an open-source image annotation model that generates category labels for each object instance in an image through multi-granular text supervision.

7. The zero-sample reference image segmentation method based on hierarchical cues and directional clues according to claim 2, characterized in that, The Grounding DINO is an open-set object detector that can accurately locate and detect object instances based on input category label cues, generating detailed bounding boxes for specific objects.

8. The zero-sample reference image segmentation method based on hierarchical cues and directional clues according to claim 2, characterized in that, The HQ-SAM is an improved version of SAM that can generate more detailed masks for target areas based on input points, boxes, or text prompts.