Open vocabulary image semantic segmentation method, system, device and storage medium

By extracting visual features and masks, generating text-aware features, and performing visual language projection, the problem of insufficient expression ability of traditional methods in open scenes is solved, and the content of local visual areas is automatically segmented and recognized, and diverse recognition results are generated.

CN119723096BActive Publication Date: 2025-05-13UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510227653.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-13
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

The traditional open vocabulary semantic segmentation method relies on predefined vocabulary, lacks sufficient expressive ability to understand open world scenarios, and is limited to unknown open scenarios.

Method used

By extracting visual features and masks from the input image, aggregating local information in combination with attention mechanisms, acquiring proposal features, and generating text-aware features based on these features. These features are projected visually into the language and inputted to the language model to obtain the category names and area descriptions contained in each area in the input image.

Benefits of technology

It realizes that without obtaining a vocabulary list, automatically segment and identify the content of visual local areas, generate accurate and diverse recognition results, cover attributes and multi-level semantic information, and has stronger generalization capabilities, and can handle different visual content more flexibly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119723096B_ABST
    Figure CN119723096B_ABST
Patent Text Reader

Abstract

The present invention discloses an open vocabulary image semantic segmentation method, system, device and storage medium, which are one-to-one corresponding schemes. The relevant schemes are different from traditional methods and can not only generate accurate and diverse recognition results, covering attributes and multi-level semantic information, but also have stronger generalization ability through vision-to-language learning, can process different visual contents more flexibly, and can effectively identify targets in open scenes; experimental results show that the scheme of the present invention can improve the open vocabulary image semantic segmentation performance on multiple data sets. In addition, the scheme of the present invention is also highly scalable and has the potential to be used as an automated system for automated annotation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of open vocabulary image semantic segmentation, and in particular to an open vocabulary image semantic segmentation method, system, device and storage medium. Background Art

[0002] With the explosive growth of image data, traditional image recognition methods face great challenges. Traditional image recognition technology relies on supervised learning and requires a large amount of manually labeled data, which not only has high requirements on the labeling cost, but also the labeling workload increases exponentially with the increase in the number of categories. Especially in the real world, there are many types of images, and some object categories even require expert knowledge for labeling, which further increases the difficulty and cost of labeling.

[0003] In recent years, visual recognition methods based on visual language models, such as the CLIP (Contrastive Language–Image Pre-training) model, have gradually attracted attention. Among them, OVSS (Open-Vocabulary Semantic Segmentation) is an emerging technology that can significantly reduce the need for labeled data and expand to unseen target categories. The traditional CLIP-based open-vocabulary semantic segmentation method realizes visual recognition of open scenes in a retrieval manner. During the model reasoning process, a fixed vocabulary list needs to be defined in advance. This method has great limitations in unknown open scenes, and the limited vocabulary set limits the expressive power of the model.

[0004] In view of this, the present invention is proposed. Summary of the invention

[0005] The purpose of the present invention is to provide an open vocabulary image semantic segmentation method, system, device and storage medium, which can automatically segment and identify the content of a visual local area without obtaining a vocabulary list.

[0006] The objective of the present invention is achieved through the following technical solutions:

[0007] An open vocabulary image semantic segmentation method, comprising:

[0008] Step 1: extract visual features and masks from the input image respectively, use the masks as attention masks, combine the attention mechanism to aggregate local information of visual features, obtain proposal features, and obtain features corresponding to the text based on the proposal features, which are called text-aware features;

[0009] Step 2: Project the text perception features from vision to language, and then input them into the language model to obtain the category name and region description contained in each region in the input image.

[0010] An open vocabulary image semantic segmentation system includes: an open vocabulary image semantic segmentation model, through which the open vocabulary image semantic segmentation model and the above-mentioned method are used to implement open vocabulary image semantic segmentation; the open vocabulary image semantic segmentation model includes:

[0011] The text-aware visual feature extraction module is used to extract visual features and masks from the input image respectively, use the masks as attention masks, and combine the attention mechanism to aggregate the local information of the visual features to obtain proposal features, and obtain the features corresponding to the text based on the proposal features, which are called text-aware features;

[0012] The category and description generation module based on the language model is used to project the text perception features from vision to language, and then input them into the language model to obtain the category name and region description contained in each region in the input image.

[0013] A processing device, comprising: one or more processors; a memory for storing one or more programs;

[0014] When the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned method.

[0015] A readable storage medium stores a computer program, which implements the above method when the computer program is executed by a processor.

[0016] It can be seen from the technical solution provided by the present invention that traditional open vocabulary recognition methods rely on predefined vocabulary sets and often lack sufficient expressive power to understand open world scenes. Unlike traditional methods, the present invention can generate accurate and diverse recognition results, covering attributes and multi-level semantic information, and has stronger generalization capabilities, can more flexibly handle different visual content, and can effectively identify targets in open scenes. Experimental results show that the present invention can improve the open vocabulary image semantic segmentation performance on multiple data sets. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0018] Figure 1A flowchart of an open vocabulary image semantic segmentation method provided by an embodiment of the present invention;

[0019] Figure 2 A schematic diagram of an open vocabulary image semantic segmentation model provided by an embodiment of the present invention;

[0020] Figure 3 A schematic diagram of the visualization of semantic segmentation results provided by an embodiment of the present invention;

[0021] Figure 4 A schematic diagram of verification of scalability provided by an embodiment of the present invention;

[0022] Figure 5 A schematic diagram of a processing device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0023] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the protection scope of the present invention.

[0024] First, the terms that may be used in this article are explained as follows:

[0025] The terms "include", "comprises", "contains", "has" or other descriptions with similar semantics should be interpreted as non-exclusive inclusion. For example, including certain technical feature elements (such as raw materials, components, ingredients, carriers, dosage forms, materials, dimensions, parts, components, mechanisms, devices, steps, procedures, methods, reaction conditions, processing conditions, parameters, algorithms, signals, data, products or products, etc.) should be interpreted as including not only certain technical feature elements explicitly listed, but also other technical feature elements known in the art that are not explicitly listed.

[0026] The term "consisting of..." means excluding any technical feature elements not explicitly listed. If this term is used in a claim, it will make the claim closed, so that it does not contain technical feature elements other than the technical feature elements explicitly listed, except for the conventional impurities related to them. If this term only appears in a clause of a claim, it only limits the elements explicitly listed in the clause, and the elements recorded in other clauses are not excluded from the overall claim.

[0027] The following is a detailed description of an open vocabulary image semantic segmentation method, system, device and storage medium provided by the present invention. The contents not described in detail in the embodiments of the present invention belong to the prior art known to professional and technical personnel in the field. If no specific conditions are specified in the embodiments of the present invention, the conventional conditions in the field or the conditions recommended by the manufacturer shall be followed. The instruments used in the embodiments of the present invention, if the manufacturer is not specified, are all conventional products that can be purchased commercially.

[0028] Embodiment 1

[0029] The embodiment of the present invention provides an open vocabulary image semantic segmentation method, such as Figure 1 As shown, it mainly includes the following steps:

[0030] Step 1: Extract text-aware features.

[0031] In an embodiment of the present invention, visual features and masks are extracted from the input image respectively, the mask is used as an attention mask, and the local information of the visual features is aggregated in combination with the attention mechanism to obtain proposal features, and based on the proposal features, features corresponding to the text are obtained, which are called text-aware features.

[0032] In the embodiment of the present invention, the purpose is to generate text corresponding to the input image region, that is, the text in the text corresponding feature here.

[0033] Preferably, this step can be implemented by a text-aware visual feature extraction module, in which a visual encoder and a pre-trained mask extractor are provided. The visual encoder is divided into layers, the front part of the layers is used to extract visual features from the input image, and the back part is used to extract category tags from the input image; the mask is extracted from the input image using a pre-trained mask extractor, and binarized, and the binarized mask is used as an attention mask, and the category tag is updated by combining the attention mechanism with the visual feature. Finally, the local information is aggregated through the updated category tag to obtain the proposal feature, and then the feature corresponding to the text is obtained from the proposal feature.

[0034] Step 2: Generate categories and descriptions.

[0035] In the embodiment of the present invention, the text perception features are projected from vision to language and then input into a language model to obtain the category name and region description contained in each region in the input image.

[0036] Preferably, this step can be implemented by a category and description generation module based on a language model, which includes a visual to language model projection module and a language model, the former being responsible for projecting the text perception features from vision to language, and the latter being responsible for using the projection results to generate the category name and region description contained in each region in the input image.

[0037] In the embodiment of the present invention, the text-aware visual feature extraction module and the category and description generation module based on the language model constitute an open vocabulary image semantic segmentation model; a training data set is pre-selected and constructed to train the model, specifically:

[0038] (1) Using data with text descriptions, a training dataset is constructed and the open vocabulary image semantic segmentation model is trained; the text descriptions include category names and sentence descriptions.

[0039] (2) During the training process, the input image is an image in the training data set. After obtaining the proposal features, the text-aware feature aggregation or selection technology is used according to the text description corresponding to the image to obtain the corresponding text-aware features; the text-aware features are input into the category and description generation module based on the language model, and the obtained semantic segmentation results containing the recognition category and description area are used to construct the language modeling loss to train the open vocabulary image semantic segmentation model.

[0040] After the training is completed, the open vocabulary image semantic segmentation is completed according to the above steps 1 and 2. In particular, after the training is completed, the proposal features are directly used as text-aware features.

[0041] The above-mentioned solution of the embodiment of the present invention can be deployed on a computer or a server to segment an image and accurately and diversely identify the segmented areas, and can output the category nouns and detailed sentence descriptions of the areas.

[0042] In order to more clearly demonstrate the technical solution and technical effects provided by the present invention, the open vocabulary image semantic segmentation model provided by the embodiment of the present invention is described in detail below with a specific embodiment. Figure 2 As shown in the figure, the open vocabulary image semantic segmentation model mainly includes: a text-aware visual feature extraction module and a category and description generation module based on a language model. In addition, considering that the model training is data-driven, it also includes a data engine part. The following mainly introduces the above three parts.

[0043] 1. Data engine part.

[0044] In the absence of a vocabulary set, the classification ability of the open vocabulary image semantic segmentation model depends largely on the available data. In the embodiment of the present invention, data with text descriptions are used to construct a training data set, specifically: a data set with category names and a data set with image-level sentence descriptions are collected.

[0045] For example, the datasets with category labels can be: a detection dataset (Visual Genome), which contains about 100,000 images; a segmentation dataset (COCO-Stuff), which contains about 100,000 images. The above detection datasets and segmentation datasets provide category labels that can be used directly.

[0046] Exemplarily, the dataset with image-level sentence descriptions can be an image-text pair dataset (CC3M subset), which includes approximately 600,000 images and corresponding image descriptions.

[0047] For datasets with image-level sentence descriptions, potential noun phrases (category names) are extracted from these descriptions by parsing and filtering.

[0048] (1) Parsing. Several noun phrases are extracted from the corresponding image-level sentence descriptions using NLP (Natural Language Processing) tools; however, this will introduce a large number of invalid nouns because, on the one hand, the titles from the Internet may contain noise and not completely match the images; on the other hand, many non-indicative terms such as "project" or "night" will also be included.

[0049] Exemplarily, the NLP tool may be spacy, which provides a set of efficient tools and models to process text data. The specific usage scheme may refer to conventional technologies, which will not be described in detail in the present invention.

[0050] (2) Filtering. Calculate the similarity between each noun phrase and the corresponding image, and sort the corresponding phrases in descending order of similarity. Filter out some noun phrases at the bottom of the sorting, and keep the remaining noun phrases.

[0051] Exemplarily, an image is marked into multiple (e.g., 100) image regions, which generally cover all semantic targets in the image. Then, the CLIP text embedding (i.e., the text features obtained by inputting the text encoder of the CLIP module) can be used to calculate the similarity between each noun phrase and each image region. For each noun phrase, the maximum value is selected from the similarities between it and all image regions as the similarity score (confidence) between the corresponding noun phrase and the image. Finally, according to the similarity scores of all noun phrases and the image, the noun phrases are sorted in descending order, and the bottom part with the lowest similarity is discarded. The proportion of the bottom part is set to .

[0052] 2. Text-aware visual feature extraction module.

[0053] In the text-aware visual feature extraction module, the visual encoder in the visual language model is first used to extract regional features adapted to the corresponding visual language model, and then text-aware feature aggregation or selection techniques are applied to obtain the required text-aware visual features (referred to as text-aware features for short).

[0054] In this embodiment, the visual language model can be selected as needed, and the CLIP model is used as an example for introduction. Figure 2 As shown in the figure, the visual encoder in the CLIP model is selected. During the training process, multiple images are input at a time (for example, 198 training set data set images are taken at a time and input into the model). The shallow layer of the visual encoder in the CLIP model is first used to extract visual features, and the deep layer [CLS] token is reused as a query for proposal feature extraction. That is, [CLS] token is used as a query for proposal feature extraction to extract proposal features from visual features. In short, the output corresponding to the input [CLS] token is the proposal feature, and the [CLS] token here is the category label. For example, in a 12-layer ViT-B / 16 CLIP model, the first 9 layers are unchanged (responsible for extracting visual features), and the last 3 layers are used to extract proposal features. ViT-B / 16 is a variant of the ViT (Vision Transformer) model. B indicates that it is a basic model with a medium-sized number of parameters, and 16 refers to the input of non-overlapping image blocks of 16x16 pixels.

[0055] To ensure compatibility with any segmentation model, this embodiment combines a binary mask to extract CLIP-adapted regional features. After extracting the mask from the input image using a pre-trained mask extractor, it is binarized to obtain N binary masks. (i.e., proposal mask), H and W represent the height and width of the mask, respectively. is the symbol of the real number set; for example, N=100 can be set.

[0056] Afterwards, the N binary masks are As an attention mask, it is input into the visual encoder in the CLIP model to mask the multi-head attention function and guide the [CLS] token to extract proposal features.

[0057] To match the attention mechanism function, the mask is updated as follows:

[0058] ;

[0059] Where i corresponds to the i-th binary mask and also corresponds to the i-th [CLS] token, i=1,…,N, j corresponds to the j-th visual feature, j=1,…,HW, represents the element corresponding to the jth visual feature in the i-th binary mask, is the element corresponding to the jth visual feature in the updated i-th binary mask. The above settings The purpose is to ignore the visual content of a specific area when subsequently calculating the softmax.

[0060] Repeat the [CLS] token N times, each [CLS] token corresponds to A binary mask in; for the i-th [CLS] token , combined with the attention mechanism for update:

[0061] ;

[0062] in, is the jth visual feature (image token), is the updated i-th binary mask, which contains HW elements, for The jth element of , and They are the query, key and value projections, which are trainable parameters (which can be understood as three different fully connected layers, or three learnable matrix parameters), and softmax is a normalized exponential function; is the assignment symbol, which means that the result of the calculation on the right is assigned to the left. Update: The update of all image tokens is the same as that in the standard CLIP, so it is not repeated here.

[0063] In the embodiment of the present invention, the number of [CLS] tokens is equal to the number of binary masks, and each [CLS] token (e.g., the i-th) interacts with the visual token and is affected by the corresponding Constraints,Finally, the updated [CLS] token can aggregate local information and obtain proposal features Specifically, the N updated category labels obtained by the method described above constitute the proposal features , where C is the dimension of the proposal feature.

[0064] During the training process, it is necessary to extract features corresponding to the text from the proposal features, so as to facilitate the subsequent language model training. In this embodiment, different extraction strategies are designed for different data set formats.

[0065] (1) Text-aware feature aggregation.

[0066] If the category description corresponding to the image does not specify the location area in the input image, such as image-level title data and extracted noun phrases, in this case, it is necessary to use text-aware feature aggregation technology to obtain the corresponding text-aware features. The steps include:

[0067] (1.1) Use a text encoder in a visual language model (e.g., CLIP text encoder) to extract embedding vectors from category descriptions , the category description here includes: sentence description, and sentences constructed by sentence templates (This is an image of a [noun phrase]) from noun phrases extracted from sentence descriptions, and M represents the number of texts associated with the image. Taking the CLIP model as an example, the text input into the text encoder of CLIP contains two types, one is the sentence description that comes with the dataset, and the other is the noun phrase. At this time, the noun phrase is also constructed into the corresponding sentence based on a fixed template. For example, the sentence description is The cat is running on a table [Sentence 1], the extracted noun phrases are cat and table, and the constructed sentences are This is an image of a cat [Sentence 2] and This is animage of a table [Sentence 3]; then [Sentences 1, 2, 3] will all be input into the CLIP text encoder.

[0068] Thanks to the alignment space of the visual language model, the embedding vector E and the proposal features can be calculated The similarity between:

[0069] ;

[0070] Among them, softmax is a normalized exponential function, temp is a temperature coefficient (for example, it can be set to 5), S is a similarity score, and T is a transposed sign.

[0071] The proposal features are aggregated into text-aware features using the similarity score S:

[0072] ;

[0073] Among them, K is the text perception feature and normalize is the normalization function.

[0074] The above process guides the aggregation of visual features through text data, so that the subsequent stage can generate text from these text-aware visual features (referred to as text-aware features for short).

[0075] (2) Text-aware feature selection.

[0076] There are two types of region-level text data, including text data associated with masks (segmentation datasets) or bounding boxes (detection datasets). That is, the category description corresponding to the image refers to the area selected by the mask or bounding box. In this case, text-aware feature selection technology is used to obtain the corresponding text-aware features.

[0077] (2.1) How to handle the area selected by the mask: the included mask is the real mask , similar to Maskformer (a deep learning model for visual tasks), the true mask is binary matched with the extracted mask. Each extracted mask corresponds to a proposal feature, and the proposal feature corresponding to the mask that best matches the true mask is used as the text-aware feature.

[0078] Since there is no visual language model classifier involved, only and The matching cost is calculated by the intersection over union (IoU) between them.

[0079] (2.2) How to deal with the area selected by the bounding box: the bounding box is the true bounding box Each extracted mask corresponds to a proposal feature. Based on each extracted mask, the corresponding bounding rectangle Z is determined, and the relationship between each bounding rectangle Z and the true bounding box is calculated. The intersection-and-union ratio is used to perform a binary matching-based allocation to find the boundary box that matches the real boundary box. The best matching mask,the proposal features corresponding to the best matching mask are used as,text-aware features.

[0080] The text-aware features and selection techniques introduced above are only applied in the training phase to facilitate the learning of the subsequent language model. In the testing phase, since the model is dataset-independent, all proposal features are treated as text-aware features, i.e. .

[0081] 3. Category and description generation module based on language model.

[0082] In this embodiment, the category and description generation module based on the language model mainly includes: a projection module from vision to language model and a language model.

[0083] After obtaining the text-aware features, the projection module of the vision-to-language model is used to adjust them to fit the language model as follows:

[0084] ;

[0085] ;

[0086] in, and are all projection matrices, K is the text perception feature, ReLU is the linear rectification function, H is the intermediate result, is the projection result from vision to language, is the dimension of the projection result, is the layer normalization function.

[0087] In order to enable the language model to distinguish between the output category label and description, this embodiment constructs different language model input instruction templates: "Identify category: [projection result ]" and "Describe area: [Projection results ]”.

[0088] Exemplarily, FlanT5-base may be used as a language model, which is a language model based on the T5 (Text-to-Text Transfer Transformer) architecture and has an encoder-decoder architecture.

[0089] In this embodiment, the language model is optimized using the standard language modeling loss, cross entropy loss, to ensure that it can effectively understand the provided text and visual information.

[0090] In this embodiment, the parameters of the entire open vocabulary image semantic segmentation model only adjust the category and description generation module based on the language model, including the projection module of the vision-to-language model and the language model (marked with an additional flame symbol). The training loss is only the language modeling loss (cross entropy loss) of the aforementioned language model.

[0091] The above solution provided by the embodiment of the present invention mainly achieves the following beneficial effects:

[0092] The present invention can generate accurate and diverse recognition results, covering attributes and multi-level semantic information. Compared with related methods, the present method achieves an average performance improvement of 13.7% on five datasets including VOC, PC-59, ADE-150, PC-459, and ADE-847. The sentences generated by the present invention can understand high-level semantics such as location and relationship, which promotes a more complete understanding of the target. In addition, the present invention is also highly scalable and has the potential to be used as an automated system for automated annotation.

[0093] In order to illustrate the effect of the above-mentioned scheme of the present invention, Figure 3The visualization results shown are Figure 3 The labels in the first row are nouns or noun phrases, and the labels in the second row are sentence descriptions. The open vocabulary image semantic segmentation model provided by the present invention can achieve accurate and diverse expression of regional content. At the same time, Figure 4 The verification results of the scalability shown can also show that the present invention has high scalability. Figure 4 In the paper, the open vocabulary image semantic segmentation model and SAM (SegmentAnything Model) are combined to realize interactive segmentation and recognition. The open vocabulary image semantic segmentation model can recognize the area selected by the user and perform accurate perception and expression.

[0094] It should be noted that the solution provided in the embodiment of the present invention can be applicable to different language types, for example, Figure 2 The Chinese language shown in the lower right corner can also be used for Figure 3~Figure 4 The English language shown can, of course, be expanded to other languages ​​as needed, which will not be elaborated here.

[0095] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above embodiments can be implemented by software, or by means of software plus necessary general hardware platforms. Based on such understanding, the technical solutions of the above embodiments can be embodied in the form of software products, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0096] Embodiment 2

[0097] The present invention also provides an open vocabulary image semantic segmentation system, which is mainly used to implement the method provided in the above embodiment. The system mainly includes: an open vocabulary image semantic segmentation model, and the open vocabulary image semantic segmentation is implemented by using the open vocabulary image semantic segmentation model and the method described above. Figure 2 , the open vocabulary image semantic segmentation model includes:

[0098] The text-aware visual feature extraction module is used to extract visual features and masks from the input image respectively, use the masks as attention masks, and combine the attention mechanism to aggregate the local information of the visual features to obtain proposal features, and obtain the features corresponding to the text based on the proposal features, which are called text-aware features;

[0099] The category and description generation module based on the language model is used to project the text perception features from vision to language, and then input them into the language model to obtain the category name and region description contained in each region in the input image.

[0100] Considering that the main technical details involved in the system have been introduced in detail in the previous embodiments, they will not be repeated here; in addition, the open vocabulary image semantic segmentation model in the system needs to be trained in advance, and some of the work processes involved in the training can be implemented by configuring corresponding modules in the system, which will not be repeated here.

[0101] Technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the system can be divided into different functional modules to complete all or part of the functions described above.

[0102] Embodiment 3

[0103] The present invention also provides a processing device, such as Figure 5 As shown, it mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the methods provided in the aforementioned embodiments.

[0104] Furthermore, the processing device also includes at least one input device and at least one output device; in the processing device, the processor, memory, input device, and output device are connected via a bus.

[0105] In the embodiment of the present invention, the specific types of the memory, input device and output device are not limited; for example:

[0106] The input device may be a touch screen, an image acquisition device, a physical button or a mouse, etc.;

[0107] The output device may be a display terminal;

[0108] The memory may be a random access memory (RAM) or a non-volatile memory such as a disk storage.

[0109] Embodiment 4

[0110] The present invention also provides a readable storage medium storing a computer program, which implements the method provided in the above embodiment when the computer program is executed by a processor.

[0111] In the embodiment of the present invention, the readable storage medium as a computer-readable storage medium can be set in the aforementioned processing device, for example, as a memory in the processing device. In addition, the readable storage medium can also be a U disk, a mobile hard disk, a read-only memory (ROM), a disk or an optical disk, etc., which can store program codes.

[0112] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by any technician familiar with the technical field within the technical scope disclosed in the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims. The information disclosed in the background technology section of this article is only intended to deepen the understanding of the overall background technology of the present invention, and should not be regarded as an admission or in any form that the information constitutes prior art known to those skilled in the art.

Claims

1. An open vocabulary image semantic segmentation method, characterized in that: include: Step 1: extract visual features and masks from the input image respectively, use the masks as attention masks, combine the attention mechanism to aggregate local information of visual features, obtain proposal features, and obtain features corresponding to the text based on the proposal features, which are called text-aware features; Step 2: Project the text perception features from vision to language, and then input them into the language model to obtain the category name and region description contained in each region in the input image; The extraction of visual features and masks from the input image, taking the masks as attention masks, and combining the attention mechanism to aggregate local information of visual features, obtains proposal features including: The visual encoder in the visual language model is introduced and divided into layers. The front part of the layers is used to extract visual features from the input image, and the back part is used to extract category tags from the input image. The mask is extracted from the input image using the pre-trained mask extractor and binarized to obtain N binary masks. , H and W represent the height and width of the mask respectively, is the symbol of the real number set; The N binary masks As the attention mask, the mask update is performed as follows: ; Among them, i corresponds to the i-th binary mask and also corresponds to the i-th category label, i=1,…,N; j corresponds to the j-th visual feature, j=1,…,HW, represents the element corresponding to the jth visual feature in the i-th binary mask, is the element corresponding to the jth visual feature in the updated i-th binary mask; Each category label corresponds to A binary mask in ; for the i-th category label , combined with the attention mechanism for update: ; in, is the jth visual feature, is the updated i-th binary mask, which contains HW elements, for The jth element of , and are the projections of query, key and value respectively, and softmax is the normalized exponential function; is the assignment symbol, which means that the result of the calculation on the right is assigned to the left. renew; Finally, the proposal features are obtained by aggregating local information through the updated N category labels , where C is the dimension of the proposal feature.

2. The open vocabulary image semantic segmentation method according to claim 1, characterized in that: The step 1 is implemented by a text-aware visual feature extraction module, and the step 2 is implemented by a category and description generation module based on a language model, wherein the text-aware visual feature extraction module and the category and description generation module based on a language model constitute an open vocabulary image semantic segmentation model; Using data with text descriptions, a training data set is constructed and the open vocabulary image semantic segmentation model is trained; the text descriptions include category names and category names; During the training process, the input image is an image in the training data set. After obtaining the proposal features, text-aware feature aggregation or selection technology is used according to the category description corresponding to the image to obtain the corresponding text-aware features. After the training is completed, the proposal features will be used as text-aware features; the text-aware features are input into the category and description generation module based on the language model, and the obtained semantic segmentation results containing the recognition category and description area are used to construct a language modeling loss to train the open vocabulary image semantic segmentation model.

3. The open vocabulary image semantic segmentation method according to claim 2, characterized in that: The method of constructing a training data set using data with text descriptions includes: Collect data sets with category names and data sets with image-level sentence descriptions; obtain corresponding noun phrases by parsing and filtering the data sets with image-level sentence descriptions; During the parsing process, a number of noun phrases are extracted from the corresponding image-level sentence descriptions through natural language processing tools; during the filtering process, the similarity between each noun phrase and the corresponding image is calculated respectively, and the corresponding phrases are arranged in order from high to low similarity, and some noun phrases at the bottom of the order are filtered out, and the remaining noun phrases are retained.

4. The open vocabulary image semantic segmentation method according to claim 2 or 3, characterized in that: The method of using a text-aware feature aggregation or selection technique according to the category description corresponding to the image to obtain the corresponding text-aware feature includes: If the category description corresponding to the image does not specify the location area in the input image, a text-aware feature aggregation technique is used to obtain the corresponding text-aware feature, the steps comprising: Use the text encoder to extract the embedding vector E from the category description and calculate the embedding vector E and the proposal feature The similarity between: ; Among them, softmax is the normalized exponential function, temp is the temperature coefficient, S is the similarity score, and T is the transposed sign; The proposal features are aggregated into text-aware features using the similarity score S: ; Among them, K is the text perception feature and normalize is the normalization function.

5. The open vocabulary image semantic segmentation method according to claim 2 or 3, characterized in that: The method of using a text-aware feature aggregation or selection technique according to the category description corresponding to the image to obtain the corresponding text-aware feature includes: If the category description corresponding to the image refers to the area selected by the mask or bounding box, the text-aware feature selection technique is used to obtain the corresponding text-aware features, including: How to handle when referring to the area selected by the mask: the included mask is the real mask , perform binary matching between the true mask and the extracted mask, each extracted mask corresponds to a proposal feature, and the proposal feature corresponding to the mask that best matches the true mask is used as the text-aware feature; How to handle when referring to the area selected by the bounding box: the included bounding box is the true bounding box Each extracted mask corresponds to a proposal feature. Based on each extracted mask, the corresponding bounding rectangle Z is determined, and the relationship between each bounding rectangle Z and the true bounding box is calculated. The intersection-and-union ratio is used to perform a binary matching-based allocation to find the boundary box that matches the real boundary box. The best matching mask,the proposal features corresponding to the best matching mask are used as,text-aware features.

6. The open vocabulary image semantic segmentation method according to claim 1, characterized in that: The visual-to-linguistic projection of the text-aware features is expressed as: ; ; in, and are all projection matrices, K is the text perception feature, ReLU is the linear rectification function, is the intermediate result, is the projection result from vision to language, is the layer normalization function.

7. An open vocabulary image semantic segmentation system, characterized in that include: An open vocabulary image semantic segmentation model, wherein the open vocabulary image semantic segmentation is achieved by using the open vocabulary image semantic segmentation model and the method according to any one of claims 1 to 6; The open vocabulary image semantic segmentation model includes: The text-aware visual feature extraction module is used to extract visual features and masks from the input image respectively, use the masks as attention masks, and combine the attention mechanism to aggregate the local information of the visual features to obtain proposal features, and obtain the features corresponding to the text based on the proposal features, which are called text-aware features; The category and description generation module based on the language model is used to project the text perception features from vision to language, and then input them into the language model to obtain the category name and region description contained in each region in the input image.

8. A processing device, characterized in that: include: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.

9. A readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Scene character recognition method based on visual language modeling network

    CN112541501A

  • Reference image segmentation method and device based on progressive visual features and storage medium

    CN117788814A