A model generation method, an image processing method, and related devices
By using pre-trained visual language models and learnable initial prompt features, a zero-sample semantic segmentation model is generated, solving the problem that existing models cannot perform semantic segmentation of unlabeled categories, and achieving high accuracy and extensive one-pixel-level semantic segmentation.
Patent Information
- Application Number
- CN202311224084.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-21
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-09-21
AI Technical Summary
The existing semantic segmentation model cannot perform semantic segmentation on unlabeled categories, and the application of visual language models in semantic segmentation tasks is limited, so it is impossible to achieve pixel-level semantic segmentation.
The pre-trained visual language model obtains the feature image of the sample image annotated with the real segmentation mask, and based on the preset multiple learnable initial prompt features and category features, the feature image is convolutional, and the predicted segmentation mask is generated, and the initial prompt features are adjusted to generate a zero-sample semantic segmentation model.
Zero-sample semantic segmentation is realized, and the images with unlabeled categories can be semantic segmented, covering the entire target area, improving the accuracy and breadth of semantic segmentation.
Smart Images

Figure CN117726808B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to a model generation method, an image processing method, and related devices. Background Art
[0002] The semantic segmentation task refers to inputting an image and assigning a class label to each pixel point in the image. Existing semantic segmentation models can often only perform semantic segmentation on images containing classes with annotations in the training dataset, and cannot perform semantic segmentation on images of unannotated classes. In order to enable the semantic segmentation model to have the ability to recognize more classes, it is often necessary to collect more images of different classes for semantic annotation and then retrain the model. During the data collection process, for some classes, there may also be problems with not being able to collect enough datasets, which hinders the popularization and application of semantic segmentation models in reality.
[0003] However, visual language models (Contrastive Language-Image Pre-training (CLIP) model, ALarge-scaleImaGe and Noisy-text embedding (ALIGN) model) have been proposed in the prior art, which can perform zero-shot and open-ended classification on any object. However, since the training text of the visual language model mainly describes the global context of the image and is mainly applicable to image-level classification, it cannot achieve pixel-level semantic segmentation tasks. Therefore, how to apply the visual language model to the semantic segmentation task to achieve zero-shot semantic segmentation has become an urgent technical problem to be solved. Summary of the Invention
[0004] Based on the above research, the present application provides a model generation method, an image processing method, and related devices, which can apply the visual language model to the semantic segmentation task to achieve zero-shot semantic segmentation.
[0005] The embodiments of the present application can be implemented in the following ways:
[0006] In a first aspect, an embodiment of the present application provides a model generation method, and the method includes:
[0007] Obtain a first training sample, where the first training sample is a sample image with a labeled true segmentation mask;
[0008] Based on a pre-trained visual language model, obtain a feature image of the first training sample;
[0009] Performing convolution processing on the feature image of the first training sample based on a plurality of preset learnable first initial prompt features and the category features of each preset category to obtain a predicted segmentation mask of the first training sample; the first initial prompt features are used to form a natural language statement in combination with a preset category;
[0010] Based on the predicted segmentation mask and the ground truth segmentation mask of the first training sample, adjusting the first initial prompt features to obtain first target prompt features;
[0011] Generating a zero-shot semantic segmentation model based on the first target prompt features.
[0012] In an alternative embodiment, the performing convolution processing on the feature image of the first training sample based on a plurality of preset learnable first initial prompt features and the category features of each preset category to obtain a predicted segmentation mask of the first training sample includes:
[0013] For each of the first initial prompt features, performing convolution processing on the feature image of the first training sample according to the first initial prompt feature and the category features of each preset category to obtain a first segmentation mask of the first training sample;
[0014] Performing convolution processing on the feature image of the first training sample based on a preset learnable second initial prompt feature and each of the category features to obtain the probabilities of each of the preset categories appearing in the first training sample; the second initial prompt feature is used to represent the probability of each of the preset categories appearing in each image;
[0015] Based on the first segmentation mask and the probability of each of the preset categories appearing, obtaining the predicted segmentation mask of the first training sample.
[0016] In an alternative embodiment, the method further includes:
[0017] Calculating a segmentation loss function and a classification loss function respectively according to the ground truth segmentation mask and the predicted segmentation mask of the first training sample;
[0018] Based on the segmentation loss function and the classification loss function, adjusting the second initial prompt features to obtain second target prompt features.
[0019] In an alternative embodiment, the performing convolution processing on the feature image of the first training sample based on a preset learnable second initial prompt feature and each of the category features to obtain the probabilities of each of the preset categories appearing in the first training sample includes:
[0020] Concatenate the second initial prompt features with the category features of each of the preset categories respectively to obtain a plurality of classification features;
[0021] Perform convolution processing on the feature image of the first training sample based on each of the classification features to obtain the probabilities of each of the preset categories appearing in the first training sample.
[0022] In an alternative embodiment, the performing convolution processing on the feature image of the first training sample based on a plurality of preset learnable first initial prompt features and the category features of each preset category to obtain the predicted segmentation mask of the first training sample includes:
[0023] Concatenate each of the first initial prompt features with the category features of each of the preset categories to obtain a plurality of segmentation features;
[0024] Perform convolution processing on the feature image of the first training sample based on each of the segmentation features to obtain a plurality of second segmentation masks;
[0025] Add the second segmentation masks corresponding to each of the first initial prompt features according to pixel points to obtain the predicted segmentation mask of the first training sample.
[0026] In an alternative embodiment, the adjusting the first initial prompt features based on the predicted segmentation mask and the true segmentation mask of the first training sample to obtain the first target prompt features includes:
[0027] Adjust the first initial prompt features based on the predicted segmentation mask and the true segmentation mask of the first training sample to obtain the adjusted first initial prompt features;
[0028] If the adjusted first initial prompt features meet the first preset condition and the preset orthogonal constraint condition, set the adjusted first initial prompt features as the first target prompt features;
[0029] Wherein, the preset orthogonal constraint condition is that the average product of feature points of each of the first initial prompt features is 0.
[0030] In an alternative embodiment, the generating a zero-shot semantic segmentation model based on the first target prompt features includes:
[0031] Obtain a sample image with an unlabeled or partially labeled true segmentation mask;
[0032] Based on the first target prompt features and the vision-language model, obtain a pseudo-segmentation mask of the sample image with the unlabeled or partially labeled true segmentation mask;
[0033] Set the sample image with the pseudo-segmentation mask as the second training sample;
[0034] Generate a zero-shot semantic segmentation model based on the first training sample, the second training sample, and the first target prompt feature.
[0035] In an alternative embodiment, the generating a zero-shot semantic segmentation model based on the first training sample, the second training sample, and the first target prompt feature includes:
[0036] Generate a target classifier based on the first target prompt feature and the category features of each preset category; and obtain the feature extraction network of the pre-trained semantic segmentation model;
[0037] Combine the target classifier and the feature extraction network to form an initial semantic segmentation model;
[0038] According to the initial semantic segmentation model, perform segmentation masks on the first training sample and the second training sample respectively to determine the predicted segmentation masks of the first training sample and the second training sample;
[0039] Based on the predicted segmentation masks of the first training sample and the second training sample, adjust the parameters of the feature extraction network to obtain a semantic segmentation backbone network;
[0040] Construct the zero-shot semantic segmentation model according to the target classifier and the semantic segmentation backbone network.
[0041] In an alternative embodiment, before obtaining the feature image of the first training sample based on the image encoder in the pre-trained vision-language model, the method further includes:
[0042] Obtain the feature dimension output by the image encoder;
[0043] If the feature dimension is less than a preset threshold, delete the attention pooling layer in the image encoder.
[0044] In a second aspect, an embodiment of the present application provides an image processing method, and the image processing method includes:
[0045] Obtain an image to be processed;
[0046] Perform semantic segmentation on the image to be processed based on a zero-shot semantic segmentation model to obtain a predicted segmentation mask of the image to be processed;
[0047] Wherein, the zero-shot semantic segmentation model is obtained by the model generation method described in any one of the above.
[0048] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the model generation method or the image processing method described in any of the above embodiments is implemented.
[0049] In a fourth aspect, an embodiment of the present application further provides a readable storage medium, where the readable storage medium includes a computer program, and when the computer program runs, it controls the electronic device where the readable storage medium is located to execute the model generation method or the image processing method described in any of the above embodiments.
[0050] For the model generation method, image processing method and related devices provided by the present application, through a pre-trained vision-language model, a feature image of a first training sample is obtained, and then through a plurality of first initial prompt features preset to be able to form natural language statements in combination with preset categories and category features of each preset category, convolutional processing is performed on the feature image of the first training sample to obtain a predicted segmentation mask of the first training sample. Then, through the predicted segmentation mask and the true segmentation mask of the first training sample, the first initial prompt feature is adjusted to obtain a first target prompt feature, and based on the first target prompt feature, a zero-shot semantic segmentation model is generated. Based on the above embodiments, in the first stage, a plurality of first target prompt features are first trained using a vision-language model with zero-shot ability, and then in the second stage, a zero-shot semantic segmentation model is generated based on the first target prompt information. On the one hand, the vision-language model is applied to the semantic segmentation task to achieve zero-shot semantic segmentation; on the other hand, compared with directly using artificially designed prompts in the CLIP model for pixel-level segmentation tasks, the zero-shot semantic segmentation model provided in this embodiment does not only focus on the most different regions in the image, but can make the generated segmentation mask cover the entire target. Description of the Drawings
[0051] The following will make the technical solutions and other beneficial effects of the present application obvious by describing the specific embodiments of the present application in detail in conjunction with the drawings.
[0052] Figure 1 It is a schematic flowchart of the model generation method provided by an embodiment of the present application;
[0053] Figure 2 It is an application scenario diagram provided by an embodiment of the present application;
[0054] Figure 3 It is provided by an embodiment of the present application Figure 1 The schematic flowchart of step S103 in
[0055] Figure 4 It is provided by an embodiment of the present application Figure 1Flow diagram of step S105 in
[0056] Figure 5 Provided by an embodiment of the present application Figure 4 Flow diagram of step S404 in
[0057] Figure 6 Provided by an embodiment of the present application Figure 1 Another flow diagram of step S103 in
[0058] Figure 7 Schematic diagram of the model generation method taking the CLIP model as an example in an embodiment of the present application
[0059] Figure 8 Flow diagram of the image processing method provided by an embodiment of the present application
[0060] Figure 9 Schematic diagram of the structure of the electronic device provided by an embodiment of the present application Detailed implementation manners
[0061] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0062] In the description of the present application, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the present application, "a plurality of" means two or more, unless otherwise specifically defined.
[0063] The following disclosure provides many different implementation manners or examples to implement different structures of the present application. To simplify the disclosure of the present application, the components and settings of specific examples are described below. Of course, they are only examples and are not intended to limit the present application. In addition, the present application may repeat reference numerals and / or reference letters in different examples. This repetition is for the purpose of simplification and clarity, and does not itself indicate the relationship between the various implementation manners and / or settings discussed. In addition, the present application provides examples of various specific processes and materials, but those skilled in the art can be aware of the application of other processes and / or the use of other materials.
[0064] As described in the background art, in complex downstream application scenarios, a semantic segmentation model is required to complete the semantic segmentation task of the open set without specific category annotation. For example, semantic segmentation is performed on the captured images in the terminal shooting scenario, and semantic segmentation is performed on the images collected during the automatic driving process. However, the current speech segmentation models often complete the semantic segmentation task with specific category annotation and cannot achieve zero-shot semantic segmentation.
[0065] Due to the remarkable progress of vision-language models (such as CLIP model, ALIGN model), vision-language models can perform zero-shot, open-set classification on arbitrary objects with astonishingly high accuracy. However, it is not easy to apply vision-language models to semantic segmentation. Vision-language models are image-level classification, while semantic segmentation is pixel-level classification. Directly calculating and using the artificially designed prompt features for vision-language models for semantic segmentation will also cause them to pay more attention to the regions with the most category differences in the image, resulting in the segmentation mask not covering the entire target and affecting the effect of zero-shot semantic segmentation.
[0066] To solve the above technical problems, the embodiments of the present application provide a model generation method, an image processing method and related devices. The feature image of the sample image labeled with the real segmentation mask is obtained through a pre-trained vision-language model, and then the feature image of the first training sample is convolved based on the category features of each preset category and a plurality of pre-set first initial prompt features that are learnable and can form natural language sentences when combined with the preset categories to determine the predicted segmentation mask of the first training sample. Then, based on the real segmentation mask and the predicted segmentation mask of the first training sample, the first initial prompt feature is optimized and adjusted to obtain a first target prompt feature that meets the first preset condition. Finally, a zero-shot semantic segmentation model can be generated based on the first target prompt feature, and the segmentation mask of the image to be processed can be obtained by inputting the image to be processed into the zero-shot semantic segmentation model. In this way, a plurality of learnable first target prompt features are first obtained through the vision-language model with zero-shot ability and the first training sample, and then a zero-shot semantic segmentation model is constructed using the first target prompt feature, adapting the zero-shot image-level classification of the vision-language model to the zero-shot pixel-level feature, so as to be applicable to more downstream application scenarios of zero-shot semantic segmentation and promote the development of downstream industries.
[0067] The model generation method, the image processing method and related devices provided by the present application are described in detail below through specific embodiments.
[0068] Please refer to Figure 1 , Figure 1The flowchart of the model generation method provided by the embodiment of the present application. The model generation method can be executed by a model generation device, which can be implemented in software and / or hardware, and the model generation device can be configured in an electronic device. Specifically, the electronic device can be a server or a terminal device. Among them, the server can be an independent server or a server network or server cluster composed of servers, including but not limited to a computer, a network host, a single network server, multiple network servers, or a cloud server composed of multiple servers. Among them, the cloud server is composed of a large number of computers or network servers based on cloud computing. The terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, a game console, or a personal computer (PC), etc. As Figure 2 shown, this application scenario includes a terminal 210 and a server 220. The server 220 can be used to execute the model generation method described in this embodiment to obtain a zero-shot semantic segmentation model for zero-shot semantic segmentation. Further, the zero-shot semantic segmentation model obtained through the model generation method is deployed to the terminal 210, so that the terminal 210 can perform image semantic segmentation according to the zero-shot semantic segmentation model, or the terminal 210 can send the image to be processed to the server 220, so that the server 220 performs image semantic segmentation on the image to be processed; or, the server 220 can perform image recognition on the images to be recognized from local, other terminals, or other servers.
[0069] As Figure 1 shown, the model generation method of this embodiment at least includes the following steps:
[0070] Step S101, obtain the first training sample.
[0071] Among them, the above first training sample is a sample image with a labeled true segmentation mask. The true segmentation mask of the sample image can be manually labeled or obtained by performing semantic segmentation on the sample image through an existing semantic segmentation model that can only perform semantic segmentation on labeled categories. There is no specific limitation in this embodiment.
[0072] For example, multiple first training samples can be obtained from a pre-set sample dataset, and several sample images with labeled true segmentation masks are pre-stored in the sample dataset.
[0073] S102, based on the pre-trained vision-language model, obtain the feature image of the first training sample.
[0074] The above-mentioned pre-trained vision-language model refers to, for example, the CLIP model, the ALIGN model, etc., all of which include an image encoder and a text encoder. The image encoder is used to extract image features and map the image to a feature space; the text encoder is used to extract text features and map the text to the same feature space.
[0075] In this embodiment, the first training sample can be input into the image encoder of the pre-trained vision-language model to extract features of the first training sample and output a feature image of the first training sample.
[0076] In some alternative embodiments, before step S102, the model generation method further includes: obtaining the feature dimension output by the image encoder; if the feature dimension is less than a preset threshold, deleting the attention pooling layer in the image encoder.
[0077] Since most of the currently publicly available vision-language models often adopt an attention pooling layer on the last feature map of their image encoders to obtain a representation vector (i.e., a feature image) of the input image in order to achieve relatively accurate image-level classification of images, for the implementation of the pixel-level segmentation task in this embodiment, if the dimension of extracting image features from the first training sample is small, it will affect the accuracy of the pixel-level segmentation task. Therefore, before obtaining the feature image of the first training sample by using the image encoder of the vision encoder, it is determined whether to delete its attention pooling layer through the feature dimension output by the image encoder, so as to further improve the accuracy of semantic segmentation of the generated zero-shot semantic segmentation model.
[0078] S103, perform convolution processing on the feature image of the first training sample based on a plurality of preset learnable first initial prompt features and the category features of each preset category to obtain a predicted segmentation mask of the first training sample.
[0079] The above-mentioned preset categories are the categories that each pixel point in the preset image may have, such as: dog, cat, tree, car, etc.
[0080] The above-mentioned learnable first initial prompt features are used to form a natural language statement in combination with the preset categories. Specifically, the prompt text corresponding to the first initial prompt feature is combined with the category words of the preset categories to form a natural language statement, such as: X X X X dog, where X is the prompt text corresponding to the above-mentioned first initial extraction feature.
[0081] In this embodiment, k learnable prompts can be randomly generated as the first initial extraction feature p i , where i = 1, 2,... k. Each first initial prompt feature is multi-dimensional, for example: p i = R32×512 。
[0082] It can be understood that randomly generating k learnable prompts as the first initial extraction features only needs to be executed before step S103.
[0083] Optionally, as Figure 3 shown, the above step S103 can be implemented through the following steps:
[0084] S301, concatenate each first initial prompt feature with the category features of each preset category to obtain multiple segmentation features.
[0085] In this embodiment, for each preset category, the category words of the preset category can be first subjected to feature extraction to obtain the category features of the preset category, and then the category features of each preset category c are concatenated with each first initial prompt feature to obtain the segmentation features Specifically, as shown in the following formula:
[0086]
[0087] where CONCAT(·) represents concatenation, cls c represents the category word of the preset category c, and W(·) represents mapping the category word cls c to the feature vector.
[0088] Optionally, the category words of the preset category can be subjected to feature extraction through the text encoder of the vision-language model, and the category features of each preset category can be obtained by inputting the category words of each preset category into the text encoder of the vision-language model.
[0089] S302, perform convolution processing on the feature images of the first training samples based on each segmentation feature to obtain multiple second segmentation images.
[0090] In this embodiment, a corresponding number of convolution kernels can be selected from a plurality of pre-set convolution kernels as the above-mentioned preset convolution kernels. The number of the preset convolution kernels is the same as the number of the segmentation features. For example: 3 segmentation features: Select 3 convolution kernels from a plurality of pre-set convolution kernels as the above-mentioned preset convolution kernels (including: preset convolution kernel A, preset convolution kernel B, and preset convolution kernel C). And, the segmentation features correspond to the preset convolution kernels one by one. For each segmentation feature, the segmentation feature is used as the parameter of the corresponding preset convolution kernel. For example, the segmentation feature is used as the parameter of convolution kernel A, is used as the parameter of convolution kernel B, is used as the parameter of convolution kernel C. The above-mentioned preset convolution kernels can be 1×1 convolution kernels.
[0091] Then, the feature images of the first training samples are respectively input into each preset convolution kernel with the segmentation feature as a parameter for convolution calculation, and the obtained segmentation mask is set as the second segmentation mask. The second segmentation masks are grouped according to the first initial prompt features corresponding to the segmentation features, and multiple second segmentation masks corresponding to each first initial prompt feature are obtained.
[0092] Optionally, each segmentation feature can be first input into the text encoder of the vision-language model for feature extraction to obtain the segmentation feature after feature extraction, and then the feature images of the first training samples are convolved based on the segmentation feature after feature extraction to obtain multiple second segmentation masks.
[0093] S303. Add the second segmentation masks corresponding to each first initial prompt feature pixel by pixel to obtain the predicted segmentation mask of the first training sample.
[0094] In this embodiment, for the multiple second segmentation masks corresponding to each first initial prompt feature, the multiple second segmentation masks corresponding to the first initial prompt feature are added pixel by pixel to obtain a to-be-determined segmentation mask; then, through the class probability of each pixel, the to-be-determined segmentation mask with the largest class probability is selected as the predicted segmentation mask of the first training sample, so as to obtain the predicted segmentation mask of the first training sample.
[0095] S104. Based on the predicted segmentation mask and the true segmentation mask of the first training sample, adjust the first initial prompt feature to obtain the first target prompt feature.
[0096] In this embodiment, based on the predicted segmentation mask and the true segmentation mask of the first training sample, the parameters of each first initial prompt feature are adjusted to obtain the adjusted first initial prompt feature, and referring to the above steps S102 to S103, a new predicted segmentation mask of the first training sample is obtained based on the adjusted first initial prompt feature, and the parameters of the first initial prompt feature are continuously adjusted according to the predicted segmentation mask and the true segmentation mask of the first training sample, and so on, until the adjusted first initial prompt feature meets the first preset condition, and the first initial prompt feature that meets the first preset condition is used as the above first target extraction feature.
[0097] The above first preset condition is to meet the preset number of training times and / or meet the preset loss function value. Among them, meeting the preset number of training times means that the number of times of adjusting the first initial prompt feature is greater than the preset number of training times, and meeting the preset loss function value means that the loss function is less than the preset loss function value. The loss function mentioned here can be, for example, the cross-entropy loss function. This loss function can calculate the corresponding loss function according to the predicted segmentation mask and the true segmentation mask of the first training sample, and this loss function can be, for example, the cross-loss function.
[0098] In some alternative embodiments, step S104 may also be implemented in the following manner:
[0099] Based on the predicted segmentation mask and the ground truth segmentation mask of the first training sample, adjust the first initial prompt feature to obtain the adjusted first initial prompt feature; if the adjusted first initial prompt feature meets the first preset condition and the preset orthogonality constraint condition, then set the adjusted first initial prompt feature as the first target prompt feature.
[0100] The above-mentioned preset orthogonality constraint condition is that the average feature point product of each first initial prompt feature is 0. As can be seen from the above embodiments, each first initial prompt feature is multi-dimensional. Therefore, by averaging each dimension of the first initial prompt feature, the above-mentioned average feature is obtained.
[0101] During the process of adjusting the first initial prompt feature, in addition to the first preset condition, this embodiment adds a constraint condition, namely the preset orthogonality constraint condition. During the adjustment process, the average feature point multiplications between the first initial prompt features are all 0. Through this preset orthogonality constraint condition, each first initial prompt feature can focus on different parts of the object in the image, and together they can completely cover the target object in the image, thereby further improving the accuracy of zero-shot semantic segmentation.
[0102] S105. Generate a zero-shot semantic segmentation model based on the first target prompt feature.
[0103] Specifically, for each first target prompt feature, concatenate the first target prompt feature with the category features of each preset category to obtain a segmentation feature, and use the segmentation feature as the parameter of a preset convolution kernel. These preset convolution kernels with the segmentation feature as the parameter form a target classifier; and extract the feature extraction network in the semantic segmentation model that can currently only perform semantic segmentation on labeled categories; combine the feature extraction network and the above-mentioned target classifier to form an initial semantic segmentation model, and train the initial semantic segmentation model with the first training sample to obtain the corresponding zero-shot semantic segmentation model.
[0104] It should be noted that during the training process of the initial semantic segmentation model, the target classifier remains unchanged, and only the parameters of the feature extraction network are adjusted.
[0105] In some alternative embodiments, as Figure 4 shown, step S105 may also be implemented through the following steps:
[0106] S401. Obtain a sample image with an unlabeled or partially labeled ground truth segmentation mask.
[0107] In this embodiment, sample images without labeled true segmentation masks or sample images with partially labeled true segmentation masks can be obtained from existing publicly available datasets.
[0108] It can be understood that the above-mentioned sample images with partially labeled true segmentation masks refer to those in which some pixel points are labeled with object categories. For example, a sample image contains a cat and a dog. The pixel points corresponding to the cat in the sample image are labeled with object categories, while the other pixel points except the cat are not labeled with object categories. Then, this sample image is a sample image with partially labeled true segmentation masks.
[0109] S402. Based on the first target prompt feature and the vision-language model, obtain the pseudo-segmentation mask of the above-mentioned sample images without labels or with partially labeled true segmentation masks.
[0110] For the convenience of description, in this embodiment, the above-mentioned sample images without labels or with partially labeled true segmentation masks are used as unlabeled sample images.
[0111] In this embodiment, referring to the above step S102, the feature image of the unlabeled sample image is obtained through the vision-language model. Then, referring to the above step S103, based on each first target prompt feature and the category features of each preset category, the feature image of the unlabeled sample image is subjected to convolution processing to obtain the predicted segmentation mask of the unlabeled sample image, and this predicted segmentation mask of the unlabeled sample image is used as the pseudo-segmentation mask.
[0112] It can be understood that the above step S402 can refer to the solution provided in the above embodiment and will not be elaborated here.
[0113] S403. Set the sample images without labels or with partially labeled true segmentation masks and having pseudo-segmentation masks as the second training samples.
[0114] S404. Generate a zero-shot semantic segmentation model based on the first training samples, the second training samples, and the first target prompt feature.
[0115] Furthermore, as Figure 5 shown, the above step S404 can be implemented at least through the following steps:
[0116] S501. Generate a target classifier based on the first target prompt feature and various category features.
[0117] S502. Obtain the feature extraction network of the pre-trained semantic segmentation model.
[0118] Among them, the pre-trained semantic segmentation model is a semantic segmentation model that only performs semantic segmentation on labeled categories.
[0119] It can be understood that the above step S501 can be executed first and then step S502, or step S502 can be executed first and then step S501, or step S501 and step S502 can be executed simultaneously, which is not specifically limited in this embodiment.
[0120] S503. Combine the above target classifier and the above feature extraction network to form an initial semantic segmentation model.
[0121] S504. Use the above initial semantic segmentation model to perform segmentation masks on the first training sample and the second training sample respectively, and determine the predicted segmentation masks of the first training sample and the second training sample.
[0122] S505. According to the predicted segmentation masks of the first training sample and the second training sample, adjust the parameters of the feature extraction network until the feature extraction network meets the third preset condition, and obtain a semantic segmentation backbone network.
[0123] Among them, the above third preset condition refers to the above first preset condition, which will not be elaborated in this embodiment.
[0124] S506. Generate a zero-shot semantic segmentation model according to the target classifier and the semantic segmentation backbone network.
[0125] In this embodiment, by using the first training sample of visible categories, the second training sample of invisible categories, and the first target prompt feature for model training to generate a zero-shot semantic segmentation model, the performance of the generated zero-shot semantic segmentation model can be further improved.
[0126] In some alternative embodiments, as Figure 6 shown, the above step S103 can be implemented at least through the following steps:
[0127] S601. For each first initial prompt feature, perform convolutional processing on the feature image of the first training sample according to the first initial prompt feature and the category features of each preset category to obtain the first segmentation mask of the first training sample.
[0128] Please refer to steps S601 to S603 in the above embodiment for step S601 to obtain the first segmentation mask. The first segmentation mask mentioned here is the predicted segmentation mask obtained in step S603 above.
[0129] S602. Perform convolutional processing on the sample image of the first training sample with the preset learnable second initial prompt feature and the category features of each preset category to obtain the probabilities of each preset category appearing in the first training sample.
[0130] Optionally, randomly generate k learnable prompts as the first initial extraction feature p iMeanwhile, a second initial extraction feature p is randomly generated 0 This second initial prompt feature is used to characterize the probability of occurrence of each preset category in each image.
[0131] That is to say, in this embodiment, the first initial prompt feature is used to form a natural language statement with the preset category, so the first initial prompt feature can be called a segmentation prompt; the second initial prompt feature is used to characterize the probability of occurrence of each preset category in each image, so the first initial prompt feature can be called a classification prompt.
[0132] In this embodiment, the second initial prompt feature can be concatenated with the category features of each preset category to obtain a plurality of classification features; based on each classification feature, convolution processing is performed on the feature image of the first training sample to obtain the probability of occurrence of each preset category in the first training sample.
[0133] It can be understood that the above-mentioned concatenation of the second initial prompt feature with each preset category can refer to the above-mentioned embodiment of concatenating each first initial prompt feature with the category features of each preset category, and will not be elaborated here.
[0134] In addition, referring to "performing convolution processing on the feature image of the first training sample based on each segmentation feature" in the above embodiment, first select a preset convolution kernel for each segmentation feature from a plurality of preset convolution kernels, and then use the segmentation feature as the parameter of its corresponding preset convolution kernel, and perform convolution calculations on the feature map of the first training sample through the preset convolution kernel with the segmentation feature as the parameter respectively, to obtain the probability of occurrence of each preset category in the first training sample.
[0135] S603. Based on the first segmentation mask and the probability of occurrence of each preset category, obtain the predicted segmentation mask of the first training sample.
[0136] Specifically, take the probability of occurrence of each preset category as the weight, and multiply the weight by the category probability of each pixel point in the first segmentation mask to obtain the predicted segmentation mask of the first training sample.
[0137] In this embodiment, in addition to randomly generating k first initial prompt features, 1 second initial prompt feature is also generated to obtain the probability of occurrence of each preset category in each image, and taking the probability of occurrence of each preset category obtained based on the second initial prompt feature as the weight to adjust the first segmentation mask obtained only through the first initial prompt feature can effectively reduce the noise of the categories not appearing in the image, and further improve the semantic segmentation performance of the generated zero-shot semantic segmentation model.
[0138] In some alternative embodiments, the model generation method may further include: calculating a segmentation loss function and a classification loss function respectively according to the ground truth segmentation mask and the predicted segmentation mask of the first training sample; and adjusting the second initial prompt feature based on the segmentation loss function and the classification loss function to obtain a second target prompt feature.
[0139] In this embodiment, the corresponding segmentation loss function and classification loss function are calculated through the ground truth segmentation mask and the predicted segmentation mask of the first training sample, and then the parameters of the second initial prompt feature are continuously adjusted based on the classification loss function and the segmentation loss function to obtain the adjusted second initial prompt feature. Then, based on the adjusted second initial prompt feature, the ground truth segmentation mask and the predicted segmentation mask of the next first training sample are obtained, and the corresponding segmentation loss function and classification loss function are calculated, and it is determined whether they meet the second preset condition. This process is repeated until the second preset condition is met, and the second initial prompt feature that meets the second preset condition is used as the second target prompt feature.
[0140] The above-mentioned second preset condition may be to meet a preset loss function value and / or reach the number of training times, which will not be elaborated here.
[0141] In some alternative embodiments, in the case where the model generation method includes both the first initial prompt feature and the second initial prompt feature, after obtaining the second target prompt feature, the second target prompt feature can also be concatenated with the category features of each preset category with reference to the first target prompt feature to obtain the corresponding classification feature, and this classification feature is used as a part of the above-mentioned target classifier, so that the zero-shot semantic segmentation model generated based on the first target prompt feature and the second target prompt feature can effectively reduce the noise of the categories not appearing in the image during semantic segmentation, and further improve the accuracy of the zero-shot semantic segmentation model.
[0142] It can be understood that in the model generation method provided in this embodiment, the image encoder and the text encoder of the vision-language model are both fixed, that is, their parameters and structures do not need to be adjusted.
[0143] See Figure 7 As shown, taking the vision-language model as the CLIP model as an example, the model generation method of this embodiment is briefly described as follows: As Figure 7 shown, the model generation method is divided into a first stage and a second stage.
[0144] First stage: For a randomly generated learnable second initial prompt feature p 0 and three first initial prompt features p 1 , p 2 , p 3, the first initial prompt feature is concatenated with the category features of each preset category to obtain corresponding segmentation features The second initial prompt feature is concatenated with the category features of each preset category to obtain corresponding classification features Subsequently, each classification feature and segmentation feature are respectively input into the text encoder of the CLIP model for feature extraction to obtain the segmentation features after feature extraction And classification features Meanwhile, the first training sample A is input into the image encoder of the CLIP model to obtain the feature image of the first training sample A; subsequently, based on the segmentation features after feature extraction And classification features An initial classifier is constructed; then, the feature image of the first training sample A is input into this initial classifier to obtain the segmentation masks SM1, SM2, SM3 of each first initial prompt feature, and the probability P of the first training sample A belonging to each preset category; the probability P is normalized by the normalization function softmax, and then the normalized probability P and the probability P before normalization are weighted and summed to obtain the adjusted probability P of the first training sample A belonging to each preset category ′ ; meanwhile, And Are added element by element (i.e., added according to pixel points) to obtain the first segmentation mask SM of the first training sample A; the first segmentation mask SM of the first training sample A is multiplied by the probability P of the first training sample A belonging to each preset category element by element ′ To adjust the first segmentation mask SM to obtain the predicted segmentation mask SM ′ ; then, through the predicted segmentation mask SM ′ Of the first training sample A and the true segmentation mask, the first initial prompt feature and the second initial prompt feature are adjusted to obtain the first target prompt feature and the second target prompt feature.
[0145] Second stage: Based on the target classifier generated by the first target prompt feature and the second target prompt feature and the above CLIP model, the predicted segmentation mask of the unlabeled or partially labeled sample image is obtained as the pseudo segmentation mask, and the sample image labeled with the pseudo segmentation mask is used as the second training sample; the parameter of the feature extraction network is adjusted by the second training sample and the first training sample to obtain the segmentation backbone network, and a zero-shot semantic segmentation model is composed of the segmentation backbone network and the target classifier generated by the first target prompt feature and the second target prompt feature.
[0146] The model generation method provided by the embodiments of the present application obtains the feature image of the first training sample through a pre-trained vision-language model, and then performs convolutional processing on the feature image of the first training sample through a plurality of first initial prompt features and category features of each preset category that can form natural language statements when combined with the preset category to obtain the predicted segmentation mask of the first training sample. Then, the first initial prompt feature is adjusted through the predicted segmentation mask and the true segmentation mask of the first training sample to obtain the first target prompt feature, and a zero-shot semantic segmentation model is generated based on the first target prompt feature. Based on the above embodiments, in the first stage, a plurality of first target prompt features are trained using a vision-language model with zero-shot capabilities, and then in the second stage, a zero-shot semantic segmentation model is generated based on the first target prompt information. On the one hand, the vision-language model is applied to the semantic segmentation task to achieve zero-shot semantic segmentation; on the other hand, compared with directly using manually designed prompts in the CLIP model for pixel-level segmentation tasks, the zero-shot semantic segmentation model provided by this embodiment does not only focus on the most different regions in the image, but can make the generated segmentation mask cover the entire target.
[0147] The embodiments of the present application also provide an image processing method. Figure 8 It is a schematic flowchart of the image processing method provided by this embodiment, as Figure 8 shown. The image processing method at least includes the following steps:
[0148] S801, obtain the image to be processed.
[0149] S802, perform semantic segmentation on the image to be processed based on the zero-shot semantic segmentation model to obtain the predicted segmentation mask of the image to be processed.
[0150] The above zero-shot semantic segmentation model is obtained through the model generation method described in any of the above embodiments.
[0151] Correspondingly, the embodiments of the present application also provide an electronic device. The electronic device can be a terminal, and the terminal can be a smart phone, a tablet computer, a laptop computer, a touch screen, a game console, a personal computer (PC, Personal Computer), a personal digital assistant (Personal Digital Assistant, PDA), and other terminal devices. Alternatively, the electronic device can be a server.
[0152] Please refer to Figure 9 , Figure 9Schematic diagram of the structure of the electronic device provided in the embodiment of the present application. The electronic device 900 includes a processor 910 with one or more processing cores, a memory 920 with one or more computer-readable storage media, and a computer program stored on the memory 920 and executable on the processor 910. Among them, the processor 910 is electrically connected to the memory 920. Those skilled in the art can understand that the structure of the electronic device shown in the figure does not constitute a limitation on the electronic device, and it may include more or fewer components than shown, or combine some components, or have different component arrangements.
[0153] The processor 910 is the control center of the electronic device 900, connecting various parts of the entire electronic device 900 through various interfaces and lines. By running or loading software programs and / or units stored in the memory 920, and calling data stored in the memory 920, it executes various functions of the electronic device 900 and processes data, thereby monitoring the entire electronic device 900. The processor 910 can be a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), etc., and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present disclosure.
[0154] In the embodiment of the present application, the processor 910 in the electronic device 900 will, according to the following steps, load the instructions corresponding to one or more application programs into the memory 920, and the processor 910 will run the application programs stored in the memory 920 to implement various functions, such as:
[0155] Obtain a first training sample, where the first training sample is a sample image with a labeled true segmentation mask;
[0156] Based on a pre-trained vision-language model, obtain a feature image of the first training sample;
[0157] Based on a plurality of preset learnable first initial prompt features and the category features of each preset category, perform convolution processing on the feature image of the first training sample to obtain a predicted segmentation mask of the first training sample; the first initial prompt features are used to form natural language sentences in combination with preset categories;
[0158] Based on the predicted segmentation mask and the true segmentation mask of the first training sample, adjust the first initial prompt features to obtain first target prompt features;
[0159] Based on the first target prompt features, generate a zero-shot semantic segmentation model.
[0160] In an optional example, based on a plurality of preset learnable first initial prompt features and the category features of each preset category, performing convolution processing on the feature image of the first training sample to obtain the predicted segmentation mask of the first training sample, including:
[0161] For each first initial prompt feature, performing convolution processing on the feature image of the first training sample according to the first initial prompt feature and the category features of each preset category to obtain the first segmentation mask of the first training sample;
[0162] Based on a preset learnable second initial prompt feature and each category feature, performing convolution processing on the feature image of the first training sample to obtain the probability of each preset category appearing in the first training sample; the second initial prompt feature is used to represent the probability of each preset category appearing in each image;
[0163] Based on the first segmentation mask and the probability of each preset category appearing, obtain the predicted segmentation mask of the first training sample.
[0164] In an optional example, it further includes:
[0165] According to the true segmentation mask and the predicted segmentation mask of the first training sample, calculate the segmentation loss function and the classification loss function respectively;
[0166] Based on the segmentation loss function and the classification loss function, adjust the second initial prompt feature to obtain the second target prompt feature.
[0167] In an optional example, based on a preset learnable second initial prompt feature and each category feature, performing convolution processing on the feature image of the first training sample to obtain the probability of each preset category appearing in the first training sample, including:
[0168] Concatenate the second initial prompt feature with the category features of each preset category respectively to obtain a plurality of classification features;
[0169] Based on each classification feature, perform convolution processing on the feature image of the first training sample to obtain the probability of each preset category appearing in the first training sample.
[0170] In an optional example, based on a plurality of preset learnable first initial prompt features and the category features of each preset category, performing convolution processing on the feature image of the first training sample to obtain the predicted segmentation mask of the first training sample, including:
[0171] Concatenate each first initial prompt feature with the category features of each preset category to obtain a plurality of segmentation features;
[0172] Based on each segmentation feature, perform convolution processing on the feature image of the first training sample to obtain a plurality of second segmentation masks;
[0173] Add the second segmentation masks corresponding to each first initial prompt feature pixel by pixel to obtain the predicted segmentation mask of the first training sample.
[0174] In an optional example, based on the predicted segmentation mask and the ground truth segmentation mask of the first training sample, adjust the first initial prompt feature until the first initial prompt feature meets the preset conditions to obtain the first target prompt feature, including:
[0175] Based on the predicted segmentation mask and the ground truth segmentation mask of the first training sample, adjust the first initial prompt feature to obtain the adjusted first initial prompt feature;
[0176] If the adjusted first initial prompt feature meets the first preset condition and the preset orthogonality constraint condition, set the adjusted first initial prompt feature as the first target prompt feature;
[0177] Among them, the preset orthogonality constraint condition is that the average feature point product of each first initial prompt feature is 0.
[0178] In an optional example, based on the first target prompt feature, generate a zero-shot semantic segmentation model, including:
[0179] Obtain a sample image with an unannotated or partially annotated ground truth segmentation mask;
[0180] Based on the first target prompt feature and the vision-language model, obtain the pseudo-segmentation mask of the sample image with an unannotated or partially annotated ground truth segmentation mask;
[0181] Set the sample image with the pseudo-segmentation mask as the second training sample;
[0182] Based on the first training sample, the second training sample, and the first target prompt feature, generate a zero-shot semantic segmentation model.
[0183] In an optional example, based on the first training sample, the second training sample, and the first target prompt feature, generate a zero-shot semantic segmentation model for image semantic segmentation, including:
[0184] Based on the first target prompt feature and the category features of each preset category, generate a target classifier; and obtain the feature extraction network of the pre-trained semantic segmentation model;
[0185] Combine the target classifier and the feature extraction network to form an initial semantic segmentation model;
[0186] According to the initial semantic segmentation model, perform segmentation masks on the first training sample and the second training sample respectively to determine the predicted segmentation masks of the first training sample and the second training sample;
[0187] Based on the predicted segmentation masks of the first training sample and the second training sample, the parameters of the feature extraction network are adjusted to obtain a semantic segmentation backbone network;
[0188] According to the target classifier and the semantic segmentation backbone network, a zero-shot semantic segmentation model is constructed.
[0189] In an optional example, before obtaining the feature image of the first training sample based on the image encoder in the pre-trained vision-language model, the method further includes:
[0190] Obtain the feature dimension output by the image encoder;
[0191] If the feature dimension is less than a preset threshold, the attention pooling layer in the image encoder is deleted.
[0192] For another example:
[0193] Obtain the image to be processed;
[0194] Based on the zero-shot semantic segmentation model, perform semantic segmentation on the image to be processed to obtain the predicted segmentation mask of the image to be processed;
[0195] Wherein, the zero-shot semantic segmentation model is obtained by the model generation method described in any one of the above.
[0196] The specific implementation manners of the above operations can be referred to the previous embodiments and will not be elaborated herein.
[0197] Optionally, as Figure 9 shown, the electronic device 900 further includes: a touch display screen 930, a radio frequency circuit 940, an audio circuit 950, an input unit 960, and a power supply 970. Among them, the processor 910 is electrically connected to the touch display screen 930, the radio frequency circuit 940, the audio circuit 950, the input unit 960, and the power supply 970 respectively. Those skilled in the art can understand that Figure 9 the structure of the electronic device shown in
[0198] The touch display screen 930 can be used to display a graphical user interface and receive operation instructions generated by a user's interaction with the graphical user interface. The touch display screen 930 may include a display panel and a touch panel. Among them, the display panel can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the electronic device. These graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. Optionally, the display panel can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. The touch panel can be used to collect touch operations of the user on or near it (such as operations of the user using any suitable object or accessory such as a finger or a stylus on or near the touch panel), and generate corresponding operation instructions, and the operation instructions execute the corresponding program. Optionally, the touch panel can include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user and detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 910, and can receive and execute commands sent by the processor 910. The touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it is transmitted to the processor 910 to determine the type of touch event. Subsequently, the processor 910 provides a corresponding visual output on the display panel according to the type of touch event. In the embodiments of the present disclosure, the touch panel and the display panel can be integrated into the touch display screen 930 to implement input and output functions. However, in some embodiments, the touch panel and the touch panel can be implemented as two independent components to implement input and output functions. That is, the touch display screen 930 can also be used as a part of the input unit 960 to implement the input function.
[0199] The radio frequency circuit 940 can be used to transmit and receive radio frequency signals to establish wireless communication with a network device or other electronic devices through wireless communication, and transmit and receive signals with the network device or other electronic devices.
[0200] The audio circuit 950 can be used to provide an audio interface between the user and the electronic device through a speaker and a microphone. The audio circuit 950 can transmit the electrical signal converted from the received audio data to the speaker, and the speaker converts it into a sound signal for output; on the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit 950 and then converted into audio data. After the audio data is output to the processor 910 for processing, it is sent through the radio frequency circuit 940 to another electronic device, for example, or the audio data is output to the memory 920 for further processing. The audio circuit 950 may also include an earphone jack to provide communication between an external earphone and the electronic device.
[0201] The input unit 960 can be used to receive input digital, character information or user feature information (such as fingerprints, irises, facial information, etc.), and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
[0202] The power supply 970 is used to supply power to each component of the electronic device 900. Optionally, the power supply 970 can be logically connected to the processor 910 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 970 can also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0203] Although Figure 9 not shown in [description], the electronic device 900 may further include a camera, a sensor, a Wi-Fi module, a Bluetooth module, etc., which will not be elaborated here.
[0204] In the above embodiments, the descriptions of each embodiment have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0205] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions or by controlling relevant hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0206] For this reason, an embodiment of the present application provides a computer-readable storage medium, in which multiple computer programs are stored. The computer programs can be loaded by a processor to execute any model generation method and image processing method provided by the embodiments of the present disclosure. The computer programs can execute the steps of the following model generation method:
[0207] Obtain a first training sample, where the first training sample is a sample image labeled with a true segmentation mask;
[0208] Based on a pre-trained vision-language model, obtain a feature image of the first training sample;
[0209] Based on a plurality of preset learnable first initial prompt features and the category features of each preset category, perform convolution processing on the feature image of the first training sample to obtain a predicted segmentation mask of the first training sample; the first initial prompt features are used to form natural language statements in combination with preset categories;
[0210] Based on the predicted segmentation mask and the true segmentation mask of the first training sample, adjust the first initial prompt features to obtain first target prompt features;
[0211] Generate a zero-shot semantic segmentation model based on the first target prompt feature.
[0212] In an optional example, perform convolution processing on the feature image of the first training sample based on a plurality of preset learnable first initial prompt features and the category features of each preset category to obtain the predicted segmentation mask of the first training sample, including:
[0213] For each first initial prompt feature, perform convolution processing on the feature image of the first training sample according to the first initial prompt feature and the category features of each preset category to obtain the first segmentation mask of the first training sample;
[0214] Perform convolution processing on the feature image of the first training sample based on a preset learnable second initial prompt feature and each category feature to obtain the probability of each preset category appearing in the first training sample; the second initial prompt feature is used to represent the probability of each preset category appearing in each image;
[0215] Based on the first segmentation mask and the probability of each preset category appearing, obtain the predicted segmentation mask of the first training sample.
[0216] In an optional example, it further includes:
[0217] Calculate the segmentation loss function and the classification loss function respectively according to the true segmentation mask and the predicted segmentation mask of the first training sample;
[0218] Based on the segmentation loss function and the classification loss function, adjust the second initial prompt feature to obtain the second target prompt feature.
[0219] In an optional example, perform convolution processing on the feature image of the first training sample based on a preset learnable second initial prompt feature and each category feature to obtain the probability of each preset category appearing in the first training sample, including:
[0220] Concatenate the second initial prompt feature with the category features of each preset category respectively to obtain a plurality of classification features;
[0221] Perform convolution processing on the feature image of the first training sample based on each classification feature to obtain the probability of each preset category appearing in the first training sample.
[0222] In an optional example, perform convolution processing on the feature image based on a plurality of preset learnable first initial prompt features and the category features of each preset category to obtain the predicted segmentation mask of the first training sample, including:
[0223] Concatenate each first initial prompt feature with the category features of each preset category to obtain a plurality of segmentation features;
[0224] Perform convolution processing on the feature image of the first training sample based on each segmentation feature to obtain multiple second segmentation masks;
[0225] Add the second segmentation masks corresponding to each first initial prompt feature according to pixel points to obtain the predicted segmentation mask of the first training sample.
[0226] In an optional example, based on the predicted segmentation mask and the true segmentation mask of the first training sample, adjust the first initial prompt feature until the first initial prompt feature meets the preset conditions to obtain the first target prompt feature, including:
[0227] Based on the predicted segmentation mask and the true segmentation mask of the first training sample, adjust the first initial prompt feature to obtain the adjusted first initial prompt feature;
[0228] If the adjusted first initial prompt feature meets the first preset condition and the preset orthogonality constraint condition, set the adjusted first initial prompt feature as the first target prompt feature;
[0229] Among them, the preset orthogonality constraint condition is that the average feature point product of each first initial prompt feature is 0.
[0230] In an optional example, based on the first target prompt feature, generate a zero-shot semantic segmentation model, including:
[0231] Obtain a sample image with an unlabeled or partially labeled true segmentation mask;
[0232] Based on the first target prompt feature and the vision-language model, obtain the pseudo-segmentation mask of the sample image with an unlabeled or partially labeled true segmentation mask;
[0233] Set the sample image with an unlabeled or partially labeled true segmentation mask and a pseudo-segmentation mask as the second training sample;
[0234] Based on the first training sample, the second training sample, and the first target prompt feature, generate a zero-shot semantic segmentation model.
[0235] In an optional example, based on the first training sample, the second training sample, and the first target prompt feature, generate a zero-shot semantic segmentation model, including:
[0236] Based on the first target prompt feature and the category features of each preset category, generate a target classifier; and obtain the feature extraction network of the pre-trained semantic segmentation model;
[0237] Form an initial semantic segmentation model with the target classifier and the feature extraction network;
[0238] Segment the first training sample and the second training sample respectively according to the initial semantic segmentation model to determine the predicted segmentation masks of the first training sample and the second training sample;
[0239] Based on the predicted segmentation masks of the first training sample and the second training sample, adjust the parameters of the feature extraction network to obtain the semantic segmentation backbone network;
[0240] Construct a zero-shot semantic segmentation model according to the target classifier and the semantic segmentation backbone network.
[0241] In an optional example, before obtaining the feature image of the first training sample based on the image encoder in the pre-trained vision-language model, the method further includes:
[0242] Obtain the feature dimension output by the image encoder;
[0243] If the feature dimension is less than the preset threshold, delete the attention pooling layer in the image encoder.
[0244] The computer program can also execute the steps of the following image processing method:
[0245] Obtain the image to be processed;
[0246] Perform semantic segmentation on the image to be processed based on the zero-shot semantic segmentation model to obtain the predicted segmentation mask of the image to be processed;
[0247] Among them, the zero-shot semantic segmentation model is obtained by the model generation method described in any one of the above.
[0248] For the specific implementation of each of the above operations, reference may be made to the previous embodiments, which will not be elaborated here.
[0249] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.
[0250] Since the computer program stored in the computer-readable storage medium can execute any one of the model generation methods or image processing methods provided by the embodiments of the present disclosure, the beneficial effects that can be achieved by any one of the model generation methods or image processing methods provided by the embodiments of the present disclosure can be realized. For details, reference may be made to the previous embodiments, which will not be elaborated here.
[0251] According to an aspect of the present disclosure, there is also provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the methods provided in the various alternative implementations in the above embodiments.
[0252] In the above embodiments of the image processing method, the electronic device, the computer-readable storage medium, and the computer program product, the descriptions of the respective embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes and the beneficial effects brought about by the above-described game interaction device, computer-readable storage medium, computer program product, electronic device, and their corresponding units can refer to the description of the game interaction method in the above embodiments, and will not be elaborated herein specifically.
[0253] The above has introduced in detail a model generation method, an image processing method, an electronic device, a computer-readable storage medium, and a computer program product provided by the embodiments of the present disclosure. Specific examples are used in this article to elaborate on the principles and implementation manners of the present disclosure. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present disclosure; at the same time, for those skilled in the art, according to the idea of the present disclosure, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present disclosure.
Claims
1. A model generation method, characterized in that, the method includes: Obtain a first training sample, where the first training sample is a sample image with an annotated true segmentation mask; Based on a pre-trained vision-language model, obtain a feature image of the first training sample; Based on a plurality of preset learnable first initial prompt features and the category features of each preset category, perform convolution processing on the feature image of the first training sample to obtain a predicted segmentation mask of the first training sample; the first initial prompt features are used to form natural language sentences in combination with preset categories; Based on the predicted segmentation mask and the true segmentation mask of the first training sample, adjust the first initial prompt features to obtain first target prompt features; Based on the first target prompt features, generate a zero-shot semantic segmentation model; the zero-shot semantic segmentation model includes a target classifier generated based on the first target prompt features and category features of each class; The step of performing convolution processing on the feature image of the first training sample based on a plurality of preset learnable first initial prompt features and the category features of each preset category to obtain a predicted segmentation mask of the first training sample includes: Concatenate each of the first initial prompt features with the category features of each of the preset categories to obtain a plurality of segmentation features; For each segmentation feature, use the segmentation feature as the parameter of the corresponding preset convolution kernel; Input the feature image of the first training sample into each preset convolution kernel with a segmentation feature as a parameter for convolution calculation.
2. The model generation method according to claim 1, characterized in that, the step of performing convolution processing on the feature image of the first training sample based on a plurality of preset learnable first initial prompt features and the category features of each preset category to obtain a predicted segmentation mask of the first training sample includes: For each of the first initial prompt features, perform convolution processing on the feature image of the first training sample according to the first initial prompt feature and the category features of each preset category to obtain a first segmentation mask of the first training sample; Based on a preset learnable second initial prompt feature and the category features of each of the categories, perform convolution processing on the feature image of the first training sample to obtain the probability of each of the preset categories appearing in the first training sample; the second initial prompt feature is used to represent the probability of each of the preset categories appearing in each image; Based on the first segmentation mask and the probability of each of the preset categories appearing, obtain the predicted segmentation mask of the first training sample.
3. The model generation method according to claim 2, characterized in that, the method further includes: Calculate a segmentation loss function and a classification loss function respectively according to the true segmentation mask and the predicted segmentation mask of the first training sample; Based on the segmentation loss function and the classification loss function, adjust the second initial prompt features to obtain second target prompt features.
4. The model generation method according to claim 2, characterized in that, Performing convolution processing on the feature image of the first training sample based on the preset learnable second initial prompt feature and each of the category features to obtain the probabilities of each of the preset categories appearing in the first training sample includes: Concatenating the second initial prompt feature with the category features of each of the preset categories respectively to obtain a plurality of classification features; Performing convolution processing on the feature image of the first training sample based on each of the classification features respectively to obtain the probabilities of each of the preset categories appearing in the first training sample.
5. The model generation method according to claim 1, wherein, Adjusting the first initial prompt feature based on the predicted segmentation mask and the true segmentation mask of the first training sample to obtain a first target prompt feature includes: Adjusting the first initial prompt feature based on the predicted segmentation mask and the true segmentation mask of the first training sample to obtain the adjusted first initial prompt feature; If the adjusted first initial prompt feature satisfies a first preset condition and a preset orthogonality constraint condition, setting the adjusted first initial prompt feature as the first target prompt feature; wherein, the preset orthogonality constraint condition is that the average feature point product of each of the first initial prompt features is 0.
6. The model generation method according to claim 1, wherein, Generating a zero-shot semantic segmentation model based on the first target prompt feature includes: Obtaining a sample image with an unlabeled or partially labeled true segmentation mask; Obtaining a pseudo-segmentation mask of the sample image with the unlabeled or partially labeled true segmentation mask based on the first target prompt feature and the vision-language model; Setting the sample image with the pseudo-segmentation mask as a second training sample; Generating a zero-shot semantic segmentation model based on the first training sample, the second training sample, and the first target prompt feature.
7. The model generation method according to claim 6, wherein, Generating a zero-shot semantic segmentation model based on the first training sample, the second training sample, and the first target prompt feature includes: Generating a target classifier based on the first target prompt feature and the category features of each of the preset categories; and obtaining a feature extraction network of a pre-trained semantic segmentation model; Forming an initial semantic segmentation model with the target classifier and the feature extraction network; Determining the predicted segmentation masks of the first training sample and the second training sample according to the initial semantic segmentation model respectively performing segmentation masks on the first training sample and the second training sample; Adjusting the parameters of the feature extraction network based on the predicted segmentation masks of the first training sample and the second training sample to obtain a semantic segmentation backbone network; Constructing the zero-shot semantic segmentation model according to the target classifier and the semantic segmentation backbone network.
8. The model generation method according to claim 1, wherein, Before obtaining the feature image of the first training sample by the image encoder in the pre-trained vision-language model, the method further includes: Obtaining the feature dimension output by the image encoder; If the feature dimension is less than a preset threshold, deleting the attention pooling layer in the image encoder.
9. An image processing method, Characterized in that, The image processing method includes: Obtaining an image to be processed; Performing semantic segmentation on the image to be processed based on a zero-shot semantic segmentation model to obtain a predicted segmentation mask of the image to be processed; Wherein, the zero-shot semantic segmentation model is obtained by the model generation method according to any one of claims 1 to 8.
10. An electronic device, Characterized in that, Comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, it implements the model generation method according to any one of claims 1 to 8 or implements the image processing method according to claim 9.
11. A readable storage medium, Characterized in that, The readable storage medium includes a computer program, and when the computer program runs, it controls the electronic device where the readable storage medium is located to execute the model generation method according to any one of claims 1 to 8 or the image processing method according to claim 9.
Citation Information
Patent Citations
Image processing method and terminal equipment
CN115294150A
Image segmentation and model training method, device and equipment
CN115631205A
Image segmentation method, and method and device for training image segmentation model
CN116433899A