Image segmentation method, device, electronic device and computer-readable storage medium

By building a knowledge base and using target prompt features for image segmentation, the problems of high difficulty and low segmentation efficiency of remote sensing images are solved, and synchronous segmentation and efficient segmentation of multiple geographic categories are achieved.

CN119516368BActive Publication Date: 2025-07-29BEIJING DATA INTELLIGENCE INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411564410.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-05
Publication Date
2025-07-29
Estimated Expiration
2044-11-05

AI Technical Summary

Technical Problem

Remote sensing image segmentation is difficult, the existing segmentation model is insufficient in generalization ability, and it is impossible to synchronously segment multiple geographic categories, and the segmentation efficiency is low.

Method used

By building a knowledge base, obtaining target prompt features and performing feature enhancement processing, and using image features and enhanced target prompt features for mask decoding, so as to achieve synchronous segmentation of multiple geographic categories of the original image.

Benefits of technology

It realizes accurate segmentation of images of various landform types, improves segmentation efficiency and accuracy, and enhances the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119516368B_ABST
    Figure CN119516368B_ABST
Patent Text Reader

Abstract

The present invention provides an image segmentation method, apparatus, electronic device, and computer-readable storage medium, including: extracting features from an original image to obtain image features and an image class identifier; using the image class identifier to match a target prompt feature corresponding to the original image from a knowledge base, where the knowledge base is pre-constructed and contains multiple pairs of positive prompt features and negative prompt features; using the image features to perform feature enhancement processing on the target prompt feature to obtain an enhanced target prompt feature; performing mask decoding processing on the image features and the enhanced prompt features to obtain a segmentation mask for the original image. The present invention guides the segmentation of the original image by obtaining the target prompt feature corresponding to the original image from the knowledge base, achieving the effect of accurately segmenting the original image in practical applications. Moreover, by using the target prompt features corresponding to multiple ground object classes in the original image, synchronous segmentation of multiple ground object classes of the original image is realized, improving the segmentation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of image processing and artificial intelligence, and in particular, to an image segmentation method, apparatus, electronic device, and computer-readable storage medium. Background Art

[0002] Image segmentation refers to the technology and process of dividing an image into several specific regions with unique properties and extracting the target of interest. Remote sensing images have a large number of ground object categories, changing scenes, and large sizes, which makes remote sensing image segmentation difficult. The accurate segmentation of remote sensing images has an important impact on the subsequent processing and application of remote sensing images.

[0003] For the segmentation of remote sensing images, an image segmentation model is usually trained using a training data set. However, in actual applications, since the image data to be processed often differs greatly from the training data set, when using the segmentation model to segment the image data to be processed, some ground objects are still difficult to clearly distinguish, the segmentation effect is poor, and the generalization ability of the segmentation model is insufficient. In addition, in related technologies, the target semantics are segmented by manually inputting prompt information or setting prompt information, and multiple ground object categories cannot be segmented synchronously, resulting in low segmentation efficiency.

[0004] Therefore, it is necessary to develop an image segmentation method that can accurately and efficiently segment various images to be segmented in actual applications. Summary of the Invention

[0005] Provided is an image segmentation method, which realizes the purpose of automatically and accurately segmenting images of various ground object types by obtaining target prompt features from a knowledge base to guide the segmentation of the original image.

[0006] In a first aspect, an embodiment of the present application provides an image segmentation method, which includes:

[0007] S1: Obtain an original image and a knowledge base, where the knowledge base is pre-constructed and contains multiple pairs of positive prompt features and negative prompt features;

[0008] S2: Extract features from the original image to obtain image features and image class identifiers;

[0009] S3: Use the image class identifier to match the target prompt features corresponding to the original image from the knowledge base;

[0010] S4: Perform feature enhancement processing on the target prompt features using the image features to obtain enhanced target prompt features;

[0011] S5: Perform mask decoding processing on the image features and the enhanced target prompt features to obtain a segmentation mask for the original image.

[0012] Further, S3 includes:

[0013] Calculate the similarity between the image class identifier and the positive hint features in the knowledge base, and use the top k positive hint features sorted by similarity and their associated k negative hint features as the target hint features.

[0014] The construction process of the knowledge base includes:

[0015] Obtain knowledge images, annotate various ground object categories in the knowledge images with masks and set labels respectively to obtain labeled knowledge images;

[0016] For the labeled knowledge images, obtain sub-images within the mask region of a certain ground object category, and use a pre-trained ViT network to obtain the class identifiers of the sub-images to obtain the positive hint features;

[0017] Obtain sub-images not within the mask region of a certain ground object category, use a pre-trained ViT network to obtain the class identifiers of the sub-images to obtain the negative hint features, and store the positive hint features, the negative hint features, and the label of the certain ground object category in an associated manner.

[0018] S2 includes:

[0019] Perform multi-scale feature extraction and feature fusion on the original image to obtain a first image feature, a second image feature, and the image class identifier, and use the first image feature and the second image feature as the image features; wherein, the scale of the first image feature is larger than the scale of the second image feature.

[0020] Furthermore, the method for obtaining the first image feature, the second image feature, and the image class identifier includes:

[0021] Use a pre-trained CNN network to perform multi-scale feature extraction on the original image to obtain first sub-features, second sub-features, third sub-features, and fourth sub-features with gradually decreasing scales;

[0022] Use a pre-trained ViT network to perform feature extraction on the original image to obtain the image class identifier and a fifth sub-feature;

[0023] Use a feature fusion network to perform feature fusion on the fourth sub-feature, the third sub-feature, and the fifth sub-feature to obtain the second image feature;

[0024] Use a feature fusion network to perform feature fusion on the second image feature, the second sub-feature, and the first sub-feature to obtain the first image feature.

[0025] S4 includes:

[0026] Input the second image feature and the target prompt feature into a Transformer network, and use the cross-attention mechanism to perform feature enhancement processing on the target prompt feature to obtain the enhanced target prompt feature.

[0027] The S5 includes:

[0028] Perform matrix multiplication on the first image feature and the enhanced target prompt feature to obtain a segmentation mask for the original image, where the segmentation mask corresponds to multiple land cover classes of the original image.

[0029] In a second aspect, an embodiment of the present application provides an image segmentation device, which includes:

[0030] An input unit for obtaining an original image and a knowledge base, where the knowledge base is pre-constructed and contains multiple pairs of positive prompt features and negative prompt features;

[0031] A feature extraction unit for performing feature extraction on the original image to obtain an image feature and an image class identifier;

[0032] A prompt matching unit for using the image class identifier to match the target prompt feature corresponding to the original image from the knowledge base;

[0033] A prompt embedding unit for performing feature enhancement processing on the target prompt feature by using the image feature to obtain an enhanced target prompt feature;

[0034] A mask decoding unit for performing mask decoding processing on the image feature and the enhanced target prompt feature to obtain a segmentation mask for the original image.

[0035] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, it can execute the method of the first aspect.

[0036] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, where when the computer program is executed by a processor, it can execute the method of the first aspect.

[0037] The beneficial effects of the present invention at least include:

[0038] 1. In the image segmentation method provided by the present invention, by using the image class identifier of the original image to obtain the positive prompt feature and the negative prompt feature corresponding to the original image from the knowledge base to guide the segmentation of the original image, the effect of accurately segmenting various original images in practical applications is achieved;

[0039] 2. By utilizing the target hint features corresponding to multiple ground object categories of the original image, the synchronous segmentation of multiple ground object categories of the original image is realized, and the segmentation efficiency is improved;

[0040] 3. By using a dual feature extraction network to extract the first image feature and the second image feature, where the first image feature is the fusion of multiple scale features of the original image, and by using the first image feature for mask decoding, the multi-scale features are fully utilized, and the segmentation accuracy can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0042] Figure 1 It is a flowchart of the image segmentation method provided by the embodiment of the present application;

[0043] Figure 2 It is a flowchart of constructing a knowledge base provided by the embodiment of the present application;

[0044] Figure 3 It is a schematic structural diagram of an image segmentation model provided by the embodiment of the present application;

[0045] Figure 4 It is a schematic structural diagram of another image segmentation model provided by the embodiment of the present application;

[0046] Figure 5 It is a schematic block diagram of an image segmentation device provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present invention.

[0048] Figure 1 It is a flowchart of the image segmentation method provided by the embodiment of the present application, as Figure 1 shown, the method includes:

[0049] S1: Obtain an original image and a knowledge base, where the knowledge base is pre-constructed and contains multiple pairs of positive hint features and negative hint features;

[0050] S2: Extract features from the original image to obtain image features and an image class identifier;

[0051] S3: Use the image class identifier to match the target hint features corresponding to the original image from the knowledge base;

[0052] S4: Perform feature enhancement processing on the target hint features using the image features to obtain enhanced target hint features;

[0053] S5: Perform mask decoding processing on the image features and the enhanced target hint features to obtain a segmentation mask for the original image.

[0054] According to the above process, it can be seen that this application first obtains the image features and image class identifier (ClassToken, Cls Token) of the original image. Among them, the image class identifier represents the overall distribution of the ground object categories of the original image. Then, the corresponding target hint features are automatically matched from the knowledge base according to the image class identifier of the original image. Next, the target hint features are enhanced, and finally, mask decoding is performed according to the image features and the enhanced target hint features to obtain a segmentation mask for the original image. Through the method provided by the present invention, it is possible to use the guidance of the knowledge base to segment the original images of multiple ground object categories.

[0055] Next, each of the above steps will be described in detail in combination with embodiments. First, step S1 will be described in detail in combination with embodiments.

[0056] In this application, the original image is remote sensing image data to be segmented. The original image may include data of multiple bands. For example, it includes data of three bands: red, green, and blue, or may also include data of the near-infrared band. Optionally, the original image may also include data of other bands.

[0057] In this application, the knowledge base is pre-constructed according to knowledge images. The knowledge base contains multiple positive hint features and multiple negative hint features, and the positive hint features and negative hint features exist in pairs.

[0058] Optionally, the knowledge image may be remote sensing image data of three bands or four bands. Optionally, the knowledge image may be a remote sensing image similar to the original image in actual application. Optionally, the knowledge image may be a remote sensing image not similar to the original image in actual application.

[0059] Figure 2 For the flowchart of constructing the knowledge base provided by the embodiments of this application, as Figure 2 shown, the process of constructing the knowledge base includes:

[0060] Step a: Obtain a knowledge image, label multiple ground object categories in the knowledge image with masks and set labels respectively to obtain a labeled knowledge image.

[0061] Specifically, first obtain a knowledge image, then label all ground object categories in the knowledge image with masks, and set different labels for each ground object category to obtain a labeled knowledge image.

[0062] Optionally, the labeling can be manual labeling or automated labeling.

[0063] Optionally, some ground object categories in the knowledge image can be labeled with masks, for example, label the ground object categories in the regions with a larger occupancy ratio.

[0064] Step b: For the labeled knowledge image, obtain a sub-image within the mask region of a certain ground object category, and use a pre-trained ViT network to obtain the class label of the sub-image to obtain the positive prompt feature.

[0065] Specifically, divide the labeled knowledge image into multiple image patches. For a certain ground object category, select multiple image patches within the corresponding mask range and input them into the pre-trained ViT network. The class label output by the network is the positive prompt feature. Optionally, the ViT network can use the ViT network in the pre-trained self-distillation with no labels (DINO) model. Among them, the pre-trained DINO model is a model trained through self-supervised learning on a remote sensing image dataset, and the backbone network of the DINO model uses the ViT network. DINO is an existing model and will not be elaborated here.

[0066] Divide the labeled knowledge image into multiple image patches, select the image patches within the mask according to the annotation of a certain ground object category, and add the initial class label and send them into the pre-trained ViT network. The class label output by the network is used as the positive prompt feature of the knowledge image. Among them, the initial class label is the class label carried by the pre-trained ViT.

[0067] Step c: Obtain a sub-image that is not within the mask region of a certain ground object category, use a pre-trained ViT network to obtain the class label of the sub-image to obtain the negative prompt feature, and store the positive prompt feature, the negative prompt feature, and the label of the certain ground object category in an associated manner.

[0068] Specifically, select the image patches that are not within the mask according to the annotation of a certain ground object category, and add the initial class label and send them into the pre-trained ViT network. The class label output by the network is used as the negative prompt feature of the knowledge image.

[0069] According to step b-step c, corresponding positive hint features and negative hint features can be made based on other ground object categories in the same knowledge image.

[0070] Other knowledge images can be processed in the same way as step a-step d, continuously increasing positive hint features and negative hint features until the preset conditions are met to obtain the said knowledge base. The finally obtained knowledge base includes multiple positive hint features and multiple negative hint features corresponding to m ground object categories. Correspondingly, there are m types of labels in the knowledge base, and each pair of positive hint features and negative hint features is stored associatively with one of the labels.

[0071] Among them, the preset conditions can be set according to actual needs, including but not limited to that the types of ground object categories corresponding to the positive hint features and negative hint features in the knowledge base can meet the segmentation requirements of the current original image, or the number of positive hint features and negative hint features in the knowledge base meets the usage requirements. In addition, in actual use, the knowledge base can be updated in real time.

[0072] Next, the above steps S2-S5 will be described in detail in combination with embodiments. Each step implemented in this process can be realized by using an image segmentation model.

[0073] Based on an embodiment of the present application, the structure of the image segmentation model can be referred to Figure 3 . As Figure 3 shown, the image segmentation model includes a feature extraction module, a hint retrieval module, a hint embedding module, and a mask decoding module.

[0074] In step S2, the feature extraction module is used to extract features from the said original image to obtain image features and image class identifiers. The feature extraction module can be, for example, ViT in the pre-trained DINO model, which has been introduced above and will not be elaborated here. In step S3, the hint retrieval module is used to match the target hint features corresponding to the original image from the knowledge base according to the image class identifier. The hint retrieval module can be, for example, to obtain the target hint features by calculating the similarity between the image class identifier and the positive hint features in the knowledge base. Correspondingly, in step S4, the hint embedding module uses the image features to perform feature enhancement processing on the target hint features to obtain enhanced target hint features. The hint embedding module can be, for example, to use a cross-attention network. In step S5, the mask decoding module performs mask decoding processing on the image features and the enhanced hint features to obtain a segmentation mask for the said original image. The mask decoding module performs matrix multiplication on the image features and the enhanced target hint features in the feature dimension to obtain an image segmentation mask, and this segmentation mask includes the position information of multiple ground object categories in the original image.

[0075] In the image segmentation method provided in this embodiment, the image segmentation model guides the segmentation of the original image by obtaining target prompt features corresponding to the original image from the knowledge base. For original images that were difficult to segment in the past, accurate segmentation can also be achieved with the help of the prompt information in the knowledge base, realizing the effect that the original images of various ground object types can be accurately segmented, and the image segmentation model has strong generalization ability.

[0076] To further improve the segmentation accuracy, multi-scale features of the original image can be extracted, and image features can be obtained based on the multi-scale features. Based on an embodiment of the present application, the structure of the image segmentation model can be referred to Figure 4 .

[0077] As Figure 4 shown, in step S2, multi-scale feature extraction and feature fusion are performed on the original image to obtain the first image feature, the second image feature, and the image class identifier, and the first image feature and the second image feature are used as the image features, where the scale of the first image feature is larger than the scale of the second image feature.

[0078] Optionally, the feature extraction module can use a dual feature extraction network. The dual feature extraction network includes a CNN network, a ViT network, and a feature fusion network. The CNN network can use a pre-trained CNN network, such as the ResNeSt50d network pre-trained using ImageNet. ResNeSt50d is a type of ResNeSt series model, and its attention module adopts a decentralized attention mechanism, aiming to improve the representation ability and generalization ability of the model. The ViT network can use the ViT in the pre-trained DINO model. Regarding ViT, it can be referred to the previous description in the present invention and will not be elaborated here.

[0079] The process of using the dual feature extraction network to perform multi-scale feature extraction and feature fusion on the original image can include the following steps S21 - S25:

[0080] S21: Use the pre-trained CNN network to perform multi-scale feature extraction on the original image to obtain the first sub-feature, the second sub-feature, the third sub-feature, and the fourth sub-feature.

[0081] Specifically, use the pre-trained ResNeSt50d network. The ResNeSt50d network includes a STEM layer and four cascaded residual blocks, and the residual block uses a decentralized attention block.

[0082] Let the resolution of the original image be A. Input the original image into the ResNeSt50d network. After being processed by the STEM layer, it is input into the first residual block, and sequentially passes through four cascaded residual blocks to extract image features with resolutions of A / 4, A / 8, A / 16, and A / 32 respectively. The four cascaded residual blocks output the first sub-feature, the second sub-feature, the third sub-feature, and the fourth sub-feature respectively.

[0083] S22: Use the pre-trained ViT network to extract features from the original image to obtain the image class label and the fifth sub-feature.

[0084] Optionally, the Transformer encoder of the ViT network includes 12 repetitively stacked encoder blocks.

[0085] Input the original image into the ViT network. After being processed by the Patch Embedding layer, it is input into the first encoder block. Each encoder block extracts image features of A / 16 and inputs the result into the next encoder block. Finally, after 12 encoder blocks, the fifth sub-feature with a resolution of A / 16 is output.

[0086] S23: Use the feature fusion network to fuse the fourth sub-feature, the third sub-feature, and the fifth sub-feature to obtain the second image feature.

[0087] Specifically, the sampling layer has three layers. Use the third sampling layer to upsample the fourth sub-feature by 2 times and fuse it with the third sub-feature to obtain the fused third sub-feature. After performing a channel dimension transformation on the fused third sub-feature, add it to the fifth sub-feature to obtain the second image feature.

[0088] S24: Use the feature fusion network to fuse the second image feature, the second sub-feature, and the first sub-feature to obtain the first image feature.

[0089] Specifically, use the second sampling layer to upsample the second image feature by 2 times and fuse it with the second sub-feature to obtain the fused second sub-feature. Then use the first sampling layer to upsample the fused second sub-feature by 2 times and fuse it with the first sub-feature to obtain the first image feature.

[0090] In this step, use the dual feature extraction network to extract the first image feature and the second image feature. The first image feature is the fusion of multi-scale features of the original image, making full use of multi-scale features and being able to improve the segmentation accuracy. The second image feature can express the semantic features of the original image, and enhancing the prompt feature using the second image feature can obtain a more accurate prompt feature.

[0091] Next, step S3 will be introduced in detail.

[0092] In step S3, calculate the similarity between the image class identifier and the positive hint features in the knowledge base, obtain the top k positive hint features with the highest similarity ranking, and use the k positive hint features and their associated k negative hint features as the target hint features.

[0093] Specifically, for step S3, one possible implementation is to retrieve from the knowledge base using the image class identifier as the retrieval condition. By calculating the similarity between the image class identifier and the positive hint features in the knowledge base, according to the similarity ranking, select the top k positive hint features, and use the k positive hint features and their associated k negative hint features as the target hint features. The k positive hint features and negative hint features correspond to t types of ground object categories. Among them, each type of ground object category corresponds to k1, k2,..., kt positive hint features and negative hint features respectively, and the sum of k1, k2,..., kt is k.

[0094] The target hint features can correspond to multiple types of ground object categories. By performing subsequent step processing based on the target hint features, segmentation masks of multiple types of ground object categories in the original image can be obtained.

[0095] However, sometimes the k positive hint features obtained by searching according to the image class identifier in this method may not cover all the ground object categories in the original image. Therefore, in some cases, it is a synchronous segmentation of partial elements.

[0096] To segment all the ground object categories in the original image at one time, that is, synchronous segmentation of all elements, in step S3, a preferred implementation can be adopted. Specifically, assume that the m types of ground object categories in the current knowledge base cover all the types of ground object categories in the original image. Using the image class identifier and the labels in the knowledge base as parallel search conditions, for each label, obtain the top r positive hint features with the highest similarity ranking retrieved from the knowledge base, and associate the corresponding r negative hint features. Then, for the m labels, a total of m×r positive hint features and m×r negative hint features are obtained, and use the m×r positive hint features and m×r negative hint features as the target hint features, where m×r is equal to k.

[0097] For example, assume that the total number of ground feature categories in the current knowledge base includes three categories: water area, grassland, and building, with labels 1, 2, and 3 respectively. For any original image, the image class identifier and labels 1, 2, and 3 are used as parallel conditions. For each label, the top 5 positive prompt features with the highest similarity ranking are searched from the knowledge base. Thus, there are 5 positive prompt features corresponding to the water area, grassland, and building respectively, a total of 15 positive prompt features, and a total of 15 negative prompt features are associated. Among them, if some prompt features cannot be matched or the quantity is insufficient, they are supplemented with zeros. In this way, the obtained target prompt features can match all ground feature categories in the original image. By performing subsequent steps based on the target prompt features, the segmentation masks corresponding to all ground feature categories in the original image can be obtained, realizing full-element segmentation.

[0098] In this embodiment, the target prompt features corresponding to multiple ground feature categories of the original image can be obtained through the image class identifier, enabling the model to simultaneously segment multiple ground feature types of the original image and improving the segmentation efficiency.

[0099] Next, in step S4, the prompt embedding module shown in FIG. 4 can be used to perform feature enhancement processing on the prompt features using the second image feature to obtain enhanced prompt features.

[0100] Specifically, the prompt embedding module adopts a cross-attention network. The second image feature and the target prompt feature are input into the cross-attention network. The cross-attention network uses the cross-attention mechanism to perform feature enhancement processing on the target prompt feature to obtain an enhanced target prompt feature. The target prompt feature is mapped to a query vector (Query), the second image feature is mapped to a key vector (Key) and a value vector (Value). The cross-attention mechanism is used to continuously update Query, while Key and Value remain unchanged, to obtain the enhanced target prompt feature.

[0101] In the embodiment of the present invention, the second image feature represents the semantic feature of the original image. The prompt embedding module only updates the target prompt feature and does not update the second image feature, which can reduce the computational complexity and will not interfere with the second image feature.

[0102] In step S5, mask decoding processing is performed on the first image feature and the enhanced target prompt feature to obtain a segmentation mask for the original image. This step can be executed by Figure 4 the mask decoding module in the image segmentation model shown.

[0103] Specifically, the mask decoding module performs matrix multiplication on the first image feature and the enhanced target hint feature in the feature dimension to obtain a plurality of positive feature maps and a plurality of negative feature maps, corresponding to multiple ground object categories. For each ground object category, a positive final feature map and a negative final feature map are obtained respectively based on the corresponding several positive feature maps and several negative feature maps. Then, the softmax function is used to normalize the positive final feature map and the negative final feature map to obtain a probability map, and the segmentation mask for one ground object category in the original image can be obtained from the probability map. The above process is synchronously performed on multiple ground object categories in the original image to obtain a plurality of probability maps, and then the segmentation mask for the original image can be obtained. The segmentation mask contains the position information of multiple ground object categories in the original image. Finally, the segmentation mask is upsampled to the resolution of the original image to complete the hint segmentation.

[0104] Optionally, a corresponding positive final feature map is obtained by calculating the average value of several positive feature maps, and a corresponding negative final feature map is obtained by calculating the average value of several negative feature maps. Optionally, a corresponding positive final feature map and a corresponding negative final feature map can also be obtained by respectively finding the maximum values of several positive feature maps and several negative feature maps. The average value method is preferred in this application.

[0105] Corresponding to step S3, in an optional manner, the mask decoding module performs matrix multiplication on the first image feature and the enhanced target hint feature in the feature dimension to obtain k positive feature maps and k negative feature maps, corresponding to t ground object categories. Each ground object category corresponds to k1, k2,..., kt positive feature maps and negative feature maps respectively. The average value of k1 positive feature maps is calculated to obtain the positive final feature map of the first ground object category. Similarly, the average values of k2,..., kt positive feature maps are calculated respectively to obtain the positive final feature maps of the second to the t-th ground object categories. Similarly, the negative final feature maps corresponding to each ground object category are obtained. Then, the softmax function is used to normalize each pair of positive final feature maps and negative final feature maps to obtain probability maps corresponding to t ground object categories.

[0106] Corresponding to step S3, in a preferred manner, the mask decoding module performs matrix multiplication on the first image feature and the enhanced target hint feature in the feature dimension to obtain m×r positive feature maps and m×r negative feature maps, corresponding to m ground object categories. Each ground object category has r positive feature maps and r negative feature maps. The average values of the r positive feature maps and the r negative feature maps of each ground object category are calculated to obtain m positive final feature maps and m negative final feature maps. Then, the softmax function is used to normalize each pair of positive final feature maps and negative final feature maps to obtain probability maps corresponding to m ground object categories.

[0107] Figure 5Schematic block diagram of the image segmentation device provided by the embodiment of the present application. As shown in Figure 5 shown in , the device may include: an input unit 501, a feature extraction unit 502, a hint matching unit 503, a hint embedding unit 504, and a mask decoding unit 505. The main functions of each component unit are as follows:

[0108] The input unit 501 is configured to obtain an original image and a knowledge base, where the knowledge base contains multiple pairs of positive hint features and negative hint features;

[0109] The feature extraction unit 502 is configured to perform feature extraction on the original image to obtain image features and an image class identifier;

[0110] The hint matching unit 503 is configured to match, using the image class identifier, target hint features corresponding to the original image from the knowledge base;

[0111] The hint embedding unit 504 is configured to perform feature enhancement processing on the target hint features using the image features to obtain enhanced target hint features;

[0112] The mask decoding unit 505 is configured to perform mask decoding processing on the image features and the enhanced hint features to obtain a segmentation mask for the original image.

[0113] The embodiment of the present invention further provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method described in any one of the foregoing method embodiments are implemented.

[0114] Further, the electronic device further includes: at least one input device; at least one output device. The above-mentioned memory, processor, input device, and output device are connected through a bus. Among them, the input device may specifically be a camera, a touch panel, a physical button, or a mouse, etc. The output device may specifically be a display screen. The memory may be a high-speed random access memory (RAM, Random Access Memory) or a non-volatile memory, such as a disk memory. The memory is used to store a set of executable program codes, and the processor is coupled to the memory.

[0115] The embodiment of the present invention further provides a computer-readable storage medium, which may be disposed in the electronic device in each of the foregoing embodiments, and the computer-readable storage medium may be the electronic device in the foregoing embodiments. A computer program is stored on the computer-readable storage medium, and when the program is executed by a processor, the steps of the method described in any one of the foregoing method embodiments are implemented.

[0116] Furthermore, the computer-readable storage medium may also be various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0117] It should be noted that in each embodiment of the present invention, the functional modules may be integrated into one processing module, or each module may exist physically alone, or two or more modules may be integrated into one module. The above-mentioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product.

[0118] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps may be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0119] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0120] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An image segmentation method, characterized in that, The method includes: S1: Obtain the original image and the knowledge base. The knowledge base is pre - constructed and contains multiple pairs of positive prompt features and negative prompt features; S2: Extract features from the original image to obtain image features and an image class identifier; S3: Use the image class identifier to match the target prompt features corresponding to the original image from the knowledge base; S4: Use the image features to perform feature enhancement processing on the target prompt features to obtain enhanced target prompt features; S5: Perform mask decoding processing on the image features and the enhanced target prompt features to obtain a segmentation mask for the original image.

2. The method according to claim 1, wherein The S3 includes: Calculate the similarity between the image class identifier and the positive prompt features in the knowledge base, and use the top k positive prompt features sorted by similarity and their associated k negative prompt features as the target prompt features.

3. The method according to claim 1, characterized in that The construction process of the knowledge base includes: Obtain knowledge images, and use masks to label multiple ground object categories in the knowledge images and set labels respectively to obtain labeled knowledge images; For the labeled knowledge images, obtain sub - images within the mask region of a certain ground object category, and use a pre - trained ViT network to obtain the class identifier of the sub - images to obtain the positive prompt features; Obtain sub - images not within the mask region of a certain ground object category, use a pre - trained ViT network to obtain the class identifier of the sub - images to obtain the negative prompt features, and store the positive prompt features, the negative prompt features, and the label of the certain ground object category in an associated manner.

4. The method according to claim 1, characterized in that, The S2 includes: Perform multi - scale feature extraction and feature fusion on the original image to obtain a first image feature, a second image feature, and the image class identifier, and use the first image feature and the second image feature as the image features; where the scale of the first image feature is larger than the scale of the second image feature.

5. The method according to claim 4, characterized in that, The method for obtaining the first image feature, the second image feature, and the image class identifier includes: Use a pre - trained CNN network to perform multi - scale feature extraction on the original image to obtain first sub - features, second sub - features, third sub - features, and fourth sub - features with gradually decreasing scales; Use a pre - trained ViT network to perform feature extraction on the original image to obtain the image class identifier and a fifth sub - feature; Use a feature fusion network to fuse the fourth sub - feature, the third sub - feature, and the fifth sub - feature to obtain the second image feature; Use a feature fusion network to fuse the second image feature, the second sub - feature, and the first sub - feature to obtain the first image feature.

6. The method according to any one of claims 4-5, characterized in that The S4 includes: Input the second image feature and the target prompt features into a Transformer network, and use the cross - attention mechanism to perform feature enhancement processing on the target prompt features to obtain the enhanced target prompt features.

7. The method according to claim 6, wherein The S5 includes: Perform matrix multiplication on the first image feature and the enhanced target prompt features to obtain a segmentation mask for the original image, where the segmentation mask corresponds to multiple ground object categories of the original image.

8. An image segmentation device, characterized in that, The device includes: an input unit configured to obtain an original image and a knowledge base, where the knowledge base is pre-constructed and contains multiple pairs of positive hint features and negative hint features; a feature extraction unit configured to perform feature extraction on the original image to obtain image features and an image class identifier; a hint matching unit configured to match, using the image class identifier, target hint features corresponding to the original image from the knowledge base; a hint embedding unit configured to perform feature enhancement processing on the target hint features using the image features to obtain enhanced target hint features; a mask decoding unit configured to perform mask decoding processing on the image features and the enhanced target hint features to obtain a segmentation mask for the original image.

9. An electronic device, characterized in that, The electronic device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the computer program is executed by the processor, the steps of the method according to any one of claims 1-7 are implemented.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, where when the computer program is executed by the processor, the steps of the method according to any one of claims 1-7 are implemented.