Feature screening method, feature screening device and feature screening system

Through the deep learning model, the target image and prompt image are extracted and screened, and the attention weight is determined based on cosine similarity and classification scores. The problems of long inference time and low segmentation accuracy under the prompt of multiple visual images in the prior art are solved, efficient feature screening and fusion are achieved, and high-precision segmentation results are maintained.

CN120164050APending Publication Date: 2025-06-17BEIJING LUSTER LIGHTTECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510191198.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

When facing multiple visual images as prompts, the existing segmentation method usually uses multiple inferences or no choice to fuse the prompt image features, resulting in a longer inference time or a reduced segmentation accuracy.

Method used

A feature screening method is proposed, which uses deep learning model to extract the target image and prompt image, calculates the cosine similarity and classification score, determines the attention weight, filters out important prompt features, and performs feature fusion to generate fusion prompt features.

Benefits of technology

This method can eliminate redundant prompt feature information without adding additional inference time and video memory to maintain high-precision segmentation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164050A_ABST
    Figure CN120164050A_ABST
Patent Text Reader

Abstract

The invention discloses a feature screening method, a feature screening device and a feature screening system. The feature screening method comprises the steps of performing feature extraction on a target image through a deep learning model to obtain a target feature; for each prompt image, performing feature extraction on the prompt image through a deep learning model to obtain a first prompt feature; calculating cosine similarity based on the target feature, the first prompt feature and the corresponding prompt mask; calculating classification scores based on the target features, the first prompt features and the corresponding prompt masks through a deep learning model; determining an attention weight according to the cosine similarity and the classification score to screen the first prompt feature to obtain a second prompt feature; and performing feature fusion on the plurality of second prompt features corresponding to the plurality of prompt images to obtain a fused prompt feature. Therefore, extra reasoning duration and reasoning video memory are not increased, redundant prompt feature information can be eliminated, and a high-precision segmentation result is kept.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of machine vision technology, and particularly to a feature screening method, a feature screening device, and a feature screening system. Background Art

[0002] Segmentation based on visual image prompts is a computer vision task, whose goal is to guide the model to segment the corresponding target object or region in the image by providing the visual image and its corresponding mask as prompt information. In current segmentation methods based on visual image prompts, when faced with multiple visual images as prompts, they often use multiple inferences or fuse the prompt images without selection into one feature. These operations will result in a longer inference time or too much redundant information, thus affecting the final segmentation accuracy. Summary of the Invention

[0003] Embodiments of this application provide a feature screening method, a feature screening device, a feature screening system, and a computer-readable storage medium to solve at least one of the above technical problems.

[0004] The feature screening method of the embodiments of this application includes:

[0005] Obtain a target image, multiple prompt images, and their corresponding multiple prompt masks;

[0006] Extract features from the target image through a deep learning model to obtain target features;

[0007] For each of the prompt images, extract features from the prompt image through the deep learning model to obtain first prompt features;

[0008] Calculate the cosine similarity based on the target features, the first prompt features, and the corresponding prompt masks;

[0009] Calculate classification scores based on the target features, the first prompt features, and the corresponding prompt masks through the deep learning model;

[0010] Determine attention weights according to the cosine similarity and the classification scores to screen the first prompt features to obtain second prompt features;

[0011] Fuse the multiple second prompt features corresponding to the multiple prompt images to obtain fused prompt features.

[0012] In some embodiments, the step of extracting features from the target image through a deep learning model to obtain target features includes:

[0013] Use a first feature encoder to perform feature encoding on the target image to obtain the target features;

[0014] Performing feature extraction on the prompt image through the deep learning model to obtain a first prompt feature, including:

[0015] Performing feature encoding on the prompt image using a second feature encoder to obtain the first prompt feature;

[0016] Wherein, the first feature encoder and the second feature encoder have the same weights.

[0017] In some embodiments, before performing feature extraction on the target image through the deep learning model to obtain target features, the feature screening method further includes:

[0018] For each of the prompt images, labeling the prompt image to obtain a prompt label according to whether the prompt category of the prompt mask corresponding to the prompt image exists in the target image;

[0019] Training the deep learning model according to the target image, multiple prompt images, corresponding multiple prompt labels and multiple prompt masks.

[0020] In some embodiments, calculating the cosine similarity based on the target feature, the first prompt feature and the corresponding prompt mask includes:

[0021] Adjusting the size of the prompt mask to have the same spatial dimension as the first prompt feature;

[0022] Multiplying the first prompt feature by the size-adjusted prompt mask to obtain a third prompt feature;

[0023] Calculating the cosine similarity according to the third prompt feature and the target feature.

[0024] In some embodiments, calculating a classification score through the deep learning model based on the target feature, the first prompt feature and the corresponding prompt mask includes:

[0025] Adjusting the size of the prompt mask to have the same spatial dimension as the first prompt feature;

[0026] Multiplying the first prompt feature by the size-adjusted prompt mask to obtain a third prompt feature;

[0027] Concatenating and fusing the third prompt feature and the target feature to obtain a fourth prompt feature;

[0028] Performing global average pooling on the fourth prompt feature to obtain the classification score.

[0029] In some embodiments, determining the attention weights according to the cosine similarity and the classification scores to filter the first prompt features to obtain second prompt features includes:

[0030] Adding the cosine similarity and the classification scores and passing through an activation function to output the attention weights;

[0031] Multiplying the attention weights by the first prompt features to obtain the second prompt features.

[0032] In some embodiments, fusing the multiple second prompt features corresponding to the multiple prompt images to obtain a fused prompt feature includes:

[0033] Performing feature fusion on the multiple second prompt features through at least two cross-attention mechanisms to obtain the fused prompt feature.

[0034] In some embodiments, performing feature fusion on the multiple second prompt features through at least two cross-attention mechanisms to obtain the fused prompt feature includes:

[0035] Determining a first cross-attention result according to the multiple second prompt features and the first prompt feature;

[0036] Determining a second cross-attention result according to the first cross-attention result and the target feature to obtain the fused prompt feature.

[0037] The feature screening device according to the embodiment of the present application includes:

[0038] An acquisition module, configured to acquire a target image, multiple prompt images, and corresponding multiple prompt masks;

[0039] A first extraction module, configured to extract features from the target image through a deep learning model to obtain target features;

[0040] A second extraction module, configured to, for each of the prompt images, extract features from the prompt image through the deep learning model to obtain first prompt features;

[0041] A first calculation module, configured to calculate a cosine similarity based on the target features, the first prompt features, and the corresponding prompt masks;

[0042] A second calculation module, configured to calculate a classification score through the deep learning model based on the target features, the first prompt features, and the corresponding prompt masks;

[0043] A screening module, configured to determine an attention weight according to the cosine similarity and the classification score, so as to screen the first prompt feature to obtain a second prompt feature;

[0044] A fusion module, configured to perform feature fusion on multiple second prompt features corresponding to multiple prompt images to obtain a fused prompt feature.

[0045] The feature screening system according to the embodiment of the present application includes one or more processors and a memory. The memory stores a computer program, and when the computer program is executed by the processor, the feature screening method according to any of the above embodiments is implemented.

[0046] The computer-readable storage medium according to the embodiment of the present application, on which a computer program is stored, and when the program is executed by a processor, the feature screening method according to any of the above embodiments is implemented.

[0047] The feature screening method, feature screening device, feature screening system, and computer-readable storage medium according to the embodiments of the present application effectively combine traditional methods with deep learning methods to screen features of multiple prompt images, can avoid increasing additional inference time and inference video memory, and can also eliminate redundant prompt feature information while maintaining a high-precision segmentation result.

[0048] Additional aspects and advantages of the embodiments of the present application will be partially given in the following description, will become apparent from the following description, or will be understood through the practice of the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, where:

[0050] Figure 1 is one of the flow diagrams of the feature screening method according to some embodiments of the present application;

[0051] Figure 2 is one of the working process diagrams of the feature screening method according to some embodiments of the present application;

[0052] Figure 3 is one of the working process diagrams of the feature screening method according to some embodiments of the present application;

[0053] Figure 4 is one of the flow diagrams of the feature screening method according to some embodiments of the present application;

[0054] Figure 5 is one of the flow diagrams of the feature screening method according to some embodiments of the present application;

[0055] Figure 6It is the fourth flow schematic diagram of the feature screening method in some embodiments of the present application;

[0056] Figure 7 It is the fifth flow schematic diagram of the feature screening method in some embodiments of the present application;

[0057] Figure 8 It is the sixth flow schematic diagram of the feature screening method in some embodiments of the present application;

[0058] Figure 9 It is the seventh flow schematic diagram of the feature screening method in some embodiments of the present application;

[0059] Figure 10 It is the eighth flow schematic diagram of the feature screening method in some embodiments of the present application;

[0060] Figure 11 It is one of the module schematic diagrams of the feature screening device in some embodiments of the present application;

[0061] Figure 12 It is the second module schematic diagram of the feature screening device in some embodiments of the present application;

[0062] Figure 13 It is the module schematic diagram of the feature screening system in some embodiments of the present application;

[0063] Figure 14 It is the schematic diagram of the connection state between the computer-readable storage medium and the processor in some embodiments of the present application.

[0064] Explanation of reference numerals:

[0065] Feature screening device 100, acquisition module 10, first extraction module 20, second extraction module 30, first calculation module 40, second calculation module 50, screening module 60, fusion module 70, annotation module 80, training module 90, feature screening system 200, processor 210, memory 220, computer-readable storage medium 300, computer program 310, processor 320. Specific embodiments

[0066] The following further describes the embodiments of the present application with reference to the accompanying drawings. The same or similar reference numerals in the drawings denote the same or similar elements or elements with the same or similar functions throughout. In addition, the embodiments of the present application described below with reference to the accompanying drawings are exemplary and are only used to explain the embodiments of the present application and should not be construed as a limitation of the present application.

[0067] Visual image prompt-based segmentation is a computer vision task, whose goal is to guide the model to segment the corresponding target object or region in the image by using the provided visual image and its corresponding mask as prompt information. In current visual image prompt-based segmentation methods, when faced with multiple visual images as prompts, they often adopt multiple inferences or fuse the prompt images without selection into one feature. These operations will lead to longer inference time or excessive redundant information, thus affecting the final segmentation accuracy.

[0068] Therefore, there is an urgent need for a feature screening method for prompt images to extract the most relevant features for the target task from the high-dimensional prompt image data, while reducing redundant information to improve the efficiency and performance of the model. Through research, it is found that feature screening techniques can be roughly divided into traditional methods and deep learning-based methods. Among them, traditional methods mainly rely on manually designed features and statistical analysis, including mainly filter methods (evaluating the importance of features through statistical metrics such as information gain, mutual information, variance, chi-square test, etc.), wrapper methods (evaluating the contribution of features based on the prediction performance of a specific model, such as classification accuracy, and common methods include forward selection, backward elimination, and recursive feature elimination, etc.), and embedding methods (integrating feature selection into the model training process, such as LASSO regression based on L1 regularization, feature importance evaluation in decision trees, etc.). In recent years, thanks to the booming development of deep learning, it has provided more powerful method support for feature screening. For example, automatic feature extraction (deep neural networks, such as convolutional neural networks, can automatically learn the hierarchical features of images through multi-layer structures, thus avoiding the limitations of traditional manually designed features), attention mechanism (the attention mechanism assigns different weights to different parts of the input data to screen out the features more relevant to the target task), sparse coding and feature dimensionality reduction (using sparse constraints to achieve feature screening and dimensionality reduction in deep networks), and methods based on generative adversarial networks (generative adversarial networks can not only generate data, but also select the features most relevant to the task through discriminators) and other methods are all based on deep learning methods.

[0069] As described above, in the current segmentation method based on visual image prompts, when faced with multiple visual images as prompts, operations such as multiple inferences or non-selective fusion are often adopted, resulting in an increase in the overall inference time of the model or a decrease in segmentation accuracy, making it difficult to be applied to scenarios such as industry that pursue high precision and high efficiency. Therefore, there is an urgent need for a feature selection method to remove redundant information and maintain low computational latency. However, traditional feature screening methods, although having strong interpretability, being simple and fast (such as the filtering method), not relying on specific machine learning models, and being suitable for early-stage feature screening. However, when faced with high-dimensional complex data, it is difficult to capture deep feature patterns. And the feature screening method based on deep learning, although significantly improving the automation ability and adaptability of feature screening, has weak interpretability and still has great room for improvement in terms of adaptability and robustness in specific tasks such as prompt segmentation.

[0070] To address the above deficiencies, the embodiments of the present application propose a feature selection method for prompt images. This method combines traditional methods and deep learning methods to perform feature selection on multiple input prompt images, so as to remove redundant information without increasing additional inference time. The embodiments of the present application mainly include the feature selection process of prompt images and the processing process of the features of the selected prompt images.

[0071] Please refer to Figure 1 , the feature selection method of the embodiments of the present application includes:

[0072] 010: Obtain a target image, multiple prompt images, and corresponding multiple prompt masks;

[0073] 020: Extract features from the target image through a deep learning model to obtain target features;

[0074] 030: For each prompt image, extract features from the prompt image through a deep learning model to obtain first prompt features;

[0075] 040: Calculate the cosine similarity based on the target features, the first prompt features, and the corresponding prompt masks;

[0076] 050: Calculate the classification score based on the target features, the first prompt features, and the corresponding prompt masks through a deep learning model;

[0077] 060: Determine the attention weights according to the cosine similarity and the classification score to screen the first prompt features to obtain second prompt features;

[0078] 070: Perform feature fusion on the multiple second prompt features corresponding to the multiple prompt images to obtain fused prompt features.

[0079] The feature screening method of the embodiment of the present application effectively combines traditional methods with deep learning methods to screen features of multiple prompt images, which can avoid increasing additional inference time and inference video memory, while also removing redundant prompt feature information and maintaining a high-precision segmentation result.

[0080] Specifically, please combine Figure 2 and Figure 3 , first, obtain a target image, multiple prompt images, and corresponding multiple prompt masks. The target image refers to the image for which the target task needs to be performed, such as the image from which the target object or region needs to be segmented. The prompt image refers to the image used to provide prompt information to accurately segment the target object or region from the target image. The prompt image contains relevant information about the target object or region to be segmented, etc. The prompt mask is used to clearly define the target object or region to be segmented in the selected target image. The prompt mask can, for example, define the contour of the target object or region. Each prompt image has a corresponding prompt mask, and the prompt mask can be a binary or Boolean image of the same size as the prompt image. The prompt mask can be obtained by manual selection or automatic generation.

[0081] Then, extract features of the target image through a deep learning model to obtain target features. In addition, for each prompt image, extract features of the prompt image through a deep learning model to obtain first prompt features. For example, feature extraction can be performed through a convolutional neural network or a deep neural network.

[0082] Next, on the one hand, calculate the cosine similarity based on the target features, the first prompt features, and the corresponding prompt mask. This process is a traditional method. On the other hand, calculate the classification score through a deep learning model based on the target features, the first prompt features, and the corresponding prompt mask. This process is a deep learning method. After obtaining the cosine similarity and the classification score, determine the attention weight according to the cosine similarity and the classification score, and screen the first prompt features based on the attention weight to obtain second prompt features to highlight important features and suppress redundant information.

[0083] Finally, perform feature fusion on the multiple second prompt features corresponding to the multiple prompt images to obtain fused prompt features. It should be noted that for each of the multiple prompt images, the steps 030-060 above need to be executed, so that for each prompt image, the corresponding second prompt features are screened. In 070, the multiple second prompt features corresponding to the multiple prompt images are fused to obtain fused prompt features to further enhance the feature detail information.

[0084] It is found through research that if a pure deep learning method is used, a more complex module structure with a larger number of parameters needs to be designed to achieve the effect of suppressing the same redundant features; if a pure traditional method is used, due to the complex feature information of multiple prompt images and the aforementioned disadvantages of the traditional method, it is difficult to bring a considerable screening effect, and the final segmentation accuracy will also be greatly reduced.

[0085] The feature screening method of the embodiment of the present application effectively combines the traditional method and the deep learning method to screen the features of multiple prompt images. Through effective verification, it can not increase the additional inference time and inference video memory, and at the same time can eliminate redundant prompt feature information and maintain a high-precision segmentation result. In addition, the feature screening method of the embodiment of the present application also has the characteristics of high robustness and high precision of deep learning, as well as simplicity and high efficiency of the traditional method. The feature screening method of the embodiment of the present application still does not require multiple inferences and multi-batch inferences in the scenario with multiple prompt images of different categories, and can maintain the accuracy of multiple inferences.

[0086] Please refer to Figure 4 , in some embodiments, the target image is subjected to feature extraction through a deep learning model to obtain target features (i.e., 020), including:

[0087] 021: Using a first feature encoder to perform feature encoding on the target image to obtain target features;

[0088] The prompt image is subjected to feature extraction through a deep learning model to obtain first prompt features (i.e., 030), including:

[0089] 031: Using a second feature encoder to perform feature encoding on the prompt image to obtain first prompt features;

[0090] Among them, the first feature encoder and the second feature encoder have the same weight.

[0091] Specifically, please refer to Figure 3 , the feature encoder is a neural network model. Through feature encoding, high-dimensional features can be extracted from image data. Compared with the original image data, the high-dimensional features have higher information density, stronger expression ability, and are more discriminative and representative, so as to facilitate subsequent classification, recognition or segmentation and other tasks. Using a first feature encoder to perform feature encoding on the target image to obtain target features; using a second feature encoder to perform feature encoding on the prompt image to obtain first prompt features.

[0092] The first feature encoder and the second feature encoder have the same weights. At this time, when the target image passes through the first feature encoder, it shares the same network parameters with the prompt image when passing through the second feature encoder. That is to say, each layer of the feature encoder (such as convolutional layer, fully connected layer, etc.) applies the same transformation rules to the two images. In this way, it helps to ensure the consistency of the feature encoder when processing different images and makes the features extracted from the two images comparable.

[0093] In the embodiments of the present application, the first feature encoder and the second feature encoder can adopt any structure, including but not limited to ResNet series convolutional neural network structures, ViT series Transformer structures, etc. The ResNet series convolutional neural network structures can improve the training speed and performance of the model while maintaining the model complexity. The ViT series Transformer structures have strong global feature extraction capabilities and scalability. In practical applications, a suitable structure can be adopted according to requirements, which is not limited here. Generally speaking, the higher the accuracy of the feature encoder, the more beneficial it is to improve the accuracy of the segmentation result.

[0094] In one embodiment, the first feature encoder and the second feature encoder may have the same structure. For example, both are composed of the same type of layers (such as convolutional layer, pooling layer, fully connected layer, etc.) in the same order and configuration, so as to follow the same logic and transformation rules when processing input data. Of course, the first feature encoder and the second feature encoder can also select exactly the same feature encoder to have the same weights and the same structure.

[0095] Please refer to Figure 5 , in some embodiments, calculating the cosine similarity (i.e., 040) based on the target feature, the first prompt feature, and the corresponding prompt mask includes:

[0096] 041: Adjust the size of the prompt mask to have the same spatial dimension as the first prompt feature;

[0097] 042: Multiply the first prompt feature by the size-adjusted prompt mask to obtain a third prompt feature;

[0098] 043: Calculate the cosine similarity according to the third prompt feature and the target feature.

[0099] In order to highlight as much as possible the images similar to the target image and suppress the redundant or significantly different images from the target image, the embodiments of the present application calculate the cosine similarity based on the target feature, the first prompt feature, and the corresponding prompt mask without adding additional parameters.

[0100] Specifically, please refer to Figure 3, first, adjust the size of the prompt mask to the same spatial dimension as the first prompt feature to facilitate their multiplication. Specifically, interpolation methods such as bilinear interpolation, nearest neighbor interpolation, etc. can be used for size adjustment. Then, multiply the first prompt feature by the size-adjusted prompt mask to extract the most important regional features in the first prompt feature, obtaining the third prompt feature. Since the first prompt feature and the prompt mask already have the same spatial dimension, the multiplication method can be element-wise multiplication. Finally, directly calculate the cosine similarity based on the third prompt feature and the target feature. The cosine similarity is based on the angle between feature vectors and evaluates the similarity between the prompt image and the target image. In this way, the purpose of similarity matching can be achieved across the entire target image according to the labeled region of the prompt mask.

[0101] Please refer to Figure 6 , in some embodiments, calculating a classification score (i.e., 050) based on the target feature, the first prompt feature, and the corresponding prompt mask through a deep learning model includes:

[0102] 051: Adjust the size of the prompt mask to have the same spatial dimension as the first prompt feature;

[0103] 052: Multiply the first prompt feature by the size-adjusted prompt mask to obtain the third prompt feature;

[0104] 053: Concatenate and fuse the third prompt feature and the target feature to obtain the fourth prompt feature;

[0105] 054: Perform global average pooling on the fourth prompt feature to obtain the classification score.

[0106] Specifically, the specific process of 051 is the same as that of 041, and the specific process of 052 is the same as that of 042. The explanations for 041 and 042 in the foregoing embodiments also apply to 051 and 052 of the embodiments of the present application, and will not be elaborated here.

[0107] Please refer to Figure 3 , after obtaining the third prompt feature, concatenate the third prompt feature and the target feature, and use 1x1 convolution for fusion to obtain the fourth prompt feature. Among them, the concatenation can be along the channel dimension. 1x1 convolution refers to performing a convolution operation using a 1x1 convolution kernel to fuse the concatenated features. Finally, use the global average pooling operation to compress the fourth prompt feature into a one-dimensional vector as the classification score.

[0108] After separately determining the cosine similarity and the classification score, the cosine similarity and the classification score can be combined to determine the attention weight for feature screening. This process is similar to the calculation method of the attention mechanism, which can suppress redundant feature information and strengthen the useful features beneficial to target localization and segmentation, so as to achieve high-precision results without increasing the inference time.

[0109] Please refer to Figure 7 , in some embodiments, determining the attention weight according to the cosine similarity and the classification score to screen the first hint feature to obtain the second hint feature (i.e., 060) includes:

[0110] 061: Add the cosine similarity and the classification score, and pass through an activation function to output the attention weight;

[0111] 062: Multiply the attention weight by the first hint feature to obtain the second hint feature.

[0112] Specifically, please refer to Figure 3 , first, add the cosine similarity and the classification score to obtain a comprehensive score. Then, input the comprehensive score into the activation function. By converting the comprehensive score through the activation function, an attention weight between 0 and 1 can be obtained. Among them, the activation function can be the Sigmoid activation function, and the Sigmoid activation function can map any real value to the interval (0, 1), which is suitable for generating the attention weight. Finally, multiply the attention weight by the first hint feature to obtain the second hint feature. In this way, the feature screening of the hint image is realized to highlight important features and suppress redundant information.

[0113] Please refer to Figure 8 , in some embodiments, before the target features (i.e., 020) are obtained by feature extraction of the target image through a deep learning model, the feature screening method further includes:

[0114] 080: For each hint image, label the hint image to obtain a hint label according to whether the hint category corresponding to the hint mask of the hint image exists in the target image;

[0115] 090: Train the deep learning model according to the target image, multiple hint images, and the corresponding multiple hint labels and multiple hint masks.

[0116] Specifically, please refer to Figure 3, first, annotate the prompt image to obtain a prompt label. The prompt label is, for example, a 0 or 1 label, which is used for loss supervision of the deep learning model. The way to annotate the prompt image can be: according to whether the prompt category corresponding to the prompt mask of the prompt image appears in the target image, if it appears, assign 1 to the prompt image as the prompt label; if it does not appear, assign 0 to the prompt image as the prompt label. In one example, the prompt category is a prompt category in industrial inspection defects. Of course, in other examples, the prompt category can also be a prompt category in other aspects, which is not limited here.

[0117] Then, train the deep learning model according to the target image, multiple prompt images, and the corresponding multiple prompt labels and multiple prompt masks. The specific process can be: use the first feature encoder to perform feature encoding on the target image to obtain target features. For each prompt image, use the second feature encoder to perform feature encoding on the prompt image to obtain the first prompt feature; adjust the size of the prompt mask corresponding to the prompt image to have the same spatial dimension as the first prompt feature; multiply the first prompt feature by the size-adjusted prompt mask to obtain the third prompt feature; splice and fuse the third prompt feature with the target feature to obtain the fourth prompt feature; perform global average pooling on the fourth prompt feature to obtain a classification score. Among them, in the process of performing global average pooling on the fourth prompt feature to obtain a classification score, a loss function is also calculated, and loss supervision is performed through the aforementioned prompt label. In this way, the embodiment of the present application uses the target image, multiple prompt images, and the corresponding multiple prompt labels and multiple prompt masks as inputs, and updates the parameters of the deep learning model through loss calculation and gradient backpropagation to complete the model training of the deep learning model.

[0118] It should be noted that in the embodiment of the present application, except for the step of calculating the cosine similarity, other processing processes can be implemented by the deep learning model. The process of model training is basically the same as the application process of feature screening. The difference is that in the process of model training, loss supervision is combined with the prompt label to realize the model training of the deep learning model.

[0119] Please refer to Figure 9 , in some embodiments, perform feature fusion on the multiple second prompt features corresponding to the multiple prompt images to obtain a fused prompt feature (i.e., 070), including:

[0120] 071: Perform feature fusion on the multiple second prompt features through at least two cross-attention mechanisms to obtain a fused prompt feature.

[0121] Specifically, the feature fusion operation includes at least two Cross - Attention mechanisms. In the implementation manner of this application, the feature fusion of multiple second prompt features is performed through at least two Cross - Attention mechanisms to obtain the fused prompt feature, so as to further enhance the feature detail information.

[0122] Please refer to Figure 10 , in some implementation manners, the feature fusion of multiple second prompt features is performed through at least two Cross - Attention mechanisms to obtain the fused prompt feature (i.e., 071), including:

[0123] 0711: Determine the first Cross - Attention result according to multiple second prompt features and the first prompt feature;

[0124] 0712: Determine the second Cross - Attention result according to the first Cross - Attention result and the target feature to obtain the fused prompt feature.

[0125] It can be understood that if the number of Cross - Attention mechanisms is too large, redundant operations will be caused; while if the number of Cross - Attention mechanisms is too small, useful features cannot be comprehensively extracted and fused. In the implementation manner of this application, the feature fusion of multiple second prompt features is performed through two Cross - Attention mechanisms to obtain the fused prompt feature. In this way, useful features can be comprehensively extracted and fused without causing redundant operations.

[0126] Specifically, the second prompt feature (in the form of a one - dimensional vector) obtained after feature screening is first multiplied by the first prompt feature, and the result is then added to the first prompt feature. The added result is input into the first Cross - Attention module and used as the key (K) and value (V) in the first Cross - Attention module, while the query (Q) in the first Cross - Attention module is initialized by a trainable vector. This operation not only helps to further suppress redundant information but also effectively enhances the regional information in the target image corresponding to the prompt category. Therefore, the first Cross - Attention result is used as the Q in the second Cross - Attention module, and the target feature is used as the K and V in the second Cross - Attention module. After passing through two consecutive Cross - Attention modules, the model can accurately locate the corresponding target position information in the target image based on the prompt image and the corresponding prompt mask.

[0127] In an example, after obtaining the fused prompt feature, a fused prompt mask corresponding to the fused prompt feature can also be determined, so as to use the fused prompt feature and the fused prompt mask as prompt information to guide the model to perform a segmentation task on the target image.

[0128] Please refer to Figure 11, the feature screening device 100 according to the embodiments of the present application includes an acquisition module 10, a first extraction module 20, a second extraction module 30, a first calculation module 40, a second calculation module 50, a screening module 60, and a fusion module 70. The feature screening method according to the embodiments of the present application can be implemented by the feature screening device 100 according to the embodiments of the present application. For example, the acquisition module 10 is used to acquire a target image, a plurality of prompt images, and corresponding plurality of prompt masks. The first extraction module 20 is used to extract features from the target image through a deep learning model to obtain target features. The second extraction module 30 is used to, for each prompt image, extract features from the prompt image through a deep learning model to obtain first prompt features. The first calculation module 40 is used to calculate the cosine similarity based on the target features, the first prompt features, and the corresponding prompt masks. The second calculation module 50 calculates a classification score through a deep learning model based on the target features, the first prompt features, and the corresponding prompt masks. The screening module 60 is used to determine attention weights according to the cosine similarity and the classification score to screen the first prompt features to obtain second prompt features. The fusion module 70 is used to perform feature fusion on the plurality of second prompt features corresponding to the plurality of prompt images to obtain fused prompt features.

[0129] In some embodiments, the first extraction module 20 is specifically configured to perform feature encoding on the target image by using a first feature encoder to obtain target features. The second extraction module 30 is specifically configured to perform feature encoding on the prompt image by using a second feature encoder to obtain first prompt features. Wherein, the first feature encoder and the second feature encoder have the same weights.

[0130] Please refer to Figure 12 , in some embodiments, the feature screening device 100 further includes a labeling module 80 and a training module 90. The labeling module 80 is used to, for each prompt image, label the prompt image according to whether the prompt category corresponding to the prompt mask of the prompt image exists in the target image to obtain a prompt label. The training module 90 is used to perform model training on the deep learning model according to the target image, the plurality of prompt images, the corresponding plurality of prompt labels, and the plurality of prompt masks.

[0131] In some embodiments, the first calculation module 40 is specifically configured to: adjust the size of the prompt mask to have the same spatial dimension as the first prompt feature; multiply the first prompt feature by the size-adjusted prompt mask to obtain a third prompt feature; calculate the cosine similarity according to the third prompt feature and the target feature.

[0132] In some embodiments, the second computing module 50 is specifically configured to: adjust the size of the prompt mask to have the same spatial dimension as the first prompt feature; multiply the first prompt feature by the size-adjusted prompt mask to obtain a third prompt feature; splice and fuse the third prompt feature with the target feature to obtain a fourth prompt feature; perform global average pooling on the fourth prompt feature to obtain a classification score.

[0133] In some embodiments, the screening module 60 is specifically configured to: add the cosine similarity to the classification score and output an attention weight through an activation function; multiply the attention weight by the first prompt feature to obtain a second prompt feature.

[0134] In some embodiments, the fusion module 70 is specifically configured to perform feature fusion on multiple second prompt features through at least two cross-attention mechanisms to obtain a fused prompt feature.

[0135] In some embodiments, the fusion module 70 is specifically configured to: determine a first cross-attention result according to multiple second prompt features and the first prompt feature; determine a second cross-attention result according to the first cross-attention result and the target feature to obtain a fused prompt feature.

[0136] It should be noted that the foregoing explanations of the feature screening method in the embodiments are equally applicable to the feature screening device 100 in the embodiments of the present application, and will not be elaborated herein.

[0137] Please refer to Figure 13 , the feature screening system 200 of the embodiments of the present application includes one or more processors 210 and a memory 220, and the memory 220 stores a computer program. When the computer program is executed by the processor 210, the feature screening method of any of the foregoing embodiments is implemented.

[0138] For example, when the computer program is executed by the processor 210, the following feature screening method is implemented:

[0139] 010: Obtain a target image, multiple prompt images, and corresponding multiple prompt masks;

[0140] 020: Extract features from the target image through a deep learning model to obtain a target feature;

[0141] 030: For each prompt image, extract features from the prompt image through a deep learning model to obtain a first prompt feature;

[0142] 040: Calculate the cosine similarity based on the target feature, the first prompt feature, and the corresponding prompt mask;

[0143] 050: Calculate a classification score based on a target feature, a first hint feature, and a corresponding hint mask through a deep learning model;

[0144] 060: Determine an attention weight based on a cosine similarity and the classification score to filter the first hint feature to obtain a second hint feature;

[0145] 070: Perform feature fusion on multiple second hint features corresponding to multiple hint images to obtain a fused hint feature.

[0146] For another example, when a computer program is executed by a processor 210, the following feature screening method is implemented:

[0147] 021: Use a first feature encoder to perform feature encoding on a target image to obtain a target feature;

[0148] 031: Use a second feature encoder to perform feature encoding on a hint image to obtain a first hint feature;

[0149] Among them, the first feature encoder and the second feature encoder have the same weights.

[0150] It should be noted that the explanations of the feature screening method in the foregoing embodiments are equally applicable to the feature screening system 200 of the embodiments of the present application, and will not be elaborated here.

[0151] Please refer to Figure 14 , a computer-readable storage medium 300 of the embodiments of the present application, on which a computer program 310 is stored. When the program is executed by a processor 320, the feature screening method of any of the foregoing embodiments is implemented.

[0152] For example, when the program is executed by a processor 320, the following feature screening method is implemented:

[0153] 010: Obtain a target image, multiple hint images, and corresponding multiple hint masks;

[0154] 020: Perform feature extraction on the target image through a deep learning model to obtain a target feature;

[0155] 030: For each hint image, perform feature extraction on the hint image through a deep learning model to obtain a first hint feature;

[0156] 040: Calculate a cosine similarity based on the target feature, the first hint feature, and the corresponding hint mask;

[0157] 050: Calculate a classification score based on the target feature, the first hint feature, and the corresponding hint mask through a deep learning model;

[0158] 060: Determine the attention weights based on the cosine similarity and classification scores to filter the first prompt features to obtain the second prompt features;

[0159] 070: Perform feature fusion on the multiple second prompt features corresponding to the multiple prompt images to obtain the fused prompt features.

[0160] For another example, when the program is executed by the processor 320, the following feature screening method is implemented:

[0161] 021: Use the first feature encoder to perform feature encoding on the target image to obtain the target features;

[0162] 031: Use the second feature encoder to perform feature encoding on the prompt image to obtain the first prompt features;

[0163] Among them, the first feature encoder and the second feature encoder have the same weights.

[0164] It should be noted that the explanations of the feature screening method in the foregoing embodiments are equally applicable to the computer-readable storage medium 300 of the embodiments of the present application, and will not be elaborated here.

[0165] In summary, the feature screening method, the feature screening device 100, the feature screening system 200, and the computer-readable storage medium 300 of the embodiments of the present application effectively combine traditional methods with deep learning methods to perform feature screening on multiple prompt images, can not increase the additional inference duration and inference video memory, and at the same time can eliminate redundant prompt feature information and maintain a high-precision segmentation result.

[0166] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0167] Any process or method description represented in a flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of the present application includes additional implementations where functions may be executed in a manner that is not shown or discussed, including in a substantially simultaneous manner according to the relevant functions or in a reverse order, which should be understood by those skilled in the technical field to which the embodiments of the present application pertain.

[0168] Logic and / or steps represented in a flowchart or otherwise described herein, for example, can be considered a sequenced list of executable instructions for implementing a logical function and can be specifically implemented in any computer-readable storage medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with such instruction execution systems, apparatus, or devices. For the purposes of this specification, a computer-readable storage medium can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable storage media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, a computer-readable storage medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing as necessary, and then stored in a computer memory.

[0169] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0170] Those of ordinary skill in the art can understand that all or part of the steps carried out in implementing the above-described embodiment methods can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment. In addition, in each of the embodiments of the present application, each functional unit can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of a software functional module. If the above-mentioned integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, or the like.

[0171] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application. The scope of the present application is defined by the claims and their equivalents.

Claims

1. A feature screening method, characterized in that: include: Acquire a target image, a plurality of prompt images, and a plurality of corresponding prompt masks; Extracting features from the target image using a deep learning model to obtain target features; For each of the prompt images, extracting features of the prompt image by using the deep learning model to obtain a first prompt feature; Calculating cosine similarity based on the target feature, the first hint feature, and the corresponding hint mask; Calculating a classification score based on the target feature, the first cue feature, and the corresponding cue mask by the deep learning model; determining an attention weight according to the cosine similarity and the classification score to filter the first prompt feature to obtain a second prompt feature; A plurality of the second prompt features corresponding to the plurality of the prompt images are fused to obtain a fused prompt feature.

2. The feature screening method according to claim 1, characterized in that: The step of extracting features from the target image using a deep learning model to obtain target features includes: Using a first feature encoder to perform feature encoding on the target image to obtain the target feature; The extracting features of the prompt image by using the deep learning model to obtain a first prompt feature includes: Using a second feature encoder to perform feature encoding on the prompt image to obtain the first prompt feature; The first feature encoder and the second feature encoder have the same weight.

3. The feature screening method according to claim 1, characterized in that: Before extracting features from the target image using a deep learning model to obtain target features, the feature screening method further includes: For each of the prompt images, according to whether the prompt category of the prompt mask corresponding to the prompt image exists in the target image, the prompt image is annotated to obtain a prompt label; The deep learning model is trained according to the target image, the plurality of prompt images, the corresponding plurality of prompt labels and the plurality of prompt masks.

4. The feature screening method according to claim 1, characterized in that: The calculating the cosine similarity based on the target feature, the first hint feature and the corresponding hint mask comprises: resizing the hint mask to have the same spatial dimension as the first hint feature; Multiplying the first hint feature by the hint mask after resizing to obtain a third hint feature; The cosine similarity is calculated based on the third hint feature and the target feature.

5. The feature screening method according to claim 1, characterized in that: The calculating the classification score based on the target feature, the first prompt feature and the corresponding prompt mask by the deep learning model includes: resizing the hint mask to have the same spatial dimension as the first hint feature; Multiplying the first hint feature by the hint mask after resizing to obtain a third hint feature; The third prompt feature is combined with the target feature to obtain a fourth prompt feature; Perform global average pooling processing on the fourth prompt feature to obtain the classification score.

6. The feature screening method according to claim 1, characterized in that: The determining of the attention weight according to the cosine similarity and the classification score to filter the first prompt feature to obtain the second prompt feature includes: Add the cosine similarity to the classification score and pass it through an activation function to output the attention weight; The attention weight is multiplied by the first prompt feature to obtain the second prompt feature.

7. The feature screening method according to claim 1, characterized in that: performing feature fusion on a plurality of the second prompt features corresponding to the plurality of the prompt images, Get fusion hint features, including: The plurality of the second prompt features are subjected to feature fusion through at least two cross-attention mechanisms to obtain the fused prompt feature.

8. The feature screening method according to claim 7, characterized in that: The step of fusing the plurality of second prompt features by at least two cross attention mechanisms to obtain the fused prompt feature comprises: determining a first cross-attention result according to a plurality of the second prompt features and the first prompt features; A second cross-attention result is determined according to the first cross-attention result and the target feature to obtain the fusion prompt feature.

9. A feature screening device, characterized in that: include: An acquisition module, used for acquiring a target image, a plurality of prompt images and a corresponding plurality of prompt masks; A first extraction module is used to extract features of the target image through a deep learning model to obtain target features; A second extraction module is used for extracting features of each of the prompt images by using the deep learning model to obtain a first prompt feature; A first calculation module, configured to calculate a cosine similarity based on the target feature, the first hint feature and the corresponding hint mask; A second calculation module calculates a classification score based on the target feature, the first prompt feature and the corresponding prompt mask through the deep learning model; A screening module, used for determining an attention weight according to the cosine similarity and the classification score, so as to screen the first prompt feature to obtain a second prompt feature; The fusion module is used to fuse the plurality of second prompt features corresponding to the plurality of prompt images to obtain a fused prompt feature.

10. A feature screening system, characterized in that: The feature screening system includes one or more processors and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the feature screening method according to any one of claims 1 to 8 is implemented.