A collaborative saliency detection method, product, device and storage medium

By combining pre-trained language models and segmenting all models to generate high-quality semantic information, the problem in the prior art that it is difficult to distinguish synergistic significant goals and backgrounds in the absence of high-quality semantics is solved, and higher recognition capabilities and detection accuracy are achieved.

CN119723114BActive Publication Date: 2025-05-30LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510229137.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-30
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

Existing collaborative significance detection methods are difficult to effectively distinguish synergistic significance targets from interference backgrounds in the absence of high-quality semantics, resulting in limited recognition capabilities and accuracy.

Method used

The synergistic significance detection method based on pre-trained language model and segmentation all models is adopted. By acquiring multimodal features of images and text, high-quality semantic information is generated, which is used to guide the segmentation model to segment the synergistic significance targets and backgrounds.

Benefits of technology

It improves the ability to identify synergistic significant targets and interference backgrounds, breaks through the limitation of semantic-independent binary masks as supervision signals, alleviates the problem of restricted characterization learning caused by training scenario limitations, and thus improves the accuracy of synergistic significance detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119723114B_ABST
    Figure CN119723114B_ABST
Patent Text Reader

Abstract

The present application discloses a collaborative saliency detection method, product, device, and storage medium, which relates to the field of computer vision and includes: obtaining at least two target images to be detected currently to obtain the current image group; inputting the current image group into a trained collaborative saliency detection model to extract target text information related to the collaborative saliency between the target images in the current image group based on a pre-trained language model, and guiding a segmentation model to segment the collaborative salient targets and the background in the current image group based on the target text information to obtain the current collaborative saliency detection result. The present application can improve the recognition ability of collaborative salient targets and interfering backgrounds and enhance the accuracy of collaborative saliency detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision, and particularly to a co-saliency detection method, product, device, and storage medium. Background Art

[0002] Currently, since the co-saliency object detection (CoSOD) technology can help people better extract main information from a large amount of image or video data, it has been widely applied in tasks such as image segmentation, image retrieval, and image quality evaluation. However, the semantic category information required for co-salient objects is unknown, and such information depends on the specific content of the given image group and needs to be inferred through designed algorithms. Therefore, how to effectively distinguish co-salient objects from interfering backgrounds in the case of high-quality semantic loss is a major challenge faced by current co-saliency detection technologies.

[0003] With the rapid development of deep learning, deep learning technologies with powerful image representation extraction capabilities have been widely applied in co-saliency detection tasks and have made remarkable progress. The current mainstream co-saliency detection methods include single-group co-saliency detection methods and multi-group co-saliency detection methods. Among them, the single-group co-saliency detection method is to input a group of images, mine the intra-group relationship to learn discriminative representations, and extract a binary mask that distinguishes co-salient objects from the background for this group of images; while the multi-group co-saliency detection method specifically inputs two or more groups of images, and on the basis of mining the intra-group relationship, introduces inter-group differences, and then learns discriminative representations through intra-group and inter-group relationships, and uses this to provide binary masks for multiple groups of images respectively.

[0004] However, both of the above two co-saliency detection methods use semantically irrelevant binary masks as supervision information, lack the guidance of semantic labels, and are difficult to obtain high-quality semantic information, resulting in limited discriminative representations learned, making the model unable to effectively distinguish co-salient objects from interfering backgrounds within the image. Summary of the Invention

[0005] In view of this, the purpose of this application is to provide a co-saliency detection method, product, device, and storage medium, which can improve the recognition ability of co-salient objects and interfering backgrounds and enhance the accuracy of co-saliency detection. The specific solutions are as follows:

[0006] In a first aspect, this application discloses a co-saliency detection method, including:

[0007] Obtain at least two target images to be currently detected to obtain the current image group;

[0008] Input the current image group into the trained co-saliency detection model; the co-saliency detection model is a model obtained by training a pre-trained language model and a segment-anything model based on an image data set containing co-saliency objects;

[0009] Obtain the current co-saliency detection result corresponding to the current image group output by the co-saliency detection model; wherein, the current co-saliency detection result is obtained by extracting target text information related to the co-saliency between each target image in the current image group based on the pre-trained language model, and guiding the segment-anything model to segment the co-salient objects and the background in the current image group based on the target text information, and the current co-saliency detection result contains a mask result generated based on the target text information.

[0010] Optionally, the co-saliency detection method further includes:

[0011] Obtain an image data set containing co-saliency objects, and group the images in the image data set according to different co-saliency objects to obtain multiple co-salient image groups;

[0012] Input the images in each co-salient image group into the pre-trained language model in turn to extract features of the images in the co-salient image group to obtain image encoding features, and extract features of the preset interactive text to obtain text encoding features;

[0013] Decode the image encoding features and text encoding features corresponding to each image in the co-salient image group to obtain a scene structure knowledge description text containing the interrelationships between different co-saliency objects corresponding to each image;

[0014] Encode each scene structure knowledge description text into text features respectively, and extract keywords in the text features to obtain prompt keyword features corresponding to each scene structure knowledge description text;

[0015] Explore the connections between different prompt keyword features to generate co-salient text prompts;

[0016] Input the co-salient text prompts into the segment-anything model to convert the co-salient text prompts into binary masks, and obtain a co-saliency detection model containing historical co-saliency detection results.

[0017] Optionally, extracting features of each image in the co-salient image group to obtain image encoding features, and extracting features of the preset interactive text to obtain text encoding features includes:

[0018] Use an image feature extractor to extract features from each image in the co-salient image group, obtain image encoding features, and convert the image encoding features into vector representations to obtain image feature vectors;

[0019] Use a text feature extractor to extract features from the preset interactive text, obtain text encoding features, and convert the text encoding features into vector representations to obtain text feature vectors;

[0020] Among them, the image feature extractor and the text feature extractor are located in the pre-trained language model.

[0021] Optionally, decode the image encoding features and text encoding features corresponding to each image in the co-salient image group to obtain a scene structure knowledge description text containing the mutual relationships between different co-salient objects corresponding to each image, including:

[0022] Combine the image feature vectors and text feature vectors corresponding to each image in the co-salient image group to obtain combined feature vectors;

[0023] Input the combined feature vectors into the first decoder located in the pre-trained language model to decode the combined feature vectors and obtain a scene structure knowledge description text containing the mutual relationships between different co-salient objects corresponding to each image.

[0024] Optionally, encode each scene structure knowledge description text into text features and extract keywords from the text features to obtain prompt keyword features corresponding to each scene structure knowledge description text, including:

[0025] Use a text encoder to encode each scene structure knowledge description text respectively to obtain text features, and perform a linear mapping on the text features to obtain mapped features;

[0026] Use a multi-head self-attention mechanism to enhance the features of the mapped features to obtain enhanced features;

[0027] Use a keyword extractor to extract keywords from the enhanced features to obtain multiple prompt keyword features corresponding to each scene structure knowledge description text.

[0028] Optionally, mine the connections between different prompt keyword features to generate co-salient text prompt words, including:

[0029] Use a co-attention mechanism to mine the connections between different prompt keyword features to obtain co-text prompt features;

[0030] Decode the co-text prompt features to obtain co-salient text prompt words.

[0031] Optionally, a co-attention mechanism is used to mine the relationships between different prompt keyword features to obtain co-text prompt features, including:

[0032] Combine multiple prompt keyword features corresponding to the text descriptions of each scenario structure knowledge to obtain a keyword group;

[0033] Input the keyword group into different linear layers respectively, and multiply the output results of different linear layers to obtain an attention affinity graph of the keyword group;

[0034] Extract the first number of values in the attention affinity graph row by row, and calculate the sum of the extracted first number of values to obtain a first sum value;

[0035] Calculate the ratio of the first sum value to the first number to obtain keyword weights, and multiply the keyword weights element-wise with the prompt keyword features in the keyword group to obtain multiple calculation results;

[0036] Extract the first second number of values from the multiple calculation results, and calculate the sum of the extracted first second number of values to obtain a second sum value;

[0037] Calculate the ratio of the second sum value to the second number to obtain co-text prompt features.

[0038] Optionally, during the training of the pre-trained language model and the segment-everything model, it also includes:

[0039] Use the real co-salient category text as a supervision signal to supervise the learning processes of the keyword extractor and the co-attention mechanism.

[0040] Optionally, during the training of the pre-trained language model and the segment-everything model, it also includes:

[0041] Obtain co-salient text labels from the image dataset, and perform high-dimensional mapping on the co-salient text labels to obtain real co-prompt features;

[0042] Use the distance between the real co-prompt features and the co-text prompt features as a loss, and supervise the loss.

[0043] Optionally, convert the co-salient text prompt words into binary masks to obtain a co-salient detection model including historical co-salient detection results, including:

[0044] Perform feature encoding on the co-salient text prompt words to obtain text feature encodings;

[0045] Perform feature encoding on the images in each co-salient image group to obtain image feature encodings;

[0046] Decode the encoded image features and the corresponding encoded text features to convert the co-salient text prompts into a binary mask, obtaining a historical co-saliency detection result, so as to complete the training of the pre-trained language model and the Segment Anything model, and obtain a co-saliency detection model.

[0047] Optionally, perform feature encoding on the co-salient text prompts to obtain encoded text features, including:

[0048] Input the co-salient text prompts into a text prompt encoder to perform feature encoding on the co-salient text prompts and obtain encoded text features;

[0049] Correspondingly, perform feature encoding on the images in each co-salient image group to obtain encoded image features, including:

[0050] Input the images in each co-salient image group into an image encoder in sequence to perform feature encoding on each image in the co-salient image group and obtain encoded image features;

[0051] Among them, the text prompt encoder and the image encoder are located in the Segment Anything model.

[0052] Optionally, decode the encoded image features and the corresponding encoded text features to convert the co-salient text prompts into a binary mask, obtaining a historical co-saliency detection result, including:

[0053] Perform vector representations on the encoded image features and the corresponding encoded text features respectively to obtain a text prompt vector and an image encoding vector;

[0054] Combine the text prompt vector and the corresponding image encoding vector to obtain a combined vector;

[0055] Input the combined vector into a second decoder located in the Segment Anything model to decode the combined vector and obtain a historical co-saliency detection result including a co-salient mask.

[0056] In a second aspect, the present application discloses a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the foregoing co-saliency detection method is implemented.

[0057] In a third aspect, the present application discloses an electronic device, including a processor and a memory; wherein, when the processor executes the computer program stored in the memory, the foregoing co-saliency detection method is implemented.

[0058] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the foregoing co-saliency detection method is implemented.

[0059] It can be seen that in this application, at least two target images to be detected currently are first obtained to get the current image group; the current image group is input into the trained collaborative saliency detection model; the collaborative saliency detection model is a model obtained by training a pre-trained language model and a segment-anything model based on an image dataset containing collaborative saliency objects; the current collaborative saliency detection result corresponding to the current image group output by the collaborative saliency detection model is obtained; wherein, the current collaborative saliency detection result is a result obtained by extracting target text information related to the collaborative saliency between each target image in the current image group based on the pre-trained language model, and guiding the segment-anything model to segment the collaborative salient objects and the background in the current image group based on the target text information, and the current collaborative saliency detection result contains a mask result generated based on the target text information. In this application, a pre-trained language model and a segment-anything model are pre-trained based on an image dataset containing collaborative saliency objects to obtain a model for performing collaborative saliency detection. When collaborative saliency detection is required, the current image group to be detected is input into this model, so as to extract target text information related to the collaborative saliency between each target image in the current image group by using the pre-trained language model, and guide the segment-anything model to segment the collaborative salient objects and the background in the current image group based on the extracted target text information, so as to obtain a collaborative saliency detection result containing a mask result generated based on the target text information; this application uses a pre-trained language model and a segment-anything model, and can solve the problem of extracting collaborative salient information from a cross-modal perspective. This is because the pre-trained language model can perform multimodal interaction, establish cross-modal implicit semantic relationships, and has the ability to understand images and texts, so it can generate credible text scene descriptions for images, while the segment-anything model has excellent instance segmentation effects. Therefore, by combining the pre-trained language model and the segment-anything model, the problem of extracting difficult image collaborative relationships can be transformed into a simpler problem of extracting text collaborative relationships, and high-quality saliency-related text information can improve the ability to distinguish collaborative salient objects and interfering backgrounds in collaborative saliency detection. In addition, using high-quality saliency-related text information to guide the segment-anything model to segment collaborative salient objects can break through the limitation of using semantically irrelevant binary masks as supervision signals, which helps to alleviate the problem of limited representation learning caused by training scene limitations, thereby further improving the ability to identify collaborative salient objects and interfering backgrounds, and further enhancing the accuracy of collaborative saliency detection. Description of the Drawings

[0060] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained according to the provided accompanying drawings.

[0061] Figure 1 Flowchart of a collaborative saliency detection method disclosed in the present application;

[0062] Figure 2 Flowchart of a specific collaborative saliency detection method disclosed in the present application;

[0063] Figure 3 Schematic diagram of the training of a specific collaborative saliency detection model disclosed in the present application;

[0064] Figure 4 Schematic diagram of the structure of a specific multi-head self-attention disclosed in the present application;

[0065] Figure 5 Schematic diagram of the extraction of specific collaborative text prompt features disclosed in the present application;

[0066] Figure 6 Schematic diagram of the calculation process of specific keyword weights disclosed in the present application;

[0067] Figure 7 Schematic diagram of loss supervision during the training of a specific model disclosed in the present application;

[0068] Figure 8 Schematic diagram of the collaborative saliency detection result disclosed in the present application;

[0069] Figure 9 Schematic diagram of the structure of an electronic device disclosed in the present application. Detailed implementation manners

[0070] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0071] The embodiments of the present application disclose a collaborative saliency detection method. Refer to Figure 1 As shown, the method includes:

[0072] Step S11: Obtain at least two target images to be detected currently, and obtain the current image group.

[0073] In this embodiment, when collaborative saliency detection is required, at least two target images to be currently subjected to collaborative saliency detection are first obtained to obtain a current image group.

[0074] Step S12: Input the current image group into the trained collaborative saliency detection model; the collaborative saliency detection model is a model obtained by training a pre-trained language model and a Segment Anything Model (SAM) based on an image data set containing collaborative saliency objects.

[0075] In this embodiment, after obtaining the current image group to be currently detected, further, the images in the current image group are input into a collaborative saliency detection model obtained by training a pre-trained language model and a Segment Anything Model (SAM, a vision model based on Transformer that can take an image and a user's prompt as input and output a segmentation mask indicating the label of each pixel) based on an image data set containing collaborative saliency objects for collaborative saliency detection, that is, to detect the collaborative saliency objects in the current image group. Among them, the pre-trained language model can be a Multimodal Large Language Model (MLLM), such as models like mPlug, GPT (Generative Pre-trained Transformer, a language model based on artificial intelligence technology), etc., which are used to establish cross-modal implicit semantic relationships, have the ability to understand image-text, and can generate credible text scene descriptions for images; the SAM model combines prompt learning and can achieve cross-modal fusion of text-images and obtain excellent Instance Segmentation results.

[0076] It should be noted that the pre-trained language model has powerful implicit semantic knowledge, which is a deep understanding of the relationships between different objects in the image. Using this semantic knowledge in the collaborative saliency detection task is beneficial to breaking through the limitation of the binary mask irrelevant to semantics as the supervision signal; in addition, using the SAM model helps to alleviate the problem of limited representation learning caused by training scene limitations. In summary, performing collaborative saliency detection using the collaborative saliency detection model based on the pre-trained language model and the Segment Anything Model can learn high-quality semantic knowledge, and high-quality semantic knowledge has an important impact on the discrimination between collaborative salient objects and the background. Therefore, by using the SAM model and based on high-quality semantic knowledge, instance segmentation of the collaborative salient objects and interfering backgrounds in the image group can be performed, which can improve the accuracy of recognizing collaborative salient objects and backgrounds, thereby improving the effect of collaborative saliency detection.

[0077] Specifically, refer to Figure 2 As shown, the creation process of the co-saliency detection model may include:

[0078] Step S21: Obtain an image dataset containing co-saliency objects, and group the images in the image dataset according to different co-saliency objects to obtain multiple co-saliency image groups;

[0079] In this embodiment, first collect images containing co-saliency objects to obtain an image dataset, and then group the images in the image dataset according to different co-saliency objects to obtain multiple co-saliency image groups, where each group contains N images, and each co-saliency image group can be expressed as:

[0080] ;

[0081] In the formula, n is a variable, H represents height, W represents width, is the image in the co-saliency image group.

[0082] Step S22: Input the images in each co-saliency image group into the pre-trained language model in turn to extract features from each image in the co-saliency image group to obtain image encoding features, and extract features from the preset interactive text to obtain text encoding features;

[0083] In this embodiment, after obtaining multiple co-saliency image groups, input the images in each co-saliency image group into the pre-trained language model (such as the mPlug model) in turn for model training. When the mPlug model receives the input image, first extract features from each image in the co-saliency image group to obtain the image encoding features corresponding to each image. Then, in order to control the mPlug model to output the scene description desired by the user, an appropriate interactive text can be selected as a prompt input, and then extract features from the interactive text to obtain text encoding features.

[0084] Step S23: Decode the image encoding features and text encoding features corresponding to each image in the co-saliency image group to obtain a scene structure knowledge description text containing the mutual relationship between different co-saliency objects corresponding to each image;

[0085] In this embodiment, after extracting features from each image in the co-saliency image group and the preset interactive text through the pre-trained language model, further, decode the image encoding features and text encoding features corresponding to each image in the co-saliency image group to obtain a scene structure knowledge description text containing the mutual relationship between different co-saliency objects corresponding to each image in the co-saliency image group , the scene structure knowledge description texts corresponding to all co - significantly image groups can be expressed as .

[0086] Step S24: Encode each scene structure knowledge description text into a text feature respectively, and extract the keywords in the text feature to obtain the prompt keyword feature corresponding to each scene structure knowledge description text;

[0087] In this embodiment, each scene structure knowledge description text is encoded into a text feature respectively, and the keywords in each text feature are extracted to obtain each scene structure knowledge description text corresponding prompt keyword feature, thereby realizing the conversion of the scene structure knowledge description text in sentence units into the prompt keyword feature .

[0088] Step S25: Mine the connections between different prompt keyword features to generate a co - significant text prompt;

[0089] Next, mine the connections between different prompt keyword features, so as to generate a reliable co - significant text prompt .

[0090] Step S26: Input the co - significant text prompt into the Segment - Anything model to convert the co - significant text prompt into a binary mask, and obtain a co - significant detection model containing historical co - significant detection results.

[0091] Furthermore, input the co - significant text prompt into the Segment - Anything model (i.e., SAM) to convert the co - significant text prompt into a binary mask, and obtain a co - significant detection model containing historical co - significant detection results , where represents the co - significant result of any image in a single co - significant image group. That is, take the Segment - Anything model (i.e., SAM) as the bridge connecting the co - significant text prompt and the co - significant mask (i.e., binary mask), and convert the co - significant text prompt into a binary mask.

[0092] Specifically, for feature extraction of each image in the co-salient image group in step S22 to obtain image coding features, and for feature extraction of the preset interactive text to obtain text coding features, it may specifically include: using an image feature extractor to perform feature extraction on each image in the co-salient image group to obtain image coding features, and converting the image coding features into vector representations to obtain image feature vectors; using a text feature extractor to perform feature extraction on the preset interactive text to obtain text coding features, and converting the text coding features into vector representations to obtain text feature vectors; where the image feature extractor and the text feature extractor are located in the pre-trained language model. In this embodiment, referring to Figure 3 as shown, the image feature extractor in the pre-trained language model (mPlug model) can be used to perform feature extraction on the images in the co-salient image group to obtain image coding features, and then convert the image coding features into vector representations to obtain image feature vectors ; then, use the text feature extractor in the pre-trained language model (such as the mPlug model) to perform feature extraction on the preset interactive text, such as what is Saliency in this image? (What is salient in this image?) to obtain text coding features, and convert the text coding features into vector representations to obtain text feature vectors . By performing vector transformation on the text coding features, the interactive text can be converted from a computer-inaccessible form to a computer-accessible form. Specifically, the process of compressing the interactive text into a feature vector can be expressed as:

[0093] ;

[0094] In the formula, text represents the interactive text, is the text feature extractor.

[0095] Specifically, for decoding the image coding features and text coding features corresponding to each image in the co-salient image group in step S23 to obtain a scene structure knowledge description text containing the mutual relationships between different co-salient objects corresponding to each image, it may specifically include: combining the image feature vectors and text feature vectors corresponding to each image in the co-salient image group to obtain a combined feature vector; inputting the combined feature vector into the first decoder located in the pre-trained language model to decode the combined feature vector to obtain a scene structure knowledge description text containing the mutual relationships between different co-salient objects corresponding to each image. In this embodiment, first, the image feature vectors corresponding to each image in the co-salient image group and the text feature vectors Combine them to obtain the combined feature vector, and then input the combined feature vector into the decoder of the pre-trained language model to decode the combined feature vector, so as to obtain the scene structure knowledge description text D corresponding to each image, which contains the mutual relationship between different co-salient objects (i.e., Figure 3 the text description in

[0096] ;

[0097] In the formula, , I represents the input image, represents the image feature extractor, represents the feature decoder.

[0098] Specifically, the process of obtaining the scene structure knowledge description text of a single co-salient image group can be expressed as:

[0099] ;

[0100] In the formula, represents the large image understanding model, represents the input image, text represents the interactive text, represents the scene structure knowledge description text corresponding to the input image.

[0101] It should be noted that the scene structure knowledge description text not only covers the description of the salient object itself, but also includes the logical relationship between the salient object and the surrounding interfering background, and this relationship constitutes the scene structure knowledge.

[0102] Specifically, in step S24, each scene structure knowledge description text is encoded into text features respectively, and the keywords in the text features are extracted to obtain the prompt keyword features corresponding to each scene structure knowledge description text. Specifically, it can include: using the text encoder to encode each scene structure knowledge description text respectively to obtain text features, and performing linear mapping on the text features to obtain the mapped features; using the multi-head self-attention mechanism to enhance the features of the mapped features to obtain the enhanced features; using the keyword extractor to extract the keywords in the enhanced features to obtain multiple prompt keyword features corresponding to each scene structure knowledge description text. In this embodiment, the text encoder, such as the BERT (Bidirectional Encoder Representations from Transformers, a pre-trained language representation model based on the Transformer architecture) text encoder, can be used to encode each scene structure knowledge description text to encode the scene structure knowledge description text Convert it into a computer-recognizable coding form to obtain text features , text features Specifically, it can be expressed as:

[0103] ;

[0104] In the formula, is the pre-trained BERT text encoder.

[0105] Next, the obtained text features are linearly mapped to obtain , , , and then the multi-head self-attention mechanism (Multi-Head Self-Attention) is used to mine the relationship between words in a sentence to strengthen the text features , and the specific formula is:

[0106] ;

[0107] ;

[0108] ;

[0109] ;

[0110] In the formula, , , respectively represent three different linear layers of queries, keys, and values, represents multi-head self-attention.

[0111] In addition, as shown in Figure 4 shown, Figure 4 shows the structure of multi-head self-attention, including a self-attention module, a multi-layer perceptron (MLP, Multilayer Perceptron), an Add (feature fusion) layer, and a normalization layer (Normalization), and the input is Q, K, V. In this embodiment, the enhanced text features can also be enhanced. Specifically, the enhanced text features can be further passed through a fully connected layer, and the specific formula is as follows:

[0112] ;

[0113] ;

[0114] In the formula, represents the enhanced text features The intermediate state features after residual connection and layer normalization denotes layer normalization is the final enhanced feature after scene knowledge enhancement denotes the fully connected layer

[0115] Furthermore, the enhanced features are input into a keyword extractor to extract keywords, obtaining text descriptions of scene structure knowledge corresponding to multiple prompt keyword features. In this embodiment, the multi-head self-attention mechanism is used to enhance the text features after linear mapping, and the keyword extractor is used to extract keywords from the enhanced features, which can make the obtained prompt keyword features more accurate and is beneficial to improving the accuracy of real-time collaborative saliency detection

[0116] Specifically, mining the connections between different prompt keyword features in step S25 to generate collaborative saliency text prompt words can specifically include: using the co-attention mechanism to mine the connections between different prompt keyword features to obtain co-text prompt features; decoding the co-text prompt features to obtain collaborative saliency text prompt words. In this embodiment, as shown in Figure 5 , the co-attention mechanism can be used to obtain co-text prompt features for the connections between different prompt keyword features , and then the co-text prompt features are decoded by the BERT decoder to obtain collaborative saliency text prompt words

[0117] In a specific implementation, using the co-attention mechanism to mine the connections between different prompt keyword features to obtain co-text prompt features includes: combining multiple prompt keyword features corresponding to each scene structure knowledge description text to obtain a keyword group; inputting the keyword group into different linear layers respectively, and multiplying the output results of different linear layers to obtain the attention affinity graph of the keyword group; extracting the first number of values in the attention affinity graph by row and calculating the sum of the extracted first number of values to obtain the first sum value; calculating the ratio of the first sum value to the first number to obtain the keyword weight, and multiplying the keyword weight element-wise with the prompt keyword features in the keyword group to obtain multiple calculation results; extracting the second number of values in the multiple calculation results and calculating the sum of the extracted second number of values to obtain the second sum value; calculating the ratio of the second sum value to the second number to obtain the co-text prompt feature. In this embodiment, first, multiple prompt keyword features corresponding to each scene structure knowledge description text can be combined to obtain a keyword group T, and then the keyword group T is input into two different linear layers respectively, and the output results of the two linear layers are multiplied to obtain the attention affinity graph of the keyword group T. The specific calculation formula is

[0118] ;

[0119] In the formula, and represent the linear mapping functions corresponding to two different linear layers, which are used to linearly map the keyword features. represents the attention affinity graph, and k represents the total number of keyword features in the keyword group.

[0120] Next, as shown in Figure 6 , extract the first n maximum values in the attention affinity graph row by row, calculate the sum of the extracted first n maximum values to obtain the first sum value, and calculate the ratio of the first sum value to n to obtain the keyword weight, that is, sum the maximum values and take the average to obtain the keyword weight. The specific calculation formula is:

[0121] ;

[0122] In the formula, represents taking the first n maximum values in a vector of length j, represents outputting a vector with a length of 1 for the average value of a vector of length n, represents the keyword weight.

[0123] Furthermore, multiply the keyword weight element-wise with the hint keyword features in the keyword group T to obtain multiple calculation results, extract the first k maximum values from the multiple calculation results, calculate the sum of the extracted first k maximum values to obtain the second sum value, and then calculate the ratio of the second sum value to k to obtain the collaborative text hint feature , and the specific calculation formula is:

[0124] ;

[0125] In the formula, represents the keyword group, represents element-wise multiplication, represents the output collaborative text hint feature.

[0126] It should be noted that during the training of the pre-trained language model and the segment everything model, it can also include: using the real collaborative significant category text as a supervision signal to supervise the learning process of the keyword extractor and the collaborative attention mechanism. In this embodiment, in order to ensure the accuracy of the extraction of the collaborative significant text prompt words, during the training of the pre-trained language model and the segment everything model, the real collaborative significant category text can be used as a supervision signal to supervise the learning process of the keyword extractor and the collaborative attention mechanism.

[0127] In addition, during the training of the pre-trained language model and the Segment Anything Model, it can also include: obtaining co-salient text labels from the image dataset, performing high-dimensional mapping on the co-salient text labels to obtain true co-prompt features; using the distance between the true co-prompt features and the co-text prompt features as a loss, and supervising the loss. Refer to Figure 7 As shown, during the model training, the true co-prompt features obtained by compiling the true co-salient text labels through BERT can be used as the distance from the co-text prompt features as the loss for supervision. The specific calculation formula is:

[0128] ;

[0129] In the formula, represents the predicted co-text prompt features , represents the true co-prompt features obtained by mapping the true label through BERT, and k represents the dimension of the features.

[0130] In addition, for the convenience of subsequent calculation and processing, the co-text prompt features can be converted into actual text. Specifically, the co-text prompt features can be decoded into actual co-salient text prompt words through the BERT decoder. The specific calculation formula is:

[0131] ;

[0132] In the formula, is the BERT decoder.

[0133] In this embodiment, converting the co-salient text prompt words into a binary mask to obtain a co-salient detection model including historical co-salient detection results can specifically include: performing feature encoding on the co-salient text prompt words to obtain text feature encodings; performing feature encoding on the images in each co-salient image group to obtain image feature encodings; decoding the image feature encodings and the corresponding text feature encodings to convert the co-salient text prompt words into a binary mask to obtain historical co-salient detection results, so as to complete the training of the pre-trained language model and the Segment Anything Model and obtain the co-salient detection model. In this embodiment, first, the co-salient text prompt words Perform feature encoding to obtain text feature encoding, then perform feature encoding on the images in each co-salient image group to obtain image feature encoding, and then decode the image feature encoding and the corresponding text feature encoding, so as to convert the co-salient text prompt into a binary mask, and obtain the corresponding historical co-salient detection result, so as to complete the training of the entire co-salient detection model.

[0134] Specifically, performing feature encoding on the co-salient text prompt to obtain text feature encoding may include: inputting the co-salient text prompt into a text prompt encoder to perform feature encoding on the co-salient text prompt to obtain text feature encoding; correspondingly, performing feature encoding on the images in each co-salient image group to obtain image feature encoding, including: sequentially inputting the images in each co-salient image group into an image encoder to perform feature encoding on each image in the co-salient image group to obtain image feature encoding; wherein, the text prompt encoder and the image encoder are located in the Segment Everything model. In this embodiment, the co-salient text prompt can be input into the text prompt encoder located in the Segment Everything model to perform feature encoding on the co-salient text prompt to obtain text feature encoding. Then, the images in each co-salient image group are sequentially input into the image encoder located in the Segment Everything model, so as to perform feature encoding on each image in the co-salient image group to obtain image feature encoding.

[0135] Specifically, decoding the image feature encoding and the corresponding text feature encoding to convert the co-salient text prompt into a binary mask and obtain the historical co-salient detection result may include: respectively performing vector representation on the image feature encoding and the corresponding text feature encoding to obtain a text prompt vector and an image encoding vector; combining the text prompt vector and the corresponding image encoding vector to obtain a combined vector; inputting the combined vector into a second decoder located in the Segment Everything model to decode the combined vector to obtain the historical co-salient detection result including the co-salient mask. In this embodiment, the text feature encoding and the image feature encoding are respectively compressed into a feature vector by the text prompt encoder and the image encoder, so as to convert the co-salient text prompt from a form that cannot be processed by a computer into a form that can be processed by a computer, and obtain the corresponding text prompt vector and the image encoding vector ; then, combining the text prompt vector and the corresponding image encoding vector to obtain a combined vector, and inputting the combined vector into the decoder of the Segment Everything model to decode the combined vector to obtain a co-salient mask (i.e., a binary mask) with high-quality semantic guidance. The specific calculation formula is:

[0136] ;

[0137] In the formula, , ; among which, represents the input image, represents the image encoder, represents the text prompt encoder, is the feature decoder, and M is the mask result of a single image.

[0138] The mask results of all images in a single image group can be expressed as:

[0139] ;

[0140] In the formula, SAM represents the Segment Anything Model, represents the input image, where , represents the co-salient text prompt. Repeat the above model training process until co-salient detection has been performed on all co-salient image groups in step S21. All co-salient masks generated by guiding with the co-salient text prompt obtained through the SAM model can be post-processed to obtain the final historical co-salient detection result. .

[0141] Step S13: Obtain the current co-salient detection result corresponding to the current image group output by the co-salient detection model; among which, the current co-salient detection result is obtained by extracting target text information related to the co-salience between each target image in the current image group based on a pre-trained language model, and guiding the Segment Anything Model to segment the co-salient targets and the background in the current image group based on the target text information. The current co-salient detection result contains the mask result generated based on the target text information.

[0142] It should be noted that when the co-salient detection model performs co-salient detection on multiple target images in the current image group, specifically, it first extracts the text information related to the co-salience between each target image in the current image group to obtain the target text information, and then guides the Segment Anything Model to segment the co-salient targets and the interfering background in the current image group based on the target text information to obtain the current co-salient detection result; among which, the current co-salient detection result contains the mask result generated based on the target text information, that is, it contains the set of co-salient masks for each image in the current image group, and the co-salient mask is a binary mask.

[0143] Specifically, as shown in Figure 8 . Figure 8The results of performing collaborative saliency detection on three groups of images respectively using the collaborative saliency detection model proposed in this application are shown. From the results of collaborative saliency detection, it can be seen that the first group is an avocado, the second group is a bird, and the third group is an axe. It can be seen that this solution can accurately extract the semantic information of the target in a complex and variable background, and obtain a relatively accurate mask result under the guidance of high-quality semantic information, thereby improving the accuracy of collaborative saliency detection.

[0144] It can be seen that in the embodiment of the present application, at least two target images to be detected currently are first obtained to obtain the current image group; the current image group is input into the trained collaborative saliency detection model; the collaborative saliency detection model is a model obtained by training a pre-trained language model and a segment-anything model based on an image data set containing collaborative saliency objects; the current collaborative saliency detection result corresponding to the current image group output by the collaborative saliency detection model is obtained; wherein, the current collaborative saliency detection result is a result obtained by extracting target text information related to the collaborative saliency between each target image in the current image group based on the pre-trained language model, and guiding the segment-anything model to segment the collaborative saliency target and the background in the current image group based on the target text information, and the current collaborative saliency detection result contains a mask result generated based on the target text information. The present application pre-trains the pre-trained language model and the segment-anything model based on an image data set containing collaborative saliency objects to obtain a model for performing collaborative saliency detection. When collaborative saliency detection is required, the current image group to be detected is input into this model, so as to extract target text information related to the collaborative saliency between each target image in the current image group by using the pre-trained language model, and guide the segment-anything model to segment the collaborative saliency target and the background in the current image group based on the extracted target text information, so as to obtain a collaborative saliency detection result containing a mask result generated based on the target text information; the embodiment of the present application uses the pre-trained language model and the segment-anything model to solve the problem of extracting collaborative saliency information from a cross-modal perspective. This is because the pre-trained language model can perform multi-modal interaction, establish cross-modal implicit semantic relationships, and has the ability to understand images and texts. Therefore, it can generate credible text scene descriptions for images, while the segment-anything model has excellent instance segmentation effects. Therefore, through the combination of the pre-trained language model and the segment-anything model, the problem of extracting difficult image collaborative relationships can be transformed into a simpler problem of extracting text collaborative relationships, and high-quality saliency-related text information can improve the ability to distinguish collaborative saliency targets and interfering backgrounds in collaborative saliency detection. In addition, using high-quality saliency-related text information to guide the segment-anything model to segment the collaborative saliency target can break through the limitation of using a semantic-irrelevant binary mask as a supervision signal, which helps to alleviate the problem of limited representation learning caused by training scene limitations, thereby further improving the ability to identify collaborative saliency targets and interfering backgrounds, and further enhancing the accuracy of collaborative saliency detection.

[0145] Further, the embodiment of the present application also discloses an electronic device. Figure 9 It is a structural diagram of the electronic device 20 shown according to an exemplary embodiment, and the content in the figure should not be regarded as any limitation on the scope of use of the present application.

[0146] Figure 9 FIG. 1 is a schematic structural diagram of an electronic device 20 provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the collaborative saliency detection method disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0147] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed thereon here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application requirements, and no specific limitation is made here.

[0148] In addition, as a carrier for resource storage, the memory 22 may be a read-only memory, a random access memory, a disk, or an optical disc, etc., and the resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be temporary storage or permanent storage.

[0149] Among them, the operating system 221 is used to manage and control each hardware device and the computer program 222 on the electronic device 20, and it may be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the collaborative saliency detection method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 may further include a computer program that can be used to complete other specific tasks.

[0150] Further, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the collaborative saliency detection method disclosed above is implemented. For the specific steps of this method, reference may be made to the corresponding content disclosed in the foregoing embodiments, and details are not repeated here.

[0151] Further, an embodiment of the present application also discloses a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the collaborative saliency detection method disclosed above are implemented.

[0152] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the products disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.

[0153] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0154] The steps of the methods or algorithms described in combination with the embodiments disclosed in this article can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium known in the technical field.

[0155] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0156] The above has introduced in detail a collaborative saliency detection method, product, device and storage medium provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A collaborative saliency detection method, characterized in that: include: Acquire at least two target images to be detected to obtain a current image group; Inputting the current image group into a trained co-saliency detection model; the co-saliency detection model is a model obtained by training a pre-trained language model and a segmentation model based on an image dataset containing co-saliency objects; Acquire a current co-saliency detection result corresponding to the current image group output by the co-saliency detection model; wherein the current co-saliency detection result is a result obtained by extracting target text information related to the co-saliency between target images in the current image group based on the pre-trained language model, and guiding the segmentation model to segment the co-saliency target and background in the current image group based on the target text information, and the current co-saliency detection result includes a mask result generated based on the target text information; The method also includes: obtaining an image data set containing co-salient objects, and grouping the images in the image data set according to different co-salient objects to obtain multiple co-salient image groups; inputting the images in each of the co-salient image groups into a pre-trained language model in turn to extract features of each image in the co-salient image group to obtain image coding features, and extracting features of preset interactive text to obtain text coding features; decoding the image coding features and the text coding features corresponding to each image in the co-salient image group to obtain the scene structure knowledge description text containing the relationship between different co-salient objects corresponding to each image; encoding each of the scene structure knowledge description texts into text features respectively, and extracting keywords from the text features to obtain prompt keyword features corresponding to each of the scene structure knowledge description texts; mining the connection between different prompt keyword features to generate co-salient text prompt words; inputting the co-salient text prompt words into a segmentation model to convert the co-salient text prompt words into binary masks to obtain the co-salient detection model containing historical co-salient detection results.

2. The collaborative saliency detection method according to claim 1, characterized in that: The extracting features of each image in the collaborative salient image group to obtain image coding features, and extracting features of the preset interactive text to obtain text coding features, includes: Using an image feature extractor to extract features from each image in the collaborative salient image group to obtain image coding features, and converting the image coding features into vector representations to obtain image feature vectors; Using a text feature extractor to extract features from a preset interactive text to obtain text encoding features, and converting the text encoding features into vector representations to obtain text feature vectors; Wherein, the image feature extractor and the text feature extractor are located in the pre-trained language model.

3. The collaborative saliency detection method according to claim 2, characterized in that: The decoding of the image encoding features and the text encoding features corresponding to each image in the co-salient image group to obtain the scene structure knowledge description text corresponding to each image containing the relationship between different co-salient objects includes: Combining the image feature vector and the text feature vector corresponding to each image in the collaborative salient image group to obtain a combined feature vector; The combined feature vector is input into a first decoder located in the pre-trained language model to decode the combined feature vector to obtain a scene structure knowledge description text corresponding to each image containing the relationship between different co-saliency objects.

4. The collaborative saliency detection method according to claim 1, characterized in that: The encoding of each of the scene structure knowledge description texts into text features and extracting keywords from the text features to obtain prompt keyword features corresponding to each of the scene structure knowledge description texts includes: Using a text encoder to encode each of the scene structure knowledge description texts to obtain text features, and linearly mapping the text features to obtain mapped features; Using a multi-head self-attention mechanism to enhance the mapped features to obtain enhanced features; A keyword extractor is used to extract keywords from the enhanced features to obtain a plurality of prompt keyword features corresponding to each of the scene structure knowledge description texts.

5. The collaborative saliency detection method according to claim 4, characterized in that: The mining of the connection between the different prompt keyword features to generate collaboratively significant text prompt words includes: The connection between the different prompt keyword features is mined using the collaborative attention mechanism to obtain collaborative text prompt features; The collaborative text prompt feature is decoded to obtain a collaborative salient text prompt word.

6. The collaborative saliency detection method according to claim 5, characterized in that: The collaborative attention mechanism is used to mine the connections between the different prompt keyword features to obtain collaborative text prompt features, including: Combining multiple prompt keyword features corresponding to each of the scene structure knowledge description texts to obtain a keyword group; Inputting the keyword groups into different linear layers respectively, and multiplying the output results of different linear layers to obtain an attention affinity graph of the keyword group; Extracting a first number of values ​​in the attention affinity graph row by row, and calculating a sum of the first number of values ​​extracted to obtain a first sum; Calculating a ratio of the first sum value to the first number to obtain a keyword weight, and multiplying the keyword weight by the prompt keyword feature in the keyword group element by element to obtain a plurality of calculation results; Extracting first second number of values ​​from the plurality of calculation results, and calculating the sum of the extracted second number of values ​​to obtain a second sum; The ratio of the second sum value to the second quantity is calculated to obtain a collaborative text prompt feature.

7. The collaborative saliency detection method according to claim 5, characterized in that: In the process of training the pre-trained language model and the segmentation model, it also includes: The real co-salient category text is used as a supervision signal to supervise the learning process of the keyword extractor and the co-attention mechanism.

8. The collaborative saliency detection method according to claim 5, characterized in that: In the process of training the pre-trained language model and the segmentation model, it also includes: Acquire collaborative salient text labels from the image dataset, and perform high-dimensional mapping on the collaborative salient text labels to obtain true collaborative prompt features; The distance between the true collaborative prompt feature and the collaborative text prompt feature is used as a loss, and the loss is supervised.

9. The collaborative saliency detection method according to any one of claims 1 to 8, characterized in that: The converting the co-salient text prompt words into binary masks to obtain the co-salient detection model including historical co-salient detection results includes: Performing feature coding on the collaboratively significant text prompt words to obtain text feature coding; Performing feature coding on the images in each of the collaborative salient image groups to obtain image feature coding; The image feature code and the corresponding text feature code are decoded to convert the co-salient text prompt word into a binary mask to obtain a historical co-saliency detection result, so as to complete the training of the pre-trained language model and the segmentation model and obtain the co-saliency detection model.

10. The collaborative saliency detection method according to claim 9, characterized in that: The step of performing feature coding on the collaborative salient text prompt words to obtain text feature coding includes: Inputting the collaboratively significant text prompt words into a text prompt encoder to perform feature encoding on the collaboratively significant text prompt words to obtain text feature encoding; Accordingly, the performing feature coding on the images in each of the collaborative salient image groups to obtain image feature coding includes: Inputting the images in each of the co-salient image groups into an image encoder in sequence to perform feature encoding on each of the images in the co-salient image groups to obtain image feature encoding; Wherein, the text prompt encoder and the image encoder are located in the segmentation model.

11. The collaborative saliency detection method according to claim 10, characterized in that: The decoding of the image feature code and the corresponding text feature code to convert the co-salient text prompt word into a binary mask to obtain a historical co-salient detection result includes: Respectively performing vector representation on the image feature code and the corresponding text feature code to obtain a text prompt vector and an image code vector; Combining the text prompt vector and the corresponding image encoding vector to obtain a combined vector; The combined vector is input into a second decoder located in the segmentation model to decode the combined vector to obtain a historical co-saliency detection result including a co-saliency mask.

12. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the collaborative saliency detection method according to any one of claims 1 to 11 is implemented.

13. An electronic device, characterized in that: The method comprises a processor and a memory; wherein, when the processor executes the computer program stored in the memory, the collaborative saliency detection method according to any one of claims 1 to 11 is implemented.

14. A computer-readable storage medium, characterized in that: Used to store a computer program; wherein, when the computer program is executed by a processor, the collaborative saliency detection method according to any one of claims 1 to 11 is implemented.

Citation Information

Patent Citations

  • Zero sample image segmentation model training method and device based on multiple modes

    CN117788981A

  • Infrared small target detection method based on scene text information guidance

    CN118762364A