A zero-shot combined image retrieval method based on symmetric inversion network

By employing self-supervised learning and implicit triplet training through symmetric inversion networks, the problems of insufficient data and training discrepancies in zero-shot combination image retrieval are solved, thereby improving the overall performance of image retrieval and making it suitable for multimodal retrieval scenarios.

CN120492651BActive Publication Date: 2025-10-28DATA INTELLIGENCE HUINENG (DALIAN) TECH DEV CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510585444.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-10-28
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

Existing zero-shot combined image retrieval methods suffer from performance limitations due to insufficient data and training discrepancies, making it difficult to effectively improve the overall performance of image retrieval.

Method used

A method based on symmetric inversion networks is adopted. Image inversion networks and text inversion networks are trained through self-supervised learning to construct implicit triples. A combined network is trained using large-scale image-text pair data, including preprocessing, image-text scrambling, symmetric inversion network training and combined network fusion, to achieve effective fusion of image and text information.

Benefits of technology

It improves the overall performance of zero-shot combined image retrieval, alleviates the problems of insufficient data and training discrepancies by constructing implicit triples at low cost, and achieves a high balance between performance and efficiency, making it suitable for multimodal retrieval scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492651B_ABST
    Figure CN120492651B_ABST
Patent Text Reader

Abstract

This invention provides a zero-shot combined image retrieval method based on a symmetric inversion network, comprising the following steps: collecting image-text pair data containing both image and text information, and preprocessing the image-text pair data; selecting arbitrary portions of the image-text pair data and performing image-text scrambling to make the image information and text information in the selected image-text pair data mismatched; training a symmetric inversion network based on self-supervised learning, constructing implicit triples by selecting image and text information from target image features, original image-text features, and fused image-text features; training a combined network based on the implicit triples, the combined network including a projection layer that projects image and text information into a unified embedding space and a fusion layer that fuses the projected information; inputting the image to be retrieved, obtaining combined features through the symmetric inversion network and the combined network, and outputting the retrieval results according to similarity values. This invention addresses the problems of insufficient data and training discrepancies in zero-shot combined image retrieval, thereby improving the overall performance of image retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image retrieval technology, and in particular to a zero-sample combined image retrieval method based on a symmetric inversion network. Background Technology

[0002] In recent years, with the rapid development of digital technology, users can easily build large image databases using smartphones or personal computers. However, with the rapid growth in data volume and data types, efficiently retrieving images that match user intent from large databases has become increasingly difficult, posing new challenges to image retrieval algorithms. Single-modal retrieval methods cannot fully capture user intent. Combining multiple modalities for retrieval better meets user needs and future trends. Therefore, some researchers have turned their attention to multimodal retrieval, with combined image retrieval being one such approach.

[0003] Combinatorial image retrieval aims to combine a reference image and modified text to form a set of multimodal queries, retrieving target images from image databases that are visually similar to the reference image and also contain modified text content. However, the difficulty in collecting multimodal data, along with its small scale and poor diversity, has become the biggest challenge restricting the development of combinatorial image retrieval. To address this challenge, researchers have proposed the zero-shot combinatorial image retrieval task. This involves not directly training the model using publicly available combinatorial image retrieval datasets, but instead leveraging the capabilities of pre-trained large models in conjunction with training on some proxy tasks to indirectly enable the model to acquire combinatorial image retrieval capabilities.

[0004] To date, zero-shot combined image retrieval methods can be broadly categorized into three types: (1) Text-inversion-based methods. These methods utilize inversion networks to project image features into the text embedding space, transforming the combined image retrieval task into a text-based retrieval task. However, due to significant differences in data and tasks between training and testing, these methods generally exhibit poor performance.

[0005] In summary, the technical problem that this invention actually solves is how to improve the overall performance of image retrieval. Summary of the Invention

[0006] To overcome the aforementioned technical shortcomings, the present invention aims to provide a zero-shot combined image retrieval method based on a symmetric inversion network. This method utilizes a large amount of existing image-text pair data to construct implicit triples at a lower cost for training the combined network. This alleviates the problems of insufficient data and training discrepancies in zero-shot combined image retrieval, thereby improving the overall performance of image retrieval.

[0007] This invention discloses a zero-shot combined image retrieval method based on a symmetric inversion network, comprising the following steps:

[0008] Step S100: Collect image-text pairs containing image and text information, and preprocess the image-text pairs data;

[0009] Step S200: Select any part of the image and text pair and perform image and text disorder processing on the data so that the image information and text information in the selected image and text pair data do not match;

[0010] Step S300: Train a symmetric inversion network based on self-supervised learning. The symmetric inversion network includes an image inversion network and a text inversion network with the same structure. Both the image inversion network and the text inversion network consist of multiple layers of modules with the same structure. The structure of each module layer is as follows:

[0011] MLP(x) = ReLU(Dropout(FC(x)))

[0012] Where x refers to the input feature, FC refers to the fully connected layer, Dropout refers to the temporary de-extension layer, and ReLU refers to the corrected linear unit; image-text pairs with mismatched image and text information are input into the image inversion branch and the text inversion branch respectively to obtain different fusion features in order to train the symmetric inversion network;

[0013] Step S400: Select any two pairs of image-text pairs and exchange the text information in the two pairs of image-text pairs to form exchanged image-text pairs. Input one of the exchanged image-text pairs into the image inversion network to obtain the target image features. Input the other exchanged image-text pair into the symmetric inversion network to obtain the original image-text features and the fused image-text features. Select image information and text information from the target image features, the original image-text features, and the fused image-text features to construct implicit triples.

[0014] Step S500: Train a fusion network based on implicit triples. The fusion network includes a projection layer that projects image information and text information into a unified embedding space and a fusion layer that fuses the projected information. The fusion layer includes a direct fusion part and a weighted combination part.

[0015] Step S600: Input the image to be retrieved and obtain combined features by passing it through a symmetric inversion network and a combined network in sequence. Then, perform cosine similarity matching on the candidate images extracted by the image encoder and output the retrieval results according to the similarity values.

[0016] Preferably, the acquired image-text pair data is large-scale image-text pair data.

[0017] Preferably, in step S300, the backbone network is a model that can achieve image encoder alignment.

[0018] Preferably, the image inversion network and the text inversion network are preset values, and the preset values ​​of the image inversion network and the text inversion network are ≥2.

[0019] Preferably, the projection layer includes normalization, full connection, activation, and random deactivation processing of image information and / or text information.

[0020] Preferably, the direct fusion portion of the fusion layer performs at least two fully connected, activation, and random deactivation operations on the image information and / or text information.

[0021] Preferably, the weighted combination part of the fusion layer generates weights based on image information and text information, and performs a weighted combination of image information and text information.

[0022] Preferably, the results of processing image information and / or text in the direct fusion part and the weighted combination part of the fusion layer are normalized to obtain the overall fusion result.

[0023] Preferably, the preprocessing of the image-text pair data includes resizing and standardizing the image information in the image-text pair data, and segmenting and embedding the text information in the image-text pair data.

[0024] Compared with existing technologies, the advantages of this invention, achieved by adopting the above technical solution, lie in its ability to construct implicit triples at a lower cost by utilizing a large amount of existing image-text pair data for training the ensemble network. This alleviates the problems of insufficient data and training discrepancies in zero-shot ensemble image retrieval, thereby improving the overall performance of image retrieval. Furthermore, since no additional generative model is used to generate data, and no large model is employed to process the image-text pair data, it achieves a high balance between performance and efficiency, and allows for better market application scenarios with lower deployment costs. Attached Figure Description

[0025] Figure 1 This is a schematic diagram illustrating the steps of a zero-sample combined image retrieval method based on a symmetric inversion network according to the present invention. Detailed Implementation

[0026] The advantages of the present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments.

[0027] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0028] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. The singular forms “a,” “the,” and “the” as used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0029] It should be understood that although the terms first, second, third, etc., may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0030] In the description of this invention, it should be understood that the terms "longitudinal", "lateral", "up", "down", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.

[0031] In the description of the present invention, unless otherwise specified and limited, it should be noted that the terms "installed", "connected" and "connected" should be understood in a broad sense. For example, it can be a mechanical connection or an electrical connection, or it can be the internal communication between two components. It can be a direct connection or an indirect connection through an intermediate medium. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to the specific circumstances.

[0032] In the following description, suffixes such as "module," "part," or "unit" used to denote elements are used only for the convenience of the description of the invention and have no specific meaning in themselves. Therefore, "module" and "part" can be used interchangeably.

[0033] This invention discloses a zero-shot combined image retrieval method based on a symmetric inversion network, comprising the following steps: Step S100: Collect image-text pairs containing image and text information, and preprocess the image-text pairs; Step S200: Select any part of the image-text pairs and perform image-text mismatch processing, so that the image information and text information in the selected image-text pairs do not match; Step S300: Train a symmetric inversion network based on self-supervised learning, the symmetric inversion network including an image inversion network and a text inversion network with the same structure, both the image inversion network and the text inversion network consisting of multiple layers of modules with the same structure, each layer having the structure: MLP(x) = ReLU(Dropout(FC(x))), where x refers to the input feature, FC refers to the fully connected layer, Dropout refers to the temporary de-extension layer, and ReLU refers to the corrected linear unit; Input the image-text pairs with mismatched image and text information into the image inversion branch and the text inversion branch respectively to obtain Different fusion features are used to train the symmetric inversion network; Step S400: Select any two pairs of image-text pairs and exchange the text information in the two pairs of image-text pairs to form exchanged image-text pairs. Input one of the exchanged image-text pairs into the image inversion network to obtain the target image features, and input the other exchanged image-text pair into the symmetric inversion network to obtain the original image-text features and the fused image-text features. Select image information and text information from the target image features, original image-text features, and fused image-text features to construct implicit triples; Step S500: Train the combination network based on the implicit triples. The combination network includes a projection layer that projects image information and text information into a unified embedding space and a fusion layer that fuses the projected information. The fusion layer includes a direct fusion part and a weighted combination part; Step S600: Input the image to be retrieved and sequentially pass it through the symmetric inversion network and the combination network to obtain the combined features. Perform cosine similarity matching on the candidate images extracted by the image encoder and output the retrieval results according to the similarity values.

[0034] See Figure 1 As shown in this embodiment, a zero-shot combined image retrieval method based on a symmetric inversion network will be described in detail. The retrieval method specifically includes the following steps:

[0035] Step S100: Collect image-text pair data containing image and text information, and preprocess the image-text pair data. A large amount of open-source, large-scale image-text pair data exists online, such as CC3M and CC12M. The corresponding data is downloaded using a provided script to obtain image-text pair data containing both image and text information. Preprocessing includes checking the format of the acquired image-text pair data and determining the appropriate backbone model based on the format. Preprocessing is then performed according to the format of the image-text pair data and the input requirements of the backbone model. In this embodiment, the preprocessing process will be illustrated using CC3M as the image-text pair data and CLIP ViT-B / 32 as the backbone model. CC3M originates from image-text pair data on the internet. After downloading via a script, a large number of text files containing image and text information are obtained. The script processes these text files, matching them one by one and storing them in a CSV file. The CLIP ViT-B / 32 backbone model has its own image and text information preprocessing pipeline. For image information, the default process involves first randomly cropping the image within a scale range of 0.9 to 1.0, then adjusting the image to a uniform width and height of 224 pixels, and finally applying ImageNet's regularization parameters for regularization. For text information, the core model's corresponding word segmenter and word embedding matrix are used for word segmentation and embedding.

[0036] Step S200: Select any portion of the image-text pair data and shuffle it to make the image and text information mismatched. This step creates training data to allow the backbone model to learn the potential relationship between image and text information, enabling it to better understand the feature representations in the image and text information and how to associate them. For example, consider a set of image-text pairs where the image is a dog in a park and the text is "beautiful flowers." Shuffling these pairs creates a mismatched combination. During training, the backbone model attempts to learn the characteristics of the dog image and the "beautiful flowers" text from this mismatch, laying the foundation for subsequent image-text matching.

[0037] Step S300: Train a symmetric inversion network based on self-supervised learning. The symmetric inversion network includes an image inversion network and a text inversion network with identical structures. Both the image inversion network and the text inversion network consist of multiple layers of modules with the same structure. The structure of each module is: MLP(x) = ReLU(Dropout(FC(x))), where x refers to the input feature, FC refers to the fully connected layer, Dropout refers to the temporary de-extension layer, and ReLU refers to the corrected linear unit. Image-text pairs that do not match image information and text information are input into the image inversion branch and the text inversion branch respectively to obtain different fusion features to train the symmetric inversion network. The symmetric inversion network is one of the core components of this invention. Self-supervised learning is a training method that allows the backbone model to discover patterns from image-text pairs without requiring a large amount of manually labeled data. The image inversion network and the text inversion network have the same structure, consisting of multiple layers of stacked MLP modules. This structure can effectively convert the features of one modality (image information or text information) into the embedding representation of another modality. Taking mismatched image-text pairs (the image of a dog and "beautiful flowers" mentioned above) as an example, the image information is processed by an image encoder to obtain features, which are then transformed by an image inversion network. The text information is processed by a text encoder to obtain features, which are then transformed by a text inversion network, resulting in two different fused features. Contrastive learning calculates the cosine similarity between these two fused features and trains the backbone model with the goal of reducing the distance between them, allowing the backbone model to learn how to better associate image features and text information.

[0038] It should be noted that, assuming image information is represented by i, text information by t, the corresponding image set by I, and the text set by T, the CLIP ViT-B / 32 backbone network contains an image encoder ψ. i and text encoder ψ t Given an image information i, use an image encoder to encode it into a feature v = ψ i (i)∈R d Where d is the embedding dimension of the CLIP ViT-B / 32 backbone network. Given a piece of text information t, it is also encoded into a feature u = ψ. ti (t)∈R d The image inversion network and the text inversion network are denoted as follows: and It adopts the same structure, consisting of multiple identical layers, each denoted as MLP, with the specific structure as follows:

[0039] MLP(x) = ReLU(Dropout(FC(x)))

[0040] Where x represents the input feature, FC represents a fully connected layer, Dropout represents a temporary de-extension layer, and ReLU represents a corrected linear unit. After inversion, the embedding of another modality is obtained:

[0041]

[0042] Where L is the number of stacked MLP layers, which is 2 in this embodiment.

[0043] Symmetric inversion networks project features encoded by a backbone network into the embedding space of another modality, and then perform pre-fusion between the embedding space and the other modality to achieve interactive fusion of multimodal features. The fused features are primarily composed of information from the uninverted modality, while also incorporating information from both modalities. By inputting a pair of mismatched image and text information into the image inversion branch and the text inversion branch, respectively, two different fused features are obtained (the image inversion branch yields fused features primarily composed of image information, and the text inversion branch yields fused features primarily composed of text information). These two different fused features supervise each other, and through comparative learning, the distance between them is narrowed, thereby training the symmetric inversion network.

[0044] Step S400: Select any two pairs of image-text pairs and exchange the text information in the two pairs to form exchanged image-text pairs. Input one of the exchanged image-text pairs into an image inversion network to obtain target image features, and input the other exchanged image-text pair into a symmetric inversion network to obtain original image-text features and fused image-text features. Select image information and text information from the target image features, original image-text features, and fused image-text features to construct implicit triples. In this step S400, two pairs of image-text pairs are selected, and the image information or text information in the two pairs of image-text pairs is exchanged to obtain two pairs of mismatched text pairs. Input one set of mismatched image-text pairs into the image inversion branch to obtain target image features, and input the other set of mismatched image-text pairs into the symmetric inversion network to obtain original image-text features and fused image-text features. Select one set of image information and text information from these features as a reference to construct implicit triples.

[0045] For example, in some embodiments, there are two pairs of image-text pairs. One pair contains image information A (a cat on a sofa) and text information a (feature description of the cat), while the other pair contains image information B (an apple on a table) and text information b (feature description of the apple). After exchanging the text information, we obtain image-text pairs A+b and B+a. The A+b pair is input into the image inversion branch to obtain target image features, and the B+a pair is input into the symmetric inversion network to obtain various features. Appropriate image and text information are selected from these to construct implicit triples. These implicit triples contain reference image information, modified text information, and target image information, and are used to train the ensemble network, enabling it to learn how to find the target image based on the reference image information and modified text information.

[0046] More specifically, randomly select two pairs of image-text data pairs i p ,t p and i k ,t k They exchange their text information and reorganize it into i p ,t k and i k ,t p . will i p ,t k Inputting the data into a symmetric inversion network yields two image information v. p , and two text messages u k , Any combination of image and text information can be used to form a tuple. Let i... k ,t p Inputting the image into an image inversion network can also yield a fused image feature. use As features of the target image, the above operations can construct four different implicit triples. and Experimental verification has shown that this embodiment uses one of the following. and Train the combinatorial network. After implicit triple construction, the following is obtained: and Will and Inputting each into a shared-weight ensemble network yields two combined features, c1 and c2:

[0047]

[0048] Here, π refers to the combined network, and the weighted summation of these networks yields the final combined feature c = λc1 + (1-λ)c2, where λ is a manually set hyperparameter. The learning of each combined feature is supervised, and the loss function used is the cross-entropy loss function L. CE and KL divergence loss function L KL .

[0049] The loss function for the two-stage process is:

[0050] L stage2 =β1lc1 + β2lc2 + β3lc,

[0051]

[0052] β1, β2, and β3 are weighting coefficients.

[0053] Step S500: Train a fusion network based on implicit triples. The fusion network includes a projection layer that projects image information and text information into a unified embedding space, and a fusion layer that fuses the projected information. The fusion layer includes a direct fusion part and a weighted combination part. In this step S500, the fusion network is trained based on the implicit triples created in step S400 above. The fusion network is responsible for fusing reference image information and modified text information to predict the target image, much like translating information from different languages ​​into the same language to facilitate subsequent fusion. The direct fusion part and the weighted combination part of the fusion layer fuse from different perspectives. The direct fusion part fuses image information and text information through a linear layer, while the weighted combination part generates weights based on the main information of the image information and text information for fusion. In some embodiments, when processing an image retrieval task that requires finding images with similar scenery to the reference image and containing the text description "with a river," the fusion network projects the information of this image retrieval task into the same space through the projection layer, and the fusion is performed through the fusion layer. By training with two sets of implicit triples and weighting the fusion results, the performance of the CLIP ViT-B / 32 backbone model can be further optimized, thereby improving the accuracy of retrieval.

[0054] It should be noted that the fusion network consists of a projection layer P and a fusion layer φ. The projection layer projects image and text information into the same embedding space, thereby better generating fusion weights and performing fusion. The image and text information components have the same structure, which can be described as follows:

[0055] P(x)=Dropout(ReLU(FC(Norm(x))))

[0056] The projection results corresponding to the image information and text information are as follows:

[0057] vp =P(v)

[0058] u p =P(u)

[0059] Norm refers to L2 normalization, used to ensure that the input image and text information are scaled uniformly. The projection results are concatenated to obtain the joint information w = v. p ,u p The joint information is used as input to the fusion layer. The fusion layer can be further divided into a direct fusion part H and a weighted combination part Ω. The direct fusion part fuses image information and text information through a linear layer:

[0060] H(w)=FC(Dropout(ReLU(FC(w))))

[0061] The weighted combination part first generates weights σ, and then uses σ to weight the original features:

[0062] σ=Sigmoid(FC(dROPOUT(ReLU(FC(w))))

[0063] Ω(v,u,w)=σ⊙Norm(u)+(1-σ)⊙Norm(v)

[0064] Where Sigmoid(.) refers to the Sigmoid function, used to map the neural network's predictions to the range 0 to 1, ⊙ indicating multiplication. The fusion result is expressed as:

[0065] φv,u,w)=Norm(H(w)+Ω(v,u,w)

[0066] To improve the model's performance, after experimentation, two sets of implicit triples were used for training, and the fusion results were weighted using the hyperparameter λ.

[0067] Step S600: The input image to be retrieved is sequentially processed through a symmetric inversion network and a fusion network to obtain combined features. Cosine similarity matching is then performed on candidate images extracted by the image encoder, and the retrieval results are output according to the similarity values. In practical applications, when a user inputs reference image and text information, the trained CLIP ViT-B / 32 backbone model begins operation. The symmetric inversion network first performs preliminary processing on the input information to extract relevant information. The fusion network further fuses this information to obtain combined information. Then, the cosine similarity between the combined information and the candidate images is calculated to determine the degree of similarity between them. In some embodiments, if a user wants to find an image similar to a snow-capped mountain photo that contains the description "has a lake," the CLIP ViT-B / 32 backbone model will input the snow-capped mountain photo and the text "has a lake," obtain combined information, calculate the cosine similarity with the features of candidate images in the database, and output images with high similarity first, thus realizing the image retrieval function.

[0068] The acquired image-text pair data is a large-scale image-text pair data.

[0069] Large-scale open-source image-text pair datasets, such as CC3M and CC12M, offer advantages in terms of large data volume and rich diversity. This abundant data allows the CLIP ViT-B / 32 backbone model to learn a wider range of image and text information relationships, improving the model's generalization ability. For example, the CC3M dataset contains a diverse range of images, from everyday scenes to natural landscapes, animals, and architecture, while the text descriptions cover different language styles and content. Using such data for training, the CLIP ViT-B / 32 backbone model can encounter more image-text combination patterns, enabling it to better handle various situations when faced with new retrieval tasks, thus improving retrieval accuracy and adaptability.

[0070] In step S300, a backbone network is used as a model that can achieve image-text encoder alignment.

[0071] The CLIP ViT-B / 32 or CLIP ViT-L / 14 backbone models, with their image-text encoder-aligned interfaces, offer the advantage of encoding image and text information into vectors with good alignment in the feature space. Taking the CLIP ViT-B / 32 backbone model as an example, its image and text encoders can encode image and text information into feature vectors of the same dimension, and these vectors have a certain semantic correspondence. Thus, during the training of the symmetric inversion network, image and text information can be more effectively converted and fused because their representations in the feature space are similar and comparable. This helps the CLIP ViT-B / 32 backbone model learn more accurate image-text conversion relationships, improving its performance.

[0072] The image inversion network and text inversion network are preset values, and the preset values ​​of the image inversion network and text inversion network are ≥2.

[0073] The preset number of module layers is a parameter obtained through experimental verification. Taking a layer count of 2 as an example, when the number of module layers is small, the network complexity is low, and it may not be able to fully learn the complex transformation relationships between image and text information; while when the number of layers is too large, the network is prone to overfitting, that is, the model performs well on the training data, but performs poorly on new data. The 2-layer setting performs well in balancing model complexity and learning ability. It allows the network to learn sufficient feature transformation information while avoiding overfitting, ensuring that the CLIP ViT-B / 32 backbone model has good generalization performance on different datasets, thereby effectively improving the training effect and overall performance of the symmetric inversion network.

[0074] The projection layer includes normalization, full connection, activation, and random deactivation processing of image and / or text information.

[0075] Normalization in the projection layer ensures that input information has a uniform scale, preventing excessively large or small values ​​from negatively impacting subsequent calculations and guaranteeing the stability of the CLIP ViT-B / 32 backbone model training. Fully connected layers perform linear transformations on input features, uncovering latent linear relationships between them. Activation functions (such as ReLU) introduce non-linearity, allowing the CLIP ViT-B / 32 backbone model to learn more complex non-linear relationships and enhancing its expressive power. Dropout randomly discards the outputs of some neurons during training, preventing overfitting and making the model more robust. For example, for input image and text information, the projection layer first normalizes them, then performs a linear transformation through a fully connected layer, followed by a ReLU activation function to add non-linearity, and finally uses dropout to prevent overfitting, projecting the processed information into a unified embedding space, preparing it for the fusion operations of subsequent fusion layers.

[0076] The direct fusion portion of the fusion layer performs at least two fully connected, activation, and random deactivation operations on the image and / or text information.

[0077] The direct fusion part of the fusion layer, through two fully connected, activation, and random deactivation operations, can more deeply mine the information in the joint data. The first fully connected operation performs an initial linear combination of the joint features, and the activation function (such as ReLU) gives it non-linear expressive power. Random deactivation prevents overfitting. The second fully connected operation further linearly transforms the information processed in the first operation, and the non-linearity is enhanced again by the activation function. Random deactivation ensures the robustness of the CLIP ViT-B / 32 backbone model. For example, when the ensemble network processes joint features containing a person's image and descriptive text, after these two processes, the direct fusion part can better extract and fuse key information from the image and text information, providing richer and more valuable feature representations for the subsequent weighted combination part, thereby improving the ensemble network's fusion effect and retrieval performance for multimodal information.

[0078] The weighted combination part of the fusion layer generates weights based on image and text information, and performs a weighted combination of image and text information.

[0079] The weighted combination component dynamically adjusts the proportions of image and text information in the fusion result based on the importance of the input information. By generating weights (e.g., using the sigmoid function combined with fully connected layers to generate weights between 0 and 1), the CLIP ViT-B / 32 backbone model can automatically determine the relative importance of image and text features according to different retrieval tasks. For example, in a retrieval task, if the text description contains crucial information (such as a specific object name), the weights generated by the weighted combination component will give the text information a larger proportion in the fusion result, highlighting the guiding role of the text information in the retrieval; if image features (such as unique image colors, shapes, etc.) are more important, the weights will make the image features dominate in the fusion result, thereby improving the matching degree between the retrieval results and user needs.

[0080] The results of processing image information and / or text in the direct fusion part and the weighted combination part of the fusion layer are normalized to obtain the overall fusion result.

[0081] Adding the results of the direct fusion and weighted combination allows for the integration of the advantages of both fusion methods, resulting in more comprehensive fusion information. Normalization, on the other hand, ensures that the fusion results have a uniform scale, facilitating subsequent calculations and comparisons. For example, when calculating the cosine similarity between the fusion information and candidate image features, normalized fusion information guarantees the accuracy and comparability of the calculation results. Without normalization, the potentially different scales of the direct fusion and weighted combination results can lead to deviations in the cosine similarity calculation, affecting the accuracy of the retrieval results. Normalization allows the fusion information of different samples to be compared under the same standard, improving the reliability and stability of the retrieval process.

[0082] Preprocessing of image-text pairs includes resizing and standardizing the image information, and segmenting and embedding the text information.

[0083] Preprocessing includes resizing image information to fit the input of the backbone network and normalizing it; for text information, it involves word segmentation and embedding.

[0084] It should be noted that the embodiments of the present invention have better implementability and are not intended to limit the present invention in any way. Any person skilled in the art may use the above-disclosed technical content to change or modify it into equivalent effective embodiments. However, any modifications or equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of the technical solution of the present invention.

Claims

1. A zero-sample combined image retrieval method based on a symmetric inversion network, characterized in that, Including the following steps: Step S100: Collect image-text pairs containing image and text information, and preprocess the image-text pairs; Step S200: Select any part of the image-text pair data and perform image-text mismatch processing to make the image information and text information in the selected image-text pair data mismatch. Step S300: Using a backbone network capable of aligning the image and text encoders, a symmetric inversion network is trained based on self-supervised learning. This symmetric inversion network includes an image inversion network and a text inversion network with identical structures. Both the image inversion network and the text inversion network consist of multiple layers of modules with the same structure. Each layer has the following structure: , in Input features Refers to fully connected layers. Refers to the temporary retreat layer. Refers to the modified linear unit; For image-text pairs where the image information and the text information do not match, the image information is processed by an image encoder to obtain features, and then processed by an image inversion network to obtain fused features with the image information as the main component. The text information is processed by a text encoder to obtain features, and then processed by a text inversion network to obtain fused features with the text information as the main component. The two fused features are compared and learned to train the symmetric inversion network. Step S400: Select any two pairs of image-text pairs and exchange the text information in the two pairs of image-text pairs to form an exchanged image-text pair. Input one of the exchanged image-text pairs into the image inversion network to obtain target image features. Input the other exchanged image-text pair into the symmetric inversion network to obtain original image-text features and fused image-text features. Select image information and text information to construct implicit triples based on the target image features, original image-text features, and fused image-text features. Step S500: Train a fusion network based on the implicit triples. The fusion network includes a projection layer that projects the image information and text information into a unified embedding space and a fusion layer that fuses the projected information. The fusion layer includes a direct fusion part and a weighted combination part. Training the fusion network includes inputting the implicit triples into a fusion network with shared weights to obtain the corresponding fusion features, and using the target image features to supervise the learning of each fusion feature. Step S600: Input the image to be retrieved and sequentially pass it through the symmetric inversion network and the combined network to obtain combined features. Then, perform cosine similarity matching with the features of the candidate images extracted by the image encoder, and output the retrieval results according to the similarity value.

2. The zero-shot combined image retrieval method based on a symmetric inversion network according to claim 1, characterized in that, The acquired image-text pair data is large-scale image-text pair data.

3. The zero-sample combined image retrieval method based on a symmetric inversion network according to claim 1, wherein the image inversion network and the text inversion network are preset values, and the preset values ​​of the image inversion network and the text inversion network are ≥2.

4. The zero-shot combined image retrieval method based on a symmetric inversion network according to claim 1, characterized in that, The projection layer includes normalization, full connection, activation, and random deactivation processing of the image information and / or the text information.

5. The zero-shot combined image retrieval method based on a symmetric inversion network according to claim 4, characterized in that, The direct fusion portion of the fusion layer performs at least two full-connection, activation, and random deactivation processes on the image information and / or the text information.

6. The zero-shot combined image retrieval method based on a symmetric inversion network according to claim 5, characterized in that, The weighted combination portion of the fusion layer generates weights based on the image information and text information, and performs a weighted combination of the image information and text information.

7. The zero-shot combined image retrieval method based on a symmetric inversion network according to claim 6, characterized in that, The results of processing the image information and / or the text by the direct fusion part and the weighted combination part of the fusion layer are normalized to obtain the overall fusion result.

8. The zero-shot combined image retrieval method based on a symmetric inversion network according to claim 1, characterized in that, Preprocessing the image-text pair data includes resizing and standardizing the image information in the image-text pair data, and segmenting and embedding the text information in the image-text pair data.

Citation Information

Patent Citations

  • Compression of sift vectors for image matching

    AU2011254041A1

  • Combined zero sample image classification method based on progressive mutual guidance

    CN118379562A