Cross-modal condition controllable retrieval method and system based on hybrid expert system

Through a cross-modal conditional controllable search method based on a hybrid expert system, combining content and style information, and dynamically adjusting the search weight, the problem of difficult to combine content and style information in the existing technology is solved, and flexible multi-dimensional image retrieval is realized, improving the accuracy and user experience of the search.

CN120234438AActive Publication Date: 2025-07-01ZHEJIANG UNIV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510704240.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-01
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

Existing image retrieval methods are difficult to effectively combine content and style information, and cannot flexibly adjust adaptively according to task needs, and cannot meet the multi-dimensional retrieval needs of users.

Method used

A cross-modal conditional controllable search method based on a hybrid expert system is adopted. By obtaining an image training data set with style labels and content labels, a sample pair is constructed, and a mixed expert system and prompt word enhancement vector group is trained, and the content and style search weights are dynamically adjusted to realize joint retrieval of content, style and content-style.

Benefits of technology

It realizes the flexibility of image retrieval and the satisfaction of multi-dimensional requirements, and can dynamically adjust the search mode according to the query text entered by the user, improving the accuracy and user experience of the search.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234438A_ABST
    Figure CN120234438A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-modal condition controllable retrieval method and system based on a hybrid expert system, and belongs to the technical field of image retrieval. An image training data set with a style label and a content label is utilized, and a mixed expert system and a cue word enhancement vector group based on a gated network are trained by combining comparison learning loss and cue regularization loss; an addition result of fusion features output by the hybrid expert system and corresponding prompt vectors in the prompt word enhancement vector group is used as an image embedding vector for similarity retrieval; in the training stage, fusion features output by the hybrid expert system are always weighted image features, and in the retrieval stage, retrieval modes are determined according to a query text input by a user, and image retrieval is performed in different retrieval modes. According to the method, the advantages of text and image modalities are combined, the content and style retrieval weight can be dynamically adjusted, the model can adaptively select task-related prompts based on a new prompt learning strategy, and different retrieval task requirements can be flexibly met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image retrieval technology, and in particular to a cross-modal conditional controllable retrieval method and system based on a hybrid expert system. Background Art

[0002] With the widespread application of multimodal data, how to effectively extract and match relevant information from different types of media (such as text, images, etc.) has become a major challenge in the field of information retrieval. Most of the existing image retrieval methods focus on a single modality (for example, content-based image retrieval or style-based image retrieval), but in many practical applications, users need to consider both content and style information for retrieval. For example, in the retrieval of works of art, users may pay attention to the specific content of the work, or they may pay more attention to the artistic style of the work. Therefore, how to effectively combine content and style in image retrieval to meet the multi-dimensional needs of users is a hot topic in current research. Existing multimodal retrieval methods mostly rely on a single retrieval model or a fixed input method, and cannot be flexibly adjusted according to the different requirements of the task. Therefore, how to design a retrieval method that can process multimodal information and can be flexibly adjusted according to user needs has become a difficult problem that needs to be solved in this field of technology. Summary of the invention

[0003] In order to address the defects in the prior art, the present invention proposes a cross-modal conditional controllable retrieval method and system based on a hybrid expert system, which combines the advantages of text and image modalities, can dynamically adjust the content and style retrieval weights, so that the retrieval can flexibly adapt to different task requirements, and is suitable for joint retrieval between image and text data, including content retrieval, style retrieval and content-style retrieval, and has wide applications in the fields of artistic style, design materials and cross-modal content retrieval.

[0004] In order to achieve the above object, the technical solution adopted by the present invention is as follows:

[0005] In a first aspect, a cross-modal conditional controllable retrieval method based on a hybrid expert system is used for image retrieval, comprising:

[0006] Get a training dataset of images with style and content labels;

[0007] Constructing sample pairs according to the image training data set, training a hybrid expert system based on a gated network and a cue word enhanced vector group with a joint contrastive learning loss and a cue regularization loss, and adding the fusion features output by the hybrid expert system and the corresponding cue vector in the cue word enhanced vector group as the image embedding vector for similarity retrieval;

[0008] In the training phase, the fusion features output by the hybrid expert system are weighted image features;

[0009] In the retrieval stage, first determine the retrieval mode according to the query text input by the user. In the content retrieval mode and the style retrieval mode, the fusion features output by the hybrid expert system are the image features output by the corresponding expert models. In the content-style joint retrieval mode, the fusion features output by the hybrid expert system are weighted image features;

[0010] The hybrid expert system includes a pre-trained content expert model and a pre-trained style expert model for generating content image features and style image features respectively; the weighted image features refer to the weighted result of the fusion weights output by the gating network and the content image features and style image features.

[0011] Furthermore, the prompt enhancement vector group includes a content prompt vector, a style prompt vector, and a content-style joint prompt vector. The three prompt vectors correspond to three retrieval modes respectively, and the initial dimension of the prompt vector is the same as the dimension of the fusion feature.

[0012] Furthermore, the content-style joint prompt vector is fixed and does not participate in training.

[0013] Furthermore, the first term of the prompt regularization loss is the negative value of the L1 norm of the style prompt vector and the content prompt vector, and the second term is a regularization term representing the sum of the squares of the L2 norms of each dimension of the style prompt vector and the content prompt vector.

[0014] Furthermore, the gating network takes the text features of the query text as input and generates fusion weights based on a probability distribution; the fusion weights include content weights, style weights, and content-style joint weights.

[0015] Furthermore, the gating network consists of an encoder based on the Transformer structure and a softmax function. The encoder takes the text features of the query text as input to generate encoded features, and the softmax function processes the encoded features into a probability distribution.

[0016] Furthermore, the corresponding prompt vectors in the prompt enhancement vector group are determined according to the fusion weights output by the gating network, and the prompt vector of the retrieval mode corresponding to the maximum value in the fusion weights is taken.

[0017] Furthermore, in the training stage, the query text of each sample pair is randomly selected from content query texts, style query texts, or content-style joint query texts. If the sample pair matches the query text, the sample pair label is a positive label, otherwise it is a negative label.

[0018] Furthermore, in the retrieval stage, three databases are constructed:

[0019] A content database, applied to a content retrieval mode, stores content image features after candidate images are superimposed with content prompt vectors;

[0020] A style database, applied to a style retrieval mode, stores style image features after candidate images are superimposed with style prompt vectors;

[0021] A content-style joint database, applied to a content-style joint retrieval mode, stores weighted image features of candidate images, and the weighted fusion weights are generated by a gating network.

[0022] In a second aspect, the present invention proposes a cross-modal conditional controllable retrieval system based on a hybrid expert system for implementing the above-mentioned cross-modal conditional controllable retrieval method based on a hybrid expert system.

[0023] The beneficial effects of the present invention are as follows:

[0024] The present invention proposes a hybrid expert system based on a gating network and a group of prompt word enhancement vectors. The hybrid expert system includes a pre-trained content expert model and a pre-trained style expert model respectively used to generate content image features and style image features. The gating network can generate fusion weights for image features generated by different expert models, and dynamically select prompt vectors based on the weights. The added result of the fusion feature and the corresponding prompt vector in the group of prompt word enhancement vectors is used as the image embedding vector for similarity retrieval. The fusion feature output by the hybrid expert system in the training stage is always the weighted image feature. In the retrieval stage, the retrieval mode is first determined according to the query text input by the user, and image retrieval is performed under different retrieval modes.

[0025] The present invention combines information of two modalities, text and image, overcomes the limitations of single-modal input, applies the contrastive learning method to the conditional retrieval task, and introduces a hybrid expert system to achieve collaborative work among multiple expert models. In addition, a new prompt learning strategy is adopted, enabling the model to adaptively select task-related prompts, thereby enhancing the task focusing ability.

[0026] The present invention not only supports content retrieval and style retrieval, but also can achieve content-style joint retrieval, has strong controllability, and users can adjust the content and style preferences of the retrieval according to the input text; the present invention has excellent effects under multiple retrieval modes, especially outstanding in complex scenarios such as art work retrieval. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 is a flowchart of the cross-modal conditional controllable retrieval method based on a hybrid expert system proposed by the present invention.

[0028] Figure 2 is a framework diagram of the training stage.

[0029] Figure 3 It is a simplified framework diagram of the retrieval stage (the hint part is omitted).

[0030] Figure 4 It is a framework diagram of the hint learning mechanism.

[0031] Figure 5 It is a schematic diagram of the application scenario of the present invention.

[0032] Figure 6 It is a schematic diagram of the construction of the training set data.

[0033] Figure 7 It is a comparison result diagram between the present invention and a typical retrieval model.

[0034] Figure 8 It is a schematic diagram of the perceived evaluation results of experts in the art field on the performance of different retrieval models. Detailed implementation manners

[0035] The present invention will be further described and illustrated below in conjunction with the detailed implementation manners. The embodiments are only examples of the present disclosure and do not delimit the scope of limitation. The technical features of each implementation manner in the present invention can be combined correspondingly without conflict.

[0036] The accompanying drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0037] The flowcharts shown in the accompanying drawings are only illustrative and do not necessarily include all steps. For example, some steps can be decomposed, while some steps can be combined or partially combined. Therefore, the actual execution order may be changed according to the actual situation.

[0038] As Figure 1 shown, the present invention proposes a cross-modal conditional controllable retrieval method based on a hybrid expert system for image retrieval, which mainly includes the following steps:

[0039] S01, obtain an image training data set with style labels and content labels.

[0040] S02, construct sample pairs according to the image training data set, and jointly train a hybrid expert system based on a gated network and a group of prompt word enhancement vectors with a contrastive learning loss and a prompt regularization loss.

[0041] Here, the addition result of the fused features output by the hybrid expert system and the corresponding prompt vectors in the prompt word enhancement vector group is used as the image embedding vector for similarity retrieval; in the training stage, the fused features output by the hybrid expert system are weighted image features.

[0042] S03. Determine the retrieval mode according to the query text input by the user, and perform image retrieval under different retrieval modes.

[0043] Here, in the content retrieval mode and the style retrieval mode, the fused features output by the hybrid expert system are the image features output by the corresponding expert models; in the content-style joint retrieval mode, the fused features output by the hybrid expert system are weighted image features.

[0044] In the present invention, the hybrid expert system includes a pre-trained content expert model and a pre-trained style expert model respectively used for generating content image features and style image features; the weighted image features refer to the weighted result of the fusion weights output by the gating network and the content image features and style image features.

[0045] The training stage and the inference stage are introduced separately below.

[0046] (1) Training stage

[0047] As Figure 2 shown, the query text generates text features through the text encoder, and the image respectively generates corresponding content features and style features by the content expert model and the style expert model. Subsequently, the gating network dynamically calculates weights according to the text features, combines the content features and style features to generate fused features, and enhances the model's focusing ability on the task through the prompt learning mechanism to obtain the image embedding vector for similarity retrieval.

[0048] In a specific implementation of the present invention, the implementation process of the training stage is as follows:

[0049] S11. Construct a training set

[0050] Consider including N labeled image samples , where each image corresponding label is defined as . In this embodiment, the label satisfies:

[0051] Style label , indicating 27 style category numbers, such as classicism, impressionism, abstractionism, minimalism, oil painting, Chinese style, sketch, watercolor, documentary photography, etc.

[0052] Content label , representing 80 object category numbers, such as buildings, vehicles (cars, airplanes, ships, bicycles), animals and plants (trees, flowers, cats, birds, marine life), natural phenomena (mountains, rivers, starry sky, lightning, volcanoes), furniture, books, backpacks, beaches, etc.

[0053] S12, each image is input into the expert models in the MOE system for processing, that is, the image is respectively input into the content expert model and the style expert model to extract the content features and style features .

[0054] The finally output fused features are:

[0055]

[0056] Among them, and are dynamically generated by the gating network, generating different weights according to different query texts, and the query texts are randomly selected from the content query text, style query text or content-style joint query text.

[0057] For example, the content query text "I want a picture with a similar content"; the style query text "I want a picture with a similar style"; the content-style joint query text "I want a picture with a similar style and content".

[0058] S13, introduce a prompt word enhancement vector to the fused features finally output by the MOE system , to obtain the embedding vector for query :

[0059]

[0060] Here, the prompt word enhancement vector is a learnable vector and is dynamically selected based on the output of the gating network; to avoid prompt overfitting, a prompt regularization loss is introduced, which will be described below.

[0061] S14, based on this embedding vector , calculate the normalized cosine similarity of the image pair:

[0062]

[0063] Calculate the contrastive learning loss according to the cosine similarity:

[0064]

[0065] Among them, represents the contrastive loss, N represents the number of training samples, Represents a temperature parameter used to adjust the contrast sensitivity; Respectively represent the similarity between training samples i and k, and between training samples i and j. Represents a label indicating whether two training samples match, as follows:

[0066]

[0067] The above loss is used to promote higher similarity between positive samples and greater distance between negative samples.

[0068] S15, jointly prompt the regularization loss and the contrast learning loss, and update the parameters of the gating network and the prompt vector until the parameters converge.

[0069] In a specific implementation of the present invention, the gating network is implemented using a model based on the Transformer architecture, and its function is to parse the text representation of the input query text and dynamically determine the weight contributions of each expert model. Specifically, the text representation of the query text Is used as the input of the gating network, and the result generated by the Transformer structure of the gating network undergoes softmax processing to generate a probability distribution and output fusion weights, including content weights , style weights And content-style joint weights , content weights And style weights Are used to weight the expert output features to obtain the fused features .

[0070] As Figure 4 Shown, to enhance task discriminability, a prompt word enhancement vector Is introduced, which is selected from the prompt vector . The process of the gating network selecting the matching prompt is expressed as:

[0071]

[0072] Among them, Is a trainable vector, which is respectively defined as a content prompt vector, a style prompt vector, and a content-style joint prompt vector. In this embodiment, the dimension of the prompt vector is 218, which is the same as the dimension of the fused feature; Represents the finally selected prompt vector; Represents the content weight , style weight And the maximum value among the content-style joint weights . The logic of the gating network for selecting the matching prompt is that when Is equal to Select the content prompt vector when, when equals select the style prompt vector when, when equals select the content-style joint prompt vector.

[0073] The gating network adaptively selects the corresponding prompt vector, enhances the model's focusing ability on tasks, and clarifies the differences between different tasks to avoid task competition among expert models.

[0074] Prompt vector Initialization and update rules: Are all initialized to zero vectors. Under the content-style joint condition, since the training and inference paradigms are the same, Do not update (remain zero vectors), freezing Forces the model to fully rely on the weights dynamically generated by the gating network under the joint condition and , rather than learning fixed joint prompts, avoiding the model from possibly directly "memorizing" joint condition features instead of truly understanding the independent semantics of content and style, can enhance the modeling ability of independent content / style features and avoid over-parameterization of joint prompts.

[0075] The present invention selects the corresponding prompt vector according to the result to enhance the model's focusing ability, adds the selected prompt vector to the fused features for the next step of matching. To avoid prompt overfitting, a prompt regularization loss is introduced:

[0076]

[0077]

[0078] Where, represents the prompt regularization loss, represents the style prompt vector, represents the i-th dimension in the style prompt vector, represents the content prompt vector, represents the i-th dimension in the content prompt vector, D represents the dimension of the prompt vector, represents the regularization term, represents the scaling factor, represents the L1 norm, represents the square of the L2 norm. In the expression of the prompt regularization loss, the first term encourages the difference between the content prompt and the style prompt, and the second term is the regularization term, which controls the range of the regularization term through the scaling factor, stabilizes the model and reduces the risk of overfitting.

[0079] In a specific implementation of the present invention, the training set data is based on the WikiArt art style dataset and the COCO dataset. Figure 6 Figure 1 shows the process of constructing the training set data. Using the style transfer method, images with clear content labels in the COCO dataset are selected as content reference images, and images with clear style labels in the WikiArt dataset are selected as style reference images. The InST style transfer model is used to perform style transfer operations on the content images to generate synthetic images that contain both clear content information and style information. Each generated image is attached with two category labels, namely content and style. After construction, images with mismatched subjects and labels can be removed. For example, images with the label "apple" but the apple only occupies a very small area of the picture are removed to ensure that the number of samples in each content category and style category is balanced. A total of 21,600 high-quality images are generated, including 27 styles and 80 categories of content. K-means clustering is performed on the style and content features, and the results show that the intra-class distance is significantly smaller than the inter-class distance, verifying the category discrimination and annotation consistency of the dataset.

[0080] The content expert model for extracting image features uses the ImageBind model, the style expert model uses the InternVL model, and the text encoder for extracting text representations uses the T5 text encoder. The ImageBind model, InternVL model, and T5 text encoder are all pre-trained models.

[0081] (2) Inference stage

[0082] As shown in Figure 3 , the query text generates text features through the text encoder. The gating network determines the retrieval mode based on the text features. In the content retrieval mode and style retrieval mode, the fusion features output by the hybrid expert system are the image features output by the corresponding expert models. In the content-style joint retrieval mode, the fusion features output by the hybrid expert system are weighted image features. Image retrieval is performed in the corresponding database under different retrieval modes. It should be noted that Figure 3 the prompt part is omitted in Figure 2. In fact, the sum of the fusion features and the corresponding prompt vectors in the prompt word enhancement vector group is used as the image embedding vector for similarity retrieval.

[0083] In a specific implementation of the present invention, the process of performing retrieval using the model trained above is as follows:

[0084] Step S21: Multimodal input and retrieval mode discrimination

[0085] The user inputs a query text and an example image , and the gating network takes the text representation of the query text as input and generates a content weight , a style weight and the content-style combined weight , automatically determine the retrieval mode as content retrieval, style retrieval, or content-style combined retrieval, and access the corresponding image database for similarity retrieval. For example, if the query text is "a cat in the impressionist style", according to the intention, the retrieval mode is determined as content-style combined retrieval, and then access the content-style combined database. In addition, it also includes a content database and a style database.

[0086] Different databases store different features of candidate images. For example, in the content database, it stores the fused features of candidate images after superimposing the content prompt vector ; in the style database, it stores the fused features of candidate images after superimposing the style prompt vector ; in the content-style combined database, it stores the fused features of candidate images , where the is updated by the weight generated by the gating network, generated through the above training stage, and the parameters are fixed in the inference stage. Similarly, the parameters of the gating network are also generated and fixed through the above training stage.

[0087] In this embodiment, the process of automatically determining the retrieval mode is as follows: when is equal to , it is the content retrieval mode; when is equal to , it is the style retrieval mode; when is equal to , it is the content-style combined retrieval mode.

[0088] Step S22: Dynamically select the expert model of the MOE system according to the retrieval mode to obtain the fused feature finally output by the MOE system :

[0089]

[0090] Step S23: Introduce a prompt word enhancement vector to the fused feature finally output by the MOE system to obtain the embedding vector for query :

[0091]

[0092] Step S24: Perform similarity retrieval in the corresponding database according to the embedding vector for query , sort according to the cosine similarity, and return the set of recommended images that best match the user input, and give a matching score.

[0093] Figure 5Three typical user demand scenarios are shown, namely the cases of retrieving only image content, retrieving only image style, and jointly retrieving both image content and style. The present invention is abbreviated as CCSR. Users can express their retrieval needs through text and example images. The present invention can accurately match the image database according to the input intention, reflecting the retrieval advantages of cross-modal and conditional control of the present invention.

[0094] This embodiment verifies the advantages of the present invention in content retrieval, style retrieval, and content-style retrieval tasks.

[0095] Table 1 shows the performance comparison in style retrieval and content retrieval tasks. The method CCSR of the present invention achieves the best in all four indicators, indicating its excellent performance in handling single tasks. Mean Average Precision @k represents the ranking quality of relevant results among the top k retrieval results; Recall @k represents the proportion of "relevant" items covered among the top k returned results. The higher the two indicators, the better.

[0096] Table 1: Performance Comparison of Different Models in Style Retrieval and Content Retrieval Tasks

[0097]

[0098] Table 2 shows the main experimental results of content-style conditional retrieval, which is used to evaluate the retrieval performance of the model under "joint conditions" (including both content and style). The present invention significantly outperforms other models again, indicating that its multi-condition fusion ability is significantly better than traditional models.

[0099] Table 2: Performance Comparison of Different Models in Content-Style Joint Retrieval

[0100]

[0101] Figure 7 It is a comparison result graph of the present invention and typical retrieval models. Several query examples are randomly selected in the figure, and the performances of each model in the simultaneous retrieval task of content and style are compared respectively. The green tick represents that the retrieval result is accurate, and the model successfully matches the target image whose content and style both meet the query conditions; the red cross represents that the retrieval fails, that is, the returned result does not meet the content or style specified by the user. Through this figure, the advantages and performance of the model of the present invention in the actual application scenario can be intuitively understood, verifying the effectiveness and advancement of the method of the present invention.

[0102] Figure 8It is a schematic diagram of the perceptual evaluation results of the performance of different retrieval models by experts in the art field. In this evaluation experiment, experts in the professional art field were invited to participate. Based on the cross-modal conditional retrieval task proposed by the present invention, the actual accuracy and user satisfaction scores of the retrieval results of different models were respectively evaluated. The horizontal axis in the figure represents each model method, and the vertical axis represents the accuracy score evaluated by experts. The CCSR model proposed by the present invention obtained the highest score in the evaluation, significantly leading other methods, reflecting the outstanding advantages of the present invention in matching actual applications with the needs of professional users.

[0103] Table 3 shows the results of the ablation experiment. The present invention analyzed the performance degradation of the mean average precision @k and recall @k after removing the prompt learning strategy, contrast learning strategy, and regularization strategy.

[0104] Table 3: Results of the ablation experiment

[0105]

[0106] Based on the same inventive concept, the present invention also proposes a cross-modal conditional controllable retrieval system based on a hybrid expert system, and the system includes:

[0107] A dataset acquisition module, which is used to acquire an image training dataset with style labels and content labels;

[0108] A hybrid expert system based on a gated network, which is composed of a gated network, a pre-trained content expert model, and a pre-trained style expert model. The gated network is used to output fusion weights; in the training stage, the fusion feature output by the hybrid expert system is the weighted image feature; in the retrieval stage, first, the retrieval mode is determined according to the query text input by the user. In the content retrieval mode and the style retrieval mode, the fusion feature output by the hybrid expert system is the image feature output by the corresponding expert model. In the content-style joint retrieval mode, the fusion feature output by the hybrid expert system is the weighted image feature; the weighted image feature refers to the weighted result of the fusion weight output by the gated network and the content image feature and the style image feature;

[0109] A prompt word enhancement module, which includes a group of learnable prompt word enhancement vectors. The addition result of the fusion feature output by the hybrid expert system and the corresponding prompt vector in the group of prompt word enhancement vectors is used as the image embedding vector for similarity retrieval;

[0110] A training module, which is used to construct sample pairs according to the image training dataset and jointly train the hybrid expert system based on the gated network and the group of prompt word enhancement vectors with the contrastive learning loss and the prompt regularization loss;

[0111] A retrieval module that performs similarity retrieval in a corresponding database based on the image embedding vector for querying, and returns a set of recommended images that best match the user input sorted by cosine similarity.

[0112] For the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment, and the implementation methods of the remaining modules will not be elaborated here. The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0113] The embodiments of the system of the present invention can be applied to any device with data processing capabilities, and the any device with data processing capabilities can be a device or apparatus such as a computer. The system embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory for operation.

[0114] The above-described embodiments only represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the present invention. For those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention.

Claims

1. A cross-modal conditional controllable retrieval method based on a mixture of experts system for image retrieval, characterized in that Including: Obtain an image training dataset with style tags and content tags; Construct sample pairs based on the image training dataset, and jointly train a mixture-of-experts system based on a gated network and a group of prompt-enhanced vectors with contrastive learning loss and prompt regularization loss. The sum of the fused features output by the mixture-of-experts system and the corresponding prompt vectors in the group of prompt-enhanced vectors is used as the image embedding vector for similarity retrieval; In the training stage, the fused features output by the mixture-of-experts system are weighted image features; In the retrieval stage, first determine the retrieval mode according to the query text input by the user. In the content retrieval mode and the style retrieval mode, the fused features output by the mixture-of-experts system are the image features output by the corresponding expert models. In the content-style joint retrieval mode, the fused features output by the mixture-of-experts system are weighted image features; The mixture-of-experts system includes a pre-trained content expert model and a pre-trained style expert model for generating content image features and style image features respectively; the weighted image features refer to the weighted result of the fused weights output by the gated network and the content image features and style image features.

2. The cross-modal conditional controllable retrieval method based on a hybrid expert system according to claim 1, wherein The group of prompt-enhanced vectors includes a content prompt vector, a style prompt vector, and a content-style joint prompt vector. The three prompt vectors correspond to three retrieval modes respectively, and the initial dimension of the prompt vectors is the same as the dimension of the fused features.

3. The cross-modal conditional controllable retrieval method based on a hybrid expert system according to claim 2, wherein, The content-style joint prompt vector is fixed and does not participate in training.

4. The cross-modal conditional controllable retrieval method based on a hybrid expert system according to claim 2, wherein The first term of the prompt regularization loss is the negative value of the L1 norm of the style prompt vector and the content prompt vector, and the second term is a regularization term representing the sum of the squares of the L2 norms of each dimension of the style prompt vector and the content prompt vector.

5. The cross-modal conditional controllable retrieval method based on a hybrid expert system according to claim 2, wherein The gated network takes the text features of the query text as input and generates fused weights based on a probability distribution; the fused weights include a content weight, a style weight, and a content-style joint weight.

6. The cross-modal conditional controllable retrieval method based on a hybrid expert system according to claim 5, wherein The gated network consists of an encoder based on the Transformer structure and a softmax function. The encoder takes the text features of the query text as input to generate encoded features, and the softmax function processes the encoded features into a probability distribution.

7. The cross-modal conditional controllable retrieval method based on a hybrid expert system according to claim 5, wherein The corresponding prompt vector in the group of prompt-enhanced vectors is determined according to the fused weights output by the gated network, and the prompt vector of the retrieval mode corresponding to the maximum value in the fused weights is taken.

8. The cross-modal conditional controllable retrieval method based on a hybrid expert system according to claim 5, wherein, In the training stage, the query text of each sample pair is randomly selected from content query texts, style query texts, or content-style joint query texts. If the sample pair matches the query text, the sample pair label is a positive label, otherwise it is a negative label.

9. The cross-modal conditional controllable retrieval method based on a hybrid expert system according to claim 2, wherein, In the retrieval stage, three databases are constructed: A content database, which is applied to the content retrieval mode and stores the content image features after the candidate images are superimposed with the content prompt vectors; A style database, which is applied to the style retrieval mode and stores the style image features after the candidate images are superimposed with the style prompt vectors; A content-style joint database, which is applied to the content-style joint retrieval mode and stores the weighted image features of the candidate images. The fused weights for weighting are generated by the gated network.

10. A cross-modal conditional controllable retrieval system based on a hybrid expert system, which is used to implement the cross-modal conditional controllable retrieval method described in claim 1, characterized in that, The system includes: A dataset acquisition module, which is used to obtain an image training dataset with style tags and content tags; A hybrid expert system based on a gating network, which consists of a gating network, a pre-trained content expert model, and a pre-trained style expert model. The gating network is used to output fusion weights. In the training stage, the fusion feature output by the hybrid expert system is the weighted image feature. In the retrieval stage, first, the retrieval mode is determined according to the query text input by the user. In the content retrieval mode and the style retrieval mode, the fusion feature output by the hybrid expert system is the image feature output by the corresponding expert model. In the content-style joint retrieval mode, the fusion feature output by the hybrid expert system is the weighted image feature. The weighted image feature refers to the weighted result of the fusion weight output by the gating network and the content image feature and the style image feature. A prompt enhancement module, which includes a group of learnable prompt enhancement vectors. The addition result of the fusion feature output by the hybrid expert system and the corresponding prompt vector in the group of prompt enhancement vectors is used as the image embedding vector for similarity retrieval. A training module, which is used to construct sample pairs according to the image training dataset and jointly train the hybrid expert system based on the gating network and the group of prompt enhancement vectors with the contrastive learning loss and the prompt regularization loss. A retrieval module, which performs similarity retrieval in the corresponding database according to the image embedding vector for querying and returns the recommended image set that best matches the user input according to the cosine similarity ranking.

Citation Information

Patent Citations

  • Training and retrieval method of cross-modal Hash model for coping with label part missing

    CN118245524A

  • Advertising word generation method and system based on multi-modal large model, and electronic equipment

    CN118506346A

  • Self-adaptive multi-modal false message detection method and model based on dual features

    CN119272106A

  • Cross-modal hash retrieval method based on prompt learning and adaptive Mama gating selection fusion

    CN119322871A

  • Intelligent question answering method and system based on semantic vector optimization and dynamic prompt

    CN119557407A