Multi-modal data semantic retrieval method and device, equipment and storage medium

By optimizing the combination of loss functions and task decomposition strategies, the semantic retrieval capabilities of the CLIP model under multiple conditions and negative descriptions are improved, solving the problem of insufficient understanding of exclusionary conditions in existing technologies and achieving more efficient and accurate semantic retrieval.

CN120950705APending Publication Date: 2025-11-14SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511100209.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing CLIP models have limited support for exclusionary/negative conditions in handling complex semantic retrieval, especially when multiple conditions are combined, making it difficult to achieve good performance.

Method used

By introducing a pre-defined combination of optimized loss functions, including bimodal opposition loss function, mutual information loss function and conditional temperature regulation function, a contrastive language-image pre-trained model is trained. By combining multi-layer sample sets and confusion attribute sets, a target sample set is generated, and the model is trained and the task is decomposed. Images are then gradually selected to determine the target image set.

Benefits of technology

The CLIP model's ability to understand exclusionary/negative conditions has been improved, enhancing the adaptability, accuracy, and reliability of semantic retrieval and improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950705A_ABST
    Figure CN120950705A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal data semantic retrieval method and device, equipment and a storage medium, and relates to the technical field of deep learning, and the method comprises the steps: completing the model training operation of a comparison language-image pre-training model based on a preset optimization loss function combination, receiving a multi-modal data semantic retrieval task to be processed based on the trained target comparison language-image pre-training model; analyzing the to-be-processed multi-modal data semantic retrieval task based on the target contrast language-image pre-training model to determine a task decomposition result; performing gradual image screening based on a target comparison language-image pre-training model, each sub-condition in the task decomposition result and a preset similarity measurement strategy; and determining a target semantic retrieval result based on the target score of each candidate image in the screened target image set and a target comparison language-image pre-training model. According to the method and the device, semantic retrieval of the CLIP model under multi-condition and negative description can be efficiently realized, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of deep learning technology, and in particular to a multimodal data semantic retrieval method, apparatus, device, and storage medium. Background Technology

[0002] Currently, CLIP (Contrastive Language-Image Pre-training) models mainly focus on positive similarity matching between text and images. They have limited support for exclusionary / negative conditions included in complex semantic retrieval (such as "the clothes are not red" or "the background does not contain the ocean") and lack the ability to understand exclusionary / negative conditions.

[0003] Some existing studies have attempted to handle exclusionary conditions by adding textual hints for "not" conditions or defining new feature distances. However, these approaches still have bottlenecks in terms of computational complexity and adaptability, making it difficult to achieve good performance in retrieval with multiple conditions. Summary of the Invention

[0004] In view of this, the purpose of this invention is to provide a multimodal data semantic retrieval method, apparatus, device, and storage medium, which can efficiently realize semantic retrieval of CLIP models under multiple conditions and negative descriptions, thereby improving the CLIP model's ability to understand exclusionary / negative conditions, enhancing the CLIP model's adaptability, accuracy, and reliability in semantic retrieval, and ultimately improving the user experience. The specific solution is as follows:

[0005] Firstly, this application provides a multimodal data semantic retrieval method, including:

[0006] The model training operation corresponding to the contrastive language-image pre-trained model is completed based on the preset optimized loss function combination, and the multimodal data semantic retrieval task to be processed is received based on the trained target contrastive language-image pre-trained model.

[0007] The semantic parsing of the multimodal data semantic retrieval task to be processed is performed based on the target contrastive language-image pre-trained model, and the task decomposition result is determined according to the corresponding semantic parsing result; the task decomposition result includes multiple positive sub-conditions and / or negative sub-conditions;

[0008] Based on the target contrast language-image pre-trained model, the sub-conditions in the task decomposition results, and the preset similarity measurement strategy, a step-by-step image screening is performed to determine the selected target image set;

[0009] Based on the target contrastive language-image pre-trained model, the target score corresponding to each candidate image in the target image set is determined, and the corresponding target semantic retrieval result is determined based on the target score.

[0010] Optionally, the step of completing the model training operation corresponding to the contrastive language-image pre-trained model based on a preset optimized loss function combination includes:

[0011] During the training of the contrastive language-image pre-training model, the model is guided by a preset bimodal opposition loss function, a mutual information loss function, and a preset conditional temperature adjustment function.

[0012] The preset bimodal opposition loss function is used to guide the model to distinguish between positive sub-conditions and negative sub-conditions in the feature space; the mutual information loss function is used to maximize the mutual information between the positive sub-conditions and image features, and minimize the mutual information between the negative sub-conditions and image features; the preset condition temperature adjustment function is used to set different temperature coefficients for the positive sub-conditions and the negative sub-conditions respectively when determining similarity.

[0013] Optionally, the method further includes:

[0014] By analyzing whether the initial task samples in the initial task sample set have single-attribute negation descriptions, multi-attribute joint negation descriptions, and nested negation descriptions, the negation complexity corresponding to each initial task sample can be determined.

[0015] Based on the multi-level sample set construction rules and the negation complexity, sample stratification is performed to determine the target sample sets with different negation complexity levels.

[0016] By designing image pairs that contain visually similar but semantically mutually exclusive attributes, a set of confusing attributes is constructed;

[0017] Based on the target sample set with different negation complexity levels, the confusion attribute set, the preset model training evaluation index, and the preset optimized loss function combination, the contrastive language-image pre-trained model is trained to determine the trained target contrastive language-image pre-trained model.

[0018] Optionally, the method further includes:

[0019] During the process of training the contrastive language-image pre-training model based on the target sample set with different levels of negation complexity, the confusion attribute set, the preset model training evaluation index, and the preset optimized loss function combination, the target negative sample is determined based on the preset false positive sample screening rule.

[0020] Based on the diffusion model and target negative samples, a target image that satisfies a partially positive description and includes excluded attributes is generated;

[0021] The target image is added to the training sample set of the target sample set.

[0022] Optionally, the step of performing semantic parsing on the multimodal data semantic retrieval task based on the target contrastive language-image pre-trained model, and determining the task decomposition result according to the corresponding semantic parsing result, includes:

[0023] Based on the target contrastive language-image pre-trained model and the preset multi-granularity conditional decoupling rules, the query text in the multimodal data semantic retrieval task to be processed is semantically parsed to determine the semantic parsing result;

[0024] The query text is decomposed based on the semantic parsing results to determine the corresponding task decomposition results.

[0025] Optionally, the stepwise image filtering based on the target contrastive language-image pre-trained model, the sub-conditions in the task decomposition results, and the preset similarity measurement strategy includes:

[0026] For any of the positive sub-conditions in the task decomposition results, based on the target contrast language-image pre-trained model and the preset similarity measurement strategy, the similarity between each image to be matched and the current positive sub-condition is determined, and the first image set corresponding to the current positive sub-condition is determined based on the corresponding first similarity.

[0027] And / or, for any of the negative sub-conditions in the task decomposition results, based on the target contrastive language-image pre-trained model and the preset similarity measurement strategy, determine the reverse similarity between each of the images to be matched and the current negative sub-condition, and determine the second image set corresponding to the current negative sub-condition based on the corresponding second similarity;

[0028] A candidate image set is determined based on the first image set and / or the second image set;

[0029] Based on the principle of non-repetition, the current sub-condition is selected from each of the positive sub-conditions and / or each of the negative sub-conditions in the task decomposition result;

[0030] The candidate image set pairs are filtered based on the current sub-conditions and the preset similarity measurement strategy to determine the current image set after filtering;

[0031] The process then jumps back to the step of selecting the current sub-condition from each of the positive sub-conditions and / or each of the negative sub-conditions in the task decomposition result based on the principle of non-repetition, until the image filtering operation corresponding to all sub-conditions in the task decomposition result is completed, and the target image set is determined.

[0032] Optionally, determining the target score corresponding to each candidate image in the target image set based on the target contrastive language-image pre-trained model, and determining the corresponding target semantic retrieval result based on the target score, includes:

[0033] Based on the target contrast language-image pre-trained model, the dynamic condition weights corresponding to each of the sub-conditions in the task decomposition results, and the target similarity between the candidate images in the target image set and each of the sub-conditions, the target score corresponding to the candidate image is determined.

[0034] The candidate images are sorted based on the target score, and the corresponding target semantic retrieval results are determined using the image sorting results.

[0035] Secondly, this application provides a multimodal data semantic retrieval device, comprising:

[0036] The training completion module is used to complete the model training operation corresponding to the contrastive language-image pre-trained model based on the preset optimized loss function combination, and to receive the multimodal data semantic retrieval task to be processed based on the trained target contrastive language-image pre-trained model;

[0037] The task decomposition module is used to perform semantic parsing on the semantic retrieval task of the multimodal data to be processed based on the target contrastive language-image pre-trained model, and to determine the task decomposition result according to the corresponding semantic parsing result; the task decomposition result includes multiple positive sub-conditions and / or negative sub-conditions;

[0038] The stepwise filtering module is used to perform stepwise image filtering based on the target contrast language-image pre-trained model, the sub-conditions in the task decomposition results, and the preset similarity measurement strategy, so as to determine the filtered target image set.

[0039] The result determination module is used to determine the target score corresponding to each candidate image in the target image set based on the target contrast language-image pre-trained model, and to determine the corresponding target semantic retrieval result based on the target score.

[0040] Thirdly, this application provides an electronic device, comprising:

[0041] Memory, used to store computer programs;

[0042] A processor is used to execute the computer program to implement the steps of the aforementioned multimodal data semantic retrieval method.

[0043] Fourthly, this application provides a computer-readable storage medium for storing a computer program, which, when executed by a processor, implements the steps of the aforementioned multimodal data semantic retrieval method.

[0044] As can be seen, in this application, model training operations corresponding to the contrastive language-image pre-trained model are completed based on a preset optimized loss function combination, and the multimodal data semantic retrieval task to be processed is received based on the trained target contrastive language-image pre-trained model; semantic parsing of the multimodal data semantic retrieval task to be processed is performed based on the target contrastive language-image pre-trained model, and the task decomposition result is determined according to the corresponding semantic parsing result; the task decomposition result includes multiple positive sub-conditions and / or negative sub-conditions; image screening is performed step by step based on the target contrastive language-image pre-trained model, each sub-condition in the task decomposition result, and a preset similarity measurement strategy to determine the selected target image set; the target score corresponding to each candidate image in the target image set is determined based on the target contrastive language-image pre-trained model, and the corresponding target semantic retrieval result is determined according to the target score. In other words, this application completes model training operations corresponding to the contrastive language-image pre-trained model based on a preset optimized loss function combination. Then, using the trained model, the retrieval task is parsed and decomposed into multiple positive and / or negative sub-conditions. Based on the task decomposition results and a preset similarity measurement strategy, images are progressively filtered, and the retrieval results are determined using the scores corresponding to the filtered target image set. This enables efficient semantic retrieval of the CLIP model under multiple conditions and negative descriptions, thereby improving the CLIP model's understanding of exclusionary / negative conditions, enhancing the CLIP model's adaptability, accuracy, and reliability in semantic retrieval, and ultimately improving the user experience. Attached Figure Description

[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0046] Figure 1 A flowchart of a multimodal data semantic retrieval method provided in this application;

[0047] Figure 2 A flowchart of a specific multimodal data semantic retrieval method provided in this application;

[0048] Figure 3A schematic diagram of a multimodal data semantic retrieval device provided in this application;

[0049] Figure 4 This application provides a structural diagram of an electronic device. Detailed Implementation

[0050] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0051] Some existing studies have attempted to handle exclusionary conditions by adding textual hints for "not" conditions or defining new feature distances. However, these approaches still have bottlenecks in terms of computational complexity and adaptability, making it difficult to achieve good performance in retrieval with multiple conditions.

[0052] To this end, this application provides a multimodal data semantic retrieval scheme that can efficiently realize semantic retrieval of CLIP models under multiple conditions and negative descriptions, thereby improving the CLIP model's ability to understand exclusionary / negative conditions.

[0053] See Figure 1 As shown in the figure, an embodiment of the present invention discloses a multimodal data semantic retrieval method, including:

[0054] Step S11: Based on the preset optimized loss function combination, complete the model training operation corresponding to the contrastive language-image pre-trained model, and receive the multimodal data semantic retrieval task to be processed based on the trained target contrastive language-image pre-trained model.

[0055] In this embodiment, to enable the model to better understand and handle complex multi-condition queries, especially exclusionary descriptions, a new set of loss functions is introduced when training the contrastive language-image pre-trained model. Specifically, during the training process, the model is guided by a preset bimodal opposition loss function, a mutual information loss function, and a preset conditional temperature adjustment function. The preset bimodal opposition loss function guides the model to distinguish between positive and negative sub-conditions in the feature space. The mutual information loss function maximizes the mutual information between the positive sub-condition and image features, and minimizes the mutual information between the negative sub-condition and image features. The preset conditional temperature adjustment function sets different temperature coefficients for the positive and negative sub-conditions when determining similarity.

[0056] Regarding the combination of preset optimization loss functions, (1) the preset bimodal opposition loss function. This is used to ensure that the model maximizes similarity under positive descriptions and minimizes similarity under negative descriptions, enabling the model to distinguish between positive and negative descriptions in the feature space. Its expression is as follows:

[0057] ;

[0058] In the formula, Indicates positive text features; This indicates a negative text feature; N represents the number of samples used in training. This represents the i-th sample; The model represents the sample Output feature values; Indicates for sample The output positive eigenvalues; Indicates for sample The output is the negative text feature value; ||| represents the norm. This loss term encourages the model to learn clearer oppositions in the feature space.

[0059] (2) Mutual information loss function, which is used to improve the model’s ability to distinguish between positive and negative conditions. It maximizes the mutual information between positive conditions and image features, while minimizing the mutual information between negative conditions and image features.

[0060] (3) Preset Condition Temperature Adjustment Function: Different temperature coefficients are set for positive and negative descriptions to control the smoothness of the model in similarity calculation, making the model more flexible in handling exclusionary conditions. Different temperature coefficients are set for positive and negative descriptions. and When dealing with negative descriptions, using smaller temperature values ​​enhances the model's ability to distinguish images that do not meet the criteria. For example:

[0061] ;

[0062] In the formula, This represents the probability or normalized similarity score. It represents the calculated probability that, given the i-th sample in a query, the j-th sample is a correct match. During training, the model aims to maximize this value (when j is a correct positive sample). This represents the feature value output for the j-th sample; This represents the feature value output for the k-th sample; This indicates the temperature coefficient corresponding to a positive description; This indicates a negative description of the corresponding temperature coefficient.

[0063] These loss functions guide the model during training, enabling it to exhibit high responsiveness under different types of query conditions.

[0064] In addition, an energy-based adversarial loss function can be applied. :

[0065] ;

[0066] In the formula, This indicates a positive core description; This represents an image that conforms to the affirmative description; Indicates a complete negative description; This indicates an image that matches the excluded description; a forced affirmative. The similarity is much higher than the negation. The similarity, and the difference is greater than the boundary value. This is to further enhance the model's ability to understand exclusionary / negative conditions.

[0067] Furthermore, when training the CLIP model, this embodiment can first analyze whether the initial task samples in the initial task sample set have single-attribute negation descriptions, multi-attribute joint negation descriptions, and nested negation descriptions to determine the negation complexity corresponding to each initial task sample; then, based on the multi-layer sample set construction rules and the negation complexity, the samples are stratified to determine target sample sets with different negation complexity levels; by designing image pairs containing visually similar but semantically mutually exclusive attributes, a confusion attribute set is constructed; based on the target sample sets with different negation complexity levels, the confusion attribute set, the preset model training evaluation index, and the preset optimization loss function combination, the contrastive language-image pre-trained model is trained to determine the trained target contrastive language-image pre-trained model.

[0068] Meanwhile, during the training of the contrastive language-image pre-training model based on the target sample set with different negation complexity levels, the confusion attribute set, the preset model training evaluation index, and the preset optimization loss function combination, target negative samples are determined based on preset false positive sample screening rules; based on the diffusion model and the target negative samples, target images that satisfy partial positive descriptions and contain excluded attributes are generated; and the target images are added to the training sample set of the target sample set.

[0069] In this way, by synthesizing sample sets with different levels of negation complexity on a large scale and training them with specially designed sets of confusing attributes, and then dynamically identifying difficult samples during training, false positive samples are taken as key samples, and GAN (Generative Adversarial Networks) or diffusion models are used to generate deceptive images that meet the partial affirmative conditions but contain excluded attributes, and then these images are added to the training set for training, the granularity and accuracy of CLIP model training can be effectively improved.

[0070] Step S12: Perform semantic parsing on the semantic retrieval task of the multimodal data to be processed based on the target contrastive language-image pre-trained model, and determine the task decomposition result according to the corresponding semantic parsing result; the task decomposition result includes multiple positive sub-conditions and / or negative sub-conditions.

[0071] In this embodiment, after model training is completed, the trained target contrastive language-image pre-trained model can be used to receive and parse the multimodal data semantic retrieval task to be processed, so as to decompose the query text in the task into multiple independent sub-conditions. That is, based on the target contrastive language-image pre-trained model and the preset multi-granularity condition decoupling rules, the query text in the multimodal data semantic retrieval task to be processed is semantically parsed to determine the semantic parsing result; based on the semantic parsing result, the query text is decomposed to determine the corresponding task decomposition result. Among them, the decomposed sub-conditions can be positive sub-conditions (e.g., "wearing a bear hat") or negative sub-conditions (e.g., "clothes are not red"). These sub-conditions will serve as independent units for subsequent processing, providing support for the filtering and matching of each condition.

[0072] Step S13: Based on the target contrast language-image pre-trained model, the sub-conditions in the task decomposition results, and the preset similarity measurement strategy, perform stepwise image filtering to determine the selected target image set.

[0073] In this embodiment, combined with Figure 2As shown, after decomposing the task, a sub-condition processing module is constructed using the decomposed sub-conditions. In this module, for each positive sub-condition, a standard text encoder is used to calculate the similarity score between the image and that sub-condition; for negative sub-conditions, a negative text encoder is used to calculate the reverse similarity. For images that meet the positive sub-conditions, the top m1% of images with the highest scores are added to the candidate set; for images that meet the negative sub-conditions, the top m2% of images with the lowest scores are added to the candidate set. After this step, a preliminary candidate set is generated, in which the images satisfy different sub-conditions. Then, based on the preliminary candidate set, condition filtering is performed step-by-step. For each sub-condition, the images in the candidate set are further filtered. For the similarity score of each sub-condition, images that meet the current condition are retained, and images that do not meet the condition are removed. This operation is performed on all sub-conditions one by one to ensure that the final retained images simultaneously meet all conditions. In this way, through step-by-step filtering, the accuracy of image retrieval is effectively improved.

[0074] Specifically, regarding the process of progressively filtering images, in this embodiment, for any positive sub-condition in the task decomposition result, based on the target contrastive language-image pre-trained model and the preset similarity measurement strategy, the similarity between each image to be matched and the current positive sub-condition is determined, and a first image set corresponding to the current positive sub-condition is determined based on the corresponding first similarity; and / or, for any negative sub-condition in the task decomposition result, based on the target contrastive language-image pre-trained model and the preset similarity measurement strategy, the reverse similarity between each image to be matched and the current negative sub-condition is determined, and a second similarity is determined based on the corresponding second similarity. The process involves: first, determining a second image set corresponding to a sub-condition; second, determining a candidate image set based on the first image set and / or the second image set; third, selecting the current sub-condition from each of the positive and / or negative sub-conditions in the task decomposition result based on the principle of non-repetition; fourth, filtering the candidate image set pairs based on the current sub-condition and the preset similarity measurement strategy to determine the current image set after filtering; fifth, returning to the step of selecting the current sub-condition from each of the positive and / or negative sub-conditions in the task decomposition result based on the principle of non-repetition, until the image filtering operation corresponding to all sub-conditions in the task decomposition result is completed, and finally determining the target image set.

[0075] Step S14: Determine the target score corresponding to each candidate image in the target image set based on the target contrast language-image pre-training model, and determine the corresponding target semantic retrieval result based on the target score.

[0076] In this embodiment, combined with Figure 2As shown, after determining the target image set, the target semantic retrieval result is determined based on the target score corresponding to each candidate image. That is, the target score corresponding to the candidate image is determined based on the target contrast language-image pre-trained model, the dynamic condition weights corresponding to each sub-condition in the task decomposition result, and the target similarity between the candidate image in the target image set and each sub-condition. The candidate images are sorted based on the target score, and the corresponding target semantic retrieval result is determined using the corresponding image sorting result.

[0077] Specifically, for any candidate image that passes the conditional filtering, a comprehensive score is calculated based on its similarity score with each sub-condition. Weights are assigned to each sub-condition. And score the results for each of the n sub-conditions. A weighted average is calculated to obtain the final score of the candidate image. :

[0078] .

[0079] Overall score This represents the degree to which the candidate image matches the complete query. Different weights can be assigned to each sub-condition according to user needs to reflect the importance or priority of the condition. The overall score is then used to determine the final result. The candidate images that meet the criteria are sorted. The top N sorted images are output as the final search results. This sorting process provides users with an intuitive and clear explanation of the results by displaying the overall score of each image and its matching degree on each sub-criteria, thus improving the user's search experience.

[0080] Therefore, in this application, model training corresponding to the contrastive language-image pre-trained model is completed based on a preset optimized loss function combination. Then, the trained model is used to parse the retrieval task and decompose it into multiple positive and / or negative sub-conditions. Based on the task decomposition results and a preset similarity measurement strategy, images are progressively filtered, and the retrieval results are determined using the scores corresponding to the filtered target image set. This enables efficient semantic retrieval of the CLIP model under multiple conditions and negative descriptions, thereby improving the CLIP model's understanding of exclusionary / negative conditions, enhancing the CLIP model's adaptability, accuracy, and reliability in semantic retrieval, and ultimately improving the user experience.

[0081] See Figure 3 As shown in the embodiments, this application also discloses a multimodal data semantic retrieval device, including:

[0082] The training completion module 11 is used to complete the model training operation corresponding to the contrastive language-image pre-trained model based on the preset optimized loss function combination, and to receive the multimodal data semantic retrieval task to be processed based on the trained target contrastive language-image pre-trained model.

[0083] The task decomposition module 12 is used to perform semantic parsing on the semantic retrieval task of the multimodal data to be processed based on the target contrastive language-image pre-trained model, and to determine the task decomposition result according to the corresponding semantic parsing result; the task decomposition result includes multiple positive sub-conditions and / or negative sub-conditions.

[0084] The stepwise filtering module 13 is used to perform stepwise image filtering based on the target contrast language-image pre-trained model, the sub-conditions in the task decomposition results and the preset similarity measurement strategy, so as to determine the filtered target image set.

[0085] The result determination module 14 is used to determine the target score corresponding to each candidate image in the target image set based on the target contrast language-image pre-trained model, and to determine the corresponding target semantic retrieval result based on the target score.

[0086] Therefore, in this application, model training corresponding to the contrastive language-image pre-trained model is completed based on a preset optimized loss function combination. Then, the trained model is used to parse the retrieval task and decompose it into multiple positive and / or negative sub-conditions. Based on the task decomposition results and a preset similarity measurement strategy, images are progressively filtered, and the retrieval results are determined using the scores corresponding to the filtered target image set. This enables efficient semantic retrieval of the CLIP model under multiple conditions and negative descriptions, thereby improving the CLIP model's understanding of exclusionary / negative conditions, enhancing the CLIP model's adaptability, accuracy, and reliability in semantic retrieval, and ultimately improving the user experience.

[0087] In some specific embodiments, the training completion module can be used to: guide the model based on a preset bimodal opposition loss function, a mutual information loss function, and a preset conditional temperature adjustment function during the training of the contrastive language-image pre-trained model; wherein, the preset bimodal opposition loss function is used to guide the model to distinguish between positive sub-conditions and negative sub-conditions in the feature space; the mutual information loss function is used to maximize the mutual information between the positive sub-conditions and image features, and minimize the mutual information between the negative sub-conditions and image features; the preset conditional temperature adjustment function is used to set different temperature coefficients for the positive sub-conditions and the negative sub-conditions respectively when determining similarity.

[0088] In some specific embodiments, the multimodal data semantic retrieval device can also be used to: determine the negation complexity corresponding to each initial task sample by analyzing whether there are single-attribute negation descriptions, multi-attribute joint negation descriptions, and nested negation descriptions in the initial task sample set; perform sample stratification based on multi-level sample set construction rules and the negation complexity to determine target sample sets with different negation complexity levels; construct a confusion attribute set by designing image pairs containing visually similar but semantically mutually exclusive attributes; and train the contrastive language-image pre-trained model based on the target sample sets with different negation complexity levels, the confusion attribute set, preset model training evaluation indicators, and the preset optimization loss function combination to determine the trained target contrastive language-image pre-trained model.

[0089] In some specific embodiments, the multimodal data semantic retrieval device can also be used to: determine target negative samples based on preset false positive sample screening rules during the process of training the contrastive language-image pre-training model based on the target sample set with different negation complexity levels, the confusion attribute set, the preset model training evaluation index, and the preset optimization loss function combination; generate target images that satisfy partial positive descriptions and contain excluded attributes based on the diffusion model and the target negative samples; and add the target images to the training sample set of the target sample set.

[0090] In some specific embodiments, the task decomposition module 12 can be used to: perform semantic parsing on the query text in the multimodal data semantic retrieval task to be processed based on the target contrastive language-image pre-trained model and preset multi-granularity conditional decoupling rules, so as to determine the semantic parsing result; and decompose the query text based on the semantic parsing result to determine the corresponding task decomposition result.

[0091] In some specific embodiments, the stepwise screening module 13 can be specifically used to: for any of the positive sub-conditions in the task decomposition results, based on the target contrastive language-image pre-training model and the preset similarity measurement strategy, determine the similarity between each image to be matched and the current positive sub-condition, and determine the first image set corresponding to the current positive sub-condition based on the corresponding first similarity; and / or, for any of the negative sub-conditions in the task decomposition results, based on the target contrastive language-image pre-training model and the preset similarity measurement strategy, determine the reverse similarity between each image to be matched and the current negative sub-condition, and determine the reverse similarity between the image to be matched and the current negative sub-condition based on the corresponding second similarity. The process involves: first, determining a second image set corresponding to a negative sub-condition; second, determining a candidate image set based on the first image set and / or the second image set; third, selecting the current sub-condition from each of the positive sub-conditions and / or each of the negative sub-conditions in the task decomposition result based on the non-repetition principle; fourth, filtering the candidate image set pairs based on the current sub-condition and the preset similarity measurement strategy to determine the current image set after filtering; fifth, returning to the step of selecting the current sub-condition from each of the positive sub-conditions and / or each of the negative sub-conditions in the task decomposition result based on the non-repetition principle, until the image filtering operation corresponding to all sub-conditions in the task decomposition result is completed, and finally determining the target image set.

[0092] In some specific embodiments, the result determination module 14 can be used to: determine the target score corresponding to the candidate image based on the target contrast language-image pre-trained model, the dynamic condition weights corresponding to each of the sub-conditions in the task decomposition result, and the target similarity between the candidate image in the target image set and each of the sub-conditions; sort the candidate images based on the target score, and determine the corresponding target semantic retrieval result using the corresponding image sorting result.

[0093] Furthermore, embodiments of this application also disclose an electronic device, Figure 4 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.

[0094] Figure 4 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of this application. Specifically, the electronic device 20 may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the multimodal data semantic retrieval method disclosed in any of the foregoing embodiments. Alternatively, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0095] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0096] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0097] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the multimodal data semantic retrieval method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, the computer program 222 may further include a computer program capable of performing other specific tasks.

[0098] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned multimodal data semantic retrieval method. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.

[0099] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0100] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0101] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0102] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0103] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A multimodal data semantic retrieval method, characterized in that, include: The model training operation corresponding to the contrastive language-image pre-trained model is completed based on the preset optimized loss function combination, and the multimodal data semantic retrieval task to be processed is received based on the trained target contrastive language-image pre-trained model. The semantic parsing of the multimodal data semantic retrieval task to be processed is performed based on the target contrastive language-image pre-trained model, and the task decomposition result is determined according to the corresponding semantic parsing result; The task decomposition result includes multiple positive sub-conditions and / or negative sub-conditions; Based on the target contrast language-image pre-trained model, the sub-conditions in the task decomposition results, and the preset similarity measurement strategy, a step-by-step image screening is performed to determine the selected target image set; Based on the target contrastive language-image pre-trained model, the target score corresponding to each candidate image in the target image set is determined, and the corresponding target semantic retrieval result is determined based on the target score.

2. The multimodal data semantic retrieval method according to claim 1, characterized in that, The model training operation based on the preset optimized loss function combination to complete the comparison language-image pre-trained model includes: During the training of the contrastive language-image pre-training model, the model is guided by a preset bimodal opposition loss function, a mutual information loss function, and a preset conditional temperature adjustment function. The preset bimodal opposition loss function is used to guide the model to distinguish between positive sub-conditions and negative sub-conditions in the feature space; the mutual information loss function is used to maximize the mutual information between the positive sub-conditions and image features, and minimize the mutual information between the negative sub-conditions and image features; the preset condition temperature adjustment function is used to set different temperature coefficients for the positive sub-conditions and the negative sub-conditions respectively when determining similarity.

3. The multimodal data semantic retrieval method according to claim 1, characterized in that, Also includes: By analyzing whether the initial task samples in the initial task sample set have single-attribute negation descriptions, multi-attribute joint negation descriptions, and nested negation descriptions, the negation complexity corresponding to each initial task sample can be determined. Based on the multi-level sample set construction rules and the negation complexity, sample stratification is performed to determine the target sample sets with different negation complexity levels. By designing image pairs that contain visually similar but semantically mutually exclusive attributes, a set of confusing attributes is constructed; Based on the target sample set with different negation complexity levels, the confusion attribute set, the preset model training evaluation index, and the preset optimization loss function combination, the contrastive language-image pre-trained model is trained to determine the trained target contrastive language-image pre-trained model.

4. The multimodal data semantic retrieval method according to claim 3, characterized in that, Also includes: During the process of training the contrastive language-image pre-training model based on the target sample set with different levels of negation complexity, the confusion attribute set, the preset model training evaluation index, and the preset optimization loss function combination, the target negative sample is determined based on the preset false positive sample screening rule. Based on the diffusion model and target negative samples, a target image that satisfies a partially positive description and includes excluded attributes is generated; The target image is added to the training sample set of the target sample set.

5. The multimodal data semantic retrieval method according to claim 1, characterized in that, The semantic parsing of the multimodal data semantic retrieval task based on the target contrastive language-image pre-trained model, and the determination of task decomposition results based on the corresponding semantic parsing results, include: Based on the target contrastive language-image pre-trained model and the preset multi-granularity conditional decoupling rules, the query text in the multimodal data semantic retrieval task to be processed is semantically parsed to determine the semantic parsing result; The query text is decomposed based on the semantic parsing results to determine the corresponding task decomposition results.

6. The multimodal data semantic retrieval method according to any one of claims 1 to 5, characterized in that, The stepwise image filtering based on the target contrastive language-image pre-trained model, the sub-conditions in the task decomposition results, and the preset similarity measurement strategy includes: For any of the positive sub-conditions in the task decomposition results, based on the target contrast language-image pre-trained model and the preset similarity measurement strategy, the similarity between each image to be matched and the current positive sub-condition is determined, and the first image set corresponding to the current positive sub-condition is determined based on the corresponding first similarity. And / or, for any of the negative sub-conditions in the task decomposition results, based on the target contrastive language-image pre-trained model and the preset similarity measurement strategy, determine the reverse similarity between each of the images to be matched and the current negative sub-condition, and determine the second image set corresponding to the current negative sub-condition based on the corresponding second similarity; A candidate image set is determined based on the first image set and / or the second image set; Based on the principle of non-repetition, the current sub-condition is selected from each of the positive sub-conditions and / or each of the negative sub-conditions in the task decomposition result; The candidate image set pairs are filtered based on the current sub-conditions and the preset similarity measurement strategy to determine the current image set after filtering; The process then jumps back to the step of selecting the current sub-condition from each of the positive sub-conditions and / or each of the negative sub-conditions in the task decomposition result based on the principle of non-repetition, until the image filtering operation corresponding to all sub-conditions in the task decomposition result is completed, and the target image set is determined.

7. The multimodal data semantic retrieval method according to claim 1, characterized in that, The step of determining the target score corresponding to each candidate image in the target image set based on the target contrastive language-image pre-trained model, and determining the corresponding target semantic retrieval result based on the target score, includes: Based on the target contrast language-image pre-trained model, the dynamic condition weights corresponding to each of the sub-conditions in the task decomposition results, and the target similarity between the candidate images in the target image set and each of the sub-conditions, the target score corresponding to the candidate image is determined. The candidate images are sorted based on the target score, and the corresponding target semantic retrieval results are determined using the image sorting results.

8. A multimodal data semantic retrieval device, characterized in that, include: The training completion module is used to complete the model training operation corresponding to the contrastive language-image pre-trained model based on the preset optimized loss function combination, and to receive the multimodal data semantic retrieval task to be processed based on the trained target contrastive language-image pre-trained model; The task decomposition module is used to perform semantic parsing on the semantic retrieval task of the multimodal data to be processed based on the target contrastive language-image pre-trained model, and to determine the task decomposition result based on the corresponding semantic parsing result; The task decomposition result includes multiple positive sub-conditions and / or negative sub-conditions; The stepwise filtering module is used to perform stepwise image filtering based on the target contrast language-image pre-trained model, the sub-conditions in the task decomposition results, and the preset similarity measurement strategy, so as to determine the filtered target image set. The result determination module is used to determine the target score corresponding to each candidate image in the target image set based on the target contrast language-image pre-trained model, and to determine the corresponding target semantic retrieval result based on the target score.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the multimodal data semantic retrieval method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Used to store a computer program, which, when executed by a processor, implements the multimodal data semantic retrieval method as described in any one of claims 1 to 7.