Target retrieval method and device, equipment and storage medium

By introducing detailed description of the graphic and text pairs and positive and negative sample pools in the training data set of the target retrieval model, the problem of poor performance of the model in the prior art when the graphic and text pairs are limited is solved, and higher target retrieval accuracy is achieved.

CN120179848AActive Publication Date: 2025-06-20HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510655451.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-06-20
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

The existing cross-modal object retrieval model based on CLIP is poor in performance when the graphics and text are limited, resulting in low accuracy in target retrieval.

Method used

A target search method is proposed. By obtaining a pre-trained target search model, the training data set of the model includes simple graphic pairs corresponding to the business data, detailed description graphic pairs, positive sample pool graphic pairs, and negative sample pool graphic pairs. This method uses template application and graphical text model construction and detailed description, and expands training data through positive and negative sample pools to improve the training effect of the model.

Benefits of technology

By extending the training data set, the effectiveness of model training is improved, and the accuracy of the target retrieval model is improved, especially suitable for situations where the number of data is small and the quality is limited.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179848A_ABST
    Figure CN120179848A_ABST
Patent Text Reader

Abstract

The invention discloses a target retrieval method, device and equipment and a storage medium, and relates to the technical field of electronic information, and the target retrieval method comprises the steps that a pre-trained target retrieval model is obtained; wherein the training data set of the target retrieval model comprises a simple image-text pair corresponding to the business data, and a detailed description image-text pair, a positive sample pool image-text pair and a negative sample pool image-text pair which are constructed based on the business data; inputting to-be-retrieved content input by the user into the target retrieval model to obtain a retrieval result output by the target retrieval model; wherein the target retrieval model is used for encoding the to-be-retrieved content into a first feature, and screening out content corresponding to at least one second feature matched with the first feature from an underlying database for target retrieval as a retrieval result. According to the method, the data quantity in the data set of model training is effectively expanded, the model training effect is improved, and then the retrieval accuracy of the target retrieval model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of electronic information technology, and particularly to a target retrieval method, apparatus, device, and storage medium. Background Art

[0002] Cross-modal retrieval refers to a technology that retrieves and displays information in other modalities through various forms of input (such as text, images, audio). It has many applications in daily life: whether it is the search engines on various common social media, the product queries on e-commerce platforms, the retrieval and diagnosis in the medical field, or smart home devices and various control departments, efficient and accurate cross-modal retrieval is required to achieve their main functions, helping people obtain information conveniently, interact, and thus solve problems.

[0003] Currently, existing cross-modal target retrieval technologies are usually based on CLIP (Contrastive Language-Image Pre-training), and different target retrieval models are optimized and designed according to the task characteristics.

[0004] However, such models often perform poorly when the information of text-image pairs is limited, resulting in low accuracy of target retrieval.

[0005] The above content is only used to assist in understanding the technical solution of this application, and does not represent an admission that the above content is prior art. Summary of the Invention

[0006] The main purpose of this application is to provide a target retrieval method, apparatus, device, and storage medium, aiming to solve the technical problem that in related technologies, models based on CLIP often perform poorly when the information of text-image pairs is limited, resulting in low accuracy of target retrieval.

[0007] To achieve the above purpose, this application proposes a target retrieval method, and the method includes: Obtain a pre-trained target retrieval model; wherein, the training data set of the target retrieval model includes simple text-image pairs corresponding to business data, as well as detailed description text-image pairs, positive sample pool text-image pairs, and negative sample pool text-image pairs constructed based on business data; Input the content to be retrieved entered by the user into the target retrieval model to obtain the retrieval result output by the target retrieval model; wherein, the target retrieval model is used to encode the content to be retrieved into a first feature, and screen out at least one content corresponding to a second feature that matches the first feature in the underlying database for target retrieval as the retrieval result.

[0008] In one embodiment, before the step of obtaining the pre-trained target retrieval model, the method further includes: Generating a simple graphic-text pair corresponding to the service data by means of template application; Constructing the detailed description graphic-text pair by passing the service data through a pre-set image-to-text large model; Constructing the positive sample pool graphic-text pair by means of template application or positive example construction for the service data; wherein, the positive example construction method is to first generate seed data by passing the service data through a pre-set large language model, and then generate the positive sample pool graphic-text pair by passing the seed data through a pre-set fine-tuning language model; Generating and matching a negative sample pool graphic-text pair with a similarity within a preset similarity range to the positive sample pool graphic-text pair by passing the positive sample pool graphic-text pair through a pre-set text matching model; Training a pre-set candidate model by using the simple graphic-text pair, the detailed description graphic-text pair, the positive sample pool graphic-text pair and the negative sample pool graphic-text pair to obtain the trained target retrieval model.

[0009] In one embodiment, the candidate model includes an image encoding module, a text encoding module and an image-to-text decoding module; the step of training the pre-set candidate model includes: For each graphic-text pair in the training dataset, perform the following steps: Inputting the image in the graphic-text pair into the image encoding module to obtain a third feature; Inputting the third feature into the image-to-text decoding module to obtain the first text description information corresponding to the third feature; Inputting the second text description information in the graphic-text pair into the text encoding module to obtain a fourth feature; Determining the loss function corresponding to the graphic-text pair based on the third feature, the fourth feature, the first text description information and the second text description information; Training the candidate model based on the loss function corresponding to each graphic-text pair to obtain the target retrieval model.

[0010] In one embodiment, the loss function includes at least one of the following: A first loss function for characterizing the graphic-text contrastive learning loss corresponding to the simple graphic-text pair; A second loss function for characterizing the description loss corresponding to the detailed description graphic-text pair; A third loss function for characterizing the loss function of the positive sample pool corresponding to the positive sample pool graphic-text pair; The fourth loss function is used to characterize the loss function of the negative sample pool corresponding to the text-image pairs in the negative sample pool.

[0011] In one embodiment, the step of inputting the content to be retrieved entered by the user into the target retrieval model to obtain the retrieval result output by the target retrieval model includes: Input the content to be retrieved entered by the user into the target retrieval model, and the target retrieval model performs the following steps: Encode the content to be retrieved into a first feature; Calculate the similarity between the first feature corresponding to the content to be retrieved and at least one second feature in the underlying database; Select the pictures corresponding to a preset number of second features as the retrieval result in the order of decreasing similarity.

[0012] In one embodiment, before the step of inputting the content to be retrieved entered by the user into the target retrieval model to obtain the retrieval result output by the target retrieval model, the method further includes: Input the pictures in the underlying database into the picture encoding module of the target retrieval model to obtain the at least one second feature corresponding to the pictures in the underlying database; The step of encoding the content to be retrieved into a first feature includes: When the retrieval type of the content to be retrieved is a picture, input the content to be retrieved into the picture encoding module of the target retrieval model for encoding to obtain a picture feature as the first feature corresponding to the content to be retrieved; When the retrieval type of the content to be retrieved is a text description, input the content to be retrieved into the text encoding module of the target retrieval model for encoding to obtain a text feature as the first feature corresponding to the content to be retrieved.

[0013] In one embodiment, the step of, when the retrieval type of the content to be retrieved is a picture, inputting the content to be retrieved into the picture encoding module of the target retrieval model for encoding to obtain a picture feature as the first feature corresponding to the content to be retrieved includes: When the retrieval type of the content to be retrieved is a picture, determine whether the content to be retrieved can be read by the target retrieval model; When the content to be retrieved can be read by the target retrieval model, input the content to be retrieved into the picture encoding module of the target retrieval model for encoding to obtain a picture feature as the first feature corresponding to the content to be retrieved.

[0014] In one embodiment, the step of inputting the content to be retrieved into the text encoding module of the target retrieval model for encoding to obtain text features as the first features corresponding to the content to be retrieved when the retrieval type of the content to be retrieved is text description includes: When the retrieval type of the content to be retrieved is text description, determine whether the content to be retrieved is a normal query statement; When the content to be retrieved is a normal query statement, input the content to be retrieved into the text encoding module of the target retrieval model to obtain text features as the first features corresponding to the content to be retrieved.

[0015] In addition, to achieve the above object, the present application also proposes a target retrieval device, which includes: An acquisition module, configured to acquire a pre-trained target retrieval model; wherein, the training data set of the target retrieval model includes simple graphic-text pairs corresponding to business data, as well as detailed description graphic-text pairs, positive sample pool graphic-text pairs, and negative sample pool graphic-text pairs constructed based on business data; A retrieval module, configured to input the content to be retrieved input by a user into the target retrieval model to obtain a retrieval result output by the target retrieval model; wherein, the target retrieval model is configured to encode the content to be retrieved into first features and screen out, in a bottom-layer database for target retrieval, content corresponding to at least one second feature that matches the first features as the retrieval result.

[0016] In addition, to achieve the above object, the present application also proposes a target retrieval device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the computer program is configured to implement the steps of the target retrieval method as described above.

[0017] In addition, to achieve the above object, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the target retrieval method as described above.

[0018] In addition, to achieve the above object, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the target retrieval method as described above.

[0019] One or more technical solutions proposed by the present application have at least the following technical effects: When the business data is limited, in addition to the simple data pairs corresponding to the business data, the training data set used for model training in this application also includes other training data constructed from different perspectives based on the business data. Specifically, it also includes detailed description graphic-text pairs, positive sample pool graphic-text pairs, and negative sample pool graphic-text pairs, effectively expanding the data volume in the data set for model training from different perspectives, effectively improving the effect of model training, and further effectively improving the accuracy of retrieval by the target retrieval model. Compared with the existing problems of noise in training data and low data utilization rate, this application makes relatively full use of business data, and this method is especially suitable for the situation where the data volume is small and the quality is limited. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0021] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0022] Figure 1 is one of the schematic flowcharts of the target retrieval method provided by the present application; Figure 2 is a schematic diagram for distinguishing between simple graphic-text pairs, detailed description graphic-text pairs, positive sample pool graphic-text pairs, and negative sample pool graphic-text pairs in the target retrieval method provided by the present application; Figure 3 is the second schematic flowchart of the target retrieval method provided by the present application; Figure 4 is the schematic flowchart of the training data processing in the target retrieval method provided by the present application; Figure 5 is the schematic diagram of the model structure in the target retrieval method provided by the present application; Figure 6 is the schematic flowchart of the model inference in the target retrieval method provided by the present application; Figure 7 is the third schematic flowchart of the target retrieval method provided by the present application; Figure 8 is the schematic diagram of the structure of the target retrieval device provided by the present application; Figure 9 is the schematic diagram of the structure of the electronic device provided by the present application.

[0023] The implementation, functional features, and advantages of the present application will be further described with reference to the embodiments and the drawings. Detailed implementation manners

[0024] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not used to limit this application.

[0025] Unless otherwise defined, all technical and scientific terms used in this application have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used in this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion.

[0026] In the description of the embodiments of this application, technical terms such as "first" and "second" are only used to distinguish different objects and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity, specific order or primary-secondary relationship of the indicated technical features. In the description of the embodiments of this application, "a plurality of" means two or more unless otherwise specifically defined.

[0027] Referring to "embodiments" in this application means that specific features, structures or characteristics described in connection with the embodiments can be included in at least one embodiment of this application. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described in this application can be combined with other embodiments.

[0028] In the description of the embodiments of this application, the term "and / or" is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally means that the associated objects before and after are in an "or" relationship.

[0029] To better understand the technical solutions of this application, the following will be described in detail in conjunction with the drawings of the specification and specific implementation manners.

[0030] In related technologies, cross-modal search refers to the retrieval behavior between different modalities, such as retrieving pictures, sounds or videos based on text information.

[0031] Existing cross-modal object retrieval technologies are usually based on CLIP, and different object retrieval models are optimized and designed according to the task characteristics. Among them, CLIP is a text-image pre-training model that uses contrastive learning. Such models often result in suboptimal model performance in the case of limited text-image pair information.

[0032] In view of the above problems, the present application provides a target retrieval method, device, equipment and storage medium, aiming to use the idea of BLIP (Bootstrapping Language-Image Pre-training), add a text generation task on the basis of the contrast training of the original text-image pairs, and then combine the optimization of the training data from different angles, such as optimizing the detailed text description (dense caption) corresponding to the picture, positive and negative sample pools, etc., thereby improving the algorithm performance. All training data is utilized more fully, especially suitable for the training of the target retrieval model when the quantity of data is small and the quality is limited.

[0033] Among them, BLIP is an improvement based on CLIP, unifying visual language understanding and generation, and has stronger capabilities in understanding long texts and complex semantics, and can achieve more accurate and in-depth language understanding.

[0034] It should be noted that the execution subject of the embodiments of the present application can be a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, a target retrieval device, etc. that can implement the above functions. Hereinafter, the embodiments of the present application and the following embodiments will be described by taking the target retrieval device as an example.

[0035] The following will specifically describe the embodiments of the present application and the following embodiments.

[0036] The embodiments of the present application provide a target retrieval method, referring to Figure 1 , Figure 1 which is one of the flow diagrams of the target retrieval method provided by the present application. The method includes steps S101 to S102: Step S101, obtain a pre-trained target retrieval model; Among them, the training data set of the target retrieval model includes simple text-image pairs corresponding to business data, as well as detailed description text-image pairs, positive sample pool text-image pairs and negative sample pool text-image pairs constructed based on business data.

[0037] It should be noted that the positive and negative sample pools where the positive sample pool text-image pairs and negative sample pool text-image pairs are located can be understood as disassembling the dense caption corresponding to a certain picture into different attributes, further assembling to form different text positive samples, and further modifying the corresponding attributes to form / assemble different text negative samples. All text positive samples form a positive sample pool, all text negative samples form a negative sample pool, and the whole is called a positive and negative sample pool.

[0038] Taking a certain picture as an example, the person in the picture is a man wearing a red scarf. The text-image pairs in the positive sample pool corresponding to this sample can be: "a person wearing a red scarf", "a man"; the text-image pairs in the negative sample pool corresponding to this sample can be: "a person wearing a black scarf", "a woman". That is, the text-image pairs in the positive sample pool refer to the text that is consistent with the target display in the picture, while the text-image pairs in the negative sample pool are the text that is inconsistent with the target display in the picture.

[0039] Step S102: Input the content to be retrieved entered by the user into the target retrieval model, and obtain the retrieval result output by the target retrieval model. Among them, the target retrieval model is used to encode the content to be retrieved into a first feature, and screen out the content corresponding to at least one second feature that matches the first feature in the underlying database for target retrieval as the retrieval result.

[0040] It should be noted that the content to be retrieved entered by the user can be content such as pictures, texts, audios, etc., and this application does not limit this.

[0041] It should also be noted that a large number of pictures, texts, audios, etc. can be stored in advance in the underlying database, so as to retrieve the retrieval result that matches the content to be retrieved from the underlying database according to the content to be retrieved entered by the user and return it to the user.

[0042] Specifically, the target retrieval model can be pre-trained based on the simple text-image pairs corresponding to the business data, as well as the detailed description text-image pairs, positive sample pool text-image pairs, and negative sample pool text-image pairs constructed based on the business data. The training data involves different angles, which is convenient for improving the model training effect. Then, the trained target retrieval model is used to retrieve the content to be retrieved entered by the user, specifically for feature-level retrieval in the underlying database. The target retrieval model first encodes the content to be retrieved entered by the user into a first feature, and screens out the content corresponding to at least one second feature that matches the first feature in the underlying database as the retrieval result and returns it to the user.

[0043] It should be noted that in order to achieve cross-modal retrieval, this application can encode the content to be retrieved entered by the user into a first feature through the target retrieval model, and perform feature-level matching with the second features that have been pre-encoded in the underlying database, avoiding situations where it is difficult to match the text entered by the user with the pictures in the underlying database. By matching features, it can effectively adapt to cross-modal retrieval.

[0044] In some embodiments, the training dataset may further include general knowledge data, which represents common sense knowledge in the existing world, while the business data represents data in the fields of interest of the target retrieval task, such as security, medical, agriculture, etc. The data type of the general knowledge data is image-text pairs, and the data can be collected from the Internet and is usually large in quantity; while the business data mainly comes from business, usually small in quantity and stored in the format of an attribute dictionary. This application adds general knowledge data for image-text contrast learning in model training, and task-related data for image-text contrast learning, positive and negative sample pool contrast learning, and detailed description generation tasks, enabling the model to learn more purposefully while maintaining generality.

[0045] An embodiment of this application provides a target retrieval method. When the business data is limited, the training dataset used for model training in this application includes not only the simple data pairs corresponding to the business data, but also other training data constructed from different perspectives based on the business data, specifically including detailed description image-text pairs, positive sample pool image-text pairs, and negative sample pool image-text pairs. This effectively expands the data quantity in the model training dataset from different perspectives, effectively improves the effect of model training, and thus effectively improves the accuracy of the target retrieval model for retrieval; compared with the problems of noise in existing training data and low data utilization rate, this application makes more full use of business data, and this method is especially suitable for the situation where the data quantity is small and the quality is limited.

[0046] In some embodiments, a specific implementation manner of constructing a training dataset and training a model is provided. Before step S101, the following steps may be included: S1-1, generating simple image-text pairs corresponding to the business data by means of template application; It should be noted that a text description containing only attribute information can be obtained through template application and combined with a picture to form a simple image-text pair. The template used in template application can be generated by a large model according to specific needs.

[0047] S1-2, constructing the detailed description image-text pairs by passing the business data through a pre-set image-to-text large model; In some embodiments, the image-to-text large model is, for example, the Intern-vl (Intern Vision-Language) model, which is a multimodal large model mainly used for vision-language understanding and generation tasks. The image-to-text large model can also be BLIP-2 (Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models), PaLI (A Jointly-Scaled Multilingual Language-Image Model), etc. This application does not limit this.

[0048] Specifically, a detailed text description of the picture can be obtained through the Intern-vl model, and a detailed description picture-text pair is formed with the picture.

[0049] S1-3. Construct the positive sample pool picture-text pairs by applying templates or constructing positive examples to the service data; wherein, the way of constructing positive examples is to first generate seed data from the service data through a pre-set large language model, and then generate the positive sample pool picture-text pairs from the seed data through a pre-set fine-tuning language model. In some embodiments, the large language model is, for example, models such as GPT-4v, Deepseek, Doubao, Wenxin Yiyan, Qianwen, etc., and the fine-tuning language model is, for example, the qwen model. This application does not limit this.

[0050] Specifically, for the positive and negative sample pools, the following two methods can be used to construct the positive sample pool picture-text pairs: <1> Obtain a text description containing only attribute information by applying a template, and form a positive sample pool picture-text pair with the picture.

[0051] It should be noted that the granularity of the template used to construct the positive sample pool picture-text pairs here can be different from the template used to generate simple picture-text pairs in S1-1, so that the text description in the simple picture-text pair is mainly used to describe the attribute information of the whole picture, while the positive sample pool picture-text pair can be used to describe the attribute information of some details in the picture.

[0052] <2> Generate seed data through, for example, the GPT-4v model, then generate the required detailed text description through, for example, the qwen model, etc., and then disassemble the detailed text description into different attribute information, and further assemble it with the picture to form different positive sample pool picture-text pairs.

[0053] S1-4. Generate and match negative sample pool text-image pairs with a similarity to the positive sample pool text-image pairs within a preset similarity range through a preset text matching model. In some embodiments, the text matching model is, for example, BGE-M3, BERT (Bidirectional Encoder Representations from Transformers), RoBERTa, ALBERT (A Lite BERT), DistilBERT, ELECTRA, T5 model, etc., and the present application does not limit this.

[0054] It should be noted that the preset similarity range can be set according to the actual situation. After constructing the sample pool text-image pairs, for example, set the sample pool text-image pairs with a similarity less than a certain threshold to the positive sample pool text-image pairs as negative sample pool text-image pairs. For example, set the sample pool text-image pairs with a similarity less than 0.4 as negative sample pool text-image pairs, or select the sample pool text-image pairs with a similarity less than 0.75 as its negative sample pool text-image pairs, to avoid the similarity between the constructed negative sample pool text-image pairs and the positive sample pool text-image pairs being too high, resulting in the constructed negative sample pool text-image pairs being difficult to effectively improve the effect of model training.

[0055] S1-5. Use the simple text-image pairs, the detailed description text-image pairs, the positive sample pool text-image pairs, and the negative sample pool text-image pairs to train a preset candidate model to obtain a trained target retrieval model.

[0056] It should be noted that the present application embodiment does not limit the order of execution of S1-1, S1-2, S1-3, and S1-4 above, that is, it does not limit the order of obtaining various data in the training dataset.

[0057] Figure 2 is a schematic diagram of the distinction between simple text-image pairs, detailed description text-image pairs, positive sample pool text-image pairs, and negative sample pool text-image pairs in the target retrieval method provided by the present application. As Figure 2 shown, the following examples illustrate the differences between simple text-image pairs, detailed description text-image pairs, positive sample pool text-image pairs, and negative sample pool text-image pairs: 1) Simple text-image pair: A man wearing a brown polo shirt on the upper body and white shorts on the lower body. 2) Detailed description text-image pair: A man wearing a brown polo shirt on the upper body and white shorts on the lower body. His right hand is in his pocket, his left hand is above his pants, and he is wearing a pair of white shoes. He is wearing a pair of sunglasses and has short hair. There are green trees in the background, and it seems he is standing on the road.

[0058] 3) Positive sample pool image-text pairs: A person wearing a brown polo shirt on the upper body,... (other content), a man wearing white shorts on the lower body.

[0059] 4) Negative sample pool image-text pairs: A person wearing a black polo shirt on the upper body,... (other content), a woman wearing white shorts on the lower body.

[0060] In the embodiments of the present application, although the business data has only one source, the present application makes the best use of its information according to different construction methods, effectively alleviating the problem that the lack of business data affects the model training effect, improving the performance of the trained target retrieval model, and contributing to more accurate target retrieval.

[0061] In some embodiments, a specific implementation manner of the candidate model is provided. The candidate model may include an image encoding module, a text encoding module, and an image-to-text decoding module; Correspondingly, the steps of training the pre-set candidate model may include the following sub-steps: S2-1. For each image-text pair in the training dataset, perform the following steps: Input the image in the image-text pair into the image encoding module to obtain a third feature; Input the third feature into the image-to-text decoding module to obtain the first text description information corresponding to the third feature; Input the second text description information in the image-text pair into the text encoding module to obtain a fourth feature; Based on the third feature, the fourth feature, the first text description information, and the second text description information, determine the loss function corresponding to the image-text pair; S2-2. Train the candidate model based on the loss function corresponding to each image-text pair to obtain the target retrieval model.

[0062] In some embodiments, the candidate model is, for example, a VIT (Vision Transformer) model, a BERT model, a transformer model, etc., and the present application does not limit this.

[0063] In the related art, the candidate model usually only includes an image encoding module and a text encoding module. During the model training process, for an image-text pair, the image is encoded into an image feature through the image encoding module, and the corresponding text description is encoded into a text feature through the text encoding module. Then, based on the image-text comparison of the image feature and the text feature, the loss function is determined to perform supervised training on the model, expecting the feature encoding between the image and the text to be mapped to a unified space and aligned in this space.

[0064] In this application, the candidate model further includes an image-to-text decoding module. During the model training process, for each image-text pair in the training dataset, the third feature output by the image encoding module can be input into the image-to-text decoding module to obtain the first text description information corresponding to the third feature. Based on the first text description information and the second text description information corresponding to the image, the loss function corresponding to the image-text pair is determined to train the candidate model. The image-to-text decoding module is mainly used to assist the image encoding module and the text encoding module in training during the model training stage. After the target retrieval model is trained, only the image encoding module and the text encoding module are used for feature encoding and similarity comparison.

[0065] In the embodiments of this application, the trained target retrieval model is an end-to-end cross-modal retrieval. Based on the contrastive learning task constructed by the two-tower model (i.e., the structure of image encoding and text encoding), an image-to-text decoding module is added to implement the text generation task, enabling the image encoding module and the text encoding module of the two-tower model to further learn the fine-grained alignment relationship between images and texts. The addition of this type of task makes the alignment between images and texts more fine-grained, and at the same time solves various problems such as poor multi-attribute constraints and poor coherent semantic modeling.

[0066] In some embodiments, a specific implementation manner of the loss function is provided. The loss function may include at least one of the following: 1) The first loss function, used to represent the image-text contrastive learning loss corresponding to the simple image-text pair; 2) The second loss function, used to represent the description loss (caption loss) corresponding to the detailed description image-text pair; 3) The third loss function, used to represent the loss function of the positive sample pool corresponding to the positive sample pool image-text pair; 4) The fourth loss function, used to represent the loss function of the negative sample pool corresponding to the negative sample pool image-text pair.

[0067] In the embodiments of this application, compared with the related technology that only uses simple image-text pairs to train the model, this application adds multiple groups of image-to-text losses for the newly added positive and negative example sample pool image-text pairs, effectively improving the training effect. At the same time, due to the addition of the loss function corresponding to the negative example samples, problems such as poor negative query of related models are also solved, and the target retrieval can be more accurately performed for the negative type of content to be retrieved input by the user.

[0068] In some embodiments, a specific implementation manner of the target retrieval model for target retrieval is provided. The above step S102 may include the following sub-steps: S3-1, input the content to be retrieved input by the user into the target retrieval model, and the target retrieval model performs the following steps: Encode the content to be retrieved into a first feature; Calculate the similarity between the first feature corresponding to the content to be retrieved and at least one second feature in the underlying database; Select pictures corresponding to a preset number of second features as the retrieval results in the order of decreasing similarity.

[0069] Specifically, after receiving the content to be retrieved input by the user, the target retrieval model can first encode the content to be retrieved into a first feature, compare it with at least one second feature in the underlying database, calculate the similarity between the first feature and the at least one second feature, and select pictures corresponding to a preset number of second features in the order of decreasing similarity. For example, select the top 100 pictures and return them to the terminal device on the user side for display.

[0070] It should be noted that the preset number of the above-mentioned returned second features can be set according to actual situations and is not limited in this application.

[0071] In some embodiments, the calculated similarity is, for example, cosine similarity.

[0072] In some embodiments, a specific implementation manner of image or text search for images is provided. Before the above step S102, the following steps may further be included: Input the pictures in the underlying database into the picture encoding module of the target retrieval model to obtain the at least one second feature corresponding to the pictures in the underlying database; A specific implementation manner of encoding the content to be retrieved into a first feature may include: 1) When the retrieval type of the content to be retrieved is a picture, input the content to be retrieved into the picture encoding module of the target retrieval model for encoding to obtain a picture feature as the first feature corresponding to the content to be retrieved; Specifically, if the content to be retrieved is a picture, it can be encoded through the picture encoding module of the target retrieval model, and after being encoded into a picture feature, it is matched with at least one second feature in the underlying database, so as to perform target retrieval.

[0073] In some embodiments, a specific implementation manner of, when the retrieval type of the content to be retrieved is a picture, inputting the content to be retrieved into the picture encoding module of the target retrieval model for encoding to obtain a picture feature as the first feature corresponding to the content to be retrieved may include: When the retrieval type of the content to be retrieved is a picture, determine whether the content to be retrieved can be read by the target retrieval model; When the content to be retrieved can be read by the target retrieval model, input the content to be retrieved into the image encoding module of the target retrieval model for encoding to obtain image features, which serve as the first features corresponding to the content to be retrieved.

[0074] Specifically, if the content to be retrieved is an image, it is also necessary to determine the legality of the content to be retrieved, specifically determine whether the image input by the user can be read by the target retrieval model. If it is determined that it can be read by the target retrieval model, the subsequent target retrieval steps can be executed by the target retrieval model; In some embodiments, if it is determined that it is not supported to be read by the target retrieval model, a prompt message for re-input can be returned to the user until the content to be retrieved input by the user is legal, and then the subsequent target retrieval steps are carried out.

[0075] 2) When the retrieval type of the content to be retrieved is a text description, input the content to be retrieved into the text encoding module of the target retrieval model for encoding to obtain text features, which serve as the first features corresponding to the content to be retrieved.

[0076] Specifically, if the content to be retrieved is a text description, it can be encoded through the text encoding module of the target retrieval model, and after being encoded into text features, it is matched with at least one second feature in the underlying database, thereby performing target retrieval.

[0077] In some embodiments, a specific implementation manner of "when the retrieval type of the content to be retrieved is a text description, input the content to be retrieved into the text encoding module of the target retrieval model for encoding to obtain text features, which serve as the first features corresponding to the content to be retrieved" may include: When the retrieval type of the content to be retrieved is a text description, determine whether the content to be retrieved is a normal query statement; When the content to be retrieved is a normal query statement, input the content to be retrieved into the text encoding module of the target retrieval model to obtain text features, which serve as the first features corresponding to the content to be retrieved.

[0078] Specifically, if the content to be retrieved is a text description, it is also necessary to determine the legality of the content to be retrieved, specifically determine whether the text description input by the user is a normal query statement. If it is determined to be a normal query statement, the subsequent target retrieval steps can be executed by the target retrieval model; In some embodiments, if it is determined that the text description input by the user is not a normal query statement, a prompt message for re-input can be returned to the user until the content to be retrieved input by the user is legal, and then the subsequent target retrieval steps are carried out.

[0079] The following is an example to illustrate the target retrieval method provided by the embodiments of the present application. The overall goal of the method of the present application is to perform retrieval according to user needs through a system built by a trained model. For example, retrieve pictures in the underlying database that match the user's description according to the text description provided by the user.

[0080] Generally speaking, first, the training data is processed, then the corresponding loss function is designed for model training, and finally the model is applied, that is, model inference is performed. Before building the retrieval system, all pictures will be encoded using the picture encoding module in the model according to the underlying database provided by the user, so as to finally match the text features encoded by the text encoding module in the model provided by the user.

[0081] Figure 3 is the second flow chart of the target retrieval method provided by the present application. As Figure 3 shown, it mainly includes the following stages: 1) Training data processing: The business data concerned by the task is constructed into simple graphic-text pairs, detailed description graphic-text pairs, positive and negative sample pool graphic-text pairs (including positive sample pool graphic-text pairs and negative sample pool graphic-text pairs) corresponding to the business data through the business attribute data processing module, and the graphic-text pairs combined with the general knowledge data are input into the model for training.

[0082] 2) Model training: The training data uses simple graphic-text pairs, positive and negative sample pool graphic-text pairs, detailed description graphic-text pairs and general knowledge data graphic-text pairs for training. The model adopts a two-tower structure similar to BLIP plus a text understanding and generation module, that is, the above-mentioned image-to-text decoding module. The model loss adopts the original graphic-text contrastive learning loss and the self-developed loss function based on business simple graphic-text pairs and positive and negative sample pools, as well as the caption loss corresponding to the next token predict task. Here, the next token predict task refers to the task of generating the next word, and the data used is the detailed description graphic-text pair. According to Figure 1 each word to generate the detailed description text (dense caption). The task is called ntp (that is, next token predict), and the corresponding loss is called the image-to-text loss (caption loss, that is, the above-mentioned description loss).

[0083] 3) Model inference: In the application stage, the input text or picture content to be queried is incorporated into the text or image encoding module of the model for encoding, and then the similarity with the pre-encoded picture features in the underlying database is calculated, and finally the retrieval result is obtained.

[0084] The following is an example to illustrate the specific implementation details of the above parts: <1> Training data processing: This application divides the overall training data into two parts: general knowledge data and business data. General knowledge data represents common sense knowledge in the existing world, while business data represents data in the fields of interest of the task, such as security, healthcare, agriculture, etc. The data type of general knowledge data is image-text pairs, and the data is collected from the Internet, usually in a large quantity; business data mainly comes from business, usually in a small quantity and is stored in the format of an attribute dictionary. The goal of this stage is to obtain the training data for the model.

[0085] Figure 4 It is a schematic diagram of the process for processing training data in the target retrieval method provided by this application. As Figure 4 shown, for business data with attribute data, this application uses two training transformation methods, namely image-text pairs and positive and negative sample pools.

[0086] Among them, the image-text pairs include two categories. One is the text description that only contains attribute information obtained by template application, generating simple image-text pairs; the other is the detailed text description of the picture obtained through an image-to-text large model (such as Intern-vl), constructing detailed description image-text pairs.

[0087] The positive and negative sample pools also include two categories, namely the image-Attr positive and negative example sample pool and the image-Dense positive and negative example sample pool. Among them, the acquisition of positive samples is different: one is the text description that only contains attribute information obtained by template application (the positive example sample in the above image-Attr positive and negative example sample pool), and the other is that the positive sample generates seed data through a large language model (such as GPT-4v), and then fine-tunes the language model (such as qwen, etc.) to generate the required detailed text description (the positive example sample in the above image-Dense positive and negative example sample pool); the negative samples are matched using a text matching model (such as BGE-M3), and texts within a certain similarity threshold range are selected as negative sample texts.

[0088] Through the above training data processing process, general knowledge data in the form of image-text pairs and business data of two types of image-text pairs and two types of positive and negative sample pools can be obtained. Although the business data has only one source, according to different construction methods of this application, its information is maximally utilized, alleviating the problem of less business data.

[0089] <2> Model training: Figure 5 It is a schematic diagram of the model structure in the target retrieval method provided by this application. As Figure 5 shown, the goal of this stage is to obtain a model for image-text alignment to complete the retrieval task.

[0090] The input of the model is the two major types of data obtained above (general knowledge data and the text-image pairs and structured data related to business data). The model consists of three main modules, including an image encoding module, a text encoding module, and an image-to-text decoding module. The models used can be VIT, Bert, transformer, etc. There are mainly two types of task types, namely text generation tasks and image-text comparison tasks. Among them, the setting of the image-text comparison task is to expect the feature encoding between images and texts to be mapped to a unified space and aligned in this space. The text generation task is an image-to-text task. The addition of this type of task makes the alignment between images and texts more fine-grained, and at the same time solves various problems such as poor multi-attribute constraints and poor coherent semantic modeling.

[0091] To meet the new data form - the addition of positive and negative example sample pools, this application designs a loss calculation method, that is, calculating the positive and negative sample losses within the samples. For example, for a single image, there are 5 positive and 5 negative text positive and negative example corresponding losses respectively. Compared with before, multiple groups of image-to-text losses are added, improving the training effect, and at the same time solving problems such as poor negative query.

[0092] <3> Model Inference: Figure 6 It is a schematic diagram of the model inference process in the target retrieval method provided by this application. As Figure 6 shown, the goal of this stage is to use the model obtained in <2> above, through the user's query input text, and then through model inference to obtain the final retrieval result.

[0093] Figure 7 It is the third schematic diagram of the process of the target retrieval method provided by this application. As Figure 7 shown, the specific process is as follows: a) Feature extraction of the bottom database images: When building the system, the images in the original bottom database (i.e., the above-mentioned underlying database) are encoded using the image encoding module in the model. When the bottom database is expanded, the images in the expanded part can also be further encoded using the same image encoding module to add to the image feature bottom database.

[0094] b) The terminal device obtains the text or image to be retrieved: The system obtains the text description or similar image that the user inputs for retrieval, and then judges the legality of its input retrieval. For text, it judges whether it is a normal query statement, and for images, it judges whether the image is readable by the model.

[0095] c) Feature extraction of the text or image to be retrieved: When it is judged in b) that the input is legal, the text is encoded using the text encoding module in the model to obtain text features, or the image to be retrieved is encoded using the image encoding module in the model to obtain image features. If it is illegal, it is necessary to return to b) to repeat the input judgment until it is legal.

[0096] d) Calculate the similarity between the retrieval feature and the image features in the bottom database: Calculate the similarity, such as cosine similarity, between the retrieval feature obtained in c) and all the image features in the bottom database.

[0097] e) Select candidate images and return the top 100 in the candidate images as the retrieval result to the terminal device: Sort the similarity values obtained in d), and finally return the top 100 images in the similarity ranking and display them on the terminal.

[0098] In the embodiments of the present application, there are at least the following beneficial effects: 1) Aiming at the problems of noise in the current existing data and low utilization rate of data information, a data processing method for constructing positive and negative sample pools for each sample and generating detailed descriptions of each sample is proposed, and different loss functions are designed for the model according to different data forms, maximizing the utilization of the information of the existing training data, adding more positive and negative samples on the basis of the original positive and negative samples, and improving the data utilization rate and training efficiency.

[0099] 2) The present application is an end-to-end cross-modal retrieval method. A text generation task is added on the basis of the contrast learning task constructed by the two-tower model, so that the graph and text encoding modules of the two-tower model further learn the fine-grained alignment relationship between graphics and text, and the knowledge learned by the graph and text encoding modules is more comprehensive.

[0100] 3) The general knowledge data proposed in the present application is used for graphic and text contrast learning. The task-related data is used for graphic and text contrast learning, positive and negative sample pool contrast learning, and detailed description generation tasks, so that the learning of the model is more purposeful on the basis of not losing generality, and the general knowledge that the model maintains generality also takes into account the task itself more.

[0101] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the target retrieval method of the present application. Based on this technical concept, more forms of simple transformations are within the protection scope of the present application.

[0102] The present application also provides a target retrieval device. Figure 8 It is a schematic structural diagram of the target retrieval device provided by the present application, as Figure 8 shown. The target retrieval device includes: An acquisition module 801, configured to acquire a pre-trained target retrieval model; wherein, the training data set of the target retrieval model includes simple graphic and text pairs corresponding to service data, as well as detailed description graphic and text pairs, positive sample pool graphic and text pairs, and negative sample pool graphic and text pairs constructed based on service data. A retrieval module 802 is configured to input the content to be retrieved entered by a user into the target retrieval model, and obtain a retrieval result output by the target retrieval model. The target retrieval model is used to encode the content to be retrieved into a first feature, and screen out, from a bottom-layer database for target retrieval, content corresponding to at least one second feature that matches the first feature as the retrieval result.

[0103] The target retrieval device provided in this application adopts the target retrieval method in the above embodiment, and can solve the technical problem that in the related art, the model based on CLIP often has suboptimal performance when the graphic and text pair information is limited, resulting in low accuracy of target retrieval. Compared with the prior art, the beneficial effects of the target retrieval device provided in this application are the same as those of the target retrieval method provided in the above embodiment, and other technical features in the target retrieval device are the same as those disclosed in the method of the above embodiment, and will not be elaborated here.

[0104] This application provides a target retrieval device, which includes: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the target retrieval method in the above embodiment.

[0105] Next, refer to Figure 9 , Figure 9 FIG. is a schematic structural diagram of an electronic device provided in this application, which shows a schematic structural diagram of a target retrieval device suitable for implementing the embodiment of this application. The target retrieval device in the embodiment of this application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistant), PADs (Portable Application Description: tablet computers), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 9 The target retrieval device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiment of this application.

[0106] As Figure 9As shown, the target retrieval device may include a processing device 901 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 902 or a program loaded from a storage device 903 into a random access memory (RAM: Random Access Memory) 904. In the RAM 904, various programs and data required for the operation of the target retrieval device are also stored. The processing device 901, the ROM 902, and the RAM 904 are connected to each other through a bus 905. An input / output (I / O) interface 906 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 906: an input device 907 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 908 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 903 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 909. The communication device 909 may allow the target retrieval device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a target retrieval device having various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be implemented or had alternatively.

[0107] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device, or installed from the storage device 903, or installed from the ROM 902. When the computer program is executed by the processing device 901, the above functions defined in the methods of the embodiments disclosed in the present application are executed.

[0108] The target retrieval device provided in the present application adopts the target retrieval method in the above embodiments, and can solve the technical problem that in the related art, the CLIP-based model often has sub-optimal performance in the case of limited text-image pair information, resulting in low accuracy of target retrieval. Compared with the prior art, the beneficial effects of the target retrieval device provided in the present application are the same as those of the target retrieval method provided in the above embodiments, and other technical features in the target retrieval device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.

[0109] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0110] As described above, the above are only specific embodiments of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all of them should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

[0111] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the target retrieval method in the above embodiments.

[0112] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electricity, magnetism, light, electromagnetic, infrared, or systems or devices, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM: Random Access Memory), read-only memory (ROM: Read Only Memory), erasable programmable read-only memory (EPROM: Erasable Programmable Read Only Memory or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM: CD-Read Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0113] The above computer-readable storage medium can be included in the target retrieval device; it can also exist separately without being assembled into the target retrieval device.

[0114] The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed by the target retrieval device, the target retrieval device is caused to perform the following steps: Obtain a pre-trained target retrieval model; wherein, the training data set of the target retrieval model includes simple graphic-text pairs corresponding to business data, as well as detailed description graphic-text pairs, positive sample pool graphic-text pairs, and negative sample pool graphic-text pairs constructed based on the business data; Input the content to be retrieved entered by the user into the target retrieval model, and obtain the retrieval result output by the target retrieval model; wherein, the target retrieval model is used to encode the content to be retrieved into a first feature, and screen out at least one content corresponding to a second feature that matches the first feature in the underlying database for target retrieval as the retrieval result.

[0115] Computer program code for performing the operations of the present application can be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0116] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0117] The modules involved in the embodiments of the present application can be implemented in software or in hardware. In some cases, the name of the module does not constitute a limitation on the unit itself.

[0118] The readable storage medium provided by the present application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned target retrieval method, which can solve the technical problem that in the related art, the CLIP-based model often has suboptimal performance when the text-image pair information is limited, resulting in low accuracy of target retrieval. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as those of the target retrieval method provided by the above embodiment, and will not be elaborated here.

[0119] The present application also provides a computer program product, including a computer program, and the steps of the above-mentioned target retrieval method are implemented when the computer program is executed by a processor.

[0120] The computer program product provided by the present application can solve the technical problem that in the related art, the CLIP-based model often has suboptimal performance when the text-image pair information is limited, resulting in low accuracy of target retrieval. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the target retrieval method provided by the above embodiment, and will not be elaborated here.

[0121] The above are only some embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made by using the content of the specification and drawings of the present application under the technical concept of the present application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.

Claims

1. A target retrieval method, characterized in that: The method comprises: Obtain a pre-trained target retrieval model; wherein the training data set of the target retrieval model includes simple image-text pairs corresponding to the business data, as well as detailed description image-text pairs constructed based on the business data, positive sample pool image-text pairs, and negative sample pool image-text pairs; The content to be retrieved input by the user is input into the target retrieval model to obtain the retrieval result output by the target retrieval model; wherein the target retrieval model is used to encode the content to be retrieved into a first feature, and filter out content corresponding to at least one second feature matching the first feature in the underlying database used for target retrieval as the retrieval result.

2. The method according to claim 1, characterized in that Before the step of obtaining the pre-trained target retrieval model, the method further includes: Generate a simple picture-text pair corresponding to the business data by applying a template; The business data is constructed into the detailed description picture and text pair through a preset picture-text model; The business data is applied to the template or constructed as a positive example to construct the positive sample pool image-text pair; wherein the positive example construction method is to first generate seed data from the business data through a preset large language model, and then generate the positive sample pool image-text pair from the seed data through a preset fine-tuning language model; The positive sample pool image-text pairs are subjected to a preset text matching model to generate and match negative sample pool image-text pairs whose similarity with the positive sample pool image-text pairs is within a preset similarity range; The simple image-text pairs, the detailed description image-text pairs, the positive sample pool image-text pairs and the negative sample pool image-text pairs are used to train a pre-set candidate model to obtain a trained target retrieval model.

3. The method according to claim 2, characterized in that The candidate model includes a picture encoding module, a text encoding module and a picture-to-text decoding module; the step of training the preset candidate model includes: For each image-text pair in the training data set, perform the following steps: Inputting the picture in the picture-text pair into the picture encoding module to obtain a third feature; Inputting the third feature into the image-to-text decoding module to obtain first text description information corresponding to the third feature; Inputting the second text description information in the image-text pair into the text encoding module to obtain a fourth feature; Determine a loss function corresponding to the image-text pair based on the third feature, the fourth feature, the first text description information, and the second text description information; The candidate model is trained based on the loss function corresponding to each image-text pair to obtain the target retrieval model.

4. The method according to claim 3, characterized in that The loss function includes at least one of the following: A first loss function, used to characterize the image-text contrast learning loss corresponding to the simple image-text pair; A second loss function is used to characterize the description loss corresponding to the detailed description image-text pair; A third loss function, used to characterize the loss function of the positive sample pool corresponding to the image-text pair in the positive sample pool; A fourth loss function is used to characterize the loss function of the negative sample pool corresponding to the negative sample pool image-text pair.

5. The method according to any one of claims 1 to 4, characterized in that: The step of inputting the to-be-searched content input by the user into the target search model to obtain the search results output by the target search model includes: The content to be searched input by the user is input into the target search model, and the target search model performs the following steps: Encoding the to-be-retrieved content into a first feature; Calculating the similarity between the first feature corresponding to the content to be retrieved and at least one second feature in the underlying database; In descending order of similarity, pictures corresponding to a preset number of second features are selected as the search results.

6. The method according to claim 5, characterized in that Before the step of inputting the to-be-searched content input by the user into the target search model to obtain the search results output by the target search model, the method further includes: Inputting the picture in the underlying database into the picture encoding module of the target retrieval model to obtain the at least one second feature corresponding to the picture in the underlying database; The step of encoding the content to be retrieved into a first feature comprises: In the case where the search type of the content to be searched is a picture, the content to be searched is input into the picture encoding module of the target search model for encoding to obtain a picture feature as a first feature corresponding to the content to be searched; In the case where the search type of the content to be searched is text description, the content to be searched is input into the text encoding module of the target search model for encoding to obtain text features as the first features corresponding to the content to be searched.

7. The method according to claim 6, characterized in that When the search type of the content to be searched is a picture, the step of inputting the content to be searched into the picture encoding module of the target search model for encoding to obtain a picture feature as the first feature corresponding to the content to be searched includes: When the search type of the content to be searched is a picture, determining whether the content to be searched can be read by the target search model; In the case that the content to be retrieved can be read by the target retrieval model, the content to be retrieved is input into the image encoding module of the target retrieval model for encoding to obtain image features as the first features corresponding to the content to be retrieved.

8. The method according to claim 6 or 7, characterized in that When the search type of the content to be searched is text description, the step of inputting the content to be searched into the text encoding module of the target search model for encoding to obtain a text feature as the first feature corresponding to the content to be searched includes: In the case where the search type of the content to be searched is a text description, determining whether the content to be searched is a normal query statement; In the case that the content to be retrieved is a normal query statement, the content to be retrieved is input into the text encoding module of the target retrieval model to obtain a text feature as the first feature corresponding to the content to be retrieved.

9. A target retrieval device, characterized in that: The device comprises: An acquisition module is used to acquire a pre-trained target retrieval model; wherein the training data set of the target retrieval model includes simple image-text pairs corresponding to the business data, as well as detailed description image-text pairs constructed based on the business data, positive sample pool image-text pairs, and negative sample pool image-text pairs; A retrieval module is used to input the content to be retrieved input by the user into the target retrieval model to obtain the retrieval results output by the target retrieval model; wherein the target retrieval model is used to encode the content to be retrieved into a first feature, and filter out content corresponding to at least one second feature matching the first feature in the underlying database used for target retrieval as the retrieval result.

10. A target retrieval device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the target retrieval method according to any one of claims 1 to 8.

11. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the target retrieval method according to any one of claims 1 to 8 are implemented.

12. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the target retrieval method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Chinese image-text retrieval model training method and device based on CLIP, equipment and medium

    CN115221276A

  • Image-text retrieval model training method, image-text retrieval method, image-text retrieval device and image-text retrieval equipment

    CN116226353A

  • Retrieval model pre-training method, text retrieval method and system

    CN118210877A

  • Multi-modal large model assisted unsupervised cross-modal video retrieval method and device

    CN118427396A

  • Cross-modal retrieval model training method and remote sensing image text retrieval method

    CN119646255A