Data retrieval method and device and nonvolatile storage medium

By using a multimodal feature extraction model to perform feature extraction processing on the input data in a multimodal search system, it supports single-modal or multimodal data as input, and solves the problem that existing systems cannot meet the needs of complex and diverse users, and realizes complex modal retrieval with high accuracy.

CN120104650AActive Publication Date: 2025-06-06CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510174716.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-06
Estimated Expiration
2045-02-17

AI Technical Summary

Technical Problem

The existing multimodal retrieval system has limitations in flexibility, inter-target interactive understanding and long-tail task processing, especially requiring users to provide information in both image and text modalities as query input, which cannot meet the complex and diverse user needs.

Method used

The multimodal feature extraction model is used to perform feature extraction processing on the input data, and supports the multimodal combined data obtained by combining single-modal data as input. By fusing multiple different types of single-modal feature extraction models, data retrieval of any modality is realized.

Benefits of technology

The function of supporting arbitrary modal as input and performing arbitrary modal retrieval is realized, which improves the accuracy of search results in complex modal retrieval scenarios, and solves the problem that multimodal retrieval technology limits the input information format and makes it impossible to achieve complex modal retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104650A_ABST
    Figure CN120104650A_ABST
Patent Text Reader

Abstract

The invention discloses a data retrieval method and device and a nonvolatile storage medium. The method comprises the steps that input data are received, and the input data comprise at least one of single-mode data and multi-mode combined data obtained by combining multiple kinds of single-mode data; a multi-modal feature extraction model is adopted to perform feature extraction processing on the input data to obtain feature information of the input data, and the multi-modal feature extraction model is obtained by performing fusion training on a plurality of different types of single-modal feature extraction models; each single-mode feature extraction model is used for performing feature extraction processing on one type of single-mode data; target data with the similarity with the feature information larger than a preset similarity value is determined in a retrieval library, and the target data serves as a retrieval result. According to the method and the device, the technical problem that complex modal retrieval cannot be realized due to the fact that the format of the input information is limited by a multi-modal retrieval technology in related technologies is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and more specifically, to a data retrieval method and device, and a non-volatile storage medium. Background Art

[0002] With the rapid development of deep learning technology, multimodal retrieval systems have emerged in the field of artificial intelligence and have become an important tool for processing cross-media information retrieval. Such systems can integrate information from multiple modalities such as text, images, and voice to provide users with more accurate and rich content retrieval services. For example, the emergence of the Contrastive Language-Image Pre-training (CLIP) model enables the system to recognize any category in an image based on text prompts, significantly improving the performance of cross-modal retrieval. However, related technologies still have obvious limitations in terms of the flexibility of multimodal retrieval, interactive understanding between targets, and processing of long-tail tasks; for example, most current multimodal retrieval systems require users to provide information in both image and text modalities as query input, which cannot meet the complex and diverse needs of users.

[0003] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention

[0004] The embodiments of the present application provide a data retrieval method and device, and a non-volatile storage medium to at least solve the technical problem that complex modal retrieval cannot be achieved due to the multimodal retrieval technology in the related art limiting the format of input information.

[0005] According to one aspect of an embodiment of the present application, a data retrieval method is provided, comprising: receiving input data, wherein the input data comprises at least one of the following: unimodal data, and multimodal combined data obtained by combining multiple unimodal data; performing feature extraction processing on the input data using a multimodal feature extraction model to obtain feature information of the input data, wherein the multimodal feature extraction model is obtained by fusion training of multiple different types of unimodal feature extraction models, and each unimodal feature extraction model is used to perform feature extraction processing on one type of unimodal data; determining target data in a retrieval library whose similarity with the feature information is greater than a preset similarity value, and using the target data as a retrieval result.

[0006] Optionally, the multimodal feature extraction model is trained in the following manner: obtaining training data, wherein the training data includes: multiple types of unimodal data, multimodal data generated based on unimodal data, unimodal data includes: image data, text data, voice data, and multimodal data includes: any combination of unimodal data; determining a training data set corresponding to a multimodal training task, wherein the training data set includes: a first type of training data set composed of unimodal data of the same type, and a second type of training data set generated based on different types of unimodal data; using a unimodal feature extraction model corresponding to the training data set to extract features from the training data set corresponding to the training data set to obtain multiple types of unimodal features, wherein the type of unimodal features corresponds to the type of unimodal data; training a unimodal feature extraction model based on the multiple types of unimodal features, and updating the weights of the multiple types of unimodal features.

[0007] Optionally, during the process of training the multimodal feature extraction model, the method also includes: determining a total loss function based on multiple categories of unimodal features, and updating the weights of a class of unimodal features in the multiple categories of unimodal features based on the total loss function, until the total change in the total loss function is less than or equal to a preset change, determining that the training is completed, and obtaining a multimodal feature extraction model, wherein the total change in the total loss function is the difference between the previous total loss function and the next total loss function, the previous total loss function is the total loss function determined by the Nth training, and the next total loss function is the total loss function determined by the N+1th training, and when updating the weights of one class of unimodal features, other classes of unimodal features in the multiple classes of unimodal features remain unchanged.

[0008] Optionally, in the case where the multimodal training task is an image-text retrieval task that indicates determining a retrieval result based on image data and text data, determining a training data set corresponding to the multimodal training task includes: determining a second type of training data set generated based on image data and text data as a training data set corresponding to the multimodal training task; performing feature extraction on the training data set corresponding to the training data set using a unimodal feature extraction model corresponding to the training data set, including: performing feature extraction processing on the text data using a text feature extraction model to obtain text features, performing feature extraction processing on the image data using an image feature extraction model to obtain first image features, and performing feature extraction processing on a mask image using an image feature extraction model to obtain second image features, wherein the mask image is obtained by masking the image data based on the text data.

[0009] Optionally, when the multimodal training task is a picture and text retrieval task, the total loss function is determined according to multiple categories of unimodal features, including: determining a first loss function according to text features and first image features; fusing the first image features and the second image features to obtain a third image feature, and determining a second loss function according to the text features and the third image features; determining the total loss function according to the first loss function, the second loss function and the model parameters.

[0010] Optionally, the method also includes: determining the feature extraction results corresponding to each type of unimodal data, determining the independent loss function of the unimodal feature extraction model corresponding to each type of unimodal data according to the feature extraction results, and determining the fusion loss function according to multiple feature extraction results; determining a first change in the independent loss function and a second change in the fusion loss function, and when the first change is less than or equal to a preset change, and the second change is less than or equal to a preset change, determining to trigger the execution of the step of determining the total loss function based on multiple unimodal feature extraction results.

[0011] Optionally, the retrieval library is generated by the following method: using different unimodal feature extraction models to extract features from input data to obtain multiple unimodal features; and using an all-things detection model to analyze the input data, determine the objects to be detected contained in the input data, and extract feature information of the objects to be detected, wherein the feature information at least includes: boundary information of the objects to be detected; and generate a retrieval library based on multiple unimodal features and feature information.

[0012] According to another aspect of an embodiment of the present application, a training method for a multimodal feature extraction model is also provided, including: obtaining training data, wherein the training data includes: multiple types of unimodal data, multimodal data generated based on the unimodal data, the unimodal data includes: image data, text data, voice data, and the multimodal data includes: a combination of any unimodal data; using multiple types of unimodal data as training data to perform fusion training on multiple different types of unimodal feature extraction models to obtain a multimodal feature extraction model, wherein, during the fusion training process, each unimodal feature extraction model is used to perform feature extraction processing on one type of unimodal data.

[0013] According to another aspect of an embodiment of the present application, a data retrieval device is also provided, including: a receiving module, used to receive input data, wherein the input data includes at least one of the following: single-modal data, multimodal combined data obtained by combining multiple single-modal data; a feature extraction module, used to use a multimodal feature extraction model to perform feature extraction processing on the input data to obtain feature information of the input data, wherein the multimodal feature extraction model is obtained by fusion training of multiple different types of single-modal feature extraction models, and each single-modal feature extraction model is used to perform feature extraction processing on one type of single-modal data; a retrieval module, used to determine target data in a retrieval library whose similarity with the feature information is greater than a preset similarity value, and use the target data as a retrieval result.

[0014] According to another aspect of an embodiment of the present application, a non-volatile storage medium is provided, in which a computer program is stored, wherein the above-mentioned data retrieval method is executed by running the computer program on a device where the non-volatile storage medium is located.

[0015] According to another aspect of an embodiment of the present application, an electronic device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the above-mentioned data retrieval method through the computer program.

[0016] According to another aspect of an embodiment of the present application, a computer program product is also provided, including computer instructions, which implement the steps of the above-mentioned data retrieval method when executed by a processor.

[0017] In an embodiment of the present application, input data is received, wherein the input data includes at least one of the following: single-modal data, multimodal combined data obtained by combining multiple single-modal data; a multimodal feature extraction model is used to perform feature extraction processing on the input data to obtain feature information of the input data, wherein the multimodal feature extraction model is obtained by fusion training of multiple different types of single-modal feature extraction models, and each single-modal feature extraction model is used to perform feature extraction processing on one type of single-modal data; in a retrieval library, target data having a similarity with the feature information greater than a preset similarity value is determined, and the target data is used as a retrieval result. By providing a multimodal retrieval model generated by the fusion of multiple single-modal feature extraction models, data retrieval is performed based on the multimodal retrieval model, and the purpose of supporting arbitrary modalities as input and performing arbitrary modal retrieval is achieved, thereby achieving a retrieval method that is adapted to various retrieval tasks and improving the technical effect of the accuracy of retrieval results in complex modal retrieval scenarios, thereby solving the technical problem of being unable to implement complex modal retrieval due to the multimodal retrieval technology in the related art limiting the format of input information. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0019] Figure 1 is a hardware structure block diagram of a computer terminal for implementing a method for data retrieval according to an embodiment of the present application;

[0020] Figure 2 is a flowchart of the steps of a data retrieval method according to an embodiment of the present application;

[0021] Figure 3 is a schematic diagram of determining a total loss function of a multimodal feature extraction model according to an embodiment of the present application;

[0022] Figure 4 is a flow chart of creating a search library according to an embodiment of the present application;

[0023] Figure 5 is a flowchart of the steps of a training method for a multimodal feature extraction model according to an embodiment of the present application;

[0024] Figure 6 is a schematic diagram of a training framework of a multimodal feature extraction model according to an embodiment of the present application;

[0025] Figure 7 It is a structural diagram of a data retrieval device according to an embodiment of the present application. DETAILED DESCRIPTION

[0026] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0028] In order to better understand the embodiments of the present application, the technical terms involved in the embodiments of the present application are explained as follows:

[0029] Everything Detection Model: A deep learning model used to detect multiple targets or objects in images, videos, or other data. This model usually uses convolutional neural networks (CNN) or other deep learning techniques to extract features and classify input data to determine the location and category of different objects or targets in the data; it has a wide range of applications in computer vision, autonomous driving, security monitoring, and other fields. Common everything detection models include the target detection model (You Only Look Once, YOLO), the fast convolutional neural network model (Region-based Convolutional Neural Networks, Faster R-CNN), etc.

[0030] In the related art, the multimodal retrieval system requires users to input image and text pairs for retrieval. This input method limits the flexibility of retrieval and cannot meet the needs of data retrieval under complex modalities. In addition, the machine learning model for implementing data retrieval mainly focuses on global features and local features in image understanding, and lacks enhanced learning of interactive features between targets. This leads to limited performance of the model in understanding the relationship between targets in complex scenes. Therefore, there is a problem of being unable to meet the needs of data retrieval under complex modalities. In order to solve this problem, a relevant solution is provided in the embodiments of the present application, which is described in detail below.

[0031] According to an embodiment of the present application, a method embodiment of data retrieval is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0032] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 FIG. 1 shows a hardware structure block diagram of a computer terminal for implementing a method for data retrieval. Figure 1 As shown, the computer terminal 10 may include one or more (102a, 102b, ..., 102n are used to illustrate) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.

[0033] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuits". The data processing circuits may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuit may be a single independent processing module, or may be incorporated in whole or in part into any of the other components in the computer terminal 10. As described in the embodiments of the present application, the data processing circuit acts as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0034] The memory 104 can be used to store software programs and modules of application software, such as program instructions / data storage devices corresponding to the data retrieval method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, realizing the above-mentioned data retrieval method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0035] The transmission device 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0036] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 .

[0037] The present application embodiment provides a data retrieval method that can be applied in the above operating environment. Figure 2 is a flowchart of the steps of the data retrieval method provided in the embodiment of the present application, such as Figure 2 As shown, the method comprises the following steps:

[0038] Step S202: receiving input data, wherein the input data includes at least one of the following: single-modal data, and multi-modal combined data obtained by combining multiple single-modal data.

[0039] The method provided in the embodiment of the present application supports using data of any modality as input data. Therefore, when performing data retrieval, the input data received in step S202 can be any type of single-modal data, for example, data of one type among images (including pictures and videos), text or audio; or it can be multi-modal combined data generated by combining different types of single-modal data, for example, an image-text pair composed of images and text, an image-audio pair composed of images and audio, and an image-audio pair composed of audio and images; for example, an image-text pair composed of images and text can be in the following form: the text is "find a picture of two dogs running on the beach", and the image is a picture containing an image of two dogs running on the beach. Since the retrieval method provided in the embodiment of the present application supports retrieval of any modality, the method provided in the embodiment of the present application can be widely used in scenarios of complex modal retrieval such as electric energy retrieval.

[0040] Step S204: Use a multimodal feature extraction model to perform feature extraction processing on the input data to obtain feature information of the input data, wherein the multimodal feature extraction model is obtained by fusion training of multiple different types of single-modal feature extraction models, and each single-modal feature extraction model is used to perform feature extraction processing on one type of single-modal data.

[0041] The input data is a retrieval request, which can be a single task request or a batch task request. When the input data contains only one search request, it is a single task request, and when the input data contains multiple search requests, it is a batch task request. In step S204, the general features are extracted by the multimodal feature extraction model, and the data that best matches the retrieval request is screened according to the general features. The multimodal large model applied in the embodiment of the present application is obtained by fusion training of different types of single-modal feature extraction data, wherein different types of single-modal feature extraction models refer to these single-modal feature extraction models used to extract features from different types of single-modal data. For example, when the input data is the text "Find a picture of two dogs running on the beach" and a picture of two dogs running on the beach, the text encoder in the multimodal feature extraction model (a unimodal feature extraction model that extracts features from text data) is used to extract features from the text, and at the same time, the image feature extraction model in the multimodal feature extraction model (a unimodal feature extraction model that extracts features from image data) is used to extract features; the features extracted from the text data and the features extracted from the image are further aligned. In addition, if there is voice data in the input data, the multimodal feature extraction model also uses the voice encoder (a unimodal feature extraction model that extracts features from voice data) to extract features.

[0042] In step S204, the multimodal feature extraction model may be loaded into the memory, for example, the raw data of the multimodal feature extraction model may be loaded from the non-volatile memory into the volatile memory, so that the processor runs the multimodal feature extraction model. The raw data of the multimodal feature extraction model refers to unprocessed data, and generally includes parameters and structural data of the multimodal feature extraction model. The structural data may be a calculation relationship based on the parameters, such as a forward propagation calculation relationship between intermediate layers and neurons. Specifically, the structural data may include code related to the structure of the first neural network, such as code for performing related calculations between intermediate layers and neurons.

[0043] In one embodiment, an area for loading a multimodal feature extraction model can be divided in the memory, which may include a structure data storage area and a parameter storage area. The structure data storage area is used to store structure-related codes, and the parameters referenced by it can point to the address of a specific parameter in the parameter storage area through a pointer. During the training process of the multimodal feature extraction model, it may be necessary to frequently update the parameters, and the parameter values ​​in the parameter storage area can be updated.

[0044] Optionally, the multimodal feature extraction model is trained in the following manner: obtaining training data, wherein the training data includes: multiple types of unimodal data, multimodal data generated based on unimodal data, unimodal data includes: image data, text data, voice data, and multimodal data includes: any combination of unimodal data; determining a training data set corresponding to a multimodal training task, wherein the training data set includes: a first type of training data set composed of unimodal data of the same type, and a second type of training data set generated based on different types of unimodal data; using a unimodal feature extraction model corresponding to the training data set to extract features from the training data set corresponding to the training data set to obtain multiple types of unimodal features, wherein the type of unimodal features corresponds to the type of unimodal data; training a unimodal feature extraction model based on the multiple types of unimodal features, and updating the weights of the multiple types of unimodal features.

[0045] As mentioned in step S204, the multimodal feature extraction model in the embodiment of the present application is obtained by fusion training of different types of single-modal feature extraction models. Before the fusion training begins, a large amount of image data, text data and voice data, as well as any combination thereof, are collected as training data. In the embodiment of the present application, the similarity relationship transfer theory is used for fusion training to ensure that the trained multimodal feature extraction model can be extended to any modality. When the similarity relationship transfer theory is used for fusion training, the training process includes two training stages: the comparison learning stage and the transfer learning stage. In the contrastive learning stage, different types of unimodal data are used to train different types of unimodal feature extraction models; in the transfer learning stage, different multimodal combination data are used to perform overall training on the multimodal feature extraction model; therefore, in the fusion training process, it is necessary to distinguish the training data, and classify all the training data into unimodal feature extraction model training data used in the training process of different types of unimodal feature extraction models, and multimodal feature extraction model training data used in the training process of multimodal feature extraction models. For example, different types of unimodal data are classified into one data set (i.e., the first type of training set), and the first type of training set is used as the training data of different unimodal feature extraction models in the contrastive learning stage; the multimodal data generated according to different types of unimodal data are classified into another data set (i.e., the second type of training data set), and the second type of training set is used to train the multimodal feature extraction model in the transfer learning stage. During the training process, in the contrastive learning stage, different types of unimodal data are used to train the corresponding unimodal feature extraction model, and the weight of each type of unimodal feature during retrieval is updated when executing the retrieval method; for example, during the training process, the image feature extraction model is used to extract image features, the text feature extraction model is used to extract text features, and the speech feature extraction model is used to extract speech features. Then, the model is trained through contrastive loss (the loss function of the contrastive learning stage) to minimize the distance between features corresponding to different modal data.

[0046] According to some optional embodiments of the present application, during the process of training the multimodal feature extraction model, the method also includes: determining a total loss function based on multiple categories of unimodal features, and updating the weights of a class of unimodal features in the multiple categories of unimodal features based on the total loss function, until the total change of the total loss function is less than or equal to a preset change, determining that the training is completed, and obtaining a multimodal feature extraction model, wherein the total change of the total loss function is the difference between the previous total loss function and the next total loss function, the previous total loss function is the total loss function determined by the Nth training, and the next total loss function is the total loss function determined by the N+1th training, and when updating the weights of one class of unimodal features, other classes of unimodal features in the multiple classes of unimodal features remain unchanged.

[0047] As mentioned in the previous embodiment, when training the model based on the similarity relationship transfer theory, the training process is divided into a comparative learning stage and a transfer learning stage. In this embodiment, in the transfer learning stage, when multimodal data generated by combining different types of single-modal data is used for training in the transfer learning stage, the degree of change of the overall loss function (i.e., the total loss function) of the multimodal feature extraction model is observed to determine whether the training is completed. Specifically, the degree of change of the total loss function is represented by the difference in the total loss function generated by two adjacent trainings (i.e., the Nth training and the (N+1)th training). When the difference between the two total loss functions in two adjacent trainings (i.e., the total change) is less than or equal to the preset change, it indicates that continued training has little effect on the accuracy of the model output results, that is, the model converges, and the model training is terminated at this time. When the total loss function difference (i.e., the total change) is greater than the preset change, the weight of a certain type of unimodal feature is adjusted, and the weights of other types of unimodal features are kept unchanged. By adjusting the weights of the unimodal features, the influence of the unimodal features in the retrieval process is changed. For example, the matching priority of the unimodal features can be defined by the weights of the unimodal features. During the retrieval process, a preliminary retrieval result is first matched based on a unimodal feature with a higher weight, and then the retrieval result is further matched based on another unimodal feature with a lower weight in the preliminary retrieval results obtained above, until each unimodal feature has been used as a retrieval condition once, the retrieval is ended, and the final retrieval result is output. As mentioned in the previous embodiment, when adjusting the weights of unimodal features in the stage of transfer learning, the weights of one type of unimodal features are adjusted, and the weights of other types of unimodal features are fixed. For example, the image feature extraction model is selected for weight update, while the weights of the text and speech feature extraction models are kept unchanged; then, in this process, only the weights of the image features are updated, while the weights of the text features and the weights of the speech features are fixed; this ensures that while optimizing the image feature extraction, the accuracy of the text and speech features is not affected; it should be noted that in the method provided in the embodiment of the present application, the weights of the text features are preferably fixed, therefore, the multimodal feature extraction model's understanding of semantics can be improved, thereby improving the accuracy of the retrieval results.

[0048] According to some optional embodiments of the present application, the method also includes: determining the feature extraction results corresponding to each type of unimodal data, determining the independent loss function of the unimodal feature extraction model corresponding to each type of unimodal data according to the feature extraction results, and determining the fusion loss function according to multiple feature extraction results; determining a first change in the independent loss function and a second change in the fusion loss function, and when the first change is less than or equal to a preset change, and the second change is less than or equal to a preset change, determining to trigger the execution of the step of determining the total loss function based on multiple unimodal feature extraction results.

[0049] In order to effectively train a multimodal feature extraction model so that it can achieve performance balance and overall optimization when processing different modal data such as images, texts, and voices, the embodiments of the present application provide a training strategy for dynamically adjusting independent loss functions and fusion loss functions, wherein, as mentioned in the above embodiments, the training process of the multimodal feature extraction model is divided into two stages: contrastive learning and transfer learning. In the training of the multimodal feature extraction model, the contrastive learning stage includes an accuracy learning stage and a similarity learning stage. In the accuracy learning stage, the accuracy of feature extraction of each single-modal feature extraction model is determined. In the alignment learning stage (i.e., the similarity learning stage), the similarity of specific modal pairs (multimodal data) is learned; after transitioning to the transfer learning stage, the similarity of other modalities with the modality is learned by fixing the weight of one modality. In order to improve the model's understanding of semantics, the method provided in the embodiments of the present application fixes the weights of data containing semantic information (text data and voice data) in the transition learning stage, and learns the similarity between image data, a single-modal data, and data containing semantic information. In this embodiment, whether the transition from the comparative learning stage to the transfer learning stage can be made is determined by evaluating the accuracy of feature extraction of each unimodal feature extraction model and the similarity of features extracted by multiple unimodal feature extraction models. When evaluating the accuracy of feature extraction of each unimodal feature extraction model, the loss function (i.e., independent loss function) of each unimodal feature extraction model is evaluated separately, and whether the unimodal feature extraction model has converged is determined by whether the difference (i.e., the first change) between the two independent loss functions of the unimodal feature extraction model in two adjacent iterative trainings is less than or equal to a preset change. When evaluating the similarity of features extracted by multiple unimodal feature extraction models, the semantic similarity of multiple unimodal features (i.e., the feature extraction results corresponding to each type of unimodal data) is determined. Specifically, it can be determined by comparing the retrieval results obtained by the model executing the multimodal training task with the actual retrieval results contained in the training data. Therefore, the overall loss function of the multimodal feature extraction model (i.e., the fusion loss function) can be used to evaluate the similarity of the features extracted by the multiple unimodal feature extraction models; the multimodal feature extraction model is determined based on the independent loss functions of multiple unimodal feature extraction models. Therefore, in this embodiment, the overall loss function of the multimodal feature extraction model will also be referred to as the fusion loss function.During the training process, by judging the difference (i.e., the second change) between the overall loss function (i.e., the fusion loss function) of the multimodal feature extraction model in two adjacent iterations and the preset change, it is judged whether the multiple unimodal feature extraction models have converged in learning the similarity of different unimodal feature data; if the unimodal feature extraction model converges in accuracy (i.e., the (first) change of the independent loss function is less than or equal to the preset change), and converges in the similarity of different modal data (i.e., the (second) change of the fusion loss function is less than or equal to the preset change), it is determined that the training of the model can enter the transfer learning stage. In the transfer learning stage, the total loss function is determined according to the multiple unimodal feature extraction models to continue to judge whether the multimodal feature extraction model has converged.

[0050] According to some other optional embodiments of the present application, in the case where the multimodal training task is an image-text retrieval task that indicates determining a retrieval result based on image data and text data, determining a training data set corresponding to the multimodal training task includes: determining a second type of training data set generated based on image data and text data as a training data set corresponding to the multimodal training task; performing feature extraction on the training data set corresponding to the training data set using a unimodal feature extraction model corresponding to the training data set, including: performing feature extraction processing on text data using a text feature extraction model to obtain text features, performing feature extraction processing on image data using an image feature extraction model to obtain first image features, and performing feature extraction processing on a mask image using an image feature extraction model to obtain second image features, wherein the mask image is obtained by masking the image data based on the text data.

[0051] In order to improve the multimodal feature extraction model's ability to understand semantic features, the method provided in the embodiment of the present application adds other modal data processed according to text data as training data during the training process of the model, that is, the generation of multimodal data includes the following two methods: combining different unimodal data (in this embodiment, the multimodal feature data generated according to this combination method is recorded as the first type of multimodal feature data); processing another type of unimodal data according to unimodal data with semantic information to extract multimodal data that conforms to the semantic information (in this embodiment, the multimodal feature data generated according to this processing method is recorded as the second type of multimodal feature data); that is, the training data contains the following types of data: unimodal data (image data, text data, etc.), the above-mentioned first type of multimodal data, and the second type of multimodal data. The above-mentioned unimodal data with semantic information is not only text data, but also speech data has semantic information. The semantic information contained in the speech data can be extracted by converting the speech data into text. In the embodiment of the present application, taking the multimodal training scenario of comparing training image data and text data as an example, the process of training a multimodal feature extraction model using similarity relationship transfer theory is illustrated. As mentioned in the above embodiment, in the training process of the model, the training data set corresponding to the multimodal training task is first determined, and the multimodal feature extraction model uses the training data in the training set to implement the multimodal task. When the multimodal training task is a graphic retrieval task in which the matching degree of the input image data and the matching degree of the text data are both greater than the preset matching degree, the data contained in the corresponding training set is a data set containing multimodal data generated according to the image data and the text data (i.e., the second type of training set). Specifically, the first type of multimodal data contained in the training set corresponding to the graphic retrieval task is a combination of text data and image data expressing the same information; the second type of multimodal data contained is a mask map generated according to the text data and the image data; that is, the data contained in the training set corresponding to the graphic retrieval task is: text data and image data expressing the same meaning, and mask image data generated according to the above text data and image data expressing the same meaning, and only the target area related to the description of the text data is retained in the mask image (data). As can be seen from the above, when the multimodal training task is a text-image retrieval task, a mask map (data) is added to the training data. The mask map is obtained by masking the image data according to the text data. The generation process includes: extracting all masks related to the text using a large segmentation model (for example, the semantic activation model Semantic-SAM) according to the text, setting some pixels outside the mask to zero, and leaving only the masked part of the picture as the text guidance mask map. The text guidance mask map only retains the target information related to the text, discards some irrelevant background noise, so that the extracted features only focus on the mask part information, thereby enhancing the target interaction information in the picture and improving the text understanding ability of the multimodal feature extraction model.In this embodiment, when the multimodal feature extraction model uses the data in its corresponding training set to perform the image and text retrieval task, in the stage of comparative learning, each unimodal feature extraction model extracts features for the unimodal data with the same modality as its own, and obtains the features of the unimodal data. The text feature extraction model (the unimodal feature extraction model corresponding to the text data) only performs feature extraction on the text data, and the result obtained is the text feature; the image feature extraction model (the unimodal feature extraction model corresponding to the image data) only performs feature extraction processing on the image data. In this embodiment of the application, both image data and mask image data belong to image data. The image feature extraction model performs feature extraction processing on the image data to obtain the features of the image data (i.e., the first image features), and performs feature extraction processing on the mask image data to obtain the features of the mask image data (i.e., the second image features).

[0052] Optionally, when the multimodal training task is a picture and text retrieval task, the total loss function is determined according to multiple categories of unimodal features, including: determining a first loss function according to text features and first image features; fusing the first image features and the second image features to obtain a third image feature, and determining a second loss function according to the text features and the third image features; determining the total loss function according to the first loss function, the second loss function and the model parameters.

[0053] Next, still taking the multimodal training scenario of comparing training image data and text data as an example, the method of determining and updating the overall loss function (i.e., the total loss function) of the multimodal feature extraction model in the process of training the multimodal feature extraction model using the similarity transfer theory is illustrated. As mentioned in the above embodiment, the total loss function of the multimodal feature extraction model is jointly determined based on the loss functions of different single modal features; Figure 3 It is a schematic diagram for determining the total loss function of the multimodal feature extraction model, such as Figure 3 As shown, specifically in the scenario where the multimodal feature extraction model performs image and text retrieval tasks, the method for determining the total loss function can be expressed by the following formula: Loss = Loss1 + a*Loss2, where Loss represents the total loss function, Loss1 is the loss calculated by aligning the text feature (T) with the feature (F1) of the original image data (i.e., the first loss function); Loss2 is the loss calculated by aligning the image feature (i.e., the third image feature) obtained by fusion of the feature F2 of the mask image (data) and the feature F1 of the original image data with the text feature (T), and a is a learnable parameter used to balance the contribution of different losses to the total loss function. The value of a can be updated during the training process of the model. First loss function Among them, S i Represents the features of the original image data (equivalent to F1), Sj is the text feature (equivalent to T), N is the number of training samples, i represents the training data used to calculate Loss1; σ is a hyperparameter, and the value of σ is adjusted during the training process to balance the impact of different unimodal features on the output results. Second loss function Among them, k represents the training data used to calculate Loss2, S k It is a feature extracted from the mask image data (equivalent to F2). Through the method provided in this embodiment, the mask image assists image-text matching in the training phase, so that the model pays more attention to features related to the text and improves the understanding of the text. However, in the reasoning phase (i.e., when the multimodal feature extraction model is applied), the input of the mask image is removed, so that the multimodal large model enhances the overall understanding of the text without increasing the time consumption of reasoning, and improves the retrieval effect of the user input text in the retrieval system.

[0054] Step S206, determining target data in the search database whose similarity with the feature information is greater than a preset similarity value, and taking the target data as a search result.

[0055] In an embodiment of the present application, after the multimodal feature extraction model performs feature extraction processing, the data in the search library is intelligently sorted according to the relevance between the data contained in the search library and the search request, and the most matching data is displayed as the search result in priority. The type of data can be graphic content or other types of data (determined according to the search request), so as to improve the accuracy of the search and user satisfaction. In the search library, quantified image features, text features and possible voice features are stored; the relevance between the data contained in the above search library and the search request can be determined by calculating the similarity between the data in the search library and the feature information extracted by the multimodal feature extraction model, and the larger the similarity value, the higher the relevance; when performing intelligent sorting, only the (target) data whose similarity with the feature information extracted by the multimodal feature extraction model is greater than the preset similarity value can be sorted, or all the data in the search library can be sorted; wherein, the preset similarity value can be adjusted according to actual needs to balance the accuracy and recall rate of the search.

[0056] Optionally, the retrieval library is generated by the following method: using different unimodal feature extraction models to extract features from input data to obtain multiple unimodal features; and using an all-things detection model to analyze the input data, determine the objects to be detected contained in the input data, and extract feature information of the objects to be detected, wherein the feature information at least includes: boundary information of the objects to be detected; and generate a retrieval library based on multiple unimodal features and feature information.

[0057] The search library mentioned in step S206 is also constructed based on the input data. The search library can be created when the system / device executing the data search method provided in the embodiment of the present application is in an offline state (it can also be created when it is in an online state). In order to meet the needs of arbitrary target retrieval, the all-things detection model and the unimodal feature extraction model are used in combination in the storage stage of generating the search library. Figure 4 This is a flowchart for creating a search library, such as Figure 4 As shown in the figure, after receiving the input data, the object detection model and the specific task detection models M1 and M2 (i.e., the unimodal feature extraction model) process the input data respectively to obtain their respective detection results (boxes1 and boxes2), wherein boxes1 is the set of target detection frames output by the object detection model for the input data, and each target detection frame contains an object to be detected in the input data; boxes2 is the set of target detection frames output by the unimodal feature extraction model for the input data; by comparing the overlap rate of the target detection frames in boxes1 and boxes2, it is determined whether the input data is correctly extracted; wherein, the overlap rate of any two target detection frames can be determined based on the feature information of the target detection frame in boxes2 (such as the coordinates of the bounding box (i.e., the boundary information), the size of the target detection frame) and the feature information of the target detection frame in boxes1 (i.e., the unimodal feature). If the overlap rate is greater than the preset overlap rate, it is considered that the feature extraction is correct and the target detection frame is retained. If the target detection frame in boxes2 is not in boxes1, the target detection frame and the corresponding feature information are added, and finally all the target detection frames of the image are obtained. Furthermore, when there is image data in the input data, the mask area of ​​the image is determined according to the image data and other types of data in the input data, and the intersection of the mask area and the overlap rate screening result (boxes) is taken to obtain the target image, such as the human body detection box. After merging with the mask area, only the pixel values ​​of the human body are obtained, and the pixel values ​​outside the human body are 0. Using this target image (the target image is the data with feature information stored in the detection library) can reduce the noise caused by background pixels, improve the feature's attention to the target, and thereby improve the intra-class similarity score in the retrieval.

[0058] Through the above steps, it is possible to achieve the technical effect of supporting the input of any modality of user voice, picture and text, and supporting the retrieval of data in any modality. In addition, the retrieved visual information can be output as a whole image or a target image, and the target information in the image can also be output: coordinates and categories, for downstream tasks.

[0059] The data retrieval method provided in the embodiment of the present application can be applied to a variety of scenarios such as image search (input an image, and the system retrieves images with similar content), text search (input text description, and the system retrieves images semantically related to the text description), text + image search (input text and image at the same time, and the system retrieves images related to both), video retrieval (input video or video description, and the system retrieves video clips or videos related to the input) and multi-image search (input multiple images, and the system retrieves images or image sets related to these image sets). For single task requests or batch task requests, common features are extracted through a multimodal feature extraction model, and intelligent sorting is performed based on the relevance of the user's intention and the query results. The most matching image and text content is displayed first, thereby improving the accuracy of retrieval and user satisfaction. In the scenario of text search for images: search for images that meet the text semantics, extract features from the text, project it into the image feature space, return the image with the highest feature similarity, and give a similarity score; it is suitable for precise retrieval by name and cross-modal retrieval. In the scenario of image search for images: find a set of images with semantic similarity to the searched image in the self-built image library, and give a similarity score (comprehensive scene, specific target and other features); it is applicable to various similar image search, related scene search, similar target search and other scenarios. The search results are merged and de-duplicated, and the merging rules are: merge targets (including the whole image and small targets) with the same image name, sort the targets in the image in descending order according to the similarity score, and take the highest confidence score as the image score, and sort the images in descending order according to the score. In other search modal scenarios, the flying search process is the same as the text search or image search process.

[0060] Figure 5 is a flowchart of the steps of the training method of the multimodal feature extraction model provided in the embodiment of the present application, such as Figure 5 As shown, the method comprises the following steps:

[0061] Step S502: Acquire training data, wherein the training data includes: multiple types of unimodal data, and multimodal data generated based on the unimodal data. The unimodal data includes: image data, text data, and voice data. The multimodal data includes: any combination of unimodal data.

[0062] In step S502, the training data of the model is obtained. The multimodal feature extraction model in the embodiment of the present application is obtained by fusion training of different types of single-modal feature extraction models. Before the fusion training begins, a large amount of image data, text data and voice data, as well as any combination thereof, are collected as training data. In the embodiment of the present application, the similarity relationship transfer theory is used for fusion training to ensure that the trained multimodal feature extraction model can be extended to any modality.

[0063] Step S504: Use multiple types of unimodal data as training data to perform fusion training on multiple different types of unimodal feature extraction models to obtain a multimodal feature extraction model, wherein during the fusion training process, each unimodal feature extraction model is used to perform feature extraction processing on one type of unimodal data.

[0064] In step S504, the model is trained according to the similarity relationship transfer theory. Figure 6 It is a schematic diagram of the training framework of the multimodal feature extraction model, such as Figure 6 As shown in the figure, the model is trained using the similarity transfer theory to achieve the expansion of any modality. For example, when comparing the similarity between images and texts, the text weights are fixed to train the similarity between speech and text. The model can also be trained to learn the similarity between image and speech features through the similarity transfer theory. Figure 6 As shown in the figure, a mask map is added to the training image-text pair. The mask map generation process is as follows: all masks related to the text are extracted based on the text using a segmentation model (such as SAM), and some pixels outside the mask are set to zero, leaving only the masked part of the image as the text guidance mask map. The text guidance mask map only retains the target information related to the text, discards some irrelevant background noise, and makes the extracted features focus only on the masked part information, thereby enhancing the target interaction information in the image and improving the text comprehension ability.

[0065] Through the above steps, in the training stage of the multimodal feature extraction model, the step of filtering the feature extraction results of the input data of other modes based on text is added to realize the supervised learning of the multimodal large model features, and improve the ability to understand the multi-target spatial interaction between pictures and texts, thereby improving the ability to understand the multi-target text description in the retrieval and improving the retrieval accuracy. In addition, using the similarity relationship transfer theory to train the multimodal large model can also achieve the technical effect of the model to achieve the expansion of any modality.

[0066] Figure 7 is a structural diagram of a data retrieval device provided according to an embodiment of the present application, such as Figure 7 As shown, the device includes: a receiving module 70, used to receive input data, wherein the input data includes at least one of the following: single-modal data, multimodal combination data obtained by combining multiple single-modal data; a feature extraction module 72, used to use a multimodal feature extraction model to perform feature extraction processing on the input data to obtain feature information of the input data, wherein the multimodal feature extraction model is obtained by fusion training of multiple different types of single-modal feature extraction models, and each single-modal feature extraction model is used to perform feature extraction processing on one type of single-modal data; a retrieval module 74, used to determine target data in the retrieval library whose similarity with the feature information is greater than a preset similarity value, and use the target data as the retrieval result.

[0067] When a data retrieval device is used for retrieval, input data is received by a receiving module 70. Since the method provided in the embodiment of the present application supports retrieval of any modality, the input data can be single-modal data of a single type, or multi-modal combined data obtained by combining multiple types of single-modal data; the receiving module 70 transmits the input data to a feature extraction module 72, and the feature extraction module 72 calls a multi-modal feature extraction model to perform feature extraction processing on the input data to obtain feature information of the input data; the feature extraction module transmits the extracted feature information to a retrieval module 74, and the retrieval module 74 retrieves data matching the feature information in the retrieval library as a retrieval result.

[0068] It should be noted that Figure 7 The preferred implementation of the illustrated embodiment can be found in Figure 2 The relevant description of the illustrated embodiment will not be repeated here.

[0069] An embodiment of the present application further provides a non-volatile storage medium, in which a computer program is stored, wherein the above data retrieval method is executed by running the computer program on a device where the non-volatile storage medium is located.

[0070] The above-mentioned non-volatile storage medium is used to store a program that performs the following functions: receiving input data, wherein the input data includes at least one of the following: single-modal data, multimodal combination data obtained by combining multiple single-modal data; using a multimodal feature extraction model to perform feature extraction processing on the input data to obtain feature information of the input data, wherein the multimodal feature extraction model is obtained by fusion training of multiple different types of single-modal feature extraction models, and each single-modal feature extraction model is used to perform feature extraction processing on one type of single-modal data; determining target data in a retrieval library whose similarity with the feature information is greater than a preset similarity value, and using the target data as a retrieval result.

[0071] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the above data retrieval method through the computer program.

[0072] The processor in the above-mentioned electronic device is used to run a program that performs the following functions: receiving input data, wherein the input data includes at least one of the following: single-modal data, multimodal combination data obtained by combining multiple single-modal data; using a multimodal feature extraction model to perform feature extraction processing on the input data to obtain feature information of the input data, wherein the multimodal feature extraction model is obtained by fusion training of multiple different types of single-modal feature extraction models, and each single-modal feature extraction model is used to perform feature extraction processing on one type of single-modal data; determining target data in a retrieval library whose similarity with the feature information is greater than a preset similarity value, and using the target data as a retrieval result.

[0073] The embodiment of the present application also provides a computer program product, including computer instructions, which implement the steps of the above data retrieval method when executed by a processor.

[0074] It should be noted that the various modules in the above-mentioned data retrieval device can be program modules (for example, a set of program instructions that implement a certain specific function) or hardware modules. For the latter, it can be expressed in the following forms, but is not limited to this: the expression form of each of the above-mentioned modules is a processor, or the functions of each of the above-mentioned modules are implemented by a processor.

[0075] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0076] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0077] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units can be a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0078] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0079] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0080] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the relevant technology or all or part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk, etc. Various media that can store program codes.

[0081] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A data retrieval method, characterized in that: include: Receiving input data, wherein the input data includes at least one of the following: single-modal data, and multi-modal combined data obtained by combining multiple types of single-modal data; Using a multimodal feature extraction model to perform feature extraction processing on the input data to obtain feature information of the input data, wherein the multimodal feature extraction model is obtained by fusion training of multiple different types of single-modal feature extraction models, and each single-modal feature extraction model is used to perform feature extraction processing on one type of single-modal data; Target data having a similarity with the feature information greater than a preset similarity value is determined in the search database, and the target data is used as a search result.

2. The method according to claim 1, characterized in that The multimodal feature extraction model is trained in the following way: Acquire training data, wherein the training data includes: multiple types of the unimodal data, and multimodal data generated according to the unimodal data, the unimodal data includes: image data, text data, and voice data, and the multimodal data includes: any combination of the unimodal data; Determine a training data set corresponding to a multimodal training task, wherein the training data set includes: a first type of training data set consisting of the unimodal data of the same type, and a second type of training data set generated according to different types of the unimodal data; Using the unimodal feature extraction model corresponding to the training data set to extract features from the training data set, a plurality of types of unimodal features are obtained, wherein the types of the unimodal features correspond to the types of the unimodal data; The unimodal feature extraction model is trained according to the multiple categories of unimodal features, and the weights of the multiple categories of unimodal features are updated.

3. The method according to claim 2, characterized in that In the process of training the multimodal feature extraction model, the method further includes: A total loss function is determined based on multiple categories of unimodal features, and the weight of one category of the unimodal features among the multiple categories of unimodal features is updated based on the total loss function until the total change of the total loss function is less than or equal to the preset change, and the training is determined to be completed, so as to obtain the multimodal feature extraction model, wherein the total change of the total loss function is the difference between the previous total loss function and the next total loss function, the previous total loss function is the total loss function determined by the Nth training, and the next total loss function is the total loss function determined by the N+1th training, and when updating the weight of one category of the unimodal features, other categories of the unimodal features among the multiple categories of the unimodal features remain unchanged.

4. The method according to claim 2, characterized in that: In the case where the multimodal training task is an image-text retrieval task indicating determining a retrieval result based on the image data and the text data, determining a training data set corresponding to the multimodal training task includes: determining a second type of training data set generated based on the image data and the text data as a training data set corresponding to the multimodal training task; The unimodal feature extraction model corresponding to the training data set is used to perform feature extraction on the training data set corresponding to the training data set, including: using a text feature extraction model to perform feature extraction processing on the text data to obtain text features, using an image feature extraction model to perform feature extraction processing on the image data to obtain first image features, and using the image feature extraction model to perform feature extraction processing on a mask image to obtain second image features, wherein the mask image is obtained by masking the image data based on the text data.

5. The method according to claim 4, characterized in that When the multimodal training task is a picture-text retrieval task, the total loss function is determined according to multiple types of single-modal features, including: Determine a first loss function according to the text feature and the first image feature; fusing the first image feature and the second image feature to obtain a third image feature, and determining a second loss function according to the text feature and the third image feature; The total loss function is determined according to the first loss function, the second loss function and model parameters.

6. The method according to claim 3, characterized in that The method further comprises: Determine a feature extraction result corresponding to each type of the unimodal data, determine an independent loss function of a unimodal feature extraction model corresponding to each type of the unimodal data according to the feature extraction result, and determine a fusion loss function according to a plurality of the feature extraction results; Determine a first change in the independent loss function and a second change in the fusion loss function. When the first change is less than or equal to a preset change, and the second change is less than or equal to the preset change, determine to trigger the step of determining the total loss function based on the extraction results of the multiple single-modal features.

7. The method according to claim 1, characterized in that The search library is generated by the following method: Using different unimodal feature extraction models to extract features from the input data to obtain multiple unimodal features; and, The input data is analyzed using an all-things detection model to determine the objects to be detected contained in the input data, and feature information of the objects to be detected is extracted, wherein the feature information at least includes: boundary information of the objects to be detected; The search library is generated according to the plurality of the unimodal features and the feature information.

8. A training method for a multimodal feature extraction model, characterized in that: include: Acquire training data, wherein the training data includes: multiple types of unimodal data, and multimodal data generated from the unimodal data, the unimodal data includes: image data, text data, and voice data, and the multimodal data includes: any combination of the unimodal data; The multimodal feature extraction model is obtained by using multiple types of the unimodal data as training data to perform fusion training on multiple different types of unimodal feature extraction models, wherein during the fusion training, each of the unimodal feature extraction models is used to perform feature extraction processing on one type of unimodal data.

9. A data retrieval device, characterized in that: include: A receiving module, configured to receive input data, wherein the input data includes at least one of the following: single-modal data, and multi-modal combined data obtained by combining a plurality of single-modal data; A feature extraction module, used to perform feature extraction processing on the input data using a multimodal feature extraction model to obtain feature information of the input data, wherein the multimodal feature extraction model is obtained by fusion training of multiple different types of single-modal feature extraction models, and each single-modal feature extraction model is used to perform feature extraction processing on one type of single-modal data; The retrieval module is used to determine target data in the retrieval library whose similarity with the feature information is greater than a preset similarity value, and use the target data as the retrieval result.

10. A non-volatile storage medium, characterized in that: The non-volatile storage medium stores a computer program, wherein the data retrieval method according to any one of claims 1 to 7 is executed by running the computer program on the device where the non-volatile storage medium is located.

11. An electronic device comprising a memory and a processor, characterized in that: The memory stores a computer program, and the processor is configured to execute the data retrieval method according to any one of claims 1 to 7 through the computer program.

12. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by a processor, the steps of the data retrieval method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Cross-modal retrieval method, system and equipment based on supervised comparison

    CN113239214A

  • Cross-modal retrieval method and device and storage medium

    CN114861016A

  • Encoder training method, encoder searching method, encoder training device, encoder searching device and computer equipment

    CN116956008A

  • Information retrieval method and device, equipment, program product and storage medium

    CN116975340A

  • Cross-modal molecular information retrieval method, system, equipment and medium

    CN118069900A