Multimodal retrieval method, device and storage medium
By using multiple feature extraction models to extract and retrieve features from multimodal data, the problem of users struggling to describe complex search intents on online merchants and short video websites is solved, resulting in more accurate and richer search results and personalized feedback.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ALIBABA (CHINA) CO LTD
- Filing Date
- 2022-11-02
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies are insufficient to meet users' multimodal retrieval needs, especially in online merchants and short video websites. Users find it difficult to describe complex search intents with a single query, and one-time recall results do not meet personalized needs and lack real-time feedback.
It employs multiple feature extraction models to extract features from data of different modalities, combines multiple query data to describe retrieval intent, extracts feature vectors from query and retrieval data through multiple feature extraction models, and retrieves the target dataset from the retrieval database, supporting interactive multimodal retrieval with multiple representations and multiple queries.
It achieves more accurate and richer search results, better meets users' multimodal search needs, and provides real-time personalized feedback.
Smart Images

Figure CN115878874B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a multimodal retrieval method, device, and storage medium. Background Technology
[0002] With the development of technologies such as mobile internet, more and more businesses are providing services to users online, such as online merchants, short video websites or short video publishers, and so on.
[0003] Users often have data search needs when using these applications. Now, users are no longer limited to single-modal text-to-image searches; in more scenarios, they have multi-modal search needs such as image-to-image and image-to-video searches. How to meet these search needs and provide search solutions that yield more accurate results is a pressing issue that needs to be addressed. Summary of the Invention
[0004] This invention provides a multimodal retrieval method, apparatus, device, and storage medium for achieving accurate data retrieval.
[0005] In a first aspect, embodiments of the present invention provide a multimodal retrieval method, the method comprising:
[0006] Obtain multiple query data corresponding to different modalities, wherein the multiple query data correspond to the same retrieval intent;
[0007] Multiple feature vectors corresponding to each query data are extracted by using multiple feature extraction models. The feature extraction models are used to extract features from data of different modalities. The multiple feature extraction models are different neural network models.
[0008] Multiple feature vectors corresponding to each piece of search data in the search database are obtained, and the multiple feature vectors corresponding to each piece of search data are extracted by the various feature extraction models respectively.
[0009] Based on the multiple feature vectors corresponding to each of the multiple query data and the multiple feature vectors corresponding to each of the multiple search data in the search database, the target search dataset corresponding to the multiple query data is retrieved from the search database.
[0010] Secondly, embodiments of the present invention provide a multimodal retrieval device, the device comprising:
[0011] The query acquisition module is used to acquire multiple query data corresponding to different modalities, wherein the multiple query data correspond to the same retrieval intent;
[0012] The feature extraction module is used to extract multiple feature vectors corresponding to each query data through multiple feature extraction models; and to obtain multiple feature vectors corresponding to each search data in the search database. The multiple feature vectors corresponding to each search data are extracted through the multiple feature extraction models, wherein the feature extraction models are used to extract features from data of different modalities, and the multiple feature extraction models are different neural network models.
[0013] The data retrieval module is used to retrieve the target retrieval dataset corresponding to the multiple query data in the retrieval database based on the multiple feature vectors corresponding to each of the multiple query data and the multiple feature vectors corresponding to each of the multiple retrieval data in the retrieval database.
[0014] Thirdly, embodiments of the present invention provide an electronic device, including: a memory, a processor, and a communication interface; wherein, the memory stores executable code, and when the executable code is executed by the processor, the processor performs the multimodal retrieval method as described in the first aspect.
[0015] Fourthly, embodiments of the present invention provide a non-transitory machine-readable storage medium storing executable code, wherein when the executable code is executed by a processor of an electronic device, the processor performs the multimodal retrieval method as described in the first aspect.
[0016] Fifthly, embodiments of the present invention provide a multimodal retrieval method, the method comprising:
[0017] The receiving terminal device triggers a request by calling a set service, the request including multiple query data corresponding to different modalities, the multiple query data corresponding to the same search intent;
[0018] The following steps are performed using the processing resources corresponding to the configured service:
[0019] Multiple feature vectors corresponding to each query data are extracted by using multiple feature extraction models. The feature extraction models are used to extract features from data of different modalities. The multiple feature extraction models are different neural network models.
[0020] Multiple feature vectors corresponding to each search data in the retrieval database are obtained, and the multiple feature vectors corresponding to each search data are extracted respectively through the various feature extraction models.
[0021] Based on the multiple feature vectors corresponding to each of the multiple query data and the multiple feature vectors corresponding to each of the multiple search data in the retrieval database, the target retrieval dataset corresponding to the multiple query data is retrieved from the retrieval database.
[0022] The target retrieval dataset is sent to the terminal device for display.
[0023] Sixthly, embodiments of the present invention provide a multimodal retrieval method applied to a virtual reality device, the method comprising:
[0024] The search input interface is displayed, which is used to input query data in different modalities;
[0025] Acquire multiple query data corresponding to different modalities input through the search input interface, wherein the multiple query data correspond to the same search intent;
[0026] Multiple feature vectors corresponding to each query data are extracted by using multiple feature extraction models. The feature extraction models are used to extract features from data of different modalities. The multiple feature extraction models are different neural network models.
[0027] Multiple feature vectors corresponding to each search data in the retrieval database are obtained, and the multiple feature vectors corresponding to each search data are extracted respectively through the various feature extraction models.
[0028] Based on the multiple feature vectors corresponding to each of the multiple query data and the multiple feature vectors corresponding to each of the multiple search data in the retrieval database, the target retrieval dataset corresponding to the multiple query data is retrieved from the retrieval database.
[0029] Displays a search recommendation interface containing the target search dataset.
[0030] In the retrieval scheme provided in this invention, for the same retrieval intent, a user can input multiple query data in various modalities, such as text, images, and videos. Furthermore, various feature extraction models of different types are employed to extract features from the query data and the retrieval data in the retrieval database, thereby extracting rich semantic representations of the same data (e.g., a query data and a retrieval data) from different dimensions. Finally, based on the multiple feature vectors corresponding to each of the multiple query data and the multiple feature vectors corresponding to each of the multiple retrieval data, the target retrieval dataset corresponding to the aforementioned multiple query data is retrieved from the retrieval database. This scheme can obtain more accurate retrieval results. Attached Figure Description
[0031] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0032] Figure 1 A flowchart of a multimodal retrieval method provided in an embodiment of the present invention;
[0033] Figure 2 This is a schematic diagram illustrating the application of a multimodal retrieval method provided in an embodiment of the present invention;
[0034] Figure 3 A flowchart illustrating a method for determining a target retrieval dataset according to an embodiment of the present invention;
[0035] Figure 4 A flowchart of a multimodal retrieval method provided in an embodiment of the present invention;
[0036] Figure 5 A flowchart of a multimodal retrieval method provided in an embodiment of the present invention;
[0037] Figure 6 A flowchart of a multimodal retrieval method provided in an embodiment of the present invention;
[0038] Figure 7 This is a schematic diagram illustrating the application of a multimodal retrieval method provided in an embodiment of the present invention;
[0039] Figure 8 This is a schematic diagram of the structure of a multimodal retrieval device provided in an embodiment of the present invention;
[0040] Figure 9 This is a schematic diagram of the structure of an electronic device provided in this embodiment. Detailed Implementation
[0041] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0042] Furthermore, the timing of the steps in the following method embodiments is merely an example and not a strict limitation.
[0043] In current multimodal retrieval scenarios, for complex retrieval needs, users find it difficult to describe their retrieval intent with a simple, single query. Furthermore, one-time recall of search results often fails to meet user needs, and personalized user feedback cannot be responded to in real time. Currently, no suitable retrieval system meets these requirements. To address these issues, this invention provides an interactive multimodal retrieval scheme with multiple representations and multiple queries. The core features are as follows: 1. Multiple different types of input queries can be used to express the same retrieval intent; 2. Multiple feature extraction models are introduced to extract features from query and retrieval data from different dimensions, achieving richer semantic representations; 3. Real-time capture of user interactive feedback and provision of corresponding optimized retrieval results.
[0044] The retrieval scheme provided in this embodiment of the invention can be executed by an electronic device, which can be a user terminal, a user-side server, or a cloud server or virtual machine.
[0045] Figure 1 A flowchart of a multimodal retrieval method provided in an embodiment of the present invention is shown below. Figure 1 As shown, the method includes the following steps:
[0046] 101. Obtain multiple query data corresponding to different modalities, with multiple query data corresponding to the same retrieval intent.
[0047] 102. Multiple feature vectors corresponding to each query data are extracted through various feature extraction models. The feature extraction models are used to extract features from data of different modalities.
[0048] 103. Obtain multiple feature vectors corresponding to each retrieved data in the database. The multiple feature vectors corresponding to each retrieved data are extracted using various feature extraction models.
[0049] 104. Based on the multiple feature vectors corresponding to each of the multiple query data and the multiple feature vectors corresponding to each of the multiple search data in the retrieval database, retrieve the target retrieval dataset corresponding to the multiple query data in the retrieval database.
[0050] The retrieval scheme provided in this embodiment of the invention can be applied to some application scenarios for merchants and enterprises, such as online merchants, short video publishers, video websites, etc.
[0051] To use this search solution, an offline search engine must first be built. The main purpose of building the offline search engine is to construct a search engine from user (such as the aforementioned merchants and video publishers) search data such as images, text, and videos, to facilitate subsequent online searches.
[0052] In summary, the construction of this offline search engine includes: determining multiple feature extraction models from a pre-defined feature extraction model database; for any search data i contained in the search database, extracting features from that search data i using these multiple feature extraction models to obtain multiple feature vectors corresponding to that search data i, where each feature vector corresponds one-to-one with one of the aforementioned feature extraction models. Therefore, based on the multiple feature vectors corresponding to each search data in the search database, a search engine can be constructed.
[0053] As mentioned above, the retrieval database can store retrieval data of different modalities, such as images, text, and videos (i.e., modal data). In this embodiment of the invention, to meet the needs of multimodal retrieval, the aforementioned multiple feature extraction models are neural network models that can be used to extract features from data of different modalities. For example, in e-commerce application scenarios, the aforementioned retrieval data of different modalities includes textual descriptions of products, product images, and videos introducing the products.
[0054] Of course, the search database can also include search data of a certain modality, such as images or videos.
[0055] In a feature extraction model database, one or more feature extraction models applicable to a specific modality of data can be set. Based on the data types involved in the current user's search database, suitable feature extraction models can be selected. For example, if the search database contains both text and image data, the multiple feature extraction models can include at least one feature extraction model corresponding to text data, at least one feature extraction model corresponding to image data, or at least one feature extraction model applicable to both text and image data. Furthermore, if the search database also includes video data, the multiple feature extraction models can also include at least one feature extraction model corresponding to video data, or at least one feature extraction model applicable to both image and video data.
[0056] In practical applications, from a structural perspective, the aforementioned feature extraction models can include various model structures such as convolutional neural network models, recurrent neural network models, transformer models, and 3D convolutional neural network models. From a functional perspective, these feature extraction models can include multimodal feature extraction models, action recognition models, self-supervised feature extraction models, and cross-modal interactive feature extraction models. In practical applications, suitable feature extraction models can be selected from the numerous existing neural network models according to specific needs.
[0057] Each feature extraction model is trained from different dimensions, with the aim of extracting multiple feature vectors from the same data to represent the data with richer information, thereby achieving more accurate and richer retrieval results.
[0058] It should be noted that the multiple feature extraction models in this embodiment of the invention are not strictly limited to being applicable to every piece of retrieved data simultaneously. For example, feature extraction model A may be applicable to both text and image data. If a retrieved data is video data, feature extraction model A may not be used for feature extraction; instead, other feature extraction models suitable for video data may be used. Based on this, the following optional rules can be followed when determining the multiple feature extraction models: for a certain modality of data, at least two feature extraction models can be used for feature extraction; and at least one feature extraction model can be selected to be applicable to data of different modalities simultaneously.
[0059] In this embodiment of the invention, for ease of description, it is assumed that a piece of data can be feature extracted by the above-mentioned multiple (e.g., 8) feature extraction models. In this case, a retrieval data can extract 8 feature vectors.
[0060] Assuming the database currently contains 10,000 search results, and each search result is used to extract 8 feature vectors through 8 feature extraction models, then 80,000 feature vectors will be obtained.
[0061] In an alternative embodiment, the search engine described above can be simply implemented as a vector library consisting of these 80,000 feature vectors.
[0062] In another optional embodiment, to further improve data storage and retrieval efficiency, an algorithm such as Hierarchical Navigable Small World graphs (HNSW) can be used to construct the vector index. Specifically, the HNSW algorithm is used to build a corresponding vector index library for each feature extraction model. To this end, firstly, for any feature extraction model j, the feature vectors extracted from each retrieved data in the retrieval database through that feature extraction model j are obtained, and these feature vectors constitute the feature vector set corresponding to that feature extraction model j. Then, the HNSW algorithm is applied to the feature vector set corresponding to that feature extraction model j to construct the vector index library corresponding to that feature extraction model j. The implementation process of the HNSW algorithm is based on existing related technologies and will not be elaborated here. It is understood that the above method of establishing the retrieval index is only one optional implementation method; in fact, other algorithms can also be used to construct it.
[0063] The above describes the process of building a search engine offline. The key point of this process is to extract and store multiple feature vectors for each search data based on various feature extraction models.
[0064] Then, the online search process can be executed. This online search process mainly retrieves the corresponding search results from the search database based on the query data entered by the user.
[0065] For more complex search needs, it is difficult for users to describe their search intent using a simple, single query. Therefore, the search scheme provided in this invention supports users using multiple query data to describe their search intent. These multiple query data can be of different modalities, such as text, images, videos, and other different types of query data. For example, suppose a user wants to search for a certain style of red dress. The user can input two query data: a reference image showing a blue dress of a certain style, and a text description stating "search for a red dress of this style." Thus, combining these two query data expresses the user's search intent to find a red dress of this style.
[0066] Allowing users to input multiple query data to describe the same search intent is intended to enable users to describe their search intent more accurately with richer query data.
[0067] Similar to data retrieval, for each query data, multiple feature vectors corresponding to that query data are extracted using various feature extraction models, thereby obtaining multiple feature vectors corresponding to each of the multiple query data.
[0068] Then, based on the multiple feature vectors corresponding to each of the multiple query data and the multiple feature vectors corresponding to each of the multiple search data in the search database, the target search dataset corresponding to the multiple query data is retrieved from the search database.
[0069] In an optional embodiment, the process of determining the target retrieval dataset can be implemented as follows:
[0070] Determine the feature vector set corresponding to each of the multiple feature extraction models, wherein the feature vector set corresponding to any feature extraction model includes feature vectors extracted from multiple retrieved data by the respective feature extraction model;
[0071] Based on the similarity between the target feature vector of any query data and the feature vector in the feature vector set corresponding to the target feature extraction model, a retrieval dataset corresponding to the target feature vector of the query data is determined, wherein the target feature vector is any one of the multiple feature vectors corresponding to the query data, and the target feature vector is extracted by the target feature extraction model;
[0072] The target retrieval dataset is determined based on multiple retrieval datasets corresponding to multiple feature vectors of multiple query data, wherein each retrieval dataset corresponds one-to-one with a feature vector of one query data.
[0073] For ease of understanding, combined with Figure 2 For example, suppose there are 8 feature extraction models, and the retrieval database contains 10,000 search results. Each search result extracts 8 feature vectors using the 8 feature extraction models. For each model, a feature vector set consisting of 10,000 corresponding feature vectors can be obtained. For each query, the 8 feature extraction models will extract 8 feature vectors, and each feature vector will generate a recall result: the retrieval dataset corresponding to that feature vector. If a user inputs 4 queries, and each query extracts 8 feature vectors, this will result in 32 feature vectors, leading to 32 retrieved results: 32 retrieval datasets. Ultimately, the target retrieval dataset can include the search results contained in these 32 datasets.
[0074] Taking the target feature vector *fi* of any query data as an example, assuming that *fi* is extracted by the target feature extraction model *Mi*, we can optionally calculate the similarity between the target feature vector *fi* and 10,000 feature vectors in the feature vector set corresponding to the target feature extraction model *Mi*, obtaining 10,000 similarity scores. Then, by setting a top-N threshold, such as N=10, we can determine the top N (10) similarity scores from these 10,000 similarity scores. The retrieval data corresponding to these N similarity scores constitutes the retrieval dataset corresponding to the target feature vector *fi*. In the example where N=10 and multiple query data correspond to 32 feature vectors, we can obtain 32 retrieval datasets, each containing 10 retrieval data. Without considering redundancy, the target retrieval dataset will contain 32*10=320 retrieval data. After obtaining the target retrieval dataset, the retrieval data in the target retrieval dataset can be displayed in the retrieval recommendation interface.
[0075] In another alternative embodiment, such as Figure 3 As shown, the process of determining the target retrieval dataset may include the following steps:
[0076] 301. Determine the feature vector set corresponding to each of the multiple feature extraction models, wherein the feature vector set corresponding to any feature extraction model includes feature vectors extracted from multiple retrieved data by the respective feature extraction model.
[0077] 302. Perform clustering processing on the feature vector set corresponding to the target feature extraction model to determine the central feature vector of multiple clusters corresponding to the target feature extraction model.
[0078] 303. Based on the similarity scores between the target feature vector of any query data and the central feature vectors of the plurality of clusters, determine the target cluster corresponding to the target feature vector; wherein, the target feature vector is any one of the plurality of feature vectors corresponding to any query data, and the target feature vector is extracted by the target feature extraction model.
[0079] 304. Based on the similarity scores between the target feature vector and each feature vector in the target cluster, determine the retrieval dataset corresponding to the target feature vector of any query data.
[0080] 305. Based on multiple retrieval datasets corresponding to multiple feature vectors of multiple query data, determine the target retrieval dataset, wherein each retrieval dataset corresponds one-to-one with a feature vector of a query data.
[0081] In this embodiment, for each feature vector in the feature vector set corresponding to a feature extraction model, a set clustering algorithm (such as k-means algorithm) can be used for clustering to obtain multiple clusters, where each cluster corresponds to a central feature vector. Optionally, the central feature vector can be obtained by calculating the mean of multiple feature vectors contained in a cluster.
[0082] Then, the target feature vector fi of a certain query data is extracted by the target feature extraction model Mi. First, the similarity score between the target feature vector fi and the central feature vectors of each cluster corresponding to the target feature extraction model Mi is calculated, and the target cluster corresponding to the maximum similarity score is determined. Next, the similarity score between the target feature vector fi and each feature vector contained in the target cluster is calculated, and the top N retrieved data or retrieved data with similarity scores greater than a set threshold are determined to constitute the retrieved dataset corresponding to the target feature vector fi.
[0083] This retrieval method can reduce the computational cost of similarity calculation.
[0084] In the solutions provided in the above embodiments, by using multiple representations and multiple queries, and combining the rich semantic information of the query data and the rich semantic information of the retrieved data, it is helpful to achieve more accurate and richer data retrieval results.
[0085] Figure 4 A flowchart of a multimodal retrieval method provided in an embodiment of the present invention is shown below. Figure 4 As shown, the method includes the following steps:
[0086] 401. Obtain multiple search data from the search database, including video data.
[0087] 402. Perform scene category identification on the video data to determine at least one scene category contained in the video data, and segment the video data to obtain video segments corresponding to the at least one scene category respectively.
[0088] 403. Extract multiple feature vectors corresponding to each of the multiple retrieval data through multiple feature extraction models. The multiple retrieval data includes the segmented video segments.
[0089] 404. Obtain multiple query data corresponding to different modalities with the same retrieval intent, perform similar semantic expansion processing on the multiple query data, and extract multiple feature vectors corresponding to the multiple expanded query data through multiple feature extraction models.
[0090] 405. Based on the multiple feature vectors corresponding to the multiple expanded query data and the multiple feature vectors corresponding to the multiple retrieval data, retrieve the target retrieval dataset corresponding to the multiple query data from the retrieval database.
[0091] In this embodiment, when retrieving data from the database, since a single video segment may involve different scenes, to improve retrieval accuracy, scene classification and recognition processing can be performed on the original complete video segment to determine the scene categories contained in the video segment and the timestamp range corresponding to each scene category. This allows the video data to be segmented according to scene categories into individual video segments. Subsequently, for each video segment, features are extracted using multiple feature extraction models.
[0092] When retrieving data from a database that includes text and images, various feature extraction models can be directly used to extract features from it.
[0093] Furthermore, in this embodiment, semantic expansion can be performed based on multiple query data originally input by the user to automatically generate more query data to enrich the description of the user's same search intent. For example, for text-type query data, semantic expansion can be performed using methods such as synonym replacement and similar word replacement.
[0094] By segmenting and semantically expanding the scenarios described above, we can enrich users' query descriptions and improve the accuracy of search results.
[0095] Steps not described in this embodiment can be referred to the relevant descriptions in the foregoing embodiments, and will not be repeated here.
[0096] Figure 5A flowchart of a multimodal retrieval method provided in an embodiment of the present invention is shown below. Figure 5 As shown, the method includes the following steps:
[0097] 501. Obtain multiple query data corresponding to different modalities, with multiple query data corresponding to the same retrieval intent.
[0098] 502. Multiple feature vectors corresponding to each of the multiple query data are extracted through various feature extraction models. The feature extraction models are used to extract features from data of different modalities.
[0099] 503. Obtain multiple feature vectors corresponding to each of the multiple search data in the retrieval database. The multiple feature vectors corresponding to each of the multiple search data are extracted using various feature extraction models.
[0100] 504. Determine the feature vector set corresponding to each of the multiple feature extraction models, wherein the feature vector set corresponding to any feature extraction model includes feature vectors extracted from multiple retrieved data by the respective feature extraction model.
[0101] 505. Based on the similarity between the target feature vector of any query data and the feature vector in the feature vector set corresponding to the target feature extraction model, determine the retrieval dataset corresponding to the target feature vector of the query data, wherein the target feature vector is any one of the multiple feature vectors corresponding to the query data, and the target feature vector is extracted by the target feature extraction model.
[0102] 506. For any retrieval data contained in multiple retrieval datasets corresponding to multiple feature vectors of multiple query data, determine the sorting factor of the retrieval data.
[0103] 507. Based on the ranking factors of each search data in multiple search datasets, determine the ranking score of each search data in multiple search datasets, and sort the search data according to the ranking scores of each search data in multiple search datasets to obtain a target search dataset with ranking results.
[0104] This embodiment provides a scheme for sorting the search data contained in the target search dataset.
[0105] Specifically, after obtaining multiple retrieval datasets corresponding to multiple feature vectors of multiple query data (e.g., 32 retrieval datasets corresponding to 32 feature vectors in the example above, with each retrieval dataset containing 10 retrieval data), the ranking score of each retrieval data contained in these retrieval datasets is determined. Thus, the retrieval data (i.e., 320 retrieval data in the example above) are sorted from high to low according to their respective ranking scores to obtain the target retrieval dataset with the ranking result. The retrieval data is then displayed in the retrieval recommendation interface according to this ranking.
[0106] The ranking score for each retrieved data is determined based on at least one ranking factor.
[0107] The ranking factor for any search data includes at least one of the following: the number of times the search data appears in the plurality of search datasets, the similarity score ranking of the search data in the plurality of search datasets, and the similarity score of the search data in the plurality of search datasets.
[0108] To facilitate understanding, let's take an example. Continuing from the example above, suppose we obtain 32 search datasets based on multiple feature vectors corresponding to various query data. Any search data x exists in 3 of these search datasets, meaning search data x appears 3 times across all search datasets. Furthermore, suppose the similarity scores (normalized scores) of search data x in these 3 search datasets are 0.7, 0.8, and 0.6, respectively. Also suppose the rankings of search data x's similarity scores in these 3 search datasets are 4th, 3rd, and 8th, respectively. Optionally, we can calculate the average similarity score as (0.7 + 0.8 + 0.6) / 3 = 0.7, and the average ranking as (4 + 3 + 8) / 3 = 5. As mentioned above, the similarity score of search data x in a particular search dataset is determined during the similarity score calculation process described above.
[0109] After obtaining the occurrence frequency, average ranking, and average similarity score of the search data x, a dynamic weighting algorithm can be used to determine the weights of each of these factors. Then, the factors are weighted and summed according to the determined weights, and the result is used as the ranking score of the search data x.
[0110] Optionally, a weight determination model can be pre-trained using a neural network model of a certain structure. This weight determination model is used to determine the weights corresponding to various input information (such as the occurrence frequency, average ranking, and average similarity score of the search data x mentioned above).
[0111] The above sorting process allows search data that appears more frequently and better matches the user's search needs to be displayed more prominently in the multi-channel retrieval data.
[0112] Figure 6 A flowchart of a multimodal retrieval method provided in an embodiment of the present invention is shown below. Figure 6 As shown, the method includes the following steps:
[0113] 601. Obtain multiple query data corresponding to different modalities, with multiple query data corresponding to the same retrieval intent.
[0114] 602. Multiple feature vectors corresponding to each query data are extracted through various feature extraction models. The feature extraction models are used to extract features from data of different modalities.
[0115] 603. Obtain multiple feature vectors corresponding to each retrieved data in the retrieval database. The multiple feature vectors corresponding to each retrieved data are extracted using various feature extraction models.
[0116] 604. Based on the multiple feature vectors corresponding to each of the multiple query data and the multiple feature vectors corresponding to each of the multiple search data in the retrieval database, retrieve the target retrieval dataset corresponding to the multiple query data in the retrieval database.
[0117] 605. Display the search recommendation interface containing the target search dataset.
[0118] 606. In response to a reordering operation triggered for target retrieval data in the target retrieval dataset, the target retrieval data is used as query supplementary data. Multiple feature vectors of the query supplementary data are obtained by extracting them through multiple feature extraction models. Based on the multiple feature vectors corresponding to the query supplementary data and the multiple feature vectors corresponding to the multiple retrieval data, the retrieval dataset corresponding to the query supplementary data is retrieved from the retrieval database. The target retrieval dataset is updated with the retrieval dataset corresponding to the query supplementary data.
[0119] In this embodiment, each search data in the search recommendation interface can be associated with a specific operation item: the rearrangement operation. When a user selects a search data (referred to as the target search data) and triggers the rearrangement operation, it means that the user wants to use this search data as new query data to supplement the multiple query data initially input. At this time, the target search data is used as query supplementary data, and multiple feature vectors of the query supplementary data are extracted by various feature extraction models. Then, based on the multiple feature vectors corresponding to the query supplementary data and the multiple feature vectors corresponding to each of the multiple search data, the search dataset corresponding to the query supplementary data is retrieved from the search database. The retrieval process is similar to the process in the previous embodiment of "determining the search dataset corresponding to the target feature vector of any query data based on the similarity between the target feature vector of any query data and the feature vector set corresponding to the target feature extraction model", and will not be described again here.
[0120] After obtaining the search dataset corresponding to the supplementary data of the query, update the target search dataset with the search dataset corresponding to the supplementary data of the query.
[0121] In one optional embodiment, if the retrieved data in the target retrieval dataset is not sorted as described above, the update can be performed by simply adding it to the target retrieval dataset.
[0122] In another optional embodiment, if the above sorting process is taken into account, the update is manifested as: recalculating the sorting scores of each retrieval data in the multiple retrieval datasets obtained based on the original multiple query data and the retrieval datasets obtained based on the query supplementary data, and then sorting according to the calculated sorting scores to obtain the target retrieval dataset after resorting.
[0123] The multimodal retrieval method provided in this invention can be executed in the cloud, where multiple computing nodes (cloud servers) can be deployed. Each computing node has processing resources such as computing and storage. In the cloud, multiple computing nodes can be organized to provide a certain service; of course, a single computing node can also provide one or more services. The cloud provides this service by providing an external service interface, which users call to use the corresponding service.
[0124] According to the solution provided in this embodiment of the invention, the cloud can provide a service interface with a defined service (multimodal retrieval service). Users invoke this service interface through their terminal devices to trigger a multimodal retrieval request to the cloud. This request includes multiple query data corresponding to different modalities, and these multiple query data correspond to the same retrieval intent. The cloud determines the computing node that responds to the request and utilizes the processing resources in that computing node to perform the following steps:
[0125] Multiple feature vectors corresponding to each query data are extracted by using multiple feature extraction models. The feature extraction models are used to extract features from data of different modalities. The multiple feature extraction models are different neural network models.
[0126] Multiple feature vectors corresponding to each search data in the retrieval database are obtained, and the multiple feature vectors corresponding to each search data are extracted respectively through the various feature extraction models.
[0127] Based on the multiple feature vectors corresponding to each of the multiple query data and the multiple feature vectors corresponding to each of the multiple search data in the retrieval database, the target retrieval dataset corresponding to the multiple query data is retrieved from the retrieval database.
[0128] The target retrieval dataset is sent to the terminal device for display.
[0129] The above execution process can be referred to the relevant descriptions in the other embodiments mentioned above, and will not be repeated here.
[0130] For ease of understanding, combined with Figure 7 To illustrate with an example. Users can... Figure 7 The terminal device E1 illustrated in the diagram calls a multimodal retrieval service to upload multiple query data in different modalities to indicate a user's search intent. The service interface for users to call this service can take the form of a Software Development Kit (SDK) or an Application Programming Interface (API). Figure 7 The diagram illustrates the API interface scenario. In the cloud, as shown, assume a multimodal retrieval service is provided by service cluster E2, which includes at least one compute node. Upon receiving the request, service cluster E2 executes the steps described in the preceding embodiments to obtain the target retrieval dataset and sends it back to terminal device E1.
[0131] The terminal device E1 displays the received target retrieval dataset on the interface. It can also receive user interactions and respond accordingly.
[0132] The following describes in detail one or more embodiments of the multimodal retrieval apparatus of the present invention. Those skilled in the art will understand that these apparatuses can be configured using commercially available hardware components through the steps taught in this solution.
[0133] Figure 8 This is a schematic diagram of the structure of a multimodal retrieval device provided in an embodiment of the present invention, as shown below. Figure 8As shown, the device includes: a query acquisition module 11, a feature extraction module 12, and a data retrieval module 13.
[0134] The query acquisition module 11 is used to acquire multiple query data corresponding to different modalities, wherein the multiple query data correspond to the same retrieval intent.
[0135] The feature extraction module 12 is used to extract multiple feature vectors corresponding to each query data through multiple feature extraction models; and to obtain multiple feature vectors corresponding to each search data in the search database. The multiple feature vectors corresponding to each search data are extracted through the multiple feature extraction models, wherein the feature extraction models are used to extract features from data of different modalities, and the multiple feature extraction models are different neural network models.
[0136] Data retrieval module 13 is used to retrieve the target retrieval dataset corresponding to the multiple query data in the retrieval database based on the multiple feature vectors corresponding to each of the multiple query data and the multiple feature vectors corresponding to each of the multiple retrieval data in the retrieval database.
[0137] Optionally, the data retrieval module 13 is specifically used for: determining the feature vector sets corresponding to each of the multiple feature extraction models, wherein the feature vector set corresponding to any feature extraction model includes feature vectors extracted from the multiple retrieved data by the respective feature extraction model; determining the retrieval dataset corresponding to the target feature vector of any query data based on the similarity between the target feature vector of any query data and the feature vectors in the feature vector set corresponding to the target feature extraction model, wherein the target feature vector is any one of the multiple feature vectors corresponding to the query data, and the target feature vector is extracted by the target feature extraction model; and determining the target retrieval dataset based on the multiple retrieval datasets corresponding to the multiple feature vectors of the multiple query data, wherein one retrieval dataset corresponds to one feature vector of one query data.
[0138] Optionally, the data retrieval module 13 is specifically used to: perform clustering processing on the feature vector set corresponding to the target feature extraction model to determine the central feature vectors of multiple clusters corresponding to the target feature extraction model; determine the target cluster corresponding to the target feature vector based on the similarity scores between the target feature vector and the central feature vectors of the multiple clusters; and determine the retrieval dataset corresponding to the target feature vector of any query data based on the similarity scores between the target feature vector and each feature vector in the target cluster.
[0139] Optionally, the data retrieval module 13 is specifically configured to: for any retrieval data included in the plurality of retrieval datasets, determine the ranking factor of the any retrieval data, wherein the ranking factor includes at least one of the following: the number of times the any retrieval data appears in the plurality of retrieval datasets, the similarity score ranking of the any retrieval data in the plurality of retrieval datasets, and the similarity score of the any retrieval data in the plurality of retrieval datasets; determine the ranking score of each retrieval data in the plurality of retrieval datasets based on the ranking factor of each retrieval data in the plurality of retrieval datasets; and sort the retrieval data according to the ranking score of each retrieval data in the plurality of retrieval datasets to obtain a target retrieval dataset with the sorting result.
[0140] Optionally, the device further includes a display module for displaying a search recommendation interface containing the target search dataset.
[0141] The feature extraction module 12 is further configured to: in response to a rearrangement operation triggered for the target retrieval data in the target retrieval dataset, use the target retrieval data as query supplementary data, and obtain multiple feature vectors of the query supplementary data extracted by the various feature extraction models respectively. The data retrieval module 13 is further configured to: retrieve the retrieval dataset corresponding to the query supplementary data from the retrieval database based on the multiple feature vectors corresponding to the query supplementary data and the multiple feature vectors corresponding to each of the multiple retrieval data; and update the target retrieval dataset with the retrieval dataset corresponding to the query supplementary data.
[0142] Optionally, the plurality of retrieved data includes video data. The apparatus further includes a segmentation module, configured to perform scene category recognition on the video data to determine at least one scene category contained in the video data; and to segment the video data to obtain video segments corresponding to the at least one scene category, so as to extract multiple feature vectors corresponding to each video segment.
[0143] Optionally, the query acquisition module 11 is further configured to: receive multiple query data from the original input; and perform semantic expansion processing on the multiple query data.
[0144] Figure 8 The device shown can perform the steps in the foregoing embodiments. For detailed execution process and technical effects, please refer to the description in the foregoing embodiments, which will not be repeated here.
[0145] In one possible design, the above Figure 8 The structure of the multimodal retrieval device shown can be implemented as an electronic device. For example... Figure 9As shown, the electronic device may include: a processor 21, a memory 22, and a communication interface 23. The memory 22 stores executable code, which, when executed by the processor 21, enables the processor 21 to at least implement the multimodal retrieval method provided in the foregoing embodiments.
[0146] In an alternative embodiment, the electronic device described above may be a virtual reality device. This virtual reality device can perform the following methods:
[0147] The search input interface is displayed, which is used to input query data in different modalities;
[0148] Acquire multiple query data corresponding to different modalities input through the search input interface, wherein the multiple query data correspond to the same search intent;
[0149] Multiple feature vectors corresponding to each query data are extracted by using multiple feature extraction models. The feature extraction models are used to extract features from data of different modalities. The multiple feature extraction models are different neural network models.
[0150] Multiple feature vectors corresponding to each search data in the retrieval database are obtained. The retrieval database stores search data of different modalities. The multiple feature vectors corresponding to each search data are extracted by the various feature extraction models.
[0151] Based on the multiple feature vectors corresponding to each of the multiple query data and the multiple feature vectors corresponding to each of the multiple search data in the retrieval database, the target retrieval dataset corresponding to the multiple query data is retrieved from the retrieval database.
[0152] Displays a search recommendation interface containing the target search dataset.
[0153] The aforementioned retrieval database can be a large number of image and video data stored in virtual reality devices.
[0154] In addition, embodiments of the present invention provide a non-transitory machine-readable storage medium storing executable code, which, when executed by a processor of an electronic device, enables the processor to at least implement the multimodal retrieval method provided in the foregoing embodiments.
[0155] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separate. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0156] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of a necessary general-purpose hardware platform, or by a combination of hardware and software. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a computer product. The present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0157] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multimodal retrieval method, characterized in that, include: Obtain multiple query data corresponding to different modalities, wherein the multiple query data correspond to the same retrieval intent; Multiple feature vectors corresponding to each query data are extracted using various feature extraction models. Each feature extraction model is used to extract features from the data of the modality to which each feature extraction model is applicable in the different modalities. The various feature extraction models are different neural network models. Multiple feature vectors corresponding to each piece of search data in the search database are obtained, and the multiple feature vectors corresponding to each piece of search data are extracted by the various feature extraction models respectively. Based on the multiple feature vectors corresponding to each of the multiple query data and the multiple feature vectors corresponding to each of the multiple search data in the search database, the target search dataset corresponding to the multiple query data is retrieved from the search database.
2. The method according to claim 1, characterized in that, The method further includes: Determine the feature vector set corresponding to each of the multiple feature extraction models, wherein the feature vector set corresponding to any feature extraction model includes feature vectors extracted from the multiple retrieved data by the respective feature extraction model; The step of retrieving the target retrieval dataset corresponding to the multiple query data based on the multiple feature vectors corresponding to each of the multiple query data and the multiple feature vectors corresponding to each of the multiple retrieval data in the retrieval database includes: Based on the similarity between the target feature vector of any query data and the feature vector in the feature vector set corresponding to the target feature extraction model, a retrieval dataset corresponding to the target feature vector of the query data is determined. The target feature vector is any one of the multiple feature vectors corresponding to the query data, and the target feature vector is extracted by the target feature extraction model. The target retrieval dataset is determined based on multiple retrieval datasets corresponding to multiple feature vectors of the multiple query data, wherein one retrieval dataset corresponds to one feature vector of one query data.
3. The method according to claim 2, characterized in that, The step of determining the retrieval dataset corresponding to the target feature vector of any query data based on the similarity between the target feature vector of any query data and the feature vectors in the feature vector set corresponding to the target feature extraction model includes: Clustering is performed on the feature vector set corresponding to the target feature extraction model to determine the center feature vectors of multiple clusters corresponding to the target feature extraction model. Based on the similarity scores between the target feature vector and the center feature vectors of the plurality of clusters, the target cluster corresponding to the target feature vector is determined; Based on the similarity scores between the target feature vector and each feature vector in the target cluster, the retrieval dataset corresponding to the target feature vector of any query data is determined.
4. The method according to claim 3, characterized in that, The step of determining the target retrieval dataset based on multiple retrieval datasets corresponding to multiple feature vectors of the multiple query data includes: For any retrieval data contained in the plurality of retrieval datasets, a ranking factor for the retrieval data is determined, wherein the ranking factor includes at least one of the following: the number of times the retrieval data appears in the plurality of retrieval datasets, the similarity score ranking of the retrieval data in the plurality of retrieval datasets, and the similarity score of the retrieval data in the plurality of retrieval datasets. Based on the ranking factors of each search data in the multiple search datasets, determine the ranking score of each search data in the multiple search datasets; Based on the ranking scores of each retrieval data in the multiple retrieval datasets, the retrieval data is sorted to obtain a target retrieval dataset with the sorting results.
5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Display a search recommendation interface containing the target search dataset; In response to a rearrangement operation triggered for target retrieval data in the target retrieval dataset, the target retrieval data is used as query supplementary data, and multiple feature vectors of the query supplementary data are obtained by extracting them through the various feature extraction models respectively. Based on the multiple feature vectors corresponding to the query supplementary data and the multiple feature vectors corresponding to each of the multiple retrieval data, retrieve the retrieval dataset corresponding to the query supplementary data from the retrieval database; The target retrieval dataset is updated with the retrieval dataset corresponding to the query supplementary data.
6. The method according to any one of claims 1 to 4, characterized in that, The multiple retrieval data include: video data; The method further includes: Scene category identification is performed on the video data to determine at least one scene category contained in the video data; The video data is segmented to obtain video segments corresponding to the at least one scene category, so as to extract multiple feature vectors corresponding to each video segment.
7. The method according to any one of claims 1 to 4, characterized in that, The acquisition of multiple query data corresponding to different modalities includes: Receives multiple query data from the original input; Perform semantic expansion processing on the multiple query data.
8. A multimodal retrieval method, characterized in that, include: The receiving terminal device triggers a request by calling a set service, the request including multiple query data corresponding to different modalities, the multiple query data corresponding to the same search intent; The following steps are performed using the processing resources corresponding to the configured service: Multiple feature vectors corresponding to each query data are extracted using various feature extraction models. Each feature extraction model is used to extract features from the data of the modality to which each feature extraction model is applicable in the different modalities. The various feature extraction models are different neural network models. Multiple feature vectors corresponding to each search data in the retrieval database are obtained, and the multiple feature vectors corresponding to each search data are extracted respectively through the various feature extraction models. Based on the multiple feature vectors corresponding to each of the multiple query data and the multiple feature vectors corresponding to each of the multiple search data in the retrieval database, the target retrieval dataset corresponding to the multiple query data is retrieved from the retrieval database. The target retrieval dataset is sent to the terminal device for display.
9. An electronic device, characterized in that, include: The system includes a memory, a processor, and a communication interface; wherein the memory stores executable code, and when the executable code is executed by the processor, the processor performs the multimodal retrieval method as described in any one of claims 1 to 7.
10. A non-transitory machine-readable storage medium, characterized in that, The non-transitory machine-readable storage medium stores executable code, which, when executed by the processor of the cloud server, causes the processor to perform the multimodal retrieval method as described in any one of claims 1 to 7.
11. A multimodal retrieval method, characterized in that, Applied to virtual reality devices, the method includes: The search input interface is displayed, which is used to input query data in different modalities; Acquire multiple query data corresponding to different modalities input through the search input interface, wherein the multiple query data correspond to the same search intent; Multiple feature vectors corresponding to each query data are extracted using various feature extraction models. Each feature extraction model is used to extract features from the data of the modality to which each feature extraction model is applicable in the different modalities. The various feature extraction models are different neural network models. Multiple feature vectors corresponding to each search data in the retrieval database are obtained, and the multiple feature vectors corresponding to each search data are extracted respectively through the various feature extraction models. Based on the multiple feature vectors corresponding to each of the multiple query data and the multiple feature vectors corresponding to each of the multiple search data in the retrieval database, the target retrieval dataset corresponding to the multiple query data is retrieved from the retrieval database. Displays a search recommendation interface containing the target search dataset.
Citation Information
Patent Citations
Cross-modal retrieval method and device and storage medium
CN114861016A