Cross-modal retrieval system and method based on pre-training model and recall ranking

By combining pre-trained models and recall sorting technology in the cross-modal search system, the problem of computing performance bottlenecks in the prior art and the inability to independently represent image and text input signals is solved, and efficient and fast multimodal search is achieved, suitable for large-scale event consultation and news search.

CN114419387BActive Publication Date: 2025-06-20BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111229288.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-21
Publication Date
2025-06-20
Estimated Expiration
2041-10-21

AI Technical Summary

Technical Problem

When existing multimodal retrieval technology processes massive multi-source and multimodal data, there are problems such as computing performance bottlenecks and inability to effectively represent image and text input signals independently.

Method used

A cross-modal retrieval system based on pre-trained model and recall sorting is proposed. Combining the advantages of vector embedding model and cross-encoder model, a rough recall and precise sorting strategy is adopted, and multi-dimensional text information extraction and intelligent image retrieval technology is used to achieve rapid retrieval between multimodals.

Benefits of technology

It realizes efficient and fast cross-modal retrieval, reduces information management costs, improves information search accuracy and efficiency, and is suitable for multimodal automated information retrieval for large-scale event consultation and news search.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114419387B_ABST
    Figure CN114419387B_ABST
Patent Text Reader

Abstract

The present invention proposes a cross-modal retrieval system and method based on a pre-trained model and recall ranking. Among them, the system includes: a multi-dimensional text information extraction module, which is used to provide information support on the text side for the cross-modal retrieval system, expand the semantic representation of text information through different dimensions, and increase the text sample size; an intelligent image retrieval module, which is used for a video intelligent frame extraction module and an image search by image module. Among them, the video intelligent frame extraction module is used to extract several pictures that can best represent the video content from a video, and the image search by image module is used to complete large-scale and high-efficiency image retrieval tasks; a cross-modal retrieval module, which is used to generate a roughly relevant candidate set according to a query item, accurately rank the candidate set, and finally return relevant retrieval results. This system is used to reduce the information management cost, improve the information search accuracy and efficiency, and support multi-modal automated information retrieval for large-scale event consulting and news search.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence. Background Art

[0002] With the development of the Internet, the information in the network no longer presents in a single text form, but develops in a diversified direction. Nowadays, in addition to a large amount of text data, the network also contains data in multiple modalities such as images, videos, and audios that are no less than the number of texts. Facing the huge amount of data generated by the rapidly developing Internet industry, how to quickly and effectively retrieve relevant information among different modality data according to user intentions has great practical value. Currently, the mainstream multi-modal retrieval technologies can be mainly divided into two types. One is the cross-encoder model based on matching function learning. Its main idea is that the text and image features are first fused, and then passed through a hidden layer (neural network), allowing the hidden layer to learn a cross-modal distance function, and finally obtaining the text-image relationship score. This model mainly focuses on fine-grained attention and cross features, and its structure is as Figure 3 ; the other is the vector embedding model based on representation learning. Its main idea is that the text and image features are separately calculated to obtain the final top-level embeddings, and then an interpretable distance function (cosine function, L2 function, etc.) is used to constrain the text-image relationship. This model pays more attention to the representation methods of two different modality signals in the same mapping space, and its structure is as Figure 4 .

[0003] Generally speaking, the model effect of the cross-encoder model is better than that of the vector embedding model because the combination of text and image features can provide more cross-feature information for the hidden layer of the model. However, the main problem of the cross-encoder model is that it cannot use the top-level embeddings to independently represent the input signals of images and texts. In a retrieval and recall scenario with N pictures and M texts as inputs, N*M combined inputs are required to be input into the model to obtain the results of image-to-text or text-to-image search; in addition, during online use, the computing performance is also a major bottleneck. After the feature combination, the hidden layer needs to perform online calculations; due to the extremely large number of cross combinations, it is impossible to pre-store the embedding vectors of text and image signals and use caching for calculations. Therefore, although the cross-encoder model has good effects, it is not the mainstream in practical applications.

[0004] The vector embedding model structure is the current mainstream retrieval structure. Since it separates the signals of two different modalities, images and text, the top-level embeddings of each modality can be calculated separately in the offline stage. When using the stored embeddings online, only the distance between the two modality vectors needs to be calculated. If it is the relevance filtering of sample pairs, only the cosine / Euclidean distance between the two vectors needs to be calculated. If it is online retrieval and recall, the embedding set of one modality needs to be constructed into a retrieval space in advance, and the nearest neighbor retrieval algorithm (such as ANN and other algorithms) is used for searching. The core of the vector embedding model is to obtain high-quality embeddings. However, although the vector embedding model is simple, effective, and widely used, its disadvantages are also obvious. From the model structure, it can be seen that there is basically no interaction between the signals of different modalities. Therefore, it is difficult to learn high-quality embeddings that represent the semantics of the signals, and the accuracy of the corresponding metric space / distance also needs to be improved.

[0005] This proposal aims at the characteristics of dynamic, multi-source, and multi-modal data in the current Internet, and proposes a cross-modal retrieval system based on pre-trained models and recall ranking, which is used to reduce the information management cost, improve the accuracy and efficiency of information search, and support the multi-modal automatic information retrieval of large-scale event consulting and news search. Summary of the Invention

[0006] The present invention aims to solve at least one of the technical problems in the related art to some extent.

[0007] To this end, the first object of the present invention is to propose a cross-modal retrieval system based on pre-trained models and recall ranking, which is used to reduce the information management cost, improve the accuracy and efficiency of information search, and support the multi-modal automatic information retrieval of large-scale event consulting and news search.

[0008] The second object of the present invention is to propose a cross-modal retrieval method based on pre-trained models and recall ranking.

[0009] To achieve the above object, the first aspect embodiment of the present invention proposes a cross-modal retrieval system based on pre-trained models and recall ranking, including: a multi-dimensional text information extraction module, which is used to provide information support on the text side for the cross-modal retrieval system, expand the semantic representation of text information through different dimensions, and increase the text sample size; an intelligent image retrieval module, including a video intelligent frame extraction module and an image search by image module. Among them, the video intelligent frame extraction module is used to extract several pictures that can best represent the video content from a video, and the image search by image module is used to complete the large-scale and high-efficiency image retrieval task; a cross-modal retrieval module, which is used to generate a candidate set that is roughly relevant according to the query item, accurately rank the candidate set, and finally return the relevant retrieval results.

[0010] The cross-modal retrieval system based on pre-trained models and recall ranking proposed in the embodiments of the present invention addresses the characteristics of cross-modal retrieval data such as being dynamic, multi-source, and multi-modal, as well as the problems existing in the current two mainstream modeling methods. It organically combines the two modeling methods, adopts the idea of rough recall and precise ranking, integrates the advantages of the two solutions, and realizes efficient and fast cross-modal retrieval. In addition, this solution proposes a text query based on inverted index retrieval and a high-dimensional image feature retrieval technology based on color and texture to achieve fast retrieval between multiple modalities and provide a good user experience for users.

[0011] In addition, the cross-modal retrieval system based on pre-trained models and recall ranking according to the above embodiments of the present invention may also have the following additional technical features:

[0012] Further, in an embodiment of the present invention, the multi-dimensional text information extraction module includes:

[0013] A voice data processing module for audio extraction and speech recognition based on deep learning;

[0014] A natural language text extension module for obtaining semantic descriptions of the current sentence in different word orders and languages, expanding the existing text data from multiple aspects, and also for obtaining a large number of negative sample data according to fine-grained text analysis.

[0015] Further, in an embodiment of the present invention, the video intelligent frame extraction module is used to extract several pictures that best represent the video content from a video, specifically including:

[0016] Extract each frame of the video to obtain several pictures;

[0017] Map the pictures to a unified LUV color space and calculate the absolute distance between each frame and the previous frame;

[0018] Sort all the extracted frames according to the absolute distance, and the several frames ranked at the front are regarded as the several pictures that best represent the video content.

[0019] Further, in an embodiment of the present invention, the image search by image module is used to complete large-scale and high-efficiency image retrieval tasks, specifically including:

[0020] Extract features from pictures using an image feature extraction technology based on the comparison gap of average gray levels;

[0021] Quickly retrieve the same or similar pictures from the picture database through the fuzzy query function provided by ElasticSearch.

[0022] Further, in an embodiment of the present invention, the cross-modal retrieval module includes:

[0023] The rough recall module uses a Transformer-based multi-modal pre-training model as a sub-model of the vector embedding model for fast rough recall.

[0024] The precise ranking module uses a Transformer-based multi-modal pre-training model as a sub-model of the cross-encoder model for precise ranking.

[0025] To achieve the above object, another embodiment of the present invention proposes a cross-modal retrieval method based on pre-training models and recall ranking, including the following steps: extracting text information, expanding the semantic representation of the text information through different dimensions, and increasing the text sample size; extracting image information, extracting several pictures that best represent the video content from a video, and retrieving the same or similar pictures from the database; generating a roughly relevant candidate set according to the query item, precisely ranking the candidate set, and finally returning the relevant retrieval result.

[0026] The cross-modal retrieval method based on pre-training models and recall ranking proposed by the embodiments of the present invention aims at the characteristics of cross-modal retrieval data such as dynamics, multiple sources, and multiple modalities, as well as the problems existing in the current two mainstream modeling methods. It organically combines the two modeling methods, adopts the idea of rough recall and precise ranking, integrates the advantages of the two schemes, and realizes efficient and fast cross-modal retrieval; in addition, this scheme proposes a text query based on inverted index retrieval and a high-dimensional image feature retrieval technology based on color and texture to realize fast retrieval between multiple modalities and provide a good user experience for users.

[0027] Further, in an embodiment of the present invention, the extracting text information includes:

[0028] Audio extraction and deep learning-based speech recognition;

[0029] Obtaining semantic descriptions of the current sentence in different word orders and different languages, expanding the existing text data from multiple aspects, and also used to obtain a large amount of negative sample data according to fine-grained text analysis.

[0030] Further, in an embodiment of the present invention, the extracting several pictures that best represent the video content from a video includes:

[0031] Extracting each frame of the video to obtain several pictures;

[0032] Mapping the pictures to a unified LUV color space and calculating the absolute distance between each frame and the previous frame;

[0033] Sorting all the extracted frames according to the absolute distance, and the several frames ranked at the front are regarded as the several pictures that best represent the video content.

[0034] Further, in an embodiment of the present invention, retrieving the same or similar pictures from the database includes:

[0035] Performing feature extraction on the pictures by using a picture feature extraction technology based on the comparison gap of average gray levels;

[0036] Retrieving the same or similar pictures from the picture database quickly through the fuzzy query function provided by ElasticSearch.

[0037] Further, in an embodiment of the present invention, generating a roughly relevant candidate set according to the query item and performing precise sorting on the candidate set includes:

[0038] Adopting a multi-modal pre-training model based on transformer as a sub-model of the vector embedding model for quick rough recall;

[0039] Utilizing a multi-modal pre-training model based on transformer as a sub-model of the cross-encoder model for precise sorting. Description of the Drawings

[0040] The above and / or additional aspects and advantages of the present invention will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where:

[0041] Figure 1 It is a schematic flowchart of a cross-modal retrieval system based on a pre-training model and recall sorting provided by an embodiment of the present invention.

[0042] Figure 2 It is a schematic flowchart of a cross-modal retrieval method based on a pre-training model and recall sorting provided by an embodiment of the present invention.

[0043] Figure 3 It is a schematic diagram of a cross-encoding model provided by an embodiment of the present invention.

[0044] Figure 4 It is a schematic diagram of a vector embedding model provided by an embodiment of the present invention.

[0045] Figure 5 It is a schematic diagram of the technical solution provided by an embodiment of the present invention.

[0046] Figure 6 It is a schematic diagram of a voice data processing module provided by an embodiment of the present invention.

[0047] Figure 7 It is a schematic diagram of a natural language text expansion module provided by an embodiment of the present invention.

[0048] Figure 8 Schematic diagram of the video intelligent key-frame extraction module provided by the embodiment of the present invention.

[0049] Figure 9 Schematic diagram of the picture feature extraction provided by the embodiment of the present invention.

[0050] Figure 10 Schematic diagram of the retrieval architecture diagram provided by the embodiment of the present invention.

[0051] Figure 11 Schematic diagram of the rough recall module provided by the embodiment of the present invention.

[0052] Figure 12 Schematic diagram of the precise sorting module provided by the embodiment of the present invention. Detailed implementation manners

[0053] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.

[0054] The cross-modal retrieval system and method based on a pre-trained model and recall sorting according to the embodiments of the present invention will be described below with reference to the accompanying drawings.

[0055] Figure 1 Schematic flow diagram of a cross-modal retrieval system based on a pre-trained model and recall sorting provided by the embodiment of the present invention.

[0056] As Figure 1 shown, the cross-modal retrieval system based on a pre-trained model and recall sorting includes the following modules: a multi-dimensional text information extraction module 10, an intelligent image retrieval module 20, and a cross-modal retrieval module 30.

[0057] Among them, the multi-dimensional text information extraction module 10 is used to provide information support on the text side for the cross-modal retrieval system, expand the semantic representation of text information through different dimensions, and increase the text sample size; the intelligent image retrieval module 20 includes a video intelligent key-frame extraction module 201 and an image search by image module 202. Among them, the video intelligent key-frame extraction module is used to extract several pictures that best represent the video content from a video, and the image search by image module is used to complete a large-scale and high-efficiency picture retrieval task; the cross-modal retrieval module 30 is used to generate a roughly relevant candidate set according to the query item, precisely sort the candidate set, and finally return the relevant retrieval result. The processing flow of this solution is as Figure 5 shown.

[0058] Further, in an embodiment of the present invention, the multi-dimensional text information extraction module 10 includes:

[0059] A voice data processing module 101 for audio extraction and deep learning-based speech recognition;

[0060] A natural language text extension module 102 for obtaining semantic descriptions of the current sentence in different word orders and different languages, expanding the existing text data from multiple aspects, and also for obtaining a large number of negative sample data according to fine-grained text analysis.

[0061] It can be understood that the multi-dimensional text information extraction module provides text-side information support for the multi-modal retrieval system, mainly expanding the semantic representation of text information through different dimensions and increasing the text sample size. In addition, this module provides sufficient data support for text unimodal retrieval, enriching the data content of the text modality on the one hand and enhancing the correlation between multi-modalities on the other hand.

[0062] Different from conventional text information extraction, the multi-dimensional text information extraction module adopts a method combining text translation and speech recognition, making full use of the advantages of multi-modal data. It performs speech recognition on the audio data in the video and the originally audio data to obtain paired training data; then it performs text translation processing on the overall text data, uses the text semantic information to improve the overall quality of the data, expands the quantity of paired multi-modal associated data, and at the same time, through multi-dimensional natural language analysis, randomly replaces the components in the sentence to form a rich negative sample space and improve the robustness of the model.

[0063] The multi-dimensional text information extraction module can be subdivided into a voice data processing sub-module and a natural language text extension sub-module.

[0064] The voice data processing sub-module mainly includes audio extraction and deep learning-based speech recognition, and its structure is as Figure 6 .

[0065] High-dimensional modalities have high information content. Projecting them onto low dimensions can greatly expand the low-dimensional modality data. Converting high-dimensional modalities (such as video, audio) into low-dimensional modality (text) data can provide a large amount of paired associated data content. Audio extraction can efficiently strip the audio data in the video and quickly provide it to subsequent functions.

[0066] Deep learning-based speech recognition uses the attention mechanism to achieve end-to-end training, performs unified speech recognition on the audio data obtained from each modality to obtain low-modal (i.e., text modality) information. The end-to-end model can well form a complete pipeline to provide a large amount of paired data for subsequent text feature extraction. At the same time, the audio features obtained during the deep learning process can support the audio feature content required for the final cross-modal retrieval.

[0067] The data used for cross-modal retrieval model training are all paired related data. Currently, most of this data is obtained by manual labeling, and the public complete data set is difficult to meet the training data volume required for deep learning. The multi-dimensional text information extraction module converts natural language text into multi-language text information through deep learning-based translation to obtain the multi-dimensional semantic representation of the current text data, and then converts it back to the original language to achieve the purpose of unified training language.

[0068] The natural language text expansion submodule mainly obtains the semantic description of the current sentence in different word orders and languages ​​through multi-language translation results, and expands the existing text data from multiple aspects. In addition, natural language processing can also obtain a large amount of negative sample data based on fine-grained text analysis, making the final cross-modal retrieval model more robust and improving the robustness of the model. Its structure is as follows: Figure 7 .

[0069] Furthermore, in one embodiment of the present invention, the video intelligent frame extraction module 201 is used to extract a number of pictures that best represent the video content from a video, specifically including:

[0070] Extract each frame of the video and get several pictures;

[0071] Map the image to a unified LUV color space and calculate the absolute distance between each frame and the previous frame;

[0072] All extracted frames are sorted according to the absolute distance, and the top frames are regarded as the pictures that best represent the video content.

[0073] It is understandable that videos are composed of picture frames, and there is a natural connection between video modality data and picture modality data. To achieve the transition from video modality to picture modality, it is only necessary to extract several representative pictures from the video.

[0074] To complete the intelligent video frame extraction, first extract each frame of the video to obtain several pictures; then map the pictures to the unified LUV color space, calculate the absolute distance between each frame and the previous frame, the larger the distance, the more drastic the change of the frame compared to the previous frame; finally, sort all the extracted frames according to the calculated absolute distance, and the top frames are regarded as the pictures that best represent the video content. Figure 8 shown.

[0075] Furthermore, in one embodiment of the present invention, the image search module 202 is used to complete large-scale and efficient image retrieval tasks, specifically including:

[0076] Extract features from pictures using the picture feature extraction technology based on the comparison gap of average gray levels;

[0077] Through the fuzzy query function provided by ElasticSearch, quickly retrieve the same or similar pictures from the picture database.

[0078] To meet the need to quickly retrieve and return the same or similar pictures according to the pictures input by users, the picture search technology is indispensable. At present, there are problems such as insufficient retrieval speed and narrow retrieval range in a large number of picture retrieval technologies. This solution proposes a picture feature extraction method based on the comparison of average gray levels, and a picture search technology empowered by the Elasticsearch search engine to accelerate, so as to complete large-scale and high-efficiency picture retrieval tasks.

[0079] The retrieval speed greatly affects the retrieval experience. Picture retrieval is different from keyword retrieval, and the computational complexity is significantly increased. To speed up the picture retrieval speed, this solution first converts the RGB three-color picture into a gray picture with 255 gray levels; then appropriately crops the picture to cut off the part that is unlikely to show the picture features, and obtains a gray picture as shown in Figure 9 the figure. To calculate the similarity between pictures, the method of extracting picture features is particularly important. In this solution, 9*9 grid points and their surrounding areas are selected in the picture as shown in Figure 9 the figure, calculate the comparison gap of the average gray level based on these rectangular areas, and quantify the comparison gap as the picture feature for storage. This picture feature extraction method can represent a picture with only an 81*8 matrix, so the speed is relatively fast when calculating the similarity between pictures; and because the storage space required for a single picture is small, large-scale picture search tasks can be realized.

[0080] To further improve the retrieval speed, this solution implements the picture retrieval task based on ElasticSearch. Using the above picture feature extraction method, store the picture features in ElasticSearch to build a picture retrieval database. The picture database based on ElasticSearch is different from the traditional database. It uses the inverted index mechanism, which greatly improves the retrieval speed. When the user inputs a picture or a picture obtained by intelligent frame extraction of a video, first extract the features, and then through the fuzzy query function provided by ElasticSearch, quickly retrieve the same or similar pictures from the picture database.

[0081] Furthermore, in an embodiment of the present invention, the cross-modal retrieval module 30 includes:

[0082] The rough recall module 301 uses a multi-modal pre-trained model based on Transformer as a sub-model of the vector embedding model to perform fast rough recall;

[0083] The precise ranking module 302 uses a multi-modal pre-trained model based on Transformer as a sub-model of the cross-encoder model to perform precise ranking.

[0084] As mentioned above, both existing mainstream modeling schemes have deficiencies. This scheme combines the two schemes organically for the first time, adopting the innovative ideas of rough recall and precise ranking, which can improve the retrieval efficiency while ensuring the retrieval effect. This scheme uses a vector embedding model to perform rough information recall, and then uses a cross-encoder model to precisely rank the recalled information, and finally returns the options with higher rankings that best meet the retrieval requirements. This architecture can utilize existing cross-modal pre-trained models and share parameters between the two models to improve the parameter efficiency of the model. The retrieval architecture is as Figure 10 shown.

[0085] Among them, the rough recall part uses a multi-modal pre-trained model based on Transformer, such as OSCAR, as a sub-model of the vector embedding model to perform fast rough recall.

[0086] As Figure 11 can be seen, the vector embedding model contains two pre-trained sub-models, which respectively process text signals and image signals, but share parameters. Through the two sub-models, signals of different modalities are respectively encoded; then they are mapped into the same high-dimensional multi-modal feature space; finally, standard distance measurement methods, such as Euclidean distance and cosine distance, are used to calculate the similarity between the two signals, and the top-k most similar candidates are selected for precise ranking by the cross-encoder model.

[0087] To make the distributions of the two modalities of the input image i and the text caption c closer in the high-dimensional multi-modal feature space, the corresponding image-text pairs are closely placed in the feature space during training, while the irrelevant sample pairs are placed farther away (the distance is at least more than the boundary value α). Therefore, the triplet loss function is used to represent (the distance measurement method uses cosine distance):

[0088] L EMB (i, c) = max(0, cos(i, c′) - cos(i, c) + α) + max(0, cos(i′, c) - cos(i, c) + α)

[0089] where (i, c) is a positive image-text pair from the training corpus, and c' and t' are negative samples sampled from the training corpus such that the image-text pairs (i, c') and (i', c) do not appear in the corpus.

[0090] Since the model encodes text and image signals independently, during retrieval, only the query text or image needs to be mapped to the same feature space for distance calculation. Therefore, the data in the database can be encoded offline to ensure the efficiency during online retrieval, enabling it to be applied to large-scale data retrieval. However, since the model is not required to learn the fine-grained features of the input, it is only used for quickly recalling the candidate target set, and the cross-encoder model is used for precise ranking.

[0091] For the precise ranking part, a multi-modal pre-trained model based on transformer, such as OSCAR, is used as a sub-model of the cross-encoder model for precise ranking. The precise ranking is as Figure 9 shown.

[0092] As Figure 9 can be seen, the cross-encoder model only uses one pre-trained sub-model, and it is necessary to concatenate the text and image signals and then judge their similarity through a neural network. This solution uses a binary classifier to judge whether the text and image are relevant, and the cross-entropy loss function is used to represent it as:

[0093] L CE (i, c) = -(y log p(i, c) + (1 - y) log(1 - p(i, c)))

[0094] p(i, c) represents the probability that the combination of the input image i and text c is a positive sample (whether it is the correct image-text combination). When (i, c) is a positive sample pair, y = 1; when (i, c) is a negative sample pair, y = 0.

[0095] During retrieval, the top-k candidates roughly recalled are concatenated with the query item in turn, and the similarity probability of each image-text pair is obtained respectively to complete the precise ranking.

[0096] Although the above method usually has high performance and can learn more information from the interaction of the two signals, its computational cost is high because it is necessary to pass each combination through the entire network to obtain the similarity score p(i, c), that is, this method does not utilize any pre-computed representations during retrieval and it is difficult to perform fast retrieval on large-scale data.

[0097] Therefore, the overall process of this sub-module is as Figure 12As shown, first, the vector embedding model is used to quickly select the top-k roughly relevant candidates according to the user's query terms, and then the cross-encoding model is used to accurately rank the candidate set according to the query terms, and finally, the relevant retrieval results are returned to the user. This solution retains both the efficiency of the vector embedding model in retrieving on large-scale data sets and the retrieval accuracy of the cross-encoding model.

[0098] This solution makes full use of the advantages of multi-modal data, adopts the method of text translation combined with speech recognition to improve the overall quality of the data, the quantity of multi-modal associated data, and the robustness of the model; uses the picture feature extraction technology based on the comparison gap of average gray levels to extract features from pictures, and combines with the ElasticSearch search engine to quickly retrieve the picture features, realizing large-scale and efficient picture retrieval; combines the respective advantages of the vector embedding model and the cross-encoder model, and innovatively adopts the strategies of rough recall and accurate ranking during retrieval, realizing fast and effective cross-modal retrieval on large-scale data.

[0099] Compared with the current mainstream cross-modal retrieval technologies, the advantages of this solution are as follows: First, a joint retrieval framework is proposed, which combines the advantages of the fast retrieval speed of the vector embedding model and the good retrieval effect of the cross-encoder model, adopts the strategies of rough recall and accurate ranking during retrieval, realizes fast and effective cross-modal retrieval on large-scale data, and shares the parameters of the two models at the same time to improve the parameter efficiency. This framework is applicable to any cross-modal pre-trained model, enabling this framework to use existing models without the need to train from scratch and having a wide range of application scenarios. Second, by combining multi-dimensional text information extraction and intelligent image retrieval, fast retrieval of single modality is realized, solving the shortcoming that the current mainstream cross-modal retrieval models cannot achieve the retrieval of the same modality information; on the one hand, multi-dimensional text information extraction enriches the information content of the text modality, on the other hand, it enhances the correlation between multi-modalities, and at the same time realizes the conversion from speech to text; intelligent image retrieval realizes the conversion from video modality data to picture modality data, can extract picture features according to information such as the pixels, colors, and textures of pictures, and efficiently retrieves the same or highly similar pictures in the database.

[0100] The cross-modal retrieval system based on pre-trained models and recall ranking proposed in the embodiments of the present invention, aiming at the characteristics of cross-modal retrieval data such as dynamic, multi-source, and multi-modal, as well as the problems existing in the current two mainstream modeling methods, organically combines the two modeling methods, adopts the idea of rough recall and accurate ranking, integrates the advantages of the two solutions, and realizes efficient and fast cross-modal retrieval; in addition, this solution proposes text query based on inverted index retrieval and high-dimensional image feature retrieval technology based on color and texture to realize fast retrieval between multiple modalities, providing a good user experience for users.

[0101] To implement the above embodiments, the present invention also provides a cross-modal retrieval method based on a pre-trained model and recall ranking.

[0102] Figure 2 Schematic diagram of a cross-modal retrieval method based on a pre-trained model and recall ranking provided by an embodiment of the present invention.

[0103] As Figure 2 shown, the cross-modal retrieval method based on a pre-trained model and recall ranking includes the following steps: S101, extract text information, expand the semantic representation of the text information through different dimensions, and increase the text sample size; S102, extract image information, extract several pictures that best represent the video content from a video, and retrieve the same or similar pictures from the database; S103, generate a roughly relevant candidate set according to the query item, accurately rank the candidate set, and finally return the relevant retrieval results.

[0104] Further, in an embodiment of the present invention, the extracting text information includes:

[0105] Audio extraction and speech recognition based on deep learning;

[0106] Obtain semantic descriptions of the current sentence in different word orders and languages, expand the existing text data from multiple aspects, and also be used to obtain a large amount of negative sample data according to fine-grained text analysis.

[0107] Further, in an embodiment of the present invention, the extracting several pictures that best represent the video content from a video includes:

[0108] Extract each frame of the video to obtain several pictures;

[0109] Map the pictures to a unified LUV color space, and calculate the absolute distance between each frame and the previous frame;

[0110] Sort all the extracted frames according to the absolute distance, and the several frames ranked at the front are regarded as the several pictures that best represent the video content.

[0111] Further, in an embodiment of the present invention, the retrieving the same or similar pictures from the database includes:

[0112] Extract features of the pictures based on the picture feature extraction technology comparing the average gray level difference;

[0113] Through the fuzzy query function provided by ElasticSearch, quickly retrieve the same or similar pictures from the picture database.

[0114] Further, in an embodiment of the present invention, generating a candidate set that is roughly relevant according to the query item and precisely sorting the candidate set includes:

[0115] Using a multi-modal pre-training model based on transformer as a sub-model of the vector embedding model for fast rough recall;

[0116] Using a multi-modal pre-training model based on transformer as a sub-model of the cross-encoder model for precise sorting.

[0117] In the description of this specification, the description with reference to terms such as "an embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0118] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of these features. In the description of the present invention, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically and clearly defined.

[0119] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A cross-modal retrieval system based on a pre-trained model and recall ranking, characterized in that, It includes the following modules: The multi-dimensional text information extraction module is used to provide information support on the text side for the cross-modal retrieval system, expand the semantic representation of text information through different dimensions, and increase the text sample size; the multi-dimensional text information extraction module adopts the method of text translation combined with speech recognition, utilizes the advantages of multi-modal data, performs speech recognition on the audio data in the video and the data that is originally audio, and obtains paired training data; performs text translation processing on the overall text data, utilizes the text semantic information, expands the quantity of paired multi-modal associated data, and at the same time, through multi-dimensional natural language analysis, randomly replaces the components in the sentence to form a rich negative sample space; The intelligent image retrieval module includes a video intelligent frame extraction module and an image search by image module. Among them, the video intelligent frame extraction module is used to extract several pictures that can best represent the video content from a video, and the image search by image module is used to complete the large-scale and high-efficiency image retrieval task; The cross-modal retrieval module is used to generate a roughly relevant candidate set according to the query item, accurately sort the candidate set, and finally return the relevant retrieval result; The multi-dimensional text information extraction module includes: The voice data processing module is used for audio extraction and speech recognition based on deep learning. The speech recognition based on deep learning uses the attention mechanism to achieve end-to-end training, and performs unified speech recognition on the audio data obtained from each modality to obtain text modality information; The natural language text extension module is used to obtain semantic descriptions of the current sentence in different word orders and different languages, expand the existing text data from multiple aspects, and is also used to obtain a large amount of negative sample data according to fine-grained text analysis.

2. The system according to claim 1, characterized in that, The video intelligent frame extraction module is used to extract several pictures that can best represent the video content from a video, specifically including: Extract each frame of the video to obtain several pictures; Map the pictures to a unified LUV color space, and calculate the absolute distance between each frame and the previous frame; Sort all the extracted frames according to the absolute distance, and the several frames ranked at the top are regarded as the several pictures that can best represent the video content.

3. The system according to claim 1, characterized in that, The image search by image module is used to complete the large-scale and high-efficiency image retrieval task, specifically including: Extract features of the pictures by using the picture feature extraction technology based on the comparison gap of the average gray level; Through the fuzzy query function provided by ElasticSearch, quickly retrieve the same or similar pictures from the picture database.

4. The system according to claim 1, characterized in that, The cross-modal retrieval module includes: The rough recall module adopts a multi-modal pre-training model based on transformer as a sub-model of the vector embedding model for fast rough recall; The accurate sorting module uses a multi-modal pre-training model based on transformer as a sub-model of the cross-encoder model for accurate sorting.

5. A cross-modal retrieval method based on a pre-trained model and recall ranking, characterized in that, It includes the following steps: Extract text information, expand the semantic representation of text information through different dimensions, and increase the text sample size; adopt the method of text translation combined with speech recognition, utilize the advantages of multimodal data, perform speech recognition on the audio data in the video and the data that is originally audio, and obtain paired training data; perform text translation processing on the overall text data, utilize the text semantic information to expand the quantity of paired multimodal associated data, and at the same time, through multi-dimensional natural language analysis, randomly replace the components in the sentence to form a rich negative sample space; Extract image information, extract several pictures that best represent the video content from a video, and retrieve the same or similar pictures from the database; Generate a roughly relevant candidate set according to the query item, accurately sort the candidate set, and finally return the relevant retrieval results; The extraction of the text information includes: Audio extraction and speech recognition based on deep learning. The speech recognition based on deep learning uses the attention mechanism to achieve end-to-end training, and performs unified speech recognition on the audio data obtained from each modality to obtain text modality information; Obtain semantic descriptions of the current sentence in different word orders and different languages, expand the existing text data from multiple aspects, and is also used to obtain a large amount of negative sample data according to fine-grained text analysis.

6. The method according to claim 5, characterized in that, The extraction of several pictures that best represent the video content from a video includes: Extract each frame of the video to obtain several pictures; Map the pictures to a unified LUV color space, and calculate the absolute distance between each frame and the previous frame; Sort all the extracted frames according to the absolute distance, and several frames with a higher ranking are regarded as several pictures that best represent the video content.

7. The method according to claim 5, characterized in that, The retrieval of the same or similar pictures from the database includes: Extract features of the pictures using the picture feature extraction technology based on the comparison gap of the average gray level; Through the fuzzy query function provided by ElasticSearch, quickly retrieve the same or similar pictures from the picture database.

8. The method according to claim 5, wherein generating a candidate set that is roughly relevant according to the query item and precisely sorting the candidate set includes: Adopt a multimodal pre-training model based on transformer as a sub-model of the vector embedding model for fast rough recall; Utilize a multimodal pre-training model based on transformer as a sub-model of the cross-encoder model for accurate sorting.

Citation Information

Patent Citations

  • Cross-media retrieval method based on Resnet-Bert network model

    CN111949806A

  • Multi-modal data retrieval method and system, terminal and storage medium

    CN112015923A