Multi-modal data retrieval method and device, equipment and storage medium

By using multimodal vector matching technology in data retrieval, target modal vectors matching the target search vector are matched from the vector database, which solves the problem of low data retrieval accuracy in the prior art, and achieves a more efficient and diversified data retrieval effect.

CN120067417APending Publication Date: 2025-05-30RUIJIE NETWORKS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311643996.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-30
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, the data retrieval method is relatively single, with low accuracy, and cannot meet the diverse retrieval needs of users.

Method used

By obtaining the target search text and determining its corresponding target search vector, the target modal vector matching the target search vector is matched from the modal vector of the vector database, and the target video clip corresponding to the target modal vector is then determined. The vector database includes modal vectors with multiple modalities of preset video clips.

Benefits of technology

Multimodal data retrieval is realized, the accuracy of data retrieval is improved, and the diverse needs of users are met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067417A_ABST
    Figure CN120067417A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal data retrieval method and device, equipment and a storage medium, and the method comprises the steps: obtaining a target search text, and determining a target search vector corresponding to the target search text; determining at least one target modal vector matched with the target search vector from modal vectors of a vector database; the vector database comprises modal vectors of a plurality of modals of the preset video clip; the number of the preset video clips is multiple; and determining a target video clip corresponding to the target modal vector from preset video clips. Thus, the electronic equipment compares the target search vector of the user with the modal vectors corresponding to the multiple modals of the preset video clip, data retrieval based on the multiple modal vectors is achieved, the accuracy of data retrieval can be improved, and diversified retrieval requirements of the user can also be met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data retrieval, and particularly to a multi-modal data retrieval method, apparatus, device, and storage medium. Background Art

[0002] With the rapid development of Internet technology, the number of network resources has grown rapidly, and the types of network resources are also gradually increasing, such as videos, audios, images, and texts. To quickly obtain network resources, users usually achieve resource search and acquisition through data retrieval.

[0003] In related technologies, electronic devices usually perform retrieval based on a single modality, such as text search based on keywords input by users. This retrieval method is relatively single and has low accuracy, unable to meet the diverse retrieval needs of users. Summary of the Invention

[0004] This application provides a multi-modal data retrieval method, apparatus, device, and storage medium, which can achieve multi-modal data retrieval, improve the accuracy of data retrieval, and meet the diverse needs of users.

[0005] In a first aspect, an embodiment of this application provides a multi-modal data retrieval method, including:

[0006] Obtain a target search text and determine a target search vector corresponding to the target search text;

[0007] Determine at least one target modality vector that matches the target search vector from the modality vectors in the vector database; the vector database includes modality vectors of multiple modalities of a preset video segment; the number of the preset video segments is multiple;

[0008] Determine a target video segment corresponding to the target modality vector from the preset video segments.

[0009] In a possible implementation manner, the method further includes:

[0010] Obtain a preset video, and based on a preset speech recognition algorithm, determine subtitle text corresponding to the preset video;

[0011] Determine a slicing point of the preset video according to the subtitle text, and based on the slicing point, slice the preset video to obtain multiple preset video segments;

[0012] For each preset video segment, determine modality vectors of multiple modalities corresponding to the preset video segment;

[0013] Store the modality vectors of multiple modalities of each preset video segment into the vector database.

[0014] In a possible implementation, determining the modal vector corresponding to the preset video segment includes:

[0015] Encoding the subtitle text of the preset video segment based on a preset multi-modal encoder to obtain the text modal vector corresponding to the preset video segment.

[0016] In a possible implementation, determining the modal vector corresponding to the preset video segment further includes:

[0017] Determining the key image frames of the preset video segment;

[0018] Determining the image text information corresponding to the key image frames based on a preset image recognition algorithm;

[0019] Encoding the image text information based on a preset multi-modal encoder to obtain the image modal vector corresponding to the preset video segment.

[0020] In a possible implementation, determining the modal vector corresponding to the preset video segment further includes:

[0021] Encoding the sequence of image frames in the preset video segment based on a preset multi-modal encoder to obtain the video modal vector corresponding to the preset video segment; the individual image frames in the sequence of image frames are arranged in the playback order of the image frames in the preset video segment.

[0022] In a possible implementation, determining at least one target modal vector matching the target search vector from the modal vectors in the vector database includes:

[0023] Determining the similarity between the target search vector and each of the modal vectors in the vector database;

[0024] Determining the modal vectors whose similarity meets a preset condition as the target modal vectors.

[0025] In a possible implementation, the method further includes:

[0026] Displaying the identification information corresponding to at least one of the target video segments;

[0027] In response to a user interaction operation on the identification information, determining the target video segment corresponding to the identification information selected by the interaction operation as the video segment to be played;

[0028] Determining the preset video to which the video segment to be played belongs;

[0029] Jump to the target position corresponding to the to-be-played video segment in the preset video, and play the to-be-played video segment.

[0030] In a second aspect, an embodiment of the present application provides a multimodal data retrieval device, including:

[0031] A first determination module, configured to obtain a target search text and determine a target search vector corresponding to the target search text;

[0032] A second determination module, configured to determine at least one target modal vector matching the target search vector from the modal vectors of the vector database; the vector database includes modal vectors of multiple modalities of preset video segments; the number of the preset video segments is multiple;

[0033] A third determination module, configured to determine a target video segment corresponding to the target modal vector from the preset video segments.

[0034] In a possible implementation manner, the device is further configured to:

[0035] Obtain a preset video, and determine subtitle text corresponding to the preset video based on a preset speech recognition algorithm;

[0036] Determine a slicing point of the preset video according to the subtitle text, and slice the preset video based on the slicing point to obtain a plurality of preset video segments;

[0037] For each preset video segment, determine modal vectors of multiple modalities corresponding to the preset video segment;

[0038] Store modal vectors of multiple modalities of each preset video segment into the vector database.

[0039] In a possible implementation manner, the device is further configured to:

[0040] Based on a preset multimodal encoder, perform encoding processing on the subtitle text of the preset video segment to obtain a text modal vector corresponding to the preset video segment.

[0041] In a possible implementation manner, the device is further configured to:

[0042] Determine a key image frame of the preset video segment;

[0043] Based on a preset image recognition algorithm, determine image text information corresponding to the key image frame;

[0044] Based on a preset multimodal encoder, perform encoding processing on the image text information to obtain an image modal vector corresponding to the preset video segment.

[0045] In a possible implementation, the device is further configured to:

[0046] Based on a preset multi-modal encoder, encode the sequence of image frames in the preset video segment to obtain a video modal vector corresponding to the preset video segment; each image frame in the sequence of image frames is arranged according to the playback order of the image frame in the preset video segment.

[0047] In a possible implementation, the second determination module is specifically configured to:

[0048] In the vector database, determine the similarity between the target search vector and each of the modal vectors;

[0049] Determine the modal vectors whose similarity meets the preset conditions as the target modal vectors.

[0050] In a possible implementation, the device is further configured to:

[0051] Display the identification information corresponding to at least one of the target video segments;

[0052] In response to an interaction operation by the user on the identification information, determine the target video segment corresponding to the identification information selected by the interaction operation as the video segment to be played;

[0053] Determine the preset video to which the video segment to be played belongs;

[0054] Jump to the target position corresponding to the video segment to be played in the preset video and play the video segment to be played.

[0055] In a third aspect, an embodiment of the present application provides a multi-modal data retrieval device, including: a processor and a memory;

[0056] The memory stores computer-executable instructions;

[0057] The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of the first aspects.

[0058] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed, they are used to implement the method according to any one of the first aspects.

[0059] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed, it implements the method according to any one of the first aspects.

[0060] Sixth aspect, an embodiment of the present application provides a chip, on which a computer program is stored, and when the computer program is executed by the chip, the method described in any item of the first aspect is implemented.

[0061] The multi-modal data retrieval method, device, equipment, and storage medium provided by the embodiments of the present application obtain a target search text and determine a target search vector corresponding to the target search text; determine at least one target modal vector matching the target search vector from the modal vectors in the vector database; the vector database includes modal vectors of multiple modalities of a preset video segment; the number of preset video segments is multiple; determine a target video segment corresponding to the target modal vector from the preset video segments. In the present application, the electronic device generates a target search vector corresponding to the user's target search text, and then performs retrieval and matching in the vector database, which includes modal vectors of multiple modalities of multiple preset video segments; in the vector database, the electronic device can determine a target modal vector that matches the user's target search vector, and then can determine a target video segment corresponding to the target modal vector. In this way, by comparing the user's target search vector with the modal vectors corresponding to multiple modalities of the preset video segment respectively, the electronic device realizes data retrieval based on multiple modal vectors, which can improve the accuracy of data retrieval and meet the diverse retrieval needs of users. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 It is a schematic diagram of an application scenario provided by an embodiment of the present application;

[0063] Figure 2 It is a schematic flowchart of a multi-modal data retrieval method provided by an embodiment of the present application;

[0064] Figure 3 It is a schematic flowchart of another multi-modal data retrieval method provided by an embodiment of the present application;

[0065] Figure 4 It is a schematic logical diagram of a multi-modal data retrieval method provided by an embodiment of the present application;

[0066] Figure 5 It is a schematic structural diagram of a multi-modal data retrieval device provided by an embodiment of the present application;

[0067] Figure 6 It is a schematic structural diagram of a multi-modal data retrieval device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0068] To enable those skilled in the art to better understand the technical solutions of this application, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments and drawings described herein are only used to explain this application, rather than limiting this application. It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards, and corresponding operation entrances are provided for users to choose to authorize or reject.

[0069] With the continuous development and popularization of Internet technology, the number of network resources has increased rapidly, and the types have gradually become rich, such as videos, audios, images, and texts. To quickly obtain network resources, users usually achieve resource acquisition based on data retrieval methods.

[0070] In related technologies, for multi-modal network resources, search engines and recommendation systems in electronic devices usually analyze only one modality independently to process single-type data. For example, when performing video retrieval, after the user enters a search term, the electronic device usually retrieves based on the text and determines videos whose titles (or names) include the search term as search results. This retrieval method is relatively single and inaccurate, ignoring the diverse information needs of users and unable to meet the diverse needs and preferences of users. Exemplarily, in an educational scenario, when students search for "linear algebra", some students may be more concerned about video explanations, while other students may want to find the content about "linear algebra" on the PPT in the video.

[0071] In this application, the electronic device generates a target search vector for the user based on the user's target search text, and then performs a retrieval match in the vector database, which includes modality vectors of multiple modalities of multiple preset video segments; in the vector database, the electronic device can determine a target modality vector that matches the user's target search vector, and then can determine the target video segment corresponding to the target modality vector. In this way, by comparing the user's target search vector with the modality vectors corresponding to multiple modalities of the preset video segments, the electronic device realizes data retrieval based on multiple modality vectors, which can improve the accuracy of data retrieval and meet the diverse retrieval needs of users.

[0072] Figure 1 Schematic diagram of the application scenario provided for the embodiments of this application. Please refer to Figure 1, including user 101 and electronic device 102. The electronic device 102 can be a user terminal, for example, devices such as mobile phones, computers, servers, etc. In the related art, when performing data retrieval, the electronic device 102 obtains the search term of the user 101, and performs text retrieval based on this search term, and determines the target video whose title or name includes this search term as the retrieval result. This single-modal data retrieval method has low accuracy and cannot meet the diverse needs of users.

[0073] In the embodiment of the present application, after the electronic device 102 obtains the search term of the user 101, it can determine the search vector corresponding to the search term of the user, and then perform vector matching in the vector database. The vector database includes modal vectors of multiple modalities for each video segment. After that, the electronic device 102 can use the target video segment corresponding to the modal vector that matches the search vector as the search result. In this way, the retrieval matching between the user search vector and the multi-modal modal vectors can be realized, improving the accuracy of data retrieval and meeting the diverse information needs of users.

[0074] The following details the solution shown in the present application through specific embodiments. It should be noted that the following several embodiments can exist independently or be combined with each other. For the same or similar content, it will not be repeated in different embodiments.

[0075] Figure 2 It is a schematic flowchart of a multi-modal data retrieval method provided by an embodiment of the present application. Please refer to Figure 2 This multi-modal data retrieval method may include:

[0076] S201. Obtain the target search text and determine the target search vector corresponding to the target search text.

[0077] The execution subject of the embodiment of the present application can be an electronic device or a multi-modal data retrieval device provided in the electronic device. The multi-modal data retrieval device can be implemented by software or by a combination of software and hardware. For the sake of easy understanding, in the following, the case where the execution subject is an electronic device is taken as an example for description.

[0078] In the embodiments of the present application, the target search text may refer to the search text input by the user, including one or more search terms, etc. When the user inputs the target search text, text input can be performed through a touch operation, and voice input can be performed through a voice collection module such as a microphone. The embodiments of the present application do not limit the specific input method of the target search text, and the electronic device can determine the target search text in response to the user's input operation. The target search vector may refer to the vectorized representation of the target search text, which can be determined through a multimodal alignment encoder or a text word vector (Embedding) model, etc. The multimodal alignment encoder may include a pre-trained model (Bidirectional Encoder Representation from Transformers, BERT), etc., and can be specifically composed based on deep learning algorithms and natural language processing algorithms, etc. The embodiments of the present application do not limit the specific type of the multimodal alignment encoder.

[0079] In this step, when the user needs to perform a retrieval, the target search text can be input in the electronic device. In response to the user's input operation, the electronic device can determine the target search text corresponding to the input operation and input the target search text into the multimodal alignment encoder to determine the target search vector corresponding to the target search text. In this way, by vectorizing the user's search terms, the electronic device can improve the accuracy of subsequent retrieval matching.

[0080] S202. Determine at least one target modal vector that matches the target search vector from the modal vectors in the vector database; the vector database includes modal vectors of multiple modalities of preset video segments; the number of preset video segments is multiple.

[0081] In the embodiments of the present application, the vector database may refer to a pre-configured database, which may include modal vectors of multiple modalities corresponding to preset video segments. The multiple modalities may include text modality, image modality, and video modality, etc., and of course may also include other modalities. The embodiments of the present application do not limit this. The preset video segment may be a video segment obtained by splitting a preset video. The vector database may include modal vectors of multiple modalities corresponding to a large number of preset video segments obtained by splitting the preset video.

[0082] The target modal vector may refer to a modal vector that matches the user's target search vector. The number of the target modal vectors may be one or multiple; when the number of the target modal vectors is multiple, the multiple target modal vectors may be modal vectors corresponding to the same preset video segment or modal vectors corresponding to different preset video segments. The embodiments of the present application do not limit this.

[0083] In this step, before the electronic device performs data retrieval, it can first preprocess the preset videos included in the network resources, slice each preset video to obtain multiple preset video segments, and then determine the modal vectors of multiple modalities for each preset video segment, such as text modal vectors, image modal vectors, and video modal vectors. Then, the modal vectors of multiple modalities are stored in the vector database. Subsequently, the target search vector of the user can be retrieved and matched in the vector database to achieve multi-modal retrieval of data. Specifically, during retrieval, the electronic device can determine the similarity between each modal vector in the vector database and the target search vector of the user, and then determine the modal vectors with higher similarity as the target modal vectors that match the target search vector. In this way, the vector database stores the modal vectors of multiple modalities for each preset video segment, and the electronic device can compare and match the target search vector of the user with the modal vectors of multiple modalities, which can achieve multi-modal data retrieval of network resources, improve the accuracy of data retrieval, and at the same time, the multi-modal data retrieval is more comprehensive and can meet the diverse information needs of users.

[0084] S203. Determine the target video segment corresponding to the target modal vector from the preset video segments.

[0085] In the embodiment of the present application, the target video segment may refer to the video segment finally retrieved by the electronic device based on the target search text of the user. In a possible implementation manner, the vector database may store the modal vectors of multiple modalities for each preset video segment, and each modal vector may include video label information (or the vector database may store the video label information corresponding to each modal vector). This video label information is used to identify the preset video segment corresponding to the modal vector, and can also be used to identify the original preset video corresponding to the preset video segment and the position of the preset video segment in the preset video, etc. After the electronic device determines the target modal vector, it can determine the target video segment corresponding to the target modal vector based on the video label in the target modal vector.

[0086] In another possible implementation manner, the vector database may store multiple preset video segments, and at the same time store the modal vectors of multiple modalities corresponding to each video segment. After the electronic device determines the target modal vector, it can determine the target video segment corresponding to the target modal vector based on the correspondence between the preset video segment and the modal vector as the search result of the user. Subsequently, at least one video segment can be displayed to the user for easy viewing. Of course, the electronic device can use other methods to determine the target search result corresponding to the target search vector, and the embodiment of the present application does not limit this.

[0087] The multi-modal data retrieval method provided by the embodiment of the present application, the electronic device obtains the target search text and determines the target search vector corresponding to the target search text; determines at least one target modal vector matching the target search vector from the modal vectors in the vector database; the vector database includes modal vectors of multiple modalities of a preset video segment; the number of preset video segments is multiple; determines the target video segment corresponding to the target modal vector from the preset video segments. In the present application, the electronic device generates the target search vector corresponding to the user's target search text based on the user's target search text, and then performs retrieval and matching in the vector database, and the vector database includes modal vectors of multiple modalities of multiple preset video segments; in the vector database, the electronic device can determine the target modal vector matching the user's target search vector, and further can determine the target video segment corresponding to the target modal vector. In this way, the electronic device realizes data retrieval based on multiple modal vectors by comparing the user's target search vector with the modal vectors corresponding to multiple modalities of the preset video segment, which can improve the accuracy of data retrieval and meet the diverse retrieval needs of users.

[0088] Based on the above embodiment, Figure 3 is a schematic flowchart of another multi-modal data retrieval method provided by the embodiment of the present application. Please refer to Figure 3 This multi-modal data retrieval method may include:

[0089] S301. Obtain a preset video, and based on a preset speech recognition algorithm, determine the subtitle text corresponding to the preset video.

[0090] In the embodiment of the present application, the preset video may refer to various videos in network resources. The preset speech recognition algorithm may refer to a pre-trained speech recognition algorithm, for example, it may refer to an Automatic Speech Recognition (ASR) algorithm, etc. The embodiment of the present application does not limit the specific type of the preset speech algorithm. The subtitle text may refer to the subtitle text information of the preset video.

[0091] In this step, before the user performs retrieval, the electronic device may first process the preset video and add the modal vectors of each preset video segment to the vector database to implement data preprocessing and data augmentation. The electronic device may first obtain the preset video, and then based on the preset speech recognition algorithm, perform conversion processing on the audio information (or speech information, etc.) in the preset video to obtain the subtitle text corresponding to the preset video.

[0092] S302. Determine the slicing points of the preset video according to the subtitle text, and based on the slicing points, split the preset video to obtain multiple preset video segments.

[0093] In the embodiments of the present application, the slicing point may refer to the segmentation point or the splitting point of a preset video. When the electronic device determines the subtitle text corresponding to the preset video, it may simultaneously determine the punctuation marks in the subtitle text, and then use punctuation marks such as full stops as the splitting points of the preset video. After that, based on the time points corresponding to the label symbols, the preset video can be split to obtain multiple preset video segments corresponding to the preset video. Subsequently, multiple modal vectors of various modalities corresponding to each preset video segment can be further determined. In this way, in the embodiments of the present application, the electronic device splits the preset video into multiple preset video segments, which can reduce the granularity of data retrieval. Compared with the search method in the related art that directly returns the entire video as the search result, in the embodiments of the present application, the preset video segments can be used as the search results, with higher accuracy. The user does not need to watch the entire video completely, and can quickly view the information, saving the user's time.

[0094] S303. For each preset video segment, determine multiple modal vectors corresponding to the preset video segment.

[0095] In the embodiments of the present application, after splitting the preset video to obtain multiple preset video segments, for each preset video segment, the electronic device can determine multiple modal vectors corresponding to each preset video segment. The modal vector may include a text modal vector, an image modal vector, a video modal vector, etc. The embodiments of the present application do not limit the specific type of the modal vector.

[0096] In a possible implementation manner, the modal vector includes a text modal vector, and the text modal vector of the preset video segment can be determined in the following manner:

[0097] Based on a preset multi-modal encoder, encode the subtitle text of the preset video segment to obtain the text modal vector corresponding to the preset video segment.

[0098] In the embodiments of the present application, the preset multi-modal encoder may refer to an encoder for multi-modal alignment, etc. The text modal vector may refer to a vector generated based on the subtitle text corresponding to the preset video segment. Specifically, the electronic device can input the subtitle text included in the preset video segment into the preset multi-modal encoder, perform processing such as word vector conversion, and obtain the text modal vector corresponding to the preset video segment, so as to realize the vector representation of the subtitle text of the preset video segment.

[0099] In another possible implementation manner, the modal vector further includes an image modal vector, and the image modal vector of the preset video segment can be determined in the following manner:

[0100] Determine the key image frames of the preset video segment; based on the preset image recognition algorithm, determine the image text information corresponding to the key image frames; based on the preset multi-modal encoder, perform encoding processing on the image text information to obtain the image modal vector corresponding to the preset video segment.

[0101] In the embodiments of the present application, the key image frame may refer to an image frame with relatively rich content included in the preset video segment. The electronic device may input each image frame included in the preset video segment into the preset image comparison algorithm in time series. The preset image comparison algorithm is used to compare the differences between adjacent two image frames. When the difference between a certain image frame and its front and rear image frames is greater than the preset difference value, it can be determined that the difference between this image frame and its front and rear image frames is relatively large and can be used as a key image frame.

[0102] The preset image recognition algorithm may refer to a pre-trained image recognition and detection algorithm, which can be used to recognize information such as text and objects included in the image. Specifically, it may include optical character recognition (OCR), etc. The embodiments of the present application do not limit the specific type of the preset image recognition algorithm. The image text information may refer to the text information included in the image frame. For example, in an educational scenario, the image text information may be the text information in the key image frame, specifically, the text on the blackboard and the text information displayed on the document page, etc. The image modal vector may refer to the vector representation corresponding to the image content of the preset video segment.

[0103] In this step, when determining the image modal vector corresponding to the preset video segment, the electronic device may first determine the key image frames in the preset video segment, then determine the image text information in the key image frames based on the preset image recognition algorithm, and then perform encoding processing on the image text information based on the preset multi-modal encoder (such as a text Embedding model, etc.) to obtain the image modal information corresponding to the preset video segment. In this way, by encoding and vectorizing the image text information in the key image frames of the preset video segment, the electronic device obtains the image modal information, which can enrich the types of modal vectors, improve the comprehensiveness of subsequent data retrieval, and at the same time only process the key image frames, which can reduce the calculation amount to a certain extent.

[0104] It should be noted that the electronic device may also determine the image text information included in all the image frames in the preset video segment based on the preset image recognition algorithm, and then perform encoding processing on the image text information of all the image frames to obtain the image modal vector corresponding to the preset video segment. This can improve the comprehensiveness of the image modal vector and improve the accuracy of subsequent data retrieval to a certain extent.

[0105] In another possible implementation, the modal vector further includes a video modal vector, and the video modal vector of the preset video segment can be determined in the following manner:

[0106] Based on the preset multi-modal encoder, encode the sequence of image frames in the preset video segment to obtain the video modal vector corresponding to the preset video segment; each image frame in the sequence of image frames is arranged according to the playback order of the image frames in the preset video segment.

[0107] In the embodiments of the present application, the sequence of image frames may refer to a plurality of consecutive image frames arranged according to the spatio-temporal relationship between the image frames in the preset video segment. The sequence of image frames may include all the image frames in the preset video segment, or may be image frames determined according to a preset time period. The embodiments of the present application do not limit this. Each image frame in the sequence of image frames may be arranged according to the playback order in the preset video segment. For example, when the preset video segment is a continuous action, the sequence of image frames may refer to a plurality of consecutive image frames from the start to the end of the action.

[0108] In this step, the electronic device may determine the sequence of image frames corresponding to the preset video segment based on the time sequence relationship between the respective image frames in the preset video segment, and then may input the sequence of image frames into the preset multi-modal encoder (such as a multi-frame Embedding model, etc.), and encode the sequence of image frames through the preset multi-modal encoder to obtain the video modal vector corresponding to the preset video segment. In this way, the electronic device encodes the sequence of image frames in combination with the front-to-back timing information between the image frames, which can make the information of the video modal vector more abundant, and thus can improve the accuracy of subsequent data retrieval.

[0109] Of course, it should be noted that the modal vectors of multiple modalities of the preset video segment may further include other types of vectors, or may be determined in other ways. The embodiments of the present application do not limit this.

[0110] S304. Store the modal vectors of multiple modalities of each preset video segment in the vector database.

[0111] In the embodiments of the present application, after determining the modal vectors of multiple modalities of each preset time-offset segment, the electronic device may store the modal vectors of multiple modalities corresponding to each preset video segment in the vector database, and may also store video label information such as the time corresponding to the preset video segment. Specifically, the modal vectors of multiple modalities corresponding to the same preset video segment may be stored corresponding to the video label information of the preset video segment, or the video label information may be embedded into the modal vector. The embodiments of the present application do not limit the specific storage method of the video label information.

[0112] In addition, the electronic device may only store the modal vectors of multiple modalities of each preset video segment and the video tag information corresponding to the preset video segment, without storing the preset video segment itself, which can reduce the storage space occupancy; or, the electronic device may store the preset video segment correspondingly while storing the modal vectors of multiple modalities of each preset video segment and the video tag information corresponding to the preset video segment, which can simplify the subsequent search process and improve the efficiency of determining the target video segment. The embodiments of the present application do not limit this.

[0113] In the embodiments of the present application, the electronic device comprehensively performs all-round information retrieval based on multiple modalities (video frames, timing information, audio, subtitles, etc.), and conducts comprehensive analysis based on the internal connections and mutual influences among the multi-modal information, which can improve the retrieval accuracy and efficiency; and through a preset multi-modal encoder model composed of algorithms such as deep learning and natural language processing technologies for vector encoding processing, intelligent recommendation can be performed according to the personalized needs and queries of users, and the diverse information needs of users can be met.

[0114] S305. Obtain a target search text and determine a target search vector corresponding to the target search text.

[0115] S306. Determine at least one target modal vector that matches the target search vector from the modal vectors in the vector database; the vector database includes the modal vectors of multiple modalities of the preset video segments; the number of the preset video segments is multiple.

[0116] In the embodiments of the present application, during the retrieval process, the electronic device may obtain the target search text input by the user and determine the target search vector corresponding to the target search text based on a word vector model or the like. Then, the electronic device may match the target search vector with each modal vector in the vector database to determine at least one target modal vector that matches the target search vector.

[0117] In a possible implementation manner, the determination process of the target modal vector may be implemented in the following manner:

[0118] In the vector database, determine the similarity between the target search vector and each modal vector; determine the modal vectors whose similarity meets the preset conditions as the target modal vectors.

[0119] In the embodiments of the present application, the similarity may refer to the degree of similarity, specifically, it may refer to the cosine distance, Euclidean distance, etc. between two vectors. The preset condition may refer to the judgment condition for determining whether two vectors match set in advance. For example, the preset condition may include that the similarity is greater than 95% or the similarity is ranked among the top N, where N is a positive integer, for example, N may be 5, 8, 10, etc. Of course, the preset condition may also be of other types, and the embodiments of the present application do not limit this. Specifically, the electronic device may determine the similarity between the target search vector of the user and each modal vector in the vector database, and determine the modal vector whose similarity meets the preset condition as the target modal vector. In this way, the electronic device can improve the accuracy of data retrieval through the similarity calculation between vectors.

[0120] In another possible implementation manner, the electronic device may also directly calculate the similarity between the target search vector and the preset video segment. Specifically, for each preset video segment, the electronic device may first determine the similarity between each modal vector of multiple modalities of the preset video segment and the target search vector respectively; then comprehensively calculate the similarities corresponding to the modal vectors of these multiple modalities, specifically, it may refer to weighted calculation, etc., to determine the similarity between the target search vector and each preset video segment; then the electronic device may determine the target video segment corresponding to the target search vector based on whether the similarity between the target search vector and the preset video segment meets the preset condition. Among them, when performing weighted calculation, the weight values corresponding to the modal vectors of different modalities may be set based on the type of the preset video to which the preset video segment belongs, or may also be flexibly set according to the actual needs of the user, and the embodiments of the present application do not limit this.

[0121] S307. Determine the target video segment corresponding to the target modal vector from the preset video segments.

[0122] In the embodiments of the present application, after determining the target modal vector, the electronic device may determine the target video segment corresponding to the target modal vector based on the corresponding relationship between the target modal vector and the video label information, or based on the corresponding relationship between the target modal vector and the preset video segment, obtain the search result corresponding to the user's target search text, and then display it to the user later to implement multi-modal data search.

[0123] S308. Display the identification information corresponding to at least one target video segment.

[0124] S309. In response to the user's interaction operation on the identification information, determine the target video segment corresponding to the identification information selected by the interaction operation as the video segment to be played.

[0125] In the embodiments of the present application, the identification information may refer to various display information corresponding to the target video segment, specifically including thumbnail images, names, and brief descriptions, etc. The specific form of the identification information may be flexibly set based on actual needs, and the embodiments of the present application do not limit this. The interaction operation may refer to the touch operation used by the user to play the target video segment, specifically may refer to click operations or swipe operations, etc. The embodiments of the present application do not limit the specific type of the interaction operation. The video segment to be played may refer to the target video segment that the user selects to play through the interaction operation.

[0126] Specifically, after the electronic device determines at least one target video segment, it may display the identification information corresponding to at least one target video segment on the retrieval result display page. For example, it may display thumbnail images and names corresponding to each video segment. When the user needs to play a certain target video segment, the user may perform interaction operations such as clicking on the screen of the electronic device. In response to the user's interaction operation, the electronic device may determine the target video segment corresponding to the identification information selected by the interaction operation as the video segment to be played, and subsequent playback can be performed.

[0127] S310. Determine the preset video to which the video segment to be played belongs; jump to the target position corresponding to the video segment to be played in the preset video, and play the video segment to be played

[0128] In the embodiments of the present application, the target position may refer to the start time point of playing the target video segment in the preset video. Specifically, after determining the video segment to be played, the electronic device may first determine the preset video to which the video segment to be played belongs, specifically based on the video tag information of the video segment to be played, etc.; then the electronic device may automatically jump to the target position corresponding to the video segment to be played in the preset video and play the video segment to be played. In this way, compared with the retrieval method in the related art that directly returns a complete video to the user, in the embodiments of the present application, the electronic device can perform retrieval for the target video segment, with a smaller search granularity, and at the same time can achieve jump playback of the target video segment, without the user having to watch the entire video, nor having to manually adjust the playback progress, simplifying the user's operation and meeting the diverse needs of the user.

[0129] Of course, in another possible implementation manner, the electronic device may load the entire preset video and highlight the target video segment in the progress bar of the preset video to facilitate the user to manually jump and view. The embodiments of the present application do not limit the specific display and playback methods of the target video segment.

[0130] Based on the above embodiments, Figure 4 is a logical schematic diagram of multimodal data retrieval provided by the embodiments of the present application. As Figure 4As shown, after obtaining a preset video, an electronic device can convert the audio information in the preset video into subtitle text based on a preset speech recognition algorithm such as ASR, determine the slice points of the video according to the subtitle text, and then slice the preset video into multiple independent preset video segments according to the slice points.

[0131] After that, the electronic device can use a preset multi-modal encoder to encode the subtitle text to obtain a text modal vector corresponding to the preset video segment; the electronic device can simultaneously determine the key image frames of each preset video segment, and determine the image text information in the key image frames through a preset image recognition algorithm such as OCR. After that, the electronic device can use the preset multi-modal encoder to encode the image text information to obtain an image modal vector corresponding to the preset video segment; the electronic device can simultaneously encode the image frame sequence with a time sequence in the preset video segment through the preset multi-modal encoder to obtain a video modal vector corresponding to the preset video segment. The electronic device can store the modal vectors of multiple modalities of the preset video segment and the video label information corresponding to the preset video segment in a vector database.

[0132] During the data retrieval process, the electronic device can obtain the target search text of the user and determine the target search vector corresponding to the target search text. After that, the electronic device can determine at least one target modal vector that matches the target search vector in the vector database, and then can determine the target video segment corresponding to the target modal vector, realizing multi-modal data retrieval.

[0133] Compared with the method in the related art where a search engine analyzes and retrieves a single modality, the electronic device in the embodiment of the present application comprehensively analyzes data of different modalities through a preset multi-modal encoder composed of deep learning and natural language processing algorithms, obtains modal vectors of different modalities for retrieval and matching, and can achieve more accurate and efficient information retrieval, better meeting the diverse information needs of users. For example, when retrieving "linear algebra", the electronic device can comprehensively analyze video images, subtitles, and video image sequences to provide a more comprehensive and personalized search result for the user. Another example is in a meeting scenario. When a user wants to review all the information about the discussion on "marketing promotion strategies", the user can enter this keyword for retrieval. The electronic device can comprehensively analyze the related audio transcriptions, whiteboard videos, meeting-related attachments, etc., and provide a comprehensive and highly relevant search result.

[0134] In the embodiments of the present application, the electronic device generates a high-dimensional embedding vector, i.e., a modality vector, based on video image content (such as image text), video caption text, image frame sequence, and other possible modalities, and can implement multi-modal data retrieval. This can not only improve the accuracy and efficiency of multi-modal resource retrieval, but also provide more personalized and context-related search results, effectively enhancing the user experience. It can be widely applied to scenarios rich in multi-modal information such as education, multimedia, conferences, and enterprise information management.

[0135] Figure 5 FIG. is a schematic structural diagram of a multi-modal data retrieval device provided by an embodiment of the present application. Please refer to Figure 5 The multi-modal data retrieval device 50 may include:

[0136] A first determination module 51, configured to obtain a target search text and determine a target search vector corresponding to the target search text;

[0137] A second determination module 52, configured to determine at least one target modality vector matching the target search vector from the modality vectors in the vector database; the vector database includes modality vectors of multiple modalities of a preset video segment; the number of preset video segments is multiple;

[0138] A third determination module 53, configured to determine a target video segment corresponding to the target modality vector from the preset video segments.

[0139] In a possible implementation manner, the device 50 is further configured to:

[0140] Obtain a preset video, and determine the caption text corresponding to the preset video based on a preset speech recognition algorithm;

[0141] Determine the slicing points of the preset video according to the caption text, and perform slicing on the preset video based on the slicing points to obtain multiple preset video segments;

[0142] For each preset video segment, determine the modality vectors of multiple modalities corresponding to the preset video segment;

[0143] Store the modality vectors of multiple modalities of each preset video segment into the vector database.

[0144] In a possible implementation manner, the device 50 is further configured to:

[0145] Based on a preset multi-modal encoder, perform encoding processing on the caption text of the preset video segment to obtain a text modality vector corresponding to the preset video segment.

[0146] In a possible implementation manner, the device 50 is further configured to:

[0147] Determine the key image frames of the preset video segment;

[0148] Based on a preset image recognition algorithm, determine the image text information corresponding to the key image frames;

[0149] Based on a preset multimodal encoder, perform encoding processing on the image text information to obtain an image modality vector corresponding to a preset video segment.

[0150] In a possible implementation manner, the apparatus 50 is further configured to:

[0151] Based on a preset multimodal encoder, perform encoding processing on the sequence of image frames in a preset video segment to obtain a video modality vector corresponding to the preset video segment; each image frame in the sequence of image frames is arranged in the playing order of the image frames in the preset video segment.

[0152] In a possible implementation manner, the second determination module 52 is specifically configured to:

[0153] In the vector database, determine the similarity between the target search vector and each modality vector;

[0154] Determine the modality vectors whose similarities meet the preset conditions as target modality vectors.

[0155] In a possible implementation manner, the apparatus 50 is further configured to:

[0156] Display the identification information corresponding to at least one target video segment;

[0157] In response to a user's interaction operation on the identification information, determine the target video segment corresponding to the identification information selected by the interaction operation as the video segment to be played;

[0158] Determine the preset video to which the video segment to be played belongs;

[0159] Jump to the target position corresponding to the video segment to be played in the preset video and play the video segment to be played.

[0160] The multimodal data retrieval apparatus 50 provided in the embodiments of the present application can execute the technical solutions shown in the above method embodiments, and its implementation principles and beneficial effects are similar, and will not be elaborated here. The multimodal data retrieval apparatus 50 can specifically be a chip, a chip module, etc., and the embodiments of the present application do not make any limitations in this regard.

[0161] Figure 6 This is a schematic structural diagram of a multimodal data retrieval device provided in an embodiment of the present application. Please refer to Figure 6 , the multimodal data retrieval device 60 may include: a memory 61 and a processor 62. Exemplarily, the memory 61 and the processor 62 are interconnected with each other through a bus 63.

[0162] The memory 61 is used to store program instructions;

[0163] The processor 62 is used to execute the program instructions stored in the memory, so as to implement the multi-modal data retrieval method shown in the above embodiments.

[0164] Figure 6 The multi-modal data retrieval device 60 shown in the embodiments can execute the technical solutions shown in the above method embodiments. The implementation principle and beneficial effects are similar, and will not be elaborated here.

[0165] The embodiments of the present application provide a computer-readable storage medium. Computer-executable instructions are stored in the computer-readable storage medium. When the computer-executable instructions are executed by a processor, they are used to implement the above multi-modal data retrieval method.

[0166] The embodiments of the present application can also provide a computer program product, including a computer program. When the computer program is executed by a processor, the above multi-modal data retrieval method can be implemented.

[0167] The embodiments of the present application provide a chip. A computer program is stored on the chip. When the computer program is executed by the chip, the above multi-modal data retrieval method is implemented.

[0168] It should be noted that the processor mentioned in the embodiments of the present application may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0169] It should be understood that the memory mentioned in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synch link DRAM (SLDRAM), and direct ram bus RAM (DR RAM). It should be noted that when the processor is a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic device, discrete gate, or transistor logic device, discrete hardware component, the memory (storage module) is integrated in the processor. It should be noted that the memory described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0170] It should be understood that in various embodiments of the present application, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0171] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processing unit of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processing unit of the computer or other programmable data processing devices generate for implementation in the processFigure 1 one process or multiple processes and / or boxes Figure 1 a device for the functions specified in one box or multiple boxes.

[0172] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device that implements the functions in the process Figure 1 one process or multiple processes and / or boxes Figure 1 specified in one box or multiple boxes.

[0173] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the process Figure 1 one process or multiple processes and / or boxes Figure 1 specified in one box or multiple boxes.

[0174] Regarding each device and each module / unit included in the product described in the above embodiments, it can be a software module / unit, a hardware module / unit, or it can also be partially a software module / unit and partially a hardware module / unit. Each device and product can be applied to or integrated into a chip, a chip module, or a terminal device. Exemplarily, for each device and product applied to or integrated into a chip, each module / chip included therein can all be implemented in a hardware manner such as a circuit, or at least some modules / units can be implemented in a software program manner, and the software program runs on a processor integrated inside the chip, and the remaining part of the modules / units can be implemented in a hardware manner such as a circuit.

[0175] In this application, the term "including" and its variants can refer to non-restrictive inclusion; the term "or" and its variants can refer to "and / or". In this application, terms such as "first" and "second" are used to distinguish similar objects and do not necessarily have to describe a specific order or sequence. In this application, "multiple" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0176] The above are only some embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A multimodal data retrieval method, characterized in that, comprising: Obtain a target search text and determine a target search vector corresponding to the target search text; Determine at least one target modal vector that matches the target search vector from the modal vectors in the vector database; the vector database includes modal vectors of multiple modalities of a preset video segment; the number of the preset video segments is multiple; Determine a target video segment corresponding to the target modal vector from the preset video segments.

2. The method according to claim 1, characterized in that, the method further comprises: Obtain a preset video, and based on a preset speech recognition algorithm, determine subtitle text corresponding to the preset video; Determine a slicing point of the preset video according to the subtitle text, and based on the slicing point, slice the preset video to obtain multiple preset video segments; For each preset video segment, determine modal vectors of multiple modalities corresponding to the preset video segment; Store the modal vectors of multiple modalities of each preset video segment into the vector database.

3. The method according to claim 2, characterized in that, the determining the modal vector corresponding to the preset video segment includes: Based on a preset multimodal encoder, perform encoding processing on the subtitle text of the preset video segment to obtain a text modal vector corresponding to the preset video segment.

4. The method according to claim 3, characterized in that, the determining the modal vector corresponding to the preset video segment further includes: Determine a key image frame of the preset video segment; Based on a preset image recognition algorithm, determine image text information corresponding to the key image frame; Based on a preset multimodal encoder, perform encoding processing on the image text information to obtain an image modal vector corresponding to the preset video segment.

5. The method according to claim 4, characterized in that, the determining the modal vector corresponding to the preset video segment further includes: Based on a preset multimodal encoder, perform encoding processing on an image frame sequence in the preset video segment to obtain a video modal vector corresponding to the preset video segment; each image frame in the image frame sequence is arranged according to the playing order of the image frame in the preset video segment.

6. The method according to claim 1, characterized in that, the determining at least one target modal vector that matches the target search vector from the modal vectors in the vector database includes: In the vector database, determine the similarity between the target search vector and each of the modal vectors; Determine the modal vectors whose similarity meets a preset condition as the target modal vectors.

7. The method according to claim 1, characterized in that, the method further comprises: Display identification information corresponding to at least one of the target video segments; In response to an interaction operation of the user on the identification information, determine a target video segment corresponding to the identification information selected by the interaction operation as a video segment to be played; Determine the preset video to which the video segment to be played belongs; Jump to the target position corresponding to the to-be-played video segment in the preset video and play the to-be-played video segment.

8. A multimodal data retrieval device, characterized in that it includes: A first determination module, configured to obtain a target search text and determine a target search vector corresponding to the target search text; A second determination module, configured to determine at least one target modal vector that matches the target search vector from the modal vectors in the vector database; the vector database includes modal vectors of multiple modalities of preset video segments; the number of the preset video segments is multiple; A third determination module, configured to determine a target video segment corresponding to the target modal vector from the preset video segments.

9. A multimodal data retrieval device, characterized in that it includes: A processor and a memory; The memory stores computer execution instructions; The processor executes the computer execution instructions stored in the memory to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer execution instructions, which are used to implement the method according to any one of claims 1 to 7 when the computer execution instructions are executed.