A method and system for searching based on a multi-modal model

By segmenting and keyframe extraction of videos, combining image and text features fusion, multimodal feature vectors are generated, which solves the complex and costly multimodal model search in the prior art, and realizes efficient video search.

CN119938986BActive Publication Date: 2025-07-22ZHEJIANG PROVINCIAL PUBLIC SECURITY SCIENCE & TECHNOLOGY RESEARCH INSTITUTE +3
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510422711.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-22
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

The prior art uses multimodal models to search, the methods are complex and costly, making it difficult to efficiently process video data.

Method used

By segmenting and keyframe extraction of videos, combining image and text features fusion, multimodal feature vectors are generated using multimodal models, and mapping relationships are established in the database to quickly retrieve and sort scene units.

Benefits of technology

Reduces the computing volume of video processing, simplifies methods, reduces costs, and improves search efficiency, allowing the most similar video clips to be quickly presented.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_3
    Figure SMS_3
  • Figure SMS_19
    Figure SMS_19
  • Figure SMS_27
    Figure SMS_27
Patent Text Reader

Abstract

The present invention discloses a method and system for searching based on a multi-modal model. The method includes: segmenting a video, and taking a set of frames that are coherent before and after and have a similarity higher than a threshold in each frame of the video as a scene unit; extracting key frames from the scene unit; extracting image features and text features from the key frames; fusing the key frame image features and text vector features to obtain multi-modal features reflecting the content of the scene unit; performing semantic understanding on a natural language query input by a user and converting it into a corresponding query feature vector; calculating the similarity between the query feature vector and the multi-modal feature vectors in a database, sorting the scene units according to the similarity, and returning the most similar scene unit; presenting the retrieved scene units to the user after sorting them according to the similarity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent search, and particularly relates to a method and a system for searching based on a multimodal model. Background Art

[0002] Existing technologies have technical solutions for searching pictures and videos using text. However, most of the existing technologies preprocess pictures and videos to obtain a single-modal data set for subsequent search. Even when using a multimodal model for search, the methods are relatively complex, increasing the implementation cost. Summary of the Invention

[0003] The present invention provides a method for searching based on a multimodal model, the method comprising the following steps:

[0004] S1. Segment the video, and use a set of frames that are consecutive and have a similarity higher than a threshold in each frame of the video as a scene unit;

[0005] S2. Extract key frames from the scene unit, and extract a key frame from each scene unit, where the key frame is the frame with the highest similarity to other frames in this scene unit:

[0006] For each image in the scene unit, calculate its cumulative difference degree from all other images, and select the image with the smallest cumulative difference degree as the key frame ,

[0007]

[0008] Scene unit contains images, denoted as , … , is the th frame image in the scene unit, is the th frame image at the position, and are the width and height of the image, is the subscript of the image with the smallest cumulative difference, is the extracted key frame;

[0009] S3. Extract image features from the key frames;

[0010] S4. Extract text features from key frames. If the scene unit to which the key frame belongs includes subtitles or dialogues, use optical character recognition / speech recognition to convert them into text, and convert the text into text feature vectors through a language model. If the scene unit to which the key frame belongs has no subtitles or dialogues, generate text feature vectors of the key frame through the multimodal model Qwen2-VL;

[0011] S5. Fuse the key frame image features and text vector features to obtain multimodal features reflecting the content of the scene unit :

[0012]

[0013] 、 is the weight matrix parameter, is the key frame image feature vector, is the text feature vector, is the softmax function, is the multimodal feature;

[0014] S6. Embed and store the extracted multimodal features into a vector database that supports fast vector similarity search for quickly retrieving video feature vectors similar to the query vector;

[0015] S7. Establish a mapping relationship between the multimodal feature vectors and the scene units in the database. After retrieving similar feature vectors, quickly locate the corresponding video segments;

[0016] S8. Semantically understand the natural language query input by the user and convert it into a corresponding query feature vector;

[0017] S9. Calculate the similarity between the query feature vector and the multimodal feature vectors in the database, sort the scene units according to the similarity, and return the most similar scene units;

[0018] S10. Present the retrieved scene units to the user after sorting them according to the similarity, including displaying preview images and timestamp information of the video segments.

[0019] Furthermore, step S1 specifically includes:

[0020] Use the SSIM algorithm to analyze each frame of the video to calculate the similarity between frames. When the similarity is greater than the threshold T, divide these frames into a scene unit ,

[0021]

[0022] Among them, represents two images of adjacent frames The similarity between the image and is the pixel mean value of the image , is the pixel mean value of the image . is the pixel standard deviation of the image , is the pixel standard deviation of the image . is the covariance between the image and the image , , are constants, T is the threshold constant, is a scene unit obtained.

[0023] Furthermore, step S3 specifically includes:

[0024] Extract the image feature vector based on the ResNet network,

[0025]

[0026]

[0027]

[0028] Among them, is the extracted image feature vector, is the input key-frame image, is the activation function, is the normalization function, , , are the weight matrix parameters, is the bias parameter, is the number of residual block parameters, is the initial input, is the th output local feature map of the residual block, is the th output local feature map of the residual block, is the th output local feature map of the residual block at the position's pixel value, , is the th output local feature map of the residual block 's height and width, is the function to obtain the feature vector, which converts the obtained feature image into a feature vector.

[0029] Furthermore, step S6 specifically includes:

[0030] Embed and store the multi-modal feature vectors into the Milvus database,

[0031]

[0032] where, is the multi-modal feature vector stored in the milvus database, is the multi-modal feature vector for embedded storage, is its corresponding unique identifier, is the database operation function to implement the operation of storing the feature vectors into the database;

[0033] Step S7 specifically includes:

[0034] Set a unique identifier for each embedded multi-modal feature vector in the database and associate it with the corresponding scenario unit. Record the mapping relationship between the multi-modal feature vectors and the scenario units by creating a mapping table,

[0035]

[0036] where, is the mapping relationship between the multi-modal feature vectors stored in the database and the scenario units, is the multi-modal feature vector, is the unique identifier of the multi-modal feature vector, is the corresponding scenario unit, is the timestamp, is the mapping function.

[0037] Furthermore, step S8 specifically includes:

[0038] Use the BERT language model to perform semantic understanding on the user's natural language query and obtain the query feature vector ,

[0039]

[0040] where, is the query feature vector, () is the feed-forward neural network function, is the excessive self-attention mechanism function, is the input of the th layer of the BERT model, that is, the output of the th layer, is the input of the first layer, () is the word embedding function, , , is the parameter matrix of the l-th layer.

[0041] The present invention also relates to a system for searching based on a multimodal model. When the system is used, the method for searching based on a multimodal model described above is adopted. The system includes:

[0042] A video segmentation unit, configured to segment a video, and use a set of frames that are coherent before and after and have a similarity higher than a threshold in each frame of the video as a scene unit;

[0043] A key frame extraction unit, configured to extract key frames from the scene unit, and extract a key frame from each scene unit. The key frame is the frame with the highest similarity to other frames in this scene unit;

[0044] A key frame image feature extraction unit, configured to extract image features from the key frames;

[0045] A key frame text feature extraction unit, configured to extract text features from the key frames. If the scene unit to which the key frame belongs includes subtitles or dialogues, use optical character recognition / speech recognition to convert them into text, and then convert the text into a text feature vector through a language model; if the scene unit to which the key frame belongs does not include subtitles or dialogues, generate a text feature vector of the key frame through the multimodal model Qwen2-VL;

[0046] A fusion unit, configured to fuse the key frame image features and text vector features to obtain multimodal features reflecting the content of the scene unit;

[0047] A retrieval unit, configured to embed the extracted multimodal features into a vector database that supports fast vector similarity search, and use it to quickly retrieve video feature vectors similar to the query vector;

[0048] A mapping unit, configured to establish a mapping relationship between the multimodal feature vectors and the scene units in the database, so as to quickly locate the corresponding video segment after retrieving similar feature vectors;

[0049] A query unit, configured to perform semantic understanding on the natural language query input by the user and convert it into a corresponding query feature vector;

[0050] A searching unit, configured to perform similarity calculation between the query feature vector and the multimodal feature vectors in the database, sort the scene units according to the similarity, and return the most similar scene unit;

[0051] A display unit, configured to present the retrieved scene units to the user after sorting them according to the similarity, including displaying preview images and timestamp information of the video segments.

[0052] The present invention also relates to a computer program product, which includes a computer program. The computer program is executed by a processor and is used to execute the above-mentioned method for searching based on a multi-modal model.

[0053] The present invention also relates to a computer-readable storage medium, which is used to store a computer program. The computer program is executed by a processor and is used to execute the above-mentioned method for searching based on a multi-modal model.

[0054] The technical solution of the present invention reduces the computational complexity of video processing by fusing the multi-modal features of video key frames, simplifies the method, reduces costs, and semantically understands the natural language query input by the user, calculates the similarity between the obtained query feature vector and the video multi-modal feature vector, sorts the scene units according to the similarity, and returns the most similar scene units to be presented to the user, including displaying preview images and timestamp information of video segments, thereby improving the search efficiency. Detailed implementation manners

[0055] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention and should not be used to limit the protection scope of the present invention. It should be noted that the following detailed descriptions are all exemplary and are intended to provide further explanations of the present invention.

[0056] Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "include" and / or "comprise" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0057] Embodiment 1 of the present invention relates to a method for searching based on a multi-modal model. The method includes the following steps:

[0058] S1. Segment the video, and use the set of frames that are coherent before and after and have a similarity higher than a threshold in each frame of the video as a scene unit.

[0059] Specifically, it includes using the SSIM algorithm to analyze each frame of the video to calculate the similarity between frames. When the similarity is greater than the threshold T, these frames are divided into a scene unit. ,

[0060]

[0061] Among them, two images representing adjacent frames and the image the similarity between is the average pixel value of the image ; is the average pixel value of the image ; is the standard deviation of the pixels of the image ; is the standard deviation of the pixels of the image ; is the covariance of the image and the image ; , are constants, and T is a threshold constant. is a scene unit obtained.

[0062] S2. Extract key frames from the scene unit. Extract key frames from each scene unit. The key frame is the frame with the highest similarity to other frames in this scene unit.

[0063] Specifically, for each image in the scene unit, calculate its cumulative difference degree from all other images, and select the image with the smallest cumulative difference degree as the key frame. ,

[0064]

[0065] Among them, the number of images contained in the scene unit is , denoted as , … , then is the th frame image in the scene unit, is the th frame image at the position, and are the width and height of the image, is the subscript of the image with the smallest cumulative difference, is the extracted key frame.

[0066] S3. Extract image features from the key frame.

[0067] Specifically, extract the image feature vector based on the ResNet network.

[0068]

[0069]

[0070]

[0071] Among them, is the extracted image feature vector, is the input key-frame image, is the activation function, is the normalization function, 、 、 are the weight matrix parameters, is the bias parameter, is the number of residual block parameter, is the initial input, is the th output local feature map of the residual block, is the th output local feature map of the residual block, is the th output local feature map of the residual block at the position of the pixel value, , is the th output local feature map of the residual block height and width, is the function to obtain the feature vector, which converts the obtained feature image into a feature vector.

[0072] S4. Extract text features from the key frame. If the scene unit to which the key frame belongs includes subtitles or dialogues, use optical character recognition / speech recognition to convert them into text, and then convert the text into a text feature vector through a language model; if the scene unit to which the key frame belongs does not have subtitles or dialogues, generate the text feature vector of the key frame through the multimodal model Qwen2-VL.

[0073] Specifically, if the scene unit to which the key frame belongs includes subtitles or dialogues, use optical character recognition / speech recognition to convert them into text, and then convert the text into a text feature vector through a language model ,

[0074]

[0075] Among them, is the obtained text feature vector, is the BERT language model, is the optical character recognition function, is the speech recognition function to extract speech into text, is the scene unit;

[0076] When there is no subtitle or dialogue in the scene unit, the text description of the key frame is generated by Qwen2-VL, and then the text description is input into the BERT model to obtain the text feature vector ,

[0077]

[0078] Among them, is the obtained text feature vector, is the BERT language model, is the scene unit, is the multimodal model.

[0079] S5. Fuse the key frame image features and text vector features to obtain multimodal features reflecting the content of the scene unit.

[0080] Specifically, the multimodal features ,

[0081]

[0082] Among them, , are the weight matrix parameters, is the key frame image feature vector, is the text feature vector, is the softmax function, is the multimodal feature.

[0083] S6. Embed and store the extracted multimodal features into a vector database that supports fast vector similarity search for quickly retrieving video feature vectors similar to the query vector.

[0084] Specifically, embed and store the multimodal feature vector into the Milvus database,

[0085]

[0086] Among them, is the multimodal feature vector stored in the milvus database, is the multimodal feature vector embedded and stored, is its corresponding unique identifier, is the database operation function to implement the operation of storing the feature vector into the database.

[0087] S7. Establish a mapping relationship between the multimodal feature vector and the scene unit in the database, so that after retrieving the similar feature vector, the corresponding video segment can be quickly located.

[0088] Specifically, a unique identifier is set for each multi-modal feature vector stored in the database, and it is associated with the corresponding scene unit. The mapping relationship between the multi-modal feature vector and the scene unit is recorded by creating a mapping table.

[0089]

[0090] Among them, is the mapping relationship between the multi-modal feature vector stored in the database and the scene unit, is the multi-modal feature vector, is the unique identifier of the multi-modal feature vector, is the corresponding scene unit, is the timestamp, is the mapping function.

[0091] S8. Perform semantic understanding on the natural language query input by the user and convert it into the corresponding query feature vector.

[0092] Specifically, use the BERT language model to perform semantic understanding on the natural language query of the user and obtain the query feature vector ,

[0093]

[0094] Among them, is the query feature vector, () is the feed-forward neural network function, is the excessive self-attention mechanism function, is the input of the th layer of the BERT model, that is, the output of the th layer, is the input of the first layer, is the natural language query of the user, , , are the parameter matrices of the

[0095] S9. Calculate the similarity between the query feature vector and the multi-modal feature vector in the database, sort the scene units according to the similarity, and return the most similar scene unit.

[0096] Specifically, it includes,

[0097]

[0098] Among them, is the query feature vector, is the th value of the query feature vector, For the th stored multi-modal feature vector, is the th value of the th multi-modal feature vector, is the number of multi-modal feature vectors, is a sorting function that sorts according to the similarity value, ) is a mapping function that maps to find the mapping table information stored in the corresponding database according to the corresponding multi-modal feature vector returned by the sorting function.

[0099] S10. Present the retrieved scene units to the user after sorting by similarity, including displaying preview images and timestamp information of video clips, so as to facilitate the user to quickly browse and select.

[0100] Embodiment 2 of the present invention relates to a system for searching based on a multi-modal model. The system uses the method for searching based on a multi-modal model described in Embodiment 1. The system includes:

[0101] A video segmentation unit for segmenting a video and taking a set of frames that are coherent before and after and have a similarity higher than a threshold in each frame of the video as a scene unit;

[0102] A key frame extraction unit for extracting key frames from the scene units. The key frame is the frame with the highest similarity to other frames in this scene unit;

[0103] A key frame image feature extraction unit for extracting image features from the key frames;

[0104] A key frame text feature extraction unit for extracting text features from the key frames. If the scene unit to which the key frame belongs includes subtitles or dialogues, use optical character recognition / speech recognition to convert them into text, and then convert the text into text feature vectors through a language model; if the scene unit to which the key frame belongs does not have subtitles or dialogues, generate text feature vectors of the key frames through the multi-modal model Qwen2-VL;

[0105] A fusion unit for fusing the key frame image features and text vector features to obtain multi-modal features reflecting the content of the scene unit;

[0106] A retrieval unit for embedding the extracted multi-modal features into a vector database that supports fast vector similarity search to quickly retrieve video feature vectors similar to the query vector;

[0107] A mapping unit, configured to establish a mapping relationship between multi-modal feature vectors and scene units in a database, so that after similar feature vectors are retrieved, the corresponding video segments can be quickly located;

[0108] A query unit, configured to perform semantic understanding on the natural language query input by a user and convert it into a corresponding query feature vector;

[0109] A search unit, configured to perform similarity calculation between the query feature vector and the multi-modal feature vectors in the database, sort the scene units according to the similarity, and return the most similar scene unit;

[0110] A display unit, configured to present the retrieved scene units to the user after sorting them according to the similarity, including displaying preview images and timestamp information of the video segments.

[0111] Embodiment 3 of the present invention relates to a computer program product, the computer program product includes a computer program, and the computer program is executed by a processor for executing a method for searching based on a multi-modal model in Embodiment 2.

[0112] Embodiment 4 of the present invention relates to a computer-readable storage medium, the computer-readable storage medium is used to store a computer program, and the computer program is executed by a processor for executing a method for searching based on a multi-modal model in Embodiment 2.

[0113] As described above, the above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A method for search based on a multimodal model, characterized in that: S1. Segment the video, and use the set of frames that are coherent before and after and have a similarity higher than a threshold in each frame of the video as a scene unit; S2. Extract key frames from the scene units, and extract key frames from each scene unit. The key frame is the frame with the highest similarity to other frames in this scene unit: For each image in the scene unit, calculate its cumulative difference degree with all other images, and select the image with the smallest cumulative difference degree as the key frame , ; Scene unit contains images, denoted as , … , is the -th frame image in the scene unit, is the pixel value of the -th frame image at the position, and are the width and height of the image, is the subscript of the image with the smallest cumulative difference, is the extracted key frame; S3. Extract image features from the key frames; S4. Extract text features from the key frames. If the scene unit to which the key frame belongs includes subtitles or dialogues, use optical character recognition / speech recognition to convert them into text, and convert the text into text feature vectors through a language model; if the scene unit to which the key frame belongs has no subtitles or dialogues, generate text feature vectors of the key frames through the multimodal model Qwen2-VL; S5. Fuse the key-frame image features and text vector features to obtain multi-modal features reflecting the content of the scene unit : ; , is the weight matrix parameter, is the key frame image feature vector, is the text feature vector, is the softmax function, is the multi-modal feature; S6. Embed and store the extracted multimodal features into a vector database that supports fast vector similarity search for quickly retrieving video feature vectors similar to the query vector; S7. Establish a mapping relationship between the multimodal feature vectors and the scene units in the database. After retrieving similar feature vectors, quickly locate the corresponding video segments; S8. Semantically understand the natural language query input by the user and convert it into a corresponding query feature vector; S9. Calculate the similarity between the query feature vector and the multimodal feature vectors in the database, sort the scene units according to the similarity, and return the most similar scene unit; S10. Present the retrieved scene units to the user after sorting them according to the similarity, including displaying preview images and timestamp information of the video segments.

2. The method for search based on a multimodal model according to claim 1, wherein Step S1 specifically includes: Use the SSIM algorithm to analyze each frame of the video and calculate the similarity between frames. When the similarity is greater than the threshold T, these frames are divided into a scene unit , ; Among them, Two images representing adjacent frames and the image The similarity between is the pixel mean of the image ; is the pixel mean of the image ; is the pixel standard deviation of the image ; is the pixel standard deviation of the image ; is the covariance of the image and the image ; , are constants, T is a threshold constant, is a scene unit obtained.

3. A method for search based on a multi-modal model according to claim 1, characterized in that, Step S3 specifically includes: Extract image feature vectors based on the ResNet network, ; ; ; Among them, is the extracted image feature vector, is the input key-frame image, is the activation function, is the normalization function, , , are the weight matrix parameters, is the bias parameter, is the number of residual block parameter, is the initial input, is the th residual block output local feature map, is the th residual block output local feature map, is the th residual block output local feature map at the position pixel value, , is the th residual block output local feature map height and width, is the function to obtain the feature vector, which converts the obtained feature image into a feature vector.

4. A method for search based on a multimodal model according to claim 1, wherein Step S6 specifically includes: Embed and store the multimodal feature vectors into the Milvus database, ; Among them, is the multi-modal feature vector stored in the Milvus database, is the multi-modal feature vector stored by embedding, is its corresponding unique identifier, is the database operation function, which realizes the operation of storing the feature vector into the database; Step S7 specifically includes: Set a unique identifier for each embedded multimodal feature vector in the database and associate it with the corresponding scene unit, and record the mapping relationship between the multimodal feature vectors and the scene units by creating a mapping table, ; Among them, is the mapping relationship between the multi-modal feature vector stored in the database and the scene unit, is the multi-modal feature vector, is the unique identifier of the multi-modal feature vector, is the corresponding scene unit, is the timestamp, is the mapping function.

5. A method for search based on a multimodal model according to claim 1, wherein Step S8 specifically includes: Using the BERT language model, semantically understand the user's natural language query to obtain a query feature vector , ; Among them, is the query feature vector, () is the feedforward neural network function, is the excessive self-attention mechanism function, is the input of the th layer of the BERT model, that is, the output of the th layer, is the input of the first layer, is the natural language query of the user, , , are the parameter matrices of the lth layer.

6. A system for performing searches based on a multi-modal model, the system using a method for performing searches based on a multi-modal model according to any one of claims 1-5, characterized in that, The system includes: A video segmentation unit for segmenting the video and using the set of frames that are coherent before and after and have a similarity higher than a threshold in each frame of the video as a scene unit; A key frame extraction unit for extracting key frames from the scene units and extracting key frames from each scene unit. The key frame is the frame with the highest similarity to other frames in this scene unit; A key frame image feature extraction unit for extracting image features from the key frames; A key frame text feature extraction unit for extracting text features from the key frames. If the scene unit to which the key frame belongs includes subtitles or dialogues, use optical character recognition / speech recognition to convert them into text, and then convert the text into text feature vectors through a language model; if the scene unit to which the key frame belongs has no subtitles or dialogues, generate text feature vectors of the key frames through the multimodal model Qwen2-VL; A fusion unit for fusing the key frame image features and text vector features to obtain multimodal features reflecting the content of the scene unit; A retrieval unit, configured to embed and store the extracted multi-modal features into a vector database that supports fast vector similarity search, for quickly retrieving video feature vectors similar to a query vector; A mapping unit, configured to establish a mapping relationship between the multi-modal feature vectors and the scene units in the database, so as to quickly locate the corresponding video segments after retrieving the similar feature vectors; A query unit, configured to perform semantic understanding on the natural language query input by the user and convert it into a corresponding query feature vector; A search unit, configured to perform similarity calculation between the query feature vector and the multi-modal feature vectors in the database, sort the scene units according to the similarity, and return the most similar scene units; A display unit, configured to present the retrieved scene units sorted by similarity to the user, including displaying preview images and timestamp information of the video segments.

7. A computer program product, characterized in that, The computer program product includes a computer program, which is executed by a processor and is used to execute a method for search based on a multi-modal model according to any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, which is executed by a processor and is used to execute a method for search based on a multi-modal model according to any one of claims 1-5.

Citation Information

Patent Citations

  • Video retrieval method

    CN117251598A

  • Video processing method and video retrieval enhancement method for multi-modal large model

    CN119339284A