Method and system for searching based on multi-modal model
By segmenting and keyframe extraction of videos, and combining image and text features to generate multimodal feature vectors, the problems of complex and costly searching methods of multimodal model in the prior art are solved, and efficient video search and positioning are achieved.
Patent Information
- Application Number
- CN202510422711.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-07
AI Technical Summary
When searching using multimodal models, the prior art methods are complex and costly, making it difficult to achieve efficient video search.
By segmenting the video, extracting keyframes, and fusing it with image features and text features, multimodal feature vectors are generated and stored in a database that supports fast vector similarity search, achieving rapid retrieval and positioning.
Reduces the computational volume of video processing, simplifies methods, reduces costs, and improves search efficiency, enabling fast positioning and presenting the most similar video clips.
Smart Images

Figure SMS_2 
Figure SMS_21 
Figure SMS_36
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent search technology, and specifically relates to a method and system for searching based on a multimodal model. Background Art
[0002] There are technical solutions for searching pictures and videos using text in the prior art, but most of the prior art processes pictures and videos in advance to obtain a single-modal data set and then implement the search. Even if a multi-modal model is used for search, the method is relatively complicated and increases the implementation cost. Summary of the invention
[0003] The present invention provides a method for searching based on a multimodal model, the method comprising the following steps: S101. Segment the video, and take a set of frames in each frame of the video that are coherent and have a similarity higher than a threshold as a scene unit; S102. Extract key frames from the scene units, extract key frames from each scene unit, and the key frame is the frame with the highest similarity to other frames in the scene unit; S103. Extract image features from key frames; S104. Extract text features from the key frame. If the scene unit to which the key frame belongs includes subtitles or dialogues, convert it into text using text recognition / speech recognition, and then convert the text into a text feature vector using a language model. If the scene unit to which the key frame belongs does not have subtitles or dialogues, generate a text feature vector for the key frame using the multimodal model Qwen2-VL. S105. Fusing the key frame image features and the text vector features to obtain multimodal features reflecting the content of the scene unit; S106. Embedding and storing the extracted multimodal features into a vector database supporting fast vector similarity search, so as to quickly retrieve video feature vectors similar to the query vector; S107. Establishing a mapping relationship between multimodal feature vectors and scene units in the database, so as to quickly locate the corresponding video clip after retrieving similar feature vectors; S108. Perform semantic understanding on the natural language query input by the user and convert it into a corresponding query feature vector; S109. Calculate the similarity between the query feature vector and the multimodal feature vector in the database, sort the scene units according to the similarity, and return the most similar scene unit; S110. Sort the retrieved scene units by similarity and present them to the user, including displaying a preview image of the video clip and timestamp information.
[0004] Further, step S101 specifically includes: Use the SSIM algorithm to analyze the video frame by frame and calculate the similarity between frames. When the similarity is greater than the threshold T, these frames are divided into a scene unit. ,
[0005] in, Two images representing adjacent frames With image The similarity between For images The pixel mean, For images The pixel mean, For images The pixel standard deviation is For images The pixel standard deviation is For images and images The covariance of , is a constant, T is a threshold constant, A scene unit is obtained.
[0006] Further, step S102 specifically includes: For each image in the scene unit, calculate the cumulative difference between it and all other images, and select the image with the smallest cumulative difference as the key frame ,
[0007] Among them, the scene unit The number of images contained in , denoted as , … ,but For the scene unit Frame image, For the Frame image in The pixel value at position, and are the width and height of the image, is the image index with the smallest cumulative difference, The extracted key frames.
[0008] Further, step S103 specifically includes: Extract image feature vectors based on ResNet network.
[0009]
[0010]
[0011] in, is the extracted image feature vector, is the input key frame image, is the activation function, is the normalization function, , , is the weight matrix parameter, is the bias parameter, is the number of residual blocks, is the initial input, For the The residual block outputs a local feature map, For the The residual block outputs a local feature map, For the The residual block outputs a local feature map exist The pixel value at position, , For the The residual block outputs a local feature map The height and width of To obtain the eigenvector function, the obtained eigenimage is converted into a eigenvector.
[0012] Further, step S105 specifically includes: Get multimodal features ,
[0013] in, , is the weight matrix parameter, is the key frame image feature vector, is the text feature vector, is the softmax function, It is a multimodal feature.
[0014] Further, step S106 specifically includes: The multimodal feature vectors are embedded and stored in the Milvus database.
[0015] in, is the multimodal feature vector stored in the milvus database, is the multimodal feature vector stored in the embedding, is its corresponding unique identifier, It is a database operation function that implements the operation of storing feature vectors in the database; Step S107 specifically includes: A unique identifier is set for each embedded and stored multimodal feature vector in the database, and it is associated with the corresponding scene unit. A mapping table is created to record the mapping relationship between the multimodal feature vector and the scene unit.
[0016] in, is the mapping relationship between the multimodal feature vector and the scene unit stored in the database, is the multimodal feature vector, is the unique identifier of the multimodal feature vector, is the corresponding scene unit, is the timestamp, is the mapping function.
[0017] Further, step S108 specifically includes: Use the BERT language model to semantically understand the user's natural language query and obtain the query feature vector ,
[0018] in, is the query feature vector, () is the feedforward neural network function, is the excessive self-attention mechanism function, For the BERT model The input of the layer is The output of the layer, is the input of the first layer, For users’ natural language queries, () is the word embedding function, , , is the parameter matrix of the lth layer.
[0019] The present invention also relates to a system for searching based on a multimodal model, wherein the system uses the method for searching based on a multimodal model as described above, and the system comprises: A video segmentation unit is used to segment the video, and a set of frames in each frame of the video that are coherent and have a similarity higher than a threshold is taken as a scene unit; A key frame extraction unit is used to extract key frames from scene units, and extract key frames from each scene unit. The key frame is a frame with the highest similarity to other frames in the scene unit. A key frame image feature extraction unit, used for extracting image features from key frames; The key frame text feature extraction unit is used to extract text features from the key frame. If the scene unit to which the key frame belongs includes subtitles or dialogues, text recognition / speech recognition is used to convert them into text, and then the text is converted into a text feature vector through a language model. If the scene unit to which the key frame belongs does not have subtitles or dialogues, the text feature vector of the key frame is generated through the multimodal model Qwen2-VL. A fusion unit is used to fuse key frame image features and text vector features to obtain multimodal features reflecting the content of scene units; A retrieval unit, used to embed and store the extracted multimodal features into a vector database that supports fast vector similarity search, so as to quickly retrieve video feature vectors similar to the query vector; A mapping unit, used to establish a mapping relationship between multimodal feature vectors and scene units in a database, so as to quickly locate the corresponding video clip after retrieving similar feature vectors; A query unit, used to understand the semantics of the natural language query input by the user and convert it into a corresponding query feature vector; A search unit is used to calculate the similarity between the query feature vector and the multimodal feature vector in the database, sort the scene units according to the similarity, and return the most similar scene unit; The display unit is used to sort the retrieved scene units according to similarity and present them to the user, including displaying a preview image of the video clip and timestamp information.
[0020] The present invention also relates to a computer program product, which comprises a computer program, and the computer program is executed by a processor to execute the above-mentioned method for searching based on a multimodal model.
[0021] The present invention also relates to a computer-readable storage medium, wherein the computer-readable storage medium is used to store a computer program, and the computer program is executed by a processor to perform the above-mentioned method of searching based on a multimodal model.
[0022] The technical solution of the present invention reduces the amount of computation required for video processing, simplifies the method, and reduces costs by integrating the multimodal features of video key frames. It also performs semantic understanding on natural language queries input by users, calculates similarity between the acquired query feature vector and the video multimodal feature vector, sorts scene units according to similarity, and returns the most similar scene unit to present to the user, including displaying a preview image of the video clip and timestamp information, thereby improving search efficiency. DETAILED DESCRIPTION
[0023] The following examples are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present invention. It should be pointed out that the following detailed descriptions are all exemplary and are intended to provide further explanation of the present invention.
[0024] Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those generally understood by those of ordinary skill in the art to which the present invention belongs. It should be noted that the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.
[0025] Embodiment 1 of the present invention relates to a method for searching based on a multimodal model, the method comprising the following steps:
[0026] S101. Segment the video, and take a set of frames in the video that are coherent and have a similarity higher than a threshold as a scene unit.
[0027] Specifically, the SSIM algorithm is used to analyze the video frame by frame to calculate the similarity between frames, and when the similarity is greater than a threshold T, these frames are divided into a scene unit. ,
[0028]
[0029] in, Two images representing adjacent frames With image The similarity between For images The pixel mean, For images The pixel mean, For images The pixel standard deviation is For images The pixel standard deviation is For images and images The covariance of , is a constant, T is a threshold constant, A scene unit is obtained.
[0030] S102. Extract key frames from the scene units. Extract key frames from each scene unit. The key frame is a frame with the highest similarity to other frames in the scene unit.
[0031] Specifically, for each image in the scene unit, calculate the cumulative difference between it and all other images, and select the image with the smallest cumulative difference as the key frame. ,
[0032]
[0033] Among them, the scene unit The number of images contained in , denoted as , … ,but For the scene unit Frame image, For the Frame image in The pixel value at position, and are the width and height of the image, is the image index with the smallest cumulative difference, The extracted key frames.
[0034] S103. Extract image features from key frames.
[0035] Specifically, it includes extracting image feature vectors based on the ResNet network,
[0036]
[0037]
[0038] in, is the extracted image feature vector, is the input key frame image, is the activation function, is the normalization function, , , is the weight matrix parameter, is the bias parameter, is the number of residual blocks, is the initial input, For the The residual block outputs a local feature map, For the The residual block outputs a local feature map, For the The residual block outputs a local feature map exist The pixel value at position, , For the The residual block outputs a local feature map The height and width of To obtain the eigenvector function, the obtained eigenimage is converted into a eigenvector.
[0039] S104. Perform text feature extraction on the key frame. If the scene unit to which the key frame belongs includes subtitles or dialogues, convert them into text using text recognition / speech recognition, and then convert the text into a text feature vector using a language model. If the scene unit to which the key frame belongs does not have subtitles or dialogues, generate a text feature vector for the key frame using the multimodal model Qwen2-VL.
[0040] Specifically, if the scene unit to which the key frame belongs includes subtitles or dialogues, text recognition / speech recognition is used to convert them into text, and then the text is converted into a text feature vector through a language model. ,
[0041] in, is the obtained text feature vector, is the BERT language model, is the text recognition function, Extract speech into text for speech recognition function. It is a scene unit; If there is no subtitle or dialogue in the scene unit, Qwen2-VL is used to generate a text description of the key frame, and then the text description is input into the BERT model to obtain a text feature vector. ,
[0042] in, is the obtained text feature vector, is the BERT language model, is the scene unit, It is a multimodal model.
[0043] S105. Fuse the key frame image features and the text vector features to obtain multimodal features reflecting the content of the scene unit.
[0044] Specifically, we obtain multimodal features ,
[0045] in, , is the weight matrix parameter, is the key frame image feature vector, is the text feature vector, is the softmax function, It is a multimodal feature.
[0046] S106. The extracted multimodal features are embedded and stored in a vector database that supports fast vector similarity search, so as to quickly retrieve video feature vectors that are similar to the query vector.
[0047] Specifically, the multimodal feature vectors are embedded and stored in the Milvus database.
[0048]
[0049] in, is the multimodal feature vector stored in the milvus database, is the multimodal feature vector stored in the embedding, is its corresponding unique identifier, It is a database operation function that implements the operation of storing feature vectors in the database.
[0050] S107. A mapping relationship between multimodal feature vectors and scene units is established in the database, so that after similar feature vectors are retrieved, the corresponding video clips can be quickly located.
[0051] Specifically, a unique identifier is set in the database for each embedded and stored multimodal feature vector, and it is associated with the corresponding scene unit. A mapping table is created to record the mapping relationship between the multimodal feature vector and the scene unit.
[0052]
[0053] in, is the mapping relationship between the multimodal feature vectors and scene units stored in the database, is the multimodal feature vector, is the unique identifier of the multimodal feature vector, is the corresponding scene unit, is the timestamp, is the mapping function.
[0054] S108. Perform semantic understanding on the natural language query input by the user and convert it into a corresponding query feature vector.
[0055] Specifically, the BERT language model is used to semantically understand the user's natural language query and obtain the query feature vector ,
[0056]
[0057] in, is the query feature vector, () is the feedforward neural network function, is the excessive self-attention mechanism function, For the BERT model The input of the layer is The output of the layer, is the input of the first layer, For users’ natural language queries, () is the word embedding function, , , is the parameter matrix of the lth layer.
[0058] S109. Calculate the similarity between the query feature vector and the multimodal feature vector in the database, sort the scene units according to the similarity, and return the most similar scene unit.
[0059] Specifically include:
[0060]
[0061] in, is the query feature vector, is the first feature vector of the query values, For storage multimodal feature vectors, For the The first multimodal feature vector values, is the number of multimodal feature vectors, is a sorting function that sorts according to the similarity value. ) is a mapping function, which maps the corresponding multimodal feature vector returned by the sorting function to find the mapping table information stored in the corresponding database.
[0062] S110. The retrieved scene units are sorted according to similarity and presented to the user, including displaying a preview image of the video clip and timestamp information, so as to facilitate quick browsing and selection by the user.
[0063] Embodiment 2 of the present invention relates to a system for searching based on a multimodal model, the system using the method for searching based on a multimodal model described in Embodiment 1, the system comprising: A video segmentation unit is used to segment the video, and a set of frames in each frame of the video that are coherent and have a similarity higher than a threshold is taken as a scene unit; A key frame extraction unit is used to extract key frames from scene units, and extract key frames from each scene unit. The key frame is a frame with the highest similarity to other frames in the scene unit. A key frame image feature extraction unit, used for extracting image features from key frames; The key frame text feature extraction unit is used to extract text features from the key frame. If the scene unit to which the key frame belongs includes subtitles or dialogues, text recognition / speech recognition is used to convert them into text, and then the text is converted into a text feature vector through a language model. If the scene unit to which the key frame belongs does not have subtitles or dialogues, the text feature vector of the key frame is generated through the multimodal model Qwen2-VL. A fusion unit is used to fuse key frame image features and text vector features to obtain multimodal features reflecting the content of scene units; A retrieval unit, used to embed and store the extracted multimodal features into a vector database that supports fast vector similarity search, so as to quickly retrieve video feature vectors similar to the query vector; A mapping unit, used to establish a mapping relationship between multimodal feature vectors and scene units in a database, so as to quickly locate the corresponding video clip after retrieving similar feature vectors; A query unit, used to understand the semantics of the natural language query input by the user and convert it into a corresponding query feature vector; A search unit is used to calculate the similarity between the query feature vector and the multimodal feature vector in the database, sort the scene units according to the similarity, and return the most similar scene unit; The display unit is used to sort the retrieved scene units according to similarity and present them to the user, including displaying a preview image of the video clip and timestamp information.
[0064] Embodiment 3 of the present invention relates to a computer program product, which includes a computer program. The computer program is executed by a processor and is used to execute an electronically controlled differential lock control and method for autonomously escaping an unmanned vehicle according to embodiment 2.
[0065] Embodiment 4 of the present invention relates to a computer-readable storage medium, which is used to store a computer program. The computer program is executed by a processor to execute an electronically controlled differential lock control method for autonomously escaping an unmanned vehicle according to embodiment 2.
[0066] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A method for searching based on a multimodal model, characterized in that: The method comprises the following steps: S101. Segment the video, and take a set of frames in each frame of the video that are coherent and have a similarity higher than a threshold as a scene unit; S102. Extract key frames from the scene units, extract key frames from each scene unit, and the key frame is the frame with the highest similarity to other frames in the scene unit; S103. Extract image features from key frames; S104. Extract text features from the key frame. If the scene unit to which the key frame belongs includes subtitles or dialogues, convert it into text using text recognition / speech recognition, and then convert the text into a text feature vector using a language model. If the scene unit to which the key frame belongs does not have subtitles or dialogues, generate a text feature vector for the key frame using the multimodal model Qwen2-VL. S105. Fusing the key frame image features and the text vector features to obtain multimodal features reflecting the content of the scene unit; S106. Embedding and storing the extracted multimodal features into a vector database supporting fast vector similarity search, so as to quickly retrieve video feature vectors similar to the query vector; S107. Establishing a mapping relationship between multimodal feature vectors and scene units in the database, so as to quickly locate the corresponding video clip after retrieving similar feature vectors; S108. Perform semantic understanding on the natural language query input by the user and convert it into a corresponding query feature vector; S109. Calculate the similarity between the query feature vector and the multimodal feature vector in the database, sort the scene units according to the similarity, and return the most similar scene unit; S110. Sort the retrieved scene units by similarity and present them to the user, including displaying a preview image of the video clip and timestamp information.
2. The method for searching based on a multimodal model according to claim 1, characterized in that: Step S101 specifically includes: Use the SSIM algorithm to analyze the video frame by frame and calculate the similarity between frames. When the similarity is greater than the threshold T, these frames are divided into a scene unit. , ; in, Two images representing adjacent frames With image The similarity between For images The pixel mean, For images The pixel mean, For images The pixel standard deviation is For images The pixel standard deviation is For images and images The covariance of , is a constant, T is a threshold constant, A scene unit is obtained.
3. The method for searching based on a multimodal model according to claim 1, characterized in that: Step S102 specifically includes: For each image in the scene unit, calculate the cumulative difference between it and all other images, and select the image with the smallest cumulative difference as the key frame , ; Among them, the scene unit The number of images contained in , denoted as , … ,but For the scene unit Frame image, For the Frame image in The pixel value at position, and are the width and height of the image, is the image index with the smallest cumulative difference, The extracted key frames.
4. The method for searching based on a multimodal model according to claim 1, characterized in that: Step S103 specifically includes: Extract image feature vectors based on ResNet network. ; ; ; in, is the extracted image feature vector, is the input key frame image, is the activation function, is the normalization function, , , is the weight matrix parameter, is the bias parameter, is the number of residual blocks, is the initial input, For the The residual block outputs a local feature map, For the The residual block outputs a local feature map, For the The residual block outputs a local feature map exist The pixel value at position, , For the The residual block outputs a local feature map The height and width of To obtain the eigenvector function, the obtained eigenimage is converted into a eigenvector.
5. The method for searching based on a multimodal model according to claim 1, characterized in that: Step S105 specifically includes: Get multimodal features , ; in, , is the weight matrix parameter, is the key frame image feature vector, is the text feature vector, is the softmax function, It is a multimodal feature.
6. The method for searching based on a multimodal model according to claim 1, characterized in that: Step S106 specifically includes: The multimodal feature vectors are embedded and stored in the Milvus database. ; in, is the multimodal feature vector stored in the milvus database, is the multimodal feature vector stored in the embedding, is its corresponding unique identifier, It is a database operation function that implements the operation of storing feature vectors in the database; Step S107 specifically includes: A unique identifier is set for each embedded and stored multimodal feature vector in the database, and it is associated with the corresponding scene unit. A mapping table is created to record the mapping relationship between the multimodal feature vector and the scene unit. ; in, is the mapping relationship between the multimodal feature vector and the scene unit stored in the database, is the multimodal feature vector, is the unique identifier of the multimodal feature vector, is the corresponding scene unit, is the timestamp, is the mapping function.
7. The method for searching based on a multimodal model according to claim 1, characterized in that: Step S108 specifically includes: Use the BERT language model to semantically understand the user's natural language query and obtain the query feature vector , ; in, is the query feature vector, () is the feedforward neural network function, is the excessive self-attention mechanism function, For the BERT model The input of the layer is The output of the layer, is the input of the first layer, For users’ natural language queries, () is the word embedding function, , , is the parameter matrix of the lth layer.
8. A system for searching based on a multimodal model, the system using a method for searching based on a multimodal model as claimed in any one of claims 1 to 7, characterized in that: The system comprises: A video segmentation unit is used to segment the video, and a set of frames in each frame of the video that are coherent and have a similarity higher than a threshold is taken as a scene unit; A key frame extraction unit is used to extract key frames from scene units, and extract key frames from each scene unit. The key frame is a frame with the highest similarity to other frames in the scene unit. A key frame image feature extraction unit, used for extracting image features from key frames; The key frame text feature extraction unit is used to extract text features from the key frame. If the scene unit to which the key frame belongs includes subtitles or dialogues, text recognition / speech recognition is used to convert them into text, and then the text is converted into a text feature vector through a language model. If the scene unit to which the key frame belongs does not have subtitles or dialogues, the text feature vector of the key frame is generated through the multimodal model Qwen2-VL. A fusion unit is used to fuse key frame image features and text vector features to obtain multimodal features reflecting the content of scene units; A retrieval unit, used to embed and store the extracted multimodal features into a vector database that supports fast vector similarity search, so as to quickly retrieve video feature vectors similar to the query vector; A mapping unit, used to establish a mapping relationship between multimodal feature vectors and scene units in a database, so as to quickly locate the corresponding video clip after retrieving similar feature vectors; A query unit, used to understand the semantics of the natural language query input by the user and convert it into a corresponding query feature vector; A search unit is used to calculate the similarity between the query feature vector and the multimodal feature vector in the database, sort the scene units according to the similarity, and return the most similar scene unit; The display unit is used to sort the retrieved scene units according to similarity and present them to the user, including displaying a preview image of the video clip and timestamp information.
9. A computer program product, characterized in that The computer program product comprises a computer program, which is executed by a processor and is used to execute a method for searching based on a multimodal model as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store a computer program, and the computer program is executed by a processor to execute a method for searching based on a multimodal model as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Image search video method and device based on scene dictionary tree and computer readable storage medium
CN110427517A
Text video cross-modal retrieval method and device based on fine granularity perception
CN116166843A
Video cross-modal search model training method, search method and device
CN116955699A
Video retrieval method
CN117251598A
Video similarity judgment method and device based on multi-modal feature fusion
CN118411536A
Cited By
Multi-modal data retrieval method based on semantic association
CN120849655A
Video retrieval method and server based on semantic embedding and video memory coding
CN121524397A