Method and system for searching based on multi-modal model

By segmenting and keyframe extraction of videos, and combining image and text features to generate multimodal feature vectors, the problems of complex and costly searching methods of multimodal model in the prior art are solved, and efficient video search and positioning are achieved.

CN119938986AActive Publication Date: 2025-05-06ZHEJIANG PROVINCIAL PUBLIC SECURITY SCIENCE & TECHNOLOGY RESEARCH INSTITUTE +3

Patent Information

Application Number
CN202510422711.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-05-06
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

When searching using multimodal models, the prior art methods are complex and costly, making it difficult to achieve efficient video search.

Method used

By segmenting the video, extracting keyframes, and fusing it with image features and text features, multimodal feature vectors are generated and stored in a database that supports fast vector similarity search, achieving rapid retrieval and positioning.

Benefits of technology

Reduces the computational volume of video processing, simplifies methods, reduces costs, and improves search efficiency, enabling fast positioning and presenting the most similar video clips.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_2
    Figure SMS_2
  • Figure SMS_21
    Figure SMS_21
  • Figure SMS_36
    Figure SMS_36
Patent Text Reader

Abstract

The invention discloses a method and a system for searching based on a multi-modal model. The method comprises the following steps: segmenting a video, and taking a frame set which is coherent front and back and has a similarity higher than a threshold value in each frame of the video as a scene unit; carrying out key frame extraction on the scene unit; carrying out image feature and text feature extraction on the key frame; fusing the key frame image features and the text vector features to obtain multi-modal features reflecting scene unit contents; performing semantic understanding on the natural language query input by the user, and converting the natural language query into a corresponding query feature vector; performing similarity calculation on the query feature vector and a multi-modal feature vector in a database, sorting the scene units according to the similarity, and returning the most similar scene unit; and sequencing the retrieved scene units according to the similarity and then presenting the scene units to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent search technology, and specifically relates to a method and system for searching based on a multimodal model. Background Art

[0002] There are technical solutions for searching pictures and videos using text in the prior art, but most of the prior art processes pictures and videos in advance to obtain a single-modal data set and then implement the search. Even if a multi-modal model is used for search, the method is relatively complicated and increases the implementation cost. Summary of the invention

[0003] The present invention provides a method for searching based on a multimodal model, the method comprising the following steps: S101. Segment the video, and take a set of frames in each frame of the video that are coherent and have a similarity higher than a threshold as a scene unit; S102. Extract key frames from the scene units, extract key frames from each scene unit, and the key frame is the frame with the highest similarity to other frames in the scene unit; S103. Extract image features from key frames; S104. Extract text features from the key frame. If the scene unit to which the key frame belongs includes subtitles or dialogues, convert it into text using text recognition / speech recognition, and then convert the text into a text feature vector using a language model. If the scene unit to which the key frame belongs does not have subtitles or dialogues, generate a text feature vector for the key frame using the multimodal model Qwen2-VL. S105. Fusing the key frame image features and the text vector features to obtain multimodal features reflecting the content of the scene unit; S106. Embedding and storing the extracted multimodal features into a vector database supporting fast vector similarity search, so as to quickly retrieve video feature vectors similar to the query vector; S107. Establishing a mapping relationship between multimodal feature vectors and scene units in the database, so as to quickly locate the corresponding video clip after retrieving similar feature vectors; S108. Perform semantic understanding on the natural language query input by the user and convert it into a corresponding query feature vector; S109. Calculate the similarity between the query feature vector and the multimodal feature vector in the database, sort the scene units according to the similarity, and return the most similar scene unit; S110. Sort the retrieved scene units by similarity and present them to the user, including displaying a preview image of the video clip and timestamp information.

[0004] Further, step S101 specifically includes: Use the SSIM algorithm to analyze the video frame by frame and calculate the similarity between frames. When the similarity is greater than the threshold T, these frames are divided into a scene unit. ,

[0005] in, Two images representing adjacent frames With image The similarity between For images The pixel mean, For images The pixel mean, For images The pixel standard deviation is For images The pixel standard deviation is For images and images The covariance of , is a constant, T is a threshold constant, A scene unit is obtained.

[0006] Further, step S102 specifically includes: For each image in the scene unit, calculate the cumulative difference between it and all other images, and select the image with the smallest cumulative difference as the key frame ,

[0007] Among them, the scene unit The number of images contained in , denoted as , … ,but For the scene unit Frame image, For the Frame image in The pixel value at position, and are the width and height of the image, is the image index with the smallest cumulative difference, The extracted key frames.

[0008] Further, step S103 specifically includes: Extract image feature vectors based on ResNet network.

[0009]

[0010]

[0011] in, is the extracted image feature vector, is the input key frame image, is the activation function, is the normalization function, , , is the weight matrix parameter, is the bias parameter, is the number of residual blocks, is the initial input, For the The residual block outputs a local feature map, For the The residual block outputs a local feature map, For the The residual block outputs a local feature map exist The pixel value at position, , For the The residual block outputs a local feature map The height and width of To obtain the eigenvector function, the obtained eigenimage is converted into a eigenvector.

[0012] Further, step S105 specifically includes: Get multimodal features ,

[0013] in, , is the weight matrix parameter, is the key frame image feature vector, is the text feature vector, is the softmax function, It is a multimodal feature.

[0014] Further, step S106 specifically includes: The multimodal feature vectors are embedded and stored in the Milvus database.

[0015] in, is the multimodal feature vector stored in the milvus database, is the multimodal feature vector stored in the embedding, is its corresponding unique identifier, It is a database operation function that implements the operation of storing feature vectors in the database; Step S107 specifically includes: A unique identifier is set for each embedded and stored multimodal feature vector in the database, and it is associated with the corresponding scene unit. A mapping table is created to record the mapping relationship between the multimodal feature vector and the scene unit.

[0016] in, is the mapping relationship between the multimodal feature vector and the scene unit stored in the database, is the multimodal feature vector, is the unique identifier of the multimodal feature vector, is the corresponding scene unit, is the timestamp, is the mapping function.

[0017] Further, step S108 specifically includes: Use the BERT language model to semantically understand the user's natural language query and obtain the query feature vector ,

[0018] in, is the query feature vector, () is the feedforward neural network function, is the excessive self-attention mechanism function, For the BERT model The input of the layer is The output of the layer, is the input of the first layer, For users’ natural language queries, () is the word embedding function, , , is the parameter matrix of the lth layer.

[0019] The present invention also relates to a system for searching based on a multimodal model, wherein the system uses the method for searching based on a multimodal model as described above, and the system comprises: A video segmentation unit is used to segment the video, and a set of frames in each frame of the video that are coherent and have a similarity higher than a threshold is taken as a scene unit; A key frame extraction unit is used to extract key frames from scene units, and extract key frames from each scene unit. The key frame is a frame with the highest similarity to other frames in the scene unit. A key frame image feature extraction unit, used for extracting image features from key frames; The key frame text feature extraction unit is used to extract text features from the key frame. If the scene unit to which the key frame belongs includes subtitles or dialogues, text recognition / speech recognition is used to convert them into text, and then the text is converted into a text feature vector through a language model. If the scene unit to which the key frame belongs does not have subtitles or dialogues, the text feature vector of the key frame is generated through the multimodal model Qwen2-VL. A fusion unit is used to fuse key frame image features and text vector features to obtain multimodal features reflecting the content of scene units; A retrieval unit, used to embed and store the extracted multimodal features into a vector database that supports fast vector similarity search, so as to quickly retrieve video feature vectors similar to the query vector; A mapping unit, used to establish a mapping relationship between multimodal feature vectors and scene units in a database, so as to quickly locate the corresponding video clip after retrieving similar feature vectors; A query unit, used to understand the semantics of the natural language query input by the user and convert it into a corresponding query feature vector; A search unit is used to calculate the similarity between the query feature vector and the multimodal feature vector in the database, sort the scene units according to the similarity, and return the most similar scene unit; The display unit is used to sort the retrieved scene units according to similarity and present them to the user, including displaying a preview image of the video clip and timestamp information.

[0020] The present invention also relates to a computer program product, which comprises a computer program, and the computer program is executed by a processor to execute the above-mentioned method for searching based on a multimodal model.

[0021] The present invention also relates to a computer-readable storage medium, wherein the computer-readable storage medium is used to store a computer program, and the computer program is executed by a processor to perform the above-mentioned method of searching based on a multimodal model.

[0022] The technical solution of the present invention reduces the amount of computation required for video processing, simplifies the method, and reduces costs by integrating the multimodal features of video key frames. It also performs semantic understanding on natural language queries input by users, calculates similarity between the acquired query feature vector and the video multimodal feature vector, sorts scene units according to similarity, and returns the most similar scene unit to present to the user, including displaying a preview image of the video clip and timestamp information, thereby improving search efficiency. DETAILED DESCRIPTION

[0023] The following examples are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present invention. It should be pointed out that the following detailed descriptions are all exemplary and are intended to provide further explanation of the present invention.

[0024] Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those generally understood by those of ordinary skill in the art to which the present invention belongs. It should be noted that the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.

[0025] Embodiment 1 of the present invention relates to a method for searching based on a multimodal model, the method comprising the following steps:

[0026] S101. Segment the video, and take a set of frames in the video that are coherent and have a similarity higher than a threshold as a scene unit.

[0027] Specifically, the SSIM algorithm is used to analyze the video frame by frame to calculate the similarity between frames, and when the similarity is greater than a threshold T, these frames are divided into a scene unit. ,

[0028]

[0029] in, Two images representing adjacent frames With image The similarity between For images The pixel mean, For images The pixel mean, For images The pixel standard deviation is For images The pixel standard deviation is For images and images The covariance of , is a constant, T is a threshold constant, A scene unit is obtained.

[0030] S102. Extract key frames from the scene units. Extract key frames from each scene unit. The key frame is a frame with the highest similarity to other frames in the scene unit.

[0031] Specifically, for each image in the scene unit, calculate the cumulative difference between it and all other images, and select the image with the smallest cumulative difference as the key frame. ,

[0032]

[0033] Among them, the scene unit The number of images contained in , denoted as , … ,but For the scene unit Frame image, For the Frame image in The pixel value at position, and are the width and height of the image, is the image index with the smallest cumulative difference, The extracted key frames.

[0034] S103. Extract image features from key frames.

[0035] Specifically, it includes extracting image feature vectors based on the ResNet network,

[0036]

[0037]

[0038] in, is the extracted image feature vector, is the input key frame image, is the activation function, is the normalization function, , , is the weight matrix parameter, is the bias parameter, is the number of residual blocks, is the initial input, For the The residual block outputs a local feature map, For the The residual block outputs a local feature map, For the The residual block outputs a local feature map exist The pixel value at position, , For the The residual block outputs a local feature map The height and width of To obtain the eigenvector function, the obtained eigenimage is converted into a eigenvector.

[0039] S104. Perform text feature extraction on the key frame. If the scene unit to which the key frame belongs includes subtitles or dialogues, convert them into text using text recognition / speech recognition, and then convert the text into a text feature vector using a language model. If the scene unit to which the key frame belongs does not have subtitles or dialogues, generate a text feature vector for the key frame using the multimodal model Qwen2-VL.

[0040] Specifically, if the scene unit to which the key frame belongs includes subtitles or dialogues, text recognition / speech recognition is used to convert them into text, and then the text is converted into a text feature vector through a language model. ,

[0041] in, is the obtained text feature vector, is the BERT language model, is the text recognition function, Extract speech into text for speech recognition function. It is a scene unit; If there is no subtitle or dialogue in the scene unit, Qwen2-VL is used to generate a text description of the key frame, and then the text description is input into the BERT model to obtain a text feature vector. ,

[0042] in, is the obtained text feature vector, is the BERT language model, is the scene unit, It is a multimodal model.

[0043] S105. Fuse the key frame image features and the text vector features to obtain multimodal features reflecting the content of the scene unit.

[0044] Specifically, we obtain multimodal features ,

[0045] in, , is the weight matrix parameter, is the key frame image feature vector, is the text feature vector, is the softmax function, It is a multimodal feature.

[0046] S106. The extracted multimodal features are embedded and stored in a vector database that supports fast vector similarity search, so as to quickly retrieve video feature vectors that are similar to the query vector.

[0047] Specifically, the multimodal feature vectors are embedded and stored in the Milvus database.

[0048]

[0049] in, is the multimodal feature vector stored in the milvus database, is the multimodal feature vector stored in the embedding, is its corresponding unique identifier, It is a database operation function that implements the operation of storing feature vectors in the database.

[0050] S107. A mapping relationship between multimodal feature vectors and scene units is established in the database, so that after similar feature vectors are retrieved, the corresponding video clips can be quickly located.

[0051] Specifically, a unique identifier is set in the database for each embedded and stored multimodal feature vector, and it is associated with the corresponding scene unit. A mapping table is created to record the mapping relationship between the multimodal feature vector and the scene unit.

[0052]

[0053] in, is the mapping relationship between the multimodal feature vectors and scene units stored in the database, is the multimodal feature vector, is the unique identifier of the multimodal feature vector, is the corresponding scene unit, is the timestamp, is the mapping function.

[0054] S108. Perform semantic understanding on the natural language query input by the user and convert it into a corresponding query feature vector.

[0055] Specifically, the BERT language model is used to semantically understand the user's natural language query and obtain the query feature vector ,

[0056]

[0057] in, is the query feature vector, () is the feedforward neural network function, is the excessive self-attention mechanism function, For the BERT model The input of the layer is The output of the layer, is the input of the first layer, For users’ natural language queries, () is the word embedding function, , , is the parameter matrix of the lth layer.

[0058] S109. Calculate the similarity between the query feature vector and the multimodal feature vector in the database, sort the scene units according to the similarity, and return the most similar scene unit.

[0059] Specifically include:

[0060]

[0061] in, is the query feature vector, is the first feature vector of the query values, For storage multimodal feature vectors, For the The first multimodal feature vector values, is the number of multimodal feature vectors, is a sorting function that sorts according to the similarity value. ) is a mapping function, which maps the corresponding multimodal feature vector returned by the sorting function to find the mapping table information stored in the corresponding database.

[0062] S110. The retrieved scene units are sorted according to similarity and presented to the user, including displaying a preview image of the video clip and timestamp information, so as to facilitate quick browsing and selection by the user.

[0063] Embodiment 2 of the present invention relates to a system for searching based on a multimodal model, the system using the method for searching based on a multimodal model described in Embodiment 1, the system comprising: A video segmentation unit is used to segment the video, and a set of frames in each frame of the video that are coherent and have a similarity higher than a threshold is taken as a scene unit; A key frame extraction unit is used to extract key frames from scene units, and extract key frames from each scene unit. The key frame is a frame with the highest similarity to other frames in the scene unit. A key frame image feature extraction unit, used for extracting image features from key frames; The key frame text feature extraction unit is used to extract text features from the key frame. If the scene unit to which the key frame belongs includes subtitles or dialogues, text recognition / speech recognition is used to convert them into text, and then the text is converted into a text feature vector through a language model. If the scene unit to which the key frame belongs does not have subtitles or dialogues, the text feature vector of the key frame is generated through the multimodal model Qwen2-VL. A fusion unit is used to fuse key frame image features and text vector features to obtain multimodal features reflecting the content of scene units; A retrieval unit, used to embed and store the extracted multimodal features into a vector database that supports fast vector similarity search, so as to quickly retrieve video feature vectors similar to the query vector; A mapping unit, used to establish a mapping relationship between multimodal feature vectors and scene units in a database, so as to quickly locate the corresponding video clip after retrieving similar feature vectors; A query unit, used to understand the semantics of the natural language query input by the user and convert it into a corresponding query feature vector; A search unit is used to calculate the similarity between the query feature vector and the multimodal feature vector in the database, sort the scene units according to the similarity, and return the most similar scene unit; The display unit is used to sort the retrieved scene units according to similarity and present them to the user, including displaying a preview image of the video clip and timestamp information.

[0064] Embodiment 3 of the present invention relates to a computer program product, which includes a computer program. The computer program is executed by a processor and is used to execute an electronically controlled differential lock control and method for autonomously escaping an unmanned vehicle according to embodiment 2.

[0065] Embodiment 4 of the present invention relates to a computer-readable storage medium, which is used to store a computer program. The computer program is executed by a processor to execute an electronically controlled differential lock control method for autonomously escaping an unmanned vehicle according to embodiment 2.

[0066] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A method for searching based on a multimodal model, characterized in that: The method comprises the following steps: S101. Segment the video, and take a set of frames in each frame of the video that are coherent and have a similarity higher than a threshold as a scene unit; S102. Extract key frames from the scene units, extract key frames from each scene unit, and the key frame is the frame with the highest similarity to other frames in the scene unit; S103. Extract image features from key frames; S104. Extract text features from the key frame. If the scene unit to which the key frame belongs includes subtitles or dialogues, convert it into text using text recognition / speech recognition, and then convert the text into a text feature vector using a language model. If the scene unit to which the key frame belongs does not have subtitles or dialogues, generate a text feature vector for the key frame using the multimodal model Qwen2-VL. S105. Fusing the key frame image features and the text vector features to obtain multimodal features reflecting the content of the scene unit; S106. Embedding and storing the extracted multimodal features into a vector database supporting fast vector similarity search, so as to quickly retrieve video feature vectors similar to the query vector; S107. Establishing a mapping relationship between multimodal feature vectors and scene units in the database, so as to quickly locate the corresponding video clip after retrieving similar feature vectors; S108. Perform semantic understanding on the natural language query input by the user and convert it into a corresponding query feature vector; S109. Calculate the similarity between the query feature vector and the multimodal feature vector in the database, sort the scene units according to the similarity, and return the most similar scene unit; S110. Sort the retrieved scene units by similarity and present them to the user, including displaying a preview image of the video clip and timestamp information.

2. The method for searching based on a multimodal model according to claim 1, characterized in that: Step S101 specifically includes: Use the SSIM algorithm to analyze the video frame by frame and calculate the similarity between frames. When the similarity is greater than the threshold T, these frames are divided into a scene unit. , ; in, Two images representing adjacent frames With image The similarity between For images The pixel mean, For images The pixel mean, For images The pixel standard deviation is For images The pixel standard deviation is For images and images The covariance of , is a constant, T is a threshold constant, A scene unit is obtained.

3. The method for searching based on a multimodal model according to claim 1, characterized in that: Step S102 specifically includes: For each image in the scene unit, calculate the cumulative difference between it and all other images, and select the image with the smallest cumulative difference as the key frame , ; Among them, the scene unit The number of images contained in , denoted as , … ,but For the scene unit Frame image, For the Frame image in The pixel value at position, and are the width and height of the image, is the image index with the smallest cumulative difference, The extracted key frames.

4. The method for searching based on a multimodal model according to claim 1, characterized in that: Step S103 specifically includes: Extract image feature vectors based on ResNet network. ; ; ; in, is the extracted image feature vector, is the input key frame image, is the activation function, is the normalization function, , , is the weight matrix parameter, is the bias parameter, is the number of residual blocks, is the initial input, For the The residual block outputs a local feature map, For the The residual block outputs a local feature map, For the The residual block outputs a local feature map exist The pixel value at position, , For the The residual block outputs a local feature map The height and width of To obtain the eigenvector function, the obtained eigenimage is converted into a eigenvector.

5. The method for searching based on a multimodal model according to claim 1, characterized in that: Step S105 specifically includes: Get multimodal features , ; in, , is the weight matrix parameter, is the key frame image feature vector, is the text feature vector, is the softmax function, It is a multimodal feature.

6. The method for searching based on a multimodal model according to claim 1, characterized in that: Step S106 specifically includes: The multimodal feature vectors are embedded and stored in the Milvus database. ; in, is the multimodal feature vector stored in the milvus database, is the multimodal feature vector stored in the embedding, is its corresponding unique identifier, It is a database operation function that implements the operation of storing feature vectors in the database; Step S107 specifically includes: A unique identifier is set for each embedded and stored multimodal feature vector in the database, and it is associated with the corresponding scene unit. A mapping table is created to record the mapping relationship between the multimodal feature vector and the scene unit. ; in, is the mapping relationship between the multimodal feature vector and the scene unit stored in the database, is the multimodal feature vector, is the unique identifier of the multimodal feature vector, is the corresponding scene unit, is the timestamp, is the mapping function.

7. The method for searching based on a multimodal model according to claim 1, characterized in that: Step S108 specifically includes: Use the BERT language model to semantically understand the user's natural language query and obtain the query feature vector , ; in, is the query feature vector, () is the feedforward neural network function, is the excessive self-attention mechanism function, For the BERT model The input of the layer is The output of the layer, is the input of the first layer, For users’ natural language queries, () is the word embedding function, , , is the parameter matrix of the lth layer.

8. A system for searching based on a multimodal model, the system using a method for searching based on a multimodal model as claimed in any one of claims 1 to 7, characterized in that: The system comprises: A video segmentation unit is used to segment the video, and a set of frames in each frame of the video that are coherent and have a similarity higher than a threshold is taken as a scene unit; A key frame extraction unit is used to extract key frames from scene units, and extract key frames from each scene unit. The key frame is a frame with the highest similarity to other frames in the scene unit. A key frame image feature extraction unit, used for extracting image features from key frames; The key frame text feature extraction unit is used to extract text features from the key frame. If the scene unit to which the key frame belongs includes subtitles or dialogues, text recognition / speech recognition is used to convert them into text, and then the text is converted into a text feature vector through a language model. If the scene unit to which the key frame belongs does not have subtitles or dialogues, the text feature vector of the key frame is generated through the multimodal model Qwen2-VL. A fusion unit is used to fuse key frame image features and text vector features to obtain multimodal features reflecting the content of scene units; A retrieval unit, used to embed and store the extracted multimodal features into a vector database that supports fast vector similarity search, so as to quickly retrieve video feature vectors similar to the query vector; A mapping unit, used to establish a mapping relationship between multimodal feature vectors and scene units in a database, so as to quickly locate the corresponding video clip after retrieving similar feature vectors; A query unit, used to understand the semantics of the natural language query input by the user and convert it into a corresponding query feature vector; A search unit is used to calculate the similarity between the query feature vector and the multimodal feature vector in the database, sort the scene units according to the similarity, and return the most similar scene unit; The display unit is used to sort the retrieved scene units according to similarity and present them to the user, including displaying a preview image of the video clip and timestamp information.

9. A computer program product, characterized in that The computer program product comprises a computer program, which is executed by a processor and is used to execute a method for searching based on a multimodal model as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store a computer program, and the computer program is executed by a processor to execute a method for searching based on a multimodal model as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Image search video method and device based on scene dictionary tree and computer readable storage medium

    CN110427517A

  • Text video cross-modal retrieval method and device based on fine granularity perception

    CN116166843A

  • Video cross-modal search model training method, search method and device

    CN116955699A

  • Video retrieval method

    CN117251598A

  • Video similarity judgment method and device based on multi-modal feature fusion

    CN118411536A

Cited By

  • Multi-modal data retrieval method based on semantic association

    CN120849655A

  • Video retrieval method and server based on semantic embedding and video memory coding

    CN121524397A