Video retrieval method based on video image vectorization
Through video image vectorization technology, keyframes in video are extracted and processed, and vector databases are established for rapid search, which solves the problem of inefficiency of traditional video retrieval methods and achieves efficient and accurate video retrieval.
Patent Information
- Application Number
- CN202510139759.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-06-06
AI Technical Summary
Traditional video retrieval methods are inefficient and rely on manual annotation and metadata, which have problems such as inefficiency and strong subjectivity, and cannot quickly and accurately find key information from massive video data.
Using a video search method based on video image vectorization, the video content is converted into vector form through keyframe extraction and vectorization processing, and a vector database is established for rapid retrieval, so as to realize the rapid positioning and playback of video content.
It improves the efficiency and accuracy of video retrieval, reduces manual intervention, and can quickly locate and play target videos. It is suitable for a variety of video retrieval scenarios, reducing the possibility of labeling errors.
Smart Images

Figure CN120104832A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video retrieval, and in particular to a video retrieval method based on video image vectorization. Background Art
[0002] Video retrieval can be understood as searching for useful or necessary information from a video. The main algorithms for achieving real-time detection, recognition, classification, and multi-target tracking of mobile targets through intelligent video technology are divided into the following five categories: target detection, target tracking, target recognition, behavior analysis, content-based video retrieval, and data fusion. In the field of social public security, video surveillance systems have become an important part of maintaining social order and strengthening social management. With the continuous deepening of the "Skynet Project" and "Safe City" construction, the upgrading of video security monitoring technology and the replacement of new technologies, video retrieval technology will receive more and more attention in the future.
[0003] In the actual application of security video surveillance systems, traditional video retrieval methods are often inefficient and mainly rely on "human sea tactics". There are many limitations, such as low attention caused by human physiological limitations, missing target clues caused by human visual fatigue, and too long time to obtain effective information. These problems make it impossible for traditional methods to quickly and accurately find key information from massive video data. Therefore, the development and application of intelligent video retrieval technology has become the key to solving this problem. Intelligent video retrieval technology can greatly improve the efficiency and accuracy of video surveillance systems through automation and intelligent processing, reduce manpower input, and improve review efficiency, so as to discover and solve problems faster. In addition, intelligent video retrieval technology can also focus on more details, such as character data, event key information, etc., making the application of video surveillance systems more extensive and in-depth. Therefore, the necessity of video retrieval lies in improving the efficiency and accuracy of social public security management and adapting to the needs of modern society for efficient and intelligent monitoring systems.
[0004] At present, the more traditional video retrieval method is based on manual video annotation, specifically:
[0005] (1) Add text tags to the video
[0006] (2) Save the tags to a database or search engine
[0007] (3) When a user searches, we search for tags related to the user’s search keywords to find the video with the highest matching degree.
[0008] With the explosive growth of video content, traditional retrieval methods based on manual video annotation can no longer meet users' needs for fast and accurate retrieval of video content. Existing video retrieval technologies mostly rely on video metadata or manual annotation, which have problems such as low efficiency and strong subjectivity.
[0009] At present, there are still methods of searching by image, which can retrieve images at a certain time point in the video. Some methods use similar image vectors to search. This method does not solve the problem of similarity of image sequences over a period of time. However, the similarity of images at a certain time point is only a representation of this time point, without continuity and trend. It is not possible to make the representation of things more specific through image vectors over a period of time. Summary of the invention
[0010] The present invention provides a video retrieval method based on video image vectorization to achieve the purpose of high video retrieval efficiency, strong objectivity and accurate video retrieval.
[0011] To achieve the above object, the technical solution of the present invention is to provide a video retrieval method based on video image vectorization, which is characterized by comprising the following steps:
[0012] (1) Establishing a video library: manually collect videos to form a video library. The videos in the video library include videos collected online, surveillance videos, and videos recorded by mobile phones or cameras;
[0013] (2) Vectorizing and storing the videos in the video library: extracting key frames from all videos in the video library, vectorizing the extracted key frame images, and composing the vectorized key frames into matrices in time segments, which are stored to form a vector database;
[0014] (3) Establishing a sample video library: extracting and copying the key frames extracted in step (2), and performing the same vectorization processing on the extracted key frames as in step (2), and constructing the vectorized data of these key frames into a pre-searched sample feature video clip vector library, the extracted key frame images are sample feature videos, and all the extracted key frame images form a sample video library;
[0015] (4) Sample feature video labeling: Classify the sample feature videos in the sample video library according to the features of each video, and annotate the sample feature videos of the same category with labels corresponding to the features;
[0016] (5) Sample feature video selection: confirm the features of the video to be searched, and select sample feature videos corresponding to the features of the video to be searched in the sample video library according to the tags, and match the corresponding vectors in the sample feature video clip vector library according to the selected sample feature videos;
[0017] (6) Performing video retrieval in the video library: searching the sample feature video selected in step (5) in the video library, and retrieving video vectors similar to the sample feature video selected in step (5) in the vector database according to the vector similarity calculation method;
[0018] (7) Video positioning: Based on the search results, locate the video clips and time points similar to the sample feature video to be searched in the video library, play the video, and complete the video retrieval.
[0019] Furthermore, the specific process of extracting key frames from all videos in the video library in step (2) is as follows: using the FFmpeg tool to automatically extract key frames from the videos, and recording the time corresponding to the extracted key frames.
[0020] Furthermore, the specific process of vectorizing the extracted key frame image in the step (2) is: converting the extracted key frame into a vector form through image embedding technology, and using a pre-trained deep learning model to generate an embedding vector, that is, using a deep learning model to extract features of the image content to obtain a vector that can express the image features.
[0021] Furthermore, the deep learning model is ResNet or VGG.
[0022] Furthermore, the extracted key frames are converted into vector form through image embedding technology. The specific example is as follows:
[0023]
[0024] Among them, each line on the left side of the formula is a vector corresponding to a key frame, and multiple key frame vectors are combined into the matrix shown on the left. The matrix on the left is flattened into the vector form on the right using NumPy, which is to vectorize the key frame image.
[0025] Furthermore, in the step (2), the vectorized key frames are organized into a matrix with time segments, and the specific process for storage is: the vectors of multiple key frames of the same video are organized according to the timeline of the video, the image vectors of multiple time series are organized into a matrix, the matrix is flattened into a long vector for storage, and all the long vectors formed by the vectorization of the videos in the video library form a vector database.
[0026] Furthermore, the sample feature videos selected in step (5) are one or more sample feature videos with the same features.
[0027] Furthermore, the vector similarity calculation method in step (6) includes a cosine similarity and a Euclidean distance calculation method.
[0028] Furthermore, the time length of the sample feature videos in the sample video library is 3-5 seconds, and the length range of the videos in the video library is greater than 5 minutes.
[0029] Compared with the prior art, the present invention has the following beneficial effects:
[0030] (1) The video retrieval method based on video image vectorization of the present invention is more efficient in video retrieval: it can automatically extract key frames in the video and perform vectorization processing, reducing human intervention, and through rapid retrieval of the vector database, it can achieve rapid positioning and playback of video content, with fast retrieval response and improved retrieval efficiency.
[0031] (2) The video retrieval method based on video image vectorization of the present invention can retrieve videos accurately: using image vectorization to search videos can search for more abstract images and video parts. At the same time, the video clips used for searching also come from the video library, so the video to be searched can be found more accurately.
[0032] (3) The video retrieval method based on video image vectorization of the present invention can retrieve videos flexibly: it is applicable to a variety of video retrieval scenarios, such as fire fighting scenarios, wind power generation scenarios, etc., and can be customized according to different usage requirements.
[0033] (4) Traditionally, searching for videos using text tags requires labeling and describing each video before performing a video search. However, the video retrieval method based on video image vectorization of the present invention does not require labeling each video one by one. Instead, it only requires classifying and labeling sample feature videos of the same type with the same features, thereby reducing the workload and the possibility of labeling errors or inaccurate descriptions affecting the search accuracy.
[0034] (5) The video retrieval method based on video image vectorization of the present invention can retrieve similar continuous image features and can match and search for similar video content over a period of time. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0036] Figure 1 The present invention is a flowchart of a video retrieval method based on video image vectorization.
[0037] Figure 2It is a key frame of flame image features contained in a video library in an implementation of a video retrieval method based on video image vectorization of the present invention.
[0038] Figure 3 It is a key frame of flame image features contained in a sample video library in an implementation of a video retrieval method based on video image vectorization of the present invention.
[0039] Figure 4 It is a key frame of the flame image feature that is not included in the sample video library in an implementation of a video retrieval method based on video image vectorization of the present invention.
[0040] Figure 5-10 For the present invention Figure 2 The vector data of the video where the image in is located.
[0041] Figure 11-16 For the present invention Figure 3 The vector data of the video where the image in is located.
[0042] Figure 17-22 For the present invention Figure 4 The vector data of the video where the image in is located. DETAILED DESCRIPTION
[0043] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0044] This embodiment 1
[0045] This embodiment provides a video retrieval method based on video image vectorization, and its specific process is as follows: Figure 1 As shown, the specific steps include:
[0046] (1) Establishing a video library: manually collect videos to form a video library. The videos in the video library include videos collected online, surveillance videos, and videos recorded by mobile phones or cameras;
[0047] (2) Vectorizing and storing the videos in the video library: extracting key frames from all videos in the video library, vectorizing the extracted key frame images, and composing the vectorized key frames into matrices in time segments, which are stored to form a vector database;
[0048] (3) Establishing a sample video library: extracting and copying the key frames extracted in step (2), and performing the same vectorization processing as in step (2) on the extracted key frames. After processing, the vectorized data of these key frames are constructed into a pre-searched sample feature video clip vector library, the extracted key frame images are sample feature videos, and all the extracted key frame images form a sample video library.
[0049] The above-mentioned method of establishing the sample video library is only one of them. The sample video library can also be collected from the Internet or shot by shooting tools.
[0050] (4) Sample feature video labeling: Classify the sample feature videos in the sample video library according to the features of each video, and annotate the sample feature videos of the same category with labels corresponding to the features;
[0051] (5) Sample feature video selection: confirm the features of the video to be searched, and select sample feature videos corresponding to the features of the video to be searched in the sample video library according to the tags, and match the corresponding vectors in the sample feature video clip vector library according to the selected sample feature videos; the selected sample feature videos are one or more sample feature videos with the same features. When in use, one or more sample feature videos can be searched one by one, so that the target video found will be more accurate.
[0052] (6) Performing video retrieval in the video library: searching the sample feature video selected in step (5) in the video library, and retrieving video vectors similar to the sample feature video selected in step (5) in the vector database according to the vector similarity calculation method;
[0053] (7) Video positioning: Based on the search results, locate the video clips and time points similar to the sample feature video to be searched in the video library, play the video, and complete the video retrieval.
[0054] Example 2
[0055] In this embodiment, based on the embodiment 1, the specific process of extracting key frames from all videos in the video library in step (2) is as follows: using the FFmpeg tool to automatically extract key frames from the videos, and recording the time corresponding to the extracted key frames. The rest is the same as the embodiment 1.
[0056] FFmpeg is an open source computer program that can be used to record and convert digital audio and video and convert them into streams. As a powerful multimedia processing tool, FFmpeg can be used to extract key frames from videos. By extracting key frames, you can achieve operations such as precise video cutting, fast preview, and key information recognition, which provides a basis for subsequent video processing and analysis.
[0057] FFmpeg supports various mainstream audio and video formats, including but not limited to MP4, FLV, AVI, MOV, MKV, etc. It is a powerful and flexible tool that is widely used in video processing, multimedia transcoding, streaming media transmission and other fields.
[0058] Example 3
[0059] In this embodiment, based on Embodiment 1, the specific process of vectorizing the extracted key frame image in step (2) is as follows: the extracted key frame is converted into a vector form through image embedding technology, and a pre-trained deep learning model is used to generate an embedding vector, that is, the deep learning model is used to extract features of the image content to obtain a vector that can express the image features.
[0060] The extracted key frames are converted into vector form through image embedding technology. The specific example is as follows:
[0061]
[0062] Each row on the left side of the formula is a vector corresponding to a key frame, and multiple key frame vectors are combined into the matrix shown on the left side. The matrix on the left side is flattened into the vector form on the right side using NumPy, that is, the key frame image is vectorized. The rest is the same as in Example 1.
[0063] Embedding technology is a method of mapping high-dimensional data to low-dimensional space. It is usually used to convert discrete, non-continuous data into continuous vector representations for computer processing. It helps to reduce the complexity of data and the demand for computing resources. Using image embedding technology to vectorize key frames improves the accuracy of video content expression.
[0064] NumPy is an open source numerical computing extension for Python. This tool can be used to store and process large matrices, supports a large number of dimensional arrays and matrix operations, and also provides a large number of mathematical function libraries for array operations.
[0065] Example 4
[0066] The deep learning model of this embodiment is ResNet or VGG. The rest is the same as in Embodiment 3.
[0067] ResNet or VGG are both existing deep learning models. Specifically:
[0068] In terms of feature extraction, ResNet (residual network) can extract rich features from the input image. These features include not only basic features such as color and shape of the image, but also more advanced semantic information. ResNet gradually extracts the features of the image through a series of convolution operations and pooling operations. These features are further processed and integrated in the subsequent layers of the network, and finally form a high-level description of the image.
[0069] The process of feature extraction using the VGG model is as follows:
[0070] Loading pre-trained VGG models: Load pre-trained VGG models in deep learning frameworks such as Keras or PyTorch. These models have been trained on a large number of images, so they can extract common visual features.
[0071] Prepare image data: Convert the image to be processed into the input format required by the VGG model, which usually includes steps such as resizing and normalizing the image.
[0072] Feature extraction: The processed image is input into the VGG model and the feature representation of the image is obtained through forward propagation.
[0073] Example 5
[0074] On the basis of Example 1, the key frames after vectorization in step (2) are organized into a matrix with time segments, and the specific process for storage is: the vectors of multiple key frames of the same video are organized according to the timeline of the video, the image vectors of multiple time series are organized into a matrix, the matrix is flattened into a long vector for storage, and all the long vectors formed by vectorizing the videos in the video library form a vector database.
[0075] The rest is the same as in Example 1.
[0076] Example 6
[0077] In this embodiment, based on the first embodiment, the vector similarity calculation method in step (6) includes the calculation method of cosine similarity and Euclidean distance. The rest is the same as the first embodiment.
[0078] Preferably, the time length of the sample feature videos in the sample video library of the present invention is 3-5 seconds, and the length range of the videos in the video library is greater than 5 minutes.
[0079] Image features represent different characteristics in different scenarios. For example, in a firefighting scenario, image features represent the source of fire or fire point or smoke; in a wind power generation scenario, image features represent cracks in wind turbine blades; or in other scenarios, image features represent features that may cause danger or risk.
[0080] In a video, there are multiple continuous still images per second, and these still images are called frames. Among these frames, key frames refer to frames with important information and changes. Key frames are clips in the video that can represent image features and usually contain important visual information. In the firefighting scene described above, key frames refer to images with fire sources, fire points or smoke, or in wind power generation scenes, key frames refer to images with cracks on wind turbine blades. In actual use, different points are selected as image features according to different scenes and actual needs, so as to extract different key frames, which can adapt to more scenes and is more flexible to use.
[0081] Based on the above technical solution, the specific implementation process of the present invention is as follows: according to steps (1) and (2), after the video library is established, the key frames of the videos in the video library are extracted, vectorized and stored, wherein the video library includes the key frames Figure 2 The video in which the image feature is flame; Figure 2 The vector data of the video as a key frame is Figure 5-10 shown.
[0082] The key frames extracted from the video library are extracted and copied, and vectorized to form a sample video library. The videos in the sample video library are labeled. For example, the images with the characteristics of flame are classified into one category and labeled as flame, and the images with the characteristics of smoke are classified into one category and labeled as smoke. The sample video library includes Figure 3 The image of the flame feature shown is a sample feature video with key frames;
[0083] The video retrieved in this implementation process is a video with image features of flames. One or more sample feature videos containing image features of flames are selected from the sample video library. In this implementation process, a sample feature video is selected, namely, a video including the following: Figure 3 Sample feature video of flame image features shown. Figure 3 The vector data of the sample feature video as a key frame is Figure 11-16 shown.
[0084] The selected sample feature video containing flame image features is searched in the video library. According to the vector similarity calculation method, the video vector similar to the selected sample feature video is retrieved in the vector database, and the video clips and time points similar to the sample feature video to be searched are located. The searched videos all contain flame image features, including those containing Figure 2 Video of the flame image features shown. And Figure 3 The video where the key frame image is located and the retrieved Figure 2 The vector similarity of the video containing the key frame image shown is 0.5348542433047829.
[0085] In addition, another sample feature video without flame image features is selected from the sample video library as a control (such as Figure 4 As shown, the included keyframes do not have the image features of flames. Figure 4 The vector data of the sample feature video as a key frame is Figure 17-22 As shown in FIG. 1 , the sample feature video without the flame image feature is also searched in the video library, and no video containing the flame image feature is found. Figure 2 The vector similarity of the video containing the key frame image shown is: 0.005269117736720452.
[0086] The similarity calculation method used in this implementation process is cosine similarity.
[0087] The embodiments described above are merely descriptions of preferred implementation modes of the present invention, and are not intended to limit the concept and scope of the present invention. Without departing from the design concept of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary engineering technicians in the field should fall within the protection scope of the present invention. The technical contents for which protection is sought in the present invention have been fully recorded in the technical requirements.
Claims
1. A video retrieval method based on video image vectorization, characterized in that: The specific steps include: (1) Establishing a video library: manually collect videos to form a video library. The videos in the video library include videos collected online, surveillance videos, and videos recorded by mobile phones or cameras; (2) Vectorizing and storing the videos in the video library: extracting key frames from all videos in the video library, vectorizing the extracted key frame images, and composing the vectorized key frames into matrices in time segments, which are stored to form a vector database; (3) Establishing a sample video library: extracting and copying the key frames extracted in step (2), and performing the same vectorization processing on the extracted key frames as in step (2), and constructing the vectorized data of these key frames into a pre-searched sample feature video clip vector library, the extracted key frame images are sample feature videos, and all the extracted key frame images form a sample video library; (4) Sample feature video labeling: Classify the sample feature videos in the sample video library according to the features of each video, and annotate the sample feature videos of the same category with labels corresponding to the features; (5) Sample feature video selection: confirm the features of the video to be searched, and select sample feature videos corresponding to the features of the video to be searched in the sample video library according to the tags, and match the corresponding vectors in the sample feature video clip vector library according to the selected sample feature videos; (6) Performing video retrieval in the video library: searching the sample feature video selected in step (5) in the video library, and retrieving video vectors similar to the sample feature video selected in step (5) in the vector database according to the vector similarity calculation method; (7) Video positioning: Based on the search results, locate the video clips and time points similar to the sample feature video to be searched in the video library, play the video, and complete the video retrieval.
2. The video retrieval method based on video image vectorization according to claim 1, characterized in that: The specific process of extracting key frames from all videos in the video library in step (2) is as follows: using the FFmpeg tool to automatically extract key frames from the videos, and recording the time corresponding to the extracted key frames.
3. The video retrieval method based on video image vectorization according to claim 1, characterized in that: The specific process of vectorizing the extracted key frame image in the step (2) is as follows: converting the extracted key frame into a vector form through image embedding technology, and using a pre-trained deep learning model to generate an embedding vector, that is, using a deep learning model to extract features of the image content to obtain a vector that can express the image features.
4. The video retrieval method based on video image vectorization according to claim 3, characterized in that: The deep learning model is ResNet or VGG.
5. The video retrieval method based on video image vectorization according to claim 3 is characterized in that: The extracted key frames are converted into vector form through image embedding technology, for example: Among them, each line on the left side of the formula is a vector corresponding to a key frame, and multiple key frame vectors are combined into the matrix shown on the left. The matrix on the left is flattened into the vector form on the right using NumPy, which is to vectorize the key frame image.
6. The video retrieval method based on video image vectorization according to claim 1, characterized in that: In the step (2), the vectorized key frames are organized into matrices with time segments, and the specific process for storage is: the vectors of multiple key frames of the same video are organized according to the timeline of the video, the image vectors of multiple time series are organized into a matrix, the matrix is flattened into a long vector for storage, and all the long vectors formed by vectorizing the videos in the video library form a vector database.
7. The video retrieval method based on video image vectorization according to claim 1, characterized in that: The sample feature videos selected in step (5) are one or more sample feature videos with the same features.
8. The video retrieval method based on video image vectorization according to claim 1, characterized in that: The vector similarity calculation method in step (6) includes the calculation method of cosine similarity and Euclidean distance.
9. The video retrieval method based on video image vectorization according to claim 1, characterized in that: The time length of the sample feature videos in the sample video library is 3-5s, and the length of the videos in the video library is greater than 5 minutes.