CNN (Convolutional Neural Network)-based video vectorization retrieval method, system, equipment and medium

Through the CNN-based video vector retrieval method, the features of video frames are extracted and converted into vector representations, which solves the problems of low video retrieval efficiency and low accuracy, and achieves efficient and accurate video retrieval.

CN120011594APending Publication Date: 2025-05-16SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510157903.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the prior art, video retrieval efficiency and low accuracy are particularly facing the challenges of tight computing resources and diversified user needs when processing large-scale video data.

Method used

Using CNN-based video vector retrieval method, the CNN convolutional neural network model is constructed and trained, the video frame is deeply analyzed, and the video frame is extracted and converted into a fixed-length vector representation is formed, and the search performance is optimized through aggregation and normalization processing.

Benefits of technology

It significantly improves the speed and accuracy of video retrieval, realizes efficient content understanding and feature extraction, supports real-time or near-real-time retrieval, and is suitable for multi-field applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011594A_ABST
    Figure CN120011594A_ABST
Patent Text Reader

Abstract

The invention discloses a CNN-based video vectorization retrieval method, system and device and a medium, belongs to the technical field of artificial intelligence application, and aims to solve the technical problems of low video retrieval efficiency and low accuracy in the prior art. The technical scheme is as follows: constructing and training a CNN convolutional neural network model: constructing the CNN convolutional neural network model; inputting the channel data of each frame of picture into a CNN convolutional neural network model for training, and obtaining a trained CNN convolutional neural network model; video frame sampling and preprocessing: uniformly extracting picture frames from an original video, preprocessing each frame of picture, and obtaining the preprocessed picture frames; entity feature extraction and vectorization: performing deep analysis on each frame of image through a trained CNN convolutional neural network model, and capturing and extracting key visual features in the image; performing visual feature aggregation and normalization; and video storage and multi-mode retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence application technology, and in particular to a CNN-based video vectorization retrieval method, system, device and medium. Background Art

[0002] In the digital age of information explosion, video, as an extremely rich and intuitive information carrier, is growing at an unprecedented rate. From social media sharing, online education courses to security surveillance videos, video content has penetrated almost every corner of social life. However, compared with text and pictures, the management and retrieval of video data faces more severe challenges.

[0003] Traditional video retrieval mostly relies on metadata (such as title, tags, upload time, etc.) for keyword matching. This approach has obvious limitations: on the one hand, metadata is often manually entered, which is prone to errors or incompleteness; on the other hand, metadata is difficult to fully reflect the actual content of the video, resulting in a significant reduction in the relevance of the retrieval results. Therefore, for large-scale video libraries, how to automatically and accurately extract content features from the visual information of the video itself, and then achieve efficient retrieval, has become a key issue that needs to be solved urgently.

[0004] In recent years, with the rapid development of deep learning, especially Convolutional Neural Networks (CNN) technology, the field of computer vision has made breakthrough progress. With its powerful feature extraction capabilities, CNN has demonstrated excellent performance in image classification and object recognition, which provides new ideas for understanding and retrieving video content. By applying CNN to video frame analysis, visual patterns within the frame can be automatically captured, but directly applying this technology to the entire video sequence still faces the obstacles of high computational complexity and time consumption. Summary of the invention

[0005] The technical task of the present invention is to provide a CNN-based video vectorization retrieval method, system, device and medium to solve the problems of low efficiency and low accuracy of video retrieval in the prior art.

[0006] The technical task of the present invention is achieved in the following way: a video vectorization retrieval method based on CNN, the method is as follows:

[0007] Build and train a CNN convolutional neural network model: Build a CNN convolutional neural network model, and input the channel data of each frame of the image into the CNN convolutional neural network model for training, and obtain the trained CNN convolutional neural network model;

[0008] Video frame sampling and preprocessing: Evenly extract picture frames from the original video to reduce computational complexity while maintaining the integrity of the video content. Preprocess each picture frame to obtain the preprocessed picture frame to ensure the consistency and efficiency of subsequent processing.

[0009] Entity feature extraction and vectorization: Use the trained CNN convolutional neural network model to perform in-depth analysis on each frame of the image, capture and extract the key visual features in the image, and then convert the key visual features into a fixed-length vector representation through the trained CNN convolutional neural network model to achieve quantification of the image frame content;

[0010] Aggregation and normalization of visual features: The vector representations extracted from the uniformly time-sampled images are averaged to integrate the key video information of the entire video to form a comprehensive video feature vector. The noise effect is reduced through aggregation to enhance the expression of the main content of the video. The video feature vector is then normalized to ensure the comparability of features between different videos and further optimize the retrieval performance.

[0011] Video storage and multimodal retrieval: The feature vector of each video is stored in the database as an index, and video retrieval is performed based on the index.

[0012] Preferably, the CNN convolutional neural network model includes an input layer, a pooling layer 1, a feature extraction convolution layer, a pooling layer 2, a fully connected layer, and an output layer. The input layer inputs the scaled picture frame channel data, and the output layer outputs a vector. Each dimension of the vector indicates whether the picture frame contains a specified object:

[0013] If the specified object is included, output 1;

[0014] If the specified object is not contained, output 0.

[0015] Preferably, the video frame sampling and preprocessing are as follows:

[0016] Decode the video file into frames;

[0017] Uniformly extract picture frames according to the configured frame interval;

[0018] Scaling, gray-scaling or color-standardizing the image frame to obtain a pre-processed image frame;

[0019] The preprocessed image frames are input into the trained CNN convolutional neural network model for feature extraction.

[0020] More preferably, the video feature vector is specifically: each dimension of the image feature vector indicates whether it corresponds to a type of item. After summing, each dimension represents the number of occurrences of the corresponding type of item in the corresponding dimension; after normalization, the final feature vector represents the distribution vector of the number of times all types of items appear in the corresponding video; and then the video feature vector is expressed as: video feature vector = normalization (sum of uniformly extracted frame image feature vectors).

[0021] Preferably, video storage and multimodal retrieval are as follows:

[0022] When a user initiates a search request, a corresponding query vector is generated based on the subject description of the input text, image, and video clip;

[0023] Through the efficient cosine similarity or Euclidean distance similarity measurement algorithm, the video feature vectors in the database are quickly matched to find the video content closest to the query, thus achieving accurate and efficient video retrieval.

[0024] Preferably, the text query to obtain the video list is specifically as follows: construct the query item into a query vector, obtain the dimension corresponding to the corresponding item, and perform normalization processing to obtain the normalized query vector; then use the vector cosine distance to represent the similarity, query the most similar vector from the vector database, and obtain the similar video list; wherein the storage format of the vector database is: key: video feature vector; value: video;

[0025] The specific steps of querying similar videos from images are as follows: converting the image into a feature vector through a trained CNN convolutional neural network model; then searching the vector database for all videos that are similar to the feature vector of the corresponding image;

[0026] The specific steps of video query similar video list are as follows: the videos are uniformly sampled and trained through CNN convolutional neural network model to obtain the feature vector of the corresponding video, and then all videos similar to the feature vector of the corresponding video are retrieved from the vector database.

[0027] A video vectorization retrieval system based on CNN, the system comprising:

[0028] The model building and training module is used to build a CNN convolutional neural network model, and input the channel data of each frame of the picture into the CNN convolutional neural network model for training, and obtain the trained CNN convolutional neural network model;

[0029] The video feature extraction module is used to receive the video data stream, periodically sample the video stream into video frames, and extract key visual features through the CNN convolutional neural network model;

[0030] The video feature fusion processing module is used to integrate key visual features in the time dimension to generate a comprehensive video feature vector, which can summarize the overall content of the video;

[0031] A query processing module, used to receive a query instruction input by a user and convert it into a search request; wherein the query instruction includes a keyword, an image sample or other multimedia information;

[0032] The matching and sorting module is used to apply to search requests, efficiently compare with video indexes, and sort the search results according to relevance;

[0033] The output module is used to display the sorted video retrieval results to users, achieving high-efficiency and high-accuracy video retrieval under limited computing resources.

[0034] Preferably, the video feature vector is specifically as follows: each dimension of the image feature vector indicates whether it corresponds to a type of item. After summing, each dimension represents the number of occurrences of the corresponding type of item in the corresponding dimension; after normalization, the final feature vector represents the distribution vector of the number of times all types of items appear in the corresponding video; and the video feature vector is further expressed as follows: video feature vector = normalization (sum of uniformly extracted frame image feature vectors).

[0035] An electronic device comprising: a memory and at least one processor;

[0036] Wherein, the memory stores computer-executable instructions;

[0037] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the CNN-based video vectorization retrieval method as described above.

[0038] A computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the above-mentioned CNN-based video vectorization retrieval method is implemented.

[0039] The CNN-based video vectorization retrieval method, system, device and medium of the present invention have the following advantages:

[0040] (1) The present invention combines convolutional neural network technology to achieve effective structured expression and efficient retrieval of video content. Through CNN deep analysis of uniformly sampled video frames, high-dimensional visual features are extracted and converted into vectors, and then averaged and normalized to form feature vectors that can represent the theme of the video as retrieval indexes, which significantly improves retrieval speed and accuracy, optimizes video management and analysis processes, and is suitable for applications in multiple fields.

[0041] (ii) The method of the present invention that can effectively extract key video information and meet the real-time or near real-time retrieval requirements is particularly important. It is necessary not only to be able to deeply explore the intrinsic features of the video frame, but also to design a reasonable strategy to integrate these features to form a highly generalized description of the entire video content, that is, a video vector representation, so as to support a fast and accurate video retrieval system; the present invention came into being in this context, and aims to solve this core problem in the field of video retrieval through an innovative CNN application solution;

[0042] (III) The main purpose of the present invention is to overcome the problems of low efficiency and low accuracy of video retrieval in the prior art, especially the challenges of tight computing resources and diversified user needs when processing large-scale video data; specifically, the present invention aims to achieve the following goals:

[0043] ① Efficient content understanding and feature extraction: Develop an efficient algorithm that uses a deep learning model to automatically extract key visual features from video frames;

[0044] ② Comprehensive vectorization of video content: Design a method that can effectively fuse the features extracted from each video frame to generate one or a series of compact and representative video vectors, which can not only capture the detailed information of a single frame, but also span the time dimension, comprehensively consider the overall context and dynamic changes of the video, and form a high-level abstraction of the video content;

[0045] ③ Real-time or near-real-time retrieval capability: Build a video indexing system based on the above vectorization to achieve fast retrieval of the video database; through the vectorized indexing structure, ensure that users can obtain video content that highly matches the query conditions almost instantly, improve user experience and adapt to real-time information needs in the big data environment;

[0046] In summary, the present invention is committed to promoting the progress of video retrieval technology, making it more intelligent and efficient, meeting the growing demand for multimedia information processing and analysis, and promoting the rapid circulation and utilization of information;

[0047] (IV) The present invention converts unstructured video data into structured feature vectors by introducing the powerful visual understanding ability of CNN, thereby greatly improving the efficiency and accuracy of video retrieval; users can quickly find relevant video content based on specific topics, which not only improves the user experience, but also provides strong technical support for large-scale video content management and analysis; in addition, the present invention has good scalability and flexibility, and can be adapted to a variety of application scenarios, including but not limited to video sharing platforms, surveillance video analysis, educational content retrieval and other fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The present invention is further described below in conjunction with the accompanying drawings.

[0049] Attached Figure 1 Schematic diagram for building and training CNN convolutional neural network model;

[0050] Attached Figure 2 It is a schematic diagram of video frame sampling and preprocessing;

[0051] Attached Figure 3 It is a schematic diagram of entity feature extraction and vectorization;

[0052] Attached Figure 4 It is a schematic diagram of obtaining a video feature vector by summarizing the feature vectors of picture frames;

[0053] Attached Figure 5 Flow chart of getting video list for text query;

[0054] Attached Figure 6 Schematic diagram of the process of searching for similar videos for an image;

[0055] Attached Figure 7 A flowchart of querying a list of similar videos for a video;

[0056] Attached Figure 8 A flowchart of storing the processed video into a vector database;

[0057] Attached Fig. 9 A flowchart of multimodal retrieval. DETAILED DESCRIPTION

[0058] The CNN-based video vectorization retrieval method, system, device and medium of the present invention are described in detail below with reference to the accompanying drawings and specific embodiments of the specification.

[0059] Embodiment 1:

[0060] As attached Figure 8 and 9 As shown, this embodiment provides a video vectorization retrieval method based on CNN, and the method is specifically as follows:

[0061] S1. Build and train a CNN convolutional neural network model: Build a CNN convolutional neural network model, and input the channel data of each frame of the image into the CNN convolutional neural network model for training, and obtain a trained CNN convolutional neural network model;

[0062] S2, video frame sampling and preprocessing: evenly extract picture frames from the original video to reduce computational complexity while maintaining the integrity of the video content, preprocess each picture frame, obtain the preprocessed picture frame, and ensure the consistency and efficiency of subsequent processing;

[0063] S3. Entity feature extraction and vectorization: as shown in the attached Figure 3As shown in the figure, each frame of the picture is deeply analyzed through the trained CNN convolutional neural network model to capture and extract the key visual features in the image, and then the key visual features are converted into a vector representation of a fixed length through the trained CNN convolutional neural network model to achieve quantification of the picture frame content;

[0064] S4. Visual feature aggregation and normalization: As shown in the attached Figure 4 As shown in the figure, the vector representation extracted from the uniformly time-sampled pictures is averaged, and the key video information of the entire video is integrated to form a comprehensive video feature vector. The noise effect is reduced through aggregation, and the expression of the main content of the video is enhanced. The video feature vector is then normalized to ensure the comparability of features between different videos and further optimize the retrieval performance.

[0065] S5. Video storage and multimodal retrieval: The feature vector of each video is stored in the database as an index, and video retrieval is performed based on the index.

[0066] As attached Figure 1 As shown, the CNN convolutional neural network model in step S1 of this embodiment includes an input layer, a pooling layer 1, a feature extraction convolution layer, a pooling layer 2, a fully connected layer and an output layer. The input layer inputs the scaled picture frame channel data, and the output layer outputs a vector. Each dimension of the vector indicates whether the picture frame contains a specified object:

[0067] If the specified object is included, output 1;

[0068] If the specified object is not contained, output 0.

[0069] As attached Figure 2 As shown, the video frame sampling and preprocessing in step S2 of this embodiment are specifically as follows:

[0070] S201, decoding and dividing the video file into frames;

[0071] S202, uniformly extracting picture frames according to the configured frame interval;

[0072] S203, scaling, graying or color standardizing the picture frame to obtain a pre-processed picture frame;

[0073] S204: input the preprocessed image frame into the trained CNN convolutional neural network model for feature extraction.

[0074] The video feature vector in step S4 of this embodiment is specifically: each dimension of the image feature vector indicates whether it corresponds to a type of item. After summing, each dimension represents the number of occurrences of the corresponding type of item in the corresponding dimension; after normalization, the final feature vector represents the distribution vector of the number of times all types of items appear in the corresponding video; and then the video feature vector is expressed as: video feature vector = normalization (sum of uniformly extracted frame image feature vectors).

[0075] For example, if item A appears in every frame of the video, the value of the final feature vector in the dimension corresponding to A is larger than the values ​​in other dimensions. If item A does not appear in another video, the dimension corresponding to A is approximately equal to 0. The distance between the feature vectors of two videos calculated by Euclidean distance or cosine distance will also be relatively far, so it can be judged that the information difference between the two videos (from the dimension corresponding to A) is relatively large. Therefore, the feature vector of the video represented in this way has the ability to summarize and express video information.

[0076] The video storage and multimodal retrieval in step S5 of this embodiment are specifically as follows:

[0077] S501, when a user initiates a search request, a corresponding query vector is generated according to the subject description of the input text, image and video clip;

[0078] S502: quickly match the video feature vectors in the database through an efficient cosine similarity or Euclidean distance similarity measurement algorithm, so as to find the video content closest to the query and achieve accurate and efficient video retrieval.

[0079] As attached Figure 5 As shown in the figure, the specific steps of obtaining a video list through a text query are as follows: constructing the query item into a query vector, obtaining the dimension corresponding to the corresponding item, and performing normalization processing to obtain a normalized query vector; then expressing the similarity through the vector cosine distance, querying the most similar vector from the vector database, and obtaining a similar video list; wherein, the storage format of the vector database is: key: video feature vector; value: video.

[0080] As attached Figure 6 As shown in FIG. 1 , the specific steps of image query similar videos are as follows: the image is converted into a feature vector of the image through a trained CNN convolutional neural network model; and then all videos similar to the feature vector of the corresponding image are retrieved from the vector database.

[0081] As attached Figure 7 As shown in FIG. 1 , the video query similar video list is specifically as follows: the videos are uniformly sampled and trained through the CNN convolutional neural network model to obtain the feature vector of the corresponding video, and then the vector database is retrieved for all videos that are similar to the feature vector of the corresponding video.

[0082] Embodiment 2:

[0083] This embodiment provides a CNN-based video vectorization retrieval system, which includes:

[0084] The model building and training module is used to build a CNN convolutional neural network model, and input the channel data of each frame of the picture into the CNN convolutional neural network model for training, and obtain the trained CNN convolutional neural network model;

[0085] The video feature extraction module is used to receive the video data stream, periodically sample the video stream into video frames, and extract key visual features through the CNN convolutional neural network model;

[0086] The video feature fusion processing module is used to integrate key visual features in the time dimension to generate a comprehensive video feature vector, which can summarize the overall content of the video;

[0087] A query processing module, used to receive a query instruction input by a user and convert it into a search request; wherein the query instruction includes a keyword, an image sample or other multimedia information;

[0088] The matching and sorting module is used to apply to search requests, efficiently compare with video indexes, and sort the search results according to relevance;

[0089] The output module is used to display the sorted video retrieval results to users, achieving high-efficiency and high-accuracy video retrieval under limited computing resources.

[0090] The video feature vector in this embodiment is specifically: each dimension of the image feature vector indicates whether it corresponds to a type of item. After summing, each dimension represents the number of occurrences of the corresponding type of item in the corresponding dimension; after normalization, the final feature vector represents the distribution vector of the number of times all types of items appear in the corresponding video; and then the video feature vector is expressed as: video feature vector = normalization (sum of uniformly extracted frame image feature vectors).

[0091] Embodiment 3:

[0092] This embodiment also provides an electronic device, including: a memory and at least one processor;

[0093] Wherein, the memory stores computer-executable instructions;

[0094] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the CNN-based video vectorization retrieval method described in any one of the present inventions.

[0095] The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor may be a microprocessor or any conventional processor, etc.

[0096] The memory can be used to store computer programs and / or modules. The processor realizes various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory can also include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash memory card, at least one disk storage period, a flash memory device, or other volatile solid-state storage devices.

[0097] Embodiment 4:

[0098] This embodiment also provides a computer-readable storage medium, in which a plurality of instructions are stored, and the instructions are loaded by a processor, so that the processor executes the CNN-based video vectorization retrieval method in any embodiment of the present invention. Specifically, a system or device equipped with a storage medium can be provided, on which a software program code that implements the functions of any of the above embodiments is stored, and a computer (or CPU or MPU) of the system or device reads and executes the program code stored in the storage medium.

[0099] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute a part of the present invention.

[0100] The storage medium embodiments for providing the program code include a floppy disk, a hard disk, a magneto-optical disk, an optical disk (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), a magnetic tape, a non-volatile memory card, and a ROM. Alternatively, the program code can be downloaded from a server computer by a communication network.

[0101] In addition, it should be clear that the functions of any of the above embodiments can be implemented not only by executing the program code read by the computer, but also by enabling an operating system operating on the computer to complete part or all of the actual operations based on instructions from the program code.

[0102] In addition, it can be understood that the program code read from the storage medium is written to a memory provided in an expansion board inserted into the computer or written to a memory provided in an expansion unit connected to the computer, and then based on the instructions of the program code, a CPU installed on the expansion board or the expansion unit is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above-mentioned embodiments.

[0103] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A video vectorization retrieval method based on CNN, characterized in that: The method is as follows: Build and train a CNN convolutional neural network model: Build a CNN convolutional neural network model, and input the channel data of each frame of the image into the CNN convolutional neural network model for training, and obtain the trained CNN convolutional neural network model; Video frame sampling and preprocessing: evenly extract picture frames from the original video, preprocess each frame, and obtain the preprocessed picture frames; Entity feature extraction and vectorization: Use the trained CNN convolutional neural network model to perform in-depth analysis on each frame of the image, capture and extract the key visual features in the image, and then convert the key visual features into a fixed-length vector representation through the trained CNN convolutional neural network model to achieve quantification of the image frame content; Visual feature aggregation and normalization: The vector representations extracted from the uniformly time-sampled images are averaged to integrate the key video information of the entire video to form a comprehensive video feature vector. The video feature vector is then normalized to ensure feature comparability between different videos and further optimize retrieval performance. Video storage and multimodal retrieval: The feature vector of each video is stored in the database as an index, and video retrieval is performed based on the index.

2. The CNN-based video vectorization retrieval method according to claim 1, characterized in that: The CNN convolutional neural network model includes an input layer, a pooling layer 1, a feature extraction convolution layer, a pooling layer 2, a fully connected layer, and an output layer. The input layer inputs the scaled image frame channel data, and the output layer outputs a vector. Each dimension of the vector indicates whether the image frame contains a specified object: If the specified object is included, output 1; If the specified object is not contained, output 0.

3. The CNN-based video vectorization retrieval method according to claim 1 or 2, characterized in that: Video frame sampling and preprocessing are as follows: Decode the video file into frames; Uniformly extract picture frames according to the configured frame interval; Scaling, gray-scaling or color-standardizing the image frame to obtain a pre-processed image frame; The preprocessed image frames are input into the trained CNN convolutional neural network model for feature extraction.

4. The CNN-based video vectorization retrieval method according to claim 3, characterized in that: The video feature vector is specifically: each dimension of the image feature vector indicates whether it corresponds to a type of item. After summing, each dimension represents the number of occurrences of the corresponding type of item in the corresponding dimension; after normalization, the final feature vector represents the distribution vector of the number of times all types of items appear in the corresponding video; and then the video feature vector is expressed as: video feature vector = normalization (sum of uniformly extracted frame image feature vectors).

5. The CNN-based video vectorization retrieval method according to claim 4, characterized in that: Video storage and multimodal retrieval are as follows: When a user initiates a search request, a corresponding query vector is generated based on the subject description of the input text, image, and video clip; Through the efficient cosine similarity or Euclidean distance similarity measurement algorithm, the video feature vectors in the database are quickly matched to find the video content closest to the query, thus achieving accurate and efficient video retrieval.

6. The CNN-based video vectorization retrieval method according to claim 5, characterized in that: The specific steps of obtaining a video list by text query are as follows: construct the query item into a query vector, obtain the dimension corresponding to the corresponding item, and perform normalization to obtain the normalized query vector; then use the vector cosine distance to represent the similarity, query the most similar vector from the vector database, and obtain a similar video list; the storage format of the vector database is: key: video feature vector; value: video; The specific steps of querying similar videos from images are as follows: converting the image into a feature vector through a trained CNN convolutional neural network model; then searching the vector database for all videos that are similar to the feature vector of the corresponding image; The specific steps of video query similar video list are as follows: the videos are uniformly sampled and trained through CNN convolutional neural network model to obtain the feature vector of the corresponding video, and then all videos similar to the feature vector of the corresponding video are retrieved from the vector database.

7. A video vectorization retrieval system based on CNN, characterized in that: The system includes: The model building and training module is used to build a CNN convolutional neural network model, and input the channel data of each frame of the picture into the CNN convolutional neural network model for training, and obtain the trained CNN convolutional neural network model; The video feature extraction module is used to receive the video data stream, periodically sample the video stream into video frames, and extract key visual features through the CNN convolutional neural network model; The video feature fusion processing module is used to integrate key visual features in the time dimension to generate a comprehensive video feature vector, which can summarize the overall content of the video; A query processing module, used to receive a query instruction input by a user and convert it into a search request; wherein the query instruction includes a keyword, an image sample or other multimedia information; The matching and sorting module is used to apply to search requests, efficiently compare with video indexes, and sort the search results according to relevance; The output module is used to display the sorted video retrieval results to users, achieving high-efficiency and high-accuracy video retrieval under limited computing resources.

8. The CNN-based video vectorization retrieval system according to claim 7, characterized in that: The video feature vector is specifically: each dimension of the image feature vector indicates whether it corresponds to a type of item. After summing, each dimension represents the number of occurrences of the corresponding type of item in the corresponding dimension; after normalization, the final feature vector represents the distribution vector of the number of times all types of items appear in the corresponding video; and then the video feature vector is expressed as: video feature vector = normalization (sum of uniformly extracted frame image feature vectors).

9. An electronic device, characterized in that: include: memory and at least one processor; Wherein, the memory stores computer-executable instructions; The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the CNN-based video vectorization retrieval method as described in any one of claims 1 to 5.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the processor executes the computer-executable instructions, the CNN-based video vectorization retrieval method as described in any one of claims 1 to 5 is implemented.

Citation Information

Cited By

  • Method and device for retrieving video and electronic equipment

    CN121166973A