Video knowledge base construction and efficient retrieval method and system
Through Milvus and MinIO, the video knowledge base storage architecture is built, and combined with video preprocessing and text vectorization technology, the problem of low retrieval accuracy in video knowledge base construction is solved, efficient and accurate video retrieval is achieved, and user experience is improved.
Patent Information
- Application Number
- CN202510373898.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
During the construction of existing video knowledge bases, the search classification requires low accuracy in abstract induction, and the output of related video results is difficult to accurately, and it is difficult to directly search from the video content, resulting in too low retrieval efficiency, time-consuming and labor-intensive, and reduced user usability.
The vector database Milvus and object storage service MinIO are used to build a video knowledge base storage architecture. Through video preprocessing, audio separation and speech recognition, text information sorting and time stamp association, a video-text-time stamp correspondence relationship is formed, and the text information is converted into vector representations and stored in the Milvus vector library, combining the search results for optimization of user behavior.
An efficient and accurate video retrieval process has been achieved, retrieval efficiency and accuracy have been improved, the waste of human resources has been reduced, and the user experience has been improved.
Smart Images

Figure CN120336583A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of knowledge base construction and retrieval, and specifically to a method and system for constructing and efficiently retrieving a video knowledge base. Background Art
[0002] Currently, the construction of video knowledge bases is to pre-select the classification of videos and then perform simple storage according to the general classification. When users retrieve and use them, they often face the following problems:
[0003] 1. Retrieving classifications requires abstract induction, and the accuracy of induction is relatively low
[0004] 2. A large number of relevant video results are output, and it is difficult to accurately achieve the expected effect
[0005] 3. It is difficult to directly retrieve from video content.
[0006] The above problems will lead to too low video knowledge retrieval efficiency, which will waste time and manpower and is difficult to meet the needs of users, ultimately resulting in a decrease in user usability. Summary of the Invention
[0007] The purpose of the present invention is to provide a method and system for constructing and efficiently retrieving a video knowledge base to solve the problems raised in the above background art.
[0008] To achieve the above purpose, the present invention provides the following technical solution: A method for constructing and efficiently retrieving a video knowledge base, the method comprising the following steps:
[0009] Deploy storage service: Use the vector database Milvus and the object storage service MinIO to construct the storage architecture of the video knowledge base, where Milvus is used to store the feature vectors of video content, and MinIO is used to store video files;
[0010] Video source file processing: Receive and upload video files, preprocess the video files, including format conversion and key frame extraction; Separate the audio information in the video and convert it into text information; Organize the text information, including removing noise, word segmentation, part-of-speech tagging, and associate the text information with the time points in the video;
[0011] Video file storage: Store the processed video files in the MinIO object storage service, set up independent storage buckets for each user or video classification, and generate unique access addresses;
[0012] Vectorization and storage of text data: Convert the extracted text information into vector representation, and store the vector, metadata of the text information, and video access address in the Milvus vector library.
[0013] Preferably, the processing steps of the video source file further include:
[0014] Provide upload interfaces for multiple video formats and automatically preprocess the video files;
[0015] Use an audio separation tool to separate the audio information from the video file and convert the audio into text information through a speech recognition model;
[0016] Organize the text information to form the corresponding relationship of video - text - timestamp.
[0017] Preferably, the video knowledge base constructed by the video knowledge base construction method includes the following steps:
[0018] Receive the retrieval keyword input by the user;
[0019] Use the same text vectorization tool as when constructing the video knowledge base to vectorize the retrieval keyword input by the user;
[0020] Send the obtained vector as a query input to the Milvus vector library for similarity retrieval;
[0021] According to the similarity retrieval result returned by Milvus, obtain the video address associated with the retrieval keyword.
[0022] Preferably, the method further includes the following steps:
[0023] According to the obtained video address, obtain the corresponding video file from the MinIO object storage service and provide it to the user for viewing or downloading;
[0024] Collect the user's viewing behavior and feedback information, and adjust the vector similarity of the video library according to the information to optimize the accuracy of the retrieval result.
[0025] Preferably, the method further includes the following dynamic update and optimization steps:
[0026] When the user is watching the video, the background re - performs audio separation and text vectorization processing;
[0027] Update the newly generated text vectors and related information to the Milvus vector library to dynamically update and optimize the video knowledge base and improve the accuracy and efficiency of the retrieval.
[0028] A system for a video knowledge base construction and efficient retrieval method, the system includes:
[0029] Vector library and file storage service deployment module, which is used to integrate the Milvus vector database and the MinIO object storage service to build the storage architecture of the video knowledge base. Milvus is used to store the feature vectors of video content to achieve content-based retrieval, and MinIO is used for the secure storage and efficient access of video files;
[0030] Video source file analysis and processing module, including:
[0031] Video upload and preprocessing sub-module, which is used to receive the upload of videos in multiple formats and perform preprocessing such as format conversion and key frame extraction;
[0032] Audio separation and speech recognition sub-module, which is used to separate the video audio and convert it into text information;
[0033] Text information collation and timestamp association sub-module, which is used to collate the text information and associate it with the time points in the video;
[0034] File storage module, which is used to set up independent buckets in MinIO for each user or video classification, store video files and generate unique access addresses;
[0035] Video text data vectorization and storage module, which is used to convert text information into vector representations using text vectorization techniques, and store it together with the metadata of the text information and the video access address in the Milvus vector library.
[0036] Preferably, the video upload and preprocessing sub-module supports upload interfaces for multiple video formats and automatically performs preprocessing of video files to ensure the accuracy of subsequent analysis.
[0037] Preferably, the audio separation and speech recognition sub-module uses open-source tools such as FFMPEG to separate the video audio and converts the audio into text information through a deep learning model; the text information collation and timestamp association sub-module performs denoising, word segmentation, and part-of-speech tagging on the extracted text information and associates it with the time points in the video to form a correspondence between video-text-timestamp.
[0038] Preferably, the buckets set up by the file storage module in MinIO have appropriate access permissions to ensure the secure storage and access of video files.
[0039] Preferably, the text vectorization techniques adopted by the video text data vectorization and storage module include Word2Vec and BERT to convert text information into efficient vector representations for subsequent retrieval.
[0040] Compared with the prior art, the beneficial effects of the present invention are:
[0041] The method and system for building a video knowledge base and efficient retrieval proposed by the present invention upload the video source file to the object file storage and return the access address; then separate the audio from the video source file, and then input the audio into the audio recognition and text model. Finally, the vectorized content of the text after audio recognition and the access address of the video source file are associated and uploaded to the vector library. When retrieving, the retrieved question or relevant content of the video can be vectorized by a vectorization tool, and then matched with the relevant content in the vector library to obtain the video access address, completing the efficient retrieval process of the video. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a schematic flowchart of the method of the present invention;
[0043] Figure 2 It is a flowchart for building a video knowledge base of the present invention;
[0044] Figure 3 It is a flowchart for efficient retrieval of videos of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] In order to clearly and completely describe the purpose, technical solution of the present invention and make the advantages more clear, the following further details the embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are part of the embodiments of the present invention, rather than all of the embodiments, and are only used to explain the embodiments of the present invention, not to limit the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present invention.
[0046] Embodiment 1, please refer to Figures 1 to 3 , the present invention provides a technical solution: a method for building a video knowledge base and efficient retrieval, the method includes the following steps:
[0047] I. Construction of video knowledge base
[0048] (1) Deploy a vector library and a file storage service
[0049] The advanced vector database Milvus and object storage service MinIO will be used to build the storage architecture of the video knowledge base. The core advantage of Milvus lies in its efficient vector retrieval algorithm and flexible expansion ability. In the construction of the video knowledge base, video content is often converted into feature vectors for storage to enable content-based retrieval. In addition, as the scale of the video knowledge base grows, Milvus does not require complex reconstruction or migration of the existing system, thus ensuring the long-term stability and scalability of the system. MinIO, as a high-performance and scalable object storage service, provides a solid foundation for the secure storage and efficient access of video files. Even in the face of a single point of failure, the system can ensure the continuous availability and integrity of data through automatic failover and data replication mechanisms.
[0050] Integrating the Milvus vector database and MinIO object storage service into the storage architecture of the video knowledge base can give full play to the advantages of both, forming a powerful data processing and storage capacity. Specifically, through Milvus, unstructured retrieval of video content can be achieved, and users can quickly find the required video resources; while MinIO is responsible for the efficient storage and access of these video files, ensuring that users can smoothly watch and download video content.
[0051] (2) Analysis and processing of video source files
[0052] Video upload and preprocessing: The system provides an upload interface that can transmit various video formats for users. After the users upload the video, the system will automatically perform preprocessing of the video file, including video format conversion, key frame extraction, etc., to ensure the accuracy of subsequent analysis.
[0053] Audio separation and speech recognition: After the preprocessing of the video file is completed, the system uses open-source tools such as FFMPEG to separate the audio information in the video file, and converts the separated audio into text information through an advanced speech recognition model (such as a deep learning model). This process will ensure the accurate extraction of audio information and provide a basis for subsequent text analysis.
[0054] Text information collation and timestamp association: The system collates the extracted text information, including noise removal, word segmentation, part-of-speech tagging, etc., and associates the text information with the time points in the video to form a correspondence of video-text-timestamp.
[0055] (3) Storing video source files in file storage
[0056] In the file storage service MinIO, we set up independent buckets for each user or video category and set appropriate access permissions. Video files will be uploaded to the corresponding buckets according to preset rules and a unique access address will be generated. This address will be stored in the vector library together with the video text information for convenient subsequent retrieval and access.
[0057] (4) Vectorization and storage of video text data
[0058] We adopt advanced text vectorization technologies (such as Word2Vec, BERT, etc.) to convert the extracted text information into vector representations. These vectors will be used as the feature representations of the video content and stored in the Milvus vector library. At the same time, the system will also store the metadata of the text information (such as title, description, tags, etc.) and the corresponding video access address in the vector library for subsequent retrieval and access.
[0059] II. Efficient video retrieval
[0060] (1) Vectorized retrieval of retrieval content
[0061] Users can initiate a retrieval request by entering relevant content of the video or classification retrieval keywords. The system first vectorizes the user input using the same text vectorization tool, and then sends the resulting vector as a query input to the Milvus vector library for similarity retrieval. Milvus will return the associated video addresses according to the vector similarity.
[0062] (2) Video access
[0063] According to the associated results returned by Milvus, the system will obtain the corresponding video file access address from the MinIO object storage service and provide it to the user for video viewing or downloading. In addition, the system can also adjust the vector similarity of the video library according to the user's viewing behavior and feedback information, so as to be able to more accurately return the video results selected by most users. At the same time, when the user is watching the video, the background will re-perform the audio separation and text vectorization processes to continuously dynamically update and optimize the video knowledge base to improve the accuracy and efficiency of retrieval.
[0064] Embodiment 2, based on Embodiment 1, proposes a system for a video knowledge base construction and efficient retrieval method, and the system includes:
[0065] A vector library and a file storage service deployment module, used to integrate the Milvus vector database and the MinIO object storage service to build the storage architecture of the video knowledge base, where Milvus is used to store the feature vectors of the video content to achieve content-based retrieval, and MinIO is used for the secure storage and efficient access of video files;
[0066] Video source file analysis and processing module, including:
[0067] Video upload and preprocessing sub-module, which is used to receive video uploads in multiple formats and perform preprocessing such as format conversion and key frame extraction; the video upload and preprocessing sub-module supports upload interfaces for multiple video formats and automatically performs preprocessing of video files to ensure the accuracy of subsequent analysis;
[0068] Audio separation and speech recognition sub-module, which is used to separate video audio and convert it into text information; the audio separation and speech recognition sub-module uses open-source tools such as FFMPEG to separate video audio and converts the audio into text information through a deep learning model;
[0069] Text information collation and timestamp association sub-module, which is used to collate text information and associate it with time points in the video; the text information collation and timestamp association sub-module performs denoising, word segmentation, and part-of-speech tagging on the extracted text information and associates it with time points in the video to form a correspondence between video-text-timestamp;
[0070] File storage module, which is used to set up independent buckets in MinIO for each user or video classification, store video files, and generate unique access addresses; the buckets set up in MinIO by the file storage module have appropriate access permissions to ensure the secure storage and access of video files;
[0071] Video text data vectorization and storage module, which is used to convert text information into vector representations using text vectorization techniques and store them together with the metadata of the text information and the video access address in the Milvus vector library; the text vectorization techniques used by the video text data vectorization and storage module include Word2Vec and BERT to convert text information into efficient vector representations for subsequent retrieval.
[0072] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for constructing a video knowledge base and efficient retrieval, characterized in that: The method includes the following steps: Deploying storage services: Build the storage architecture of the video knowledge base using the vector database Milvus and the object storage service MinIO, where Milvus is used to store the feature vectors of video content, and MinIO is used to store video files; Processing video source files: Receive and upload video files, and preprocess the video files, including format conversion and key frame extraction; Separate the audio information in the video and convert it into text information; Organize the text information, including removing noise, word segmentation, and part-of-speech tagging, and associate the text information with the time points in the video; Storing video files: Store the processed video files in the MinIO object storage service, set up independent storage buckets for each user or video classification, and generate unique access addresses; Vectorizing and storing text data: Convert the extracted text information into vector representations, and store the vectors, metadata of the text information, and video access addresses in the Milvus vector library.
2. The video knowledge base construction and efficient retrieval method according to claim 1, characterized in that: The processing steps of the video source files further include: Providing upload interfaces for multiple video formats and automatically preprocessing the video files; Using an audio separation tool to separate the audio information in the video file and converting the audio into text information through a speech recognition model; Organizing the text information to form the corresponding relationship of video-text-timestamp.
3. The video knowledge base construction and efficient retrieval method according to claim 2, characterized in that: The video knowledge base constructed by the video knowledge base construction method includes the following steps: Receiving the retrieval keywords input by the user; Using the same text vectorization tool as when constructing the video knowledge base to vectorize the retrieval keywords input by the user; Sending the obtained vector as a query input to the Milvus vector library for similarity retrieval; Obtaining the video addresses associated with the retrieval keywords according to the similarity retrieval results returned by Milvus.
4. A method for constructing and efficiently retrieving a video knowledge base according to claim 3, characterized in that: The method further includes the following steps: According to the obtained video addresses, obtaining the corresponding video files from the MinIO object storage service and providing them to the user for viewing or downloading; Collecting the user's viewing behaviors and feedback information, and adjusting the vector similarity of the video library according to the information to optimize the accuracy of the retrieval results.
5. A method for constructing and efficiently retrieving a video knowledge base according to claim 4, characterized in that: The method further includes the following dynamic update and optimization steps: When the user is watching a video, the background re-performs audio separation and text vectorization processing; Updating the newly generated text vectors and related information to the Milvus vector library to dynamically update and optimize the video knowledge base and improve the accuracy and efficiency of the retrieval.
6. A system for the method of constructing and efficiently retrieving a video knowledge base according to claim 5, characterized in that: The system includes: A vector library and file storage service deployment module, which is used to integrate the Milvus vector database and the MinIO object storage service to build the storage architecture of the video knowledge base, where Milvus is used to store the feature vectors of video content to achieve content-based retrieval, and MinIO is used for the secure storage and efficient access of video files; A video source file analysis and processing module, including: A video upload and preprocessing sub-module, which is used to receive the upload of videos in multiple formats and perform preprocessing such as format conversion and key frame extraction; An audio separation and speech recognition sub-module, which is used to separate the video audio and convert it into text information; The text information sorting and timestamp association sub-module is used to sort text information and associate it with time points in the video; The file storage module is used to set up independent storage buckets in MinIO for each user or video classification, store video files, and generate unique access addresses; The video text data vectorization and storage module is used to convert text information into vector representations using text vectorization techniques, and store them together with the metadata of the text information and the video access address in the Milvus vector database.
7. A system according to claim 6, characterized in that: The video upload and preprocessing sub-module supports upload interfaces for multiple video formats and automatically preprocesses video files to ensure the accuracy of subsequent analysis.
8. A system according to claim 7, characterized in that: The audio separation and speech recognition sub-module uses open-source tools such as FFMPEG to separate video audio and converts the audio into text information through a deep learning model; the text information sorting and timestamp association sub-module performs denoising, word segmentation, and part-of-speech tagging on the extracted text information and associates it with time points in the video to form a correspondence between video-text-timestamp.
9. A system according to claim 8, wherein: The storage buckets set up by the file storage module in MinIO have appropriate access permissions to ensure the secure storage and access of video files.
10. A system according to claim 9, characterized in that: The text vectorization techniques adopted by the video text data vectorization and storage module include Word2Vec and BERT to convert text information into efficient vector representations for subsequent retrieval.
Citation Information
Patent Citations
Method for realizing quick retrieval of mass videos
CN104050247A
Short video content label knowledge base quick retrieval method based on natural language processing
CN117009461A
ES retrieval knowledge base method based on BERT enhancement
CN118885565A
Video semantic retrieval method and device based on deep learning, equipment and medium
CN119311915A