Video retrieval method and device, electronic equipment and storage medium
By keyframe extraction of the original video to generate description text and matching based on the query text, the problem of low video retrieval accuracy in the prior art is solved, and more efficient and accurate video retrieval is achieved.
Patent Information
- Application Number
- CN202411971708.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-16
AI Technical Summary
The existing video search methods are based on video feature information, resulting in a single matching method and low retrieval accuracy, especially in specific video recognition scenarios.
The original video is generated by keyframe extraction and matching the description text based on the query text to determine the target video.
It enriches the types of feature information, enhances the adaptability of video retrieval, improves the accuracy of video retrieval, and makes the analysis of video content more concise and intuitive.
Smart Images

Figure CN120011591A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a video retrieval method, device, electronic device and storage medium. Background Art
[0002] In some specific application scenarios, video retrieval is required. The current video retrieval method is generally based on the feature information in the video. However, due to the single type of feature information, the matching method is not extensive, resulting in low video retrieval accuracy in some specific video recognition scenarios. Summary of the invention
[0003] In view of this, the purpose of the present disclosure is to provide a video retrieval method, device, electronic device and storage medium, which generate corresponding description text after key frame extraction from the original video, and query the description text based on the query text to determine the target video, and can use the description text to find the video, so that the analysis of the video content is more concise and intuitive, enriches the types of feature information, enhances the adaptability of video frequency retrieval, and improves the accuracy of video retrieval.
[0004] In a first aspect, an embodiment of the present disclosure provides a video retrieval method, the video retrieval method comprising:
[0005] Preprocessing each original video in the target video set to obtain each preprocessed video;
[0006] Extract key frames from each original video in the target video set to obtain a target key frame sequence for each original video;
[0007] Generate description text corresponding to each target key frame in the target key frame sequence of each original video, and associate each description text with the corresponding original video to obtain a description text sequence of each original video;
[0008] In response to the retrieval instruction, the description text sequences of the original videos in the target video set are matched based on the query text, and the matched original videos are determined as the target videos.
[0009] In a second aspect, an embodiment of the present disclosure provides a video retrieval device, the video retrieval device comprising:
[0010] A preprocessing module, used to preprocess each original video in the target video set to obtain each preprocessed video;
[0011] An extraction module is used to extract key frames from each original video in the target video set to obtain a target key frame sequence of each original video;
[0012] A generation module, used to generate description texts corresponding to each target key frame in the target key frame sequence of each original video, and associate each description text with the corresponding original video to obtain a description text sequence of each original video;
[0013] The matching module is used to match the description text sequence of each original video in the target video set based on the query text in response to the retrieval instruction, and determine the matched original video as the target video.
[0014] In a third aspect, an embodiment of the present disclosure provides an electronic device, including a processor and a memory, wherein the memory stores machine executable instructions that can be executed by the processor, and the processor executes the machine executable instructions to implement the above-mentioned video retrieval method.
[0015] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the above-mentioned video retrieval method.
[0016] The embodiments of the present disclosure bring the following beneficial effects:
[0017] The above-mentioned video retrieval method, device, electronic device and storage medium preprocess each original video in the target video set to obtain each preprocessed video; extract key frames from each original video in the target video set to obtain a target key frame sequence of each original video; generate description texts corresponding to each target key frame in the target key frame sequence of each original video, and associate each description text with the corresponding original video to obtain a description text sequence of each original video; in response to a retrieval instruction, match the description text sequences of each original video in the target video set based on the query text, and determine the matched original video as the target video. In this method, by extracting key frames from the original video to generate corresponding description texts, and querying the description texts based on the query text to determine the target video, the description text can be used to search for the video, making the analysis of the video content more concise and intuitive, enriching the types of feature information, enhancing the adaptability of video frequency retrieval, and improving the accuracy of video retrieval.
[0018] Other features and advantages of the present disclosure will be described in the following description, and partly become apparent from the description, or understood by practicing the present disclosure. The purpose and other advantages of the present disclosure are realized and obtained by the structures particularly pointed out in the description, claims and drawings.
[0019] In order to make the above-mentioned objectives, features and advantages of the present disclosure more obvious and easy to understand, preferred embodiments are specifically cited below and described in detail with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the drawings required for use in the specific embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.
[0021] Figure 1 A schematic diagram of an embodiment of a video retrieval method provided by an embodiment of the present disclosure;
[0022] Figure 2 A schematic diagram of another embodiment of the video retrieval method provided by the embodiment of the present disclosure;
[0023] Figure 3 A schematic diagram of an embodiment of a viewing angle effect diagram of six cameras provided in an embodiment of the present disclosure;
[0024] Figure 4 A schematic diagram of a video retrieval device provided in an embodiment of the present disclosure;
[0025] Figure 5 A schematic diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0026] In order to make the purpose, technical solution and advantages of this embodiment clearer, the technical solution of this disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of this disclosure, rather than all the embodiments. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this disclosure.
[0027] This embodiment provides a video retrieval method, device, electronic device, and storage medium, which can be applied to video retrieval scenarios in any field, especially in the field of autonomous driving.
[0028] In one embodiment of the present disclosure, the video retrieval method can be run on a terminal device or a server. The terminal device can be a local terminal device. When the video retrieval method is run on a server, the method can be implemented and executed based on a cloud interaction system, wherein the cloud interaction system includes a server and a client device.
[0029] For ease of understanding, the specific process of this embodiment is described below. Figure 1 In this embodiment, an embodiment of the video retrieval method includes the following steps:
[0030] Step 101, preprocessing each original video in the target video set to obtain each preprocessed video;
[0031] Among them, as an example but not limitation, each original video in the target video set is a video initially shot by at least one camera on a target object in a specific field, and may also be a video shot by at least one camera on a target object in a specific field and subjected to video data preprocessing, wherein video data preprocessing includes but is not limited to video format conversion, resolution adjustment, frame rate adjustment, color correction and enhanced security detection, and stabilization processing (using algorithms to compensate for camera jitter or unstable movement to make the video picture smoother).
[0032] Among them, as an example but not limitation, the target video set includes videos of at least one target device. For example, in the field of autonomous driving, the target video set may include videos of an autonomous driving vehicle (target device) or videos of more than one autonomous driving vehicles. Each original video in the target video set is a first-perspective video, and the target video set may also include videos from the first perspective and outside the first perspective (third perspective). When the target video set includes videos from the first perspective and outside the first perspective (third perspective), it is necessary to fuse the first perspective and videos outside the first perspective (third perspective) corresponding to the target device in the target video set to obtain a video mainly based on the target device.
[0033] The process of preprocessing each original video in the target video set is a video processing process other than video data preprocessing, for example, video denoising processing.
[0034] Step 102, extracting key frames from each original video in the target video set to obtain a target key frame sequence for each original video;
[0035] As an example but not a limitation, when extracting key frames from each original video in the target video set to obtain the initial key frame sequence of each original video, the following can be done: obtaining continuous frames of each original video in the target video set; extracting continuous frames of each original video based on a preset time interval to obtain an image set of each original video after initial screening; detecting the difference (e.g., brightness, color, texture, etc.) between adjacent frames in the image set after initial screening of each original video, if the difference is greater than a preset threshold, selecting the image of the corresponding frame to obtain an image set after re-screening of each original video; obtaining the image features of each image in the image set after re-screening of each original video, evaluating and extracting the image set after re-screening of each original video based on the image features, to obtain the initial key frame sequence of each original video, wherein the image features include but are not limited to edge and object information. This achieves multi-angle and multi-level screening of key frames, and improves the accuracy and reliability of the initial key frame sequence of each original video.
[0036] Among them, in one implementable manner, it is possible to: extract key frames from each original video in the target video set to obtain an initial key frame sequence of each original video; screen the initial key frame sequence of each original video to obtain a target key frame sequence of each original video. Specifically, when screening the initial key frame sequence of each original video to obtain the target key frame sequence of each original video, it is possible to: extract features from the initial key frame sequence of each original video to obtain feature information corresponding to each initial key frame; calculate the similarity between each pair of frames based on the feature information; perform clustering analysis on the initial key frame sequence of each original video based on the similarity between each pair of frames using a preset clustering algorithm to classify similar frames into one category to obtain a clustered key frame sequence; calculate the distance between each frame and the cluster center; extract the clustered key frame sequence based on the distance to obtain a candidate key frame sequence of each original video; post-process the candidate key frame sequence of each original video (for example, remove redundant frames, adjust the order of key frames) to obtain a target key frame sequence of each original video. The accuracy, reliability and efficiency of screening are improved, thereby improving the accuracy and reliability of the target key frame sequence of each original video.
[0037] Step 103, generating description texts corresponding to each target key frame in the target key frame sequence of each original video, and associating each description text with the corresponding original video to obtain a description text sequence of each original video;
[0038] As an example but not a limitation, the following can be used: by using a first preset model, feature extraction is performed on each target key frame in the target key frame sequence of each original video to obtain feature information of each target key frame; by using a second preset model, corresponding description text is generated based on the feature information of each target key frame; a correspondence is created between each description text and the corresponding original video to obtain a description text sequence of each original video. The description text is used to express the information of the field scene corresponding to the image, for example, the traffic behavior of the autonomous driving vehicle in the driving scene of the autonomous driving field.
[0039] Step 104 , in response to the search instruction, matching the description text sequences of the original videos in the target video set based on the query text, and determining the matched original videos as the target videos.
[0040] Among them, as an example but not limitation, the query text is a piece of text input into the search engine, database or information retrieval system to trigger the search instruction; the query text can be a keyword, a phrase or a complex query statement, for example, a question, a keyword, a sentence or a longer paragraph, depending on the user's needs and the characteristics of the system used, and there is no specific limitation; the type of query text can be a keyword query, a natural language query or a Boolean query.
[0041] The above-mentioned video retrieval method generates corresponding description text after extracting key frames from the original video, and queries the description text based on the query text to determine the target video. It can use the description text to find the video, making the analysis of the video content more concise and intuitive, enriching the types of feature information, enhancing the adaptability of video frequency retrieval, and improving the accuracy of video retrieval.
[0042] See also Figure 2 Another embodiment of the video retrieval method in this embodiment includes:
[0043] Step 201: pre-process each original video in the target video set to obtain each pre-processed video;
[0044] In one implementation, when preprocessing each original video in the target video set to obtain each preprocessed video, you can: identify each original video in the target video set to obtain an identification result, wherein each original video is a video from the first perspective; if the identification result indicates that the original video contains a target object, then remove the target object in the original video to obtain each preprocessed video.
[0045] As an example but not a limitation, the target object information to be identified is determined; each original video in the target video set is identified based on the target object information by a preset recognition model to obtain a recognition result; if the recognition result indicates that the original video does not contain the target object, no processing is performed. For example, the target video set is a video of the first-person perspective of an autonomous driving vehicle, and the target object information is vehicle information. If the original video contains a vehicle, the vehicle in the original video is deleted to obtain each pre-processed video.
[0046] Among them, as an example but not limitation, when removing the target object in the original video, you can: remove the target object in the original video based on a mask method, a chroma cutout method, a background replacement method or a mask method; the mask method, for example, responds to a mask adding instruction and adds a mask to the original video; responds to an adjustment instruction and adjusts the size and position of the mask so that the mask completely covers the target object, and performs edge feathering on the mask.
[0047] By identifying and deleting the target objects in the first-person video, the purity of the original video is improved, the quality of the original video is optimized, and the accuracy of data analysis is improved, which is conducive to improving the robustness of each pre-processed video processing and improving the accuracy and reliability of each pre-processed video.
[0048] In another implementation, when preprocessing each original video in the target video set to obtain each preprocessed video, it is also possible to: identify each original video in the target video set to obtain identification information, wherein the identification information includes time period and device information; obtain multiple other videos corresponding to each original video based on the identification information, wherein the multiple other videos are videos corresponding to multiple cameras other than the original video; process each original video and the multiple other videos corresponding to each original video to obtain each preprocessed video.
[0049] Take an autonomous vehicle as an example: Since an autonomous vehicle needs to pay attention to various places in all directions during driving, the front main camera alone is far from being able to fully reflect the situations that a vehicle may encounter during driving. Therefore, in order to describe the vehicle's situation more accurately, it is necessary to use the capabilities of the sensors on the autonomous vehicle as much as possible. For example, an autonomous vehicle is equipped with six cameras. The perspective effect diagram of the six cameras can be shown as follows: Figure 3 shown.
[0050] Among them, as an example but not limitation, the time period in the identification information can be understood as the duration between the initial moment and the end moment of the original video, and the device information can be understood as the information of the device where the camera is configured, for example, the autonomous driving vehicle where the camera is located.
[0051] As an example but not a limitation, a preset fine-tuning model (for example, a Robustly Optimized BERT Approach, 2.
[0052] RoBERTa), Enhanced Representation through Knowledge Integration (ERNIE), etc.) are used to generate description texts corresponding to target key frames in the target key frame sequences of each original video. In the training process of the fine-tuning (Finetune) model, the pre-training data can be fused with videos corresponding to multiple cameras, and the fine-tuning (Finetune) model can be trained for text reasoning through the fused multi-camera videos.
[0053] By processing each original video and multiple other videos corresponding to each original video, the video content is optimized and enhanced, which helps to conduct a deeper analysis and understanding of the video content, helps to improve the accuracy of describing behaviors in blind spots / blind spots that the main camera cannot see clearly, helps to achieve rapid video retrieval, and thus helps to improve the accuracy of video retrieval.
[0054] In one implementation, when processing each original video and multiple other videos corresponding to each original video to obtain each pre-processed video, it is possible to: weight each original video and multiple other videos corresponding to each original video to obtain a configured video set corresponding to each original video, wherein each original video is a video from a first perspective; and group the configured video sets corresponding to each original video into videos from a target perspective to obtain each pre-processed video.
[0055] Among them, as an example but not limitation, among the weights of each original video and multiple other videos corresponding to each original video, the weight hyperparameter of the main camera is higher, and the weight hyperparameters corresponding to other cameras are lower; the target perspective is the third perspective.
[0056] As an example but not a limitation, when weighting each original video and a plurality of other videos corresponding to each original video to obtain a configured video set corresponding to each original video, it is possible to: set the weights of each original video and a plurality of other videos corresponding to each original video, and establish a corresponding video format based on the weighted video, thereby obtaining a configured video set corresponding to each original video, wherein the video format may be as shown in Table 1 below:
[0057] Table 1 Video formats
[0058] left front Main camera right front Left rear Post Right rear
[0059] Wherein, as an example but not limitation, when the configured video sets corresponding to each original video are set to become the target perspective video, and each pre-processed video is obtained, the configured video sets corresponding to each original video can be spliced based on a preset splicing mode to obtain the to-be-processed video corresponding to each original video, wherein the preset splicing mode can be horizontal splicing, vertical splicing or grid splicing; or, the configured video sets corresponding to each original video are fused together by a preset seam-based local weighting algorithm or video fusion algorithm to form a smooth and coherent composite video to obtain the to-be-processed video corresponding to each original video; according to the relative position and angle of the camera, the transformation matrix from the first perspective to the third perspective is calculated; and the to-be-processed video corresponding to each original video is converted from the first perspective to the third perspective (i.e., the target perspective) by the transformation matrix to obtain each pre-processed video. The efficiency and accuracy of information processing are improved, and it is helpful to improve the accuracy and reliability of the description text generated by the model used to generate the description text.
[0060] By weighting the first-person perspective videos of multiple cameras and synthesizing the target perspective, the perspective coverage is improved, the information integration and processing capabilities are enhanced, multi-scenario applications and flexible configuration are supported, the accuracy and reliability of identifying objects or events in the video are improved, and it helps to improve the accuracy and reliability of the description text generated by the model used to generate the description text.
[0061] Step 202: Perform video rendering on each pre-processed video to obtain each target simulation video; wherein each target simulation video is a simulation video that retains key information;
[0062] Among them, as an example but not limitation, when rendering each preprocessed video to obtain each target simulation video, you can: obtain the target information of each preprocessed video, wherein the target information includes but is not limited to a bag / package identifier (bag_id), a start timestamp (start_timestamp) and an end timestamp (end_timestamp); record each preprocessed video based on the target information through a preset visualization rendering / interaction engine (for example, VIZ tool) to obtain each initial simulation video, and store the recorded video in a preset path, wherein sensor information is added to each initial simulation video, and the sensor information includes but is not limited to the current driving direction, speed, acceleration and steering angle; denoise each initial simulation video based on a preset noise type through a preset visualization rendering / interaction engine (for example, Viz-Vehicle L4 visualization rendering / interaction engine) to obtain each target simulation video, wherein the noise type can be determined according to the field scene corresponding to the target video set.
[0063] Step 203: extract key frames from each target simulation video to obtain an initial key frame sequence of each original video;
[0064] In one implementation, when performing key frame extraction on each target simulation video to obtain an initial key frame sequence of each original video, the following steps can be performed: obtaining continuous frames of each target simulation video; and determining the key frames in the continuous frames of each target simulation video based on target indicator information in each target simulation video to obtain an initial key frame sequence of each original video, wherein the target indicator information includes point cloud data and sensor information added during video rendering.
[0065] Among them, as an example but not limitation, the point cloud data in the target indicator information may be the obstacle distance, and the sensor information added during video rendering includes but is not limited to speed, acceleration and steering angle; based on the target indicator information in each target simulation video, the key frames in the continuous frames of each target simulation video are determined to obtain the initial key frame sequence of each original video, it can be: judged whether the target indicator information in each target simulation video is greater than a first preset threshold, if the target indicator information is greater than the first preset threshold, the corresponding frame is determined as a key frame, and the initial key frame sequence of each original video is obtained; or, the rate of change between the target indicator information of two consecutive frames is calculated, if the rate of change exceeds the second preset threshold, the corresponding frame is determined as a key frame, and the initial key frame sequence of each original video is obtained.
[0066] By determining the key frames through target indicator information, the representativeness of the key frames is improved, the adaptability to application scenarios is enhanced, the consumption of computing resources is reduced, the processing efficiency of subsequent description text generation is optimized, and the accuracy and reliability of the key frames are improved.
[0067] Step 204: Screen the initial key frame sequence of each original video to obtain the target key frame sequence of each original video;
[0068] By performing video rendering, key frame extraction and key frame screening on each original video in the target video set, the video content is optimized and enhanced, the video quality is improved, redundant information is reduced, complex scenes are supported, the data volume is reduced, and the retrieval capability is enhanced, which helps to improve the efficiency, accuracy and reliability of video retrieval.
[0069] In one implementation, when screening the initial key frame sequence of each original video to obtain the target key frame sequence of each original video, the initial key frame sequence of each original video can be selected based on the target indicator information in the image to obtain the target key frame sequence of each original video.
[0070] As an example but not a limitation, target indicator information is combined or fused according to specific requirements to obtain processed target indicator information; initial key frame sequences of each original video are selected based on the processed target indicator information to obtain initial key frame sequences of each original video. By selecting the initial key frame sequences of each original video based on the target indicator information, the relevance and accuracy of the key frames are improved.
[0071] In one implementation, when the initial key frame sequence of each original video is selected based on the target indicator information in the image to obtain the target key frame sequence of each original video, the following can be done: based on the weight of the target indicator information in the image, the initial key frame sequence of each original video is sorted to obtain the candidate key frame sequence of each original video; the indicator difference between the current frame and the previous frame in the candidate key frame sequence of each original video is obtained, wherein the indicator difference is the difference between the numerical values measured by the target indicator information; based on the indicator difference between the current frame and the previous frame, the candidate key frame sequence of each original video is sorted to obtain the sorted key frame sequence of each original video; and the target key frame sequence of each original video is determined from the sorted key frame sequence of each original video.
[0072] Among them, as an example but not limitation, the preset weight corresponding to the matching target indicator information is a weight value customized according to the scene elements that need to be paid attention to in the field, for example, in the field of autonomous driving, the weight of acceleration jump is the highest and the weight of steering angle jump is lower; according to the preset weight corresponding to the target indicator information, the initial key frame sequence of each original video is sorted to obtain the candidate key frame sequence of each original video, or, according to the preset weight corresponding to the target indicator information, the target indicator information is fused to obtain the fused target indicator information, and the initial key frame sequence of each original video is sorted according to the fused target indicator information to obtain the candidate key frame sequence of each original video; determine the target model for generating the description text, and determine the selection strategy according to the target model, wherein the selection strategy includes but is not limited to the selection quantity and the selected sorting position, for example, taking the top 12 to 36 frames in the sorting; based on the selection strategy, the sorted key frame sequence of each original video is selected to obtain the target key frame sequence of each original video.
[0073] By sorting the initial key frame sequence of each original video based on the weight of the target indicator information and selecting it according to the indicator difference, it has wide applicability, optimizes the selection of key frames, improves the quality of video processing, and makes subsequent video analysis more accurate and efficient.
[0074] Step 205: Generate description texts corresponding to each target key frame in the target key frame sequence of each original video, and associate each description text with the corresponding original video to obtain a description text sequence of each original video;
[0075] In one implementation, when generating descriptive text corresponding to each target key frame in the target key frame sequence of each original video, it is possible to: infer each target key frame in the target key frame sequence of each original video and generate text to obtain an initial text corresponding to each target key frame; adjust the initial text corresponding to each target key frame to obtain a descriptive text corresponding to each target key frame.
[0076] Among them, as an example but not limitation, through a preset inference model, each target key frame in the target key frame sequence of each original video is inferred, and text is generated to obtain the initial text corresponding to each target key frame, wherein the inference model can be an embedding (embedding) combined with the K-Nearest Neighbors (KNN) algorithm; the initial text corresponding to each target key frame is divided into clauses to obtain the divided text corresponding to each target key frame; the divided text corresponding to each target key frame is deduplicated and cleaned to obtain the description text corresponding to each target key frame.
[0077] Among them, in one implementation method, when adjusting the initial text corresponding to each target key frame to obtain the description text corresponding to each target key frame, it is possible to: based on at least one strategy in the preset optimization strategy set, adjust the initial text corresponding to each target key frame to obtain the description text corresponding to each target key frame, wherein the preset optimization strategy set includes but is not limited to adjustment rules, manual review and correction strategies, context-based optimization strategies and application scenario demand adjustment strategies, the adjustment rules include but are not limited to keyword replacement, sentence structure optimization, grammar checking and emotion adjustment, the context-based optimization strategies include but are not limited to timeline correlation analysis, event chain description analysis and context comparison analysis, the application scenario demand adjustment strategies include but are not limited to personalized application requirements and multiple language expressions.
[0078] By generating description text for each target key frame and adjusting the text, the understanding, analysis and application of video content are improved, making the description text corresponding to each target key frame more accurate, efficient and flexible.
[0079] Step 206 : In response to the search instruction, the description text sequences of the original videos in the target video set are matched based on the query text, and the matched original videos are determined as the target videos.
[0080] Among them, as an example but not limitation, the target video can be a specific original video. For example, based on a preset similarity algorithm, the similarity between the query text and each description text in the description text sequence of each original video is calculated; the description texts are sorted in order of similarity from large to small to obtain a target description text sequence; the original videos corresponding to the target description text sequence are deduplicated to obtain the target video; the target video can also be more than one different original videos. For example, based on a preset similarity algorithm, the similarity between the query text and each description text in the description text sequence of each original video is calculated; the original video corresponding to the maximum similarity is determined as the target video.
[0081] The above-mentioned video retrieval method generates corresponding description text after extracting key frames from the original video, and queries the description text based on the query text to determine the target video. It can use the description text to find the video, making the analysis of the video content more concise and intuitive, enriching the types of feature information, enhancing the adaptability of video frequency retrieval, and improving the accuracy of video retrieval.
[0082] Corresponding to the above method embodiment, see Figure 4 A schematic diagram of a video retrieval device is shown, the device comprising:
[0083] The preprocessing module 401 is used to preprocess each original video in the target video set to obtain each preprocessed video;
[0084] An extraction module 402 is used to extract key frames from each original video in the target video set to obtain a target key frame sequence of each original video;
[0085] A generating module 403 is used to generate description texts corresponding to each target key frame in the target key frame sequence of each original video, and associate each description text with the corresponding original video to obtain a description text sequence of each original video;
[0086] The matching module 404 is used to match the description text sequence of each original video in the target video set based on the query text in response to the search instruction, and determine the matched original video as the target video.
[0087] The above-mentioned video retrieval device generates corresponding description text after extracting key frames from the original video, and queries the description text based on the query text to determine the target video. It can use the description text to find the video, making the analysis of the video content more concise and intuitive, enriching the types of feature information, enhancing the adaptability of video frequency retrieval, and improving the accuracy of video retrieval.
[0088] Optionally, the preprocessing module 401 may also be used for:
[0089] Recognize each original video in the target video set to obtain a recognition result, wherein each original video is a first-person perspective video;
[0090] If the recognition result indicates that the original video contains the target object, the target object in the original video is removed to obtain each preprocessed video.
[0091] Optionally, the preprocessing module 401 may also be used for:
[0092] Identify each original video in the target video set to obtain identification information, wherein the identification information includes time period and device information;
[0093] Acquire multiple other videos corresponding to each original video based on the identification information, wherein the multiple other videos are videos corresponding to multiple cameras other than the original video;
[0094] Each original video and a plurality of other videos corresponding to each original video are processed to obtain each pre-processed video.
[0095] Optionally, the preprocessing module 401 may also be used for:
[0096] Performing weight configuration on each original video and a plurality of other videos corresponding to each original video to obtain a configured video set corresponding to each original video, wherein each original video is a video from a first-person perspective;
[0097] The configured videos corresponding to the original videos are aggregated into a video of the target perspective to obtain the preprocessed videos.
[0098] Optionally, the extraction module 402 may also be used to:
[0099] Performing video rendering on each preprocessed video to obtain each target simulation video; wherein each target simulation video is a simulation video that retains key information;
[0100] Extract key frames from each target simulation video to obtain an initial key frame sequence of each original video;
[0101] The initial key frame sequence of each original video is screened to obtain the target key frame sequence of each original video.
[0102] Optionally, the extraction module 402 may also be used to:
[0103] Acquire continuous frames of each target simulation video;
[0104] Based on the target indicator information in each target simulation video, the key frames in the continuous frames of each target simulation video are determined to obtain the initial key frame sequence of each original video, wherein the target indicator information includes the point cloud data and the sensor information added during the video rendering;
[0105] The step of screening the initial key frame sequence of each original video to obtain the target key frame sequence of each original video includes:
[0106] Based on the target indicator information in the image, the initial key frame sequence of each original video is selected to obtain the target key frame sequence of each original video.
[0107] Optionally, the extraction module 402 may also be used to:
[0108] Based on the weight of the target index information in the image, the initial key frame sequence of each original video is sorted to obtain the candidate key frame sequence of each original video;
[0109] Obtaining the index difference between the current frame and the previous frame in the candidate key frame sequence of each original video, wherein the index difference is the difference between the values measured by the target index information;
[0110] Based on the difference between the index of the current frame and the previous frame, the candidate key frame sequences of each original video are sorted to obtain the sorted key frame sequences of each original video;
[0111] A target key frame sequence for each original video is determined from the sorted key frame sequence of each original video.
[0112] Optionally, the generating module 403 may also be used to:
[0113] Inferring each target key frame in the target key frame sequence of each original video and generating text to obtain an initial text corresponding to each target key frame;
[0114] The initial text corresponding to each target key frame is adjusted to obtain the description text corresponding to each target key frame.
[0115] This embodiment also provides an electronic device, including a processor and a memory, wherein the memory stores machine executable instructions that can be executed by the processor, and the processor executes the machine executable instructions to implement the above video retrieval method. The electronic device can be a server or a terminal device.
[0116] See also Figure 5 As shown, the electronic device includes a processor 500 and a memory 501 , wherein the memory 501 stores machine executable instructions that can be executed by the processor 500 , and the processor 500 executes the machine executable instructions to implement the above-mentioned video retrieval method.
[0117] Further, Figure 5 The electronic device shown further includes a bus 502 and a communication interface 503 , and the processor 500 , the communication interface 503 and the memory 501 are connected via the bus 502 .
[0118] The memory 501 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk storage. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 503 (which may be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. may be used. The bus 502 may be an ISA bus, a PCI bus, or an EISA bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0119] The processor 500 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 500. The above processor 500 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in this embodiment can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in conjunction with this embodiment can be directly embodied as a hardware decoding processor for execution, or a combination of hardware and software modules in the decoding processor for execution. The software module can be located in a storage medium mature in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 501, and the processor 500 reads the information in the memory 501 and completes the steps of the video retrieval method in combination with its hardware.
[0120] This embodiment also provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the above-mentioned video retrieval method.
[0121] The computer program product of the video retrieval method, device, electronic device and storage medium provided in this embodiment includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the previous method embodiments. The specific implementation can be found in the method embodiments, which will not be repeated here.
[0122] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system and device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0123] In addition, in the description of this embodiment, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, or it can be the internal communication of two components. For those skilled in the art, the specific meanings of the above terms in this disclosure can be understood according to specific circumstances.
[0124] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present disclosure. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0125] In the description of the present disclosure, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present disclosure and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present disclosure. In addition, the terms "first", "second", and "third" are used for descriptive purposes only, and cannot be understood as indicating or implying relative importance.
[0126] Finally, it should be noted that the above embodiments are only specific implementation methods of the present disclosure, which are used to illustrate the technical solutions of the present disclosure, rather than to limit them. The protection scope of the present disclosure is not limited thereto. Although the present disclosure is described in detail with reference to the above embodiments, those skilled in the art should understand that any person skilled in the art who is familiar with the technical field can still modify the technical solutions recorded in the above embodiments within the technical scope disclosed in the present disclosure, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the present embodiments, and should be included in the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure shall be based on the protection scope of the claims.
Claims
1. A video retrieval method, characterized in that: The method comprises: Preprocessing each original video in the target video set to obtain each preprocessed video; Extract key frames from each original video in the target video set to obtain a target key frame sequence for each original video; Generate description text corresponding to each target key frame in the target key frame sequence of each original video, and associate each description text with the corresponding original video to obtain a description text sequence of each original video; In response to the retrieval instruction, the description text sequences of the original videos in the target video set are matched based on the query text, and the matched original videos are determined as the target videos.
2. The method according to claim 1, characterized in that The step of preprocessing each original video in the target video set to obtain each preprocessed video includes: Recognize each original video in the target video set to obtain a recognition result, wherein each original video is a first-person perspective video; If the recognition result indicates that the original video contains the target object, the target object in the original video is removed to obtain preprocessed videos.
3. The method according to claim 1, characterized in that The step of preprocessing each original video in the target video set to obtain each preprocessed video includes: Identify each original video in the target video set to obtain identification information, wherein the identification information includes time period and device information; Based on the identification information, a plurality of other videos corresponding to each original video are obtained, wherein the plurality of other videos are videos corresponding to a plurality of cameras other than the camera of the original video; Each original video and a plurality of other videos corresponding to each original video are processed to obtain each pre-processed video.
4. The method according to claim 3, characterized in that The step of processing each original video and a plurality of other videos corresponding to each original video to obtain each pre-processed video includes: Performing weight configuration on each original video and a plurality of other videos corresponding to each original video to obtain a configured video set corresponding to each original video, wherein each original video is a video from a first-person perspective; The configured videos corresponding to the original videos are aggregated into a video of the target perspective to obtain the preprocessed videos.
5. The method according to claim 1, characterized in that The step of extracting key frames from each original video in the target video set to obtain a target key frame sequence for each original video includes: Performing video rendering on each preprocessed video to obtain each target simulation video; wherein each target simulation video is a simulation video that retains key information; Extract key frames from each target simulation video to obtain an initial key frame sequence of each original video; The initial key frame sequence of each original video is screened to obtain the target key frame sequence of each original video.
6. The method according to claim 5, characterized in that The step of extracting key frames from each target simulation video to obtain an initial key frame sequence of each original video includes: Acquire continuous frames of each target simulation video; Based on the target indicator information in each target simulation video, key frames in the continuous frames of each target simulation video are determined to obtain an initial key frame sequence of each original video, wherein the target indicator information includes point cloud data and sensor information added during video rendering; The step of screening the initial key frame sequence of each original video to obtain the target key frame sequence of each original video includes: Based on the target indicator information in the image, the initial key frame sequence of each original video is selected to obtain the target key frame sequence of each original video.
7. The method according to claim 6, characterized in that The step of selecting the initial key frame sequence of each original video based on the target indicator information in the image to obtain the target key frame sequence of each original video includes: Based on the weight of the target index information in the image, the initial key frame sequence of each original video is sorted to obtain the candidate key frame sequence of each original video; Obtaining an index difference between a current frame and a previous frame in a candidate key frame sequence of each original video, wherein the index difference is a difference between values measured by the target index information; Based on the difference between the index of the current frame and the previous frame, the candidate key frame sequences of each original video are sorted to obtain the sorted key frame sequences of each original video; A target key frame sequence for each original video is determined from the sorted key frame sequence of each original video.
8. The method according to claim 1, characterized in that The step of generating a description text corresponding to each target key frame in the target key frame sequence of each original video includes: Inferring each target key frame in the target key frame sequence of each original video and generating text to obtain an initial text corresponding to each target key frame; The initial text corresponding to each target key frame is adjusted to obtain the description text corresponding to each target key frame.
9. A video retrieval device, characterized in that: The video retrieval device comprises: A preprocessing module, used to preprocess each original video in the target video set to obtain each preprocessed video; An extraction module is used to extract key frames from each original video in the target video set to obtain a target key frame sequence of each original video; A generation module, used to generate description texts corresponding to each target key frame in the target key frame sequence of each original video, and associate each description text with the corresponding original video to obtain a description text sequence of each original video; The matching module is used to match the description text sequence of each original video in the target video set based on the query text in response to the retrieval instruction, and determine the matched original video as the target video.
10. An electronic device, characterized in that: The invention comprises a processor and a memory, wherein the memory stores machine executable instructions that can be executed by the processor, and the processor executes the machine executable instructions to implement the video retrieval method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the video retrieval method according to any one of claims 1 to 8.
Citation Information
Cited By
Multi-mode video content retrieval method and device based on shot frame sampling and medium
CN121256087A