Method for quickly searching pictures based on video data

By segmenting video lenses, extracting and vectorizing keyframes, and combining text similarity calculations, the problems of slow retrieval speed and low accuracy in massive video data are solved, and fast and accurate video picture positioning is achieved.

CN120407847APending Publication Date: 2025-08-01LINKER
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510363743.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The prior art has slow retrieval speed and low accuracy in massive video data, and cannot effectively perform cross-modal retrieval of text and video content.

Method used

By detecting video transition points, segmenting the lens segment, extracting keyframes and vectorizing them, using the FAISS library to build index-accelerated search, combining the BERT model for text conversion and similarity calculation, and integrating the weight adjustment of the lens start frame, end frame and keyframe to improve matching accuracy.

Benefits of technology

It realizes the rapid and accurate positioning of the pictures that users are interested in in the video, reduces the resource consumption and labor costs of frame-by-frame query, and improves the retrieval efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407847A_ABST
    Figure CN120407847A_ABST
Patent Text Reader

Abstract

The invention discloses a method for quickly searching a picture based on video data, which comprises the following steps of: S1, searching transition points in a video, segmenting the video into shot segments according to the transition points, and extracting key frames in each shot segment; s2, vectorizing the key frame to generate a key frame vector; s3, storing the key frame vector and the time point of the corresponding key frame into a database; s4, converting the query text into a text vector; and S5, calculating the similarity between the text vector and all key frame vectors in the database, and returning a matching result. And the user can click the returned cover frame and directly jump to a corresponding position of the video. According to the scheme, the corresponding picture can be quickly and accurately queried, the high resource consumption of frame-by-frame query and the labor cost of manual query are avoided, and the method is suitable for the field of video retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video data processing, and in particular to a method for quickly searching for video images through video transition segmentation, key frame vectorization, and text-image cross-modal retrieval technology. Background Art

[0002] With the explosive growth of Internet video content, how to quickly and accurately find the images that users are interested in from a large number of videos has become an important issue. Traditional video retrieval methods rely on manual marking or keyword-based search, with low efficiency and difficulty in processing a large amount of video data. Existing content-based video retrieval technologies (such as color histograms and motion features) can partially improve efficiency, but there are the following problems: 1. Slow retrieval speed: Frame-by-frame comparison leads to high computational complexity. 2. Difficulty in cross-modal retrieval: It is impossible to directly match video content through text queries. 3. Inaccurate key frame extraction: Traditional methods are easily interfered by in-shot motion. Summary of the Invention

[0003] The present invention mainly solves the technical problems existing in the prior art, such as slow retrieval speed and low accuracy, and provides a method for quickly searching for images based on video data with high retrieval efficiency and accurate positioning.

[0004] The present invention mainly solves the above technical problems through the following technical solutions: A method for quickly searching for images based on video data includes the following steps: S1. Find the transition points in the video, segment the video into shot segments according to the transition points, and extract the key frames in each shot segment; S2. Vectorize the key frames to generate key frame vectors; S3. Store the key frame vectors and the corresponding time points of the key frames in a database; This solution uses the FAISS library to build an IVF (Inverted File System) index to accelerate the nearest neighbor search; S4. Convert the query text into a text vector; S5. Calculate the similarity between the text vector and all the key frame vectors in the database, and return the matching result.

[0005] The user can click on the returned cover frame to directly jump to the corresponding position in the video.

[0006] Preferably, the specific step 1 is as follows: S101. Calculate the distance between the RGB color histograms of two consecutive frames, and the formula is as follows: ; In the formula, N is the number of histogram intervals set in advance, H t(i) represents the histogram value of the i-th interval in the t-th frame; that is, first divide the histogram into intervals according to the set N, and then calculate the distance according to the above formula. N is empirical data and is determined according to the actual usage scenario, and can be one of 16, 32, and 256; S102. If D t is greater than the preset L1 distance threshold, it is determined that there is a transition point between these two frames; S103. Repeat steps S101 and S102 until all transition points in the video are found; S104. Segment the video into several shot segments according to the transition points, that is, the video segment between two adjacent transition points is a shot segment; S105. For each shot segment, select the middle frame as the key frame. If the total number of frames T in the shot segment is even, select the T / 2-th frame as the key frame.

[0007] The transition detection adopts a mutation detection method based on the histogram. The L1 distance threshold can be 0.3.

[0008] Preferably, the step S2 is specifically: input the key frame into the feature extraction model, and use the output of the last fully connected layer of the feature extraction model as the feature vector, and then perform L2 normalization on the feature vector to obtain the key frame vector.

[0009] Preferably, in the step S2, the feature extraction model is a pre-trained ResNet-50 model.

[0010] Preferably, the step S4 is specifically: input the query text into the BERT model, extract the hidden layer output of the [CLS] token as the text feature vector, and then perform L2 normalization processing on the text feature vector to obtain the text vector corresponding to the query text.

[0011] Preferably, the step S5 is specifically: S501. Calculate the matching degree S0 between the text vector and the key frame vector using the cosine similarity: S0=(V text •V image ) / (||V text ||2•||V text ||2); In the formula, V text is the text vector, and V image is the key frame vector; S502. Select the top K results with the highest matching degree and sort them in descending order; S503. Return the key frames corresponding to the results selected in step S502 and their corresponding time points to the user.

[0012] Preferably, step S2 further includes inputting the start frame and the end frame of each shot segment into the feature extraction model respectively, and performing L2 normalization on the output of the last fully connected layer of the feature extraction model to obtain a start frame vector and an end frame vector; in step S3, the start frame vector, the end frame vector, the key frame vector, and the time point corresponding to the key frame are stored together. In step S501, after obtaining the matching degree between the text vector and the key frame vector, it is corrected in the following way: Calculate the similarity S between the text vector and the start frame vector s and the similarity S between the text vector and the end frame vector e , and the formula is the same as the calculation formula of S0; Fuse S s and S e into S0 in the way of weighted fusion to obtain the corrected matching degree S k : S k =a×S0 + b×S s + c×S e ; In the formula, a is the weight of the similarity S0 between the text vector and the key frame vector, b is the weight of the similarity S s between the text vector and the start frame vector, and c is the weight of the similarity S e between the text vector and the end frame vector; when sorting later, the corrected matching degree S k is used for sorting.

[0013] The start frame usually contains the introduction of a new scene, new object, or new information after a shot transition, which is crucial for understanding the beginning of the shot. The end frame usually contains the last frame before the end of the shot, which may indicate the start of the next shot or contain the final result of some actions or events within the shot. The middle frames represent the main content of the shot. Combining these three can capture the content changes and key information of the shot more completely, avoiding information loss that may occur when relying solely on a single frame (such as the middle frame).

[0014] Preferably, the weights a, b, and c are determined in the following way: Input the key frame, the start frame, and the end frame into the pre-trained image recognition model respectively to obtain the key frame text description, the start frame text description, and the end frame text description; Calculate the similarity S k-s between the key frame text description and the start frame text description, and calculate the similarity S k-e between the key frame text description and the end frame text description; Determine the weights a, b, and c by the following formula: a = 1 / (S k-s + Sk-e + 1); b = S k-s / (S k-s + S k-e + 1); c = S k-e / (S k-s + S k-e + 1).

[0015] Semantic difference reflects the change in the meaning of the scene or object between frames. In this solution, the weights of the starting frame and the ending frame are determined through the semantic differences between frames, so as to more comprehensively understand the content of the video shot, and the changes in the query keywords can be better incorporated into the query to obtain more accurate query results.

[0016] The substantial effects brought by the present invention are as follows: The corresponding picture can be quickly and accurately queried, avoiding the high resource consumption of frame-by-frame query and the labor cost of manual query; The information can be utilized more comprehensively, not only considering the key frames, but also considering the starting frame and the ending frame of the lens segment, which can capture the lens content more comprehensively and improve the search accuracy. By dynamically adjusting the weight ratio of each frame in the final matching degree, the matching degree is made more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a flowchart of a method for quickly finding pictures based on video data according to the present invention; Figure 2 is another flowchart of a method for quickly finding pictures based on video data according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] The technical solution of the present invention will be further specifically described below through embodiments and in conjunction with the drawings.

[0019] Embodiment 1: A method for quickly finding pictures based on video data in this embodiment, as Figure 1 shown, includes the following steps: S1. Find the transition points in the video, segment the video into lens segments according to the transition points, and extract the key frames in each lens segment; S2. Vectorize the key frames to generate key frame vectors; S3. Store the key frame vectors and the corresponding time points of the key frames in the database; This solution uses the FAISS library to build an IVF (Inverted File System) index to accelerate the nearest neighbor search; S4. Convert the query text into a text vector; S5. Calculate the similarity between the text vector and all the key frame vectors in the database and return the matching result.

[0020] The user can click on the returned cover frame to directly jump to the corresponding position of the video.

[0021] The specific steps of step 1 are as follows: S101. Calculate the L1 distance between the RGB color histograms of two consecutive frames. The formula is as follows: ; In the formula, N is the preset number of histogram intervals, and H t (i) represents the histogram value of the i-th interval of the t-th frame; that is, first divide the histogram into intervals according to the set N, and then calculate the distance according to the above formula. N is empirical data and is determined according to the actual usage scenario, and can be one of 16, 32, and 256; S102. If D t is greater than the preset L1 distance threshold, it is determined that there is a transition point between these two frames; S103. Repeat steps S101 and SIO2 until all transition points in the video are found; S104. Segment the video into several shot segments according to the transition points, that is, the video segment between two adjacent transition points is one shot segment; S105. For each shot segment, select the middle frame as the key frame. If the total number of frames T in the shot segment is even, select the T / 2-th frame as the key frame.

[0022] The transition detection uses a mutation detection method based on histograms. The L1 distance threshold is 0.3.

[0023] The specific steps of step S2 are as follows: Input the key frame into the feature extraction model, and use the output of the last fully connected layer of the feature extraction model as the feature vector, and then perform L2 normalization on the feature vector to obtain the key frame vector.

[0024] In step S2, the feature extraction model is a pre-trained ResNet-50 model.

[0025] The specific steps of step S4 are as follows: Input the query text into the BERT model, extract the hidden layer output of the [CLS] token as the text feature vector, and then perform L2 normalization on the text feature vector to obtain the text vector corresponding to the query text.

[0026] The specific steps of step S5 are as follows: S501. Calculate the matching degree S0 between the text vector and the key frame vector using cosine similarity: S0=(V text •V image ) / (||V text ||2•||Vtext ||2); Wherein, V text is the text vector, and V image is the key frame vector; S502. Select the top K results with the highest matching degree and sort them in descending order; S503. Return the key frames corresponding to the results selected in step S502 and their corresponding time points to the user.

[0027] Embodiment 2: A method for quickly searching for a picture based on video data in this embodiment, as Figure 2 shown, includes the following steps: S1. Search for the transition points in the video, segment the video into shot segments according to the transition points, and extract the key frames, start frames, and end frames in each shot segment; S2. Vectorize the key frames, start frames, and end frames to generate key frame vectors, start frame vectors, and end frame vectors; S3. Store the key frame vectors, start frame vectors, and end frame vectors and the time points where the corresponding key frames are located in the database; In this solution, the FAISS library is used to construct an IVF (Inverted File System) index to accelerate the nearest neighbor search; S4. Convert the query text into a text vector; S5. Calculate the similarity between the text vector and all the key frame vectors in the database, then correct the similarity, and return the matching results.

[0028] The user can click on the returned cover frame to directly jump to the corresponding position of the video.

[0029] In step S1, the method for searching for transition points and segmenting shot segments is the same as that in Embodiment 1.

[0030] In step S2, the method for generating key frame vectors, start frame vectors, and end frame vectors is the same as the method for generating key frame vectors in Embodiment 1.

[0031] Step S4 is the same as step S4 in Embodiment 1.

[0032] In step S5, the process of calculating the similarity between the text vector and all the key frame vectors in the database and the process of sorting according to the matching degree and returning the results are the same as those in Embodiment 1.

[0033] After obtaining the matching degree between the text vector and the key frame vector, it is corrected in the following way: Calculate the similarity S s between the text vector and the start frame vector and the similarity S e between the text vector and the end frame vector. The formula is the same as the calculation formula of S0; Fuse S s and S e into S0 according to the weighted fusion method to obtain the corrected matching degree S k : S k = a × S0 + b × S s + c × S e ; In the formula, a is the weight of the similarity S0 between the text vector and the key frame vector, b is the similarity S s between the text vector and the starting frame vector, and c is the similarity S e between the text vector and the ending frame vector; use the corrected matching degree S k for sorting in subsequent sorting.

[0034] The starting frame usually contains the introduction of a new scene, new object, or new information after a shot transition, which is crucial for understanding the beginning of the shot. The ending frame usually contains the last frame before the end of the shot, which may indicate the start of the next shot or contain the final result of certain actions or events within the shot. The middle frames represent the main content of the shot. Combining these three can capture the content changes and key information of the shot more completely, avoiding information loss that may occur when relying solely on a single frame (such as the middle frame).

[0035] The weights a, b, and c are determined as follows: Input the key frame, starting frame, and ending frame into the pre-trained image recognition model respectively to obtain the key frame text description, starting frame text description, and ending frame text description; Calculate the similarity S k-s between the key frame text description and the starting frame text description, and calculate the similarity S k-e between the key frame text description and the ending frame text description; Determine the weights a, b, and c by the following formula: a = 1 / (S k-s + S k-e + 1); b = S k-s / (S k-s + S k-e + 1); c = S k-e / (S k-s + S k-e + 1).

[0036] Semantic differences reflect changes in the meaning of scenes or objects between frames. In this solution, the weights of the starting frame and the ending frame are determined based on the semantic differences between frames, so as to more comprehensively understand the content of the video shot, better incorporate the changes in the query keywords into the query, and obtain more accurate query results.

[0037] The specific embodiments described in this article are only illustrative of the spirit of the present invention. Those skilled in the art to which the present invention pertains can make various modifications or supplements to the described specific embodiments or use similar methods for substitution, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.

[0038] Although terms such as transition points, key frames, and similarity are used more frequently in this article, the possibility of using other terms is not excluded. The use of these terms is only to more conveniently describe and explain the essence of the present invention; interpreting them as any additional limitation is contrary to the spirit of the present invention.

Claims

1. A method for quickly searching for a picture based on video data, characterized in that, It includes the following steps: S1. Locate the transition points in the video, segment the video into shot fragments according to the transition points, and extract the key frames in each shot fragment; S2. Vectorize the key frames to generate key frame vectors; S3. Store the key frame vectors and the corresponding time points of the key frames in the database; S4. Convert the query text into a text vector; S5. Calculate the similarity between the text vector and all the key frame vectors in the database, and return the matching results.

2. The method for quickly finding a picture based on video data according to claim 1, wherein The specific steps of step 1 are as follows: S101. Calculate the distance between the RGB color histograms of two consecutive frames. The formula is as follows: ; where N is the number of histogram intervals set in advance, and H t (i) represents the histogram value of the i-th interval in the t-th frame; S102. If D t is greater than a preset L1 distance threshold, it is determined that there is a transition point between these two frames; S103. Repeat steps S101 and S102 until all the transition points in the video are found; S104. Segment the video into several shot fragments according to the transition points; S105. For each shot fragment, select the middle frame as the key frame.

3. A method for quickly searching for a picture based on video data according to claim 1, characterized in that, The specific steps of step S2 are as follows: Input the key frames into the feature extraction model, and take the output of the last fully connected layer of the feature extraction model as the feature vector. Then perform L2 normalization on the feature vector to obtain the key frame vector.

4. The method for quickly finding a picture based on video data according to claim 2, wherein In step S2, the feature extraction model is a pre-trained ResNet-50 model.

5. A method for quickly finding a picture based on video data according to claim 1, characterized in that The specific steps of step S4 are as follows: Input the query text into the BERT model, extract the hidden layer output marked with [CLS] as the text feature vector, and then perform L2 normalization on the text feature vector to obtain the text vector corresponding to the query text.

6. The method for quickly finding a picture based on video data according to claim 1, wherein, The specific steps of step S5 are as follows: S501. Use cosine similarity to calculate the matching degree S0 between the text vector and the key frame vector; S0=(V text •V image ) / (||V text ||2•||V text ||2); Wherein, V text is the text vector, and V image is the key frame vector; S502. Select the top K results with the highest matching degree and sort them in descending order; S503. Return the key frames corresponding to the results selected in step S502 and their corresponding time points to the user.

7. A method for quickly finding a picture based on video data according to claim 6, characterized in that, Step S2 also includes inputting the starting frame and ending frame of each shot fragment into the feature extraction model respectively, and performing L2 normalization on the output of the last fully connected layer of the feature extraction model to obtain the starting frame vector and the ending frame vector. In step S3, store the starting frame vector, the ending frame vector, the key frame vector and the time point corresponding to the key frame together; In step S501, after obtaining the matching degree between the text vector and the key frame vector, it is corrected in the following way: Calculate the similarity S between the text vector and the starting frame vector s and the similarity S between the text vector and the ending frame vector e , and the formula is the same as the calculation formula of S0; Fuse S s and S e into S0 in a weighted fusion manner to obtain the corrected matching degree S k : S k = a × S0 + b × S s + c × S e ; Wherein, a is the weight of the similarity S0 between the text vector and the key frame vector, b is the weight of the similarity S s between the text vector and the start frame vector, and c is the weight of the similarity S e between the text vector and the end frame vector; during subsequent sorting, the corrected matching degree S k is used for sorting.

8. A method for quickly searching for a picture based on video data according to claim 7, characterized in that, The weights a, b, and c are determined in the following way: Input the key frame, the starting frame, and the ending frame into the pre-trained image recognition model respectively to obtain the key frame text description, the starting frame text description, and the ending frame text description; Calculate the similarity S between the text description of the key frame and the text description of the starting frame k-s and calculate the similarity S between the text description of the key frame and the text description of the ending frame k-e ; The weights a, b, and c are determined by the following formula: a = 1 / (S k-s + S k-e + 1); b = S k-s / (S k-s + S k-e + 1); c = S k-e / (S k-s + S k-e + 1).

Citation Information

Patent Citations

  • Video semantic scene segmentation method based on convolutional neural network

    CN107590442A

  • Video processing and searching method and device, electronic equipment and storage medium

    CN116010655A

  • Video data processing method, device and equipment and computer readable storage medium

    CN116975364A

  • Picture-based video retrieval method, system and equipment and storage medium

    CN117540047A

  • Fusion movie and television play content retrieval method and device based on large language model, face recognition, target detection and cross-modal vector, medium and product

    CN119669518A