A Video Retrieval Method Based on Frame Index and Cross-Modal Representation

By adopting a method based on frame index and cross-modal representation in video retrieval, and using hadoop cluster and multi-task optimization of the cross-modal representation method enhanced by the following diagram structure, the problem of video retrieval in massive video libraries is solved, and a more efficient and stable video retrieval effect is achieved.

CN117972135BActive Publication Date: 2025-05-27SHANDONG BAIMENG INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202311873221.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-05-27
Estimated Expiration
2043-12-29

AI Technical Summary

Technical Problem

The existing video search methods are difficult to quickly and stably retrieve target videos in massive video libraries, with high computational complexity, and the cross-modal feature representation method is difficult and unstable to train, ignoring the correlation between internal data of the modality.

Method used

Using video retrieval methods based on frame index and cross-modal representation, a massive video library is preprocessed through the hadoop cluster to build a video retrieval database, and a multi-task optimization cross-modal representation method is used to map video frame images and text data to a unified feature space, and search using the Frobenius norm similarity algorithm.

Benefits of technology

The ability to mine data correlation between different modes and within the same mode is improved, the training difficulty is reduced, the search accuracy and stability is improved, and more efficient video retrieval is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117972135B_ABST
    Figure CN117972135B_ABST
Patent Text Reader

Abstract

The present invention relates to a video retrieval method based on frame index and cross-modal representation. The method includes the following steps: preprocessing a massive video library by using the distributed computing power of a hadoop cluster to construct a video retrieval database; performing frame segmentation on the video segment to be retrieved; using the proposed cross-modal representation method enhanced by a graph structure under multi-task optimization to map the two-modal data of video frame images and text obtained in step S1 to a unified cross-modal feature space; using the Frobenius norm similarity algorithm to find the position of the first frame that meets the similarity threshold in the video library and record the time sequence of the frame in the video where it is located; calculating the similarity of each video by using the similarity algorithm in step S3 and sorting according to the similarity of the videos, and selecting the top ten videos as the final retrieval results. The present invention solves the problem that it is difficult to quickly and stably retrieve the target video in a massive video source based on the existing title and keyword retrieval strategies in video retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of video retrieval and image-text cross-modal recognition, and specifically to a video retrieval method based on frame indexing and cross-modal representation. Background Art

[0002] Video is a very convenient medium for describing and storing the spatial, temporal, spectral, and physical components of information contained in various fields (such as satellite images, medical images, online education, etc.). Due to the low cost of storage devices, large storage space, and the progress of compression algorithms further reducing the cost of using the video medium, video plays an important role in depicting and disseminating image information. However, due to the huge amount and rapid growth of video data, it is difficult for users to browse each video and quickly retrieve the target video. Therefore, an efficient and stable method is needed to assist users in quickly retrieving videos from a video library.

[0003] In traditional methods, text features such as file names, titles, and keywords are generally used to annotate and retrieve videos. However, these methods have some problems. First, it is necessary to manually describe and label the content of the video according to a selected set of titles and keywords. In most videos, each frame may have many objects, and each object has its own set of attributes. Therefore, it is also necessary to express the spatial relationships of the various objects. However, as the video data grows, the keywords become more and more complex and difficult to express the actual content of the video.

[0004] Therefore, retrieving video content is more preferable in video retrieval, but cross-modal retrieval needs to be implemented. All existing cross-modal retrieval methods have a common goal, which is to achieve cross-modal feature representation, that is, to extract and transform the data features of different modalities into the same feature space. In this same space, similarity calculation can be performed to achieve the purpose of cross-modal retrieval. In this scenario, existing methods take the entire video as the retrieval object, resulting in a high computational complexity of the algorithm, which affects the retrieval speed. Extracting key frames in the video as the retrieval target is an effective strategy. At the same time, how to map the data of the two modalities of text and video frame images into a unified space for similarity calculation is an important step in achieving accurate, efficient, and stable video retrieval. Currently, for cross-modal feature representation methods, training requires a large number of sample pairs of data, and the acquisition of these data is difficult, and the marking is time-consuming and laborious. Secondly, some generative adversarial methods (GAN) can alleviate the demand for sample pairs, but existing GAN training schemes are difficult and unstable, affecting the final application effect. Finally, current methods mainly focus on the correlation between modalities and ignore the correlation between data within modalities (such as between data of the same category and different categories). Therefore, new methods need to be explored to simultaneously mine the intra-modal and inter-modal correlations, rely on multiple training tasks, reduce the training difficulty, and improve the stability. Summary of the Invention

[0005] Aiming at the problems proposed in the background technology, the purpose of the present invention is to propose a video retrieval method based on frame indexing and cross-modal representation.

[0006] To achieve the above purpose, the present invention provides the following technical solutions:

[0007] A video retrieval method based on frame indexing and cross-modal representation, the method includes the following steps:

[0008] S1. Establish a massive video processing module based on Hadoop, and use the distributed computing power of the Hadoop cluster to preprocess the massive video library to construct a video retrieval database; perform frame segmentation on the video segment to be retrieved and split it into individual picture frames; wherein, the process of preprocessing the massive video library is to extract and store the colors of video frames, and what is actually stored is the information obtained by normalizing and standardizing the RGB colors of video frames, rather than the video;

[0009] S2. Use the proposed cross-modal representation method enhanced by graph structure under multi-task optimization to map the video frame images and text data of the two modalities obtained in step S1 into a unified cross-modal feature space, so that the two modalities of data can perform feature similarity matching calculation in the unified space;

[0010] S3. Based on the mapping feature representation of the video frame images and text obtained in step S2, use the Frobenius norm similarity algorithm to find the position of the first frame that meets the similarity threshold in the video library, and record the time sequence of this frame in the video where it is located;

[0011] S4. Use the similarity algorithm in step S3 to calculate the similarity of each video and sort according to the similarity of the videos, select the top ten videos as the final retrieval results, and submit them to the user for manual review.

[0012] As a further technical solution of the present invention, in step S1, the method for extracting the color of the video frame in the preprocessing process is: represent the color information of the video frame through a separate three-dimensional histogram, and standardize the three-dimensional histogram; let H(i) be the three-dimensional histogram of the image, and i represents the bin of the histogram, then the standardized three-dimensional histogram can be expressed as formula 1:

[0013] As a further technical solution of the present invention, in step S2, in the proposed cross-modal representation method with enhanced graph structure under multi-task optimization, a structure encoding model under graph constraints and a multi-task stable training scheme are adopted. The structure encoding model under graph constraints uses the classification label information to establish a relational subgraph adjacency matrix within each modality, and uses a graph neural network model to encode information based on the constructed graph structure to obtain a pseudo cross-modal representation. The multi-task stable training scheme constructs two tasks using the label classification loss and the discrimination loss of different data under the pseudo cross-modal representation to achieve stable model training.

[0014] As a further technical solution of the present invention, for the structure encoding model: there is associated information between different samples in the database. Based on the video frame images and text modalities respectively, establish a graph to characterize the hidden relationships; treat each sample as a node, and the edges between them reflect the relationships between them; if the labels Label between two nodes i and j are the same, then we connect these two nodes with an edge, that is, e m1 (v i +v j ) = 1, where e m1 represents the edge under modality m 1 , otherwise e m1 (v i , v j ) = 0; each point will add a self-connecting edge pointing to the node itself, that is, e m1 (v i , v i ) = 1; through the above construction scheme, obtain the subgraph G m1 or G m2, based on the dual-branch graph convolutional network (GCN), new information aggregation for samples is performed within each modality for each graph branch by following the standard hierarchical propagation rules:

[0015] Equation 2:

[0016] where ε m1 is the adjacency matrix of the subgraph G m1 ; represents the node representation at the l-th layer of the neural network in modality m1; is the degree matrix, is the parameter of the scientific department of the l-th layer of the graph neural network in modality m1, σ(·) represents the non-linear activation function, and after passing through L layers of the graph neural network, the pseudo cross-modal representation

[0017] As a further technical solution of the present invention, the multi-task stable training scheme:

[0018] By using the label classification loss and the mutual information loss between the pseudo cross-modal representation and the mapped cross-modal representation to construct two tasks, the stable model training is realized collaboratively; first, the pseudo cross-modal representation obtained previously is used for the label classification task:

[0019] Equation 3:

[0020] where, ||·|| F represents the Frobenius norm, that is, the square root of the sum of the absolute value squares of the elements of the matrix, Y is the classification label, k represents the number of samples, and f class is the classifier model, which can be any classifier and is shared in each modality; through Equation 3, the data with the same semantic label has the same representation, thereby increasing the discrimination ability and reducing the cross-modal difference at the same time;

[0021] The second task is constructed through the discrimination loss of different data under the pseudo cross-modal representation; specifically, if the features of different modalities belong to the same category, the distance between them is minimized, and the definition of the distance loss L dis is:

[0022] Equation 4:

[0023] Integrating the above two tasks, the fused loss function is obtained:

[0024] Equation 5: L = L class + λ·L diss

[0025] where λ is the combination coefficient, representing Ldiss The weight of the task.

[0026] As a further technical solution of the present invention, in step S3, the step of using the Frobenius norm similarity algorithm to find the position of the first frame that meets the similarity threshold in the video library and record the timing of this frame in the video where it is located includes:

[0027] Use the Frobenius norm similarity algorithm to calculate the similarity with each frame image in each video of the video library and filter according to the set similarity threshold. The similarity threshold set in the present invention is 0.3; after finding the position of the first frame in the video library, record the timing of this frame in the video where it is located; then calculate and record the timing of the remaining frames in this video. Finally, calculate the similarity of the video using the similarity of the frames. The calculation method is: the sum of the products of the similarity of each frame and the corresponding timing difference is divided by the number of frames; each node of the hadoop cluster can calculate the similarity of a video, and finally summarize the similarities to the reduce node; the total number of nodes used in the present invention is 1000.

[0028] Compared with the prior art, the beneficial effects of the present invention are as follows: Starting from content-based video retrieval, according to the characteristics of the massive and rapidly growing video, first preprocess the massive video library based on hadoop and establish a massive video retrieval library, and then perform frame segmentation on the input video to be retrieved, splitting it into individual frames. A cross-modal retrieval method enhanced by a graph structure under multi-task optimization is proposed. First, use classification labels to construct relationship subgraphs within each modality, and use graph neural networks for encoding to capture the structural information within the modality; secondly, use the label classification task and the discriminative loss task between different spatial representations to establish a multi-task to improve the training stability. Through this method, the ability to mine the data correlation between different modalities and within the same modality can be improved, the performance and training stability in the case of scarce training data can be improved, and more accurate cross-modal feature representation can be achieved. Use the similarity algorithm to find the position of the first frame in the video library and record the timing of this frame in the video where it is located. Then calculate and record the timing of the remaining frames in this video, and use the timing difference from the first frame as the weight of each frame to calculate the weighted sum of the similarities of all frames as the similarity of the video. Finally, calculate the similarity of each video using the similarity of the frames. Then sort according to the similarity and select the top ten videos as the final retrieval results for manual review by the user. Through this frame-index-based video retrieval method, the problem that it is difficult to quickly and stably retrieve the target video in the massive video source based on the title and keyword retrieval strategies in the existing video retrieval can be effectively solved. Description of the Drawings

[0029] Figure 1Flowchart of a video retrieval method based on frame index and cross-modal representation.

[0030] Figure 2 Process model diagram of video library construction and image frame acquisition in a video retrieval method based on frame index and cross-modal representation.

[0031] Figure 3 Model diagram of cross-modal retrieval method with enhanced graph structure under multi-task optimization in a video retrieval method based on frame index and cross-modal representation. Detailed implementation manners

[0032] The technical solutions of this patent will be further described in detail below in combination with the specific implementation manners.

[0033] As an embodiment of the present invention, please refer to Figure 1 , a video retrieval method based on frame index and cross-modal representation, the method includes the following steps:

[0034] S1. Establish a massive video processing module based on Hadoop, and use the distributed computing power of the Hadoop cluster to preprocess the massive video library to construct a video retrieval database; perform frame segmentation on the video segment to be retrieved and split it into individual picture frames; the process model diagram of video library construction and image frame acquisition is as Figure 2 shown, wherein, the process of preprocessing the massive video library is to extract and store the colors of video frames, and what is actually stored is the information obtained by standardizing and normalizing the RGB colors of video frames, rather than the videos;

[0035] Since colors are generally represented in the RGB format, the color information of video frames can be represented by a single three-dimensional histogram. These color representations are basically invariant under image rotation and translation, and by normalizing the three-dimensional histogram, it can also be ensured that they are basically invariant under changes in the image frame size; let H(i) be the three-dimensional histogram of the image, and i represents the bin of the histogram, then the normalized three-dimensional histogram can be represented as

[0036] Formula 1:

[0037] S2. Use the proposed cross-modal representation method with enhanced graph structure under multi-task optimization to map the video frame images and text two-modal data obtained in step S1 to a unified cross-modal feature space, so that the two-modal data can perform feature similarity matching calculations in the unified space.

[0038] Aiming at the problems that existing methods ignore the correlation between data within a modality and have poor stability in the case of scarce data labels. First, use classification labels to construct a relational subgraph within each modality, and use a graph neural network for encoding to capture the structural information within the modality. Second, use the label classification task and the discriminative loss task between different spatial representations to establish a multi-task to improve training stability. Through this method, the ability to mine the correlation between data across different modalities and within the same modality can be improved, the performance and training stability in the case of scarce training data can be enhanced, and more accurate cross-modal retrieval can be achieved.

[0039] The schematic diagram of the model is as Figure 3 shown. It mainly includes two core parts. The first part is the structure encoding model under graph constraints, which uses label information to establish the adjacency matrix of the relational subgraph within each modality, and uses the graph neural network model to encode information based on the constructed graph structure to obtain pseudo cross-modal representations. The second part is the multi-task stable training scheme, which constructs two tasks using the label classification loss and the discriminative loss of different data under the pseudo cross-modal representation to achieve stable model training.

[0040] Structure encoding model under graph constraints:

[0041] There is correlation information between different samples in the database. However, they are largely ignored by existing methods. To mine valuable correlated semantic information from the data, a graph is established based on video frame images and text modalities respectively to represent the hidden relationships. Each sample is treated as a node, and the edges between them reflect the relationships between them. That is to say, if they contain the same semantics, they are connected together. If the labels Label between two nodes i and j are the same, we connect these two nodes with an edge, that is, e m1 (v i +v j ) = 1, where e m1 represents the edge under modality m 1 , otherwise e m1 (v i , v j ) = 0. A self-connecting edge pointing to the node itself is added to each point, that is, e m1 (v i , v i ) = 1. Through the above construction scheme, the subgraph G m1 or G m2, the message propagation iteration process is performed through a graph neural network to effectively encode the structural information within a rich modality in a set of training samples, which helps to better capture high-level semantic information from a global view and obtain discriminative representation embeddings. Taking two modalities as an example, based on a dual-branch graph convolutional network (GCN), each graph branch aggregates new information for samples within each modality by following the standard hierarchical propagation rules:

[0042] Equation 2:

[0043] where ε m1 is the adjacency matrix of the subgraph G m1 , represents the node representation at the l-th layer of the neural network for the node under modality m1; is the degree matrix, is the parameter of the scientific department of the l-th layer graph neural network under modality m1, and σ(·) represents the non-linear activation function. After passing through L layers of the graph neural network, the pseudo cross-modal representation

[0044] Multi-task stable training scheme:

[0045] Two tasks are constructed by utilizing the label classification loss and the mutual information loss between the pseudo cross-modal representation and the mapped cross-modal representation to jointly achieve stable model training; First, the pseudo cross-modal representation obtained previously is used for the label classification task:

[0046] Equation 3:

[0047] where, ||·|| F represents the Frobenius norm, that is, the square root of the sum of the absolute value squares of each element of the matrix, Y is the classification label, k represents the number of samples, and f class is the classifier model, which can be any classifier and is shared among each modality; Through Equation 3, data with the same semantic label has the same representation, thus increasing the discriminative ability and reducing the cross-modal differences at the same time;

[0048] The second task is constructed through the discriminative loss of different data under the pseudo cross-modal representation; Specifically, if the features of different modalities belong to the same category, we can also directly minimize the distance between them, and the definition of the distance loss L dis is:

[0049] Equation 4:

[0050] Integrating the above two tasks, the fused loss function is obtained:

[0051] Formula 5: L = L class + λ·L diss

[0052] where λ is the combination coefficient, representing the weight of task L diss .

[0053] S3. Based on the mapping feature representation of the video frame images and text obtained in step S2, use the Frobenius norm similarity algorithm to find the position of the first frame that meets the similarity threshold in the video library, and record the time sequence of this frame in the video where it is located.

[0054] Use the Frobenius norm similarity algorithm to calculate the similarity with each frame image in each video in the video library and filter according to the set similarity threshold. The similarity threshold set in the present invention is 0.3; after finding the position of the first frame in the video library, record the time sequence of this frame in the video where it is located; then calculate and record the time sequences of the remaining frames in this video. Finally, calculate the similarity of the video using the similarity of the frames. The calculation method is: the sum of the products of the similarity of each frame and the corresponding time sequence difference is divided by the number of frames; each node of the hadoop cluster can calculate the similarity of a video, and finally summarize the similarities to the reduce node; the total number of nodes used in the present invention is 1000.

[0055] S4. Use the similarity algorithm in step S3 to calculate the similarity of each video and sort according to the similarity of the videos, select the top ten videos as the final retrieval results, and submit them to the user for manual review.

[0056] In summary, the present invention starts from content-based video retrieval. According to the characteristics of massive and rapidly growing videos, it first preprocesses the massive video library based on Hadoop and establishes a massive video retrieval library. Then, it performs frame segmentation on the input video to be retrieved, splitting it into individual frames. A cross-modal retrieval method enhanced by graph structure under multi-task optimization is proposed. First, relationship subgraphs are constructed within each modality using classification labels, and graph neural networks are used for encoding to capture the structural information within the modality. Second, a multi-task is established using the label classification task and the discriminative loss task between different spatial representations to improve the training stability. Through this method, the ability to mine the data correlation between different modalities and within the same modality can be improved, the performance and training stability in the case of scarce training data can be enhanced, and more accurate cross-modal feature representation can be achieved. The similarity algorithm is used to find the position of the first frame in the video library and record the time sequence of this frame in the video where it is located. Then, the time sequences of the remaining frames in this video are calculated and recorded, and the weighted sum of the similarities of all frames is calculated with the time sequence difference from the first frame as the weight for each frame as the similarity of the video. Finally, the similarity of each video is calculated using the similarity of the frames. Then, the videos are sorted according to the similarity, and the top ten videos are selected as the final retrieval results for manual review by the user.

[0057] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be encompassed within the present invention. Any reference signs in the claims should not be regarded as limiting the claimed rights. In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A video retrieval method based on frame index and cross-modal representation, characterized in that, the method comprises the following steps: S1. Establish a massive video processing module based on Hadoop, and use the distributed computing power of the Hadoop cluster to preprocess the massive video library, and construct a video retrieval database; perform frame segmentation on the video segment to be retrieved, and segment it into individual picture frames; wherein, the process of preprocessing the massive video library is to extract and store the colors of video frames, and actually store the information obtained by standardizing and normalizing the RGB colors of video frames; S2. Use the proposed cross-modal representation method enhanced by graph structure under multi-task optimization to map the video frame images and text two-modal data obtained in step S1 to a unified cross-modal feature space, so that the two-modal data perform feature similarity matching calculation in the unified space; S3. Based on the mapped feature representations of the video frame images and text obtained in step S2, use the Frobenius norm similarity algorithm to find the position of the first frame that meets the similarity threshold in the video library, and record the time sequence of the frame in the video where it is located; then calculate and record the time sequences of the remaining frames in the video, and use the time sequence difference from the first frame as the weight of each frame to calculate the weighted sum of the similarities of all frames as the similarity of the video, and finally calculate the similarity of each video using the similarity of the frames; S4. Use the similarity algorithm in step S3 to calculate the similarity of each video, sort according to the similarity of the video, select the top ten videos as the final retrieval results, and submit them to the user for manual review.

2. The video retrieval method based on frame index and cross-modal representation according to claim 1, characterized in that, in step S2, in the proposed cross-modal representation method enhanced by graph structure under multi-task optimization, a structure encoding model under graph constraint and a multi-task stable training scheme are adopted. The structure encoding model under graph constraint is to establish a relational subgraph adjacency matrix within each modality using classification label information, and use a graph neural network model to encode information based on the constructed graph structure to obtain a pseudo cross-modal representation. The multi-task stable training scheme is to construct two tasks using the label classification loss and the discrimination loss of different data under the pseudo cross-modal representation to achieve stable model training.

Citation Information

Patent Citations

  • Method for realizing quick retrieval of mass videos

    CN104050247A

  • Cross-modal retrieval method for fusion graph convolution

    CN113536016A

  • Video retrieval method, computer storage medium, electronic equipment and computer program product

    CN114357248A