Weakly Supervised Video Localization Method and System Based on Multi-Level and Multi-Modal Alignment

Through the multi-layer multi-modal alignment method, combining the multi-layer feature alignment of video and text, the video clip positioning is optimized, and the problem of relying on data labeling deviation in the existing technology is solved, and video clip positioning is achieved with higher precision.

CN117093748BActive Publication Date: 2025-07-22SHANDONG UNIV +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310744391.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-21
Publication Date
2025-07-22
Estimated Expiration
2043-06-21

AI Technical Summary

Technical Problem

The existing natural language video positioning methods rely on data labeling deviations and cannot perform fine-grained predictions, resulting in insufficient positioning of video clips. Especially in weak supervision tasks, the lack of necessary supervision information makes it difficult to accurately obtain the timestamp of video clips.

Method used

The multi-layer multimodal alignment method is adopted to encode video and text information through the encoder, and feature fusion is used to generate masked text information through video clip features. The alignment reconstruction loss function between text and video is calculated, and the modal alignment reconstruction loss function at the noun object, action and event levels is combined to optimize the feature positioning of video clips.

Benefits of technology

It improves the accuracy of video clip positioning, suppresses data-dependent annotation deviation, and improves prediction accuracy, especially in weak supervision tasks, which can better locate the beginning and end positions of video clips.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117093748B_ABST
    Figure CN117093748B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of video localization, and provides a weakly supervised video localization method and system based on multi-level multi-modal alignment. The localization method includes encoding text and several segments of video information and fusing them, mapping to obtain several node pairs, masking the encoded representations of the video information with the same number as the node pairs to obtain video segment features; sequentially masking, encoding the text information and performing modal alignment with the video segment features; generating the masked text information, calculating the text and video alignment reconstruction loss function, and then combining the modal alignment reconstruction loss functions at the noun object level, action level and event level to obtain the total reconstruction loss function; using the video segment features located with the strategy of minimizing the total reconstruction loss function as the video segments corresponding to the text information. While considering the alignment of the overall text information with the video, it also performs more fine-grained modal alignment from levels such as verbs, nouns, and events, so as to achieve the purpose of improving the prediction accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of video localization, and in particular, relates to a weakly supervised video localization method and system based on multi-level multi-modal alignment. Background Art

[0002] The statements in this part only provide background technical information related to the present invention, and do not necessarily constitute prior art.

[0003] With the rapid development of intelligent devices and information technology, videos have become an important medium for people to share daily life, obtain information, entertain, and even obtain clues for crimes. Due to the leapfrog development of the Internet, mobile intelligent devices, and the communication field, combined with the rise of major video platforms, the number of videos that people come into contact with has also increased exponentially. The emergence of a huge video library has also increased people's demand for more intelligent and efficient video retrieval technologies. The existing video retrieval technologies mainly match through keywords, and find the videos that may meet the requirements by matching the keywords provided by the user with the video titles, tags, etc. There are many problems with this technology. For example, the lack of video tags or the irrelevance of the tags to the video content will seriously affect the retrieval results, and the form of keywords cannot flexibly describe the video content like natural language.

[0004] In recent years, deep learning has developed rapidly and achieved remarkable results in image recognition, object detection, natural language processing, etc., providing strong support for more complex downstream application tasks. Combining the above video retrieval requirements and the achievements of deep learning in multiple fields, researchers have proposed the natural language video localization task. As the name implies, the natural language video localization task is to find the segment in an untrimmed video that is most similar in semantics to the given natural language text. The natural language video localization technology has great application value and potential. For example, in the field of video recommendation, it can more accurately recommend the video segments that users want to watch; in the field of security, it can extract the most useful segments from the surveillance to assist criminal investigation, etc. In addition, the development of this technology is also beneficial to the development of other fields of video understanding. In recent years, this technology has received more and more attention from researchers and has achieved relatively rapid development. Especially in the fully supervised task, natural language video localization has achieved good results; however, the high cost of manual annotation in the fully supervised task limits its development in practical applications. In order to solve the problem of high manual annotation cost, people have proposed the weakly supervised natural language video localization task.

[0005] Weakly-supervised natural language video localization refers to using natural language descriptions to locate targets in videos in the absence of explicit video annotations or with only partial video annotations. Under weak supervision, the natural language video localization task requires the use of some weak annotation techniques, such as using visual processing techniques in the video to automatically identify targets and annotating them with natural language descriptions. These weakly annotated information can be imprecise or incomplete, but can still help the computer learn to recognize targets. Weakly-supervised natural language video localization methods are very useful in practical applications because they can reduce the workload of manual annotation and can use a large amount of unannotated data to train models. However, these methods also have some challenges, such as imprecise annotations, inaccurate target localization, etc. Therefore, further research is needed to improve the accuracy and robustness of weakly-supervised natural language video localization. The natural language video localization task needs to obtain and process features from two modalities, and it is difficult to correctly align the semantic information of different modalities during this process; in addition, how to obtain the timestamps of video segments from the aligned relationships is also a difficult task. Although existing methods have proposed many good strategies to address the above problems, these methods still have certain problems:

[0006] Existing natural language video localization methods often rely on data annotation biases rather than true multi-modal alignment reasoning. The methods tend to predict the start and end endpoint moments of the video. Existing methods mainly obtain features from the entire natural language text and then use these features for video segment localization. Since fine-grained prediction cannot be performed, the prediction is not precise enough. Summary of the Invention

[0007] In order to solve the technical problems existing in the above background art, the present invention provides a weakly-supervised video localization method and system based on multi-level multi-modal alignment, which, while considering the alignment of the overall text information and the video, also performs more fine-grained modal alignment from aspects such as verbs, nouns, events, etc., to achieve the purpose of improving the prediction accuracy.

[0008] In order to achieve the above object, the present invention adopts the following technical solutions:

[0009] The first aspect of the present invention provides a weakly-supervised video localization method based on multi-level multi-modal alignment.

[0010] A weakly-supervised video localization method based on multi-level multi-modal alignment, which includes:

[0011] Correspondingly encode and represent the text and several video information, and perform fusion to obtain a fusion feature;

[0012] Obtain a number of node pairs from the fusion feature through a preset mapping relationship, where each node pair represents the start position and the end position of the video;

[0013] Mask the encoded representations of video information with the same number of node pairs obtained using the above mapping to obtain video segment features; successively mask, encode the text information, and perform modality alignment with the video segment features.

[0014] Generate the masked text information from the video segment features, calculate the text-video alignment reconstruction loss function, and then combine the modality alignment reconstruction loss functions at the noun object level, action level, and event level to obtain the total reconstruction loss function; use the video segment features located with the strategy of minimizing the total reconstruction loss function as the video segment corresponding to the text information.

[0015] As an implementation manner, the length of the encoded representation of the text information is equal to that of the encoded representation of each video information.

[0016] As an implementation manner, the process of generating the fused information is as follows:

[0017] Concatenate the text encoded representation, the encoded representations of several video information segments, and a trainable preset marker to obtain a concatenated feature, and then perform feature interaction through Transformer to obtain a fused feature.

[0018] As an implementation manner, the text reconstruction loss function is the cross-entropy loss function, which is used to calculate the error between the reconstructed text and the original text.

[0019] As an implementation manner, in the process of calculating the modality alignment reconstruction loss function at the noun object level, given a positive example video and the corresponding text information, and a negative example video, randomly mask a noun object in the text information, and then generate the masked text information from the video segment features to calculate the modality alignment reconstruction loss function at the noun object level.

[0020] As an implementation manner, in the process of calculating the modality alignment reconstruction loss function at the action level, given a video and the corresponding text information, randomly mask a verb in the text information, and then generate the masked text information from the video segment features to calculate the modality alignment reconstruction loss function at the action level.

[0021] As an implementation, during the process of calculating the modal alignment reconstruction loss function at the event level, a positive example video and corresponding text information, as well as a negative example video, are selected. Some words in the text are randomly masked, and then the negative example video segment is spliced with the positive example video. Then, one of the spliced videos is randomly selected and the video segment features are obtained. At the same time, the video segment features of the positive example video are obtained. The video information is used to reconstruct the missing verb information in the text and calculate the modal alignment reconstruction loss function at the event level.

[0022] The second aspect of the present invention provides a weakly supervised video localization system based on multi-level multi-modal alignment.

[0023] A weakly supervised video localization system based on multi-level multi-modal alignment includes:

[0024] An encoding fusion module, which is used to perform corresponding encoding representations on the text and several video information segments, and perform fusion to obtain fusion features;

[0025] An information mapping module, which is used to obtain several node pairs through a preset mapping relationship for the fusion features, where each node pair represents the start position and end position of the video;

[0026] A modal alignment module, which is used to mask the encoded representations of the video information with the same number of node pairs obtained by the above mapping to obtain video segment features; mask, encode the text information in sequence, and perform modal alignment with the video segment features;

[0027] A video localization module, which is used to generate the masked text information through the video segment features, calculate the text and video alignment reconstruction loss function, and then combine the modal alignment reconstruction loss function at the noun object level, the modal alignment reconstruction loss function at the action level, and the modal alignment reconstruction loss function at the event level to obtain the total reconstruction loss function; use the video segment features located with the strategy of minimizing the total reconstruction loss function as the video segment corresponding to the text information.

[0028] The third aspect of the present invention provides a computer-readable storage medium.

[0029] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in the weakly supervised video localization method based on multi-level multi-modal alignment as described above.

[0030] The fourth aspect of the present invention provides a computer device.

[0031] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the above-mentioned weakly supervised video localization method based on multi-level multi-modal alignment.

[0032] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0033] The present invention generates masked text information through video segment features, calculates the text and video alignment reconstruction loss function, and then combines the modal alignment reconstruction loss function at the noun object level, the modal alignment reconstruction loss function at the action level, and the modal alignment reconstruction loss function at the event level to obtain the total reconstruction loss function. Finally, the video segment features located with the strategy of minimizing the total reconstruction loss function are used as the video segments corresponding to the text information. While considering the alignment of the overall text information with the video, the present invention also performs more fine-grained modal alignment from aspects such as verbs, nouns, and events, improving the prediction accuracy and, to a certain extent, suppressing the problem that previous methods rely on data annotation biases and tend to predict the beginning and ending positions of the video.

[0034] The advantages of the additional aspects of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The specification drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0036] Figure 1 is a flowchart of the weakly supervised video localization method based on multi-level multi-modal alignment according to an embodiment of the present invention;

[0037] Figure 2 is a schematic diagram of the modal alignment method at the concept level according to an embodiment of the present invention;

[0038] Figure 3 is a schematic diagram of the modal alignment method at the action level according to an embodiment of the present invention;

[0039] Figure 4 is a schematic diagram of the modal alignment method at the event level according to an embodiment of the present invention;

[0040] Figure 5 is a partial explanatory diagram of the dR@n, IoU@m evaluation index according to an embodiment of the present invention;

[0041] Figure 6 is a partial result example according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0042] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0043] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0044] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0045] Term Explanation:

[0046] Natural language video localization: Refers to the task of locating specific objects, scenes, actions, and other targets in a video according to natural language descriptions.

[0047] Weakly-supervised natural language video localization: Refers to the use of natural language descriptions to locate targets in a video without explicit video annotations or with only partial video annotations.

[0048] I3D: The full name of I3D is Inflated 3D Convnet, which is extended based on a two-dimensional convolutional neural network (CNN). The I3D model is trained on the pre-trained ImageNet dataset and learns the ability to extract spatial and temporal features from video data.

[0049] Glove: The full name is Global Vectors for Word Representation, which is a word embedding model for converting words into vector representations. The Glove model obtains the vector representation of each word through training on a large-scale corpus, and these vectors capture the semantic relationships between words.

[0050] Transformer: It is a neural network architecture for natural language processing (NLP) tasks, consisting of an encoder and a decoder. The encoder is used to convert the input text into a series of vector representations, and the decoder is used to decode these vectors into the corresponding output text. Generally speaking, Transformer is a neural network architecture based on the self-attention mechanism, which can better process sequence data and is widely used in natural language processing tasks.

[0051] FC: The fully connected layer is usually abbreviated as the FC layer. It is a common hierarchical structure in neural networks and is also known as the dense layer or the linear layer.

[0052] SOTA: SOTA is an abbreviation for "state-of-the-art", referring to the most advanced technology or method in a certain field at present. In the fields of machine learning and artificial intelligence, SOTA usually refers to the model or algorithm with the best performance on a specific task or dataset.

[0053] Existing natural language video localization methods rely too much on data annotation bias, do not truly achieve multimodal data alignment, and tend to predict the start and end positions of videos. At the same time, the feature extraction of natural language texts tends to focus on global information and does not make good use of local information, which limits the accuracy of video localization. Also, in weakly supervised tasks, there is a lack of necessary supervision information, that is, the training data does not have the start and end node labels of video segments. How to establish a suitable supervision mechanism is also a problem that needs to be solved.

[0054] To address the existing problems, this paper proposes a weakly supervised natural language video localization method based on multi-level modal alignment. This method first encodes video and text information through an encoder, then obtains the segment features of the video through data fusion and attention mechanisms, and then uses the video segment features to reverse-complement the masked text information to establish a supervision mechanism for method optimization. In this method, by reasonably designing the modal alignment module, problems existing in existing methods such as "tending to predict the start and end positions of videos" and "insufficient combination of global and local information of video and text information" can be well addressed.

[0055] Example 1

[0056] Refer to Figure 1 and combine with Figure 1, first, the text and video information are respectively obtained with their respective encoded representations through a text encoder and a video encoder, then the two types of features are passed through a data fusion module (Fusion Module) to obtain fused information, and then mapped through a fully connected layer (FC) to obtain N node pairs (the start and end positions of the video), and then N node pairs are used to mask N copies of the original encoded representations of the video to obtain video segment features P. At the same time, the input text information is masked, and then the masked text is encoded and modality-aligned with the above video segment features, and then through a reconstruction module, the masked text information is generated using the video segment features P, and the generation loss is calculated, thereby realizing the supervised optimization of the model. In addition, in the inference stage, a loss minimization strategy is adopted to select the video segment as the final result. This embodiment provides a weakly supervised video localization method based on multi-level multi-modal alignment, which specifically includes the following steps:

[0057] Step 1: Correspondingly encode and represent the text and several video information, and perform fusion to obtain fused features.

[0058] First, define the input text as where w i represents the i-th word, and n w represents the sentence length; define the input video as where f i represents the i-th sampled frame in the video, and n v is the number of sampled frames.

[0059] Encoder part: In this paper, the pre-trained models I3D and Glove are respectively used as the video and text encoders, and I3D and Glove are used to obtain the encoded representations of the video and the encoded representation of the text where and respectively represent the encoded features of the i-th sampled frame and the encoded features of the i-th word, as shown in Equation (1-1).

[0060]

[0061] In the above formula dv and dw respectively represent the lengths of the feature vectors, and the two are equal, and R represents the vector space.

[0062] It should be noted here that for video encoding, TSN, C3D, R2+1D can also be used; for text encoding, BERT, ELMo can also be used.

[0063] Among them:

[0064] TSN (Temporal Segment Networks) is a network architecture for video understanding tasks and is pre-trained on the large-scale video dataset Kinetics.

[0065] C3D (Convolutional 3D): C3D is a network architecture based on 3D convolution that performs 3D convolution operations on video data to capture dynamic features in the time dimension.

[0066] R2+1D (ResNet2+1D): R2+1D is a network architecture that combines 2D and 3D convolutions. It extracts spatial and temporal features by applying 2D convolution and 3D convolution separately in the time dimension.

[0067] BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained language model based on the Transformer model. Through large-scale unsupervised training, it has learned rich language representations, including word, phrase, and sentence-level representations.

[0068] ELMo is a context-based word embedding model proposed by Stanford University. It learns the vector representations of words from large-scale text data by using a bidirectional language model.

[0069] Data fusion part: Concatenate the video encoded features, text encoded features, and a trainable token to obtain feature F, and then perform feature interaction through Transformer to obtain the fused feature as shown in Equation (1-2).

[0070]

[0071] In the above formula, d h represents the dimension of the hidden fused feature.

[0072] Step 2: Obtain a number of node pairs from the fused feature through a preset mapping relationship, where each node pair represents the start position and end position of the video.

[0073] Use the fully connected layer FC to map the hidden fused feature h k into N pairs of time nodes N represents the number of candidate prediction time pairs.

[0074] Step 3: Use the node pairs obtained from the above mapping to mask the video information encoded representations with the same number as the node pairs to obtain video segment features; mask, encode the text information in turn, and perform modality alignment with the video segment features.

[0075] Video encoding feature mask part: Use the prediction time to generate a time series mask Denote the i-th mask sequence, and then fuse M and the video encoding feature F v Obtain the candidate video segment feature Denote the i-th video segment feature, as shown in Equation (1-3).

[0076] p i = m i ⊙ F v , m i ∈ M, p i ∈ P (1-3)

[0077] In the above formula, ⊙ represents the dot product operation of matrices.

[0078] Step 4: Generate the masked text information through the video segment feature, calculate the text and video alignment reconstruction loss function, and then combine the modal alignment reconstruction loss function at the noun object level, the modal alignment reconstruction loss function at the action level, and the modal alignment reconstruction loss function at the event level to obtain the total reconstruction loss function; The video segment feature located with the strategy of minimizing the total reconstruction loss function is used as the video segment corresponding to the text information.

[0079] Text mask completion part: Randomly mask a part of the words (about 1 / 3) of the text information, and then obtain the masked text encoding feature through the text encoder Then, through the text mask completion module, combine the acquisition of the video segment feature P and the text encoding feature Complete the completion of the text information mask to obtain Denote the predicted i-th word. This module is mainly based on Transformer-decoder, as shown in Equation (1-4).

[0080]

[0081] To better align the video segment feature P and the text encoding feature In this paper, the cross-entropy loss function is used to calculate the error between the reconstructed text and the original text for model optimization. The process of the loss function is shown in Equation (1-5).

[0082]

[0083] In the above formula Represents the conditional probability, that is, when the i-th word is Predict the i+1-th word as Probability.

[0084] However, the above-mentioned main text mask completion model is insufficiently designed, without considering temporal bias, and the multi-modal alignment is performed at a coarse-grained level. To better align visual and text information and address temporal bias during prediction, this paper also designs and adds three additional levels of modal alignment methods to improve the model performance.

[0085] Modal alignment method at the noun object level: To enhance the alignment between the noun concept information from the text and the video segments, and encourage the model to capture the video segments specific to the noun object, this paper introduces a noun concept level alignment module. This module helps reconstruct the masked noun concept using the original video segments, and also adopts a contrastive learning method to minimize the correlation between the masked noun concept and irrelevant video segments from other videos.

[0086] As Figure 2 shown, given a positive example video and corresponding text information, as well as a negative example video, randomly mask a noun object in the text information, and obtain the corresponding text encoding information through the text encoder. Meanwhile, use the same method as formulas (1-1, 1-2, 1-3) in the above text to obtain video segment features P1 and P2 for the positive and negative example videos. Subsequently, use the positive and negative example video segment features in the same way as formula (1-4) to generate the masked noun object in the text, and then calculate the reconstruction loss using formula (1-5). Then calculate the final loss through formula (1-6).

[0087]

[0088] Modal alignment method at the action level: Some verbs in the text information often contain some temporal information, and this part of the information often corresponds to some segments in the video. To better align this information, this paper adds an action level alignment module. This module can enhance the model's ability to reconstruct the masked verb information using video segments, and also adopts a contrastive learning method to minimize the correlation between the masked verb and the scrambled features of the same segment video.

[0089] As Figure 3 shown, given a video and corresponding text information, randomly mask a verb in the text information, and obtain the corresponding text encoding information through the text encoder. Meanwhile, process the video to obtain video segment information P1 and scrambled video segment information P2, and then use the video information to reconstruct the missing verb information in the text. Subsequently, calculate the reconstruction loss in the same way as the above process. Then calculate the final loss through formula (1-7).

[0090]

[0091] Modal alignment method at the event level: Video content usually involves event-level context. Existing methods only focus on aligning complete unrelated videos with text information, and tend to predict the start or end time points of the entire video or long video segments. To enable the model to focus on cross-modal semantic alignment at the video event level and mitigate the time bias problem, this paper introduces an event-level alignment module. This module obtains negative examples for predicting the start or end time points of the entire video by splicing positive and negative example videos. Similarly, by leveraging the idea of contrastive learning, it mitigates the time bias problem of the model and better realizes the alignment between video event segments and text information.

[0092] As Figure 4 shown, select a positive example video and corresponding text information, as well as a negative example video. Randomly mask some words in the text, and then splice the negative example video segment with the positive example video. There are three splicing methods: splicing the segment to the beginning, end, beginning and end of the positive example. Then randomly select one from the spliced videos and obtain the video segment feature P2, and at the same time obtain the video segment feature P1 of the positive example video. Subsequently, use the video information to reconstruct the missing verb information in the text and calculate the reconstruction loss. Finally, calculate the total loss L through Equation (1-8). e .

[0093]

[0094] Training and inference parts: The model proposed in this paper includes four loss components, namely the reconstruction loss (Equation 1-5) in the main prediction branch, which is used to enhance the reconstruction loss for aligning noun concepts from text and video segments (Equation 1-6), which is used to enhance the reconstruction loss for aligning verb information from text and video segment time (Equation 1-7) and the reconstruction loss for mitigating the time bias problem of model prediction (Equation 1-8).

[0095] During model training, this paper uses the loss L j Equation (1-9) to optimize the entire model.

[0096]

[0097] During model inference, this paper selects the optimal segment through the following Equation (1-10)

[0098]

[0099] In the above formula, the argmin method represents obtaining the index of the minimum value, which represents the index value.

[0100] To verify the effectiveness of the method, this paper conducts tests on the publicly available dataset Charades-CD. The data situation of this dataset is shown in Table 1. Charades-CD is a dataset re-split based on Charades-STA, mainly containing indoor activities and related text descriptions. This paper mainly selects 11,071 video-text pairs as the training set, 859 video-text pairs as the validation set, 823 video-text pairs as the in-distribution test set, and 3,375 video-text pairs as the out-of-distribution test set. Here, in-distribution means that the probability distribution of the data is the same as that of the training set.

[0101] Table 1. Situation of Charades-CD Dataset

[0102]

[0103] The model is mainly implemented using python + pytorch. The experimental environment is Ubuntu20.04 + RTX3090. The Transformer in the model is set to 3 layers, containing 4 attention heads. The hyperparameter N is set to 8, the learning rate is 0.0004, and the optimizer is selected as Adam.

[0104] This paper uses the improved "R@n, IoU@m" evaluation criterion as the model evaluation index. As shown in Equation (1-11), for each retrieval q i , first calculate the intersection IoU between the predicted time interval and the target time interval. If at least one of the top n predictions has an IoU greater than the IoU of the threshold m, then r(n, m, q i ) is 1, otherwise it is equal to 0. N q is the total number of all retrievals.

[0105]

[0106] In addition, in the above formula represents the prediction time and the real time the absolute distance between the boundaries, as Figure 5 shown. and are both normalized to the range (0, 1) by dividing the entire video length

[0107] The method proposed in this paper is compared with the existing fully supervised SOTA methods CTRL, 2D-TAN, SCDM, and the existing weakly supervised algorithms WSSL, CNM, and CPL. The dR@n and IoU@m results on the out-of-distribution test set are shown in Table 2, where n = 1 and m = 0.1, 0.3, 0.5, 0.7 are selected in this paper.

[0108] Table 2. Experimental results on the Charades-CD out-of-distribution test set

[0109]

[0110] As can be seen from the above table, the dR@n and IoU@m metrics of the method in this paper on the Charades-CD dataset are higher than those of other weakly supervised methods, and even higher than those of fully supervised methods when m = 0.1 and 0.3.

[0111] In addition, Figure 6 shows some result prediction cases. Here, the method in this paper and the CPL method are mainly compared. It can be seen that the method in this paper not only has higher prediction accuracy, but also suppresses the deviation at the beginning and end positions of the predicted video to a certain extent.

[0112] Embodiment 2

[0113] This embodiment provides a weakly supervised video localization system based on multi-level and multi-modal alignment, which specifically includes the following modules:

[0114] An encoding and fusion module, which is used to perform corresponding encoding representations on the text and several video information segments, and perform fusion to obtain fusion features;

[0115] An information mapping module, which is used to obtain several node pairs through a preset mapping relationship for the fusion features, where each node pair represents the start position and end position of the video;

[0116] A modal alignment module, which is used to mask the encoded representations of video information with the same number of node pairs as the obtained node pairs to obtain video segment features; mask, encode the text information in turn, and perform modal alignment with the video segment features;

[0117] A video localization module, which is used to generate the masked text information through the video segment features, calculate the text and video alignment reconstruction loss function, and then combine the modal alignment reconstruction loss function at the noun object level, the modal alignment reconstruction loss function at the action level, and the modal alignment reconstruction loss function at the event level to obtain the total reconstruction loss function; the video segment features located with the strategy of minimizing the total reconstruction loss function are used as the video segments corresponding to the text information.

[0118] It should be noted here that each module in this embodiment corresponds to each step in the first embodiment one by one, and the specific implementation process is the same, so it will not be repeated here.

[0119] Embodiment Three

[0120] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps in the above-mentioned weakly supervised video localization method based on multi-level multi-modal alignment.

[0121] Embodiment Four

[0122] This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the above-mentioned weakly supervised video localization method based on multi-level multi-modal alignment.

[0123] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories and optical memories, etc.) containing computer-usable program code.

[0124] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in one Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0125] Those of ordinary skill in the art can understand that all or part of the processes of implementing the above-mentioned method embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-mentioned method embodiments. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0126] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A weakly supervised video localization method based on multi-level and multi-modal alignment, characterized in that Including: Correspondingly encoding and representing text and several segments of video information, and fusing them to obtain fused features; Obtaining several pairs of nodes from the fused features through a preset mapping relationship, where each pair of nodes represents the start position and end position of a video; Using the pairs of nodes obtained by the above mapping to mask the encoding representations of video information with the same number as the pairs of nodes to obtain video segment features; successively masking, encoding the text information, and performing modality alignment with the video segment features; Generating the masked text information from the video segment features, calculating the text-video alignment reconstruction loss function, and then combining the modality alignment reconstruction loss function at the noun object level, the modality alignment reconstruction loss function at the action level, and the modality alignment reconstruction loss function at the event level to obtain the total reconstruction loss function; using the video segment features located with the strategy of minimizing the total reconstruction loss function as the video segment corresponding to the text information.

2. The weakly supervised video localization method based on multi-level and multi-modal alignment according to claim 1, wherein The length of the encoding representation of the text information is equal to the length of the encoding representation of each segment of video information.

3. The weakly supervised video localization method based on multi-level multi-modal alignment according to claim 1, wherein The process of generating fused information is as follows: Concatenating the text encoding representation, several segments of video information encoding representations, and a trainable preset token to obtain a concatenated feature, and then performing feature interaction through a Transformer to obtain fused features.

4. The weakly supervised video localization method based on multi-level and multi-modal alignment according to claim 1, characterized in that The text reconstruction loss function is a cross-entropy loss function used to calculate the error between the reconstructed text and the original text.

5. The weakly supervised video localization method based on multi-level and multi-modal alignment according to claim 1, wherein In the process of calculating the modality alignment reconstruction loss function at the noun object level, given a positive example video and the corresponding text information, and a negative example video, randomly masking a noun object in the text information, and then generating the masked text information from the video segment features to calculate the modality alignment reconstruction loss function at the noun object level.

6. The weakly supervised video localization method based on multi-level multi-modal alignment according to claim 1, wherein, In the process of calculating the modality alignment reconstruction loss function at the action level, given a video and the corresponding text information, randomly masking a verb in the text information, and then generating the masked text information from the video segment features to calculate the modality alignment reconstruction loss function at the action level.

7. The weakly supervised video localization method based on multi-level and multi-modal alignment according to claim 1, wherein In the process of calculating the modality alignment reconstruction loss function at the event level, select a positive example video and the corresponding text information, and a negative example video, randomly mask some words of the text, then concatenate the negative example video segment with the positive example video, and then randomly select one from the concatenated video and obtain the video segment feature, and at the same time obtain the video segment feature of the positive example video, and use the video information to reconstruct the missing verb information in the text and calculate the modality alignment reconstruction loss function at the event level.

8. A weakly supervised video localization system based on multi-level and multi-modal alignment, characterized in that Including: An encoding and fusion module for correspondingly encoding and representing text and several segments of video information, and fusing them to obtain fused features; An information mapping module for obtaining several pairs of nodes from the fused features through a preset mapping relationship, where each pair of nodes represents the start position and end position of a video; A modality alignment module for using the pairs of nodes obtained by the above mapping to mask the encoding representations of video information with the same number as the pairs of nodes to obtain video segment features; successively masking, encoding the text information, and performing modality alignment with the video segment features; A video localization module, which is used to generate the masked text information through video clip features, calculate the text and video alignment reconstruction loss function, and then combine the modal alignment reconstruction loss function at the noun object level, the modal alignment reconstruction loss function at the action level, and the modal alignment reconstruction loss function at the event level to obtain the total reconstruction loss function; the video clip features located with the strategy of minimizing the total reconstruction loss function are used as the video clips corresponding to the text information.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps in the weakly supervised video localization method based on multi-level multi-modal alignment described in any one of claims 1-7.

10. A computer device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the weakly supervised video localization method based on multi-level multi-modal alignment described in any one of claims 1-7.

Citation Information

Patent Citations

  • Weak supervision video representation learning method without aligned text in sequence video

    CN116052054A

  • Cross-modal video time period retrieval method based on weak supervision

    CN116091958A