Video clip positioning method
By building a positioning model combining video and audio information, the problem of inaccurate positioning caused by ignoring audio information in the prior art is solved, and more accurate video clip positioning is achieved.
Patent Information
- Application Number
- CN202411826819.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-12
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-12
AI Technical Summary
Existing video clip positioning methods ignore audio information when paying attention to visual information, resulting in inaccurate positioning results.
A video clip positioning method is proposed, by constructing a positioning model, which includes feature encoding module, fusion graph module, progressive dynamic interaction module and fragment positioning module. The model combines the video encoding module and the query encoding module, uses audio diagrams and video diagrams to be fused, and features are analyzed and positioned through multi-layer perceptron MLP and Transformer encoder.
By combining visual and audio information, the model can optimize the use of visual features through feedback of audio information in an environment where visual information is blurred or incomplete, thereby improving the accuracy of video clip positioning.
Smart Images

Figure CN119938981A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multimodal machine learning, and in particular to a video segment positioning method. Background Art
[0002] Video clip localization enables users to quickly and accurately find video clips of interest through natural language descriptions, thereby improving the efficiency and accuracy of video retrieval. It has a wide range of applications in scenarios such as video content management, video editing, and video surveillance. Combining natural language query technology with user preference information, personalized video recommendations can also be achieved to improve user satisfaction with video content and viewing experience. At the academic level, video clip localization has promoted cross-disciplinary research in the fields of natural language processing and computer vision, and promoted the development of multimodal machine learning. Researchers need to solve problems such as semantic understanding, video content analysis, and cross-modal interaction, which has promoted the deepening of theories in these fields.
[0003] At present, the commonly used video segment localization method without candidate video segment localization has the following implementation process: First, the input video and text query are preprocessed to extract the corresponding features. Then, the video features are cross-modally aligned or fused with the text features to establish a semantic association between the two. Then, the potential time period in the video is identified through a temporal convolutional network to capture the temporal information of the segment. Finally, based on the fused features, boundary regression is used to predict the start and end time of the target segment.
[0004] The defects of the above-mentioned prior art are: in the video clip positioning task, the visual information in the video is focused on, the accompanying audio information is ignored, and the factors considered are not comprehensive enough, which leads to inaccurate positioning results. Summary of the invention
[0005] Based on this, it is necessary to provide a video clip positioning method to address the above technical issues.
[0006] An embodiment of the present invention provides a video segment positioning method, including:
[0007] Constructing a positioning model, the positioning model includes a feature encoding module, a fusion graph module, a progressive dynamic interaction module and a fragment positioning module; the feature encoding module includes a video encoding module and a query encoding module, the fusion graph module includes an audio graph module, a video graph module and a graph fusion module, the progressive dynamic interaction module includes a shallow interaction module and multiple deep interaction modules, and the fragment positioning module includes a Transformer encoder and a multi-layer perceptron MLP;
[0008] Among them, the video encoding module is connected to the input end of the audio graph module and the video graph module, and the query encoding module is connected to the input end of the shallow interaction module; the output end of the audio graph module and the video graph module is connected to the input end of the graph fusion module, and the output end of the graph fusion module is also connected to the input end of the shallow interaction module; the output end of the shallow interaction module is connected to multiple deep interaction modules, and the output end of the previous deep interaction module is connected to the input end of the next deep interaction module; the output end of the last deep interaction module is connected to the input end of the Transformer encoder, and the output end of the Transformer encoder is connected to the multi-layer perceptron MLP;
[0009] The positioning model is trained through historical videos, the target video is input into the trained positioning model, the visual features and audio features of the target video are extracted through the video encoding module, and the query features of the target video are extracted through the query encoding module;
[0010] The audio graph module is used to construct an audio graph for the audio features, the video graph module is used to construct a video graph for the visual features, and then the audio graph and the video graph are fused through the graph fusion module to obtain the fused features;
[0011] The fusion features and query features are initially interacted through the shallow interaction module to obtain the initial fusion features. Then, the initial fusion features are deeply interacted through multiple deep interaction modules, and the output of the previous deep interaction module is used as the input of the next deep interaction module to obtain the final fusion features.
[0012] The final fusion module is modeled through the Transformer encoder to obtain context information; the context information is analyzed through the multi-layer perceptron MLP to obtain the start time and end time of the positioning segment (τ s ,τ e ).
[0013] Optionally, the visual features and audio features of the target video are extracted by a video encoding module, which specifically includes:
[0014] The original video V is divided into a series of segments with fixed lengths, and the visual features X of each segment are extracted using the pre-trained 3DCNN v ∈R T×d , use the pre-trained VGGish to extract the audio features X of each clip a ∈R T×d ; Its formula is:
[0015] X v / a =Conv1d(W seg (f v (V)));
[0016] Where f(·) represents 3DCNN, W seg represents the learnable fragment feature embedding matrix, Conv1d(·) represents 1D convolution;
[0017] When the input video is short and the number of clips is less than T, the missing parts are filled with zeros; position encoding is introduced in the features of each input clip and mapped to dimension d through 1D convolution.
[0018] Optionally, the query features of the target video are extracted by a query encoding module, which specifically includes:
[0019] For queries containing N words, word-level features are generated through word embedding and character embedding
[0020] Pass the word-level features to the self-weighted pooling layer to obtain the sentence-level features Q s ;
[0021] The sentence-level feature Q s It is concatenated with the n-1th semantic phrase feature and then projected into a mapping space to obtain the guide vector The calculation formula is:
[0022] g n =ReLU(W g ([W gq Q s ;e n-1 ]));
[0023] The guide vector Q g As query vector, semantic entity-level features are extracted through attention mechanism Get the nth semantic phrase feature e n , and its calculation formula is:
[0024] c n =softmax(w cT (tanh(W cg g n +W cq Q T )));
[0025]
[0026] in, W g ∈R d×2d and W gq ∈R d×d All are learnable weight matrices;
[0027] The semantic phrase feature e n As the query feature Qe .
[0028] Optionally, constructing an audio graph for the audio feature through an audio graph module, and constructing a video graph for the visual feature through a video graph module, which specifically includes:
[0029] In the audio graph, each node represents an audio clip, and each edge represents the similarity or correlation between audio clips; in the video graph, each node represents a video clip, and each edge represents the dependency between clips;
[0030] Cosine similarity is used to measure the similarity between fragments, and the cosine similarity score is used as the weight of the edge. The calculation formula is:
[0031]
[0032] Among them, cos(·) is the cosine similarity function.
[0033] Optionally, the audio graph and the video graph are fused through a graph fusion module, which specifically includes:
[0034] Calculating audio graph features by cosine similarity and video graph features The similarity between them is obtained by aligning the matrix M∈R T×T , each element of the alignment matrix represents the correspondence between the audio graph node and the video graph node, and its calculation formula is:
[0035]
[0036] Among them, sim(·) is the calculation and The node similarity score between them, softmax(·) is a column-by-column operation;
[0037] By aligning the matrix M Convert to The features of the nodes are weighted and combined according to the alignment matrix M, and the calculation formula is:
[0038]
[0039] X G a With X v G Fusion into fusion graph through gating mechanism The calculation formula is:
[0040]
[0041] Here, λ is a hyperparameter.
[0042] Optionally, the initial fusion features are deeply interacted through multiple deep interaction modules, which specifically include:
[0043] Initial input of the first layer of shallow interaction and Fusion graph and query feature Q e , the interaction process of the i-th layer is expressed as:
[0044]
[0045] The feedback mechanism of the i-th layer is expressed as:
[0046]
[0047] Among them, f i-1 is the fusion feature of the i-1th layer, which is used as the fusion graph and query feature Q e Additional input to update the features, G va (·) and G q (·) are fusion graphs and query feature Q e The feature update function.
[0048] Optionally, get the start and end time of the positioning segment (τ s ,τ e ), which specifically include:
[0049] O = attn(Conv1d(X vaq ));
[0050] τs,τe=MLP(FFN(O));
[0051] Among them, attn(·) represents a multi-head self-attention layer, Conv1d(·) represents a channel-separable 1D convolution, and FFN(·) represents a feed-forward network.
[0052] Compared with the prior art, the video clip positioning method provided by the embodiment of the present invention has the following beneficial effects:
[0053] The present invention accurately extracts the visual features, audio features and query features of the target video through the video encoding module and the query encoding module, and uses the audio graph module and the video graph module to construct intuitive audio graphs and video graphs for the audio features and visual features respectively, and then fuses the two through the graph fusion module to form fusion features. In this process, not only the audio information is closely combined with the visual information, but also the possible noise interference between the two is fully considered, so that the model can optimize the use of visual features through the feedback of audio information in complex environments where the visual information is blurred or incomplete. Therefore, the present invention takes comprehensive factors into consideration and the positioning result of the video is more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 A diagram of the overall model architecture of a video clip positioning method provided in an embodiment;
[0055] Figure 2 A visual example diagram of a video clip positioning method provided in an embodiment. DETAILED DESCRIPTION
[0056] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0057] In one embodiment, a video clip positioning method is provided, the method comprising:
[0058] 1 Overall model
[0059] Construct a positioning model, such as Figure 1 As shown, the positioning model includes a feature encoding module, a fusion graph module, a progressive dynamic interaction module and a fragment positioning module. The feature encoding module includes a video encoding module and a query encoding module, the fusion graph module includes an audio graph module, a video graph module and a graph fusion module, the progressive dynamic interaction module includes a shallow interaction module and multiple deep interaction modules, and the fragment positioning module includes a Transformer encoder and a multi-layer perceptron MLP.
[0060] Among them, the video encoding module is connected to the input of the audio graph module and the video graph module, and the query encoding module is connected to the input of the shallow interaction module; the output of the audio graph module and the video graph module are connected to the input of the graph fusion module, and the output of the graph fusion module is also connected to the input of the shallow interaction module; the output of the shallow interaction module is connected to multiple deep interaction modules, and the output of the previous deep interaction module is connected to the input of the next deep interaction module; the output of the last deep interaction module is connected to the input of the Transformer encoder, and the output of the Transformer encoder is connected to the multi-layer perceptron MLP.
[0061] 2 A video segment localization technology based on progressive dynamic interaction with audio supplement (Progressive Dynamic Interaction Network with Audio Supplement for Video Moment Localization, PDIN) aims to solve the problem of complementarity and model flexibility of audio and visual information. It mainly consists of four parts: (1) The feature encoding module extracts visual features, audio features and query features; (2) The fusion graph module complements and fuses the video graph and the audio graph; (3) The progressive interaction module captures the semantic relationship between the query and the video; (4) The segment localization module determines the start time and end time (τ s ,τ e ).
[0062] The localization model is trained through historical videos, and the target video is input into the trained localization model. The visual features and audio features of the target video are extracted through the video encoding module, and the query features of the target video are extracted through the query encoding module. The audio graph module is used to construct an audio graph for the audio features, and the video graph module is used to construct a video graph for the visual features. The audio graph and the video graph are then fused through the graph fusion module to obtain fused features. The fused features and query features are preliminarily interacted with each other through the shallow interaction module to obtain the initial fused features. The initial fused features are then deeply interacted with through multiple deep interaction modules, and the output of the previous deep interaction module is used as the input of the next deep interaction module to obtain the final fused features. The final fusion module is modeled through the Transformer encoder to obtain context information. The context information is analyzed through the multi-layer perceptron MLP to obtain the start time and end time of the localization segment (τ s ,τ e ).
[0063] The specific implementation process includes:
[0064] (1) Data preprocessing
[0065] For a given original video, it is divided into two modalities V = {X v ,X a},in and If the number of segments in the video is less than T, the missing parts are padded with zeros. For a given query text, the words in the query text are first represented as word embedding vectors using the pre-trained GloVe where q n is the nth word embedding vector, and N represents the length of the query text.
[0066] (2) Feature Coding
[0067] A. Video Coding
[0068] In order to extract the visual and audio features of the video, the original video V is segmented into a series of segments with fixed lengths, and the pre-trained 3DCNN is used to extract the visual features X of each segment. v ∈R T×d , use the pre-trained VGGish to extract the audio features X of each clip a ∈R T×d ; Its formula is:
[0069] X v / a =Conv1d(W seg (f v (V)));
[0070] Where f(·) represents 3DCNN, W seg Denotes a learnable segment feature embedding matrix, and Conv1d(·) denotes a 1D convolution. When the input video is short and the number of segments is less than T, the missing parts are filled with zeros; position encoding is introduced in the features of each input segment and mapped to dimension d through 1D convolution.
[0071] B. Query Code
[0072] For queries containing N words, we first use pre-trained GloVe to generate an embedding representation of the query. Then, we perform multi-granular encoding on the query to extract information at different levels and scales, thereby improving the performance of the model in complex scenarios.
[0073] Specifically, for a query containing N words, word-level features are generated through word embedding and character embedding The word-level features are then passed to the self-weighted pooling layer to obtain the sentence-level features Q s In order to extract N semantic phrase features from the query, the sentence-level feature Q s It is concatenated with the n-1th semantic phrase feature to convert the sentence-level feature Q sIt is concatenated with the n-1th semantic phrase feature and then projected into a mapping space to obtain the guide vector The calculation formula is:
[0074] g n =ReLU(W g ([W gq Q s ;e n-1 ]));
[0075] The guide vector Q g As query vector, semantic entity-level features are extracted through attention mechanism Get the nth semantic phrase feature e n , and its calculation formula is:
[0076] c n =softmax(w cT (tanh(W cg g n +W cq Q T )));
[0077]
[0078] in, W g ∈R d×2d and W gq ∈R d×d are all learnable weight matrices. n As the query feature Q e .
[0079] (3) Fusion graph
[0080] In order to deeply understand and process audio and visual features, the obtained audio feature X a and visual feature X v Construct audio graph and video graph respectively. In the audio graph, each node represents an audio segment, and each edge represents the similarity or correlation between audio segments; in the video graph, each node represents a video segment, and each edge represents the dependency between segments. In this way, the complex structures and relationships in audio and video data can be effectively analyzed and processed. Finally, the audio graph features are obtained. and video graph features
[0081]
[0082] Specifically, in order to construct the audio and visual features into a graph, cosine similarity is used to measure the similarity between segments, and the cosine similarity score is used as the weight of the edge, which is calculated as:
[0083]
[0084] Where cos(·) represents the cosine similarity function.
[0085] However, since the value range of cosine similarity is between -1 and 1, and the edge weight should be non-negative, the edge weights less than 0 are removed to ensure the rationality of the graph.
[0086] Audio information contains rich information, such as background music, environmental sound effects, etc., which can reflect the emotions, plot development and rhythm of the video content. When visual information is missing or unrecognizable, audio information provides discriminative clues as a supplement. Through the complementarity of audio and visual information, the model can better capture and understand the dynamic changes of events, character emotions and background environment, and improve the model's understanding and recognition ability of the scene. Therefore, a graph fusion method is used to supplement and fuse the video graph and audio graph to enhance the model's understanding ability in complex scenes.
[0087] Specifically, the audio graph features are calculated by cosine similarity and video graph features The similarity between them is obtained, thus obtaining the alignment matrix M∈R T×T , each element of the alignment matrix represents the correspondence between the audio graph node and the video graph node, and its calculation formula is:
[0088]
[0089] Where sim(·) represents the calculation and The node similarity score between them, softmax(·) is a column-by-column operation.
[0090] Then, the alignment matrix M is used to Convert to The features of the nodes are weighted and combined according to the alignment matrix M. Specifically, The node characteristics in are The weighted average of all node features in , where the weight is determined by the value corresponding to the alignment matrix M. The calculation formula is:
[0091]
[0092]
[0093] Afterwards, and Fusion into fusion graph through gating mechanism It not only integrates visual and auditory information, but also optimizes the expression of this information, making it suitable for more complex analysis tasks. The calculation formula is:
[0094]
[0095] where λ is a hyperparameter.
[0096] (4) Progressive dynamic interaction
[0097] In order to improve the effect of cross-modal interaction, a progressive dynamic interaction method is proposed. This method first performs preliminary interaction on each modal data at the shallow level, and then gradually introduces feedback information from previous interactions at the deep level, so that the model can optimize multimodal representation in the process of continuous learning. The interaction of each layer depends on the interaction results of the previous layer, thus forming a progressive interaction process. Each step in this process adaptively adjusts the fusion weights and methods to capture the most relevant information between modalities.
[0098] Specifically, the purpose of shallow interaction is to capture the basic associations between modalities and ensure the effect of preliminary fusion. The dot product attention mechanism is used to perform preliminary interactions on video and query features, and the interaction weights are generated by calculating the similarity between them. The output of shallow interaction is passed to the deep interaction module as the initial fusion feature. In the deep interaction stage, the shallow interaction mechanism is continued, and the reverse connection and feedback mechanism is introduced at the same time. The features obtained from the previous round of interaction are fed back to the input layer of the model as context vectors. Through this process, the attention weights are dynamically adjusted to ensure that the interaction of each layer can accurately capture the association between modalities. In the layer-by-layer adjustment process, the model adaptively determines the focus of each layer of interaction, so as to more accurately locate the target segment.
[0099] Initial input of the first layer of shallow interaction and are the fusion graph and query feature Q respectively e , where the interaction process of the i-th layer can be expressed as:
[0100]
[0101] Among them, W vaq , W va and W q are all learnable weight matrices.
[0102] The feedback mechanism of the i-th layer can be expressed as:
[0103]
[0104] Among them, f i-1 is the fusion feature of the i-1th layer, which is used as the fusion graph and query feature Qe Additional input to update the features, G va (·) and G q (·) are fusion graphs and query feature Q e The feature update function.
[0105] (5) Fragment positioning
[0106] After obtaining the fusion feature X vaq After that, it is input into a Transformer encoder, which models the features through the self-attention mechanism to capture long-distance dependencies and contextual information. Finally, the fused features are processed by MLP to predict the start time and end time of the target segment (τ s ,τ e ), which specifically include:
[0107] O = attn(Conv1d(X vaq ));
[0108] τs,τe=MLP(FFN(O));
[0109] Among them, attn(·) represents a multi-head self-attention layer, Conv1d(·) represents a channel-separable 1D convolution, and FFN(·) represents a feed-forward network.
[0110] As an example, comparative experiments of the present invention are provided.
[0111] 1. Training
[0112] Two loss functions are used to train the network during training, namely segment localization loss and regularization loss. The segment localization loss is used to guide the model to accurately locate the video segment corresponding to the query.
[0113] L reg =L1(τ' s -τ s )+L1(τ' e -τ e );
[0114] Where L1 represents the SmoothL1 distance, (τ s ',τ e ') indicates the start time and end time of the real video segment.
[0115] In order to extract more diverse semantic entity-level features, regularization technology is used as loss, so that the model can better capture the association between semantic entities during training and further improve the generalization ability of the model:
[0116]
[0117] in Is to generate semantic level entity features Q e The attention weight when ||·|| represents the Frobenius norm, η is a hyperparameter, represents the identity matrix. The total loss function is L reg and L e Addition of:
[0118]
[0119] 2. Experimental Setup
[0120] To verify the video clip localization performance of PDIN, ActivityNet Captions and Charades-STA datasets were selected for experiments and analysis. In order to make a fair comparison with previous work, for the ActivityNet Captions dataset, a pre-trained C3D network was used to extract visual features, and VGGish was used to extract audio features. For the Charades-STA dataset, a pre-trained I3D network was used to extract visual features, and PANN was used to extract audio features. The feature dimension d was set to 512. For query encoding, all query words were converted to lowercase and labeled, and GloVe with a dimension of 300 was used to generate an embedded representation of the query. The initial learning rates of Charades-STA and ActivityNet Captions were set to 0.00015 and 0.0005, respectively. During the experiment, all experiments were performed on an NVIDIA GeForce RTX 4090 graphics card, in the Pytorch1.12 environment, and with the Adam optimizer for parameter optimization.
[0121] 3. Ablation experiment
[0122] In order to evaluate the effectiveness of different modules in PDIN, we conducted in-depth ablation studies on two challenging datasets, ActivityNet Captions and Charades-STA, and analyzed the contribution of different components to the model to verify its effectiveness. Specifically, each time a module is removed, the ablation variant of the PDIN model is generated as follows:
[0123] ①: The graph fusion module is removed, and the supplement of audio features to visual features is not considered. Only visual features and query features are used for progressive dynamic interaction.
[0124] ②: The audio features are removed, and the visual features are converted into a graph structure, which interacts with the query features in a progressive and dynamic manner.
[0125] ③: The progressive dynamic interaction module is removed. After passing through the graph fusion module, the video features and query features are directly fused through splicing.
[0126] ④: The graph fusion module and the progressive dynamic interaction module are removed, and the audio features and visual features are directly added and then directly fused with the query features through splicing.
[0127] ⑤: Video clip localization network for progressive dynamic interaction with complete audio supplementation.
[0128] Table 1 shows the results of the ablation experiment. According to the results, on the ActivityNet Captions and Charades-STA datasets, the complete model outperforms other variant models in all evaluation indicators, which shows that the audio features, graph fusion module, and progressive dynamic interaction module play a positive role in the video segment localization task.
[0129] First, by comparing model ⑤ with model ①, we can observe that the R@1 and IoU=0.7 indicators of the two datasets are improved by 1.43% and 2.29% respectively. This finding shows that the graph fusion module plays an important role in cross-modal information fusion. By fusing visual features and audio features, the module enhances the complementarity of cross-modal information, thereby improving the overall performance of the model.
[0130] Secondly, compared with model ①, model ② improves by 0.47% and 1.2% on the R@1, IoU=0.7 indicators of the two datasets, respectively. However, compared with model ⑤, it decreases by 0.96% and 1.13% on the two datasets, respectively. This shows that audio features can indeed provide meaningful complementary information to visual features, making the overall performance more superior and helping to better capture the correlation between video and query.
[0131] In addition, comparing model ⑤ with model ③, the R@1, IoU=0.7 indicators of the two datasets were improved by 1.12% and 2.18% respectively. This finding shows that the progressive dynamic interaction module plays an important role in improving the precise matching and semantic association of multimodal information. This module gradually enhances the matching degree between video and query by deepening multimodal interaction layer by layer and optimizing feature fusion, thus significantly improving the performance of the model.
[0132] Finally, the comparison between Model ④ and Model ⑤ shows that the progressive dynamic interaction module and the graph fusion module are crucial for retrieval. They are improved by 1.8% and 3.23% on the two datasets respectively. This is because the progressive dynamic interaction module can better capture the deep semantic relationship between the query and the video, and the graph fusion module effectively combines visual and audio features, thereby enhancing the expressive power of multimodal features.
[0133] Table 1 Evaluation results of ablation test indicators
[0134]
[0135] 4. Qualitative results analysis
[0136] like Figure 2 As shown in the figure, PDIN shows excellent performance and can accurately retrieve the moments most relevant to the language query, even if these moments are very similar visually. The qualitative results of model ①, model ③ and the complete model mentioned in the ablation experiment are also compared. It can be observed that there are large errors in model ① and model ③. After removing the graph fusion module, model ① no longer considers the complement of audio features to visual features, resulting in the inability to accurately distinguish the differences between different objects in the video. Model ③ has insufficient semantic understanding due to the removal of the progressive dynamic interaction module, making it difficult to effectively capture complex cross-modal associations. Overall, the experimental results show that the joint application of the progressive dynamic interaction module and the graph fusion module is crucial for fragment retrieval. PDIN not only performs well in various scenarios, but also maintains high accuracy when dealing with moments that are visually similar but semantically different.
[0137] The above-mentioned embodiments only express several implementation methods of the present invention, and the description is relatively specific and detailed, but it cannot be understood as limiting the scope of the invention patent. It should be pointed out that for ordinary technicians in this field, several modifications and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention.
Claims
1. A video segment positioning method, characterized in that: include: Constructing a positioning model, the positioning model includes a feature encoding module, a fusion graph module, a progressive dynamic interaction module and a fragment positioning module; the feature encoding module includes a video encoding module and a query encoding module, the fusion graph module includes an audio graph module, a video graph module and a graph fusion module, the progressive dynamic interaction module includes a shallow interaction module and multiple deep interaction modules, and the fragment positioning module includes a Transformer encoder and a multi-layer perceptron MLP; Among them, the video encoding module is connected to the input end of the audio graph module and the video graph module, and the query encoding module is connected to the input end of the shallow interaction module; the output end of the audio graph module and the video graph module is connected to the input end of the graph fusion module, and the output end of the graph fusion module is also connected to the input end of the shallow interaction module; the output end of the shallow interaction module is connected to multiple deep interaction modules, and the output end of the previous deep interaction module is connected to the input end of the next deep interaction module; the output end of the last deep interaction module is connected to the input end of the Transformer encoder, and the output end of the Transformer encoder is connected to the multi-layer perceptron MLP; The positioning model is trained through historical videos, the target video is input into the trained positioning model, the visual features and audio features of the target video are extracted through the video encoding module, and the query features of the target video are extracted through the query encoding module; The audio graph module is used to construct an audio graph for the audio features, the video graph module is used to construct a video graph for the visual features, and then the audio graph and the video graph are fused through the graph fusion module to obtain the fused features; The fusion features and query features are initially interacted through the shallow interaction module to obtain the initial fusion features. Then, the initial fusion features are deeply interacted through multiple deep interaction modules, and the output of the previous deep interaction module is used as the input of the next deep interaction module to obtain the final fusion features. The final fusion module is modeled through the Transformer encoder to obtain context information; the context information is analyzed through the multi-layer perceptron MLP to obtain the start time and end time of the positioning segment (τ s ,τ e ).
2. A video segment positioning method as claimed in claim 1, characterized in that: The extracting of visual features and audio features of the target video through the video encoding module specifically includes: The original video V is divided into a series of segments with fixed lengths, and the visual features X of each segment are extracted using the pre-trained 3DCNN v ∈R T×d , use the pre-trained VGGish to extract the audio features X of each clip a ∈R T×d ; Its formula is: X v / a =Conv1d(W seg (f v (V))); Where f(·) represents 3DCNN, W seg represents the learnable fragment feature embedding matrix, Conv1d(·) represents 1D convolution; When the input video is short and the number of clips is less than T, the missing parts are filled with zeros; position encoding is introduced in the features of each input clip and mapped to dimension d through 1D convolution.
3. A video segment positioning method as claimed in claim 1, characterized in that: The query feature of the target video is extracted by the query encoding module, which specifically includes: For queries containing N words, word-level features are generated through word embedding and character embedding Pass the word-level features to the self-weighted pooling layer to obtain the sentence-level features Q s ; The sentence-level feature Q s It is concatenated with the n-1th semantic phrase feature and then projected into a mapping space to obtain the guide vector The calculation formula is: g n =ReLU(W g ([IN gq Q s ;e n-1 ])); The guide vector Q g As query vector, semantic entity-level features are extracted through attention mechanism Get the nth semantic phrase feature e n , and its calculation formula is: c n =softmax(w cT (tanh(W cg g n +W cq Q T ))); in, and W gq ∈R d×d All are learnable weight matrices; The semantic phrase feature e n As the query feature Q e .
4. A video segment positioning method as claimed in claim 1, characterized in that: The step of constructing an audio graph for audio features by using an audio graph module and constructing a video graph for visual features by using a video graph module specifically includes: In the audio graph, each node represents an audio clip, and each edge represents the similarity or correlation between audio clips; in the video graph, each node represents a video clip, and each edge represents the dependency between clips; Cosine similarity is used to measure the similarity between fragments, and the cosine similarity score is used as the weight of the edge. The calculation formula is: Among them, cos(·) is the cosine similarity function.
5. A video segment positioning method as claimed in claim 1, characterized in that: The step of fusing the audio graph and the video graph through the graph fusion module specifically includes: Calculate audio graph feature X by cosine similarity G a and video graph feature X v G The similarity between them is obtained by aligning the matrix M∈R T×T , each element of the alignment matrix represents the correspondence between the audio graph node and the video graph node, and its calculation formula is: Among them, sim(·) is the calculation and The node similarity score between them, softmax(·) is a column-by-column operation; By aligning the matrix M Convert to The features of the nodes are weighted and combined according to the alignment matrix M, and the calculation formula is: Will and Fusion into fusion graph through gating mechanism The calculation formula is: Here, λ is a hyperparameter.
6. A video segment positioning method as claimed in claim 1, characterized in that: The deep interaction of the initial fusion features through multiple deep interaction modules specifically includes: Initial input of the first layer of shallow interaction and Fusion graph and query feature Q e , the interaction process of the i-th layer is expressed as: Among them, W vaq , W va and W q All are learnable weight matrices; The feedback mechanism of the i-th layer is expressed as: Among them, f i-1 is the fusion feature of the i-1th layer, which is used as the fusion graph and query feature Q e Additional input to update the features, G va (·) and G q (·) are fusion graphs and query feature Q e The feature update function.
7. A video segment positioning method as claimed in claim 1, characterized in that: The start time and end time (τ s ,τ e ), which specifically include: O=attn(Conv1d(X vaq )); τs,τe=MLP(FFN(O)); Among them, attn(·) represents a multi-head self-attention layer, Conv1d(·) represents a channel-separable 1D convolution, FFN(·) represents a feed-forward network, and X vaq Represents fusion features.
Citation Information
Patent Citations
Cross-modal time domain video positioning method under text segment question and answer framework
CN114925232A
Video clip positioning system based on space-time semantic decomposition
CN115309939A
Multi-modal data fusion method and system and storage medium
CN115545093A
Video positioning method and device
CN117453949A
Cross-modal sensitive information identification method
CN117668292A