A progressive multi-modal semantic alignment teaching video content extraction method
By employing a progressive multimodal semantic alignment method, which combines visual and textual feature extraction with a two-layer attention and graph attention module, the teaching video summarization is optimized. This addresses the problem of insufficient long-term, multimodal semantic alignment in teaching videos, generating high-quality, fine-grained teaching video summaries and improving learning efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG UNIV OF TECH
- Filing Date
- 2025-09-25
- Publication Date
- 2026-05-15
AI Technical Summary
Existing multimodal video content extraction methods lack optimization for the characteristics of long time sequence, strong logic and redundancy in teaching video scenarios. They also lack multimodal semantic alignment capabilities, making it difficult to achieve fine-grained knowledge point extraction, resulting in coarse-grained summarization results.
A progressive multimodal semantic alignment method is adopted. By combining visual feature extraction and text feature extraction models with a two-layer attention module and a graph attention module, coarse-grained and fine-grained video frame summaries and text summaries are generated. Cosine similarity and graph attention network are used to optimize model parameters to achieve cross-modal semantic alignment.
It improves the accuracy and coherence of teaching video summaries, generating high-quality summaries with complete structure and highlighting key points, helping users quickly acquire core knowledge and improve learning efficiency.
Smart Images

Figure CN121214299B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of video summarization technology, specifically relating to a method for extracting teaching video content through progressive multimodal semantic alignment. Background Technology
[0002] With the rapid development of internet technology and the widespread adoption of online education, educational videos have become an important way for learners to acquire knowledge. However, educational videos generally suffer from being too long, containing redundant content and a large amount of information, making it difficult for learners to extract core knowledge points in a short time, thus affecting learning efficiency.
[0003] Existing technologies have proposed methods for multimodal content extraction. For example, Chinese patent application CN118708716A discloses a method for generating multimodal summaries based on semantic planning guidance. This method can generate relatively accurate multimodal summaries from news data by constructing textual semantic units and combining them with a large language model. However, this method is mainly geared towards the news field and is difficult to effectively model the complex semantic relationships present in educational videos, thus limiting its application in the extraction of educational video content.
[0004] Furthermore, Chinese patent application CN118520417A discloses a multimodal summarization method based on visually enhanced entity-level interaction networks. By introducing cross-modal entity interaction and visual enhancement mechanisms, it improves the depth and expressive power of image-text fusion. However, this method focuses more on entity and object-level interaction modeling and still has shortcomings in addressing the long temporal dependencies and multimodal semantic alignment problems commonly found in teaching videos.
[0005] In summary, existing multimodal video content extraction methods still have the following technical shortcomings in the context of teaching videos: 1) They lack optimization for the content characteristics of teaching videos, which are characterized by "long time sequence, strong logic, and a lot of redundancy"; 2) Their multimodal semantic alignment capability is insufficient, making it difficult to simultaneously take into account the accurate matching of video frames and text semantics; 3) The summarization results often remain at the coarse-grained level, making it difficult to achieve fine-grained extraction of knowledge points. Summary of the Invention
[0006] The purpose of this invention is to address the above-mentioned problems by proposing a progressive multimodal semantic alignment method for extracting teaching video content, which can effectively utilize the semantic associations between multiple modalities to generate more accurate and fluent teaching video summary content.
[0007] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0008] The present invention proposes a method for extracting teaching video content through progressive multimodal semantic alignment, comprising the following steps:
[0009] S1. Sample the educational video to obtain the video frame sequence and the corresponding sentence sequence;
[0010] S2. Input the video frame sequence into the visual feature extraction model to obtain the video frame feature representation, and input the sentence sequence into the text feature extraction model to obtain the sentence feature representation;
[0011] S3. Establish a progressive multimodal semantic alignment model to perform the following operations:
[0012] S31. Use a two-layer attention module to extract features from the video frame feature representation and the sentence feature representation respectively, and generate a coarse-grained video frame summary set and a coarse-grained text summary set accordingly.
[0013] S32. Based on the coarse-grained video frame summary set and the coarse-grained text summary set, construct the corresponding video frame graph set and sentence graph set;
[0014] S33. The graph attention module is used to extract features from the video frame graph set and the sentence graph set respectively, and corresponding coarse-grained video frame graph embedding representation set and coarse-grained sentence embedding representation set are obtained. The graph attention module includes three layers of graph attention network connected in sequence.
[0015] S34. Based on the coarse-grained video frame embedding representation set and the coarse-grained sentence embedding representation set, respectively, extract the fine-grained video frame summary set and the fine-grained text summary set according to the cosine similarity principle.
[0016] S4. Input the video frame sequence, sentence sequence, video frame feature representation, and sentence feature representation obtained from the teaching video to be refined into the progressive multimodal semantic alignment model to obtain a fine-grained video frame summary set and a fine-grained text summary set, which are the summary contents extracted from the teaching video.
[0017] Preferably, the visual feature extraction model is one of the pre-trained ResNet model, DenseNet model, GoogleNet model, or Vision Transformer model; the text feature extraction model is one of the pre-trained RoBERTa model, BERT model, GPT model, XLNet model, or a language model based on the Transformer architecture.
[0018] Preferably, a dual-layer attention module is used to extract features from the video frame feature representation and the sentence feature representation respectively, thereby generating a coarse-grained video frame summary set and a coarse-grained text summary set, as follows:
[0019] S311. Constructing the video modality attention mask matrix Text modality attention mask matrix and cross-modal mask matrix In this context, an element in each mask matrix being 1 indicates that attention operations are performed; otherwise, attention operations are not performed. N represents the total number of video frames in the video frame sequence. The total number of sentences in the sentence sequence;
[0020] S312. Establish a self-attention module consisting of a parallel first self-attention layer and a second self-attention layer. Input the video frame feature representation and the video modality attention mask matrix into the first self-attention layer, and input the sentence feature representation and the text modality attention mask matrix into the second self-attention layer to obtain the enhanced video frame feature representation and the enhanced sentence feature representation respectively.
[0021] S313. Input the enhanced video frame feature representation, enhanced sentence feature representation, and cross-modal attention mask matrix into the cross-modal attention module to obtain the cross-modal video frame feature representation and the cross-modal sentence feature representation;
[0022] S314. Input the cross-modal video frame feature representation and the cross-modal sentence feature representation into the feedforward neural network respectively to obtain the aligned frame feature representation and the aligned sentence feature representation accordingly;
[0023] S315. Input the aligned frame feature representation and the aligned sentence feature representation into the score prediction model to obtain the video frame importance score and sentence importance score respectively. The score prediction model is a linear layer.
[0024] S316. Select the top-scoring video frames from the importance scores. The top-scoring video frames and sentences in importance scores Each sentence corresponds to a coarse-grained video frame summary set and a coarse-grained text summary set. , .
[0025] Preferably, the feedforward neural network includes a linear layer, a GeLU activation function, and another linear layer connected in sequence.
[0026] Preferably, a set of video frames ,in, For the first A coarse-grained video frame diagram, For the first Coarse-grained video frames The corresponding set of video frames , For the first Coarse-grained video frames The corresponding set of video frame edges, The total number of coarse-grained video frames in the coarse-grained video frame summary set; sentence graph set. ,in, For the first A coarse-grained sentence diagram, For the first A coarse-grained sentence The corresponding set of sentences For the first A coarse-grained sentence The corresponding set of sentence edges, The total number of coarse-grained sentences in the coarse-grained sentence summary set, where:
[0027] 1) The set of video frames and the set of video frame edges corresponding to each coarse-grained video frame are constructed as follows:
[0028] Centered on the corresponding video frame of the current coarse-grained video frame in the video frame sequence, a first preset number of video frames are taken forward and backward in the video frame sequence. If the video frame sequence is exceeded, the actual range is truncated to obtain the corresponding video frame set.
[0029] Select the aligned frame feature representations corresponding to the video frame set from the aligned frame feature representations to form a feature subset, and calculate the cosine similarity between any two aligned frame feature representations in the feature subset;
[0030] If the cosine similarity between the feature representations of two aligned frames is greater than or equal to the first threshold, it is considered that there is a connection between the nodes where the corresponding two video frames are located, and it is recorded as 1; otherwise, it is considered that there is no connection between the nodes where the corresponding two video frames are located, and it is recorded as 0.
[0031] The connection relationships between all the nodes of the video frames are represented as an adjacency matrix, which is the set of video frame edges.
[0032] 2) The set of sentences and the set of sentence edges corresponding to each coarse-grained sentence are constructed as follows:
[0033] Centered on the corresponding sentence of the current coarse-grained sentence in the sentence sequence, a second preset number of sentences are taken forward and backward in the sentence sequence. If the sentence exceeds the sentence sequence, the actual range is truncated to obtain the corresponding sentence set.
[0034] Select the aligned sentence feature representations corresponding to the sentence set from the aligned sentence feature representations to form a feature subset, and calculate the cosine similarity between any two aligned sentence feature representations in the feature subset;
[0035] If the cosine similarity between the feature representations of two aligned sentences is greater than or equal to the first threshold, it is considered that there is a connection between the nodes where the corresponding two sentences are located, and it is recorded as 1; otherwise, it is considered that there is no connection between the nodes where the corresponding two sentences are located, and it is recorded as 0.
[0036] The connection relationships between all nodes containing sentences are represented as an adjacency matrix, which serves as the set of sentence edges.
[0037] Preferably, the fine-grained video frame summary set is extracted as follows:
[0038] For the coarse-grained video frame embedding representation set, the first... The embedded representation set of the first coarse-grained video frame graph is selected from the corresponding video frame set, excluding the first one. The coarse-grained video frame is compared with the other video frames in the video frame sequence, and the embedding representation of each of the other video frames is calculated. Cosine similarity between the embedded representations of corresponding video frames in a video frame sequence for each coarse-grained video frame;
[0039] Before selection The video frames with the highest cosine similarity constitute the corresponding set of fine-grained video frames. It is a positive integer;
[0040] Will The fine-grained video frame sets are merged to form a fine-grained video frame summary set;
[0041] The fine-grained text summarization collection is extracted as follows:
[0042] For the coarse-grained sentence graph embedding representation set, the first... The embedding representation set of the first coarse-grained sentence graph, selected from the corresponding sentence set except for the first... The coarse-grained sentence is compared with the other sentences in the sentence sequence, and the embedding representations of the other sentences are calculated respectively. Cosine similarity between the embedded representations of the corresponding sentences in the sentence sequence of each coarse-grained sentence;
[0043] Before selection The sentences with the highest cosine similarity constitute the corresponding fine-grained sentence set. It is a positive integer;
[0044] Will The sets of fine-grained sentences are merged to form a set of fine-grained text summaries.
[0045] Preferably, the progressive multimodal semantic alignment teaching video content extraction method further trains the progressive multimodal semantic alignment model using a training set and backpropagates the joint total loss to optimize the model parameters of the progressive multimodal semantic alignment model before extracting the teaching video to be extracted.
[0046] Preferably, the combined total loss is calculated as follows:
[0047] The final feature representations of fine-grained video frame summaries and fine-grained sentence summaries are respectively input into the prediction layer to obtain the prediction outputs of fine-grained video frames and fine-grained sentences. The prediction layer includes a linear transform layer and a... The final feature representation of fine-grained video frame summarization is the embedding representation of all fine-grained video frames. When a fine-grained video frame appears in multiple fine-grained video frame sets, the corresponding embedding representation in the embedding representation set of the coarse-grained video frame with the largest sequence number is taken as the embedding representation of that fine-grained video frame. Similarly, the final feature representation of fine-grained sentence summarization is the embedding representation of all fine-grained sentences. When a fine-grained sentence appears in multiple fine-grained sentence sets, the corresponding embedding representation in the embedding representation set of the coarse-grained sentence with the largest sequence number is taken as the embedding representation of that fine-grained sentence.
[0048] Based on the predicted output of fine-grained video frames and the predicted output of fine-grained sentences, the classification loss of video frames and sentences is calculated using binary cross-entropy loss, and the classification loss of video frames and sentences is added together as the multimodal classification loss.
[0049] Based on the cross-modal attention mask matrix, coarse-grained positive sample sets and coarse-grained negative sample sets, as well as fine-grained positive sample sets and fine-grained negative sample sets are formed.
[0050] Multiple coarse-grained cross-modal triples and multiple fine-grained cross-modal triples are constructed based on the entire set of positive and negative samples, and the cross-modal contrast loss of each cross-modal triple is calculated using the InfoNCE Loss function.
[0051] The coarse-grained cross-modal contrast loss is formed by summing the cross-modal contrast losses of each coarse-grained cross-modal triplet, and the fine-grained cross-modal contrast loss is formed by summing the cross-modal contrast losses of each fine-grained cross-modal triplet.
[0052] The joint total loss is formed by weighted summation of the multimodal classification loss, coarse-grained cross-modal contrast loss, and fine-grained cross-modal contrast loss.
[0053] Preferably, the coarse-grained positive sample set includes a coarse-grained video frame positive sample set. and coarse-grained sentence positive sample set The coarse-grained negative sample set includes a coarse-grained set of video frame negative samples. and coarse-grained sentence negative sample set ,in:
[0054] coarse-grained positive sample set of video frames Let represent the set of coarse-grained video frames that are associated with at least one coarse-grained sentence in the cross-modal attention mask matrix, and the set of negative samples of coarse-grained video frames. This represents the set of video frames that are not associated with any of the coarse-grained sentences in the cross-modal attention mask matrix. In the cross-modal attention mask matrix, a value of 1 indicates association, and a value of 0 indicates no association.
[0055] coarse-grained sentence positive sample set Let represent the set of coarse-grained sentences that are associated with at least one coarse-grained video frame in the cross-modal attention mask matrix, and the set of negative samples for each coarse-grained sentence. This represents the set of sentences that are not associated with any of the coarse-grained video frames in the cross-modal attention mask matrix;
[0056] The fine-grained positive sample set includes a fine-grained set of positive samples from video frames. and fine-grained sentence positive sample set The fine-grained negative sample set includes a fine-grained set of video frame negative samples. and fine-grained sentence negative sample set ,in:
[0057] Fine-grained positive sample set of video frames Represents the set of fine-grained video frames that are associated with at least one fine-grained sentence in the cross-modal attention mask matrix; the set of negative samples of fine-grained video frames. This represents the set of video frames that are not associated with any fine-grained sentences in the cross-modal attention mask matrix;
[0058] Fine-grained sentence positive sample set Let f(x) represent the set of fine-grained sentences that are associated with at least one fine-grained video frame in the cross-modal attention mask matrix, and the set of negative samples for each fine-grained sentence. This represents the set of sentences that are not associated with any of the fine-grained video frames in the cross-modal attention mask matrix;
[0059] Coarse-grained transmodal triplets include the first coarse-grained transmodal triplet. Second coarse-grained transmodal triplet Fine-grained transmodal triplets include the first fine-grained transmodal triplet. Second fine-grained cross-modal triplet .
[0060] Preferably, the training set is the TVSum dataset.
[0061] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0062] This invention extracts video frames and sentences from educational videos to generate video frame feature representations and sentence feature representations. It employs a two-layer attention module to accurately model intra-modal and inter-modal relationships, thus highlighting key video frames and sentences in the coarse-grained stage. In the fine-grained stage, it introduces a graph attention module to aggregate node and neighbor information, expanding the coarse-grained summary into a fine-grained summary. This effectively improves the accuracy and coherence of cross-modal semantic alignment, achieving progressive multimodal semantic alignment. Simultaneously, by constructing multiple cross-modal triples from positive and negative sample sets, it optimizes model parameters from coarse to fine learning, maximizing the retention of key information and minimizing redundant information. This results in accurate, fluent, and information-rich summary content, particularly suitable for complex scenarios involving long-term, multimodal teaching videos. It can generate high-quality educational video summaries with complete structure and prominent key points, helping users quickly acquire core knowledge and improve learning efficiency. Attached Figure Description
[0063] Figure 1 This is a flowchart illustrating the teaching video content extraction method for progressive multimodal semantic alignment according to the present invention. Detailed Implementation
[0064] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0065] It should be noted that, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application.
[0066] like Figure 1 As shown, a method for progressive multimodal semantic alignment in extracting teaching video content includes the following steps:
[0067] S1. Sample the educational video to obtain a video frame sequence. and corresponding sentence sequence .
[0068] Extract frame-by-frame images (video frames) from educational videos and sort them chronologically to form a video frame sequence. ,in, For the first One video frame, N is the total number of video frames. Simultaneously, video frame sequences are extracted from the educational video. After the corresponding audio is received, it is transcribed into text using Automatic Speech Recognition (ASR) technology to form a sentence sequence. ,in, For the first One sentence. , This represents the total number of sentences. By sampling and extracting data from educational videos, multimodal (video and text) separation of educational videos is achieved.
[0069] S2, Sequence of video frames Input visual feature extraction model to obtain video frame feature representation , sentence sequence Input text feature extraction model obtains sentence feature representation .
[0070] Specifically, a visual feature extraction model is used to extract features from the video frame sequence X to obtain video frame feature representations. ,in, For the first video frames Feature representation, This represents the dimension of the feature representation of a video frame. Simultaneously, a pre-trained text feature extraction model is used to extract features from the sentence sequence Y, yielding the sentence feature representation. ,in, For the first a sentence Feature representation, This represents the dimension of the sentence's feature representation. Both the visual feature extraction model and the text feature extraction model are pre-trained models. The visual feature extraction model is one of the following: ResNet, DenseNet, GoogleNet, or Vision Transformer, including but not limited to these models. This embodiment uses the GoogleNet model. The text feature extraction model is one of the following: RoBERTa, BERT, GPT, XLNet, or a language model based on the Transformer architecture, including but not limited to these models. This embodiment uses the RoBERTa model.
[0071] S3. Establish a progressive multimodal semantic alignment model to perform the following operations:
[0072] S31. Representation of video frame features Sentence feature representation A two-layer attention module is used for enhancement and alignment, and key frames and key sentences are selected by score prediction to generate a coarse-grained video frame summary set. and coarse-grained text summarization collection The details are as follows:
[0073] S311. Construct the video modal attention mask matrix Text modality attention mask matrix These represent the dependencies within each single modality of the video frame and the sentence, respectively; simultaneously, a cross-modal mask matrix is constructed. , used to represent the interaction relationship between video frames and sentences, where:
[0074] Video modal attention mask matrix The definition is as follows:
[0075] ,
[0076] in, Indicates the first The video frame and the first The mask value of each video frame, when When, it indicates the first The video frame and the first Attention operations are performed on each video frame, when... When, it indicates the first The video frame and the first No attention calculation is performed on any video frame. The possible values of are defined as follows:
[0077]
[0078] in, The size of the local attention window for the video frame. , This is for floor function.
[0079] Specifically, this includes the following three situations: 1) The middle region: when At that time, with the first Centered on a video frame, the area before and after it is covered. Video frames, within this range All are 1, the rest are 0; 2) Initial boundary region: when When the left side of the local attention window of a video frame extends beyond the start of the video frame sequence (i.e., the first video frame), the effective range is truncated. Within this range 3) Terminal boundary region: when At that time, the right side of the local attention window of the video frame extends beyond the end of the video frame sequence (the last video frame), and the effective range is truncated. Within this range The rest are 0;
[0080] Text modality attention mask matrix The definition is as follows:
[0081]
[0082] in, Indicates the first The sentence and the first The mask value of each sentence, when When, it indicates the first The sentence and the first Attention operations are performed on each sentence, when... When, it indicates the first The sentence and the first No attention calculation is performed on these sentences. The possible values of are defined as follows:
[0083]
[0084] in, The size of the local attention window for the sentence. .
[0085] Specifically, this includes the following three situations: 1) The middle region: when At that time, with the first Centered on this sentence, covering both the preceding and following sentences. Sentences within this range All are 1, the rest are 0; 2) Initial boundary region: when When the left side of the local attention window for a sentence extends beyond the starting point of the sentence sequence (i.e., the first sentence), the effective range is clipped. Within this range 3) Terminal boundary region: when When the right side of the sentence's local attention window extends beyond the end of the sentence sequence (i.e., the last sentence), the effective range is clipped. Within this range The rest are 0.
[0086] Cross-modal attention mask matrix The definition is as follows:
[0087] = =
[0088] in, Indicates the first The video frame and the first The mask value of each sentence, when When, it indicates the first The occurrence time of the video frame is at the... Within the start and end time window of the sentence, it is considered that the first sentence... The video frame and the first The sentences are aligned at this point. The video frame and the first Attention operations are performed on each sentence; when When, it indicates the first The occurrence time of the video frame is not in the 1st Within the start and end time window of the sentence, it is considered that the first sentence... The video frame and the first The sentence is not aligned; at this point, the first sentence... The video frame and the first No attention calculation is performed on these sentences.
[0089] S312. Establish a self-attention module consisting of a first self-attention layer and a second self-attention layer in parallel to represent video frame features. With video modality attention mask matrix Input the first self-attention layer and simultaneously represent the sentence features. Text modality attention mask matrix The input to the second self-attention layer yields the corresponding enhanced video frame feature representation. and enhanced sentence feature representation ;
[0090] Enhanced video frame feature representation The calculation is as follows:
[0091]
[0092]
[0093] in, , , These represent the query weight matrix, key weight matrix, and value weight matrix of the first self-attention layer, respectively. For predefined dimensions, , , The first self-attention layer is represented by its query feature representation, key feature representation, and value feature representation, respectively. For the softmax function, This represents element-wise matrix multiplication. This represents the matrix transpose operation;
[0094] Enhanced sentence feature representation The calculation is as follows:
[0095]
[0096]
[0097] in, , , These represent the query weight matrix, key weight matrix, and value weight matrix of the second self-attention layer, respectively. , , The query feature representation, key feature representation, and value feature representation of the second self-attention layer are represented in sequence.
[0098] S313, Enhance video frame feature representation Enhance sentence feature representation Cross-modal attention mask matrix Inputting the cross-modal attention module yields the corresponding cross-modal video frame feature representations. and cross-modal sentence feature representation The cross-modal attention module includes a parallel video cross-modal attention layer and a text cross-modal attention layer;
[0099] Cross-modal video frame feature representation The calculation is as follows:
[0100]
[0101]
[0102] in, , , These represent the query weight matrix, key weight matrix, and value weight matrix of the video cross-modal attention layer, respectively. , , The query feature representation, key feature representation, and value feature representation of the video cross-modal attention layer are represented sequentially.
[0103] Cross-modal sentence feature representation The calculation is as follows:
[0104]
[0105]
[0106] in, , , These represent the query weight matrix, key weight matrix, and value weight matrix of the cross-modal attention layer, respectively. , , The following are the query feature representations, key feature representations, and value feature representations of the cross-modal attention layer of the text, respectively.
[0107] S314. Represent cross-modal video frame features and cross-modal sentence feature representation Further feature enhancement is performed in the feedforward neural network (FFN) to obtain the aligned frame feature representation. And aligned sentence feature representation :
[0108] ;
[0109] in, It is a feedforward neural network, consisting of sequentially connected linear layers, GeLU activation functions, and linear layers. These are the learnable parameters in a feedforward neural network. For the first video frames The alignment frame feature representation, For the first a sentence The alignment of sentence features is represented.
[0110] S315. Represent the alignment frame features. And aligned sentence feature representation In the input score prediction model, the importance score of the corresponding video frame is calculated. Sentence importance score ;
[0111]
[0112] in, For a score prediction model consisting of a single linear layer, the transformation matrix of the linear layer is... This is used to map input features to importance scores for corresponding video frames or sentences.
[0113] S316. Scoring is done based on the importance of each video frame. Sentence importance score The corresponding selection is based on the highest importance score. Each video frame and the preceding 1 sentence, construct a coarse-grained video frame summary set. and coarse-grained text summarization collection ,in, For the first A coarse-grained video frame, For the first A coarse-grained sentence.
[0114] S32. Based on the coarse-grained video frame summary set With coarse-grained text summarization sets Construct video frame sets respectively and sentence diagram set .
[0115] Video frame collection Depend on Composed of several coarse-grained video frame diagrams, the first... Coarse-grained video frame diagram By the Coarse-grained video frames The corresponding set of video frames and video frame edge set Its composition is as follows:
[0116] 1) No. Coarse-grained video frames The corresponding set of video frames It reflects the preceding and following video frames associated with the corresponding coarse-grained video frame, and is constructed as follows:
[0117] 1.1. Locate the index. In video frame sequence The corresponding video frame ;
[0118] 1.2, with In video frame sequence The corresponding video frame Centered on, in the video frame sequence The middle also takes forward and backward. One video frame (if it exceeds the video frame sequence) The sequence boundary is then truncated to the actual range, thus obtaining the first... Coarse-grained video frames The corresponding set of video frames :
[0119] ,
[0120] in, For the first One video frame, The video window value is a predefined value, and satisfies , Represents a set of video frames The total number of video frames in the medium video. ;
[0121] 1.3. Representation of Alignment Frame Features Selecting a set of video frames The aligned frame feature representation of the corresponding video frame forms the first... Feature subset ;
[0122] 2) No. Coarse-grained video frames The corresponding set of video frame edges Indicates the first Coarse-grained video frames The corresponding set of video frames The connection relationships between corresponding nodes in each video frame are constructed as follows:
[0123] 2.1 Each video frame is considered as a node, and the calculation of the... Feature subset The cosine similarity between the aligned frame feature representations of any two video frames is then... video frames Alignment frame feature representation and the video frames Alignment frame feature representation The cosine similarity between them is :
[0124]
[0125] in , Represents the L2 norm, The cosine similarity function;
[0126] 2.2 Setting the first threshold If the cosine similarity between two video frames is greater than or equal to the first threshold, then the two video frames are considered adjacent, and there is a connection (edge) between their nodes. Otherwise, the two video frames are considered not adjacent, and there is no connection between their nodes. Let the first threshold be denoted as . video frames and the video frames edge for:
[0127]
[0128] 2.3, the first Coarse-grained video frames The corresponding set of video frame edges Represented as an adjacency matrix: .
[0129] Edges, to some extent, reflect the connectivity between all nodes, and can be represented by an adjacency matrix {0, 1}. Treating each video frame as a node, if there are a total of... If each video frame corresponds to a node, then the dimension of the adjacency matrix is... When the value of an element in the adjacency matrix is 1, it means that there is an edge between the two nodes; when the value is 0, it means that there is no edge between the two nodes.
[0130] Sentence Image Collection Depend on A coarse-grained sentence diagram Composition, then the first A coarse-grained sentence diagram By the A coarse-grained sentence The corresponding set of sentences and sentence edge set Its composition is as follows:
[0131] 1) The sentence set reflects the sentences preceding and following the coarse-grained sentence, and is constructed as follows:
[0132] 1.1. Locate the index. In sentence sequence The corresponding sentence in ;
[0133] 1.2, with In sentence sequence The corresponding sentence in Centered on the sentence sequence The middle and the front each take By analyzing the sentences (truncating the actual range if they exceed the sequence boundaries), we can obtain the sentence set. :
[0134] ,
[0135] in, For the first One sentence. For predefined sentence window values, , For a set of sentences The total number of sentences in the text;
[0136] 1.3. Representation of Aligned Sentence Features Selecting a set of sentences The aligned sentence feature representation of the corresponding sentences forms the first... Feature subset ;
[0137] 2) Sentence edge set Represents a set of sentences The connection relationships between all sentence-corresponding nodes are constructed as follows:
[0138] 2.1 Each sentence is considered as a node, and the calculation of the first... Feature subset The cosine similarity between the aligned sentence feature representations of any two sentences in the given text is then determined by the first cosine similarity. a sentence Alignment of sentence features and the video frames Alignment of sentence features The cosine similarity between them is ,in, ;
[0139] 2.2 Setting the first threshold If the cosine similarity between two sentences is greater than or equal to the first threshold, then the two sentences are considered adjacent, and there is a connection (edge) between their nodes. Otherwise, the two sentences are not adjacent, and there is no connection between their nodes. Let the first threshold be denoted as . a sentence and the a sentence edge for:
[0140]
[0141] 2.3, the first A coarse-grained sentence The corresponding set of sentence edges Represented as an adjacency matrix: .
[0142] S33, Collect video frames and sentence diagram set The inputs are fed into the graph attention module, which consists of a three-layer graph attention network (GAT) connected in sequence, and correspondingly obtains a coarse-grained video frame graph embedding representation set with enhanced structural information. With coarse-grained sentence graph embedding representation set For the set of video frames and sentence diagram set Perform the following operations on each graph in the dataset:
[0143] For the graph ,in, Given a set of video frames or sentences, and a corresponding set of edges. Represented by the adjacency matrix: , Let be the total number of video frames in the video frame set or the total number of sentences in the sentence set, and denote . This is either a set of video frame nodes or a set of sentence nodes, meaning it corresponds to the nodes formed by all video frames in the video frame set or the nodes formed by all sentences in the sentence set. Then the... Nodes The initial features are denoted as ;No. Nodes The set of neighbor nodes is defined as }, Indicates the first Nodes and the Nodes The connection relationship, Indicates the first Nodes and the Nodes There are edges between them. Indicates the first Nodes and the Nodes There is no edge between them. , This represents the total number of nodes in the video frame node set or the total number of nodes in the sentence node set. Then the... ( The computation process of a graph attention network with 1 layer is as follows:
[0144] (1) Feature transformation: For the first Layer Input features of each node Perform a linear transformation (for the 0th level) Input features of each node That is, the aligned frame feature representation or aligned sentence feature representation of the corresponding node, to obtain the first... The first layer Nodes The middle representation :
[0145]
[0146] in, For the first The trainable weight matrix of the layer;
[0147] (2) Attention score calculation: Calculate the attention score of neighbor node pairs Between the middle nodes Layer attention score :
[0148]
[0149] in, This represents a vector concatenation operation. For the first The trainable attention parameters of the layer For the first The first layer Nodes The middle representation, , The value in parentheses;
[0150] (3) Attention normalization: Use the Softmax function to normalize the attention of neighbor node pairs. The The layer attention scores are normalized.
[0151]
[0152] in, For neighbor node pairs Attention score between nodes For the first 1 node For normalized neighbor node pairs The Layer attention score;
[0153] (4) Feature aggregation: Based on the normalized neighbor node pairs, the first feature aggregation is performed. Layer attention score for the first Nodes The intermediate representations of the neighboring nodes are weighted and summed, and then the result is updated using a non-linear activation function. The first layer Input features of each node :
[0154]
[0155] in, It is a non-linear activation function.
[0156] Repeat the above graph attention network calculation steps three times. This yields the third layer after enhancement using a three-layer graph attention network structure. Input features of each node ; will the first Coarse-grained video frame diagram The input is fed into the graph attention module calculation process described above (i.e. , This allows us to obtain the coarse-grained video frame embedding representation set after structural enhancement. , among which, the The set of embedded representations of each coarse-grained video frame is , Indicates the 0th layer. The input features of each node, Indicates the first The first coarse-grained video frame video frames The embedded representation. Similarly, the first A coarse-grained sentence diagram The input is fed into the graph attention module calculation process described above (i.e. , This will yield the coarse-grained sentence embedding representation set after structural enhancement. , among which, the The set of embedded representations for each coarse-grained sentence graph is as follows: , Indicates the first The first coarse-grained sentence in the diagram a sentence Embedded representation.
[0157] S34. Embed the representation set based on the coarse-grained video frame graph. and coarse-grained sentence embedding representation set Extraction and coarse-grained video frame summarization set based on cosine similarity principle A collection of highly relevant fine-grained video frame summaries and coarse-grained text summarization sets A collection of highly relevant fine-grained text summaries .
[0158] Regarding the first A set of embedded representations of coarse-grained video frames First, in its corresponding video frame set In, select except the first Coarse-grained video frames Other video frames in the video frame sequence besides the corresponding video frame. , For the first The nth video frame, and calculate the nth Embedding representation of video frames With the Coarse-grained video frames Embedding representation of corresponding video frames in a video frame sequence cosine similarity Then, the cosine similarity between the embedded representations of all coarse-grained video frames is sorted, and the top [number] are selected. The video frame with the highest score (cosine similarity) (including the corresponding video frame of the coarse-grained video frame in the video frame sequence and other eligible video frames) constitutes the first... Coarse-grained video frames The corresponding set of fine-grained video frames Finally, all of them By merging the sets of fine-grained video frames corresponding to each coarse-grained video frame, we can obtain the overall set of fine-grained video frame summaries. ,and For the first Fine-grained video frames , , This represents the total number of video frames in the fine-grained video frame summary set, and its corresponding embedding representation. Obtained from the embedded representation set of coarse-grained video frame graphs, specifically defined as:
[0159]
[0160] That is, when the first Fine-grained video frames In multiple fine-grained video frame sets When a given element appears in the fine-grained video frame graph, the embedding representation of that element is taken from the set of embedding representations of the coarse-grained video frame with the highest index as the final representation. Therefore, the final feature representation of the overall fine-grained video frame summary is... .
[0161] Regarding the first A set of embedded representations of coarse-grained sentence graphs First, in its corresponding sentence set In, select except the first A coarse-grained sentence Other sentences besides the corresponding sentence in the sentence sequence , For the first The sentence is ____, and the calculation is ____. Embedded representation of sentences With the A coarse-grained sentence Embedding representation of corresponding sentences in a sentence sequence cosine similarity Then, the cosine similarity between the embedded representations of all coarse-grained sentences is ranked, and the top [number] are selected. The sentence with the highest cosine similarity (including the corresponding sentence of the coarse-grained sentence in the sentence sequence and other sentences that meet the requirements) is used as the first... A coarse-grained sentence The corresponding set of fine-grained sentences Finally, By merging the sets of fine-grained sentences corresponding to each coarse-grained sentence, we can obtain the overall set of fine-grained text summaries. ,and For the first Fine-grained sentences , , This represents the total number of sentences in the fine-grained text summarization set, and its corresponding embedding representation. Obtained from the embedding representation set of coarse-grained sentence graphs, specifically defined as:
[0162]
[0163] That is, the first Fine-grained sentences In multiple fine-grained sentence sets When a sentence appears in the sentence graph, its embedding features are taken from the set of embedding representations of the coarse-grained sentence graph with the highest index as the final representation. Therefore, the final feature representation of the overall fine-grained sentence summary is... .
[0164] S35. The progressive multimodal semantic alignment model is trained using the training set, and the joint total loss is backpropagated to optimize the model parameters. That is, the model training is jointly optimized by combining classification supervision and cross-modal contrastive learning to further improve the consistency of cross-modal summaries. The joint total loss is calculated as follows:
[0165] S351, Represent the final features of fine-grained video frame summarization The final feature representation of fine-grained sentence summarization Input the prediction layer to obtain the corresponding [number]th [layer]. Fine-grained video frames Predicted output and the Fine-grained sentences Predicted output The prediction layer consists of one function and a linear transformation matrix The implementation and calculation process are as follows:
[0166]
[0167]
[0168] in, for function, It is an exponential function;
[0169] S352. Based on the predicted output of each fine-grained video frame and the predicted output of the fine-grained sentence, calculate the classification loss for each video frame. And sentence classification loss ;
[0170] Classification loss of video frames It is achieved using binary cross-entropy loss, and its calculation process is as follows:
[0171]
[0172] in, Hyperparameters Used to balance the weights of correctly predicted samples and incorrectly predicted samples. Hyperparameters used to adjust the degree of attention given to difficult samples. Indicates the first Fine-grained video frames The label (1 indicates that the fine-grained video frame meets the summary requirements, and 0 indicates that the fine-grained video frame does not meet the summary requirements).
[0173] Sentence classification loss It is also implemented using binary cross-entropy loss, and its calculation process is as follows:
[0174]
[0175] in, Indicates the first Fine-grained sentences The tag (1 indicates that the fine-grained sentence meets the abstract requirements, 0 indicates that the fine-grained sentence does not meet the abstract requirements);
[0176] The classification loss for multimodal modes can be expressed as: ;
[0177] S353, Based on the cross-modal attention mask matrix To define coarse-grained sets of positive and negative samples, and fine-grained sets of positive and negative samples;
[0178] For coarse-grained video frames, their set of positive samples is... With coarse-grained video frame negative sample set They are defined as follows:
[0179]
[0180] in, Indicates at least one coarse-grained sentence In cross-modal attention mask matrix There is a relationship between them (i.e.) A collection of coarse-grained video frames; and This represents the set of video frames that are unrelated to all coarse-grained sentences. Additionally, it includes the set of positive samples of coarse-grained video frames. With coarse-grained video frame negative sample set The feature representations of each element are taken from the feature representations of the aligned frames. And obtain it through index mapping.
[0181] For coarse-grained sentences, its set of positive coarse-grained sentence samples With coarse-grained sentence negative sample set They are defined as follows:
[0182]
[0183] in, Indicates at least one coarse-grained video frame In cross-modal attention mask matrix There is a relationship between them (i.e.) A collection of coarse-grained sentences; while This represents the set of sentences that are unrelated to all coarse-grained video frames. Additionally, it includes a set of positive samples for coarse-grained sentences. With coarse-grained sentence negative sample set The feature representations of each element are all taken from the aligned sentence feature representation. And obtain it through index mapping.
[0184] For fine-grained video frames, its set of positive samples is... With fine-grained video frame negative sample set The definition is as follows:
[0185]
[0186] in, Indicates at least one fine-grained sentence In cross-modal attention mask matrix There is a relationship between them (i.e.) A collection of fine-grained video frames; and This represents the set of video frames that are unrelated to all fine-grained sentences. It is also the set of positive samples of fine-grained video frames. The feature representations of each element are taken from the final feature representation of the overall fine-grained video frame summary. fine-grained video frame negative sample set The feature representation is taken from the aligned frame feature representation And it is mapped to the corresponding video frame through index mapping.
[0187] For fine-grained sentences, the set of positive samples of fine-grained sentences is... and fine-grained sentence negative sample set :
[0188]
[0189] in, Indicates at least one fine-grained video frame In cross-modal attention mask matrix There is a relationship between them (i.e.) A collection of fine-grained sentences; and This represents the set of sentences that are unrelated to all fine-grained video frames. The set of positive samples for fine-grained sentences. The feature representations of each element are derived from the final feature representation of the fine-grained sentence summary as a whole. The fine-grained sentence negative sample set The feature representations of each element are taken from the final aligned sentence feature representation. And it is mapped to the corresponding sentences through index mapping.
[0190] S354. Based on the complete set of positive samples and the set of negative samples, construct the triplet for cross-modal interaction and calculate the cross-modal contrast loss;
[0191] Construct cross-modal triples from all positive and negative sample sets. In the form of ), where, This represents a set of positive samples (video frames or sentences) representing a modality. Describes the set of positive samples of another modality. This represents the set of negative samples for another modality; therefore, two coarse-grained cross-modal triples can be constructed respectively: including the first coarse-grained cross-modal triple. Second coarse-grained transmodal triplet And two fine-grained cross-modal triples: including the first fine-grained cross-modal triplet. Second fine-grained cross-modal triplet ;
[0192] Using the constructed cross-modal triples, the cross-modal contrastive loss is calculated based on the InfoNCE Loss function. The calculation process for the cross-modal contrastive loss of each cross-modal triple is as follows:
[0193]
[0194] in, For cross-modal triples ( The cross-modal contrast loss can be obtained by substituting the parameters from each cross-modal triplet. for The number of elements in for The number of elements in for The number of elements in For temperature parameters, Represents the mathematical expectation. express The Middle The feature representation corresponding to each element express The Middle The feature representation corresponding to each element express The Middle The feature representation corresponding to each element For transpose;
[0195] Therefore, based on the above calculation of cross-modal contrast loss, coarse-grained cross-modal contrast loss can be obtained respectively. Compared with fine-grained cross-modal loss :
[0196]
[0197]
[0198] in, For the first coarse-grained transmodal triplet Cross-modal contrast loss, For the second coarse-grained transmodal triplet Cross-modal contrast loss, For the first fine-grained cross-modal triplet Cross-modal contrast loss, For the second fine-grained cross-modal triplet Cross-modal contrast loss;
[0199] S355. Construct the joint total loss and update the model parameters of the progressive multimodal semantic alignment model based on the backpropagation mechanism. The joint total loss is defined as:
[0200]
[0201] in, and These are preset hyperparameters, all greater than 0, used to adjust the weight ratio of each part of the loss in the overall optimization process. The training set can be an existing video dataset familiar to those skilled in the art, such as the TVSum dataset.
[0202] S4. The video frame sequence obtained from the teaching video to be refined. Sentence sequence Video frame feature representation Sentence feature representation Input a progressive multimodal semantic alignment model to obtain a set of fine-grained video frame summaries. and fine-grained text summarization collection This is a summary extracted from the instructional video.
[0203] Specifically, this method is further validated through experiments. The TVSum dataset was used for performance verification in this embodiment. The TVSum dataset contains 50 videos covering 10 themes including news, education, sports, music, and documentaries, exhibiting rich thematic diversity and content variation. This allows for a comprehensive and effective verification of the effectiveness and generalization ability of this method in extracting video content across different scenarios. To scientifically and rigorously evaluate the performance advantages of this method, four existing video summarization methods—A2Summ, VJMHT, MMAN, and AMFM—were selected, along with the coarse-grained scheme proposed in this embodiment (i.e., obtaining a coarse-grained set). ) and fine-grained schemes (i.e., obtaining a set of fine-grained video frame summaries) and fine-grained text summarization collection Systematic comparative experiments were conducted, including: 1) A2Summ reference: He, Bo, et al. "Align and attend: Multimodal summarization with dual contrastive losses." Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2023.; 2) VJMHT reference: Li, Haopeng, et al. "Video joint modelling based on hierarchical transformer for co-summarization." IEEE Transactions on Pattern Analysis and Machine Intelligence 45.3 (2022): 3904-3917.; 3) MMAN reference: Han, Tingting, et al. "Effective video summarization by extracting parameter-free motion attention." ACM Transactions on Multimedia Computing, Communications and Applications 20.7 (2024): 1-20.; 4) AMFM reference: Singh, Aditya Kumar, DhruvSrivastava, and Makarand Tapaswi. "Previously on..." From Recaps to StorySummarization." Proceedings of the IEEE / CVF Conference on Computer Vision andPattern Recognition. 2024.The experiment used three core evaluation metrics: F1 score, Kendall correlation coefficient, and Spearman correlation coefficient. The F1 score comprehensively measures the precision and recall when selecting the summary set (i.e., extracting key video frames and sentences), thus measuring the accuracy and completeness of the method in extracting key video information. The Kendall and Spearman correlation coefficients, from the perspective of ranking consistency, quantify the correlation between the importance ranking of video frames and sentences predicted by the method and the manually labeled ranking, providing a stable and reliable reflection of the method's performance in prioritizing video frames and sentences. See Table 1 for details.
[0204] Table 1
[0205]
[0206] The experimental data in Table 1 show that both the coarse-grained and fine-grained schemes proposed in this embodiment significantly outperform the comparative methods in all three core metrics. The fine-grained scheme is particularly outstanding, achieving a dual improvement in the accuracy and ranking consistency of key video frame and sentence selection with an F1 score of 65.8, a Kendall correlation coefficient of 0.194, and a Spearman correlation coefficient of 0.261. Compared to AMFM, the best-performing comparative method, the fine-grained scheme improves the F1 score by 4.8 percentage points, the Kendall correlation coefficient from 0.179 to 0.194, and the Spearman correlation coefficient from 0.233 to 0.261. These experimental results fully validate the significant advantages of this method in video summarization performance, enabling more efficient and accurate extraction of video content summaries and ranking of key video frames and sentences.
[0207] It should be understood that, unless explicitly stated herein, there is no strict order in which these steps are executed; they can be performed in any other order. Furthermore, each step may include multiple sub-steps or stages, which do not necessarily complete simultaneously but can be executed at different times. The execution order of these sub-steps or stages is also not necessarily sequential; they can be executed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps. These steps can be stored in software form in the memory of a computer device so that the processor can invoke and execute the operations corresponding to these steps.
[0208] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0209] The embodiments described above are merely specific and detailed examples of the embodiments described in this application, and should not be construed as limiting the scope of the application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these modifications and improvements all fall within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the appended claims.
Claims
1. A method for extracting teaching video content through progressive multimodal semantic alignment, characterized in that: Includes the following steps: S1. Sample the educational video to obtain the video frame sequence and the corresponding sentence sequence; S2. Input the video frame sequence into the visual feature extraction model to obtain the video frame feature representation, and input the sentence sequence into the text feature extraction model to obtain the sentence feature representation; S3. Establish a progressive multimodal semantic alignment model to perform the following operations: S31. Use a two-layer attention module to extract features from the video frame feature representation and the sentence feature representation respectively, and generate a coarse-grained video frame summary set and a coarse-grained text summary set accordingly. S32. Based on the coarse-grained video frame summary set and the coarse-grained text summary set, construct the corresponding video frame graph set and sentence graph set; S33. The graph attention module is used to extract features from the video frame graph set and the sentence graph set respectively, thereby obtaining a coarse-grained video frame graph embedding representation set and a coarse-grained sentence embedding representation set. The graph attention module includes a three-layer graph attention network connected in sequence. S34. Based on the coarse-grained video frame embedding representation set and the coarse-grained sentence embedding representation set, respectively, extract the fine-grained video frame summary set and the fine-grained text summary set according to the cosine similarity principle. S4. Input the video frame sequence, sentence sequence, video frame feature representation, and sentence feature representation obtained from the teaching video to be refined into the progressive multimodal semantic alignment model to obtain a fine-grained video frame summary set and a fine-grained text summary set, which are the summary contents extracted from the teaching video.
2. The method for extracting teaching video content through progressive multimodal semantic alignment as described in claim 1, characterized in that: The visual feature extraction model is one of the pre-trained ResNet, DenseNet, GoogleNet, or VisionTransformer models; the text feature extraction model is one of the pre-trained RoBERTa, BERT, GPT, XLNet, or a language model based on the Transformer architecture.
3. The method for extracting teaching video content through progressive multimodal semantic alignment as described in claim 1, characterized in that: The method utilizes a dual-layer attention module to extract features from both video frame and sentence feature representations, generating coarse-grained video frame summary sets and coarse-grained text summary sets, as detailed below: S311. Constructing the video modal attention mask matrix Text modality attention mask matrix and cross-modal mask matrix In this context, an element in each mask matrix being 1 indicates that attention operations are performed; otherwise, attention operations are not performed. N represents the total number of video frames in the video frame sequence. The total number of sentences in the sentence sequence; S312. Establish a self-attention module consisting of a parallel first self-attention layer and a second self-attention layer. Input the video frame feature representation and the video modality attention mask matrix into the first self-attention layer, and input the sentence feature representation and the text modality attention mask matrix into the second self-attention layer to obtain the enhanced video frame feature representation and the enhanced sentence feature representation respectively. S313. Input the enhanced video frame feature representation, enhanced sentence feature representation, and cross-modal attention mask matrix into the cross-modal attention module to obtain the cross-modal video frame feature representation and the cross-modal sentence feature representation; S314. Input the cross-modal video frame feature representation and the cross-modal sentence feature representation into the feedforward neural network respectively to obtain the aligned frame feature representation and the aligned sentence feature representation accordingly; S315. Input the aligned frame feature representation and the aligned sentence feature representation into the score prediction model to obtain the video frame importance score and sentence importance score, respectively. The score prediction model is a linear layer. S316. Select the top-scoring video frames from the importance scores. The top-scoring video frames and sentences in importance scores Each sentence corresponds to a coarse-grained video frame summary set and a coarse-grained text summary set. , .
4. The method for extracting teaching video content through progressive multimodal semantic alignment as described in claim 3, characterized in that: The feedforward neural network comprises a linear layer, a GeLU activation function, and another linear layer connected in sequence.
5. The method for extracting teaching video content through progressive multimodal semantic alignment as described in claim 1, characterized in that: The video frame set ,in, For the first A coarse-grained video frame diagram, For the first Coarse-grained video frames The corresponding set of video frames , For the first Coarse-grained video frames The corresponding set of video frame edges, The total number of coarse-grained video frames in the coarse-grained video frame summary set; the sentence graph set ,in, For the first A coarse-grained sentence diagram, For the first A coarse-grained sentence The corresponding set of sentences For the first A coarse-grained sentence The corresponding set of sentence edges, The total number of coarse-grained sentences in the coarse-grained sentence summary set, where: 1) The set of video frames and the set of video frame edges corresponding to each coarse-grained video frame are constructed as follows: Centered on the corresponding video frame of the current coarse-grained video frame in the video frame sequence, a first preset number of video frames are taken forward and backward in the video frame sequence. If the video frame sequence is exceeded, the actual range is truncated to obtain the corresponding video frame set. Select the aligned frame feature representations corresponding to the video frame set from the aligned frame feature representations to form a feature subset, and calculate the cosine similarity between any two aligned frame feature representations in the feature subset; If the cosine similarity between the feature representations of two aligned frames is greater than or equal to the first threshold, it is considered that there is a connection between the nodes where the corresponding two video frames are located, and it is recorded as 1; otherwise, it is considered that there is no connection between the nodes where the corresponding two video frames are located, and it is recorded as 0. The connection relationships between all the nodes of the video frames are represented as an adjacency matrix, which is the set of video frame edges. 2) The set of sentences and the set of sentence edges corresponding to each coarse-grained sentence are constructed as follows: Centered on the corresponding sentence of the current coarse-grained sentence in the sentence sequence, a second preset number of sentences are taken forward and backward in the sentence sequence. When the sentence exceeds the sentence sequence, the actual range is truncated to obtain the corresponding sentence set. Select the aligned sentence feature representations corresponding to the sentence set from the aligned sentence feature representations to form a feature subset, and calculate the cosine similarity between any two aligned sentence feature representations in the feature subset; If the cosine similarity between the feature representations of two aligned sentences is greater than or equal to the first threshold, it is considered that there is a connection between the nodes where the corresponding two sentences are located, and it is recorded as 1; otherwise, it is considered that there is no connection between the nodes where the corresponding two sentences are located, and it is recorded as 0. The connection relationships between all nodes containing sentences are represented as an adjacency matrix, which serves as the set of sentence edges.
6. The method for extracting teaching video content through progressive multimodal semantic alignment as described in claim 5, characterized in that: The fine-grained video frame summary set is extracted as follows: For the coarse-grained video frame embedding representation set, the first... The embedded representation set of the first coarse-grained video frame graph is selected from the corresponding video frame set, excluding the first one. The coarse-grained video frame is compared with the other video frames in the video frame sequence, and the embedding representation of each of the other video frames is calculated. Cosine similarity between the embedded representations of corresponding video frames in a video frame sequence for each coarse-grained video frame; Before selection The video frames with the highest cosine similarity constitute the corresponding set of fine-grained video frames. It is a positive integer; Will The fine-grained video frame sets are merged to form a fine-grained video frame summary set; The fine-grained text summary set is extracted as follows: For the coarse-grained sentence graph embedding representation set, the first... The embedding representation set of the first coarse-grained sentence graph, selected from the corresponding sentence set except for the first... The coarse-grained sentence is compared with the other sentences in the sentence sequence, and the embedding representations of the other sentences are calculated respectively. Cosine similarity between the embedded representations of the corresponding sentences in the sentence sequence of each coarse-grained sentence; Before selection The sentences with the highest cosine similarity constitute the corresponding fine-grained sentence set. It is a positive integer; Will The sets of fine-grained sentences are merged to form a set of fine-grained text summaries.
7. The method for extracting teaching video content through progressive multimodal semantic alignment as described in claim 6, characterized in that: The progressive multimodal semantic alignment teaching video content extraction method further trains the progressive multimodal semantic alignment model using a training set and backpropagates the joint total loss to optimize the model parameters of the progressive multimodal semantic alignment model before extracting the teaching video content.
8. The method for extracting teaching video content through progressive multimodal semantic alignment as described in claim 7, characterized in that: The combined total loss is calculated as follows: The final feature representations of the fine-grained video frame summary and the fine-grained sentence summary are respectively input into the prediction layer to obtain the prediction outputs of the fine-grained video frames and the fine-grained sentences. The prediction layer includes a linear transform layer and a... The function, wherein the final feature representation of the fine-grained video frame summary is the embedding representation of all fine-grained video frames, when a fine-grained video frame appears in multiple fine-grained video frame sets, takes the corresponding embedding representation in the embedding representation set of the coarse-grained video frame with the largest sequence number as the embedding representation of that fine-grained video frame; the final feature representation of the fine-grained sentence summary is the embedding representation of all fine-grained sentences, when a fine-grained sentence appears in multiple fine-grained sentence sets, takes the corresponding embedding representation in the embedding representation set of the coarse-grained sentence with the largest sequence number as the embedding representation of that fine-grained sentence; Based on the predicted output of fine-grained video frames and the predicted output of fine-grained sentences, the classification loss of video frames and sentences is calculated using binary cross-entropy loss, and the classification loss of video frames and sentences is added together as the multimodal classification loss. Based on the cross-modal attention mask matrix, coarse-grained positive sample sets and coarse-grained negative sample sets, as well as fine-grained positive sample sets and fine-grained negative sample sets are formed. Multiple coarse-grained cross-modal triples and multiple fine-grained cross-modal triples are constructed based on the entire set of positive and negative samples, and the cross-modal contrast loss of each cross-modal triple is calculated using the InfoNCE Loss function. The coarse-grained cross-modal contrast loss is formed by summing the cross-modal contrast losses of each coarse-grained cross-modal triplet, and the fine-grained cross-modal contrast loss is formed by summing the cross-modal contrast losses of each fine-grained cross-modal triplet. The joint total loss is formed by weighted summation of the multimodal classification loss, coarse-grained cross-modal contrast loss, and fine-grained cross-modal contrast loss.
9. The method for extracting teaching video content through progressive multimodal semantic alignment as described in claim 8, characterized in that: The coarse-grained positive sample set includes a coarse-grained video frame positive sample set. and coarse-grained sentence positive sample set The coarse-grained negative sample set includes a coarse-grained video frame negative sample set. and coarse-grained sentence negative sample set ,in: coarse-grained positive sample set of video frames Let represent the set of coarse-grained video frames that are associated with at least one coarse-grained sentence in the cross-modal attention mask matrix, and the set of negative samples of coarse-grained video frames. This represents the set of video frames that are not associated with any of the coarse-grained sentences in the cross-modal attention mask matrix. In the cross-modal attention mask matrix, a value of 1 indicates association, and a value of 0 indicates no association. coarse-grained sentence positive sample set Let represent the set of coarse-grained sentences that are associated with at least one coarse-grained video frame in the cross-modal attention mask matrix, and the set of negative samples for each coarse-grained sentence. This represents the set of sentences that are not associated with any of the coarse-grained video frames in the cross-modal attention mask matrix; The fine-grained positive sample set includes a fine-grained set of positive samples from video frames. and fine-grained sentence positive sample set The fine-grained negative sample set includes a fine-grained video frame negative sample set. and fine-grained sentence negative sample set ,in: Fine-grained positive sample set of video frames Represents the set of fine-grained video frames that are associated with at least one fine-grained sentence in the cross-modal attention mask matrix; the set of negative samples of fine-grained video frames. This represents the set of video frames that are not associated with any fine-grained sentences in the cross-modal attention mask matrix; Fine-grained sentence positive sample set Let f(x) represent the set of fine-grained sentences that are associated with at least one fine-grained video frame in the cross-modal attention mask matrix, and the set of negative samples for each fine-grained sentence. This represents the set of sentences that are not associated with any of the fine-grained video frames in the cross-modal attention mask matrix; The coarse-grained transmodal triplet includes a first coarse-grained transmodal triplet. Second coarse-grained transmodal triplet The fine-grained transmodal triplet includes a first fine-grained transmodal triplet. Second fine-grained cross-modal triplet .
10. The method for extracting teaching video content through progressive multimodal semantic alignment as described in claim 7, characterized in that: The training set is the TVSum dataset.