Video key frame extraction method and device, equipment and medium

By using a multimodal encoder and feature space alignment technology, the problems of modal uniformity and insufficient temporal coherence in video keyframe extraction are solved, achieving accurate keyframe extraction and semantic integrity, and improving the accuracy and efficiency of video analysis.

CN121921700APending Publication Date: 2026-04-24PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2026-01-09
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing technologies for extracting keyframes in videos suffer from modal singularity, lack of multimodal feature alignment, and insufficient temporal coherence, resulting in insufficient accuracy and semantic integrity in keyframe extraction in complex scenes.

Method used

A multimodal encoder is used to extract feature vectors from visual, audio and text data respectively, and these vectors are mapped to a unified semantic space through feature space alignment. By combining temporal clustering and saliency aggregator, key video frames are determined.

Benefits of technology

It significantly improves the accuracy and semantic integrity of keyframes in complex semantic scenarios, providing efficient and high-quality input support for video summarization and content retrieval tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121921700A_ABST
    Figure CN121921700A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of artificial intelligence, and relates to a video key frame extraction method, device and equipment and a medium, and the method comprises the steps: receiving visual, audio and text multi-modal data of a target video, extracting visual, audio and text feature vectors through a multi-modal encoder, and carrying out the spatial alignment of the three types of feature vectors, and obtaining a target feature vector set. And based on the aligned target feature vectors, performing time sequence clustering on the video frames, dividing a plurality of time sequence semantic stages, and splicing visual, audio and text features corresponding to each stage to obtain stage multi-modal splicing vectors. And calculating a significance score of each stage by using a time sequence significance aggregator and combining the splicing vector, screening out a target stage of which the score reaches a threshold value, and finally extracting a key video frame from the video frames corresponding to the target stage. The method can be applied to the business fields of financial science and technology, medical treatment and the like, and can improve the extraction accuracy of the video key frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology and is applied to online processing business scenarios such as fintech and healthcare. In particular, it relates to a method, apparatus, device, and medium for extracting keyframes from videos. Background Technology

[0002] Video keyframe extraction is a fundamental task in the field of video content analysis. Its core objective is to select frames with representative semantic information from highly redundant video sequences, providing efficient input support for downstream tasks such as video summarization, content retrieval, and data compression. It has wide application value in scenarios such as multimedia processing, intelligent monitoring, online education, and finance.

[0003] In financial scenarios, the need for keyframe extraction is particularly prominent. For example, banks need to extract frames for key processes such as "customer identity verification" and "document signing" from massive amounts of counter business video recordings for business compliance audits; securities institutions need to extract core frames such as "instruction issuance" and "transaction confirmation" from transaction process videos to assist in transaction retrospection and risk verification. In existing technologies, video keyframe extraction methods are mainly divided into two categories: one is the traditional method based on manually designed features, which relies on manually defined visual features such as color histograms, edge gradients, and optical flow fields, combined with inter-frame difference thresholds or clustering algorithms to select keyframes; the other is the method based on single-modal deep learning, which uses models such as convolutional neural networks and visual Transformers to perform end-to-end feature extraction on visual data to complete the selection of keyframes.

[0004] However, the above methods have significant technical limitations: First, they suffer from modality singularity, utilizing only visual modality information and failing to mine the complementary semantics of multimodal data such as audio and text, resulting in insufficient accuracy in keyframe extraction in complex semantic scenarios; second, they lack multimodal feature alignment, although some recent studies have attempted to fuse multimodal information through cross-modal contrastive learning, they have failed to achieve effective alignment of different modal features in a unified space, making it difficult to accurately capture multimodal semantic associations; third, they lack temporal coherence, as existing methods mostly use independent frames as processing units, failing to consider the temporal semantic dependencies of videos, and are unable to aggregate and evaluate video stages with continuous semantics, resulting in a lack of semantic completeness in the keyframe extraction results. Summary of the Invention

[0005] The purpose of this application is to propose a video keyframe extraction method, apparatus, computer device, and storage medium to solve the problems of insufficient accuracy and semantic integrity of keyframes in complex scenes caused by the reliance on single-modal evaluation, lack of multimodal semantic integration, and lack of temporal aggregation processing in the prior art.

[0006] Firstly, a method for extracting keyframes from video is provided, which employs the following technical solution: The system receives multimodal data from the target video, including visual, audio, and text data. A multimodal encoder is used to extract features from each of the visual, audio, and text data, resulting in multiple visual, audio, and text feature vectors. Feature space alignment is performed on these vectors to obtain aligned target visual, audio, and text feature vectors. Based on these vectors, temporal clustering is performed on the video frames of the visual data to determine multiple temporal semantic stages and the corresponding multimodal concatenation vector for each stage. A temporal saliency score is calculated for each stage using a temporal saliency aggregator. Temporal semantic stages with saliency scores greater than or equal to a preset threshold are designated as target temporal semantic stages, and multiple key video frames are extracted from the corresponding video frames.

[0007] Secondly, a video keyframe extraction device is provided, which adopts the following technical solution: The receiving module is used to receive multimodal data of the target video, which includes visual data, audio data, and text data. The feature extraction module is used to extract features from visual data, audio data, and text data using a multimodal encoder, resulting in multiple visual feature vectors, multiple audio feature vectors, and multiple text feature vectors. The alignment module is used to perform feature space alignment processing on multiple visual feature vectors, multiple audio feature vectors and multiple text feature vectors to obtain multiple aligned target visual feature vectors, multiple target audio feature vectors and multiple target text feature vectors. The clustering module is used to perform temporal clustering of video frames of visual data based on multiple target visual feature vectors, multiple target audio feature vectors, and multiple target text feature vectors, to determine multiple temporal semantic stages and the multimodal concatenation vector corresponding to each temporal semantic stage. The computation module is used to calculate the saliency score of each temporal semantic stage based on the multimodal concatenation vector of each temporal semantic stage through the temporal saliency aggregator; The acquisition module is used to identify temporal semantic stages with saliency scores greater than or equal to a preset score threshold as target temporal semantic stages, and to acquire multiple key video frames from multiple video frames corresponding to the target temporal semantic stages.

[0008] Thirdly, a computer device is provided, which adopts the following technical solution: The system receives multimodal data from the target video, including visual, audio, and text data. A multimodal encoder is used to extract features from each of the visual, audio, and text data, resulting in multiple visual, audio, and text feature vectors. Feature space alignment is performed on these vectors to obtain aligned target visual, audio, and text feature vectors. Based on these vectors, temporal clustering is performed on the video frames of the visual data to determine multiple temporal semantic stages and the corresponding multimodal concatenation vector for each stage. A temporal saliency score is calculated for each stage using a temporal saliency aggregator. Temporal semantic stages with saliency scores greater than or equal to a preset threshold are designated as target temporal semantic stages, and multiple key video frames are extracted from the corresponding video frames.

[0009] Fourthly, a computer-readable storage medium is provided, which adopts the following technical solution: The system receives multimodal data from the target video, including visual, audio, and text data. A multimodal encoder is used to extract features from each of the visual, audio, and text data, resulting in multiple visual, audio, and text feature vectors. Feature space alignment is performed on these vectors to obtain aligned target visual, audio, and text feature vectors. Based on these vectors, temporal clustering is performed on the video frames of the visual data to determine multiple temporal semantic stages and the corresponding multimodal concatenation vector for each stage. A temporal saliency score is calculated for each stage using a temporal saliency aggregator. Temporal semantic stages with saliency scores greater than or equal to a preset threshold are designated as target temporal semantic stages, and multiple key video frames are extracted from the corresponding video frames.

[0010] Compared with existing technologies, the embodiments of this application have the following main advantages: By simultaneously receiving visual, audio, and text data and extracting features separately through a dedicated encoder, the complementary semantics of audio and text are fully utilized, overcoming the information limitations of single-modal evaluation and significantly improving the accuracy of keyframe extraction in complex semantic scenarios. An additional feature space alignment step is added, mapping the features of the three modalities to a unified semantic space, achieving accurate capture of multimodal semantic associations, solving the problem of feature misalignment in existing cross-modal fusion, and strengthening the semantic consistency of keyframe selection. Video frames are aggregated into temporal semantic stages through temporal clustering, and stage-level saliency evaluation is performed by combining multimodal concatenation vectors and a temporal saliency aggregator, fully considering the temporal dependencies of the video, avoiding semantic fragmentation caused by independent frame evaluation, and ensuring the semantic integrity and logical coherence of keyframes. Keyframes are extracted based on threshold-based selection of target semantic stages, ensuring the representativeness of keyframes while reducing redundant frame retention through stage aggregation, providing efficient and high-quality input support for downstream tasks such as video summarization and content retrieval. Attached Figure Description

[0011] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is an exemplary system architecture diagram to which this application can be applied; Figure 2 A flowchart of an embodiment of the video keyframe extraction method according to this application; Figure 3 This is a schematic diagram of the structure of one embodiment of the video keyframe extraction apparatus according to this application; Figure 4 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation

[0013] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.

[0014] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0015] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0016] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables.

[0017] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.

[0018] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptop computer 1011, tablet computer 1012 or mobile phone 1013, terminal device 101 can also be e-book reader, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer and desktop computer, etc.

[0019] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.

[0020] It should be noted that the video keyframe extraction method provided in this application embodiment is generally executed by a server / terminal device, and correspondingly, the video keyframe extraction device is generally set in the server / terminal device.

[0021] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0022] Continue to refer to Figure 2 A flowchart of an embodiment of the video keyframe extraction method according to this application is shown. The video keyframe extraction method includes the following steps: Step S201: Receive multimodal data of the target video, which includes visual data, audio data, and text data.

[0023] In this context, "target video" refers to the original video footage from which keyframes are extracted. Typical examples include "cake baking tutorial videos" and "machine tool assembly operation videos," which can originate from recordings on multimedia storage systems, online platforms, or monitoring equipment. Multimodal data, on the other hand, is a collection of various types of information parsed and correlated from the target video. This data can include visual, audio, and textual data. Visual data comes from the frame sequence of the target video, audio data from the video's audio track, and textual data typically comes from embedded subtitles, step-by-step instructions, or scene description documents. This type of data characterizes the semantic content of the target video across different perceptual dimensions. Its core purpose is to provide multi-dimensional input for the multimodal encoder, enabling the discovery of complementary semantics that cannot be covered by a single modality.

[0024] Visual data consists of a series of consecutive image frames arranged chronologically from the target video. These frames are directly derived from video sampling, and each frame represents the visual content of the video at a specific moment (e.g., object shape, motion state, scene layout). Audio data comprises sound signals synchronized with the visual content, originating from the video's audio track (including speech, sound effects, background music, etc.). It exists as continuous sound waves, representing the auditory semantics of the video content. Text data is textual information associated with the target video content. Sources may include embedded subtitles, manually annotated step-by-step instructions, and automatically generated speech-to-text results. It exists as a sequence of characters, representing the textual semantics of the video content.

[0025] In step S202, a multimodal encoder is used to extract features from the visual data, audio data, and text data respectively, resulting in multiple visual feature vectors, multiple audio feature vectors, and multiple text feature vectors.

[0026] The multimodal encoder is the core functional component for extracting multimodal features. It consists of three parallel feature extraction branches: visual, audio, and text. Each branch is adapted to the processing requirements of a specific modality of data. For example, the visual branch of the multimodal encoder uses a visual feature extraction model, the audio branch uses an audio feature extraction model, and the text branch uses a text feature extraction model, which can process the three types of data and output the corresponding features respectively.

[0027] Feature extraction is the process of transforming raw multimodal data (visual data, audio data, and text data) into low-dimensional, quantized feature vectors. The core of this process is to filter and abstract representative semantic information from the raw data.

[0028] Each visual feature vector corresponds to an image frame in the visual data. These vectors represent the high-dimensional semantic information of each image frame (such as features of objects, actions, and scenes). Each audio feature vector corresponds to an audio segment in the audio data, and its data source is the result of the audio segment being converted and encoded using a spectrogram. These vectors represent the frequency domain semantic information of each audio segment. Each text feature vector corresponds to a segment of text in the text data (such as a subtitle), and its data source is the result of the text sequence being semantically encoded. These vectors represent the deep semantic information of the text content (such as the meaning of instructions and descriptions).

[0029] Step S203: Perform feature space alignment processing on multiple visual feature vectors, multiple audio feature vectors and multiple text feature vectors to obtain multiple aligned target visual feature vectors, multiple target audio feature vectors and multiple target text feature vectors.

[0030] Feature space alignment refers to the process of semantic matching of visual, audio and text feature vectors in a unified semantic space through cross-modal contrastive learning. Its core logic is based on the construction of positive / negative sample pairs and loss optimization.

[0031] Here, multiple target visual feature vectors refer to the set of visual feature vectors output by the updated visual feature extraction model after feature space alignment processing. The data source is the image patches of the original visual data, which are then re-encoded by the updated model. The core characteristic of this vector set is that it semantically matches the target audio and text feature vectors representing the same semantics in a unified space; that is, visual features with the same semantics are closer to other modal features, while those with different semantics are farther apart.

[0032] Here, multiple target audio feature vectors refer to the set of audio feature vectors output by the updated audio feature extraction model after feature space alignment processing. The data source is the result of the original audio segmentation and conversion spectrogram, which is then re-encoded by the updated model. The core characteristic of this vector set is that it achieves semantic matching with target visual and text feature vectors representing the same semantics in a unified space.

[0033] Among them, multiple target text feature vectors refer to the set of text feature vectors output by the text feature extraction model after feature space alignment processing, and their data source is the result of the original text data (such as subtitles) being re-encoded by the updated model. The core feature of this vector set is that it is semantically aligned with the target visual and audio feature vectors representing the same semantics in a unified space, that is, the text semantics and the features of the corresponding visual images and audio content are matched in space.

[0034] Step S204: Based on multiple target visual feature vectors, multiple target audio feature vectors, and multiple target text feature vectors, perform temporal clustering on the video frames of the visual data to determine multiple temporal semantic stages and the multimodal concatenation vector corresponding to each temporal semantic stage. Temporal clustering is a process that divides consecutive video frames into several groups based on the temporal correlation and semantic similarity of video frames. Multiple temporal semantic stages are sets of video segments obtained after temporal clustering. Each stage consists of consecutive video frames (i.e., image frames), and its data source is the temporal and semantic aggregation of video frames. These stages represent continuous intervals with unified semantics in the target video. For example, a "car repair video" can be divided into four temporal semantic stages: "removing tires," "inspecting the rim," "installing a new tire," and "tightening screws."

[0035] The multimodal concatenation vector is a high-dimensional vector obtained by concatenating the visual, audio, and text feature vectors of a target within a certain temporal semantic stage. This vector represents the comprehensive multimodal semantic information (integrating visual, audio, and text features) of the corresponding temporal semantic stage.

[0036] Step S205: Based on the multimodal concatenation vector of each temporal semantic stage, calculate the saliency score of each temporal semantic stage through a temporal saliency aggregator.

[0037] The temporal saliency aggregator is used to process multimodal concatenation vectors from temporal semantic stages. Its core capability is to calculate the degree of correlation between stages through a self-attention mechanism and to perform weighted aggregation of multimodal features. Its purpose is to output the saliency score of each temporal semantic stage to evaluate the semantic importance of that stage.

[0038] The saliency score is the evaluation result of the temporal saliency aggregator for each temporal semantic stage. Its value ranges from 0 to 1, and the data source is the output of the multimodal concatenation vector after attention weighting and activation function processing. This score characterizes the semantic importance of the corresponding temporal semantic stage in the entire video; a higher value indicates a more important stage. For example, the saliency score of the "installing a new tire" stage is 0.95, indicating that it is a core stage in the video.

[0039] Step S206: The temporal semantic stage with a saliency score greater than or equal to a preset score threshold is taken as the target temporal semantic stage, and multiple key video frames are obtained from the multiple video frames corresponding to the target temporal semantic stage.

[0040] The score threshold is a pre-defined critical value for saliency score, the value of which can be determined based on the requirements of downstream tasks (such as the number of keyframes and semantic integrity). This threshold represents the standard for distinguishing between "critical stages" and "non-critical stages," and its core purpose is to filter out target temporal semantic stages that meet the saliency score criteria. For example, if the score threshold is set to 0.7, stages with a saliency score greater than or equal to 0.7 will be judged as critical stages (target temporal semantic stages).

[0041] The target temporal semantic stage is defined as a temporal semantic stage with a saliency score greater than or equal to a score threshold. This stage represents a continuous interval with high semantic importance in the target video, and its core purpose is to extract key video frames. For example, in a "car repair video," the saliency scores of the "installing a new tire" and "tightening screws" stages are both greater than or equal to 0.7, and are therefore identified as the target temporal semantic stage.

[0042] Among them, the multiple video frames corresponding to the target temporal semantic stage are a collection of video frames (i.e. image frames) contained in the target temporal semantic stage. These video frames originate from the image frames corresponding to the target temporal semantic stage within the visual data.

[0043] Several key video frames are representative frames selected from the video frames corresponding to the target temporal semantic stage. Their data source is the selection of video frames within the target stage, such as choosing the frame with the richest semantics within that stage. Their application can support downstream tasks such as video summarization, content retrieval, and data compression.

[0044] In one example, the target video is a 10-minute video titled "Laparoscopic Appendectomy Teaching Video," with a frame rate of 25 frames per second. Multimodal data is acquired, including visual data consisting of surgical field frames captured by the laparoscopy, encompassing scenes of "umbilical puncture," "appendix localization," "mesovascular ligation," and "appendectomy," totaling 15,000 frames; audio data consisting of the surgeon's narration, such as "Identify the ileocecal junction here" and "Use an electrocautery hook to treat the mesentery"; and text data consisting of seven subtitles, including "Step 2: Establish pneumoperitoneum" and "Step 4: Appendix transection." The multimodal encoder's visual branch uses a visual feature extraction model to segment each frame into 16×16 image blocks for encoding, resulting in 15,000 768-dimensional visual feature vectors. Audio is divided into 15,000 segments according to frame time sequence, converted to Mel spectrograms, and then processed by an audio feature extraction model to obtain 15,000 256-dimensional audio feature vectors. The text branch uses a text feature extraction model to encode subtitles, resulting in 7 768-dimensional text feature vectors. After feature space alignment, based on inter-frame feature similarity, the segments are temporally clustered into five stages: "pneumoperitoneum establishment," "appendix localization," "mesopleural ligation," "appendix resection," and "incision suturing." The features from each stage are then concatenated to obtain 5 multimodal concatenated vectors. The scores were calculated using a temporal saliency aggregator: "Initiation of pneumoperitoneum" scored 0.58 points, "Appendix localization" scored 0.85 points, "Mesentery ligation" scored 0.92 points, "Appendicectomy" scored 0.88 points, and "Incision suturing" scored 0.63 points. Three target temporal semantic stages were selected with a threshold of 0.7. Ten keyframes for "Precise Appendix Localization" and "Safe Mesentery Ligation" were extracted from these three temporal semantic stages.

[0045] This application's embodiments simultaneously receive visual, audio, and text data and extract features separately using a dedicated encoder. By fully utilizing the complementary semantics of audio and text, it overcomes the information limitations of single-modal evaluation and significantly improves the accuracy of keyframe extraction in complex semantic scenarios. An additional feature space alignment step maps the features of the three modalities to a unified semantic space, enabling accurate capture of multimodal semantic associations and resolving the feature misalignment problem in existing cross-modal fusion, thus strengthening the semantic consistency of keyframe selection. Temporal clustering aggregates video frames into temporal semantic stages, and stage-level saliency evaluation is performed using a multimodal concatenation vector and a temporal saliency aggregator. This fully considers the temporal dependencies of the video, avoids semantic fragmentation caused by independent frame evaluation, and ensures the semantic integrity and logical coherence of keyframes. Keyframe extraction is based on threshold-based selection of target semantic stages, ensuring the representativeness of keyframes while reducing redundant frame retention through stage aggregation, providing efficient and high-quality input support for downstream tasks such as video summarization and content retrieval.

[0046] In some optional implementations of this embodiment, the multimodal encoder includes a visual feature extraction model, an audio feature extraction model, and a text feature extraction model. Step 202 involves using the multimodal encoder to extract features from the visual data, audio data, and text data respectively, obtaining multiple visual feature vectors, multiple audio feature vectors, and multiple text feature vectors. Specifically, this includes the following steps: Each video frame of the visual data is segmented to obtain multiple image blocks. A visual feature extraction model is used to encode these image blocks, outputting a visual feature vector for each video frame, resulting in multiple visual feature vectors. Based on the temporal sequence of the video frames in the visual data, the audio data is segmented to obtain multiple audio segments. Each audio segment is converted into a spectrogram, resulting in multiple spectrograms. An audio feature extraction model is used to extract features from these spectrograms, resulting in multiple audio feature vectors. Finally, a text feature extraction model is used to encode the text data, resulting in multiple text feature vectors.

[0047] The visual feature extraction model is the part of the multimodal encoder responsible for processing visual data. Its core capability is to segment visual frames into image patches and encode them into visual feature vectors. The audio feature extraction model is the part of the multimodal encoder responsible for processing audio data. Its core capability is to extract features from audio spectrograms and convert them into audio feature vectors. The text feature extraction model is the part of the multimodal encoder responsible for processing text data. Its core capability is to semantically encode text sequences and convert them into text feature vectors.

[0048] In this system, multiple image patches are the basic units of a visual frame after spatial partitioning. Their data source is the uniform segmentation of the visual frame, and these image patches represent local region information of the visual frame. The core of encoding multiple image patches is to transform them into high-dimensional visual feature vectors, which are used to aggregate the local information of the visual frame into global semantic features. The video frame temporal sequence of the visual data refers to the chronological order of the video frames in the visual data. Its purpose is to provide a synchronization benchmark for the segmentation of audio data, ensuring a one-to-one correspondence between audio segments and visual frames.

[0049] The segmentation process involves dividing continuous audio data into multiple segments based on the temporal sequence of video frames in the visual data. The core of this operation is to synchronize the audio segments with the visual frames in time, resulting in audio segments that correspond one-to-one with each visual frame. Multiple audio segments are the result set after segmentation, with each segment corresponding to the accompanying audio of a visual frame (video frame). Multiple spectrograms are the results of frequency domain transformation of the audio segments, which can be in Mel spectrogram format. These spectrograms characterize the frequency domain features of the audio segments, such as frequency distribution and intensity variations. Encoding the text data involves transforming the text sequence into a high-dimensional text feature vector. This step aims to extract deep semantic information from the text data.

[0050] In one embodiment, taking the recorded video of the online course "Python Programming Basics - Loop Structures" as the target video, multimodal data is acquired, including visual data, audio data, and text data. The visual data consists of 360 video frames (25 frames / second, 14.4 seconds long) of the teacher's whiteboard writing and code demonstration. The audio data is the teacher's explanation, and the text data consists of 8 subtitles, such as "for loop syntax" and "range function usage." The visual branch of the multimodal encoder uses a visual feature extraction model to segment the frames into 16×16 image blocks for encoding, resulting in 360 768-dimensional visual feature vectors. The audio data is divided into 360 segments according to frame time sequence. After conversion to Mel spectrograms, the audio branch of the multimodal encoder uses an audio feature extraction model to obtain 360 256-dimensional audio feature vectors. The text branch uses a text feature extraction model to encode the subtitles, resulting in 8 768-dimensional text feature vectors. After feature alignment, the data is temporally clustered into three stages: "grammar explanation," "code demonstration," and "case practice." The multimodal features from each stage are then concatenated to obtain a concatenated vector. The saliency score of each temporal semantic stage is calculated by the temporal saliency aggregator. The saliency score is 0.56 for the "syntax explanation" stage, 0.91 for the "code demonstration" stage, and 0.65 for the "case practice" stage. The "code demonstration" stage is selected as the target temporal semantic stage with a threshold of 0.7. Five keyframes, including "loop structure code writing" and "running result display", are extracted from the "code demonstration" stage.

[0051] This application's embodiments segment each video frame of visual data into multiple image blocks and encode these image blocks using a visual feature extraction model. This allows for the detailed capture of visual feature information in local areas within the video frame. Compared to traditional methods that only extract overall visual features, this approach more accurately reflects the visual semantic content of the video frame, laying the foundation for accurate keyframe extraction. Based on the temporal sequence of the video frames, audio data is segmented and converted into a spectrogram. Then, an audio feature extraction model extracts feature vectors, closely linking audio features with the temporal sequence of the video frames, enabling the discovery of semantic changes in audio over time. A text feature extraction model is used to encode text data, obtaining text feature vectors that effectively extract the semantic information contained within the text. In summary, these operations fully utilize the characteristics of multimodal data, acquiring rich semantic information from different dimensions. This overcomes the limitations of existing technologies with their single modality, providing a comprehensive and accurate feature foundation for subsequent multimodal feature space alignment, temporal clustering, and other processing, thus improving the accuracy of keyframe extraction in complex scenes.

[0052] In some optional implementations, step 203 involves performing feature space alignment on multiple visual feature vectors, multiple audio feature vectors, and multiple text feature vectors to obtain aligned multiple target visual feature vectors, multiple target audio feature vectors, and multiple target text feature vectors. This specifically includes the following steps: For multiple visual feature vectors, multiple audio feature vectors, and multiple text feature vectors, visual feature vectors and audio feature vectors, visual feature vectors and text feature vectors, and audio feature vectors and text feature vectors with the same semantic content are treated as positive sample pairs, resulting in multiple visual-audio positive sample pairs, multiple visual-text positive sample pairs, and multiple audio-text positive sample pairs, respectively. Similarly, for multiple visual feature vectors, multiple audio feature vectors, and multiple text feature vectors, visual feature vectors and audio feature vectors, visual feature vectors and text feature vectors, and audio feature vectors and text feature vectors with different semantic content are treated as negative sample pairs, resulting in multiple visual-audio negative sample pairs, multiple visual-text negative sample pairs, and multiple... Audio-text negative sample pairs; based on multiple visual-audio positive sample pairs, multiple visual-text positive sample pairs, multiple audio-text positive sample pairs, multiple visual-audio negative sample pairs, multiple visual-text negative sample pairs, and multiple audio-text negative sample pairs, calculate the total cross-modal contrast loss; through the total cross-modal contrast loss, update the model parameters of the visual feature extraction model, audio feature extraction model, and text feature extraction model in reverse; use the updated model parameters of the visual feature extraction model, audio feature extraction model, and text feature extraction model to extract features from multiple image patches, multiple spectrograms, and text data, respectively, to obtain multiple aligned target visual feature vectors, multiple target audio feature vectors, and multiple target text feature vectors.

[0053] In cross-modal contrastive learning, positive sample pairs are pairings of feature vectors from different modalities containing the same semantic content. These pairs originate from semantically related visual, audio, and textual information within the target video. They represent the semantic matching relationship between multiple modalities and guide the model to bring multimodal features with the same semantic meaning closer together in space during loss calculation. Multiple visual-audio positive sample pairs are sets of positive sample pairs consisting of visual and audio feature vectors, with each pair corresponding to the same semantic content in the target video. They support contrastive learning between visual and audio modalities. Multiple visual-text positive sample pairs are sets of positive sample pairs consisting of visual and text feature vectors, with each pair corresponding to the same semantic content in the target video. They support contrastive learning between visual and text modalities, ensuring the model learns the semantic consistency between visual images and textual descriptions. Multiple audio-text positive sample pairs are sets of positive sample pairs consisting of audio and text feature vectors, with each pair corresponding to the same semantic content in the target video. These pairs originate from semantic matching and filtering of audio and text feature vectors. They support contrastive learning between audio and text modalities, guiding the model to optimize audio and text feature extraction model parameters.

[0054] In cross-modal contrastive learning, negative sample pairs are pairings composed of feature vectors from different modalities with different semantic content. They originate from semantically unrelated visual, audio, and textual information in the target video. They represent semantic mismatches between multimodalities, and their core purpose is to guide the model to keep multimodal features with different semantics spatially separate during loss calculation. Multiple visual-audio negative sample pairs are sets of negative sample pairs composed of visual and audio feature vectors. Each pair corresponds to different semantic content in the target video, and their origin is the selection of semantic differences between visual and synchronized audio feature vectors. Multiple visual-text negative sample pairs are sets of negative sample pairs composed of visual and text feature vectors. Each pair corresponds to different semantic content in the target video, and their origin is the matching of semantic differences between visual and text feature vectors. These are used to support contrastive learning between visual and text modalities, preventing the model from mistakenly associating irrelevant visual images with textual descriptions. Multiple audio-text negative sample pairs are sets of negative sample pairs composed of audio and text feature vectors. Each pair corresponds to different semantic content in the target video. Used to support comparative learning of audio and text modalities, ensuring that the model can distinguish audio and text features with different semantics.

[0055] The total cross-modal contrastive loss is a comprehensive loss calculated using the cross-modal contrastive loss function based on three classes of positive / negative sample pairs: visual and audio, visual and text, and audio and text. The model parameters are learnable variables from the visual, audio, and text feature extraction models. Initial values ​​are pre-trained weights or random values, and they are dynamically updated subsequently through backpropagation of the total cross-modal contrastive loss.

[0056] In one example, we take the keyframe extraction task from a cooking tutorial video as an example. The video is 10 minutes long and includes visual data of the cooking process, audio data of the narration explaining the cooking steps, and text data of the video subtitles and step descriptions. First, the visual feature extraction model extracts features from each frame, capturing visual information such as ingredient preparation and cooking actions. The audio feature extraction model analyzes the time points of the narration to identify the key steps being explained. The text feature extraction model processes the subtitles to obtain semantic information. Next, multimodal positive and negative sample pairs are constructed, and the total cross-modal contrastive loss is calculated to update the model parameters in reverse. The updated model re-extracts features, aligning the visual cooking actions, audio narration, and text descriptions in the feature space. Finally, a temporal saliency aggregator synthesizes the information and selects frames showing the ingredients, key operations, and close-ups of the dish from the ingredient preparation, cooking, and finished product display stages as keyframes. Compared to traditional single-modal methods, the extracted keyframes are more representative and coherent.

[0057] This application's embodiments optimize the model by carefully constructing positive and negative sample pairs and calculating the total cross-modal contrastive loss, achieving significant technical effects. By forming positive and negative sample pairs from multimodal feature vectors with the same and different semantic content, the semantic relationships between multimodal features can be comprehensively captured. Calculating the total cross-modal contrastive loss quantifies the alignment degree of different modal features in space. This loss is used to update the model parameters in reverse, enabling visual, audio, and text feature extraction models to adjust their parameters according to semantic relevance, thereby extracting features more accurately. Finally, using the parameter-updated model to extract features effectively achieves the alignment of multimodal features in a unified space, overcoming the problem of missing multimodal feature alignment in existing technologies.

[0058] In some optional implementations, the step "calculate the total cross-modal contrast loss based on multiple visual-audio positive sample pairs, multiple visual-text positive sample pairs, multiple audio-text positive sample pairs, multiple visual-audio negative sample pairs, multiple visual-text negative sample pairs, and multiple audio-text negative sample pairs" specifically includes the following steps: For each positive visual-audio sample pair, calculate the first feature inner product of each positive visual-audio sample pair and the second feature inner product of the corresponding negative visual-audio sample pair. Based on the first and second feature inner products, calculate the visual-audio loss for each positive visual-audio sample pair. Sum all visual-audio losses to obtain the total visual-audio loss. For each positive visual-text sample pair, calculate the third feature inner product of each positive visual-text sample pair and the fourth feature inner product of the corresponding negative visual-text sample pair. Based on the third and fourth feature inner products, calculate the total visual-audio loss for each positive visual-text sample pair. For each pair of positive visual text samples, calculate the visual text loss corresponding to each pair. Sum all the visual text losses to obtain the total visual text loss. For each pair of positive audio text samples, calculate the inner product of the fifth feature of each pair and the inner product of the sixth feature of each pair of negative audio text samples. Based on the inner products of the fifth and sixth features, calculate the audio text loss corresponding to each pair of positive audio text samples. Sum all the audio text losses to obtain the total audio text loss. Fuse the total visual-audio loss, the total visual text loss, and the total audio text loss to obtain the total cross-modal contrast loss.

[0059] The first feature inner product refers to the vector inner product of the visual and audio feature vectors in a positive visual-audio sample pair. It originates from the dot product of visual and audio feature vectors with the same semantic meaning and represents the semantic association of each positive sample pair in the feature space. A larger inner product value indicates a higher semantic matching degree. The second feature inner product refers to the vector inner product of the visual and audio feature vectors in a corresponding negative visual-audio sample pair. It originates from the dot product of visual and audio feature vectors with different semantic meanings and represents the semantic association of each negative sample pair in the feature space. A smaller inner product value indicates a higher semantic discriminative degree. The visual-audio loss is the loss value calculated for a single positive visual-audio sample pair based on the first and second feature inner products using a cross-modal contrastive loss function. The total visual-audio loss is the comprehensive loss value obtained by summing the visual-audio losses corresponding to all positive visual-audio sample pairs. It represents the overall alignment deviation of all positive sample pairs in the visual and audio modalities and is used as the loss optimization target for the visual and audio modalities.

[0060] The third feature inner product refers to the vector inner product of the visual feature vector and the text feature vector in a positive visual-text sample pair. It originates from the dot product of visual and text feature vectors with the same semantic meaning and represents the semantic correlation of each positive sample pair in the feature space. The fourth feature inner product refers to the vector inner product of the visual feature vector and the text feature vector in the corresponding negative visual-text sample pair. It originates from the dot product of visual and text feature vectors with different semantic meanings and represents the semantic correlation of each negative sample pair in the feature space. The visual-text loss refers to the loss value calculated for a single positive visual-text sample pair based on the third and fourth feature inner products using the same contrastive loss function as the visual-audio loss. It originates from the inner product operation of the positive sample pair and the corresponding negative sample pair, followed by logarithmic normalization, and represents the alignment deviation of the positive sample pair in the visual and text modalities. The total visual text loss refers to the comprehensive loss value obtained by summing the visual text losses of all positive visual text sample pairs. It is derived from the sum of the losses of all individual visual text samples and represents the overall alignment deviation of all positive sample pairs in the visual and text modalities. It is used as the loss optimization target for the visual and text modalities.

[0061] The fifth feature inner product refers to the inner product of the audio feature vector and the text feature vector in a positive audio-text sample pair, representing the semantic correlation of each positive sample pair in the feature space. The sixth feature inner product refers to the inner product of the audio feature vector and the text feature vector in a negative audio-text sample pair corresponding to a positive audio-text sample pair, representing the semantic correlation of each negative sample pair in the feature space. The audio-text loss is the loss value calculated for a single positive audio-text sample pair based on the fifth and sixth feature inner products using the same contrastive loss function. The total audio-text loss is the comprehensive loss value obtained by summing the audio-text losses corresponding to all positive audio-text sample pairs, representing the overall alignment deviation of all positive sample pairs in the audio and text modalities.

[0062] In one example, taking the feature space alignment processing of the "Coffee Latte Art Heart Pattern Tutorial" video as an example, the total cross-modal contrast loss is calculated as follows: First, the visual-audio modality is processed: For a certain visual-audio positive sample pair corresponding to "tilting the milk pitcher", its visual feature vector V is taken. i With audio feature vector a j Calculate the first feature inner product; then take the second feature inner product of each of the five negative visual-audio sample pairs corresponding to the positive sample pair. Based on formula (1): The visual audio loss of the positive sample is calculated. This is the temperature hyperparameter (which can be set to 0.07), and T is the vector dimension of the visual feature vector. First, calculate the molecular exp(first inner product / ), in calculating the denominator exp[(first inner product / ) + exp(second inner product / The logarithm of the obtained value is then negative to obtain the visual-audio loss for that positive sample pair. The losses of 200 positive visual-audio sample pairs are summed to obtain the total visual-audio loss. Similarly, the total visual-text loss and total audio-text loss are calculated using the same method. Finally, the total visual-audio loss, the total visual-text loss, and the total audio-text loss are added together to obtain the total cross-modal contrast loss.

[0063] This application's embodiments accurately quantify the semantic correlation between positive / negative sample pairs in different modalities by calculating the inner product of features in each modality. This avoids the fuzzy representation of cross-modal feature correlation, providing a more granular semantic basis for loss calculation and improving the accuracy of the loss value in characterizing multimodal alignment deviations. Calculating the single-sample loss in each modality and summing them to obtain the total modal loss enables independent quantification of the alignment effect for each modality. This preserves the optimization requirements specific to each modality and provides a clear modality-level optimization objective for subsequent multimodal loss fusion. The fusion design of the total multimodal loss achieves collaborative optimization of alignments between visual and audio, visual and text, and audio and text modalities, ensuring consistent parameter update directions for feature extraction models across different modalities.

[0064] In some optional implementations, step 204 involves performing temporal clustering on video frames of the visual data based on multiple target visual feature vectors, multiple target audio feature vectors, and multiple target text feature vectors to determine multiple temporal semantic stages and the multimodal concatenation vector corresponding to each temporal semantic stage. This specifically includes the following steps: Based on multiple target visual feature vectors, the visual feature similarity between the target visual feature vector of each video frame in the visual data and the target visual feature vectors of adjacent video frames is calculated. Based on multiple target audio feature vectors, the audio feature similarity between the target audio feature vector of each video frame and the target audio feature vectors of adjacent video frames is calculated. Based on multiple target text feature vectors, the text feature similarity between the target text feature vector of each video frame and the target text feature vectors of adjacent video frames is calculated. Based on visual feature similarity, audio feature similarity, and text feature similarity, multiple temporal semantic stages are determined. The target visual feature vector, target audio feature vector, and target text feature vector corresponding to each temporal semantic stage are obtained. The target visual feature vector, target audio feature vector, and target text feature vector corresponding to each temporal semantic stage are concatenated to obtain the multimodal concatenation vector corresponding to each temporal semantic stage.

[0065] Visual feature similarity refers to the quantified semantic association between the target visual feature vector of each video frame and the target visual feature vectors of adjacent video frames. It can be calculated using metrics such as cosine similarity and Euclidean distance. Since these vectors are aligned in a unified semantic space, the similarity value accurately reflects the semantic continuity of visual content in adjacent frames. For example, consecutive frames of "coffee latte art with a tilted milk pitcher" will maintain a high level of visual feature similarity.

[0066] Audio feature similarity refers to the quantified value of the semantic association between the target audio feature vector corresponding to each video frame and the target audio feature vector corresponding to adjacent video frames. It can also be calculated using indicators such as cosine similarity.

[0067] Among them, text feature similarity refers to the quantification value of the semantic association between the target text feature vector corresponding to each video frame and the target text feature vector corresponding to adjacent video frames, which can be calculated by indicators such as cosine similarity.

[0068] In one example, keyframe extraction from an online financial training video in the fintech field is used. The system receives multimodal data from a 10-minute online financial training video, including visual data (such as the instructor's image and PPT slides), audio data (the instructor's voice), and text data (text information and subtitles on the PPT). A pre-trained multimodal encoder is used to extract features from the visual, audio, and text data respectively. For visual data, multiple visual feature vectors are extracted using a convolutional neural network; for audio data, multiple audio feature vectors are obtained using an audio feature extraction model; and for text data, multiple text feature vectors are obtained through a natural language processing model. These feature vectors are then aligned in a feature space to ensure effective alignment of features from different modalities within a unified space. Next, based on the obtained target visual feature vectors, the visual feature similarity between the target visual feature vector of each video frame and the target visual feature vectors of adjacent frames is calculated; similarly, the audio and text feature similarities are calculated. Based on these similarities, multiple temporal semantic stages are determined; for example, explaining "stock investment strategies" is one stage, and explaining "fund classification" is another. Obtain the target feature vector corresponding to each stage and concatenate them into a multimodal concatenation vector.

[0069] This application's embodiments calculate the visual feature similarity between adjacent video frames based on multiple target visual feature vectors, enabling precise capture of temporal changes at the visual level. Similarly, similarity calculations on audio and text feature vectors comprehensively uncover the temporal semantic relationships between audio and text. Combining the similarity of these three elements to determine multiple temporal semantic stages fully considers the temporal coherence of multimodal data. Obtaining the target feature vectors corresponding to each stage and concatenating them into a multimodal concatenation vector achieves deep fusion and effective aggregation of multimodal features at the temporal semantic stage.

[0070] In some optional implementations, step S205, based on the multimodal concatenation vector of each temporal semantic stage, calculates the saliency score of each temporal semantic stage using a temporal saliency aggregator, specifically including the following steps: Based on the multimodal concatenation vector of each temporal semantic stage, the degree of association between each temporal semantic stage and all stages is calculated using the query projection matrix and key projection matrix within the temporal saliency aggregator, thus obtaining the self-attention weight. Based on the self-attention weight and the multimodal concatenation vectors of all temporal semantic stages, the saliency score of each temporal semantic stage is calculated using the saliency feature projection matrix and activation function of the temporal saliency aggregator.

[0071] The query projection matrix is ​​a learnable parameter matrix used in the temporal saliency aggregator to map multimodal concatenation vectors to "query vectors," and its dimension is the same as the dimension of the multimodal concatenation vectors. It is used to perform a linear transformation on the input temporal semantic stage multimodal concatenation vectors, generating a "query" representation for subsequent calculation of association degree. The parameters of this matrix are optimized through model training to accurately capture the semantic information in the multimodal concatenation vectors.

[0072] The key projection matrix, in the temporal saliency aggregator, is a learnable parameter matrix that works in conjunction with the query projection matrix to map multimodal concatenated vectors into "key vectors." It is used to perform a linear transformation on the multimodal concatenated vectors at each temporal semantic stage, generating a "key" representation. The inner product operation between the query vector and the key vector quantifies the semantic association between different stages, while parameter optimization of the key projection matrix allows the key vectors to more accurately match the semantic features of the query vector.

[0073] Among them, the self-attention weight is a numerical value that quantifies the degree of semantic association between each temporal semantic stage and all stages (including itself) based on the output of the query projection matrix and the key projection matrix.

[0074] The saliency feature projection matrix is ​​a learnable parameter matrix used in the temporal saliency aggregator to map the "self-attention-weighted multimodal features" to saliency features. Its dimension is usually the same as that of the multimodal concatenation vector. It is used to perform a linear transformation on the "self-attention-weighted sum of multimodal concatenation vectors across all stages" to extract features relevant to saliency evaluation.

[0075] The activation function is a non-linear function used in the temporal saliency aggregator to compress the output of the saliency feature projection matrix to a specific range. It is used to transform high-dimensional feature vectors into normalized scores that can be directly used for "critical / non-critical" judgments.

[0076] In one example, in a keyframe extraction scenario from financial news videos in the fintech field, multimodal data (visual, audio, and text) of the video is received. After feature extraction and alignment, multiple temporal semantic stages and their corresponding multimodal concatenation vectors are determined. Taking temporal saliency aggregator processing as an example, based on the multimodal concatenation vector of each stage, the correlation between each temporal semantic stage and all stages is calculated using the internal query projection matrix and key projection matrix, resulting in self-attention weights. Then, based on the self-attention weights and the multimodal concatenation vectors of all stages, the saliency score of each stage is calculated using the saliency feature projection matrix and an activation function. For example, when analyzing stock trend explanation videos, this method can accurately capture key information stages.

[0077] This application's embodiments, based on the multimodal concatenation vector of each temporal semantic stage, utilize the query projection matrix and key projection matrix within the temporal saliency aggregator to calculate the correlation with all stages to obtain self-attention weights. This process fully considers the intrinsic connections between different temporal semantic stages, accurately capturing the dependencies of video temporal semantics. Then, based on the self-attention weights and multimodal concatenation vectors, saliency scores are calculated using the saliency feature projection matrix and activation function, effectively aggregating multimodal information and temporal features. This allows keyframe extraction to comprehensively consider multimodal semantics and temporal coherence, significantly improving the accuracy and semantic completeness of keyframe extraction.

[0078] In some optional implementations, after step 206, where the temporal semantic stage with a saliency score greater than or equal to a preset score threshold is taken as the target temporal semantic stage, and multiple key video frames are obtained from the multiple video frames corresponding to the target temporal semantic stage, the following steps are further included: Obtain the multimodal stitching vector corresponding to each key video frame; construct a support set and a query set based on the multimodal stitching vector corresponding to each key video frame. The support set includes multimodal stitching vectors with semantic labels; for the support set, take the mean of all labeled multimodal stitching vectors with semantic labels as key frames to obtain the key frame prototype vector, and take the mean of all labeled multimodal stitching vectors with semantic labels as non-key frames to obtain the non-key frame prototype vector; calculate the key frame prototype similarity between each multimodal stitching vector and the key frame prototype vector in the query set, and perform exponential normalization on the key frame prototype similarity to obtain the key frame probability; calculate the non-key frame prototype similarity between each multimodal stitching vector and the non-key frame prototype vector in the query set, and perform exponential normalization on the non-key frame prototype similarity to obtain the non-key frame probability; based on the true label, key frame probability, and non-key frame probability of each multimodal stitching vector in the query set, calculate the key frame prototype classification loss to update the parameters of the multimodal encoder and the temporal saliency aggregator.

[0079] The support set is a set of "reference samples" that are divided from the multimodal splicing vectors corresponding to key video frames, and is used to provide labeled benchmark data for subsequent prototype vector construction.

[0080] The query set is a set of multimodal spliced ​​vectors to be evaluated, corresponding to the support set. Its source is also the multimodal spliced ​​vectors corresponding to key video frames. These are used as "samples to be judged," and their probability of belonging to a key frame or non-key frame is obtained by calculating their similarity with the prototype vectors constructed from the support set.

[0081] Among them, the semantically labeled multimodal concatenation vector is the core element in the support set, referring to multimodal concatenation vectors with attached "keyframe / non-keyframe" semantic labels. The semantic label "keyframe" is one of the label attributes of the semantically labeled multimodal concatenation vector, used to identify that the temporal semantic stage corresponding to the vector belongs to a core, highly important stage in the video (i.e., the stage corresponding to the keyframe). The keyframe prototype vector is a feature vector obtained by taking the average value of all multimodal concatenation vectors with "semantic label: keyframe" in the support set.

[0082] The semantic label "non-keyframe" is another label attribute of the semantically labeled multimodal concatenation vector, used to identify that the temporal semantic stage corresponding to the vector belongs to a transitional or low-importance stage in the video (i.e., the stage corresponding to a non-keyframe). The non-keyframe prototype vector is a feature vector obtained by taking the average value of all multimodal concatenation vectors with "semantic label 'non-keyframe'" in the support set.

[0083] Keyframe prototype similarity is a quantified value of the semantic association between the multimodal concatenation vector in the query set and the keyframe prototype vector, which can be calculated using metrics such as cosine similarity and Euclidean distance. Keyframe probability is a value obtained by exponentially normalizing the keyframe prototype similarity, ranging from 0 to 1.

[0084] Among them, non-keyframe prototype similarity is a quantified value of the semantic association between the multimodal concatenation vector and the non-keyframe prototype vector in the query set, and its calculation method is the same as that of keyframe prototype similarity. Non-keyframe probability is a value obtained by exponentially normalizing the non-keyframe prototype similarity, and its range is 0 to 1.

[0085] The ground truth label is the actual semantic attribute label (keyframe / non-keyframe) corresponding to the multimodal concatenation vector in the query set, and serves as the benchmark for evaluating the prediction accuracy of "keyframe probability / non-keyframe probability". The keyframe prototype classification loss is a loss value calculated based on the ground truth label, keyframe probability, and non-keyframe probability of the query set vector. It is used to quantify the deviation between the predicted probability and the ground truth label. The smaller the loss value, the closer the prediction result is to the actual situation.

[0086] In one example, in keyframe extraction from investment lecture videos in the fintech field, the multimodal concatenation vectors corresponding to the key video frames obtained through preprocessing are first acquired. Based on this, a support set and a query set are constructed. Vectors in the support set carry semantic labels such as "keyframe" and "non-keyframe." For the support set, the mean of the labeled multimodal concatenation vectors with semantic labels "keyframe" and "non-keyframe" is calculated separately to obtain keyframe prototype vectors and non-keyframe prototype vectors. Next, the similarity between each vector in the query set and these two prototype vectors is calculated, and exponential normalization is applied to obtain the keyframe probability and non-keyframe probability. For example, when analyzing a stock investment strategy lecture video, this method allows for more accurate judgment based on the semantic characteristics of the video content. The keyframe prototype classification loss is calculated based on the true labels of the vectors in the query set and the two probabilities mentioned above, thereby updating the parameters of the multimodal encoder and the temporal saliency aggregator. This allows the model to better adapt to the complex semantic scenarios of financial lecture videos, improve the accuracy of keyframe extraction, provide higher-quality input for subsequent video summarization, and facilitate the efficient dissemination and utilization of financial knowledge.

[0087] This application's embodiments construct a support and query set by concatenating keyframe multimodal vectors. Based on semantically labeled vectors, the average value of similar labeled vectors is used to obtain keyframe / non-keyframe prototype vectors. This design allows the prototype vectors to accurately aggregate the semantic features of key / non-keyframes, providing a unified benchmark for subsequent similarity evaluation. The similarity between the query vector and the two types of prototype vectors is calculated and normalized to obtain the probability. This probability is then combined with the real labels to calculate the classification loss and back-update the model parameters. This makes the features extracted by the multimodal encoder more closely match the semantic patterns of keyframes, while also allowing the attention weights of the temporal saliency aggregator to more accurately capture inter-frame dependencies. This process relies on meta-learning prototype optimization, strengthening the model's semantic discrimination ability in low-sample scenarios. It improves the accuracy of keyframe probability prediction and further optimizes the consistency and generalization of cross-modal features, ultimately making keyframe extraction more adaptable to low-sample scenarios while ensuring the coherence of temporal semantics.

[0088] It should be emphasized that, in order to further ensure the privacy and security of the aforementioned visual data, audio data, text data, saliency scores, and multiple key video frames, these data can also be stored in a blockchain node.

[0089] The blockchain referred to in this application is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.

[0090] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0091] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).

[0092] Further reference Figure 3 As a response to the above Figure 2 The implementation of the method shown in this application provides an embodiment of a video keyframe extraction device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0093] like Figure 3 As shown, the video keyframe extraction device 400 of this embodiment includes: a receiving module 401, a feature extraction module 402, an alignment module 403, a clustering module 404, a calculation module 405, and an acquisition module 406. Wherein: The receiving module 401 is used to receive multimodal data of the target video, including visual data, audio data and text data; The feature extraction module 402 is used to extract features from visual data, audio data and text data respectively using a multimodal encoder to obtain multiple visual feature vectors, multiple audio feature vectors and multiple text feature vectors; Alignment module 403 is used to perform feature space alignment processing on multiple visual feature vectors, multiple audio feature vectors and multiple text feature vectors to obtain multiple aligned target visual feature vectors, multiple target audio feature vectors and multiple target text feature vectors. Clustering module 404 is used to perform temporal clustering of video frames of visual data based on multiple target visual feature vectors, multiple target audio feature vectors and multiple target text feature vectors, to determine multiple temporal semantic stages and the multimodal concatenation vector corresponding to each temporal semantic stage. The calculation module 405 is used to calculate the saliency score of each temporal semantic stage based on the multimodal concatenation vector of each temporal semantic stage through a temporal saliency aggregator. The acquisition module 406 is used to take the temporal semantic stage with a saliency score greater than or equal to a preset score threshold as the target temporal semantic stage, and acquire multiple key video frames from multiple video frames corresponding to the target temporal semantic stage.

[0094] This application's embodiments simultaneously receive visual, audio, and text data and extract features separately using a dedicated encoder. By fully utilizing the complementary semantics of audio and text, it overcomes the information limitations of single-modal evaluation and significantly improves the accuracy of keyframe extraction in complex semantic scenarios. An additional feature space alignment step maps the features of the three modalities to a unified semantic space, enabling accurate capture of multimodal semantic associations and resolving the feature misalignment problem in existing cross-modal fusion, thus strengthening the semantic consistency of keyframe selection. Temporal clustering aggregates video frames into temporal semantic stages, and stage-level saliency evaluation is performed using a multimodal concatenation vector and a temporal saliency aggregator. This fully considers the temporal dependencies of the video, avoids semantic fragmentation caused by independent frame evaluation, and ensures the semantic integrity and logical coherence of keyframes. Keyframe extraction is based on threshold-based selection of target semantic stages, ensuring the representativeness of keyframes while reducing redundant frame retention through stage aggregation, providing efficient and high-quality input support for downstream tasks such as video summarization and content retrieval.

[0095] In one embodiment, the feature extraction module 402 includes: The cutting submodule is used to cut each video frame of the visual data into multiple image blocks; The first encoding submodule is used to encode multiple image blocks using a visual feature extraction model, output the visual feature vector of each video frame, and obtain multiple visual feature vectors. The segmentation submodule is used to segment audio data based on the video frame timing of visual data, resulting in multiple audio segments. The first extraction submodule is used to convert each audio segment into a spectrogram, resulting in multiple spectrograms. Then, through an audio feature extraction model, feature extraction is performed on the multiple spectrograms to obtain multiple audio feature vectors. The second encoding submodule is used to encode text data using a text feature extraction model to obtain multiple text feature vectors.

[0096] In one embodiment, the alignment module 403 includes: The first submodule is used to take multiple visual feature vectors, multiple audio feature vectors and multiple text feature vectors as positive sample pairs, respectively, and to obtain multiple visual-audio positive sample pairs, multiple visual-text positive sample pairs and multiple audio-text positive sample pairs. The second submodule is used to take visual feature vectors and audio feature vectors, visual feature vectors and text feature vectors, and audio feature vectors and text feature vectors with different semantic content as negative sample pairs to obtain multiple visual-audio negative sample pairs, multiple visual-text negative sample pairs, and multiple audio-text negative sample pairs, respectively. The first calculation submodule is used to calculate the total cross-modal contrast loss based on multiple visual-audio positive sample pairs, multiple visual-text positive sample pairs, multiple audio-text positive sample pairs, multiple visual-audio negative sample pairs, multiple visual-text negative sample pairs, and multiple audio-text negative sample pairs. The update submodule is used to update the model parameters of the visual feature extraction model, audio feature extraction model, and text feature extraction model in reverse using the total cross-modal contrast loss. The second extraction submodule is used to extract features from multiple image blocks, multiple spectrograms and text data using the updated visual feature extraction model, audio feature extraction model and text feature extraction model, respectively, to obtain multiple aligned target visual feature vectors, multiple target audio feature vectors and multiple target text feature vectors.

[0097] In one embodiment, the first calculation submodule is further configured to: calculate, for each visual-audio positive sample pair, a first feature inner product of each visual-audio positive sample pair and a second feature inner product of the visual-audio negative sample pair corresponding to each visual-audio positive sample pair; calculate the visual-audio loss corresponding to each visual-audio positive sample pair based on the first and second feature inner products, and sum all visual-audio losses to obtain the total visual-audio loss; calculate, for each visual-text positive sample pair, a third feature inner product of each visual-text positive sample pair and a fourth feature inner product of the visual-text negative sample pair corresponding to each visual-text positive sample pair; and calculate, based on the third feature inner product and the second feature inner product, calculate the visual-audio loss corresponding to each visual-audio positive sample pair and sum all visual-audio losses to obtain the total visual-audio loss; calculate, for each visual-text positive sample pair, a third feature inner product of each visual-text positive sample pair and a fourth feature inner product of the visual-text negative sample pair corresponding to each visual-text positive sample pair; and sum all visual-audio losses to obtain the total visual-audio loss. The four-feature inner product is used to calculate the visual text loss for each positive visual text sample pair. All visual text losses are summed to obtain the total visual text loss. For each positive audio text sample pair, the fifth feature inner product and the sixth feature inner product of the corresponding negative audio text sample pair are calculated. Based on the fifth and sixth feature inner products, the audio text loss for each positive audio text sample pair is calculated. All audio text losses are summed to obtain the total audio text loss. The total visual-audio loss, the total visual text loss, and the total audio text loss are fused to obtain the total cross-modal contrast loss.

[0098] In one embodiment, the clustering module 404 includes: The second calculation submodule is used to calculate the visual feature similarity between the target visual feature vector of each video frame in the visual data and the target visual feature vector of the adjacent video frames, based on multiple target visual feature vectors. The third calculation submodule is used to calculate the audio feature similarity between the target audio feature vector corresponding to each video frame and the adjacent target audio feature vectors based on multiple target audio feature vectors. The fourth calculation submodule is used to calculate the text feature similarity between the target text feature vector corresponding to each video frame and the adjacent target text feature vectors based on multiple target text feature vectors. The determination submodule is used to determine multiple temporal semantic stages based on visual feature similarity, audio feature similarity, and text feature similarity; The acquisition submodule is used to acquire the target visual feature vector, target audio feature vector, and target text feature vector corresponding to each temporal semantic stage; The splicing submodule is used to splice the target visual feature vector, target audio feature vector, and target text feature vector corresponding to each temporal semantic stage to obtain the multimodal splicing vector corresponding to each temporal semantic stage.

[0099] In one embodiment, the calculation module 405 includes: The fifth calculation submodule is used to calculate the degree of association between each temporal semantic stage and all stages based on the multimodal concatenation vector of each temporal semantic stage, through the query projection matrix and key projection matrix in the temporal saliency aggregator, and obtain the self-attention weight; The sixth computational submodule is used to calculate the saliency score of each temporal semantic stage based on the self-attention weights and the multimodal concatenation vectors of all temporal semantic stages, through the saliency feature projection matrix of the temporal saliency aggregator and the activation function.

[0100] In one embodiment, the video keyframe extraction device 400 further includes: The vector acquisition module is used to acquire the multimodal stitching vector corresponding to each key video frame; The building module is used to construct a support set and a query set based on the multimodal splicing vector corresponding to each key video frame. The support set includes multimodal splicing vectors with semantic labels. The module is used to obtain the keyframe prototype vector by taking the mean of all labeled multimodal concatenation vectors with semantic labels as keyframes for the support set, and to obtain the non-keyframe prototype vector by taking the mean of all labeled multimodal concatenation vectors with semantic labels as non-keyframes. The first similarity calculation module is used to calculate the key frame prototype similarity between each multimodal splicing vector and the key frame prototype vector in the query set, and to perform exponential normalization on the key frame prototype similarity to obtain the key frame probability. The second similarity calculation module is used to calculate the non-key frame prototype similarity between each multimodal concatenation vector and non-key frame prototype vector in the query set, and to perform exponential normalization on the non-key frame prototype similarity to obtain the non-key frame probability. The loss calculation module is used to calculate the keyframe prototype classification loss based on the true label, keyframe probability and non-keyframe probability of each multimodal concatenation vector in the query set, so as to update the parameters of the multimodal encoder and the temporal saliency aggregator.

[0101] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.

[0102] Computer device 6 includes a memory 61, a processor 62, and a network interface 63 that are interconnected via a system bus. It should be noted that only computer device 6 with memory 61, processor 62, and network interface 63 is shown in the figure; however, it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described herein is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0103] Computer devices can include desktop computers, laptops, handheld computers, and cloud servers. These devices allow for human-computer interaction with users through keyboards, mice, remote controls, touchpads, or voice-activated devices.

[0104] The memory 61 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 61 may be an internal storage unit of the computer device 6, such as the hard disk or memory of the computer device 6. In other embodiments, the memory 61 may also be an external storage device of the computer device 6, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 6. Of course, the memory 61 may include both the internal storage unit and the external storage device of the computer device 6. In this embodiment, the memory 61 is typically used to store the operating system and various application software installed on the computer device 6, such as computer-readable instructions for video keyframe extraction methods. In addition, the memory 61 may also be used to temporarily store various types of data that have been output or will be output.

[0105] In some embodiments, processor 62 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. This processor 62 is typically used to control the overall operation of the computer device 6. In this embodiment, processor 62 is used to execute computer-readable instructions stored in memory 61 or to process data, such as computer-readable instructions for executing a video keyframe extraction method.

[0106] The network interface 63 may include a wireless network interface or a wired network interface, which is typically used to establish a communication connection between the computer device 6 and other electronic devices.

[0107] This application's embodiments simultaneously receive visual, audio, and text data and extract features separately using a dedicated encoder. By fully utilizing the complementary semantics of audio and text, it overcomes the information limitations of single-modal evaluation and significantly improves the accuracy of keyframe extraction in complex semantic scenarios. An additional feature space alignment step maps the features of the three modalities to a unified semantic space, enabling accurate capture of multimodal semantic associations and resolving the feature misalignment problem in existing cross-modal fusion, thus strengthening the semantic consistency of keyframe selection. Temporal clustering aggregates video frames into temporal semantic stages, and stage-level saliency evaluation is performed using a multimodal concatenation vector and a temporal saliency aggregator. This fully considers the temporal dependencies of the video, avoids semantic fragmentation caused by independent frame evaluation, and ensures the semantic integrity and logical coherence of keyframes. Keyframe extraction is based on threshold-based selection of target semantic stages, ensuring the representativeness of keyframes while reducing redundant frame retention through stage aggregation, providing efficient and high-quality input support for downstream tasks such as video summarization and content retrieval.

[0108] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the video keyframe extraction method described above.

[0109] This application's embodiments simultaneously receive visual, audio, and text data and extract features separately using a dedicated encoder. By fully utilizing the complementary semantics of audio and text, it overcomes the information limitations of single-modal evaluation and significantly improves the accuracy of keyframe extraction in complex semantic scenarios. An additional feature space alignment step maps the features of the three modalities to a unified semantic space, enabling accurate capture of multimodal semantic associations and resolving the feature misalignment problem in existing cross-modal fusion, thus strengthening the semantic consistency of keyframe selection. Temporal clustering aggregates video frames into temporal semantic stages, and stage-level saliency evaluation is performed using a multimodal concatenation vector and a temporal saliency aggregator. This fully considers the temporal dependencies of the video, avoids semantic fragmentation caused by independent frame evaluation, and ensures the semantic integrity and logical coherence of keyframes. Keyframe extraction is based on threshold-based selection of target semantic stages, ensuring the representativeness of keyframes while reducing redundant frame retention through stage aggregation, providing efficient and high-quality input support for downstream tasks such as video summarization and content retrieval.

[0110] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of this application.

[0111] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.

[0112] The software tools or components not belonging to our company that appear in the embodiments of this application are merely examples and do not represent actual use.

Claims

1. A method for extracting keyframes from a video, characterized in that, Includes the following steps: Receive multimodal data of the target video, wherein the multimodal data includes visual data, audio data, and text data; A multimodal encoder is used to extract features from the visual data, the audio data, and the text data respectively, resulting in multiple visual feature vectors, multiple audio feature vectors, and multiple text feature vectors. The plurality of visual feature vectors, the plurality of audio feature vectors, and the plurality of text feature vectors are subjected to feature space alignment processing to obtain the aligned plurality of target visual feature vectors, the plurality of target audio feature vectors, and the plurality of target text feature vectors. Based on multiple target visual feature vectors, multiple target audio feature vectors, and multiple target text feature vectors, temporal clustering is performed on the video frames of the visual data to determine multiple temporal semantic stages and the multimodal concatenation vector corresponding to each temporal semantic stage; Based on the multimodal concatenation vector of each temporal semantic stage, the saliency score of each temporal semantic stage is calculated by a temporal saliency aggregator; The temporal semantic stage with a saliency score greater than or equal to a preset score threshold is taken as the target temporal semantic stage, and multiple key video frames are obtained from the multiple video frames corresponding to the target temporal semantic stage.

2. The method according to claim 1, characterized in that, The multimodal encoder includes a visual feature extraction model, an audio feature extraction model, and a text feature extraction model. The step of using the multimodal encoder to extract features from the visual data, the audio data, and the text data respectively to obtain multiple visual feature vectors, multiple audio feature vectors, and multiple text feature vectors specifically includes: Each video frame of the visual data is segmented to obtain multiple image blocks; The visual feature extraction model is used to encode the multiple image blocks and output the visual feature vector of each video frame to obtain multiple visual feature vectors. Based on the video frame timing of the visual data, the audio data is segmented to obtain multiple audio segments; Each audio segment is converted into a spectrogram, resulting in multiple spectrograms. The audio feature extraction model is then used to extract features from these multiple spectrograms, resulting in multiple audio feature vectors. The text feature extraction model is used to encode the text data to obtain multiple text feature vectors.

3. The method according to claim 2, characterized in that, The step of performing feature space alignment processing on the plurality of visual feature vectors, the plurality of audio feature vectors, and the plurality of text feature vectors to obtain aligned plurality of target visual feature vectors, plurality of target audio feature vectors, and plurality of target text feature vectors specifically includes: For the multiple visual feature vectors, the multiple audio feature vectors, and the multiple text feature vectors, visual feature vectors and audio feature vectors, visual feature vectors and text feature vectors, and audio feature vectors and text feature vectors with the same semantic content are taken as positive sample pairs, resulting in multiple visual-audio positive sample pairs, multiple visual-text positive sample pairs, and multiple audio-text positive sample pairs, respectively. For the multiple visual feature vectors, the multiple audio feature vectors, and the multiple text feature vectors, visual feature vectors and audio feature vectors, visual feature vectors and text feature vectors, and audio feature vectors and text feature vectors with different semantic content are used as negative sample pairs to obtain multiple visual-audio negative sample pairs, multiple visual-text negative sample pairs, and multiple audio-text negative sample pairs, respectively. Based on the multiple visual-audio positive sample pairs, multiple visual-text positive sample pairs, multiple audio-text positive sample pairs, the multiple visual-audio negative sample pairs, multiple visual-text negative sample pairs, and multiple audio-text negative sample pairs, calculate the total cross-modal contrast loss; The model parameters of the visual feature extraction model, the audio feature extraction model, and the text feature extraction model are updated in reverse using the total cross-modal contrast loss. The visual feature extraction model, audio feature extraction model, and text feature extraction model with updated model parameters are used to extract features from the multiple image blocks, the multiple spectrograms, and the text data, respectively, to obtain multiple aligned target visual feature vectors, multiple target audio feature vectors, and multiple target text feature vectors.

4. The method according to claim 3, characterized in that, The step of calculating the total cross-modal contrast loss based on the multiple visual-audio positive sample pairs, multiple visual-text positive sample pairs, multiple audio-text positive sample pairs, multiple visual-audio negative sample pairs, multiple visual-text negative sample pairs, and multiple audio-text negative sample pairs specifically includes: For each positive visual-audio sample pair, calculate the first feature inner product of each positive visual-audio sample pair and the second feature inner product of the corresponding negative visual-audio sample pair. Based on the first feature inner product and the second feature inner product, calculate the visual audio loss corresponding to each visual audio positive sample pair, and add all the visual audio losses to obtain the total visual audio loss; For each pair of positive visual text samples, calculate the third feature inner product of each pair of positive visual text samples and the fourth feature inner product of the corresponding pair of negative visual text samples. Based on the third feature inner product and the fourth feature inner product, calculate the visual text loss corresponding to each visual text positive sample pair, and add up all the visual text losses to obtain the total visual text loss. For each positive audio text sample pair, calculate the fifth feature inner product of each positive audio text sample pair and the sixth feature inner product of the corresponding negative audio text sample pair. Based on the inner product of the fifth feature and the inner product of the sixth feature, calculate the audio text loss corresponding to each positive audio text sample pair, and sum all the audio text losses to obtain the total audio text loss; The total visual-audio loss, the total visual-text loss, and the total audio-text loss are fused to obtain the total cross-modal contrast loss.

5. The method according to claim 1, characterized in that, The step of performing temporal clustering on video frames of the visual data based on multiple target visual feature vectors, multiple target audio feature vectors, and multiple target text feature vectors to determine multiple temporal semantic stages and the multimodal concatenation vector corresponding to each temporal semantic stage specifically includes: Based on the multiple target visual feature vectors, calculate the target visual feature vector of each video frame in the visual data and the visual feature similarity between it and the target visual feature vectors of adjacent video frames; Based on the multiple target audio feature vectors, calculate the audio feature similarity between the target audio feature vector corresponding to each video frame and the adjacent target audio feature vectors. Based on the multiple target text feature vectors, calculate the text feature similarity between the target text feature vector corresponding to each video frame and the adjacent target text feature vectors; Based on the visual feature similarity, the audio feature similarity, and the text feature similarity, multiple temporal semantic stages are determined; Obtain the target visual feature vector, target audio feature vector, and target text feature vector corresponding to each temporal semantic stage; The target visual feature vector, target audio feature vector, and target text feature vector corresponding to each temporal semantic stage are concatenated to obtain the multimodal concatenation vector corresponding to each temporal semantic stage.

6. The method according to claim 1, characterized in that, The step of calculating the saliency score of each temporal semantic stage based on the multimodal concatenation vector of each temporal semantic stage through a temporal saliency aggregator specifically includes: Based on the multimodal concatenation vector of each temporal semantic stage, the degree of association between each temporal semantic stage and all stages is calculated through the query projection matrix and key projection matrix in the temporal saliency aggregator, and the self-attention weight is obtained. Based on the self-attention weights and the multimodal concatenation vectors of all temporal semantic stages, the saliency score of each temporal semantic stage is calculated using the saliency feature projection matrix of the temporal saliency aggregator and the activation function.

7. The method according to claim 1, characterized in that, After the step of taking the temporal semantic stage with a saliency score greater than or equal to a preset score threshold as the target temporal semantic stage and obtaining multiple key video frames from multiple video frames corresponding to the target temporal semantic stage, the method further includes: Obtain the multimodal stitching vector corresponding to each key video frame; Based on the multimodal splicing vector corresponding to each key video frame, a support set and a query set are constructed, wherein the support set includes multimodal splicing vectors with semantic labels; For the support set, the average of all labeled multimodal splicing vectors with semantic tags as keyframes is taken to obtain the keyframe prototype vector, and the average of all labeled multimodal splicing vectors with semantic tags as non-keyframes is taken to obtain the non-keyframe prototype vector. Calculate the keyframe prototype similarity between each multimodal concatenation vector in the query set and the keyframe prototype vector, and perform exponential normalization on the keyframe prototype similarity to obtain the keyframe probability. Calculate the non-key frame prototype similarity between each multimodal concatenation vector in the query set and the non-key frame prototype vector, and perform exponential normalization on the non-key frame prototype similarity to obtain the non-key frame probability. Based on the true label of each multimodal concatenation vector in the query set, the keyframe probability, and the non-keyframe probability, the keyframe prototype classification loss is calculated to update the parameters of the multimodal encoder and the temporal saliency aggregator.

8. A video keyframe extraction device, characterized in that, include: A receiving module is used to receive multimodal data of the target video, wherein the multimodal data includes visual data, audio data, and text data; The feature extraction module is used to extract features from the visual data, the audio data, and the text data using a multimodal encoder, respectively, to obtain multiple visual feature vectors, multiple audio feature vectors, and multiple text feature vectors. The alignment module is used to perform feature space alignment processing on the multiple visual feature vectors, the multiple audio feature vectors and the multiple text feature vectors to obtain multiple aligned target visual feature vectors, multiple target audio feature vectors and multiple target text feature vectors. The clustering module is used to perform temporal clustering on video frames of the visual data based on multiple target visual feature vectors, multiple target audio feature vectors, and multiple target text feature vectors, to determine multiple temporal semantic stages and the multimodal concatenation vector corresponding to each temporal semantic stage. The calculation module is used to calculate the saliency score of each temporal semantic stage based on the multimodal concatenation vector of each temporal semantic stage through a temporal saliency aggregator; The acquisition module is used to take the temporal semantic stage with a saliency score greater than or equal to a preset score threshold as the target temporal semantic stage, and acquire multiple key video frames from multiple video frames corresponding to the target temporal semantic stage.

9. A computer device, characterized in that, The device includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the video keyframe extraction method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the video keyframe extraction method as described in any one of claims 1 to 7.