Video key frame extraction method and device, equipment and medium
By combining dynamic sampling and size normalization with pre-trained models and attention mechanisms, this method addresses the shortcomings of existing video keyframe extraction methods in terms of domain adaptability and dynamic modeling efficiency, achieving high-quality keyframe extraction and adapting to video analysis in complex scenarios.
Patent Information
- Application Number
- CN202511058809.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-18
AI Technical Summary
Existing video keyframe extraction methods are insufficient in terms of domain adaptability and dynamic modeling efficiency, especially in scenarios with few samples where performance is limited and they are difficult to adapt to the dynamic changes of complex scenarios.
By generating a uniform-sized test image through dynamic sampling and size normalization, weights are allocated using a pre-trained detection model and attention mechanism to enhance the attention weight of the target region and extract key frames that meet the conditions.
It improves the quality and efficiency of keyframe extraction, ensuring that the extracted keyframes can truly reflect the core content and semantic information of the video, and adapt to the video analysis needs of different fields.
Smart Images

Figure CN120976822A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular to a video key frame extraction method and device, equipment and medium. BACKGROUND
[0002] In the processing of analyzing and identifying video content in the fields of financial technology, medicine, health and the elderly, video key frame extraction is a basic task of video analysis and content understanding, and its goal is to select the most representative frames from video sequences to reduce storage costs and retain core semantic information. Traditional methods mainly select key frames based on handcrafted features (such as color histograms, optical flow) or clustering algorithms, but such methods are difficult to adapt to dynamic changes in complex scenes. In recent years, deep learning technology has significantly improved the performance of key frame extraction through convolutional neural networks (CNN) and time series modeling methods (such as LSTM), but this method still has two major limitations: 1. Insufficient domain adaptability: Most methods rely on specific data sets for training, and when applied to new scenarios (such as transferring from surveillance videos to medical endoscope videos), the entire model needs to be retrained. Although domain adaptation techniques (such as attention matching) can alleviate this problem, they usually require a large amount of labeled data in the target domain; 2. Low efficiency of dynamic modeling: Existing methods that integrate spatial and temporal features (such as 3D-CNN or CNN-LSTM) often use fixed architectures, making it difficult to balance computational efficiency and long-range dependency capture capability. Although the Transformer model improves time series modeling through self-attention mechanisms, its demand for data volume limits its application in small sample scenarios. SUMMARY
[0003] The embodiments of the present application provide a video key frame extraction method, device, equipment and medium, aiming to solve the problem of poor feature extraction quality in the prior art.
[0004] In a first aspect, the embodiments of the present application provide a video key frame extraction method, which comprises:
[0005] Receiving a to-be-tested video captured by a camera;
[0006] Dynamically sampling the to-be-tested video frames to obtain a plurality of to-be-tested frames with uniform sizes;
[0007] Identifying each region of each to-be-tested frame to obtain all key frame pictures containing target regions;
[0008] According to the preset detection model, the attention weights of each target region of all key frame pictures are distributed;
[0009] The attention weights of the specified target region are processed to obtain key frame pictures;
[0010] extracting all key frame pictures meeting preset weight conditions.
[0011] In a second aspect, the embodiment of the present application further provides a video key frame extraction device, which comprises:
[0012] a data preprocessing unit configured to receive a to-be-tested video captured by a camera;
[0013] a scale normalization unit configured to perform dynamic sampling on pictures of the to-be-tested video to obtain a plurality of to-be-tested pictures with uniform sizes;
[0014] a meta-learning basic attention unit configured to identify each region of each to-be-tested picture to obtain all key frame pictures containing target regions, and to assign attention weights of the target regions of all the key frame pictures according to a preset detection model;
[0015] a meta-control unit configured to perform promotion processing on the attention weights of specified target regions to obtain key frame pictures;
[0016] a post-processing unit configured to extract all key frame pictures meeting preset weight conditions.
[0017] In a third aspect, the embodiment of the present application further provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the method of the first aspect when executing the computer program.
[0018] In a fourth aspect, the embodiment of the present application further provides a computer readable storage medium, wherein the storage medium stores a computer program, and the computer program comprises program instructions, and the program instructions can implement the method of the first aspect when executed by a processor.
[0019] The present application receives a to-be-tested video, performs dynamic sampling on pictures to generate to-be-tested pictures with uniform sizes, identifies each region of each to-be-tested picture to extract all key frame pictures containing target regions, assigns attention weights of each target region of each key frame picture by using a preset detection model, performs promotion processing on the attention weights of specified target regions to obtain key frame pictures, and finally extracts key frame pictures meeting preset weight conditions. In this way, the semantic relevance of pictures is preserved during the extraction of key frames, and the extracted key frames can truly and completely reflect the core content and semantic information of a video, thereby providing high-quality key frame data for video processing and analysis. BRIEF DESCRIPTION OF DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0021] Figure 1 The flowchart of the video key frame extraction method provided by the embodiment of the present application is shown in the figure.
[0022] Figure 2 The schematic block diagram of the video key frame extraction device provided by the embodiment of the present application is shown in the figure.
[0023] Figure 3 The schematic block diagram of the electronic device provided by the embodiment of the present application is shown in the figure.
[0024] Figure 4 The application environment schematic diagram of the video key frame extraction method provided by the embodiment of the present application is shown in the figure. DETAILED DESCRIPTION
[0025] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.
[0026] It should be understood that when used in the specification and the appended claims, the terms "comprise" and "include" indicate the presence of described features, integers, steps, operations, elements, and / or components, but do not exclude one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0027] It should also be understood that the terms used in the present application specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the present application specification and the appended claims, unless otherwise clearly indicated by the context, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0028] It should be further understood that the term "and / or" used in the present application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes these combinations.
[0029] The embodiments of the present application provide a video key frame extraction method, device, equipment and medium. The video key frame extraction method please refer to Figure 4, Figure 4 The application environment schematic diagram of the video key frame extraction method provided by the embodiment of the present application. The video key frame extraction method is applied in the application environment as shown in Figure 4 The computer device communicates with at least one host device through a network; the computer device can be a server or a personal computer, wherein the server can be implemented by an independent server or a server cluster composed of multiple servers, and the host device can be, but is not limited to, a server, a smart phone, a tablet computer, a desktop computer, a camera, and the like. The host device sends a to-be-tested video to the computer device, and the computer device executes the video key frame extraction method to extract the video key frame. In the embodiment of the present application, the video key frame extraction method is applied in the field of smart city to extract and analyze the key frame of the device inspection video collected by the device needing to be safety inspected; in addition to the field of smart city, the video key frame extraction method in the embodiment of the present application can also be applied in the fields of finance and insurance and medical treatment, such as constructing a medical image analysis system corresponding to the video key frame extraction method in the medical field to extract and analyze the key frame of the dynamic video image information obtained by medical detection, and constructing a portrait analysis system corresponding to the video key frame extraction method in the field of finance and insurance to extract and analyze the key frame of the video information containing a portrait. The present application will be described in detail through specific embodiments.
[0030] Figure 1 The flowchart of the video key frame extraction method provided by the embodiment of the present application. As shown in Figure 1 The extraction method includes the following steps S110-S160.
[0031] In steps S110 and S120, a to-be-tested video captured by a camera is received; and the to-be-tested video is dynamically sampled to obtain a plurality of to-be-tested frames with uniform size.
[0032] The sampling rate is dynamically adjusted according to the dynamic complexity of the video content captured by the camera, so as to reduce the data amount and retain important information, and the sampled frames are uniformly scaled to the resolution required by the model to ensure the consistency of the input. The to-be-tested video can be video information collected in various application fields, such as a road surface monitoring video collected in the field of smart city or a device inspection video collected by inspecting the device, and the like. The to-be-tested video can also be a customer portrait video collected in the financial business scenario, or a medical detection video collected in the medical business scenario, and the like.
[0033] For example, in the field of smart city, a 10-minute power transmission line inspection video is taken by a camera, where the frame rate of the inspection video is 30 fps, and the video frame contains components such as wires, insulators, tower body, etc. The goal of this time is to extract the key frames required for insulator defect detection. In order to retain key information while reducing data volume, the sampling rate is dynamically adjusted according to the dynamic complexity of the video content.
[0034] First, the dynamic complexity is evaluated by calculating the differential entropy of adjacent frames, that is, each frame image is grayed to simplify the calculation, the gray histogram of the current frame and the previous frame is calculated, the entropy of each frame is calculated using the histogram, and finally the differential entropy between adjacent frames is calculated. If the differential entropy is low, it means that the video content changes little (such as static tower body picture), the sampling rate can be reduced, for example, from 30 fps to 1 fps, if the differential entropy is high, it means that the content of the video changes dramatically (such as fast moving or vibrating insulators), so a higher sampling rate is needed, for example, 30 fps.
[0035] In this way, according to the above rules, 1800 frames are dynamically sampled from the 10-minute video.
[0036] Next, the sampled frames are uniformly scaled to the resolution required by the model to ensure consistency of input. For example, the sampled frames are uniformly scaled to a resolution such as 224x224, that is, using bilinear interpolation or other image scaling algorithms, the resolution of each frame is scaled from the original size (such as 1920x1080) to 224x224 to ensure that the scaled image is not distorted while retaining key information. By converting the original 10-minute power transmission line inspection video into 1800 frames, each with a resolution of 224x224, the attention mechanism adjustment for subsequent use is provided.
[0037] By this setting, through adaptive sampling, the number of frames of the original video is reduced from 18000 frames (10 minutes x 30 fps) to 1800 frames, which reduces the amount of calculation and also ensures that high-speed motion or dynamic changes (such as insulator defects) can be retained at a high sampling rate, providing high-quality input for subsequent feature extraction and key frame extraction.
[0038] In steps S130 and S140, each region of each to-be-detected picture is identified to obtain all key frame pictures containing target regions; and the attention weights of each target region of all key frame pictures are distributed according to a preset detection model
[0039] Specifically, in order to be able to quickly identify the target region in the new video, for example, the insulation defect, there is also a technical effect of automatically adjusting the attention degree to different regions according to the task requirement. Therefore, a plurality of tasks related to the target task are selected for pre-training. For example, in the insulator defect detection, insulator defect video clips containing different scenes (such as different light conditions, different backgrounds) can be used, and then annotation data is prepared to mark whether there is a defect in each frame and the position of the defect. By using the MAML algorithm to pre-train the model, the model only needs a small amount of gradient update to quickly adapt to the new task, for example, a task is randomly sampled from the task set, the model parameters are updated on the training set of the sampled task, the task-specific parameters are obtained, the loss is calculated on the validation set of the sampled task, and the global parameters are updated, so that the model can quickly adapt to the new task. When the pre-trained MAML parameters are loaded into the model as the initial attention weight matrix. This parameter has the ability to quickly adapt to similar tasks, and then a few example videos containing insulator damage are obtained as a support set, and a task embedding vector is generated through the support set. The task embedding vector is used to guide the model to focus on the features related to the current task (insulator defect detection), and when processing each frame, the model will dynamically adjust the attention weight according to the features of the current frame and the task embedding vector. For example, if the current frame contains the features of the insulator crack, the model will increase the attention weight of the frame according to the task embedding vector.
[0040] Through this dynamic adjustment mechanism, the model can flexibly focus on the importance of different frames according to the task requirement, and finally, the model generates an attention weight for each frame to represent the importance of the frame in the current task.
[0041] It can be understood that the generation method of the task embedding vector of the embodiment of the present application is to extract the features of each frame in the support set, and obtain the task embedding vector through aggregation (such as average pooling, maximum pooling or attention pooling).
[0042] It should be noted that the attention weight of the embodiment of the present application is obtained by the following formula:
[0043]
[0044] The weight reflects the general visual importance prior shared across domains, and the calculated attention weight and V are weighted and averaged to obtain the output vector v h t , as the subsequent input.
[0045] In step S150, the attention weight of the specified target region is increased to obtain the key frame picture.
[0046] Specifically, by identifying the specified target region, i.e., the feature of the insulator region, for all key frame pictures, the attention weight of the target region can be improved at this time. This approach can facilitate the model to pay more attention to the features of the insulator region, especially the defect features such as cracks and damage, in the subsequent key frame extraction process. It should be noted that for non-insulator regions (such as conductors, tower bodies, etc.), their attention weights can be adjusted appropriately to make them as key frame pictures. For example, abnormal vibration of the conductor in the key frame picture is also related to insulator defect detection, so the weight of the conductor region will be appropriately increased; otherwise, the weight can be reduced to reduce the attention to irrelevant features.
[0047] Specifically, in the process of this intelligent distribution, as follows:
[0048] 1. The target domain examples can be mapped to low-dimensional vectors using the hypernetwork H:
[0049]
[0050] where m is the example frame.
[0051] 2. Calculate the weight distribution coefficient of CNN and transformer:
[0052]
[0053] where, is a learnable matrix, and σ is the Sigmoid function.
[0054] 3. Fusion of basic attention and domain-specific attention:
[0055] α t ′=g t ·α t +(1-g t )·MLP(z t )
[0056] where MLP is a two-layer perceptron, and the output is the attention score after domain correction. The higher the attention score, the more important the current frame t.
[0057] In step S160, all key frame pictures that meet the pre-set weight condition are extracted.
[0058] Finally, the final output key frame index set and its confidence score. For example, in this example, 42 key frames are finally extracted, of which 38 frames clearly show the insulator surface crack, and 4 frames capture the abnormal vibration of the conductor; at the same time, the extracted key frames are clustered, and the key frames of different categories (such as cracks, vibrations, etc.) are further analyzed. For example, all key frames showing insulator cracks are classified into one category, and key frames showing abnormal conductor vibration are classified into another category, and finally a video summary is generated according to the key frames to facilitate quick viewing of important information in the video.
[0059] In some specific embodiments, the step is to assign attention weights of each target region of all key frame pictures according to a preset detection model, specifically including the following steps:
[0060] Obtain the source domain features corresponding to the region features of each key frame picture from the detection model.
[0061] Specifically, the source domain features refer to features learned from a large amount of labeled data in the pre-training stage, which are related to a specific task (such as insulator defect detection). Such features are extracted in a fixed and known source domain (such as a substation monitoring video) and can describe the normal state and defect state of the target object (such as an insulator). The detection model of the embodiment of the application is a pre-trained deep learning model (such as CNN or Transformer), which can extract features of input images or video frames for representation of high-dimensional vectors and can capture semantic information in the image. By inputting the extracted key frames (such as key frames required for insulator defect detection) into the pre-trained detection model, the detection model performs feature extraction on each frame and outputs the feature representation of each frame. These feature representations can be the output of the convolutional layer of CNN or the embedding vector of Transformer, and each frame is divided into multiple regions (such as grid division or region of interest based on target detection). For each region, a corresponding feature vector is extracted, and for each region feature, a source domain feature corresponding thereto is obtained from the pre-trained model. The source domain feature can be a pre-defined feature library containing normal state and defect state.
[0062] Wherein, the CNN time sequence is obtained by using an inflation causal convolution to process the frame sequence
[0063]
[0064] Wherein, d is the inflation factor (exponentially increasing with network depth), k d is the convolution kernel size, w i is the learnable weight.
[0065] Align each region feature of each keyframe picture with the corresponding source domain feature to obtain the target region.
[0066] Specifically, the region features in the keyframe picture are matched with the source domain features, so that the model can identify which regions are most similar to the known source domain features, thereby determining the target region. Distance metrics such as Euclidean distance, cosine similarity, or Wasserstein distance can be used to calculate the similarity between each region feature in the keyframe and the source domain feature. The region features in the keyframe can be adjusted to the closest state to the source domain feature according to the similarity, for example, by feature interpolation, feature mapping, or feature transformation. If the similarity of a certain region feature to the source domain feature is low, it can be adjusted to be closer to the distribution of the source domain feature through feature transformation such as linear transformation or nonlinear mapping.
[0067] The Wasserstein distance is obtained in the following way:
[0068]
[0069] where μ s ,μ t are the frame embedding distributions of the source domain and the target domain, respectively.
[0070] Assign attention weights to the target regions of each keyframe picture.
[0071] Specifically, according to the similarity calculated in the feature alignment stage, an attention weight is assigned to each region. The higher the similarity, the greater the attention weight. For example, if the similarity of a certain region feature to the source domain feature is 0.9 (close to 1), a higher attention weight is assigned; if the similarity is 0.3, a lower weight is assigned. Finally, the attention weights of all regions are normalized to ensure that the sum of the weights is 1, so that the rationality of the weight distribution can be guaranteed.
[0072] In summary, the region features of the keyframes are extracted from the detection model, the features of the insulator region are aligned with the source domain features, and the similarity is calculated. If the similarity of the insulator region to the source domain feature is 0.8 and the similarity of the background region is 0.2, the initially assigned attention weights are 0.8 and 0.2, respectively.
[0073] In further embodiments, the step of aligning each region feature of each keyframe picture with the corresponding source domain feature to obtain the target region specifically includes the following steps:
[0074] Calculate the distribution difference value between each region feature of each keyframe picture and the corresponding source domain feature.
[0075] Specifically, for each key frame, it is divided into multiple regions (e.g. by grid division or region of interest division based on object detection). Feature vectors are extracted for each region, which can be the output of the convolutional layer of CNN, the embedding vector of Transformer or other feature representations. The source domain features are predefined, usually obtained by training on a large amount of labeled data, which can well describe the normal state and defect state of the target object (such as insulator). A suitable metric method is selected to calculate the distribution difference between the key frame region features and the source domain features. For example, the metric method is Wasserstein distance, which measures the similarity between two probability distributions, is suitable for high-dimensional feature space, and can capture the global difference of feature distribution. For each region feature fregion of each key frame, the distribution difference value D between it and the source domain feature fsource is calculated.
[0076] For example, using Wasserstein distance:
[0077] D = W (fregion, fsource).
[0078] According to the distribution difference value of the two, the region features of each key frame picture are optimized to obtain the target region.
[0079] Specifically, through optimization, the region features of the key frame are closer to the source domain features, so as to obtain the target region. The optimization goal is to reduce the distribution difference value D, so that the region features in the key frame and the source domain features are more consistent in distribution. Through a certain mapping function φ, the region features fregion of the key frame are mapped to the distribution of the source domain features. The mapping function can be linear (such as linear transformation) or nonlinear (such as neural network).
[0080] Taking linear mapping as an example, assume that the mapping function is a linear transformation φ (fregion) = Wfregion + b, where W and b are parameters learned by optimization; or, taking nonlinear mapping as an example, using neural network (such as MLP or multi-layer Transformer) as mapping function, optimizing network parameters through back propagation, so that the distribution difference between the mapped features and the source domain features is minimized.
[0081] The specific optimization process is to initialize the parameters θ (such as the weights and biases of the neural network) of the mapping function, input the region features fregion of the key frame into the mapping function φ, obtain the mapped features fregion' = φ (fregion; θ), calculate the distribution difference value D between the mapped features fregion' and the source domain features fsource, update the parameters θ through back propagation to minimize the loss D, and finally repeat the above steps until the distribution difference value D converges to a smaller value. In this way, after optimization, the region features in the key frame and the source domain features are more consistent in distribution. At this time, the similarity of the optimized region features and the source domain features can be used as a judgment basis to select the region with the highest similarity as the target region. For each region, the similarity of the optimized features and the source domain features is calculated as the confidence of the region.
[0082] For example, by extracting the region features fregion of the key frame and the source domain features fsource, using the Wasserstein distance to calculate the distribution difference value D between each region feature and the source domain feature, using a neural network as the mapping function φ to map the region features fregion of the key frame to the distribution of the source domain features, optimizing the network parameters through back propagation to minimize the distribution difference between the mapped features fregion' and the source domain features fsource, calculating the similarity of the optimized features fregion' and the source domain features fsource, and selecting the region with the highest similarity as the target region, for example, the similarity of the insulator region is 0.9, and the similarity of the background region is 0.2, so the insulator region is determined as the target region.
[0083] Through the above steps, the model can effectively align the region features in the key frame with the source domain features and highlight the target region, thereby improving the accuracy and efficiency of key frame extraction.
[0084] In some specific embodiments, the step of enhancing the attention weight of the specified target region includes the following steps:
[0085] Receiving a task embedding vector; the task embedding vector contains semantic information corresponding to the current task.
[0086] Specifically, the task embedding vector is a high-dimensional vector that contains semantic information of the current task, which can guide the model to focus on the most relevant features to the current task. For example, in the insulator defect detection task, the task embedding vector contains semantic information such as the shape, texture, and common defect patterns of the insulator. The generation of the task embedding vector requires a small amount of labeled data, such as a video clip containing an example of a broken insulator. These labeled data are input into a pre-trained model to extract feature representations, and then these feature representations are compressed into a high-dimensional vector, i.e., the task embedding vector eT, through some aggregation methods such as average pooling, max pooling, or a small neural network.
[0087] According to the task embedding vector, a corresponding dynamic gating instruction is generated; the dynamic gating instruction is an instruction for adjusting the attention weight.
[0088] Specifically, since the dynamic gating instruction is an instruction for adjusting the attention weight, it is used to guide the model to allocate attention in different tasks, and it can be dynamically generated according to the task embedding vector to adapt to different task requirements. For example, by setting a gating module, the input is the task embedding vector eT, and the output is the dynamic gating instruction gt. In the training phase, the gating module adjusts its parameters by learning the relationship between the task embedding vector and the attention weight, for example, if the task embedding vector indicates that the current task is to detect surface cracks of the insulator, the gating module will learn to increase the attention weight of the insulator region.
[0089] The dynamic gating instruction generation is to input the task embedding vector eT into the gating module to generate the dynamic gating instruction gt, which is a vector, and each element corresponds to an attention weight adjustment factor for a region. For example, for the insulator region, the adjustment factor may be a larger value (such as 2.3), and for the background region, the adjustment factor may be close to 1.
[0090] According to the dynamic gating instruction, the attention weight of the specified target region is increased, focusing on the key frame picture to enhance the attention to the target region.
[0091] Specifically, after the feature alignment and optimization step, each region already has an initial attention weight, and then the dynamic gating instruction gt is used to adjust the initial attention weight. Specifically, the initial attention weight αregion of each region is multiplied by the corresponding adjustment factor gregion, for example, if the initial attention weight of the insulator region is 0.5 and the adjustment factor in the dynamic gating instruction is 2.3, then the adjusted attention weight is:
[0092] αregion′=αregion×gregion=0.5×2.3=1.15;
[0093] The adjusted attention weights need to be normalized to ensure that the sum of the weights is 1. By adjusting the attention weights in this way, the attention weight of the target region (such as the insulator region) is significantly increased, thereby enhancing the model's focus on the target region.
[0094] For example, if one of the key frames contains an insulator region and a background region, the features are extracted from the small amount of labeled insulator damage video segments to generate the task embedding vector eT. The task embedding vector eT is input into the gating module to generate the dynamic gating instruction gt. The adjustment factor output by the gating module is: the insulator region is 2.3, and the background region is 1. The initial attention weight is 0.5 for the insulator region and 0.5 for the background region. The adjusted attention weight is as follows:
[0095] Insulator region: 0.5 x 2.3 = 1.15, background region: 0.5 x 1 = 0.5.
[0096] The normalized attention weight is:
[0097]
[0098] As can be seen, by the above steps, the attention weight of the model is dynamically adjusted, and the weight of the insulator region is significantly increased, thereby enhancing the focus on the target region. This step can flexibly adjust the attention distribution of the model according to different task requirements, thereby improving the accuracy of key frame extraction.
[0099] In further embodiments, the attention weight of the specified target region is increased according to the dynamic gating instruction to obtain a key frame picture, specifically including the following steps:
[0100] Confirm the position of the target region in the key frame picture according to the dynamic gating instruction.
[0101] Specifically, since the dynamic gating instruction `gt` is a vector containing information about the location and importance of the target region, it is generated from the task embedding vector and includes attention weight adjustment factors for different regions. If the dynamic gating instruction specifies the location information of the target region, or indirectly indicates the location of the target region through a mapping relationship, then when confirming the location of the target region, the keyframe image is first divided into multiple regions, and the weight vector in the dynamic gating instruction is mapped to the regions in the keyframe. For example, if the dynamic gating instruction is a vector of length N, and the keyframe is also divided into N regions, with each weight value corresponding to one region, the region with the highest weight is determined as the target region by analyzing the weight values in the dynamic gating instruction. For example, if the dynamic gating instruction indicates that the weight of the insulator region is 2.3, while the weight of other regions is 1, then the insulator region is identified as the target region.
[0102] The attention weight of the corresponding target area in the keyframe is increased by a preset multiple.
[0103] Specifically, the preset multiplier is a fixed value used to increase the attention weight of the target region. This multiplier can be preset according to task requirements. For example, in an insulator defect detection task, the preset multiplier might be set to 2.3 to significantly increase the attention given to parts of the insulator. Before the dynamic gating command adjustment, each region already has an initial attention weight αregion. For parts identified as the corresponding target region, their initial attention weight αregion is multiplied by the preset multiplier. For example, if the initial weight of the target region is 0.5 and the preset multiplier is 2.3, the increased weight will be as follows:
[0104] αregion′=αregion×preset multiplier=0.5×2.3=1.15;
[0105] For non-target areas, the attention weight remains unchanged, or is adjusted slightly as needed. The adjusted attention weights need to be normalized to ensure that the sum of the weights is 1.
[0106] Through the above adjustments, the attention weights of the corresponding target regions are significantly increased, thereby enhancing the model's attention to the target regions.
[0107] In some specific embodiments, after the step of increasing the attention weight of the specified target region to obtain the keyframe image with attention, the following steps are also included:
[0108] All keyframes are suppressed to obtain deduplicated keyframes.
[0109] Specifically, after boosting the attention weight of the target region, the attention weights of the entire keyframe picture need to be redistributed to suppress redundant information, for example, by non-maximum suppression (NMS) or other suppression strategies, to reduce the weights of similar or repetitive keyframes and retain the most representative keyframes.
[0110] Taking non-maximum suppression (NMS) as an example, all keyframes are sorted according to the boosted attention weights, with the keyframes with higher weights placed in front. The keyframe with the highest weight is selected as the current "maximum value", and the similarity (such as Euclidean distance, cosine similarity, etc.) between the current maximum value keyframe and other keyframes is calculated. If the similarity exceeds a certain threshold (for example, 0.8), the weights of these adjacent keyframes are reduced to a smaller value (for example, 0 or a value close to 0), and the keyframe with the highest weight is selected from the remaining keyframes, and the above suppression process is repeated until all keyframes are processed.
[0111] After the suppression process, a set of simplified and effective keyframes is obtained. These keyframes not only focus on the target region but also remove redundant information, so that each keyframe has high representativeness. For each retained keyframe, its attention weight can be retained as a confidence score for subsequent analysis or display.
[0112] In some specific embodiments, the step of dynamically sampling the video pictures to be tested to obtain a plurality of uniformly sized video pictures to be tested includes the following steps:
[0113] Calculate the differential entropy of adjacent video pictures to be tested.
[0114] Specifically, since differential entropy is a measure of video picture dynamic change, it is used to evaluate the difference between adjacent frames. It can help us judge the dynamic complexity of the video picture. The higher the differential entropy, the more dramatic the picture change and the higher the dynamic complexity, which can be obtained by calculating the pixel difference or feature difference between adjacent frames.
[0115] According to the dynamic complexity of each video picture to be tested, the sampling rate of each video picture to be tested is adjusted to obtain the sampled video picture to be tested.
[0116] Specifically, the dynamic complexity is measured by the differential entropy. The higher the differential entropy, the higher the dynamic complexity, indicating that the picture changes more dramatically. In the UAV inspection video, the area with high dynamic complexity may include a fast-moving UAV, a swaying conductor, or a rapidly changing insulator. For areas with high dynamic complexity (such as fast-moving insulators or conductors), maintain a high sampling rate (e.g., 30 fps) to capture more details. For areas with low dynamic complexity (such as static towers or smooth backgrounds), reduce the sampling rate (e.g., 1 fps) to reduce redundant information. For example, by setting a dynamic complexity threshold θ. For example, when the differential entropy Ht> θ, maintain a high sampling rate; when Ht< θ, reduce the sampling rate; where the threshold θ = 0.5, if Ht> 0.5, the sampling rate is 30 fps, and if Ht< 0.5, the sampling rate is 1 fps.
[0117] The normalized processing is performed on each sampled video frame to obtain a plurality of frames of uniform size of the test frame.
[0118] Specifically, all the sampled video frames are normalized to a uniform size for subsequent processing. For example, all frames are scaled to 224x224 pixels, and through scaling and cropping, it is ensured that all frames have the same resolution and proportion. By scaling each sampled frame to a uniform size. For example, using bilinear interpolation or other image scaling algorithms, the frame is scaled to 224x224 pixels, and if the aspect ratio of the frame does not match the target size, the frame proportion can be adjusted by cropping or padding. After normalization, all frames have the same size and resolution, facilitating subsequent feature extraction and analysis.
[0119] For example, for a 10-minute power line inspection video with a frame rate of 30 fps, containing conductors, insulators, and tower components, by calculating the pixel difference Dt for each frame Ft and its previous frame Ft-1, and normalizing it to differential entropy Ht, the differential entropy value of each frame is obtained, for example, the differential entropy of the static tower area is low (e.g., 0.1), and the differential entropy of the fast-moving insulator area is high (e.g., 0.8). By setting the dynamic complexity threshold θ = 0.5, for frames with differential entropy Ht> 0.5 (such as the fast-moving insulator area), maintain a sampling rate of 30 fps, and for frames with differential entropy Ht< 0.5 (such as the static tower area), reduce the sampling rate to 1 fps. If 1800 frames are dynamically sampled from a 10-minute video, with fewer static areas and more dynamic areas, the 1800 frames of sampled video frames are scaled to 224x224 pixels, and finally a plurality of frames of uniform size of the test frame are obtained, i.e., each frame has a size of 224x224 pixels.
[0120] Therefore, through the above steps, the dynamic sampling and normalization processing of the to-be-tested video picture are realized, so that the redundant information is effectively reduced, and the details of the key area are retained.
[0121] Figure 3 A schematic block diagram of a video key frame extraction device provided by an embodiment of the present application is shown in Figure 2 Corresponding to the above video key frame extraction method, the present application also provides a video key frame extraction device, which is configured in the application environment as shown in Figure 4 The computer device can be a server or a personal computer, wherein the server can be implemented by an independent server or a server cluster composed of multiple servers, and the host device can be but not limited to a server, a smart phone, a tablet computer, a desktop computer and the like electronic devices.
[0122] Specifically, referring to Figure 2 The video key frame extraction device 700 comprises:
[0123] A data preprocessing unit 701 configured to receive a to-be-tested video captured by a camera;
[0124] A scale normalization unit 702 configured to perform dynamic sampling on the to-be-tested video picture to obtain a plurality of to-be-tested pictures with uniform sizes;
[0125] A meta-learning basic attention unit 703 configured to identify each region of each to-be-tested picture to obtain all key frame pictures containing target regions, and distribute attention weights of each target region of all key frame pictures according to a preset detection model;
[0126] A meta-control unit 704 configured to perform promotion processing on the attention weights of the specified target regions to obtain key key frame pictures;
[0127] A post-processing unit 705 configured to extract all key key frame pictures satisfying a preset weight condition.
[0128] In some embodiments, when performing the step of distributing the attention weights of each target region of all key frame pictures according to the preset detection model, the meta-learning basic attention unit 703 is specifically configured to:
[0129] Obtain source domain features corresponding to the region features of each key frame picture from the detection model;
[0130] Align the region features of each key frame picture with the corresponding source domain features to obtain the target regions;
[0131] Distribute the attention weights of the target regions of each key frame picture.
[0132] In some embodiments, the meta-learning base attention unit 703, when performing the step of aligning the region features of each key frame picture with the corresponding source domain features to obtain the target region, is specifically configured to:
[0133] calculate the distribution difference value between the region features of each key frame picture and the corresponding source domain features;
[0134] optimize the region features of each key frame picture according to the distribution difference value between them to obtain the target region.
[0135] In some embodiments, the meta-control unit 704, when performing the step of performing attention weight promotion processing on the specified target region to obtain the key frame picture, is specifically configured to:
[0136] receive a task embedding vector; the task embedding vector contains semantic information corresponding to the current task;
[0137] generate a corresponding dynamic gating instruction according to the task embedding vector; the dynamic gating instruction is an instruction for adjusting the attention weight;
[0138] perform attention weight promotion processing on the specified target region according to the dynamic gating instruction to obtain the key frame picture.
[0139] In some embodiments, the meta-control unit 704, when performing the step of performing attention weight promotion processing on the specified target region according to the dynamic gating instruction, is specifically configured to:
[0140] confirm the position of the target region in the key frame picture according to the dynamic gating instruction;
[0141] promote the attention weight of the position corresponding to the target region in the key frame picture by a preset multiple.
[0142] In some specific embodiments, after the meta-control unit 704 performs the attention weight promotion processing on the specified target region to obtain the key frame picture, it further performs the following steps:
[0143] suppress all key frame pictures to obtain the non-redundant key frame picture.
[0144] In some specific embodiments, the scale normalization unit 702, when performing the step of dynamically sampling the test video picture to obtain multiple frames of test pictures with uniform size, is specifically configured to:
[0145] calculate the differential entropy of adjacent test video pictures;
[0146] According to the dynamic complexity of each to-be-tested video picture, the sampling rate of each to-be-tested video picture is adjusted to obtain a sampled to-be-tested video picture.
[0147] Each sampled to-be-tested video picture is normalized to obtain a plurality of to-be-tested pictures with uniform sizes.
[0148] It should be noted that the specific implementation process of the above-mentioned video key frame extraction device and each unit can be clearly understood by those skilled in the art, and can refer to the corresponding description in the foregoing method embodiments. For the convenience and brevity of description, it will not be repeated here.
[0149] The above-mentioned video key frame extraction device can be realized in the form of a computer program, which can run on an electronic device as shown in Figure 3 .
[0150] Please refer to Figure 3 , Figure 3 is a schematic block diagram of an electronic device provided by an embodiment of the present application. The electronic device 800 can be a terminal or a server, wherein the terminal can be an electronic device with communication function. The server can be a stand-alone server or a server cluster composed of multiple servers.
[0151] Please refer to Figure 3 , the electronic device 800 includes a processor 802, a memory and a network interface 805 connected through a system bus 801, wherein the memory can include a non-volatile storage medium 803 and an internal memory 804.
[0152] The non-volatile storage medium 803 can store an operating system 8031 and a computer program 8032. The computer program 8032 includes program instructions which, when executed, can cause the processor 802 to perform a video key frame extraction method.
[0153] The processor 802 is configured to provide computing and control capabilities to support the operation of the entire electronic device 800.
[0154] The internal memory 804 provides an environment for the running of the computer program 8032 in the non-volatile storage medium 803, which, when executed by the processor 802, can cause the processor 802 to perform a video key frame extraction method.
[0155] The network interface 805 is configured to perform network communication with other devices. Those skilled in the art can understand that Figure 3The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the electronic device 800 to which the scheme of the present application is applied. The specific electronic device 800 can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0156] The processor 802 is configured to run the computer program 8032 stored in the memory to implement the following steps:
[0157] receiving a to-be-tested video captured by a camera;
[0158] performing dynamic sampling on a to-be-tested video frame to obtain a plurality of to-be-tested frames with uniform sizes;
[0159] identifying each region of each to-be-tested frame to obtain all key frame pictures containing target regions;
[0160] distributing attention weights of each target region of all key frame pictures according to a preset detection model;
[0161] performing promotion processing on the attention weights of the specified target region to obtain key key frame pictures;
[0162] extracting all key key frame pictures satisfying a preset weight condition.
[0163] In some embodiments, when implementing the step of distributing the attention weights of each target region of all key frame pictures according to the preset detection model, the processor 802 specifically implements the following steps:
[0164] obtaining source domain features corresponding to region features of each key frame picture from the detection model;
[0165] aligning the region features of each key frame picture with the corresponding source domain features to obtain target regions;
[0166] distributing the attention weights of the target regions of each key frame picture.
[0167] In some embodiments, when implementing the step of aligning the region features of each key frame picture with the corresponding source domain features to obtain target regions, the processor 802 specifically implements the following steps:
[0168] calculating a distribution difference value between the region features of each key frame picture and the corresponding source domain features;
[0169] optimizing the region features of each key frame picture according to the distribution difference value therebetween to obtain target regions.
[0170] In some embodiments, the processor 802, when implementing the step of performing attention weight promotion on the specified target region to obtain the key frame picture focusing on the key frame picture, specifically implements the following steps:
[0171] receiving a task embedding vector; the task embedding vector comprises semantic information corresponding to the current task;
[0172] generating a corresponding dynamic gating instruction according to the task embedding vector; the dynamic gating instruction is an instruction for adjusting the attention weight;
[0173] performing attention weight promotion on the specified target region according to the dynamic gating instruction to obtain the key frame picture focusing on the key frame picture.
[0174] In some embodiments, the processor 802, when implementing the step of performing attention weight promotion on the specified target region according to the dynamic gating instruction, further implements the following steps:
[0175] confirming the position of the target region in the key frame picture according to the dynamic gating instruction;
[0176] promoting the attention weight of the position corresponding to the target region in the key frame picture by a preset multiple.
[0177] In some embodiments, the processor 802, after implementing the step of performing attention weight promotion on the specified target region to obtain the key frame picture focusing on the key frame picture, specifically implements the following steps:
[0178] performing suppression processing on all key frame pictures focusing on the key frame picture to obtain the key frame picture focusing on the key frame picture after redundancy.
[0179] In some embodiments, the processor 802, when implementing the step of performing dynamic sampling on the to-be-tested video picture to obtain the plurality of to-be-tested pictures with uniform sizes, specifically implements the following steps:
[0180] calculating the differential entropy of adjacent to-be-tested video pictures;
[0181] adjusting the sampling rate of each to-be-tested video picture according to the dynamic complexity of each to-be-tested video picture to obtain the sampled to-be-tested video picture;
[0182] performing normalization processing on each sampled to-be-tested video picture to obtain the plurality of to-be-tested pictures with uniform sizes.
[0183] It should be appreciated that in the embodiments of the present application, the processor 802 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0184] It can be understood by those skilled in the art that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware. The computer program includes program instructions, and the computer program can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the above-mentioned embodiments.
[0185] Therefore, the present application also provides a storage medium. The storage medium can be a computer-readable storage medium. The storage medium stores a computer program, wherein the computer program includes program instructions. The program instructions are executed by a processor to make the processor perform the following steps:
[0186] receiving a to-be-tested video captured by a camera;
[0187] performing dynamic sampling on a to-be-tested video frame to obtain a plurality of to-be-tested frames with uniform sizes;
[0188] identifying each region of each to-be-tested frame to obtain all key frame pictures containing target regions;
[0189] distributing attention weights of each target region of all key frame pictures according to a preset detection model;
[0190] performing promotion processing on the attention weights of the specified target region to obtain key key frame pictures;
[0191] extracting all key key frame pictures satisfying a preset weight condition.
[0192] In an embodiment, when the processor performs the step of distributing attention weights of each target region of all key frame pictures according to a preset detection model, the processor specifically implements the following steps:
[0193] obtaining source domain features corresponding to the region features of each key frame picture from the detection model;
[0194] aligning the region features of each key frame picture with the corresponding source domain features to obtain target regions;
[0195] allocating attention weights to the target regions of each key frame picture.
[0196] In an embodiment, when the processor performs the step of aligning the region features of each key frame picture with the corresponding source domain features to obtain target regions, the following steps are implemented:
[0197] calculating distribution difference values between the region features of each key frame picture and the corresponding source domain features;
[0198] optimizing the region features of each key frame picture according to the distribution difference values to obtain target regions.
[0199] In an embodiment, when the processor performs the step of enhancing the attention weights of the specified target regions to obtain key frame pictures, the following steps are implemented:
[0200] receiving a task embedding vector; the task embedding vector contains semantic information corresponding to the current task;
[0201] generating a corresponding dynamic gating instruction according to the task embedding vector; the dynamic gating instruction is an instruction for adjusting the attention weights;
[0202] enhancing the attention weights of the specified target regions according to the dynamic gating instruction to obtain key frame pictures.
[0203] In an embodiment, when the processor performs the step of enhancing the attention weights of the specified target regions according to the dynamic gating instruction, the following steps are implemented:
[0204] confirming the positions of the target regions in the key frame pictures according to the dynamic gating instruction;
[0205] enhancing the attention weights of the positions corresponding to the target regions in the key frame pictures by a preset multiple.
[0206] In an embodiment, after the processor performs the step of enhancing the attention weights of the specified target regions to obtain key frame pictures, the following steps are implemented:
[0207] suppressing all key frame pictures to obtain key frame pictures after redundancy.
[0208] In an embodiment, the processor, when performing the step of dynamically sampling the to-be-tested video pictures to obtain the to-be-tested pictures with uniform sizes, implements the following steps:
[0209] calculating the differential entropy of the adjacent to-be-tested video pictures;
[0210] adjusting the sampling rate of each to-be-tested video picture according to the dynamic complexity of each to-be-tested video picture to obtain the sampled to-be-tested video pictures;
[0211] normalizing each sampled to-be-tested video picture to obtain the to-be-tested pictures with uniform sizes.
[0212] The storage medium can be a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk, or various computer readable storage media that can store program codes.
[0213] Those skilled in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized in electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in general terms in the above description. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0214] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of each unit is only a logical function division, and actual implementation can have another division manner. For example, a plurality of units or components can be combined or integrated into another system, or some features can be omitted or not executed.
[0215] The steps in the method embodiments of the present application can be adjusted, combined and deleted in sequence according to actual needs. The units in the device embodiments of the present application can be combined, divided and deleted according to actual needs. In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit.
[0216] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a storage medium. Based on such an understanding, the technical solutions of the present application essentially or say the part that contributes to the prior art, or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing an electronic device (which can be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application.
[0217] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical range disclosed by the present application, and these modifications or replacements should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for extracting key frames from a video, characterized in that, The extraction method comprises: receiving a to-be-tested video captured by a camera; performing dynamic sampling on pictures of the to-be-tested video to obtain a plurality of to-be-tested pictures with uniform sizes; identifying each region of each to-be-tested picture to obtain all key frame pictures containing target regions; allocating attention weights of each target region of all key frame pictures according to a preset detection model; performing promotion processing on the attention weights of the specified target region to obtain key frame pictures; extracting all key frame pictures satisfying a preset weight condition.
2. The method of claim 1, wherein, The allocation of the attention weights of each target region of all key frame pictures according to the preset detection model comprises: obtaining source domain features corresponding to region features of each key frame picture from the detection model; aligning each region feature of each key frame picture with the corresponding source domain feature to obtain a target region; allocating attention weights of the target region of each key frame picture.
3. The method of claim 2, wherein, The alignment of each region feature of each key frame picture with the corresponding source domain feature to obtain a target region comprises: calculating a distribution difference value between each region feature of each key frame picture and the corresponding source domain feature; optimizing each region feature of each key frame picture according to the distribution difference value therebetween to obtain a target region.
4. The method of claim 1, wherein, The promotion processing on the attention weights of the specified target region to obtain key frame pictures comprises: receiving a task embedding vector; the task embedding vector contains semantic information corresponding to a current task; generating a corresponding dynamic gating instruction according to the task embedding vector; the dynamic gating instruction is an instruction for adjusting attention weights; performing promotion processing on the attention weights of the specified target region according to the dynamic gating instruction to obtain key frame pictures.
5. The method of claim 4, wherein, The promotion processing on the attention weights of the specified target region according to the dynamic gating instruction comprises: confirming a position of the target region in the key frame picture according to the dynamic gating instruction; performing promotion processing on the attention weights of the position corresponding to the target region in the key frame picture according to a preset multiple.
6. The method of claim 1, wherein, After the promotion processing on the attention weights of the specified target region to obtain key frame pictures, the method further comprises: performing suppression processing on all key frame pictures to obtain the key frame pictures after redundancy elimination.
7. The method of claim 1, wherein, The dynamic sampling on pictures of the to-be-tested video to obtain a plurality of to-be-tested pictures with uniform sizes comprises: calculating differential entropy of adjacent to-be-tested video pictures; adjusting a sampling rate of each to-be-tested video picture according to the dynamic complexity of each to-be-tested video picture to obtain a to-be-tested video picture after sampling; performing normalization processing on each to-be-tested video picture after sampling to obtain a plurality of to-be-tested pictures with uniform sizes.
8. An apparatus for extracting key frames of a video, characterized by comprising: The device is used for performing the extraction of video key frames according to any one of claims 1-7, and the device comprises: a data preprocessing unit configured to receive a to-be-tested video captured by a camera; a scale normalization unit, configured to perform dynamic sampling on the to-be-tested video pictures to obtain a plurality of to-be-tested pictures with uniform sizes; a meta-learning basic attention unit, configured to identify each region of each to-be-tested picture to obtain all key frame pictures containing target regions; and distribute attention weights of each target region of all key frame pictures according to a preset detection model; a meta-control unit, configured to perform promotion processing on the attention weights of the specified target regions to obtain key frame pictures; a post-processing unit, configured to extract all key frame pictures satisfying a preset weight condition.
9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the video key frame extraction method in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and the computer program includes program instructions, which, when executed by a processor, cause the processor to execute the video key frame extraction method in any one of claims 1-7.