Video abstract generation method, electronic device, storage medium and program product

By extracting semantic features and image features of the video, and fusion of attention features to generate video abstracts, the problem of lack of semantic information in traditional video abstract generation methods is solved, and a more comprehensive and accurate video abstract generation is achieved.

CN120014513APending Publication Date: 2025-05-16CHONGQING CHANGAN AUTOMOBILE CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510086838.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The traditional video digest generation method lacks attention to semantic information, resulting in the generated video digests being incomplete and accurate enough.

Method used

By extracting semantic features of the original video based on the first preset strategy, semantic theme features are obtained; extracting image features of the original video based on the second preset strategy to obtain frame-level image features; then fusion of attention features of the semantic theme features and frame-level image features are carried out to obtain graphic and text fusion features; finally filtering video frames based on the graphic and text fusion features to generate a video summary.

Benefits of technology

By extracting and combining subtitle semantic features and video frame image features, the generation of video abstracts takes into account both semantic and image information, improving the shortcomings of traditional video abstract generation methods, and the generated video abstracts are more comprehensive and more accurate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014513A_ABST
    Figure CN120014513A_ABST
Patent Text Reader

Abstract

The invention provides a video abstract generation method, electronic equipment, a storage medium and a program product. The method comprises the steps of performing semantic feature extraction on an original video based on a first preset strategy to obtain semantic theme features; based on a second preset strategy, image feature extraction is carried out on the original video, and frame-level image features corresponding to video frames in the original video are obtained; carrying out attention feature fusion on the semantic theme features and the frame-level image features to obtain image-text fusion features; and according to the image-text fusion features, screening the video frames to obtain a set composed of target video frames as a video abstract, the target video frames representing video frames satisfying a first preset condition. Therefore, by extracting and combining the subtitle semantic features and the video frame image features, the generation of the video abstract gives consideration to semantic and image information, and the problems that the generated video abstract is not comprehensive enough and not accurate enough due to the fact that a traditional video abstract generation mode lacks attention to semantic information are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a video summary generation method, electronic equipment, storage medium and program product. Background Art

[0002] With the widespread deployment of surveillance cameras, surveillance systems have become an indispensable part of urban management, traffic supervision, security monitoring, health care, etc. Large-scale surveillance video data requires efficient management, analysis and understanding, so surveillance video summary generation has become an important technology to solve this problem.

[0003] Among them, surveillance video summary generation relies on the analysis and understanding of video content. This includes research in the fields of target detection and tracking, event recognition, image processing, and computer vision. Through these technologies, the system can automatically extract and identify key information in the video.

[0004] For example, through the six steps of background reconstruction, target detection, target tracking, target trajectory post-processing, target trajectory generation, and video summary generation, the three-dimensional information of the target's activity pipeline is extracted from the original video, which is then synthesized with the background video to condense into a shorter video clip to generate a surveillance video summary. However, the traditional video summary generation method lacks the extraction of semantic information. The video summary is obtained simply by image input, ignoring important semantic information. Summary of the invention

[0005] In view of this, the purpose of the embodiments of the present application is to provide a video summary generation method, electronic device, storage medium and program product, which can improve the problem that traditional video summary generation methods lack attention to semantic information, resulting in the generated video summary being not comprehensive and accurate enough.

[0006] In order to achieve the above technical objectives, the technical solutions adopted in this application are as follows:

[0007] In a first aspect, an embodiment of the present application provides a method for generating a video summary, the method comprising:

[0008] Based on the first preset strategy, semantic features are extracted from the original video for which the summary is to be generated, so as to obtain semantic topic features;

[0009] Based on a second preset strategy, image feature extraction is performed on the original video to obtain frame-level image features corresponding to video frames in the original video;

[0010] Performing attention feature fusion on the semantic topic feature and the frame-level image feature to obtain an image-text fusion feature;

[0011] According to the image-text fusion feature, the video frames are screened to obtain a set consisting of target video frames as a video summary, where the target video frames represent the video frames that meet a first preset condition.

[0012] In combination with the first aspect, in some optional implementations, based on the first preset strategy, extracting semantic features from the original video for which a summary is to be generated to obtain semantic topic features includes:

[0013] The bimodal transformer is used to extract semantic features from the original video to obtain the semantic theme features.

[0014] In combination with the first aspect, in some optional implementations, the bimodal converter includes a bimodal encoder, a bimodal decoder, and a proposal generator;

[0015] The step of extracting semantic features from the original video using a bimodal converter to obtain the semantic theme features includes:

[0016] A1. Generate a word vector library for matching subtitles for the original video through a first pre-trained model;

[0017] A2. extracting visual features from the original video through a second pre-trained model;

[0018] A3. Extracting audio features from the original video through a third pre-trained model;

[0019] A4. Using the bimodal encoder to perform feature fusion on the visual feature and the audio feature to obtain a bimodal fusion feature, wherein the bimodal fusion feature includes a first feature sequence corresponding to the visual feature and a second feature sequence corresponding to the audio feature;

[0020] A5. generating, by the proposal generator, a plurality of prediction proposals and a confidence level corresponding to each prediction proposal in the plurality of prediction proposals according to the first feature sequence and the second feature sequence;

[0021] A6. Cutting the prediction proposal whose confidence satisfies the second preset condition from the multiple prediction proposals as the target proposal, wherein the target proposal includes a first feature sequence that satisfies the second preset condition and a second feature sequence that satisfies the second preset condition;

[0022] A7. Input the target proposal into the bimodal encoder and repeat step A4 to obtain the cut bimodal fusion feature;

[0023] A8. Input the cut bimodal fusion features into the bimodal decoder to match subtitles from the word vector library according to the cut bimodal fusion features as the semantic topic features;

[0024] A9. Repeat the above steps A2 to A8 until a preset mark appears in the video frame.

[0025] In combination with the first aspect, in some optional implementations, the extracting image features from the original video based on the second preset strategy to obtain frame-level image features corresponding to video frames in the original video includes:

[0026] Preprocessing the original video to downsample the video frames and remove useless frames and blurred frames in the original video to obtain an image input, wherein the image input includes a plurality of preprocessed video frames;

[0027] The image input is classified using a CLIP model to obtain a category corresponding to each of the multiple preprocessed video frames as the frame-level image feature.

[0028] In conjunction with the first aspect, in some optional implementations, the CLIP model includes a text encoder and an image encoder;

[0029] The using the CLIP model to classify the image input to obtain a category corresponding to each of the plurality of preprocessed video frames as the frame-level image feature includes:

[0030] Using the text encoder to convert multiple preset category texts into multiple text feature vectors;

[0031] Converting the image input into an image feature vector using the image encoder;

[0032] According to the cosine similarities between the multiple text feature vectors and the image feature vector, a preset category text corresponding to the text feature vector corresponding to the largest cosine similarity is determined as the frame-level image feature.

[0033] In combination with the first aspect, in some optional implementations, performing attention feature fusion on the semantic topic feature and the frame-level image feature to obtain the image-text fusion feature includes:

[0034] Based on the cross attention calculation formula, the semantic theme feature and the frame-level image feature are fused with attention features to obtain the image-text fusion feature:

[0035]

[0036] Where, Q = V 主题 , K=V 图像 , V = V 图像 , V 主题 、V图像 denote semantic topic features and frame-level image features respectively, and d k represents the dimensions of Q and K, K T represents the transposed matrix of K, and v1 represents the image and text fusion features.

[0037] In combination with the first aspect, in some optional implementations, the video frames are screened according to the image-text fusion feature to obtain a set of target video frames as a video summary, including:

[0038] According to the image-text fusion feature, a correlation index corresponding to each of the video frames is determined, where the correlation index represents the correlation between the video frame and the semantic theme feature:

[0039] v2=Linear(ReLu(Linear(v1)))

[0040] y = sigmoid(v2)

[0041] In the formula, v1 represents the image-text fusion feature, y represents the correlation index, ReLu represents the first activation function, sigmoid represents the second activation function, and Linear represents linear transformation;

[0042] The video frame whose correlation index satisfies the first preset condition is taken as the target video frame, and a set of all the target video frames is taken as the video summary.

[0043] In a second aspect, an embodiment of the present application further provides an electronic device, comprising a processor and a memory coupled to each other, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the electronic device executes the above method.

[0044] In a third aspect, an embodiment of the present application further provides a computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium. When the computer program is executed on a computer, the computer executes the above method.

[0045] In a fourth aspect, an embodiment of the present application further provides a computer program product, including a computer program, which implements the above method when executed by a processor.

[0046] The invention adopting the above technical solution has the following advantages:

[0047] In the technical solution provided in the present application, firstly, based on the first preset strategy, semantic features are extracted from the original video for which a summary is to be generated, and semantic theme features are obtained. Then, based on the second preset strategy, image features are extracted from the original video, and frame-level image features corresponding to the video frames in the original video are obtained. Then, attention features are fused between the semantic theme features and the frame-level image features to obtain image-text fusion features. Finally, based on the image-text fusion features, the video frames are screened to obtain a set consisting of target video frames as the video summary, and the target video frames represent video frames that meet the first preset conditions. In this way, by extracting and combining the semantic features of subtitles and the image features of video frames, the generation of video summaries takes into account both semantic and image information, thereby improving the problem that the traditional video summary generation method lacks attention to semantic information, resulting in the generated video summary being not comprehensive enough and not accurate enough. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The present application may be further described by the non-limiting embodiments given in the accompanying drawings. It should be understood that the following drawings only illustrate certain embodiments of the present application and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other relevant drawings may be obtained based on these drawings without creative effort.

[0049] Figure 1 A structural block diagram of an electronic device provided in an embodiment of the present application.

[0050] Figure 2 This is one of the flowcharts of the video summary generation method provided in the embodiment of the present application.

[0051] Figure 3 A schematic diagram of the structure of a dual-mode converter provided in an embodiment of the present application.

[0052] Figure 4 A schematic diagram of the structure of the CLIP model provided in the embodiments of the present application.

[0053] Figure 5 The second flowchart of the video summary generation method provided in the embodiment of the present application is shown in FIG.

[0054] Icon: 100 - electronic device; 101 - processor; 102 - memory. DETAILED DESCRIPTION

[0055] The present application will be described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that in the drawings or descriptions, similar or identical parts use the same figure numbers, and the implementation methods not shown or described in the drawings are forms known to ordinary technicians in the relevant technical field. In the description of this application, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.

[0056] Please refer to Figure 1 , an embodiment of the present application provides an electronic device 100 that may include a processor 101 and a memory 102. The memory 102 stores a computer program, and when the computer program is executed by the processor 101, the electronic device 100 can perform the corresponding steps in the following video summary generation method.

[0057] In this embodiment, the electronic device 100 may be a personal computer, a laptop computer, a cloud server, etc. It is used to extract semantic features from the original video for which a summary is to be generated based on a first preset strategy to obtain semantic theme features. Then, based on a second preset strategy, image features are extracted from the original video to obtain frame-level image features corresponding to the video frames in the original video. Then, attention features are fused on the semantic theme features and the frame-level image features to obtain image-text fusion features. Finally, according to the image-text fusion features, the video frames are screened to obtain a set consisting of target video frames as a video summary, and the target video frames represent video frames that meet the first preset condition.

[0058] In this embodiment, the processor 101 may be an integrated circuit chip having signal processing capability. The processor 101 may be a general-purpose processor 101. For example, the processor 101 may be a central processing unit 101 (CPU), a digital signal processor 101 (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and may implement or execute the methods, steps and logic block diagrams disclosed in the embodiments of the present application.

[0059] The memory 102 may be, but is not limited to, a random access memory 102, a read-only memory 102, a programmable read-only memory 102, an erasable programmable read-only memory 102, an electrically erasable programmable read-only memory 102, etc. In this embodiment, the memory 102 may be used to store the first preset strategy, the second preset strategy, semantic theme features, frame-level image features, image-text fusion features, the first preset condition, the second preset condition, the video summary, etc. Of course, the memory 102 may also be used to store a program, and the processor 101 executes the program after receiving the execution instruction.

[0060] Please refer to Figure 2The present application also provides a video summary generation method, which can be applied to the above electronic device 100, and each step in the method is executed or implemented by the electronic device 100. The video summary generation method may include the following steps:

[0061] Step 210, based on a first preset strategy, extracting semantic features from the original video for which a summary is to be generated, to obtain semantic topic features;

[0062] Step 220: extract image features from the original video based on a second preset strategy to obtain frame-level image features corresponding to video frames in the original video;

[0063] Step 230, performing attention feature fusion on the semantic topic feature and the frame-level image feature to obtain an image-text fusion feature;

[0064] Step 240 , screening the video frames according to the image-text fusion feature to obtain a set of target video frames as a video summary, wherein the target video frames represent the video frames that meet a first preset condition.

[0065] In the above-mentioned implementation, firstly, based on the first preset strategy, semantic features are extracted from the original video for which a summary is to be generated, and semantic theme features are obtained. Then, based on the second preset strategy, image features are extracted from the original video, and frame-level image features corresponding to the video frames in the original video are obtained. Then, attention features are fused between the semantic theme features and the frame-level image features to obtain image-text fusion features. Finally, based on the image-text fusion features, the video frames are screened to obtain a set consisting of target video frames as the video summary, and the target video frames represent video frames that meet the first preset conditions. In this way, by extracting and combining the semantic features of subtitles and the image features of video frames, the generation of video summaries takes into account both semantic and image information, thereby improving the problem that the traditional video summary generation method lacks attention to semantic information, resulting in the generated video summary being not comprehensive enough and not accurate enough.

[0066] The following is a detailed description of the steps of the video summary generation method:

[0067] In step 210, based on the first preset strategy, extracting semantic features from the original video for which summary is to be generated to obtain semantic topic features may include:

[0068] The bimodal transformer is used to extract semantic features from the original video to obtain the semantic theme features.

[0069] In this embodiment, the bimodal converter includes a bimodal encoder, a bimodal decoder, and a proposal generator;

[0070] The step of extracting semantic features from the original video using a bimodal converter to obtain the semantic theme features may include:

[0071] A1. Generate a word vector library for matching subtitles for the original video through a first pre-trained model;

[0072] A2. extracting visual features from the original video through a second pre-trained model;

[0073] A3. Extracting audio features from the original video through a third pre-trained model;

[0074] A4. Using the bimodal encoder to perform feature fusion on the visual feature and the audio feature to obtain a bimodal fusion feature, wherein the bimodal fusion feature includes a first feature sequence corresponding to the visual feature and a second feature sequence corresponding to the audio feature;

[0075] A5. generating, by the proposal generator, a plurality of prediction proposals and a confidence level corresponding to each prediction proposal in the plurality of prediction proposals according to the first feature sequence and the second feature sequence;

[0076] A6. Cutting the prediction proposal whose confidence satisfies the second preset condition from the multiple prediction proposals as the target proposal, wherein the target proposal includes a first feature sequence that satisfies the second preset condition and a second feature sequence that satisfies the second preset condition;

[0077] A7. Input the target proposal into the bimodal encoder and repeat step A4 to obtain the cut bimodal fusion feature;

[0078] A8. Input the cut bimodal fusion features into the bimodal decoder to match subtitles from the word vector library according to the cut bimodal fusion features as the semantic topic features;

[0079] Repeat the above steps A2 to A8 until a preset mark appears in the video frame.

[0080] In this embodiment, refer to Figure 3 The bimodal encoder may include two layers of encoding structures with the same structure, each layer of the encoding structure may include a self-attention module, a bimodal attention module and a fully connected layer; the bimodal decoder may include a self-attention module, two layers of the same decoding structure, a bridge layer, a fully connected layer and a maximum activation (softmax) fully connected layer, and each layer of the decoding structure includes a bimodal attention module; the proposal generator may include a proposal generation head, a common pooling layer and a screening layer.

[0081] In this embodiment, the first pre-trained model, the second pre-trained model, and the third pre-trained model can be any pre-trained model that has word vector generation function, visual feature extraction, and audio feature extraction, respectively. In this embodiment, the first pre-trained model can be a GloVe pre-trained model, the second pre-trained model can be an I3D pre-trained model, and the third pre-trained model can be a VGGish pre-trained model.

[0082] The bimodal encoder is used to fuse the visual features and audio features to obtain the bimodal fusion features:

[0083]

[0084]

[0085]

[0086]

[0087]

[0088]

[0089] In the formula, A represents the audio features extracted by the third preset model, V represents the video features extracted by the second preset model, n represents the number of layers of the encoding structure, and FullyConnected represents full connection;

[0090] In the formula, MultiHeadAttention represents multi-head attention operation:

[0091]

[0092] Where q, k, and v represent the three inputs of the multi-head attention operation, Attention represents the scaled dot product attention operation, and H represents the number of attention heads. W out They are different trainable weight matrices respectively.

[0093] In this way, the visual features and audio features are fused through the bimodal encoder to obtain the first feature sequence and the second characteristic sequence

[0094] In this embodiment, according to the first feature sequence and the second feature sequence, the corresponding prediction proposals (which can be understood as subtitles) and the confidence corresponding to each prediction proposal are generated by the proposal generation head. Then, a common pool (i.e., common pooling) is constructed according to the prediction proposals corresponding to the first feature sequence and the second feature sequence, and the prediction proposals in the common pool are sorted according to the confidence, and then the prediction proposals that meet the second preset condition are selected from the sorted prediction proposals through the screening layer as the target proposal. Among them, the second preset condition can be flexibly set according to user needs. For example, the preset condition can be the top one hundred prediction proposals with confidence from large to small, and the number of prediction proposals corresponding to the first feature sequence and the second feature sequence are equal.

[0095] In this embodiment, please refer to Figure 3 After the screening layer completes the screening of the target proposal, it determines whether to generate subtitles (i.e., whether the returned target proposal is a null value). If it is not a null value, the visual features and audio features initially input are cut through the target proposal, and the cut visual features and audio features are input into the bimodal encoder for feature fusion again to obtain the cut bimodal fusion features. In this way, it can be more accurately determined whether the subtitles or background interference sounds generated by the bimodal encoder are generated.

[0096] In this embodiment, the cut dual-mode fusion features may include cut audio features and cut video features. Then, the cut dual-mode fusion features are converted into subtitle features through a dual-mode decoder:

[0097]

[0098]

[0099]

[0100]

[0101]

[0102] In the formula, A v represents the audio features after clipping, V a Represents the features of the cut video, represents the set of subtitles generated before the current round of loops in the process of looping through step A9 to generate subtitles, t represents the number of layers of the decoding structure, Represents the generated caption features.

[0103] In this embodiment, after the subtitle features are generated, the subtitle features are mapped to the same dimension as the word vector library extracted by the first pre-trained model through the maximum activation fully connected layer in the dual-mode decoder, so as to match the corresponding subtitle words from the word vector library through the subtitle features as semantic topic features.

[0104] In this embodiment, for the continuously input video stream and audio, the above steps A2 to A8 are repeated to continuously generate subtitles until a preset mark (which can be understood as an end mark) appears in the sampled video stream.

[0105] In this way, by generating subtitles based on the original video through the above step 210, the generation of the video summary focuses on the semantic information in the original video, and the generation of the video summary is more comprehensive and more accurate.

[0106] In step 220, the extracting image features from the original video based on the second preset strategy to obtain frame-level image features corresponding to the video frames in the original video may include:

[0107] Preprocessing the original video to downsample the video frames and remove useless frames and blurred frames in the original video to obtain an image input, wherein the image input includes a plurality of preprocessed video frames;

[0108] The image input is classified using a CLIP model to obtain a category corresponding to each of the multiple preprocessed video frames as the frame-level image feature.

[0109] In this embodiment, video frames are obtained by performing frame sampling on the original video, and image size adjustment (downsampling, also known as downsampling, which can be understood as image reduction), color space conversion, data enhancement, resolution screening and other operations are performed on the video frames to remove useless frames and blurred frames in the video frames, thereby improving the robustness and generalization ability of subsequent video summary generation.

[0110] In this embodiment, refer to Figure 4 , the CLIP model may include a text encoder and an image encoder;

[0111] The classifying the image input by using the CLIP model to obtain a category corresponding to each of the plurality of preprocessed video frames as the frame-level image feature may include:

[0112] Using the text encoder to convert multiple preset category texts into multiple text feature vectors;

[0113] Converting the image input into an image feature vector using the image encoder;

[0114] According to the cosine similarities between the multiple text feature vectors and the image feature vector, a preset category text corresponding to the text feature vector corresponding to the largest cosine similarity is determined as the frame-level image feature.

[0115] In this embodiment, the preset category text (such as dog, pedestrian, apple, kicking ball, etc.) is converted into a plurality of text feature vectors (T1, T2...T N ), and then convert the preprocessed original video (i.e., image input, after preprocessing, the original video is converted into multiple images constituting the original video) into an image feature vector I corresponding to each frame of the image. For each frame of the image, the image feature vector I corresponding to the image is calculated with the text feature vector for cosine similarity (for ease of description, the figure shows the image feature vector I*T N ), and determine the preset category text corresponding to the largest item among all cosine similarity values ​​as the frame-level image feature. For example, if Figure 4 If the cosine similarity of the last item in is greater than that of the other items, the frame-level image feature is the corresponding preset category text: "kick a ball". It can be understood that in practical applications, the preset text category and frame-level image features can be expressed in a non-text format, such as a feature matrix representing various categories, a digital code, etc.

[0116] In this embodiment, it can be understood that the semantic feature extraction and image feature extraction proposed in step 210 and step 220 are independent execution actions, that is, step 210 and step 220 can be executed successively in any order, or can be executed independently and simultaneously, and the execution order of step 210 and step 220 is not specifically limited here.

[0117] In step 230, the performing attention feature fusion on the semantic topic feature and the frame-level image feature to obtain the image-text fusion feature may include:

[0118] Based on the cross attention calculation formula, the semantic theme feature and the frame-level image feature are fused with attention features to obtain the image-text fusion feature:

[0119]

[0120] Where, Q = V 主题 , K=V 图像 , V = V 图像 , V 主题 、V 图像 denote semantic topic features and frame-level image features respectively, and d k represents the dimensions of Q and K, K T represents the transposed matrix of K, and v1 represents the image and text fusion features.

[0121] In step 240, the video frames are screened according to the image-text fusion feature to obtain a set of target video frames as a video summary, which may include:

[0122] According to the image-text fusion feature, a correlation index corresponding to each of the video frames is determined, where the correlation index represents the correlation between the video frame and the semantic theme feature:

[0123] v2=Linear(ReLu(Linear(v1)))

[0124] y = sigmoid(v2)

[0125] In the formula, v1 represents the image-text fusion feature, y represents the correlation index, ReLu represents the first activation function, sigmoid represents the second activation function, and Linear represents linear transformation;

[0126] The video frame whose correlation index satisfies the first preset condition is taken as the target video frame, and a set of all the target video frames is taken as the video summary.

[0127] In this embodiment, after determining the image-text fusion feature, two linear regressions are performed on the image-text fusion feature, and the image-text fusion feature after linear regression is normalized by the sigmoid activation function to obtain a correlation index. Then, video frames whose correlation index meets the first preset condition are selected from multiple video frames in the original video as key frames, and the key frames are spliced ​​into a video summary. Among them, the first preset condition can be flexibly set according to user needs. For example, a correlation threshold can be set, and when any video frame is greater than or equal to the correlation threshold, it is retained, and if it is less than, it is eliminated. Among them, the correlation threshold can adjust the length of the video summary (or the degree of simplification). The higher the correlation threshold, the longer the video summary.

[0128] It is understandable that in practical applications, the above-mentioned video summary generation method can be implemented by encapsulating steps 210 to 240 into a subtitle generation module, an image feature extraction module, an attention feature fusion module, and a summary generation module respectively. The encapsulated modules are combined into a video summary generation network model, and the video summary generation neural network model is trained and the training is completed when its loss function converges, so as to realize the generation of the video summary through the trained video summary generation network model. Among them, the loss function of the video summary generation network model can be as follows:

[0129]

[0130] In the formula, x i represents the estimated value generated during the training process, y irepresents the true value marked in advance, and O (capital letter O, not 0) represents the total sample size.

[0131] In summary, refer to Figure 5 , a video summary generation method provided in this embodiment, inputs the original video into the above-mentioned subtitle generation module, and generates subtitles (i.e., semantic theme features) based on the self-attention mechanism through the BMT subtitle generator (i.e., bimodal transformer) in the subtitle generation module. Input the original video (which can be decomposed into multiple video frames) into the image feature extraction module, preprocess the video frames through the image feature extraction module, and extract features from the preprocessed video frames through the CLIP model in the image feature extraction module to obtain frame-level image features. Then, the attention feature fusion module (i.e., cross-attention mechanism module) is used to fuse the semantic theme features and the frame-level image features to obtain the image-text fusion features. Finally, the image-text fusion features are linearly regressed twice and normalized through the summary generation module, so as to select the video frames that meet the first preset condition from the multiple video frames in the original video as key frames (i.e., recognition results / target video frames), and then the key frames are spliced ​​to obtain key frame summaries, i.e., video summaries.

[0132] Understandably, Figure 1 The structure of the electronic device 100 shown in FIG. 1 is only a schematic diagram of a structure. The electronic device 100 may also include Figure 1 More components are shown. Figure 1 Each component shown in the figure can be implemented by hardware, software or a combination thereof.

[0133] It should be noted that those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the electronic device 100 described above can refer to the corresponding process of each step in the aforementioned method, and will not be elaborated herein.

[0134] The embodiment of the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed on a computer, the computer executes the video summary generation method described in the above embodiment.

[0135] The embodiment of the present application further provides a computer program product, including a computer program. When the computer program is executed by the processor 101, the video summary generation method described in the above embodiment can be implemented.

[0136] Through the description of the above implementation methods, technical personnel in this field can clearly understand that the present application can be implemented by hardware, and can also be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each implementation scenario of the present application.

[0137] In summary, the embodiments of the present application provide a video summary generation method, an electronic device, a storage medium and a program product. In the technical scheme, firstly, based on the first preset strategy, the semantic feature extraction is performed on the original video for which the summary is to be generated to obtain the semantic theme feature. Then, based on the second preset strategy, the image feature extraction is performed on the original video to obtain the frame-level image feature corresponding to the video frame constituting the original video. Then, the attention feature fusion is performed on the semantic theme feature and the frame-level image feature to obtain the image-text fusion feature. Finally, according to the image-text fusion feature, the video frames are screened to obtain a set consisting of target video frames as the video summary, and the target video frames represent the video frames that meet the first preset condition. In this way, by extracting and combining the semantic features of subtitles and the image features of video frames, the generation of the video summary takes into account both semantic and image information, thereby improving the problem that the traditional video summary generation method lacks attention to semantic information, resulting in the generated video summary being not comprehensive enough and inaccurate.

[0138] In the embodiments provided by the present application, it should be understood that the disclosed method can also be implemented in other ways. The method embodiments described above are merely schematic, for example, the flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the implementation of the methods and computer program products according to the multiple embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a part of a module, a program segment or a code, and a part of the module, a program segment or a code includes one or more executable instructions for implementing the specified logical function. It should also be noted that each box in the block diagram and / or the flow chart, and the combination of the boxes in the block diagram and / or the flow chart can be implemented by a dedicated hardware-based system that performs the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions. In addition, each functional module in each embodiment of the present application can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part.

[0139] The above description is only an embodiment of the present application and is not intended to limit the protection scope of the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A video summary generation method, characterized in that: The method comprises: Based on the first preset strategy, semantic features are extracted from the original video for which the summary is to be generated, so as to obtain semantic topic features; Based on a second preset strategy, image feature extraction is performed on the original video to obtain frame-level image features corresponding to video frames in the original video; Performing attention feature fusion on the semantic topic feature and the frame-level image feature to obtain an image-text fusion feature; According to the image-text fusion feature, the video frames are screened to obtain a set consisting of target video frames as a video summary, where the target video frames represent the video frames that meet a first preset condition.

2. The method according to claim 1, characterized in that Based on the first preset strategy, the original video for which the summary is to be generated is subjected to semantic feature extraction to obtain semantic topic features, including: The bimodal transformer is used to extract semantic features from the original video to obtain the semantic theme features.

3. The method according to claim 2, characterized in that The bimodal transformer includes a bimodal encoder, a bimodal decoder and a proposal generator; The step of extracting semantic features from the original video using a bimodal converter to obtain the semantic theme features includes: A1. Generate a word vector library for matching subtitles for the original video through a first pre-trained model; A2. extracting visual features from the original video through a second pre-trained model; A3. Extracting audio features from the original video through a third pre-trained model; A4. Using the bimodal encoder to perform feature fusion on the visual feature and the audio feature to obtain a bimodal fusion feature, wherein the bimodal fusion feature includes a first feature sequence corresponding to the visual feature and a second feature sequence corresponding to the audio feature; A5. generating, by the proposal generator, a plurality of prediction proposals and a confidence level corresponding to each prediction proposal in the plurality of prediction proposals according to the first feature sequence and the second feature sequence; A6. Cutting the prediction proposal whose confidence satisfies the second preset condition from the multiple prediction proposals as the target proposal, wherein the target proposal includes a first feature sequence that satisfies the second preset condition and a second feature sequence that satisfies the second preset condition; A7. Input the target proposal into the bimodal encoder and repeat step A4 to obtain the cut bimodal fusion feature; A8. Input the cut bimodal fusion features into the bimodal decoder to match subtitles from the word vector library according to the cut bimodal fusion features as the semantic topic features; A9. Repeat the above steps A2 to A8 until a preset mark appears in the video frame.

4. The method according to claim 1, characterized in that: The extracting image features of the original video based on the second preset strategy to obtain frame-level image features corresponding to the video frames in the original video includes: Preprocessing the original video to downsample the video frames and remove useless frames and blurred frames in the original video to obtain an image input, wherein the image input includes a plurality of preprocessed video frames; The image input is classified using a CLIP model to obtain a category corresponding to each of the multiple preprocessed video frames as the frame-level image feature.

5. The method according to claim 4, characterized in that The CLIP model includes a text encoder and an image encoder; The using the CLIP model to classify the image input to obtain a category corresponding to each of the plurality of preprocessed video frames as the frame-level image feature includes: Using the text encoder to convert multiple preset category texts into multiple text feature vectors; Converting the image input into an image feature vector using the image encoder; According to the cosine similarities between the multiple text feature vectors and the image feature vector, a preset category text corresponding to the text feature vector corresponding to the largest cosine similarity is determined as the frame-level image feature.

6. The method according to claim 1, characterized in that The performing attention feature fusion on the semantic theme feature and the frame-level image feature to obtain the image-text fusion feature includes: Based on the cross attention calculation formula, the semantic theme feature and the frame-level image feature are fused with attention features to obtain the image-text fusion feature: Where, Q = V 主题 , K=V 图像 , V = V 图像 , V 主题 、V 图像 denote semantic topic features and frame-level image features respectively, and d k represents the dimensions of Q and K, K T represents the transposed matrix of K, and v1 represents the image and text fusion features.

7. The method according to claim 1, characterized in that The step of screening the video frames according to the image-text fusion feature to obtain a set of target video frames as a video summary includes: According to the image-text fusion feature, a correlation index corresponding to each of the video frames is determined, where the correlation index represents the correlation between the video frame and the semantic theme feature: v2=Linear(ReLu(Linear(v1))) y = sigmoid(v2) In the formula, v1 represents the image-text fusion feature, y represents the correlation index, ReLu represents the first activation function, sigmoid represents the second activation function, and Linear represents linear transformation; The video frame whose correlation index satisfies the first preset condition is taken as the target video frame, and a set of all the target video frames is taken as the video summary.

8. An electronic device, characterized in that: The electronic device comprises a processor and a memory coupled to each other, wherein the memory stores a computer program. When the computer program is executed by the processor, the electronic device executes the method as claimed in any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed on a computer, the computer is enabled to execute the method according to any one of claims 1 to 7.

10. A computer program product, characterized in that The invention comprises a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • New media image processing method and system based on AI large model

    CN120563656A