Video outline generation method and device, computer equipment and medium

By extracting key video clips and audio highlight clips, combining the motion trajectory of the moving target, a video outline is generated, and the problems of redundancy and low computing efficiency of video abstracts in the prior art are solved, and the refinement and readability of information are improved.

CN120281995AActive Publication Date: 2025-07-08ZHEJIANG HUAZHI WANXIANG TECHNOLOGY CO LTD
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202510766021.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

The summary video generated by existing video digest technology is relatively redundant, has low computing efficiency, and cannot efficiently reflect the key information, logic and time chaos in the original video.

Method used

By extracting key video clips and audio highlights, combining the motion trajectory of the moving target, a video outline is generated, including obtaining the audio feature vector of the initial video for emotional feature detection, determining the audio highlights, and filtering out the target video highlights based on the motion trajectory of the moving target in the video highlights.

Benefits of technology

The generated video outline information is refined and readable, and can efficiently extract and recombine key clips in the video, ensuring the comprehensiveness and readability of the video outline.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281995A_ABST
    Figure CN120281995A_ABST
Patent Text Reader

Abstract

The invention relates to a video outline generation method and device, computer equipment and a medium. Comprising the following steps: acquiring an initial video, and determining video key segments in the initial video; extracting an audio feature vector of the initial video, performing emotional feature detection on the audio feature vector, determining a basic audio clip in the audio feature vector based on an emotional feature detection result, and detecting rhythm intensity and audio short-time energy in the basic audio clip, jointly determining an audio highlight segment according to a rhythm intensity detection result and an audio short-time energy detection result; determining at least one video highlight segment in the initial video according to the video key segment and the audio highlight segment; and screening out a target video highlight clip from the video highlight clips based on the motion trail of the motion target in the video highlight clips, and generating a video outline based on the target video highlight clip. According to the invention, a high-quality video outline can be generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video processing technology, and in particular to a method, device, computer equipment and medium for generating a video outline. Background Art

[0002] With the development of the digital information age, digital video content has experienced explosive growth. Video summarization technology analyzes and processes the structure and content of the video, extracts representative video frames or video clips from the original video, and arranges and combines these meaningful video frames or clips to eventually form a shorter and more compact summary video. This summary video contains a large amount of semantic information in the original video and can express the meaning and emotion of the original video to a certain extent.

[0003] However, the summary video extracted by the current video clip extraction technology is still relatively redundant and has a relatively low computational efficiency, and cannot effectively reflect the key information in the original video. Furthermore, the summary video generally has a relatively chaotic logic and time.

[0004] At present, there is no effective solution to the problem of how to efficiently extract key segments from the original video and generate a video segment with high readability and viewing value in the existing technology. Summary of the invention

[0005] Based on this, it is necessary to provide a video outline generation method, device, computer equipment and medium to address the above technical problems.

[0006] In a first aspect, the present application provides a method for generating a video outline. The method comprises:

[0007] Obtain an initial video and determine a key video segment in the initial video;

[0008] Extracting the audio feature vector of the initial video, and performing emotional feature detection on the audio feature vector, determining the basic audio segment in the audio feature vector based on the result of the emotional feature detection, detecting the rhythm intensity and audio short-time energy in the basic audio segment, and jointly determining the audio highlight segment based on the rhythm intensity detection result and the audio short-time energy detection result;

[0009] Determine at least one video highlight segment in the initial video according to the video key segment and the audio highlight segment;

[0010] Based on the motion trajectory of the moving target in the video highlight clips, the target video highlight clips are screened out from the video highlight clips, and the video outline is generated based on the target video highlight clips.

[0011] In one embodiment, the rhythm intensity and short-time audio energy in the base audio segment are detected, and the audio highlight segment is jointly determined according to the detection results of the rhythm intensity and the short-time audio energy, including:

[0012] Perform short-time audio energy detection on the base audio segment, calculate the root mean square of the short-time audio energy detection result, and determine the volume of each frame of audio in the base audio segment based on the root mean square calculation result;

[0013] Determine the rhythm intensity of each frame of audio according to the difference between the spectral amplitude of each frame of audio in the base audio segment and the spectral amplitude of the historical frame;

[0014] Perform fusion weighted calculation on the volume of each frame of audio and the rhythm intensity of each frame of audio to obtain the highlight score of each frame of audio in the base audio segment;

[0015] Based on all audio frames with highlight scores greater than a preset score threshold, obtain the audio highlight segment.

[0016] In one embodiment, determining the video key segments in the initial video includes:

[0017] Obtain the video key frames in the initial video, and divide the video material in the initial video into multiple video segments based on the video key frames;

[0018] Traverse all video segments, and determine the video key segments based on the video feature vectors of each video segment.

[0019] In one embodiment, determining at least one video highlight segment in the initial video according to the video key segments and the audio highlight segments includes:

[0020] In the time dimension, based on the timestamps corresponding to the video key segments and the timestamps corresponding to the audio highlight segments, extract the initial video highlight segments from the initial video;

[0021] Perform fusion processing on the image feature vector and the audio feature vector corresponding to the initial video highlight segment to obtain the fusion feature vector of the initial video highlight segment;

[0022] Obtain the prompt words corresponding to the video outline, and extract the text feature vector of the prompt words;

[0023] Perform similarity matching between the text feature vector and the fusion feature vector to obtain the similarity corresponding to each fusion feature vector, and determine the video highlight segment based on the fusion feature vectors with similarities greater than a preset similarity threshold.

[0024] In one embodiment, based on the timestamps corresponding to the key video segments and the timestamps corresponding to the highlight audio segments, extracting the initial video highlight segments from the initial video includes:

[0025] Restoring the timestamps corresponding to the highlight audio segments to the time axis of the initial video to obtain the audio highlight indexes corresponding to each highlight audio segment;

[0026] Restoring the timestamps corresponding to the key video segments to the time axis of the initial video to obtain the key video frame indexes corresponding to each key video segment;

[0027] Based on the audio highlight indexes and the key video frame indexes, obtaining at least one initial video highlight segment in the initial video.

[0028] In one embodiment, based on the motion trajectories of the moving objects in the video highlight segments, screening out the target video highlight segments from the video highlight segments includes:

[0029] Extracting the motion trajectories of the moving objects in each video highlight segment;

[0030] Performing anomaly detection on each motion trajectory, and determining the first motion trajectory in the motion trajectory based on the anomaly detection result; wherein, the first motion trajectory includes a plurality of first sub-trajectories;

[0031] Detecting the moving objects in the first motion trajectory, and rearranging the first sub-trajectories included in the first motion trajectory based on the moving objects to obtain at least one second motion trajectory, wherein each second motion trajectory corresponds to a moving object;

[0032] Calculating the scores of the second motion trajectories based on the dynamic programming algorithm, and determining the motion trajectories with scores greater than the preset second score threshold as the target motion trajectories; wherein, based on the video highlight scores, audio highlight scores and motion trajectory highlight scores of the second motion trajectories, the score calculation of the second motion trajectories is completed;

[0033] Based on the target motion trajectories, determining the corresponding target video highlight segments in the initial video.

[0034] In one embodiment, based on the target motion trajectories, determining the corresponding target video highlight segments in the initial video includes:

[0035] Determining the target sub-motion trajectories in the target motion trajectories, and rearranging the target sub-motion trajectories in chronological order to obtain the third motion trajectories corresponding one-to-one to the target sub-motion trajectories;

[0036] Based on the third motion trajectories, obtaining the target video highlight segments corresponding to the initial video.

[0037] Second aspect, the present application also provides a video outline generation device. The device includes:

[0038] An acquisition module, configured to acquire an initial video and determine video key segments in the initial video;

[0039] A calculation module, configured to extract an audio feature vector of the initial video, perform emotion feature detection on the audio feature vector, determine a basic audio segment in the audio feature vector based on the result of the emotion feature detection, detect the rhythm intensity and audio short-time energy in the basic audio segment, and jointly determine an audio highlight segment according to the rhythm intensity detection result and the audio short-time energy detection result; combine the video key segments and the audio highlight segments to determine at least one video highlight segment in the initial video;

[0040] A generation module, configured to screen out target video highlight segments from the video highlight segments based on the motion trajectories of moving objects in the video highlight segments, and generate a video outline based on the target video highlight segments, where the video outline characterizes the motion trajectories of the moving objects in the target video highlight segments.

[0041] Third aspect, the present application also provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0042] Acquire an initial video and determine video key segments in the initial video;

[0043] Extract an audio feature vector of the initial video, perform emotion feature detection on the audio feature vector, determine a basic audio segment in the audio feature vector based on the result of the emotion feature detection, detect the rhythm intensity and audio short-time energy in the basic audio segment, and jointly determine an audio highlight segment according to the rhythm intensity detection result and the audio short-time energy detection result;

[0044] Determine at least one video highlight segment in the initial video according to the video key segments and the audio highlight segments;

[0045] Screen out target video highlight segments from the video highlight segments based on the motion trajectories of moving objects in the video highlight segments, and generate a video outline based on the target video highlight segments.

[0046] Fourth aspect, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0047] Acquire an initial video and determine video key segments in the initial video;

[0048] Extract the audio feature vectors of the initial video, perform emotion feature detection on the audio feature vectors, determine the basic audio segments in the audio feature vectors based on the results of the emotion feature detection, detect the rhythm intensity and short-time audio energy in the basic audio segments, and jointly determine the audio highlight segments according to the rhythm intensity detection results and the short-time audio energy detection results;

[0049] Determine at least one video highlight segment in the initial video according to the video key segments and the audio highlight segments;

[0050] Based on the motion trajectories of the moving objects in the video highlight segments, screen out the target video highlight segments from the video highlight segments, and generate a video outline based on the target video highlight segments.

[0051] The above video outline generation method, device, computer device and medium extract video key segments based on the initial video, determine the audio highlight segments according to emotion feature detection, rhythm intensity detection and short-time audio energy detection, determine multiple video highlight segments in the initial video according to the video key segments and the audio highlight segments, and finally screen out the target video highlight segments from the video highlight segments according to the motion trajectories of the moving objects in the video highlight segments, and generate a video outline. The video outline generated by this application ensures the refinement of information, and also ensures the comprehensiveness and readability of the video outline. Description of the Drawings

[0052] Figure 1 It is an application environment diagram of the video outline generation method in an embodiment;

[0053] Figure 2 It is a flowchart of the video outline generation method in an embodiment;

[0054] Figure 3 It is a flowchart of generating a fusion feature vector in an embodiment;

[0055] Figure 4 It is a flowchart of the video outline generation method in a preferred embodiment;

[0056] Figure 5 It is a structural block diagram of the video outline generation device in an embodiment;

[0057] Figure 6 It is an internal structure diagram of a computer device in an embodiment. Detailed Embodiments

[0058] In order to make the purpose, technical solutions and advantages of this application clearer, the following further describes this application in detail with reference to the drawings and embodiments. It should be understood that the specific embodiments described here are only used to explain this application, and are not used to limit this application.

[0059] The video outline generation method provided by the embodiments of the present application can be applied to an application environment as shown in Figure 1 . Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or can be placed in the cloud or on other network servers. Obtain the initial video, determine the video key segments in the initial video, extract the audio feature vectors in the initial video, perform emotion feature detection on the audio feature vectors, determine the basic audio segments in the audio feature vectors based on the results of the emotion feature detection, detect the rhythm intensity and audio short-time energy in the basic audio segments, and jointly determine the audio highlight segments according to the rhythm intensity detection results and the audio short-time energy detection results. Combine the video key segments and the audio highlight segments to determine at least one video highlight segment in the initial video; based on the motion trajectories of the moving objects in the video highlight segments, screen out the target video highlight segments from the video highlight segments, and generate a video outline based on the target video highlight segments, where the video outline characterizes the motion trajectories of the moving objects in the target video highlight segments. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers.

[0060] In one embodiment, as shown in Figure 2 , a video outline generation method is provided. Taking the server in Figure 1 as an example for description, the method includes the following steps:

[0061] Step S210, obtain the initial video and determine the video key segments in the initial video.

[0062] Specifically, in the present application, the video outline refers to extracting some core segments from the initial video, and rearranging the key target motion trajectories in the core segments according to a certain method and a certain logic, so as to obtain the video outline. The video outline carries the key information in the original video and the important visual information that the user wants to obtain. Compared with the existing video summary extraction technology, the video outline extraction technology in the present application can analyze and reorganize the extracted video materials based on the key segments in the original video, combined with target tracking and video content reorganization technology, and finally obtain video segments with higher readability, ornamental value, and more compact content, that is, the above-mentioned video outline.

[0063] Among them, the key video segments refer to the key and core video material (only video without audio) in the initial video, which can be one or more video segments (only video without audio), or one or more images. There are various methods for extracting the key video segments from the initial video. In some embodiments, the video material in the initial video can be extracted first, and then the video material is segmented. Detection is performed on each video material segment. For example, it can be detected whether there are significant changes in the content of the video segment; whether there are important people, objects, scenes, etc. in the video material segment can be detected through a neural network; whether there is violent movement in the video material segment can be judged through an optical flow algorithm, etc. Through the above methods, the key video segments can be extracted from the video material.

[0064] Step 220: Extract the audio feature vectors of the initial video, perform emotion feature detection on the audio feature vectors, determine the basic audio segments in the audio feature vectors based on the results of the emotion feature detection, detect the rhythm intensity and short-time audio energy in the basic audio segments, and jointly determine the audio highlight segments according to the rhythm intensity detection results and the short-time audio energy detection results; determine at least one video highlight segment in the initial video according to the image highlight segment and the audio highlight segment.

[0065] Specifically, frame-level audio feature extraction is performed on the audio of the initial video, that is, only the audio is processed at this time without processing the image. The audio feature vectors of the initial video are obtained, and emotion detection is performed on the audio feature vectors. It can be understood that in this field, emotion detection is a technology for inferring the emotional state of the speaker through audio features (such as pitch, volume, speech rate, etc.). Common emotion detection methods include, but are not limited to, traditional feature extraction based on Support Vector Machine (SVM), and feature extraction completed by a neural network based on deep learning, so as to determine the emotional state of the speaker, etc. Among them, this embodiment provides a specific emotion feature extraction method, including performing frame-level audio feature extraction on the audio of the initial video through the HuBERT (Hidden-Unit BERT) model fine-tuned on the Chinese emotion dataset, and then performing mutation detection through the cosine distance to find the mutation points in the audio. Segment the original audio according to the detected mutation points, and then perform emotion recognition on the segmented audio through the trained emotion classifier to extract the above basic audio segments with drastic emotional changes. The emotions here can include various emotions such as anger, happiness, and surprise.

[0066] After detecting the basic audio segment, detect the rhythm intensity and short-time audio energy in the basic audio segment. Among them, the audio can be converted into a spectrum, and the rhythm intensity of each frame of audio can be calculated according to the spectral amplitude of each frame. After calculating the rhythm intensity of each frame, the audio rhythm mutation points in the basic audio segment can be determined according to the rhythm intensity. It can be understood that the time period where the audio rhythm mutation points are located is probably the key segment, which can also be called the highlight segment. And the short-time energy of each frame can be calculated. The short-time energy can be understood as the audio signal value, and the change of the audio volume can be reflected based on the audio signal value. Furthermore, the peaks with higher short-time energy in the basic audio segment can be determined according to the short-time audio energy. Similarly, it can be understood that the segment where the short-time energy peak is located is probably the highlight segment. In summary, the monitoring results of the rhythm intensity and the detection results of the short-time energy can be combined to determine the audio highlight segment. For example, the segments corresponding to each audio rhythm mutation point and the segments corresponding to the peaks with higher short-time energy can be used as the above-mentioned audio highlight segments. Or the intersection between the segments corresponding to the audio rhythm mutation points and the segments corresponding to the peaks with higher short-time energy can be used as the above-mentioned audio highlight segment, and so on.

[0067] It can be understood that the above video highlight segment is a video segment that includes both audio and the corresponding image. Therefore, it is necessary to combine the video key segment with the audio highlight segment to obtain the video highlight segment. Specifically, the corresponding video segment can be extracted from the initial video based on the time corresponding to the video key segment (this video segment is intercepted from the initial video and includes the audio and the corresponding video), and the corresponding video segment can be extracted from the initial video based on the time corresponding to the audio highlight segment (similarly, this video segment is intercepted from the initial video and includes the audio and the corresponding video). Combining all the video segments can obtain the above video highlight segment. Similarly, the intersection between the video segment corresponding to the video key segment in the original video and the video segment corresponding to the audio highlight segment in the original video can be calculated, and the video segment that is the intersection can be determined as the above video highlight segment.

[0068] Step S230, based on the motion trajectories of the moving objects in the video highlight segment, screen out the target video highlight segment from the video highlight segment, and generate a video outline based on the target video highlight segment.

[0069] Specifically, generally multiple video highlight segments will be obtained through the above method. To ensure the readability of the finally generated video outline, these video highlight segments need to be reorganized. Specifically, determine the motion trajectories of the moving objects included in each video highlight segment, screen out the segments including the moving objects from the video highlight segments to obtain the target video highlight segments, so as to further delete the redundant pictures, and then rearrange the target video highlight segments in chronological order, and then the above video outline can be obtained.

[0070] Through steps S210 to S230, a video outline of any initial video can be efficiently generated. On the basis of ensuring the extraction of key segments in the initial video, the extracted materials can be rearranged according to the moving objects and the time sequence of video segments, which not only ensures the refinement and comprehensiveness of the information in the finally generated video outline, but also ensures the readability of the video outline, enabling the audience to more smoothly understand the content of the initial video through the video outline.

[0071] In some embodiments, the rhythm intensity and short-time audio energy in the basic audio segment are detected, and the audio highlight segment is jointly determined according to the rhythm intensity detection result and the short-time audio energy detection result, including:

[0072] Performing short-time audio energy detection on the basic audio segment, calculating the root mean square of the short-time audio energy detection result, and determining the volume of each frame of audio in the basic audio segment based on the root mean square calculation result;

[0073] Determining the rhythm intensity of each frame of audio according to the difference between the spectral amplitude of each frame of audio in the basic audio segment and the spectral amplitude of the historical frame;

[0074] Performing fusion weighted calculation on the volume of each frame of audio and the rhythm intensity of each frame of audio to obtain the highlight score of each frame of audio in the basic audio segment;

[0075] Based on all audio frames with highlight scores greater than a preset score threshold, the audio highlight segment is obtained.

[0076] Specifically, performing short-time audio energy detection on the basic audio segment, the short-time audio energy (Short-Term Energy, STE) is calculated as follows:

[0077] ;

[0078] where, E n refers to the short-time energy of the nth frame, x(n) refers to the audio signal value, m refers to the local sampling point index within the frame, N refers to the frame length, and then the root mean square (Root Mean Square, RMS) of the short-time audio energy detection result is calculated to measure the change in volume, and the calculation is as follows:

[0079] ;

[0080] Thus, the volume of the nth frame of audio can be calculated.

[0081] Further, perform spectral conversion processing on the basic audio segment. Determine the rhythm intensity of each frame according to the difference between the spectral amplitude value of each frame of audio after spectral conversion and the spectral amplitude value of the historical frame (preferably the previous frame of the current frame). The calculation is as follows:

[0082] ;

[0083] Among them, O(n) refers to the rhythm intensity of the nth frame, S(n,k) refers to the spectral amplitude value of the kth frequency, and ω(k) refers to the weight. This weight can be used to control the influence of the changes of each frequency component (such as sounds in different frequency bands) on the rhythm intensity (or the perception of audio changes). The importance of different frequency bands in rhythm analysis can be adjusted according to actual needs. In order to treat all frequency points equally, this weight can be set to 1. At this time, it is considered that all frequency components have equal contributions to the rhythm. Similarly, in some complex scenarios, ω(k) can be adjusted to more accurately capture the rhythm changes or emotional fluctuations in the audio.

[0084] After calculating the volume of each frame of audio and the rhythm intensity of each frame of audio, for the convenience of calculation, they can be normalized first. The calculation is as follows:

[0085] ;

[0086] ;

[0087] Among them, R(n) is the RMS energy of the nth frame, and O(n) is the rhythm intensity of the nth frame. and are the values after normalization calculation, and the range is [0,1]. Further, perform fusion weighted calculation on the volume of each frame of audio and the rhythm intensity of each frame of audio to obtain the highlight score of each frame of audio in the basic audio segment. The combined highlight score of each frame of audio is calculated as follows:

[0088] ;

[0089] Among them, H(n) refers to the highlight score of the nth frame, α∈[0,1] is the weight parameter, which represents the importance balance between volume and rhythm. In practical applications, the value of α can be adjusted according to actual needs to correspondingly adjust the importance degree of volume and rhythm. In some embodiments, α can be set to 0.5, that is, the importance of volume and rhythm is set to the same level.

[0090] Finally, the audio frames with highlight scores greater than the preset score threshold can be used as audio highlight segments. In some preferred embodiments, the highlight points can be obtained by the threshold median method. The calculation is as follows:

[0091] ;

[0092] ;

[0093] Among them, T H is the 90th percentile of all high - light score values, that is, the set high - light detection threshold. HighlightIndices represents the frame indices of audio frames whose all high - light detection scores are greater than the threshold. Finally, by converting the frame indices into timestamps, the corresponding audio high - light segments can be obtained.

[0094] In some of these embodiments, determining the video key segments in the initial video includes:

[0095] Obtaining the video key frames in the initial video, and dividing the video material in the initial video into multiple video segments based on the video key frames;

[0096] Traversing all video segments, and determining the video key segments based on the video feature vectors of each video segment.

[0097] Specifically, this embodiment provides a method for obtaining video key segments, including first determining the video key frames in the initial video. These video key frames generally have scene switches, so the video material of the initial video can be divided into multiple video segments according to different scenes. In practical applications, ffmpeg (Fast Forward Moving Picture Experts Group) can be used to separate the video and audio materials from the original video material. Then, for the video material, the image key frames can be extracted based on scene changes, and the calculation is as follows:

[0098] ;

[0099] Among them, I t represents the video key frame at time t with a large scene change.

[0100] Then, feature extraction is performed for each video segment, and then the video key segments in the initial video are determined. Among them, there are methods for determining video key segments. In some embodiments, high - light segments can be screened out by scoring according to the features of different video segments. The methods for calculating the scores of each video segment are: calculating the high - light score of each video segment, and determining the above - mentioned video key segments as the video segments with higher high - light scores. The calculation of the high - light score is as follows:

[0101]

[0102] Among them, highlight_score is the final score of each video clip, visual_score is the visual rating of the video clip, which can be generated by relevant technical personnel or through a trained neural network for each video clip. Similarly, the above emotion_score is the emotion rating, which can be generated by relevant technical personnel or through methods such as neural networks for the emotion rating of the people in each video clip. Similarly, the above object_score is the object rating, that is, for the objects appearing in the video clip, it can be rated by relevant technical personnel or neural networks. In some other embodiments, it is possible to judge whether there are drastic changes in the video clip according to the inter-frame image changes. If so, the video clip is determined as a key video clip, and so on. The above gives two methods for determining key video clips, and relevant technical personnel can also design methods for determining key video clips according to actual needs.

[0103] In some of the embodiments, determining at least one video highlight clip in the initial video according to the key video clips and audio highlight clips includes:

[0104] In the time dimension, based on the timestamps corresponding to the key video clips and the timestamps corresponding to the audio highlight clips, extract the initial video highlight clips from the initial video;

[0105] Perform fusion processing on the image feature vectors and audio feature vectors corresponding to the initial video highlight clips to obtain the fusion feature vectors of the initial video highlight clips;

[0106] Obtain the prompt words corresponding to the video outline and extract the text feature vectors of the prompt words;

[0107] Perform similarity matching between the text feature vectors and the fusion feature vectors to obtain the similarity corresponding to each fusion feature vector, and determine the video highlight clips based on the fusion feature vectors with a similarity greater than a preset similarity threshold.

[0108] Specifically, according to the time corresponding to the key video clips and the time corresponding to the audio highlight clips, extract the initial video highlight clips from the initial video, where the initial video highlight clips are segments intercepted from the initial video according to the above time information and contain audio and corresponding videos.

[0109] Then, the image feature vectors and audio feature vectors of each frame in the initial video highlight segment are fused. In the time dimension, the extracted audio feature vectors and image feature vectors are fused to obtain the above-mentioned fused feature vectors, that is, the audio feature vectors and image feature vectors at the same moment, and the fused feature vectors corresponding to the corresponding moments are fused. Specifically, the corresponding feature vectors of the audio and video are extracted respectively, that is, the above-mentioned image feature vectors and audio feature vectors. The calculation of the image feature vectors and audio feature vectors is as follows:

[0110] ;

[0111] ;

[0112] Among them, the image feature vector F V can be extracted through the CLIP4Clip model, and the audio feature vector F A can be extracted through the Mel spectrogram. Further, in order to be consistent in actual calculation later, the audio feature vector needs to be unified and normalized to the time axis corresponding to the original video frame rate and aligned with the video frame feature vector. As Figure 3 shown, the video frame feature vector and the audio feature vector are fused in the time dimension to obtain the fused feature vector. The fused feature vector at the final t moment can be denoted as: Ft = [F V ’, F A ’]. It can be understood that the video frame feature vector is generally two-dimensional, and the audio feature vector is one-dimensional.

[0113] Then, the text feature vector and the fused feature vector are subjected to similarity matching, and the cosine similarity between the text feature vector F T and the fused feature vector Ft is calculated. Among them, the text feature vector can be obtained according to the prompt words input by the user. For example, in a football game, the user can input words such as “football” and “shoot” as prompt words. Specifically, the text feature vector of the input prompt words can also be determined through the CLIP4Clip model. The above similarity threshold can be set dynamically. For example, it can be set as the average value of the scores ranked in the top n of the similarity scores. n can be set manually, and is generally set to 3. The video segment corresponding to the fused feature vector with a similarity greater than the above similarity threshold is recognized as the video highlight segment. In this application, the fused feature vector in the initial video highlight segment is further matched with the prompt words corresponding to the video outline, so that on the basis of extracting the key segments in the initial video, and then matching again based on the prompt words input by the user, a video highlight segment that is more appropriate for the required video outline can be obtained.

[0114] In some of these embodiments, based on the timestamps corresponding to the key video segments and the timestamps corresponding to the audio highlight segments, extracting an initial video highlight segment from the initial video includes:

[0115] Restoring the timestamps corresponding to the audio highlight segments to the timeline of the initial video to obtain the audio highlight indices corresponding to each audio highlight segment;

[0116] Restoring the timestamps corresponding to the key video segments to the timeline of the initial video to obtain the key video frame indices corresponding to each key video segment;

[0117] Based on the audio highlight indices and the key video frame indices, obtaining at least one initial video highlight segment in the initial video.

[0118] Specifically, this embodiment provides a method for extracting the corresponding video highlight segment after separately calculating the audio highlight segment and the key video segment.

[0119] Converting the timestamps corresponding to the key video segments and the audio highlight segments into frame indices, and then the corresponding segments in the original video or audio can be extracted according to these frame indices for fusion to obtain the above-mentioned video highlight segment. In some embodiments, the extracted audio highlight segments and image highlight segments may not be exactly the same. Therefore, the union of the audio highlight segments and the image highlight segments can be taken as the candidate highlight segments. For example, if the extracted audio highlight segment is from the 2nd second to the 5th second in the original audio, and the extracted video highlight segment is from the 3rd second to the 6th second in the original video, at this time, the video segment from the 2nd second to the 5th second can be extracted from the original video according to the audio highlight segment, and the video segment from the 3rd second to the 6th second can be extracted from the original video according to the video highlight segment. Combining these two video segments, that is, taking the video segment from the 2nd second to the 6th second as the above-mentioned video highlight segment.

[0120] In some of these embodiments, based on the motion trajectories of the moving objects in the video highlight segment, screening out the target video highlight segments from the video highlight segment includes:

[0121] Extracting the motion trajectories of the moving objects in each video highlight segment;

[0122] Performing anomaly detection on each motion trajectory, and determining the first motion trajectory in the motion trajectory based on the anomaly detection result; wherein, the first motion trajectory includes a plurality of first sub-trajectories;

[0123] Detecting the moving objects in the first motion trajectory, and rearranging the first sub-trajectories included in the first motion trajectory based on the moving objects to obtain at least one second motion trajectory, wherein each second motion trajectory corresponds to a moving object;

[0124] Calculate the score of the second motion trajectory based on the dynamic programming algorithm, and determine the target motion trajectory whose score is greater than the preset second score threshold; wherein, based on the video highlight score, audio highlight score and motion trajectory highlight score of the second motion trajectory, complete the score calculation of the second motion trajectory;

[0125] Based on the target motion trajectory, determine the corresponding target video highlight segment in the initial video.

[0126] Specifically, after obtaining multiple video highlight segments, in order to improve the readability of the subsequent generated video outline and also to further refine the extracted highlight video, in this embodiment, according to the motion trajectory of the moving target, the target video highlight segments are screened out from the video highlight segments. Specifically, the above video highlight segments usually contain rich moving objects. First, extract the motion trajectory of the moving target in each video highlight segment. An existing neural network, such as YOLO-v11, can be used to detect and track the moving target in each video highlight segment, and according to the tracking results, the motion trajectory of the moving target can be obtained. And in practical applications, since the motion trajectory of the moving target may have phenomena such as jitter, breakage or even drift, in some preferred embodiments, the motion trajectory of the moving target can be optimized. And since the optimized motion trajectory may belong to multiple targets or different video segments, trajectory rearrangement and time rearrangement are required. Among them, trajectory optimization includes trajectory smoothing. When performing trajectory smoothing, in order to reduce jitter, double exponential smoothing can be used, and the calculation is as follows:

[0127] ;

[0128] ;

[0129] Wherein, and are the coordinates after smoothing, x t and y t are the original coordinates, and α is the smoothing factor, which can usually be set as 0.1 ≤ α ≤ 0.3.

[0130] After extracting the motion trajectory of the moving target corresponding to each video highlight segment, abnormal detection can be performed on each motion trajectory to obtain the first motion trajectory in each motion trajectory. It can be understood that in this embodiment, through abnormal detection, the sudden and violently changing moments during the movement of the moving target can be found, and these moments are often the key points of the highlight moments. If the motion abnormality is ignored, wonderful video segments may be missed. Among them, the above abnormal detection can screen out the above first motion trajectory with abnormalities through the acceleration a t and the calculation is as follows:

[0131] ;

[0132] Among them, is the speed at time t, v t-1 is the speed at time t-1, and △t is the time difference between time t-1 and time t. If a t is greater than the preset anomaly detection threshold, it can be considered that the target has a sudden movement (such as sudden stop, sprint, fall, etc.) at time t, and this time period can be recorded as a high-light segment of the abnormal trajectory.

[0133] In practical applications, in the segments corresponding to the above first motion trajectory, there may be multiple moving targets, and there are different first sub-trajectories corresponding to different moving targets. The relationship between the moving target and the first sub-trajectory is one-to-many or one-to-one, and the moving targets corresponding to each motion trajectory are not clear. Therefore, in order to improve the quality of the subsequent generated video outline and to improve the logic of video display, it is necessary to rearrange the first sub-trajectories based on the moving target, that is, to regularize the sub-trajectories belonging to the same target into the same segment to obtain the second motion trajectory. At this time, the relationship between the second motion trajectory and the moving target is a one-to-one correspondence.

[0134] Then, in order to ensure the quality of the highlight video, the score of the second motion trajectory can be calculated according to the dynamic programming algorithm, and the motion trajectory with a score greater than the preset second score threshold is determined as the target motion trajectory. The calculation of the trajectory score is as follows:

[0135] ;

[0136] Among them, represents the highlight score of trajectory i, △T i represents the duration of the video segment corresponding to trajectory i, T max represents the maximum duration of the video outline. Among them, the highlight score of trajectory i can be calculated by the following formula.

[0137] ;

[0138] Among them, represents the visual highlight score of the i-th trajectory, represents the audio highlight score of the i-th trajectory, represents the motion highlight score of the i-th trajectory (which can be obtained through the above trajectory anomaly detection), ω V , ω A , ω M are weight coefficients, satisfying ω V +ω A +ω M =1, and can be set according to the relative importance of video, audio and motion.

[0139] Specifically, the sports highlight score is calculated as follows:

[0140] ;

[0141] where a t represents the acceleration at the t-th frame, θ represents the set abnormal acceleration threshold, L represents the indicator function, which is 1 if the condition is satisfied and 0 otherwise, and T i represents the total number of frames of the trajectory.

[0142] The visual highlight score is calculated as follows:

[0143] ;

[0144] where c t represents the confidence of the detection box, and △s t represents the change rate of the detection box area, and there is:

[0145] ;

[0146] and d t represents the distance between the normalized object center and the center of the screen. a1, a2, and a3 are weight coefficients, satisfying a1 + a2 + a3 = 1, represents the total number of frames of the trajectory.

[0147] The audio highlight score is calculated as follows:

[0148] ;

[0149] where RMS t represents the root mean square of the short-time energy detection result of the audio at time t, and O(t) represents the rhythm intensity at time t, is a parameter that can adjust the importance degree of the rhythm and is less than 1.

[0150] In summary, through the above method, the extracted video highlight segments can be further refined to obtain the target video highlight segments after trajectory rearrangement. Moreover, since the trajectories are rearranged in this embodiment, the readability of the finally generated video outline can be effectively improved.

[0151] In some of the embodiments, based on the target motion trajectory, determining the corresponding target video highlight segments in the initial video includes:

[0152] Determining the target sub-motion trajectories in the target motion trajectory, rearranging the target sub-motion trajectories in chronological order to obtain the third motion trajectory corresponding to the target sub-motion trajectories;

[0153] Based on the third motion trajectory, obtaining the target video highlight segments corresponding to the initial video.

[0154] After obtaining the target motion trajectory, the target sub-motion trajectories included in the target motion trajectory can be determined. Each target sub-motion trajectory corresponds to a motion target. To improve the logic of subsequent video display, the target sub-motion trajectories can be rearranged in chronological order to obtain a third motion trajectory corresponding to the target sub-motion trajectories. Chronological rearrangement can make the order of highlight events more coherent, enhance the viewing experience, and at the same time reduce redundancy and avoid repeated appearance of the same scene. It can be understood that when rearranging the motion trajectory chronologically, in essence, the video segments corresponding to the motion trajectory are rearranged in chronological order.

[0155] In some preferred embodiments, when rearranging the target sub-motion trajectories chronologically, the optimal path algorithm is used for sorting. Specifically:

[0156] ;

[0157] Among them, represents the similarity between segment and . The smaller the better. N is the number of segments. By optimizing the sorting, the transition between each target sub-motion trajectory can be made smoother. Finally, after obtaining the third motion trajectory, based on the segments corresponding to the third motion trajectory in the original video, the above-mentioned target video highlight segments are obtained.

[0158] In summary, the target video highlight segments corresponding to the original video can be obtained. By integrating the target video highlight segments, the above-mentioned video outline can be obtained.

[0159] This application also provides a preferred embodiment for generating a video outline, Figure 4 which is a method for generating a video outline in a preferred embodiment.

[0160] Step S410, obtain the original video and complete the preprocessing of the original video, such as noise reduction, feature extraction, etc.

[0161] Step S420, perform highlight detection on the original video, that is, perform highlight detection on the audio of the original video respectively to obtain audio highlight segments. Similarly, perform highlight detection on the video frame materials of the original video to obtain the above-mentioned video key segments. Finally, multiple video highlight segments are obtained comprehensively.

[0162] Step S430, detect and track the motion targets in the video highlight segments to obtain a first motion trajectory. Among them, the first motion trajectory includes multiple first sub-trajectories.

[0163] Step S440: Optimize the trajectory and rearrange the time of the first sub-trajectory, including: first, rearrange the first sub-trajectory according to the motion target to obtain multiple second motion trajectories, which can achieve a one-to-one correspondence between the second motion trajectories and the motion target. At the same time, calculate the highlight scores for the second motion trajectories, and use the trajectories with scores greater than the second score threshold as the target motion trajectories. Then, rearrange the sub-trajectories in the target motion trajectories in chronological order to obtain the third motion trajectory, that is, the final target motion trajectory, thereby improving the readability of the subsequent generated video outline and making the video more logical. Through this step, based on time rearrangement, trajectory rearrangement, and the calculation of scores for multiple motion trajectories, the quality of the final motion trajectory is ensured.

[0164] Step S450: Obtain the target video highlight segments from the final motion trajectory, and then obtain the video outline corresponding to the original video.

[0165] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in rotation with at least a part of other steps or steps or stages in other steps.

[0166] Based on the same inventive concept, the embodiments of the present application also provide a video outline generation device for implementing the above-mentioned video outline generation method. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the video outline generation device provided below can refer to the limitations on the video outline generation method in the above text, and will not be repeated here.

[0167] In one embodiment, as Figure 5 shown, a video outline generation device is provided, including: an acquisition module 51, a calculation module 52, and a generation module 53, where:

[0168] The acquisition module 51 is used to acquire the initial video and determine the video key segments in the initial video;

[0169] A calculation module 52 is configured to extract the audio feature vectors of the initial video, perform emotion feature detection on the audio feature vectors, determine the basic audio segments in the audio feature vectors based on the results of the emotion feature detection, detect the rhythm intensity and short-time audio energy in the basic audio segments, and jointly determine the audio highlight segments according to the rhythm intensity detection results and the short-time audio energy detection results; determine at least one video highlight segment in the initial video according to the video key segments and the audio highlight segments.

[0170] A generation module 53 is configured to screen out target video highlight segments from the video highlight segments based on the motion trajectories of the moving objects in the video highlight segments, and generate the video outline based on the target video highlight segments.

[0171] Each module in the above video outline generation device can be implemented in whole or in part by software, hardware, and their combinations. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0172] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 6 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data related to video outline generation. The network interface of the computer device is used to communicate with external terminals through a network. When the computer program is executed by the processor, it implements a video outline generation method.

[0173] Those skilled in the art can understand that Figure 6 the structure shown in

[0174] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0175] Obtain an initial video and determine the video key segments in the initial video;

[0176] Extract the audio feature vectors of the initial video, perform emotion feature detection on the audio feature vectors, determine the basic audio segments in the audio feature vectors based on the results of the emotion feature detection, detect the rhythm intensity and short-time audio energy in the basic audio segments, and jointly determine the audio highlight segments according to the rhythm intensity detection results and the short-time audio energy detection results; determine at least one video highlight segment in the initial video according to the video key segments and the audio highlight segments;

[0177] Based on the motion trajectories of the moving objects in the video highlight segments, screen out the target video highlight segments from the video highlight segments, and generate a video outline based on the target video highlight segments.

[0178] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0179] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0180] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0181] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A method for generating a video outline, characterized in that, The method includes: Obtaining an initial video and determining video key segments in the initial video; Extracting an audio feature vector of the initial video, performing emotion feature detection on the audio feature vector, determining a basic audio segment in the audio feature vector based on the result of the emotion feature detection, detecting the rhythm intensity and short-time audio energy in the basic audio segment, and jointly determining an audio highlight segment according to the rhythm intensity detection result and the short-time audio energy detection result; determining at least one video highlight segment in the initial video according to the video key segment and the audio highlight segment; Based on the motion trajectory of the moving object in the video highlight segment, screening out a target video highlight segment from the video highlight segment, and generating the video outline based on the target video highlight segment.

2. The method according to claim 1, characterized in that, The detecting the rhythm intensity and short-time audio energy in the basic audio segment, and jointly determining the audio highlight segment according to the rhythm intensity detection result and the short-time audio energy detection result includes: Performing short-time audio energy detection on the basic audio segment, calculating the root mean square of the short-time audio energy detection result, and determining the volume of each frame of audio in the basic audio segment based on the root mean square calculation result; Determining the rhythm intensity of each frame of audio according to the difference between the spectral amplitude of each frame of audio in the basic audio segment and the spectral amplitude of the historical frame; Performing fusion weighted calculation on the volume of each frame of audio and the rhythm intensity of each frame of audio to obtain the highlight score of each frame of audio in the basic audio segment; Based on all audio frames with highlight scores greater than a preset score threshold, obtaining the audio highlight segment.

3. The method according to claim 1, wherein The determining the video key segments in the initial video includes: Obtaining video key frames in the initial video, and dividing the video material in the initial video into multiple video segments based on the video key frames; Traversing all the video segments, and determining the video key segments based on the video feature vectors of the respective video segments.

4. The method according to claim 1, characterized in that, The determining at least one video highlight segment in the initial video according to the video key segment and the audio highlight segment includes: In the time dimension, extracting an initial video highlight segment from the initial video based on the time stamp corresponding to the video key segment and the time stamp corresponding to the audio highlight segment; Performing fusion processing on the image feature vector and the audio feature vector corresponding to the initial video highlight segment to obtain a fusion feature vector of the initial video highlight segment; Obtaining a prompt word corresponding to the video outline, and extracting a text feature vector of the prompt word; Performing similarity matching between the text feature vector and the fusion feature vector to obtain the similarity corresponding to each fusion feature vector, and determining the video highlight segment based on the fusion feature vectors with similarities greater than a preset similarity threshold.

5. The method according to claim 4, characterized in that The extracting an initial video highlight segment from the initial video based on the time stamp corresponding to the video key segment and the time stamp corresponding to the audio highlight segment includes: Restore the timestamps corresponding to the audio highlight segments to the timeline of the initial video to obtain an audio highlight index corresponding to each audio highlight segment; Restore the timestamps corresponding to the video key segments to the timeline of the initial video to obtain a video key frame index corresponding to each video key segment; Based on the audio highlight index and the video key frame index, obtain at least one initial video highlight segment in the initial video.

6. The method according to claim 1, characterized in that The screening of the target video highlight segment from the video highlight segment based on the motion trajectory of the moving target in the video highlight segment includes: Extract the motion trajectory of the moving target in each video highlight segment; Perform anomaly detection on each motion trajectory, and determine the first motion trajectory in the motion trajectory based on the anomaly detection result; wherein, the first motion trajectory includes a plurality of first sub-trajectories; Detect the moving target in the first motion trajectory, and rearrange the first sub-trajectories included in the first motion trajectory based on the moving target to obtain at least one second motion trajectory, wherein each second motion trajectory corresponds to one moving target; Calculate the score of the second motion trajectory based on the dynamic programming algorithm, and determine the motion trajectory with a score greater than the preset second score threshold as the target motion trajectory; wherein, based on the video highlight score, audio highlight score and motion trajectory highlight score of the second motion trajectory, complete the score calculation of the second motion trajectory; Based on the target motion trajectory, determine the corresponding target video highlight segment in the initial video.

7. The method according to claim 6, characterized in that, The determining the corresponding target video highlight segment in the initial video based on the target motion trajectory includes: Determine the target sub-motion trajectory in the target motion trajectory, and rearrange the target sub-motion trajectory in chronological order to obtain a third motion trajectory corresponding one-to-one to the target sub-motion trajectory; Based on the third motion trajectory, obtain the target video highlight segment corresponding to the initial video.

8. A video outline generation device, characterized in that, The device includes: An acquisition module, configured to acquire an initial video and determine video key segments in the initial video; A calculation module, configured to extract an audio feature vector of the initial video, perform emotion feature detection on the audio feature vector, determine a basic audio segment in the audio feature vector based on the result of the emotion feature detection, detect the rhythm intensity and audio short-time energy in the basic audio segment, and jointly determine an audio highlight segment according to the rhythm intensity detection result and the audio short-time energy detection result; determine at least one video highlight segment in the initial video according to the video key segment and the audio highlight segment; A generation module, configured to screen out a target video highlight segment from the video highlight segment based on the motion trajectory of the moving target in the video highlight segment, and generate the video outline based on the target video highlight segment.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Surveillance video abstraction extraction method based on moving object detection

    CN104331905A

  • Video data processing method and electronic equipment

    CN112765399A

  • Video splitting method and device, computer equipment and storage medium

    CN115359409A

  • Method, system and storage medium for generating highlight moment video from video and text input

    CN116170651A

  • Transition point determination method and device, equipment and storage medium

    CN116189708A