Video highlight clip editing method and device, electronic equipment and storage medium
By applying multi-scale convolutional neural network and adaptive attention mechanism in video clips, audio events are detected and aligned to video data, the problem of inability to accurately edit video highlighting content in the prior art is solved, and efficient and accurate video high-scaling clips are achieved.
Patent Information
- Application Number
- CN202510286087.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art cannot accurately edit high-quality video highlighted content clips, especially in multi-camera video data. Direct editing according to audio time cannot ensure accuracy and quality.
By obtaining the audio data of the target video, using a multi-scale convolutional neural network and an adaptive attention mechanism to detect the event category of the audio frame data, determine the target audio event and mark the timestamp, then align the audio event to the video data of each camera, analyze the best camera and edit it, and obtain a highlighted video clip.
It realizes accurate editing of video highlighted content, improves editing efficiency and quality, can adapt to a variety of audio scenes and reduces the deviation between audio event positioning and video content.
Smart Images

Figure CN120050464A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of audio and video editing, and in particular to a method and device for editing a video highlight segment, an electronic device, and a storage medium. Background Art
[0002] In the process of editing and processing audio and video content, especially when it involves long-term live streaming, conference recordings, variety shows and sports events, it is usually necessary to edit out the more important highlight clips.
[0003] Currently, the editing methods of highlight video clips, such as manually identifying and editing or using models to identify the content of each frame of the video, are relatively inefficient and prone to errors. Therefore, a relatively efficient method is to directly identify the audio clips related to the highlight video in the audio data through a classifier, and then edit the video data of the same time range in the video data according to the time range of the audio clip, and finally splice the edited clips to obtain the highlight video clips.
[0004] However, recognition only through classifiers cannot adapt to various audio situations and cannot guarantee the accuracy of the results. In addition, there will be deviations between the positioning of audio events and video content, especially when the picture changes frequently, the deviation is more obvious. In addition, there will be multiple camera positions in the video, and the content shot by different camera positions is different. Direct editing will directly edit the data of one camera position, so the existing method directly edits according to the audio time, which cannot guarantee the effective and accurate editing of high-quality video highlight content segments. Summary of the invention
[0005] Based on the above-mentioned deficiencies of the prior art, the present application provides a method and device for editing video highlight clips, an electronic device, and a storage medium to solve the problem that the prior art cannot accurately edit highlight content clips of high-quality videos.
[0006] In order to achieve the above objectives, this application provides the following technical solutions:
[0007] The first aspect of the present application provides a method for editing a video highlight segment, comprising:
[0008] Get the audio data of the target video;
[0009] Detecting the event category to which each frame of audio data of the target video belongs through a multi-scale convolutional neural network and an adaptive attention mechanism respectively;
[0010] Determine the target audio event of the target video based on the event category to which each frame of the audio frame data belongs, and mark the time stamp of the target audio event; wherein, the target audio event refers to an event of audio reflecting the highlight content of the video.
[0011] Based on the time stamp of the target audio event, align the target audio event to the video data of each camera position of the target video.
[0012] Based on the audio features of the audio data of each camera position of the target video and the video features of the video data of each camera position, analyze the best camera position for each period of time within the range of the time stamp of the target audio event.
[0013] Clip the audio data and video data of the best camera position for each period of time within the range of the time stamp of the target audio event respectively to obtain the highlight video segment corresponding to the target audio event.
[0014] Optionally, in the above video highlight segment clipping method, before respectively detecting the event category to which each frame of the audio frame data of the audio data of the target video belongs through a multi-scale convolutional neural network and an adaptive attention mechanism, it further includes:
[0015] Preprocess the audio data of the target video.
[0016] Slide a time window on the preprocessed audio data of the target video with a preset sliding step length to obtain each frame of audio frame data of the audio data of the target video; wherein, the audio data included in one slide of the time window is one frame of audio frame data.
[0017] Respectively use the short-time Fourier transform to extract the frequency domain features of each frame of the audio frame data, and at least extract the Mel frequency cepstral coefficients and time-frequency diagrams of the audio frame data therefrom, and construct a multi-dimensional feature matrix for each frame of the audio frame data.
[0018] Optionally, in the above video highlight segment clipping method, the step of respectively detecting the event category to which each frame of the audio frame data of the audio data of the target video belongs through a multi-scale convolutional neural network and an adaptive attention mechanism includes:
[0019] For each frame of the audio frame data respectively, input the multi-dimensional feature matrix of the audio frame data extracted from the audio frame data into the multi-scale convolutional neural network.
[0020] Perform multi-scale convolution on the multi-dimensional feature matrix of the audio frame data through the multi-scale convolutional neural network to obtain the high-level feature matrix of the audio frame data.
[0021] Optimize the high-level feature matrix of the audio frame data by using an adaptive attention mechanism;
[0022] Based on the optimized high-level feature matrix of the audio frame data, the classifier calculates the probabilities of the audio frame data belonging to each event category, and determines the event category with the highest probability as the event category to which the audio frame data belongs.
[0023] Optionally, in the above video highlight clip method, determining the target audio event of the target video based on the event category to which each frame of the audio frame data belongs, and marking the time stamp of the target audio event includes:
[0024] Analyze whether there is a target event category based on the probabilities of consecutive N frames of the audio frame data belonging to the event categories to which they belong; wherein, the target event category is the category of the target audio event, and the sum of the probabilities of each frame of the audio frame data belonging to the target event category among the consecutive N frames of the audio frame data is greater than a probability threshold;
[0025] If it is analyzed that there is a target event category based on the probabilities of consecutive N frames of the audio frame data belonging to the event categories to which they belong, then it is determined that there is a target audio event of the target video belonging to the target event category among the consecutive N frames of the audio frame data;
[0026] Determine the start time stamp of the target audio event as the time stamp of the first frame of the audio frame data belonging to the target event category among the consecutive N frames of the audio frame data, and determine the end time stamp of the target audio event as the time stamp of the last frame of the audio frame data belonging to the target event category.
[0027] Optionally, in the above video highlight clip method, aligning the target audio event to the video data of each camera position of the target video based on the time stamp of the target audio event includes:
[0028] Find the key frame with the smallest time difference from the time stamp of the target audio event in the video data of any camera position of the target video;
[0029] Align the target audio event with the video data of the camera position according to the time difference between the time stamp of the target audio event and the time stamp of the found key frame;
[0030] Align the video data of each camera position of the target video by using the dynamic time warping algorithm.
[0031] Optionally, in the above video highlight clip method, analyzing the best camera positions for each period of time within the time stamp range of the target audio event based on the audio features of the audio data of each camera position of the target video and the video features of the video data of each camera position includes:
[0032] Subtracting the first adjustment time from and adding the second adjustment time to the time stamp of the target audio event to obtain the minimum value and the maximum value of the time range corresponding to the target audio event;
[0033] For each camera position of the target video, weighting the volume intensity and the source direction of the audio data of the camera position for each unit time within the time range corresponding to the target audio event to obtain the audio feature score for each unit time corresponding to the camera position;
[0034] Weighting the amplitude of the human action and the emotional reaction information of the video data of the camera position for each unit time within the time range corresponding to the target audio event to obtain the video feature score for each unit time corresponding to the camera position;
[0035] Adding the audio feature score and the video feature score for each unit time corresponding to the camera position to obtain the fused feature score for each unit time corresponding to the camera position;
[0036] Selecting the camera position with the highest fused feature score for each unit time within the time range corresponding to the target audio event as the best camera position for each unit time.
[0037] Optionally, in the above video highlight clip method, respectively editing the audio data and the video data of the best camera positions for each period of time within the time stamp range of the target audio event to obtain the highlight video segment corresponding to the target audio event includes:
[0038] Polling each unit time within the time range corresponding to the target audio event in sequence;
[0039] If the best camera position for the currently polled unit time is different from the current editing camera position and the time difference from the time point when switching to the current editing camera position is greater than the preset time difference, then switching the current editing camera position to the best camera position for the currently polled unit time;
[0040] Editing the audio data and the video data of the current editing camera position until all unit times within the time range corresponding to the target audio event are polled.
[0041] The second aspect of the present application provides a video highlight clip device, including:
[0042] A data acquisition unit for acquiring audio data of a target video;
[0043] A category detection unit for detecting the event category to which each frame of audio frame data of the audio data of the target video belongs respectively through a multi-scale convolutional neural network and an adaptive attention mechanism;
[0044] A time marking unit for determining a target audio event of the target video based on the event category to which each frame of the audio frame data belongs, and marking the timestamp of the target audio event; wherein, the target audio event refers to an event of audio reflecting the highlight content of the video;
[0045] An alignment unit for aligning the target audio event to the video data of each camera position of the target video based on the timestamp of the target audio event;
[0046] A camera position analysis unit for analyzing the best camera position for each period of time within the range of the timestamp of the target audio event based on the audio features of the audio data of each camera position of the target video and the video features of the video data of each camera position;
[0047] An editing unit for editing the audio data and video data of the best camera position for each period of time within the range of the timestamp of the target audio event respectively to obtain a highlight video segment corresponding to the target audio event.
[0048] Optionally, in the above video highlight segment editing device, it further includes:
[0049] A preprocessing unit for preprocessing the audio data of the target video;
[0050] A sliding unit for sliding a time window on the preprocessed audio data of the target video with a preset sliding step length to obtain each frame of audio frame data of the audio data of the target video; wherein, the audio data included in one sliding of the time window is one frame of audio frame data;
[0051] A feature extraction unit for respectively using the short-time Fourier transform to extract the frequency domain features of each frame of the audio frame data, and at least extracting the Mel frequency cepstral coefficients and the time-frequency diagram of the audio frame data therefrom, and constructing a multi-dimensional feature matrix for each frame of the audio frame data.
[0052] Optionally, in the above video highlight segment editing device, the category detection unit includes:
[0053] An input unit for respectively inputting the multi-dimensional feature matrix of each frame of the audio frame data extracted from the audio frame data into the multi-scale convolutional neural network for each frame of the audio frame data;
[0054] A convolutional unit, configured to perform multi-scale convolution on the multi-dimensional feature matrix of the audio frame data through the multi-scale convolutional neural network to obtain the high-level feature matrix of the audio frame data;
[0055] An optimization unit, configured to optimize the high-level feature matrix of the audio frame data by using an adaptive attention mechanism;
[0056] A category determination unit, configured to calculate the probabilities of the audio frame data belonging to each event category based on the optimized high-level feature matrix of the audio frame data through a classifier, and determine the event category with the highest probability as the event category to which the audio frame data belongs.
[0057] Optionally, in the above video highlight clip device, the time marking unit includes:
[0058] An event judgment unit, configured to analyze whether there is a target event category based on the probabilities of continuous N frames of the audio frame data belonging to the event category to which it belongs; wherein, the target event category is the category of the target audio event, and the sum of the probabilities of each frame of the audio frame data belonging to the target event category among the continuous N frames of the audio frame data is greater than a probability threshold;
[0059] An event determination unit, configured to determine that there is a target audio event of the target video belonging to the target event category among the continuous N frames of the audio frame data when it is analyzed based on the probabilities of the continuous N frames of the audio frame data belonging to the event category to which it belongs that there is a target event category;
[0060] A time determination unit, configured to determine the timestamp of the first frame of the audio frame data belonging to the target event category among the continuous N frames of the audio frame data as the start timestamp of the target audio event, and determine the timestamp of the last frame of the audio frame data belonging to the target event category as the end timestamp of the target audio event.
[0061] Optionally, in the above video highlight clip device, the alignment unit includes:
[0062] A key frame search unit, configured to search for the key frame with the smallest time difference from the video data of any camera position of the target video to the timestamp of the target audio event;
[0063] A video data alignment unit, configured to align the target audio event with the video data of the camera position according to the time difference between the timestamp of the target audio event and the timestamp of the searched key frame;
[0064] A camera position data alignment unit, configured to align the video data of each camera position of the target video by using a dynamic time warping algorithm.
[0065] Optionally, in the above video highlight clip device, the camera position analysis unit includes:
[0066] A time adjustment unit, configured to subtract a first adjustment time and add a second adjustment time to the time stamp of the target audio event to obtain the minimum value and the maximum value of the time range corresponding to the target audio event;
[0067] An audio scoring unit, configured to weight the volume intensity and the source direction of the audio data of each camera position of the target video for each unit time within the time range corresponding to the target audio event, to obtain the audio feature score corresponding to each unit time of each camera position;
[0068] A video scoring unit, configured to weight the amplitude of the human action and the emotional reaction information of the video data of each camera position of the target video for each unit time within the time range corresponding to the target audio event, to obtain the video feature score corresponding to each unit time of each camera position;
[0069] A scoring fusion unit, configured to add the audio feature score and the video feature score corresponding to each unit time of each camera position to obtain the fusion feature score corresponding to each unit time of each camera position;
[0070] A camera position selection unit, configured to select the camera position with the highest fusion feature score corresponding to each unit time within the time range corresponding to the target audio event as the best camera position for each unit time.
[0071] Optionally, in the above video highlight clip device, the clip unit includes:
[0072] A polling unit, configured to sequentially poll each unit time within the time range corresponding to the target audio event;
[0073] A camera position switching unit, configured to switch the current clip camera position to the best camera position of the currently polled unit time when the best camera position of the currently polled unit time is different from the current clip camera position and the time difference from the time point when switching to the current clip camera position is greater than a preset time difference;
[0074] A data clip unit, configured to clip the audio data and the video data of the current clip camera position until all unit times within the time range corresponding to the target audio event are polled.
[0075] The third aspect of the present application provides an electronic device, including:
[0076] A memory and a processor;
[0077] wherein, the memory is used for storing a program;
[0078] The processor is used for executing the program, and when the program is executed, it is specifically used for implementing the video highlight clip method described in any one of the above.
[0079] The fourth aspect of the present application provides a computer storage medium for storing a computer program, and when the computer program is executed by a processor, it is used for implementing the video highlight clip method described in any one of the above.
[0080] A video highlight clip method provided by the present application first obtains the audio data of a target video. Respectively through a multi-scale convolutional neural network and an adaptive attention mechanism, the event category to which each frame of audio frame data of the audio data of the target video belongs is detected, so that through the multi-scale convolutional neural network and the adaptive attention mechanism, various complex scenarios can be adapted, and it can be accurately determined whether there is an audio event reflecting the high-frequency content of the video in the audio frame data. Therefore, then based on the event category to which each frame of audio frame data belongs, the target audio event of the target video is determined, and the time stamp of the target audio event is marked. Then specifically based on the time stamp of the target audio event, the target audio event is aligned to the video data of each camera position of the target video, so as to correct the deviation between the audio data and the video data, as well as the deviation of the video data of each camera position, facilitating accurate clipping of the corresponding video content from each camera position. Then based on the audio features of the audio data of each camera position of the target video and the video features of the video data of each camera position, the best camera position for each period of time within the range of the time stamp of the target audio event is analyzed, so that through the integrated analysis of the video data and audio data of each camera position, the camera position of the highlight video content that most accurately captures the highest-quality video in each time period is determined, facilitating subsequent accurate clipping of high-quality highlight video segments. The audio data and video data of the best camera position for each period of time within the range of the time stamp of the target audio event are respectively clipped to obtain the highlight video segment corresponding to the target audio event, thus realizing a method for accurately clipping high-quality highlight video segments in a video. Description of the Drawings
[0081] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0082] Figure 1 Flowchart of a method for video highlight segment editing provided by an embodiment of the present application;
[0083] Figure 2 Flowchart of a method for extracting features of audio frame data provided by an embodiment of the present application;
[0084] Figure 3 Flowchart of a method for detecting the event category to which audio frame data belongs provided by an embodiment of the present application;
[0085] Figure 4 Flowchart of a method for identifying a target audio event and marking its timestamp provided by an embodiment of the present application;
[0086] Figure 5 Flowchart of a method for aligning a target audio event to video data provided by an embodiment of the present application;
[0087] Figure 6 Flowchart of a method for analyzing the best camera position provided by an embodiment of the present application;
[0088] Figure 7 Flowchart of a method for editing audio-visual data of the best camera position provided by an embodiment of the present application;
[0089] Figure 8 Schematic diagram of the architecture of a video highlight segment editing device provided by an embodiment of the present application;
[0090] Figure 9 Schematic diagram of the architecture of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0091] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0092] In this application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0093] An embodiment of this application provides a method for editing video highlight segments, as Figure 1 shown, including the following steps:
[0094] S101. Obtain the audio data of the target video.
[0095] Wherein, the target video refers to the video that needs to be edited for highlight segments.
[0096] Since different situations of video content can be reflected by different audio data of the video, the position of high-frequency video segments can be analyzed through the audio related to the highlighted video content in the audio data of the video, such as analyzing the occurrence of events such as applause, laughter, cheers, etc., so as to determine the high-frequency video content. Therefore, it is necessary to obtain the audio data of the target video to analyze its audio data.
[0097] S102. Detect the event category to which each frame of audio frame data of the audio data of the target video belongs through a multi-scale convolutional neural network and an adaptive attention mechanism respectively.
[0098] In order to analyze the specific situation of the audio data and perform real-time analysis, in the embodiment of this application, the audio data will be divided into multiple frames of audio frame data.
[0099] Optionally, in order to accurately analyze each complete height segment and avoid the analysis results being too discrete or missing, in the embodiments of the present application, the audio frames are not directly divided into multiple segments at a certain time interval, but a sliding window method is adopted to logically divide multiple frames of audio frame data. Specifically, set the size of the sliding window, then start sliding from the starting position of the audio data of the target video, and slide according to the set step size, that is, the time length of each slide is the set step size. Among them, each time it slides, the audio data within the current sliding window is taken as one frame of audio frame data. And the set step size is usually set smaller than the sliding window, so there will be overlapping audio data between each frame of audio frame data, so that more accurate analysis can be carried out.
[0100] In order to analyze the audio data at multiple scales and improve the accuracy of the analysis results, in the embodiments of the present application, multi-scale feature extraction is performed through a multi-scale convolutional neural network, and the weights of different features are adaptively adjusted through an adaptive attention mechanism to adapt to various complex scenarios, so as to ensure the accuracy of the final analysis results.
[0101] When analyzing the audio frame data, it is mainly to analyze whether it contains audio events corresponding to the highlight content, that is, the target audio events of the audio reflecting the highlight content of the video. For example, audio events such as applause, cheering, and singing. Therefore, specifically analyze the event category to which the audio frame data belongs. Optionally, the event category may include non-target audio events and target audio events. And the target audio events can specifically be further divided into multiple categories according to different events.
[0102] In the embodiments of the present application, there is video data of multiple camera positions for the target video. In order to obtain high-quality highlight segments, the video data of each camera position will be analyzed later to clip the video data of the optimal camera position. And there will be corresponding audio data for different camera positions, and the audio data of different camera positions will be different due to factors such as the distance of the acquisition device, the environment, and the device settings.
[0103] So optionally, when analyzing the audio data of the target video, it can specifically be to analyze the audio data of each camera position separately, and then fuse the analysis results of the audio data of each camera position to obtain the final result, for example, using a weighted average or maximum voting strategy for fusion.
[0104] It can also be to select the audio of the camera position with the best sound quality and the clearest sound pickup as the global audio source and analyze its audio data. Or select the audio data collected by a dedicated main microphone, with the highest signal-to-noise ratio, or the audio data with the highest ability level for analysis, and use its analysis result as the final detection result.
[0105] When analyzing the audio frame data, it is mainly through analyzing its features. Therefore, in another embodiment of the present application, before performing step S102, it further includes extracting the features of the audio frame data. As Figure 2 shown, a method for extracting the features of audio frame data provided by an embodiment of the present application includes:
[0106] S201. Preprocess the audio data of the target video.
[0107] In order to improve the quality of the audio data, in the embodiment of the present application, the audio data is first denoised, filtered, and normalized, etc., so as to remove background noise and redundant information, etc.
[0108] S202. Slide a time window on the preprocessed audio data of the target video with a preset sliding step length to obtain each frame of audio frame data of the audio data of the target video.
[0109] Wherein, the audio data included in one sliding of the time window is one frame of audio frame data.
[0110] S203. Respectively use the short-time Fourier transform to extract the frequency domain features of each frame of audio frame data, and at least extract the Mel frequency cepstral coefficients and time-frequency diagrams of the audio frame data from them, and construct a multi-dimensional feature matrix for each frame of audio frame data.
[0111] It should be noted that in the embodiment of the present application, the short-time Fourier transform is mainly used to convert each frame of audio frame data into frequency domain features, and in order to meet the recognition requirements of different types of events, the Mel frequency cepstral coefficients and time-frequency diagrams are further extracted, and finally the extracted features are used to form a multi-dimensional feature matrix. Of course, other features, such as timbre features, etc., can also be further extracted.
[0112] Optionally, in another embodiment of the present application, a specific implementation manner of step S102, as Figure 3 shown, includes the following steps:
[0113] S301. For each frame of audio frame data respectively, input the multi-dimensional feature matrix of the audio frame data extracted from the audio frame data into a multi-scale convolutional neural network.
[0114] S302. Perform multi-scale convolution on the multi-dimensional feature matrix of the audio frame data through the multi-scale convolutional neural network to obtain a high-level feature matrix of the audio frame data.
[0115] Specifically, the multi-scale convolutional neural network can specifically use convolutional kernels of different sizes to perform convolutional operations on the multi-dimensional feature matrix, achieve multi-scale feature extraction, thereby efficiently capturing the time features of short and persistent events, and correct the time sensitivity of the system to audio events. The process of convolution can be specifically expressed as:
[0116]
[0117] Among them, k s is a convolutional kernel of size s; and X represents the multi-dimensional feature matrix of the input audio frame data.
[0118] S303. Optimize the high-level feature matrix of the audio frame data by using the adaptive attention mechanism.
[0119] Specifically, project the high-level feature matrix of the extracted audio frame data into the query, key, and value spaces to obtain the corresponding matrix. Then use this matrix to calculate the attention scores of the features, thereby determining the matrix of the corresponding attention weights, and obtaining the weights of various features. Specifically, the calculation method of calculating the attention scores of the features by the matrix is specifically:
[0120]
[0121] Among them, Q is the query matrix; K is the key matrix; V is the value matrix; d k is the dimension of the key matrix.
[0122] S304. Based on the optimized high-level feature matrix of the audio frame data, the classifier calculates the probabilities that the audio frame data belongs to each event category, and determines the event category with the highest probability as the event category to which the audio frame data belongs.
[0123] Specifically, the classifier calculates the probability that it belongs to an event category through the following formula:
[0124]
[0125] Among them, W i is the vector of weights corresponding to the event category C i ; and W j is the set of weight vectors of all time categories; X is the input feature vector.
[0126] S103. Based on the event categories to which each frame of audio frame data belongs, determine the target audio event of the target video and mark the time stamp of the target audio event.
[0127] Among them, the target audio event refers to the event of the audio reflecting the highlight content of the video.
[0128] It should be noted that the target audio event reflecting the highlighted content of the video lasts for a certain period of time, and the event category to which the audio frame data belongs can reflect whether there is a target audio event. Therefore, each target audio event that appears can be determined according to the event category to which each frame of audio frame data belongs, and then the timestamps of each target audio event are edited, that is, the time range of the target audio event is marked.
[0129] Optionally, in another embodiment of the present application, a specific implementation manner of step S103 is as Figure 4 shown, including the following steps:
[0130] S401. Analyze whether there is a target event category based on the probabilities that consecutive N frames of audio frame data belong to their respective event categories.
[0131] Among them, the target event category is the category of the target audio event, and the sum of the probabilities that each frame of audio frame data belonging to the target event category among consecutive N frames of audio frame data belongs to the target event category is greater than the probability threshold.
[0132] It should be noted that the length of one frame of audio frame data is small, which is convenient for accurately analyzing the event category to which it belongs. However, when there is high-light video content, it lasts for a long time, so a target event that appears will also last for a relatively long time accordingly. Correspondingly, the target event will appear in multiple frames of audio frame data. Therefore, in the embodiment of the present application, consecutive multiple frames of audio frame data are summarized and analyzed. When the sum of the probabilities that each frame of audio frame data belonging to a target event category among consecutive N frames of audio frame data belongs to the target event category is greater than the probability threshold, it is determined that there is a target event category. Among them, the selection of the N frames of audio data for analysis can also be made at a certain step. Therefore, if it is analyzed based on the probabilities that consecutive N frames of audio frame data belong to their respective event categories that there is a target event category, then step S402 is executed.
[0133] S402. Determine that there is a target audio event of a target video belonging to the target event category among consecutive N frames of audio frame data.
[0134] S403. Determine the start timestamp of the target audio event as the timestamp of the first frame of audio frame data belonging to the target event category among consecutive N frames of audio frame data, and determine the end timestamp of the target audio event as the timestamp of the last frame of audio frame data belonging to the target event category.
[0135] It should be noted that since the target audio event is a continuous event, the timestamps of the marked target audio event include its start timestamp and end timestamp.
[0136] Optionally, the timestamp of a frame of audio frame data may specifically be the start time of the frame of audio frame data.
[0137] S104. Align the target audio event to the video data of each camera position of the target video based on the timestamp of the target audio event.
[0138] Considering that there may be differences between the rhythm of the audio event and the rhythm of visual changes, and when the picture changes frequently, the original alignment accuracy between the audio frame and the video frame will decrease, resulting in incorrect event positioning and thus incorrect content in the clip. Therefore, in the embodiments of the present application, the target audio event will be specifically aligned to the video data of the target video based on the timestamp of the target audio event, that is, determine the corresponding time of the target audio event in the video data, so as to facilitate accurate subsequent editing, rather than directly editing according to the timestamp of the target audio event. And in order to facilitate editing from the video data of each camera position, it is necessary to align the target audio event to the video data of each camera position of the target video.
[0139] Optionally, in another embodiment of the present application, a specific implementation manner of step S104 is as Figure 5 shown, and includes the following steps:
[0140] S501. From the video data of any one camera position of the target video, find the key frame with the smallest difference in time from the timestamp of the target audio event.
[0141] It should be noted that the video data includes key frames and predicted frames. A key frame refers to a video frame that contains complete picture information and can be independently decoded. A predicted frame is a video frame that depends on the information of the previous and next frames and needs to refer to the key frame for decoding. Therefore, precise cutting can only be performed at the key frame during editing, so it is specifically necessary to align with the key frame of the video data.
[0142] Specifically, the difference between the timestamp of each frame of video frame data in the video data and the timestamp of the target audio event can be calculated. Specifically, it can be the difference between the start time of the timestamp of the video frame data and the start time of the timestamp of the target audio event, and then the key frame with the smallest difference is selected.
[0143] S502. Align the target audio event with the video data of the camera position according to the difference between the timestamp of the target audio event and the timestamp of the found key frame.
[0144] The difference between the timestamp of the target audio event and the timestamp of the found key frame is the deviation between the two. Therefore, based on this deviation, the timestamp of the target audio event can be aligned to the video data, so that its timestamp changes from the audio timestamp to the corresponding video data timestamp.
[0145] S503. Align the video data of each camera position of the target video using the dynamic time warping algorithm.
[0146] Since there may be certain latency problems in the video data of each camera position, etc., in order to align the target audio event to the video data of one camera position, which can be achieved by aligning with the video data of each camera position, it is necessary to align the video data of each camera position. In the embodiment of the present application, the dynamic time warping algorithm is used for alignment, which can be specifically expressed as:
[0147]
[0148] where T x and T y respectively represent the video frame time series of camera position x and camera position y, and T xi and T yi respectively represent the timestamps of the i-th video frame of camera position x and camera position y.
[0149] S105. Analyze the best camera position for each period within the time range of the timestamp of the target audio event based on the audio features of the audio data of each camera position of the target video and the video features of the video data of each camera position.
[0150] It should be noted that the content and content quality captured by different camera positions at different times are different, and the corresponding audio data is also different. Therefore, in order to obtain accurate and high-quality highlighted content, during the process of editing the video data and audio data within the time range of the timestamp of the target audio event, it is necessary to continuously switch to the content captured by the best camera position for editing. Therefore, it is necessary to first determine the best camera position for each period within the time range of the timestamp of the target audio event.
[0151] Since both the video data and the audio data can reflect the content and quality of the video, etc., in the embodiment of the present application, by analyzing the audio features of the audio data and the video features of the video data of each camera position within each period of the time range of the timestamp of the target audio event, and then based on the audio features and video features, evaluating the scores of each period, and then selecting the best camera position for each period according to the scores.
[0152] Optionally, in another embodiment of the present application, a specific implementation manner of step S105, as Figure 6 shown, includes the following steps:
[0153] S601. Subtract the first adjustment time from the timestamp of the target audio event and add the second adjustment time to obtain the minimum value and the maximum value of the time range corresponding to the target audio event.
[0154] In order to make the clipped video segment have better starting and ending points, that is, a video of a relatively complete scene, rather than suddenly playing from the middle position of a certain piece of content. Also, in order to make the length of the clipped video segment meet the user's preferences, in the embodiments of the present application, the time range of the target audio event will be correspondingly expanded by setting the first adjustment time and the second adjustment time, that is, subtracting the first adjustment time Tp from the time stamp of the target audio event and adding the second adjustment time Tq to obtain the minimum and maximum values of the time range corresponding to the target audio event. Therefore, the range of the time stamp Te of the target audio event is: [Te - Tp, Te + Tq]. Specifically, it can be subtracting the first adjustment time from the minimum value of the time stamps of the target audio event and adding the second adjustment time to the maximum value of the time stamps of the target audio event.
[0155] Therefore, optionally, on the premise of user authorization, the video viewing behavior data of the user can be collected, and then their preferences can be analyzed based on their behavior data, and their real-time feedback can be received. According to the user's preferences and real-time feedback, the first adjustment time and the second adjustment time can be dynamically adjusted, so as to clip out a highlight segment that meets the personalized needs of the user.
[0156] S602. For each camera position of the target video, weight the volume intensity and the source direction of the audio data of each unit time within the time range corresponding to the target audio event to obtain the audio feature score of each unit time corresponding to this camera position.
[0157] It should be noted that since the volume intensity can reflect the distance between the camera position and the shooting scene, and the source direction can reflect the shooting angle of the camera position, both can reflect its importance relative to the current shooting scene and whether it is the best camera position.
[0158] Among them, the unit time is set to a relatively short time to be able to analyze the change of the best camera position in a timely manner.
[0159] Specifically, use the weights corresponding to the volume intensity and the source direction to weight the two to obtain the audio feature score, which can be specifically expressed as:
[0160]
[0161] Among them, A is the volume intensity; D is the source direction; α and β are the weights corresponding to the volume intensity and the source direction. Optionally, the corresponding weights of the two can be continuously learned and optimized through machine learning.
[0162] S603. Weight the amplitude of the human actions and the emotional response information in the video data of this camera position for each unit time within the time range corresponding to the target audio event, to obtain the video feature score for each unit time corresponding to this camera position.
[0163] Similarly, weight the amplitude of the human actions and the emotional response information through corresponding weights to obtain the video feature score, which can be specifically expressed as:
[0164]
[0165] where M is the amplitude of the human actions; E is the emotional response information; and are the weights corresponding to the amplitude of the human actions and the emotional response information.
[0166] S604. Add the audio feature score and the video feature score for each unit time corresponding to the camera position to obtain the fused feature score for each unit time corresponding to the camera position.
[0167] S605. Select the camera position with the highest fused feature score for each unit time within the time range corresponding to the target audio event as the best camera position for each unit time.
[0168] S106. Clip the audio data and video data of the best camera position for each period of time within the time range of the time stamps of the target audio event respectively, to obtain the highlighted video segment corresponding to the target audio event.
[0169] Optionally, after determining the best camera position for each time period, during the clipping process, continuously switch to the audio data and video data of the best camera position for clipping, so that a highlighted video segment corresponding to the target audio event composed of the data of the best camera position for each time period can be obtained.
[0170] It should be noted that considering that there may be multiple target audio events in the target video. Therefore, optionally, if adjacent target audio events overlap or the time interval is less than a certain threshold, they can be merged into one target audio event and then processed. And during the clipping process, for the data without target audio events, Bezier curves can be used for smooth transition.
[0171] Optionally, in order to avoid overly frequent switching of camera positions, thus affecting the user's viewing experience, so in another embodiment of the present application, a specific implementation manner of step S106 is as Figure 7 shown, including the following steps:
[0172] S701. Poll each unit time within the time range corresponding to the target audio event in sequence.
[0173] S702. Determine whether the best camera position for the current polling unit time is different from the current editing camera position, and whether the time difference from the time point when switching to the current editing camera position is greater than the preset time difference.
[0174] That is to say, after determining the best camera position for the current polling unit time, it is judged whether it is consistent with the current editing camera position being used, that is, whether the best camera position has changed. If they are consistent, the current editing camera position is not changed, and the editing continues according to the current editing camera position, that is, step S704 is executed. If it is judged that it is inconsistent with the current editing camera position being used, then it is judged whether the time difference from the time point when switching to the current editing camera position is greater than the preset time difference, that is, it is judged whether the time since the last camera position switch is greater than the preset time difference. If it is greater than the preset time difference, it is selected not to switch the camera position, and the editing continues according to the current editing camera position, that is, step S704 is executed. If it is greater than the preset time difference, it is necessary to first execute step S703 to use the best camera position for the current polling unit time as the latest current editing camera position.
[0175] S703. Switch the current editing camera position to the best camera position for the current polling unit time.
[0176] S704. Edit the audio data and video data of the current editing camera position until all unit times within the time range corresponding to the target audio event are polled.
[0177] Optionally, in order to facilitate subsequent search and management, etc., in another embodiment of the present application, after executing step S106, it may further include:
[0178] Mark the labels of the highlighted segments of each target video based on the event category of each target audio event and the highlighted video segment corresponding to each target audio event respectively.
[0179] After marking the labels, the user can retrieve according to the labels, and personalized push can be performed for the user according to the labels.
[0180] Optionally, in another embodiment of the present application, after obtaining the highlighted video segments corresponding to each target audio event, the highlighted video segments corresponding to each target audio event can be further spliced to obtain a combined video of a high-frequency video segment of the target video.
[0181] An embodiment of the present application provides a method for editing video highlight segments. First, the audio data of the target video is obtained. By using a multi-scale convolutional neural network and an adaptive attention mechanism respectively, the event category to which each frame of audio frame data in the audio data of the target video belongs is detected, so that it can be determined whether there is an audio event reflecting the high-frequency content of the video in the audio frame data. Therefore, then based on the event category to which each frame of audio frame data belongs, the target audio event of the target video is determined, and the time stamp of the target audio event is marked. Then, specifically based on the time stamp of the target audio event, the target audio event is aligned to the video data of each camera position of the target video, so as to correct the deviation between the audio data and the video data, as well as the deviation of the video data of each camera position, facilitating accurate editing of the corresponding video content from each camera position. Then, based on the audio features of the audio data of each camera position of the target video and the video features of the video data of each camera position, the best camera position for each period within the time range of the time stamp of the target audio event is analyzed, so that by integrating and analyzing the video data and audio data of each camera position, the camera position of the highlight video content that most accurately captures the highest-quality video in each time period is determined, facilitating subsequent accurate editing of high-quality highlight video segments. The audio data and video data of the best camera position for each period within the time range of the time stamp of the target audio event are respectively edited to obtain the highlight video segment corresponding to the target audio event, thus realizing a method for accurately editing high-quality highlight video segments in a video.
[0182] Another embodiment of the present application provides a video highlight segment editing device, as Figure 8 shown, including:
[0183] A data acquisition unit 801, configured to acquire the audio data of the target video.
[0184] A category detection unit 802, configured to detect the event category to which each frame of audio frame data in the audio data of the target video belongs by using a multi-scale convolutional neural network and an adaptive attention mechanism respectively.
[0185] A time marking unit 803, configured to determine the target audio event of the target video based on the event category to which each frame of audio frame data belongs, and mark the time stamp of the target audio event. Wherein, the target audio event refers to an event of audio reflecting the highlight content of the video.
[0186] An alignment unit 804, configured to align the target audio event to the video data of each camera position of the target video based on the time stamp of the target audio event.
[0187] The camera position analysis unit 805 is configured to analyze the best camera positions for each period of time within the time stamp range of the target audio event based on the audio features of the audio data of each camera position of the target video and the video features of the video data of each camera position.
[0188] The editing unit 806 is configured to edit the audio data and video data of the best camera positions for each period of time within the time stamp range of the target audio event respectively, so as to obtain a highlighted video segment corresponding to the target audio event.
[0189] Optionally, in the video highlighted segment editing device provided in another embodiment of the present application, it further includes:
[0190] The preprocessing unit is configured to preprocess the audio data of the target video.
[0191] The sliding unit is configured to slide a time window on the preprocessed audio data of the target video with a preset sliding step length, so as to obtain each frame of audio frame data of the audio data of the target video. Wherein, the audio data included in one sliding of the time window is one frame of audio frame data.
[0192] The feature extraction unit is configured to respectively extract the frequency domain features of each frame of audio frame data by using the short-time Fourier transform, and at least extract the Mel frequency cepstral coefficients and time-frequency diagrams of the audio frame data therefrom, and construct a multi-dimensional feature matrix of each frame of audio frame data.
[0193] Optionally, in the video highlighted segment editing device provided in another embodiment of the present application, the category detection unit includes:
[0194] The input unit is configured to respectively input the multi-dimensional feature matrix of the audio frame data extracted from the audio frame data into a multi-scale convolutional neural network for each frame of audio frame data.
[0195] The convolutional unit is configured to perform multi-scale convolution on the multi-dimensional feature matrix of the audio frame data through the multi-scale convolutional neural network to obtain a high-level feature matrix of the audio frame data.
[0196] The optimization unit is configured to optimize the high-level feature matrix of the audio frame data by using an adaptive attention mechanism.
[0197] The category determination unit is configured to calculate the probabilities that the audio frame data belongs to each event category based on the optimized high-level feature matrix of the audio frame data through a classifier, and determine the event category with the highest probability as the event category to which the audio frame data belongs.
[0198] Optionally, in the video highlighted segment editing device provided in another embodiment of the present application, the time marking unit includes:
[0199] An event judgment unit for analyzing whether there is a target event category based on the probabilities that consecutive N audio frame data belong to their respective event categories. Wherein, the target event category is the category of the target audio event, and the sum of the probabilities that each frame of the consecutive N audio frame data belonging to the target event category belongs to the target event category is greater than a probability threshold.
[0200] An event determination unit for determining that there is a target audio event of a target video belonging to the target event category in the consecutive N audio frame data when it is analyzed that there is a target event category based on the probabilities that the consecutive N audio frame data belong to their respective event categories.
[0201] A time determination unit for determining the timestamp of the first audio frame data belonging to the target event category in the consecutive N audio frame data as the start timestamp of the target audio event, and determining the timestamp of the last audio frame data belonging to the target event category as the end timestamp of the target audio event.
[0202] Optionally, in the video highlight segment editing device provided in another embodiment of the present application, the alignment unit includes:
[0203] A key frame search unit for searching for the key frame with the smallest time difference from the timestamps of the target audio event in the video data of any camera position of the target video.
[0204] A video data alignment unit for aligning the target audio event with the video data of the camera position according to the time difference between the timestamp of the target audio event and the timestamp of the searched key frame.
[0205] A camera position data alignment unit for aligning the video data of each camera position of the target video by using the dynamic time warping algorithm.
[0206] Optionally, in the video highlight segment editing device provided in another embodiment of the present application, the camera position analysis unit includes:
[0207] A time adjustment unit for subtracting the first adjustment time and adding the second adjustment time to the timestamp of the target audio event to obtain the minimum value and the maximum value of the time range corresponding to the target audio event.
[0208] An audio scoring unit for respectively weighting the volume intensity and the source direction of the audio data of each camera position per unit time within the time range corresponding to the target audio event for each camera position of the target video to obtain the audio feature score per unit time corresponding to the camera position.
[0209] A video scoring unit, configured to weight the amplitude of human actions and the emotional response information in the video data of each camera position within each unit time within the time range corresponding to the target audio event, so as to obtain the video feature score corresponding to each unit time of the camera position.
[0210] A scoring fusion unit, configured to add the audio feature score and the video feature score corresponding to each unit time of the camera position to obtain the fusion feature score corresponding to each unit time of the camera position.
[0211] A camera position selection unit, configured to select the camera position with the highest fusion feature score corresponding to each unit time within the time range corresponding to the target audio event as the best camera position for each unit time.
[0212] Optionally, in the video highlight clip device provided in another embodiment of the present application, the clip unit includes:
[0213] A polling unit, configured to sequentially poll each unit time within the time range corresponding to the target audio event.
[0214] A camera position switching unit, configured to switch the current clip camera position to the best camera position of the currently polled unit time when the best camera position of the currently polled unit time is different from the current clip camera position and the time difference from the time point when switching to the current clip camera position is greater than a preset time difference.
[0215] A data clip unit, configured to clip the audio data and video data of the current clip camera position until all unit times within the time range corresponding to the target audio event are polled.
[0216] It should be noted that for the specific working processes of the various units provided in the above embodiments of the present application, reference may specifically be made to the specific implementation manners provided in the above method embodiments, and details are not elaborated herein.
[0217] Another embodiment of the present application provides an electronic device, as Figure 9 shown, including:
[0218] A memory 901 and a processor 902.
[0219] Among them, the memory 901 is used to store programs.
[0220] The processor 902 is configured to execute the programs stored in the memory 901, and when the programs are executed, it is specifically configured to implement the video highlight clip method provided in any one of the above embodiments.
[0221] Another embodiment of the present application provides a computer storage medium, configured to store a computer program, and when the computer program is executed by a processor, it is used to implement the video highlight clip method provided in any one of the above embodiments.
[0222] Computer storage media includes both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0223] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0224] The above description of the disclosed embodiments enables those skilled in the art to implement or use this application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A video highlight clip editing method, characterized in that: include: Get the audio data of the target video; Detecting the event category to which each frame of audio data of the target video belongs through a multi-scale convolutional neural network and an adaptive attention mechanism respectively; Based on the event category to which the audio frame data of each frame belongs, a target audio event of the target video is determined, and a timestamp of the target audio event is marked; wherein the target audio event refers to an event of the audio reflecting the highlighted content of the video; Based on the timestamp of the target audio event, align the target audio event to the video data of each camera position of the target video; Analyzing the best camera positions for each time period within the time stamp of the target audio event based on the audio features of the audio data of each camera position of the target video and the video features of the video data of each camera position; The audio data and video data of the best camera position of each time period within the time stamp range of the target audio event are respectively edited to obtain a highlighted video clip corresponding to the target audio event.
2. The method according to claim 1, characterized in that Before detecting the event category to which each frame of audio data of the target video belongs by using a multi-scale convolutional neural network and an adaptive attention mechanism, the method further includes: Preprocessing the audio data of the target video; Sliding the time window on the preprocessed audio data of the target video by a preset sliding step length to obtain audio frame data of each frame of the audio data of the target video; wherein the audio data included in one sliding of the time window is one frame of audio frame data; The frequency domain features of the audio frame data of each frame are extracted by short-time Fourier transform, and at least the Mel-frequency cepstral coefficients and the time-frequency diagram of the audio frame data are extracted therefrom to construct a multi-dimensional feature matrix of the audio frame data of each frame.
3. The method according to claim 1, characterized in that The event category to which each frame of audio data of the target video belongs is detected by using a multi-scale convolutional neural network and an adaptive attention mechanism, including: For each frame of the audio frame data, respectively, inputting the multi-dimensional feature matrix of the audio frame data extracted from the audio frame data into the multi-scale convolutional neural network; Performing multi-scale convolution on the multi-dimensional feature matrix of the audio frame data through the multi-scale convolutional neural network to obtain a high-level feature matrix of the audio frame data; Optimizing the high-level feature matrix of the audio frame data using an adaptive attention mechanism; The classifier calculates the probability that the audio frame data belongs to each event category based on the optimized high-level feature matrix of the audio frame data, and determines the event category with the highest probability as the event category to which the audio frame data belongs.
4. The method according to claim 1, characterized in that The step of determining a target audio event of the target video based on the event category to which the audio frame data of each frame belongs, and marking the timestamp of the target audio event includes: Based on the probability that the audio frame data of N consecutive frames belong to the event category to which they belong, analyzing whether there is a target event category; wherein the target event category is the category of the target audio event, and the sum of the probabilities that the audio frame data of each frame belonging to the target event category in the N consecutive frames of the audio frame data belongs to the target event category is greater than a probability threshold; If the target event category exists based on the probability that the continuous N frames of the audio frame data belong to the event category to which they belong, it is determined that there is a target audio event of the target video belonging to the target event category in the continuous N frames of the audio frame data; The timestamp of the first frame of the audio frame data belonging to the target event category among the consecutive N frames of audio frame data is determined as the start timestamp of the target audio event, and the timestamp of the last frame of the audio frame data belonging to the target event category is determined as the end timestamp of the target audio event.
5. The method according to claim 1, characterized in that The step of aligning the target audio event to the video data of each camera position of the target video based on the timestamp of the target audio event includes: Finding a key frame having the smallest time difference with the timestamp of the target audio event from the video data of any camera position of the target video; Aligning the target audio event with the video data of the camera position according to the difference between the timestamp of the target audio event and the timestamp of the found key frame; The video data of each camera position of the target video are aligned using a dynamic time warping algorithm.
6. The method according to claim 1, characterized in that The step of analyzing the best camera positions for each time period within the time stamp range of the target audio event based on the audio features of the audio data of each camera position of the target video and the video features of the video data of each camera position includes: Subtract the first adjustment time from the timestamp of the target audio event and add the second adjustment time to obtain the minimum value and the maximum value of the time range corresponding to the target audio event; For each camera position of the target video, weight the volume intensity and source direction of the audio data of the camera position at each unit time within the time range corresponding to the target audio event to obtain an audio feature score corresponding to each unit time of the camera position; Weighting the character action amplitude and emotional response information of the video data of the camera position at each unit time within the time range corresponding to the target audio event to obtain a video feature score for each unit time corresponding to the camera position; Adding the audio feature score and the video feature score of each unit time corresponding to the camera position to obtain a fusion feature score of each unit time corresponding to the camera position; The camera position with the highest fusion feature score corresponding to each unit time within the time range corresponding to the target audio event is selected as the best camera position for each unit time.
7. The method according to claim 6, characterized in that The step of editing the audio data and video data of the best camera position at each time period within the time stamp range of the target audio event to obtain a highlighted video clip corresponding to the target audio event includes: Polling each unit time in the time range corresponding to the target audio event in sequence; If the currently polled best camera position for the unit time is different from the current editing camera position, and the time difference from the time point of switching to the current editing camera position is greater than the preset time difference, the current editing camera position is switched to the currently polled best camera position for the unit time; The audio data and the video data of the current editing position are edited until all the unit times within the time range corresponding to the target audio event are polled.
8. A video highlight segment editing device, characterized in that: include: A data acquisition unit, used to acquire audio data of a target video; A category detection unit, used to detect the event category to which each frame of audio data of the target video belongs through a multi-scale convolutional neural network and an adaptive attention mechanism; A time marking unit, used to determine the target audio event of the target video based on the event category to which the audio frame data of each frame belongs, and mark the timestamp of the target audio event; wherein the target audio event refers to an event of the audio reflecting the highlighted content of the video; an alignment unit, configured to align the target audio event to the video data of each camera position of the target video based on the timestamp of the target audio event; A camera position analysis unit, configured to analyze the best camera position for each time period within the range of the timestamp of the target audio event based on the audio features of the audio data of each camera position of the target video and the video features of the video data of each camera position; The editing unit is used to edit the audio data and video data of the best camera position in each time period within the time stamp range of the target audio event to obtain a highlight video clip corresponding to the target audio event.
9. An electronic device, characterized in that: include: Memory and processor; Wherein, the memory is used to store programs; The processor is used to execute the program, and when the program is executed, it is specifically used to implement the video highlight segment editing method as described in any one of claims 1 to 7.
10. A computer storage medium, characterized in that: Used to store a computer program, which, when executed by a processor, is used to implement the video highlight segment editing method as described in any one of claims 1 to 7.