Dance video special effect adding method and device, computer equipment and storage medium

By automatically identifying music beats and action special effects nodes in dance videos and efficiently adding screen special effects, the problem of time-consuming addition of video special effects in the existing technology is solved, and automated editing and dynamic effects are improved.

CN120017774APending Publication Date: 2025-05-16ARASHI VISION INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311537815.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-16
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the prior art, the addition of video special effects requires users to manually adjust and repeatedly adjust, which takes a long time. Especially for dance videos with music, ordinary users have low editing efficiency.

Method used

By automatically determining the beat effect nodes and action effect nodes based on the music beat and action effect nodes in the dance video, the beat and action information are used to efficiently add screen effects to generate special effects dance videos.

Benefits of technology

It realizes efficient and automatic addition of dance video special effects, reduces the time and complexity of user manual operation, and improves the dynamic effect and viewing of the video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017774A_ABST
    Figure CN120017774A_ABST
Patent Text Reader

Abstract

The invention relates to a dance video special effect adding method and device, computer equipment and a storage medium. The method comprises the following steps: determining beat special effect nodes according to music beats in a dance video; determining an action special effect node according to the condition that the action in the dance video accords with a special effect adding type; and according to at least one special effect node in the beat special effect node and the action special effect node, adding a picture special effect to the dance video to obtain a special effect dance video. By adopting the method, the dance video special effect adding efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a method, device, computer equipment, storage medium and computer program product for adding special effects to a dance video. Background Art

[0002] Existing video effects are generally created by users pre-shooting videos and then adding special effects through post-processing software. Taking camera effects as an example, it is necessary to add effects such as zooming in, zooming out, left and right movement / shaking / jittering, up and down movement / shaking / jittering, and image rotation through post-editing to simulate camera movements during filming.

[0003] In the prior art, the use of post-processing software (Adobe Premiere Pro, Final Cut Pro, etc.) to add special effects requires users to manually mark and repeatedly adjust, and the entire editing process is time-consuming and requires a certain level of professionalism and operating equipment from the user. Especially for dance videos with music, the requirements for users are higher, and it takes a long time for ordinary users to edit videos.

[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present application, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention

[0005] Based on this, it is necessary to provide a method, device, computer equipment, computer-readable storage medium and computer program product that can more efficiently add dance video special effects to the above-mentioned technical problems.

[0006] In a first aspect, the present application provides a method for adding special effects to a dance video, comprising:

[0007] Determine the beat special effect node according to the music beat in the dance video;

[0008] Determining an action special effect node according to whether an action in the dance video meets the type of special effect addition;

[0009] According to at least one of the beat special effect nodes and the action special effect nodes, a screen special effect is added to the dance video to obtain a special effect dance video.

[0010] In one embodiment, determining the beat special effect node according to the music beat in the dance video includes:

[0011] Generate music amplitude based on audio from dance videos;

[0012] The music amplitude is subjected to feature extraction according to a scale that increases step by step, so as to obtain the amplitude features of the scale at each level; the input features for feature extraction of the amplitude features of the scale at each level are obtained by downsampling the output features of the previous level according to a preset step size;

[0013] Determining the strong beat position for characterizing the accented beat according to the amplitude characteristics of the scales at each level;

[0014] A beat special effect node is determined according to the strong beat position.

[0015] In one embodiment, determining the action special effect node according to whether the action in the dance video meets the type of special effect addition includes:

[0016] Detecting a target person based on multiple frames of the dance video;

[0017] Detecting the human body posture of the target person in the multiple frames to obtain a human body key point sequence;

[0018] Identify the action of the target person in each frame of the picture according to the human body key point sequence to obtain the action type;

[0019] If the action type meets the preset dance special effect action type, the action special effect node is determined according to the time point corresponding to the action.

[0020] In one embodiment, adding a screen special effect to the dance video according to at least one of the beat special effect node and the action special effect node to obtain a special effect dance video includes:

[0021] Determine the note strength at the time point corresponding to the beat special effect node; the note strength is used to represent the amplitude change at the time point corresponding to the beat special effect node;

[0022] Determine a target beat special effect node among the beat special effect nodes according to the note intensity;

[0023] According to the target beat time point corresponding to the target beat special effect node, a camera movement special effect is added to the dance video to obtain a special effect dance video.

[0024] In one embodiment, adding a screen special effect to the dance video according to at least one of the beat special effect node and the action special effect node to obtain a special effect dance video includes:

[0025] Determining the camera movement special effect of the action special effect node according to the special effect adding type that the action matches;

[0026] According to the action time point corresponding to the action special effect node, the camera movement special effect of the action special effect node is added to the dance video to obtain a special effect dance video.

[0027] In one embodiment, adding a screen special effect to the dance video according to at least one of the beat special effect node and the action special effect node to obtain a special effect dance video includes:

[0028] Determine the target beat time point according to the beat special effect node;

[0029] Determine the action time point corresponding to the action special effect node;

[0030] Determine the intersection time point of the target beat time point and the action time point;

[0031] According to the intersection time point, a special effect is added to the dance video to obtain a special effect dance video.

[0032] In one embodiment, adding a special effect to the dance video according to the intersection time point to obtain a special effect dance video includes:

[0033] Determine whether the number of the intersection time points meets the intersection special effect condition;

[0034] If yes, then at the intersection time point, adding a special effect to the dance video to obtain a special effect dance video;

[0035] If not, add screen special effects to the dance video according to the high priority nodes among the beat special effect nodes and the action special effect nodes to obtain a special effect dance video.

[0036] In one embodiment, adding special effects to the dance video to obtain a special effects dance video includes:

[0037] When the screen special effect is a camera movement special effect, the pan / tilt movement is controlled according to the camera movement special effect;

[0038] During the movement of the pan / tilt platform, the dance video is captured to obtain a special effects dance video.

[0039] In a second aspect, the present application also provides a dance video special effects adding device, comprising:

[0040] A beat detection module is used to determine the beat special effect node according to the music beat in the dance video;

[0041] An action detection module is used to determine an action special effect node according to whether the action in the dance video meets the type of special effect addition;

[0042] The special effect adding module is used to add screen special effects to the dance video according to at least one of the beat special effect nodes and the action special effect nodes to obtain a special effect dance video.

[0043] In a third aspect, the present application also provides a pan-tilt head, comprising a motor and a processor, wherein the motor is used to control the rotation of the pan-tilt head, and the processor implements the steps of video generation in any of the above embodiments when executing the computer program.

[0044] In a fourth aspect, the present application further provides a computer device, wherein the computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of adding special effects to a dance video in any of the above embodiments are implemented.

[0045] In a fifth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of adding dance video special effects in any of the above embodiments are implemented.

[0046] In a sixth aspect, the present application further provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps of adding dance video special effects in any of the above embodiments are implemented.

[0047] The above-mentioned method, device, computer equipment, storage medium and computer program product for adding special effects to dance videos determine the beat special effects node according to the music beat in the dance video, thereby using the characteristics of the music beat in representing the amplitude strength law to accurately determine the beat special effects node, and can accurately control the process of adding picture special effects from the audio perspective; determine the action special effects node according to the situation that the action in the dance video meets the type of special effects addition, so that the results of human body movement analysis can be displayed and the corresponding dance movements can be accurately identified; finally, according to at least one special effect node among the beat special effects node and the action special effects node, add picture special effects to the dance video to obtain a special effects dance video, and can efficiently and automatically add video picture special effects to the accurately identified specific music nodes and specific dance movements, so that the video has a dynamic effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0049] Figure 1A diagram of an application environment of a method for adding special effects to a dance video in one embodiment;

[0050] Figure 2 A schematic diagram of a flow chart of a method for adding special effects to a dance video in one embodiment;

[0051] Figure 3 A schematic diagram of a visualization sample of the output of the WaveBeat model in one embodiment;

[0052] Figure 4 A schematic diagram of the structure of skeleton information of each frame in an embodiment;

[0053] Figure 5 A schematic diagram of performing action recognition on skeleton information of multiple frames in one embodiment;

[0054] Figure 6 A schematic diagram of the effect of a method for adding special effects to a dance video in one embodiment;

[0055] Figure 7 It is a structural block diagram of a device for adding special effects to a dance video in one embodiment;

[0056] Figure 8 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0057] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0058] The video clip classification method of the motion event provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the terminal 102 can be, but is not limited to, various cameras, video cameras, panoramic cameras, sports cameras, personal computers, laptops, smart phones, tablet computers, pan / tilt bodies and portable wearable devices, and the portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The terminal 102 can be fixed to the pan / tilt body by welding or the like, and can also be detachably connected or rotatably connected to the pan / tilt body.

[0059] In one embodiment, Figure 2 As shown, a method for adding special effects to a dance video is provided, and the method is applied to Figure 1 The terminal 102 in the example is used as an example to illustrate, including the following steps 202 to 206. Among them:

[0060] Step 202, determining a beat special effect node according to the music beat in the dance video.

[0061] A dance video is a video that includes at least one target object's dance movements. A dance video can be used to record a target object in the real world, or can be used to record a target object in a video. A dance video includes a picture of the dance movement, and also includes audio used to match the dance movement. The audio used to match the dance movement has a music beat.

[0062] The music beat is used to characterize the amplitude strength pattern. Optionally, the music beat in the audio is composed of music bars, and a three-beat music beat or a four-beat music beat can be set for the music bar; under the three-beat music beat, the amplitude strength pattern is strong beat-weak beat-weak beat; under the four-beat music beat, the amplitude strength pattern is strong beat-weak beat-second strong beat-weak beat.

[0063] Optionally, the music beat is used to represent the unit duration, and is used to represent the fixed duration of each amplitude. Optionally, the music beat is used to represent that the strong beat at a certain moment is a quarter note, and the weak beat at another moment is an eighth note. Optionally, when the music beat is used to represent the amplitude strength law and the unit duration, the music amplitude at different moments can be directly represented by the music beat, and the music amplitude at different moments can be input into the neural network for processing to obtain the beat special effect node.

[0064] The beat effect node is a candidate node for adding visual effects in a music video. Optionally, the beat effect node can be a time point or a visual mark for indicating a time point. Optionally, each dance video can have multiple beat effect nodes, and a target beat effect node for adding visual effects to the dance video can be selected from the multiple beat effect nodes according to the conditions met by the multiple beat effect nodes.

[0065] In one embodiment, a beat effect node is determined based on the music beat in a dance video, including: generating music amplitude based on the audio in the dance video; performing accent beat detection on the music amplitude to obtain a strong beat position; and determining a beat effect node based on the strong beat position; wherein the accent beat is a position where the amplitude meets the strong beat condition.

[0066] In an optional embodiment, the music amplitude is subjected to accent beat detection, including: performing accent beat detection on the music amplitude through a neural network model for accent beat detection; wherein the neural network model for accent beat detection may be a wave beat detection model (WaveBeat), a recurrent neural network (RNN) for accent beat detection, a transformer model (Transformer) and other models.

[0067] In another embodiment, determining a beat effect node according to the music beat in a dance video includes: searching for an audio score of the dance video according to the music beat in the dance video; determining a strong beat position according to the audio score; and determining a beat effect node according to the strong beat position.

[0068] From the perspective of hearing perception obtained from the physiological structure of the human body, strong beats are usually more bass-oriented, heavier and more full-bodied. However, after research, it was found that there is a certain inherent correspondence between the choreographed dance movements and the beats, and thus the beat special effects nodes of the dance video are determined by the beat of the music. Optionally, explosive / large-amplitude dance movements are choreographed on strong beats, so that the dance movements and music appear to have a regular response and a corresponding relationship, and the effect of integrating audio and picture appears to be more harmonious and more ornamental. Based on the principles of the dance field, it is necessary to choreograph explosive / large-amplitude pictures on strong beats. This field uses strong beats as a factor in adding special effects, which can make the visual and auditory effects consistent and more harmonious in terms of senses.

[0069] In an optional embodiment, the beat effect node is determined according to the music beat in the dance video, including: determining the note intensity envelope of the audio according to the audio music amplitude of the dance video; determining the beat effect node according to the note intensity envelope; wherein the note intensity envelope (Amplitude Envelope) refers to the curve of the amplitude of the note changing with time. Specifically, the amplitude of the music is short-time Fourier transformed, and the logarithmic Mel spectrum is calculated; the first-order time difference of the logarithmic Mel spectrum is calculated, and the mean of the logarithmic spectrum difference degree in each frame is calculated; the mean of the differential amplitude of all frames is normalized and smoothed to obtain the note intensity envelope, and then the beat effect node is determined according to the note intensity envelope. Optionally, all the peaks on each envelope can be used as the note intensity peaks, and the position of the note intensity peak is the strong beat position; the note intensity peaks can be screened out based on all the peaks on each envelope and the screening conditions, and the position of the note intensity peak is used as the strong beat position.

[0070] like Figure 3 As shown in the visualization sample diagram of WaveBeat output, the note intensity envelope is represented by a dotted line, the horizontal axis represents time, the vertical axis represents amplitude, the points with larger radius are the time of strong beats (accents, downbeats), and the points with smaller radius are the time of weak beats (upbeats).

[0071] Step 204, determining an action special effect node according to whether the action in the dance video meets the type of special effect addition.

[0072] The action is the result obtained by recognizing the posture of the target object in the dance video. Some actions belong to the special effect adding type, and some actions belong to the non-special effect adding type. Optionally, the action in the dance video can be recognized according to the posture of the target object in the dance video; the posture of the target object in the dance video can also be detected according to the parameters required by the template of the special effect adding type, so as to specifically determine whether the action in the dance video meets the special effect adding type.

[0073] The special effect adding type is a preset action type, which is an indicator of whether a dance video needs to add special effects. If there is an action at a certain moment in the dance video that meets the special effect adding type, the time corresponding to the action at that moment is used to determine the action special effect node; if there is an action at a certain moment in the dance video that does not meet the special effect adding type, the time corresponding to the action at that moment is not used to determine the action special effect node; wherein, the time corresponding to the action can be a time point for adding special effects, or it can contain a time period for which the special effects need to last. Optionally, the preset action type can be left / right head shaking, left / right hand raising, left / right hand shaking, left / right leg kicking, left / right hip swinging, left / right turning, squatting, standing up, body wave and other action types.

[0074] Optionally, it can be determined whether the posture of the target object in the dance video is of a special effect adding type based on the posture of the target object in the dance video; if so, the action meets the special effect adding type; if not, the action does not meet the special effect adding type.

[0075] The action effect node is a candidate node for adding visual effects in a music video. Optionally, the action effect node can be a time point or a screen identifier for indicating a time point. Optionally, each dance video can have multiple action effect nodes, and a target action effect node for adding visual effects to the dance video can be selected from the multiple action effect nodes according to the conditions met by the multiple action effect nodes.

[0076] In one embodiment, determining an action special effect node according to whether an action in a dance video meets the type of special effect addition includes: determining an action posture in the dance video; determining an action posture that meets the type of special effect addition; and determining an action special effect node in the dance video according to a time point corresponding to the action posture that meets the type of special effect addition. The time point corresponding to the action posture may be a time point when the action posture is executed, a time point before the action posture is executed, or a time point after the action posture is executed for a preset period of time.

[0077] In another embodiment, determining action special effect nodes according to whether actions in a dance video meet the type of special effect addition includes: determining a human key point sequence in the dance video; and determining action special effect nodes in the dance video according to time points corresponding to the human key point sequence that meets the type of special effect addition.

[0078] Step 206, adding a screen special effect to the dance video according to at least one of a beat special effect node and an action special effect node, to obtain a special effect dance video.

[0079] The screen special effect is a screen processing method set for the screen indicated by the special effect node. Optionally, the screen special effects include but are not limited to camera movement special effects, special effects that change the image characteristics of the screen, and special effects of adding elements. Camera movement special effects are screen processing methods for screen motion effects. Camera movement special effects include but are not limited to special effects such as screen zooming in, zooming out, moving left and right / shaking / shaking, moving up and down / shaking / shaking, and screen rotation. Special effects that change the image characteristics of the screen include but are not limited to gradient special effects, mutation special effects, and flashing special effects, which can change one or more of the image characteristics such as color, brightness, contrast, saturation, and edge strength. In the gradient special effect, it includes effects such as the screen brightness gradually dimming and the screen brightness gradually brightening. For the special effects of adding elements, it includes other animation elements such as photoelectric surround, lightning surround, graffiti stroke, angel wings, heart shape, and thumbs-up hand shape. Taking the photoelectric surround effect as an example, the human body edge detection and segmentation are performed on the target object to obtain the edge of the portrait, and then the photoelectric surround effect is added around the portrait.

[0080] The special effect dance video is a dance video with added visual effects. Optionally, visual effects can be added to the dance video according to the beat special effect node alone, or according to the action special effect node alone. Optionally, at least one of the beat special effect node and the action special effect node can be selected according to the conditions met by the beat special effect node and the action special effect node to add visual effects to the dance video.

[0081] Optionally, you can add visual effects to the dance video according to the intersection time point of the beat special effect node and the action special effect node; you can also select the intersection time point, the beat special effect node and the action special effect node according to the conditions met by the intersection time point to obtain the time point for adding visual effects, and add the visual effects at the time point for adding visual effects.

[0082] In an optional embodiment, adding visual special effects to a dance video according to at least one of a beat special effect node and an action special effect node includes: adding visual special effects to a dance video according to at least one of a beat special effect node; and / or, adding visual special effects to a dance video according to at least one of a action special effect node.

[0083] In an optional embodiment, according to at least one of the beat special effect nodes and the action special effect node, a special effect is added to a dance video to obtain a special effect dance video, including: according to the special effect node priority between the beat special effect node and the action special effect node, a high priority node is determined in the beat special effect node and the action special effect node; according to the high priority node, a special effect is added to the dance video to obtain a special effect dance video; wherein, the special effect node priority can be set for the environment of the dance video, or can be set according to the received instruction; the high priority node is a special effect node with a higher priority in the beat special effect node and the action special effect node. Therefore, by selecting the special effect node according to the priority, the detection process of the two special effect nodes can be avoided from conflicting, so as to ensure the efficiency of adding special effects.

[0084] In an optional embodiment, adding a special effect to the dance video to obtain a special effect dance video includes: when the special effect is a camera movement effect, controlling the movement of the pan / tilt head according to the camera movement effect; and during the movement of the pan / tilt head, performing video capture on the dance video to obtain a special effect dance video.

[0085] Optionally, controlling the movement of the pan / tilt head according to the camera movement special effect includes: controlling the pan / tilt head to move continuously according to a pan / tilt movement method for realizing the camera movement special effect.

[0086] Correspondingly, during the movement of the pan-tilt platform, the dance video is captured to obtain a special-effect dance video, including: during the process of controlling the continuous movement of the pan-tilt platform, the dance video is continuously captured to obtain a special-effect dance video.

[0087] Specifically, when following, zooming in, zooming out, moving left and right / shaking / jittering, moving up and down / shaking / jittering, rotating the screen and other camera movement effects are used, the PTZ can be controlled to move according to the camera movement effects. When the camera movement effect is a follow effect, the PTZ is controlled to follow a certain person in the dance video to collect video. When the camera movement effect is a left shake effect, the PTZ is controlled to shake to its left side and collect video of the dance video while shaking to the left.

[0088] In the above-mentioned method for adding special effects to dance videos, the beat special effects node is determined according to the music beat in the dance video, and the characteristic of the music beat in representing the amplitude strength law is used to accurately determine the beat special effects node, which can accurately control the process of adding picture special effects from the audio perspective; the action special effects node is determined according to the situation that the action in the dance video meets the type of special effects addition, which can make the result of human body action analysis appear and accurately identify the corresponding dance action; finally, according to at least one special effect node among the beat special effects node and the action special effects node, the dance video is added with picture special effects to obtain a special effects dance video, which can efficiently and automatically add picture special effects such as video camera effects to the accurately identified specific music nodes and specific dance actions, so that the video has a dynamic effect.

[0089] In one embodiment, a beat effect node is determined based on the music beat in a dance video, including: generating music amplitude based on the audio in the dance video; performing feature extraction on the music amplitude at a scale that increases step by step to obtain amplitude features at each scale; input features for feature extraction of amplitude features at each scale are obtained by downsampling output features of the previous level at a preset step; determining a strong beat position for characterizing an accented beat based on the amplitude features at each scale; and determining a beat effect node based on the strong beat position.

[0090] The amplitude features of each scale have a progressively larger scale. Optionally, each scale can be adjusted by using a more rapidly increasing dilation factor to achieve a larger receptive field without requiring too many parameters.

[0091] The preset stride is used to characterize the number of pixels skipped during feature extraction. Optionally, the preset stride is a stride value. When the stride value is greater than 1, the convolution kernel will skip a certain number of pixels each time it slides, so that the size of the output feature map is reduced to achieve downsampling.

[0092] In an optional embodiment, music amplitude is generated based on audio in a dance video, including: determining sampling points at a fixed sampling rate or a sampling rate of the original audio; reading audio data in the dance video at each sampling point to obtain the music amplitude of the audio data at each sampling point; wherein the fixed sampling rate may be a sampling rate preset to 22050 Hz.

[0093] In a specific implementation, feature extraction is performed on the music amplitude at progressively increasing scales to obtain amplitude features at each scale, including: determining a current amplitude feature of the music amplitude at a current scale; downsampling the current amplitude feature at a preset step size to obtain a downsampled current amplitude feature; performing feature extraction based on the downsampled current amplitude feature to obtain an amplitude feature at a next scale; wherein the next scale is a scale smaller than the current scale.

[0094] Optionally, feature extraction is performed based on the downsampled current amplitude features to obtain amplitude features at the next scale, including: feature extraction is performed based on the downsampled current amplitude features to obtain amplitude features at the next scale, until feature extraction is performed at each scale to obtain amplitude features at each scale.

[0095] In another specific embodiment, the music amplitude is feature extracted according to the scales increasing step by step to obtain the amplitude features of each scale; including: inputting the music amplitude into the fluctuation beat detection model of the convolution and residual network structure to extract the amplitude features of each scale through the fluctuation beat detection model. Correspondingly, the strong beat position for characterizing the accent beat is determined according to the amplitude features of each scale, including: receiving the positions of each accent beat output by the fluctuation beat detection model according to the amplitude features.

[0096] In an optional implementation, the beat effect node is determined according to the strong beat position, including: using the time point where the strong beat position is located as the beat effect node; or, filtering the time point where the strong beat position is located, and using the filtered time point as the beat effect node.

[0097] In this embodiment, by using an increasing scale, the model can achieve a larger receptive field without too many parameters, so as to achieve a larger receptive field while maintaining computational efficiency; and sampling the features at each scale according to the preset stride is a stride convolution process with dynamically changing windows, and the signal is downsampled when passing through the network. This downsampling can reduce computational and memory requirements while still maintaining sufficient time resolution for beat tracking. Therefore, by fusing the increasing scale and downsampling, it is possible to obtain a large amount of context, and the computational efficiency is high, thereby more efficiently estimating the strong beat position in the music.

[0098] In one embodiment, an action special effect node is determined based on whether an action in a dance video meets the type of special effect addition, including: detecting a target person based on multiple frames of the dance video; detecting the body posture of the target person in the multiple frames to obtain a body key point sequence; identifying the action of the target person in each frame based on the body key point sequence to obtain an action type; if the action type meets a preset dance special effect action type, determining the action special effect node based on a time point corresponding to the action.

[0099] The human body key points are set for each target person to reflect the skeleton information of the target person. Optionally, the human body key points are positions that can produce curvature changes, including but not limited to key points of the human body such as the neck, shoulder, elbow, wrist, waist, knee, ankle, etc. The human body key point sequence is the human body key points arranged in chronological order. The human body key points at the same time can be connected to the skeleton information at this time. The human body posture can be determined through the skeleton information at multiple times.

[0100] Detecting the target person based on multiple frames of the dance video is a process of target detection. All the target objects of interest in the image are identified, and the target person and the target person's position are determined. Unlike image classification, which only focuses on whether there is a target in the image, target detection also needs to determine the position and bounding box of the target person. Optionally, after determining the target person, target tracking is performed based on the target person; target tracking means that after the target position of the target person in the first frame of the image is given, the position of the target person in subsequent frames is predicted based on the tracking algorithm.

[0101] In an optional implementation, detecting a target person based on multiple frames of a dance video includes: performing person detection on the multiple frames of the dance video according to a traditional detection method, and determining the detected person as the target person; wherein the traditional detection method may be a real-time object detection algorithm based on feature faces (Viola-Jones), histogram of oriented gradients (HOG), scale-invariant feature transform (SIFT), edge detection algorithm, template matching, color features, and the like; for the histogram of oriented gradients, the feature is a feature descriptor used for object detection in computer vision and image processing.

[0102] In another optional embodiment, the target person is detected based on multiple frames of a dance video, including: detecting multiple frames of a dance video through a neural network model for human body detection; determining the person to whom the detected human body belongs as the target person; wherein the neural network model can be a faster region-based convolutional neural network (Faster Region-based Convolutional Neural Networks, Faster R-CNN), a single-stage dense frame detector (Single Shot MultiBox, SSD), a retina network (RetinaNet), and an Exceeding YOLO series in 2021 (Exceeding YOLO series in 2021, YOLOX), etc.; wherein the faster region-based convolutional neural network is a two-stage target detection algorithm. Optionally, the detector used to detect the target person can be a single-task detector or a multi-task detector.

[0103] In an optional embodiment, the human key point sequence is implemented by a deep learning-based model. The deep learning-based model can be a real-time human posture recognition algorithm. The real-time human posture recognition algorithm includes, but is not limited to, OpenPose, DensePose, and EfficientPose. OpenPose is a human posture recognition project developed by Carnegie Mellon University (CMU) in the United States based on convolutional neural networks and supervised learning and developed with Caffe as the framework. OpenPose is an open source library developed by Facebook researchers Natalia Neverova, Iasonas Kokkinos, and INRIA in France. An amazing real-time human pose recognition system developed by Alp Guler; EfficientPose, which is an efficient, accurate and scalable end-to-end 6D multi object pose estimation approach, is a new method for 6D target pose estimation. The input of the network is an RGB image, and the output of the network is the 2D Bounding Box of all objects to be detected in the image and the 6D Pose (x, y, z, roll, yaw, pitch) in three-dimensional space. For example, OpenPose extracts 18 body joints and 17 lines connecting the joints to obtain the skeleton. These skeleton information will be stored in chronological order for input features for subsequent action recognition.

[0104] In an optional implementation, the action of the target person in each frame is identified according to the human body key point sequence to obtain the action type, including: determining the human body key points at the same moment according to the human body key point sequence; connecting the key points at each moment into multiple frames of skeleton information; performing action recognition based on the multiple frames of skeleton information in time sequence to obtain the action type of the target person.

[0105] By arranging the skeleton key points of each target person in time sequence, the multi-frame skeleton information of each target person can be formed, so as to represent the action of the corresponding target person through the multi-frame skeleton information. The multi-frame skeleton information is connected in sequence according to the human body key points of each target person in the video screen. The multi-frame skeleton information is composed of the skeleton information of each frame. The skeleton information of each frame is as follows: Figure 4 As shown in (1), (2) and (3) in FIG, a schematic diagram of action recognition based on the skeleton information of multiple frames of images based on the temporal sequence is shown in FIG. Figure 5 As shown; among them, Figure 4 The a-th frame, the b-th frame and the c-th frame are arranged in chronological order.

[0106] In a feasible implementation, the action of the target person in each frame is identified according to the human body key point sequence to obtain the action type, including: based on the time sequence, performing action recognition on the skeleton information of multiple frames composed of the human body key point sequence to obtain the action characteristics of the target person; determining the action type of the target person according to the action type that the action characteristics of the target person conform to.

[0107] In another feasible implementation, the action of the target person in each frame is identified according to the sequence of human key points to obtain the action type, including: standardizing the spatiotemporal graph composed of the human key points in the video screen to obtain the skeleton information of multiple frames in time sequence; substituting the skeleton information of multiple frames in time sequence into a certain action recognition algorithm for encoding to obtain the action characteristics of the target person; inputting the action characteristics of the target person into a classification network for action recognition to realize action recognition, and outputting the defined action type. Optionally, the algorithm for obtaining the action type is a skeleton-based behavior recognition algorithm. Skeleton-based behavior recognition algorithms include but are not limited to spatial temporal graph convolutional networks (Spatial Temporal Graph Convolutional Networks, ST-GCN), adaptive graph convolutional networks (Adaptive Graph Convolutional Neural Networks, AGCN), and channel topology refinement graph convolution networks (Channel-wise Topology Refinement Graph Convolution, CTR-GCN).

[0108] Optionally, since the size of the human body in the picture is variable, the detected human body key points are at different scales, so the scales of the key points of different frames of video need to be unified. In addition, due to the viewing angle, the human body may rotate, so the connection lines of some human body key points need to be consistent with the direction of the reference line. Position standardization generally involves stringing together human body frames of different frames according to a point (such as the center point of the spine) for standardization on the time axis after completing scale normalization and angle normalization.

[0109] For example, the model input dimension of the classification network for action recognition is usually (N, C, T, V, M), where:

[0110] N represents the number of videos. Usually a batch has 256 videos. The number of videos should preferably be a power of 2.

[0111] C represents the characteristics of the joint. Usually a joint contains three features: x, y, and acc (if it is a three-dimensional skeleton, it is four). x and y are the position coordinates of the node joint, and acc is the confidence.

[0112] T represents the number of key frames in the video. Generally, a video has 150 frames.

[0113] V represents the number of joints, and usually one person labels 18 joints.

[0114] M represents the number of people in a frame, and generally the two people with the highest average confidence are selected.

[0115] In an optional embodiment, an action special effect node is determined according to the time point corresponding to the action, including: determining a reference time point for executing the action of the target character in each frame; determining the reference time point as the action special effect node; or, setting an identifier of the action special effect node for the reference time point to obtain the action special effect node; or, determining a forward time point with a preset time length between the reference time point and the action special effect node; or, determining a backward time point with a preset time length between the reference time point and the action special effect node; wherein the reference time point may be the starting time point when the action is executed, may be a point after one or several frames of the action execution, or may be an action receiving time point; the forward time point is a time point before the reference time point, and the backward time point is a time point before the reference time point.

[0116] By detecting the target person based on multiple frames of the dance video, target detection is achieved, which lays a foundation for high accuracy of posture detection. Therefore, the human body posture of the target person in multiple frames is detected to obtain a human body key point sequence. The human body key point sequence is used to reflect the movement changes in a certain time period. The movement of the target person in each frame is identified according to the human body key point sequence to obtain the movement type, thereby achieving accurate recognition of the movement type; therefore, when the movement type conforms to the preset dance special effect movement type, the movement special effect node is efficiently and accurately determined according to the time point corresponding to the movement.

[0117] In one embodiment, according to at least one of the beat special effect nodes and the action special effect node, a visual special effect is added to a dance video to obtain a special effect dance video, including: determining the note strength at the time point corresponding to the beat special effect node; the note strength is used to characterize the amplitude change at the time point corresponding to the beat special effect node; determining a target beat special effect node in the beat special effect node according to the note strength; adding a camera movement special effect to the dance video according to the target beat time point corresponding to the target beat special effect node to obtain a special effect dance video.

[0118] Note intensity is used to characterize the amplitude change at the time point corresponding to the beat effect node. Optionally, note intensity can be characterized according to the note intensity envelope to characterize the amplitude at a certain time point. The beat effect node can be multiple time points, and the note intensity at multiple time points can characterize the volume change at the start, duration and end stages of the note. Therefore, the peak value of each strong beat node can be determined by the note intensity envelope to screen out the target beat effect node. Optionally, the note intensity at multiple time points forms a note intensity envelope (Amplitude Envelope), which is a curve of the amplitude of the note changing over time.

[0119] The time point corresponding to the beat effect node and the beat time point corresponding to the target beat effect node. Optionally, the beat time point may be the time point where the beat effect node is located; or the time point represented by the beat effect node.

[0120] The target beat special effect node is a beat special effect node used to add camera movement special effects to the dance video; the beat time point corresponding to the target beat special effect node is the target beat time point. Optionally, the number of target beat special effect nodes is preset.

[0121] In a feasible implementation, the beat effect node is determined according to the music beat in the dance video, including: performing a short-time Fourier transform on the amplitude of the music to obtain a spectrum diagram of the amplitude; generating a corresponding logarithmic Mel spectrum based on the spectrum diagram of the amplitude; calculating the first-order time difference of the logarithmic Mel spectrum, and calculating the mean of the logarithmic spectrum difference degree in each frame; normalizing and smoothing the mean of the differential amplitude of all frames to obtain a note intensity envelope; correspondingly, determining the note intensity at the time point corresponding to the beat effect node, including: determining the note intensity represented by the note intensity envelope at the time point corresponding to the beat effect node. Optionally, the note intensity represented by the beat effect node at the time point corresponding to the peak value of the note intensity envelope can be used to represent the note intensity represented by the peak value of the note intensity envelope.

[0122] In another feasible implementation, determining the note strength at the time point where the beat effect node is located includes: determining the note strength at the time point where the beat effect node is located according to the music score corresponding to the time point where the beat effect node is located.

[0123] In a feasible implementation, determining the target beat effect node in the beat effect node according to the note strength includes: selecting a preset number of beat effect nodes from the note strength of the beat effect node as the target beat effect node according to the note strength from large to small. Exemplarily, first determine the step note strength envelope and peak value, and select from large to small according to the note strength peak value at the time point where the strong beat is located, until N strong beat nodes are selected as the target beat effect nodes.

[0124] In another feasible implementation, determining a target beat effect node among the beat effect nodes according to the note strength includes: selecting a beat effect node whose note strength is a preset value from the note strengths of the beat effect nodes as the target beat effect node.

[0125] In a specific embodiment, the first beat special effect node and the second beat special effect node are adjacent and continuous beat special effect nodes, the first beat special effect node corresponds to the first beat time point, the second beat special effect node corresponds to the second beat time point, and the first beat special effect node has been selected as the target beat special effect node. Determining the target beat special effect node among the beat special effect nodes according to the note intensity includes: determining the time difference between the first beat time point and the second beat time point; if the time difference is greater than a preset duration, determining the second beat special effect node as the target beat special effect node for adding camera movement special effects to the dance video.

[0126] Optionally, after determining the time difference between the first beat time point and the second beat time point, it also includes: if the time difference is less than a preset duration, the second beat special effect node is not a target beat special effect node for adding camera movement special effects to the dance video.

[0127] Optionally, after determining the time difference between the first beat time point and the second beat time point, it also includes: if the time difference is less than the preset duration, then the time difference between the first beat time point and the third beat time point is calculated to obtain the time difference after the jump point; the third beat time point corresponds to the second beat special effect node, and the second beat special effect and the third beat special effect nodes are adjacent and continuous beat special effect nodes; if the time difference after the jump point is greater than the preset duration, the third beat special effect node is determined as the target beat special effect node for adding camera movement special effects to the dance video. Exemplarily, if the distance between the second strong beat node currently being judged and the first strong beat node that has been selected is less than 5 seconds, the current node is skipped, and the next strong beat node is continuously judged to be the target beat special effect node.

[0128] In a feasible implementation, according to the target beat time point corresponding to the target beat special effect node, a camera movement special effect is added to a dance video to obtain a special effect dance video, including: according to the target beat time point corresponding to the target beat special effect node, randomly adding camera movement special effects to the dance video to obtain a special effect dance video; or, according to the target beat time point corresponding to the target beat special effect node and the amplitude size of the target beat special effect node, randomly adding camera movement special effects with amplitude matching the dance video to obtain a special effect dance video.

[0129] Specifically, the method further includes: controlling the movement of the pan / tilt platform according to the camera movement special effect; and during the movement of the pan / tilt platform, capturing a dance video to obtain a special effect dance video.

[0130] In this embodiment, based on the amplitude strength law of the audio, the target beat special effect node is screened out according to the amplitude change at the time point represented by the note intensity. This screening process is based on the volume changes in the starting, continuing and ending stages of the note level, and can more accurately control the process of adding screen special effects.

[0131] In one embodiment, according to at least one of the beat special effect nodes and the action special effect nodes, visual special effects are added to a dance video to obtain a special effect dance video, including: determining the camera movement special effects of the action special effect node according to the special effect adding type that the action conforms to; adding the camera movement special effects of the action special effect node to the dance video according to the action time point corresponding to the action special effect node to obtain the special effect dance video.

[0132] In an optional implementation, determining the camera movement special effect of the action special effect node according to the special effect adding type that the action conforms to includes: determining a preset camera movement special effect corresponding to the special effect adding type according to the special effect adding type that the action conforms to; and determining the preset camera movement special effect as the camera movement special effect of the action special effect node.

[0133] In another optional embodiment, the camera movement special effect of the action special effect node is determined according to the special effect adding type that the action matches, including: determining the camera movement special effect matched by the special effect adding type that the action matches as the camera movement special effect of the action special effect node. Optionally, the special effect adding type and the camera movement special effect bound to it have the same movement direction. Optionally, the camera movement special effect bound to the special effect adding type can be used as the camera movement special effect matched by the special effect adding type by pre-setting a matching table of actions and camera movement effects; the special effect adding type and the camera movement special effect can also be randomly matched to obtain the camera movement special effect matched by the special effect adding type. Exemplarily, when the target object performs a leftward action, the picture suddenly turns to the left to indicate that the camera is moving to the left; when the target object performs a squatting action, the picture first moves downward and then moves upward to indicate that the camera moves downward first and then upward, and so on.

[0134] Specifically, after the camera movement special effect that matches the special effect adding type that matches the action is determined as the camera movement special effect of the action special effect node, it includes: controlling the movement of the pan-tilt head according to the camera movement special effect; during the movement of the pan-tilt head, capturing the dance video to obtain the special effect dance video.

[0135] In an optional implementation, determining the camera movement special effects of the action special effects node according to the special effects adding type that the action special effects node conforms to includes: determining the camera movement special effects of each action special effects node according to the special effects adding type that each action special effects node conforms to. Correspondingly, adding the camera movement special effects of the action special effects node to the dance video according to the action time point corresponding to the action special effects node to obtain the special effects dance video includes: adding the camera movement special effects of each action special effects node to the dance video at the action time point corresponding to each action special effects node to obtain the special effects dance video.

[0136] In this implementation, the camera movement special effects of the action special effect node are determined according to the special effect adding type that the action special effect node conforms to, so that the action special effect node can have diversified camera movement special effects; and according to the action time point corresponding to the action special effect node, the camera movement special effects of the action special effect node are added to the dance video to obtain a special effects dance video, so that the camera movement special effects in the special effects dance video are added to the dance video according to human body movements, so that the special effects adding process is accurate and efficient.

[0137] In one embodiment, according to at least one of the beat special effect nodes and the action special effect node, a special effect is added to a dance video to obtain a special effect dance video, including: determining a target beat time point according to the beat special effect node; determining an action time point corresponding to the action special effect node; determining an intersection time point of the target beat time point and the action time point; according to the intersection time point, adding a special effect to the dance video to obtain a special effect dance video.

[0138] The intersection time point is the time point where the target beat time point and the action time point intersect. Optionally, the target beat time point that is the same as the action time point can be found, and the intersection time point is determined based on the found target beat time point; the action time point that is the same as the target beat time point can be found, and the intersection time point is determined based on the found action time point.

[0139] In an optional embodiment, determining the target beat time point according to the beat special effect node includes: determining the note strength of the time point corresponding to the beat special effect node; the note strength is used to characterize the amplitude change of the time point corresponding to the beat special effect node; and determining the target beat special effect node in the beat special effect node according to the note strength.

[0140] In another optional implementation, the target beat time point is determined according to the beat special effect node, including: judging whether the current beat special effect node satisfies the time interval, if so, determining the current beat special effect node as the target beat time point, and so on, until judging whether the last beat special effect node satisfies the time interval, then stopping determining the target beat time point.

[0141] In an optional implementation, determining the action time point corresponding to the action special effect node includes: determining the corresponding action time point according to the reference time point of the action special effect node.

[0142] In an optional implementation, special effects are added to a dance video according to an intersection time point to obtain a special effects dance video, including: determining whether the intersection time point is valid; if valid, adding special effects to the dance video at the intersection time point to obtain a special effects dance video; if invalid, not adding special effects to the dance video according to the intersection time point to obtain a special effects dance video.

[0143] In this embodiment, the target beat time point is determined according to the beat special effect node to clarify the time point where the picture special effect needs to be added from the audio dimension; the action time point corresponding to the action special effect node is determined to clarify the time point where the picture special effect needs to be added from the picture dimension; on this basis, the intersection time point of the target beat time point and the action time point is determined, which realizes the comprehensive consideration of the audio and picture dimensions and reduces the time point for adding the camera movement special effect; finally, according to the intersection time point, the picture special effect is added to the dance video to obtain the special effect dance video, so that the number of time points for adding the picture special effect to the special effect dance video is appropriate, thereby improving the efficiency of adding the picture special effect.

[0144] In one embodiment, according to the intersection time points, special effects are added to a dance video to obtain a special effects dance video, including: determining whether the number of intersection time points meets the intersection special effects condition; if so, adding special effects to the dance video at the intersection time points to obtain a special effects dance video; if not, adding special effects to the dance video according to high priority nodes among the beat special effects nodes and the action special effects nodes to obtain a special effects dance video.

[0145] The intersection special effect condition is an evaluation index of the number of intersection time points, which is used to determine whether the number of intersection time points is too small. Optionally, when the number of intersection time points is not too small, the intersection special effect condition is met; when the number of intersection time points is too small, the intersection special effect condition is not met. Optionally, the intersection special effect condition can be a threshold of the number of intersection time points, or a ratio threshold. A high priority node is a special effect node among a beat special effect node and an action special effect node, and a high priority node is a special effect node that is effectively used.

[0146] In a feasible implementation, determining whether the number of intersection time points satisfies the intersection special effect condition includes: determining whether the number of intersection time points is greater than a threshold value of the number of intersection time points; if so, the intersection special effect condition is satisfied, and if not, the intersection special effect condition is not satisfied.

[0147] In another feasible implementation, the intersection time point number threshold is a ratio threshold; judging whether the number of intersection time points meets the intersection special effect condition includes: determining the ratio between the number of intersection time points and the dance video; if the ratio is greater than the ratio threshold, the intersection special effect condition is met; if the ratio is less than the ratio threshold, the intersection special effect condition is not met. Thus, the ratio threshold can be adaptively adjusted according to the time length of the dance video to more accurately judge whether the number of intersection time points meets the intersection special effect condition.

[0148] In one feasible implementation, at intersection time points, visual special effects are added to a dance video to obtain a special effects dance video, including: at each intersection time point, the movements in the dance video meet the special effects adding type, and the camera movement special effects at each intersection time point are determined; at each intersection time point, the camera movement special effects at each intersection time point are added to the dance video to obtain a special effects dance video.

[0149] In another feasible implementation, at the intersection time points, visual special effects are added to the dance video to obtain a special effects dance video, including: at each intersection time point, determining the camera movement special effects at each intersection time point according to the note intensity; at each intersection time point, adding the camera movement special effects at each intersection time point to the dance video to obtain a special effects dance video.

[0150] In a feasible implementation, according to the high-priority nodes among the beat special effect nodes and the action special effect nodes, visual special effects are added to a dance video to obtain a special effect dance video, including: according to the node priority acquired in advance, determining one of the beat special effect nodes and the action special effect node as a high-priority node; according to the high-priority node, adding visual special effects to the dance video to obtain a special effect dance video.

[0151] In another feasible implementation, according to high-priority nodes among beat special effect nodes and action special effect nodes, visual special effects are added to a dance video to obtain a special effect dance video, including: according to the number of beat special effect nodes and the number of action special effect nodes, determining a special effect node with a small number as a high-priority node; adding visual special effects to the dance video through the high-priority nodes to obtain a special effect dance video.

[0152] In this implementation, it is determined whether the number of intersection time points meets the intersection special effects condition, and the number of intersection time points is used to estimate the effect of adding visual special effects to the dance video at the intersection time points; if the effect is better, a smaller number of intersection time points corresponding to both audio and picture angles are used to add visual special effects to the dance video; otherwise, visual special effects are added to the dance video based on high-priority nodes in the beat special effects nodes and action special effects nodes to avoid data conflicts and adding too many special effects, and can be adaptively adjusted according to different situations, so as to more efficiently add visual special effects.

[0153] In an exemplary embodiment, Figure 6 As shown in Figure 6 As shown in (a) of the figure, the target person turns right and a rotation effect is added to the picture; Figure 6 As shown in (b) in the figure, the target person image is shaken, and a shaking effect is added to the picture; Figure 6 As shown in (c) in the figure, the target person turns right and shakes, and the picture adds shaking and zoom processing in the left direction; Figure 6 As shown in (d), the target person turns left and shakes, and the picture adds shaking and zoom processing in the left direction.

[0154] Based on this, the present application proposes an automatic editing solution for dance videos. For dance videos with music, the music beat is automatically analyzed, the human body movements are analyzed on the video screen, and video camera effects are automatically added to specific music nodes and dance movements, so that the video has a dynamic effect and increases the watchability of the video. The present invention not only saves the time and energy cost of users to learn editing software, but also replaces the lengthy process of users manually editing using video editing software, and has high editing efficiency.

[0155] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0156] Based on the same inventive concept, the embodiment of the present application also provides a dance video special effect adding device for implementing the dance video special effect adding method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more dance video special effect adding device embodiments provided below can refer to the limitations of the dance video special effect adding method above, and will not be repeated here.

[0157] In an exemplary embodiment, Figure 7 As shown, a dance video special effect adding device is provided, comprising:

[0158] A beat detection module 702, used to determine a beat special effect node according to the music beat in the dance video;

[0159] The action detection module 704 is used to determine the action special effect node according to whether the action in the dance video meets the special effect adding type;

[0160] The special effect adding module 706 is used to add visual special effects to the dance video according to at least one of the beat special effect nodes and the action special effect nodes to obtain a special effect dance video.

[0161] In one embodiment, the beat detection module 702 is used to:

[0162] Generate music amplitude based on audio from dance videos;

[0163] The music amplitude is subjected to feature extraction according to a scale that increases step by step, so as to obtain the amplitude features of the scale at each level; the input features for feature extraction of the amplitude features of the scale at each level are obtained by downsampling the output features of the previous level according to a preset step size;

[0164] Determining the strong beat position for characterizing the accented beat according to the amplitude characteristics of the scales at each level;

[0165] A beat special effect node is determined according to the strong beat position.

[0166] In one embodiment, the action detection module 704 is used to:

[0167] Detecting a target person based on multiple frames of the dance video;

[0168] Detecting the human body posture of the target person in the multiple frames to obtain a human body key point sequence;

[0169] Identify the action of the target person in each frame of the picture according to the human body key point sequence to obtain the action type;

[0170] If the action type meets the preset dance special effect action type, the action special effect node is determined according to the time point corresponding to the action.

[0171] In one embodiment, the special effect adding module 706 is used to:

[0172] Determine the note strength at the time point corresponding to the beat special effect node; the note strength is used to represent the amplitude change at the time point corresponding to the beat special effect node;

[0173] Determine a target beat special effect node among the beat special effect nodes according to the note intensity;

[0174] According to the target beat time point corresponding to the target beat special effect node, a camera movement special effect is added to the dance video to obtain a special effect dance video.

[0175] In one embodiment, the special effect adding module 706 is used to:

[0176] Determining the camera movement special effect of the action special effect node according to the special effect adding type that the action matches;

[0177] According to the action time point corresponding to the action special effect node, the camera movement special effect of the action special effect node is added to the dance video to obtain a special effect dance video.

[0178] In one embodiment, the special effect adding module 706 is used to:

[0179] Determine the target beat time point according to the beat special effect node;

[0180] Determine the action time point corresponding to the action special effect node;

[0181] Determine the intersection time point of the target beat time point and the action time point;

[0182] According to the intersection time point, a special effect is added to the dance video to obtain a special effect dance video.

[0183] In one embodiment, the special effect adding module 706 is used to:

[0184] Determine whether the number of the intersection time points meets the intersection special effect condition;

[0185] If yes, then at the intersection time point, adding a special effect to the dance video to obtain a special effect dance video;

[0186] If not, add screen special effects to the dance video according to the high priority nodes among the beat special effect nodes and the action special effect nodes to obtain a special effect dance video.

[0187] In one embodiment, the special effect adding module 706 is used to:

[0188] When the screen special effect is a camera movement special effect, the pan / tilt movement is controlled according to the camera movement special effect;

[0189] During the movement of the pan / tilt platform, the dance video is captured to obtain a special effects dance video.

[0190] Each module in the dance video special effects adding device can be implemented in whole or in part by software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each module.

[0191] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 8 As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected via a system bus, and the communication interface, the display unit and the input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a method for adding special effects to a dance video is implemented. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device shell, or an external keyboard, touchpad or mouse.

[0192] Those skilled in the art will understand that Figure 8 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0193] In one embodiment, the present application further provides a gimbal, including a motor and a processor, wherein the motor is used to control the rotation of the gimbal, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program; wherein the gimbal can be a handheld gimbal of a real person, or a virtual gimbal; when the handheld gimbal can be a handheld gimbal of a real person, the motor is a real motor, and the processor of the handheld gimbal can be the real processor of the gimbal itself, or a processor of a mobile phone installed on the handheld gimbal. When the above-mentioned gimbal is a virtual gimbal, the motor and the processor are both virtual.

[0194] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.

[0195] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0196] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0197] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0198] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.

[0199] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0200] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A method for adding special effects to a dance video, characterized in that: The method comprises: Determine the beat special effect node according to the music beat in the dance video; Determining an action special effect node according to whether an action in the dance video meets the type of special effect addition; According to at least one of the beat special effect nodes and the action special effect nodes, a screen special effect is added to the dance video to obtain a special effect dance video.

2. The method according to claim 1, characterized in that Determining the beat special effect node according to the music beat in the dance video includes: Generate music amplitude based on audio from dance videos; The music amplitude is subjected to feature extraction according to a scale that increases step by step, so as to obtain the amplitude features of the scale at each level; the input features for feature extraction of the amplitude features of the scale at each level are obtained by downsampling the output features of the previous level according to a preset step size; Determining the strong beat position for characterizing the accented beat according to the amplitude characteristics of the scales at each level; A beat special effect node is determined according to the strong beat position.

3. The method according to claim 1, characterized in that The step of determining the action special effect node according to the situation that the action in the dance video meets the special effect adding type includes: Detecting a target person based on multiple frames of the dance video; Detecting the human body posture of the target person in the multiple frames to obtain a human body key point sequence; Identify the action of the target person in each frame of the picture according to the human body key point sequence to obtain the action type; If the action type meets the preset dance special effect action type, the action special effect node is determined according to the time point corresponding to the action.

4. The method according to claim 1, characterized in that The step of adding a screen special effect to the dance video according to at least one of the beat special effect node and the action special effect node to obtain a special effect dance video comprises: Determine the note strength at the time point corresponding to the beat special effect node; the note strength is used to represent the amplitude change at the time point corresponding to the beat special effect node; Determine a target beat special effect node among the beat special effect nodes according to the note intensity; According to the target beat time point corresponding to the target beat special effect node, a camera movement special effect is added to the dance video to obtain a special effect dance video.

5. The method according to claim 1, characterized in that The step of adding a screen special effect to the dance video according to at least one of the beat special effect node and the action special effect node to obtain a special effect dance video comprises: Determining the camera movement special effect of the action special effect node according to the special effect adding type that the action matches; According to the action time point corresponding to the action special effect node, the camera movement special effect of the action special effect node is added to the dance video to obtain a special effect dance video.

6. The method according to claim 1, characterized in that The step of adding a screen special effect to the dance video according to at least one of the beat special effect node and the action special effect node to obtain a special effect dance video comprises: Determine the target beat time point according to the beat special effect node; Determine the action time point corresponding to the action special effect node; Determine the intersection time point of the target beat time point and the action time point; According to the intersection time point, a special effect is added to the dance video to obtain a special effect dance video.

7. The method according to claim 6, characterized in that The step of adding a special effect to the dance video according to the intersection time point to obtain a special effect dance video comprises: Determine whether the number of the intersection time points meets the intersection special effect condition; If yes, then at the intersection time point, adding a special effect to the dance video to obtain a special effect dance video; If not, add screen special effects to the dance video according to the high priority nodes among the beat special effect nodes and the action special effect nodes to obtain a special effect dance video.

8. The method according to claim 1, characterized in that The step of adding special effects to the dance video to obtain a special effects dance video includes: When the screen special effect is a camera movement special effect, the pan / tilt movement is controlled according to the camera movement special effect; During the movement of the pan / tilt platform, the dance video is captured to obtain a special effects dance video.

9. A device for adding special effects to a dance video, characterized in that: The device comprises: A beat detection module is used to determine the beat special effect node according to the music beat in the dance video; An action detection module is used to determine an action special effect node according to whether the action in the dance video meets the type of special effect addition; The special effect adding module is used to add screen special effects to the dance video according to at least one of the beat special effect nodes and the action special effect nodes to obtain a special effect dance video.

10. A pan / tilt head, characterized in that: The invention comprises a motor and a processor, wherein the motor is used to control the rotation of the pan / tilt head, and the processor is used to implement the steps of the method according to any one of claims 1 to 8.

11. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.