Adding Animation Effects Based on Voice Features

The method synchronizes video effects with audio features by analyzing video and audio data to apply dynamic effects, addressing the lack of audio-controlled effects in existing video editing techniques and enhancing user experience.

JP7711876B2Active Publication Date: 2025-07-23LEMON CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023544220
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-06-21
Filing Date
2022-05-10
Publication Date
2025-07-23
Estimated Expiration
2042-05-10

AI Technical Summary

Technical Problem

Existing video editing techniques do not effectively synchronize visual effects with audio data, lacking the ability to control effects based on audio features.

Method used

A method and system for rendering video effects that analyze video and audio data to determine additional points and animation-related effects, applying these effects to the video based on audio features such as beat and frequency spectrum, creating synchronized video effects.

Benefits of technology

Enables real-time or near-real-time synchronization of video effects with music beats and frequency characteristics, enhancing user experience by dynamically reacting to audio elements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007711876000001
    Figure 0007711876000001
  • Figure 0007711876000002
    Figure 0007711876000002
  • Figure 0007711876000003
    Figure 0007711876000003
Patent Text Reader

Abstract

A system and method are described for rendering a video effect on a display. More specifically, video data and audio data are obtained. The video data is analyzed to determine one or more attachment points of a target object appearing in the video data. The audio data is analyzed to determine audio characteristics. Based on the audio characteristics, an animation-related video effect is determined to be added to the one or more attachment points. A rendered video is generated by applying the video effect to the video data.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Video editing techniques are widely used to provide users with various ways to edit videos. For example, a user may edit a video to add visual effects and / or music to the video. However, many of these video editing techniques do not consider controlling visual effects based on audio data. Therefore, in order to improve the user experience, it is necessary to develop video editing techniques that render the synchronization of video effects.

[0002] As related to these and other general considerations, aspects disclosed herein are described. Also, relatively specific problems will be discussed, but it should be understood that these examples should not be limited to solving the specific problems identified in the background art or other parts of the present disclosure.

Summary of the Invention

[0003] According to at least one example of the present disclosure, a method for rendering video effects on a display is provided. The method includes obtaining video data and audio data, analyzing the video data to determine one or more additional points of a target object appearing in the video data, analyzing the audio data to determine audio features, determining animation-related video effects to be added to the one or more additional points based on the audio features, and generating a rendered video by applying the video effects to the video data.

[0004] According to at least one example of the present disclosure, a computing device for rendering video effects on a display is provided. The computing device includes a processor and, when executed by the processor, causes the computing device to: obtain video data and audio data; analyze the video data to determine one or more additional points of a target object appearing in the video data; analyze the audio data to determine audio features; determine animation-related video effects to be added to the one or more additional points based on the audio features; and generate a rendered video by applying the video effects to the video data. and a memory storing a plurality of instructions for causing the above to be executed.

[0005] According to at least one example of the present disclosure, a non-transitory computer-readable medium storing instructions for rendering video effects on a display is provided. When executed by one or more processors of a computing device, the instructions cause the computing device to: obtain video data and audio data; analyze the video data to determine one or more additional points of a target object appearing in the video data; analyze the audio data to determine audio features; determine animation-related video effects to be added to the one or more additional points based on the audio features; and generate a rendered video by applying the video effects to the video data. to be executed.

[0006] Any one of the one or more aspects is combined with any other aspect of the one or more aspects. It is any one of the one or more aspects described herein.

[0007] This summary is provided to introduce, in a simplified form, a selection of concepts that are further described in the "Detailed Description of the Invention" below. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Additional aspects, features, and / or advantages of the examples will be described in part in the following description, will become apparent in part from the description, or will be understood by practice of the present disclosure. **Brief Description of the Drawings**

[0008] The following drawings are referred to in order to describe non-limiting and non-exhaustive examples.

[0009]

Figure 1

[0010]

Figure 2

[0011]

Figure 3A

Figure 3B

Figure 3C

[0012]

Figure 4

Figure 5

[0013]

Figure 6

[0014]

Figure 7A

[0015]

Figure 7B

[0016]

Figure 8

DETAILED DESCRIPTION OF THE INVENTION

[0017] In the following detailed description, reference is made to the accompanying drawings which form a part hereof and which illustrate specific aspects or examples. Without departing from the present disclosure, these aspects may be combined, other aspects may be utilized, or structural changes may be made. These aspects may be implemented as a method, system, or apparatus. Accordingly, these aspects may take the form of a hardware implementation, a complete software implementation, or an implementation combining software and hardware aspects. Therefore, the following detailed description should not be understood as limiting, and the scope of the present disclosure is limited by the appended claims and their equivalents.

[0018] According to an example of the present disclosure, a video effect synchronization system enables a user to apply video effects that respond to audio to one or more additional points within a video clip. For example, the user may select a video effect from a video effect library and add an animation to the video clip. The video effect may be defined by one or more video effect parameters associated with the animation. As an example, the video effect parameters may include one or more animated objects added to the video clip, one or more additional points for each animated object within the video clip, and one or more animation effects applied to each animated object, but are not limited thereto. It should be understood that the animated object may include a plurality of visual elements.

[0019] In an exemplary aspect, the one or more additional points may be one or more body joints of a target object that appears in and / or is tracked within the video clip. For this purpose, body joint identification may be performed to separate or identify a body joint list from one or more target subjects of the video clip. As an example, the additional points of the animated object may be determined based on the musical characteristics of the audio music within the video clip. Alternatively, the additional points may be preselected by a selected video effect or may be predefined.

[0020] Additionally, the one or more animation effects applied to the animated object may be determined based on the musical characteristics of the audio music (e.g., beat information and / or frequency information). For this purpose, a musical characteristic evaluation may be performed to determine the musical characteristics of the audio music. Alternatively, when the audio music is selected by the user from a music library, the musical characteristics may be embedded in the audio music as metadata. As an example, the animation effect may include a glow effect, and the color of the glow around the animated object may change based on the corresponding beat characteristics or spectrum of the audio music. Accordingly, the one or more video effect parameters may be updated periodically (e.g., at each beat) based on the audio music and the video clip. Video effect synchronization enables adding an animation that reacts to the music beats of the audio music to one or more body joints of the target subject within the video clip.

[0021] FIG. 1 shows a video effect synchronization system 100 for rendering one or more video effects according to an example of the present disclosure. For example, a user 102 may generate, receive, obtain, or otherwise acquire a video clip 108. Thereafter, the user may select an audio music 110 to be added to the video clip 108. The video effect synchronization system 100 enables the user 102 to create a video effect that reacts to the audio and is added to one or more target body joints of a target subject within the video clip 108 based on the music 110. For this purpose, the video effect synchronization system 100 includes a computing device 104 associated with the user 102 and a server 106 communicatively coupled to the computing device 104 via a network 114. The network 114 may include any type of computing network including, but not limited to, a wired or wireless local area network (LAN), a wired or wireless wide area network (WAN), and / or the Internet.

[0022] In one example, user 102 may obtain video clip 108 and music 110 using computing device 104. User 102 may generate video clip 108 using a camera communicatively coupled to computing device 104. In such an example, the video effect may be synchronized with music 110 in real time or near real time to enable user 102 to view the video effect around the one or more body joints on a display (e.g., display 705) when the user is shooting a video on computing device 104. Alternatively or additionally, user 102 may receive, obtain, or otherwise acquire video clip 108 on computing device 104. In some examples, user 102 may edit video clip 108 based on music 110 to add video effects. In some aspects, user 102 may use computing device 104 to transmit video clip 108 and music 110 to server 106 via network 114. Computing device 104 may be either a portable computing device or a non-portable computing device. For example, computing device 104 may be a smartphone, a notebook computer, a desktop computer, or a server. Video clip 108 may be obtained in any format and may be in a compressed and / or decompressed form.

[0023] Computing device 104 is configured to analyze each frame of video clip 108 to identify the body joints of one or more target subjects within the frame. For example, a body joint algorithm may define a list of body joints identified and extracted from video clip 108. The body joints may include, but are not limited to, the head, neck, pelvis, spine, right / left shoulders, right / left upper arms, right / left forearms, right / left hands, right / left thighs, right / left feet, right / left toes.

[0024] Computing device 104 is configured to receive audio music 110 selected by user 102 from a music library for addition to video clip 108. Alternatively, in some embodiments, audio music 110 may be associated with a video effect. In such an embodiment, the video effect may include default music added to video clip 108. In some embodiments, audio music 110 may be extracted from video clip 108. Computing device 104 is configured to analyze the audio data to determine beat information or frequency spectrum information of audio music 110. For example, as described above, computing device 104 may determine the beat characteristics of each beat using an automatic beat tracking algorithm. It should be understood that in some embodiments, the music beat characteristic evaluation result may be embedded in the music as metadata. The music beat characteristic evaluation result may include the number and relative positions of the accented beats and unaccented beats of audio music 110. For example, if audio music 110 has a 4 / 4 beat structure, each section has four beats with different beat strengths, which are the strong beat, weak beat, medium-strong beat, and weak beat.

[0025] Alternatively or additionally, computing device 104 may determine a frequency spectrum characteristic evaluation result of the audio music. For example, computing device 104 may determine the average frequency spectrum of each beat of the audio music. It should be understood that in some embodiments, the frequency spectrum characteristics may be embedded in the music as metadata.

[0026] A video effect includes video effect parameters that control the behavior of one or more animated objects added to one or more attachment points of video data. In some embodiments, the video effect parameters may define, but are not limited to, an animated object, one or more attachment points of the animated object (e.g., one or more body joints of a target object), and an animation effect for the animated object added to one or more attachment points of the video clip. In an exemplary embodiment, the parameters of the video effect may be updated periodically (e.g., at each beat) based on the audio music and the video clip. In other words, video effect synchronization enables adding an animation that reacts to the music beat to a specific target object within the video clip.

[0027] In some embodiments, the user may add the animated object to one or more body joints, along with a specific animation effect defined by the video effect parameters, by selecting the video effect applied to the video clip 108. The video effect parameters are set to control which animated object is added to which addition point within the video clip and / or which animation effect is applied to the animated object. In other words, the video effect parameters define one or more body joints to which the animated object is added, how one or more body joints are selected for video effect application throughout the video clip, and one or more animation effects applied to the animated object. For example, the video effect may be randomly applied to a specific set of body joints throughout the video clip. Alternatively, the video effect may be applied to the video clip in a specific sequence (e.g., from head to toes). Alternatively, the animation effect may be applied to a specific body joint based on the beat strength of the music. For example, if the audio music has a 4 / 4 beat structure, the pelvis may be assigned to the strong beat (e.g., FIG. 3A), the left and right thighs may be assigned to the weak beats (e.g., FIG. 3B), and the left and right feet may be assigned to the medium-strong beats (e.g., FIG. 3C). Also, the animation effect may be determined based on the frequency spectrum range. For example, the pelvis may be assigned to the high spectrum range (e.g., 4 kHz to 20 Hz), the left and right thighs may be assigned to the medium spectrum range (e.g., 500 Hz to 4 kHz), and the left and right feet may be assigned to the low spectrum range (e.g., 20 Hz to 500 Hz). In other words, the beat or spectrum of the audio music may control the position where the animation is added to the video clip.

[0028] Additionally, the musical characteristics of the audio music may control one or more animation effects applied to the animated object. For example, the animation effect may include a lighting effect, in which case the video effect parameters may control the color and / or intensity of the light emitted from the animated object. Accordingly, the computing device 104 may determine the color and / or intensity of the light emitted by the animation effect based on the beat characteristics. For example, when the audio music has a 4 / 4 beat structure, high light intensity may be assigned to the strong beats, low light intensity may be assigned to the weak beats, and medium light intensity may be assigned to the medium-strong beats.

[0029] Alternatively, the color and / or intensity of the light emitted by the animation effect may be determined based on the frequency spectrum range. For example, high light intensity may be assigned to the high spectrum range (e.g., 4 kHz to 20 Hz), medium light intensity may be assigned to the medium spectrum range (e.g., 500 Hz to 4 kHz), and low light intensity may be assigned to the low spectrum range (e.g., 20 Hz to 500 Hz).

[0030] Additionally or alternatively, the animation speed of the animation effect may be controlled based on the beat characteristics or the frequency spectrum range. For example, a fast animation speed may be assigned to the strong beats and / or the high spectrum range, a medium animation speed may be assigned to the medium-strong beats and / or the medium spectrum range, and a low animation speed may be assigned to the weak beats and / or the low spectrum range.

[0031] Once ready to add the video effect to the video clip, the computing device 104 may modify the animation sequence to blend the 2D texture of the animated object onto the 3D mesh around the attachment point. By overlaying and blending the 2D animated object on the 3D mesh, a 3D-style animation effect can be created. Thereafter, the computing device 104 may synchronize the video effect with the music beats of the audio music and generate a rendered video with the video effect that can be presented to the user on a display (e.g., display 705) communicatively coupled to the computing device 104. It should be understood that the video effect may be synchronized with the music beats in real-time or near real-time to enable the user 102 to view the video effect around one or more body joints on the display while the user is shooting the video. Alternatively or additionally, the video effect may be synchronized with the music beats by the server 106. In such a manner, when the video clip 108 is uploaded to the server 106, the video effect may be applied to the video clip 108 to render the video effect.

[0032] Referring now to FIG. 2, a computing device 202 according to an example of the present disclosure will be described. The computing device 202 may be the same as or similar to the computing device 104 described above with reference to FIG. 1. The computing device 202 may include a communication interface 204, a processor 206, and a computer-readable storage 208. In one example, the communication interface 204 may be coupled to a network and may receive the video clip 108 and the audio music 110 (FIG. 1). The video clip 108 (FIG. 1) may be stored as video frames 246, and the music 110 may be stored as audio data 248.

[0033] In some examples, one or more video effects may be received at the communication interface 204 and stored as video effect data 252. The video effect data 252 may include one or more video effect parameters associated with the video effect. The video effect parameters may define, but are not limited to, one or more animated objects added to the video clip, one or more addition points where the animated objects are added within the video clip, and one or more animation effects applied to each animated object.

[0034] In one example, one or more applications 210 may be provided by the computing device 104. The one or more applications 210 may include a video processing module 212, an audio processing module 214, a video effect module 216, and a shader 218. The video processing module 212 may include a video acquisition manager 224 and a body joint identifier 226. The video acquisition manager 224 is configured to receive, acquire, or otherwise obtain video data including one or more video frames. Additionally, the body joint identifier 226 is configured to identify one or more body joints of one or more target subjects within the frame. In an exemplary aspect, the target subject is a human. For example, a body segmentation algorithm may define a list of body joints identified and extracted from the video clip 108. The body joints may include, but are not limited to, the head, neck, pelvis, spine, right / left shoulder, right / left upper arm, right / left forearm, right / left hand, right / left thigh, right / left leg, right / left foot, and right / left toes. In some examples, the body joint list may be received at the communication interface 204 and stored as body joints 250. In some aspects, the body joint list may be received from a server (e.g., 106).

[0035] Additionally, the audio processing module 214 may include an audio collection manager 232 and an audio analysis unit 234. The audio acquisition manager 232 is configured to receive, acquire, or obtain audio data by other means. The audio analysis unit 234 is configured to determine the audio information of the audio data. For example, the audio information may include, but is not limited to, the beat information and spectral information of each beat of the audio data. As an example, an automatic beat tracking algorithm may be used to determine the beat information. In some embodiments, the beat information may be embedded in the audio data as metadata. In other embodiments, the beat information may be received at the communication interface 204 and stored as audio data 248. The beat information provides the beat characteristics of each beat. The beat characteristics include, but are not limited to, the beat structure, the repeating sequence of strong beats and weak beats, the number of accented beats and unaccented beats, and the relative positions of the accented beats and unaccented beats. For example, if the audio music has a 4 / 4 beat structure, each section has four beats with different beat strengths: strong beat, weak beat, medium-strong beat, and weak beat. Additionally, the frequency spectrum information may be extracted from the audio data at predetermined time intervals (e.g., each beat). In some embodiments, the frequency spectrum information may be embedded in the audio data as metadata. In other embodiments, the frequency spectrum information may be received at the communication interface 204 and stored as audio data 248.

[0036] Furthermore, the video effect module 216 may further include an animation determination unit 238 and a video effect synchronization unit 240. The animation determination unit 238 determines a video effect to be applied to video data based on audio data. Specifically, the animation determination unit 238 is configured to determine one or more video effect parameters. For example, the animation determination unit 238 is configured to determine one or more animated objects to be added to the video clip and one or more additional points for each animated object. It should be understood that the animated object may include a plurality of visual elements and may have a plurality of additional points defined within the video clip. In an exemplary embodiment, the target additional point is a body joint of a target subject that appears in and / or is tracked within the video clip.

[0037] Furthermore, as will be further described below, the animation determination unit 238 is further configured to determine one or more animation effects to be applied to each animated object. In an exemplary embodiment, the animation effect may change for each beat. In other words, the beats of the selected audio music may control the visual changes of the video effects on one or more body joints. For example, the animation effect may include a lighting effect, in which case the animation determination unit 238 may determine the color and / or intensity of the light emitted from the animated object. In some embodiments, the color and / or intensity of the light may depend on the audio music added to the video clip.

[0038] The animation determination unit 238 may further determine one or more additional points (e.g., body joints) in the video clip that are suitable for the addition of animated objects. In some embodiments, the animation determination unit 238 may determine that the video effect is randomly applied to a specific set of body joints throughout the video clip. Alternatively, the animation determination unit 238 may determine that the video effect is applied to the video clip in a specific sequence (e.g., from top to bottom). Alternatively, the animation determination unit 238 may determine that the video effect is applied to a specific body joint based on the beat strength of the music. For example, if the audio music has a 4 / 4 beat structure, the pelvis may be assigned to the strong beat (e.g., Figure 3A), the left and right thighs may be assigned to the weak beats (e.g., Figure 3B), and the left and right feet may be assigned to the medium-strong beats (e.g., Figure 3C). Alternatively, the additional points may be determined based on the frequency spectrum range. For example, the pelvis may be assigned to the high spectrum range (e.g., 4 kHz to 20 Hz), the left and right thighs may be assigned to the medium spectrum range (e.g., 500 Hz to 4 kHz), and the left and right feet may be assigned to the low spectrum range (e.g., 20 Hz to 500 Hz). In other words, the beat or spectrum of the audio music may control which specific body joints the animation is added to in the video clip.

[0039] The video effect synchronization unit 240 is configured to synchronize the video effect with the music beats of the selected audio music to generate a rendered video with the video effect. In some embodiments, the video effect synchronization unit 240 is configured to modify the animation sequence to blend the 2D texture of the animated object within the 3D mesh around the additional point. It should be understood that a 3D-style animation effect can be created by overlaying and blending the 2D animated object on the 3D mesh.

[0040] The video effect synchronization unit 240 includes, or communicates with, the shader 218. The shader 218 is configured to receive video effect parameters. Based on the video effect parameters, the shader 218 is configured to generate an effect or cause an effect to be rendered. For example, the shader 218 may vary visual effects associated with video effects including, but not limited to, generation of blur, light bloom (e.g., glow), illumination (light bloom, e.g., shadow, highlight, and translucency), bump mapping, and distortion.

[0041] Figures 3A - 3C are diagrams showing examples of video frames 310, 320, 330 of a video clip having video effect synchronization according to an example of the present disclosure. In an exemplary example, an animation 302 (e.g., an animated object) added at different attachment points 304 is shown.

[0042] In one example, when audio data added to a video clip is received, audio information (e.g., beat features and / or frequency spectrum) of the audio data may be determined at predetermined intervals (e.g., music beats). Based on the audio information, one or more attachment points 304 for the animation 302 may be determined. For example, if the audio music has a 4 / 4 beat structure, the pelvis may be assigned to the strong beat (e.g., Figure 3A), the left and right thighs may be assigned to the weak beats (e.g., Figure 3B), and the left and right feet may be assigned to the medium - strong beats (e.g., Figure 3C). In such an embodiment, the animation 302 is added to the pelvis at the strong beat as shown in Figure 3A, added to the left and right thighs at the weak beats as shown in Figure 3B, and added to the left and right feet at the medium - strong beats as shown in Figure 3C.

[0043] In addition, the animation effect may be determined based on the frequency spectrum range. For example, the pelvis may be assigned to a high spectrum range (e.g., 4 kHz to 20 Hz), the left and right thighs may be assigned to a medium spectrum range (e.g., 500 Hz to 4 kHz), and the left and right feet may be assigned to a low spectrum range (e.g., 20 Hz to 500 Hz). In other words, the beats or spectrum of the audio music control at which specific body joints the animation is added to the video clip.

[0044] Referring now to FIG. 4, a simplified method for rendering one or more video effects onto video data based on audio data, according to an example of the present disclosure, will be described. A general order of steps of method 400 is shown in FIG. 4. Generally, method 400 begins at 402 and ends at 460. Method 400 can include more steps or fewer steps, or can be configured with an order of steps different from the steps shown in FIG. 4. Method 400 is executed as a set of computer-executable instructions by a computer system and can be encoded or stored on a computer-readable medium. In an exemplary aspect, method 400 is executed by a computing device associated with a user (e.g., 102). However, it should be understood that aspects of method 400 may be executed by one or more processing devices, such as a computer or a server (e.g., 104, 106). Further, method 400 can be executed by gates or circuits associated with a processor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a system on chip (SOC), a neural processing unit, or other hardware device. In the following, method 400 should be described with reference to the system, components, modules, software, data structures, user interfaces, etc. described in connection with FIGS. 1 and 8.

[0045] Method 400 begins at 402, and the flow may proceed to process 404. At 404, the computing device receives video data (e.g., video clip 108) including one or more video frames. For example, user 102 may generate, receive, obtain, or otherwise acquire video clip 108 via the computing device. At 408, the computing device processes each frame of the video data to identify the body joints of one or more target subjects within the frame. For example, a body joint algorithm may define a list of body joints identified and extracted from video clip 108. Body joints may include, but are not limited to, the head, neck, pelvis, spine, right / left shoulders, right / left upper arms, right / left forearms, right / left hands, right / left thighs, right / left legs, right / left feet, and right / left toes.

[0046] Referring back to start 402, method 400 may proceed to 412. It should be understood that the computing device may perform operations 404 and 412 simultaneously. Alternatively, operation 412 may be performed after operation 404. In some aspects, operation 404 may be performed after operation 412.

[0047] At 412, the computing device receives audio data (e.g., audio music 110) that is added to the video data selected by user 102. Next, at 416, the computing device analyzes the audio data to determine the audio information of the audio music 110. For example, the audio information includes the beat characteristics and / or frequency spectrum of each beat. In some embodiments, the computing device may determine the beat characteristics of each beat by an automatic beat tracking algorithm. The beat characteristics include, but are not limited to, the beat structure, the repeating sequence of strong beats and weak beats, the number of accented beats and unaccented beats, and the relative positions of the accented beats and unaccented beats. For example, if the audio music 110 has a 4 / 4 beat structure, each section has four beats with different beat strengths that are strong beat, weak beat, medium-strong beat, and weak beat. In other embodiments, the computing device may determine the frequency of each beat so as to associate it with a specific frequency range (e.g., high range, middle range, and low range).

[0048] When video data and audio data are received and analyzed in operations 404 - 416, method 400 proceeds to 420. At 420, the computing device determines video effects to be added to the video data. For example, the user may select a video effect from a video effect library to add animation to the video clip. The video effect is defined by video effect parameters that control which animated objects are added at which additional points within the video clip. It should be understood that the animated objects may include multiple visual elements and may have multiple additional points within the video clip. Additionally, as further described below, the video effect parameters further control one or more animation effects applied to the animated objects.

[0049] At 424, the computing device determines one or more additional points for the animated object to be added to the video clip based on the voice data analysis performed at operation 416. As described above, the animation addition points are the body joints of the target subject that appear in and / or are tracked in the video clip. For example, at 428, the computing device determines a specific body joint as an animation addition point based on the beat characteristics. For example, if the voice music has a 4 / 4 beat structure, the pelvis may be assigned to the strong beat (e.g., Figure 3A), the left and right thighs may be assigned to the weak beats (e.g., Figure 3B), and the left and right feet may be assigned to the medium-strong beats (e.g., Figure 3C). Alternatively, the animation addition points may be determined based on the frequency spectrum range. For example, the pelvis may be assigned to the high spectrum range (e.g., 4 kHz to 20 Hz), the left and right thighs may be assigned to the medium spectrum range (e.g., 500 Hz to 4 kHz), and the left and right feet may be assigned to the low spectrum range (e.g., 20 Hz to 500 Hz). In other words, the beat or spectrum of the voice music controls at which specific body joints the animation is added to the video clip.

[0050] Additionally, at 432, the computing device generates a three-dimensional (3D) mesh around each of the one or more addition points of the animated object in the video data. As will be further described below, the 3D mesh is used to add the animated object (e.g., a two-dimensional (2D) animated object) to the corresponding addition point. Thereafter, as shown by the alphanumeric A in Figures 4 and 5, method 400 proceeds to 436 in Figure 5.

[0051] At 436, the computing device determines one or more animation effects to be applied to the animated object based on the voice data analysis performed at operation 416. For example, the animation effect may include a lighting effect, in which case the video effect parameters may control the color and / or intensity of the light emitted from the animated object. Thus, at 440, the computing device may determine the color and / or intensity of the light of the animation effect based on the beat characteristics. For example, if the audio music has a 4 / 4 beat structure, a high light intensity may be assigned to the strong beats, a low light intensity may be assigned to the weak beats, and a medium light intensity may be assigned to the medium-strong beats.

[0052] Alternatively, the color and / or intensity of the light of the animation effect may be determined based on the frequency spectrum range. For example, a high light intensity may be assigned to the high spectrum range (e.g., 4 kHz to 20 Hz), a medium light intensity may be assigned to the medium spectrum range (e.g., 500 Hz to 4 kHz), and a low light intensity may be assigned to the low spectrum range (e.g., 20 Hz to 500 Hz).

[0053] Additionally or alternatively, the animation speed of the animation effect may be controlled based on the beat characteristics and / or the frequency spectrum range. For example, a fast animation speed may be assigned to the strong beats and / or the high spectrum range, a medium animation speed may be assigned to the medium-strong beats and / or the medium spectrum range, and a low animation speed may be assigned to the weak beats and / or the low spectrum range.

[0054] Once the animation object and the corresponding (plural) animation effects are determined, method 400 proceeds to operation 444. In 444, the computing device adds the animation object, along with its corresponding (plural) animation effects, to one or more corresponding additional points within the video frame. In other words, the animation is added to one or more target body joints throughout the video clip based on the audio data.

[0055] Subsequently, or simultaneously, in 448, the computing device modifies the animation sequence to blend the 2D texture of the animation object within the 3D mesh generated around the additional points. By overlaying and blending the 2D animation object on the 3D mesh, a more 3D-like animation effect can be created.

[0056] Thereafter, in 452, the video effect is synchronized with the music beats or spectrum of the selected audio music to generate a rendered video with the video effect. In 456, the computing device presents the rendered video with the video effect to the user on a display (e.g., display 705). It should be understood that the video effect may be synchronized with the music beats in real-time or near real-time to enable user 102 to view the video effect around one or more body joints on the display (e.g., display 705) while the user is filming the video. The method may end at 460.

[0057] Method 400 is described as being performed by a computing device associated with a user, but it should be understood that one or more operations of method 400 may be performed by any computing device or server (e.g., server 106). For example, synchronization of video effects and music beats may be performed by a server that receives music and video clips from a computing device associated with a user.

[0058] FIG. 6 is a block diagram showing physical components (e.g., hardware) of a computing device 600 that may be utilized to implement aspects of the present disclosure. The computing device components described below may be suitable for the computing devices described above. For example, computing device 600 may correspond to computing device 104 of FIG. 1. In a basic configuration, computing device 600 may include at least one processing unit 602 and a system memory 604. Depending on the configuration and type of the computing device, system memory 604 may include volatile storage (e.g., random access memory), non-volatile storage (e.g., read only memory), flash memory, or any combination of such memories, but is not limited thereto.

[0059] The system memory 604 may include an operating system 605 and one or more program modules 606 suitable for executing the various aspects disclosed herein. For example, the operating system 605 may be suitable for controlling the operation of the computing device 600. Further, aspects of the present disclosure may be executed in relation to a graphics library, other operating systems, or any other application program, but are not limited to any particular application or system. This basic configuration is shown in FIG. 6 by those components within the dashed line 608. The computing device 600 may have additional features or functionality. For example, the computing device 600 may further include additional data storage devices (removable devices and / or non-removable devices) such as magnetic disks, optical disks, or tapes. Such additional storage is shown in FIG. 6 by a removable storage device 609 and a non-removable storage device 610.

[0060] As described above, a plurality of program modules and data files may be stored in the system memory 604. The program modules 606 may execute processes including, but not limited to, one or more aspects as described herein when executed on at least one processing unit 602. The application 620 includes a video processing module 623, an audio processing module 624, a video effect module 625, and a shader module 626, as will be described in more detail with reference to FIG. 1. Other program modules that may be used in accordance with aspects of the present disclosure may include, for example, an email and contact application, a word processing application, a spreadsheet application, a database application, a slide presentation application, a drafting program, or a computer-aided application program, and / or one or more components supported by the systems described herein.

[0061] Furthermore, aspects of the present disclosure may be implemented within an electrical circuit including discrete electronic components, a packaged chip or integrated electronic chip including logic gates, a circuit utilizing a microprocessor, or a single chip including electronic components or a microprocessor. For example, aspects of the present disclosure may be implemented via a system-on-chip (SOC) in which each component or components shown in FIG. 6 can be integrated onto a single integrated circuit. Such an SOC device may include one or more processing units, a graphics unit, a communication unit, a system virtualization unit, and various application functions, all of which are integrated (or "burned") onto a chip substrate as a single integrated circuit. When operating via an SOC, the functions described herein regarding the client's ability to switch protocols may operate via application-specific logic integrated with other components of the computing device 600 on a single integrated circuit (chip). Aspects of the present disclosure may also be implemented using other technologies including, but not limited to, mechanical, optical, fluidic, and quantum technologies that can perform logical operations such as AND, OR, and NOT. Furthermore, aspects of the present disclosure may be implemented within a general-purpose computer or within any other circuit or system.

[0062] The computing device 600 may also have one or more input devices 612, such as a keyboard, a mouse, a pen, an acoustic or voice input device, a touch or swipe input device, etc. (Multiple) output devices 614A such as a display, a speaker, a printer, etc. may also be included. An output 614B corresponding to a virtual display may also be included. The above devices are examples, and other devices may be used. The computing device 600 may include one or more communication connections 616 that enable communication with other computing devices 450. Examples of suitable communication connections 616 include, but are not limited to, radio frequency (RF) transmitter, receiver, and / or transceiver circuits, universal serial bus (USB), parallel port, and / or serial port.

[0063] As used herein, the term computer-readable medium can include computer storage media. Computer storage media can include volatile and nonvolatile removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, or program modules. System memory 604, removable storage device 609, and non-removable storage device 610 are each an example of computer storage media (e.g., memory storage). Computer storage media can include RAM, ROM, electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other product used to store information and accessible by computing device 600. Any such computer storage media can be part of computing device 600. Computer storage media does not include a carrier wave or other propagated or modulated data signal.

[0064] Communication media can embody computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave or other transport mechanism, and can include any information delivery media. The term "modulated data signal" can represent a signal having one or more characteristics set or changed to encode information in the signal. By way of example and not limitation, communication media can include wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, radio frequency (RF), infrared, and other wireless media.

[0065] Figures 7A and 7B illustrate a computing device or mobile computing device 700 such as a mobile phone, smartphone, wearable computer (such as a smartwatch), tablet computer, laptop computer, smart home appliance, etc. that can be used to implement aspects of the present disclosure. Referring to Figure 7A, one aspect of a mobile computing device 700 for implementing these aspects is shown. In a basic configuration, the mobile computing device 700 is a handheld computer having both input and output elements. The mobile computing device 700 typically includes a display 705 and one or more input buttons 709 / 710 that enable a user to input information into the mobile computing device 700. The display 705 of the mobile computing device 700 may also function as an input device (e.g., a touch screen display). An optional secondary input element 715 (if included) further enables another user input. The secondary input element 715 may be a rotary switch, button, or any other type of manual input element. In an alternative aspect, the mobile computing device 700 may incorporate more or fewer input elements. For example, in some aspects, the display 705 may not be a touch screen. In yet another alternative aspect, the mobile computing device 700 is a mobile phone system such as a cellular phone. The mobile computing device 700 may further include an optional keypad 735. The optional keypad 735 may be a physical keypad or a "soft" keypad generated on a touch screen display. In various aspects, the output elements include a display 705 for displaying a graphical user interface (GUI), visual indicators 731 (e.g., light emitting diodes), and / or an audio transducer 725 (e.g., a speaker). In some aspects, the mobile computing device 700 incorporates a vibration transducer for providing haptic feedback to the user.In yet another aspect, the mobile computing device 700 incorporates input and / or output ports 730, such as a voice input (e.g., microphone jack), a voice output (e.g., headphone jack), and a video output (e.g., HDMI (registered trademark) port), to transmit and receive signals with an external source.

[0066] FIG. 7B is a block diagram showing an architecture of one aspect of a computing device, a server, or a mobile computing device. That is, the mobile computing device 700 can implement several aspects by incorporating a system (e.g., architecture) 602. The system 702 may be implemented as a "smartphone" that can execute one or more applications (e.g., browser, email, calendar, contact manager, messaging client, game, and media client / player). In some aspects, the system 702 is integrated as a computing device such as an integrated personal digital assistant (PDA), a wireless phone, etc.

[0067] One or more application programs 766 may be loaded into the memory 762 and executed on or in relation to the operating system 764. Examples of application programs include a telephone dialing program, an email program, a personal information management (PIM) program, a word processing program, a spreadsheet program, an Internet browser program, a messaging program, and / or one or more components supported by the systems described herein. The system 702 also includes a non-volatile memory area 768 within the memory 762. The non-volatile memory area 768 may be used to store persistent information that must not be lost when the power to the system 702 is turned off. The application 766 may use the information within the non-volatile memory area 768, such as emails or other messages used by an email or email application, and store it in the non-volatile memory area 1168. A synchronization application (not shown) also resides on the system 702 and is programmed to communicate with a corresponding synchronization application residing on the host computer to maintain synchronization between the information stored in the non-volatile memory area 768 and the corresponding information stored on the host computer. It should be understood that other applications may be loaded into the memory 762 and operate on the mobile computing device 700 described herein (e.g., video processing module 623, audio processing module 624, video effect module 625, and shader module 626).

[0068] The system 702 has a power source 770 that may be implemented as one or more batteries. The power source 770 may further include an external power source such as an AC adapter or a charging stand with a power source for charging or recharging the battery.

[0069] System 702 can further include a wireless interface layer 772 that performs the function of transmitting and receiving radio frequency communications. The wireless interface layer 772 facilitates a wireless connection between System 702 and the "outside world" via a communication carrier or service provider. Transmissions to and from the wireless interface layer 772 are performed under the control of the operating system 764. In other words, communications received by the wireless interface layer 772 can be distributed to the application program 766 via the operating system 764, and vice versa.

[0070] The visual indicator 720 can be used to provide visual notifications and / or the audio interface 774 may be used to generate audible notifications via the audio transducer 725. In the illustrated configuration, the visual indicator 720 is a light-emitting diode (LED) and the audio transducer 725 is a speaker. These devices are directly coupled to the power supply 770 so that when activated, they can remain on for a duration specified by the notification mechanism even if the processor 760 / 761 and other components are turned off to conserve battery power. The LED may be programmed to continue to light indefinitely until the user performs an action indicating that the device is powered on. The audio interface 774 is used to provide audible signals to and receive audible signals from the user. For example, in addition to being coupled to the audio transducer 725, the audio interface 774 may be coupled to a microphone to receive audible input, for example, to facilitate a telephone conversation. According to aspects of the present disclosure, the microphone may function as an audio sensor to facilitate control of notifications, as described below. System 702 may further include a video interface 776 that enables recording of still images, video streams, etc. by operation of an on-board camera.

[0071] The mobile computing device 700 that implements the system 702 may have additional features or functions. For example, the mobile computing device 700 may further include additional data storage devices (removable devices and / or non-removable devices) such as magnetic disks, optical disks, or tapes. Such additional storage is represented by the non-volatile memory area 768 in FIG. 7B.

[0072] As described above, the data / information generated or captured by the mobile computing device 700 and stored via the system 702 may be stored locally on the mobile computing device 700, or the data may be stored on any number of storage media accessible from the device via the wireless interface layer 772 or via a wired connection between the mobile computing device 700 and another computing device associated with the mobile computing device 700 (e.g., a server computer in a distributed computing network such as the Internet). It should be understood that such data / information may be accessed via the mobile computing device 700, via the wireless interface layer 772, or via a distributed computing network. Similarly, such data / information may be easily transmitted between computing devices for storage and use according to known data / information transmission and storage means including electronic mail and collaborative data / information sharing systems.

[0073] FIG. 8 illustrates one aspect of the architecture of a system for processing data received in a computing system from a remote source such as a personal computer 804, a tablet computing device 806, or a mobile computing device 808 as described above. The content displayed in server device 802 may be stored in different communication channels or other storage types. For example, computing devices 804, 806, 808 may represent the computing device 104 of FIG. 1, and server device 802 may represent the server 106 of FIG. 1.

[0074] In some aspects, one or more of video processing module 823, audio processing module 824, and video effect module 825 may be implemented by server device 802. Server device 802 may provide data to client computing devices such as personal computer 804, tablet computing device 806, and / or mobile computing device 808 (e.g., smartphone) via network 812 and may be provided data from the client computing devices. As an example, the computer system described above may be implemented within personal computer 804, tablet computing device 806, and / or mobile computing device 808 (e.g., smartphone). In addition to receiving graphic data that can be used for pre-processing in a system starting from graphics or post-processing in a receiving computing system, any of these aspects of the computing device may also obtain content from store 816. Content storage may include video data 818, audio data 820, and rendered video data 822.

[0075] FIG. 8 shows an example of a mobile computing device 808 that may execute one or more aspects disclosed herein. Further, the aspects and functions described herein may operate on a distributed system (e.g., a cloud-based computing system), where application functions, memory, data storage and retrieval, and various processing functions may operate remotely from each other on a distributed computing network (e.g., the Internet or an intranet). Various types of user interfaces and information may be displayed via an on-board computing device display or via a remote display unit associated with one or more computing devices. For example, various types of user interfaces and information can be displayed and interacted with on a wall surface that projects the various types of user interfaces and information. Interaction with multiple computing systems that can be utilized to implement aspects of the present invention includes keystroke input, touch screen input, voice or other audio input, gesture input, etc., and in the case of gesture input, the associated computing device includes detection (e.g., camera) functionality for capturing and interpreting a user's gesture to control the functions of the computing device.

[0076] The phrases "at least one", "one or more", "or", and "and / or" are conjunctive and disjunctive open-ended expressions in an operation. For example, each of the expressions "at least one of A, B, and C", "at least one of A, B, or C", "one or more of A, B, and C", "one or more of A, B, or C", "A, B, and / or C", and "A, B, or C" means only A, only B, only C, A and B, A and C, B and C, or A, B, and C.

[0077] The term "one" entity means one or more of such entities. Thus, the terms "one", "one or more", and "at least one" may be used interchangeably herein. It should also be noted that the terms "comprising", "including", and "having" may be used interchangeably.

[0078] As used herein, the term "automatically" and variations thereof refer to any process or operation that is performed without significant manual input when the process or operation is executed, typically continuously or semi - continuously. However, even if significant or insignificant manual input is used, if the input is received prior to the execution of the process or operation, the execution of the process or operation can be performed automatically. Manual input is considered significant if it affects the way the process or operation is executed. Manual input for the purpose of consenting to the execution of a process or operation is not considered "substantial".

[0079] Any of the steps, functions, and operations discussed herein may be performed continuously and automatically.

[0080] The exemplary systems and methods of the present disclosure have been described in relation to computing devices. However, some known structures and devices have been omitted from the foregoing description to avoid unnecessarily obscuring the present disclosure. This omission should not be construed as a limitation. Specific details are described to provide an understanding of the present disclosure. However, it should be understood that the present disclosure may be implemented in various ways in addition to the specific details described herein.

[0081] Furthermore, in the exemplary embodiments presented herein, while various components of the system are shown as being co-located, some components of the system may be located remotely at a distal portion of a distributed network such as a LAN and / or the Internet, or may be located within a dedicated system. Accordingly, it should be understood that the components of the system may be coupled to one or more devices, such as servers, communication devices, or may be co-located on a particular node of a distributed network such as an analog and / or digital electrical communication network, a packet switched network, or a circuit switched network. As will be appreciated from the foregoing description, for reasons of computing efficiency, the components of the system may be located anywhere within the distributed network of components without affecting the operation of the system.

[0082] Furthermore, it should be understood that the various links connecting the elements may be wired or wireless links, or any combination thereof, or any other known or later developed element capable of providing data to and / or communicating data from the connected elements. These wired or wireless links may also be secure links and may be capable of communicating encrypted information. For example, the transmission medium used as a link may be any suitable carrier of electrical signals including coaxial cables, copper wire, and fiber optic cables, and may take the form of acoustic or optical waves, such as those generated during radio frequency and infrared data communications.

[0083] Although flowcharts have been discussed and illustrated in relation to a particular sequence of events, it should be understood that changes, additions, and omissions may be made to the sequence without substantially affecting the operation of the disclosed configurations and embodiments.

[0084] Some modifications and variations of the present disclosure can be used. It is also possible to provide some features of the present disclosure and not others.

[0085] In yet another configuration, the systems and methods of the present disclosure can be implemented in combination with a dedicated computer, a programmed microprocessor or microcontroller and peripheral integrated circuit elements, an ASIC or other integrated circuit, a digital signal processor, a hardwired electronic circuit or logic circuit (e.g., discrete element circuitry), a programmable logic device or gate array (e.g., PLD, PLA, FPGA, PAL), a dedicated computer, any similar device, etc. Generally, any device or means capable of implementing the methods set forth herein can be used to implement various aspects of the present disclosure. Exemplary hardware that can be used in the present disclosure includes computers, handheld devices, telephones (e.g., cellular phones, Internet-enabled telephones, digital telephones, analog telephones, hybrid telephones, etc.), and other hardware known in the art. Some of these devices include a processor (e.g., single or multiple microprocessors), memory, non-volatile storage, input devices, and output devices. Additionally, alternative software implementations including, but not limited to, distributed processing or component / object distributed processing, parallel processing, or virtual machine processing may be constructed to implement the methods described herein.

[0086] In yet another configuration, the disclosed methods can be readily implemented in combination with software that provides portable source code that is usable on various computer or workstation platforms, an object or object-oriented software development environment. Alternatively, the disclosed systems may be implemented partially or fully in hardware using standard logic circuits or VLSI designs. Whether to use software or hardware to implement the systems according to the present disclosure depends on the speed and / or efficiency requirements of the system, specific functions, and the particular software or hardware system or microprocessor or microcomputer system being used.

[0087] In yet another configuration, the disclosed method may be partially realized by software stored on a storage medium and executable on a programmed general-purpose computer, a special-purpose computer, a microprocessor, etc. that cooperate with a controller and a memory. In these instances, the systems and methods of the present disclosure may be implemented as resources resident on a server or computer workstation, as well as routines incorporated into a dedicated measurement system, system components, etc., a program incorporated into a personal computer, e.g., an applet, JAVA (registered trademark), or CGI script. The system and / or method may also be realized by physically incorporating it into a software and / or hardware system.

[0088] The present disclosure is not limited to the described standards and protocols. Other similar standards and protocols not described herein already exist and are included in the present disclosure. Further, the standards and protocols described herein, and other similar standards and protocols not described herein, are periodically replaced by faster or more efficient equivalents having substantially the same functionality. Such alternative standards and protocols having the same functionality are considered equivalents included in the present disclosure.

[0089] In various configurations and aspects, the present disclosure includes components, methods, processes, systems, and / or apparatuses as illustrated and described herein, including their various combinations, sub-combinations, and subsets. Those skilled in the art, if they understand the present disclosure, will understand how to manufacture and use the systems and methods of the present disclosure. The present disclosure includes, in various configurations and aspects, the state where there are no items that may have been used in previous devices or processes, e.g., to improve performance, achieve ease of use, and / or reduce implementation costs, and there are no items not illustrated and / or described herein or in its various configurations or aspects, and provides devices and processes.

[0090] (A1) In one aspect, some examples include a method for rendering video effects on a display. The method includes obtaining video data and audio data, analyzing the video data to determine one or more additional points of a target object appearing in the video data, analyzing the audio data to determine audio features, determining animation-related video effects to be added to the one or more additional points based on the audio features, and generating a rendered video by applying the video effects to the video data.

[0091] (A2) In some examples of A1, determining the animation-related video effects to be added to the one or more additional points based on the audio features includes determining an animated object to be added to the one or more additional points, determining the one or more additional points in the video data to which the animated object is to be added, and determining one or more animation effects to be applied to the animated object based on the audio features.

[0092] (A3) In some examples of A1 - A2, the one or more additional points are selected from one or more body points of the target object appearing in the video data.

[0093] (A4) In some examples of A1 - A3, determining the one or more additional points in the video data includes determining one or more additional points in the video data based on the audio features.

[0094] (A5) In some examples of A1 - A4, the method further includes generating a three - dimensional (3D) mesh around the one or more attachment points and modifying an animation sequence to blend the animated object within a corresponding 3D mesh around the one or more attachment points to which the animated object is added.

[0095] (A6) In some examples of A1 - A5, the voice feature includes a beat feature of each beat or a frequency spectrum value for each beat.

[0096] (A7) In some examples of A1 - A6, obtaining the voice data includes selecting the voice data from a music library, and the voice feature is embedded in the voice data as metadata.

[0097] In yet another aspect, some examples include a computing system that includes one or more processors and a memory coupled to the one or more processors, the memory storing one or more instructions that, when executed by the one or more processors, cause the one or more processors to perform any of the methods described herein (e.g., A1 - A7 above).

[0098] In yet another aspect, some examples include a non - transitory computer - readable storage medium storing one or more programs for execution by one or more processors of a storage device, the one or more programs including instructions for performing any of the methods described herein (e.g., A1 - A7 above).

[0099] (B1) In one aspect, some examples include a computing device for rendering video effects on a display. The computing device includes a processor, and when executed by the processor, causes the computing device to: obtain video data and audio data; analyze the video data to determine one or more additional points of one or more target objects appearing in the video data; analyze the audio data to determine audio features; determine animation-related video effects to be added to the one or more additional points based on the audio features; and generate a rendered video by applying the video effects to the video data.

[0100] (B2) In some examples of B1, determining the animation-related video effects to be added to the one or more additional points based on the audio features includes: determining an animated object to be added to the one or more additional points; determining the one or more additional points in the video data to which the animated object is to be added; and determining one or more animation effects to be applied to the animated object based on the audio features.

[0101] (B3) In some examples of B1 - B2, the one or more additional points are selected from one or more body points of the target object appearing in the video data.

[0102] (B4) In some examples of B1 - B3, determining the one or more additional points in the video data includes determining the one or more additional points in the video data based on the audio features.

[0103] In some examples of (B5) B1 - B4, when the plurality of instructions are executed, they cause the computing device to generate a three - dimensional (3D) mesh around the one or more attachment points, and to modify an animation sequence to blend the animated object within the corresponding 3D mesh around the one or more attachment points to which the animated object is added.

[0104] In some examples of (B6) B1 - B5, the voice features include the beat features of each beat or the frequency spectrum values for each beat.

[0105] In some examples of (B7) B1 - B6, obtaining the voice data includes selecting the voice data from a music library, and the voice features are embedded in the voice data as metadata.

[0106] In one aspect, some examples include a non - transitory computer - readable medium storing instructions for rendering a video effect on a display. When the instructions are executed by one or more processors of a computing device, they cause the computing device to obtain video data and voice data, analyze the video data to determine one or more attachment points of a target object appearing in the video data, analyze the voice data to determine voice features, determine animation - related video effects to be added to the one or more attachment points based on the voice features, and generate a rendered video by applying the video effect to the video data.

[0107] (C2) In some examples of C1, determining the video effect related to animation added to the one or more additional points based on the voice feature includes determining an animated object added to the one or more additional points, and based on the voice feature, determining the one or more additional points in the video data to which the animated object is added, and one or more animation effects applied to the animated object.

[0108] (C3) In some examples of C1 - C2, the one or more additional points are selected from one or more body points of the target object appearing in the video data.

[0109] (C4) In some examples of C1 - C3, determining the one or more additional points in the video data includes determining one or more additional points in the video data based on the voice feature.

[0110] (C5) In some examples of C1 - C4, when the instruction is executed by the one or more processors, the computing device is further caused to generate a three - dimensional (3D) mesh around the one or more additional points, and modify the animation sequence to blend the animated object into the corresponding 3D mesh around the one or more additional points to which the animated object is added.

[0111] (C6) In some examples of C1 - C5, the voice feature includes a beat feature of each beat or a frequency spectrum value for each beat.

[0112] For example, aspects of the present disclosure have been described above with reference to block diagrams and / or operation descriptions of methods, systems, and computer program products according to aspects of the present disclosure. The functions / operations recited within a block can occur in an order different from that shown in any flowchart. For example, depending on the relevant functions / operations, two blocks shown in succession may actually be executed substantially simultaneously, or these blocks may sometimes be executed in the reverse order.

[0113] The description and illustration of one or more aspects provided in this application are not intended to limit or restrict the scope of the claimed present disclosure in any way. The aspects, examples, and details described herein are considered to be sufficient to convey ownership and to enable others to form and use the best mode of the claimed present disclosure. The claimed present disclosure should not be construed as limited to the aspects, examples, or details described herein. Various features (structural features and method features), whether illustrated or described in combination or individually, are intended to be selectively included or omitted in order to form embodiments having a particular set of features. By providing the description and illustration of this application, those skilled in the art may envision modifications, variations, and alternative aspects that are included within the broader spirit of the general inventive concept realized in this application without departing from the broader scope of the claimed disclosure.

Claims

1. A method for rendering video effects on a display, the method comprising: obtaining video data and audio data; analyzing the video data to determine one or more additional points of a target object appearing in the video data; analyzing the audio data to determine audio features, the audio features including beat features of each beat or frequency spectrum values for each beat; determining animation-related video effects to be added to the one or more additional points based on the audio features; generating a rendered video by applying the video effects to the target object in a specific sequence; A method comprising the above.

2. Determining the animation-related video effects to be added to the one or more additional points based on the audio features comprises: determining an animated object to be added to the one or more additional points; determining the one or more additional points in the video data to which the animated object is to be added; determining one or more animation effects to be applied to the animated object based on the audio features; The method according to claim 1, comprising the above.

3. The one or more additional points are selected from one or more body points of the target object appearing in the video data. The method according to claim 1.

4. Determining the one or more additional points in the video data comprises determining the one or more additional points in the video data based on the audio features. The method according to claim 2.

5. generating a three-dimensional (3D) mesh around the one or more additional points; modifying an animation sequence to blend the animated object into the corresponding 3D mesh around the one or more additional points to which the animated object is added; The method according to claim 2, further comprising the above.

6. Obtaining the audio data comprises selecting the audio data from a music library, and the audio features are embedded in the audio data as metadata. The method according to claim 1.

7. A computing device for rendering video effects on a display, the computing device comprising: a processor; when executed by the processor, cause the computing device to acquire video data and audio data; analyze the video data to determine one or more additional points of a target object appearing in the video data; analyze the audio data to determine audio features, the audio features including beat features of each beat or frequency spectrum values for each beat; determine animation-related video effects to be added to the one or more additional points based on the audio features; generate a rendered video by applying the video effects to the target object in a specific sequence; a memory storing a plurality of instructions for causing the above to be executed.

8. Determining the animation-related video effects to be added to the one or more additional points based on the audio features includes: determining an animated object to be added to the one or more additional points; determining the one or more additional points in the video data to which the animated object is to be added; determining one or more animation effects to be applied to the animated object based on the audio features. The computing device according to claim 7.

9. The one or more additional points are selected from one or more body points of the target object appearing in the video data. The computing device according to claim 7.

10. Determining the one or more additional points in the video data includes determining the one or more additional points in the video data based on the audio features. The computing device according to claim 8.

11. When executed, the plurality of instructions cause the computing device to generate a three-dimensional (3D) mesh around the one or more additional points. Modifying an animation sequence to blend the animation object within corresponding 3D meshes around the one or more attachment points to which the animation object is added; The computing device according to claim 8, further causing the computing device to perform the above operations. **Claim 12** Obtaining the audio data includes selecting the audio data from a music library, and the audio features are embedded in the audio data as metadata. The computing device according to claim 7. **Claim 13** A non-transitory computer-readable medium storing instructions for rendering video effects on a display, wherein when the instructions are executed by one or more processors of a computing device, the computing device is caused to: Obtain video data and audio data; Analyze the video data to determine one or more attachment points of a target object appearing in the video data; Analyze the audio data to determine audio features, where the audio features include beat features of each beat or frequency spectrum values for each beat; Based on the audio features, determine animation-related video effects to be added to the one or more attachment points; Generate a rendered video by applying the video effects to the target object in a specific sequence; A non-transitory computer-readable medium for causing the above operations to be performed. **Claim 14** Based on the audio features, determining the animation-related video effects to be added to the one or more attachment points includes: Determining an animation object to be added to the one or more attachment points; Determining the one or more attachment points in the video data to which the animation object is added; Based on the audio features, determining one or more animation effects to be applied to the animation object; The non-transitory computer-readable medium according to claim 13, including the above steps. **Claim 15** The one or more attachment points are selected from one or more body points of the target object appearing in the video data. The non-transitory computer-readable medium according to claim 13.

16. Determining the one or more additional points in the video data includes determining one or more additional points in the video data based on the audio feature, The non-transitory computer-readable medium according to claim 14.

17. When executed by the one or more processors, the instructions cause the computing device to generate a three-dimensional (3D) mesh around the one or more additional points; modify the animation sequence to blend the animated object within the corresponding 3D mesh around the one or more additional points to which the animated object is added; The non-transitory computer-readable medium according to claim 14, which further causes the above to be executed.

Citation Information

Patent Citations

  • Virtual model processing method and device, electronic equipment and storage medium

    CN112034984A

  • Image processing method, storage medium, and computer device

    US20200380031A1

  • Method and apparatus for an interactive user interface

    US20200401372A1