Method, computing device, and program for generating virtual clothing on a display

The integration of audio analysis into video editing technologies allows for the synchronization of virtual clothing with audio characteristics, enhancing user experience through interactive and vivid visual effects.

JP7695405B2Active Publication Date: 2025-06-18LEMON CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023573678
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-06-21
Filing Date
2022-05-10
Publication Date
2025-06-18
Estimated Expiration
2042-05-10

AI Technical Summary

Technical Problem

Existing video editing technologies do not effectively synchronize visual effects with audio data, leading to a suboptimal user experience.

Method used

A method and system for generating virtual clothing on a display by analyzing video and audio data to determine body joints, generating a mesh, analyzing audio characteristics, and using texture rendering information to render virtual clothing synchronized with audio characteristics.

Benefits of technology

The solution enhances user experience by creating interactive and vivid virtual clothing that responds to audio characteristics, such as tempo and frequency, in real-time or near real-time, thereby improving the synchronization of video effects with audio data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007695405000001
    Figure 0007695405000001
  • Figure 0007695405000002
    Figure 0007695405000002
  • Figure 0007695405000003
    Figure 0007695405000003
Patent Text Reader

Abstract

Systems and methods for generating virtual clothing on a display are described. Some examples may include obtaining video data and audio data, and analyzing the video data to determine one or more body joints of a target object appearing in the video data. A mesh based on the determined one or more body joints may be generated. The audio data may be analyzed to determine audio characteristics associated with the audio data. Texture rendering information associated with the virtual clothing may be determined based on the audio characteristics. A rendered image may be generated by rendering the virtual clothing onto the generated mesh using the texture rendering information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application claims the priority of a U.S. patent application filed on June 21, 2021 (Application No. 17 / 353,296, titled "RENDERING VIRTUAL ARTICLES OF CLOTHING BASED ON AUDIO CHARACTERISTICS"), which is incorporated herein by reference in its entirety.

Background Art

[0002] Video editing technology has become widespread and provides users with various video editing methods. For example, users can edit videos to add visual effects and / or music to the videos. However, in many video editing technologies, controlling visual effects based on audio data is not considered. Therefore, in order to improve the user experience, it is necessary to develop video editing technology for rendering the synchronization of video effects.

[0003] Aspects disclosed herein are described with respect to the above and other general considerations. Also, while relatively specific problems may be discussed, it should be understood that the examples should not be limited to solving the specific problems identified in the background section or elsewhere of the present disclosure.

Summary of the Invention

[0004] According to at least one embodiment of the present disclosure, a method for generating virtual clothing on a display will be described. The method may include obtaining video data and audio data, analyzing the video data to determine one or more body joints of a target object appearing in the video data, generating a mesh based on the determined one or more body joints, analyzing the audio data to determine audio characteristics, determining texture rendering information associated with the virtual clothing based on the audio characteristics, and generating a rendered video by using the texture rendering information to render the virtual clothing on the generated mesh.

[0005] According to an embodiment of the present disclosure, a computing device for generating virtual clothing on a display will be described. The computing device may include a processor and a memory storing a plurality of instructions, and when the plurality of instructions are executed by the processor, the computing device is caused to obtain video data and audio data, analyze the video data to determine one or more body joints of a target object appearing in the video data, generate a mesh based on the determined one or more body joints, analyze the audio data to determine audio characteristics, determine texture rendering information associated with the virtual clothing based on the audio characteristics, and generate a rendered video by using the texture rendering information to render the virtual clothing on the generated mesh.

[0006] According to an embodiment of the present disclosure, a non-transitory computer-readable medium storing instructions for generating virtual clothing on a display will be described. When the instructions are executed by one or more processors of a computing device, the computing device is caused to obtain video data and audio data, analyze the video data to determine one or more body joints of a target object appearing in the video data, generate a mesh based on the determined one or more body joints, analyze the audio data to determine audio characteristics, determine texture rendering information associated with the virtual clothing based on the audio characteristics, and generate a rendered video by rendering the virtual clothing on the generated mesh using the texture rendering information.

[0007] Any one or more of the above-described aspects can be combined with any other of the one or more aspects. All of the one or more aspects are described herein.

[0008] The summary of the present invention selects and briefly introduces a part of the concept, which will be further described in the mode for carrying out the invention. The summary of the present invention is not intended to identify important features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Additional aspects, features, and / or advantages of the embodiments are partially described in the following description, will be made somewhat apparent by the description, or will be understood by practicing the present disclosure.

Brief Description of the Drawings

[0009] With reference to the following figures, non-limiting and non-exhaustive examples will be described.

[0010]

Figure 1

[0011]

Figure 2

[0012]

Figure 3A

Figure 3B

[0013]

Figure 3C

Figure 3D

Figure 3E

[0014]

Figure 4

Figure 5

[0015]

Figure 6

[0016]

Figure 7A

[0017]

Figure 7B

[0018]

Figure 8

DETAILED DESCRIPTION OF THE INVENTION

[0019] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof, and in which are shown by way of illustration specific aspects and examples. These aspects may be combined, other aspects may be utilized, and structural changes may be made without departing from the scope of the present disclosure. The aspects may be implemented as a method, system, or apparatus. Thus, the aspects may take the form of an implementation in hardware, an all-software implementation, or an implementation combining software and hardware aspects. Accordingly, the following detailed description is not to be taken in a limiting sense, and the scope of the present disclosure is defined by the appended claims and their equivalents.

[0020] According to an embodiment of the present disclosure, a video effect synchronization system enables a user to interact with clothing by applying video effects responsive to audio to one or more virtual clothing items displayed on a video clip, making the clothing more vivid and interactive. For example, the user may select a virtual clothing item from a clothing library. The virtual clothing may be rendered onto a three-dimensional mesh generated based on a combination of body joints or otherwise worn. In some examples, the virtual clothing may be any object that can be rendered onto a three-dimensional mesh generated based on a combination of body joints or otherwise worn. In some examples, the clothing may include video effects set by one or more video effect parameters associated with the virtual clothing. As an example, the video effect parameters may include the transparency of the clothing, a sparkle effect associated with the virtual clothing (Sparkle The (Effect), color of the virtual clothing, and / or reflectance associated with the virtual clothing may be included, but are not limited thereto. The virtual clothing may be associated with a specific three-dimensional mesh (e.g., a torso mesh, a head mesh, a leg mesh), and thus, the generated mesh can be dressed or otherwise rendered with the virtual clothing. It should be understood that when the user shown in the video sequence moves, the virtual clothing may appear to move. In some examples, parts of the user's body such as the hands, head, feet, etc. may hide the virtual clothing. Thus, when the virtual clothing is worn on the mesh, a part of the body may appear in front of the virtual clothing.

[0021] In an exemplary aspect, a texture associated with the virtual clothing may be mapped to the three-dimensional mesh. The three-dimensional mesh is generated based on a combination of one or more body joints of the target object tracked within the video clip. Therefore, body joint identification may be performed so as to separate or identify a list of body joints from one or more targets of the video clip. The three-dimensional mesh may be generated based on one or more joints. The texture, color, color change rate, and physically based rendering (PBR) material parameters may be based on the music characteristics of the audio music.

[0022] For example, based on the music characteristics of the audio music (e.g., tempo information and / or frequency information), one or more characteristics of the texture may be changed by a shader configured to render the texture of the virtual clothing. As an example, to determine the music characteristics of the audio music, music characterization may be performed. Alternatively, when the audio music is selected by the user from a music library, the music characteristics may be embedded in the audio music as metadata. As an example, the texture of the virtual clothing may include a rendering effect located around the virtual clothing, and such an effect may change based on the corresponding tempo characteristics or frequency spectrum of the audio music. Accordingly, one or more video effect parameters may be updated periodically (e.g., every beat) based on the audio music and the video clip. With the synchronization of the video effects, the texture of the virtual clothing becomes responsive to the tempo of the music of the audio music. In some examples, the distance between one or more body parts (e.g., the user's hands) may also affect the texture. For example, the color of the texture rendered by the shader may be changed based on the distance between the hands shown in the video. Further, a second shader such as a sparkle shader may be applied to the virtual clothing and / or the second shader may be used to create another effect such as crystals, animations, or other graphics within the video.

[0023] FIG. 1 shows a virtual clothing rendering system 100 for rendering one or more video effects according to an embodiment of the present disclosure. For example, user 102 may generate, receive, obtain, or otherwise acquire video clip 108. The user may then select audio music 110 to add to video clip 108. In some examples, music 110 may be preset by system 100. With virtual clothing rendering system 100, user 102 can virtually try on virtual clothing. The virtual clothing may be worn on one or more three-dimensional meshes created from one or more target body joints of the target within video clip 108 based on music 110. To this end, virtual clothing rendering system 100 includes a computing device 104 associated with user 102 and a server 106 communicatively coupled to computing device 104 via network 114. Network 114 may include any type of computing network including, but not limited to, a wired or wireless local area network (LAN), a wired or wireless wide area network (WAN), and / or the Internet.

[0024] As an example, user 102 may utilize computing device 104 to obtain video clip 108 and music 110. User 102 may use a camera communicatively coupled to computing device 104 to generate video clip 108. In such an example, video effects such as the color, opacity, rotation, and a plurality of other graphics (such as, but not limited to, glow) of the virtual clothing may be synchronized with music 110 in real-time or near real-time as if user 102 is actually wearing the virtual clothing, and thus the change in the texture information of the virtual clothing may change according to the music. As an example, the user may view on a display (such as, display 705) a video with the virtual clothing added while the user is taking a video with computing device 104. Alternatively, or in addition, user 102 may receive, obtain, or otherwise acquire video clip 108 on computing device 104. In some examples, user 102 may edit video clip 108 and add video effects based on music 110. In some aspects, user 102 may utilize computing device 104 to transmit video clip 108 and music 110 to server 106 via network 114. Computing device 104 may be either a portable computing device or a non-portable computing device. For example, computing device 104 may be a smartphone, laptop, desktop, or server. Video clip 108 may be obtained in any format and may be in a compressed and / or decompressed form.

[0025] Computing device 104 is configured to analyze each frame of video clip 108 and identify the body joints of one or more objects within the frame. For example, a body joint algorithm may define a list of body joints to be identified and extracted from video clip 108. The body joints may include, but are not limited to, the head, neck, pelvis, spine, left and right shoulders, left and right upper arms, left and right forearms, left and right hands, left and right thighs, left and right legs, left and right feet, and left and right toes.

[0026] Computing device 104 is configured to receive audio music 110. Audio music 110 is selected by user 102 to be added to video clip 108 from a music library. Alternatively, in some embodiments, audio music 110 may be associated with a video effect. In such an embodiment, the default music added to video clip 108 may be included in the video effect. In some embodiments, audio music 110 may be extracted from video clip 108. Computing device 104 is configured to analyze the audio data to determine the beat information or frequency spectrum information of audio music 110. For example, as described above, computing device 104 may determine the beat characteristics of each beat through an automatic beat tracking algorithm. It should be understood that in some embodiments, the characterization of the beat of the music may be embedded in the music as metadata. The characterization of the beat of the music may include the number and relative positions of the accented and unaccented beats of audio music 110. For example, if audio music 110 has a 4 / 4 time signature, each section has four beats with different beat strengths: a strong beat, a weak beat, a second strong beat, and a weak beat.

[0027] Alternatively, or in addition, computing device 104 may determine a characterization of the frequency spectrum of the audio music. For example, computing device 104 may determine the average frequency spectrum of each beat of the audio music. It should be understood that in some aspects, the characterization of the frequency spectrum may be embedded in the music as metadata.

[0028] The video effects include video effect parameters that control how the virtual clothing looks or is rendered. In some aspects, the video effect parameters may control the PBR material associated with the texture rendered by the shader. In an exemplary aspect, the parameters of the video effect may be updated periodically (e.g., every beat) based on the audio music and the video clip. In some examples, one or more video effect parameters may be set based on the distance between two body parts of the user. In other words, with the synchronization of the video effects, the virtual clothing becomes responsive to the beats of the music in the video clip.

[0029] In some embodiments, the user may select virtual clothing and / or video effects to apply to video clip 108 based on video effect parameters. The video effect parameters are set to control which virtual clothing is applied to which 3D mesh around which of the user's multiple joints within the video clip and / or how the PBR material is affected by the music. In other words, the video effect parameters define one or more meshes based on multiple body joints, and the meshes can provide points or surfaces of attachment where the virtual clothing can be rendered. For example, video effects may be applied to the virtual clothing. Alternatively, the visual characteristics of the virtual clothing may be based on the intensity of the beats of the music. For example, if the audio music has a 4 / 4 time signature and is strong, one set of video effect parameters may be selected (e.g., FIG. 3A). If the audio music has a 4 / 4 time signature and is weak, another set of video effect parameters may be selected (e.g., FIG. 3B). If the audio music has a 4 / 4 time signature and is weak, a third set of video effect parameters may be selected (e.g., FIG. 3C). As another example, the video effect parameters for a 4 / 4 time signature may be different from those for a 3 / 4 time signature. That is, the time signature and the beats being tracked may be used to select the video effect parameters, and the video effect parameters may affect one or more aspects of the shader for rendering the virtual clothing on the display. Alternatively, or in addition, the video effect parameters may be determined based on a frequency spectrum range. For example, a first video effect parameter may be selected according to a first frequency spectrum, a second video effect parameter may be selected according to a second frequency spectrum, and a third video effect parameter may be selected according to a third frequency spectrum. As an example, different frequency spectrums may correspond to a high spectrum range (e.g., 4 kHz to 20 Hz), a mid-spectrum range (e.g., 500 Hz to 4 kHz), and a low spectrum range (e.g., 20 Hz to 500 Hz). In other words, the beats or spectrum of the audio music may control the texture information associated with how the shader renders the virtual clothing on the video clip.

[0030] Also, the music characteristics of the audio music can also control one or more lighting effects applied to the virtual clothing and / or other graphics around the user shown in the video. For example, a metallic effect may be applied to the virtual clothing, and the video effect parameters may control where the lighting shader is applied to which part of the virtual clothing and / or what kind of lighting graphics (e.g., crystal) are applied to the video clip. Therefore, the computing device 104 may determine the lighting characteristics based on the tempo characteristics. For example, when the audio music has a 4 / 4 time signature, a strong lighting effect may be applied to the strong beats and a weak lighting effect may be applied to the weak beats. Alternatively, the lighting characteristics may be determined based on the frequency spectrum range.

[0031] Additionally, or alternatively, based on the tempo characteristics or the frequency spectrum range, the animation speed of the additional graphics rendered by the lighting shader may be controlled. For example, during the strong beats, the crystal may rotate quickly around the user, and during the weak beats, the crystal may rotate slowly around the user.

[0032] Once the preparation to add video effects to the video clip is complete, the computing device 104 may modify the virtual clothing to harmonize the 2D (two-dimensional) texture of the virtual clothing with the 3D (three-dimensional) mesh around the body joints. By overlaying and harmonizing 2D animation objects on the 3D mesh, a virtual clothing that looks like 3D may be created. Then, the computing device 104 may generate a rendered video with video effects by synchronizing the video effects to the beat of the audio music. This video may be shown to the user on a display (e.g., display 705) communicatively coupled to the computing device 104. It should be understood that the video effects may be synchronized to the beat of the music in real time or near real time so that the user can visually recognize the video effects around the mesh generated based on one or more body joints when the user is filming the video on the display. Alternatively, or in addition, the server 106 may synchronize the video effects to the beat of the music. In such a manner, the video effects may be applied to the video clip 108 when the video clip 108 is uploaded to the server 106 for rendering the video effects.

[0033] Next, referring to FIG. 2, a computing device 202 according to an embodiment of the present disclosure is depicted. The computing device 202 may be the same as or similar to the computing device 104 depicted in FIG. 1 above. The computing device 202 may include a communication interface 204, a processor 206, and a computer-readable storage 208. By way of example, the communication interface 204 may be coupled to a network and may receive the video clip 108 and the audio music 110 (FIG. 1). The video clip 108 (FIG. 1) may be stored as video frames 246, and the music 110 may be stored as audio data 248 of the input 242.

[0034] In some examples, one or more video effects may also be received at the communication interface 204 and stored as video effect data 252. The video effect data 252 may include one or more video effect parameters associated with the video effect. The video effect parameters can define, but are not limited to, how the texture of the virtual clothing is rendered on the mesh around one or more joints of the body within the video clip and / or how it is worn.

[0035] As an example, one or more applications 210 may be provided by the computing device 104. The one or more applications 210 may include a video processing module 212, an audio processing module 214, a video effect module 216, a shader 218, and a specular shader 220. The video processing module 212 may include a video acquisition manager 224, a body joint identifier 226, and a mesh generator 228. The video acquisition manager 224 is configured to receive, acquire, or otherwise obtain video data including one or more video frames. Further, the body joint identifier 226 is configured to identify one or more body joints of one or more objects within the frame. The mesh generator 228 is configured to generate a mesh based on one or more body joints of one or more objects within the frame, and the mesh may be determined based on a combination of one or more body joints. In an exemplary aspect, the object is a person. Thus, a body segmentation algorithm may define a list of body joints identified and extracted from the video clip 108. The body joints may include, but are not limited to, the head, neck, pelvis, spine, left and right shoulders, left and right upper arms, left and right forearms, left and right hands, left and right thighs, left and right legs, left and right feet, and left and right toes. In some examples, the list of body joints may be received at the communication interface 204 and stored as body joints 250. In some aspects, the list of body joints may be received from a server (e.g., 106).

[0036] Furthermore, the audio processing module 214 may include an audio acquisition manager 232 and an audio analyzer 234. The audio acquisition manager 232 is configured to receive, acquire, or otherwise obtain audio data. The audio analyzer 234 is configured to determine audio information of the audio data. For example, the audio information may include, but is not limited to, beat information and frequency spectrum information of each beat of the audio data. As an example, an automatic beat tracking algorithm may be used to determine the beat information. In some embodiments, the beat information may already be embedded in the audio data as metadata. In other embodiments, the beat information may be received via the communication interface 204 and stored as audio data 248. The beat information provides the beat characteristics of each beat. The beat characteristics include, but are not limited to, the composition of the beat, the repeating order of strong beats and weak beats, the number of accented beats and non-accented beats, and the relative positions of accented beats and non-accented beats. For example, if the audio music has a 4 / 4 beat composition, each section has four beats with different beat strengths: a strong beat, the first weak beat, the second strong beat, and the weak beat. Furthermore, the frequency spectrum information may be extracted from the audio data at predetermined intervals (e.g., every beat). In some embodiments, the frequency spectrum information may already be embedded in the audio data as metadata. In other embodiments, the frequency spectrum information may be received via the communication interface 204 and stored as audio data 248.

[0037] Furthermore, the video effect module 216 may further include a clothing effect generator 238 and a video effect synchronizer 240. The clothing effect generator 238 is configured to determine an effect to be applied to virtual clothing and, by extension, video data based on audio data. Specifically, the clothing effect generator 238 is configured to determine one or more video effect parameters. For example, the clothing effect generator 238 is configured to determine one or more video effect parameters based on audio data. The determined one or more video effect parameters may be passed to a shader to control the way the shader renders the texture associated with the virtual clothing. In one exemplary embodiment, the virtual clothing is worn or otherwise rendered on a mesh that surrounds a plurality of body joints of a subject. The virtual clothing may move, change, or otherwise follow the user as the user moves within the video clip. In an exemplary embodiment, the texture of the virtual clothing may change every beat. In other words, the beat of the selected audio music may control the change in the video effect that controls the appearance of the virtual clothing. As another example, one or more other attributes of the virtual clothing and / or images within the video may change in response to the audio music. For example, the amount of shine, the opacity of the texture map, the color, and the speed of movement may be based on the audio music.

[0038] The video effect synchronizer 240 is configured to synchronize the video effects to the beat of the music of the selected audio music to generate a rendered video having video effects. The rendered video is stored or otherwise provided as the rendered video frame 254 of the output 244. In some embodiments, the video effect synchronizer 240 is configured to modify the virtual clothing to harmonize the 2D texture of the virtual clothing with the 3D mesh around the body joints. It should be understood that by overlaying and harmonizing 2D animation objects on the 3D mesh, an animation effect similar to 3D can be created.

[0039] The video effect synchronizer 240 includes, or alternatively communicates with, the shader 218 and / or the glow shader 220. The shader 218 is configured to receive video effect parameters. Based on the video effect parameters, the shader 218 is configured to produce or otherwise render an effect. For example, the shader 218 may change visual effects associated with the video effect, including but not limited to color, reflection, diffusion, translucency, transparency, metallicity, Fresnel reflection, and microsurface scattering. As another example, the glow shader 220 may change visual effects associated with the video effect, including but not limited to color, reflection, diffusion, translucency, transparency, metallicity, Fresnel reflection, and microsurface scattering. Further, the glow shader may render one or more graphics to the video clip based on the video effect parameters and / or the audio music. For example, the glow shader may determine that a crystal rotating around the user should be rendered to appear metallic and reflective.

[0040] Figures 3A - 3B show an example of determining one or more body joints 302 and generating a mesh 304 associated with the one or more body joints 302. As shown in FIG. 3A, one or more body joints 302 may be identified, and each body joint 302 may be associated with a predefined list of body joints. Based on the identified body joints 302, body parts configured to receive clothing may be identified. For example, a shirt may be identified based on the position of the body joint 302, in which case the upper body torso is configured to receive the shirt. As shown in FIG. 3B, a mesh 304 may be generated, in which case the mesh 304 is based on the body joints 302 of FIG. 3A. As an example, the mesh 304 of FIG. 3B may be used to virtually dress a user. Thus, the mesh 304 assists in occlusion when a part of a person's body is behind the virtual clothing and the person's hand may not be visible.

[0041] Figures 3C - 3E show exemplary video frames 310 - 330 showing exemplary clothing 306 worn on a mesh such as mesh 304 according to an embodiment of the present disclosure. That is, the virtual clothing 306 may be a virtual representation of clothing that a user 308 can virtually wear. As an example, the display of the virtual clothing 306 within frame 310 may be responsive to audio data 312. For example, the color of the virtual clothing may change according to the beat characteristics and / or frequency spectrum of the audio data 312. That is, the characteristics of the audio data 312 may be determined and used to specify video effect parameters, which may be used to determine or otherwise change the way a shader renders the texture of the virtual clothing on the mesh. Further, a specular shader may determine other characteristics of the virtual clothing (e.g., the amount of specular) and / or how other graphics 314 (e.g., crystals) are shown in video frame 310.

[0042] As shown in frame 320 of FIG. 3D, the audio data 316 may be different from the audio data 312 of frame 310. Thus, the shader may render a texture of a virtual garment that is different from the texture of the virtual garment 306 in frame 310 in frame 320. Alternatively, or in addition, the virtual garment 318 may be the same as or similar to the virtual garment 306 of frame 310. Further, the graphic 322 and / or other effects rendered by the specular shader may be different from the graphic 314 of frame 310. As further shown in FIG. 3D, the virtual garment 318 may be worn on or otherwise rendered on a mesh generated based on body joints. Thus, when the user 308 raises an arm, the virtual garment 318 adapts to the new joint position and thus the new configuration of the body part (e.g., the torso).

[0043] As shown in frame 330 of FIG. 3E, the audio data 324 may be different from the audio data 316 of frame 320. Thus, the shader may render a texture of a virtual garment 326 that is different from the texture of the virtual garment 318 of frame 320 in frame 330. Alternatively, or in addition, the virtual garment 326 may be the same as or similar to the virtual garment 318 of frame 320. Further, the graphic 328 and / or other effects rendered by the specular shader may be different from the graphic 322 of frame 320. As further shown in FIG. 3E, the virtual garment 326 may be worn on or otherwise rendered on a mesh generated based on body joints. Thus, when the user 308 raises an arm, the virtual garment 326 adapts to the new joint position and thus the new configuration of the body part (e.g., the torso).

[0044] As described above, the virtual clothing may be responsive to audio data. For example, when receiving audio data added to a video clip, the audio information of the audio data (e.g., tempo characteristics and / or frequency spectrum) may be determined for each predetermined period (e.g., the tempo of music). Based on the audio information, the texture rendered by the shader may be changed. For example, when the audio music has a 4 / 4 tempo composition, the first texture may be rendered; when the music has a 3 / 4 tempo composition, the second texture may be rendered; and when the music has a 2 / 4 tempo composition, the third texture may be rendered. The first texture, the second texture, and the third texture are different from each other. As another example, when the audio music has a strong beat, the fourth texture may be rendered; when the music has a first weak beat, the fifth texture may be rendered; and when the music has a second weak beat, the sixth texture may be rendered. The fourth texture, the fifth texture, and the sixth texture are different from each other. Alternatively, as described above, an animation effect may be determined based on the frequency spectrum range. In other words, the tempo or spectrum of the audio music controls the texture rendered on the virtual clothing within the video clip.

[0045] Next, referring to FIG. 4, a simplified method for rendering a virtual garment based on audio data according to an embodiment of the present disclosure is described. The general order of the steps of method 400 is shown in FIGS. 4-5. Generally, method 400 begins at 402 and ends at 460. Method 400 may have more or fewer steps than those shown in FIGS. 4-5, or the order of the steps may be set to be different from that shown in FIGS. 4-5. Method 400 can be executed as a set of computer-readable instructions encoded or stored in a computer-readable medium and executed by a computer system. In an exemplary aspect, method 400 is executed by a computing device associated with a user (e.g., 102). However, it should be understood that aspects of method 400 may also be executed by one or more processing devices such as a computer or a server (e.g., 104, 106). Further, method 400 can be executed by a gate or circuit associated with a processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a system-on-chip (SOC), a neural processing unit, or other hardware device. Hereinafter, method 400 will be described with reference to the system, components, modules, software, data structures, user interfaces, etc. depicted together with FIGS. 1 and 2.

[0046] Method 400 begins at 402, and the flow proceeds to 404. At 404, the computing device receives video data (e.g., video clip 108) that includes one or more video frames. For example, user 102 may generate, receive, obtain, or otherwise acquire video clip 108 via the computing device. At 408, the computing device processes each frame of the video data and identifies one or more target body joints within the frame. For example, a body joint algorithm may define a list of body joints to be identified and extracted from video clip 108. The body joints may include, but are not limited to, the head, neck, pelvis, spine, left and right shoulders, left and right upper arms, left and right forearms, left and right hands, left and right thighs, left and right legs, left and right feet, and left and right toes.

[0047] Referring back to 402 at the start, method 400 may proceed to 412. It should be understood that the computing device may perform operations 404 and 412 simultaneously. Alternatively, operation 412 may be performed after operation 404. In some aspects, operation 404 may be performed after operation 412.

[0048] At 412, the computing device receives audio data (e.g., audio music 110) selected by user 102 for addition to the video data. Subsequently at 416, the computing device analyzes the audio data to determine the audio information of the audio music 110. For example, the audio information includes the beat characteristics and / or frequency spectrum of each beat. In some aspects, the computing device may determine the beat characteristics of each beat through an automatic beat tracking algorithm. The beat characteristics include, but are not limited to, the composition of the beat, the repeating order of strong beats and weak beats, the number of accented beats and unaccented beats, and the relative positions of accented beats and unaccented beats. For example, if the audio music 110 has a 4 / 4 beat composition, each section has four beats with different beat strengths: a strong beat, a weak beat, a second strong beat, and a weak beat. In other aspects, the computing device may determine the frequency of each beat for association with a specific frequency range (e.g., high frequency, mid frequency, low frequency).

[0049] When video data and audio data are received and analyzed in operation 404-416, method 400 proceeds to 418. At 418, the computing device may generate a mesh based on the body joints identified at 408. For example, a mesh of the upper body torso may be based on the identification of the body joints of the left and right arms, head, and neck. Method 400 may proceed to 420. Here, the computing device determines virtual clothing to be applied to the mesh generated at 418. For example, the user may select virtual clothing for virtual try-on. Here, the virtual clothing may be a virtual display of the clothing or may be responsive to the audio data. As another example, the virtual clothing may be default clothing. Thereafter, method 400 may proceed to 424 and may determine texture rendering information associated with the determined clothing. For example, the video effect parameters may be based on the audio data analyzed at 416. The video effect parameters may control how the shader renders the texture. As an example, the rendering information may include, but is not limited to, color, reflectivity, opacity, transparency, etc., and may be specific to the combination of the audio data and the virtual clothing. That is, the video effect parameters may be different for audio music with a certain beat composition and beat intensity and audio music with a different beat composition and / or a different beat intensity. Thereafter, method 400 proceeds to 444 of FIG. 5, as indicated by alphanumeric A in FIGS. 4 and 5.

[0050] In 444, a glow effect associated with the audio music and / or the virtual clothing may be determined. As an example, a glow shader may be used, which is different from the main shader for rendering the texture information of the virtual clothing. The glow shader may determine different characteristics from the main shader that are applied to the video frame and / or the graphics within the video frame. As an example, the graphic information may include a crystal rotating within the frame, or other graphics with a glittering property. In some examples, the glow effect may be a glowing effect or other effect that affects the virtual clothing and / or other graphics within the video frame.

[0051] In 424 and 444 respectively, when the texture rendering information and the glow effect are determined for the animation object and the corresponding (plural) animation effect is determined, method 400 proceeds to operation 448. In 448, the computing device renders the texture and the glow effect according to the audio data analyzed in 416. In some examples, in 452, the video effect is synchronized with the beat or spectrum of the selected audio music to generate a rendered video with virtual clothing. For example, the video effect may be related to how fast the color or other recognized properties of the virtual clothing change. In 456, the computing device shows the user the rendered video with the video effect on a display (e.g., display 705). It should be understood that when worn on a mesh generated around one or more body joints, the changes in the virtual clothing may be synchronized in real time or almost in real time with the beat of the music so that the user can visually recognize the virtual clothing. As another example, when the user shown in the video moves, the virtual clothing may move with the user. The method ends at 460.

[0052] Method 400 is described as being performed by a computing device associated with a user, but it should be understood that one or more operations of Method 400 may be performed by any computing device or server, such as server 106. For example, the synchronization of video effects to the beat of music may be performed by a server that receives the music and video clip from a computing device associated with the user.

[0053] FIG. 6 is a block diagram showing physical components (e.g., hardware) of a computing device 600 in which aspects of the present disclosure may be implemented. The components of the computing device described below may be suitable for the computing device described above. For example, computing device 600 may represent computing device 104 of FIG. 1. In a basic configuration, computing device 600 may include at least one processing unit 602 and a system memory 604. Depending on the configuration and type of the computing device, system memory 604 may include volatile storage (e.g., random access memory), non-volatile storage (e.g., read only memory), flash memory, or any combination of such memories, but is not limited thereto.

[0054] System memory 604 may include an operating system 605 and one or more program modules 606 suitable for executing the various aspects disclosed herein. The operating system 605 may be suitable, for example, for controlling the operation of computing device 600. Further, aspects of the present disclosure may be implemented in combination with a graphics library, other operating systems, or any other application program and are not limited to a particular application or system. This basic configuration is shown by the components within dashed line 608 in FIG. 6. Computing device 600 may have additional features or functionality. For example, computing device 600 may also include additional data storage devices (removable and / or non-removable), such as, for example, magnetic disks, optical disks, or tapes. Such additional storage is shown in FIG. 6 as removable storage device 609 and non-removable storage device 610.

[0055] As described above, system memory 604 may store a plurality of program modules and data files. Program module 606 may execute a process that includes, but is not limited to, one or more aspects as described herein while being executed in at least one processing unit 602. Applications 607 may include a video processing module 623, an audio processing module 624, a video effects module 625, and a shader module 626, as described in more detail with respect to FIG. 1. Other program modules that may be used in accordance with aspects of the present disclosure may include, for example, email and contact applications, word processing applications, spreadsheet applications, database applications, slide presentation applications, drawing or computer-aided application programs, and / or one or more components supported by the systems described herein.

[0056] Furthermore, aspects of the present disclosure may be implemented in an electrical circuit including individual electronic components, an electronic chip including logic gates and packaged or integrated, a circuit using a microprocessor, or a single chip including electronic components or a microprocessor. For example, aspects of the present disclosure may be implemented via a system-on-chip (SOC) in which each component or components shown in FIG. 6 may be integrated on a single integrated circuit. Such an SOC device may include one or more processing units, a graphic unit, a communication unit, a system virtualization unit, and various application functions, all of which are integrated (or "burned in") on a chip substrate as a single integrated circuit. When operating via an SOC, with respect to the ability of a client to switch protocols, the functions described herein may be operated via application-specific logic integrated with other components of computing device 600 in a single integrated circuit (chip). Aspects of the present disclosure may be implemented using other techniques that can perform logical operations such as AND, OR, and NOT. Such techniques include, but are not limited to, mechanical techniques, optical techniques, fluid techniques, and quantum techniques. Furthermore, aspects of the present disclosure may be implemented in a general-purpose computer or in other circuits or systems.

[0057] Computing device 600 may also include one or more input devices 612, such as a keyboard, a mouse, a pen, a voice input device, an input device by touch or swipe, etc. It may also include (a plurality of) output devices 614A such as a display, a speaker, a printer, etc. It may also include an output 614B corresponding to a virtual display. The devices described above are exemplary, and other devices may be used. Computing device 600 may include one or more communication connections 616 that enable communication with other computing devices 650. Examples of suitable communication connections 616 include, but are not limited to, radio frequency (RF) transmitters, receivers, and / or transceiver circuits, and universal serial bus (USB), parallel ports, and / or serial ports.

[0058] As used herein, the term computer-readable medium may include computer storage media. Computer storage media may include volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, and the like. System memory 604, removable storage device 609, and non-removable storage device 610 are all examples of computer storage media (e.g., memory storage). Computer storage media may include RAM, ROM, electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other product accessible by computing device 600 that can be used for storing information. Such computer storage media may be part of computing device 600. Computer storage media does not include carrier waves or other transmitted or modulated data signals.

[0059] Communication media may be embodied by computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and includes any information delivery media. The term "modulated data signal" may represent a signal having one or more characteristics set or changed in such a manner as to encode information in the signal. By way of example, communication media may include, but is not limited to, wired media such as a wired network or direct wired connection, and wireless media such as acoustic, radio frequency (RF), infrared, and other wireless media.

[0060] Figures 7A and 7B illustrate a computing device or mobile computing device 700, such as a mobile phone, smartphone, wearable computer (such as a smartwatch), tablet computer, laptop computer, etc., in which aspects of the present disclosure may be implemented. Referring to Figure 7A, one aspect of a mobile computing device 700 for implementing an aspect is shown. In a basic configuration, the mobile computing device 700 is a portable computer having both input and output elements. The mobile computing device 700 typically includes a display 705 and one or more input buttons 709 / 710 that enable a user to input information into the mobile computing device 700. The display 705 of the mobile computing device 700 may also function as an input device (e.g., a touchscreen display). If a side input element 715 is included as an option, this allows for additional user input. The side input element 715 may be a rotary switch, button, or other type of manual input element. In an alternative aspect, the mobile computing device 700 may incorporate more or fewer input elements. For example, in some aspects, the display 705 may not be a touchscreen. In yet another alternative aspect, the mobile computing device 700 is a mobile phone system such as a cellular phone. The mobile computing device 700 may optionally also include a keypad 735. The optional keypad 735 may be a physical keypad or a "soft" keypad generated on a touchscreen display. In various aspects, the output elements include a display 705 for displaying a graphical user interface (GUI), visual indicators 731 (e.g., light-emitting diodes), and / or an acoustic transducer 725 (e.g., a speaker). In some aspects, the mobile computing device 700 incorporates a vibration transducer for providing tactile feedback to the user.In yet another aspect, the mobile computing device 700 incorporates input and / or output ports 730, such as audio inputs (e.g., microphone jacks), audio outputs (e.g., headphone jacks), and video outputs (e.g., HDMI™ ports), for transmitting signals to or receiving signals from an external source.

[0061] FIG. 7B is a block diagram illustrating an architecture of one aspect of a computing device, server, or mobile computing device. That is, the mobile computing device 700 can incorporate a system (702) (e.g., an architecture) for implementing some aspects. The system 702 can be implemented as a “smartphone” capable of executing one or more applications (e.g., browsers, email, calendars, contact managers, messaging clients, games, and media clients / players). In some aspects, the system 702 is integrated as a computing device, such as an integrated personal digital assistant (PDA) or a wireless phone.

[0062] One or more application programs 766 may be loaded into the memory 762 and executed on or in relation to the operating system 764. Examples of application programs include a telephone dialer program, an email program, a personal information management (PIM) program, a word processing program, a spreadsheet program, an Internet browser program, a messaging program, and / or one or more components supported by the system described herein. The system 702 also includes a non-volatile memory area 768 within the memory 762. The non-volatile memory area 768 may be used to store persistent information that should not be lost even when the power of the system 702 is turned off. The application program 766 may use and store information such as emails or other messages used in an email application in the non-volatile memory area 768. A synchronization application (not shown) also resides in the system 702 and is programmed to interact with a corresponding synchronization application residing on the host computer to synchronize the information stored in the non-volatile memory area 768 with the corresponding information stored on the host computer. As will be understood, other applications may be loaded into the memory 762 and executed on the mobile computing device 700 described herein (e.g., video processing module 623, audio processing module 624, video effect module 625, and shader module 626, etc.).

[0063] The system 702 has a power source 770, which may be implemented as one or more batteries. The power source 770 may further include an AC adapter or an external power source such as a powered docking cradle that replenishes or charges the battery.

[0064] System 702 can also include a wireless interface layer 772 that performs the function of transmitting and receiving radio frequency communications. The wireless interface layer 772 facilitates a wireless connection between system 702 and the "outside world" via a communication carrier or service provider. Transmissions to or from the wireless interface layer 772 are performed under the control of the operating system 764. In other words, communications received by the wireless interface layer 772 may be conveyed to the application program 766 via the operating system 764, and vice versa.

[0065] The visual indicator 720 may be used to provide visual notifications, and / or the audio interface 774 may be used to generate notification sounds via the acoustic transducer 725. In the illustrated configuration, the visual indicator 720 is a light-emitting diode (LED), and the acoustic transducer 725 is a speaker. These devices may be directly coupled to the power supply 770. By doing so, these devices will remain on during the period indicated by the notification mechanism even if the processor 760 / 761 and other components may shut down for battery power savings when activated. The LED may be programmed to remain on indefinitely until the user takes an action indicating the device is powered on. The audio interface 774 is used to provide audible signals to the user and receive audible signals from the user. The audio interface 774 may be coupled to, for example, in addition to the acoustic transducer 725, a microphone for receiving audible input for purposes such as facilitating a telephone conversation. In accordance with aspects of the present disclosure, the microphone can also function as an audio sensor for facilitating control of notifications, as described later. System 702 may further include a video interface 776 that enables operation of an on-board camera for recording still images, video streams, and the like.

[0066] The mobile computing device 700 implementing the system 702 may have additional features or functions. For example, the mobile computing device 700 can also include additional data storage devices (removable and / or non-removable) such as magnetic disks, optical disks, tapes, etc. Such additional storage is represented by the non-volatile memory area 768 in FIG. 7B.

[0067] Data / information generated or acquired by the mobile computing device 700 and stored via the system 702 may be stored locally on the mobile computing device 700 as described above. Alternatively, the data may be stored on any number of storage media accessible by the device via a wireless interface layer 772 or via a wired connection between the mobile computing device 700 and another computing device associated with the mobile computing device 700, such as a server computer within a distributed computing network such as the Internet. As will be appreciated, such data / information may be accessed via the wireless interface layer 772 or via the distributed computing network or via the mobile computing device 700. Similarly, such data / information may be easily transferred between computing devices for storage and use according to well-known data / information transfer and storage means including email and collaborative data / information sharing systems.

[0068] As described above, FIG. 8 shows one aspect of the architecture of a system that processes data received in a computing system from a remote source such as computing device 804, tablet computing device 806, or mobile computing device 808. The content displayed on server device 802 may be stored in different communication channels or other storage types. For example, computing devices 804, 806, 808 may represent computing device 104 of FIG. 1, and server device 802 may represent server 106 of FIG. 1.

[0069] In some aspects, one or more of video processing module 823, audio processing module 824, and video effect module 825 may be employed by server device 802. Server device 802 may provide data to client computing devices such as computing device 804, tablet computing device 806, and / or mobile computing device 808 (e.g., smartphone) through network 812, and may also receive data from client computing devices. The computer system described above may be embodied, for example, by computing device 804, tablet computing device 806, and / or mobile computing device 808 (e.g., smartphone). In these aspects of the computing device, usable graphic data is received, either pre-processed by the system of the graphic generator or post-processed by the received computing system, and in addition, content may be retrieved from store 816. The content store may include video data 818, audio data 820, and rendered video data 822.

[0070] FIG. 8 shows an exemplary mobile computing device 808 that can execute one or more aspects disclosed herein. Further, the aspects and functions described herein may operate on a distributed system (e.g., a cloud-based computing system), in which case the application functions, memory, data storage and retrieval, and various processing functions may operate remotely from each other on a distributed computing network such as the Internet or an intranet. Various types of user interfaces and information may be displayed via a computing device's mounted display or via a remote display unit associated with one or more computing devices. For example, various types of user interfaces and information may be displayed and interacted with on a wall surface onto which the various types of user interfaces and information are projected. Interactions with a number of computing systems in which aspects of the present invention may be implemented include key input, touch screen input, voice input, gesture input (where the associated computing device includes a detection function (e.g., a camera) for capturing and interpreting user gestures for controlling the functions of the computing device), and the like.

[0071] The phrases “at least one,” “one or more,” “or,” and “and / or,” when used, are non-restrictive forms of expression having both a conjunctive and a disjunctive meaning. For example, the expressions “at least one of A, B, and C,” “at least one of A, B, or C,” “one or more of A, B, and C,” “one or more of A, B, or C,” “A, B, and / or C,” and “A, B, or C” mean only A, only B, only C, A and B, A and C, B and C, or A, B, and C.

[0072] The term "a" or "one" entity refers to one or more of that entity. Thus, in this specification, the terms "a" (or "one"), "one or more", and "at least one" can be used interchangeably. It should also be noted that the terms "comprising", "including", and "having" can be used interchangeably.

[0073] As used herein, the term "automated" and variations thereof refer to any process or operation that, when performed, is typically continuous or semi - continuous and is performed without significant human input. However, even if the execution of a process or operation is due to significant or non - significant human input, if input is received prior to the execution of the process or operation, the process or operation can be automated. Such input is considered significant if it affects the way the process or operation is performed. Input that merely authorizes the execution of a process or operation by a human is not considered "significant".

[0074] All of the steps, functions, and operations discussed herein can be performed continuously and automatically.

[0075] Exemplary systems and methods of the present disclosure have been described in relation to computing devices. However, to avoid unnecessarily obscuring the present disclosure, some known structures and devices have been omitted from the foregoing description. This omission should not be construed as a limitation. Also, while specific details have been set forth to facilitate understanding of the present disclosure, it should be understood that the present disclosure can be implemented in various ways other than those specifically described herein.

[0076] Furthermore, while the exemplary embodiments described herein show the arrangement of various components of the system, certain components of the system can be remotely located in remote portions of a distributed network such as a LAN and / or the Internet, or within a dedicated system. Accordingly, it should be understood that the components of the system can be combined in one or more devices such as a server, communication device, etc., or can be located at specific nodes of a distributed network such as an analog and / or digital telecommunications network, a packet switching network, or a circuit switching network. From the foregoing description, and for reasons of computer efficiency, it will be understood that the components of the system can be located anywhere within the distributed network of components without affecting the operation of the system.

[0077] Furthermore, it should be understood that the various links connecting the elements can be wired / wireless links, or any combination thereof, or any other known or later-developed element(s) capable of supplying and / or communicating data between the connected elements. These wired or wireless links can also be secure links and may be capable of communicating encrypted information. For example, the communication medium used as a link can be any carrier suitable for electrical signals, including coaxial cables, copper wires, fiber optics, and may take the form of acoustic or light waves such as those generated during radio and infrared data communication.

[0078] Although the flowcharts have been discussed and shown in relation to a particular order of events, it should be understood that changes, additions, and omissions can be made to this order without materially affecting the operation of the disclosed configurations and embodiments.

[0079] For the present disclosure, several variations and modifications are possible. It is possible to provide some of the functions of the present disclosure without providing other functions.

[0080] In yet another configuration, the systems and methods of the present disclosure can be implemented in combination with application-specific computers, programmed microprocessors or microcontrollers and peripheral integrated circuit elements, ASICs or other integrated circuits, digital signal processors, electronic circuits or logic circuits incorporated in hardware such as individual element circuits, programmable logic devices or gate arrays such as PLDs, PLAs, FPGAs, PALs, application-specific computers, any equivalent means, etc. Generally, any device or means capable of implementing the methods described herein can be used to implement various aspects of the present disclosure. Exemplary hardware that can be used for the present disclosure includes computers, portable devices, telephones (cellular, Internet-enabled, digital, analog, hybrid, etc.), and other hardware known in the art. Some of these devices include processors (such as single or multiple microprocessors, etc.), memory, non-volatile storage, input devices, and output devices. Further, alternative software implementations can be constructed to implement the methods described herein. Alternative software implementations include, but are not limited to, distributed processing or component / object distributed processing, parallel processing, or virtual machine processing.

[0081] In yet another configuration, the disclosed method may be readily implemented in combination with object-using software or an object-oriented software development environment that provides portable source code that can be used on various computer or workstation platforms. Alternatively, the disclosed system may be implemented partially or fully in hardware using standard logic circuits or VLSI designs. Which software or hardware to use to implement the system according to the present disclosure depends on the requirements of the speed and / or efficiency of the system, specific functions, and the particular software or hardware system utilized, or the microprocessor or microcomputer system.

[0082] In yet another configuration, the disclosed method may be implemented, at least in part, in software that is executable on a programmed general-purpose computer, a special-purpose computer, a microprocessor, etc., that cooperate with a controller and a memory and that can be stored on a storage medium. In such a case, the systems and methods of the present disclosure can be implemented as a program incorporated into a personal computer, such as an applet, JAVA (registered trademark), or CGI script, as a resource resident on a server or computer workstation, as a dedicated measurement system, as a routine incorporated into a system component, etc. The system can also be implemented by physically incorporating the system and / or method into a software and / or hardware system.

[0083] The present disclosure is not limited to the described standards and protocols. If there are other similar standards and protocols not mentioned in this specification, these are also included in the present disclosure. Further, the standards and protocols mentioned in this specification, as well as other similar standards and protocols not mentioned in this specification, are periodically replaced by faster or more effective equivalents that have essentially the same function. Such replaced standards and protocols with the same function are considered equivalents included in the present disclosure.

[0084] The present disclosure relates to systems and methods for generating virtual clothing on a display according to at least the embodiments shown in the following paragraphs.

[0085] (A1) In one aspect, some examples include a method for generating virtual clothing on a display. The method may include obtaining video data and audio data, analyzing the video data to determine one or more body joints of a target object appearing in the video data, generating a mesh based on the determined one or more body joints, analyzing the audio data to determine audio characteristics, determining texture rendering information associated with the virtual clothing based on the audio characteristics, and generating a rendered video by using the texture rendering information to render the virtual clothing on the generated mesh.

[0086] (A2) In some examples of A1, the method may further include determining at least one of a beat characteristic or a frequency spectrum value from the audio data, selecting at least one video effect parameter based on the at least one of the beat characteristic or the frequency spectrum value, determining texture rendering information based on the at least one video effect parameter, and using the texture rendering information to render the virtual clothing on the generated mesh. The texture rendering information is based on at least one of the beat characteristic or the frequency spectrum value.

[0087] (A3) In some examples of A1 - A2, the method may further include rendering the virtual clothing using a first shader and rendering one or more graphics using a second shader different from the first shader, where the one or more graphics are rendered according to at least one of the beat characteristic or the frequency spectrum value.

[0088] (A4)In some examples of A1 to A3, the method further includes determining at least one of a second beat characteristic or a second frequency spectrum value from audio data, selecting a second video effect parameter based on at least one of the second beat characteristic or the second frequency spectrum value, determining second texture rendering information based on at least one of the second beat characteristic or the second frequency spectrum value, and rendering virtual clothing on the generated mesh using the second texture rendering information. The second texture rendering information is based on at least one of the second beat characteristic or the second frequency spectrum value.

[0089] (A5)In some examples of A1 to A4, the method further includes synchronizing a change from the first texture rendering information to the second texture rendering information based on the audio data.

[0090] (A6)In some examples of A1 to A5, the texture rendering information includes at least one of opacity, transparency, and metallicity.

[0091] (A7)In some examples of A1 to A6, the mesh includes a three-dimensional mesh around one or more determined body joints shown in the video data.

[0092] (A8)In some examples of A1 to A7, obtaining the audio data includes selecting the audio data from a music library, and the audio characteristics are embedded in the audio data as metadata.

[0093] In yet another aspect, some examples include a system including one or more processors and a memory coupled to the one or more processors. The memory stores one or more instructions that, when executed by the one or more processors, cause the one or more processors to perform any of the methods described herein (e.g., A1 to A8 above).

[0094] In yet another aspect, some examples include a computer-readable storage medium storing one or more programs for execution by one or more processors of a device. The one or more programs include instructions for performing any of the methods described herein (e.g., A1 - A8 described above).

[0095] The present disclosure substantially includes the components, methods, processes, systems, and / or apparatuses shown and described herein in various configurations and aspects, including their various combinations, secondary combinations, and subsets. Those skilled in the art should be able to understand how to construct and use the systems and methods disclosed herein after understanding the present disclosure. The present disclosure provides devices and processes in various configurations and aspects where there are no items not described and / or recited herein or in various configurations or aspects of this specification, for example, to improve performance, seek ease, and / or reduce implementation costs, when there are no items that may have been used in previous devices or processes, etc.

Claims

A method executed by one or more processing devices, comprising: obtaining video data and audio data; analyzing the video data to determine one or more body joints of an object appearing in the video data; generating a mesh based on the determined one or more body joints; analyzing the audio data to determine audio characteristics; determining texture rendering information associated with the virtual clothing based on the audio characteristics; generating a rendered video by rendering the virtual clothing on the generated mesh using the texture rendering information; The method further includes determining at least one of a second beat characteristic or a second frequency spectrum value from the audio data, selecting a second video effect parameter based on at least one of the second beat characteristic or the second frequency spectrum value, determining second texture rendering information based on at least one of the second beat characteristic or the second frequency spectrum value, rendering the virtual clothing on the generated mesh using the second texture rendering information, The method further includes: The second texture rendering information is based on at least one of the second beat characteristic or the second frequency spectrum value, a method for generating virtual clothing on a display. Claim 2 determining at least one of a beat characteristic or a frequency spectrum value from the audio data; selecting at least one video effect parameter based on at least one of the beat characteristic or the frequency spectrum value; Determining texture rendering information based on the at least one video effect parameter; Rendering the virtual clothing on the generated mesh using the texture rendering information; further comprising; the texture rendering information is based on at least one of the tempo characteristic or the frequency spectrum value; The method according to claim 1.

3. Rendering the virtual clothing using a first shader; Rendering one or more graphics using a second shader different from the first shader; further comprising; the one or more graphics are rendered according to at least one of the tempo characteristic or the frequency spectrum value; The method according to claim 2.

4. further comprising synchronizing a change from the texture rendering information to the second texture rendering information based on the audio data; The method according to claim 1.

5. the texture rendering information includes at least one of opacity, transparency, and metallicity; The method according to claim 1.

6. the mesh includes a three-dimensional mesh around the determined one or more body joints shown in the video data; The method according to claim 1.

7. Obtaining the audio data comprises; selecting the audio data from a music library; the audio characteristics are embedded in the audio data as metadata; The method according to claim 1.

8. A computing device for generating virtual clothes on a display, comprising: a processor; and a memory storing a plurality of instructions, wherein when the plurality of instructions are executed by the processor, the computing device is caused to: obtain video data and audio data; analyze the video data to determine one or more body joints of a target object appearing in the video data; generate a mesh based on the determined one or more body joints; analyze the audio data to determine audio characteristics; determine texture rendering information associated with the virtual clothes based on the audio characteristics; generate a rendered video by rendering the virtual clothes on the generated mesh using the texture rendering information; and wherein when the plurality of instructions are executed, the computing device is further caused to: determine at least one of a second beat characteristic or a second frequency spectrum value from the audio data; select a second video effect parameter based on the at least one of the second beat characteristic or the second frequency spectrum value; determine second texture rendering information based on the at least one of the second beat characteristic or the second frequency spectrum value; render the virtual clothes on the generated mesh using the second texture rendering information; and wherein when the plurality of instructions are executed, the computing device is caused to: The computing device, wherein the second texture rendering information is based on at least one of the second beat characteristic or the second frequency spectrum value. **Claim 9** When executed, the plurality of instructions cause the computing device to determine at least one of a beat characteristic or a frequency spectrum value from the audio data; select at least one video effect parameter based on at least one of the beat characteristic or the frequency spectrum value; determine texture rendering information based on the at least one video effect parameter; render the virtual clothing on the generated mesh using the texture rendering information; and further cause the computing device to wherein the texture rendering information is based on at least one of the beat characteristic or the frequency spectrum value. The computing device according to claim 8. **Claim 10** When executed, the plurality of instructions cause the computing device to render the virtual clothing using a first shader; render one or more graphics using a second shader different from the first shader; and further cause the computing device to wherein the one or more graphics are rendered according to at least one of the beat characteristic or the frequency spectrum value. The computing device according to claim 9. **Claim 11** When executed, the plurality of instructions cause the computing device to The computing device according to claim 8, further causing a change from the texture rendering information to the second texture rendering information to be synchronized based on the audio data.

12. The texture rendering information includes at least one of opacity, transparency, and metallicity. The computing device according to claim 8.

13. The mesh is a three-dimensional mesh around one or more body joints shown in the video data. The computing device according to claim 8.

14. A computer program including instructions for generating virtual clothing on a display, When the instructions are executed by one or more processors of a computing device, the instructions cause the computing device to acquire video data and audio data, analyze the video data to determine one or more body joints of a target object appearing in the video data, generate a mesh based on the determined one or more body joints, analyze the audio data to determine audio characteristics, determine texture rendering information associated with the virtual clothing based on the audio characteristics, generate a rendered video by rendering the virtual clothing on the generated mesh using the texture rendering information, and cause the above to be performed. When the instructions are executed by the one or more processors, the instructions cause the computing device to determine at least one of a second beat characteristic or a second frequency spectrum value from the audio data, Selecting a second video effect parameter based on at least one of the second beat characteristics or the second frequency spectrum value; Determining second texture rendering information based on at least one of the second beat characteristics or the second frequency spectrum value; Rendering the virtual clothes on the generated mesh using the second texture rendering information; which is further to be performed; The second texture rendering information is a computer program based on at least one of the second beat characteristics or the second frequency spectrum value. **Claim 15** When executed by the one or more processors, the instructions cause the computing device to Determine at least one of a beat characteristic or a frequency spectrum value from the audio data; Select at least one video effect parameter based on at least one of the beat characteristic or the frequency spectrum value; Determine texture rendering information based on the at least one video effect parameter; Render the virtual clothes on the generated mesh using the texture rendering information; which is further to be performed, wherein the texture rendering information is based on at least one of the beat characteristic or the frequency spectrum value, The computer program according to claim 14. **Claim 16** When executed by the one or more processors, the instructions cause the computing device to Render the virtual clothes using a first shader; Render one or more graphics using a second shader different from the first shader; further perform the one or more graphics are rendered according to at least one of the tempo characteristics or the frequency spectrum values The computer program according to claim 15.

17. the texture rendering information includes at least one of opacity, transparency, and metallicity The computer program according to claim 14.

Citation Information

Patent Citations

  • Virtual character reloading method and device

    CN112274926A

  • Facilitating synchronization of motion imagery and audio

    US20180197578A1

  • Virtual character generation from image or video data

    US20200306640A1

  • Learning-based animation of clothing for virtual try-on

    US20210118239A1

  • Producing realistic body movement using body images

    WO2017137948A1