Method, system, computer-readable storage medium, and program for video creation and audio or visual input for interaction

Real-time audio and video analysis techniques enable enhanced control of visual and audio effects in video creation, addressing limitations in existing platforms and improving user experience.

JP2025520387AActive Publication Date: 2025-07-03LEMON CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024573292
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-17
Filing Date
2023-05-30
Publication Date
2025-07-03
Estimated Expiration
2043-05-30

AI Technical Summary

Technical Problem

Existing video creation platforms have limited technology for controlling and selecting audio and visual effects, leading to user confusion, slowed processes, and degraded experience due to inconvenient input methods.

Method used

Techniques for controlling visual effects based on audio input and audio effects based on video input in real time, using machine learning models to analyze voice and video signals for modifying visual and audio elements in video items.

Benefits of technology

Enhances user experience by allowing real-time control of visual and audio effects, simplifying the video creation process, and improving user interaction with video content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025520387000001_ABST
    Figure 2025520387000001_ABST
Patent Text Reader

Abstract

During the creation of a video item, a first type of input can be received via a first component of a computing device. The first type of input can correspond to a first type of element associated with the video item. Based on the first type of input, the characteristics of signals within the first type of input can be determined in real time. At least one modification can be made to a second type of element associated with the video item, at least partially based on the characteristics of signals within the first type of input. In some examples, the first type of input can be an audio input, and the second type of input can include visual elements. In other examples, the first type of input can be a video input, and the second type of input can include audio elements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross - Reference to Related Applications) This application claims priority to U.S. Patent No. 17 / 843,891, filed on June 17, 2022, with the title "Voice or Visual Input for Video Creation and Interaction", the disclosure of which is hereby incorporated by reference in its entirety.

Background Art

[0002] Communication using Internet - based tools is increasing. The Internet - based tools may be any software or platform. Existing social media platforms enable communication by users sharing information such as images and videos via static applications or web pages. As communication devices such as mobile phones become more and more powerful, people continue to seek new entertainment, social network, and communication methods.

Brief Description of the Drawings

[0003] The following detailed description is better understood when read in conjunction with the accompanying drawings. For the purpose of illustration, exemplary embodiments of various aspects of the present disclosure are shown in the accompanying drawings, but the present invention is not limited to the specific methods and means disclosed.

[0004]

Figure 1

[0005]

Figure 2

[0006]

Figure 3

[0007]

Figure 4

[0008]

Figure 5

[0009]

Figure 6

[0010]

Figure 7

[0011]

Figure 8

[0012]

Figure 9

[0013]

Figure 10

[0014]

Figure 11A

[0015]

Figure 11B

[0016]

Figure 12A

[0017]

Figure 12B

[0018]

Figure 13A

[0019]

Figure 13B

[0020]

Figure 14

DETAILED DESCRIPTION OF THE INVENTION

[0021] A user may use a content creation platform to generate content, such as a video item. In some examples, the video item may be created based on a video input, for example, obtained via a camera. The video input may include, for example, a video of the user including the user's face. Additionally, the video item may have a corresponding audio output. The audio output may be generated based on an audio input, for example, obtained via a microphone. The audio input may include, for example, the user's voice including the user's speech. The user may desire to add effects, such as one or more visual effects, etc., to the video item. The user may also desire to add audio effects to the output audio associated with the video item.

[0022] One drawback of existing video creation platforms is that the technology for controlling audio and visual effects can be limited. For example, users may have limited input technology for indicating the time when visual or audio effects are applied to video items, or the time applied to the output audio. Also, the technology for selecting the types of visual and audio effects, or for controlling the size and duration of visual and audio effects, may be limited. Some existing input technologies for visual and audio effects are inconvenient and may require the user to perform additional actions or steps not necessary for creating a video. This can cause confusion for the user, slow down the video creation process, cause the user to abandon desired effects, or degrade the user experience for both content creators and content viewers.

[0023] Described herein are techniques for audio and visual effects based on video creation inputs. In some examples, with the described techniques, a user can, optionally in real time, create and control visual effects based on audio input. For example, changes in the audio input from a microphone, such as changes in the user's voice, can be used to control one or more visual effects within a video item. These changes may be included in audio features, such as pitch, tone, volume, energy, or duration characteristics. Visual effects may include, for example, stretching or swaying visual elements, inserting animated graphical elements into a video item, or moving visual elements. Additionally, in some examples, with the described techniques, a user can, optionally in real time, create and control audio effects based on video input. In one example, changes in the video input from a camera, such as the movement of one or more body parts, can be used to control one or more audio effects within the audio output. Audio effects may include, for example, changing audio features in the audio output, such as pitch, tone, volume, energy, echo, duration, and the like.

[0024] The techniques for audio and visual effects based on video creation inputs described in this specification may be utilized by a system for delivering content. FIG. 1 shows an exemplary system 100 for delivering content. The system 100 may include a server 102 and a plurality of client devices 104. The server 102 and the plurality of client devices 104a-104n may communicate with each other via one or more networks 132.

[0025] The server 102 may be located in a data center such as a single building or may be distributed across different geographical locations (e.g., several buildings). The server 102 may provide services via the one or more networks 132. The one or more networks 132 may include various network devices such as routers, switches, multiplexers, hubs, modems, bridges, repeaters, firewalls, proxy devices, and / or similar devices. The one or more networks 132 may include physical links such as coaxial cable links, twisted pair cable links, fiber optic links, combinations thereof, etc. The one or more networks 132 may include wireless links such as cellular links, satellite links, Wi-Fi links, etc.

[0026] Server 102 may include a plurality of computing nodes that host various services. In one embodiment, the node hosts a video service 112. The video service 112 may include a content streaming service such as an Internet Protocol video streaming service. The video service 112 may be configured to deliver content 123 via various transmission technologies. The video service 112 is configured to provide content 123 such as video, audio, text data, combinations thereof, and the like. The content 123 may include a content stream (e.g., a video stream, an audio stream, an information stream), a content file (e.g., a video file, an audio file, a text file), and / or other data. The content 123 can be stored in a database 122. For example, the video service 112 may include a video sharing service, a video hosting platform, a content delivery platform, a collaborative game platform, and the like.

[0027] In one embodiment, the content 123 distributed or provided by the video service 112 includes video. The video may have a duration of less than a predetermined time limit, such as 1 minute, 5 minutes, or other predetermined minutes. By way of example and not limitation, the video may include at least one and up to four 15-second segments combined with each other. The short video duration can provide viewers with rapid and continuous entertainment that enables users to view a large number of videos within a short time frame. Such rapid and continuous entertainment may be popular on social media platforms.

[0028] The video may include a pre-recorded audio overlay, such as a song or audio clip pre-recorded from a TV program or movie. If the short video includes a pre-recorded audio overlay, the short video may be characterized by one or more people lip-syncing (mouthing the words), dancing, or moving their bodies in other ways along with the pre-recorded audio. For example, the short video may be characterized by an individual completing a "dance challenge" to a hit song, or the short video may be characterized by two people participating in lip-syncing or partner dancing. As another example, the short video may be characterized by an individual achieving a challenge that requires moving their body to correspond to, for example, the beat or rhythm of a pre-recorded song characterized by the pre-recorded audio overlay. Other videos may not include a pre-recorded audio overlay. For example, these videos may be characterized by an individual doing sports, playing pranks, or giving advice on beauty, fashion, cooking tips, home interior tips, etc.

[0029] In one embodiment, the content 123 may be output to different client devices 104 via the network 132. The content 123 may be streamed to the client device 104. The content stream may be a stream of video received from the video service 112. The plurality of client devices 104 may be configured to access the content 123 from the video service 112. In one embodiment, the client device 104 may include a content application 106. The content application 106 outputs (e.g., displays, renders, presents) the content 123 to the user associated with the client device 104. The content may include video, audio, comments, text data, etc.

[0030] The plurality of client devices 104 may include any type of computing device, such as a mobile device, a tablet device, a laptop computer, a desktop computer, a smart TV or other smart devices (e.g., smart watch, smart speaker, smart glasses, smart helmet), a gaming device, a set-top box, a digital streaming device, a robot, etc. The plurality of client devices 104 may be associated with one or more users. A single user may access the server 102 using one or more of the plurality of client devices 104. The plurality of client devices 104 may move to various locations and access the server 102 using different networks.

[0031] The video service 112 may be configured to receive input from a user. The user may be registered as a user of the video service 112 and may also be a user of the content application 106 operating on the client device 104. The user input may include a video created by the user, a user comment associated with the video, or a "like" associated with the video. The user input may include a connection request and user input data such as text data, digital image data, or user content. The connection request may include a request to connect to the video service 112 from the client devices 104a - d. The user input data may include information that a user connected to the video service 112, such as a video and / or a user comment, desires to share with other connected users of the video service 112.

[0032] Video service 112 may be able to receive different types of inputs from users using different types of client devices 104. For example, a user using content application 106 on a first user device such as a mobile phone or a tablet may be able to create and upload a video using content application 106. A user using content application 106 on a different mobile phone or tablet may be able to view, comment on, or "like" a video or a comment written by another user. In another example, a user using content application 106 on a smart TV, laptop, desktop, or gaming device may not be able to create and upload a video or comment on a video using content application 106. Instead, a user using content application 106 on a smart TV, laptop, desktop, or gaming device may only be able to view a video, view comments left by other users, or "like" a video using content application 106.

[0033] In one embodiment, a user may use content application 106 on client device 104 to create a video, such as a short video, and upload the video to server 102. Client device 104 may be able to access interface 108 of content application 106. Interface 108 may include input elements. For example, the input elements may be configured to enable the user to create a video. To create a short video, the user may give content application 106 permission to access an image acquisition device, such as a camera of client device 104, or a microphone. Using content application 106, the user may select the duration of the video or set the speed of the video, such as "slow motion" or "speed up".

[0034] The user may edit the video using the content application 106. After the user creates a video, the user may use the content application 106 to upload the video to the server 102 and / or locally store the video in the client devices 104a - n. When the user uploads the video to the server 102, the user may select whether to make the video viewable by all other users of the content application 106 or only by a subset of the users of the content application 106. The video service 112 may store the uploaded video and any metadata associated with the video in one or more databases 122.

[0035] In one embodiment, the user may provide input on the video using the content application 106 on the client device 104. The client device 104 may access an interface 108 of the content application 106 that enables the user to provide input associated with the video. The interface 108 may include input elements. For example, the input elements may be configured to receive input from the user such as comments or "likes" associated with a particular video. If the input is a comment, the content application 106 may allow the user to set an emoji associated with their input. The content application 106 may be able to determine time information about the input, such as when the user wrote the comment. The content application 106 may be able to send the input and the associated metadata to the server 102. For example, the content application 106 may send the comment, the identifier of the user who wrote the comment, and the time information about the comment to the server 102. The video service 112 may store the input and the associated metadata in the database 122.

[0036] Video service 112 may be configured to output uploaded videos and user inputs to other users. A user may be registered as a user of video service 112 and view videos created by other users. The user may also be a user of content application 106 operating on client device 104. Content application 106 may output (display, render, present) videos and user comments to the user associated with client device 104. Client device 104 can access interface 108 of content application 106. Interface 108 may include output elements. The output elements may be configured to display information about different videos so that the user can select and view the videos. For example, the output elements may be configured to display a plurality of cover images, subtitles, or hashtags associated with the videos. The output elements may also be configured to arrange the videos according to the categories associated with each video.

[0037] In one embodiment, user comments associated with a video may be output to other users who are viewing the same video. For example, all users accessing the video may view the comments associated with the video. Video service 112 may output the comments associated with the video simultaneously. The comments may be output by video service 112 in real time or near real time. Content application 106 may display the video and the comments on client device 104 in various ways. For example, the comments may be displayed in an overlay on the content, or in an overlay next to the content. As another example, a user who wants to view the comments of other users associated with a video may need to select a button to view the comments. The comments may be displayed with animation when displayed. For example, the comments may be scrolled across the video or across the overlay.

[0038] According to the technology described in this specification, characteristics of an input signal, such as signals within an input video and an input audio, may be used to control modifications to elements associated with a video item, such as visual or audio elements. FIG. 2 is an exemplary diagram showing modifications to audio and visual elements according to the present disclosure. As shown in FIG. 2, a system 200 generates a video item 154 based on a video input 114. The video input 114 may be obtained by a camera 107 and may be included, for example, in one or more of the client devices 104a - n of FIG. 1. In the example of FIG. 2, the video input 114 includes a video of the user 101, such as a video of the user 101 including the face and / or other body parts of the user 101. The video item 154 may include at least a portion, and in some cases all, of the content of the video input 114, but if the content of the video input 114 is included in the video item 154, it may be modified in one or more ways.

[0039] The video item 154 has a corresponding audio output 155. The audio output 155 may be generated based on an audio input 115. The audio input 115 may be obtained by a microphone 105 and may be included, for example, in one or more of the client devices 104a - n of FIG. 1. In the example of FIG. 2, the audio input 115 includes the audio of the user 101, such as an audio including the voice of the user 101. The audio output 155 may include at least a portion, and in some cases all, of the content of the audio input 115, but if the content of the audio input 115 is included in the audio output 155, it may be modified in one or more ways. During playback, the video item 154 and the audio output 155 may be played together in synchronization with each other. For example, the audio output 155 may include words spoken by the user 101, and these words may be played in the audio output 155 in synchronization with the display of the corresponding mouth movements of the user 101 within the video item 154.

[0040] As shown in FIG. 2, the voice analysis unit 125 may determine the voice signal characteristics 135 by performing voice analysis on the voice input 115. The voice analysis may be performed during the creation of the video item 154. The voice signal characteristics may change over time in the voice input. For example, the voice signal characteristics 135 may include voice features such as pitch, tone, volume, energy, and / or duration. In some examples, the voice signal characteristics 135 may include voice features of the voice of the user 101, such as pitch, tone, volume, energy, and / or duration, including the sound made by the user 101. In some cases, the voice analysis unit 125 may perform voice analysis on the voice input 115 using one or more machine learning models, such as one or more neural network models. In some examples, the voice analysis may include converting the voice input 115 into the frequency domain, for example, by performing one or more frequency domain transforms (e.g., Fourier transform or another similar transform) on the voice input 115. In some examples, the voice analysis may include determining the voice signal characteristics 135, such as pitch, tone, volume, and energy, at different times, and determining and tracking the changes over time in these characteristics.

[0041] Once determined by the voice analysis unit 125, the voice signal characteristics 135 may be provided to the modification unit 140. Then, the modification unit 140 may determine one or more modifications to add to the video item 154 by evaluating the voice signal characteristics 135 in combination with the voice-based modification instruction 145. For example, the voice-based modification instruction 145 may indicate various conditions for modifying the visual elements 164 within the video item 154. Specifically, the voice-based modification instruction 145 may indicate a change in the voice signal characteristics 135 that can trigger a modification to the visual element 164. For example, these trigger conditions may include changes in the user's voice or another sound, such as the pitch, tone, or volume of background noise. In some cases, the trigger conditions may require that the pitch, tone, or volume meet a selected criterion for at least a threshold duration. The selected criterion may include, for example, exceeding a minimum threshold, falling below a maximum threshold, remaining within a predetermined value range, matching or correlating with a predetermined pattern, and so on.

[0042] In addition to specifying the trigger conditions, the voice-based modification instruction 145 may specify the modification resulting from the trigger conditions. In some examples, the modification may include stretching, shaking, discoloring, and / or moving the visual element 164. Also, in some examples, the modification may include inserting an animated graphical element that replaces, obscures, or modifies the video element 164 into the video item 154. In some examples, the visual element 164 may include body parts of the user 101, such as the user's eyes, eyebrows, mouth, lips, nose, hands, arms, and so on. In other examples, the visual element 164 may be a different type of object, such as a controllable character in a game-related video. Also, in some examples, the visual element 164 may include all or part of an image frame within the video item.

[0043] In one specific example, the visual element 164 may extend and / or sway in one or more directions based on the pitch of the user 101's voice, for example when the pitch rises above a certain level. In another specific example, increasing the strength of the user's voice (e.g., volume, energy, etc.) may cause the user's voice to convey an emotion such as anger or rage by causing an animated graphical element, such as a fireball, to be inserted onto the visual element 164, e.g., onto the user's eyeball. Additionally, as the strength of the user's voice increases, the user's voice may gradually take on a metallic texture. In yet another specific example, a change in the pitch of the user's voice may control the movement of the visual element 164. For example, the visual element 164 may be a character within a game. As the pitch of the user's voice increases, the character may move upward, and as the pitch of the user's voice decreases, the character may move downward. Any or all of these modifications to the visual element 164 may be performed in real time when the audio signal characteristic 135 that triggers the modification is detected.

[0044] As also shown in FIG. 2, the video analysis unit 124 may determine the video signal characteristics 134 by performing video analysis on the video input 114. The video analysis may be performed during the creation of the video item 154. The video signal characteristics 134 may change over time in the video input 114. For example, the video signal characteristics 134 may include the positions and movements of different objects within the video input 114. For example, the video signal characteristics 134 may include the positions and movements of at least one body part of the user 101, such as the eyes, eyebrows, mouth, head, hands, etc. In some cases, the video analysis unit 124 may perform video analysis on the video input 114 using one or more machine learning models, such as one or more neural network models. In some examples, the video analysis may include performing object detection and / or object recognition analysis on different frames of the video input. The object detection and / or object recognition analysis may include face detection and / or face recognition analysis. In some examples, the video analysis may include determining the video signal characteristics 134, such as object positions, at different times and determining and tracking the changes over time in these positions.

[0045] Once determined by the video analysis unit 124, the video signal characteristics 134 may be provided to the modification unit 140. Then, the modification unit 140 may determine one or more modifications to apply to the audio item 155 by evaluating the video signal characteristics 134 in combination with the video-based modification instructions 144. For example, the video-based modification instructions 144 may indicate various conditions for modifying the audio elements 165 within the audio output 155. Specifically, the video-based modification instructions 144 may indicate changes in the video signal characteristics 134 that may trigger modifications to the video elements 165. For example, these trigger conditions may include the movement of an object, such as a body part of the user 101, such as the user's eyes, eyebrows, mouth, lips, nose, hands, arms, etc. In some cases, the trigger conditions may require the object to move along a selected movement pattern, exceeding a given movement speed threshold and / or a movement distance threshold, in one or more selected directions.

[0046] In addition to defining the trigger conditions, the video-based modification instruction 144 may define the modifications resulting from the trigger conditions. In some examples, the modification may include causing the audio element 165 to change. In some examples, the audio element 165 may include characteristics of the user 101's voice, such as pitch, tone, volume, energy, echo, or duration. In one specific example, the user 101 may change the pitch of the user's voice by moving their eyebrows. For example, moving the eyebrows upward may increase the pitch of the user's voice, and moving the eyebrows downward may decrease the pitch of the user's voice. In another specific example, the movement of a body part or other object may cause an echo effect. For example, raising a body part upward may enhance the echo effect (effect), such as by converting the user's voice in real time to achieve an echo of an angelic sound. Lowering the body part may reduce the echo effect. Any or all of these modifications to the audio element 165 may be executed in real time when the video signal characteristic 134 that triggers the modification is detected.

[0047] In some examples, any or all of the video analysis technology, audio analysis technology, visual element modification technology, and audio element modification technology described above may be executed by one or more of the client devices 104a - n and / or the server 102 in FIG. 1. Returning to FIG. 1, an example is shown where the video analysis unit 124, the audio analysis unit 125, and the modification unit 140 are included in the server 102 and the client devices 104a - n respectively. Therefore, in the example of FIG. 1, the client devices 104a - n and the server 102 may be capable of executing any or all of the video analysis technology, audio analysis technology, visual element modification technology, and audio element modification technology described above. In some examples, the video analysis technology, audio analysis technology, visual element modification technology, and / or audio element modification technology described above may be executed entirely on the client device 104. In some other examples, the video analysis technology, audio analysis technology, visual element modification technology, and / or audio element modification technology described above may be executed entirely on the server 102.

[0048] In still other examples, the execution of the video analysis technology, audio analysis technology, visual element modification technology, and / or audio element modification technology described above may be distributed between the client device 104 and the server 102. For the scenario where the server 102 is employed to execute any or all of these technologies, the video input 114 and / or the audio input 115 may be acquired by the client devices 104a - n (e.g., via the camera 107 and / or the microphone 105) and provided to the server 102 via one or more networks 132 for analysis, processing, and / or modification. Additionally, in some examples, the results of the analysis, processing, and / or modification may be transmitted from the server 102 to the client devices 104a - n via one or more networks 132. Specifically, in some examples, the server 102 may execute modifications on the visual element 164 and / or the audio element 165. In other examples, the server 102 may only determine when to execute the modification, and the server may send back instructions for the client device 104a - n to execute this change in order to enable the modification to be executed on the client side.

[0049] Figure 3 shows an exemplary process 300 for element modification based on video creation input according to the present disclosure. In operation 310, during the creation of a video item, a first type of input corresponding to a first type of element associated with the video item is received via a first component of the computing device. In some examples, receiving the first type of input via the first component of the computing device may include receiving voice input via a microphone. As described above with reference to FIG. 2, the first type of input may be voice input 115 that can be received via microphone 105 during the creation of video item 154. Voice input 115 may include or correspond to a first type of element, such as voice element 165. Voice element 165 may include characteristics of the user 101's voice, such as pitch, tone, volume, energy, echo, or duration.

[0050] In some other examples, receiving the first type of input via the first component of the computing device may include receiving video input via a camera. As described above, the first type of input may be video input 114 that can be received via camera 107 during the creation of video item 154. Video input 114 may include or correspond to a first type of element, such as visual element 164. In some examples, visual element 164 may include body parts of user 101, such as the user's eyes, eyebrows, mouth, lips, nose, hands, arms, etc.

[0051] In operation 312, the characteristics of the signals within the first type of input are determined in real time based on the first type of input. The characteristics may vary over time in the first type of input. In some examples, the characteristics include sound features, which may include at least one of pitch, tone, volume, energy, or duration. As described above, the speech analysis unit 125 may determine the speech signal characteristics 135 by performing speech analysis on the speech input 115. The speech signal characteristics 135 may include measurement results of at least one of pitch, tone, volume, energy, or duration. In some examples, the speech analysis unit 125 may perform speech analysis on the speech input 115 using one or more machine learning models, such as one or more neural network models. In some cases, the speech analysis may include converting the speech input 115 into the frequency domain, for example, by performing one or more frequency domain transforms (e.g., Fourier transform or another similar transform) on the speech input 115. In some examples, the speech analysis may include determining the speech signal characteristics 135, such as pitch, tone, volume, and energy, at different times, and determining and tracking the changes over time in these characteristics.

[0052] In some other examples, this characteristic may include the movement of at least one body part within the video input. As described above, the video analysis unit 124 may determine the video signal characteristic 134 by performing video analysis on the video input 114. The video signal characteristic 134 may include the calculation of the movement of at least one body part within the video input 114. In some examples, the video analysis unit 124 may perform video analysis on the video input 114 using one or more machine learning models, such as one or more neural network models. In some cases, the video analysis may include performing object detection and / or object recognition analysis for different frames of the video input. The object detection and / or object recognition analysis may include face detection and / or face recognition analysis. In some examples, the video analysis may include determining the video signal characteristic 134, such as the object position, at different times and determining and tracking the changes over time at these positions.

[0053] In operation 314, at least one modification is made to a second type of element associated with a video item, based at least in part on the characteristics of a signal within a first type of input. The at least one modification may be performed in real time when the characteristics of the signal within the first type of input that trigger the at least one modification are determined. In some examples, the second type of element may include visual elements within the video item. As described above with reference to FIG. 2, the modification unit 140 may modify the visual element 164 within the video item 154 based on the audio signal characteristics 135 and the audio-based modification instructions 145. For example, the audio-based modification instructions 145 may indicate various conditions for modifying the visual element 164 within the video item 154. Specifically, the audio-based modification instructions 145 may indicate a change in the audio signal characteristics 135 that may trigger a modification to the visual element 164. In addition to defining the trigger conditions, the audio-based modification instructions 145 may define the modification resulting from the trigger conditions. In one specific example, operation 314 may include stretching a visual element in at least one direction based on a change in at least one of the characteristics of the signal within the audio input. This example will be described in detail below with reference to FIG. 5. In another specific example, operation 314 may include swaying a visual element based on a change in at least one of the characteristics of the signal within the audio input. This example will be described in detail below with reference to FIG. 6. In another specific example, operation 314 may include inserting at least one animated graphical element into the video item based on a change in at least one of the characteristics of the signal within the audio input. This example will be described in detail below with reference to FIG. 7. In another specific example, operation 314 may include causing movement of a visual element based on a change in at least one of the characteristics of the signal within the audio input. This example will be described in detail below with reference to FIG. 8.

[0054] In some other examples, the second type of element may include an audio element within the audio output associated with the video item. As described above with reference to FIG. 2, the modification unit 140 may modify the audio element 165 within the audio output 155 based on the video signal characteristic 134 and the video-based modification instruction 144. For example, the video-based modification instruction 144 may indicate various conditions for modifying the audio element 165 within the audio output 155. Specifically, the video-based modification instruction 144 may indicate a change in the video signal characteristic 134 that can trigger a modification to the video element 165. In addition to defining the trigger condition, the video-based modification instruction 144 may define the modification resulting from the trigger condition. In one specific example, the operation 314 may include modifying an audio feature in the audio output, including at least one of pitch, tone, volume, energy, echo, or duration, based on the movement of at least one body part in the video input. This example will be described in detail below with reference to FIG. 10.

[0055] FIG. 4 shows an exemplary process 400 for modifying visual elements based on an audio input according to the present disclosure. In operation 410, during the creation of a video item, an audio input is received via a microphone, where the audio input corresponds to a voice element associated with this video item. As described above with reference to FIG. 2, the audio input 115 may be received via the microphone 105 during the creation of the video item 154. The audio input 115 may include or correspond to a first type of element, such as an audio element 165. The audio element 165 may include features of the user 101's voice, such as pitch, tone, volume, energy, echo, or duration.

[0056] In operation 412, the characteristics of the signals in the input voice are determined in real time based on the voice input. These characteristics may change over time in the voice input. In some examples, the characteristics include voice features, and the voice features may include at least one of pitch, tone, volume, energy, or duration. As described above, the voice analysis unit 125 may determine the voice signal characteristics 135 by performing voice analysis on the voice input 115. The voice signal characteristics 135 may include measurement results of at least one of pitch, tone, volume, energy, or duration. In some examples, the voice analysis unit 125 may perform voice analysis on the voice input 115 using one or more machine learning models, such as one or more neural network models. In some cases, the voice analysis may include converting the voice input 115 into the frequency domain, for example, by performing one or more frequency domain transforms (e.g., Fourier transform or another similar transform) on the voice input 115. In some examples, the voice analysis may include determining the voice signal characteristics 135, such as pitch, tone, volume, and energy, at different times, and determining and tracking the changes over time in these characteristics.

[0057] In operation 414, at least one modification is applied to the visual elements within the video item, based at least in part on the characteristics of the signals within the voice input. The at least one modification may be performed in real time when the characteristics of the signals within the voice input that trigger the at least one modification are determined. As described above with reference to FIG. 2, the modification unit 140 may modify the visual elements 164 within the video item 154 based on the voice signal characteristics 135 and the voice-based modification instructions 145. For example, the voice-based modification instructions 145 may indicate various conditions for modifying the visual elements 164 within the video item 154. Specifically, the voice-based modification instructions 145 may indicate a change in the voice signal characteristics 135 that can trigger a modification to the visual elements 164. In addition to defining the trigger conditions, the voice-based modification instructions 145 may define the modifications resulting from the trigger conditions. In some examples, audio features, such as pitch, tone, volume, energy, or duration, may cause a modification to the visual elements. The audio features may be features of the user's voice and / or other sounds, such as background sounds. Below, with reference to FIGS. 5-8, some specific examples of modifications that can be performed on the visual elements in operation 414 will be described in detail. For example, for the frames of the video input 114 and / or the video item 154, the video analysis unit 124 may perform object detection (e.g., face detection) and / or object recognition (e.g., face recognition) analysis to determine the position of the visual elements 164 associated with the frames, such as body parts (e.g., facial features) or other objects. Once the position of the visual elements within the frame is determined, the modification unit 140 may modify the visual elements, for example, by changing the pixel values at the determined position of the visual elements 164 or by modifying the frame.

[0058] Referring to FIG. 5, an exemplary process 500 is described that stretches a visual element in at least one direction based on a change in at least one characteristic of a signal within an audio input. It should be noted that operations 510 and 512 in FIG. 5 are the same as operations 410 and 412 in FIG. 4. Therefore, the descriptions of operations 410 and 412 may be considered applicable to operations 510 and 512 respectively, and these descriptions are omitted here. In operation 514 of FIG. 5, a visual element is stretched in at least one direction based on a change in at least one characteristic of a signal within the audio input. The visual element may be stretched in real time when at least one characteristic of the signal within the audio input changes. As described with reference to FIG. 2, the visual element 164 may be a body part of the user 101, such as the user's face, eyes, eyebrows, mouth, lips, nose, hands, arms, etc. For this reason, the user's body part may be stretched in one or more directions (e.g., horizontal, vertical, diagonal, etc.). For example, in some cases, based on changes in the user's voice, such as pitch, tone, volume, energy, etc., the user's face including facial features such as eyes, mouth, and nose may be stretched in one or more directions. Other objects, such as objects worn by the user (e.g., glasses, hats, etc.), other background objects or foreground objects may be stretched. In one specific example, the user's face may stretch in one or more directions based on a change in the pitch of the user 101's voice. For example, in some cases, the amount of stretch of the user's face may increase as the pitch of the user's voice increases, and the amount of stretch of the user's face may decrease as the pitch of the user's voice decreases. For example, when the pitch of the user's voice reaches a selected threshold pitch value, the user's face may begin to stretch, and the amount of stretch may increase as the pitch of the user's voice increases. When the user stops speaking or the pitch of the user's voice falls below the threshold pitch value, stretching of the user's face may be stopped. For example, the overall central part of the image frame, which often includes the user's face, may be horizontally stretched such that the left and right sides of the image frame are not displayed (e.g., cropped), and only the stretched central part of the image frame is displayed.Below, with reference to FIGS. 11A - B, several exemplary user interfaces showing exemplary stretching of visual elements will be described in detail.

[0059] Referring to FIG. 6, an exemplary process 600 for swaying a visual element based on a change in at least one characteristic of a signal within an audio input will be described. It should be noted that operations 610 and 612 in FIG. 6 are the same as operations 410 and 412 in FIG. 4. Therefore, the descriptions of operations 410 and 412 may be considered applicable to operations 610 and 612 respectively, and these descriptions will be omitted here. In operation 614 of FIG. 6, a visual element is swayed based on a change in at least one characteristic of a signal within the audio input. The visual element may be swayed in real time when at least one characteristic of the signal within the audio input changes. As described with reference to FIG. 2, the visual element 164 may be a body part of the user 101, such as the user's face, eyes, eyebrows, mouth, lips, nose, hands, arms, etc. Thus, the user's body part may be swayed. For example, in some cases, based on a change in the user's voice, such as pitch, tone, volume, energy, etc., the user's face including facial features such as eyes, mouth, and nose may be swayed. Other objects, such as objects worn by the user (e.g., glasses, hats, etc.), other background objects or foreground objects may sway. In one specific example, the user's face may sway based on a change in the pitch of the user 101's voice. For example, in some cases, the amount of sway of the user's face may increase as the pitch of the user's voice increases, and the amount of sway of the user's face may decrease as the pitch of the user's voice decreases. For example, when the pitch of the user's voice reaches a selected threshold pitch value, the user's face may start to sway, and the amount of sway may increase as the pitch of the user's voice increases. When the user stops speaking or the pitch of the user's voice falls below the threshold pitch value, the swaying of the user's face may stop. Below, with reference to FIGS. 11A - B, several exemplary user interfaces showing exemplary swaying of visual elements will be described in detail.

[0060] Referring now to FIG. 7, an exemplary process 700 will be described in which at least one animated graphical element is inserted into a video item based on a change in at least one characteristic of a signal within an audio input. It should be noted that operations 710 and 712 of FIG. 7 are the same as operations 410 and 412 of FIG. 4. Therefore, the descriptions of operations 410 and 412 may be considered applicable to operations 710 and 712 respectively, and these descriptions will be omitted here. In operation 714 of FIG. 7, at least one animated graphical element is inserted into the video item based on a change in at least one characteristic of a signal within the audio input. The at least one animated graphical element may be inserted into the video item in real time upon a change in at least one characteristic of a signal within the audio input. For example, in some cases, increasing the strength of the user's voice (e.g., volume, energy, etc.) may cause the user's voice to be inserted as an animated graphical element, such as a fireball, onto a visual element, such as the user's eyeball, thereby conveying an emotion such as anger or rage. Additionally, in some examples, other animated graphical elements may be displayed. For example, the background may be modified by making animated rays or spikes appear to emanate from the user's face, similarly conveying a sense of anger or rage. In some examples, the emission of the animated rays or spikes may be synchronized with the user's speech and may increase in magnitude as the intensity of the user's voice increases. Additionally, in some cases, as the strength of the user's voice increases, the user's voice may gradually take on a metallic quality. Below, with reference to FIGS. 12A - B, several exemplary user interfaces will be described in detail that show examples of inserting animated graphical elements into a video item.

[0061] Referring now to FIG. 8, an exemplary process 800 for moving a visual element based on a change in at least one characteristic of a signal within an audio input will be described. It should be noted that operations 810 and 812 in FIG. 8 are the same as operations 410 and 412 in FIG. 4. Therefore, the descriptions of operations 410 and 412 may be considered applicable to operations 810 and 812 respectively, and these descriptions will be omitted here. In operation 814 of FIG. 8, a visual element is moved based on a change in at least one characteristic of a signal within the audio input. The visual element may be moved in real time upon a change in at least one characteristic of the signal within the audio input. For example, in some cases, the visual element may be an object within the game, such as a character controlled by the user. In some cases, a change in the pitch of the user's voice or other voice characteristics may control the movement of the visual element 164. For example, the character may move upward as the pitch of the user's voice increases, and the character may move downward as the pitch of the user's voice decreases. As another example, the character may move upward as the volume of the user's voice increases, and the character may move downward as the volume of the user's voice decreases. Hereinafter, with reference to FIGS. 13A - B, some exemplary user interfaces showing exemplary movement of a visual element based on a change in at least one characteristic of a signal within an audio input will be described in detail.

[0062] FIG. 9 shows an exemplary process 900 for audio element modification based on video input according to the present disclosure. In operation 910, during the creation of a video item, video input is received via a camera, where the video input corresponds to a visual element associated with this video item. As described above with reference to FIG. 2, video input 114 may be received via camera 107 during the creation of video item 154. Video input 114 may include or correspond to a first type of element, such as visual element 164. In some examples, visual element 164 may include body parts of user 101, such as the user's eyes, eyebrows, mouth, lips, nose, hands, arms, and the like.

[0063] In operation 912, the characteristics of the signals in the video input are determined in real time based on the video input. These characteristics may change over time in the video input. In some examples, these characteristics may include the movement of at least one body part within the video input. As described above, video analysis unit 124 may determine video signal characteristics 134 by performing video analysis on video input 114. Video signal characteristics 134 may include the calculation of the movement of at least one body part within video input 114. In some examples, video analysis unit 124 may perform video analysis on video input 114 using one or more machine learning models, such as one or more neural network models. In some examples, the video analysis may include performing object detection and / or object recognition analysis on different frames of the video input. The object detection and / or object recognition analysis may include face detection and / or face recognition analysis. In some examples, the video analysis may include determining video signal characteristics 134, such as object positions, at different times and determining and tracking the changes over time at these positions.

[0064] In operation 914, at least one modification is applied to the audio elements in the audio output associated with the video item, based at least in part on the characteristics of the signals in the video input. The at least one modification may be performed in real time when the characteristics of the signals in the video input that trigger the at least one modification are determined. As described above with reference to FIG. 2, the modification unit 140 may modify the audio elements 165 in the audio output 155 based on the video signal characteristics 134 and the video-based modification instructions 144. For example, the video-based modification instructions 144 may indicate various conditions for modifying the audio elements 165 in the audio output 155. Specifically, the video-based modification instructions 144 may indicate a change in the video signal characteristics 134 that may trigger a modification to the video element 165. In addition to defining the trigger conditions, the video-based modification instructions 144 may define the modifications resulting from the trigger conditions. In one specific example, operation 914 may include modifying an audio characteristic in the audio output, including at least one of pitch, tone, volume, energy, echo, or duration, based on the movement of at least one body part in the video input. This example will be described in detail below with reference to FIG. 10. In some examples, the audio characteristics of the user's voice, such as pitch, tone, volume, energy, echo, or duration, may be modified by modifying the audio input 115 to produce a desired audio effect. For example, in some cases, the audio analysis unit 125 may analyze the audio input 115 to detect data in the audio input corresponding to the user's voice, for example, by filtering out background noise, other non-voice audio data. Then, the modification unit 140 may obtain a desired effect by modifying the remaining data in the audio input 115 corresponding to the user's voice.

[0065] Referring to FIG. 10, an exemplary process 1000 will be described that stretches visual elements in at least one direction based on a change in at least one characteristic of the signals within the voice input. It should be noted that operation 1010 of FIG. 10 is the same as operation 910 of FIG. 9. Therefore, the description of operation 910 may be considered applicable to operation 1010, and this description will be omitted here. In operation 1012 of FIG. 10, the characteristics of the signals within the video input are determined in real time based on the video input, where the characteristics include the movement of at least one body part within the video input. As described above, the video analysis unit 124 may determine the video signal characteristics 134 by performing video analysis on the video input 114. The video signal characteristics 134 may include the calculation of the movement of at least one body part within the video input 114. In some examples, the video analysis unit 124 may perform video analysis on the video input 114 using one or more machine learning models, such as one or more neural network models. In some cases, the video analysis may include performing object detection and / or object recognition analysis on different frames of the video input. The object detection and / or object recognition analysis may include face detection and / or face recognition analysis. In some examples, the video analysis may include determining the video signal characteristics 134, such as object positions, at different times, and determining and tracking the changes over time at these positions.

[0066] In operation 1014, based on the movement of at least one body part in the video input, at least one of the pitch, tone, volume, energy, echo, or duration in the audio output is modified. When the at least one body part in the video input moves, the audio feature may be modified in real time. In one specific example, user 101 may change the pitch of the user's voice by moving their eyebrows. For example, moving the eyebrows upward may increase the pitch of the user's voice, and moving the eyebrows downward may decrease the pitch of the user's voice. In another specific example, the movement of a body part or another object may cause an echo effect. For example, raising a body part upward may enhance the echo effect, for example, converting the user's voice in real time to achieve an echo of an angelic sound. Lowering the body part may reduce the echo effect.

[0067] The following describes several exemplary user interfaces that illustrate the above-described effects. Specifically, FIGS. 11A - B show examples related to stretching and shaking of visual elements based on changes in at least one characteristic of the signals within the voice input. Specifically, FIG. 11A shows a user interface 1100 before the execution of the stretch effect and the shake effect. As shown in the illustration, the user interface 1100 shows a frame 1120 of a video item that includes the user's face 1103. The face 1103 includes a mouth 1102 and a jaw 1104. The user is wearing glasses 1101. Referring now to FIG. 11B, a user interface 1110 is shown that includes a frame 1121 of the same video to which both the stretch effect and the shake effect have been applied. As shown in frame 1121, the face 1103 is stretched horizontally. For example, the face 1103 is wider when it is in frame 1121 than when it was in frame 1120. Specifically, in frame 1121 of FIG. 11B, the glasses 1101, the mouth 1102, and the jaw 1104 (as well as other facial features) are stretched wider than when they were visible in frame 1120 of FIG. 11A. Additionally, the shake effect is applied within frame 1121. For example, in frame 1121, the glasses 1101 and the jaw 1104 appear to have a wavy shape, for example, not straight or curved, but rather having several edges with an up - down pattern. When applied to multiple frames within the video, this wavy appearance can create a shake effect where the object appears to be shaking.

[0068] The stretching and shake effects within frame 1121 may be applied based on a change in at least one characteristic of the signals within the voice input. For example, in some instances, based on a change in the user's voice, such as pitch, tone, volume, energy, etc., visual elements, such as face 1103, glasses 1101, mouth 1102, and / or jaw 1104 may be stretched in one or more directions. In one specific example, the visual element may stretch in one or more directions based on a change in the pitch of the user's voice. For example, in some cases, the amount of stretching of the user's face may increase as the pitch of the user's voice increases, and the amount of stretching of the user's face may decrease as the pitch of the user's voice decreases. Thus, when frame 1121 is generated, the pitch of the user's voice may be higher than the pitch of the user's voice when frame 1120 was generated. This increase in pitch may cause the stretching and shake effects to be applied to frame 1121.

[0069] Figures 12A - B show examples related to inserting at least one animated graphical element into a video item based on a change in at least one characteristic of the signals within the voice input. Specifically, Figure 12A shows a user interface 1200 presenting a frame 1220 of a video item that includes a user's face 1203. The face 1203 has eyes 1201A - B. The user is positioned in front of a background 1202. Referring now to Figure 12B, a user interface 1210 is shown that includes a frame 1221 of the same video where the animated graphical element insertion has been performed. As shown in frame 1221, for example, to convey a sense of anger or fury, fireballs 1211A - B are inserted into frame 1221 at the positions of eyes 1201A - B. Thus, in Figure 12B, the user's eyes 1201A - B are obscured by the fireballs 1211A - B. In some examples, the user's eyes 1201A - B may be partially visible as they are only partially obscured by the fireballs 1211A - B. Therefore, the fireballs 1211A - B modify the visual elements by obscuring the field of view of the visual elements (e.g., eyes 1201A - B) and changing their appearance. Additionally, the background 1202 is obscured by a background animation 1212. The background animation 1212 includes animated rays that radiate from the user's face 1203 and convey a sense of anger or fury. In some examples, the emission of the animated rays or spikes may be synchronized with the user's speech and may increase in magnitude as the intensity of the user's voice increases. Both the fireballs 1211A - B and the background animation 1212 may be inserted into the video item in real - time upon a change in at least one characteristic of the signals within the voice input. In some examples, increasing the strength of the user's voice (e.g., volume, energy, etc.) may cause animated graphical elements, such as fireballs 1211A - B and background animation 1212, to be inserted into the video item.Additionally, decreasing the strength of the user's voice (e.g., volume, energy, etc.) can cause the animation of graphical elements, such as fireballs 1211A - B and background animation 1212, to stop being inserted into the video item.

[0070] Figures 13A - B show examples regarding the movement of visual elements based on changes in at least one characteristic of the signals within the voice input. Specifically, Figure 13A shows a user interface 1300 that shows frame 1320 of a video item. In this example, the video item is a video game in which the user controls the movement of a character 1301 within the video game. Character 1301 is a movable sea creature that includes an image of the user's face. The game includes a sea coral 1302 that moves from right to left across the screen, and the user can gain points in the game by interacting (e.g., making contact) character 1301 with the sea coral 1302 as the sea coral 1302 moves across the screen. Referring now to Figure 13B, a user interface 1310 including frame 1321 of the same video is shown. As shown in frame 1321, character 1301 has already been moved upward from its previous position in frame 1320. In this example, character 1301 is moved based on changes in at least one characteristic of the signals within the voice input. Character 1301 may be moved in real time when there is a change in at least one characteristic of the signals within the voice input. In some cases, changes in the pitch of the user's voice or other voice characteristics may control the movement of character 1301. For example, the character may move upward as the pitch of the user's voice increases, and the character may move downward as the pitch of the user's voice decreases. As another example, the character may move upward as the volume of the user's voice increases, and the character may move downward as the volume of the user's voice decreases.

[0071] FIG. 14 shows a computing device that may be used in various ways, such as the services, networks, modules, and / or devices shown in FIG. 1. With respect to the architecture example of FIG. 1, the message service, interface service, processing service, content service, cloud network, and client may each be implemented by one or more instances of the computing device 1400 in FIG. 14. The computer architecture shown in FIG. 14 represents a conventional server computer, workstation, desktop computer, laptop computer, tablet, network device, PDA, electronic reader, digital cellular phone, or other computing node, and may be used to perform any aspect of the computers described herein, such as implementing the methods described herein.

[0072] The computing device 1400 may include a substrate or “motherboard,” which is a printed circuit board that can be connected to a plurality of components or devices via a system bus or other electrical communication path. One or more central processing units (CPUs) 1404 may operate in conjunction with a chipset 1406. The CPU 1404 may be a standard programmable processor that performs the arithmetic and logical operations necessary for the operation of the computing device 1400.

[0073] The CPU 1404 may execute by operating switching elements that distinguish and change these states to perform the operations necessary to transition from one discrete physical state to the next. The switching elements may typically include an electronic circuit that maintains one of two binary states, such as a flip-flop, and an electronic circuit that provides an output state based on a logical combination of the states of one or more other switching elements, such as a logic gate. By combining these basic switching elements, more complex logic circuits may be created, including registers, adders and subtractors, arithmetic logic units, floating-point units, and the like.

[0074] The CPU 1404 may be augmented by, or replaced with, other processing units such as the GPU 1405. The GPU 1405 may include, but is not necessarily limited to, processing units specialized for high - degree parallel computing such as graphics and other visualization - related processing.

[0075] The chipset 1406 may provide an interface between the CPU 1404 and the remaining components and devices on the substrate. The chipset 1406 may provide an interface to the random - access memory (RAM) 1408 used as the main memory within the computing device 1400. The chipset 1406 may also provide an interface to a computer - readable storage medium, such as a read - only memory (ROM) 1420 or non - volatile RAM (NVRAM) (not shown), to store basic routines that can start up the computing device 1400 and facilitate the transmission of information between various components and devices. According to the aspects described herein, the ROM 1420 or NVRAM may store other software components necessary for the operation of the computing device 1400.

[0076] The computing device 1400 may operate in a network environment using logical connections to remote computing nodes and computer systems via a local - area network (LAN). The chipset 1406 may include functionality to provide a network connection via a network - interface controller (NIC) 1422, such as a Gigabit Ethernet (registered trademark) adapter. The NIC 1422 may be capable of connecting the computing device 1400 to other computing nodes via the network 1416. It should be understood that multiple NICs 1422 may be present within the computing device 1400 to connect the computing device to other types of networks and remote computer systems.

[0077] Computing device 1400 may be connected to a mass storage device 1428 that provides a non-volatile storage device for a computer. The mass storage device 1428 may store system programs, application programs, other program modules, and data described in more detail herein. The mass storage device 1428 may be connected to the computing device 1400 via a storage controller 1424 connected to the chipset 1406. The mass storage device 1428 may be composed of one or more physical storage units. The mass storage device 1428 may include a management unit 1410. The storage controller 1424 may interface with the physical storage unit via a Serial Attached SCSI (SAS) interface, a Serial Advanced Technology Attachment (SATA) interface, a Fibre Channel (FC) interface, or other types of interfaces for physically connecting and transmitting data between the computer and the physical storage unit.

[0078] The computing device 1400 may store data on the mass storage device 1428 by converting the physical state of the physical storage unit to reflect the stored information. The specific conversion of the physical state may depend on various factors and different embodiments of this specification. Examples of such factors include, but are not limited to, the technology for implementing the physical storage device and whether the mass storage device 1428 is characterized as a primary storage device or a secondary storage device, etc.

[0079] For example, the computing device 1400 may store information in the mass storage device 1428 by issuing commands via the storage controller 1424 to change the magnetic characteristics at a specific location within a magnetic disk drive unit, the reflective or refractive characteristics at a specific location within an optical storage unit, or the electrical characteristics of a specific capacitor, transistor, or other discrete component within a solid-state storage unit. Without departing from the scope and spirit of this specification, other conversions of the physical medium are possible, and the foregoing examples are provided solely for the purpose of facilitating its explanation. The computing device 1400 may further read information from the mass storage device 1428 by detecting the physical state or characteristics of one or more specific locations within the physical storage unit.

[0080] In addition to the mass storage device 1428 described above, the computing device 1400 may be able to access other computer-readable storage media for storing and retrieving information such as program modules, data structures, or other data. Those skilled in the art should understand that a computer-readable storage medium can provide non-transitory data storage and can be any available medium accessible by the computing device 1400.

[0081] By way of example and not limitation, computer-readable storage media may include volatile and non-volatile, temporary and non-temporary computer-readable storage media, as well as removable and non-removable media implemented in any method or technology. Computer-readable storage media includes, but is not limited to, RAM, ROM, erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory or other solid-state memory technologies, compact disc ROM (CD-ROM), digital versatile disc (DVD), high-definition DVD (“HD-DVD”), BLU-RAY (registered trademark) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage, other magnetic storage devices, or any other medium that can be used to store desired information in a non-temporary manner.

[0082] A mass storage device such as the mass storage device 1428 shown in FIG. 14 may store an operating system for controlling the operation of the computing device 1400. The operating system may include one version of the LINUX (registered trademark) operating system. The operating system may include one version of the MICROSOFT WINDOWS (registered trademark) SERVER operating system. According to another aspect, the operating system may include one version of the UNIX (registered trademark) operating system. Also, various mobile phone operating systems such as IOS (registered trademark) and ANDROID (registered trademark) may be used. It should be understood that other operating systems may also be used. The mass storage device 1428 may store other systems or applications and data used by the computing device 1400.

[0083] The mass storage device 1428 or other computer-readable storage media may also be encoded with computer-executable instructions that, when loaded into the computing device 1400, convert the computing device from a general-purpose computing system to a dedicated computer capable of implementing the aspects described herein. As described above, these computer-executable instructions convert the computing device 1400 by specifying how the CPU 1404 transitions between states. The computing device 1400 may be able to access a computer-readable storage medium that stores computer-executable instructions that can execute the methods described herein when executed by the computing device 1400.

[0084] A computing device, such as the computing device 1400 shown in FIG. 14, may further include an input / output controller 1432 for receiving and processing inputs from a plurality of input devices such as a keyboard, mouse, touchpad, touch screen, electronic stylus pen, or other types of input devices. Similarly, the input / output controller 1432 may provide outputs to a display such as a computer monitor, flat panel display, digital projector, printer, plotter, or other types of output devices. It should be understood that the computing device 1400 may not include all of the components shown in FIG. 14, may include other components not explicitly shown in FIG. 14, or may utilize an architecture that is completely different from the architecture shown in FIG. 14.

[0085] As described herein, the computing device may be a physical computing device such as the computing device 1400 of FIG. 14. The computing node may also include a virtual machine host process and one or more virtual machine instances. The computer-executable instructions may be executed indirectly by the physical hardware of the computing device by interpreting and / or executing instructions stored and executed within the context of the virtual machine.

[0086] It should be understood that the methods and systems are not limited to a particular method, particular component, or particular embodiment. It should also be understood that the terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting.

[0087] As used in the specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Ranges may be expressed herein as from "about" one particular value and / or to "about" another particular value. When such a range is expressed, another embodiment includes from the one particular value and / or to the other particular value. Similarly, when values are expressed by use of the antecedent "about," it should be understood that the particular value forms another embodiment. Furthermore, it should be understood that each end point of a range is significant both in relation to the other end point and independently of the other end point.

[0088] "Any" or "optionally" means that the subsequent event or circumstance may or may not occur, and the specification includes both the case where the event or circumstance occurs and the case where it does not occur.

[0089] Throughout the description and claims of this specification, the words "comprise" and variations of the word, such as "comprising," "include," mean "including, but not limited to," and are not intended to exclude, for example, other components, integers, or steps. "Exemplary" represents "an example of" and is not intended to convey an indication of a preferred or desired embodiment. "Such as" is used for interpretive purposes and not for a limiting meaning.

[0090] Components that can be used to implement the described methods and systems are described. When describing combinations, subsets, interactions, groups, etc. of these components, specific references to each of the various individual and collective combinations and permutations of these components may not be explicitly described, and it should be understood that each of them is specifically contemplated and described herein for all methods and systems. This applies to all aspects of the present application, including but not limited to operations in the described methods. Thus, if there are various additional executable operations, it should be understood that each of these additional operations is executable in any particular embodiment or combination of embodiments of the described method.

[0091] The present methods and systems can be more readily understood by reference to the following detailed description of the preferred embodiments and examples included therein, as well as the accompanying drawings and their descriptions.

[0092] As will be understood by those skilled in the art, the methods and systems may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Further, the present methods and systems may take the form of a computer program product on a computer-readable storage medium having computer-readable program instructions (e.g., computer software) embodied therein. More specifically, the present methods and systems may take the form of web-implemented computer software. Any suitable computer-readable storage medium, including a hard disk, CD-ROM, optical storage device, or magnetic storage device, may be utilized.

[0093] Referring to the block diagrams and flowcharts of the method, system, apparatus, and computer program product, embodiments of the method and system will be described below. It should be understood that each block of the block diagrams and flowcharts, and combinations of blocks in the block diagrams and flowcharts, may be implemented by computer program instructions. These computer program instructions may be loaded onto a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed on the computer or other programmable data processing apparatus generate means for realizing the functions specified in one or more blocks of the flowchart.

[0094] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including computer-readable instructions for realizing the functions specified in one or more blocks of the flowchart. The computer program instructions may be loaded onto a computer or other programmable data processing apparatus to provide steps for realizing the functions specified in one or more blocks of the flowchart, and to execute a series of operation steps on the computer or other programmable data processing apparatus to generate a computer-implemented process.

[0095] The various characteristics and processes described above may be used independently of each other or may be combined in various ways. All possible combinations and sub-combinations are intended to fall within the scope of this disclosure. Further, in some implementations, some methods or process blocks may be omitted. The methods and processes described herein are not limited to any particular order, and the associated blocks or states may be executed in other suitable orders. For example, the described blocks or states may be executed in an order other than the specifically described order, or multiple blocks or states may be combined within a single block or state. The exemplary blocks or states may be executed sequentially, in parallel, or in some other manner. Blocks or states may be added to or removed from the described exemplary embodiments. The exemplary systems and components described herein may be configured differently from those described. For example, elements may be added, removed, or rearranged compared to the described exemplary embodiments.

[0096] Also, various items are shown to be stored in memory or on a storage device during use, and it should be understood that these items or portions thereof may be transferred between memory and other storage devices for memory management and data integrity purposes. Alternatively, in other embodiments, software modules and / or portions or all of the system may be executed in memory on another device and communicate with the illustrated computing system via inter-computer communication. Further, in some embodiments, portions or all of the system and / or module may be implemented or provided in other ways, such as at least partially in firmware and / or hardware, where the hardware includes, but is not limited to, one or more application specific integrated circuits (ASICs), standard integrated circuits, controllers (e.g., including microcontrollers and / or embedded controllers by executing appropriate instructions), field programmable gate arrays (FPGAs), complex programmable logic devices (CPLDs), etc. Portions or all of the modules, systems, and data structures may be stored (e.g., as software instructions or structured data) on a computer-readable medium such as a hard disk, memory, network, or portable media product for reading by an appropriate device or via an appropriate connection. The systems, modules, and data structures may be transmitted as generated data signals (e.g., as part of a carrier wave or other analog or digital propagation signal) on various computer-readable transmission media including wireless-based media and wired / cable-based media, and may take various forms (e.g., as part of a single or multiplexed analog signal or as multiple discrete digital packets or frames). In other embodiments, such computer program products may take other forms. Accordingly, the present invention can be implemented in other computer system configurations.

[0097] Although methods and systems have been described in connection with preferred embodiments and specific examples, the embodiments herein are intended to be illustrative rather than limiting in all respects, and it is not intended to limit the scope to the specific embodiments.

[0098] Unless otherwise specified, the methods described herein do not require that their operations be performed in a particular order. Thus, where method claims do not actually recite an order for the operations to follow, or where the operations are not specifically recited in the claims or the specification as being limited to a particular order, no inference of order is intended in any aspect. This applies to any possible non-expressive basis for interpretation, including logical issues regarding the arrangement of steps or the operation flow, the plain meaning obtained from grammatical construction and punctuation, the number or type of embodiments described in the specification.

[0099] As will be apparent to those skilled in the art, various modifications and changes are possible without departing from the scope or spirit of the present disclosure. Considering the present specification and the practice described herein, other embodiments will be apparent to those skilled in the art. The present specification and the exemplary drawings are intended to be considered only as illustrative, and the true scope and spirit are indicated by the following claims.

Claims

1. During the creation of a video item, receiving, via a first component of a computing device, a first type of input corresponding to a first type of element associated with the video item; Based on the first type of input, determining, in real time, characteristics of signals within the first type of input that change over time in the first type of input; Applying at least one modification to a second type of element associated with the video item, at least partially based on the characteristics of the signals within the first type of input; A method comprising the above.

2. Receiving the first type of input via the first component of the computing device includes receiving audio input via a microphone, The second type of element includes visual elements within the video item, The method according to claim 1.

3. The characteristics include sound features, and the sound features include at least one of pitch, tone, volume, energy, or duration, The method according to claim 2.

4. Applying the at least one modification to the second type of element associated with the video item, at least partially based on the characteristics of the signals within the first type of input, Includes stretching the visual element in at least one direction based on a change in at least one of the characteristics of the signals within the audio input, The method according to claim 2.

5. Applying the at least one modification to the second type of element associated with the video item, at least partially based on the characteristics of the signals within the first type of input, Includes swaying the visual element based on a change in at least one of the characteristics of the signals within the audio input, The method according to claim 2.

6. Applying the at least one modification to the second type of element associated with the video item, at least partially based on the characteristics of the signals within the first type of input, Includes inserting at least one animated graphical element into the video item based on a change in at least one of the characteristics of the signals within the audio input, The method according to claim 2.

7. Based at least in part on the characteristics of the signal within the first type of input, adding the at least one modification to the second type of element associated with the video item is including moving the visual element based on a change in at least one of the characteristics of the signal within the audio input, The method according to claim 2.

8. Receiving the first type of input via the first component of the computing device includes receiving video input via a camera, The second type of element includes an audio element within an audio output associated with the video item, The method according to claim 1.

9. The characteristics include movement of at least one body part within the video input, The method according to claim 8.

10. Based at least in part on the characteristics of the signal within the first type of input, adding the at least one modification to the second type of element associated with the video item is including modifying a sound characteristic in the audio output, including at least one of pitch, tone, volume, energy, echo, or duration, based on the movement of the at least one body part within the video input, The method according to claim 9.

11. A system comprising one or more computer processors and one or more computer memories including computer-readable instructions, when the computer-readable instructions are executed by the one or more computer processors, during creation of a video item, receiving, via a first component of a computing device, a first type of input corresponding to a first type of element associated with the video item; determining, in real time, characteristics of a signal within the first type of input that change over time in the first type of input, based on the first type of input; adding at least one modification to a second type of element associated with the video item, based at least in part on the characteristics of the signal within the first type of input; A system configured to perform operations including.

12. Receiving the first type of input via the first component of the computing device includes receiving voice input via a microphone, The second type of element includes visual elements within the video item, The system according to claim 11.

13. The characteristic includes a sound characteristic, and the sound characteristic includes at least one of pitch, tone, volume, energy, or duration, The system according to claim 12.

14. Adding the at least one modification to the second type of element associated with the video item, at least in part based on the characteristic of the signal in the first type of input, Stretching the visual element in at least one direction based on a change in at least one of the characteristics of the signal in the voice input, Inserting at least one animated graphical element into the video item based on a change in at least one of the characteristics of the signal in the voice input, or Moving the visual element based on a change in at least one of the characteristics of the signal in the voice input, The system according to claim 12, including at least one of.

15. Receiving the first type of input via the first component of the computing device includes receiving video input via a camera, and the second type of element includes a voice element within a voice output associated with the video item, The system according to claim 11.

16. Adding the at least one modification to the second type of element associated with the video item, at least in part based on the characteristic of the signal in the first type of input, Modifying a sound characteristic in the voice output, including at least one of pitch, tone, volume, energy, echo, or duration, based on the movement of at least one body part in the video input, The system according to claim 15.

17. A non-transitory computer-readable storage medium storing computer-readable instructions, The computer-readable instructions, when executed by a processor, cause the processor to, During the creation of a video item, receiving, via a first component of the computing device, a first type of input corresponding to a first type of element associated with the video item; Based on the first type of input, determining, in real time, characteristics of a signal within the first type of input that vary over time in the first type of input; Applying at least one modification to a second type of element associated with the video item, based at least in part on the characteristics of the signal within the first type of input; A non-transitory computer-readable storage medium that causes an operation including the above to be executed.

18. Receiving the first type of input via the first component of the computing device includes receiving audio input via a microphone, and the second type of element includes visual elements within the video item. The non-transitory computer-readable storage medium according to claim 17.

19. Applying the at least one modification to the second type of element associated with the video item, based at least in part on the characteristics of the signal within the first type of input, includes: Stretching the visual element in at least one direction based on a change in at least one of the characteristics of the signal within the audio input; Inserting at least one animated graphical element into the video item based on a change in at least one of the characteristics of the signal within the audio input; or Moving the visual element based on a change in at least one of the characteristics of the signal within the audio input. The non-transitory computer-readable storage medium according to claim 18, including at least one of the above.

20. Receiving the first type of input via the first component of the computing device includes receiving video input via a camera, and the second type of element includes audio elements within an audio output associated with the video item. The non-transitory computer-readable storage medium according to claim 17.

21. Applying the at least one modification to the second type of element associated with the video item, based at least in part on the characteristics of the signal within the first type of input, includes: Modifying a sound feature in the voice output, including at least one of pitch, tone, volume, energy, echo, or duration, based on movement of at least one body part within the video input. The non-transitory computer-readable storage medium according to claim 20.

Citation Information

Patent Citations

  • Playing scene synthesis enhancement control method and device

    CN106648083A

  • Video recording method and electronic equipment

    CN114390341A

  • Information processing device, information processing method, information processing program, and information processing system

    JP2003125361A

  • Automatic animation production system

    JP2005346721A

  • Dynamic custom interstitial transition video for video streaming services

    JP2020523810A