Method, system, computer-readable storage medium and program for automatic and fast generation of music audio content for video

The system addresses content creator challenges by automatically generating music that matches video characteristics, enhancing efficiency and compliance with licensing, thus providing high-quality audio without user effort.

JP7766814B2Active Publication Date: 2025-11-10LEMON CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024544730
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-07-27
Filing Date
2023-07-03
Publication Date
2025-11-10
Estimated Expiration
2043-07-03

AI Technical Summary

Technical Problem

Content creators face challenges in selecting appropriate music for their videos, as existing methods are inefficient, time-consuming, and often violate music licensing agreements, while creating music without professional equipment is discouraged.

Method used

A system that automatically generates music audio matching video transitions, intensity, and movements, using models to extract information from videos and generate musical notes, vectors, and audio, ensuring royalty-free content of any length.

Benefits of technology

Efficiently generates high-quality music that fits video duration and style, reducing user effort and compliance with licensing issues.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007766814000001
    Figure 0007766814000001
  • Figure 0007766814000002
    Figure 0007766814000002
  • Figure 0007766814000003
    Figure 0007766814000003
Patent Text Reader

Abstract

This disclosure describes techniques for automatically and rapidly generating music for videos. The techniques include receiving a video from a user. The video may include a plurality of frame segments. Information may be extracted from the video including information indicative of a rate of motion in the video, information indicative of motion saliency in the video, information indicative of scene transitions in the video, and timing information associated with the video. A plurality of musical note sets may be generated matching the plurality of frame segments based at least in part on the extracted information. A plurality of vectors corresponding to the plurality of musical note sets may be generated. A plurality of musical audio corresponding to the plurality of frame segments may be generated based at least in part on the plurality of vectors.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims priority to U.S. Application No. 17 / 815,402, filed July 27, 2022, entitled "Automatic and Rapid Generation of Music Audio Content for Video," the entire contents of which are incorporated herein by reference in their entirety. [Background technology]

[0002] Increasingly, communication is conducted using internet-based tools. Internet-based tools can be any software or platform. Existing social media platforms allow users to communicate with each other by sharing information such as images and videos through static applications or web pages. As communication devices such as mobile phones become increasingly sophisticated, people continue to seek new ways of entertainment, social networking, and communication. [Brief explanation of the drawings]

[0003] The following detailed description can be better understood when read in conjunction with the accompanying drawings, in which: For purposes of illustration, there are shown in the accompanying drawings exemplary embodiments of various aspects of the disclosure; however, the invention is not limited to the specific methods and instrumentalities disclosed.

[0004] [Figure 1] FIG. 1 illustrates an exemplary system for distributing content in accordance with the present disclosure.

[0005] [Figure 2] FIG. 1 is an exemplary diagram illustrating information extracted from a video according to the present disclosure.

[0006] [Figure 3] FIG. 10 is an exemplary diagram illustrating musical notes generated based on information extracted from a video in accordance with the present disclosure.

[0007] [Figure 4] FIG. 10 is an exemplary diagram illustrating vectors corresponding to generated musical notes in accordance with the present disclosure.

[0008] [Figure 5] FIG. 10 is an exemplary diagram illustrating music audio generated based on vectors corresponding to musical notes according to the present disclosure.

[0009] [Figure 6] FIG. 1 illustrates an exemplary process for automatically generating music audio for a video in accordance with the present disclosure.

[0010] [Figure 7] FIG. 10 illustrates another exemplary process for automatically generating music audio for a video in accordance with the present disclosure.

[0011] [Figure 8] FIG. 10 illustrates another exemplary process for automatically generating music audio for a video in accordance with the present disclosure.

[0012] [Figure 9] FIG. 10 illustrates another exemplary process for automatically generating music audio for a video in accordance with the present disclosure.

[0013] [Figure 10] FIG. 10 illustrates another exemplary process for automatically generating music audio for a video in accordance with the present disclosure.

[0014] [Figure 11] FIG. 1 illustrates an exemplary computing device that can be used to perform any of the methods disclosed herein. DETAILED DESCRIPTION OF THE INVENTION

[0015] Users of content creation platforms may struggle to select music to be featured within their content. For example, a user of a content creation platform may be creating a video to be shared on the content creation platform. However, the user may find it difficult to select an appropriate song to play within the video. For example, the user may need to continually re-shoot clips of the video to match the beat and / or style of the selected song. This process may be frustrating, inefficient, and time-consuming. Additionally, it may be difficult for users to find a song that fits the length of the video. For example, a user may be creating a 10-minute video and may struggle to find a song for a video that lasts at least 10 minutes. For business-oriented content creators (i.e., creating content to promote their business), it may be difficult to find high-quality music that does not violate region-specific music licensing agreements. Alternatively, content creators may attempt to create their own music to accompany their videos. However, creators who do not have professional equipment may be discouraged from undertaking music creation or may be unable to produce high-quality music.

[0016] Therefore, improvements in content creation technology, especially technology for music generation, are desirable. This specification describes technology that enables efficient smart music generation and professional music editing. After a content creator creates a video, music or audio is automatically generated that matches the transitions, intensity, and movements in the video. Such music or audio may be royalty-free and of any length.

[0017] The music generation techniques described herein may be utilized by a system for distributing content. Figure 1 shows an exemplary system 100 for distributing content. The system 100 may include a server 102 and a plurality of client devices 104a-104n. The server 102 and the plurality of client devices 104a-104n may communicate with each other via one or more networks 132.

[0018] The server 102 may be located in a data center, such as a single building, or may be distributed across different geographic locations (e.g., several buildings). The server 102 may provide services via one or more networks 132. The network 132 may include various network devices, such as routers, switches, multiplexers, hubs, modems, bridges, repeaters, firewalls, proxy devices, and / or similar devices. The network 132 may include physical links, such as coaxial cable links, twisted pair cable links, optical fiber links, or combinations thereof. The network 132 may also include wireless links, such as cellular links, satellite links, Wi-Fi links, etc.

[0019] The server 102 may include multiple computing nodes hosting various services. In one embodiment, the nodes host a content service 112. The content service 112 may include a content streaming service, such as an Internet Protocol video streaming service. The content service 112 may be configured to deliver content 123 via various transmission technologies. The content service 112 may be configured to provide content 123, such as video, audio, text data, or a combination thereof. The content 123 may include content streams (e.g., video streams, audio streams, information streams), content files (e.g., video files, audio files, text files), and / or other data. The content 123 may be stored in a database 122. For example, the content service 112 may include a video sharing service, a video hosting platform, a content distribution platform, a collaborative gaming platform, etc.

[0020] In one embodiment, the content 123 distributed or provided by the content service 112 includes video. The video may have a duration up to a predetermined time limit, such as one minute, five minutes, or other predetermined number of minutes. By way of example and not limitation, a video may include at least one and no more than four 15-second segments joined together. Short video durations can provide viewers with entertainment in rapid succession, allowing users to watch a large amount of video within a short time frame. Such rapid succession of entertainment can be popular on social media platforms.

[0021] A video may include a pre-recorded audio overlay, such as music or sounds from a television program or movie. When a short video includes a pre-recorded audio overlay, the short video may feature one or more people lip-syncing, dancing, or otherwise moving their bodies along with the pre-recorded audio. For example, a short video may feature a "dance challenge" that an individual completes to a hit song, or the short video may feature two people participating in a lip-sync or two-person dance. As another example, a short video may feature an individual completing a challenge that requires them to move their body to correspond to the pre-recorded audio overlay, e.g., to the beat or rhythm of a pre-recorded song that is characterized by the pre-recorded audio overlay. Other videos may not include a pre-recorded audio override. For example, these videos may feature individuals playing sports, playing pranks, or giving beauty and fashion advice, cooking tips, home decorating tips, etc.

[0022] In one embodiment, the content 123 may be output to different client devices 104 over the network 132. The content 123 may be streamed to the client devices 104. The content stream may be a stream of video received from the content service 112. Multiple client devices 104 may be configured to access the content 123 from the content service 112. In one embodiment, the client devices 104 may include a content application 106. The content application 106 outputs (e.g., displays, renders, presents) the content 123 to a user associated with the client device 104. The content may include video, audio, comments, text data, etc.

[0023] The multiple client devices 104 may include any type of computing device, such as a mobile device, a tablet device, a laptop computer, a desktop computer, a smart television or other smart device (e.g., a smart watch, a smart speaker, smart glasses, a smart helmet), a gaming device, a set-top box, a digital streaming device, a robot, etc. The multiple client devices 104 may be associated with one or more users. A single user may access the server 102 using one or more of the multiple client devices 104. The multiple client devices 104 may travel to different locations and access the server 102 using different networks.

[0024] The content service 112 may be configured to receive input from users. The users may be registered users of the content service 112 or may be users of a content application 106 running on a client device 104. User input may include a video created by the user, a user comment associated with the video, or a "like" associated with the video. User input may include a connection request and user input data, such as text data, digital image data, or user content. A connection request may include a request from a client device 104a-d to connect to the content service 112. User input data may include information, such as a video and / or user comments, that a user connected to the content service 112 desires to share with other connected users of the content service 112.

[0025] The content service 112 may be able to receive different types of input from users using different types of client devices 104. For example, a user using a content application 106 on a first user device, such as a mobile phone or tablet, may be able to create and upload videos using the content application 106. A user using the content application 106 on a different mobile phone or tablet may be able to view, comment on, and "like" videos or comments written by other users. In another example, a user using the content application 106 on a smart TV, laptop, desktop, or gaming device may not be able to create and upload videos or comment on videos using the content application 106. Instead, a user using the content application 106 on a smart TV, laptop, desktop, or gaming device may only be able to use the content application 106 to watch videos, view comments left by other users, and "like" videos.

[0026] In one embodiment, a user may use a content application 106 on a client device 104 to create a video, e.g., a short video, and upload the video to the server 102. The client device 104 may access an interface 108 of the content application 106. The interface 108 may include input elements. For example, the input elements may be configured to allow the user to create the video. To create the short video, the user may grant the content application 106 permission to access an image capture device, such as a camera or microphone, of the client device 104. Using the content application 106, the user may select the duration of the video or set the speed of the video, e.g., "slow motion" or "speed up."

[0027] A user may edit a video using the content application 106. A user may add one or more text, filters, sounds, or effects, such as beauty effects, to a video. To add a pre-recorded audio overlay to a video, a user may select a song or sound clip from the content application 106's sound library. The sound library may include different songs, sound effects, or audio clips from movies, albums, and television shows. In addition to or instead of adding a pre-recorded audio overlay to a video, a user may use the content application 106 to add narration to a video. The narration may be sound recorded by the user using the client device 104's microphone. A user can add text overlays to a short video and may use the content application 106 to specify when they want the text overlay to appear in the video. A user may assign a caption, a location tag, and one or more hashtags to a video to indicate the subject of the video. The content application 106 may prompt the user to select a frame of the video to use as a "cover image" for the video.

[0028] After a user creates a video, the user may use the content application 106 to upload the video to the server 102 and / or store the video locally on the user device 104. When a user uploads a video to the server 102, the user may select whether the video should be viewable by all other users of the content application 106 or by only a subset of users of the content application 106. The content service 112 may store the uploaded video and any metadata associated with the video in one or more databases 122.

[0029] In one embodiment, a user may provide input on a video using a content application 106 on a client device 104. The client device 104 may access an interface 108 of the content application 106 that allows the user to provide input associated with the video. The interface 108 may include an input element. For example, the input element may be configured to receive input from the user, such as a comment or a "like" associated with a particular video. If the input is a comment, the content application 106 may allow the user to set an emoji to be associated with their input. The content application 106 may determine time information about the input, such as when the user wrote the comment. The content application 106 may transmit the input and associated metadata to the server 102. For example, the content application 106 may transmit the comment, an identifier of the user who wrote the comment, and time information about the comment to the server 102. The content service 112 may store the input and associated metadata in a database 122.

[0030] The content service 112 may be configured to output uploaded videos and user input to other users. A user may register as a user of the content service 112 and view videos created by other users. A user may be a user of a content application 106 running on a client device 104. The content application 106 may output (display, render, present) videos and user comments to a user associated with the client device 104. The client device 104 may access an interface 108 of the content application 106. The interface 108 may include an output element. The output element may be configured to display information about different videos so that a user can select and watch the videos. For example, the output element may be configured to display multiple cover images, subtitles, or hashtags associated with the videos. The output element may also be configured to arrange the videos according to a category associated with each video.

[0031] In one embodiment, user comments associated with a video may be output to other users viewing the same video. For example, all users accessing the video may see the comments associated with the video. The content service 112 may output the comments associated with the video simultaneously. The comments may be output by the content service 112 in real time or near real time. The content application 106 may display the video and comments on the client device 104 in various ways. For example, the comments may be displayed in an overlay on top of the content or next to the content. As another example, a user who wants to see other users' comments associated with the video may need to select a button to view the comments. The comments may be animated when displayed. For example, the comments may scroll across the video or across an overlay.

[0032] As described above, a user may create a video and upload it to the server 102 using a content application 106 on a client device 104. In one embodiment, a video created by a user via the content application 106 on a client device 104 may be a video that does not include pre-recorded audio overlays, such as pre-recorded song clips or audio from a television program or movie. Alternatively, music may be automatically generated for a video after the user has already created the video. For example, music audio may be automatically generated for a video locally on the client device 104 after the user has already created the video but before the user uploads it to the server 102. Additionally or alternatively, music audio may be automatically generated for a video by the content service 112 after the user has already uploaded the video to the server 102. The music audio may be generated using, for example, an extraction model 117, a note generation model 118, a vector generation model 119, an audio generation model 120, and / or a refinement model 121.

[0033] In one embodiment, at least one of the content service 112 or the client device 104 includes an extraction model 117. The extraction model 117 may be utilized, at least in part, to retrieve (e.g., determine, extract, etc.) information from user-created videos. For example, the extraction model 117 may be utilized to extract information associated with video motion speed, video motion saliency (i.e., the visibility of objects in the video), scene transitions, and / or video timing. For example, if a video depicts a man walking, the extraction model 117 may be utilized to extract information associated with the man's walking speed, whether / for how long he stops, etc. The information extracted from the user-created videos by the extraction model 117 may be stored in the database 124, for example, as extracted data 125.

[0034] FIG. 2 shows an example diagram 200 illustrating information extracted from a video 202. The video 202 may include multiple segments 204a-d. The segments 204a-d may include one or more video frames from the video 202. The extraction model 117 may extract information associated with video motion speed within the video 202 from the entire video 202 (e.g., from all segments 204a-d). The extraction model 117 may additionally extract information associated with video motion saliency within the video 202 from the video 202. The extraction model 117 may additionally extract information associated with scene transitions within the video 202 from the video 202. The extraction model 117 may additionally extract information associated with timing within the video 202 from the video 202. For example, the extraction model 117 may determine, for each item of information extracted from the video 202, at which time in the video 202 the information was extracted. For example, if video 202 has a time interval of 60 seconds, extraction model 117 may extract information associated with video motion saliency, video motion speed, and / or scene transitions at a particular time, such as 30 seconds. Timing information associated with this information item may indicate that it was extracted from video 202 at the 30-second mark. The extracted information associated with video motion speed, video motion saliency, scene transitions, and / or timing may collectively be referred to as extracted information 204. Extracted information 204 may be stored in database 124, for example, as extracted data 125.

[0035] Returning to FIG. 1 , in one embodiment, at least one of the content service 112 or the client device 104 includes a musical note generation model 118. The musical note generation model 118 may be utilized to automatically generate musical notes that correspond, at least in part, to the extracted data 125. For example, the musical note generation model 118 may utilize information extracted from the video by the extraction model 117 to automatically generate a set of musical notes for the video for each segment of the video. For example, the musical note generation model 118 may utilize information associated with the video's motion rate, motion saliency (i.e., the visibility of objects in the video), scene transitions, and / or timing of the video to automatically generate musical notes for the video.

[0036] In an embodiment, to generate musical notes for a video, musical note generation model 118 may retrieve extracted data 125 from database 124. Extracted data 125 may be fed into a trained model, such as a compound word transformation model. The model may be pre-trained and configured to correlate video motion speed with musical note density, video motion saliency with musical note intensity, video scene transitions with musical composition, and video timing with musical beats. For example, the model may receive extracted data 125 as input. The model may output musical notes (e.g., musical notation) associated with a particular note density, note intensity, musical composition, and / or musical beat. The generated notes may be stored in database 126, for example, as generated musical note data 127.

[0037] In an embodiment, the model is a compound word conversion model trained on a large number of MIDI files (e.g., 3000 MIDI files). Additional parameters may be added during the training process to allow for control of music generation. The compound word conversion model may utilize different feedforward heads to model different types of tokens. Different types of tokens may include, for example, note types and metric types. Using an expansion / compression technique, the compound word conversion converts music into compound word sequences by grouping adjacent tokens, significantly reducing the length of the token sequence.

[0038] FIG. 3 is an exemplary diagram 300 illustrating musical notes generated based on information extracted from a video. As described above with respect to FIG. 2, the video 202 may include one or more segments 204a-d of video frames. The extraction model 117 may extract information associated with video motion speed, video motion saliency, scene transitions, and / or timing within each of the segments 204a-d from the video 202. The extracted information associated with video motion speed, video motion saliency, scene transitions, and / or timing may collectively be referred to as extracted information 204. The musical note generation model 118 may generate musical notes having a note density correlating with each item of extracted information associated with video motion speed in the video 202. The musical notes generated by the musical note generation model 118 may additionally have a note intensity correlating with each item of extracted information associated with video motion saliency in the video 202. Additionally, the notes generated by the musical note generation model 118 may have a note configuration that correlates with each item of extracted information associated with video scene transitions in the video 202. Finally, the notes generated by the musical note generation model 118 may have a musical beat that correlates with each item of extracted information associated with timing in the video 202. Notes with corresponding note densities, note intensities, musical configurations, and / or musical beats may be collectively referred to as generated notes.

[0039] In an embodiment, the generated notes may be divided into, for example, multiple note sets 304a-d. Each of the note sets 304a-d may correspond to one of the segments 204a-d of the video 202. For example, note set 304a may correspond to segment 204a of the video 202, note set 304b may correspond to segment 204b of the video 202, note set 304c may correspond to segment 204c of the video 202, and note segment 304d may correspond to segment 204d of the video 202. The multiple note sets 304a-d may be stored in database 126 as note data 127.

[0040] Returning to FIG. 1 , in one embodiment, at least one of the content service 112 or the client device 104 includes a vector generation model 119. The vector generation model 119 may be utilized, at least in part, to automatically generate vectors corresponding to the musical note data 127. For example, the vector generation model 119 may utilize the musical note data 127 generated by the musical note generation model 118 to automatically generate vectors for each of the plurality of musical note sets 304a-d. Each vector may indicate at least one musical characteristic of the musical audio associated with the corresponding musical note set. For example, the at least one musical characteristic may include a musical style (e.g., musical genre and / or mood), a bar structure, or an instrument. A musical genre may indicate whether the musical audio associated with the corresponding musical note set is closest to pop, rock, hip hop, country, etc. A musical mood may indicate whether the musical audio associated with the corresponding musical note set has a particular energy (e.g., lively, sad, slow, upbeat, etc.).

[0041] In an embodiment, to generate the vectors, the vector generation model 119 may retrieve musical note data 127 from the database 126. The vector generation model 119 may be a trained model configured to determine musical characteristics based on musical notes. For example, the vector generation model 119 may receive the musical note data 127 as input. The vector generation model 119 may output a vector (e.g., a one-dimensional data row or a multi-dimensional vector) indicating the musical characteristics of the input musical note data 127. The vector may have a format consumable by, for example, the audio generation model 120. The generated vector may be stored in the database 130, for example, as vector data 131.

[0042] In embodiments, a user may be able to specify one or more musical preferences. Such preferences may be utilized by the vector generation model 119 when generating one or more vectors. For example, a user may indicate that they would like upbeat music to accompany their video. As such, the user may specify this preference, and the vector generation model 119 may utilize the user-defined preference in addition to, or as an alternative to, the plurality of musical note sets 304a-d. The preference may indicate, for example, the genre, mood, style, and / or one or more instruments that the user would like the generated music to reflect. The preference may harmonize or clash with the plurality of musical note sets 304a-d. For example, one set of musical notes may indicate a slow or sad energy. Nevertheless, the user may indicate that they would like musical content for the video to have an upbeat or happy energy. Thus, the vector may reflect the user-selected preference rather than the preference indicated by the musical notes.

[0043] FIG. 4 shows an example diagram 400 illustrating vectors generated based on musical notes. As described above with respect to FIG. 3, multiple sets of notes 304a-d may be generated. Each of the sets of notes 304a-d may correspond to one of the segments 204a-d of the video 202. The vector generation model 119 may generate multiple vectors 402a-d. Each of the multiple vectors 402a-d may correlate to a particular set of notes 304a-d. For example, vector 402a may correlate to set of notes 304a, vector 402b may correlate to set of notes 304b, vector 402c may correlate to set of notes 304c, and vector 402d may correlate to set of notes 304d.

[0044] The vectors 402a-d generated by the vector generation model 119 may represent at least one musical feature of the musical audio associated with the corresponding set of notes. For example, vector 402a may represent at least one musical feature of the musical audio associated with the set of notes 304a. The musical audio associated with the set of notes 304a may be a musical audio signal (e.g., an actual audio signal, not just musical notes) that may be generated based on the set of notes 304a. Similarly, vector 402b may represent at least one musical feature of the musical audio associated with the set of notes 304b, vector 402c may represent at least one musical feature of the musical audio associated with the set of notes 304c, and vector 402d may represent at least one musical feature of the musical audio associated with the set of notes 304d. The vectors 402a-d may be stored in the database 130 as vector data 131.

[0045] The vectors 402a-d may be consumable by the audio generation model 120. For example, the audio generation model 120 may not be able to read the notes from the note sets 304a-d. Thus, the vector generation model 119 may be used to translate or convert the note sets 304a-d into a format that is readable or consumable by the audio generation model 120.

[0046] As described above, each vector 402a-d may indicate at least one musical characteristic of the musical audio associated with the corresponding musical note set 304a-d. For example, the at least one musical characteristic may include a musical style (e.g., musical genre and / or mood), a bar structure, or an instrument. A musical genre may indicate whether the musical audio associated with the corresponding musical note set is closest to pop, rock, hip hop, country, etc. A musical mood may indicate whether the musical audio associated with the corresponding musical note set has a particular energy (e.g., lively, sad, slow, upbeat, etc.).

[0047] Returning to FIG. 1 , in one embodiment, at least one of the content service 112 or the client device 104 includes an audio generation model 120. The audio generation model 120 may be utilized to generate a plurality of musical pieces based at least in part on vector data 131. The audio generation model 120 may receive the vector data 131 as input and output a plurality of musical pieces. Each musical piece from the plurality of musical pieces may correspond to a particular vector from the plurality of vectors generated by the vector generation model 119. As described above, each vector from the plurality of vectors corresponds to a particular set of notes from the plurality of note sets generated by the note generation model 118. As described above, each set of notes from the plurality of note sets corresponds to a particular video segment. Thus, each musical piece from the plurality of musical pieces generated by the audio generation model 120 may correspond to a particular video segment. Each musical piece generated by the audio generation model 120 may represent or reflect the musical characteristics of the particular musical piece as indicated by the corresponding vector. The musical piece may be stored in a database 132, for example, as musical audio data 133.

[0048] In an embodiment, to generate the multiple musical audios, the audio generation model 120 may utilize one or more templates. The one or more templates may be from a plurality of templates pre-stored in a database, such as template 129 in database 128. Each of the multiple templates may include an audio file. The audio file may be associated with at least a main track (e.g., chords) and / or bass music associated with a particular musical style or type. Based on the vector data 131, the audio generation model 120 may determine one or more templates associated with the appropriate musical style or type. For example, the audio generation model 120 may retrieve the one or more templates from database 128.

[0049] In embodiments, the audio generation model 120 may modify or add features to the template(s) after the template(s) have already been retrieved. For example, the audio generation model 120 may modify or add features to the template(s) based on a corresponding note set previously generated by the note generation model 118. For example, the audio generation model 120 may generate a main melody, ornamental melodies, and / or instrumental tracks based on the corresponding note sets. The audio generation model 120 may add the main melody, ornamental melodies, and / or instrumental tracks to the audio file (e.g., main audio track) of the template retrieved from the database.

[0050] By utilizing one or more templates, the musical audio generation process is quick and does not require excessive computing power. For example, by utilizing one or more templates to provide a base track for the musical audio, the need to generate a new audio track from scratch each time musical audio content is generated for a video can be avoided. As a result, users do not have to wait long periods of time for the musical audio content to be generated. In comparison, if a new audio track had to be generated each time musical audio content was needed for a video, this would require significant computing resources and would also require users to wait longer periods of time for the musical audio content to be generated.

[0051] In embodiments, after generating the plurality of musical audios, the audio generation model 120 may generate final musical audio content for the video based at least in part on synthesizing the plurality of musical audios, which may match the movements, intensities, and transitions in the video.

[0052] FIG. 5 shows an example diagram 500 illustrating music generated at least in part based on vectors. As described above with respect to FIG. 4, multiple vectors 402a-d may be generated. Each of the vectors 402a-d may correspond to one of the musical note sets 304a-d. Each of the musical note sets 304a-d may correspond to one of the segments 204a-d of the video 202. Thus, each of the vectors 402a-d may correspond to one of the video segments 204a-d. The audio generation model 120 may generate multiple musical audio 502a-d based at least in part on the vectors 402a-d. Each of the multiple musical audio 502a-d may be correlated to a vector 402a-d. For example, music audio 502a may correlate to vector 402a, music audio 502b may correlate to vector 402b, music audio 502c may correlate to vector 402c, and music audio 502d may correlate to vector 402d. If vectors 402a-d each correspond to one of segments 204a-d of the video, music audio 502a may correlate to segment 204a, music audio 502b may correlate to segment 204b, music audio 502c may correlate to segment 204c, and music audio 502d may correlate to segment 204d. Music audio 502a-d may be synthesized to generate final music audio content for video 202 that matches the movements, intensity, and transitions within video 202. The video audio matching between video 202 and music audio content may be fine-tuned by applying video warping, for example, while generating the music audio content and / or while adding the music audio content to video 202.

[0053] Returning to FIG. 1 , in one embodiment, at least one of the content service 112 or the client device 104 includes a refinement model 121. The refinement model 121 may be utilized, at least in part, to refine or modify the final music audio content generated for the video. For example, the refinement model 121 may determine whether a user likes the generated music audio content, e.g., based on user input. For example, the refinement model 121 may receive an indication that the user likes the generated music audio content. If the refinement model 121 receives an indication that the user likes the generated music audio content, the generated music audio content may remain unchanged. Conversely, if the refinement model 121 receives an indication that the user does not like the generated music audio content, the refinement model 121 may present or cause a plurality of options to be displayed. For example, the refinement model 121 may present or cause a plurality of options to be displayed on the interface 108a-d of the content application 106. Each of the plurality of options may represent potential modifications or changes that may be made to the generated music audio content. The user may select (e.g., click) one or more of the plurality of options. Refine model 121 may receive an indication of the selected option or options. Refine model 121 may update or modify the music audio content based on the option or options selected by the user. This process may be repeated until the user indicates that they like the music audio content.

[0054] 6 illustrates an example process 600 performed by a content service (e.g., content service 112) and / or a client device (e.g., client device 104). The content service and / or client device can automatically perform this process to efficiently generate music audio content for a video. While illustrated in FIG. 6 as a series of operations, one skilled in the art will recognize that various embodiments may add, remove, reorder, or modify the described operations.

[0055] As described above, a user of a content service may create a video for distribution to other users of the content service. At 602, a video may be received from the user. The video may include multiple frame segments. The video may have already been created by the user. The video may not include background music. The user may want to automatically and efficiently generate background music to correspond to the video.

[0056] To automatically generate music corresponding to video frames, information may be extracted from the video frames. At 604, information may be extracted from the video. The extracted information may include information indicative of motion rate in the video, information indicative of motion saliency in the video, information indicative of scene transitions in the video, and timing information associated with the video. For example, if a video depicts a man walking, information associated with the man's walking speed, whether / for how long he stops, etc. may be extracted from the video. The extracted information associated with video motion rate, video motion saliency, scene transitions, and / or timing may collectively be referred to as extracted information.

[0057] A note generation model may be utilized, at least in part, to automatically generate a set of notes for each segment of the video frames. For example, the note generation model may utilize the extracted information to automatically generate a set of notes for each segment of the video frames. At 606, a model (e.g., the note generation model) may be used to generate a plurality of sets of notes matching the plurality of frame segments based at least in part on the extracted information. The model may be pre-trained and may have learned to correlate video motion speed with note density, video motion saliency with note intensity, video scene transitions with musical composition, and video timing with musical beat. As described above, in an embodiment, the model may be a compound word conversion model trained on a large number of MIDI files (e.g., 3,000 MIDI files).

[0058] The vector generation model may be utilized, at least in part, to automatically generate vectors corresponding to a plurality of sets of notes. For example, the vector generation model may utilize a plurality of sets of notes to automatically generate a vector for each of the plurality of sets of notes. Each vector may indicate at least one musical feature of musical audio associated with the corresponding set of notes.

[0059] At 608, a plurality of vectors corresponding to the plurality of note sets may be generated. Each of the plurality of vectors may indicate at least one musical characteristic of one of the plurality of musical audios. For example, the at least one musical characteristic may include a musical style (e.g., musical genre and / or mood), a bar structure, or an instrument. A musical genre may indicate whether the musical audio associated with the corresponding note set is closest to pop, rock, hip hop, country, etc. A musical mood may indicate whether the musical audio associated with the corresponding note set has a particular energy (e.g., lively, sad, slow, upbeat, etc.). The vectors may have a format consumable by an audio generation model, for example.

[0060] An audio generation model may be utilized to generate a plurality of musical audio pieces based at least in part on the plurality of vectors. The audio generation model may receive the vectors as input and output a plurality of musical pieces. At 610, a plurality of musical audio pieces corresponding to the plurality of frame segments may be generated based at least in part on the plurality of vectors. Each musical piece from the plurality of musical pieces may correspond to a particular vector from the plurality of vectors. As described above, each vector from the plurality of vectors corresponds to a particular set of notes from a plurality of note sets generated by the musical note generation model. As described above, each set of notes from the plurality of note sets corresponds to a particular video segment. Thus, each musical piece from the plurality of musical pieces generated by the audio generation model may correspond to a particular video segment. Each musical piece generated by the audio generation model may represent or reflect the musical characteristics of the particular musical audio piece indicated by the corresponding vector.

[0061] 7 illustrates an example process 700 performed by a content service (e.g., content service 112) and / or a client device (e.g., client device 104). The content service and / or client device may perform this process 700 to efficiently generate music for a video. While illustrated in FIG. 7 as a series of operations, one skilled in the art will recognize that various embodiments may add, remove, reorder, or modify the described operations.

[0062] The vector generation model may be utilized, at least in part, to automatically generate vectors corresponding to a plurality of sets of notes. For example, the vector generation model may utilize a plurality of sets of notes to automatically generate a vector for each of the plurality of sets of notes. Each vector may indicate at least one musical feature of musical audio associated with the corresponding set of notes.

[0063] At 702, a plurality of vectors corresponding to the plurality of note sets may be generated. Each of the plurality of vectors may indicate at least one musical characteristic of one of the plurality of musical audios. For example, the at least one musical characteristic may include a musical style (e.g., musical genre and / or mood), a bar structure, or an instrument. A musical genre may indicate whether the musical audio associated with the corresponding note set is closest to pop, rock, hip hop, country, etc. A musical mood may indicate whether the musical audio associated with the corresponding note set has a particular energy (e.g., lively, sad, slow, upbeat, etc.). The vectors may have a format consumable by an audio generation model, for example.

[0064] To generate the plurality of musical audios, the audio generation model may utilize one or more templates. The one or more templates may be from a plurality of templates stored in a database. Each of the plurality of templates may include an audio file. The audio file may be associated with at least a main track (e.g., chords) and / or bass music associated with a particular musical style or type. At 704, at least one template may be obtained based on each of the plurality of vectors. For example, based on musical features indicated by the plurality of vectors, the audio generation model may determine one or more templates associated with an appropriate musical style or type. Such templates may be retrieved from the database. At 706, musical audio corresponding to each of the plurality of frame segments may be generated based at least in part on the at least one audio file.

[0065] In embodiments, after the one or more templates have been retrieved and musical audio corresponding to each of the plurality of frame segments has been generated, the audio generation model may make modifications or additions to the audio file. At 708, the at least one audio file may be modified by adding a melody based on one of the plurality of note sets. For example, the audio generation model may generate a main melody, ornamental melodies, and / or instrumental tracks based on the corresponding note sets. The audio generation model may add the main melody, ornamental melodies, and / or instrumental tracks to the template's audio file (e.g., the main audio track).

[0066] 8 illustrates an example process 800 performed by a content service (e.g., content service 112) and / or a client device (e.g., client device 104). The content service and / or client device may perform this process 800 to efficiently generate music for a video. While illustrated in FIG. 8 as a series of operations, one skilled in the art will recognize that various embodiments may add, remove, reorder, or modify the described operations.

[0067] As described above, a user of a content service may create a video for distribution to other users of the content service. At 802, a video may be received from the user. The video may include multiple frame segments. The video may have already been created by the user. The video may not include background music. The user may want to automatically and efficiently generate background music to correspond to the video.

[0068] To automatically generate music corresponding to video frames, information may be extracted from the video frames. At 804, information may be extracted from the video. The extracted information may include information indicative of motion rate in the video, information indicative of motion saliency in the video, information indicative of scene transitions in the video, and timing information associated with the video. For example, if a video depicts a man walking, information associated with the man's walking speed, whether / for how long he stops, etc. may be extracted from the video. The extracted information associated with video motion rate, video motion saliency, scene transitions, and / or timing may collectively be referred to as extracted information.

[0069] A note generation model may be utilized, at least in part, to automatically generate a set of notes for each segment of the video frames. For example, the note generation model may utilize the extracted information to automatically generate a set of notes for each segment of the video frames. At 806, a model (e.g., the note generation model) may be used to generate a plurality of sets of notes matching the plurality of frame segments based at least in part on the extracted information. The model may be pre-trained and may learn to correlate video motion speed with note density, video motion saliency with note intensity, video scene transitions with musical composition, and video timing with musical beat. The model may be pre-trained and may have been trained to correlate video motion speed with note density, video motion saliency with note intensity, video scene transitions with musical composition, and video timing with musical beat. As described above, in an embodiment, the model may be a compound word conversion model trained on a large number of MIDI files (e.g., 3,000 MIDI files).

[0070] The vector generation model may be utilized, at least in part, to automatically generate vectors corresponding to a plurality of sets of notes. For example, the vector generation model may utilize a plurality of sets of notes to automatically generate a vector for each of the plurality of sets of notes. Each vector may indicate at least one musical feature of musical audio associated with the corresponding set of notes.

[0071] At 808, a plurality of vectors corresponding to the plurality of note sets may be generated. Each of the plurality of vectors may indicate at least one musical characteristic of one of the plurality of musical audios. For example, the at least one musical characteristic may include a musical style (e.g., musical genre and / or mood), a bar structure, or an instrument. A musical genre may indicate whether the musical audio associated with the corresponding note set is closest to pop, rock, hip hop, country, etc. A musical mood may indicate whether the musical audio associated with the corresponding note set has a particular energy (e.g., lively, sad, slow, upbeat, etc.). The vectors may have a format consumable by an audio generation model, for example.

[0072] An audio generation model may be utilized to generate a plurality of musical audio pieces based at least in part on the plurality of vectors. The audio generation model may receive the vectors as input and output a plurality of musical pieces. At 810, a plurality of musical audio pieces corresponding to the plurality of frame segments may be generated based at least in part on the plurality of vectors. Each musical piece from the plurality of musical pieces may correspond to a particular vector from the plurality of vectors. As described above, each vector from the plurality of vectors corresponds to a particular set of notes from a plurality of note sets generated by the musical note generation model. As described above, each set of notes from the plurality of note sets corresponds to a particular video segment. Thus, each musical piece from the plurality of musical pieces generated by the audio generation model may correspond to a particular video segment. Each musical piece generated by the audio generation model may represent or reflect musical characteristics of the particular musical piece as indicated by the corresponding vector. At 812, musical audio content may be generated for the video based at least in part on synthesizing the plurality of musical audio pieces. The musical audio content may match motion, intensity, and transitions within the video.

[0073] 9 illustrates an example process 900 performed by a content service (e.g., content service 112) and / or a client device (e.g., client device 104). The content service and / or client device may perform this process 900 to efficiently generate music for a video. While illustrated in FIG. 9 as a series of operations, one skilled in the art will recognize that various embodiments may add, remove, reorder, or modify the described operations.

[0074] As described above, musical audio content may be generated for the video. At 902, musical audio content may be generated for the video based at least in part on combining multiple musical audio pieces. The musical audio content may match the movements, intensities, and transitions in the video.

[0075] For example, based on user input, it may be determined whether the user likes the generated musical audio content. At 904, it may be determined whether the user likes the musical audio content based on user input. For example, an indication indicating that the user likes the generated musical audio content may be received. If an indication indicating that the user likes the generated musical audio content is received, the generated musical audio content may remain unchanged. Conversely, an indication indicating that the user does not like the generated musical audio content may be received. At 906, in response to determining that the user does not like the musical audio content, multiple options may be presented. For example, the multiple options may be presented or displayed on a content application interface.

[0076] At 908, the musical audio content may be updated based on one or more options selected by the user from the plurality of options. Each of the plurality of options may represent a potential modification or change that may be made to the generated musical audio content. The user may select (e.g., click) one or more of the plurality of options. An indication of the selected option or options may be received. The musical audio content may be updated, modified, or refined based on the one or more options selected by the user. This process may be repeated until the user indicates that they like the musical audio content.

[0077] 10 illustrates an example process 1000 performed by a content service (e.g., content service 112) and / or a client device (e.g., client device 104). The content service and / or client device may perform this process 1000 to efficiently generate music for a video. While illustrated in FIG. 10 as a series of operations, one skilled in the art will recognize that various embodiments may add, remove, reorder, or modify the described operations.

[0078] As described above, a user of a content service may create a video for distribution to other users of the content service. At 1002, a video may be received from the user. The video may include multiple frame segments. The video may have already been created by the user. The video may not include background music. The user may want to automatically and efficiently generate background music to correspond to the video.

[0079] An audio generation model may be utilized to generate a plurality of musical audio pieces based at least in part on the plurality of vectors. The audio generation model may receive the vectors as input and output a plurality of musical pieces. At 1004, a plurality of musical audio pieces corresponding to the plurality of frame segments may be generated based at least in part on the plurality of vectors. Each musical piece from the plurality of musical pieces may correspond to a particular vector from the plurality of vectors. As described above, each vector from the plurality of vectors corresponds to a particular set of notes from a plurality of note sets generated by the musical note generation model. As described above, each set of notes from the plurality of note sets corresponds to a particular video segment. Thus, each musical piece from the plurality of musical pieces may correspond to a particular video segment. Each musical piece may represent or reflect musical characteristics of the particular musical audio piece indicated by the corresponding vector.

[0080] At 1006, music audio content may be generated for the video based at least in part on combining the plurality of music audios. The music audio content may match motion, intensity, and transitions in the video. The video audio matching between the video and the music audio content may be fine-tuned by applying video warping, for example, while generating the music audio content and / or while adding the music audio content to the video. At 1008, the video audio matching between the video and the music audio content may be fine-tuned by applying video warping while generating the music audio content.

[0081] Figure 11 illustrates a computing device that may be used in various aspects, such as the services, networks, modules, and / or devices illustrated in Figure 1. With respect to the exemplary architecture of Figure 1, any one or all of the models, servers, content services, and client devices may each be implemented by one or more instances of computing device 1100 of Figure 11. The computer architecture illustrated in Figure 11 illustrates a conventional server computer, workstation, desktop computer, laptop computer, tablet, network appliance, PDA, e-reader, digital mobile phone, or other computing node that may be used to perform any aspect of the computer described herein, such as implementing the methods described herein.

[0082] Computing device 1100 may include a substrate or "motherboard," which is a printed circuit board that can be connected to multiple components or devices via a system bus or other electrical communication pathways. One or more central processing units (CPUs) 1104 may operate in conjunction with a chipset 1106. CPU 1104 may be a standard programmable processor that performs arithmetic and logical operations necessary for the operation of computing device 1100.

[0083] The CPU 1104 may transition from one discrete physical state to the next and perform the necessary operations by manipulating switching elements that distinguish between these states. Switching elements may typically include electronic circuits, such as flip-flops, that maintain one of two binary states, and electronic circuits, such as logic gates, that provide an output state based on a logical combination of the states of one or more other switching elements. These basic switching elements may be combined to create more complex logic circuits, including registers, adders / subtractors, arithmetic logic units, floating-point units, etc.

[0084] CPU 1104 may be augmented or replaced by other processing units, such as GPUs, which may include processing units specialized, but not necessarily limited to, highly parallel computing such as graphics and other visualization-related processing.

[0085] Chipset 1106 may provide an interface between CPU 1104 and the remaining components and devices on the board. Chipset 1106 may provide an interface to random access memory (RAM) 1108, which is used as the main memory within computing device 1100. Chipset 1106 may also provide an interface to a computer-readable storage medium, such as read-only memory (ROM) 1120 or non-volatile RAM (NVRAM) (not shown), for storing basic routines that may boot computing device 1100 and facilitate transmitting information between various components and devices. According to aspects described herein, ROM 1120 or NVRAM may store other software components necessary for the operation of computing device 1100.

[0086] Computing device 1100 may operate in a networked environment using logical connections to remote computing nodes and computer systems through a local area network (LAN). Chipset 1106 may include functionality for providing network connectivity through a network interface controller (NIC) 1122, such as a Gigabit Ethernet adapter. NIC 1122 may be able to connect computing device 1100 to other computing nodes over network 1116. It should be understood that multiple NICs 1122 may be present in computing device 1100 to connect the computing device to other types of networks and remote computer systems.

[0087] The computing device 1100 may be connected to a mass storage device 1128 that provides non-volatile storage for the computer. The mass storage device 1128 may store system programs, application programs, other program modules, and data, as described in more detail herein. The mass storage device 1128 may be connected to the computing device 1100 through a storage controller 1124 connected to the chipset 1106. The mass storage device 1128 may be comprised of one or more physical storage units. The mass storage device 1128 may include a management component 1110. The storage controller 1124 may interface with the physical storage units through a Serial Attached SCSI (SAS) interface, a Serial Advanced Technology Attachment (SATA) interface, a Fibre Channel (FC) interface, or any other type of interface for physically connecting and transmitting data between the computer and the physical storage units.

[0088] Computing device 1100 may store data on mass storage device 1128 by transforming the physical state of the physical storage units to reflect the stored information. The particular transformation of the physical state may depend on various factors and different implementations of the present specification. Examples of such factors include, but are not limited to, the technology used to implement the physical storage devices and whether mass storage device 1128 is characterized as a primary storage device, a secondary storage device, etc.

[0089] For example, computing device 1100 may store information in mass storage device 1128 by issuing instructions via storage controller 1124 to change the magnetic properties of a particular location in a magnetic disk drive unit, the reflective or refractive properties of a particular location in an optical storage unit, or the electrical properties of a particular capacitor, transistor, or other discrete component in a solid-state storage unit. Other transformations of physical media are possible without departing from the scope and spirit of this specification, and the foregoing examples are provided merely for ease of explanation. Computing device 1100 may also read information from mass storage device 1128 by detecting the physical state or characteristics of one or more particular locations in the physical storage unit.

[0090] In addition to the mass storage device 1128 discussed above, computing device 1100 may also have access to other computer-readable storage media for storing and retrieving information such as program modules, data structures, or other data. Those skilled in the art will appreciate that computer-readable storage media can be any available media that provide non-transitory data storage and that can be accessed by computing device 1100.

[0091] By way of example, and not limitation, computer-readable storage media may include volatile and non-volatile, transient and non-transitory computer-readable storage media, as well as removable and non-removable media implemented in any manner or technology, including, but not limited to, RAM, ROM, erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory or other solid-state memory technology, compact disc ROM (CD-ROM), digital versatile disc (DVD), high definition DVD ("HD-DVD"), BLU-RAY or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage, other magnetic storage devices, or any other medium that can be used to store desired information in a non-transitory manner.

[0092] A mass storage device, such as mass storage device 1128 shown in FIG. 11 , may store an operating system for controlling the operation of computing device 1100. The operating system may include a version of the LINUX® operating system. The operating system may include a version of Microsoft's WINDOWS SERVER® operating system. According to another aspect, the operating system may include a version of the UNIX® operating system. Also, various mobile phone operating systems, such as IOS® and ANDROID®, may be utilized. It should be understood that other operating systems may also be utilized. Mass storage device 1128 may also store other system or application and data used by computing device 1100.

[0093] The mass storage device 1128 or other computer-readable storage medium may also be encoded with computer-executable instructions that, when loaded into computing device 1100, transform the computing device from a general-purpose computing system into a special-purpose computer capable of implementing aspects described herein. As described above, these computer-executable instructions transform computing device 1100 by defining how CPU 1104 transitions between states. Computing device 1100 may have access to a computer-readable storage medium that stores computer-executable instructions that, when executed by computing device 1100, can perform the methods described herein.

[0094] A computing device, such as computing device 1100 shown in Figure 11, may further include an input / output controller 1132 for receiving and processing input from multiple input devices, such as a keyboard, a mouse, a touchpad, a touchscreen, an electronic stylus, or other types of input devices. Similarly, input / output controller 1132 may provide output to a display, such as a computer monitor, a flat panel display, a digital projector, a printer, a plotter, or other type of output device. It should be understood that computing device 1100 may not include all of the components shown in Figure 11, may include other components not explicitly shown in Figure 11, or may utilize an entirely different architecture than that shown in Figure 11.

[0095] As described herein, a computing device may be a physical computing device, such as computing device 1100 of Figure 11. A computing node may also include a virtual machine host process and one or more virtual machine instances. Computer-executable instructions may be executed indirectly by the physical hardware of the computing device by interpreting and / or executing instructions stored and executed within the context of a virtual machine.

[0096] It is to be understood that the methods and systems are not limited to particular methods, components, or implementations, and that the terminology used herein is used for the purpose of describing particular embodiments only, and is not intended to be limiting.

[0097] As used in the specification and the appended claims, the singular forms "a," "one," and "the" include plural referents unless the context clearly dictates otherwise. Ranges may be expressed herein as from "about" one particular value, and / or to "about" another particular value. When such a range is expressed, another embodiment includes the range from the one particular value and / or to the other particular value. Similarly, when values ​​are expressed as approximations, by use of the antecedent "about," it will be understood that the particular value forms another embodiment. It will also be understood that the endpoints of each range are significant both relative to the other endpoint, and independently of the other endpoint.

[0098] "Optional" or "optionally" means that the subsequently described event or circumstance may or may not occur, and the specification includes cases where said event or circumstance occurs and cases where said event or circumstance does not occur.

[0099] Throughout the description and claims of this specification, the word "comprise" and variations of that word, such as "including" and "including," mean "including, but not limited to," and are not intended to exclude, for example, other elements, integers, or steps. "Exemplary" denotes "an example of" and is not intended to convey indication of a preferred or desirable embodiment. "Such as" is used for purposes of interpretation and not limitation.

[0100] Components are described that can be used to implement the described methods and systems. When describing combinations, subsets, interactions, groups, etc. of these components, specific reference to each of the various individual and collective combinations and permutations of these components may not be explicitly described, and it should be understood that each is specifically contemplated and described herein for all methods and systems. This applies to all aspects of the present application, including, but not limited to, the operations in the described methods. Thus, to the extent that there are various additional operations that can be performed, it should be understood that each of these additional operations can be performed in any specific embodiment or combination of embodiments of the described methods.

[0101] The present method and system may be more readily understood by reference to the following detailed description of the preferred embodiment and examples contained therein, and the accompanying drawings and their descriptions.

[0102] As will be appreciated by those skilled in the art, the methods and systems may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the methods and systems may take the form of a computer program product on a computer-readable storage medium having computer-readable program instructions (e.g., computer software) embodied in the storage medium. More specifically, the methods and systems may take the form of web-implemented computer software. Any suitable computer-readable storage medium may be utilized, including a hard disk, a CD-ROM, an optical storage device, or a magnetic storage device.

[0103] Embodiments of the methods and systems are described below with reference to block diagrams and flowcharts of methods, systems, devices, and computer program products. It should be understood that each block of the block diagrams and flowcharts, and combinations of blocks in the block diagrams and flowcharts, respectively, may be implemented by computer program instructions. These computer program instructions may be loaded into a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute on the computer or other programmable data processing apparatus, generate means for implementing the functions specified in one or more blocks of the flowcharts.

[0104] These computer program instructions may be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including computer-readable instructions for implementing the functions defined in one or more blocks of the flowcharts. The computer program instructions may be loaded into a computer or other programmable data processing apparatus to cause the computer or other programmable data processing apparatus to perform a series of operational steps to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable data processing apparatus, provide steps for implementing the functions defined in one or more blocks of the flowcharts.

[0105] The various features and processes described above may be used independently of one another or may be combined in various ways. All possible combinations and subcombinations are intended to fall within the scope of the present disclosure. Furthermore, in some implementations, some method or process blocks may be omitted. The methods and processes described herein are not limited to any particular order, and the associated blocks or states may be performed in other orders as appropriate. For example, the described blocks or states may be performed in an order other than the order specifically described, or multiple blocks or states may be combined within a single block or state. The example blocks or states may be performed sequentially, in parallel, or in some other manner. Blocks or states may be added to or deleted from the described example embodiments. The example systems and components described herein may be configured differently from that described. For example, elements may be added, deleted, or rearranged compared to the described example embodiments.

[0106] It should also be understood that while various items are shown as being stored in memory or on a storage device during use, these items, or portions thereof, may be transferred between memory and other storage devices for purposes of memory management and data integrity. Alternatively, in other embodiments, some or all of the software modules and / or systems may execute in memory on another device and communicate with the illustrated computing system via computer-to-computer communications. Furthermore, in some embodiments, some or all of the systems and / or modules may be implemented or provided in other ways, such as at least partially in firmware and / or hardware, including, but not limited to, one or more application-specific integrated circuits (ASICs), standard integrated circuits, controllers (e.g., by executing appropriate instructions, and including microcontrollers and / or embedded controllers), field programmable gate arrays (FPGAs), complex programmable logic devices (CPLDs), etc. Some or all of the modules, systems, and data structures may be stored (e.g., as software instructions or structured data) on a computer-readable medium, such as a hard disk, memory, network, or portable media product, for reading by an appropriate device or via an appropriate connection. The systems, modules, and data structures may be transmitted as a generated data signal (e.g., as part of a carrier wave or other analog or digital propagated signal) over a variety of computer-readable transmission media, including wireless-based media and wire / cable-based media, which may take a variety of forms (e.g., as part of a single or multiplexed analog signal, or as a number of discrete digital packets or frames). In other embodiments, such computer program products may take other forms. Accordingly, the invention may be practiced with other computer system configurations.

[0107] While the methods and systems have been described in connection with preferred embodiments and specific examples, it is not intended that the scope be limited to the particular embodiments, as the embodiments herein are intended in all respects to be illustrative and not restrictive.

[0108] Unless otherwise expressly stated, methods described herein do not require that their operations be performed in a particular order. Thus, where a method claim does not actually recite the order in which its operations must be followed, or where the claims or the specification do not specifically state that the operations are limited to a particular order, no order inference is intended to be drawn in any manner. This applies to any possible non-expressive basis for interpretation, including logical issues regarding the arrangement of steps or operational flow, the simple meaning derived from grammatical construction and punctuation, or the number or type of embodiments described in the specification.

[0109] It will be apparent to those skilled in the art that various modifications and variations can be made without departing from the scope or spirit of the present disclosure. Other embodiments will be apparent to those skilled in the art from consideration of the specification and practice described herein. It is intended that the specification and illustrative drawings be considered exemplary only, with its true scope and spirit being indicated by the following claims.

Claims

1. 1. A method for automatically and efficiently generating music audio content for video, comprising: receiving a video from a user, the video comprising a plurality of frame segments; extracting information from the video, the information including information indicative of a rate of motion in the video, information indicative of a saliency of motion in the video, information indicative of a scene transition in the video, and timing information associated with the video; generating a plurality of sets of musical notes matching the plurality of frame segments based at least in part on the extracted information using a pre-trained model that learns to correlate video motion speed with musical note density, to correlate video motion saliency with musical note intensity, to correlate video scene transitions with musical composition, and to correlate video timing with musical beat; generating a plurality of vectors corresponding to the plurality of sets of musical notes, each of the plurality of vectors representing at least one musical feature of one of the plurality of musical audios; generating the plurality of music audios corresponding to the plurality of frame segments based at least in part on the plurality of vectors; generating music audio content for the video based at least in part on synthesizing the plurality of music audios, the music audio content matching motion, intensity, and transitions in the video; applying video warping during generation of the music audio content to fine-tune video audio matching between the video and the music audio content; A method comprising:

2. obtaining at least one template including at least one audio file based on each of the plurality of vectors; generating musical audio corresponding to each of the plurality of frame segments based at least in part on the at least one audio file; The method of claim 1 further comprising:

3. Obtaining at least one template based on each of the plurality of vectors includes: retrieving at least one template from at least one database containing a plurality of templates, each of the plurality of templates including an audio file having a particular musical characteristic; The method of claim 2 further comprising:

4. generating musical audio corresponding to each of the plurality of frame segments based at least in part on the at least one audio file, modifying the at least one audio file by adding a melody based on one of the plurality of note sets; The method of claim 2 further comprising:

5. the at least one musical feature of one of the plurality of musical audios includes at least one of a musical style, a bar structure, or an instrument; The method of claim 1.

6. determining whether the user likes the music audio content based on user input; The method of claim 1 further comprising:

7. presenting a plurality of options in response to the user determining that the music audio content is disliked; and updating the music audio content based on one or more options selected by the user from the plurality of options; The method of claim 6 further comprising:

8. 1. A system for automatically and efficiently generating music audio content for video, comprising: at least one processor; and at least one memory communicatively coupled to the at least one processor and storing computer-readable instructions, the system comprising: The computer-readable instructions, when executed by the at least one processor, cause the at least one processor to perform the method of any one of claims 1 to 7. system.

9. A non-transitory computer-readable storage medium storing computer-readable instructions, comprising: The computer readable instructions, when executed by a processor, cause the processor to perform the method of any one of claims 1 to 7. A non-transitory computer-readable storage medium.

10. A program for causing a computer to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Video background music automatic generation method based on audio and video rhythm relationship

    CN113889059A