Music Generator
By combining user interaction, environmental information and previously created music in the music generator module, dynamically adjusting rules and loops, the problem that existing music services cannot adjust music according to user taste and environmental changes is solved, and personalized and diversified music generation is achieved.
Patent Information
- Application Number
- CN201980034805.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-05-24
- Filing Date
- 2019-05-23
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2039-05-23
AI Technical Summary
Existing streaming music services cannot dynamically adjust music based on users' tastes, environments, behaviors, etc., causing users to feel bored when they hear the same type of music.
Custom music content is generated through the generator module combining multiple inputs such as user interaction, environmental information and previously created music. The system can modify loops, apply audio filters, and adjust rulesets in real time to achieve target music properties.
It realizes real-time generation of music based on users' personalized needs and environmental changes, enhancing the diversity and user experience of music.
Smart Images

Figure CN112189193B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to audio engineering and, more particularly, to generating music content. Background Art
[0002] Streaming music services usually provide songs to users via the Internet. Users can subscribe to these services and stream music through a web browser or application. Examples of such services include PANDORA, SPOTIFY, GROOVESHARK, etc. Often, users can choose music genres or specific artists to stream. Users can usually rate songs (e.g., using a star rating system or a "like / dislike" system), and some music services can customize which songs are streamed to users based on previous ratings. The cost of running a streaming service (which may include paying royalties for each streamed song) is usually borne by the user's subscription costs and / or advertisements played between songs.
[0003] Song selection may be limited by licensing agreements and the number of songs written for a particular genre. Users may become tired of hearing the same songs of a particular genre. Additionally, these services cannot adjust the music based on the user's taste, environment, behavior, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0004] Figure 1 is a diagram illustrating an exemplary music generator module for generating music content based on multiple different types of inputs in accordance with some embodiments.
[0005] Figure 2 is a block diagram illustrating an exemplary overview of a system with multiple user interactive applications for generating output music content according to some embodiments.
[0006] Figure 3 is a block diagram illustrating an exemplary rule set generator that generates rules based on previously composed music, according to some embodiments.
[0007] Figure 4 is a block diagram illustrating an exemplary artist graphical user interface (GUI) in accordance with some embodiments.
[0008] Figure 5 is a diagram illustrating an exemplary music generator module that generates music content based on multiple different types of inputs, including rule sets for one or more different types of instruments, in accordance with some embodiments.
[0009] Figure 6A-6B is a block diagram illustrating an exemplary storage rule set in accordance with some embodiments.
[0010] Figure 7 is a block diagram illustrating an exemplary music generator module for outputting music content for a video in accordance with some embodiments.
[0011] Figure 8 is a block diagram illustrating an exemplary music generator module that outputs music content for a video in real time, according to some embodiments.
[0012] Figures 9A-9B is a block diagram illustrating an exemplary graphical user interface in accordance with some embodiments.
[0013] Fig.10 is a block diagram illustrating an example music generator system implementing a neural network in accordance with some embodiments.
[0014] Fig.11 is a diagram illustrating example constructed portions of music content according to some embodiments.
[0015] Fig.12 is a diagram illustrating an example technique for arranging segments of musical content in accordance with some embodiments.
[0016] Fig.13 is a flow chart illustrating an example method for automatically generating music content according to some embodiments. DETAILED DESCRIPTION
[0017] This specification includes references to various embodiments to indicate that the present disclosure is not intended to refer to one particular implementation, but rather to a series of embodiments falling within the spirit of the present disclosure, including the appended claims. The particular features, structures, or characteristics may be combined in any suitable manner consistent with the present disclosure.
[0018] Within the present disclosure, different entities (which may be variously referred to as "units," "circuits," other components, etc.) may be described or claimed as "configured" to perform one or more tasks or operations. The expression "[entity] configured to [perform one or more tasks]" is used herein to refer to a structure (i.e., a physical thing, such as an electronic circuit). More specifically, the expression is used to indicate that the structure is arranged to perform one or more tasks during operation. A structure can be said to be "configured to" perform a task even if the structure is not currently being operated. "A resource negotiator module configured to generate a predicted queue graph" is intended to, for example, cover a module that performs that function during operation even if the corresponding device is not currently being used (e.g., when its battery is not connected to it). Therefore, an entity described or recited as "configured to" perform a task refers to a physical thing, such as a device, a circuit, a memory storing program instructions executable to perform the task, etc. The phrase is not used herein to refer to an intangible thing.
[0019] The term "configured to" is not intended to mean "configurable to." For example, an unprogrammed mobile computing device would not be considered "configured to" perform a particular function, although it might be "configurable to" perform that function. After appropriate programming, the mobile computing device could then be configured to perform that function.
[0020] The statement in the appended claims that a structure is "configured to" perform one or more tasks is expressly intended to define that claim element. No Invoking 35 U.S.C. § 112(f). Therefore, none of the claims in this application as filed are intended to be interpreted as having means-plus-function elements. If the applicant wishes to invoke § 112(f) during prosecution, it will recite the claim elements using the “means for [performing the function]” construction.
[0021] As used herein, the term "based on" is used to describe one or more factors that influence a determination. The term does not exclude the possibility that other factors may influence the determination. That is, a determination may be based only on specified factors, or based on specified factors and other unspecified factors. Consider the phrase "determining A based on B." The phrase specifies that B is a factor used to determine A or a factor that influences the determination of A. The phrase does not exclude that the determination of A may also be based on some other factor, such as C. The phrase is also intended to cover embodiments in which A is determined based only on B. As used herein, the phrase "based on" is synonymous with the phrase "based at least in part on."
[0022] Detailed Description
[0023] U.S. Patent Application No. 13 / 969,372, filed on August 16, 2013 (now U.S. Patent No. 8,812,144), which is incorporated herein by reference in its entirety, discusses techniques for generating music content based on one or more musical attributes. To the extent any interpretation is made based on a perceived conflict between the definitions of the '372 application and the remainder of the present disclosure, it is intended that the present disclosure control. The music attributes may be input by a user or may be determined based on environmental information such as ambient noise, lighting, etc. The '372 disclosure discusses techniques for selecting stored loops and / or tracks or generating new loops / tracks and layering the selected loops / tracks to generate output music content.
[0024] The present disclosure generally relates to a system for generating customized music content. Target music attributes can be declarative, so that a user specifies one or more goals for the music to be generated and a rule engine selects and combines loops to achieve these goals. The system can also modify loops, for example, by cutting to use only a portion of the loop or applying one or more audio filters to change the sound of the loop. The various techniques discussed below can provide more relevant custom music for different contexts, facilitate the generation of music based on specific sounds, allow users more control over how music is generated, generate music that achieves one or more specific goals, generate music in real time to accompany other content, and the like.
[0025] In some embodiments, computer learning is used to generate a "grammar" (e.g., a set of rules) for a particular artist or type of music. For example, a previous composition can be used to determine a set of rules for using the artist's style to achieve the target musical attributes. The rule set can then be used to automatically generate custom music in the style of the artist. Note that the rule set may include explicit user-understandable rules, such as "combine these types of loops together to achieve the sound of a particular artist," or may be otherwise encoded as parameters of a machine learning engine that implements the composition rules internally, where the user may not have access to the rules. In some embodiments, as discussed in more detail below, the rules are probabilistic.
[0026] In some embodiments, a music generator may be implemented using multiple different rule sets for different types of loops. For example, sets of loops corresponding to specific instruments may be stored (e.g., drum loops, bass loops, melody loops, rhythm guitar loops, etc.). Each rule set may then evaluate which loops in its corresponding set to select, and when to join with other rule sets in the overall composition. Additionally, a master rule set may be used to coordinate the output of the various rule sets.
[0027] In some embodiments, a rule engine is used to generate music based on video and / or audio data. For example, a music generator can automatically generate a soundtrack for a movie even while it is being played. In addition, different scores can be provided to different listeners, for example based on culture, language, demographics, etc. In some embodiments, the music generator can use environmental feedback to adjust the rule set in real time, for example to achieve a desired emotion in the audience. In this way, the rule set can be adjusted to achieve certain environmental goals.
[0028] This disclosure initially refers to Figure 1 and Figure 2 Describes an example music generator module and the overall system organization with multiple applications. Figure 3 and Figure 4Discusses techniques for generating rule sets for specific artists or musical styles. Figure 5 and Figure 6A-6B Techniques for using different rule sets for different sets of loops (e.g., instruments) are discussed. Figure 7 and Figure 8 Techniques for generating music based on video data are discussed. Figures 9A-9B An exemplary application interface is shown.
[0029] In general, the disclosed music generator includes loop data, metadata (e.g., information describing the loop), and a grammar for combining loops based on the metadata. The generator can use rules to create a music experience to identify loops based on metadata and target characteristics of the music experience. It can be configured to expand the set of experiences it can create by adding or modifying rules, loops, and / or metadata. Adjustments can be performed manually (e.g., an artist adds new metadata), or the music generator can expand rules / loops / metadata as it monitors the music experience and desired goals / characteristics within a given environment. For example, if the music generator observes a crowd and sees that people are smiling, it can expand its rules and / or metadata to note that certain loop combinations cause people to smile. Similarly, if cash register sales increase, the rule generator can use this feedback to expand the rules / metadata of associated loops related to sales growth.
[0030] As used herein, the term "loop" refers to the sound information of a single instrument within a specific time interval. The loop can be played in a repeated manner (e.g., a 30-second loop can be played four times in succession to generate 2 minutes of music content), but the loop can also be played once, e.g., without repetition. The various techniques discussed with reference to loops can also be performed using audio files that include multiple instruments.
[0031] Overview of an Exemplary Music Generator
[0032] Figure 1 is a diagram illustrating an exemplary music generator according to some embodiments. In the illustrated embodiment, the music generator module 160 receives various information from a number of different sources and generates output music content 140.
[0033] In the illustrated embodiment, the module 160 accesses the stored (one or more) loops and the (one or more) corresponding attributes 110 of the stored (one or more) loops, and combines these loops to generate the output music content 140. Specifically, the music generator module 160 selects loops based on the attributes of the loops, and combines the loops based on the target music attributes 130 and / or the environmental information 150. In some embodiments, the environmental information is indirectly used to determine the target music attributes 130. In some embodiments, the target music attributes 130 are explicitly specified by the user, for example, by specifying the desired energy level, mood, multiple parameters, etc. Examples of target music attributes 130 include, for example, energy, complexity, and diversity, but more specific attributes (e.g., corresponding to the attributes of the stored tracks) may also be specified. In general, when a higher level of target music attributes is specified, the system can determine a lower level of specific music attributes before generating the output music content.
[0034] Complexity can refer to many loops and / or instruments included in the composition. Energy may be related to other attributes, or may be orthogonal to other attributes. For example, changing the pitch or rhythm can affect energy. However, for a given rhythm and pitch, the energy can be changed by adjusting the instrument type (for example, by adding a high hat or white noise), complexity, volume, etc. Diversity can refer to the amount of change of the generated music over time. Diversity can be generated for a group of static other music attributes (for example, by selecting different tracks for a given rhythm and pitch), or diversity can be generated by changing the music attributes over time (for example, by changing the rhythm and pitch more often when greater diversity is desired). In some embodiments, the target music attribute can be considered to exist in a multidimensional space, and the music generator module 160 can be based on environmental changes and / or user input, for example, slowly moving through the space with course correction (if necessary).
[0035] In some embodiments, the properties stored with the loops include information about one or more loops including: tempo, volume, energy, diversity, spectrum, envelope, modulation, periodicity, rise and decay times, noise, artist, instrument, theme, etc. Note that in some embodiments, the loops are partitioned so that a set of one or more loops is specific to a particular loop type (e.g., one instrument or type of instrument).
[0036] In the illustrated embodiment, module 160 accesses (one or more) stored rule sets 120. In some embodiments, (one or more) stored rule sets 120 specify rules for how many loops overlap so that they are played simultaneously (which can correspond to the complexity of the output music), which major / minor progressions to use when switching between loops or phrases, which instruments to use together (e.g., instruments that have an affinity for each other), etc. to achieve the target music attributes. In other words, the music generator module 160 uses the stored one or more rule sets 120 to achieve one or more declarative goals defined by the target music attributes (and / or target environment information). In some embodiments, the music generator module 160 includes one or more pseudo-random number generators that are configured to introduce pseudo-randomness to avoid repetitive output music.
[0037] In some embodiments, the environmental information 150 includes one or more of the following: lighting information, ambient noise, user information (facial expressions, body posture, activity level, movement, skin temperature, performance of certain activities, clothing type, etc.), temperature information, purchasing activity in the area, time of day, day of the week, time of year, number of people present, weather conditions, etc. In some embodiments, the music generator module 160 does not receive / process environmental information. In some embodiments, the environmental information 130 is received by another module that determines the target music attributes 130 based on the environmental information. The target music attributes 130 can also be derived based on other types of content (e.g., video data). In some embodiments, the environmental information is used to adjust one or more stored rule sets 120, for example to achieve one or more environmental goals. Similarly, the music generator can use environmental information to adjust the properties of one or more stored loops, for example, to indicate target music attributes or target audience characteristics that are particularly relevant to these loops.
[0038] As used herein, the term "module" refers to a circuit configured to perform a specified operation, or refers to a physical non-transient computer-readable medium that stores information (e.g., program instructions) that instructs other circuits (e.g., processors) to perform specified operations. Modules can be implemented in a variety of ways, including as hard-wired circuits or as memories in which program instructions are stored, which can be executed by one or more processors to perform operations. Hardware circuits can, for example, include customized very large scale integration (VLSI) circuits or gate arrays, off-the-shelf semiconductors (such as logic chips, transistors), or other discrete components. Modules can also be implemented in programmable hardware devices such as field programmable gate arrays, programmable array logic, programmable logic devices, etc. Modules can also be any suitable form of non-transient computer-readable media that stores executable program instructions to perform specified operations.
[0039] As used herein, the phrase "music content" refers to both the music itself (the audible representation of the music) and the information that can be used to play the music. Thus, a song recorded as a file on a storage medium (such as, but not limited to, a compact disk, a flash drive, etc.) is an example of music content; the sound produced by outputting the recorded file or other electronic representation (e.g., through a speaker) is also an example of music content.
[0040] The term "music" includes its well-understood meaning, including sounds produced by musical instruments as well as human voices. Thus, music includes, for example: instrumental performances or recordings, a cappella performances or recordings, and performances or recordings that include both musical instruments and human voices. One of ordinary skill in the art will recognize that "music" does not include all recordings of human voices. Works that do not include musical attributes such as rhythm or cadence (e.g., speech, news broadcasts, and audiobooks) are not music.
[0041] A piece of musical "content" may be distinguished from another piece of musical content in any suitable manner. For example, a digital file corresponding to a first song may represent the first piece of musical content, while a digital file corresponding to a second song may represent the second piece of musical content. The phrase "musical content" may also be used to distinguish specific intervals within a given musical work, so that different portions of the same song may be considered different musical content. Similarly, different tracks (e.g., piano track, guitar track) within a given musical work may also correspond to different musical content. In the context of a potentially endless stream of generated music, the phrase "musical content" may be used to refer to a portion of that stream (e.g., a few bars or minutes).
[0042] The music content generated by the embodiments of the present disclosure may be "new music content" - a combination of music elements that has never been generated before. A related (but broader) concept - "original music content" is further described below. To facilitate the explanation of this term, the concept of a "controlling entity" relative to a music content generation instance is described. Unlike the phrase "original music content", the phrase "new music content" does not involve the concept of a controlling entity. Therefore, new music content refers to music content that has never been generated by any entity or computer system before.
[0043] Conceptually, the present disclosure refers to a certain "entity" as a specific instance of controlling computer-generated music content. Such an entity has any legal rights (e.g., copyright) that may correspond to the computer-generated content (to the extent that any such rights may actually exist). In one embodiment, the individual who creates (e.g., encodes various software routines) a computer-implemented music generator or operates (e.g., provides input to) a specific instance of computer-implemented music generation will be the controlling entity. In other embodiments, the computer-implemented music generator can be created by a legal entity (e.g., a company or other commercial organization) such as in the form of a software product, a computer system, or a computing device. In some cases, such a computer-implemented music generator can be deployed to many clients. According to the licensing terms associated with the distribution of the music generator, in various cases, the controlling entity can be a creator, a distributor, or a client. If there is no such explicit legal agreement, the controlling entity of the computer-implemented music generator is an entity that promotes (e.g., provides input to and thereby operates) a specific instance of computer generation of music content.
[0044] Within the meaning of the present disclosure, computer generation of "original musical content" by a controlling entity refers to: 1) a combination of musical elements that have never been generated before by the controlling entity or anyone else, and 2) a combination of musical elements that have been generated before but were originally generated by the controlling entity. Content type 1) is referred to herein as "novel musical content" and is similar to the definition of "new musical content", except that the definition of "novel musical content" involves the concept of "controlling entity", while the definition of "new musical content" does not involve the concept of "controlling entity". On the other hand, content type 2) is referred to herein as "proprietary musical content". Note that the term "proprietary" in this context does not refer to any implied legal rights in the content (although such rights may exist), but is only used to indicate that the musical content was originally generated by the controlling entity. Therefore, the "regeneration" of musical content previously originally generated by the controlling entity by a controlling entity constitutes "generation of original musical content" within the present disclosure. "Non-original musical content" with respect to a particular controlling entity is musical content that is not the "original musical content" of the controlling entity.
[0045] Some music contents may include music components from one or more other music contents. Creating music contents in this way is called "sampling" music contents, and is common in some music works, especially in some music genres. Such music contents are referred to as "music contents with sampling components", "derivative music contents" or other similar terms are used in this article. By contrast, music contents that do not include sampling components are referred to as "music contents without sampling components", "non-derivative music contents" or other similar terms are used in this article.
[0046] In applying these terms, note that if any particular musical content is reduced to a sufficient level of granularity, an argument can be made that that musical content is derivative (effectively meaning that all musical content is derivative). The terms "derivative" and "non-derivative" are not used in this sense in this disclosure. With respect to computer generation of musical content, if the computer generation selects portions of components from pre-existing musical content of an entity other than the controlling entity (e.g., a computer program selects specific portions of an audio file of a work by a popular artist to include in a piece of musical content being generated), then such computer generation is considered derivative (and produces derivative musical content). On the other hand, if the computer generation does not utilize such components of such pre-existing content, then the computer generation of musical content is considered non-derivative (and produces non-derivative musical content). Note that some "original musical content" may be derivative musical content, while some may be non-derivative musical content.
[0047] Note that the term "derivative" is intended to have a broader meaning within this disclosure than the term "derivative work" as used in U.S. copyright law. For example, derivative musical content may or may not be a derivative work under U.S. copyright law. The term "derivative" in this disclosure is not intended to convey a negative connotation; it is only used to indicate whether particular musical content "borrows" portions of content from another work.
[0048] In addition, the phrases "new musical content", "novel musical content" and "original musical content" are not intended to include musical content that is only slightly different from the pre-existing combination of musical elements. For example, only changing a few notes of a pre-existing musical work will not result in new, novel or original musical content, as those phrases are used in the present disclosure. Similarly, only changing the pitch or rhythm of a pre-existing musical work or adjusting the relative strength of its frequency (e.g., using an equalizer interface) will not produce new, novel or original musical content. In addition, the phrases "new musical content, novel musical content and original musical content" are not intended to cover those musical contents that are borderline cases between original content and non-original content; on the contrary, these terms are intended to cover musical content that is undoubtedly and provably original, including musical content that will be eligible for copyright protection for the controlling entity (referred to as "protectable" musical content in this article). In addition, as used herein, the term "available" musical content refers to musical content that does not violate the copyright of any entity other than the controlling entity. New and / or original musical content is often protectable and available. This may be advantageous in preventing the copying of musical content and / or paying royalties for musical content.
[0049] Although various embodiments discussed herein use rule-based engines, various other types of computer-implemented algorithms may be used for any of the computer learning and / or music generation techniques discussed herein. However, rule-based approaches may be particularly effective in a music environment.
[0050] Overview of Applications, Storage Elements, and Data that May be Used in an Exemplary Music System
[0051] The music generator module can interact with a plurality of different applications, modules, storage elements, etc. to generate music content. For example, an end user can install one of a plurality of types of applications for different types of computing devices (e.g., mobile devices, desktop computers, DJ equipment, etc.). Similarly, another type of application can be provided to business users. Interacting with applications when generating music content can allow the music generator to receive external information, which it can use to determine the target music attribute and / or update one or more rule sets for generating music content. In addition to interacting with one or more applications, the music generator module can also interact with other modules to receive rule sets, update rule sets, etc. Finally, the music generator module can access one or more rule sets, loops, and / or generated music content stored in one or more storage elements. In addition, the music generator module can store any of the above listed items in one or more storage elements, which can be local or accessed (e.g., cloud-based) via a network.
[0052] Figure 2 2 is a block diagram showing an exemplary overview of a system for generating output music content based on input from multiple different sources. In the illustrated embodiment, the system 200 includes a rule module 210, a user application 220, a web application 230, an enterprise application 240, an artist application 250, an artist rule generator module 260, a storage device for generated music 270, and an external input 280.
[0053] In the illustrated embodiment, the user application 220, the web application 230, and the enterprise application 240 receive external input 280. In some embodiments, the external input 280 includes: environmental input, target music attributes, user input, sensor input, and the like. In some embodiments, the user application 220 is installed on the user's mobile device and includes a graphical user interface (GUI) that allows the user to interact / communicate with the rule module 210. In some embodiments, the web application 230 is not installed on the user's device, but is configured to run within the browser of the user's device and can be accessed through a website. In some embodiments, the enterprise application 240 is an application used by large-scale entities to interact with the music generator. In some embodiments, the application 240 is used in conjunction with the user application 220 and / or the web application 230. In some embodiments, the application 240 communicates with one or more external hardware devices and / or sensors to collect information about the surrounding environment.
[0054] In the illustrated embodiment, the rule module 210 communicates with the user application 220, the web application 230, and the enterprise application 240 to generate output music content. In some embodiments, the music generator 160 is included in the rule module 210. Note that the rule module 210 may be included in one of the applications 220, 230, and 240, or may be installed on a server and accessed via a network. In some embodiments, the applications 220, 230, and 240 receive the generated output music content from the rule module 210 and cause the content to be played. In some embodiments, the rule module 210, for example, requests input about target music attributes and environmental information to the applications 220, 230, and 240, and may use the data to generate music content.
[0055] In the illustrated embodiment, the rules module 210 accesses the stored rule set(s) 120. In some embodiments, the rules module 210 modifies and / or updates the stored rule set(s) 120 based on communicating with the applications 220, 230, and 240. In some embodiments, the rules module 210 accesses the stored rule set(s) 120 to generate output music content. In the illustrated embodiment, the stored rule set(s) 120 may include rules from the artist rule generator module 260 discussed in more detail below.
[0056] In the illustrated embodiment, the artist application 250 communicates with an artist rule generator module 260 (which may be part of the same application or may be cloud-based, for example). In some embodiments, the artist application 250 allows an artist to create a rule set for their specific sound, for example based on previous compositions. Figure 3 to Figure 4 This functionality is discussed further. In some embodiments, the artist rule generator module 260 is configured to store the generated artist rule sets for use by the rule module 210. A user can purchase a rule set from a specific artist and then use it to generate output music via their specific application. A rule set for a specific artist can be referred to as a signature pack.
[0057] In the illustrated embodiment, when applying the rules to select and combine tracks to generate output music content, the module 210 accesses the stored loop(s) and the corresponding attributes(s) 110. In the illustrated embodiment, the rules module 210 stores the generated output music content in the storage element 270.
[0058] In some embodiments, Figure 2 One or more elements in the cloud are implemented on a server and accessed via a network, which can be referred to as a cloud-based implementation. For example, (one or more) stored rule sets 120, one (one or more) loops / (one or more) attributes 110, and generated music 270 can all be stored in the cloud and accessed by module 210. In another example, module 210 and / or module 260 can also be implemented in the cloud. In some embodiments, generated music 270 is stored in the cloud and digitally watermarked. This, for example, can allow detection of copying generated music, as well as generating a large amount of custom music content.
[0059] In some embodiments, one or more of the disclosed modules are configured to generate other types of content in addition to music content. For example, the system can be configured to generate output visual content based on target music attributes, determined environmental conditions, currently used rule sets, etc. For example, the system can search a database or the Internet based on the current attributes of the music being generated and display a collage of images that dynamically changes as the music changes and matches the attributes of the music.
[0060] Example rule set generator based on previously composed music
[0061] In some embodiments, the music generator is configured to generate output music content having a style similar to that of a known artist or a known style. In some embodiments, the rule set generator is configured to generate a rule set to facilitate such custom music. For example, the rule generator module can capture the specific style of an artist by determining a rule set using previously created music content from the artist. Once a rule set is determined for an artist, the music generator module can generate new music content unique to the artist's style.
[0062] Figure 3 3 is a block diagram illustrating an exemplary rule set generator that generates rules based on previously composed music according to some embodiments. In the illustrated embodiment, module 300 includes a storage device 310 of previous artist compositions, a storage device 320 of one or more artist loops, and an artist interface 330.
[0063] In the illustrated embodiment, the artist rule generator module 260 is configured to generate a rule set for a particular artist (or in other embodiments, a particular theme or musical style) and add the rule set to the stored rule set (120). In some embodiments, the artist uploads previous compositions 310 and / or artist loops 320 (e.g., for creating loops of previous compositions). In other embodiments, the artist may only upload previous compositions without uploading corresponding loops. However, uploading loops may facilitate decomposing previously created music to more accurately generate a rule set for the artist. Therefore, in the illustrated embodiment, the rule generator module 260 has access to previous artist compositions 310 and (one or more) artist loops 320.
[0064] (One or more) artist compositions 310 may include all music content generated by one or more artists. Similarly, (one or more) loops 320 may include all loops or a portion of loops used to generate (one or more) compositions 310.
[0065] In some embodiments, the artist rule generator module 260 separates one or more individual loops from (one or more) artist compositions 310. In some embodiments, knowledge of loops 320 can improve accuracy and reduce processing requirements for the decomposition. Based on the decomposition, the rule generator module determines a set of rules about how artists usually create. In some embodiments, the determined set of rules is called an artist signature package. For example, a rule can specify which instruments an artist usually combines, how an artist usually changes the pitch, the complexity and diversity of the artist, etc. The rule can be binary (e.g., true or false) or can be determined statistically (e.g., artist A changes from pitch A to pitch E 25% of the time, and artist A changes from pitch A to pitch D 60% of the time). Based on statistical rules, the music generator may try to match the specified percentage over time.
[0066] In some embodiments, artists may indicate which music matches certain target music attributes for previously composed music. For example, some compositions may be high or low energy, high or low complexity, happy, sad, etc. Based on this classification and processing of the classified compositions, the rule generator module 260 may determine rules about how artists typically compose for a particular target attribute (e.g., artist A increases tempo to achieve greater energy, while artist B may tend to increase complexity).
[0067] In the illustrated embodiment, the artist interface 330 communicates with the artist rule generator module 260. In some embodiments, the module 260 requests input from the artist through the interface 330. In some embodiments, the artist provides feedback to the artist rule generator module 260 via the interface 330. For example, the module 260 may request feedback from the artist for one or more rules in the generated artist rule set. This may allow the artist to add additional rules, modify the generated rules, etc. For example, the interface may display a rule (i.e., "Artist A transitions from pitch A to pitch E 25% of the time") and allow the artist to delete or change the rule (e.g., the artist may specify that the transition should occur 40% of the time). As another example, the module 260 may request feedback from the artist that confirms whether the module 260 has properly decomposed one or more loops from the artist composition 310.
[0068] In some embodiments, Figure 3 The various elements of may be implemented as part of the same application installed on the same device. In other embodiments, Figure 3One or more elements of the artist interface 330 may be stored separately from the artist interface 330, such as on a server. Additionally, the stored rule set(s) 120 may be provided via the cloud, such as to allow a user to download a rule set corresponding to a particular desired artist. As discussed above, a similar interface may be used to generate rule sets for themes or contexts that are not necessarily artist-specific.
[0069] Figure 4 4 is a block diagram illustrating an exemplary artist interface for generating an artist rule set according to some embodiments. In the illustrated embodiment, a graphical user interface (GUI) 400 includes: a stored loop 410, selection elements 420, 430, 440, and 450, and a display 460. Note that the illustrated embodiment is Figure 3 An example embodiment of an artist interface 330 is shown.
[0070] In the illustrated embodiment, at least a portion of the stored loops 410 are displayed as loops AN 412. In some embodiments, the loops are uploaded by artists, for example, to facilitate analysis of the artist's music to determine a rule set. In some embodiments, the interface allows the artist to select one or more loops 412 from the stored loops 410 to modify or delete. In the illustrated embodiment, a selection element 420 allows the artist to add one or more loops to the list of stored loops 410.
[0071] In the illustrated embodiment, selection element 430 allows an artist to add previously created music content. Selecting this element may result in displaying another interface to upload and otherwise manage such content. In some embodiments, this interface may allow uploading of multiple different music collections. For example, this may allow an artist to create different rule sets for different styles of the same artist. In addition, this may allow an artist to upload previously generated music that the artist considers to be suitable for certain target music attributes, which may facilitate automatic determination of the artist's (one or more) rule sets. As another example, this interface may allow an artist to listen to previous music content and mark parts of previous music content with target music attributes. For example, an artist may mark certain parts as higher energy, lower diversity, certain emotions, etc., and rule generator module 260 may use these tags as input to generate the artist's rule set. Generally speaking, rule generator module 260 may implement any of various appropriate computer learning techniques to determine one or more rule sets.
[0072] In the illustrated embodiment, selection element 440 allows the artist to initiate determining a rule set based on previously composed music (which was added, for example, using element 430). In some embodiments, in response to the artist selecting 440, the artist rule generator module analyzes and separates loops from the previously composed music. In some embodiments, the artist rule generator module generates a rule set for the artist based on the separated loops. In the illustrated embodiment, selection element 450 allows the artist to modify the generated artist rule set (e.g., it may open another GUI that displays the determined rules and allows modification).
[0073] In the illustrated embodiment, display 460 displays the artist's rule set (eg, the original set and / or the rule set modified by the artist). In other embodiments, display 460 may also display various other information disclosed herein.
[0074] In some embodiments, the rule set generator can generate a rule set for a specific user. For example, the music preferred by the user can be decomposed to determine one or more rule sets for the specific user. The user preferences can be based on explicit user input, listening history, indications of preferred artists, etc.
[0075] Example music generator module with different rule sets for different types of loops
[0076] Figure 5 is a diagram illustrating a music generator module configured to access multiple loop sets with multiple corresponding rule sets. In some embodiments, using multiple different rule sets (e.g., for different instruments) can provide greater variation, more accurate matching of music to target attributes, more similarity to real-life musicians, etc.
[0077] In the illustrated embodiment, information 510 includes loop sets for multiple loop types. Loops can be grouped into sets of the same instrument, the same type of instrument, the same type of sound, the same mood, similar attributes, etc. As discussed above, properties of each loop can also be maintained.
[0078] In the illustrated embodiment, the rule sets 520 correspond to corresponding sets of loops in the set of loops and specify rules for selecting and / or combining those loops based on the target music attributes 130 and / or the environmental information 150. These rule sets can similarly coordinate with artists in a jam session by deciding what loops to select and when to join. In some embodiments, one or more master rule sets can operate to select and / or combine outputs from other rule sets.
[0079] Fig. 6A is a block diagram illustrating multiple rule sets for different cycle types AN 612 according to some embodiments. Figure 6B is a block diagram illustrating rule sets 612 and a master rule set 614 for multiple loop sets.
[0080] For example, consider a set of loops for a certain type of drums (e.g., loop type A). The corresponding rule set 512 may indicate various loop parameters to be prioritized when selecting loops based on target music attributes (such as tempo, pitch, complexity, etc.). The corresponding rule set 612 may also indicate whether to provide drum loops (e.g., based on the expected energy level). In addition, the main rule set 614 may determine a subset of selected loops to actually be incorporated into the output stream based on the drum rule set. For example, the main rule set 614 may select from multiple loop sets for different types of drums (so that some selected loops suggested by the corresponding rule set may not actually be incorporated into the output music content 140). Similarly, for example, the main rule set 614 may indicate that drum loops below a certain specified energy level are never included, or one or more drum loops above another specified energy level are always included.
[0081] In addition, the master rule set 614 can indicate the number of loop outputs selected by the rule set 612 to be combined based on the target music attribute. For example, if seven rule sets decide to provide loops from their corresponding loop sets based on the target music attribute (e.g., out of a total of ten rule sets, because three of the rule sets decide not to provide loops at this time), the master rule set 614 can still select only five of the provided loops for combination (e.g., by ignoring or discarding loops from the other two rule sets). In addition, the master rule set 614 can change the loops provided, and / or add additional loops not provided by other rule sets.
[0082] In some embodiments, all rule sets have the same target music attribute at a given time. In other embodiments, target music attributes can be determined or specified for different rule sets respectively. In these embodiments, the main rule set may be useful for avoiding contention between other rule sets.
[0083] Exemplary music generator for video content
[0084] Generating music content for a video can be a long and tedious process. Applying rule-based machine learning using one or more rule sets can avoid this process and / or provide more relevant music content for a video. In some embodiments, a music generator uses video content as input to one or more rule sets when selecting and combining loops. For example, a music generator can generate target music attributes based on video data and / or directly use the attributes of video data as input to a rule set. In addition, when generating a soundtrack for a video, different rule sets can be used for different audiences to create a unique experience for each audience. Once the music generator has selected one or more rule sets and one or more loops for a video, the music generator generates music content and outputs the music content while the video is being watched. Further, the rule set can be adjusted in real time, for example, based on environmental information associated with the viewer of the video.
[0085] Figure 7 7 is a block diagram illustrating an exemplary music generator module configured to output music content based on video data according to some embodiments. In the illustrated embodiment, system 700 includes analysis module 710 and music generator module 160.
[0086] In the illustrated embodiment, the analysis module 710 receives video data 712 and audio data 714 corresponding to the video data. In some embodiments, the analysis module 710 does not receive audio data 714 corresponding to the video data 712, but is configured to generate music based only on the video data. In some embodiments, the analysis module 710 analyzes the data 712 and the data 714 to identify certain attributes of the data. In the illustrated embodiment, one or more attributes 716 of the video and audio content are sent to the music generator module 160.
[0087] In the illustrated embodiment, the music generator module 160 accesses the stored loop(s), the corresponding attribute(s) 110, and the stored rule set(s) 120. To generate music content for a video, the module 160 evaluates the attribute(s) 716 and uses one or more rule sets to select and combine loops to generate output music content 140. In the illustrated embodiment, the music generator module 160 outputs music content 140. In some embodiments, the music content 140 is generated by the music generator module based on both the video data 712 and the audio data 714. In some embodiments, the music content 140 is generated based only on the video data 712.
[0088] In some embodiments, music content is generated as a soundtrack for a video. For example, a soundtrack can be generated for a video based on one or more video and / or audio attributes. In this example, one or more of the following video attributes from the video can be used to update the rule set for the video: pitch of voice (e.g., whether a character in the video is angry), culture (e.g., what accents, costumes, etc. are used in a scene), objects / props in a scene, color / darkness of a scene, frequency of switching between scenes, sound effects indicated by audio data (e.g., explosions, dialogue, moving sounds), etc. Note that the disclosed technology can be used to generate music content for any type of video (e.g., 30 second clips, short films, commercials, still photos, slideshows of still photos, etc.).
[0089] In another example, multiple different tracks are generated for one or more viewers. For example, music content can be generated for two different viewers based on viewer age. For example, a first rule set for an adult audience over 30 years old can be applied, while a second rule set for a child audience under 16 years old can be applied. In this example, the music content generated for the first viewer can be more mature than the music content generated for the second viewer. Similar techniques can be used to generate different music content for various different contexts such as: different times of day, display devices used to display the video, available audio devices, country of display, language, etc.
[0090] Exemplary music generator for video content with real-time updates to rule sets
[0091] In some embodiments, using rule-based machine learning to generate music content for a video can allow real-time adjustment of a rule set (e.g., a rule set on which the music content is based) based on environmental information. This method of generating music content can produce different music for different viewers of the same video content.
[0092] Figure 8 is a block diagram illustrating an exemplary music generator module 160 configured to output music content for a video utilizing real-time adjustments to a set of rules, according to some embodiments.
[0093] In the illustrated embodiment, during display of the video, environmental information 150 is input to the music generator module 160. In the illustrated embodiment, the music generator module performs real-time adjustments to the rule set(s) based on the environmental information 810. In some embodiments, the environmental information 150 is obtained from viewers watching the video. In some embodiments, the information 150 includes one or more of the following: facial expressions (e.g., frowning, smiling, concentrating, etc.), body movements (e.g., clapping, fidgeting, concentrating, etc.), language expressions (e.g., laughing, sighing, crying, etc.), demographics, age, lighting, ambient noise, number of viewers, etc.
[0094] In various embodiments, the output music content 140 is played based on the adjusted rule set simultaneously with the audience watching the video. These techniques can generate unique music content for videos displayed to multiple different audiences at the same time. For example, two audiences in the same theater watching the same video on different screens may hear completely different music content. Similar applications of this example include different audiences on airplanes, subways, sports bars, etc. In addition, if the user has a personal audio device (e.g., headphones), a custom soundtrack can be created for each individual user.
[0095] The disclosed technology can also be used to emphasize specific desired emotions in the audience. For example, the purpose of a horror movie can be to scare the audience. Based on the audience response, the rule set can be dynamically adjusted to increase intensity, fear, etc. Similarly, for sad / happy scenes, the rule set can be adjusted based on whether the target audience is actually sad or happy (for example, the goal is to increase the desired emotion). In some embodiments, the video producer can mark certain parts of his video with certain target attributes, which can be input into the music generator module to more accurately produce the desired type of music. Generally speaking, in some embodiments, the music generator updates the rule set based on whether the attributes exhibited by the audience correspond to the previously determined attributes of the video and / or audio content. In some embodiments, these technologies provide an adaptive audience feedback control loop, wherein audience feedback is used to update the rule set or target parameters.
[0096] In some embodiments, a video may be played for multiple viewers to adjust the rule set in real time. Environmental data may be recorded and used to select a final rule set (e.g., a rule set based on one or more viewers that most closely matches the desired target audience attributes). This rule set may then be used to generate music for the video statically or dynamically without real-time updates to the final rule set.
[0097] Exemplary User and Enterprise GUIs
[0098] Figures 9A-9Bis a block diagram illustrating a graphical user interface according to some embodiments. In the illustrated embodiment, Fig. 9A contains a GUI displayed by a user application 910, and Fig. 9B Contains a GUI displayed by enterprise application 930. In some embodiments, Fig. 9A and Fig. 9B The GUI displayed in is generated by the website, not by the application. In various embodiments, any of a variety of suitable elements may be displayed, including one or more of the following elements: a dial (e.g., for controlling volume, energy, etc.), a button, a knob, a display box (e.g., for providing updated information to the user), and the like.
[0099] exist Fig. 9A 9, the user application 910 displays a GUI that includes a portion 912 for selecting one or more artist packs. In some embodiments, the packs 914 may alternatively or additionally include theme packs or packs for specific occasions (e.g., weddings, birthday parties, graduations, etc.). In some embodiments, the number of packs displayed in portion 912 is greater than the number of packs that can be displayed in portion 912 at one time. Therefore, in some embodiments, the user scrolls up and / or down in portion 912 to view one or more packs 914. In some embodiments, the user may select an artist pack 914 based on the artist pack 914 for which he / she wishes to hear the output music content. In some embodiments, for example, the artist packs may be purchased and / or downloaded.
[0100] In the illustrated embodiment, selection element 916 allows the user to adjust one or more music attributes (eg, energy level). In some embodiments, selection element 916 allows the user to add / delete / modify one or more target music attributes.
[0101] In the illustrated embodiment, selection element 920 allows a user to have a device (e.g., a mobile device) listen to the environment to determine target music attributes. In some embodiments, after the user selects selection element 920, the device uses one or more sensors (e.g., a camera, a microphone, a thermometer, etc.) to collect information about the environment. In some embodiments, application 910 also selects or suggests one or more artist packages based on the environmental information collected by the application when the user selects element 920.
[0102] In the illustrated embodiment, selection element 922 allows a user to combine multiple artist packages to generate a new rule set. In some embodiments, the new rule set is based on a user selecting one or more packages of the same artist. In other embodiments, the new rule set is based on a user selecting one or more packages of different artists. The user can indicate the weights of different rule sets, for example, so that a high-weighted rule set has more influence on the generated music than a low-weighted rule set. The music generator can combine rule sets in a variety of different ways (e.g., by switching between rules from different rule sets, averaging the values of rules from multiple different rule sets, etc.).
[0103] In the illustrated embodiment, selection element 924 allows a user to manually adjust one or more rules in one or more rule sets. For example, in some embodiments, a user wishes to adjust the music content being generated at a more granular level by adjusting one or more rules in the rule set used to generate the music content. In some embodiments, this allows a user of application 910 to manually adjust one or more rules in the rule set used to generate the music content. Fig. 9A The controls displayed in the GUI in adjust the rule set used by the music generator to generate output music content to become its own disc jockey (DJ). These embodiments can also allow for more fine-grained control over target music attributes.
[0104] exist Fig. 9B , the enterprise application 930 displays a GUI that also includes an artist package selection portion 912 having artist packages 914. In the illustrated embodiment, the enterprise GUI displayed by the application 930 also includes an element 916 to adjust / add / delete one or more music attributes. In some embodiments, Fig. 9B The GUI shown in is used in an enterprise or storefront to create a certain environment (e.g., for optimizing sales) by generating music content. In some embodiments, an employee uses application 930 to select one or more artist packages that have been previously proven to increase sales (e.g., metadata for a given rule set can indicate actual experimental results of using the rule set in a real-world environment).
[0105] In the illustrated embodiment, input hardware 940 sends information to the application or website where the enterprise application 930 is being displayed. In some embodiments, input hardware 940 is one of the following: a cash register, a heat sensor, a light sensor, a clock, a noise sensor, etc. In some embodiments, information sent from one or more of the hardware devices listed above is used to adjust the target music attributes and / or rule sets to generate output music content for a specific environment. In the illustrated embodiment, selection element 938 allows a user of application 930 to select one or more hardware devices from which to receive environmental input.
[0106] In the illustrated embodiment, display 934 displays environmental data to a user of application 930 based on information from input hardware 940. In the illustrated embodiment, display 932 displays changes to a rule set based on environmental data. In some embodiments, display 932 allows a user of application 930 to see changes made based on environmental data.
[0107] In some embodiments, Fig. 9A and Fig. 9B The elements shown in are for theme packages and / or occasion packages. That is, in some embodiments, a user or enterprise using the GUI displayed by applications 910 and 930 can select / adjust / modify a rule set to generate music content for one or more occasions and / or themes.
[0108] Detailed example music generator system
[0109] Figures 10 to 12 Details about a specific embodiment of the music generator module 160 are shown. Note that although these specific examples are disclosed for illustrative purposes, they are not intended to limit the scope of the present disclosure. In these embodiments, music is constructed according to loops by client systems such as personal computers, mobile devices, media devices, etc. The loops can be divided into professionally organized loop packages, which can be called artist packages. The loops can be analyzed for musical attributes, and these attributes can be stored as loop metadata. The audio in the constructed track can be analyzed (e.g., in real time) and filtered to mix and control the output stream. Various feedbacks can be sent to the server, including explicit feedback such as from the user's interaction with a slider or button and implicit feedback generated by a sensor based on volume changes, based on listening length, environmental information, etc. In some embodiments, the control input has a known effect (e.g., directly or indirectly specifies the target music attribute) and is used by the composition module.
[0110] The following discussion introduces the reference Figures 10 to 12 Various terms used. In some embodiments, a loop library is a master library of loops that can be stored by a server. Each loop can include audio data and metadata describing the audio data. In some embodiments, a loop package is a subset of the loop library. A loop package can be a package for a specific artist, for a specific mood, for a specific type of event, etc. A client device can download a loop package for offline listening, or download portions of a loop package on demand, such as for online listening.
[0111] In some embodiments, the generated stream is data that specifies the music content that users hear when they use the music generator system. Note that for a given generated stream, the actual output audio signal may vary slightly, for example based on the capabilities of the audio output device.
[0112] In some embodiments, the composition module constructs a composition based on the loops available in the loop package. The composition module may receive loops, loop metadata, and user input as parameters and may be executed by a client device. In some embodiments, the composition module outputs a performance script that is sent to the performance module and one or more machine learning engines. In some embodiments, the performance script outlines which loops will be played on each track of the generated stream and what effects will be applied to the stream. The performance script may use beat relative timing to indicate when an event occurs. The performance script may also encode effect parameters (e.g., for effects such as reverb, delay, compression, equalization, etc.).
[0113] In some embodiments, the performance module receives a performance script as input and renders it as a generated stream. The performance module can generate multiple tracks specified by the performance script and mix these tracks into a stream (e.g., a stereo stream, but the stream can have various encodings in various embodiments, including surround sound encoding, object-based audio encoding, multi-channel stereo, etc.). In some embodiments, when provided with a specific performance script, the performance module will always produce the same output.
[0114] In some embodiments, the analysis module is a server-implemented module that receives feedback information and configures the composition module (e.g., in real time, periodically, based on administrator commands, etc.). In some embodiments, the analysis module uses a combination of machine learning techniques to correlate user feedback with performance script and loop library metadata.
[0115] Fig.10 is a block diagram illustrating an example music generator system including analysis and composition modules according to some embodiments. In some embodiments, Fig.10 The system is configured to generate a potentially unlimited stream of music using the user's direct control over the mood and style of the music. In the illustrated embodiment, the system includes an analysis module 1010, a composition module 1020, a performance module 1030, and an audio output device 1040. In some embodiments, the analysis module 1010 is implemented by a server, and the composition module 1020 and the performance module 1030 are implemented by one or more client devices. In other embodiments, the modules 1010, 1020, and 1030 may all be implemented on the client device, or may all be implemented on the server side.
[0116] In the illustrated embodiment, the analysis module 1010 stores one or more artist packages 1012 and implements a feature extraction module 1014 , a client simulator module 1016 , and a deep neural network 1018 .
[0117] In some embodiments, the feature extraction module 1014 adds the loop to the loop library after analyzing the loop audio (but note that some loops may be received with metadata already generated and may not need analysis). For example, raw audio in a format such as wav, aiff, or FLAC can be analyzed for quantifiable musical attributes such as instrument classification, pitch transcription, beat timing, tempo, file length, and audio amplitude in multiple frequency bins. The analysis module 1010 can also store more abstract musical attributes or emotional descriptions of the loop, for example based on manual labeling by the artist or machine listening. For example, emotions can be quantified using multiple discrete categories, with a range of values for each category for a given loop.
[0118] For example, consider loop A, which is analyzed to determine that the notes G2, Bb2, and D2 are used, the first beat starts at 6 milliseconds into the file, the tempo is 122 bpm, the file is 6483 milliseconds long, and the loop has normalized amplitude values of 0.3, 0.5, 0.7, 0.3, and 0.2 over five frequency bins. The artist may label the loop as "funk genre" with the following mood values:
[0119] Beyond Peace strength joy sad tension high high Low medium none Low
[0120] The analysis module 110 may store this information in a database, and the client may download sub-segments of the information, for example, as loop packages. Although the artist package 1012 is shown for illustrative purposes, the analysis module 1010 may provide various types of loop packages to the composition module 1020.
[0121] In the illustrated embodiment, the client simulator module 1016 analyzes various types of feedback to provide feedback information in a format supported by the deep neural network 1018. In the illustrated embodiment, the deep neural network 1018 also receives a performance script generated by the composition module as input. In some embodiments, the deep neural network configures the composition module based on these inputs, for example, to improve the correlation between the type of generated music output and the desired feedback. For example, the deep neural network can periodically push updates to the client device implementing the composition module 1020. Note that the deep neural network 1018 is shown for illustrative purposes and can provide powerful machine learning performance in the disclosed embodiments, but is not intended to limit the scope of the present disclosure. In various embodiments, various types of machine learning techniques can be implemented separately or in various combinations to perform similar functions. Note that the machine learning module can be used to directly implement a rule set (e.g., a layout rule or technique) in some embodiments, or can be used to control a module that implements other types of rule sets, for example, using the deep neural network 1018 in the illustrated embodiment.
[0122] In some embodiments, the analysis module 1010 generates composition parameters for the composition module 1020 to improve the correlation between the expected feedback and the use of certain parameters. For example, actual user feedback can be used to adjust the composition parameters, for example, to try to reduce negative feedback.
[0123] As an example, consider a case where module 1010 discovers a correlation between negative feedback (e.g., clear low ranking, low volume listening, short listening time, etc.) and composition using a large number of layers. In some embodiments, module 1010 uses a technique such as back-propagation to determine that adjusting a probability parameter for adding more tracks reduces the frequency of the problem. For example, module 1010 may predict that reducing the probability parameter by 50% will reduce negative feedback by 8%, and may decide to perform the reduction and push the updated parameter to the composition module (note that the probability parameter is discussed in detail below, but any of the various parameters of the statistical model may be similarly adjusted).
[0124] As another example, consider a case where module 1010 finds that negative feedback is associated with the user setting the emotion control to high tension. A correlation may also be found between loops with low tension labels and users requesting high tension. In this case, module 1010 may add parameters such that the probability of selecting loops with high tension labels increases when the user requests high tension music. Thus, machine learning may be based on a variety of information, including composition output, feedback information, user control input, and the like.
[0125] In the illustrated embodiment, composition module 1020 includes a part sequencer 1022, a part arranger 1024, a technique implementation module 1026, and a loop selection module 1028. In some embodiments, composition module 1020 organizes and structures the parts of a composition based on loop metadata and user control inputs (e.g., mood controls).
[0126] In some embodiments, the section sequencer 1022 sequences sections of different types. In some embodiments, the section sequencer 1022 implements a finite state machine to continuously output sections of the next type during operation. For example, the composition module 1020 can be configured to use different types of sections, such as an intro, buildup, drop, breakdown, and bridge, as described below with reference to Fig.12 In addition, each section may include multiple subsections that define how the music changes throughout the section, for example, including a transition-in subsection, a main content subsection, and a transition-out subsection.
[0127] In some embodiments, part arranger 1024 constructs sub-parts according to arrangement rules. For example, one rule may specify to transfer in by gradually adding tracks. Another rule may specify to transfer in by gradually increasing the gain on a set of tracks. Another rule may specify to cut off vocal loops to create melody. In some embodiments, the probability that the loop in the loop library is attached to a track is a function of the following items: the current position in the part or sub-part, the loops overlapping in time on another track, and user input parameters such as emotional variables (which can be used to determine the target attribute of the generated music content). This function can be adjusted, for example, by adjusting the coefficient based on machine learning.
[0128] In some embodiments, the technical implementation module 1020 is configured to facilitate partial arrangements by adding rules, such as those specified by the artist or determined by analyzing the composition of a particular artist. "Techniques" can describe how a particular artist implements an arrangement rule at a technical level. For example, for an arrangement rule that specifies transitioning in by gradually adding tracks, a technique can indicate adding tracks in the order of drums, bass, pads, and then vocals, while another technique can indicate adding tracks in the order of bass, pads, and then drums. Similarly, for an arrangement rule that specifies cutting off a vocal loop to create a melody, a technique can indicate cutting off the vocals on every second beat, and repeating the looped cut twice before moving to the next cut section.
[0129] In the illustrated embodiment, loop selection module 1028 selects loops to include in the parts of part arranger 1024 according to arrangement rules and techniques. Once the parts are complete, the corresponding performance scripts can be generated and sent to performance module 1030. Performance module 1030 can receive performance script parts at various granularities. This can, for example, include the entire performance script of a certain length of performance, the performance script of each part, the performance script of each sub-part, etc. In some embodiments, arrangement rules, techniques, or loop selection is implemented in a statistical manner, for example, where different methods use different percentages of the time.
[0130] In the illustrated embodiment, the performance module 1030 includes a filter module 1031, an effects module 1032, a mixing module 1033, a mastering module 1034, and a performance module 1035. In some embodiments, these modules process a performance script and generate music data in a format supported by an audio output device 1040. The performance script may specify: which loops to play, when they should be played, what effects the module 1032 should apply (e.g., on a per-track or per-subpart basis), what filters the module 1031 should apply, and so on.
[0131] For example, a performance script may specify that a low pass filter from 1000 to 20000 Hz is applied from 0 to 5000 milliseconds on a particular track. As another example, a performance script may specify that a reverb with a 0.2 wet setting is applied from 5000 to 15000 milliseconds on a particular track.
[0132] In some embodiments, the mixing module 1033 is configured to perform automatic level control on the tracks being combined. In some embodiments, the mixing module 1033 uses frequency domain analysis of the combined tracks to measure frequencies with too much or too little energy, and applies gain to the tracks in different frequency bands to flatten the mix. In some embodiments, the mastering module 1034 is configured to perform multi-band compression, equalization (EQ), or limiting processes to generate data for final formatting by the performance module 1035. Fig.10 Embodiments may automatically generate various output music content based on user input or other feedback information, while machine learning techniques may allow for an improved user experience over time.
[0133] Fig.11 is a diagram illustrating example constructed portions of music content according to some embodiments. Fig.10 The system can compose such a part by applying arrangement rules and techniques. In the example shown, the construction part includes three sub-parts and separate tracks for vocals, pads, drums, bass and white noise.
[0134] In the example shown, the transition in the subsection includes a drum loop A, which is also repeated for the main content subsection. The transition in the subsection also includes a bass loop A. As shown, the gain of the section starts low and increases linearly throughout the section (but non-linear increases or decreases are contemplated). In the example shown, the main content and roll-out subsections include various vocals, pads, drums, and bass loops. As described above, the disclosed techniques for automatically sequencing and arranging sections and implementation techniques can generate a nearly infinite stream of output music content based on various user-adjustable parameters.
[0135] In some embodiments, the computer system displays a display similar to Fig.11 interface and allows the artist to specify the technique used to compose the part. For example, an artist can create Fig.11 The structure shown in , which can be parsed into the code of the composer module.
[0136] Fig.12 12 is a diagram illustrating an example technology of the part for arranging music content according to some embodiments. In the illustrated embodiment, the stream 1210 generated comprises a plurality of parts 1220, and each part comprises a beginning sub-part 1222, a development sub-part 1224 and a transition sub-part 1226. In the illustrated example, the multiple types of each part / sub-part are shown in the table connected via dotted lines. In the illustrated embodiment, the circular element is an example of an arrangement tool, and it can also be realized using the specific techniques discussed below. As shown in the figure, various composition decisions can be performed pseudo-randomly according to statistical percentages. For example, the type of the sub-part, the arrangement tool for a specific type or sub-part or the technology for realizing the arrangement tool can be determined statistically.
[0137] In the example shown, a given section 1220 is one of five types: prelude, build-up, destruction, weakening, and bridge, each type having a different function for controlling the intensity on the section. In this example, the state subsection is one of three types: slowly building up. Sudden transition. Or at the very least, each type has different behaviors. In this example, the development subsection is one of three types: reduce, transform, or expand. In this example, the transition subsection is one of three types: collapse, ramp, or prompt. For example, different types of sections and subsections can be selected based on rules, or different types of sections and subsections can be selected pseudo-randomly.
[0138] In the example shown, one or more arrangement tools are used to implement behaviors for different subsection types. For slow builds, in this example, a low pass filter is applied 40% of the time, and layers are added 80% of the time. For transform development subsections, in this example, loops are cut 25% of the time. Various other arrangement tools are shown, including single beats, dropping beats, applying reverb, adding pads, adding themes, removing layers, and white noise. These examples are included for illustrative purposes and are not intended to limit the scope of the present disclosure. In addition, for ease of illustration, these examples may not be complete (e.g., actual arrangements may typically involve a much larger number of arrangement rules).
[0139] In some embodiments, one or more arrangement tools may be implemented using specific techniques (which may be artist-specified or determined based on analysis of the artist's content). For example, single shots may be implemented using sound effects or vocals, loop removal may be implemented using stutter or half-cut techniques, layer removal may be implemented by removing synths or removing vocals, white noise may be implemented using ramp or pulse functions, and so on. In some embodiments, the specific technique selected for a given arrangement tool may be selected based on a statistical function (e.g., layer removal may remove synths 30% of the time, and it may remove vocals for a given artist 70% of the time). As discussed above, arrangement rules or techniques may be automatically determined by analyzing existing compositions, for example, using machine learning.
[0140] Example Method
[0141] Fig.13 is a flow chart illustrating a method for generating output music content according to some embodiments. Fig.13 The method shown can be used in combination with any of the computer circuits, systems, devices, elements or components disclosed herein. In various embodiments, some of the method elements shown can be performed simultaneously, in an order different from the order shown, or can be omitted. Other method elements can also be performed as needed.
[0142] At 1310, in the illustrated embodiment, the computer system accesses a collection of music content. For example, the collection of music content may be an album, a song, a complete work, etc. by a particular artist. As another example, the collection of music content may be associated with a particular genre, event type, mood, etc.
[0143] At 1320, in the illustrated embodiment, the system generates a composition rule set based on analyzing a combination of multiple loops in the music content set. The composition rule can be specified in a statistical manner, and a random or pseudo-random process can be utilized to meet the statistical criteria. Loops can be explicitly provided for the music content set, or the system can decompose the music content set to determine the loops. In some embodiments, in addition to or in lieu of direct artist input to the technology implementation module 1026, the analysis module 1010 can generate techniques (which can also be referred to as rule sets or grammars) for composing the music content set, and the composition module 1020 can use these techniques to generate new music content. In some embodiments, the arrangement rules can be determined at element 1320.
[0144] At 1330, in the illustrated embodiment, the system generates new output music content by selecting loops from the loop set and combining the selected loops so that multiple loops in the loop set overlap in time, wherein the selection and combination are performed based on the composition rule set and the properties of the loops in the loop set. Note that in some embodiments, different devices of the computing system can generate the output music content and generate the composition rules. In some embodiments, the client device generates the output music content based on parameters provided by the server system (e.g., generated by the deep neural network 1018).
[0145] In some embodiments, generating new output music content includes modifying one or more of the selected loops. For example, the system can cut a loop, apply a filter to a loop, and so on.
[0146] In some embodiments, loop selection and combination are performed based on target musical attributes (e.g., user control inputs to composition module 1020). In some embodiments, various system parameters can be adjusted based on environmental information. For example, the system can adjust the rules / techniques / grammar itself (e.g., using a machine learning engine such as deep neural network 1018) based on environmental information or other feedback information. As another example, the system can adjust or weight target attributes based on environmental information.
[0147] Although specific embodiments have been described above, these embodiments are not intended to limit the scope of the present disclosure, even in the case where only a single embodiment is described with respect to a particular feature. Unless otherwise stated, the examples of features provided in the present disclosure are intended to be illustrative rather than restrictive. The above description is intended to cover such alternatives, modifications, and equivalents that will be apparent to those skilled in the art having the benefit of this disclosure.
[0148] The scope of the present disclosure includes any feature or combination of features disclosed herein (explicitly or implicitly) or any generalization thereof, whether or not it mitigates any or all of the problems addressed herein. Accordingly, new claims may be made during the examination of the present application (or an application claiming priority thereto) based on any such combination of features. In particular, with reference to the attached claims, features from dependent claims may be combined with those in the independent claims, and features from individual independent claims may be combined in any appropriate manner rather than just in the specific combinations listed in the attached claims.
Claims
1. A method, include: Accessing a collection of music content by a computer system; A composition rule set is generated by the computer system based on analyzing a combination of a plurality of loops in the music content set, wherein the rules include rules for: Choose how many loops to stack so they can play simultaneously; Select the type of instrument from which to compose the above loop; and selecting the next tone for the tone progression; and New output music content is generated by the computer system by the following operations: selecting loops from a loop set and combining selected loops in the loops so that multiple loops in the loops overlap in time, wherein the selection and combination are performed based on the composition rule set and the properties of the loops in the loop set.
2. The method according to claim 1, in, The selecting and combining are also performed based on one or more target music attributes for the new output music content.
3. The method according to claim 2, further comprising: include: At least one composition rule in the composition rule set or the one or more target music attributes is adjusted based on environmental information associated with an environment in which the new output music content is played.
4. The method according to claim 1, in, Generating the composition rule set includes generating a plurality of different rule sets for corresponding different types of instruments for loops in the plurality of loops.
5. The method according to claim 1, in, The rules also include rules for: Choose one or more types of musical instruments; Select one or more parameters for chopping the vocal loop; Select one or more low-pass filter parameters; and Select one or more reverb parameters.
6. The method according to claim 1, in, The rules also include one or more rules for: Whether it’s creating energy by increasing the tempo or adding complexity; whether the transformation is done by adding tracks or increasing gain; The order of the tracks added for one or more transitions; Pitch transposition; Select the beat timing; Select one or more white noise parameters; Choose a rhythm; or Select a loop based on the amplitude of the audio in one or more frequency bins.
7. The method according to claim 1, in, Generating the new output music content further comprises modifying at least one of the loops based on the set of composition rules.
8. The method according to claim 1, in, One or more rules in the composition rule set are specified in a statistical manner.
9. The method according to claim 1, in, At least one rule in the rule set specifies a relationship between a target musical attribute and one or more cyclic attributes, wherein the one or more cyclic attributes include one or more of: tempo, volume, energy, diversity, spectrum, envelope, modulation, periodicity, rise time, decay time, or noise.
10. The method according to claim 1, in, The collection of music content includes content targeted at specific types of occasions.
11. The method according to claim 1, in, Generating the composition rule set includes: training one or more machine learning engines to implement the composition rule set, wherein the selecting and combining are performed by the one or more machine learning engines.
12. The method of claim 1, in, The composition rule sets include multiple rule sets for specific types of loops and a master rule set that specifies rules for combining different types of loops.
13. A non-transitory computer readable medium having instructions stored thereon, the instructions being executable by a computing device to perform operations, the operations include: Access a collection of music content; A composition rule set is generated based on analyzing a combination of a plurality of loops in the music content set, wherein the rules include rules for: Choose how many loops to stack so they can play simultaneously; Select the type of instrument from which to compose the above loop; and selecting the next tone for the tone progression; and New output music content is generated by selecting loops from a loop set and combining selected ones of the loops so that multiple ones of the loops overlap in time, wherein the selection and combination are performed based on the composition rule set and properties of the loops in the loop set.
14. The non-transitory computer readable medium of claim 13, in, The selecting and combining are also performed based on one or more target music attributes for the new output music content.
15. The non-transitory computer readable medium of claim 14, in, The operations also include: At least one composition rule in the composition rule set or the one or more target music attributes is adjusted based on environmental information associated with an environment in which the new output music content is played.
16. The non-transitory computer readable medium of claim 13, in, Generating the composition rule set includes generating a plurality of different rule sets for corresponding different types of instruments for loops in the plurality of loops.
17. The non-transitory computer readable medium of claim 13, in, The rules also include rules for: Choose one or more types of musical instruments; Select one or more parameters for chopping the vocal loop; Select one or more low-pass filter parameters; Select one or more reverb parameters; Whether it’s creating energy by increasing the tempo or adding complexity; whether the transformation is done by adding tracks or increasing gain; The order of the tracks added for one or more transitions; Pitch transposition; Select the beat timing; Select one or more white noise parameters; Choose a rhythm; or Select a loop based on the amplitude of the audio in one or more frequency bins.
18. The non-transitory computer readable medium of claim 13, in, At least one rule in the rule set specifies a relationship between a target musical attribute and one or more cyclic attributes, wherein the one or more cyclic attributes include one or more of: tempo, volume, energy, diversity, spectrum, envelope, modulation, periodicity, rise time, decay time, or noise.
19. The non-transitory computer readable medium of claim 13, in, Generating the composition rule set includes: training one or more machine learning engines to implement the composition rule set, wherein the selecting and combining are performed by the one or more machine learning engines.
20. A device, include: one or more processors; and One or more memories having program instructions stored thereon, the program instructions being executable by the one or more processors to perform the following operations: Access a collection of music content; A composition rule set is generated based on analyzing a combination of a plurality of loops in the music content set, wherein the rules include rules for: Choose how many loops to stack so they can play simultaneously; Select the type of instrument from which to compose the above loop; and selecting the next tone for the tone progression; and New output music content is generated by selecting loops from a loop set and combining selected ones of the loops so that multiple ones of the loops overlap in time, wherein the selection and combination are performed based on the composition rule set and properties of the loops in the loop set.
Citation Information
Patent Citations
Music generator
US8812144B2
Device, system and method for generating an accompaniment of input music data
CN104380371A
Systems and methods for creating, modifying, interacting with and playing musical compositions
US20040089141A1
Music generator
US20140052282A1