Generation of music content
A system using machine learning and user-defined controls generates personalized music tailored to user preferences and environments, addressing limitations in streaming services by enhancing music variety and user engagement.
Patent Information
- Application Number
- JP2025090760
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-08-21
- Filing Date
- 2025-05-30
- Publication Date
- 2025-08-15
AI Technical Summary
Streaming music services often lack the ability to tailor music to individual user preferences and environmental contexts, leading to user boredom and limited song selection due to licensing agreements.
A system that generates custom musical content using machine learning algorithms, incorporating user-defined controls and environmental data to create personalized music, and tracks playback data for stakeholders.
Provides personalized music experiences that adapt to user preferences and environments, while ensuring proper compensation for stakeholders, enhancing user engagement and music variety.
Smart Images

Figure 2025120210000001_ABST
Abstract
Description
[Technical Field]
[0001] TECHNICAL FIELD This disclosure relates to audio engineering, and more particularly to the creation of musical content. [Background technology]
[0002] Streaming music services typically provide music to users over the Internet. Users can subscribe to these services and stream music through a web browser or application. Examples of such services include Pandora, Spotify, and Grooveshark. Users often have the ability to select a music genre or a particular artist. Users can typically rate songs (e.g., using a star rating or a like / dislike system), and some music services can adjust which songs they stream to users based on previous ratings. The costs of running a streaming service (which may include paying royalties for each song streamed) are typically covered by user subscription fees and / or advertisements played between songs.
[0003] Song selection may be limited by the number of songs written for a particular genre and by licensing agreements. Users may become bored of listening to the same songs in a particular genre. Furthermore, these services may not be able to tailor music to a user's preferences, environment, behavior, etc. [Brief explanation of the drawings]
[0004] [Figure 1] FIG. 1 illustrates an exemplary music generator.
[0005] [Figure 2] FIG. 1 is a block diagram illustrating an exemplary overview of a system for generating output musical content based on inputs from multiple different sources, according to some embodiments.
[0006] [Figure 3] FIG. 1 is a block diagram illustrating an exemplary music generation system configured to output musical content based on an analysis of a graphical representation of an audio file, according to some embodiments.
[0007] [Figure 4] 1 shows an example of a pictorial representation of an audio file.
[0008] [Figure 5A] 1 shows an example of a grayscale image for melody image feature representation. [Figure 5B] 1 shows an example of a grayscale image for drumbeat image feature representation.
[0009] [Figure 6] FIG. 1 is a block diagram illustrating an exemplary system configured to generate a single image representation, according to some embodiments.
[0010] [Figure 7] 1 shows an example of a single image representation of multiple audio files.
[0011] [Figure 8] FIG. 1 is a block diagram illustrating an exemplary system configured to implement user-generated control in musical content generation, according to some embodiments.
[0012] [Figure 9] 1 illustrates a flowchart of a method for training a music generation module based on user-created control elements, according to some embodiments.
[0013] [Figure 10] FIG. 1 is a block diagram illustrating an exemplary teacher / student framework system, according to some embodiments.
[0014] [Figure 11]FIG. 1 is a block diagram illustrating an exemplary system configured to implement audio techniques in musical content generation, according to some embodiments.
[0015] [Figure 12] 1 shows an example of an audio signal graph.
[0016] [Figure 13] 1 shows an example of an audio signal graph.
[0017] [Figure 14] 1 illustrates an exemplary system for implementing real-time modification of musical content using a music technology music generation module, according to some embodiments.
[0018] [Figure 15] 1 illustrates a block diagram of exemplary API modules in a system for audio parameter automation, according to some embodiments.
[0019] [Figure 16] FIG. 1 illustrates a block diagram of an exemplary memory zone, according to some embodiments.
[0020] [Figure 17] 1 shows a block diagram of an exemplary system for storing new music content, according to some embodiments.
[0021] [Figure 18] FIG. 10 illustrates an example of playback data according to some embodiments.
[0022] [Figure 19] FIG. 1 is a block diagram illustrating an exemplary music composition system, according to some embodiments.
[0023] [Figure 20A] FIG. 1 is a block diagram illustrating a graphical user interface according to some embodiments. [Figure 20B] FIG. 1 is a block diagram illustrating a graphical user interface according to some embodiments.
[0024] [Figure 21] FIG. 1 is a block diagram illustrating an exemplary music generation system including an analysis module and a composition module, according to some embodiments.
[0025] [Figure 22] 1 illustrates an example of an enhanced section of music content, according to some embodiments.
[0026] [Figure 23] FIG. 1 illustrates an exemplary technique for arranging sections of musical content, according to some embodiments.
[0027] [Figure 24] 1 is a flowchart method for using a ledger according to some embodiments.
[0028] [Figure 25] 1 is a flow diagram of a method for using image representations to combine audio files, according to some embodiments.
[0029] [Figure 26] FIG. 1 is a flow diagram of a method for implementing a user-created control element, according to some embodiments.
[0030] [Figure 27] 1 is a flow diagram of a method for generating musical content by modifying audio parameters according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0031] While the embodiments disclosed herein are susceptible to various modifications and alternative forms, specific embodiments have been shown by way of example in the drawings and are herein described in detail. It should be understood, however, that the drawings and their detailed description are not intended to limit the scope of the claims to the particular forms disclosed. On the contrary, the present application is intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the disclosure of this application, as defined by the appended claims.
[0032] This disclosure includes references to "one embodiment," "a particular embodiment," "some embodiments," "various embodiments," or "one embodiment." Appearances of the phrases "in one embodiment," "a particular embodiment," "some embodiments," "various embodiments," or "one embodiment" do not necessarily refer to the same embodiment. Particular features, structures, or characteristics may be combined in any suitable manner consistent with this disclosure.
[0033] In the appended claims, the statement that an element is "configured to" perform one or more tasks is expressly intended not to invoke 35 U.S.C. §112(f) for that claim element. Accordingly, any claims in this application as filed are intended to be construed as having means-function elements. If applicant intends to invoke §112(f) during prosecution, claim elements will be recited using the "means" construction.
[0034] As used herein, the term "based on" is used to describe one or more factors that influence a determination. This term does not exclude the possibility that additional factors may influence the determination. That is, the determination may be based solely on the specified factor, or on the specified factor and other unspecified factors. Consider the phrase "determining A based on B." This phrase specifies that B is a factor used to determine A or that influences determining A. This phrase does not presuppose that the determination of A can also be based on other factors, such as C. This phrase is intended to encompass embodiments in which A is determined solely based on B.
[0035] As used herein, the phrase "responsive to" describes one or more factors that trigger an effect. This phrase does not exclude the possibility that additional factors may influence or otherwise cause the effect. That is, the effect may be responsive only to these factors, or to the specified factors and other unspecified factors.
[0036] As used herein, the terms "first," "second," etc. are used as labels for the nouns they precede and do not imply any kind of ordering (e.g., spatial, temporal, logical, etc.) unless otherwise specified. As used herein, the term "or" may be used inclusively, exclusively, or not. For example, the phrase "at least one of x, y, or z" means any one of x, y, and z, and any combination thereof (e.g., x and y, but not z). In some situations, the context of the use of the term "or" may indicate that it is used in an exclusive sense, e.g., "select one of x, y, or z" means that only one of x, y, and z is selected in that instance.
[0037] In the following description, numerous specific details are set forth to provide a thorough understanding of the disclosed embodiments. However, those skilled in the art should recognize that aspects of the disclosed embodiments may be practiced without these specific details. In some instances, well-known structures, computer program instructions, and techniques have not been shown in detail to avoid obscuring the disclosed embodiments.
[0038] U.S. Patent Application No. 13 / 969,372 (now U.S. Pat. No. 8,812,144), filed August 16, 2013, discusses techniques for generating musical content based on one or more musical attributes and is incorporated herein by reference in its entirety. To the extent any interpretation is made based on a perceived discrepancy between the definitions in the '372 application and the remainder of this disclosure, it is intended that the present disclosure apply. The musical attributes may be input by a user or may be determined based on environmental information such as ambient noise, lighting, etc. The '372 disclosure discusses techniques for selecting stored loops and / or tracks or generating new loops / tracks, and for layering selected loops / tracks to generate output musical content.
[0039] U.S. Patent Application No. 16 / 420,456 (now U.S. Patent No. 10,679,596), filed May 23, 2019, discusses techniques for generating musical content and is incorporated herein by reference in its entirety. To the extent any interpretation is made based on a perceived conflict between the definitions in the '456 application and the remainder of this disclosure, it is intended that the present disclosure apply. Music may be generated based on user input or using computer-implemented methods. The '456 disclosure discusses various music generator embodiments.
[0040] The present disclosure generally relates to systems for generating custom musical content by selecting and combining audio tracks based on various parameters. In various embodiments, machine learning algorithms (including neural networks such as deep learning neural networks) are configured to generate and customize musical content for a particular user. In some embodiments, users can create their own control elements, and a computing system can be trained to generate output musical content according to the user's intended functions of the user-defined control elements. In some embodiments, playback data of musical content generated by the techniques described herein may be recorded to record and track use of various musical content by different rights holders (e.g., copyright holders). The various techniques described below can provide custom music that is more relevant to different contexts, facilitate generating music according to specific voices, allow users more control over how music is generated, generate music that achieves one or more specific goals, generate music in real time to accompany other content, and the like.
[0041] As used herein, the term "audio file" refers to audio information for musical content. For example, the audio information may include data describing the musical content as raw audio in a format such as wav, aiff, or FLAC. Properties of the musical content may be included in the audio information. The properties may include quantifiable musical properties such as instrument classification, pitch transcription, beat timing, tempo, file length, and audio amplitude in multiple frequency bins. In some implementations, an audio file contains audio information over a specific time interval. In various embodiments, an audio file includes a loop. The term "loop," as used herein, refers to audio information for a single instrument over a specific time interval. Various techniques discussed with reference to audio files can also be implemented using loops containing a single instrument. While an audio file or loop may be played repeatedly (e.g., a 30-second audio file may be played four times in a row to generate two minutes of musical content), an audio file may also be played once, for example, without repeating.
[0042] In some embodiments, image representations of music files are generated and used to generate musical content. The image representations of audio files are generated based on data in the audio files and a MIDI representation of the audio files. The image representations may be, for example, two-dimensional image representations of pitch and rhythm determined from the MIDI representation of the audio files. Rules (e.g., composition rules) can be applied to the image representations to select audio files to use to generate new musical content. In various embodiments, machine learning / neural networks are implemented on the image representations to select audio files to combine to generate new musical content. In some embodiments, the image representations are compressed (e.g., low-resolution) versions of the audio files. Compressing the image representations can improve the speed of searching for selected musical content in the image representations.
[0043] In some embodiments, a music generator can generate new musical content based on various parametric representations of a music file. For example, an audio file typically has an audio signal that can be represented as a graph of the signal versus time (e.g., signal amplitude, frequency, or a combination thereof). However, the time-based representation is dependent on the tempo of the musical content. In various embodiments, the audio file is also represented using a graph of the signal versus beat (e.g., a signal graph). The signal graph is tempo-independent, allowing for tempo-invariant modification of the audio parameters of the musical content.
[0044] In some embodiments, the music generator allows a user to create and label user-defined controls. For example, a user can create controls that the music generator can then train to affect music according to the user's preferences. In various embodiments, the user-defined controls are high-level controls, such as controls that adjust mood, intensity, or genre. Such controls are typically subjective measures based on the listener's personal preferences. In some embodiments, a user creates and labels controls for user-defined parameters. The music generator then plays various music files and allows the music to change according to the user-defined parameters. The music generator can learn and remember the user's preferences based on the user's adjustment of the user-defined parameters. Thus, during subsequent playback, the user-defined controls for the user-defined parameters can be adjusted by the user, and the music generator will adjust the music playback according to the user's preferences. In some embodiments, the music generator can also select music content according to the user's preferences set by the user-defined parameters.
[0045] In some implementations, the musical content generated by the music generator includes music with various stakeholders (e.g., rights holders or copyright owners). In commercial applications that continuously play the generated musical content, compensation based on the playback of individual audio tracks (files) is difficult. Therefore, in various embodiments, techniques are implemented for recording playback data of the continuous musical content. The recorded playback data may include information related to the playback times of individual audio tracks within the continuous musical content matched with stakeholders for each audio track. Furthermore, techniques may be implemented to prevent tampering of the playback data information. For example, the playback data information may be stored in a publicly accessible, fixed blockchain ledger.
[0046] This disclosure first describes an example music generation module and an overall system configuration with multiple applications with reference to Figures 1 and 2. Techniques for generating musical content from visual representations are described with reference to Figures 3-7. Implementation methods for implementing user-generated control elements are described with reference to Figures 8 and 10. Techniques for realizing audio technology are described with reference to Figures 11-17. Techniques for recording information about generated music or elements on a blockchain or other cryptographic ledger are described with reference to Figures 18-19. Figures 20A-20B show example application interfaces.
[0047] In general, the disclosed music generator includes audio files, metadata (e.g., information describing the audio files), and a grammar for combining the audio files based on the metadata. The generator can generate a musical experience using rules for identifying audio files based on the metadata and target characteristics of the musical experience. It can be configured to expand the set of experiences that can be created by adding or modifying rules, audio files, and / or metadata. Adjustments can be performed manually (e.g., an artist adds new metadata), or the music generator can augment the rules / audio files / metadata as it monitors the musical experience and desired goals / characteristics within a given environment. For example, listener-defined controls can be implemented to obtain user feedback regarding musical goals or characteristics.
[0048] Overview of representative music generators 1 is a diagram illustrating an exemplary music generator according to some embodiments. In the illustrated embodiment, a music generation module 160 receives various information from multiple different sources and generates output musical content 140.
[0049] In the illustrated embodiment, module 160 accesses stored audio files and their corresponding attributes 110 and combines the audio files to generate output musical content 140. In some embodiments, music generation module 160 selects audio files based on their attributes and combines the audio files based on target musical attributes 130. In some embodiments, audio files may be selected based on environmental information 150 combined with target musical attributes 130. In some embodiments, environmental information 150 is used indirectly to determine target musical attributes 130. In some embodiments, target musical attributes 130 are explicitly specified by a user, for example, by specifying a desired energy level, mood, multiple parameters, etc. For example, listener-defined controls described herein may be implemented to specify listener preferences to be used as target musical attributes. Examples of target musical attributes 130 include energy, complexity, and variety, although more specific attributes (e.g., corresponding to attributes of stored tracks) may also be specified. In general, if higher-level target musical attributes are specified, lower-level musical attributes may be determined by the system before generating the output musical content.
[0050] Complexity refers to the number of audio files, loops, and / or instruments included in the synthesis. Energy may be related to or orthogonal to other attributes. For example, changing key or tempo may affect energy. However, for a given tempo and key, energy may be changed by adjusting the type of instrument (e.g., by adding hi-hats or white noise), complexity, volume, etc. Diversity may refer to the amount of variation in the generated music over time. Variation may be created relative to a static set of other musical attributes (e.g., by selecting different tracks for a given tempo and key) or by varying musical attributes over time (e.g., by changing tempo and key more frequently if greater variation is desired). In some embodiments, the target musical attributes are considered to exist in a multi-dimensional space, and the music generation module 160 can slowly move through that space, e.g., making course corrections as needed, based on environmental changes and / or user input.
[0051] In some implementations, the attributes stored with the audio files include information about one or more audio files including tempo, volume, energy, diversity, spectrum, envelope, modulation, periodicity, rise and decay times, noise, artist, instrument, theme, etc. Note that in some implementations, the audio files are partitioned such that a set of one or more audio files is specific to a particular audio file type (e.g., one audio source or one audio source type).
[0052] In the illustrated embodiment, module 160 accesses a stored rule set 120. In some embodiments, the stored rule set 120 specifies rules such as the number of overlaying audio files to be played simultaneously (which may correspond to the complexity of the output music), the primary / secondary key progression to be used when transitioning between audio files or musical phrases (which instruments are used together (e.g., instruments that have an affinity with each other)), etc. to achieve targeted musical attributes. Alternatively, music generation module 160 uses the stored rule set 120 to achieve one or more declarative goals defined by target musical attributes (and / or target environmental information). In some embodiments, music generation module 160 includes one or more pseudo-random number generators configured to introduce pseudo-random numbers to avoid repetitive output music.
[0053] In some examples, environmental information 150 includes one or more of lighting information, ambient noise, user information (e.g., facial expression, body position, activity level, movement, skin temperature, performance in specific activities, type of clothing), temperature information, local purchasing activity, time of day, day of the week, season, number of people present, weather, etc. In some embodiments, music generation module 160 does not receive / process environmental information. In some embodiments, environmental information 150 is received by another module that determines target musical attributes 130 based on the environmental information. Target musical attributes 130 can also be derived based on other types of content, e.g., video data. In some embodiments, environmental information is used to adjust one or more stored rule sets 120, for example, to achieve one or more environmental goals. Similarly, the music generator may use environmental information to adjust stored attributes of one or more audio files, for example, to indicate target musical attributes or target audience characteristics with which these audio files are particularly relevant.
[0054] As used herein, the term "module" refers to a circuit configured to perform specified operations or a non-transitory physical computer-readable medium storing information (e.g., program instructions) that directs other circuits (e.g., processors) to perform specified operations. A module can be implemented in multiple ways, including as hardwired circuitry or as memory storing program instructions executable by one or more processors to perform operations. Hardware circuitry can include, for example, custom very large scale integrated circuits or gate arrays, off-the-shelf semiconductors such as logic chips, transistors, or other discrete components. A module can also be implemented in programmable hardware devices such as field programmable gate arrays, programmable array logic, programmable logic devices, etc. A module can also be any suitable form of non-transitory computer-readable medium that stores executable program instructions to perform specified operations.
[0055] As used herein, the phrase "musical content" refers to both the music itself (the audible representation of the music) and information that can be used to play the music. Thus, a song recorded as a file on a storage medium (e.g., a compact disc, a flash drive, etc.) is an example of musical content, and the sound produced by playing (e.g., through a speaker) this recorded file or other electronic representation is also an example of musical content.
[0056] The term "music" has a well-understood meaning that includes sounds produced by musical instruments and vocal sounds. Thus, music includes, for example, instrumental performances or recordings, a cappella performances or recordings, and performances or recordings that include both instruments and voices. Those skilled in the art will recognize that "music" does not encompass all vocal recordings. Works that do not contain musical attributes such as rhythm or cadence, such as speeches, news, and audiobooks, are not music.
[0057] One piece of musical "content" can be appropriately distinguished from other musical content. For example, a digital file corresponding to a first song can represent the musical content of the first song, and a digital file corresponding to a second song can represent the musical content of the second song. The term "musical content" can also be used to distinguish specific intervals within a given musical work, such that different parts of the same song can be considered to have different musical content. Similarly, different tracks (e.g., piano track, guitar track) within a given musical work may correspond to different musical content. In the context of a potentially infinite stream of generated music, the term "musical content" can be used to refer to a portion of the stream (e.g., a few bars or a few minutes).
[0058] Musical content generated by embodiments of the present disclosure may be "new musical content," which is a combination of musical elements not previously generated. A related (but more expansive) concept, "original musical content," is described in more detail below. To facilitate the explanation of this term, the concept of a "controlling entity" is described for an instance of musical content generation. Unlike "original musical content," "new musical content" does not refer to the concept of a controlling entity. Thus, new musical content refers to musical content that has never been generated by any entity or computer system.
[0059] Conceptually, this disclosure refers to some "entity" as controlling a particular instance of computer-generated musical content. Such entity owns the legal rights (e.g., copyright) corresponding to the computer-generated content (to the extent that such rights actually exist). In one embodiment, the individual who creates the computer-implemented music generator (e.g., codes the various software routines) or who operates (e.g., provides inputs to) a particular instance of computer-implemented music generation is the controlling entity. In other embodiments, a computer-implemented music generator may be created by an entity (e.g., a corporation or other business organization), such as in the form of a software product, computer system, or computing device. In some cases, such a computer-implemented music generator may be deployed to many clients. Depending on the terms of the license associated with the distribution of this music generator, the controlling entity may be the creator, the distributor, or the client in various cases. In the absence of such an explicit legal agreement, the controlling entity of a computer-implemented music generator is the entity that facilitates (e.g., provides inputs to and thereby operates) a particular computer-generated instance of musical content.
[0060] For purposes of this disclosure, computer-generated "original musical content" by a controlling entity refers to 1) a combination of musical elements not previously generated by either the controlling entity or another party, and 2) a combination of musical elements generated by the controlling entity but not, in the first instance. Content type 1) is referred to herein as "new musical content," similar to the definition of "new musical content," except that the definition of "new musical content" refers to the concept of a "controlling entity" while the definition of "new musical content" does not. Content type 2, on the other hand, is referred to herein as "proprietary musical content." The term "proprietary" here does not imply any implied legal rights in the content (if such rights exist), but is used merely to indicate that the musical content was originally generated by the controlling entity. Thus, a controlling entity "playing" musical content previously and originally generated by a controlling entity constitutes "generation of original musical content" for purposes of this disclosure. "Non-original musical content" with respect to a particular controlling entity refers to musical content that is not the controlling entity's "original musical content."
[0061] A portion of musical content may include musical components from one or more other musical content. Creating musical content in this manner is called "sampling" musical content and is common in certain musical works, particularly in certain genres. Such musical content is referred to herein as "musical content with sampled components," "derived musical content," or using other similar terms. In contrast, musical content that does not include sampled components is referred to herein as "musical content without sampled components," "non-derived musical content," or using other similar terms.
[0062] In applying these terms, it should be noted that when particular musical content is reduced to a sufficient level of granularity, it can be argued that the musical content is derivative (and substantially all musical content is derivative). The terms "derivative" and "non-derivative" are not used in this sense in this disclosure. With respect to computer-generated musical content, such computer-generated content is said to be derivative (and results in derivative musical content) if the computer-generated content selects some components from pre-existing musical content of an entity other than the controlling entity (e.g., if a computer program selects specific portions of an audio file of a popular artist's work to include some of the generated musical content). On the other hand, computer-generated musical content is said to be non-derivative (and results in non-derivative musical content) if the computer-generated content does not utilize components of such pre-existing content. It should be noted that some "original musical content" may be derivative musical content, while some may be non-derivative musical content.
[0063] It should be noted that the term "derivative" is intended to have a broader meaning in this disclosure than the term "derivative" as used in U.S. copyright law. For example, derivative musical content may or may not be a derivative work under U.S. copyright law. The term "derivative" in this disclosure is not intended to convey a negative connotation, but is merely used to indicate whether a particular piece of musical content "borrows" part of the content from another copyrighted work.
[0064] Furthermore, the phrases "new musical content," "novel musical content," and "original musical content" do not encompass musical content that is entirely different from a combination of existing musical elements. For example, simply changing a few notes in an existing musical composition does not create new, novel, or original musical content, as those terms are used in this disclosure. Similarly, simply changing the key or tempo or adjusting the relative intensity of frequencies in an existing musical composition does not create new, novel, or original musical content. Furthermore, the terms "new," "novel," and "original" musical content are not intended to cover musical content that is a borderline case between original and non-original content; instead, these terms are intended to cover undeniably and unequivocally original musical content (referred to herein as "protectable" musical content), including musical content that is subject to copyright protection. Furthermore, as used herein, the term "available" musical content refers to musical content that does not infringe the copyrights of entities other than the controlling entity. New and / or original musical content is often protected and available for use. This has advantages in preventing copying of the musical content and / or paying royalties for the musical content.
[0065] Although the various embodiments discussed herein use rule-based engines, various other types of computer-implemented algorithms may be used for any of the computer learning and / or music generation techniques discussed herein, although rule-based approaches may be particularly effective in musical contexts.
[0066] Overview of applications, storage elements, and data that may be used in a typical music system The music generation module can interact with multiple different applications, modules, storage elements, etc. to generate musical content. For example, end users can install one of multiple types of applications for different types of computing devices (e.g., mobile devices, desktop computers, DJ equipment, etc.). Similarly, other types of applications may be provided to enterprise users. By interacting with an application while generating musical content, the music generator can receive external information used to determine target musical attributes in order to update one or more rule sets used to generate the musical content. In addition to interacting with one or more applications, the music generation module can interact with other modules to receive rule sets, updated rule sets, etc. Finally, the music generation module can access one or more rule sets, audio files, and / or generated musical content stored in one or more storage elements. Additionally, the music generation module can store any of the above items in one or more storage elements, which may be local or accessed over a network (e.g., cloud-based).
[0067] 2 is a block diagram illustrating an exemplary overview of a system for generating output musical content based on input from multiple disparate sources, according to some embodiments. In the illustrated embodiment, system 200 includes a rules module 210, a user application 220, a web application 230, an enterprise application 240, an artist application 250, an artist rule generation module 260, a generated music store 270, and an external input 280.
[0068] In the illustrated embodiment, user application 220, web application 230, and enterprise application 240 receive external input 280. In some embodiments, external input 280 includes environmental input, target music attributes, user input, sensor input, etc. In some embodiments, user application 220 is installed on a user's mobile device and includes a graphical user interface (GUI) that allows the user to interact / communicate with rules module 210. In some implementations, web application 230 is not installed on the user device but is configured to run within the user device's browser and can be accessed via a website. In some implementations, enterprise application 240 is an application used by a large entity to interact with music generators. In some implementations, application 240 is used in combination with user application 220 and / or web application 230. In some embodiments, application 240 communicates with one or more external hardware devices and / or sensors to gather information about the surrounding environment.
[0069] In the illustrated embodiment, rules module 210 communicates with user applications 220, web applications 230, and enterprise applications 240 to generate output musical content. In some embodiments, music generator 160 is included in rules module 210. Note that rules module 210 may be included in one of applications 220, 230, and 240, or may be installed on a server and accessed over a network. In some embodiments, applications 220, 230, and 240 receive the generated output musical content from rules module 210 and play the content. In some embodiments, rules module 210 may request input from applications 220, 230, and 240, for example, regarding target musical attributes and environmental information, and use this data to generate musical content.
[0070] Stored rule set 120, in the illustrated embodiment, is accessed by rules module 210. In some embodiments, rules module 210 modifies and / or updates stored rule set 120 based on communications with applications 220, 230, and 240. In some embodiments, rules module 210 accesses stored rule set 120 to generate output musical content. In the illustrated embodiment, stored rule set 120 may include rules from artist rule creation module 260, described in more detail below.
[0071] In the illustrated embodiment, artist application 250 communicates with artist rule generation module 260 (which may be part of the same application or may be cloud-based, for example). In some implementations, artist application 250 allows an artist to create rule sets for particular sounds, for example, based on previous compositions. This functionality is further described in U.S. Pat. No. 10,679,596. In some embodiments, artist rule generation module 260 is configured to store generated artist rule sets for use by rules module 210. A user can purchase rule sets from a particular artist and then use them to generate output music via a particular application. A rule set for a particular artist may be referred to as a signature pack.
[0072] The stored audio files and corresponding attributes 110 are accessed by module 210 in the illustrated embodiment when applying rules to select and combine tracks to generate output musical content 270. In the illustrated embodiment, rules module 210 stores the generated output musical content 270 in a storage element.
[0073] 2 are implemented on a server and accessed over a network, referred to as a cloud-based implementation. For example, stored rule sets 120, audio files / attributes 110, and generated music 270 may all be stored on the cloud and accessed by module 210. In another example, module 210 and / or module 260 may also be implemented in the cloud. In some implementations, generated music 270 is stored in the cloud and digitally watermarked. This allows, for example, not only to mass-produce custom music content but also to detect copies of the generated music.
[0074] In some embodiments, one or more of the disclosed modules are configured to generate other types of content in addition to musical content. For example, the system can be configured to generate output visual content based on target musical attributes, determined environmental conditions, a currently used rule set, etc. As another example, the system can search a database or the internet based on the current attributes of the music being generated and display a collage of images that dynamically change as the music changes and match the attributes of the music.
[0075] Example machine learning approaches As described herein, the music generation module 160 shown in FIG. 1 can implement various artificial intelligence techniques (e.g., machine learning techniques) to generate the output musical content 140. In various embodiments, the implemented AI techniques include a combination of deep neural networks (DNNs) with more traditional machine learning techniques and knowledge-based systems. This combination can tailor the strengths and weaknesses of each of these techniques to the unique challenges of music composition and personalization systems. Musical content is composed of multiple levels. For example, a song has sections, phrases, melodies, notes, and textures. DNNs are effective at analyzing and generating very high and very low levels of detail in musical content. For example, DNNs can classify the texture of a sound as belonging to a clarinet or electric guitar at a low level, or detect verses and choruses at a high level. Intermediate levels of detail in musical content, such as melodic construction, orchestration, etc., are more challenging. DNNs are typically good at capturing a wide range of styles in a single model, and thus, DNNs can be implemented as generative tools with a wide range of representations.
[0076] In some embodiments, music generation module 160 leverages expert knowledge by having human-created audio files (e.g., loops) as the basic units of musical content used by the music generation module. For example, the social context of the expert knowledge can be incorporated through rhythmic, melodic, and textural selection, and heuristics can be recorded at multiple levels of structure. Unlike the separation between DNNs and traditional machine learning based on structural levels, expert knowledge can be applied to any domain where musicality can be improved without imposing overly strong constraints on the trainability of music generation module 160.
[0077] In some embodiments, music generation module 160 uses DNNs to find patterns in how audio files or loops are combined vertically by overlapping layers of audio on top of each other, and horizontally by combining audio files or loops into a sequence. For example, music generation module 160 may implement a long short-term memory (LSTM) recurrent neural network trained on mel-frequency cepstral coefficient (MFCC) audio features of loops used in multi-track audio recordings. In some embodiments, the network is trained to predict and select audio features of a loop for an upcoming beat based on knowledge of audio features of the previous beat. For example, the network may be trained to predict audio features of a loop for the next 8 beats based on knowledge of audio features of the last 128 beats. Thus, the network is trained to utilize low-dimensional feature representations to predict upcoming beats.
[0078] In particular embodiments, music generation module 160 uses known machine learning algorithms to assemble multi-track audio sequences into musical structures with dynamics of intensity and complexity. For example, music generation module 160 can implement a hierarchical hidden Markov model, which can behave like a state machine, with state transitions determined by multiple levels of the hierarchical structure. As an example, a particular type of drop may be more likely to occur after a build-up section, but less likely if there is no drum at the end of the build-up. In various embodiments, the probabilities may be trained transparently, as opposed to DNN training, where what is being learned is more opaque.
[0079] Markov models may handle larger temporal structures, and therefore presenting example tracks may not be easily trained because the example tracks may be too long. Feedback can be provided to the music at any time using a feedback control element (such as a thumbs-up / thumbs-down user interface). In certain embodiments, the feedback control element is implemented as one of the UI control elements 830 shown in FIG. 8. The correlation between the musical structure and the feedback can then be used to update a structural model used for the composition, such as a transition table or Markov model. This feedback can also be gathered directly from measurements of heart rate, sales, or other metrics that allow the system to determine clear classifications. The expertise heuristics described above are designed to be as probabilistic as possible and to train in the same way as Markov models.
[0080] In certain embodiments, training may be performed by a composer or DJ. Such training may be separate from listener training. For example, training performed by a listener (such as a typical user) may be limited to identifying correct or incorrect classifications based on positive and negative model feedback, respectively. For composers and DJs, training may include hundreds of time steps and include details of the layers and volume controls used to more clearly represent the driving factors of musical content. For example, training performed by composers and DJs may include sequence prediction training similar to the global training of the DNN described above.
[0081] In various embodiments, a DNN is trained by incorporating multi-track audio and interface interactions to predict what a DJ or composer will do next. In some embodiments, these interactions can be recorded and used to develop new, more transparent heuristics. In some embodiments, a DNN receives many previous bars of music as input and utilizes a low-dimensional feature representation, as described above, along with additional features that describe changes to the track that the DJ or composer has applied. For example, a DNN can receive the last 32 bars of music as input and utilize a low-dimensional feature representation, along with additional features, to describe changes to the track that the DJ or composer has applied. These changes may include adjustments to the gain of a particular track, applied filters, delays, etc. For example, a DJ may use the same drum loop repeated for five minutes during a performance, but gradually increase the gain and delay on the track over time. Thus, a DNN can be trained to predict such gain and delay changes in addition to loop selection. If no loops are played for a particular instrument (e.g., no drum loops are played), the feature set may be all zeros for that instrument, allowing the DNN to learn that predicting all zeros is a successful strategy that can lead to selective layering.
[0082] In some cases, DJs or composers record live performances using mixers and devices such as TRAKTOR (Native Instruments GmbH). These recordings are typically captured in high resolution (e.g., four-track recording or MIDI). In some embodiments, the system decomposes the recording into component loops, obtaining information about the acoustic quality of each individual loop as well as the combination of loops within the component. By training a DNN (or other machine learning) with this information, the DNN provides the DNN with the ability to correlate both the composition of the loops (e.g., sequencing, succession, loop timing, etc.) and the acoustic quality to inform the music generation module 160 how to create a musical experience similar to the artist's performance, without using the actual loops used in the artist's performance.
[0083] Exemplary Music Generator Using Graphical Representations of Audio Files Widely popular music often has a combination of rhythm, texture, and pitch. When generating music note-by-note for each instrument in a song (as is done by music generators), rules can be implemented based on these combinations to create coherent music. In general, the stricter the rules, the less room there is for creative variation and the greater the likelihood of creating copies of existing music.
[0084] When creating music by combining previously recorded pieces, the combinations must consider multiple invariant combinations of notes in each phrase. However, when drawing from a library of thousands of audio recordings, searching all possible combinations can be computationally expensive. Furthermore, significant comparisons may be necessary to check for harmonically inconsistent combinations, especially with regard to beats. New rhythms created by combining multiple files can also be checked against rules regarding the rhythmic organization of the combined phrases.
[0085] It is not always possible to extract the features necessary to create a combination from an audio file. Even if it is possible, extracting the necessary features from an audio file can be computationally expensive. In various embodiments, a symbolic audio representation is used for music composition to reduce computational costs. The symbolic audio representation can rely on the music composer's storage of instrumental textures, memorized rhythmic and pitch information. A common format for symbolic music representation is MIDI. MIDI includes precise timing, pitch, and performance control information. In some embodiments, MIDI can be simplified and further compressed through a piano roll representation, where musical notes are represented as bars on a discrete time / pitch graph, typically in eight octaves of pitch.
[0086] In some embodiments, the music generator is configured to generate the output musical content by generating image representations of music files and selecting musical combinations based on an analysis of the image representations. The image representations may be further compressed from a piano roll representation. For example, the image representations may be lower-resolution representations generated based on a MIDI representation of an audio file. In various embodiments, composition rules are applied to the image representations to select musical content from music files and combine and generate the output musical content. The composition rules may be applied, for example, using rule-based methods. In some embodiments, a machine learning algorithm or model (e.g., a deep learning neural network) is implemented to select and combine audio files to generate the output musical content.
[0087] 3 is a block diagram illustrating an exemplary music generation system configured to output musical content based on an analysis of graphical representations of audio files, according to some embodiments. In the illustrated embodiment, the system 300 includes a graphical representation generation module 310, a music selection module 320, and a music generation module 160.
[0088] In the illustrated embodiment, image representation generation module 310 is configured to generate one or more image representations of audio files. In certain embodiments, image representation generation module 310 receives audio file data 312 and MIDI representation data 314. MIDI representation data 314 includes a MIDI representation of a particular audio file in audio file data 312. For example, for a particular audio file in audio file data 312, data 312 may have a corresponding MIDI representation in MIDI representation data 314. In some implementations with multiple audio files in audio file data 312, each audio file in audio file data 312 has a corresponding MIDI representation in MIDI representation data 314. In the illustrated embodiment, MIDI representation data 314 is provided to image representation generation module 310 along with audio file data 312. However, in some contemplated embodiments, image representation generation module 310 may itself generate MIDI representation data 314 from audio file data 312.
[0089] 3, image representation generation module 310 generates image representation 316 from audio file data 312 and MIDI representation data 314. MIDI representation data 314 includes pitch, time, and velocity data for notes associated with a music file, while music file data 312 includes data for playing the music itself. In particular embodiments, image representation generation module 310 generates an image representation of the audio file based on the pitch, time, and velocity data from MIDI representation data 314. The image representation may be, for example, a two-dimensional image representation of the audio file. In the 2D image representation of an audio file, the x-axis represents time (rhythm), the y-axis represents pitch (similar to a piano roll representation), and the pixel value for each x and y coordinate represents velocity.
[0090] The 2D image representation of an audio file can have a variety of image sizes, but the image size is typically selected to correspond to the musical structure. For example, in one contemplated embodiment, the 2D image display is 32 (x-axis) by 24 (y-axis). A 32-pixel wide image representation allows each pixel to represent a quarter of a beat in the time dimension. Thus, an eight-beat piece of music can be represented by a 32-pixel wide image representation. While this representation may not have enough detail to capture the expressive detail of the music in the audio file, the expressive detail is retained in the audio file itself, which is used in combination with the image representation by system 300 to generate the output musical content. However, a quarter-beat time resolution provides a fair amount of coverage for common pitch and rhythm combination rules.
[0091] FIG. 4 shows an example of an image representation 316 of an audio file. The image representation 316 is 32 pixels wide (time) and 24 pixels high (pitch). Each pixel (square) 402 has a value that represents the velocity and pitch within the audio file at that time. In various embodiments, the image representation 316 may be a grayscale image representation of the audio file, where pixel values are represented by varying intensities of gray. The gray variations based on pixel value may be small and difficult for many people to notice. FIGS. 5A and 5B show example grayscale images for a melody image feature representation and a drumbeat image feature representation, respectively. However, other representations (e.g., color or numeric values) may also be considered. In these representations, each pixel may have multiple different values corresponding to different musical attributes.
[0092] In particular embodiments, the image representation 316 is an 8-bit representation of the audio file. Thus, each pixel can have 256 possible values. MIDI representations typically have 128 possible values for velocity. In various embodiments, the details of the velocity values are less important than the task of selecting the audio files for combination. Thus, in such embodiments, the pitch axis (y-axis) may be banded to cover two sets of octaves with an eight-octave range of four octaves in each set. For example, the eight octaves can be defined as follows:
number
[0093] Within these defined ranges of octaves, the pixel row and value determine the octave and velocity of the note. For example, a pixel value of 10 in row 1 represents a note in octave 0 at velocity 10, and a pixel value of 74 in row 1 represents a note in octave 2 at velocity 10. As another example, a pixel value of 79 in row 13 represents a note in octave 3 at velocity 15, and a pixel value of 207 in row 13 represents a note in octave 7 at velocity 15. Thus, using the defined ranges of octaves described above, the first 12 rows (rows 0-11) represent the first set of four octaves (octaves 0, 2, 4, and 6), with pixel values determining which of the first four octaves will be represented (the pixel values also determine the velocity of the note). Similarly, the second 12 rows (rows 12-23) represent a second set of four octaves (octaves 1, 3, 5, and 7), with pixel values determining which one of the second four octaves is represented (the pixel values also determine the note velocity).
[0094] As described above, by banding the pitch axis to cover an eight-octave range, the velocity of each octave can be defined by 64 values rather than the 128 values of the MIDI representation. Therefore, the 2D image representation (e.g., image representation 316) can be more compressed (e.g., lower resolution) than the MIDI representation of the same audio file. In some embodiments, 64 values may be greater than the values required by system 300 to select musical combinations, so further compression of the image representation may be permitted. For example, by having odd pixel values represent the onset of a note and pixel values represent the duration of a note, the velocity resolution can be further reduced, allowing for compression in the temporal representation. This reduced resolution allows two consecutive notes played at the same velocity to be distinguished from one long note based on the odd or even pixel values.
[0095] As noted above, the compactness of the image representation reduces the size of the files required to represent the music (e.g., compared to a MIDI representation). Therefore, implementing an image representation of an audio file can reduce the amount of disk storage required. Furthermore, the compressed image representation may be stored in high-speed memory, allowing for rapid searching of possible musical combinations. For example, 8-bit image representations can be stored in graphics memory on a computing device, allowing large parallel searches to be performed together.
[0096] In various embodiments, image representations generated for multiple audio files are combined into a single image representation. For example, image representations for tens, hundreds, or thousands of audio files can be combined into a single image representation. The single image representation can be a large, searchable image that can be used for parallel searches of the multiple audio files that make up the single image. For example, the single image can be searched in a manner similar to large textures in video games using software such as MegaTextures (from id Software).
[0097] 6 is a block diagram illustrating an exemplary system configured to generate a single image representation, according to some embodiments. In the illustrated embodiment, system 600 includes a single image representation generation module 610 and a texture feature extraction module 620. In particular embodiments, single image representation generation module 610 and texture feature extraction module 620 are located within image representation generation module 310 shown in FIG. 3. However, single image representation generation module 610 or texture feature extraction module 620 may be located outside image representation generation module 310.
[0098] As shown in the exemplary embodiment of FIG. 6, multiple image representations 316A-N are generated. The image representations 316A-N may be N individual image representations corresponding to the N number of individual audio files. The single image representation generation module 610 may combine the individual image representations 316A-N into a single combined image representation 316. In some embodiments, the individual image representations combined by the single image representation generation module 610 include individual image representations for different instruments. For example, different instruments in an orchestra may be represented by individual image representations, which are then combined into a single image representation for musical exploration and selection.
[0099] In certain embodiments, the individual image representations 316A-N are combined with individual image representations 316 located adjacent to one another without overlapping. Thus, the single image representation 316 is a complete dataset representation of all individual image representations 316A-N without loss of data (e.g., without data from one image representation modifying data for another image representation). Figure 7 shows an example of a single image representation 316 of multiple audio files. In the illustrated embodiment, the single image representation 316 is a combined image created from the individual image representations 316A, 316B, 316C, and 316D.
[0100] In some embodiments, the single image representation 316 is appended with texture features 622. In the illustrated embodiment, the texture features 622 are added to the single image representation 316 as a single line. Referring to Figure 6, the texture features 622 are determined by a texture feature extraction module 620. The texture features 622 may include, for example, instrument textures of the music in the audio file. For example, the texture features may include features from different instruments such as drums, string instruments, etc.
[0101] In particular embodiments, the texture feature extraction module 620 extracts texture features 622 from the audio file data 312. The texture feature extraction module 620 may implement, for example, a rule-based method, a machine learning algorithm or model, a neural network, or other feature extraction technique to determine texture features from the audio file data 312. In some embodiments, the texture feature extraction module 620 may extract texture features 622 from the image representations 316 (e.g., multiple image representations or a single image representation). For example, the texture feature extraction module 620 may implement an image-based analysis (e.g., an image-based machine learning algorithm or model) to extract the texture features 622 from the image representations 316.
[0102] The addition of texture features 622 to single-image representation 316 provides the single-image representation with additional information that is not typically available in a MIDI or piano roll representation of an audio file. In some embodiments, the lines with texture features 622 in single-image representation 316 (shown in FIG. 7 ) may not need to be human-readable. For example, texture features 622 may only need to be machine-readable for implementation in a music generation system. In certain embodiments, texture features 622 are added to single-image representation 316 and used in image-based analysis of the single-image representation. For example, texture features 622 may be used by an image-based machine learning algorithm or model used for music selection, as described below. In some embodiments, texture features 622 may be ignored during music selection, for example, in rule-based selection, as described below.
[0103] 3 , in the illustrated embodiment, image representations 316 (e.g., multiple image representations or a single image representation) are provided to a music selection module 320. The music selection module 320 can select audio files or portions of audio files to be combined in the music generation module 160. In particular embodiments, the music selection module 320 applies a rule-based method to search for and select audio files or portions of audio files for combination by the music generation module 160. As shown in FIG. 3 , the music selection module 320 accesses rules for the rule-based method from a stored rule set 120. For example, the rules accessed by the music selection module 320 may include search and selection rules, such as, but not limited to, composition rules and note combination rules. The application of the rules to the image representation 316 can be performed using graphics processing available on a computing device.
[0104] For example, in various embodiments, note combination rules may be expressed as vector and matrix calculations. Graphics processing units are typically optimized for performing vector and matrix calculations. Note, for example, that intervals of one pitch step may typically be inconsistent and often avoided. Notes such as these may be found by searching neighboring pixels in an additionally layered image (or segment of a larger image) based on the rules. Thus, in various embodiments, the disclosed modules may invoke kernels to perform all or part of the disclosed operations on a graphics processor of a computing device.
[0105] In some embodiments, the pitch banding in the above-described graphical representation enables the use of graphics processing to embed high-pass or low-pass filtering of the audio. Removing (e.g., filtering out) pixel values below a threshold can simulate high-pass filtering, and removing pixel values above a threshold can simulate low-pass filtering. For example, in the above-described banding example, filtering out (removing) pixel values below 64 can have a similar effect to applying a high-pass filter with a shelf at B1, in the example by removing octaves 0 and 1. Thus, the use of filters on each audio file can be efficiently simulated by applying rules to the graphical representation of the audio file.
[0106] In various embodiments, when overlapping music files to create music, the pitch of certain audio files can be altered. Altering the pitch can increase the range of successful combinations and potentially broaden the combinatorial search space. For example, each audio file can be tested with 12 different pitch-shifted keys. Offsetting the row order of the image representation when analyzing the image and adjusting the octave shift as needed may enable an optimized search for these combinations.
[0107] In particular embodiments, the music selection module 320 implements machine learning algorithms or models on the image representations 316 to search for and select audio files or portions of audio files for combination by the music generation module 160. The machine learning algorithms / models may include, for example, deep learning neural networks or other machine learning algorithms that classify images based on training of the algorithm. In such embodiments, the music selection module 320 includes one or more machine learning models that are trained based on combinations and sequences of audio files that provide desired musical characteristics.
[0108] In some embodiments, the music selection module 320 includes a machine learning model that continuously learns during the selection of the output musical content. For example, the machine learning model can receive user input or other input reflecting characteristics of the output musical content that can be used to adjust classification parameters performed by the machine learning model. Similar to rule-based methods, the machine learning model can be implemented using a graphics processing unit on a computing device.
[0109] In some embodiments, the music selection module 320 implements a combination of rule-based methods and machine learning models. In one contemplated embodiment, a machine learning model is trained to initiate a search for musical content and find combinations of audio files and image representations to combine when the search is performed using rule-based methods. In some embodiments, the music selection module 320 tests harmonic and rhythmic rule coherence in music selected for combination by the music generation module 160. For example, the music selection module 320 can test harmony and rhythm in selected audio files 322 before providing the selected audio files to the music generation module 160, as described below.
[0110] 3, as described above, the music selected by the music selection module 320 is provided to the music generation module 160 as selected music files 322. The selected audio files 322 may include complete or partial audio files that are combined by the music generation module 160 to generate the output musical content 140, as described herein. In some embodiments, the music generation module 160 accesses the stored rule set 120 to retrieve the rules that were applied to the selected audio files 322 to generate the output musical content 140. The rules retrieved by the music generation module 160 may differ from the rules applied by the music selection module 320.
[0111] In some implementations, the selected audio files 322 include information for combining the selected audio files. For example, the machine learning model implemented by the music selection module 320 may provide output instructions describing how to combine the musical content in addition to selecting the music to combine. These instructions are then provided to the music generation module 160 and implemented by the music generation module to combine the selected audio files. In some implementations, the music generation module 160 tests for harmonic and rhythmic rule coherence before finalizing the output musical content 140. Such tests may be performed in addition to or instead of the tests performed by the music selection module 320.
[0112] Exemplary Control for Musical Content Generation In various embodiments, as described herein, a music generation system is configured to automatically generate output musical content by selecting and combining audio tracks based on various parameters. As described herein, machine learning models (or other AI techniques) are used to generate the musical content. In some implementations, AI techniques are implemented to customize musical content for a particular user. For example, the music generation system can implement various types of adaptive control to personalize music generation. Personalizing music generation allows for content control by the composer or listener in addition to content generation by AI techniques. In some embodiments, a user can create unique control elements that the music generation system can train (e.g., using AI techniques) to generate output musical content according to the intended function of the user-created control elements. For example, a user can create control elements that train the music generation system to influence music according to the user's preferences.
[0113] In various embodiments, the user-created control elements are high-level controls, such as controls that adjust mood, intensity, or genre. Such user-created control elements are typically subjective measures based on the listener's individual preferences. In some embodiments, the user labels the user-created control elements and defines user-specified parameters. The music generation system can play various musical content, and the user can use the control elements to modify the user-specified parameters in the musical content. The music generation system can learn and remember how the user-defined parameters change audio parameters in the musical content. Thus, during subsequent playback, the user-created control elements can be adjusted by the user, and the music generation system adjusts the audio parameters in the music playback according to the adjustment level of the user-specified parameters. In some contemplated embodiments, the music generation system can also select musical content according to the user's preferences as set by the user-specified parameters.
[0114] 8 is a block diagram illustrating an exemplary system configured to implement user-generated control over musical content generation, according to some embodiments. In the illustrated embodiment, system 800 includes music generation module 160 and user interface module 820. In various embodiments, music generation module 160 implements the techniques described herein to generate output musical content 140. For example, music generation module 160 can access stored audio files 810 and generate output musical content 140 based on stored rule sets 120.
[0115] In various embodiments, music generation module 160 modifies the musical content based on input from one or more UI control elements 830 implemented in UI module 820. For example, a user may adjust the level of a control element 830 while interacting with UI module 820. Examples of control elements include, but are not limited to, a slider, a dial, a button, or a knob. The level of the control element 830 sets the control element level 832 that is provided to music generation module 160. Music generation module 160 may then modify the output musical content 140 based on the control element level 832. For example, music generation module 160 may implement AI techniques to modify the output musical content 140 based on the control element level 830.
[0116] In certain embodiments, one or more of the control elements 830 are user-defined control elements. For example, the control elements may be defined by a composer or a listener. In such embodiments, a user may create and label UI control elements that specify parameters the user wants to implement to control the output musical content 140 (e.g., the user creates a control element to control a user-specified parameter in the control output musical content 140).
[0117] In various embodiments, music generation module 160 may be learned or trained to affect output musical content 140 in a particular way based on input from user-generated control elements. In some embodiments, music generation module 160 is trained to modify audio parameters in output musical content 140 based on levels of user-generated control elements set by the user. Training music generation module 160 may include, for example, determining a relationship between the audio parameters in output musical content 140 and the levels of the user-created control elements. The relationship between the audio parameters in output musical content 140 and the levels of the user-generated control elements may then be utilized by music generation module 160 to modify output musical content 140 based on the input levels of the user-generated control elements.
[0118] 9 shows a flowchart of a method for training music generation module 160 based on user-created control elements, according to some embodiments. Method 900 begins with a user creating and labeling control elements in 910. For example, as described above, the user may create and label UI control elements to control user-specified parameters in output musical content 140 generated by music generation module 160. In various embodiments, the labels of the UI control elements describe the user-specified parameters. For example, the user may label a control element as "Attitude" to specify that the user wishes to control a (user-defined) attitude in the generated musical content.
[0119] After creating the UI control elements, the method 900 continues with a playback session 915. The playback session 915 can be used to train the system (e.g., music generation module 160) how to modify audio parameters based on the levels of the UI control elements created by the user. During the playback session 915, an audio track is played at 920. The audio track may be a musical loop or sample from an audio file stored on or accessed by the device.
[0120] At 930, the user provides input regarding his / her interpretation of the user-specified parameters in the audio track being played. For example, in certain embodiments, the user is asked to listen to the audio track and select levels of the user-specified parameters that the user believes describe the music in the audio track. The levels of the user-specified parameters may be selected, for example, using control elements created by the user. This process may be repeated for multiple audio tracks in the playback session 915 to generate multiple data points for the levels of the user-specified parameters.
[0121] In some contemplated embodiments, a user may be required to listen to multiple audio tracks at once and comparatively rate the audio tracks based on user-defined parameters. For example, in an example of user-generated control defining "attributes," a user may listen to multiple audio tracks and select which audio tracks have more of the "attribute" and / or which audio tracks have fewer of the "attribute." Each selection made by the user may be a data point for the level of the user-specified parameter.
[0122] After the playback session 915 is completed, the levels of audio parameters in the audio tracks from the playback session are evaluated at 940. Examples of audio parameters include, but are not limited to, volume, tone, bass, treble, reverb, etc. In some embodiments, the levels of audio parameters in the audio tracks are evaluated as the audio tracks are played (e.g., during the playback session 915). In some embodiments, the audio parameters are evaluated after the playback session 915 has ended.
[0123] In various embodiments, audio parameters within an audio track are evaluated from metadata of the audio track. For example, audio analysis algorithms can be used to generate metadata or symbolic music data (e.g., MIDI) for the audio track, which may be a short, pre-recorded music file. The metadata can include, for example, notes present in the recording, the number of onsets per beat, the ratio to unpitched notes, volume level, and other quantifiable characteristics of the notes.
[0124] Correlations between user-selected levels for user-specified parameters and audio parameters are determined at 950. Because user-selected levels for user-specified parameters correspond to levels of control elements, the correlations between user-selected levels for user-specified parameters and audio parameters can be used to define relationships between levels of one or more audio parameters and levels of control elements in 960. In various embodiments, the correlations between levels of user-selected parameters and audio parameters, and the relationships between levels of one or more audio parameters and levels of control elements, are determined using AI techniques (e.g., regression models or machine learning algorithms).
[0125] Returning to FIG. 8 , the relationships between the levels of one or more audio parameters and the levels of the control elements may then be implemented by music generation module 160 to determine how to adjust the audio parameters in output musical content 140 based on control element level 832 inputs received from user-created control elements 830. In particular embodiments, music generation module 160 implements a machine learning algorithm to generate output musical content 140 based on the control element level 832 inputs and relationships received from user-created control elements 830. For example, the machine learning algorithm may analyze how the metadata description of an audio track changes throughout recording. The machine learning algorithm may include, for example, a neural network, a Markov model, or a dynamic Bayesian network.
[0126] As described herein, a machine learning algorithm can be trained to predict the metadata of an upcoming piece of music, given the metadata of the music up to that point. The music generation module 160 can implement the prediction algorithm by searching a pool of pre-recorded audio files for those with properties closest to the predicted upcoming metadata. Selecting the closest matching audio file to play next helps create output musical content with a progression of musical properties similar to the example recordings on which the prediction algorithm was trained.
[0127] In some embodiments, parametric control of the music generation module 160 that uses a prediction algorithm may be included in the prediction algorithm itself. In such embodiments, some predefined parameters may be used as input to the algorithm, along with music metadata, and the predictions vary based on these parameters. Alternatively, parametric control may be applied to the predictions to modify them. As an example, generative synthesis is performed by sequentially selecting the closest musical fragments predicted by the prediction algorithm to come next and appending the audio from the file to the end. At some point, the listener can increase the control element level (such as the beat-by-beat start control element), and the output of the prediction model is modified by increasing the predicted beat-by-beat start data field. When selecting the next audio file to add to the composition, in this scenario, an audio file with a large beat-by-beat property is more likely to be selected.
[0128] In various embodiments, a generative system such as music generation module 160 that utilizes metadata descriptions of musical content can use hundreds or thousands of data fields in the metadata of each musical fragment. To provide greater variety, multiple simultaneous tracks featuring different sound sources and sound types may be used. In such cases, the predictive model may have thousands of data fields representing musical properties, each with a distinct effect on the listening experience. In such cases, an interface may be used to modify each data field of the predictive model's output to generate thousands of control elements for the listener to control the music. Alternatively, multiple data fields may be combined and exposed as a single control element. The more musical properties are affected by a single control element, the more abstract the control element becomes from specific musical properties, and the more subjective the labeling of these controls becomes. In this manner, primary control elements and subparameter control elements (described below) can be implemented for dynamic and personalized control of output musical content 140.
[0129] As described herein, users can specify their own control elements and train the music generation module 160 on how to behave based on user adjustments of the control elements. This process reduces bias and complexity, and data fields may be completely hidden from the listener. For example, in some embodiments, the listener is provided with user-created control elements on a user interface. The listener is then shown a short music clip and asked to set the levels of the control elements that they believe best represent the music they have heard. Repeating this process generates multiple data points that can be used to recursively model the desired effect of the controls on the music. In some embodiments, these data points can be added as additional inputs in a predictive model. The predictive model then attempts to predict musical properties that will produce composition sequences similar to the trained sequences, while also matching the expected behavior of control elements set at specific levels. Alternatively, a control element mapper in the form of a regression model may be used to map predictive modifiers to control elements without retraining the predictive model.
[0130] In some embodiments, training for a given control element may include both global training (e.g., training based on feedback from multiple user accounts) and local training (e.g., training based on feedback from the current user account). In some embodiments, a set of control elements specific to a subset of musical elements provided by a composer may be created. For example, a scenario may involve an artist creating a loop pack and then using these loops to train the music generation module 160 using previously created performance or synthesis examples. Patterns in these examples can be modeled with a regression or neural network model and used to create rules for the construction of new music with similar patterns. These rules can be parameterized and exposed as control elements for the composer to manually modify offline before beginning use of the music generation module 160, or for the listener to adjust while listening. Examples that the composer perceives as the opposite of the desired effect of the control can also be used for negative reinforcement.
[0131] In some embodiments, in addition to utilizing exemplary musical patterns, music generation module 160 can generate musical patterns that correspond to input from the composer before the listener begins listening to the generated music. The composer can do this through direct feedback (described below), for example, tapping a favorable control element for positive reinforcement of the pattern or a unfavorable control element for negative reinforcement.
[0132] In various embodiments, the music generation module 160 may allow a composer to create their own sub-parameter control elements, as described below, of the control elements learned by the music generation module. For example, a control element for "intensity" may be created as a primary control element from learned patterns related to the number of notes expressed per beat and the texture quality of the instrument being played. The composer can then create two sub-parameter control elements by selecting patterns related to the expression of notes, such as a "rhythm intensity" control element and a "texture intensity" control element for a texture pattern. Examples of sub-parameter control elements include control elements for vocals, the intensity of a particular frequency range (e.g., bass), complexity, tempo, etc. These sub-parameter control elements can be used in conjunction with more abstract control elements (e.g., primary control elements) such as energy. These composer-skill control elements may be trained for the music generation module 160 by the composer, similar to the user-created controls described herein.
[0133] As described herein, training the music generation module 160 to control audio parameters based on input from user-created control elements allows individual control elements to be implemented for different users. For example, one user may associate increased attitude with increased bass content, while another user may associate increased attitude with a particular type of vocal or a particular tempo range. The music generation module 160 can modify audio parameters for different specifications of attitude based on the training of the music generation module for a particular user. In some embodiments, individualized control can be used in combination with global rules or control elements that are implemented identically for many users. The combination of global and local feedback or control can provide a quality music production that provides special control to the individuals involved.
[0134] 8 , one or more UI control elements 830 are implemented within the UI module 820. As described above, a user can use the control elements 830 to adjust a control element level 832 while interacting with the UI module 820 to modify the output musical content 140. In particular embodiments, one or more of the control elements 830 are system-defined control elements. For example, the control elements may be defined as parameters controllable by the system 800. In such embodiments, the user can adjust the system-defined control elements to modify the output musical content 140 according to system-defined parameters.
[0135] In particular embodiments, system-defined UI control elements (e.g., knobs or sliders) allow a user to control abstract parameters of the output musical content 140 automatically generated by the music generation module 160. In various embodiments, the abstract parameters act as primary control element inputs. Examples of abstract parameters include, but are not limited to, intensity, complexity, mood, genre, and energy level. In some embodiments, an intensity control element may adjust the number of incorporated low-frequency loops. A complexity control element can guide the number of overlaid tracks. Other adjustment elements, such as mood adjustment elements, range from calm to happy, affecting, for example, the musical key being played, among other attributes.
[0136] In various embodiments, a system-defined UI control element (e.g., a knob or slider) allows a user to control the energy level of output musical content 140 automatically generated by music generation module 160. In some embodiments, the label of the control element (e.g., "Energy") may change in size, color, or other characteristic to reflect the user input adjusting the energy level. In some embodiments, as the user adjusts the control element, the current level of the control element may be output until the user releases the control element (e.g., releases a mouse click or removes a finger from a touchscreen).
[0137] Energy may be an abstract parameter related to multiple, more specific musical attributes, as defined by the system. As an example, energy may be related to tempo in various embodiments. For example, changes in energy level may be related to tempo changes of a selected number of beats per minute (e.g., 6 beats per minute). In some embodiments, within a given range for one parameter (e.g., tempo), music generation module 160 can explore musical variations by modifying other parameters. For example, music generation module 160 can create buildups and drops, create tension, vary the number of simultaneously layered tracks, change the key, add or remove vocals, add or remove bass, play different melodies, etc.
[0138] In some embodiments, one or more sub-parameter control elements are implemented as control elements 830. Sub-parameter control elements may allow for more specific control of attributes incorporated into primary control elements, such as an energy control element. For example, an energy control element may change the number of percussion layers and amount of vocals used, but separate control elements may control these sub-parameters directly, so that not all control elements are necessarily independent. In this way, a user can select the level of specificity of control they wish to utilize. In some embodiments, sub-parameter control elements may be implemented for user-created control elements described above. For example, a user may create and label control elements that specify sub-parameters of other user-specified parameters.
[0139] In some embodiments, the user interface module 820 allows the user the option to extend the UI control elements 830 to reveal one or more sub-parameter user control elements. In addition, a particular artist can provide attribute information used to guide music synthesis under user control of a high-level control element (e.g., an energy slider). For example, an artist can provide an "artist pack" with tracks from that artist and rules for music composition. The artist can use the artist interface to provide values for sub-parameter user control elements. For example, a DJ might have rhythm and drums as control elements exposed to the user, allowing listeners to incorporate more or less rhythm and drums. In some embodiments, artists or users can create their own custom control elements, as described herein.
[0140] In various embodiments, a human-in-the-loop generation system can be used to generate artifacts with the aid of human intervention and control, potentially improving the quality and adaptability of the generated music to personal purposes. In some embodiments of the music generation module 160, the listener can become a listener-composer by controlling the generation process via interface control elements 830 implemented in the UI module 820. The design and implementation of these control elements can affect the balance between the individual listener and composer roles. For example, highly detailed and technical control elements reduce the influence of the generation algorithm and place more creative control in the hands of the user, while requiring more hands-on interaction and technical skill to manage.
[0141] Conversely, a higher level of control can reduce the effort and interaction time required while decreasing creative control. For example, for individuals who desire a more listener-type role, a primary control element, as described herein, may be preferred. The primary control element can be based on abstract parameters such as mood, intensity, or genre. These abstract music parameters are often subjective measures that are individually interpreted. For example, the listening environment often influences how a listener describes music. Thus, music that a listener calls "relaxing" at a party may be too energetic and tense for a meditation session.
[0142] In some implementations, one or more UI control elements 830 are implemented to receive user feedback regarding the output musical content 140. User feedback control elements may include, for example, star ratings, thumbs-up / thumbs-down, etc. In various embodiments, user feedback may be used to train the system to a user's specific tastes and / or more global tastes that apply to multiple users. In embodiments with thumbs-up / thumbs-down (e.g., positive / negative) feedback, the feedback is binary. Binary feedback, including strong positive and strong negative responses, may be effective in providing positive and negative reinforcement for the functioning of the control elements 830. In some possible embodiments, input from the thumbs-up / thumbs-down control elements may be used to control the output musical content 140 (e.g., the thumbs-up / thumbs-down control element is used to control the output itself). For example, the thumbs-up control element may be used to change the maximum number of repeats of the currently playing musical content 140.
[0143] In some implementations, a counter for each audio file tracks the number of times a section of that audio file (e.g., an 8-beat segment) has been recently played. Once a file has been used beyond a desired threshold, a bias can be applied to its selection. This bias can gradually return to zero over time. This repetition counter and bias, along with rule-defined musical sections that set the desired function of the music (e.g., buildup, dropdown, breakdown, intro, sustain), can be used to shape the music into segments with a coherent theme. For example, the music generation module 160 can increment the counter with a negative press to encourage the audio content of the output musical content 140 to change more quickly without compromising the section's musical function. Similarly, the music generation module 160 can decrement the counter with a positive press to prevent the audio content of the output musical content 140 from biasing away from repetition for longer periods of time. Other machine learning and rule-based mechanisms within the music generation module 160 can guide the selection of other audio content before the threshold is reached and the bias is applied.
[0144] In some embodiments, music generation module 160 is configured to determine various contextual information (e.g., environmental information 150 shown in FIG. 1 ) before and after user feedback is received. For example, upon receiving a "thumbs up" indication from the user, music generation module 160 can determine the time of day, location, device speed, biometric data (e.g., heart rate), etc. from environmental information 150. In some embodiments, this contextual information may be used to train a machine learning model to generate music that users will prefer in a variety of different contexts (e.g., the machine learning model is context-aware).
[0145] In various embodiments, music generation module 160 determines the current type of environment and takes different actions for the same user adjustment in different environments. For example, music generation module 160 can take environmental measurements and listener biometric measurements as the listener trains the “Attitude” control element. During training, music generation module 160 is trained to include these measurements as part of the control element. In this example, when the listener is performing high-intensity exercise in a gym, the “Attitude” control element may affect the intensity of a drum beat. When sitting at a computer, changing the “Attitude” control element may not affect the drum beat but may increase baseline distortion. In such embodiments, a single user control element may have different sets of rules or differently trained machine learning models that are used differently in different listening environments, alone or in combination.
[0146] In contrast to context-awareness, if the expected behavior of a control element is static, a large number of controls may be necessary or desirable for all uses of the listening context music generation module 160. Thus, in some embodiments, the disclosed technology may provide functionality for multiple environments with a single control element. Implementing a single control element for multiple environments can reduce the number of control elements, making the user interface simpler and faster to navigate. In some embodiments, the behavior of the control element is made dynamic. The dynamism of the control element is achieved by utilizing environmental measurements, such as sound levels recorded by a microphone, heart rate measurements, time of day, and movement speed. These measurements can be used as additional inputs to training the control element. Thus, the same listener interaction with the control element can potentially have different musical effects depending on the environmental context in which the interaction occurs.
[0147] In some embodiments, the context-aware features described above are distinct from the concept of generative music systems that alter the generative process based on environmental context. For example, these techniques can alter the effect of user-controlled elements based on environmental context, which can be used alone or in combination with the concept of generating music based on environmental context and user-controlled output.
[0148] In some embodiments, music generation module 160 is configured to control the generated output musical content 140 to achieve a predetermined goal. Examples of stated goals include, but are not limited to, sales goals, biometric goals such as heart rate or blood pressure, and ambient noise goals. Music generation module 160 can learn how to modify manually (user-created) or algorithmically (system-defined) generated control elements using the techniques described herein to generate output musical content 140 to achieve the predetermined goal.
[0149] A goal state may be a measurable environment that a listener describes that they hope to achieve while listening to music using, and with the help of, music generation module 160. These goal states may be influenced directly or mediated by psychological effects, such as specific music-encouraging focus. That is, they may be influenced by the music changing the listener's acoustic experience of the space. As an example, a listener may set a goal of lowering their heart rate while running. By recording the listener's heart rate under different states of the available control elements, music generation module 160 learned that the listener's heart rate typically decreases when a control element named "Attitude" is set to a low level. Thus, to help the listener achieve a lower heart rate, music generation module 160 may automate the "Attitude" control to a low level.
[0150] By generating the type of music a listener expects in a particular environment, music generation module 160 can help create that particular environment. Examples include heart rate, overall volume in the listener's physical space, sales in a store, etc. Some environmental sensor and state data may not be appropriate for the target state. For example, time of day may be an environmental measure used as an input to achieve a sleep-inducing target state, but music generation module 160 cannot control the time of day itself.
[0151] In various embodiments, sensor inputs may be disconnected from the control element mapper while attempting to reach a state goal, but the sensors may continue to record and instead provide measurements for comparing the actual goal state to the target state. The difference between the goal state and the actual environmental state may be formulated as a reward function for a machine learning algorithm that can adjust the mapping in the control element mapper in a mode that attempts to achieve the goal state. The algorithm can adjust the mapping to reduce the difference between the target and actual environmental states.
[0152] While musical content has many physiological and psychological effects, creating musical content that listeners expect in a particular environment does not necessarily help create that environment for the listener. In some cases, it may have no beneficial or adverse effect on reaching a target state. In some embodiments, music generation module 160 can adjust musical properties based on past results while branching in other directions if the change does not meet a threshold. For example, if decreasing the "Attitude" control element does not decrease the listener's heart rate, music generation module 160 can transition and develop a new strategy using other control elements or generate a new control element using the actual state of the target variable as positive or negative reinforcement for a regression or neural network model.
[0153] In some embodiments, if context is found to affect the expected behavior of a control element for a particular listener, it may mean that the data points (e.g., audio parameters) that are changed by the control element in a particular context are relevant to that context for that listener. In this way, these data points provide good starting points for creating music that creates an environmental change. For example, if a listener always manually increases a "rhythmic" control element when going to a train station, music generation module 160 can automatically begin increasing this control element when it detects that the listener is at a train station.
[0154] In some embodiments, as described herein, music generation module 160 is trained to implement control elements that match user expectations. If music generation module 160 were trained end-to-end for each control element (e.g., from the control element level to the output musical content 140), the training complexity for each control element would be high, potentially slowing down the training. Furthermore, it would be difficult to establish the ideal combined effect of multiple control elements. However, for each control element, music generation module 160 should ideally be trained to perform the expected musical changes based on the control element. For example, music generation module 160 may be trained by a listener for the “energy” control element to increase rhythmic density as “energy” increases. Because listeners are exposed to the final output musical content 140, as well as individual layers of musical content, music generation module 160 may be trained to influence the final output musical content with control elements. However, this can be a multi-level problem, such as: with certain control settings, music should sound like X; to create music that sounds like X, set of audio files Y should be used on each track.
[0155] In certain embodiments, a teacher / student framework is employed to address the above-mentioned problems. Figure 10 is a block diagram illustrating an exemplary teacher / student framework system, according to some embodiments. In the illustrated embodiment, the system 1000 includes a teacher model implementation module 1010 and a student model implementation module 1020.
[0156] In particular embodiments, the teacher model implementation module 1010 implements a trained teacher model. For example, the trained teacher model may be a model that learns how to predict how a final mix (e.g., a stereo mix) should sound, without any consideration of the set of loops available in the final mix. In some embodiments, the teacher model's learning process utilizes real-time analysis of the output music content 140 using a fast Fourier transform to calculate the distribution of notes across different frequencies for a sequence of short-time steps. The teacher model can search for patterns in these sequences using a time sequence prediction model, such as a recurrent neural network (RNN). In some embodiments, the teacher model in the teacher model implementation module 1010 may be trained offline on stereo recordings for which individual loops or audio files are not available.
[0157] In the illustrated embodiment, the teacher model implementation module 1010 receives the output musical content 140 and generates a compact description 1012 of the output musical content. Using the trained teacher model, the teacher model implementation module 1010 can generate the compact description 1012 without considering any audio tracks or audio files in the output musical content 140. The compact description 1012 may include a description X of what the output musical content 140 should sound like as determined by the teacher model implementation module 1010. The compact description 1012 is more compact than the output musical content 140 itself.
[0158] The compact description 1012 may be provided to a student model implementation module 1020. The student model implementation module 1020 implements a trained student model. For example, the trained student model may be a model that learns how to produce music that matches the compact description using audio file or loop Y (different from X). In the illustrated embodiment, the student model implementation module 1020 generates student output musical content 1014 that substantially matches the output musical content 140. As used herein, the phrase "substantially matches" indicates that the student output musical content 1014 sounds similar to the output musical content 140. For example, a trained listener may perceive the student output musical content 1014 and the student output musical content 140 as sounding the same.
[0159] Control elements are often expected to affect similar patterns in music. For example, a control element may affect both pitch relationships and rhythm. In some embodiments, music generation module 160 is trained for multiple control elements according to one teacher model. By training music generation module 160 for multiple control elements using a single teacher model, it may be unnecessary to relearn similar basic patterns for each control element. In such embodiments, a student model of the teacher model learns how to vary the loop selection for each track to achieve desired attributes in the final music mix. In some embodiments, loop characteristics may be pre-calculated to reduce the learning challenge and baseline performance (albeit at the expense of potentially reducing the likelihood of finding an optimal mapping of control elements).
[0160] Non-limiting examples of pre-computed musical attributes for each loop or audio file that may be used in student model training include: bass to treble frequency ratio, note onsets per second, ratio of detected to undetected notes, spectral range, and average onset intensity. In some implementations, the student model is a simple regression model trained to select the loop for each track to obtain the closest musical properties in the final stereo mix. In various embodiments, the student / teacher model framework can have several advantages. For example, if a new property is added to the loop pre-computation routine, the entire end-to-end model, i.e., the student model, does not need to be retrained.
[0161] As another example, because the characteristics of the final stereo mix that affect different control elements are likely to be common to other control elements, training music generation module 160 for each control element as an end-to-end model would mean that each model would have to learn the same thing to arrive at the best loop selection, making training slower and more difficult than necessary. Because only the stereo output needs to be analyzed in real time and the output musical content is generated in real time for the listener, music generation module 160 can obtain a "free" signal through computation. Even FFTs may already be applied for visualization or audio mixing purposes. In this way, a teacher model can be trained to predict the combined behavior of control elements, and music generation module 160 can be trained to find ways to adapt to other control elements while still producing the desired output musical content. This may encourage control element training to emphasize the unique effects of certain control elements and reduce control elements that have the effect of reducing the influence of other control elements.
[0162] Exemplary Low-Resolution Pitch Detection System Pitch detection that is robust to polyphonic musical content and diverse instrument types has traditionally been difficult to achieve. Tools that implement end-to-end music transfer may take an audio recording and attempt to generate a written score in the form of MIDI or a symbolic music representation. Without knowledge of beat placement or tempo, these tools must infer the musical rhythmic structure, instrumentation, and pitch. Results can vary, and common problems include detecting too many short notes that are not present in the audio file and detecting harmonics of notes as the fundamental pitch.
[0163] However, pitch detection is also useful in situations where end-to-end transcription is not required. For example, to create a harmonically plausible combination of musical loops, it may be sufficient to know which pitches are heard on each beat, without needing to know the exact position of the notes. If the length and tempo of the loop are known, the temporal location of the beats does not need to be inferred from the audio.
[0164] In some embodiments, the pitch detection system is configured to detect which fundamental pitches (e.g., C, C#...B) are present in short musical audio files of known beat length. By reducing the problem area and focusing on robustness to instrument texture, highly accurate results can be achieved for beat-resolution pitch detection.
[0165] In some embodiments, the pitch detection system is trained on examples for which ground truth is known. In some implementations, audio data is created from score data. MIDI and other symbolic music formats can be synthesized using a software audio synthesizer with random parameters for texture and effects. For each audio file, the system can generate a log-spectrogram 2D representation with multiple frequency bins for each pitch class. This 2D representation is used as input to a neural network or other AI technique, and multiple convolutional layers are used to create a feature representation of the frequency and time representation of the audio. The convolution stride and padding can be varied depending on the length of the audio file to generate a consistent model output shape with different tempo inputs. In some embodiments, the pitch detection system adds a recurrent layer to the convolutional layer to output a temporally dependent sequence of predictions. A category-wide entropy loss can be used to compare the neural network's logical output with a binary representation of the score.
[0166] The design of convolutional layers combined with recurrent layers is similar to speech-to-text tasks, with modifications. For example, speech-to-text typically needs to be sensitive to relative pitch changes, but not absolute pitch changes. Therefore, frequency range and resolution are typically small. Furthermore, the text may need to be invariant to velocity in a way that is undesirable for static-tempo music. For example, connectionist temporal classification (CTC) loss computation, often utilized in speech-to-text tasks, may not be required because the length of the output sequence is known a priori, reducing training complexity.
[0167] The following representation has 12 pitch classes for each beat, with 1 representing the presence of that fundamental note in the score used to synthesize the audio. (C,C#...B) Each line represents a beat, and subsequent lines represent scores on different beats:
number
[0168] In some embodiments, the neural network is trained on pseudo-randomly generated music scores of classical music and one- to four-part (or more) harmonies and polyphonies. Data augmentation can help improve robustness to effects on musical content, such as filters and reverberation, which can be a challenge for pitch detection (e.g., because part of the underlying tone remains after the original note ends). In some embodiments, the dataset can be biased, and loss weights are used, since pitch classes are much more likely not to play a note on each beat.
[0169] In some embodiments, the format of the output allows for avoiding harmonic clashes on each beat while maximizing the range of harmonic contexts in which the loop can be used. For example, a bass loop can contain only an F, descending to an E on the final beat of the loop. This loop would sound harmonically acceptable to most people in the key of F. If temporal resolution were not provided and only an E and an F were known to be included in the audio, it might be a sustained E with a short F at the end. This would not be acceptable to most people in the context of the key of F. At higher resolutions, as individual notes increase, harmonics, fretboard notes, and slides may be detected, potentially resulting in the false identification of additional notes. According to some embodiments, the complexity of the pitch detection problem can be reduced and robustness to short, insignificant pitch events increased by developing a system with optimal resolution of temporal and pitch information for combining short audio recordings of musical instruments to create harmonically sound musical mixes.
[0170] In various embodiments of the music generation system described herein, the system allows listeners to select audio content to be used to create a pool from which the system builds (generates) new music. This approach may differ from creating a playlist because the user does not need to select individual tracks or organize the selections sequentially. Furthermore, content from multiple artists can be used simultaneously. In some implementations, musical content is grouped into "packs" designed by the software provider or contributing artists. A pack contains multiple audio files with corresponding image features and feature metadata files. A single pack may contain, for example, 20 to 100 audio files that the music generation system can use to create music. In some embodiments, a single pack or a combination of multiple packs can be selected. Packs can be added or removed during playback without pausing the music.
[0171] Exemplary Audio Techniques for Musical Content Creation In various embodiments, software frameworks for managing real-time generated audio can benefit from supporting certain types of functionality. For example, audio processing software can follow the modular signal chain metaphor inherited from analog hardware, where different modules providing audio generation and audio effects are chained together in an audio signal graph. Individual modules typically expose various continuous parameters that allow real-time modification of the module's signal processing. In the early days of electronic music, parameters were often analog signals themselves, and thus the parameter processing chain and the signal processing chain coincided. Since the digital revolution, parameters have tended to be separate digital signals.
[0172] The embodiments disclosed herein recognize that a flexible control system that allows for the adjustment and combination of parameter manipulations can be advantageous for real-time music generation systems (whether the system interacts with human performers or whether the system implements machine learning or other artificial intelligence techniques to generate music). Furthermore, the present disclosure recognizes that it can also be advantageous for the effects of parameter changes to be invariant to changes in tempo.
[0173] In some implementations, the music generation system generates new musical content from the reproduced musical content based on different parameter representations of the audio signal. For example, the audio signal can be represented by both a graph of the signal versus time (e.g., an audio signal graph) and a graph of the signal versus beat (e.g., a signal graph). The signal graph is tempo-invariant, allowing for tempo-invariant modification of the audio parameters of the musical content as well as tempo-invariant modification based on the audio signal graph.
[0174] 11 is a block diagram illustrating an exemplary system configured to implement audio techniques in musical content generation, according to some embodiments. In the illustrated embodiment, system 1100 includes a graph generation module 1110 and an audio-technology music generation module 1120. Audio-technology music generation module 1120 may operate as a music generation module (e.g., audio-technology music generation module is music generation module 160 described herein), or the audio-technology music generation module may be implemented as part of a music generation module (e.g., as part of music generation module 160).
[0175] In the illustrated embodiment, music content 1112, including music file data, is accessed by a graph generation module 1110. The graph generation module 1110 can generate a first graph 1114 and a second graph 1116 for an audio signal within the accessed music content 1112. In particular embodiments, the first graph 1114 is an audio signal graph that graphs the audio signal as a function of time. The audio signal may include, for example, amplitude, frequency, or a combination of both. In particular embodiments, the second graph 1116 is a signal graph that graphs the audio signal as a function of beats.
[0176] 11 , graph generation module 1110 is located within system 1100 and generates first graph 1114 and second graph 1116. In such an embodiment, graph generation module 1110 may be co-located with audio-technical music generation module 1120. However, other embodiments are contemplated in which graph generation module 1110 is located on a separate system and audio-technical music generation module 1120 accesses the graphs from the separate system. For example, the graphs may be generated and stored on a cloud-based server accessible by audio-technical music generation module 1120.
[0177] FIG. 12 shows an example of an audio signal graph (e.g., first graph 1114). FIG. 13 shows an example of a signal graph (e.g., second graph 1116). In the graphs of FIGS. 12 and 13, each change in the audio signal is represented as a node (e.g., audio signal node 1202 in FIG. 12 and signal node 1302 in FIG. 13). Thus, the parameters of a particular node determine (e.g., define) the change in the audio signal at the particular node. Because first graph 1114 and second graph 1116 are based on the same audio signal, the graphs may have a similar structure; the difference between the graphs is the X-axis scale (time vs. beats). Having a similar structure in the graphs allows changes in parameters (described below) of a node in one graph (e.g., node 1202 in first graph 1114) that corresponds to a node in another graph (e.g., node 1302 in second graph 1116) to be determined by parameters either downstream or upstream of the node in one graph.
[0178] 11 , the first graph 1114 and the second graph 1116 are received (or accessed) by an audio-technical music generation module 1120. In particular embodiments, the audio-technical music generation module 1120 generates new musical content 1122 from the played musical content 1118 based on the audio modifier parameters selected from the first graph 1114 and the audio modifier parameters selected from the second graph 1116. For example, the audio-technical music generation module 1120 can modify the played musical content 1118 with either the audio modifier parameters from the first graph 1114, the audio modifier parameters from the second graph 1116, or a combination thereof. The new musical content 1122 is generated by modifying the played musical content 1118 based on the audio modifier parameters.
[0179] In various embodiments, the audio technical music generation module 1120 can select audio modifier parameters based on whether tempo variant modification, tempo invariant modification, or a combination thereof is desired in modifying the playback content 1118. For example, tempo variant modification may be made based on audio modifier parameters selected or determined from the first graph 1114, and tempo invariant modification may be made based on audio modifier parameters selected or determined from the second graph 1116. In embodiments where a combination of tempo variant modification and tempo invariant modification is desired, audio modifier parameters may be selected from both the first graph 1114 and the second graph 1116. In some embodiments, the audio modifier parameters from each individual graph are applied separately to different properties (e.g., amplitude or frequency) or different layers (e.g., different instrumental layers) within the playback musical content 1118. In some embodiments, the audio modifier parameters from each graph are combined into a single audio modifier parameter to apply to a single property or layer within the musical content 1118.
[0180] 14 illustrates an exemplary system for implementing real-time modification of musical content using a music technology music generation module 1420, according to some embodiments. In the illustrated embodiment, the music technology music generation module 1420 includes a first node determination module 1410, a second node determination module 1420, an audio parameter determination module 1430, and an audio parameter modification module 1440. Collectively, the first node determination module 1410, the second node determination module 1420, the audio parameter determination module 1430, and the audio parameter modification module 1440 implement the system 1400.
[0181] In the illustrated embodiment, the audio-technical music generation module 1420 receives playback music content 1418 including an audio signal. The audio-technical music generation module 1420 can process the audio signal through a first graph 1414 (e.g., a time-based audio signal graph) and a second graph 1416 (e.g., a beat-based signal graph) in the first node determination module 1410. As the audio signal passes through the first graph 1414, parameters at each node in the graph determine changes in the audio signal. In the illustrated embodiment, the second node determination module 1420 can receive information on the first node 1412 and determine information for the second node 1422. In particular embodiments, the second node determination module 1420 reads parameters in the second graph 1416 based on the location of the first node found in the first node information 1412 within the audio signal passing through the first graph 1414. Thus, as an example, an audio signal going to node 1202 in the first graph 1414 (shown in FIG. 12) can trigger a second node determination module 1420 that determines a corresponding (parallel) node 1302 in the second graph 1416 (shown in FIG. 13).
[0182] 14 , the audio parameter determination module 1430 may receive the second node information 1422 and determine (e.g., select) a particular audio parameter 1432 based on the second node information. For example, the audio parameter determination module 1430 may select an audio parameter based on a portion of the next beat (e.g., x number of next beats) of the second graph 1416 following the position of the second node as identified in the second node information 1422. In some embodiments, a beat-to-real-time conversion may be implemented to determine the portion of the second graph 1416 from which the audio parameter is read. The specified audio parameter 1432 may be provided to the audio parameter modification module 1440.
[0183] The audio parameter modification module 1440 can control the modification of musical content to generate new musical content. For example, the audio parameter modification module 1440 can modify the played musical content 1418 to generate new musical content 1122. In particular embodiments, the audio parameter modification module 1440 modifies properties of the played musical content 1418 by modifying specific audio parameters 1432 (determined by the audio parameter determination module 1430) for audio signals in the played musical content. For example, modifying the audio parameters 1432 specified for audio signals in the played musical content 1418 changes properties such as amplitude, frequency, or a combination of both within the audio signals. In various embodiments, the audio parameter modification module 1440 modifies properties of different audio signals in the played musical content 1418. For example, different audio signals in the played musical content 1418 may correspond to different instruments represented in the played musical content 1418.
[0184] In some embodiments, the audio parameter modification module 1440 uses machine learning algorithms or other AI techniques to modify the characteristics of the audio signals in the played musical content 1418. In some embodiments, the audio parameter modification module 1440 modifies the properties of the played musical content 1418 according to user input to the module, which may be provided via a user interface associated with the music generation system. Other embodiments are also contemplated in which the audio parameter modification module 1440 modifies the properties of the played musical content 1418 using a combination of AI techniques and user input. Various embodiments for modifying the characteristics of the played musical content 1418 by the audio parameter modification module 1440 enable real-time manipulation of the musical content (e.g., manipulation during playback). As described above, real-time manipulation can include applying tempo-change modifications, tempo-invariant combinations, or a combination of both to the audio signals of the played musical content 1418.
[0185] In some embodiments, the audio technical music generation module 1420 implements a two-tiered parameter system for modifying characteristics of the played musical content 1418 via the audio parameter modification module 1440. In a two-tiered parameter system, a distinction may be made between "automation," which directly controls audio parameter values (e.g., tasks automatically performed by the music generation system), and "modulation," which multiplicatively overlays audio parameter modifications on top of the automation, as described below. A two-tiered parameter system allows different parts of the music generation system (e.g., different machine learning models in the system architecture) to consider different musical aspects separately. For example, one part of the music generation system may set the volume of a particular instrument according to the intended section type of the composition, while another part may overlay periodic variations in volume for additional interest.
[0186] Exemplary Techniques for Real-Time Audio Effects in Musical Content Creation Music technology software typically allows the composer / producer to control various abstract envelopes through automation. In some embodiments, the automation is a pre-programmed temporal manipulation of some audio processing parameter (such as volume or reverberation amount). The automation is typically either a manually defined breakpoint envelope (e.g., a piecewise linear function) or a programmatic function such as a sine wave (also known as a low frequency oscillator (LFO)).
[0187] The disclosed music generation system may differ from typical music software. For example, most parameters are, in some sense, automated by default. The AI technology of the music generation system may control most or all audio parameters in various ways. At a base level, a neural network can predict appropriate settings for each audio parameter based on its training. However, it may be useful to provide the music generation system with higher-level automation rules. For example, a large-scale musical structure may dictate a slow build of volume as an extra consideration in addition to the lower-level settings that might otherwise be predicted.
[0188] This disclosure generally relates to an information architecture and procedural approach for combining multiple parametric instructions issued simultaneously by different levels of a hierarchical generative system to create a musically coherent, evolving, continuous output. The disclosed music generation system is capable of generating long-form musical experiences intended to be experienced continuously over several hours. Long-form musical experiences require the creation of a coherent musical journey to be a more satisfying experience. To do this, the music generation system can reference itself on a long time scale. These references range from direct to abstract.
[0189] In certain embodiments, to facilitate larger-scale musical rules, the music generation system (e.g., music generation module 160) exposes an automation API (application programming interface). FIG. 15 shows a block diagram of exemplary API modules in a system for audio parameter automation, according to some embodiments. In the illustrated embodiment, system 1500 includes API module 1505. In some embodiments, API module 1505 includes automation module 1510. The music generation system can support both wavetable-style LFOs and arbitrary breakpoint envelopes. Automation module 1510 can apply automation 1512 to any audio parameter 1520. In some embodiments, automation 1512 is applied recursively. For example, any programmatic automation, such as a sine wave, itself has parameters (frequency, amplitude, etc.), and automation can be applied to those parameters.
[0190] In various embodiments, automation 1512 includes a signal graph parallel to the audio signal graph, as described above. The signal graph can be handled similarly via a "pull" technique, in which API module 1505 requests recalculation from automation module 1510 as needed, and recalculation can occur such that automation 1512 recursively requests upstream automations on which it depends. In particular embodiments, the signal graph for the automation is updated at a controlled rate. For example, the signal graph can be updated after each execution of the execution engine update routine, which can be aligned with the audio block rate (e.g., after the audio signal graph renders one block (e.g., 512 samples)).
[0191] In some embodiments, it is desirable for the audio parameters 1520 themselves to change at the audio sample rate; otherwise, discontinuous parameter changes at audio block boundaries could result in audible artifacts. In particular embodiments, the music generation system manages this problem by treating automated updates as parameter value targets. As the real-time audio thread renders an audio block, it smoothly ramps the specified parameters from their current values to the supplied target values over the course of that block.
[0192] The music generation systems described herein (e.g., music generation module 160 shown in FIG. 1) may have an architecture that is hierarchical in nature. In some embodiments, different parts of the hierarchy may provide multiple suggestions for values of particular audio parameters. In certain embodiments, the music generation system provides two separate mechanisms for combining / resolving the multiple suggestions: modulation and overriding. In the illustrated embodiment of FIG. 15, modulation 1532 is achieved by modulation module 1530, and overriding 1542 is achieved by overriding module 1540.
[0193] In some embodiments, automation 1512 may be declared to be a modulation 1532. Such a declaration may mean that the automation 1512 should act multiplicatively on the current value of the audio parameter, rather than directly setting the value of the audio parameter. Thus, a large musical section may apply a long modulation 1532 to the audio parameter, and the value of the modulation will be multiplied regardless of the value dictated by other parts of the music generation system.
[0194] In various embodiments, the API module 1505 includes an override module 1540. The override module 1540 may be, for example, an override function for audio parameter automation. The override module 1540 may be intended for use by an external control interface (e.g., an artist-controlled user interface). The override module 1540 may control the audio parameters 1520 regardless of what the music generation system attempts to make the musical parameters 1520. If there are override parameters 1520 that are overwritten by an override 1542, the music generation system may generate “shadow parameters” 1522 that track the audio parameters that would otherwise be overwritten (e.g., if the override parameters are based on an override 1512 or modulation 1532). Thus, when the override 1542 is “released” (e.g., removed by the artist), the audio parameters 1520 can snap back to where they would have been according to the automation 1512 or modulation 1532.
[0195] In various embodiments, these two approaches can be combined. For example, the override 1542 can be a modulation 1532. When the override 1542 is a modulation 1532, the base value of the override parameter 1520 is set by the music generation system, but may be multiplicatively modulated by the override 1542. Each audio parameter 1520 may simultaneously have one (or zero) automation 1512 and one (or zero) modulation 1532, as well as one (or zero) of each override 1542.
[0196] In various embodiments, the abstract class hierarchy is defined as follows (note some multiple inheritance):
number
[0197] Based on a hierarchy of abstract classes, things are considered to be either Automations or Automatable. In some embodiments, any automation can be applied to anything that is automatable. Automations include LFOs, breakpoint envelopes, etc. All of these automations are tempo-locked, meaning they change over time depending on the current beat.
[0198] Automations can themselves have automatable parameters. For example, the frequency and amplitude of an LFO automation can be automatable. Thus, there are signal graphs of dependent automations and automation parameters that run in parallel to the audio signal graph, at a control rate rather than an audio rate. As mentioned above, the signal graphs use a pull model. The music generation system keeps track of any automations 1512 applied to audio parameters 1520 and updates these every "game loop." The automations 1512 then recursively request updates to their own automated audio parameters 1520. This recursive update logic may reside in a base class, Beat-Dependent, which expects to be called frequently (but not necessarily regularly). The update logic may have a prototype written as follows:
number
[0199] In certain embodiments, the BeatDependent class maintains a list of its own dependencies (e.g., other BeatDependent instances) and recursively calls their update functions. The updateCounter can be passed down the chain so that the signal graph can have cycles without double updates. This can be important because automation may apply to several different automation targets. In some embodiments, this is not an issue because the second update will have the same currentBeat as the first, and these update routines will have no effect unless the beat changes.
[0200] In various embodiments, if automation is applied to an automation, then in each cycle of the "game loop," the music generation system requests (recursively) an updated value from each automation and uses it to set the value of the automation, where "setting" may depend on the particular subclass and may also depend on whether the parameter is modulated and / or disabled.
[0201] In certain embodiments, modulation 1532 is automation 1512 that is applied multiplicatively rather than absolutely. For example, modulation 1532 can be applied to an already automated audio parameter 1520, and the effect is a percentage of the automated value. This can allow for ongoing vibrations around a vehicle, for example, multiplicatively.
[0202] In some embodiments, audio parameters 1520 are overridable, meaning that any automation 1512 or modulation 1532 applied to them, or other (less privileged) requirements, as described above, are overridden by the override value in override 1542. This override allows external control over some aspects of the music generation system, while the music generation system continues as if it were not. When audio parameters 1520 are overridden, the music generation system keeps track of what the value is (e.g., keeps track of automation / modulation and other requirements that have been applied). When the override is removed, the music generation system snaps the parameter and checks to see where it was.
[0203] To facilitate modulation 1532 and override 1542, the music generation system may abstract the setValue method of Parameter. There is also a private method _setValue that actually sets the value. An example of a public method is:
number
[0204] Public methods can reference a member variable of the Parameter class called _unmodified. This variable is an instance of the ShadowParameter mentioned above. Each audio parameter 1520 has a shadow parameter 1522 that tracks where it would be if unmodulated. If an audio parameter 1520 is currently unmodulated, both the audio parameter 1520 and its shadow parameter 1522 are updated with the requested value. Otherwise, the shadow parameter 1522 tracks the request, and the actual audio parameter value 1520 is set elsewhere (e.g., in the updateModulations routine - where the modulation coefficient is multiplied with the shadow parameter value to give the actual parameter value).
[0205] In various embodiments, large-scale structure in a long-form musical experience is achieved through a variety of mechanisms. One broad approach is the long-term use of musical self-reference. For example, a very direct self-reference is the exact repeat of a previously played audio segment. In music theory, a repeated segment is called a theme (or motif). More typically, musical content uses theme and variation, with the theme later repeated with some variation, to give a sense of coherence but maintain a sense of progression. The music generation system disclosed herein can use theme and variation to create large-scale structure in several ways, including direct repetition or the use of abstract envelopes.
[0206] An abstract envelope is the value of an audio parameter over time. Abstract envelopes can be applied to other audio parameters in abstraction from the audio parameters they control. For example, a collection of audio parameters can be coordinated and automated by a single controlling abstract envelope. This technique allows different layers to be perceptually "coupled" together for short periods of time. Abstract envelopes may also be temporarily reused and applied to different audio parameters. In this way, abstract envelopes become abstract musical themes, and this theme is repeated by applying this envelope to different audio parameters in subsequent listening experiences. In this way, a sense of structure and long-term consistency is established, while variations on this theme exist.
[0207] Abstract envelopes can be thought of as musical themes and can abstract many musical features. Examples of musical features that can be abstracted include, but are not limited to: · Build in tension (track volume, distortion level, etc.). Rhythm (volume adjustment and / or gate settings can give rhythmic effects to pads etc.). Melodic (pitch filtering can mimic the contours of a melody applied to pads etc.).
[0208] Exemplary Additional Audio Techniques for Real-Time Musical Content Creation Real-time musical content generation can present unique challenges. For example, due to strict real-time constraints, function calls and subroutines with unpredictable and unbounded execution times should be avoided. Avoiding this problem may preclude the use of most high-level programming languages, as well as the majority of low-level languages like C and C++. Anything that allocates memory from the heap (e.g., malloc under the hood) can be ruled out, as can anything that might block, such as by locking a mutex. This makes multithreaded programming particularly difficult in real-time musical content generation. Most standard memory management approaches may also be infeasible, and as a result, dynamic data structures like C++ STL containers are of limited use for real-time musical content generation.
[0209] Another area of challenge is the management of audio parameters involved in DSP (digital signal processing) functions (such as filter cutoff frequencies). For example, dynamically changing audio parameters can result in audible artifacts unless the audio parameters are continuously changed. Therefore, communication between the real-time DSP audio thread and a user-facing or programmatic interface may be required to change audio parameters.
[0210] To address these constraints, various audio software implementations can be used and various approaches exist. For example: Inter-thread communication may be handled by lock-free message queues. Functions are written in plain C and use function pointer callbacks. Memory management may be implemented via custom "zones" or "regions". A "two-speed" system may be implemented where real-time audio thread calculations run at the audio rate, and the audio thread is controlled to run at a "control rate." The controlling audio thread can set audio parameter change targets that the real-time audio thread smoothly ramps to.
[0211] In some embodiments, synchronization between control rate audio parameter manipulation and real-time audio thread-safe storage of audio parameter values for use in actual DSP routines may require some type of thread-safe communication of audio parameter targets. Most audio parameters in audio routines are continuous (rather than discrete), and therefore typically represented as floating-point data types. Various distortions to the data have historically been necessitated by the lack of lock-free atomic floating-point data types.
[0212] In certain embodiments, a simple lock-free atomic floating-point data type is implemented in the music generation system described herein. A lock-free atomic floating-point data type can be achieved by treating a floating-point type as a bit string and "tricking" the compiler to treat it as an atomic integer type of the same bit width. This approach can support atomic getting / setting suitable for the music generation system described herein. An example implementation of a lock-free atomic floating-point data type is shown below:
number
[0213] In some embodiments, dynamic memory allocation from the heap is not feasible for real-time code associated with musical content generation. For example, static stack-based allocation can make it difficult to use programming techniques such as dynamic storage containers and functional programming approaches. In particular embodiments, the music generation system described herein implements “memory zones” for memory management in a real-time context. As used herein, a “memory zone” is a region of heap-allocated memory that is proactively allocated without real-time constraints (e.g., when real-time constraints do not yet exist or are suspended). Memory storage objects are created within a region of heap-allocated memory without having to request more memory from the system, thereby making the memory real-time safe. Garbage collection can include deallocation of the entire memory zone. Additionally, the memory implementation by the music generation system may be multi-threaded, safe, real-time safe, and efficient.
[0214] 16 shows a block diagram of an exemplary memory zone 1600, according to some embodiments. In the illustrated embodiment, the memory zone 1600 includes a heap-allocated memory module 1610. In various embodiments, the heap-allocated memory module 1610 receives and stores a first graph 1114 (e.g., an audio signal graph), a second graph 1116 (e.g., a signal graph), and audio signal data 1602. Each of the stored items can be retrieved, for example, by the audio parameter modification module 1440 (shown in FIG. 14).
[0215] An example implementation of a memory zone is shown below:
number
[0216] In some embodiments, different audio threads of a music generation system need to communicate with each other. Typical thread-safety approaches (which may involve locking "mutually exclusive" data structures) may not be usable in a real-time context. In particular embodiments, dynamically routing data serialization to a pool of single-producer, single-consumer circular buffers is implemented. A circular buffer is a type of FIFO (first-in, first-out) queue data structure that requires no dynamic memory allocation after initialization. A single-producer, single-consumer thread-safe circular buffer allows one audio thread to push data into a queue while another audio thread pulls data from it. For the music generation systems described herein, circular buffers can be extended to allow for multiple-producer, single-consumer audio threads. These buffers can be realized by pre-allocating a static array of circular buffers and dynamically routing serialized data to specific "channels" (e.g., specific circular buffers) according to identifiers added to the musical content generated by the music generation system. A static array of circular buffers is accessible by a single user (eg, a single consumer).
[0217] 17 shows a block diagram of an exemplary system for storing new music content, according to some embodiments. In the illustrated embodiment, system 1700 includes a circular buffer static array module 1710. Circular buffer static array module 1710 includes multiple circular buffers, allowing storage of multiple producer, single-consumer audio threads according to thread identifiers. For example, circular buffer static array module 1710 can receive new music content 1122 and store the new music content in 1712 for access by a user.
[0218] In various embodiments, abstract data structures such as dynamic containers (vectors, queues, lists) are typically implemented in a non-real-time safe manner. However, these abstract data structures are useful for audio programming. In particular embodiments, the music generation system described herein implements a custom list data structure (e.g., a singly linked list). Many functional programming techniques can be implemented from the custom list data structure. The custom list data structure implementation may use "memory zones" (described above) for underlying memory management. In some embodiments, the custom list data structure is serializable, which can be made safe for real-time use and can be communicated between audio threads using the multi-producer, single-consumer audio thread described above.
[0219] Exemplary Blockchain Ledger Technologies The disclosed system, in some embodiments, can utilize secure recording technologies, such as blockchain or other cryptographic ledger technology, to record information about the generated music or its elements, e.g., loops or tracks. In some implementations, the system combines multiple audio files (e.g., tracks or loops) to generate output musical content. This combination can be performed by combining multiple layers of audio content so that they at least partially overlap in time. The output content can be individual pieces of music or can be continuous. Tracking the use of musical elements can be challenging in the context of continuous music, for example, to provide royalties to relevant stakeholders. Therefore, in some embodiments, the disclosed system records identifiers and usage information (e.g., timestamps or play counts) of audio files used in the composite musical content. Additionally, the disclosed system can utilize various algorithms to track play time, for example, in the context of blended audio files.
[0220] As used herein, the term "blockchain" refers to a cryptographically linked series of records (called blocks). For example, each block may include a cryptographic hash of the previous block, a timestamp, and transaction data. A blockchain can be used as a public, distributed ledger and managed by a network of computing devices that use agreed-upon protocols for communication and validation of new blocks. Some blockchain implementations are immutable, while others allow for subsequent modification of blocks. In general, a blockchain can record transactions in a verifiable and permanent manner. While blockchain ledgers are discussed herein for illustrative purposes, it should be understood that the disclosed technology can be used with other types of cryptographic ledgers in other embodiments.
[0221] FIG. 18 illustrates an example of playback data according to some embodiments. In the illustrated embodiment, the database structure includes entries for multiple files. Each illustrated entry includes a file identifier, a start timestamp, and a total time. The file identifier may uniquely identify an audio file tracked by the system. The start timestamp may indicate the first inclusion of the audio file in mixed audio content. This timestamp may be based on the playback device's local clock or, for example, an Internet clock. The total time may indicate the length of the interval in which the audio file was incorporated. Note that this may differ from the length of the audio file. For example, if only a portion of the audio file was used, if the audio file was accelerated or slowed down in the mix, etc. In some implementations, if an audio file is incorporated at multiple different times, each time results in an entry. In other embodiments, additional plays for a file may result in an increment of the time field of an existing entry if an entry for the file already exists. In still other embodiments, the data structure may track the number of times each audio file is used rather than the length of the ingestion. Additionally, other encodings of time-based usage data are contemplated.
[0222] In various embodiments, different devices may determine, store, and use ledgers to record replay data. Exemplary scenarios and topologies are described below with reference to FIG. 19. Replay data may be temporarily stored on a computing device before being committed to the ledger. Stored replay data may be encrypted, for example, to reduce or prevent manipulation of entries or insertion of false entries.
[0223] 19 is a block diagram illustrating an exemplary music composition system according to some embodiments. In the illustrated example, the system includes a playback device 1910, a computing system 1920, and a ledger 1930.
[0224] In the illustrated embodiment, the playback device 1910 receives control signals from and transmits playback data to the computing system 1920. In this embodiment, the playback device 1910 includes a playback data recording module 1912 that can record playback data based on the audio mix played by the playback device 1910. The playback device 1910 also includes a playback data storage module 1914 that is configured to store the playback data temporarily, in a ledger, or both. The playback device 1910 may report the playback data to the computing system 1920 periodically or may report the playback data in real time. The playback data may be stored for later reporting, for example, when the playback device 1910 is offline.
[0225] In the illustrated embodiment, computing system 1920 receives playback data and commits entries reflecting the playback data to ledger 1930. Computing system 1920 also sends control signals to playback device 1910. This control signaling may include various types of information in different embodiments. For example, control signaling may include configuration data, mixing parameters, audio samples, machine learning updates, etc. for use by playback device 1910 to compose musical content. In other embodiments, computing system 1920 may compose musical content and stream musical content data to playback device 1910 via control signals. In these embodiments, modules 1912 and 1914 may be included in computing system 1920. In general, the modules and functionality described with reference to FIG. 19 may be distributed among multiple devices according to various topologies.
[0226] In some embodiments, playback device 1910 is configured to commit entries directly to ledger 1930. For example, a playback device such as a mobile phone may create music content, determine playback data, and store the playback data. In this scenario, the mobile device may report playback data to a server such as computing system 1920, or may report directly to a computing system (or collection of computing nodes) that maintains ledger 1930.
[0227] In some embodiments, the system keeps a record of rights holders, for example, with a mapping to an audio file identifier or set of audio files. This record of entities can be maintained in ledger 1930, or a separate ledger, or other data structure. This allows rights holders to remain anonymous, for example, if ledger 1930 is public but contains non-identifying entity identifiers mapped to entities in some other data structure.
[0228] In some embodiments, a music synthesis algorithm can generate a new audio file from two or more existing audio files for inclusion in a mix. For example, the system can generate a new audio file C based on two audio files A and B. One implementation method for such blending uses interpolation between vector representations of the audio representations of files A and B and generates file C using an inverse transformation from vectors to audio representations. In this example, the play times of audio files A and B may both be incremented, but may be incremented less than their actual play times, for example because they are being blended.
[0229] For example, if audio file C is incorporated into the mixed content for 20 seconds, audio file A will have playback data indicating 15 seconds, and audio file B will have playback data indicating 5 seconds (and note that the sum of the blended audio files may or may not match the resulting duration of file C). In some embodiments, the playback time of each original file is based on it as well as the mixed file C. For example, in a vector embodiment, for an n-dimensional vector representation, the interpolated vector a has the following distance d from the vector representations of audio files A and B:
number
[0230] In these embodiments, the play time of each original file, i, may be determined as follows:
number
[0231] In some embodiments, forms of compensation can be incorporated into the ledger structure. For example, a particular entity can include information associating an audio file with an execution requirement, such as displaying a link or including an advertisement. In these embodiments, the composition system can provide proof of performance of the associated operation (e.g., displaying an advertisement) when including the audio file in a mix. The proof of performance can be reported according to one of a variety of appropriate reporting templates, requiring specific fields to indicate how and when the action was performed. The proof of performance can include time information and utilize cryptography to avoid false claims of performance. In these embodiments, uses of an audio file that do not also show evidence of performance of the associated requested action may require some other form of compensation, such as royalty payments. In general, different entities submitting audio files can register different forms of compensation.
[0232] As described above, the disclosed technology can provide a reliable record of audio file usage in music mixes, even if they are composed in real time. The public nature of the ledger can provide confidence in the fairness of rewards. This may encourage participation from artists and other contributors and improve the variety and quality of audio files available for automated mixing.
[0233] In some embodiments, artist packs may be created with elements used by the music engine to create a continuous soundscape. An artist pack may be a professionally (or otherwise) curated set of elements stored in one or more data structures associated with an entity such as an artist or group. Examples of these elements include, but are not limited to, loops, composition rules, heuristics, and neural network vectors. Loops may be included in a database of musical phrases. Each loop is typically a single instrument or set of related instruments playing a musical progression over a period of time. These may range from short loops (e.g., 4 bars) to long loops (e.g., 32-64 bars). Loops may be organized into layers such as melody, harmony, drums, bass, top, and FX. The loop database may also be represented as a variational autoencoder with an encoded loop representation. In this case, the loops themselves are not required, but rather the NN is used to generate the audio encoded by the NN.
[0234] Heuristics refer to parameters, rules, and data that guide the music engine, such as factors like section length, use of effects, frequency of variation techniques, complexity of the music, or generally any type of parameter that can be used to enhance the music engine's decision-making as it composes and renders music.
[0235] The ledger records transactions related to the consumption of content with associated rights holders. This could be, for example, loops, heuristics, or neural network vectors. The purpose of the ledger is to record these transactions and allow for transparent accounting. The ledger is intended to capture transactions as they occur, including the consumption of content, the use of parameters in guiding a music engine, the use of vectors on a neural network, etc. The ledger can record various transaction types, such as individual events (e.g., this loop was played at this time), this pack was played for this amount of time, or this machine learning module (e.g., a neural network module) was used for this amount of time.
[0236] The ledger can associate multiple rights holders with any artist pack, and more specifically with specific loops or other elements of an artist pack. For example, a label, artist, and composer may have rights to a given artist pack. The ledger can allow for associating payment details for the pack that specify what percentage each party receives. For example, the artist may receive 25%, the label may receive 25%, and the composer may receive 50%. Using a blockchain to manage these transactions allows for small payments to be made to each rights holder in real time, or they can be accumulated over a suitable period of time.
[0237] As mentioned above, in some implementations, loops may be replaced with a VAE, which is essentially an encoding of the loop within a machine learning module. In this case, the ledger may associate playtime with a particular artist pack that includes the machine learning module. For example, if an artist pack is played for 10% of the total playtime across all devices, the artist may receive 10% of the total revenue share.
[0238] In some embodiments, the system allows artists to create an artist profile. The profile contains information related to the artist, including a bio, profile picture, bank details, and other data necessary to verify the artist's identity. Once an artist profile is created, artists can upload and publish artist packs. These packs contain elements that the music engine uses to create soundscapes.
[0239] For each artist pack created, a rights holder can be defined and associated with the pack. Each rights holder can claim a percentage of the pack. Additionally, each rights holder can create a profile and associate a bank account and payment profile. The artist themselves can be the rights holder and own 100% of the rights associated with that pack.
[0240] In addition to recording events in the ledger that are used for revenue recognition, the ledger can manage promotions associated with artist packs. For example, an artist pack may have a free month promotion that results in different revenue than if the promotion were not running. The ledger automatically calculates these revenue entries when calculating payments to rights holders.
[0241] This rights management model allows artists to sell the rights to a pack to one or more external rights holders. For example, upon the release of a new pack, an artist could pre-fund the pack by selling 50% of their share of the pack to fans / investors. In this case, the number of investors / rights holders can be arbitrarily large. For example, an artist could sell 50% to 100K users and receive 1 / 100K of that revenue. Since all accounting is maintained on the ledger, investors receive compensation directly in this scenario, eliminating the need for audits of artist accounts.
[0242] Exemplary User and Enterprise GUI 20A-20B are block diagrams illustrating graphical user interfaces, according to some embodiments. In the illustrated embodiment, FIG. 20A includes a GUI displayed by a user application 2010, and FIG. 20B includes a GUI displayed by an enterprise application 2030. In some implementations, the GUIs shown in FIGS. 20A and 20B are generated by a website rather than by an application. In various embodiments, any of a variety of suitable elements may be displayed, including one or more elements such as dials (e.g., to control volume, energy, etc.), buttons, knobs, display boxes (e.g., to provide updated information to the user), etc.
[0243] In FIG. 20A , user application 2010 displays a GUI that includes section 2012 for selecting one or more artist packs. In some embodiments, packs 2014 may alternatively or additionally include themed packs or packs for particular occasions (e.g., weddings, birthday parties, graduations, etc.). In some embodiments, the number of packs shown in section 2012 is greater than the number of packs that can be displayed in section 2012 at one time. Thus, in some embodiments, a user may scroll up or down in section 2012 to display one or more packs 2014. In some embodiments, a user may select an artist pack 2014 from which they would like to hear the output music content. In some embodiments, artist packs may be purchased and / or downloaded, for example.
[0244] The selection element 2016, in the illustrated embodiment, allows the user to adjust one or more music attributes (e.g., energy level). In some embodiments, the selection element 2016 allows the user to add / delete / change one or more target music attributes. In various embodiments, the selection element 2016 may render one or more UI control elements (e.g., control element 830).
[0245] In the illustrated embodiment, the selection element 2020 allows a user to have a device (e.g., a mobile device) listen to the environment to determine target musical attributes. In some embodiments, the device collects information about the environment using one or more sensors (e.g., a camera, microphone, thermometer, etc.) after the user selects the selection element 2020. In some embodiments, the application 2010 also selects or suggests one or more artist packs based on the environmental information collected by the application when the user selects the element 2020.
[0246] The selection element 2022, in the illustrated embodiment, allows the user to combine multiple artist packs to generate a new rule set. In some embodiments, the new rule set is based on the user selecting one or more packs for the same artist. In other embodiments, the new rule set is based on the user selecting one or more packs for different artists. The user can indicate weights for different rule sets, for example, so that a highly weighted rule set has a greater effect on the generated music than a low-weighted rule set. The music generator can combine rule sets in a number of different ways, for example, by switching between rules from different rule sets, averaging rules from multiple different rule sets, etc.
[0247] In the illustrated embodiment, selection element 2024 allows a user to manually adjust rules in one or more rule sets. For example, in some embodiments, a user may want to adjust the generated musical content at a more granular level by adjusting one or more rules in the rule set used to generate the musical content. In some embodiments, this allows a user of application 2010 to become their own disc jockey by using the controls displayed in the GUI of FIG. 20A to adjust the rule set that the music generator uses to generate the output musical content. These embodiments may also allow for more granular control of target musical attributes.
[0248] In FIG. 20B , enterprise application 2030 displays a GUI that includes artist pack selection section 2012 and artist packs 2014. In the illustrated embodiment, the enterprise GUI displayed by application 2030 also includes elements 2016 for adjusting / adding / removing one or more music attributes. In some implementations, the GUI shown in FIG. 20B is used in a business or storefront to create a particular environment (e.g., to optimize sales) by generating music content. In some embodiments, an employee uses application 2030 to select one or more artist packs that have previously been shown to increase sales (e.g., the metadata for a given rule set can indicate actual experimental results using the rule set in a real-world context).
[0249] In the illustrated embodiment, input hardware 2040 transmits information to an application or website displaying enterprise application 2030. In some embodiments, input hardware 2040 is one of a cash register, a heat sensor, a light sensor, a clock, a noise sensor, etc. In some embodiments, information transmitted from one or more of the above hardware devices is used to adjust target musical attributes and / or rule sets for generating output musical content for a particular environment. In the illustrated embodiment, selection element 2038 allows a user of application 2030 to select one or more hardware devices for receiving environmental input.
[0250] In the illustrated embodiment, display 2034 displays environmental data to a user of application 2030 based on information from input hardware 2040. In the illustrated embodiment, display 2032 shows changes to the rule set based on the environmental data. In some embodiments, display 2032 allows a user of application 2030 to view changes made based on the environmental data.
[0251] 20A and 20B are for theme packs and / or opportunity packs, i.e., in some embodiments, a user or business using the GUI displayed by applications 2010 and 2030 can select / adjust / modify rule sets to generate musical content for one or more occasions and / or themes.
[0252] Detailed example of a music generation system Figures 21-23 show details regarding specific embodiments of the music generation module 160. Note that these specific examples are disclosed for illustrative purposes and are not intended to limit the scope of the present disclosure. In these embodiments, the construction of music from loops is performed by a client system, such as a personal computer, a mobile device, or a media device. As used in the descriptions of Figures 21-23, the term "loop" can be interchanged with the term "audio file." Generally, loops are contained in audio files, as described herein. Loops can be divided into expertly curated loop packs, which can be referred to as artist packs. Loops can be analyzed for musical properties, and the properties can be stored as loop metadata. The audio in the constructed track can be analyzed (e.g., in real time), filtered, and the output stream can be mixed and mastered. Various feedback is sent to the server, including explicit feedback, for example, from the user's interaction with sliders or buttons, and implicit feedback generated by sensors, for example, based on volume changes, listening length, environmental information, etc. In some embodiments, the control inputs have known effects (eg, to directly or indirectly identify musical attributes of interest) and are used by the composition module.
[0253] The following discussion introduces various terms used with reference to Figures 21-23. In some embodiments, a loop library is a master library of loops that may be stored by a server. Each loop may include audio data and metadata describing the audio data. In some implementations, a loop package is a subset of a loop library. A loop package may be a pack for a particular artist, a particular mood, a particular type of event, etc. A client device may download a loop pack for offline listening or download portions of a loop pack on demand, for example, for online listening.
[0254] In some embodiments, the generated stream is data that specifies the musical content that a user hears when using the music generation system. Note that the actual output audio signal may vary slightly for a given generated stream, based on, for example, the capabilities of the audio output device.
[0255] The composition module, in some embodiments, builds a composition from loops available in a loop package. The composition module can receive loops, loop metadata, and user input as parameters and can be executed by a client device. In some embodiments, the composition module outputs an execution script that is sent to an execution module and one or more machine learning engines. In some embodiments, the execution script outlines which loops will be played on each track of the generated stream and what effects will be applied to the stream. The execution script can utilize beat-relative timing to represent when events occur. The execution script can also encode effect parameters (e.g., for effects such as reverb, delay, compression, equalization, etc.).
[0256] In some embodiments, the execution module receives an execution script as input and renders it into a generated stream. The execution module may generate multiple tracks specified by the execution script and mix the tracks into a stream (e.g., a stereo stream), although the stream may have various encodings, including voice encoding, object-based audio encoding, multi-channel stereo, etc. In some embodiments, when providing a particular execution script, the execution module always produces the same output.
[0257] The analytics module, in some embodiments, is a server-implemented module that receives the feedback information and configures the configuration module (e.g., in real-time, periodically, based on administrator command, etc.) In some embodiments, the analytics module uses a combination of machine learning techniques to correlate user feedback with execution scripts and loop library metadata.
[0258] Figure 21 is a block diagram illustrating an exemplary music generation system including an analysis module and a composition module, according to some embodiments. In some embodiments, the system of Figure 21 is configured to generate a potentially infinite stream of music with direct user control of musical mood and style. In the illustrated embodiment, the system includes an analysis module 2110, a composition module 2120, an execution module 2130, and an audio output device 2140. In some embodiments, the analysis module 2110 is implemented by a server, and the composition module 2120 and the execution module 2130 are implemented by one or more client devices. In other embodiments, modules 2110, 2120, and 2130 may all be implemented on the client device or all be implemented on the server side.
[0259] In the illustrated embodiment, the analysis module 2110 stores one or more artist packs 2112 and implements a feature extraction module 2114 , a client simulator module 2116 , and a deep neural network 2118 .
[0260] In some embodiments, the feature extraction module 2114 adds loops to the loop library after analyzing the looped audio (though it should be noted that some loops may be received with already generated metadata and do not require analysis). For example, raw audio in formats such as wav, aiff, or FLAC may be analyzed for quantifiable musical characteristics such as instrument classification, pitch transcription, beat timing, tempo, file length, and audio amplitude in multiple frequency bins. The analysis module 2110 may also store more abstract musical characteristics or mood descriptions for loops, based, for example, on manual tagging by the artist or machine listening. For example, mood may be quantified using multiple discrete categories with a range of values for each category for a given loop.
[0261] For example, consider Loop A, which is analyzed to determine that the notes G2, Bb2, and D2 are used, the first beat begins 6 milliseconds into the file, the tempo is 122 bpm, the file is 6483 milliseconds long, and the loop has normalized amplitude values of 0.3, 0.5, 0.7, 0.3, and 0.2 across five frequency bins. An artist could call the loop "Funk genre" with the following mood values: [Table 1]
[0262] The analysis module 2110 can store this information in a database, and clients can download subsections of the information, for example, as loop packages. While an artist pack 2112 is shown for illustrative purposes, the analysis module 2110 can provide various types of loop packages to the composition module 2120.
[0263] The client simulator module 2116, in the illustrated embodiment, analyzes various types of feedback and provides feedback information in a format supported by the deep neural network 2118. In the illustrated embodiment, the deep neural network 2118 also receives as input the execution script generated by the composition module. In some embodiments, the deep neural network configures the composition module based on these inputs, for example, to improve the correlation between the type of musical output generated and the desired feedback. For example, the deep neural network may periodically push updates to the client device implementing the composition module 2120. Note that the deep neural network 2118 is shown for illustrative purposes and may provide powerful machine learning capabilities in the disclosed embodiments, but is not intended to limit the scope of the present disclosure. In various embodiments, various types of machine learning techniques may be implemented alone or in various combinations to perform similar functions. It should be noted that in some embodiments, the machine learning module may be used to directly implement a rule set (e.g., placement rules or techniques), or may be used to control a module that implements other types of rule sets, for example, in the illustrated embodiment using a deep neural network 2118.
[0264] In some embodiments, the analysis module 2110 generates composition parameters for the composition module 2120 to improve the correlation between desired feedback and the use of certain parameters. For example, actual user feedback may be used to adjust composition parameters, e.g., to attempt to reduce negative feedback.
[0265] As an example, consider a situation in which module 2110 discovers a correlation between negative feedback (e.g., explicit low rankings, low volume listening, short listening times, etc.) and compositions that use a large number of layers. In some embodiments, module 2110 uses techniques such as back-propagation to determine that adjusting the probability parameters used to add more tracks will reduce the frequency of this problem. For example, module 2110 may determine that decreasing the probability parameters by 50% will reduce negative feedback by 8%, and predict that it can push the updated parameters to the composition module (the probability parameters are discussed in more detail below, but note that any of the various parameters for the statistical model may be adjusted similarly).
[0266] As another example, consider a situation in which module 2110 discovers that negative feedback correlates with a user setting the mood control to high tension. A correlation may also be found between loops with low tension tags and users requesting high tension. In this case, module 2110 can increase a parameter to increase the probability of selecting a loop with a high tension tag when a user requests high tension music. Thus, machine learning can be based on a variety of information, including composition output, feedback information, user control input, etc.
[0267] In the illustrated embodiment, the composition module 2120 includes a section sequencer 2122, a section arranger 2124, a technique implementation module 2126, and a loop selection module 2128. In some embodiments, the composition module 2120 organizes and configures sections of a composition based on loop metadata and user control input (e.g., mood control).
[0268] In some embodiments, the section sequencer 2122 sequences through different types of sections. In some embodiments, the section sequencer 2122 implements a finite state machine to sequentially output sections of the next type during operation. For example, the composition module 2120 may be configured to use different types of sections, such as intros, buildups, drops, breakdowns, and bridges, which are described in more detail below with reference to FIG. 23. Furthermore, each section may include multiple subsections that define how the music changes throughout the section, including, for example, a transition subsection, a main content subsection, and a transition-out subsection.
[0269] The section arranger 2124, in some embodiments, composes subsections according to arrangement rules. For example, one rule may specify a transition-in by gradually adding tracks. Another rule may specify a transition-in by gradually increasing the gain of a set of tracks. Another rule may specify chopping vocal loops to create a melody. In some embodiments, the probability of a loop in the loop library being added to a track is a function of user-input parameters such as the current position within the section or subsection, loops that overlap in time on other tracks, and mood variables. Functions may be adjusted by adjusting coefficients, for example, based on machine learning.
[0270] In some embodiments, the technique implementation module 2120 is configured to facilitate section arrangement, for example, by adding rules specified by the artist or determined by analyzing a particular artist's compositions. A "technique" may describe how a particular artist implements an arrangement rule at a technique level. For example, in an arrangement rule that specifies a transition-in by gradually adding tracks, one technique may indicate adding tracks in the following order: drums, bass, pads, and then vocals, while another technique may indicate adding tracks in the following order: bass, pads, vocals, and then drums. Similarly, in an arrangement rule that specifies chopping a vocal loop to create a melody, a technique may indicate chopping the vocals every second beat and repeating the chopped section of the loop twice before moving on to the next chopped section.
[0271] In the illustrated embodiment, loop selection module 2128 selects loops according to placement rules and techniques for inclusion in sections by section arranger 2124. Once a section is completed, a corresponding execution script may be generated and sent to execution module 2130. Execution module 2130 may receive execution script portions at various levels of granularity, including, for example, an overall execution script for a performance of a certain length, an execution script for each section, an execution script for each subsection, etc. In some embodiments, placement rules, techniques, or loop selection are performed statistically, e.g., with different approaches using different time percentages.
[0272] In the illustrated embodiment, execution module 2130 includes a filter module 2131, an effects module 2132, a mix module 2133, a master module 2134, and an execution module 2135. In some embodiments, these modules process an execution script and generate music data in a format supported by audio output device 2140. The execution script can specify loops to be played, when they should be played, effects to be applied by module 2132 (e.g., per track or per subsection), filters to be applied by module 2131, etc.
[0273] For example, a run script can specify that a low-pass filter ramp from 1000 to 20000 Hz should be applied to a particular track. As another example, a run script can specify that reverb should be applied to a particular track with a 0.2 wet setting (5000 to 15000 ms).
[0274] In some embodiments, the mix module 2133 is configured to perform automatic level control on the combined tracks. In some embodiments, the mix module 2133 also uses frequency domain analysis of the combined tracks to determine frequencies with excess or insufficient energy and applies gain to tracks in different frequency bands to perform the mix. In some embodiments, the master module 2134 is configured to perform multi-band compression, equalization (EQ), or limiting procedures to generate data for final formatting by the execute module 2135. The embodiment of FIG. 21 can automatically generate various output musical content according to user input or other feedback information, while machine learning techniques can improve the user's experience over time.
[0275] 22 is a diagram illustrating an example of an augmentation section of musical content, according to some embodiments. The system of FIG. 21 can construct such a section by applying placement rules and techniques. In the illustrated example, the build-up section includes three subsections and separate tracks for vocals, pads, drums, bass, and white noise.
[0276] In the illustrated example, the subsection transition includes drum loop A, which is also repeated in the main content subsection. The subsection transition also includes bass loop A. As shown, the gain of the section starts low and increases linearly throughout the section (although non-linear increases and decreases are possible). In the illustrated example, the main content and transition-out subsections include various vocals, pads, drums, and bass loops. As described above, the disclosed techniques for automatically sequencing, arranging, and implementing sections are capable of generating a nearly infinite stream of output musical content based on various user-adjustable parameters.
[0277] In some embodiments, the computer system displays an interface similar to Figure 22, allowing the artist to specify the techniques used to compose a section. For example, the artist can create a structure like that shown in Figure 22, which can be parsed into code for a composition module.
[0278] FIG. 23 illustrates an exemplary technique for arranging sections of musical content, according to some embodiments. In the illustrated embodiment, a generated stream 2310 includes multiple sections 2320, each including a beginning subsection 2322, a development subsection 2324, and a transition subsection 2326. In the illustrated example, multiple types of each section / subsection are indicated in a table with dotted lines. In the illustrated embodiment, the circular elements are examples of arranging tools, which may be further implemented using specific techniques as discussed below. As shown, various composition decisions may be performed pseudo-randomly according to statistical percentages. For example, the type of subsection, the arranging tools for a particular type or subsection, or the techniques used to implement the arranging tools may be determined statistically.
[0279] In the illustrated example, a given section 2320 is one of five types: intro, buildup, drop, breakdown, and bridge, each with a different function controlling the intensity of the entire section. In this example, the state subsection is one of three types: slow build, sudden shift, or minimal, each type being different. The development subsection in this example is one of three types: reduce, transform, or augment. In this example, the transition subsection is one of three types: collapse, ramp, or hint. The different types of sections and subsections may be selected based on rules, for example, or may be selected pseudo-randomly.
[0280] In the illustrated example, the behavior of different subsection types is implemented using one or more placement tools. For the slow build, in this example, a low-pass filter is applied 40% of the time, and a time layer is added 80% of the time. For the transform evolution subsection, in this example, 25% of the time loop is chopped. Various additional placement tools are displayed, such as one-shot, drop-out beat, apply reverb, add pad, add theme, remove layer, white noise, etc. These examples are included for illustrative purposes and are not intended to limit the scope of the present disclosure. Furthermore, for ease of explanation, these examples may not be complete (e.g., actual placement may typically include a much larger number of placement rules).
[0281] In some embodiments, one or more arrangement tools may be implemented using specific techniques (which may be artist-specified or determined based on an analysis of the artist's content). For example, one-shots may be implemented using sound effects or vocals, loop chopping may be implemented using stutter or half-chopping techniques, layer removal may be implemented by synthesizing or removing vocals, white noise may be implemented using a ramp or pulse function, etc. In some embodiments, the specific technique selected for a given arrangement tool may be selected according to a statistical function (e.g., 30% of the time layer removal may remove synthesis, 70% of the time vocals may be removed for a given artist). As noted above, arrangement rules or techniques may be determined automatically by analyzing existing compositions, for example, using machine learning.
[0282] Example method 24 is a flowchart method for using a ledger according to some embodiments. The method shown in FIG. 24 may be used in conjunction with, among other things, any of the computer circuits, systems, devices, elements, or components disclosed herein. In various embodiments, some of the illustrated method elements may be performed simultaneously, in a different order than that shown, or may be omitted. Additional method elements may also be performed as desired.
[0283] In the illustrated embodiment, at 2410, the computing device determines playback data characteristic of the playback of the musical content mix. The mix may include a determined combination of multiple audio tracks (note that the track combination may be determined in real time, e.g., just prior to the output of the current portion of the musical content mix, which is a continuous stream of content). This determination may be based on the configuration of the content mix (e.g., by a server or playback device such as a mobile phone) or may be received from another device that determines which audio files to include in the mix. The playback data may be stored (e.g., in offline mode) and / or encrypted. The playback data may be reported periodically or in response to a specific event (e.g., restored connectivity to the server).
[0284] In the illustrated embodiment, the computing device records, in an electronic blockchain ledger data structure, information specifying individual playback data for one or more of the plurality of audio tracks in the music content mix. In the illustrated embodiment, the information specifying the individual playback data for the individual audio tracks includes usage data for the individual audio tracks and signature information associated with the individual audio tracks.
[0285] In one aspect, the characteristic information is an identifier for one or more entities. For example, the signature information may be a string or a unique identifier. In other embodiments, the signature information may be encrypted or otherwise obfuscated to prevent others from identifying the entity. In some embodiments, the usage data includes at least one of a duration played for the musical content mix or a number of times played for the musical content mix.
[0286] In some embodiments, data identifying individual audio tracks within the musical content mix is retrieved from a data store that also indicates actions to be performed in connection with including one or more individual audio tracks. In these embodiments, recording can include recording an indication of proof of performance of the indicated actions.
[0287] In some embodiments, the system determines rewards for multiple entities associated with multiple audio tracks based on information identifying their respective playback data recorded in the electronic blockchain ledger.
[0288] In some embodiments, the system determines usage data for a first individual audio track not included in the musical content mix in its original musical form. For example, the audio track can be modified and used to generate a new audio track, and the usage data can be adjusted to reflect this modification or usage. In some embodiments, the system generates the new audio track based on interpolation between vector representations of audio in at least two of the multiple audio tracks, and the usage data is determined based on the distance between the vector representation of the first individual audio track and the vector representation of the new audio track. In some implementations, the usage data is based on a ratio of Euclidean distances from the interpolated vector representations and vectors in at least two of the multiple audio tracks.
[0289] 25 is a flow diagram of a method for using a graphical representation to combine audio files, according to some implementations. The method illustrated in FIG. 25 may be used in conjunction with, among other things, any of the computer circuits, systems, devices, elements, or components disclosed herein. In various embodiments, some of the illustrated method elements may be performed simultaneously, in a different order than that illustrated, or may be omitted. Additional method elements may also be performed as desired.
[0290] In the illustrated embodiment, the computing device generates 2510 multiple image representations of multiple audio files, where an image representation for a particular audio file is generated based on data in the particular audio file and a MIDI representation of the particular audio file. In some implementations, pixel values in the image representation represent velocities in the audio file where the image representation is compressed at a resolution of the velocities.
[0291] In some embodiments, the image representation is a two-dimensional representation of the audio file. In some embodiments, pitch is represented by columns of the two-dimensional representation, where time is represented by the columns, and pixel values of the two-dimensional representation represent velocity. In some embodiments, pitch is represented by columns of the two-dimensional representation, where time is represented by the columns, and pixel values of the two-dimensional representation represent velocity. In some embodiments, the pitch axis is banded into two sets of octaves over an eight-octave range, where the first 12 rows of pixels represent the first four octaves that determine the pixel values of the pixels, the second 12 rows of pixels represent the second four octaves that determine the pixel values of the pixels, and the second four octaves represent one of the second four octaves. In some embodiments, odd pixel values along the time axis represent the onset of a note, and even pixel values along the time axis represent the duration of a note. In some embodiments, each pixel represents a portion of a beat in the time dimension.
[0292] In the illustrated embodiment, the computing device selects 2520 a plurality of audio files based on the plurality of image representations.
[0293] In the illustrated embodiment, the computing device combines (2530) multiple music files to generate output music content.
[0294] In some embodiments, one or more composition rules are applied to select the plurality of audio files based on the plurality of image representations. In some implementations, applying the one or more composition rules includes removing pixel values in the image representations that exceed a first threshold and removing pixel values in the image representations that are below a second threshold.
[0295] In some embodiments, one or more machine learning algorithms are applied to the image representations to select and combine multiple ones of the audio files and generate the output musical content. In some embodiments, harmonic and rhythmic coherence are tested in the output musical content.
[0296] In some embodiments, a single image representation is generated from the multiple image representations, and a description of texture features is added to the single image representation from which texture features are extracted from the multiple audio files. In some embodiments, the single image representation is stored with the multiple audio files. In some embodiments, the multiple audio files are selected by applying one or more composition rules to the single image representation.
[0297] 26 is a flow diagram of a method for implementing user-generated control elements, according to some embodiments. The method illustrated in FIG. 26 may be used, among other things, with any of the computer circuits, systems, devices, elements, or components disclosed herein. In various embodiments, some of the illustrated method elements may be performed simultaneously, in a different order than that illustrated, or may be omitted. Additional method elements may also be performed as desired.
[0298] In the illustrated embodiment, a computing device accesses a plurality of audio files at 2610. In some embodiments, the audio files are accessed from memory of a computer system and a user has rights to the accessed audio files.
[0299] In the illustrated embodiment, the computing device generates the output musical content at 2620 by combining musical content from two or more audio files using at least one trained machine learning algorithm. In some embodiments, the combination of musical content is determined by the at least one trained machine learning algorithm based on the musical content in the two or more audio files. In some embodiments, the at least one trained machine learning algorithm combines the musical content by sequentially selecting musical content from the two or more audio files based on the musical content in the two or more audio files.
[0300] In some embodiments, the at least one trained machine learning algorithm is trained to select musical content for a beat that comes after a specified time based on metadata of the musical content played up to the specified time, hi some embodiments, the at least one trained machine learning algorithm is further trained to select musical content for a beat that comes after a specified time based on a level of the control element.
[0301] In the illustrated embodiment, the computing device implements user-generated control elements on a user interface for varying user-specified parameters in the generated output musical content, the levels of the one or more audio parameters in the generated output musical content are determined based on the levels of the control elements, and the relationship between the levels of the one or more audio parameters and the levels of the control elements is determined based on user input during at least one music playback session. In some embodiments, the levels of the user-specified parameters vary based on one or more environmental conditions.
[0302] In some embodiments, the relationship between the levels of one or more audio parameters and the levels of the control element is determined by playing a plurality of audio tracks during at least one music playback session, the plurality of audio tracks having varying audio parameters, receiving, for each of the audio tracks, an input specifying a user-selected level of the parameter within the audio track, evaluating, for each of the audio tracks, the levels of one or more audio parameters within the audio track, and determining a relationship between the levels of the one or more audio parameters and the levels of the control element based on a correlation between each level of the user-selected parameter and each level of the one or more audio parameters.
[0303] In some embodiments, the relationship between the level of the one or more audio parameters and the level of the control element is determined using one or more machine learning algorithms. In some embodiments, the relationship between the level of the one or more audio parameters and the level of the control element is refined based on user variation of the level of the control element during playback of the generated output musical content. In some embodiments, the level of the one or more audio parameters in the audio track is evaluated using metadata from the audio track. In some embodiments, the relationship between the level of the one or more audio parameters and the level of the user-specified parameter is further based on additional user input during one or more additional music playback sessions.
[0304] In some embodiments, the computing device implements at least one additional user-generated control element on the user interface to vary an additional user-specified parameter in the generated output musical content, where the additional user-specified parameter is a sub-parameter of the user-specified parameter. In some embodiments, the generated output musical content is altered based on a user adjustment of a level of the control element. In some embodiments, a feedback control element is implemented on the user interface, the feedback control element allowing a user to provide positive or negative feedback on the generated output musical content during playback. In some embodiments, at least one trained machine algorithm alters the generation of subsequent generated output musical content based on feedback received during playback.
[0305] 27 is a flow diagram of a method for generating musical content by modifying audio parameters, according to some embodiments. The method shown in FIG. 27 may be used in conjunction with, among other things, any of the computer circuits, systems, devices, elements, or components disclosed herein. In various embodiments, some of the illustrated method elements may be performed simultaneously, in a different order than that shown, or may be omitted. Additional method elements may also be performed as desired.
[0306] In the illustrated embodiment, the computing device accesses a set of music content at 2710. In some embodiments.
[0307] In the illustrated embodiment, the computing device generates a first graph of an audio signal of the musical content, the first graph being a graph of an audio parameter over time.
[0308] In the illustrated embodiment, the computing device generates a second graph of the audio signal of the musical content, the second graph being a signal graph of the audio parameters versus beats, hi some embodiments, the second graph of the audio signal has a similar structure to the first graph of the audio signal.
[0309] In the illustrated embodiment, the computing device generates new music content from the played music content by modifying audio parameters in the played music content, where the audio parameters are modified based on a combination of the first graph and the second graph.
[0310] In some embodiments, the audio parameters of the first graph and the second graph are defined by nodes of the graphs that determine changes in properties of the audio signals. In some embodiments, generating new musical content includes receiving the played musical content, determining a first node in the first graph that corresponds to an audio signal in the played musical content, determining a second node in the second graph that corresponds to the first node, determining one or more specific audio parameters based on the second node, and modifying one or more properties of the audio signal in the played musical content by modifying the specific audio parameters. In some embodiments, one or more additional identified audio parameters are determined based on the first node, and one or more properties of the additional audio signal in the played musical content are modified by modifying the additional identified audio parameters.
[0311] In some embodiments, determining the one or more audio parameters includes determining a portion of the second graph to realize the audio parameters based on a position of the second node in the second graph, and selecting audio parameters from the determined portion of the second graph as the one or more audio specification parameters. In some embodiments, a portion of the played musical content corresponding to the determined portion of the second graph is modified by modifying the one or more identified audio parameters. In some embodiments, the modified property of the audio signal in the played musical content includes signal amplitude, signal frequency, or a combination thereof.
[0312] In some embodiments, one or more automations are applied to audio parameters, and at least one of the at least one automations is a pre-programmed temporal manipulation of the at least one audio parameter. In some embodiments, one or more modulations are applied to audio parameters, where at least one of the at least one modulations multiplicatively modifies the at least one audio parameter on top of the at least one automation.
[0313] The following numbered items describe various non-limiting embodiments disclosed herein. Set A (A1) 1. A method comprising: generating, by a computer system, a plurality of graphical representations of a plurality of audio files, wherein the graphical representation of a particular audio file is generated based on data within the particular audio file and a MIDI representation of the particular audio file; selecting a plurality of the audio files based on the plurality of image representations; combining said plurality of said audio files to generate output musical content; A method comprising: (A2) The method of any of the preceding clauses in set A, wherein pixel values of the image representation represent a speed of the audio file, and the image representation is compressed at a speed resolution. (A3) The method of any of the preceding clauses in set A, wherein the image representation is a two-dimensional representation of the audio file. (A4) 10. The method of any of the preceding items in set A, wherein pitch is represented by rows of the two-dimensional representation, time is represented by columns of the two-dimensional representation, and pixel values of the two-dimensional representation represent velocity. (A5) The method of any of the preceding items in set A, wherein the two-dimensional representation is 32 pixels wide by 24 pixels high, with each pixel representing a fraction of a beat in the time dimension. (A6) 10. The method of any of the preceding clauses in set A, wherein the pitch axis is banded into two octave sets over a range of eight octaves, and wherein the first 12 rows of pixels represent the first four octaves with pixel values of pixels determining which of the first four octaves are represented, and the second 12 rows of pixels represent the second four octaves with pixel values of pixels determining which of the second four octaves are represented. (A7) 10. The method of any of the preceding items in set A, wherein odd pixel values along the time axis represent the onset of a musical note and even pixel values along the time axis represent the duration of a musical note. (A8) The method of any of the preceding clauses in set A, further comprising applying any one of a set of one or more composition rules to select a plurality of audio files based on the plurality of image representations. (A9) The method of any of the preceding clauses in set A, wherein the step of applying the one or more composition rules includes removing pixel values in the image representation that exceed a first threshold and removing pixel values in the image representation that are below a second threshold. (A10) 10. The method of any of the preceding clauses in set A, further comprising the steps of selecting and combining the plurality of audio files from the audio files and applying one or more machine learning algorithms to the image representations to generate the output musical content. (A11) The method of any of the preceding clauses in set A, further comprising testing harmonic and rhythmic coherence in the output musical content. (A12) generating a single image representation from the plurality of image representations; extracting one or more texture features from the plurality of audio files; adding a description of the extracted texture features to the single image representation; The method of any of the preceding clauses in Set A, further comprising: (A13) A non-transitory computer-readable medium having instructions stored thereon, the non-transitory computer-readable medium being executable by a computing device to perform operations including any combination of the operations performed by the method of any of the preceding clauses in set A. (A14) A device, one or more processors; one or more memories containing program instructions; wherein the program instructions are executable by the one or more processors to perform any combination of the operations performed by the method of any of the preceding clauses in set A. Set B (B1) 1. A method comprising: accessing, by a computer system, a set of music content; generating, by the computer system, a first graph of an audio signal of the musical content, the first graph being a graph of an audio parameter versus time; generating, by the computer system, a second graph of the audio signal of the musical content, the second graph being a signal graph of the audio parameters relative to beats; generating new music content from the reproduced music content by modifying the audio parameters in the reproduced music content, the audio parameters being modified based on a combination of the first graph and the second graph; A method comprising: (B2) The method of any of the preceding clauses in set B, wherein the second graph of the audio signal has a similar structure to the first graph of the audio signal. (B3) The method of any of the preceding clauses in set B, wherein the audio parameters in the first graph and the second graph are defined by nodes in graphs that determine changes in characteristics of the audio signal. (B4) The step of generating new music content includes: receiving a playback musical content; and determining a first node in a first graph corresponding to an audio signal in the playback musical content; determining a second node in a second graph corresponding to the first node; and determining one or more particular audio parameters based on the second node; modifying one or more properties of an audio signal in the played music content by changing certain audio parameters; Any of the methods of the preceding items in set B, including: (B5) determining one or more additional identified audio parameters based on the first node; changing one or more properties of the additional audio signal in the played music content by modifying the additional identified audio parameters; The method of any of the preceding clauses in Set B, further comprising: (B6) determining the one or more audio parameters includes: determining a portion of the second graph that realizes the audio parameter based on a position of the second node within the second graph; selecting the audio parameters from the determined portion of the second graph as the one or more audio designation parameters; The method of any of the preceding items in Set B, comprising: (B7) The method of any of the preceding clauses in set B, wherein modifying the one or more identified audio parameters modifies a portion of the played musical content that corresponds to the determined portion of the second graph. (B8) The method of any of the preceding clauses in set B, wherein the modified property of the audio signal in the played music content includes signal amplitude, signal frequency, or a combination thereof. (B9) The method of any of the preceding clauses in set B, further comprising the step of applying one or more automations to the audio parameters, at least one of the automations being pre-programmed temporal manipulation of at least one audio parameter. (B10) The method of any of the preceding clauses in set B, further comprising applying one or more modulations to the audio parameters, at least one of the modulations multiplicatively modifying at least one audio parameter on top of at least one automation. (B11) The method of any of the preceding clauses in set B, further comprising providing, by the computer system, a heap-allocated memory for storing one or more objects associated with the set of musical content, the one or more objects including at least one of the audio signal, the first graph, and the second graph. (B12) The method of any of the preceding clauses in set B, further comprising storing the one or more objects in a data structure list in the heap-allocated memory, the data structure list comprising a serialized list of links for the objects. (B13) A non-transitory computer-readable medium having instructions stored thereon, the non-transitory computer-readable medium being executable by a computing device to perform operations including any combination of the operations performed by the method of any of the preceding clauses in Set B. (B14) A device, one or more processors; one or more memories containing program instructions; wherein the program instructions are executable by the one or more processors to perform any combination of the operations performed by the method of any of the preceding clauses in Set B. Set C (C1) 1. A method comprising: accessing a plurality of audio files by a computer system; generating output musical content by combining musical content from two or more audio files using at least one trained machine learning algorithm; implementing, on a user interface associated with said computer system, user-created control elements for varying user-specified parameters in said generated output musical content; Including, wherein levels of one or more audio parameters in the generated output musical content are determined based on a level of the control element, and a relationship between the levels of the one or more audio parameters and the levels of the control element is based on user input during at least one music playback session. (C2) The method of any of the preceding clauses in Set C, wherein the combination of musical content is determined by the at least one trained machine learning algorithm based on musical content in two or more audio files. (C3) The method of any of the preceding clauses in Set C, wherein the at least one trained machine learning algorithm combines the musical content by sequentially selecting musical content from the two or more audio files based on the musical content in the two or more audio files. (C4) The method of any of the preceding clauses in Set C, wherein the at least one trained machine learning algorithm is trained to select musical content for a beat that comes after a specified time based on metadata of musical content played up to the specified time. (C5) The method of any of the preceding clauses in set C, wherein the at least one trained machine learning algorithm is further trained to select musical content for a beat coming after the specified time based on a level of the control element. (C6) The relationship between the level of the one or more audio parameters and the level of the control element is: playing a plurality of audio tracks during the at least one music playback session, the plurality of audio tracks having varying audio parameters; receiving, for each of the audio tracks, an input specifying a user-selected level of the user-specified parameter within the audio track; for each of said audio tracks, evaluating the level of one or more audio parameters within said audio track; determining a relationship between the levels of the one or more audio parameters and the levels of the control element based on a correlation between each of the user-selected levels of the user-specified parameters and each of the evaluated levels of the one or more audio parameters; any of the preceding terms in set C, as determined by (C7) The method of any of the preceding clauses in set C, wherein the relationship between the level of the one or more audio parameters and the level of the control element is determined using one or more machine learning algorithms. (C8) The method of any of the preceding clauses in set C, further comprising fine-tuning the relationship between the level of the one or more audio parameters and the level of the control element based on user variation of the level of the control element during playback of the generated output musical content. (C9) The method of any of the preceding clauses in set C, wherein the level of the one or more audio parameters in the audio track is evaluated using metadata from the audio track. (C10) The method of any of the preceding clauses in set C, further comprising the step of varying, by the computer system, the level of the user-specified parameter based on one or more environmental conditions. (C11) The method of any of the preceding clauses in Set C, further comprising implementing on the user interface associated with the computer system at least one additional control element created by the user for variation of an additional user-specified parameter in the generated output musical content, the additional user-specified parameter being a sub-parameter of the user-specified parameter. (C12) The method of any of the preceding clauses in set C, further comprising accessing the audio file from memory of the computer system, wherein the user has rights to the accessed audio file. (C13) A non-transitory computer-readable medium having instructions stored thereon, the non-transitory computer-readable medium being executable by a computing device to perform operations including any combination of the operations performed by the method of any of the preceding clauses in Set C. (C14) A device, one or more processors; one or more memories containing program instructions; wherein the program instructions are executable by the one or more processors to perform any combination of the operations performed by the method of any of the preceding clauses in Set C. Set D (D1) 1. A method comprising: determining, by a computer system, playback data for a musical content mix, the playback data indicating playback characteristics of the musical content mix, the musical content mix including a determined combination of a plurality of audio tracks; recording, by the computer system, in an electronic blockchain ledger data structure, information specifying individual playback data for one or more of the plurality of audio tracks in the music content mix, wherein the information specifying individual playback data for each audio track includes usage data for the individual audio track and signature information associated with the individual audio track; A method comprising: (D2) The method of any of the preceding clauses in set D, wherein the usage data includes at least one of a duration of time the musical content mix has been played or a number of times the musical content mix has been played. (D3) The method of any of the preceding clauses in set D, further comprising determining the information specifying individual playback data for the individual audio tracks based on the playback data for the musical content mix and data identifying individual audio tracks within the musical content mix. (D4) The method of any of the preceding clauses in Set D, wherein data identifying individual audio tracks in the musical content mix is retrieved from a data store that also indicates actions to be performed in connection with including one or more individual audio tracks, and the recording includes recording instructions of proof of performance of the indicated actions. (D5) The method of any of the preceding clauses in set D, further comprising identifying one or more entities associated with each of the audio tracks based on signature information for a plurality of entities associated with the plurality of audio tracks and data specifying signature information. (D6) The method of any of the preceding clauses in Set D, further comprising determining rewards for the plurality of entities associated with the plurality of audio tracks based on information identifying each playback data recorded in the electronic blockchain ledger. (D7) The method of any of the preceding clauses in Set D, wherein the playback data includes usage data for one or more machine learning modules used to generate the music mix. (D8) The method of any of the preceding clauses in set D, further comprising the step of at least temporarily storing, by a playback device, the playback data of the musical content mix. (D9) The method of any of the preceding clauses in D, wherein the stored playback data is communicated by the playback device to another computer system periodically or in response to an event. (D10) The method of any of the preceding clauses in set D, further comprising the step of encrypting, by the playback device, the playback data of the music content mix. (D11) The method of any of the preceding clauses in set D, wherein the blockchain ledger is publicly accessible and immutable. (D12) The method of any of the preceding clauses in set D, further comprising determining usage data for a first individual audio track not included in the music content mix in its original music format. (D13) generating a new audio track based on interpolation between vector representations of audio in at least two of the plurality of audio tracks; The method of any of the preceding clauses in set D, wherein the determining of the usage data is based on a distance between the vector representation of the first individual audio track and the vector representation of the new audio track. (D14) A non-transitory computer-readable medium having instructions stored thereon, the non-transitory computer-readable medium being executable by a computing device to perform operations including any combination of operations performed by the method of any of the preceding clauses in set D. (D15) A device, one or more processors; one or more memories containing program instructions; wherein the program instructions are executable by the one or more processors to perform any combination of the operations performed by the method of any of the preceding clauses in Set D. Set E (E1) 1. A method comprising: storing, by a computing system, data specifying a plurality of tracks of a plurality of different entities, said data including signature information for each of said plurality of music tracks; layering, by said computing system, a plurality of tracks to generate output musical content; Recording information specifying the tracks included in the layering and signature information for each of the tracks in an electronic blockchain ledger data structure; A method comprising: (E2) 10. The method of any of the clauses in set E, wherein the blockchain ledger is publicly accessible and immutable. (E3) detecting use of one of the tracks by an entity that does not match the signature information; Any method of the items in set E further including: (E4) 1. A method comprising: storing, by a computing system, data specifying a plurality of tracks for a plurality of different entities, said data including metadata linking to other content; selecting and layering, by said computing system, a plurality of tracks to generate output musical content; retrieving content to be output in association with the output musical content based on the metadata of the selected track; A method comprising: (E5) The method of any of the clauses in set E, wherein the other content includes a visual advertisement. (E6) 1. A method comprising: storing, by a computing system, data specifying a plurality of tracks; causing output of a plurality of musical examples by said computing system; receiving, by the computing system, a user input indicating a user's opinion as to whether one or more of the musical examples exhibit a first musical parameter; storing custom rule information for the user based on the receiving step; receiving, by the computing system, a user input indicating an adjustment to the first musical parameter; selecting, by the computing system, and layer-linking a plurality of tracks based on the custom rule information to generate output musical content; A method comprising: (E7) The method of any of the clauses in set E, wherein the user specifies a name for the first musical parameter. (E8) The method of any of the clauses in set E, wherein the user specifies one or more targets for one or more ranges of the first musical parameter, and the selecting and layering steps are based on past feedback information. (E9) The method of any of the clauses in set E, further comprising, in response to determining that the selecting and layering steps based on the past feedback information did not meet one or more goals, adjusting a plurality of other musical attributes and determining an effect of adjusting the plurality of other musical attributes on one or more goals. (E10) The method of any of the terms in set E, wherein the goal is the user's heart rate. (E11) The method of any of the clauses in set E, wherein the selecting and layering steps are performed by a machine learning engine. (E12) The method of any of the clauses in set E, further comprising the steps of training the machine learning engine, the steps comprising training a teacher model based on a plurality of types of user input, and training a student model based on the received user input. (E13) 1. A method comprising: storing, by a computing system, data specifying a plurality of tracks; determining, by the computing system, that an audio device is located in a first type of environment; receiving user input adjusting a first musical parameter while the audio device is located within the first type of environment; selecting and layering, by the computing system, a plurality of tracks based on the adjusted first musical parameters and the first type of environment to generate output musical content via the audio device; determining, by the computing system, that the audio device is located in a second type of environment; receiving a user input adjusting a first musical parameter while the audio device is in the second type of environment; selecting and layering a plurality of tracks based on the adjusted first musical parameters and the second type of environment, by the computing system, to output musical content via the audio device; A method comprising: (E14) The method of any of the clauses in set E, wherein the selecting and layering step adjusts different musical attributes based on adjustments to the first musical parameters when the audio device is in the first type of environment than when the audio device is in the second type of environment. (E15) 1. A method comprising: storing, by a computing system, data specifying a plurality of tracks; storing, by said computing system, data specifying musical content; determining one or more parameters of the stored musical content; selecting and layering, by the computing system, a plurality of tracks to generate output musical content based on one or more parameters of the stored musical content; causing output of the output musical content and the stored musical content such that they overlap in time; A method comprising: (E16) 1. A method comprising: storing, by a computing system, data specifying a plurality of tracks and corresponding track attributes; storing, by said computing system, data specifying musical content; determining one or more parameters of the stored musical content to determine one or more tracks included in the stored musical content; selecting and layering, by the computing system, tracks from the stored tracks and tracks from the determined tracks to generate output musical content based on the one or more parameters; A method comprising: (E17) 1. A method comprising: storing, by a computing system, data specifying a plurality of tracks; selecting and layering, by said computing system, a plurality of tracks to generate output musical content; Including, The selection may be a neural network module configured to select tracks for future use based on recently used tracks; one or more hierarchical hidden Markov models configured to select a structure for upcoming output musical content based on the selected structure and to constrain the neural network module; This is done using both methods. (E18) The method of any of the clauses in set E, further comprising training one or more hierarchical hidden Markov models based on positive or negative user feedback regarding the output musical content. (E19) The method of any of the clauses in set E, further comprising training the neural network model based on user selections provided in response to a plurality of examples of musical content. (E20) The method of any of the clauses in set E, wherein the neural network module includes multiple layers, including at least one globally trained layer and at least one layer specifically trained based on feedback from a particular user account. (E21) 1. A method comprising: storing, by a computing system, data specifying a plurality of tracks; selecting and layering, by said computing system, a plurality of tracks to generate output musical content; wherein the selection is performed using both a globally trained machine learning module and a locally trained machine learning module. (E22) The method of any of the clauses in set E, comprising training the globally trained machine learning module using feedback from multiple user accounts and training a locally trained machine learning module based on feedback from a single user account. (E23) The method of any of the clauses in set E, wherein both the globally trained machine learning module and the locally trained machine learning module provide output information based on user adjustments of musical attributes. (E24) The method of any of the terms in set E, wherein the globally trained machine learning module and the locally trained machine learning module are included in different layers of a neural network. (E25) 1. A method comprising: receiving a first user input indicating a value of a high-level music composition parameter; automatically selecting and combining audio tracks based on values of said high-level music composition parameters, comprising using a plurality of values of one or more sub-parameters associated with said high-level music composition parameters while outputting musical content; receiving a second user input indicating one or more values of the one or more sub-parameters; said automatically selecting and combining based on said second user input; A method comprising: (E26) The method of any of the clauses in set E, wherein the high-level music composition parameter is an energy parameter and the one or more sub-parameters include tempo, number of layers, a vocal parameter, or a bass parameter. (E27) 1. A method comprising: accessing the score information by a computing system; synthesizing, by the computing system, musical content based on the score information; analyzing the music content to generate frequency information for a plurality of frequency bins at different times; training a machine learning engine, the step including inputting the frequency information and using the score information as labels for the training; A method comprising: (E28) inputting music content into the trained machine learning engine; generating score information for the musical content using the machine learning engine; wherein the generated score information includes frequency information for a plurality of frequency bins at different times. (E29) Any method of clauses in set E, wherein said different times have a constant distance between them. (E30) The method of any of the clauses in set E, wherein the different points in time correspond to beats of the musical content. (E31) The method of any of the terms in set E, wherein the machine learning engine includes convolutional layers and recurrent layers. (E32) The method of any of the clauses in set E, wherein the frequency information includes a binary indication of whether the music content contains content in a frequency bin at a given time. (E33) The method of any of the clauses in set E, wherein the musical content includes a plurality of different instruments.
[0314] Although specific embodiments have been described above, these embodiments are not intended to limit the scope of the present disclosure, even though only a single embodiment is described with respect to a particular feature. The examples of features provided in this disclosure are intended to be illustrative rather than limiting, unless otherwise specified. The above description is intended to cover such alternatives, modifications, and equivalents as will be apparent to those skilled in the art having the benefit of this disclosure.
[0315] The scope of the present disclosure includes any feature or combination of features (whether explicit or implicit) disclosed herein, or any generalization thereof, whether or not it alleviates any or all of the problems solved herein. Accordingly, new claims may be formulated during prosecution of this application (or an application claiming priority from this application) to any such combination of features. In particular, with reference to the appended claims, features from the dependent claims may be combined with features of the independent claims, and features of each independent claim may be combined in any suitable manner, not just in specific combinations not recited in the appended claims. [Explanation of symbols]
[0316] 110 Stored audio files and corresponding attributes 120 stored rule sets 130 Target Music Attributes 140 output music content 150 Environmental information 160 Music Generation Module
Claims
1. 1. A method comprising: determining, by a computer system, playback data for a musical content mix, the playback data indicating playback characteristics of the musical content mix, the musical content mix including a determined combination of a plurality of audio tracks; recording, by the computer system, in an electronic blockchain ledger data structure, information specifying individual playback data for one or more of the plurality of audio tracks in the music content mix, wherein the information specifying individual playback data for each audio track includes usage data for the individual audio track and signature information associated with the individual audio track; A method comprising:
2. The method of claim 1 , wherein the usage data includes at least one of a duration played for the musical content mix or a number of times played for the musical content mix.
3. 2. The method of claim 1, further comprising determining the information specifying individual playback data for the individual audio tracks based on the playback data for the musical content mix and data identifying individual audio tracks within the musical content mix.
4. 2. The method of claim 1, wherein data identifying individual audio tracks within the musical content mix is retrieved from a data store that also indicates actions to be performed in connection with including one or more individual audio tracks, and wherein the recording step includes recording an indication of proof of performance of the indicated actions.
5. 10. The method of claim 1, further comprising identifying one or more entities associated with each of the audio tracks based on signature information for a plurality of entities associated with the plurality of audio tracks and data specifying the signature information.
6. 6. The method of claim 5, further comprising determining rewards for the plurality of entities associated with the plurality of audio tracks based on information specifying individual playback data recorded in the electronic blockchain ledger data structure.
7. The method of claim 1 , wherein the playback data includes usage data for one or more machine learning modules used to generate the musical content mix.
8. The method of claim 1 , further comprising at least temporarily storing the playback data of the music content mix by a playback device.
9. 9. The method of claim 8, wherein the stored playback data is communicated by the playback device to another computer system periodically or in response to an event.
10. The method of claim 8 , further comprising the step of encrypting, by the playback device, the playback data of the music content mix.
11. 10. The method of claim 1, wherein the electronic blockchain ledger data structure is publicly accessible and immutable.
12. 10. The method of claim 1, further comprising determining usage data for a first individual audio track not included in the music content mix in its original music format.
13. generating a new audio track based on interpolation between vector representations of audio in at least two of the plurality of audio tracks; The method of claim 12 , wherein determining usage data is based on a distance between a vector representation of the first individual audio track and a vector representation of the new audio track.
14. A non-transitory computer-readable medium storing instructions executable by a computing device to perform operations, the operations comprising: determining playback data for a musical content mix, the playback data being recorded by a playback device, the playback data indicating playback characteristics of the musical content mix, the musical content mix including a determined combination of a plurality of audio tracks; recording information specifying individual playback data for one or more of the plurality of audio tracks in the music content mix in an electronic blockchain ledger data structure, the information specifying individual playback data for each audio track including usage data for the individual audio track and signature information associated with the individual audio track; 1. A non-transitory computer-readable medium comprising:
15. 15. The non-transitory computer-readable medium of claim 14, wherein the musical content mix is a continuous combination of the multiple audio tracks created by layering and sequencing the multiple audio tracks.
16. 16. The non-transitory computer-readable medium of claim 15, wherein recording information specifying individual playback data comprises recording information specifying playback data for a layer or sequence within the plurality of audio tracks.
17. The musical content mix includes content that is a blended combination of multiple audio tracks, the blended combination comprising: interpolating between vector representations of audio in at least two of the plurality of audio tracks; generating the music content mix using an inverse transform from the vector representation; 15. The non-transitory computer-readable medium of claim 14, generated by
18. 18. The non-transitory computer-readable medium of claim 17, wherein the playback data includes at least one of a time played for the musical content mix or a number of times played for the musical content mix, and wherein the time played for each audio track or the number of times played is based, for at least two of the plurality of audio tracks, on a ratio of Euclidean distances from a vector in at least two of the plurality of audio tracks and the interpolated vector representation.
19. A device, one or more processors; one or more memories storing program instructions; wherein the program instructions include: determining playback data for a musical content mix, the playback data being recorded by a playback device, the playback data indicating playback characteristics of the musical content mix, the musical content mix including a determined combination of a plurality of audio tracks; recording, in an electronic blockchain ledger data structure, information specifying individual playback data for one or more of the plurality of audio tracks in the music content mix, the information specifying individual playback data for each audio track including usage data for the individual audio track and signature information associated with the individual audio track; 20. The method of claim 19, wherein the method is executable by the one or more processors to:
20. 20. The device of claim 19, wherein the electronic blockchain ledger data structure is publicly accessible and immutable.