Music content generation

By applying machine learning algorithms and user-defined controls in streaming music services, the problem of the inability to adjust music according to user tastes and environment in the prior art is solved, and customized and personalized music generation is achieved, improving the user experience.

CN115066681BActive Publication Date: 2025-05-30AIMI INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202180013939.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-08-21
Filing Date
2021-02-11
Publication Date
2025-05-30
Estimated Expiration
2041-02-11

AI Technical Summary

Technical Problem

Existing streaming music services are unable to adjust music to users’ tastes, environments, and behaviors, causing users to get bored when they hear the same songs and genres.

Method used

Through machine learning algorithms, customized music content is generated based on user input and environmental information, allowing users to create and train user-defined controls to influence music based on personal preferences.

Benefits of technology

It realizes the generation of customized music based on users' personalized preferences, improves the diversity and adaptability of music, and enhances the user's music experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115066681B_ABST
    Figure CN115066681B_ABST
Patent Text Reader

Abstract

Techniques related to the automatic generation of new music content from audio file-based image representations are disclosed. Techniques related to implementing audio techniques for real-time audio generation are also disclosed. Additional techniques related to implementing user-created controls for modifying music content are disclosed. More techniques related to tracking contributions to created music content are also disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to audio engineering and, more particularly, to generating music content. Background Art

[0002] Streaming music services typically provide songs to users via the Internet. Users can subscribe to these services and stream music through a web browser or application. Examples of such services include PANDORA, SPOTIFY, GROOVESHARK, etc. Users can often select a music genre or a specific artist to stream. Users can typically rate songs (e.g., using a star rating or a like / dislike system), and some music services can customize which songs are streamed to a user based on previous ratings. The cost of running a streaming service (which may include paying royalties for each streamed song) is typically paid for by the user subscription cost and / or the advertisements played between songs.

[0003] Song selection may be limited by licensing agreements and the number of songs written for a particular genre. Users may become tired of hearing the same songs in a particular genre. Additionally, these services may not be able to adapt the music based on a user's taste, environment, behavior, etc. Brief Description of the Drawings

[0004] Figure 1 is a schematic diagram showing an exemplary music generator.

[0005] Figure 2 is a block diagram showing an exemplary overview of a system for generating output music content based on inputs from multiple different sources according to some embodiments.

[0006] Figure 3 is a block diagram showing an exemplary music generator system configured to output music content based on an analysis of an image representation of an audio file according to some embodiments.

[0007] Figure 4 depicts an example of an image representation of an audio file.

[0008] Figure 5A and Figure 5B respectively depict examples of grayscale images for a melody image feature representation and a drum beat image feature representation.

[0009] Figure 6 is a block diagram showing an exemplary system configured to generate a single image representation according to some embodiments.

[0010] Figure 7 depicts an example of a single image representation of multiple audio files.

[0011] Figure 8is a block diagram showing an exemplary system configured to implement user-created controls in music content generation.

[0012] Figure 9 Depicts a flowchart of a method for training a music generator module based on user-created control elements according to some embodiments.

[0013] Figure 10 is a block diagram showing an exemplary teacher / student framework system according to some embodiments.

[0014] Figure 11 is a block diagram showing an exemplary system configured to implement audio technology in music content generation according to some embodiments.

[0015] Figure 12 Depicts an example of an audio signal diagram.

[0016] Figure 13 Depicts an example of an audio signal diagram.

[0017] Figure 14 Depicts an exemplary system for implementing real-time modification of music content using an audio technology music generator module according to some embodiments.

[0018] Figure 15 Depicts a block diagram of an exemplary API module for audio parameter automation in a system according to some embodiments.

[0019] Figure 16 Depicts a block diagram of an exemplary storage area according to some embodiments.

[0020] Figure 17 Depicts a block diagram of an exemplary system for storing new music content according to some embodiments.

[0021] Figure 18 is a schematic diagram showing example playback data according to some embodiments.

[0022] Figure 19 is a block diagram showing an example creation system according to some embodiments.

[0023] Figures 20A - 20B is a block diagram showing a graphical user interface according to some embodiments.

[0024] Figure 21 is a block diagram showing an example music generator system including an analysis and creation module according to some embodiments.

[0025] Figure 22 is a schematic diagram showing an example enhancement section of music content according to some embodiments.

[0026] Figure 23 is a schematic diagram showing an example technique for arranging chapters of music content according to some embodiments.

[0027] Figure 24 is a flowchart method for using a ledger according to some embodiments.

[0028] Figure 25 is a flowchart method for combining audio files using an image representation according to some embodiments.

[0029] Figure 26 is a flowchart method for implementing user-created control elements according to some embodiments.

[0030] Figure 27 is a flowchart method for generating music content by modifying audio parameters according to some embodiments.

[0031] Although the embodiments disclosed herein admit of various modifications and alternative forms, specific embodiments are shown by way of example in the drawings and described in detail herein. However, it should be understood that the drawings and the detailed description thereof are not intended to limit the scope of the claims to the particular forms disclosed. On the contrary, this application is intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the disclosure of this application as defined by the appended claims.

[0032] This disclosure includes references to "one embodiment", "a particular embodiment", "some embodiments", "various embodiments", or "an embodiment". The appearances of the phrases "in one embodiment", "in a particular embodiment", "in some embodiments", "in various embodiments", or "in an embodiment" do not necessarily refer to the same embodiment. Specific features, structures, or characteristics may be combined in any suitable manner consistent with this disclosure.

[0033] The recitation in the appended claims that an element is "configured to" perform one or more tasks is expressly not intended to invoke 35 U.S.C. § 112(f) for that claim element. Accordingly, none of the claims filed in this application are intended to be construed as having "means-plus-function" elements. If the applicant wishes to invoke Section 112(f) during the patent application process, it will use the "means for [performing the function]" structure to recite the claim element.

[0034] As used herein, the term "based on" is used to describe one or more factors that influence a determination. This term does not exclude the possibility of additional factors that may influence the determination. That is, the determination may be based solely on the specified factors or on the specified factors and other unspecified factors. Consider the phrase "determine A based on B". This phrase specifies that B is a factor used to determine A or that influences the determination of A. This phrase does not exclude the possibility that the determination of A may also be based on certain other factors, such as C. This phrase is also intended to cover embodiments where A is determined based solely on B. As used herein, the phrase "based on" is synonymous with the phrase "at least partially based on".

[0035] As used herein, the phrase "in response to" describes one or more factors that trigger an effect. This phrase does not exclude the possibility of additional factors that may influence or otherwise trigger the effect. That is, the effect may be in response solely to those factors, or it may be in response to the specified factors and other unspecified factors.

[0036] As used herein, the terms "first", "second", etc. are used as labels for the nouns that precede them and do not imply any type of ordering (e.g., spatial, temporal, logical, etc.) unless otherwise stated. As used herein, the term "or" is used as an inclusive or rather than an exclusive or. For example, the phrase "at least one of x, y, or z" means any one of x, y, and z, as well as any combination thereof (e.g., x and y, but not z). In certain cases, the context in which the term "or" is used may indicate that it is being used in an exclusive sense, such as where "select one of x, y, or z" means that only one of x, y, and z is selected in that example.

[0037] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the disclosed embodiments. However, those of ordinary skill in the art will recognize that aspects of the disclosed embodiments may be practiced without these specific details. In some instances, well-known structures, computer program instructions, and techniques have not been shown in detail in order to avoid obscuring the disclosed embodiments. Detailed Description

[0038] U.S. Patent Application No. 13 / 969,372, filed on August 16, 2013 (now U.S. Patent No. 8,812,144), discusses techniques for generating music content based on one or more musical attributes, the entire content of which is incorporated herein by reference. For any interpretation of any perceived conflict between the definitions based on the above Application 372 and the rest of the present disclosure, the present disclosure is intended to govern. Musical attributes can be user input or can be determined based on environmental information such as ambient noise, lighting, etc. The above disclosure 372 discusses techniques for selecting stored loops and / or tracks or generating new loops / tracks and layering the selected loops / tracks to generate output music content.

[0039] U.S. Patent Application No. 16 / 420,456, filed on May 23, 2019 (now U.S. Patent No. 10,679,596), discusses techniques for generating music content, the entire content of which is incorporated herein by reference. For any interpretation of any perceived conflict between the definitions based on the above Application 456 and the rest of the present disclosure, the present disclosure is intended to govern. Music can be generated based on user input or using computer-implemented methods. The above disclosure 456 discusses various music generator embodiments.

[0040] The present disclosure generally relates to systems for generating customized music content by selecting and combining audio tracks based on various parameters. In various embodiments, machine learning algorithms (including neural networks such as deep learning neural networks) are configured to generate and customize music content for a particular user. In some embodiments, users can create their own control elements and can train a computing system to generate output music content according to the user-expected functions of the user-defined control elements. In some embodiments, playback data of music content generated by the techniques described herein can be recorded to record and track the use of various music content by different rights holders (e.g., copyright holders). The various techniques discussed below can provide more relevant customized music for different contexts, facilitate music generation based on specific sounds, allow users to better control the way music is generated, generate music that achieves one or more specific goals, generate music in real time along with other content, etc.

[0041] As used herein, the term "audio file" refers to sound information for music content. For example, the sound information can include data that describes the music content as raw audio in formats such as wav, aiff, or FLAC. Attributes of the music content can be included in the sound information. Attributes can include, for example, quantifiable music attributes such as instrument classification, pitch transcription, beat timing, meter, file length, and audio amplitude in multiple frequency segments. In some embodiments, the audio file includes sound information within a specific time interval. In various embodiments, the audio file includes a loop. As used herein, the term "loop" refers to sound information of a single instrument within a specific time interval. The various techniques discussed with reference to audio files can also be performed using loops that include a single instrument. The audio file or loop can be played in a repeating manner (e.g., a 30-second audio file can be played continuously four times to generate 2 minutes of music content), but the audio file can also be played once, e.g., without repetition.

[0042] In some embodiments, an image representation of the audio file is generated and used to generate music content. The image representation of the audio file can be generated based on the data in the audio file and the MIDI representation of the audio file. For example, the image representation can be a two-dimensional (2D) image representation of pitch and rhythm determined from the MIDI representation of the audio file. Rules (e.g., composition rules) can be applied to the image representation to select the audio files to be used to generate new music content. In various embodiments, machine learning / neural networks are implemented on the image representation to select audio files for combination to generate new music content. In some embodiments, the image representation is a compressed (e.g., lower resolution) version of the audio file. The compressed image representation can increase the speed of searching for the selected music content in the image representation.

[0043] In some embodiments, a music generator can generate new music content based on various parametric representations of the audio file. For example, an audio file typically has an audio signal that can be represented as a graph of a signal (e.g., signal amplitude, frequency, or a combination thereof) relative to time. However, the time-based representation depends on the meter of the music content. In various embodiments, a graph of the signal relative to the beat (e.g., a signal graph) is also used to represent the audio file. The signal graph is independent of the meter, allowing for meter-invariant modification of the audio parameters of the music content.

[0044] In some embodiments, a music generator allows a user to create and label user-defined controls. For example, the user can create a control, and then the music generator can train the control to affect music according to the user's preferences. In various embodiments, the user-defined controls are advanced controls, such as controls for adjusting mood, intensity, or genre. Such controls are typically subjective measurements based on an individual listener's preferences. In some embodiments, the user creates and labels controls for user-defined parameters. Then the music generator can play various music files and allow the user to modify the music according to the user-defined parameters. The music generator can learn and store the user's preferences based on the user's adjustments to the user-defined parameters. Thus, during a later playback, the user can adjust the user-defined controls for the user-defined parameters, and the music generator adjusts the music playback according to the user's preferences. In some embodiments, the music generator can also select music content according to the user preferences set by the user-defined parameters.

[0045] In some embodiments, the music content generated by the music generator includes music having various stakeholder entities (e.g., rights holders or copyright holders). In a commercial application of continuous playback of the generated music content, it may be difficult to remunerate based on the playback of a single audio track (file). Thus, in various embodiments, techniques are implemented for recording playback data of continuous music content. The recorded playback data can include information related to the playback time of individual audio tracks within the continuous music content that matches the stakeholders of each individual audio track. Additionally, techniques can be implemented to prevent tampering with the playback data information. For example, the playback data information can be stored in a publicly accessible immutable blockchain ledger.

[0046] This disclosure initially refers to Figure 1 and Figure 2 describe an example music generator module and the overall system organization with various applications. Referring to Figures 3 - 7 discusses techniques for generating music content from an image representation. Referring to Figure 8 and Figure 10 discusses techniques for implementing user-created control elements. Referring to Figures 11 - 17 discusses techniques for generating implemented audio techniques. Referring to Figures 18 - 19 discusses techniques for recording information about generated music or elements in a blockchain or other cryptographic ledger. Figures 20A - 20B An exemplary application interface is shown.

[0047] Generally speaking, the disclosed music generator includes audio files, metadata (e.g., information describing the audio files), and a grammar for combining the audio files based on the metadata. The generator can create a music experience using rules to identify audio files based on the metadata and target characteristics of the music experience. The generator can be configured to expand the set of experiences it can create by adding or modifying rules, audio files, and / or metadata. The adjustments can be performed manually (e.g., an artist adds new metadata), or the music generator can add rules / audio files / metadata as it monitors the music experience and desired goals / characteristics within a given environment. For example, listener-defined controls can be implemented to obtain user feedback on music goals or characteristics.

[0048] Overview of an Exemplary Music Generator

[0049] Figure 1 is a schematic diagram showing an exemplary music generator according to some embodiments. In the illustrated embodiment, the music generator module 160 receives various information from multiple different sources and generates output music content 140.

[0050] In the illustrated embodiment, the module 160 accesses the stored audio file(s) and the corresponding attributes 110 for the stored audio file(s), and combines the audio files to generate the output music content 140. In some embodiments, the music generator module 160 selects audio files based on their attributes and combines the audio files based on the target music attributes 130. In some embodiments, the audio files can be selected based on the environmental information 150 in combination with the target music attributes 130. In some embodiments, the environmental information 150 is indirectly used to determine the target music attributes 130. In some embodiments, the target music attributes 130 are explicitly specified by the user, e.g., by specifying a desired energy level, mood, multiple parameters, etc. For example, the listener-defined controls described herein can be implemented to specify listener preferences for use as the target music attributes. Examples of target music attributes 130 include energy, complexity, and diversity, but more specific attributes (e.g., attributes corresponding to the stored tracks) can also be specified. Generally speaking, when higher-level target music attributes are specified, the system can determine lower-level specific music attributes before generating the output music content.

[0051] Complexity can refer to the number of audio files, loops, and / or instruments included in a piece of work. Energy may or may not be related to other attributes. For example, changing the key or tempo may affect the energy. However, for a given tempo and key, the energy can be changed by adjusting the instrument type (e.g., by adding high hats or white noise), complexity, volume, etc. Diversity can refer to the amount of change in the generated music over time. Diversity can be generated for a set of static other music attributes (e.g., by selecting different tracks for a given tempo and key), or can be generated by changing the music attributes over time (e.g., by changing the tempo and key more frequently when greater diversity is desired). In some embodiments, it can be considered that the target music attributes exist in a multi-dimensional space, and the music generator module 160 can move slowly through this space, e.g., making route corrections based on environmental changes and / or user input if needed.

[0052] In some embodiments, the attributes stored with an audio file contain information about one or more audio files, including: tempo, volume, energy, genre, spectrum, envelope, modulation, periodicity, attack and decay times, noise, artist, instrument, theme, etc. Note that in some embodiments, the audio files are partitioned such that a group of one or more audio files is specific to a particular audio file type (e.g., a particular instrument or instrument type).

[0053] In the illustrated embodiment, module 160 accesses the stored rule set(s) 120. In some embodiments, the stored rule set(s) 120 specify rules for how many audio files to overlay such that they play simultaneously (which may correspond to the complexity of the output music), which major / minor key progression to use when transitioning between audio files or musical phrases, which instruments to use together (e.g., instruments that have an affinity for each other), etc., to achieve the target music attributes. In other words, the music generator module 160 uses the stored rule set(s) 120 to achieve one or more declarative goals defined by the target music attributes (and / or target environmental information). In some embodiments, the music generator module 160 includes one or more pseudo-random number generators that are configured to introduce pseudo-randomness to avoid repeating the output music.

[0054] In some embodiments, the environmental information 150 includes one or more of the following: lighting information, ambient noise, user information (facial expressions, body postures, activity levels, movements, skin temperature, performance of specific activities, clothing type, etc.), temperature information, purchase activities in the area, time of day, day of the week, time of year, number of people present, weather conditions, etc. In some embodiments, the music generator module 160 does not receive / process the environmental information. In some embodiments, the environmental information 150 is received by another module that determines the target music attributes 130 based on the environmental information. The target music attributes 130 can also be derived based on other types of content (such as video data). In some embodiments, the environmental information is used to adjust one or more stored rule sets 120, for example, to achieve one or more environmental goals. Similarly, the music generator can use the environmental information to adjust the storage attributes of one or more audio files, for example, to indicate the target music attributes or target audience characteristics that are particularly relevant to those audio files.

[0055] As used herein, the term "module" refers to a circuit configured to perform a specified operation or a physical non-transitory computer-readable medium that stores information (such as program instructions) indicating other circuits (such as a processor) to perform the specified operation. Modules can be implemented in various ways, including as hardwired circuits or as memories in which program instructions are stored and can be executed by one or more processors to perform the operations. The hardware circuits can include, for example, custom very large scale integration (VLSI) circuits or gate arrays, off-the-shelf semiconductors such as logic chips, transistors, or other discrete components. Modules can also be implemented in programmable hardware devices such as field programmable gate arrays, programmable array logic, programmable logic devices, or the like. A module can also be any suitable form of non-transitory computer-readable medium that stores program instructions executable to perform the specified operation.

[0056] As used herein, the phrase "music content" refers to both the music itself (the audible representation of the music) and the information that can be used to play the music. Thus, a song recorded as a file on a storage medium (such as, but not limited to, a compact disc, flash drive, etc.) is an example of music content; the sound produced by outputting that recorded file or other electronic representation (such as through a speaker) is also an example of music content.

[0057] The term "music" includes its well-known meaning, including the sounds produced by musical instruments as well as human voices. Thus, music includes, for example, instrumental performances or recordings, a cappella performances or recordings, and performances or recordings that include both musical instruments and voices. One of ordinary skill in the art will recognize that "music" does not include all sound recordings. Works that do not contain musical attributes such as rhythm or rhyme (such as speeches, news broadcasts, and audiobooks) are not music.

[0058] A piece of music "content" can be distinguished from another piece of music content in any suitable manner. For example, a digital file corresponding to a first song can represent the first piece of music content, while a digital file corresponding to a second song can represent the second piece of music content. The phrase "music content" can also be used to distinguish specific intervals within a given musical work, such that different parts of the same song can be regarded as different segments of music content. Similarly, different tracks within a given musical work (e.g., a piano track, a guitar track) can also correspond to different segments of music content. In the context of potentially endless generated music streams, the phrase "music content" can be used to refer to a specific part of the stream (e.g., several bars or minutes).

[0059] The music content generated by embodiments of the present disclosure can be "new music content" - a combination of musical elements that has never been generated before. A related (but broader) concept - "original music content" - will be described further below. For the sake of explaining this term, the concept of a "control entity" related to an instance of music content generation is described. Different from the phrase "original music content", the phrase "new music content" does not involve the concept of a control entity. Thus, new music content refers to music content that has never been generated by any entity or computer system before.

[0060] Conceptually, the present disclosure refers to some "entities" as controlling specific instances of computer-generated music content. Such entities possess any legal rights (e.g., copyright) that may correspond to the computer-generated content (to the extent that any such rights may actually exist). In one embodiment, the individual who creates (e.g., codes various software routines) a computer-implemented music generator or operates (e.g., provides input to) a specific instance of computer-implemented music generation will be the control entity. In other embodiments, a computer-implemented music generator can be created by a legal entity (e.g., a company or other business organization), such as in the form of a software product, a computer system, or a computing device. In some instances, such a computer-implemented music generator can be deployed to many clients. In various instances, according to the terms of the license associated with the distribution of the music generator, the control entity can be the creator, the distributor, or the client. If no such explicit legal agreement exists, the control entity of a computer-implemented music generator is the entity that facilitates (e.g., provides input and thereby operates) a specific instance of the computer generation of music content.

[0061] Within the meaning of the present disclosure, the computer generation of "original music content" by a controlling entity means 1) a combination of musical elements that have never been generated before by the controlling entity or anyone else, and 2) a combination of musical elements that have been generated before but were originally generated by the controlling entity. Content type 1) is referred to herein as "novel music content", which is similar to the definition of "new music content", except that the definition of "novel music content" refers to the concept of a "controlling entity", while the definition of "new music content" does not. On the other hand, content type 2) is referred to herein as "proprietary music content". Note that the term "proprietary" in context does not refer to any implicit legal rights in the content (although such rights may exist), but is only used to indicate that the music content was originally generated by the controlling entity. Thus, the "regeneration" by a controlling entity of music content that was previously and originally generated by the controlling entity constitutes the "generation of original music content" in the present disclosure. "Non-original music content" for a particular controlling entity is music content that is not the "original music content" of that controlling entity.

[0062] Some musical content segments can include musical components from one or more other musical content segments. Creating musical content in this way is referred to as "sampling" musical content and is common in certain musical works, especially in certain musical genres. Such musical content is referred to herein as "musical content with sampled components", "derivative musical content", or using other similar terms. In contrast, musical content that does not include sampled components is referred to herein as "musical content without sampled components", "non-derivative musical content", or using other similar terms.

[0063] When applying these terms, it should be noted that if any particular musical content is reduced to a fine enough granularity level, it could be argued that the musical content is derivative (which effectively means that all musical content is derivative). In the present disclosure, the terms "derivative" and "non-derivative" are not used in this sense. Regarding the computer generation of musical content, if the computer generation selects components from pre-existing musical content from an entity other than the controlling entity (e.g., a computer program selects a specific portion of an audio file of a popular artist's work to include in a piece of musical content being generated), then such computer generation is considered derivative (and produces derivative musical content). On the other hand, if the computer generation does not utilize such components of such pre-existing content, then the computer generation of musical content is considered non-derivative (and produces non-derivative musical content). Note that some segments of "original music content" may be derivative musical content, while some segments may be non-derivative musical content.

[0064] Note that the term "derivative" is intended in this disclosure to have a broader meaning than the term "derivative work" as used in U.S. copyright law. For example, under U.S. copyright law, derivative musical content may or may not be a derivative work. The term "derivative" in this disclosure is not intended to convey a negative connotation; it is merely used to indicate whether a particular piece of musical content "borrows" portions of another work.

[0065] In addition, the phrases "new musical content," "novel musical content," and "original musical content" are not intended to encompass musical content that differs only slightly from a combination of pre-existing musical elements. For example, merely changing a few notes of a pre-existing musical work does not result in new, novel, or original musical content, as those phrases are used in this disclosure. Similarly, merely changing the key or tempo of a pre-existing musical work or adjusting the relative intensities of frequencies (e.g., using an equalizer interface) does not result in new, novel, or original musical content. Further, the phrases new, novel, and original musical content are not intended to cover musical content that lies on the boundary between original and non-original content; rather, these terms are intended to cover musical content that is unquestionably and demonstrably original, including musical content that is eligible for copyright protection by a controlling entity (referred to herein as "protected" musical content). Additionally, as used herein, the term "usable" musical content refers to musical content that does not infringe the copyrights of any entity other than the controlling entity. New and / or original musical content is often protected and usable. This can be advantageous in preventing the copying of musical content and / or paying royalties for musical content.

[0066] Although the various embodiments discussed herein use a rule-based engine, various other types of computer-implemented algorithms can be used for any of the computer learning and / or music generation techniques discussed herein. However, a rule-based approach may be particularly effective in a musical context.

[0067] Overview of Applications, Storage Elements, and Data That Can Be Used in an Exemplary Music System

[0068] The music generator module can interact with multiple different applications, modules, storage elements, etc. to generate music content. For example, an end user can install one of multiple types of applications for different types of computing devices (e.g., mobile devices, desktop computers, DJ devices, etc.). Similarly, another type of application can be provided to enterprise users. Interacting with an application while generating music content can allow the music generator to receive external information, which can be used to determine target music attributes and / or update one or more rule sets for generating music content. In addition to interacting with one or more applications, the music generator module can also interact with other modules to receive rule sets, update rule sets, etc. Finally, the music generator module can access one or more rule sets, audio files, and / or generated music content stored in one or more storage elements. Additionally, the music generator module can store any of the items listed above in one or more storage elements, which can be local or accessed via a network (e.g., cloud-based).

[0069] Figure 2 is a block diagram showing an exemplary overview of a system for generating output music content based on inputs from multiple different sources. In the illustrated embodiment, system 200 includes a rule module 210, user application 220, web application 230, enterprise application 240, artist application 250, artist rule generator module 260, storage of generated music 270, and external input 280.

[0070] In the illustrated embodiment, user application 220, web application 230, and enterprise application 240 receive external input 280. In some embodiments, external input 280 includes: environmental input, target music attributes, user input, sensor input, etc. In some embodiments, user application 220 is installed on the user's mobile device and includes a graphical user interface (GUI) that allows the user to interact / communicate with rule module 210. In some embodiments, web application 230 is not installed on the user device but is configured to run within the user device's browser and can be accessed via a website. In some embodiments, enterprise application 240 is an application used by a larger-scale entity to interact with the music generator. In some embodiments, application 240 is used in combination with user application 220 and / or web application 230. In some embodiments, application 240 communicates with one or more external hardware devices and / or sensors to collect information about the surrounding environment.

[0071] In the illustrated embodiment, the rules module 210 communicates with user applications 220, web applications 230, and enterprise applications 240 to generate output music content. In some embodiments, the music generator 160 is included within the rules module 210. Note that the rules module 210 may be included within one of the applications 220, 230, and 240, or may be installed on a server and accessed via a network. In some embodiments, the applications 220, 230, and 240 receive the generated output music content from the rules module 210 and cause the content to be played. In some embodiments, the rules module 210 requests inputs from the applications 220, 230, and 240, such as regarding target music attributes and environmental information, and may use this data to generate music content.

[0072] In the illustrated embodiment, the stored rule set(s) 120 are accessed by the rules module 210. In some embodiments, the rules module 210 modifies and / or updates the stored rule set(s) 120 based on communication with the applications 220, 230, and 240. In some embodiments, the rules module 210 accesses the stored rule set(s) 120 to generate output music content. In the illustrated embodiment, the stored rule set(s) 120 may include rules from the artist rule generator module 260, which will be discussed in further detail below.

[0073] In the illustrated embodiment, the artist application 250 communicates with the artist rule generator module 260 (e.g., which may be part of the same application or may be cloud-based). In some embodiments, the artist application 250 allows artists to create rule sets for their particular sound, e.g., based on previous works. This functionality is further discussed in U.S. Patent No. 10,679,596. In some embodiments, the artist rule generator module 260 is configured to store the generated artist rule sets for use by the rules module 210. A user may purchase a rule set from a particular artist before using the particular artist to generate output music via their particular application. A rule set for a particular artist may be referred to as a signature pack.

[0074] In the illustrated embodiment, the stored audio file(s) and corresponding attribute(s) 110 are accessed by the module 210 when applying rules to select and combine tracks to generate output music content. In the illustrated embodiment, the rules module 210 stores the generated output music content 270 in a storage element.

[0075] In some embodiments, implemented on a server and accessed via a network Figure 2One or more components, which may be referred to as cloud-based implementations. For example, the stored rule set(s) 120, audio file(s) / attribute(s) 110, and generated music 270 can all be stored on the cloud and accessed by module 210. In another example, module 210 and / or module 260 can also be implemented in the cloud. In some embodiments, the generated music 270 is stored in the cloud and watermarked digitally. For example, this can allow detection of copied generated music as well as generation of a large amount of customized music content.

[0076] In some embodiments, one or more of the disclosed modules are configured to generate other types of content in addition to music content. For example, the system can be configured to generate output visual content based on target music attributes, determined environmental conditions, currently used rule sets, etc. As another example, the system can search a database or the Internet based on the current attributes of the music being generated and display an image collage that changes dynamically as the music changes and matches the attributes of the music.

[0077] Exemplary Machine Learning Method

[0078] As described herein, Figure 1 The illustrated music generator module 160 can implement a variety of artificial intelligence (AI) techniques (e.g., machine learning techniques) to generate output music content 140. In various embodiments, the implemented AI techniques include a combination of deep neural networks (DNN) with more traditional machine learning techniques and knowledge-based systems. This combination can combine the respective advantages and disadvantages of these techniques with the challenges inherent in music works and personalized systems. Music content has multiple levels of structure. For example, a song has sections, phrases, melodies, notes, and textures. DNN can effectively analyze and generate music content at very high levels and very low levels of detail. For example, DNN can be good at classifying the texture of a sound as belonging to a low-level clarinet or electric guitar, or detecting high-level verses and choruses. Intermediate levels of music content detail, such as the construction of melodies, orchestration, etc., may be more difficult. DNN is generally good at collecting a wide range of styles in a single model, and thus, DNN can be implemented as a generation tool with a large expressive range.

[0079] In some embodiments, the music generator module 160 leverages expert knowledge by using human-written audio files (e.g., loops) as the basic unit of the music content used by the music generator module. For example, the social context of the expert knowledge can be embedded through the selection of rhythm, melody, and texture to record heuristics in a multi-level structure. Different from the separation of DNN and traditional machine learning at the structural level, expert knowledge can be applied to any area that can enhance musicality without imposing overly strong restrictions on the trainability of the music generator module 160.

[0080] In some embodiments, the music generator module 160 uses a DNN to find patterns in how audio layers are combined, combining them vertically by stacking sounds on top of each other and horizontally by combining audio files or loops into sequences. For example, the music generator module 160 can implement a long short-term memory (LSTM) recurrent neural network that is trained on the Mel Frequency Cepstral Coefficient (MFCC) audio features of the loops used in multi-track audio recordings. In some embodiments, the network is trained to predict and select the audio features of the loops for upcoming beats based on knowledge of the audio features of previous beats. For example, the network can be trained to predict the audio features of the loops for the next 8 beats based on knowledge of the audio features of the last 128 beats. Thus, the network is trained to use low-dimensional feature representations to predict upcoming beats.

[0081] In a particular embodiment, the music generator module 160 uses known machine learning algorithms to assemble a sequence of multi-track audio into a musical structure with dynamic features of intensity and complexity. For example, the music generator module 160 can implement a hierarchical hidden Markov model that can perform state transitions like a state machine, with its probabilities determined by a multi-level hierarchical structure. For example, a particular kind of drop may be more likely to occur after a build-up section, but less likely to occur if there is no drum at the end of the build-up. In various embodiments, the probabilities can be transparently trained, in contrast to the more opaque training of DNNs where what is being learned is less transparent.

[0082] The Markov model can handle larger time structures and may thus not be easily trained by presenting example tracks because the examples may be too long. Feedback control elements (e.g., thumbs up / down on the user interface) can be used to provide feedback on the music at any time. In a particular embodiment, the feedback control element is implemented as one of the UI control elements 830, as Figure 8As shown. Then, the correlation between the music structure and the feedback can be used to update the structure model for composition, such as a transition table or a Markov model. This feedback can also be directly collected from measurements of heart rate, sales, or any other metric that the system can determine a clear classification for. As mentioned above, the expert knowledge heuristic is also designed to be as probabilistic as possible and is trained in the same way as a Markov model.

[0083] In a particular embodiment, the training can be performed by a composer or a DJ. Such training can be separate from the listener training. For example, the training performed by a listener (e.g., a typical user) may be limited to identifying correct or incorrect classifications based on positive and negative model feedback respectively. For composers and DJs, the training may include hundreds of time steps and include details about the layers used and volume control to provide more explicit details on the factors driving changes in the music content. For example, the training performed by composers and DJs can include sequence prediction training similar to the global training of the DNN described above.

[0084] In various embodiments, the DNN is trained with multi-track audio and interface interactions to predict what a DJ or composer will do next. In some embodiments, these interactions can be recorded and used to develop more transparent new heuristics. In some embodiments, the DNN receives many previous music metrics as input and utilizes the low-dimensional feature representation as described above, as well as additional features that describe the modifications applied to the tracks by the DJ or composer. For example, the DNN can receive the last 32 bars of music as input and utilize the low-dimensional feature representation and additional features to describe the modifications applied by the DJ or composer to the track. These modifications can include adjusting the gain of a particular track, the filters applied, the delay, etc. For example, a DJ may repeat the same drum loop for five minutes during a performance but may gradually increase the gain and delay on the track over time. Thus, in addition to loop selection, the DNN can be trained to predict such gain and delay changes. When no loop is being played for a particular instrument (e.g., no drum loop is being played), the feature set for that instrument may be all zeros, which can allow the DNN to learn that predicting all zeros may be a successful strategy, which may lead to selective layering.

[0085] In some instances, a DJ or composer records a live performance using a mixer and a device such as TRAKTOR (Native Instruments GmbH). These recordings are typically captured at a high resolution (e.g., 4-channel audio or MIDI). In some embodiments, the system breaks the recording into its constituent loops, thereby generating information about the combination of loops in the piece and the sound quality of each individual loop. Training a DNN (or other machine learning) with this information provides the DNN with the ability to correlate the piece (e.g., the sequencing, layering, timing of the loops, etc.) and the sound quality of the loops to inform the music generator module 160 how to create a music experience similar to that of an artist's performance without using the actual loops used by the artist in the performance.

[0086] Exemplary Music Generator Using an Image Representation of an Audio File

[0087] Popular music often has combinations of rhythm, texture, and pitch that are widely observed. When creating music note by note for each instrument in a piece (as can be done by a music generator), rules can be implemented based on these combinations to create coherent music. Generally, the more restrictive the rules, the less room there is for variation, and thus the more likely it is to create a copy of existing music.

[0088] When creating music by combining musical phrases that have been played and recorded as audio, it may be necessary to consider multiple unchangeable note combinations within each phrase to create the combination. However, when extracting from a library with thousands of recordings, searching for every possible combination can be computationally expensive. Additionally, note-by-note comparisons may be required to check for combinations of harmonic dissonance, especially on the beat. New rhythms created by combining multiple files can also be checked against the rhythmic composition rules for the combined phrases.

[0089] It may not always be possible to extract the necessary features from audio files for combination. Even when possible, extracting the required features from audio files can be computationally expensive. In various embodiments, symbolic audio representations are used for music composition to reduce computational overhead. Symbolic audio representations may rely on a music composer's memory of instrument textures as well as stored rhythm and pitch information. A common symbolic music representation format is MIDI. MIDI contains precise timing, pitch, and performance control information. In some embodiments, MIDI can be further simplified and compressed by a piano roll representation, where notes are shown as bars on a discrete time / pitch graph, typically with 8 octaves.

[0090] In some embodiments, the music generator is configured to generate output music content by generating an image representation of an audio file and selecting a combination of music based on an analysis of the image representation. The image representation may be a further compressed representation from a piano roll representation. For example, the image representation may be a low-resolution representation generated based on the MIDI representation of the audio file. In various embodiments, compositional rules are applied to the image representation to select music content from the audio file for combination and generation of the output music content. For example, a rule-based method may be used to apply the compositional rules. In some embodiments, a machine learning algorithm or model (such as a deep learning neural network) is implemented to select and combine audio files for generating the output music content.

[0091] Figure 3 FIG. 4 is a block diagram illustrating an exemplary music generator system configured to output music content based on an analysis of an image representation of an audio file. In the illustrated embodiment, system 300 includes an image representation generation module 310, a music selection module 320, and a music generator module 160.

[0092] In the illustrated embodiment, the image representation generation module 310 is configured to generate one or more image representations of an audio file. In a particular embodiment, the image representation generation module 310 receives audio file data 312 and MIDI representation data 314. The MIDI representation data 314 includes the MIDI representation(s) of the specified audio file(s) in the audio file data 312. For example, for a specified audio file in the audio file data 312, there may be a corresponding MIDI representation in the MIDI representation data 314. In some embodiments where there are multiple audio files in the audio file data 312, each audio file in the audio file data 312 has a corresponding MIDI representation in the MIDI representation data 314. In the illustrated embodiment, the MIDI representation data 314 is provided to the image representation generation module 310 together with the audio file data 312. However, in some contemplated embodiments, the image representation generation module 310 may itself generate the MIDI representation data 314 from the audio file data 312.

[0093] As Figure 3As shown, the image representation generation module 310 generates one or more image representations 316 from the audio file data 312 and the MIDI representation data 314. The MIDI representation data 314 may include pitch, time (or rhythm), and velocity (or note intensity) data of the notes in the music associated with the audio file, while the audio file data 312 includes data for playing back the music itself. In a particular embodiment, the image representation generation module 310 generates an image representation for the audio file based on the pitch, time, and velocity data from the MIDI representation data 314. The image representation may be, for example, a two-dimensional (2D) image representation of the audio file. In the 2D image representation of the audio file, the x-axis represents time (rhythm), the y-axis represents pitch (similar to a piano roll representation), and the pixel value at each xy coordinate represents velocity.

[0094] The 2D image representation of the audio file can have various image sizes, although the image size is typically chosen to correspond to the music structure. For example, in one envisioned embodiment, the 2D image representation is 32 (x-axis) × 24 images (y-axis). An image representation that is 32 pixels wide allows each pixel to represent a quarter beat in the time dimension. Thus, an 8-beat piece of music can be represented by an image representation that is 32 pixels wide. Although this representation may not have enough detail to capture the performance details of the music in the audio file, the performance details are retained in the audio file itself, which is combined with the image representation of the system 300 for generating the output music content. However, the quarter-beat time resolution does allow for a large coverage of common pitch and rhythm combination rules.

[0095] Figure 4 An example of the image representation 316 of the audio file is depicted. The image representation 316 is 32 pixels wide (time) and 24 pixels high (pitch). Each pixel (square) 402 has a value representing the velocity at that time and the pitch in the audio file. In various embodiments, the image representation 316 may be a grayscale image representation of the audio file, where the pixel values are represented by varying grayscale intensities. The gray variation based on the pixel values may be small and imperceptible to many people. Figure 5A and Figure 5B Examples of grayscale images for the melody image feature representation and the drum beat image feature representation are depicted, respectively. However, other representations (e.g., color or digital) may also be considered. In these representations, each pixel may have multiple different values corresponding to different music attributes.

[0096] In a particular embodiment, the image representation 316 is an 8-bit representation of an audio file. Thus, each pixel can have 256 possible values. A MIDI representation typically has 128 possible velocity values. In various embodiments, the details of the velocity values may be less important than the task of selecting the audio files for combination. Thus, in such embodiments, the pitch axis (y-axis) can be banded into two sets of four octaves within an 8-octave range. For example, the 8 octaves can be defined as follows:

[0097] Octave 0: rows 0 - 11, values 0 - 63;

[0098] Octave 1: rows 12 - 23, values 0 - 63;

[0099] Octave 2: rows 0 - 11, values 64 - 127;

[0100] Octave 3: rows 12 - 23, values 64 - 127;

[0101] Octave 4: rows 0 - 11, values 128 - 191;

[0102] Octave 5: rows 12 - 23, values 128 - 191

[0103] Octave 6: rows 0 - 11, values 192 - 255; and

[0104] Octave 7: rows 12 - 23, values 192 - 255.

[0105] With these defined octave ranges, the row and value of a pixel determine the octave and velocity of a note. For example, a pixel value of 10 in row 1 represents a note in octave 0 with a velocity of 10, while a pixel value of 74 in row 1 represents a note in octave 2 with a velocity of 10. As another example, a pixel value of 79 in row 13 represents a note in octave 3 with a velocity of 15, while a pixel value of 207 in row 13 represents a note in octave 7 with a velocity of 15. Thus, using the defined ranges of the above octaves, the first 12 rows (rows 0 - 11) represent the first set of 4 octaves (octaves 0, 2, 4, and 6), and the pixel value determines which of the first 4 octaves is represented (the pixel value also determines the velocity of the note). Similarly, the second 12 rows (rows 12 - 23) represent the second set of 4 octaves (octaves 1, 3, 5, and 7), and the pixel value determines which of the second 4 octaves is represented (the pixel value also determines the velocity of the note).

[0106] As described above, by banding the pitch axis to cover an eight octave range, the velocity for each octave can be defined by 64 values instead of the 128 values represented by MIDI. Thus, a 2D image representation (e.g., image representation 316) can be compressed (e.g., have a lower resolution) compared to a MIDI representation of the same audio file. In some embodiments, further compression of the image representation may be allowed because 64 values may be more than the system 300 needs to select musical combinations. For example, the velocity resolution can be further reduced to allow compression in the time representation by having odd pixel values represent note onsets and even pixel values represent note durations. Reducing the resolution in this way allows two notes played in quick succession at the same velocity to be distinguished from a single longer note based on odd or even pixel values.

[0107] As described above, the compactness of the image representation reduces the file size required to represent music (e.g., compared to a MIDI representation). Thus, implementing an image representation of an audio file reduces the disk storage required. In addition, a compressed image representation can be stored in a high-speed memory that allows for quick searching of possible musical combinations. For example, an 8-bit image representation can be stored in the graphics memory of a computer device, allowing for large parallel searches to be implemented together.

[0108] In various embodiments, image representations generated for multiple audio files are combined into a single image representation. For example, the image representations of dozens, hundreds, or thousands of audio files can be combined into a single image representation. The single image representation can be a large, searchable image that can be used to parallel search the multiple audio files that make up the single image. For example, software such as MegaTextures (from id Software) can be used to search the single image in a manner similar to large textures in a video game.

[0109] Figure 6 is a block diagram showing an exemplary system configured to generate a single image representation. In the illustrated embodiment, system 600 includes a single image representation generation module 610 and a texture feature extraction module 620. In a particular embodiment, the single image representation generation module 610 and the texture feature extraction module 620 are located within the image representation generation module 310, as Figure 3 shown. However, the single image representation generation module 610 or the texture feature extraction module 620 can be located outside of the image representation generation module 310.

[0110] As Figure 6As shown in the illustrated embodiment, a plurality of image representations 316A-N are generated. The image representations 316A-N can be N individual image representations of N individual audio files. A single image representation generation module 610 can combine the individual image representations 316A-N into a single combined image representation 316. In some embodiments, the individual image representations combined by the single image representation generation module 610 include individual image representations for different musical instruments. For example, different musical instruments in an orchestra can be represented by individual image representations, which are then combined into a single image representation for searching and selecting music.

[0111] In a particular embodiment, the individual image representations 316A-N are combined into a single image representation 316, where the individual image representations are placed adjacent to each other without overlap. Thus, the single image representation 316 is a complete data set representation of all the individual image representations 316A-N without data loss (e.g., no data from one image representation is modified for use in the data of another image representation). Figure 7 An example of a single image representation 316 of a plurality of audio files is depicted. In the illustrated embodiment, the single image representation 316 is a combined image generated from individual image representations 316A, 316B, 316C, and 316D.

[0112] In some embodiments, a texture feature 622 is appended to the single image representation 316. In the illustrated embodiment, the texture feature 622 is appended as a single row to the single image representation 316. Returning to Figure 6 , the texture feature 622 is determined by a texture feature extraction module 620. The texture feature 622 can include, for example, the instrumental texture of the music in the audio file. For example, the texture feature can include features from different musical instruments (such as drums, string instruments, etc.).

[0113] In a particular embodiment, the texture feature extraction module 620 extracts the texture feature 622 from the audio file data 312. The texture feature extraction module 620 can implement, for example, a rule-based method, a machine learning algorithm or model, a neural network, or other feature extraction techniques to determine the texture feature from the audio file data 312. In some embodiments, the texture feature extraction module 620 can extract the texture feature 622 from one or more image representations 316 (e.g., multiple image representations or a single image representation). For example, the texture feature extraction module 620 can implement image-based analysis (such as an image-based machine learning algorithm or model) to extract the texture feature 622 from one or more image representations 316.

[0114] Adding texture features 622 to a single image representation 316 provides additional information that is not typically available in the MIDI representation or piano roll representation of an audio file. In some embodiments, the rows in the single image representation 316( Figure 7 shown) that have the texture features 622 may not need to be human-readable. For example, the texture features 622 may only need to be machine-readable for implementation in a music generation system. In a particular embodiment, the texture features 622 are appended to the single image representation 316 for image-based analysis of the single image representation. For example, the texture features 622 can be used by an image-based machine learning algorithm or model used in music selection, as described below. In some embodiments, the texture features 622 can be ignored during music selection, such as in rule-based selection, as described below.

[0115] Returning to Figure 3 , in the illustrated embodiment, the (one or more) image representations 316 (e.g., multiple image representations or a single image representation) are provided to the music selection module 320. The music selection module 320 can select an audio file or a portion of an audio file for combination in the music generator module 160. In a particular embodiment, the music selection module 320 applies a rule-based method to search for and select an audio file or a portion of an audio file for combination by the music generator module 160. As Figure 3 shown, the music selection module 320 accesses the rules of the rule-based method from the stored (one or more) rule sets 120. For example, the rules accessed by the music selection module 320 can include rules for search and selection, such as but not limited to composition rules and note combination rules. Applying the rules to the (one or more) image representations 316 can be implemented using the graphics processing available on a computer device.

[0116] For example, in various embodiments, the note combination rules can be represented as vector and matrix calculations. Graphics processing units are typically optimized for vector and matrix calculations. For example, notes that are one pitch apart are typically dissonant and are often avoided. Notes such as these can be found by rule-based searching of adjacent pixels in an overlayed hierarchical image (or a fragment of a large image). Thus, in various embodiments, the disclosed modules can call a kernel to perform all or part of the disclosed operations on the graphics processor of a computing device.

[0117] In some embodiments, the pitch grading in the image representation as described above allows for the use of graphics processing to implement high-pass or low-pass filtering of audio. Removing (e.g., filtering out) pixel values below a threshold can simulate high-pass filtering, while removing pixel values above a threshold can simulate low-pass filtering. For example, filtering out (removing) pixel values below 64 in the grading example above may have a similar effect to applying a high-pass filter with a shelf at B1 by removing octaves 0 and 1 in the example. Thus, filtering can be effectively simulated for each audio file by applying rules to the image representation of the audio file.

[0118] In various embodiments, when audio files are layered together to create music, the pitch of the specified audio files can be changed. Changing the pitch may open up a larger range of possible successful combinations and combination search spaces. For example, each audio file can be tested in 12 different pitch transposition modes. Shifting the row order in the image representation when parsing the image and adjusting the octave transposition as necessary can allow for an optimized search through these combinations.

[0119] In a particular embodiment, the music selection module 320 implements a machine learning algorithm or model on the (one or more) image representations 316 to search for and select audio files or portions of audio files for combination by the music generator module 160. The machine learning algorithm / model can include, for example, a deep learning neural network or other machine learning algorithms that classify images based on algorithm-based training. In such embodiments, the music selection module 320 includes one or more machine learning models that are trained based on combinations and sequences of audio files that provide the desired music attributes.

[0120] In some embodiments, the music selection module 320 includes a machine learning model that continuously learns during the selection of the output music content. For example, the machine learning model can receive user input or other input that reflects the attributes of the output music content, which can be used to adjust the classification parameters implemented by the machine learning model. Similar to the rule-based method, the machine learning model can be implemented using the graphics processing unit on the computer device.

[0121] In some embodiments, the music selection module 320 implements a combination of a rule-based method and a machine learning model. In one contemplated embodiment, the machine learning model is trained to find combinations of audio files and image representations for initiating a search for music content for combination, where the search is implemented using a rule-based method. In some embodiments, the music selection module 320 tests the coherence of harmony and rhythm rules in the music selected by the music generator module 160 for combination. For example, the music selection module 320 can test the harmony and rhythm in the selected audio file 322 before providing the selected audio file to the music generator module 160, as described below.

[0122] In Figure 3 the illustrated embodiment, as described above, the music selected by the music selection module 320 is provided to the music generator module 160 as the selected audio file 322. The selected audio file 322 can include a complete or partial audio file that is combined by the music generator module 160 to generate the output music content 140, as described herein. In some embodiments, the music generator module 160 accesses the stored rule set(s) 120 to retrieve the rules applied to the selected audio file 322 for generating the output music content 140. The rules retrieved by the music generator module 160 may be different from the rules applied by the music selection module 320.

[0123] In some embodiments, the selected audio file 322 includes information for combining the selected audio file. For example, the machine learning model implemented by the music selection module 320 can provide instructions to the output that describe how to combine music content in addition to the selection of the music to be combined. These instructions can then be provided to the music generator module 160 and implemented by the music generator module for combining the selected audio file. In some embodiments, the music generator module 160 tests the coherence of harmony and rhythm rules before finalizing the output music content 140. Such tests can supplement or replace the tests implemented by the music selection module 320.

[0124] Exemplary Controls for Music Content Generation

[0125] In various embodiments, as described herein, a music generator system is configured to automatically generate output music content by selecting and combining audio tracks based on various parameters. As described herein, a machine learning model (or other AI techniques) is used to generate music content. In some embodiments, AI techniques are implemented to customize music content for a particular user. For example, the music generator system may implement various types of adaptive controls for personalized music generation. In addition to generating content via AI techniques, personalized music generation allows a composer or listener to control the content. In some embodiments, a user creates their own control elements, and the music generator system can be trained (e.g., using AI techniques) to generate output music content based on the intended functionality of the user-created control elements for the user. For example, a user can create a control element, and then the music generator system trains the control element to affect the music according to the user's preferences.

[0126] In various embodiments, the user-created control elements are high-level controls, such as controls for adjusting mood, intensity, or genre. Such user-created control elements are typically subjective measures based on the individual preferences of the listener. In some embodiments, the user tags the user-created control elements to define user-specified parameters. The music generator system can play various music content and allow the user to modify the user-specified parameters in the music content using the control elements. The music generator system can learn and store the way in which the user-defined parameters change the audio parameters in the music content. Thus, during a later playback, the user can adjust the user-created control elements, and the music generation system adjusts the audio parameters in the music playback according to the adjustment level of the user-specified parameters. In some envisioned embodiments, the music generator system can also select music content according to the user preferences set by the user-specified parameters.

[0127] Figure 8 is a block diagram showing an exemplary system configured to implement user-created controls in music content generation. In the illustrated embodiment, system 800 includes a music generator module 160 and a user interface (UI) module 820. In various embodiments, the music generator module 160 implements the techniques described herein for generating output music content 140. For example, the music generator module 160 can access stored audio file(s) 810 and generate output music content 140 based on the stored rule set(s) 120.

[0128] In various embodiments, the music generator module 160 modifies music content based on inputs from one or more UI control elements 830 implemented in the UI module 820. For example, during interaction with the UI module 820, the user can adjust the level of the (one or more) control elements 830. Examples of control elements include, but are not limited to, sliders, dials, buttons, or knobs. The level of the (one or more) control elements 830 then sets the (one or more) control element levels 832, which are provided to the music generator module 160. The music generator module 160 can then modify the output music content 140 based on the (one or more) control element levels 832. For example, the music generator module 160 can implement AI techniques to modify the output music content 140 based on the (one or more) control element levels 830.

[0129] In a particular embodiment, one or more control elements 830 are user-defined control elements. For example, the control elements can be defined by a composer or a listener. In such embodiments, the user can create and label a UI control element that specifies the parameter that the user wants to implement to control the output music content 140 (e.g., the user creates a control element for controlling a user-specified parameter in controlling the output music content 140).

[0130] In various embodiments, the music generator module 160 can learn or be trained to affect the output music content 140 in a specified manner based on inputs from user-created control elements. In some embodiments, the music generator module 160 is trained to modify audio parameters in the output music content 140 based on the levels of user-created control elements set by the user. Training the music generator module 160 can include, for example, determining the relationship between the audio parameters in the output music content 140 and the levels of the user-created control elements. Then, the music generator module 160 can utilize the relationship between the audio parameters in the output music content 140 and the levels of the user-created control elements to modify the output music content 140 based on the input levels of the user-created control elements.

[0131] Figure 9 A flowchart depicting a method for training the music generator module 160 based on user-created control elements according to some embodiments is shown. Method 900 begins with the user creating and labeling a control element in 910. For example, as described above, the user can create and label a UI control element for controlling a user-specified parameter in the output music content 140 generated by the music generator module 160. In various embodiments, the label of the UI control element describes the user-specified parameter. For example, the user can label the control element as "attitude" to specify that the user wants to control the attitude (defined by the user) in the generated music content.

[0132] After creating the UI control element, method 900 continues with playback session 915. The playback session 915 can be used to train the system (e.g., the music generator module 160) on how to modify audio parameters based on the level of the UI control element created by the user. During playback session 915, at 920, a sound track is played. The sound track can be a loop or sample of music from an audio file stored on or accessed by the device.

[0133] At 930, the user provides input regarding his / her interpretation of a user-specified parameter in the currently playing sound track. For example, in a particular embodiment, the user is asked to listen to the sound track and select the level of the user-specified parameter that the user believes describes the music in the sound track. For example, the level of the user-specified parameter can be selected using the control element created by the user. This process can be repeated for multiple sound tracks during playback session 915 to generate multiple data points for the level of the user-specified parameter.

[0134] In some contemplated embodiments, the user can be asked to listen to multiple sound tracks at once and provide a comparative evaluation of the sound tracks based on user-defined parameters. For example, in the example of a user-created control that defines "attitude", the user can listen to multiple sound tracks and select which sound tracks have more "attitude" and / or which sound tracks have less "attitude". Each selection made by the user can be a data point for the level of the user-specified parameter.

[0135] After playback session 915 is completed, at 940, the levels of the audio parameters in the sound tracks from the playback session are evaluated. Examples of audio parameters include, but are not limited to: volume, pitch, bass, treble, reverb, etc. In some embodiments, the levels of the audio parameters in the sound track are evaluated while the sound track is being played (e.g., during playback session 915). In some embodiments, the audio parameters are evaluated after playback session 915 ends.

[0136] In various embodiments, the audio parameters in the sound track are evaluated from the metadata of the sound track. For example, an audio analysis algorithm can be used to generate metadata or symbolic music data (e.g., MIDI) for the sound track (which may be a short, pre-recorded music file). The metadata can include, for example, the note pitches present in the recording, the onset per beat, the ratio of pitched to non-pitched sounds, the volume, and other quantifiable attributes of the sound.

[0137] In 950, the correlation between the user selection level of the user-specified parameter and the audio parameter is determined. Since the user selection level of the user-specified parameter corresponds to the level of the control element, the correlation between the user selection level of the user-specified parameter and the audio parameter can be used to define the relationship between the level of one or more audio parameters and the level of the control element in 960. In various embodiments, AI techniques (e.g., regression models or machine learning algorithms) are used to determine the correlation between the user selection level of the user-specified parameter and the audio parameter and the relationship between the level of one or more audio parameters and the level of the control element.

[0138] Return to Figure 8 , and then the relationship between the level of one or more audio parameters and the level of the control element can be implemented by the music generator module 160 to determine how to adjust the audio parameters in the output music content 140 based on the input of the control element level 832 received from the user-created control element 830. In a particular embodiment, the music generator module 160 implements a machine learning algorithm to generate the output music content 140 and the relationship based on the input of the control element level 832 received from the user-created control element 830. For example, the machine learning algorithm can analyze how the metadata description of the audio track changes throughout the recording. The machine learning algorithm can include, for example, neural networks, Markov models, or dynamic Bayesian networks.

[0139] As described herein, the machine learning algorithm can be trained to predict the metadata of the upcoming music segment when the music metadata provided up to that point is given. The music generator module 160 can implement the prediction algorithm by searching a pool of pre-recorded audio files for those files having attributes that most closely match the predicted metadata for what is to come next. Selecting the closest matching audio file to play next helps create an output music content with musical properties of sequential progression that is similar to the example recordings used to train the prediction algorithm.

[0140] In some embodiments, the parameter control of the music generator module 160 using the prediction algorithm can be included within the prediction algorithm itself. In such embodiments, some predefined parameters can be used as inputs to the algorithm along with the music metadata, and the prediction varies based on the parameter. Alternatively, parameter control can be applied to the predictions to modify those predictions. As an example, by sequentially selecting the closest music segment predicted by the prediction algorithm to come next and splicing the audio of the files end to end, a generated work is made. At some point, the listener can increase the control element level (e.g., the control element at the start of each beat), and modify the output of the prediction model by increasing the predicted "start of each beat" data field. When selecting the next audio file to append to the work, an audio file with a higher start-of-each-beat attribute is more likely to be selected in such a scenario.

[0141] In various embodiments, a generation system such as music generator module 160 that utilizes metadata descriptions of music content may use hundreds or thousands of data fields in the metadata of each music segment. To provide more variability, multiple concurrent tracks may be used, each characterized by different instruments and sound types. In such instances, a prediction model may have thousands of data fields representing music attributes, each of which has a distinct impact on the listening experience. To allow a listener to control the music in such instances, an interface may be used for modifying each data field of the prediction model output, thereby creating thousands of control elements. Alternatively, multiple data fields may be combined and exposed as a single control element. As the more music attributes a single control element affects, the more abstract the control element becomes due to the specific music attributes, and the labels of these controls become subjective. In this way, primary control elements and sub-parameter control elements (described below) may be implemented for dynamic and personalized control of the output music content 140.

[0142] As described herein, a user may specify their own control elements and train the music generator module 160 on how to act based on the user's adjustment of the control elements. This process may reduce bias and complexity, and the data fields may be completely hidden from the listener. For example, in some embodiments, user-created control elements are provided to the listener on a user interface. The listener is then presented with a short music clip and asked to set a level of a control element that they believe best describes the music they hear. By repeating this process, multiple data points may be created that can be used to perform regression modeling on the desired effect of the control on the music. In some embodiments, these data points may be added as additional inputs to the prediction model. The prediction model may then attempt to predict a sequence of works that will produce a sequence similar to the one it has been trained on while also matching the expected behavior of the music attributes for which the control element has been set to a particular level. Alternatively, a control element mapper in the form of a regression model may be used to map prediction modifiers to control elements without retraining the prediction model.

[0143] In some embodiments, training for a given control element can include global training (e.g., training based on feedback from multiple user accounts) and local training (e.g., training based on feedback from the current user account). In some embodiments, a set of control elements can be created that are specific to a subset of music elements provided by the composer. For example, a scenario might include an artist creating a loop pack and then training the music generator module 160 using examples of performances or works they created previously using those loops. Patterns in these examples can be modeled using regression or neural network models and used to create rules for creating new music with similar patterns. These rules can be parameterized and exposed as control elements for the composer to manually modify offline before the listener begins using the music generator module 160, or for the listener to adjust while listening. Examples that the composer feels have the opposite effect of the desired effect of the control can also be used for negative reinforcement.

[0144] In some embodiments, in addition to leveraging patterns in example music, the music generator module 160 can find patterns in the music it creates that correspond to the composer's input before the listener begins listening to the generated music. The composer can do this through direct feedback (described below), such as tapping a thumbs-up control element for positive reinforcement of a pattern, or tapping a thumbs-down control element for negative reinforcement.

[0145] In various embodiments, the music generator module 160 can allow the composer to create their own sub-parameter control elements, such as the control elements learned by the music generator module described below. For example, a control element for "intensity" might have been created as a primary control element from a learned pattern related to the number of notes at each beat and the texture quality of the playing instrument. The composer can then create two sub-parameter control elements by selecting patterns related to the note onset, such as a "rhythm intensity" control element and a "texture intensity" control element for the texture pattern. Examples of sub-parameter control elements include vocal control elements, intensity for a specific frequency range (e.g., bass), complexity, tempo, etc. These sub-parameter control elements can be combined with more abstract control elements such as those related to injected energy (e.g., primary control elements). These composer skill control elements can be trained for the music generator module 160 by the composer in a manner similar to the user-created controls described herein.

[0146] As described herein, training the music generator module 160 to control audio parameters based on input from user-created control elements allows for separate control elements to be implemented for different users. For example, one user may associate increased attitude with increased bass content, while another user may associate increased attitude with a certain type of vocal or a range of beats. The music generator module 160 may modify the audio parameters for different attitude specifications based on the training of the music generator module for a particular user. In some embodiments, the personalized controls may be used in conjunction with global rules or control elements that are implemented in the same way for many users. The combination of global and local feedback or control may provide high-quality music production with specialized control for the individuals involved.

[0147] In various embodiments, as Figure 8 shown, one or more UI control elements 830 are implemented in the UI module 820. As described above, during interaction with the UI module 820, the user may use the control element(s) 830 to adjust the control element level(s) 832 to modify the output music content 140. In a particular embodiment, one or more of the control elements 830 are system-defined control elements. For example, the control elements may be defined by the system 800 as controllable parameters. In such embodiments, the user may adjust the system-defined control elements to modify the output music content 140 according to the system-defined parameters.

[0148] In a particular embodiment, system-defined UI control elements (e.g., knobs or sliders) allow the user to control abstract parameters of the output music content 140 automatically generated by the music generator module 160. In various embodiments, the abstract parameters serve as the primary control element inputs. Examples of abstract parameters include, but are not limited to, intensity, complexity, mood, genre, and energy level. In some embodiments, the intensity control element may adjust the number of combined low-frequency loops. The complexity control element may direct the number of covered tracks. Other control elements such as the mood control element may be in a range from awake to happy and affect, for example, the key of the music being played and other attributes.

[0149] In various embodiments, system-defined UI control elements (e.g., knobs or sliders) allow the user to control the energy level of the output music content 140 automatically generated by the music generator module 160. In some embodiments, the label of the control element (e.g., "Energy") may change size, color, or other attributes to reflect the user input for adjusting the energy level. In some embodiments, when the user adjusts the control element, the current level of the control element may be output until the user releases the control element (e.g., releases a mouse click or removes a finger from a touch screen).

[0150] System-defined energy can be an abstract parameter related to multiple more specific music attributes. For example, in various embodiments, energy can be related to the beat. For example, a change in the energy level can be associated with a change in the beats per minute (e.g., ~6 beats per minute) selected. In some embodiments, within a given range of one parameter (e.g., the beat), the music generator module 160 can explore musical variations by changing other parameters. For example, the music generator module 160 can create buildups and drops, create tension, change the number of simultaneously layered tracks, change the mode, add or remove vocals, add or remove bass, play different melodies, etc.

[0151] In some embodiments, one or more sub-parameter control elements are implemented as (one or more) control elements 830. The sub-parameter control elements can allow for more specific control of the attributes incorporated into a primary control element such as an energy control element. For example, the energy control element can modify the number of layers of percussion sounds and the amount of vocals used, but separate control elements allow for direct control of these sub-parameters, such that not all control elements need to be independent. In this way, the user can select the level of control specificity they desire. In some embodiments, sub-parameter control elements can be implemented for user-created control elements as described above. For example, a user can create and label a control element that specifies a sub-parameter of another user-specified parameter.

[0152] In some embodiments, the user interface module 820 allows the user to expand the UI control element 830 to display options for one or more sub-parameter user control elements. Additionally, a particular artist can provide attribute information for guiding music creation under a user control of a high-level control element (e.g., an energy slider). For example, an artist can provide an "artist pack" that has tracks and music creation rules of that artist. The artist can use an artist interface to provide values for the sub-parameter user control elements. For example, a DJ might expose tempo and drums as control elements to allow the listener to incorporate more or less tempo and drums. In some embodiments, as described herein, an artist or user can generate their own custom control elements.

[0153] In various embodiments, a human-in-the-loop generation system can be used to generate artifacts with the help of human intervention and control to potentially improve the quality and fitness of the generated music for personal purposes. For some embodiments of the music generator module 160, a listener can become a listener-composer by controlling the generation process through interface control elements 830 implemented in the UI module 820. The design and implementation of these control elements may affect the balance between an individual's listener and composer roles. For example, highly detailed and technical control elements may reduce the influence of the generation algorithm and give more creative control to the user, while requiring more hands-on interaction and technical skills to manage.

[0154] Conversely, higher-level control elements may reduce the amount of work and interaction time required while reducing creative control. For example, for an individual who desires more of a listener-type role, the primary control elements as described herein may be advantageous. For example, the primary control elements can be based on abstract parameters such as mood, intensity, or genre. These abstract parameters of music may be subjective metrics that are often interpreted individually. For example, in many cases, the listening environment affects how a listener describes music. Thus, music that a listener may call "relaxing" at a party may be too energetic and intense for a meditation session.

[0155] In some embodiments, one or more UI control elements 830 are implemented to receive user feedback regarding the output music content 140. User feedback control elements can include, for example, star ratings, thumbs up / thumbs down, etc. In various embodiments, user feedback can be used to train the system to adapt to a user's specific taste and / or apply to more global tastes of multiple users. In embodiments with thumbs up / thumbs down (e.g., positive / negative) feedback, the feedback is binary. Binary feedback including strong positive and strong negative responses can effectively provide positive and negative reinforcement for the functionality of the (one or more) control elements 830. In some envisioned embodiments, the input from the thumbs up / thumbs down control element can be used to control the output music content 140 (e.g., the thumbs up / thumbs down control element is used to control the output itself). For example, the thumbs up control element can be used to modify the maximum repeat count of the currently playing output music content 140.

[0156] In some embodiments, a counter for each audio file keeps track of how many times a section (e.g., an 8-beat section) of that audio file has been played recently. Once a file has been used above a desired threshold, a bias can be applied to its selection. Over time, this bias may gradually return to zero. Along with the rules defining musical sections that set the desired function of the music (e.g., build, drop, breakdown, intro, sustain), this repetition counter and bias can be used to shape the music into sections with a coherent theme. For example, the music generator module 160 can increment the count when a thumbs-down is pressed, such that the audio content of the output music content 140 is encouraged to change more quickly without disrupting the musical function of the section. Similarly, the music generator module 160 can decrement the count when a thumbs-up is pressed, such that the audio content of the output music content 140 does not deviate from the repetition for a long period of time. Before the threshold is reached and the bias is applied, other machine learning and rule-based mechanisms in the music generator module 160 may still result in the selection of other audio content.

[0157] In some embodiments, the music generator module 160 is configured to determine various context information (e.g., Figure 1 the environmental information 150 shown) around the time when user feedback is received. For example, in combination with receiving a "thumbs-up" indication from the user, the music generator module 160 can determine the time of day, location, device speed, biometric data (e.g., heart rate), etc. from the environmental information 150. In some embodiments, this context information can be used to train a machine learning model to generate music that the user likes in various different contexts (e.g., the machine learning model is context-aware).

[0158] In various embodiments, the music generator module 160 determines the current environmental type and takes different actions for the same user in different environments. For example, when a listener trains the "attitude" control element, the music generator module 160 can take environmental measurements and the listener's biometrics. During training, the music generator module 160 is trained to include these metrics as part of the control element. In this example, when the listener is doing high-intensity fitness in the gym, the "attitude" control element may affect the intensity of the drum beats. When sitting in front of a computer, changing the "attitude" control element may not affect the drum beats but may increase the distortion of the bass line. In such embodiments, a single user control element can have different rule sets or differently trained machine learning models that are used differently, either individually or in combination, in different listening environments.

[0159] In contrast to context awareness, if the expected behavior of a control element is static, then for many usage listening situation music generator modules 160, many controls may likely become necessary or desirable. Thus, in some embodiments, the disclosed techniques can provide functionality for multiple environments with a single control element. Implementing a single control element for multiple environments can reduce the number of control elements, making the user interface simpler and search faster. In some embodiments, the behavior of the control element is dynamic. The power of the control element can come from leveraging environmental measurements such as: sound levels recorded by a microphone, heart rate measurements, time of day, and rate of movement, etc. These measurements can be used as additional inputs for control element training. Thus, the same listener's interaction with the control element may produce different music effects, depending on the environmental context in which the interaction occurs.

[0160] In some embodiments, the above context awareness functionality is different from the concept of a music generation system that changes the generation process based on environmental context. For example, these techniques can modify the effect of a user control element based on environmental context, which can be used alone or in combination with the concept of generating music based on environmental context and the output of user controls.

[0161] In some embodiments, the music generator module 160 is configured to control the generated output music content 140 to achieve a given goal. Examples of given goals include but are not limited to sales goals, biometric goals such as heart rate or blood pressure, and environmental noise goals. The music generator module 160 can learn how to use the techniques described herein to modify manually (user-created) or algorithmically (system-defined) generated control elements to generate the output music content 140 in order to meet the given goal.

[0162] The target state can be a measurable environmental and listener state that the listener desires when using and listening to music with the music generator module 160. These target states may be directly affected - by changing the acoustic experience of the space where the listener is located through music, or may be regulated through psychological effects, such as specific music encouraging concentration. As an example, a listener can set a goal to have a lower heart rate during running. By recording the listener's heart rate in different states of the available control elements, the music generator module 160 learns that when a control element named "Attitude" is set to a low level, the listener's heart rate typically decreases. Thus, to help the listener achieve a lower heart rate, the music generator module 160 can automate the "Attitude" control to a low level.

[0163] By creating the kind of music that the listener expects in a particular environment, the music generator module 160 can help create that particular environment. Examples include heart rate, total volume in the listener's physical space, sales in a store, etc. Some environmental sensors and status data may not be suitable for the target state. For example, the time of day can be an environmental metric that is used as an input to achieve the target state of inducing sleep, but the music generator module 160 itself cannot control the time of day.

[0164] In various embodiments, while sensor inputs can be disconnected from the control element mapper when attempting to reach a state goal, the sensors can continue to record and instead provide measurements for comparing the actual state to the target state. The difference between the target and actual environmental states can be formulated as a reward function for a machine learning algorithm that can adjust the mapping in the control element mapper in a mode that attempts to achieve the target state. The algorithm can adjust the mapping to reduce the difference between the target and actual environmental states.

[0165] While music content has many physiological and psychological effects, creating the music content that the listener expects in a particular environment may not always help create that environment for the listener. In certain cases, it may have no effect or a negative impact on reaching the target state. In some embodiments, if the change does not meet a threshold, the music generator module 160 can adjust the music attributes based on past results while branching in other directions. For example, if decreasing the "attitude" control element does not result in a decrease in the listener's heart rate, the music generator module 160 can use other control elements to transform and develop new strategies, or use the target variable as the actual state for positive or negative reinforcement of a regression or neural network model to generate new control elements.

[0166] In some embodiments, if it is found that the context affects the expected behavior of the control elements for a particular listener, it may imply that the data points (e.g., audio parameters) modified by the control elements in a particular context are relevant to that listener's context. Therefore, these data points can provide good initial points for attempting to generate music that produces environmental changes. For example, if the listener always manually increases the "rhythm" control element when going to the train station, the music generator module 160 can start automatically increasing that control element when it detects that the listener is at the train station.

[0167] In some embodiments, as described herein, the music generator module 160 is trained to implement control elements that match user expectations. If the music generator module 160 is trained end-to-end for each control element (e.g., from the control element level to the output music content 140), the complexity of training for each control element may be high, which may slow down the training. Additionally, it may be difficult to establish the desired combined effect of multiple control elements. However, for each control element, the music generator module 160 should ideally be trained to perform the expected musical variations based on the control element. For example, the music generator module 160 can be trained by a listener for the "energy" control element such that the rhythm density increases as the "energy" increases. Since the listener is exposed to the final output music content 140 rather than just the individual layers of the music content, the music generator module 160 can be trained to use the control element to affect the final output music content. However, this may become a multi-step problem. For example, for a specific control setting, the music should sound like X, and to create music that sounds like X, a set of audio files Y should be used on each track.

[0168] In a particular embodiment, a teacher / student framework is employed to address the above issues. Figure 10 FIG. 5 is a block diagram illustrating an exemplary teacher / student framework system according to some embodiments. In the illustrated embodiment, the system 1000 includes a teacher model implementation module 1010 and a student model implementation module 1020.

[0169] In a particular embodiment, the teacher model implementation module 1010 implements a trained teacher model. For example, the trained teacher model can be a model that learns how to predict how the final mix (e.g., stereo mix) should sound without considering the set of loops available in the final mix. In some embodiments, the learning process of the teacher model utilizes real-time analysis of the output music content 140 using the fast Fourier transform (FFT) to calculate the sound distribution at different frequencies for a short time-step sequence. The teacher model can utilize a time series prediction model such as a recurrent neural network (RNN) to search for patterns in these sequences. In some embodiments, the teacher model in the teacher model implementation module 1010 can be trained offline on stereo recordings where no individual loops or audio files are available.

[0170] In the illustrated embodiment, the teacher model implementation module 1010 receives the output music content 140 and generates a compact description 1012 of the output music content. Using a trained teacher model, the teacher model implementation module 1010 can generate the compact description 1012 without considering the tracks or audio files in the output music content 140. The compact description 1012 may include a description X of what the output music content 140 should sound like, which is determined by the teacher model implementation module 1010. The compact description 1012 is more compact than the output music content 140 itself.

[0171] The compact description 1012 can be provided to the student model implementation module 1020. The student model implementation module 1020 implements a trained student model. For example, a trained student model may be a model that learns how to use audio files or loops Y (different from X) to produce music that matches the compact description. In the illustrated embodiment, the student model implementation module 1020 generates student output music content 1014 that substantially matches the output music content 140. As used herein, the phrase "substantially matches" means that the student output music content 1014 sounds similar to the output music content 140. For example, a trained listener may perceive the student output music content 1014 and the output music content 140 to sound the same.

[0172] In many instances, control elements are expected to affect similar patterns in music. For example, control elements may affect pitch relationships and rhythms. In some embodiments, the music generator module 160 is trained for a large number of control elements according to a teacher model. By using a single teacher model to train the music generator module 160 for a large number of control elements, it may not be necessary to relearn similar basic patterns for each control element. In such embodiments, the student model of the teacher model then learns how to change the loop selection for each track to achieve the desired properties in the final music mix. In some embodiments, the properties of the loops can be pre-computed to reduce the learning challenge and baseline performance (although it may come at the cost of potentially reducing the likelihood of finding the optimal mapping of the control elements).

[0173] Non-limiting examples of music properties pre-computed for each loop or audio file available for student model training include the following: the ratio of bass to treble frequencies, the number of note onsets per second, the ratio of detected tonal to atonal sounds, spectral range, average onset intensity. In some embodiments, the student model is a simple regression model that is trained to select loops for each track to obtain the closest music properties in the final stereo mix. In various embodiments, the student / teacher model framework may have some advantages. For example, if a new property is added to the pre-computation routine of the loops, it is not necessary to retrain the entire end-to-end model, only the student model.

[0174] As another example, since the attributes that affect the final stereo mix for different controls may be common to other control elements, training the music generator module 160 of each control element as an end-to-end model would mean that each model would need to learn the same thing (stereo mix music features) to obtain the best loop selection, making the training slower and harder than it might need to be. Only the stereo output needs to be analyzed in real time, and since the output music content is generated in real time for the listener, the music generator module 160 can obtain "free" signals through calculations. Even FFT may have been applied for visualization and audio mixing purposes. In this way, the teacher model can be trained to predict the combined behavior of the control elements, and the music generator module 160 is trained to find ways to adapt to other control elements while still producing the desired output music content. This can encourage training of the control elements to emphasize the unique effects of specific control elements and reduce control elements with effects that weaken the influence of other control elements.

[0175] Exemplary Low - Resolution Pitch Detection System

[0176] Pitch detection that is robust to polyphonic music content and multiple instrument types has traditionally been difficult to achieve. Tools for implementing end-to-end music transcription may record audio and attempt to generate a written musical score or a symbolic music representation in MIDI form. Without knowledge of the beat position or tempo, these tools may need to infer the musical rhythm structure, instruments, and pitch. The results may vary, and a common problem is detecting too many short, non-existent notes in the audio file and detecting the harmonics of a note as the fundamental pitch.

[0177] However, pitch detection can also be useful in cases where end-to-end transcription is not required. For example, to create a reasonable combination of harmonics for a music loop, it is sufficient to know which pitches can be heard on each beat without knowing the exact position of the notes. If the length and tempo of the loop are known, it may not be necessary to infer the temporal position of the beats from the audio.

[0178] In some embodiments, a pitch detection system is configured to detect which fundamental pitches (e.g., C, C#....B) are present in a short music audio file of a known beat length. By reducing the scope of the problem and focusing on robustness to instrument textures, high-precision results for beat-resolution pitch detection can be achieved.

[0179] In some embodiments, the pitch detection system is trained on examples with known ground truth. In some embodiments, the audio data is created from score data. MIDI and other symbolic music formats can be synthesized using a software audio synthesizer with random parameters to obtain textures and effects. For each audio file, the system can generate a log spectrum with multiple frequency bins for each pitch classFigure 2 It is represented by D. This 2D representation is used as the input to a neural network or other AI techniques, where multiple convolutional layers are used to create a feature representation of the audio frequency and time representation. The convolutional stride and padding can vary according to the length of the audio file to produce a constant model output shape for inputs with different beats. In some embodiments, the pitch detection system attaches a recurrent layer to the convolutional layer to output a sequence of time-related predictions. Categorical cross-entropy loss can be used to compare the logical output of the neural network with the binary representation of the musical score.

[0180] The design combining convolutional layers and recurrent layers may be similar to speech-to-text work, but modifications are needed. For example, speech-to-text typically needs to be sensitive to relative pitch changes rather than absolute pitch. Therefore, the frequency range and resolution are usually small. In addition, the text may need to remain unchanged, which is not desirable in music with a static beat as it may accelerate in an unwanted way. Connectionist temporal classification (CTC) loss calculation, which is often used in speech-to-text tasks, may not be needed, for example, because the length of the output sequence is known in advance, which reduces the complexity of training.

[0181] The following represents 12 pitch levels for each beat, where 1 indicates the presence of that basic note in the musical score for synthesizing audio. (C, C#....B) and each row represents a beat. For example, the subsequent rows represent musical scores for different beats:

[0182]

[0183] In some embodiments, the neural network is trained on pseudo-randomly generated musical scores of classical music and 1 - 4 part (or more) harmony and polyphony. Data augmentation can help enhance the robustness of the music content through filters and effects such as reverb, which can be a difficult point in pitch detection (for example, because partial fundamental tones still exist after the original note ends). In some embodiments, the dataset may be biased and loss weights are used because it is more likely that no note is played on each beat for the pitch levels.

[0184] In some embodiments, the output format allows for avoiding harmonic conflicts on each beat while maximizing the range of harmonic context that the loop can use. For example, a bass loop can include only F and move down to E on the last beat of the loop. For most people in the key of F, this loop sounds harmonious. If no time resolution is provided and it is only known that there is an E and an F in the audio, then it might end up being a sustained E with a short F, which sounds unacceptable to most people in the context of the key of F. As the resolution increases, the chance of detecting harmonics, fretboard sounds, and slides as individual notes increases, and thus additional notes might be misidentified. According to some embodiments, by developing a system with optimal time and pitch information resolution for combining short audio recordings of musical instruments to create a music mix with a combined harmonic sound, the complexity of the pitch detection problem can be reduced, and robustness to short, less significant pitch events can be increased.

[0185] In various embodiments of the music generator system described herein, the system can allow a listener to select audio content that is used to create a pool from which the system constructs (generates) new music. This approach might be different from creating a playlist because the user does not need to select individual tracks or organize the selections in sequence. Additionally, content from multiple artists can be used together simultaneously. In some embodiments, the music content is grouped into "packages" designed by a software provider or contributing artists. A package contains multiple audio files with corresponding image characteristics and feature metadata files. For example, a single package might contain 20 to 100 audio files that the music generator system can use to create music. In some embodiments, a single package can be selected or multiple packages can be selected in combination. During playback, packages can be added or removed without stopping the music.

[0186] Exemplary Audio Techniques for Music Content Generation

[0187] In various embodiments, a software framework for managing real-time generated audio can benefit from supporting specific types of functionality. For example, audio processing software might follow a modular signal chain metaphor inherited from analog hardware, where different modules providing audio generation and audio effects are linked together to form an audio signal graph. Individual modules typically expose various continuous parameters, thus allowing for real-time modification of the module's signal processing. In the early days of electronic music, the parameters themselves were often analog signals, so the parameter processing chain and the signal processing chain coincided. Since the digital revolution, parameters tend to be a separate digital signal.

[0188] The embodiments disclosed herein recognize that for a real-time music generation system—whether the system interacts live with a human performer or the system implements machine learning or other artificial intelligence (AI) techniques to generate music—a flexible control system that allows for the coordination and combination of parameter manipulation may be advantageous. Additionally, the present disclosure recognizes that it may be advantageous for the effects of parameter changes to be invariant to changes in tempo.

[0189] In some embodiments, a music generator system generates new music content from playback music content based on different parameter representations of an audio signal. For example, the audio signal can be represented by both a signal graph relative to time (e.g., an audio signal graph) and a signal graph relative to beat (e.g., a signal graph). The signal graph is invariant to tempo, which allows for tempo-invariant modification of the audio parameters of the music content in addition to tempo-based modification of the audio signal graph.

[0190] Figure 11 is a block diagram illustrating an exemplary system configured to implement audio techniques in music content generation. In the illustrated embodiment, system 1100 includes a graph generation module 1110 and an audio-technique music generator module 1120. The audio-technique music generator module 1120 can operate as a music generator module (e.g., the audio-technique music generator module is music generator module 160, as described herein) or the audio-technique music generator module can be implemented as part of a music generator module (e.g., as part of music generator module 160).

[0191] In the illustrated embodiment, music content 1112 including audio file data is accessed by graph generation module 1110. The graph generation module 1110 can generate a first graph 1114 and a second graph 1116 for the audio signal in the accessed music content 1112. In a particular embodiment, the first graph 1114 is an audio signal graph that plots the audio signal as a function of time. The audio signal can include, for example, amplitude, frequency, or a combination of both. In a particular embodiment, the second graph 1116 is a signal graph that plots the audio signal as a function of beat.

[0192] In a particular embodiment, as Figure 11 shown in the illustrated embodiment of, the graph generation module 1110 is located in system 1100 to generate the first graph 1114 and the second graph 1116. In such embodiments, the graph generation module 1110 can be collocated with the audio-technique music generator module 1120. However, other embodiments are envisioned where the graph generation module 1110 is located in a separate system and the audio-technique music generator module 1120 accesses the graphs from the separate system. For example, the graphs can be generated and stored on a cloud-based server accessible to the audio-technique music generator module 1120.

[0193] Figure 12 Depicts an example of an audio signal diagram (e.g., the first diagram 1114). Figure 13 Depicts an example of a signal diagram (e.g., the second diagram 1116). In Figure 12 and Figure 13 In the diagrams shown, each change in the audio signal is represented as a node (e.g., Figure 12 audio signal node 1202 in Figure 13 and signal node 1302 in ). Thus, the parameters of a specified node determine (e.g., define) the change in the audio signal at the specified node. Since the first diagram 1114 and the second diagram 1116 are based on the same audio signal, the diagrams may have a similar structure, and the change between the diagrams is the x-axis scale (time vs. beat). Having a similar structure in the diagrams allows modifying the parameters of a node in one diagram (e.g., node 1202 in the first diagram 1114) corresponding to a node in another diagram (e.g., node 1302 in the second diagram 1116) determined by the parameters downstream or upstream of the node in one diagram (as described below).

[0194] Returning to Figure 11 , the first diagram 1114 and the second diagram 1116 are received (or accessed) by the audio technology music generator module 1120. In a particular embodiment, the audio technology music generator module 1120 generates new music content 1122 from the playback music content 1118 based on the audio modifier parameters selected from the first diagram 1114 and the audio modifier parameters selected from the second diagram 1116. For example, the audio technology music generator module 1120 may use the audio modifier parameters from the first diagram 1114, the audio modifier parameters from the second diagram 1116, or a combination thereof to modify the playback music content 1118. The new music content 1122 is generated by modifying the playback music content 1118 based on the audio modifier parameters.

[0195] In various embodiments, the audio technology music generator module 1120 may select audio modifier parameters to implement in the modification of the playback content 1118 based on whether a beat change modification, a beat invariant modification, or a combination thereof is required. For example, a beat change modification may be performed based on the audio modifier parameters selected or determined from the first diagram 1114, while a beat invariant modification may be performed based on the audio modifier parameters selected or determined from the second diagram 1116. In embodiments where a combination of a beat change modification and a beat invariant modification is required, the audio modifier parameters may be selected from both the first diagram 1114 and the second diagram 1116. In some embodiments, the audio modifier parameters from each individual diagram are applied separately to different attributes (e.g., amplitude or frequency) or different layers (e.g., different instrument layers) in the playback music content 1118. In some embodiments, the audio modifier parameters from each diagram are combined into a single audio modifier parameter to be applied to a single attribute or layer in the playback music content 1118.

[0196] Figure 14 depicts an exemplary system for implementing real-time modification of music content using an audio technology music generator module 1420 according to some embodiments. In the illustrated embodiment, the audio technology music generator module 1420 includes a first node determination module 1410, a second node determination module 1420, an audio parameter determination module 1430, and an audio parameter modification module 1440. The first node determination module 1410, the second node determination module 1420, the audio parameter determination module 1430, and the audio parameter modification module 1440 together implement the system 1400.

[0197] In the illustrated embodiment, the audio technology music generator module 1420 receives playback music content 1418 including an audio signal. The audio technology music generator module 1420 can process the audio signal in the first node determination module 1410 through a first graph 1414 (e.g., a time-based audio signal graph) and a second graph 1416 (e.g., a beat-based signal graph). When the audio signal passes through the first graph 1414, the parameters of each node in the graph determine the change of the audio signal. In the illustrated embodiment, the second node determination module 1420 can receive information about the first node 1412 and determine information about the second node 1422. In a particular embodiment, the second node determination module 1420 reads the parameters in the second graph 1416 based on the position of the first node in the first node information 1412 in the audio signal passing through the first graph 1414. Thus, as an example, the audio signal going to node 1202 in the first graph 1414 ( Figure 12 shown) determined by the first node determination module 1410 can trigger the second node determination module 1420 to determine the corresponding (parallel) node 1302 in the second graph 1416 ( Figure 13 shown).

[0198] As Figure 14 shown, the audio parameter determination module 1430 can receive the second node information 1422 and determine (e.g., select) a specified audio parameter 1432 based on the second node information. For example, the audio parameter determination module 1430 can select an audio parameter based on a portion of the next beat (e.g., x next beats) in the second graph 1416 after the position of the second node identified in the second node information 1422. In some embodiments, a beat-to-real-time conversion can be implemented to determine the portion of the second graph 1416 from which the audio parameter can be read. The specified audio parameter 1432 can be provided to the audio parameter modification module 1440.

[0199] The audio parameter modification module 1440 can control the modification of music content to generate new music content. For example, the audio parameter modification module 1440 can modify the playback music content 1418 to generate new music content 1122. In a particular embodiment, the audio parameter modification module 1440 modifies the properties of the playback music content 1418 by modifying the specified audio parameters 1432 (determined by the audio parameter determination module 1430) of the audio signals in the playback music content. For example, modifying the specified audio parameters 1432 of the audio signals in the playback music content 1418 modifies properties such as amplitude, frequency, or a combination of both in the audio signals. In various embodiments, the audio parameter modification module 1440 modifies the properties of different audio signals in the playback music content 1418. For example, the different audio signals in the playback music content 1418 can correspond to different musical instruments represented in the playback music content 1418.

[0200] In some embodiments, the audio parameter modification module 1440 uses machine learning algorithms or other AI techniques to modify the properties of the audio signals in the playback music content 1418. In some embodiments, the audio parameter modification module 1440 modifies the properties of the playback music content 1418 according to user input to the module, which can be provided through a user interface associated with the music generation system. Embodiments where the audio parameter modification module 1440 uses a combination of AI techniques and user input to modify the properties of the playback music content 1418 are also conceivable. The various embodiments for modifying the properties of the playback music content 1418 by the audio parameter modification module 1440 allow for real-time manipulation of the music content (e.g., manipulation during playback). As described above, real-time manipulation can include applying beat change modifications, beat-invariant combinations, or a combination of both to the audio signals in the playback music content 1418.

[0201] In some embodiments, the audio technology music generation module 1420 implements a two-layer parameter system for modifying the properties of the playback music content 1418 by the audio parameter modification module 1440. In the two-layer parameter system, a differentiation may exist between "automation" that directly controls the audio parameter values (e.g., tasks automatically performed by the music generation system) and "modulation" that multiplies the audio parameter modifications on top of the automation, as described below. The two-layer parameter system may allow different parts of the music generation system (e.g., different machine learning models in the system architecture) to consider different musical aspects separately. For example, one part of the music generation system can set the volume of a particular instrument according to the expected part type of the music piece, while another part can override the periodic changes in volume for added interest.

[0202] Exemplary Techniques for Real - Time Audio Effects in Music Content Generation

[0203] Music technology software typically allows composers / producers to control various abstract envelopes through automation. In some embodiments, automation is a pre-programmed temporal manipulation of some audio processing parameters, such as volume or reverb amount. Automation is typically a manually defined breakpoint envelope (e.g., a piecewise linear function) or a programmed function, such as a sine wave (also known as a low-frequency oscillator (LFO)).

[0204] The disclosed music generator system may be different from typical music software. For example, in a sense, most parameters are automated by default. The AI technology in the music generator system can control most or all of the audio parameters in various ways. At a basic level, a neural network can predict the appropriate settings for each audio parameter based on its training. However, it may be helpful to provide some higher-level automation rules for the music generator system. For example, a large music structure may require a slow build-up of volume as an additional consideration, on top of the low-level settings that may be predicted.

[0205] The present disclosure generally relates to an information architecture and a program method for combining multiple parameter commands simultaneously issued by different levels of a hierarchical generation system to create a continuous output that is musically coherent and varied. The disclosed music generator system can create a long-form music experience designed to be experienced continuously for several hours. A long-form music experience requires creating a coherent musical journey for a more satisfying experience. To this end, the music generator system can refer to itself on a very long time scale. These references can range from direct to abstract.

[0206] In a particular embodiment, to facilitate larger-scale music rules, the music generator system (e.g., music generator module 160) discloses an automation API (Application Programming Interface). Figure 15 A block diagram of an exemplary API module in a system for audio parameter automation according to some embodiments is depicted. In the illustrated embodiment, system 1500 includes API module 1505. In a particular embodiment, API module 1505 includes automation module 1510. The music generator system can support wavetable-style LFOs and arbitrary breakpoint envelopes. Automation module 1510 can apply automation 1512 to any audio parameter 1520. In some embodiments, automation 1512 is applied recursively. For example, any program automation, such as a sine wave, which itself has parameters (frequency, amplitude, etc.), can have automation applied to those parameters.

[0207] In various embodiments, automation 1512 includes a signal map parallel to the audio signal map, as described above. The signal map can be processed similarly: by a "pull" technique. In the "pull" technique, the API module 1505 can request the automation module 1510 to recalculate as needed and perform the recalculation such that automation 1512 recursively requests the upstream automation it depends on to do the same. In a particular embodiment, the signal map for automation is updated at a controlled rate. For example, the signal map can be updated once per run of the performance engine update routine, which can be aligned with the block rate of the audio (e.g., after an audio signal map presents a block (e.g., a block is 512 samples)).

[0208] In some embodiments, it may be desirable for the audio parameter 1520 itself to vary at the audio sampling rate, otherwise discontinuous parameter changes at the audio block boundaries can cause audible artifacts. In a particular embodiment, the music generator system manages this issue by treating automation updates as parameter value targets. When the real-time audio thread presents an audio block, the audio thread will smoothly fade a given parameter from its current value to the provided target value over the course of the block.

[0209] The music generator system described herein (e.g., Figure 1 the music generator module 160 shown in Figure 15 may have an architecture with a hierarchical nature. In some embodiments, different parts of the hierarchy can provide multiple suggestions for the value of a particular audio parameter. In a particular embodiment, the music generator system provides two separate mechanisms for combining / resolving multiple suggestions: modulation and override. In the

[0210] shown embodiment of

[0211] In various embodiments, the API module 1505 includes an override module 1540. The override module 1540 can be, for example, an override tool for audio parameter automation. The override module 1540 can be intended to be used by an external control interface (e.g., an artist control user interface). The override module 1540 can control the audio parameters 1520 regardless of what the music generator system is trying to do with them. When an override 1542 overrides the audio parameters 1520, the music generator system can create a "shadow parameter" 1522 that tracks where the audio parameters would be if they were not overridden (e.g., where the audio parameters would be based on the automation 1512 or modulation 1532). Thus, when the override 1542 is "released" (e.g., removed by the artist), the audio parameters 1520 can quickly return to where they would have been based on the automation 1512 or modulation 1532.

[0212] In various embodiments, the two methods can be combined. For example, the override 1542 can be the modulation 1532. When the override 1542 is the modulation 1532, the base value of the audio parameter 1520 can still be set by the music generator system but is then multiplicatively modulated by the override 1542 (overriding any other modulation). Each audio parameter 1520 can have one (or zero) automation 1512 and one (or zero) modulation 1532 at the same time, as well as one (or zero) in each override 1542.

[0213] In various embodiments, the abstract class hierarchy is defined as follows (note that there is some multiple inheritance):

[0214]

[0215]

[0216] Based on the abstract class hierarchy, things can be considered to be automation or automatable. In some embodiments, any automation can be applied to anything that is automatable. Automation includes things such as LFOs, breakpoint envelopes, etc. These automations are all beat-locked, which means they change over time according to the current beat.

[0217] Automation itself may have parameters that are automatable. For example, the frequency and amplitude of LFO automation are automatable. Thus, there are signal graphs that run in parallel with the audio signal graph but at the control rate rather than the audio rate and that depend on automation and automation parameters. As described above, the signal graph uses a pull model. The music generator system tracks any automation 1512 applied to the audio parameters 1520 and updates these once per "game loop". The automation 1512 in turn recursively requests updates to its own automated audio parameters 1520. This recursive update logic may reside in the base class beat dependency, which is expected to be called frequently (but not necessarily regularly). The update logic may have a prototype described as follows:

[0218] BeatDependency::update(double currentBeat, int updateCounter, bool override)

[0219] In a particular embodiment, the beat dependency class maintains a list of its own dependencies (e.g., other beat dependency instances) and recursively calls their update functions. An update counter can be passed up so that the signal graph can have loops without double updates. This may be important because automation can be applied to several different automatables. In some embodiments, this may not matter because the second update will have the same current beat as the first and these update routines should be ineffective unless the beat changes.

[0220] In various embodiments, when automation is applied to each loop of the "game loop" of an automatable, the music generator system can request (recursively) the updated value from each automation and use it to set the value of the automatable. In this case, "setting" may depend on the specific subclass and also on whether the parameter is also being modulated and / or overridden.

[0221] In a particular embodiment, modulation 1532 is automation 1512 applied multiplicatively rather than absolutely. For example, modulation 1532 can be applied to an audio parameter 1520 that is already automated, and the effect will be a percentage of the automation value. For example, this multiplicative approach may allow for continuous oscillation around a moving average.

[0222] In some embodiments, the audio parameters 1520 can be overridden, as described above, meaning that any automation 1512 or modulation 1532 applied to them, or other (less privileged) requests are overridden by the override value in 1542. This override can allow for external control of specific aspects of the music generator system while the music generator system continues. When an audio parameter 1520 is overridden, the music generator system keeps track of what the value would be (e.g., keeps track of the applied automation / modulation and other requests). When the override is released, the music generator system snaps the parameter to where it would have been.

[0223] To facilitate modulation 1532 and coverage 1542, the music generator system can abstract the setValue method of Parameter. There may also be a private method _setValue that actually sets the value. An example of the public method is as follows:

[0224]

[0225] The public method can reference a member variable of the Parameter class named _unmodulated. This variable is an instance of ShadowParameter, as described above. Each audio parameter 1520 has a shadow parameter 1522 that keeps track of where it would be if it were not modulated. If the audio parameter 1520 is not currently being modulated, both the audio parameter 1520 and its shadow parameter 1522 are updated with the requested value. Otherwise, the shadow parameter 1522 keeps track of the request, and the actual audio parameter value 1520 is set elsewhere (e.g., in the updateModulations routine - where the modulation factor is multiplied by the shadow parameter value to give the actual parameter value).

[0226] In various embodiments, large-scale structures in a long-form music experience are achieved through various mechanisms. A broad approach might be to use musical self-reference over time. For example, a very straightforward self-reference would be to exactly repeat a particular audio segment that was played previously. In music theory, a repeated segment can be called a theme (or motif). More typically, musical content uses themes and variations, whereby the theme repeats at a later time with some variations to provide a sense of coherence but maintain a sense of progression. The music generator system disclosed herein can use themes and variations to create large-scale structures in a variety of ways, including direct repetition or by using abstract envelopes.

[0227] An abstract envelope is the value of an audio parameter that changes over time. Abstracted from the audio parameter it controls, an abstract envelope can be applied to any other audio parameter. For example, a collection of audio parameters can be automated consistently by a single controlling abstract envelope. This technique can "glue" different layers together perceptually in the short term. Abstract envelopes can also be reused temporarily and applied to different audio parameters. In this way, the abstract envelope becomes an abstract musical theme, and the theme is repeated by applying the envelope to different audio parameters in a later listening experience. Thus, while a sense of structure and long-term coherence is established, there is variation in the theme.

[0228] As a musical theme, an abstract envelope can abstract many musical characteristics. Examples of musical characteristics that can be abstracted include, but are not limited to:

[0229] · Building tension (volume of any track, degree of distortion, etc.).

[0230] · Rhythm (volume adjustment and / or gating to create rhythm effects applied to pads, etc.).

[0231] · Melody (pitch filtering can mimic the melody contour applied to pads, etc.).

[0232] Exemplary Additional Audio Techniques for Real - Time Music Content Generation

[0233] Real-time music content generation can pose unique challenges. For example, due to strict real-time constraints, function calls or subroutines with unpredictable and potentially infinite execution times should be avoided. Avoiding this problem may rule out the use of most high-level programming languages as well as most low-level languages (such as C and C++). Anything that allocates memory from the heap (e.g., via the malloc function under the hood) may be excluded, as well as anything that may block, such as locking a mutex. This can make multi-threaded programming particularly difficult for real-time music content generation. Most standard memory management methods may also not be feasible, so dynamic data structures (such as C++, STL containers) have limited use in real-time music content generation.

[0234] Another area of challenge may be managing audio parameters (such as the cutoff frequency of a filter) involved in DSP (Digital Signal Processing) functions. For example, when dynamically changing an audio parameter, audible artifacts may occur unless the audio parameter changes continuously. Thus, communication between a real-time DSP audio thread and a user- or programming interface may be required to change the audio parameter.

[0235] · Various audio software can be implemented to handle these limitations, and there are various methods. For example:

[0236] · Inter-thread communication can be handled using lock-free message queues.

[0237] · Functions can be written in pure C and utilize function pointer callbacks.

[0238] · Memory management can be implemented through custom "regions" or "arenas"

[0239] · A "two-speed" system can be implemented with a real-time audio thread running at audio rate and a control audio thread running at a "control rate". The control audio thread can set the audio parameter change target, and the real-time audio thread smoothly ramps up to that target.

[0240] In some embodiments, synchronizing between controlled-rate audio parameter manipulation and thread-safe storage of audio parameter values for actual DSP routines may require some thread-safe communication of audio parameter targets. Most audio parameters of audio routines are continuous (as opposed to discrete) and are thus typically represented by floating-point data types. Due to the lack of lock-free atomic floating-point data types, various contortions of the data have historically been required.

[0241] In a particular embodiment, a simple lock-free atomic floating-point data type is implemented in the music generator system described herein. A lock-free atomic floating-point data type can be achieved by treating the floating-point type as a sequence of bits and "tricking" the compiler into treating it as an atomic integer type with the same bit width. This approach can support atomic fetch / set and is applicable to the music generator system described herein. An example implementation of the lock-free atomic floating-point data type is described below:

[0242] / / atomic float

[0243] class af32{

[0244] public:

[0245] af32(){}

[0246] af32(float x){operator()(x);}

[0247] ~af32(){}

[0248] af32(const af32&x):valueStore(x()){}

[0249] af32&operator=(const af32&x){this->operator()(x());return*this;}

[0250] float operator()()const{uint32_t voodoo = atomic_load(&valueStore);

[0251] return((float)&voodoo);}

[0252] void operator()(float value){

[0253] uint32_t voodoo = ((uint32_t)&value); atomic_store(&_valueStore,voodoo);

[0254] }

[0255] private:

[0256] std::atomic_uint32_t _valueStore{0};

[0257] };

[0258] In some embodiments, dynamic memory allocation from the heap is not feasible for real-time code associated with music content generation. For example, static stack-based allocation may make it difficult to use programming techniques such as dynamic storage containers and functional programming methods. In a particular embodiment, the music generator system described herein implements a "memory zone" for memory management in a real-time context. As used herein, a "memory zone" is a region of heap-allocated memory that is pre-allocated without real-time constraints (e.g., when real-time constraints do not yet exist or are suspended). Memory storage objects can then be created in the heap-allocated memory zone without requesting more memory from the system, thus making the memory real-time safe. Garbage collection may include deallocating the memory zone as a whole. The memory implementation of the music generator system can also be multi-thread safe, real-time safe, and efficient.

[0259] Figure 16 A block diagram depicting an exemplary memory zone 1600 in accordance with some embodiments is shown. In the illustrated embodiment, the memory zone 1600 includes a heap-allocated memory module 1610. In various embodiments, the heap-allocated memory module 1610 receives and stores a first graph 1114 (e.g., an audio signal graph), a second graph 1116 (e.g., a signal graph), and audio signal data 1602. For example, each stored item can be retrieved by the audio parameter modification module 1440 ( Figure 14 as shown).

[0260] An example implementation of the memory zone is described below:

[0261] / / Memory pool class memory zone

[0262] { public:

[0263] MemoryZone(uint64_t sz) : sz(sz), zone((char*)malloc(sz)) {}

[0264] ~MemoryZone() { free(zone);}

[0265] void* bags(size_t obj_size, size_t alignment) {

[0266] uint64_t p = atomic_load(&p);

[0267] uint64_t q = p % uint64_t(alignment);

[0268] if (p + q > sz) return nullptr;

[0269] uint64_t pp = atomic_fetch_add(&p, uint64_t(obj_size) + q);

[0270] if (pp == p) { return zone_ + p + q;}

[0271] else { return bags(obj_size, alignment);}

[0272] } uint64_t used() { return atomic_load(&p);}

[0273] uint64_t available() { return int64_t(sz) - int64_t(atomic_load(&p));}

[0274] void hose() { atomic_store(&p, 0ULL);}

[0275] private:

[0276] char zone;

[0277] uint64_t sz;

[0278] std::atomic_uint64_t p_{0};

[0279] };

[0280] In some embodiments, different audio threads of a music generator system need to communicate with each other. Typical thread-safe methods (which may include locking "mutex" data structures) may not be usable in a real-time context. In a particular embodiment, dynamic routing data serialization to a single-producer single-consumer circular buffer pool is implemented. A circular buffer is a FIFO (first-in first-out) queue data structure that generally does not require dynamic memory allocation after initialization. A single-producer, single-consumer thread-safe circular buffer may allow one audio thread to push data into the queue while another audio thread pulls the data. For the music generator system described herein, the circular buffer can be extended to allow multi-producer, single-consumer audio threads. These buffers can be implemented by preallocating a static array of circular buffers and dynamically routing serialized data to specific "channels" (e.g., specific circular buffers) based on an identifier added to the music content generated by the music generator system. A single user (e.g., a single consumer) can access the static array of circular buffers.

[0281] Figure 17 A block diagram of an exemplary system for storing new music content in accordance with some embodiments is depicted. In the illustrated embodiment, system 1700 includes a circular buffer static array module 1710. The circular buffer static array module 1710 may include multiple circular buffers that allow storage of multi-producer, single-consumer audio threads based on a thread identifier. For example, the circular buffer static array module 1710 may receive new music content 1122 and store the new music content at 1712 for user access.

[0282] In various embodiments, abstract data structures such as dynamic containers (vectors, queues, lists) are generally implemented in a non-real-time-safe manner. However, these abstract data structures may be useful for audio programming. In a particular embodiment, the music generator system described herein implements a custom list data structure (e.g., a singly-linked list). Many functional programming techniques can be implemented from the custom list data structure. The custom list data structure implementation can use a "memory region" (as described above) for underlying memory management. In some embodiments, the custom list data structure is serializable, which can make it safe for real-time use and enable communication between audio threads using the multi-producer, single-consumer audio threads described above.

[0283] Exemplary Blockchain Ledger Technology

[0284] In some embodiments, the disclosed system can utilize secure recording techniques such as blockchain or other cryptographic ledgers to record information about the generated music or its elements, such as loops or tracks. In some embodiments, the system combines multiple audio files (e.g., tracks or loops) to generate output music content. The combination can be performed by combining multiple audio content layers such that they at least partially overlap in time. The output content can be discrete music segments or continuous. In the context of continuous music, tracking the use of music elements can be challenging, e.g., for providing royalties to relevant stakeholders. Thus, in some embodiments, the disclosed system records the identifiers and usage information (e.g., timestamps or play counts) of the audio files used in the composed music content. Additionally, for example, the disclosed system can utilize various algorithms to track the playback time in the context of mixing audio files.

[0285] As used herein, the term "blockchain" refers to a cryptographically linked set of records (called blocks). For example, each block can include the cryptographic hash of the previous block, a timestamp, and transaction data. A blockchain can be used as a public distributed ledger and can be managed by a network of computing devices that communicate and verify new blocks using an agreed-upon protocol. Some blockchain implementations may be immutable, while others may allow subsequent changes to the blocks. Generally, a blockchain can record transactions in a verifiable and permanent manner. Although blockchain ledgers are discussed herein for illustrative purposes, it should be understood that in other embodiments, the disclosed techniques can be used with other types of cryptographic ledgers.

[0286] Figure 18 is a diagram showing example playback data according to some embodiments. In the illustrated embodiment, the database structure includes entries for multiple files. Each illustrated entry includes a file identifier, a start timestamp, and a total time. The file identifier can uniquely identify the audio file being tracked by the system. The start timestamp can indicate when the audio file was first included in the mixed audio content. For example, the timestamp can be based on the local clock of the playback device or on an internet clock. The total time can indicate the length of the interval during which the audio file was incorporated. Note that this may be different from the length of the audio file, e.g., if only a portion of the audio file was used, if the audio file was accelerated or decelerated in the mix, etc. In some embodiments, when an audio file is incorporated at multiple different times, an entry is created each time. In other embodiments, if an entry for the file already exists, additional playbacks of the file may result in an increase in the time field of the existing entry. In still other embodiments, the data structure can track the number of times each audio file was used rather than the length of incorporation. Additionally, other encodings of time-based usage data are envisioned.

[0287] In various embodiments, different devices can determine, store, and use a ledger to record playback data. The following refers to Figure 19 discusses example scenarios and topologies. The playback data can be temporarily stored on a computing device before being submitted to the ledger. The stored playback data can be encrypted, for example, to reduce or avoid manipulation of entries or insertion of incorrect entries.

[0288] Figure 19 is a block diagram showing an example composition system according to some embodiments. In the example shown, the system includes a playback device 1910, a computing system 1920, and a ledger 1930.

[0289] In the illustrated embodiment, the playback device 1910 receives control signaling from the computing system 1920 and sends playback data to the computing system 1920. In this embodiment, the playback device 1910 includes a playback data recording module 1912, which can record playback data based on an audio mix played by the playback device 1910. The playback device 1910 also includes a playback data storage module 1914, which is configured to temporarily store the playback data in the ledger, or both. The playback device 1910 can report the playback data to the computing system 1920 periodically or can report the playback data in real time. For example, when the playback device 1910 is offline, the playback data can be stored for later reporting.

[0290] In the illustrated embodiment, the computing system 1920 receives the playback data and submits an entry reflecting the playback data to the ledger 1930. The computing system 1920 also sends control signaling to the playback device 1910. In different embodiments, the control signaling can include various types of information. For example, the control signaling can include configuration data, mixing parameters, audio samples, machine learning updates, etc., for the playback device 1910 to use in creating music content. In other embodiments, the computing system 1920 can compose music content and stream the music content data to the playback device 1910 via the control signaling. In these embodiments, the modules 1912 and 1914 can be included in the computing system 1920. Generally, referring to Figure 19 the modules and functions discussed can be distributed among multiple devices according to various topologies.

[0291] In some embodiments, the playback device 1910 is configured to submit entries directly to the ledger 1930. For example, a playback device such as a mobile phone can compose music content, determine the playback data, and store the playback data. In this case, the mobile device can report the playback data to a server, such as the computing system 1920, or directly to the computing system (or a group of computing nodes) that maintains the ledger 1930.

[0292] In some embodiments, the system maintains a record of rights holders, e.g., having a mapping to an audio file identifier or to a set of audio files. This entity record can be stored in ledger 1930 or in a separate ledger or in some other data structure. This can allow the rights holders to remain anonymous, e.g., when ledger 1930 is public but includes non-identifying entity identifiers that map to entities in some other data structure.

[0293] In some embodiments, a music composition algorithm can generate a new audio file from two or more existing audio files for inclusion in a remix. For example, the system can generate a new audio file C based on two audio files A and B. One technique for such mixing uses interpolation between the vector representations of the audio of file A and file B and uses an inverse transform from the vector to the audio representation to generate file C. In this example, the play times of both audio file A and audio file B can be increased, but the amount of their increase may be less than their actual play times, e.g., because they are being mixed.

[0294] For example, if audio file C is incorporated into the mixed content for 20 seconds, audio file A may have playback data indicating 15 seconds, while audio file B may have playback data indicating 5 seconds (and note that the sum of the mixed audio files may or may not match the usage length of the resulting file C). In some embodiments, the playback time of each original file is based on its similarity to the mixed file C. For example, in a vector embodiment, for an n-dimensional vector representation, the interpolated vector a has the following distances d from the vector representations of audio files A and B:

[0295] d(a,c) = ((a1 - c1) 2 +(a2 - c2) 2 +...+(an - cn) 2 ) 1 / 2

[0296] d(b,c) = ((b1 - c1) 2 +(b2 - c2) 2 +...+(bn - cn) 2 ) 1 / 2

[0297] In these embodiments, the playback time i of each original file can be determined as:

[0298]

[0299]

[0300] where t represents the playback time of file C.

[0301] In some embodiments, the form of compensation can be incorporated into the ledger structure. For example, a particular entity can include information associating an audio file with performance requirements, such as displaying a link or including an advertisement. In these embodiments, when an audio file is included in a mix, the synthesis system can provide proof of performance of the associated operation (e.g., displaying an advertisement). The proof of performance can be reported according to one of various suitable reporting templates that require specific fields to show how and when the operation was performed. The proof of performance can include time information and utilize cryptography to avoid false performance assertions. In these embodiments, an audio file that does not show proof of execution of the relevant required operation may require some other form of compensation, such as a royalty payment. Generally, different entities submitting audio files can register for different forms of compensation.

[0302] As described above, the disclosed techniques can provide a reliable record of audio files used in a music mix, even during real-time synthesis. The openness of the ledger may provide confidence in the fairness of compensation. This, in turn, may encourage the participation of artists and other collaborators, which may increase the variety and quality of audio files available for automated mixing.

[0303] In some embodiments, an artist pack can be made with the elements that a music engine uses to create a continuous soundscape. An artist pack can be a professionally (or otherwise) curated set of elements that are stored in one or more data structures associated with an entity such as an artist or group. Examples of these elements include, but are not limited to, loops, compositional rules, heuristics, and neural network vectors. Loops can be included in a music phrase database. Each loop is typically a single instrument or a group of related instruments playing a musical progression over a period of time. These can range from short loops (e.g., 4 bars) to longer loops (e.g., 32 to 64 bars) and so on. Loops can be organized into layers, such as melody, harmony, drums, bass, tops, FX, etc. The loop database can also be represented as a variational autoencoder with an encoded loop representation. In this case, instead of the loops themselves, an NN is used to generate the sounds encoded in the NN.

[0304] A heuristic refers to the parameters, rules, or data that guide the music engine in creating music. Parameters guide elements such as the length of a section, the use of effects, the frequency of variational techniques, the complexity of the music, or generally, any type of parameter that can be used to enhance decision-making when the music engine creates and presents music.

[0305] The ledger records transactions related to content consumption associated with a rights holder. For example, this can be a loop, a heuristic, or a neural network vector. The goal of the ledger is to record these transactions and make them available for transparent accounting checks. The ledger is designed to capture transactions that occur, which may include content consumption, use of parameters to drive a music engine, and use of vectors on a neural network, among others. The ledger can record various types of transactions, including discrete events (e.g., this loop played at this time), this package was played for this amount of time, or this machine learning module (e.g., neural network module) was used for this amount of time.

[0306] The ledger can associate multiple rights holders with any given artist package, or more specifically, with a particular loop or other element of the artist package. For example, a label, an artist, and a composer may have rights to a given artist package. The ledger can allow them to associate payment details for the package, specifying the percentage that each party will receive. For example, the artist can receive 25%, the record label 25%, and the composer 50%. Using blockchain to manage these transactions can allow for micropayments to each rights holder in real time, or accumulation over an appropriate time period.

[0307] As described above, in some embodiments, the loop may be replaced with a VAE that is essentially an encoding of the loop in a machine learning module. In this case, the ledger can associate playtime with a particular artist package that contains the machine learning module. For example, if an artist package accounts for 10% of the total playtime across all devices, the artist can receive 10% of the total revenue distribution.

[0308] In some embodiments, the system allows an artist to create an artist profile. The profile includes relevant information about the artist, including a resume, a profile picture, bank details, and other data required to verify the artist's identity. After creating the artist profile, the artist can upload and publish artist packages. These packages include elements that the music engine uses to create soundscapes.

[0309] For each artist package created, rights holders can be defined and associated with the package. Each rights holder can claim a certain percentage of the package. Additionally, each rights holder creates a profile and associates a bank account with their profile for payment. The artist themselves is a rights holder and may own 100% of the rights associated with their work package.

[0310] In addition to recording events in the ledger that will be used for revenue recognition, the ledger can also manage promotional activities associated with artist packs. For example, an artist pack may have a free monthly promotion, and the revenue generated during this promotion will be different from the revenue generated when the promotion is not running. The ledger automatically takes these revenue inputs into account when calculating payments to rights holders.

[0311] The same copyright management model can allow an artist to sell the copyright of their pack to one or more external rights holders. For example, when launching a new pack, an artist can pre-fund their pack by selling 50% of the shares in the pack to fans or investors. In this case, the number of investors / rights holders can be arbitrarily large. For example, an artist can sell 50% of their percentage to 100,000 users, and these users will receive 1 / 100,000 of the revenue generated by the pack. Since all accounting is managed by the ledger, in this scenario, the investors will be paid directly without the need to audit the artist's account.

[0312] Exemplary User and Enterprise GUIs

[0313] Figures 20A - 20B is a block diagram showing a graphical user interface according to some embodiments. In the illustrated embodiment, Figure 20A contains the GUI displayed by the user application 2010, while Figure 20B contains the GUI displayed by the enterprise application 2030. In some embodiments, Figure 20A and Figure 20B the displayed GUIs are generated by a website rather than by an application. In various embodiments, any of a variety of suitable elements can be displayed, including one or more of the following elements: a turntable (e.g., for controlling volume, energy, etc.), buttons, knobs, display boxes (e.g., for providing update information to the user), etc.

[0314] In Figure 20A the user application 2010 displays a GUI that includes a section 2012 for selecting one or more artist packs. In some embodiments, the pack 2014 may alternatively or additionally include a themed pack or a pack for a specific occasion (e.g., wedding, birthday party, graduation ceremony, etc.). In some embodiments, the number of packs displayed in section 2012 is greater than the number that can be displayed in section 2012 at one time. Therefore, in some embodiments, the user scrolls up and / or down in section 2012 to view one or more packs 2014. In some embodiments, the user can select an artist pack 2014 based on the output music content they want to hear. In some embodiments, for example, an artist pack can be purchased and / or downloaded.

[0315] In the illustrated embodiment, selection element 2016 allows a user to adjust one or more music attributes (e.g., energy level). In some embodiments, selection element 2016 allows a user to add / delete / modify one or more target music attributes. In various embodiments, selection element 2016 may present one or more UI control elements (e.g., control element 830).

[0316] In the illustrated embodiment, selection element 2020 allows a user to have the device (e.g., a mobile device) listen to the environment to determine target music attributes. In some embodiments, the device uses one or more sensors (e.g., a camera, a microphone, a thermometer, etc.) to collect information about the environment after the user selects selection element 2020. In some embodiments, application 2010 also selects or suggests one or more artist packs based on the environmental information collected by the application when the user selects element 2020.

[0317] In the illustrated embodiment, selection element 2022 allows a user to combine multiple artist packs to generate a new rule set. In some embodiments, the new rule set is based on the user selecting one or more packs for the same artist. In other embodiments, the new rule set is based on the user selecting one or more packs for different artists. The user may indicate the weights of different rule sets, e.g., such that a rule set with a high weight has a greater impact on the generated music than a rule set with a low weight. The music generator may combine the rule sets in a variety of different ways, e.g., by switching between rules from different rule sets, averaging the values of rules from multiple different rule sets, etc.

[0318] In the illustrated embodiment, selection element 2024 allows a user to manually adjust one or more rules in a rule set. For example, in some embodiments, the user desires to adjust the music content being generated at a finer level by adjusting one or more rules in the rule set used to generate the music content. In some embodiments, this allows the user of application 2010 to act as their own disc jockey (DJ) by using the controls displayed in the Figure 2 GUI to adjust the rule set used by the music generator to generate the output music content. These embodiments may also allow for finer control of the target music attributes.

[0319] In Figure 20B the enterprise application 2030 displays a GUI that also includes an artist pack selection section 2012 with artist packs 2014. In the illustrated embodiment, the enterprise GUI displayed by application 2030 also includes an element 2016 for adjusting / adding / deleting one or more music attributes. In some embodiments, the display in Figure 20BThe GUI in [description] is used in a business or storefront to generate a specific environment (e.g., for optimizing sales) by generating music content. In some embodiments, an employee uses Application 2030 to select one or more artist packs that have been previously shown to increase sales (e.g., the metadata of a given rule set can indicate the actual experimental results of using the rule set in a real-world context).

[0320] In the illustrated embodiment, Input Hardware 2040 sends information to the application or website that is displaying Enterprise Application 2030. In some embodiments, Input Hardware 2040 is one of the following: a cash register, a heat sensor, a light sensor, a clock, a noise sensor, etc. In some embodiments, the information sent from one or more of the hardware devices listed above is used to adjust the target music attributes and / or the rule set for generating the output music content for a specific environment. In the illustrated embodiment, Selection Element 2038 allows the user of Application 2030 to select one or more hardware devices from which to receive environmental input.

[0321] In the illustrated embodiment, Display 2034 displays environmental data to the user of Application 2030 based on information from Input Hardware 2040. In the illustrated embodiment, Display 2032 shows the changes to the rule set based on the environmental data. In some embodiments, Display 2032 allows the user of Application 2030 to see the changes made based on the environmental data.

[0322] In some embodiments, Figure 20A and Figure 20B the elements shown in [description] are used for theme packs and / or occasion packs. That is, in some embodiments, the user or enterprise using the GUI displayed by Application 2010 and Application 2030 can select / adjust / modify the rule set to generate music content for one or more occasions and / or themes.

[0323] Detailed Example Music Generator System

[0324] Figures 21 - 23 Details of a specific embodiment regarding Music Generator Module 160 are shown. Note that although these specific examples are disclosed for illustrative purposes, they are not intended to limit the scope of the present disclosure. In these embodiments, building music from loops is performed by a client system such as a personal computer, a mobile device, a media device, etc. As Figures 21 - 23For the discussion herein, the term "loop" may be interchangeable with the term "audio file". Generally, as described herein, loops are included within an audio file. Loops can be divided into professionally curated loop packs, which may be referred to as artist packs. Loops can be analyzed for musical attributes and the attributes can be stored as loop metadata. Audio within a constructed track can be analyzed (e.g., in real-time) and filtered to mix and control an output stream. Various feedback can be sent to a server, including explicit feedback such as from a user's interaction with a slider or button and implicit feedback (e.g., generated by sensors based on volume changes, listening length, environmental information, etc.). In some embodiments, control inputs have known effects (e.g., directly or indirectly specifying target musical attributes) and are used by a composition module.

[0325] The following discussion introduces various terms used in reference Figures 21 - 23 In some embodiments, a loop library is a master library of loops that can be stored by a server. Each loop can include audio data and metadata that describes the audio data. In some embodiments, a loop pack is a subset of the loop library. A loop pack can be a pack for a specific artist, for a specific mood, for a specific event type, etc. A client device can download a loop pack for offline listening or download portions of a loop pack on demand, e.g., for online listening.

[0326] In some embodiments, a generated stream is data that specifies the music content that a user hears when using a music generator system. Note that for a given generated stream, the actual output audio signal may vary slightly, e.g., based on the capabilities of the audio output device.

[0327] In some embodiments, a composition module constructs a composition from loops available in a loop pack. The composition module can receive loops, loop metadata, and user input as parameters and can be executed by a client device. In some embodiments, the composition module outputs a performance script that is sent to a performance module and one or more machine learning engines. In some embodiments, the performance script outlines which loops will be played on each track of the generated stream and what effects will be applied to the stream. The performance script can utilize beat-related time to represent the time at which events occur. The performance script can also encode effect parameters (e.g., for effects such as reverb, delay, compression, equalization, etc.).

[0328] In some embodiments, a performance module receives a performance script as input and renders it as a generated stream. The performance module can produce multiple tracks specified by the performance script and mix those tracks into a single stream (e.g., a stereo stream, although the stream can have various encodings in various embodiments, including surround encoding, object-based audio encoding, multichannel stereo, etc.). In some embodiments, when provided with a specific performance script, the performance module will always produce the same output.

[0329] In some embodiments, the analysis module is a server-implemented module that receives feedback information and configures the composition module (e.g., in real time, periodically, based on an administrator command, etc.). In some embodiments, the analysis module uses a combination of machine learning techniques to correlate user feedback with performance scripts and loop library metadata.

[0330] Figure 21 is a block diagram showing an example music generator system including analysis and composition modules. In some embodiments, Figure 21 the system is configured to generate a potentially infinite stream of music that users can directly control the mood and style of. In the illustrated embodiment, the system includes an analysis module 2110, a composition module 2120, a performance module 2130, and an audio output device 2140. In some embodiments, the analysis module 2110 is implemented by a server, and the composition module 2120 and the performance module 2130 are implemented by one or more client devices. In other embodiments, modules 2110, 2120, and 2130 can all be implemented on a client device or can all be implemented on the server side.

[0331] In the illustrated embodiment, the analysis module 2110 stores one or more artist packages 2112 and implements a feature extraction module 2114, a client simulator module 2116, and a deep neural network 2118.

[0332] In some embodiments, the feature extraction module 2114 adds loops to the loop library after analyzing loop audio (although note that some loops may be received with already generated metadata and may not require analysis). For example, raw audio in formats such as wav, aiff, or FLAC can be analyzed to obtain quantifiable music attributes such as instrument classification, pitch transcription, beat timing, tempo, file length, and audio amplitude in multiple frequency bins. The analysis module 2110 can also store more abstract music attributes or mood descriptions of loops, e.g., based on manual tagging by an artist or machine listening. For example, moods can be quantified using multiple discrete categories, where the value range for each category is used for a given loop.

[0333] For example, consider loop A, which is analyzed to determine that notes G2, Bb2, and D2 are used, the first beat starts at 6 milliseconds into the file, the tempo is 122 bpm, the file length is 6483 milliseconds, and the loop has normalized amplitude values of 0.3, 0.5, 0.7, 0.3, and 0.2 in five frequency bins. The artist can label the loop as "funk genre" with the following mood values:

[0334] Beyond Peace Power Joy Sadness Tension High High Low Medium None Low

[0335] The analysis module 2110 can store this information in a database and the client can download a sub - portion of the information, e.g., as a loop pack. Although the artist pack 2112 is shown for illustrative purposes, the analysis module 2110 can provide various types of loop packs to the composition module 2120.

[0336] In the illustrated embodiment, the client simulator module 2116 analyzes various types of feedback to provide feedback information in a format supported by the deep neural network 2118. In the illustrated embodiment, the deep neural network 2118 also receives the performance script generated by the composition module as input. In some embodiments, the deep neural network configures the composition module based on these inputs, e.g., to improve the correlation between the generated music output type and the desired feedback. For example, the deep neural network can periodically push updates to the client device implementing the composition module 2120. Note that the deep neural network 2118 is shown for illustrative purposes and can provide powerful machine - learning performance in the disclosed embodiments, but is not intended to limit the scope of the present disclosure. In various embodiments, various types of machine - learning techniques can be implemented alone or in various combinations to perform similar functions. Note that the machine - learning module can be used to directly implement a rule set (e.g., arrangement rules or techniques) in some embodiments, or can be used to control a module implementing other types of rule sets, e.g., using the deep neural network 2118 in the illustrated embodiment.

[0337] In some embodiments, the analysis module 2110 generates composition parameters for the composition module 2120 to improve the correlation between the desired feedback and the use of specific parameters. For example, actual user feedback can be used to adjust the composition parameters, e.g., to try to reduce negative feedback.

[0338] As an example, consider the case where the module 2110 discovers a correlation between negative feedback (e.g., an explicit low ranking, low - volume listening, short listening time, etc.) and works that use a large number of layers. In some embodiments, the module 2110 uses techniques such as backpropagation to determine that adjusting the probability parameter used to add more tracks will reduce the frequency of this problem. For example, the module 2110 can predict that reducing the probability parameter by 50% will reduce negative feedback by 8% and can determine to perform the reduction and push the updated parameter to the composition module (note that the probability parameter is discussed in detail below, but various parameters of a statistical model can be adjusted similarly).

[0339] As another example, consider a situation where module 2110 discovers that negative feedback is associated with the user setting the mood control to highly tense. It may also discover a correlation between loops with a low-tension label and users who request high tension. In such a case, module 2110 can increase a parameter such that the probability of selecting a loop with a high-tension label when the user requests high-tension music is increased. Thus, machine learning can be based on various information, including combined outputs, feedback information, user control inputs, and so on.

[0340] In the illustrated embodiment, the combination module 2120 includes a partial sequencer 2122, a partial arranger 2124, a technology implementation module 2126, and a loop selection module 2128. In some embodiments, the combination module 2120 organizes and constructs the combined parts based on loop metadata and user control inputs (e.g., mood control).

[0341] In some embodiments, the partial sequencer 2122 sequences different types of parts. In some embodiments, the partial sequencer 2122 implements a finite state machine to continuously output the next type of part during operation. For example, the combination module 2120 can be configured to use different types of parts, such as intros, builds, drops, breakdowns, and bridges, as discussed in further detail below with reference to Figure 23 Furthermore, each part can include multiple sub-parts that define how the music changes throughout the part, e.g., including an intro sub-part, a main content sub-part, and an outro sub-part.

[0342] In some embodiments, the partial arranger 2124 constructs sub-parts according to arrangement rules. For example, one rule can specify an intro by gradually adding tracks. Another rule can specify an intro by gradually increasing the gain of a set of tracks. Another rule can specify slicing a vocal loop to create a melody. In some embodiments, the probability that a loop in the loop library is appended to a track is a function of the current position in the part or sub-part, loops that overlap in time on another track, and user input parameters such as mood variables (which can be used to determine the target attributes of the generated music content). For example, the function can be adjusted by adjusting coefficients based on machine learning.

[0343] In some embodiments, the technical implementation module 2120 is configured to facilitate partial arrangement by adding rules such as those specified by an artist or determined by analyzing the works of a particular artist. "Technical" can describe how a particular artist implements arrangement rules at the technical level. For example, for an arrangement rule that specifies transitioning by gradually adding tracks, one technique can indicate adding tracks in the order of drums, bass, pads, and then vocals, while another technique can indicate adding tracks in the order of bass, pads, vocals, and then drums. Similarly, for an arrangement rule that specifies slicing a vocal loop to create a melody, one technique can indicate slicing the vocals on every second beat and repeating the sliced portion of the loop twice before moving to the next sliced portion.

[0344] In the illustrated embodiment, the loop selection module 2128 selects loops according to the arrangement rules and techniques to be included in the parts of the partial arranger 2124. Once a part is complete, a corresponding performance script can be generated and sent to the performance module 2130. The performance module 2130 can receive parts of the performance script at various interval sizes. This can include, for example, the entire performance script for a particular length of performance, the performance script for each part, the performance script for each sub - part, etc. In some embodiments, the arrangement rules, techniques, or loop selection are statistically implemented, e.g., different methods use different percentages of time.

[0345] In the illustrated embodiment, the performance module 2130 includes a filter module 2131, an effects module 2132, a mixing module 2133, a master module 2134, and an execution module 2135. In some embodiments, these modules process the performance script and generate music data in a format supported by the audio output device 2140. The performance script can specify which loops are to be played, when they should be played, what effects the module 2132 should apply (e.g., on a per - track or per - sub - part basis), what filters the module 2131 should apply, etc.

[0346] For example, the performance script can specify applying a low - pass filter from 1000 to 20000 Hz for 5000 milliseconds on a particular track. As another example, the performance script can specify applying a reverb with a 0.2 wet setting from 5000 to 15000 milliseconds on a particular track.

[0347] In some embodiments, the mixing module 2133 is configured to perform automatic level control on the combined tracks. In some embodiments, the mixing module 2133 uses a frequency domain analysis of the combined tracks to measure frequencies with too much or too little energy and applies gain to the tracks in different frequency bands for uniform mixing. In some embodiments, the master module 2134 is configured to perform multi-band compression, equalization (EQ), or limiting procedures to generate data for final formatting by the execution module 2135. Figure 21 Embodiments of can automatically generate various output music content based on user input or other feedback information, and machine learning techniques can allow for the improvement of the user experience over time.

[0348] Figure 22 is a diagram showing an example enhancement portion of music content according to some embodiments. Figure 21 The system of can compose such portions by applying arrangement rules and techniques. In the example shown, the enhancement portion includes three sub-portions and separate tracks for vocals, pads, drums, bass, and white noise.

[0349] In the example shown, the transition in the sub-portion includes drum loop A, which is also repeated for the main content sub-portion. The transition in the sub-portion also includes bass loop A. As shown, the gain of this portion starts low and increases linearly throughout the portion (although non-linear increases or decreases can be envisioned). In the example shown, the main content and the out-transition sub-portion include various vocal, pad, drum, and bass loops. As described above, the disclosed techniques for automatically sequencing portions, arranging portions, and implementing techniques can generate a near-infinite stream of output music content based on various user-adjustable parameters.

[0350] In some embodiments, the computer system displays an interface similar to Figure 22 and allows the artist to specify the techniques for composing the portions. For example, the artist can create structures such as Figure 22 shown, which can be parsed as code for the composition module.

[0351] Figure 23 is a diagram showing an example technique for arranging portions of music content according to some embodiments. In the embodiment shown, the generated stream 2310 includes multiple portions 2320, each portion including a start sub-portion 2322, a development sub-portion 2324, and a transition sub-portion 2326. In the example shown, the various types of each portion / sub-portion are shown in a table connected by dashed lines. In the embodiment shown, the circular elements are examples of arrangement tools, which can be further implemented using specific techniques as described below. As shown, various combination decisions can be performed pseudo-randomly according to statistical percentages. For example, the type of sub-portion, a specific type or arrangement tool of the sub-portion, or the technique for implementing the arrangement tool can be determined statistically.

[0352] In the example shown, a given section 2320 is one of five types: prelude, build, drop, breakdown, and bridge, each having a different function for controlling the intensity of the section. In this example, a state sub-section is one of three types: slow build, sudden transition, or minimal, each having different behavior. In this example, a development sub-section is one of three types: decrease, transition, or increase. In this example, a transition sub-section is one of three types: fold, fade, or cue. For example, different types of sections and sub-sections can be selected based on rules, or can be selected pseudo-randomly.

[0353] In the example shown, the behavior of different sub-section types is implemented using one or more arrangement tools. For slow build, in this example, a low-pass filter is applied 40% of the time and layers are added 80% of the time. For the transition development sub-section, in this example, the loop is sliced 25% of the time. Various additional arrangement tools are shown, including single-shot, pressure differential beats, applying reverb, adding pads, adding themes, removing layers, and white noise. These examples are included for illustrative purposes and are not intended to limit the scope of the present disclosure. Additionally, for ease of illustration, these examples may not be complete (e.g., actual arrangements may typically involve a greater number of arrangement rules).

[0354] In some embodiments, one or more arrangement tools can be implemented using specific techniques (which can be artist-specified or determined based on an analysis of the artist's content). For example, single-shot can be implemented using sound effects or vocals, loop slicing can be implemented using stuttering or half-slicing techniques, removing layers can be implemented by removing a synthesizer or removing vocals, and white noise can be implemented using fade or pulse functions, etc. In some embodiments, the specific technique selected for a given arrangement tool can be selected according to a statistical function (e.g., removing a layer can remove a synthesizer 30% of the time and can remove the vocals of a given artist 70% of the time). As described above, arrangement rules or techniques can be automatically determined by analyzing existing works, such as using machine learning.

[0355] Example Method

[0356] Figure 24 is a flowchart method for using a ledger according to some embodiments. Figure 24 The method shown can be used in combination with any computer circuit, system, device, element, or component, etc. disclosed herein. In various embodiments, some of the method elements shown can be executed simultaneously in a different order than shown, or can be omitted. Additional method elements can also be executed as needed.

[0357] In 2410, in the illustrated embodiment, the computing device determines playback data indicative of playback characteristics of a music content mix. The mix can include a determined combination of multiple audio tracks (note that the combination of tracks can be determined in real time, e.g., just prior to the output of the current portion of the music content mix, which can be a continuous content stream). This determination can be based on the compositional content mix (e.g., by a server or a playback device such as a mobile phone) or can be received from another device that determines which audio files are included in the mix. The playback data can be stored (e.g., in an offline mode) and can be encrypted. The playback data can be reported periodically or in response to a specific event (e.g., regaining connection to the server).

[0358] In 2420, in the illustrated embodiment, the computing device records in an electronic blockchain ledger data structure information specifying the individual playback data of one or more of the multiple audio tracks in a music content mix. In the illustrated embodiment, the information specifying the individual playback data of a particular audio track includes usage data of the individual audio track and signature information associated with the individual audio track.

[0359] In some embodiments, the signature information is an identifier of one or more entities. For example, the signature information can be a string or a unique identifier. In other embodiments, the signature information can be encrypted or otherwise obfuscated to prevent others from identifying the entity. In some embodiments, the usage data includes at least one of the following: the play time of the music content mix or the number of times the music content mix has been played.

[0360] In some embodiments, the data identifying the individual audio tracks in a music content mix is retrieved from a data store that also indicates operations to be performed in association with the one or more individual audio tracks. In these embodiments, the recording can include recording an indication of proof of execution of the indicated operations.

[0361] In some embodiments, the system determines the remuneration of multiple entities associated with multiple audio tracks based on information specifying the individual playback data recorded in the electronic blockchain ledger.

[0362] In some embodiments, the system determines usage data for a first individual audio track that is not included in the music content mix in its original musical form. For example, the audio track can be modified, used to generate a new audio track, etc., and the usage data can be adjusted to reflect the modification or use. In some embodiments, the system generates a new audio track based on interpolation between vector representations of audio in at least two of the multiple audio tracks, and the usage data is determined based on the distance between the vector representation of the first individual audio track and the vector representation of the new audio track. In some embodiments, the usage data is based on the ratio of the Euclidean distance to the interpolated vector representation to the vectors in at least two of the multiple audio tracks.

[0363] Figure 25 A flowchart method for combining audio files using image representations according to some embodiments. Figure 25 The methods shown can be used in combination with any computer circuits, systems, devices, components, or elements disclosed herein. In various embodiments, some of the method elements shown may be executed simultaneously in a different order than shown, or may be omitted. Additional method elements may also be executed as needed.

[0364] In 2510, in the illustrated embodiment, the computing device generates multiple image representations of multiple audio files, wherein an image representation of a specified audio file is generated based on data in the specified audio file and a MIDI representation of the specified audio file). In some embodiments, the pixel values in the image representation represent the tempo in the audio file, wherein the image representation is compressed in tempo resolution.

[0365] In some embodiments, the image representation is a two-dimensional representation of the audio file. In some embodiments, pitch is represented by rows in the two-dimensional representation, wherein time is represented by columns in the two-dimensional representation and wherein the pixel values in the two-dimensional representation represent tempo. In some embodiments, pitch is represented by rows in the two-dimensional representation, wherein time is represented by columns in the two-dimensional representation and wherein the pixel values in the two-dimensional representation represent tempo. In some embodiments, the pitch axis is brought into two sets of octaves within an eight-octave range, wherein the first 12 rows of pixels represent the first 4 octaves, wherein the pixel values of the pixels determine which of the first 4 octaves is represented, and wherein the second 12 rows of pixels represent the second 4 octaves, wherein the pixel values of the pixels determine which of the second 4 octaves is represented. In some embodiments, odd pixel values along the time axis represent note starts and even pixel values along the time axis represent note durations. In some embodiments, each pixel represents a portion of a beat in the time dimension.

[0366] In 2520, in the illustrated embodiment, the computing device selects multiple audio files based on the multiple image representations.

[0367] In 2530, in the illustrated embodiment, the computing device combines the multiple audio files to generate output music content.

[0368] In some embodiments, one or more creative rules are applied to select multiple audio files based on the multiple image representations. In some embodiments, applying one or more creative rules includes removing pixel values in the image representation that are above a first threshold and removing pixel values in the image representation that are below a second threshold.

[0369] In some embodiments, one or more machine learning algorithms are applied to an image representation for selecting and combining multiple audio files and generating an output music content. In some embodiments, the coherence of harmony and rhythm is tested in the output music content.

[0370] In some embodiments, a single image representation is generated from multiple image representations and a description of texture features is appended to the single image representation from which texture features are extracted from multiple audio files. In some embodiments, the single image representation is stored together with the multiple audio files. In some embodiments, multiple audio files are selected by applying one or more creative rules to the single image representation.

[0371] Figure 26 is a flowchart method for implementing a user-created control element according to some embodiments. Figure 26 The methods shown can be used in combination with any computer circuits, systems, devices, elements, or components etc. disclosed herein. In various embodiments, some of the method elements shown can be executed simultaneously in a different order than shown, or can be omitted. Additional method elements can also be executed as needed.

[0372] In 2610, in the illustrated embodiment, the computing device accesses multiple audio files. In some embodiments, the audio files are accessed from the memory of the computer system, where the user has permissions to the accessed audio files.

[0373] In 2620, in the illustrated embodiment, the computing device generates an output music content by combining music content from two or more audio files using at least one trained machine learning algorithm. In some embodiments, the combination of the music content is determined by at least one trained machine learning algorithm based on the music content within the two or more audio files. In some embodiments, at least one trained machine learning algorithm combines the music content by sequentially selecting music content from the two or more audio files based on the music content within the two or more audio files.

[0374] In some embodiments, at least one trained machine learning algorithm has been trained to select music content for an upcoming beat after a specified time based on the metadata of the music content played up to the specified time. In some embodiments, at least one trained machine learning algorithm has been further trained to select music content for an upcoming beat after a specified time based on the level of a control element.

[0375] In the illustrated embodiment, in 2630, the computing device implements on a user interface control elements created by the user for changing user-specified parameters in the generated output music content, wherein the levels of one or more audio parameters in the generated output music content are determined based on the levels of the control elements, and wherein the relationship between the levels of the one or more audio parameters and the levels of the control elements is based on user input during at least one music playback session. In some embodiments, the levels of the user-specified parameters vary based on one or more environmental conditions.

[0376] In some embodiments, the relationship between the levels of the one or more audio parameters and the levels of the control elements is determined by: playing a plurality of audio tracks during at least one music playback session, wherein the plurality of audio tracks have varying audio parameters; for each audio track, receiving input specifying the user-selected level of the user-specified parameter in the audio track; evaluating the levels of the one or more audio parameters in the audio track for each audio track; and determining the relationship between the levels of the one or more audio parameters and the levels of the control elements based on the correlation between each user-selected level of the user-specified parameter and each evaluated level of the one or more audio parameters.

[0377] In some embodiments, one or more machine learning algorithms are used to determine the relationship between the levels of the one or more audio parameters and the levels of the control elements. In some embodiments, the relationship between the levels of the one or more audio parameters and the levels of the control elements is refined based on user changes in the levels of the control elements during playback of the generated output music content. In some embodiments, metadata from the audio tracks is used to evaluate the levels of the one or more audio parameters in the audio tracks. In some embodiments, the relationship between the levels of the one or more audio parameters and the levels of the user-specified parameters is further based on additional user input during one or more additional music playback sessions.

[0378] In some embodiments, the computing device implements on the user interface at least one additional control element created by the user for changing additional user-specified parameters in the generated output music content, wherein the additional user-specified parameters are sub-parameters of the user-specified parameters. In some embodiments, the generated output music content is modified based on the user's adjustment of the levels of the control elements. In some embodiments, a feedback control element is implemented on the user interface, wherein the feedback control element allows the user to provide positive or negative feedback on the generated output music content during playback. In some embodiments, at least one trained machine algorithm modifies the generation of subsequent generated output music content based on the feedback received during playback.

[0379] Figure 27 is a flowchart method for generating music content by modifying audio parameters according to some embodiments. Figure 27The method shown can be used in combination with any computer circuits, systems, devices, components, or assemblies disclosed herein. In various embodiments, some of the method elements shown may be executed simultaneously in a different order than shown, or may be omitted. Additional method elements may also be executed as needed.

[0380] In 2710, in the illustrated embodiment, the computing device accesses a set of music content. In some embodiments.

[0381] In 2720, in the illustrated embodiment, the computing device generates a first graph of an audio signal of the music content, where the first graph is a graph of audio parameters versus time.

[0382] In 2730, in the illustrated embodiment, the computing device generates a second graph of the audio signal of the music content, where the second graph is a signal graph of audio parameters versus beats. In some embodiments, the second graph of the audio signal has a similar structure to the first graph of the audio signal.

[0383] In 2740, in the illustrated embodiment, the computing device generates new music content from the playback music content by modifying the audio parameters in the playback music content, where the audio parameters are modified based on a combination of the first graph and the second graph.

[0384] In some embodiments, the audio parameters in the first graph and the second graph are defined by nodes that determine attribute changes of the audio signal in the graph. In some embodiments, generating new music content includes: receiving the playback music content; determining a first node in the first graph corresponding to the audio signal in the playback music content; determining a second node in the second graph corresponding to the first node; determining one or more specified audio parameters based on the second node; and modifying one or more attributes of the audio signal in the playback music content by modifying the specified audio parameters. In some embodiments, one or more additional specified audio parameters are determined based on the first node and one or more attributes of an additional audio signal in the playback music content are modified by modifying the additional specified audio parameters.

[0385] In some embodiments, determining one or more audio parameters includes: determining a portion of the second graph to implement the audio parameters based on the position of the second node in the second graph, and selecting an audio parameter from the determined portion of the second graph as one or more specified audio parameters. In some embodiments, modifying one or more specified audio parameters modifies a portion of the playback music content corresponding to the determined portion of the second graph. In some embodiments, the modified attributes of the audio signal in the playback music content include signal amplitude, signal frequency, or a combination thereof.

[0386] In some embodiments, one or more automations are applied to audio parameters, where at least one of the automations is a pre-programmed temporal manipulation of at least one audio parameter. In some embodiments, one or more modulations are applied to audio parameters, where at least one of the modulations multiplicatively modifies at least one audio parameter on top of at least one automation.

[0387] The following numbered clauses set forth various non-limiting embodiments disclosed herein:

[0388] Set A

[0389] A1. A method comprising:

[0390] A computer system generates multiple image representations of multiple audio files, where an image representation of a specified audio file is generated based on data in the specified audio file and a MIDI representation of the specified audio file;

[0391] selects multiple audio files based on the multiple image representations; and

[0392] combines the multiple audio files to generate output music content.

[0393] A2. The method according to any of the preceding clauses in Set A, where pixel values in the image representation represent tempo in the audio file, and where the image representation is compressed in tempo resolution.

[0394] A3. The method according to any of the preceding clauses in Set A, where the image representation is a two-dimensional representation of the audio file.

[0395] A4. The method according to any of the preceding clauses in Set A, where pitch is represented by rows in the two-dimensional representation, where time is represented by columns in the two-dimensional representation, and where pixel values in the two-dimensional representation represent tempo.

[0396] A5. The method according to any of the preceding clauses in Set A, where the two-dimensional representation is 32 pixels wide by 24 pixels high, and where each pixel represents a fraction of a beat in the time dimension.

[0397] A6. The method according to any of the preceding clauses in Set A, where the pitch axis is brought into two sets of octaves within an 8-octave range, where the first 12 rows of pixels represent the first 4 octaves, where the pixel value of a pixel determines which of the first 4 octaves is represented, and where the second 12 rows of pixels represent the second 4 octaves, where the pixel value of a pixel determines which of the second 4 octaves is represented.

[0398] A7. The method according to any of the preceding clauses in Set A, where odd pixel values along the time axis represent note onsets and even pixel values along the time axis represent note durations.

[0399] A8. The method according to any of the foregoing clauses in set A further includes applying one or more creative rules to select a plurality of audio files based on a plurality of image representations.

[0400] A9. The method according to any of the foregoing clauses in set A, wherein applying one or more creative rules includes removing pixel values in the image representation that are higher than a first threshold and removing pixel values in the image representation that are lower than a second threshold.

[0401] A10. The method according to any of the foregoing clauses in set A further includes applying one or more machine learning algorithms to the image representation for selecting and combining a plurality of audio files and generating output music content.

[0402] A11. The method according to any of the foregoing clauses in set A further includes testing the harmony and rhythm coherence in the output music content.

[0403] A12. The method according to any of the foregoing clauses in set A further includes:

[0404] generating a single image representation from a plurality of image representations;

[0405] extracting one or more texture features from a plurality of audio files; and

[0406] attaching a description of the extracted texture features to the single image representation.

[0407] A13. A non-transitory computer-readable medium having instructions stored thereon that are executable by a computing device to perform operations including:

[0408] any combination of operations performed according to the method of any of the foregoing clauses in set A.

[0409] A14. An apparatus, comprising:

[0410] one or more processors; and

[0411] one or more memories having program instructions stored thereon that are executable by the one or more processors to:

[0412] perform any combination of operations performed according to the method of any of the foregoing clauses in set A.

[0413] Set B

[0414] B1. A method, comprising:

[0415] A computer system accessing a set of music content;

[0416] A first diagram of an audio signal of music content generated by a computer system, where the first diagram is a diagram of audio parameters with respect to time;

[0417] A second diagram of an audio signal of music content generated by a computer system, where the second diagram is a signal diagram of audio parameters with respect to beats; and

[0418] The computer system generates new music content from the played-back music content by modifying audio parameters in the played-back music content, where the audio parameters are modified based on a combination of the first diagram and the second diagram.

[0419] B2. The method according to any of the preceding clauses in set B, where the second diagram of the audio signal has a structure similar to the first diagram of the audio signal.

[0420] B3. The method according to any of the preceding clauses in set B, where the audio parameters in the first diagram and the second diagram are defined by nodes in the diagram that determine changes in the attributes of the audio signal.

[0421] B4. The method according to any of the preceding clauses in set B, where generating new music content includes:

[0422] Receiving the played-back music content;

[0423] Determining a first node in the first diagram corresponding to the audio signal in the played-back music content;

[0424] Determining a second node in the second diagram corresponding to the first node;

[0425] Determining one or more specified audio parameters based on the second node; and

[0426] Modifying one or more attributes of the audio signal in the played-back music content by modifying the specified audio parameters.

[0427] B5. The method according to any of the preceding clauses in set B, further including:

[0428] Determining one or more additional specified audio parameters based on the first node; and

[0429] Modifying one or more attributes of the additional audio signal in the played-back music content by modifying the additional specified audio parameters.

[0430] B6. The method according to any of the preceding clauses in set B, where determining one or more audio parameters includes:

[0431] Based on the position of the second node in the second diagram, determining a part of the second diagram to implement the audio parameters; and

[0432] Selecting the audio parameters from the determined part of the second diagram as one or more audio specified parameters.

[0433] B7. A method according to any of the preceding clauses in set B, wherein modifying one or more specified audio parameters modifies a part of the playback music content corresponding to the determined part of the second figure.

[0434] B8. A method according to any of the preceding clauses in set B, wherein the modified attributes of the audio signals in the playback music content include signal amplitude, signal frequency, or a combination thereof.

[0435] B9. A method according to any of the preceding clauses in set B, further comprising applying one or more automations to the audio parameters, wherein at least one of the automations is a pre-programmed time manipulation of at least one audio parameter.

[0436] B10. A method according to any of the preceding clauses in set B, further comprising applying one or more modulations to the audio parameters, wherein at least one of the modulations multiplicatively modifies at least one audio parameter on top of at least one automation.

[0437] B11. A method according to any of the preceding clauses in set B, further comprising a computer system providing heap-allocated memory for storing one or more objects associated with the set of music content, the one or more objects including at least one of the following: audio signals, the first figure, and the second figure.

[0438] B12. A method according to any of the preceding clauses in set B, further comprising storing one or more objects in a list of data structures in the heap-allocated memory, wherein the list of data structures includes a serialized list of links for the objects.

[0439] B13. A non-transitory computer-readable medium having instructions stored thereon that are executable by a computing device to perform operations including:

[0440] Any combination of operations performed according to the method of any of the preceding clauses in set B.

[0441] B14. An apparatus, comprising:

[0442] One or more processors; and

[0443] One or more memories having program instructions stored thereon that are executable by the one or more processors to:

[0444] Perform any combination of operations performed according to the method of any of the preceding clauses in set B.

[0445] Set C

[0446] C1. A method, comprising:

[0447] A computer system accessing a plurality of audio files;

[0448] Generating output music content by using at least one trained machine learning algorithm to combine music content from two or more audio files; and

[0449] Implementing, on a user interface associated with a computer system, control elements created by a user for changing user-specified parameters in the generated output music content, wherein the level of one or more audio parameters in the generated output music content is determined based on the level of the control elements, and wherein the relationship between the level of the one or more audio parameters and the level of the control elements is based on user input during at least one music playback session.

[0450] C2. The method according to any of the preceding clauses in set C, wherein the combination of music content is determined by at least one trained machine learning algorithm based on the music content within two or more audio files.

[0451] C3. The method according to any of the preceding clauses in set C, wherein at least one trained machine learning algorithm combines music content by sequentially selecting music content from two or more audio files based on the music content within the two or more audio files.

[0452] C4. The method according to any of the preceding clauses in set C, wherein at least one trained machine learning algorithm has been trained to select music content for an upcoming beat after a specified time based on the metadata of the music content played up to the specified time.

[0453] C5. The method according to any of the preceding clauses in set C, wherein at least one trained machine learning algorithm has been further trained to select music content for an upcoming beat after a specified time based on the level of the control element.

[0454] C6. The method according to any of the preceding clauses in set C, wherein the relationship between the level of the one or more audio parameters and the level of the control element is determined by:

[0455] Playing multiple audio tracks during at least one music playback session, wherein the multiple audio tracks have varying audio parameters;

[0456] For each audio track, receiving an input specifying the user-selected level of a user-specified parameter in the audio track; evaluating the level of one or more audio parameters in the audio track for each audio track; and

[0457] Determining the relationship between the level of the one or more audio parameters and the level of the control element based on the correlation between each user-selected level of the user-specified parameter and each evaluated level of the one or more audio parameters.

[0458] C7. A method according to any of the preceding clauses in set C, wherein one or more machine learning algorithms are used to determine the relationship between the levels of one or more audio parameters and the level of a control element.

[0459] C8. A method according to any of the preceding clauses in set C, further comprising refining the relationship between the levels of one or more audio parameters and the level of a control element based on user changes in the level of the control element during playback of the generated output music content.

[0460] C9. A method according to any of the preceding clauses in set C, wherein metadata from a sound track is used to evaluate the levels of one or more audio parameters in the sound track.

[0461] C10. A method according to any of the preceding clauses in set C, further comprising the computer system changing the level of a user-specified parameter based on one or more environmental conditions.

[0462] C11. A method according to any of the preceding clauses in set C, further comprising implementing at least one additional control element created by the user on a user interface associated with the computer system for changing an additional user-specified parameter in the generated output music content, wherein the additional user-specified parameter is a sub-parameter of the user-specified parameter.

[0463] C12. A method according to any of the preceding clauses in set C, further comprising accessing an audio file from the memory of the computer system, wherein the user has permission to access the accessed audio file.

[0464] C13. A non-transitory computer-readable medium having instructions stored thereon that are executable by a computing device to perform operations including:

[0465] Any combination of operations performed according to a method according to any of the preceding clauses in set C.

[0466] C14. An apparatus comprising:

[0467] One or more processors; and

[0468] One or more memories having program instructions stored thereon that are executable by the one or more processors to:

[0469] Perform any combination of operations performed according to a method according to any of the preceding clauses in set C.

[0470] Set D

[0471] D1. A method comprising:

[0472] A computer system determines playback data for a music content mix, where the playback data indicates playback characteristics of the music content mix, and where the music content mix includes a determined combination of multiple audio tracks; and

[0473] The computer system records, in an electronic blockchain ledger data structure, information specifying the individual playback data of one or more of the multiple audio tracks in a specified music content mix, where the information specifying the individual playback data of a specified individual audio track includes usage data of the individual audio track and signature information associated with the individual audio track.

[0474] D2. The method according to any of the preceding clauses in set D, where the usage data includes at least one of the following: the play time of the music content mix or the number of times the music content mix is played.

[0475] D3. The method according to any of the preceding clauses in set D, further comprising determining information specifying the individual playback data of a specified individual audio track based on the playback data of the music content mix and data identifying the individual audio tracks in the music content mix.

[0476] D4. The method according to any of the preceding clauses in set D, where data identifying the individual audio tracks in the music content mix is retrieved from a data store that also indicates operations to be performed associated with including one or more individual audio tracks, where the recording includes recording an indication attesting to the performance of the operations.

[0477] D5. The method according to any of the preceding clauses in set D, further comprising identifying one or more entities associated with an individual audio track based on the signature information and data specifying the signature information of multiple entities associated with the multiple audio tracks.

[0478] D6. The method according to any of the preceding clauses in set D, further comprising determining the remuneration of multiple entities associated with the multiple audio tracks based on the information specifying the individual playback data recorded in the electronic blockchain ledger.

[0479] D7. The method according to any of the preceding clauses in set D, where the playback data includes usage data of one or more machine learning modules used to generate the music mix.

[0480] D8. The method according to any of the preceding clauses in set D, further comprising a playback device at least temporarily storing the playback data of the music content mix.

[0481] D9. The method according to any of the preceding clauses in set D, where the stored playback data is transmitted by the playback device periodically or in response to an event to another computer system.

[0482] D10. The method according to any of the preceding clauses in set D, further comprising the playback device encrypting the playback data of the music content mix.

[0483] D11. A method according to any of the foregoing clauses in set D, wherein the blockchain ledger is publicly accessible and immutable.

[0484] D12. A method according to any of the foregoing clauses in set D, further comprising:

[0485] Determining usage data for a first individual audio track not included in a musical content composition in its original musical form.

[0486] D13. A method according to any of the foregoing clauses in set D, further comprising:

[0487] Generating a new audio track by interpolating between vector representations of audio in at least two of a plurality of audio tracks;

[0488] Wherein determining the usage data is based on the distance between the vector representation of the first individual audio track and the vector representation of the new audio track.

[0489] D14. A non-transitory computer-readable medium having instructions stored thereon that are executable by a computing device to perform operations, including:

[0490] Any combination of operations performed according to a method according to any of the foregoing clauses in set D.

[0491] D15. An apparatus, comprising:

[0492] One or more processors; and

[0493] One or more memories having program instructions stored thereon that are executable by the one or more processors to:

[0494] Perform any combination of operations performed according to a method according to any of the foregoing clauses in set D.

[0495] Set E

[0496] E1. A method, comprising:

[0497] A computing system stores data for a plurality of audio tracks specifying a plurality of different entities, wherein the data includes signature information for each of the plurality of musical audio tracks;

[0498] The computing system hierarchically arranges the plurality of audio tracks to generate output musical content;

[0499] Recording in an electronic blockchain ledger data structure information about the audio tracks included in the specified hierarchy and the respective signature information for those audio tracks.

[0500] E2. A method according to any clause in set E, wherein the blockchain ledger is publicly accessible and immutable.

[0501] The method according to any clause in set E further includes:

[0502] Detecting the use of an entity that does not match the signature information for one of the tracks.

[0503] A method includes:

[0504] A computing system stores data specifying multiple tracks for multiple different entities, where the data includes metadata linked to other content;

[0505] The computing system selects and layers multiple tracks to generate output music content;

[0506] Retrieving content to be output together with the output music content based on the metadata of the selected tracks.

[0507] The method according to any clause in set E, wherein the other content includes visual advertisements.

[0508] A method includes:

[0509] A computing system stores data specifying multiple tracks;

[0510] The computing system causes the output of multiple music examples;

[0511] The computing system receives user input that indicates the user's opinion on whether one or more of the music examples exhibit a first music parameter;

[0512] Based on the reception, storing the user's custom rule information;

[0513] The computing system receives user input indicating an adjustment to the first music parameter;

[0514] The computing system selects and layers multiple tracks based on the custom rule information to generate output music content.

[0515] The method according to any clause in set E, wherein the user specifies the name of the first music parameter.

[0516] The method according to any clause in set E, wherein the user specifies one or more target targets for one or more ranges of the first music parameter, and wherein the selection and layering are based on past feedback information.

[0517] The method according to any clause in set E further includes: in response to determining that the selection and layering based on past feedback information do not meet one or more target targets, adjusting multiple other music attributes and determining the effect of adjusting the multiple other music attributes on the one or more target targets.

[0518] Method according to any clause of set E, wherein the target objective is the user's heart rate.

[0519] Method according to any clause of set E, wherein the selection and stratification are performed by a machine learning engine.

[0520] Method according to any clause of set E, further comprising training a machine learning engine, including training a teacher model based on multiple types of user input and training a student model based on received user input.

[0521] A method, comprising:

[0522] A computing system stores data specifying multiple tracks;

[0523] The computing system determines that the audio device is in a first type of environment;

[0524] When the audio device is in the first type of environment, receive user input to adjust a first music parameter;

[0525] The computing system selects and stratifies multiple tracks based on the adjusted first music parameter and the first type of environment to generate output music content through the audio device;

[0526] The computing system determines that the audio device is in a second type of environment;

[0527] When the audio device is in the second type of environment, receive user input to adjust the first music parameter;

[0528] The computing system selects and stratifies multiple tracks based on the adjusted first music parameter and the second type of environment to generate output music content through the audio device.

[0529] Method according to any clause of set E, wherein different music attributes are selected and stratified for adjustment based on the adjustment of the first music parameter when the audio device is in a first type of environment compared to when it is in a second type of environment.

[0530] A method, comprising:

[0531] A computing system stores data specifying multiple tracks;

[0532] The computing system stores data specifying music content;

[0533] Determine one or more parameters of the stored music content;

[0534] The computing system selects and stratifies multiple tracks based on one or more parameters of the stored music content to generate output music content; and

[0535] Cause the output of the output music content and the stored music content such that they overlap in time.

[0536] E16. A method comprising:

[0537] A computing system stores data specifying a plurality of tracks and corresponding track attributes;

[0538] A computing system stores data specifying music content;

[0539] Determine one or more parameters of the stored music content and determine one or more tracks included in the stored music content;

[0540] Based on the one or more parameters, the computing system selects and hierarchically stores the tracks and the determined tracks to generate output music content.

[0541] E17. A method comprising:

[0542] A computing system stores data specifying a plurality of tracks;

[0543] A computing system selects and hierarchically arranges a plurality of tracks to generate output music content;

[0544] Wherein the selection is performed using both:

[0545] A neural network module configured to select future tracks to be used based on most recently used tracks; and

[0546] One or more hierarchical hidden Markov models configured to select a structure for the upcoming output music content and to constrain the neural network module based on the selected structure.

[0547] E18. The method according to any clause in set E, further comprising:

[0548] Training one or more hierarchical hidden Markov models based on positive or negative user feedback regarding the output music content.

[0549] E19. The method according to any clause in set E, further comprising:

[0550] Training a neural network model based on user selections provided in response to multiple music content samples.

[0551] E20. The method according to any clause in set E, wherein the neural network module comprises a plurality of layers including at least one globally trained layer and at least one layer specifically trained based on feedback from a particular user account.

[0552] E21. A method comprising:

[0553] A computing system stores data specifying a plurality of tracks;

[0554] A computing system selects and layers multiple tracks to generate output music content, where the selection is performed using both a globally trained machine learning module and a locally trained machine learning module.

[0555] The method according to any clause in set E further includes using feedback from multiple user accounts to train the globally trained machine learning module, and training the locally trained machine learning module based on feedback from a single user account.

[0556] The method according to any clause in set E, where both the globally trained machine learning module and the locally trained machine learning module provide output information based on user adjustments to music attributes.

[0557] The method according to any clause in set E, where the globally trained machine learning module and the locally trained machine learning module are included in different layers of a neural network.

[0558] A method includes:

[0559] Receiving a first user input indicating a value of a higher-level music composition parameter;

[0560] Automatically selecting and combining audio tracks based on the value of the higher-level music composition parameter, including using multiple values of one or more sub-parameters associated with the higher-level music composition parameter when outputting music content;

[0561] Receiving a second user input indicating one or more values of one or more sub-parameters; and

[0562] Constraining the automatic selection and combination based on the second user input.

[0563] The method according to any clause in set E, where the higher-level music composition parameter is an energy parameter, and where one or more sub-parameters include tempo, number of layers, vocal parameter, or bass parameter.

[0564] A method includes:

[0565] A computing system accesses sheet music information;

[0566] The computing system synthesizes music content based on the sheet music information;

[0567] Analyzing the music content to generate frequency information for multiple frequency bins at different time points;

[0568] Training a machine learning engine, including inputting the frequency information and using the sheet music information as labels for training.

[0569] The method according to any clause in set E further includes:

[0570] Input music content into a trained machine learning engine; and

[0571] Use the machine learning engine to generate musical score information for the music content, where the generated musical score information includes frequency information of multiple frequency bins at different time points.

[0572] E29. A method according to any clause in set E, wherein there is a fixed distance between different time points.

[0573] E30. A method according to any clause in set E, wherein different time points correspond to the beats of the music content.

[0574] E31. A method according to any clause in set E, wherein the machine learning engine includes a convolutional layer and a recurrent layer.

[0575] E32. A method according to any clause in set E, wherein the frequency information includes a binary indication of whether the music content includes the content in the frequency bin at the time point.

[0576] E33. A method according to any clause in set E, wherein the music content includes multiple different musical instruments.

[0577] Although specific embodiments have been described above, these embodiments are not intended to limit the scope of the present disclosure, even if only a single embodiment is described with respect to a specific feature. Unless otherwise stated, the feature examples provided in the present disclosure are illustrative rather than restrictive. The above description is intended to cover such alternatives, modifications, and equivalents as would be apparent to those skilled in the art who benefit from the present disclosure.

[0578] The scope of the present disclosure includes any feature or combination of features (explicit or implicit) disclosed herein, or any generalization thereof, whether or not it alleviates any or all of the problems solved herein. Accordingly, new claims may be presented during the prosecution of this application (or an application claiming its priority). In particular, with reference to the appended claims, features from dependent claims may be combined with features of independent claims, and features from individual independent claims may be combined in any suitable manner, not merely in the specific combinations recited in the appended claims.

Claims

1. A method, comprising: generating, by a computer system, multiple image representations of multiple audio files, wherein an image representation of a specified audio file is generated based on a MIDI representation of playback data for the specified audio file and pitch, time, and velocity data of display notes in the specified audio file; generating a single image representation of the multiple audio files by combining the multiple image representations, wherein the single image representation is an image representation having separate image representations placed adjacent to each other from the multiple image representations without overlap; selecting several audio files based on the multiple image representations combined into the single image representation; and combining the several audio files to generate output music content.

2. The method according to claim 1, wherein, a pixel value in the image representation represents the velocity of a note in the audio file, and wherein the image representation is compressed in terms of velocity resolution.

3. The method according to claim 1, wherein, the image representation is a two-dimensional representation of the audio file.

4. The method according to claim 3, wherein pitch is represented by rows in the two-dimensional representation, time is represented by columns in the two-dimensional representation, and a pixel value in the two-dimensional representation represents the velocity of a note in the audio file.

5. The method according to claim 4, wherein, the two-dimensional representation is 32 pixels wide by 24 pixels high, and wherein each pixel represents a portion of a beat in the time dimension.

6. The method according to claim 4, wherein, the two-dimensional representation has a pitch axis on the y-axis and a time axis on the x-axis to represent coordinates along each axis, wherein the pitch axis is graded into two sets of octaves within an eight-octave range, wherein the first 12 rows of pixels along the pitch axis represent the first four octaves, wherein the pixel value of the pixel coordinates along the pitch axis determines which of the first four octaves is represented on the two-dimensional representation, and wherein the second 12 rows of pixels along the pitch axis represent another four octaves, wherein the pixel value of the pixel coordinates along the pitch axis determines which of the another four octaves is represented on the two-dimensional representation.

7. The method according to claim 4, wherein, the two-dimensional representation has a pitch axis on the y-axis and a time axis on the x-axis to represent coordinates along each axis, wherein odd pixel values along the time axis represent the start of a note in the audio file and even pixel values along the time axis represent the duration of a note in the audio file.

8. The method according to claim 4, further comprising applying one or more creative rules to select the several audio files based on the multiple image representations.

9. The method according to claim 8, wherein, applying one or more creative rules includes removing pixel values in the image representation that are higher than a first threshold and removing pixel values in the image representation that are lower than a second threshold.

10. The method according to claim 1 further comprises applying one or more machine learning algorithms to the image representation for selecting and combining the several audio files and generating the output music content.

11. The method according to claim 10 further comprises testing for harmonic and rhythmic coherence in the output music content.

12. The method according to claim 1 further comprises, after generating the single image representation from the plurality of image representations: extracting one or more texture features from the plurality of audio files; and appending a description of the extracted texture features to the single image representation.

13. A non-transitory computer-readable medium having instructions stored thereon that are executable by a computing device to perform operations, the operations comprising: generating a plurality of image representations of a plurality of audio files, wherein an image representation of a specified audio file is generated based on a MIDI representation of playback data for the specified audio file and pitch, time, and velocity data of the displayed notes in the specified audio file; generating a single image representation of the plurality of audio files by combining the plurality of image representations, wherein the single image representation is an image representation having individual image representations placed adjacent to each other from the plurality of image representations without overlap; selecting several audio files based on the plurality of image representations combined into the single image representation; and combining the several audio files to generate output music content.

14. The non-transitory computer-readable medium according to claim 13, wherein the image representation is a two-dimensional representation of the audio file, where pitch is represented by rows, time is represented by columns, and where the pixel value of a pixel in the image representation represents the velocity of a note from the MIDI representation of the specified audio file.

15. The non-transitory computer-readable medium according to claim 14, wherein each pixel represents a part of a beat in the time dimension.

16. The non-transitory computer-readable medium according to claim 13, the operations further comprising, after generating the single image representation from the plurality of image representations: appending a description of texture features to the single image representation, wherein the texture features are extracted from the plurality of audio files.

17. The non-transitory computer-readable medium according to claim 16 further comprises storing the single image representation together with the plurality of audio files.

18. The non-transitory computer-readable medium according to claim 16 further comprises selecting the several audio files by applying one or more creative rules to the single image representation.

19. An apparatus, comprising: one or more processors; and one or more memories having program instructions stored thereon that are executable by the one or more processors to perform the following operations: Generate multiple image representations of multiple audio files, wherein an image representation of a specified audio file is generated based on a MIDI representation of playback data for the specified audio file and pitch, time, and tempo data of the displayed notes in the specified audio file; Generate a single image representation of the multiple audio files by combining the multiple image representations, wherein the single image representation is an image representation having separate image representations placed adjacent to each other from the multiple image representations without overlapping; Select several audio files based on the multiple image representations combined into the single image representation; and Combine the several audio files to generate output music content.

20. The apparatus according to claim 19, wherein, the program instructions stored on the one or more memories are executable by the one or more processors to play the generated output music content.

Citation Information

Patent Citations

  • Music generator

    US10679596B2

  • Music generator

    US8812144B2

  • Music generator

    US20190362696A1