Generation of music content
The music generation system addresses the limitations of existing music streaming services by using machine learning to customize music content based on user preferences and environmental data, resulting in a unique and engaging music experience.
Patent Information
- Application Number
- JP2024096356
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-08-21
- Filing Date
- 2024-06-14
- Publication Date
- 2025-06-11
- Estimated Expiration
- 2041-02-11
AI Technical Summary
Existing music streaming services struggle to adapt music content to individual user preferences, environment, and actions, leading to user boredom and limited song selection due to licensing agreements and genre limitations.
A music generation system that uses machine learning algorithms, including neural networks, to customize music content based on user inputs, environmental data, and user-defined control elements, allowing for real-time music generation and adaptation.
The system effectively generates custom music content that aligns with user preferences and environmental contexts, providing a unique and engaging music experience while reducing user boredom and expanding music selection beyond genre limitations.
Smart Images

Figure 0007691552000011 
Figure 0007691552000012 
Figure 0007691552000013
Abstract
Description
Technical Field
[0001] The present disclosure relates to audio engineering, and more particularly to the generation of music content.
Background Art
[0002] Streaming music services typically provide music to users over the Internet. Users can subscribe to these services and stream music via a web browser or application. Examples of such services include PANDORA, SPOTIFY, GROOVESHARK, etc. In many cases, users can select a genre of music or a particular artist. Users can typically rate songs (e.g., using a star rating or like / dislike system), and some music services can adjust which songs to stream to a user based on previous ratings. The cost of operating a streaming service (which can include paying royalties for each streamed song) is typically covered by the user's subscription fee and / or advertisements played between songs.
[0003] The selection of songs may be limited by the number of songs written for a particular genre and licensing agreements. Users may become bored of listening to the same songs in a particular genre. Additionally, these services may not be able to adapt the music to a user's preferences, environment, actions, etc.
Brief Description of the Drawings
[0004]
Figure 1
[0005]
Figure 2
[0006]
Figure 3
[0007]
Figure 4
[0008]
Figure 5A
Figure 5B
[0009]
Figure 6
[0010]
Figure 7
[0011]
Figure 8
[0012]
Figure 9
[0013]
Figure 10
[0014]
Figure 11
[0015]
Figure 12
[0016]
Figure 13
[0017]
Figure 14
[0018]
Figure 15
[0019]
Figure 16
[0020]
Figure 17
[0021]
Figure 18
[0022]
Figure 19
[0023]
Figure 20A
Figure 20B
[0024]
Figure 21
[0025]
Figure 22
[0026]
Figure 23
[0027]
Figure 24
[0028]
Figure 25
[0029]
Figure 26
[0030]
Figure 27
DETAILED DESCRIPTION OF THE INVENTION
[0031] The embodiments disclosed in this specification are susceptible to various changes and alternative forms, but specific embodiments are shown by way of example in the drawings and described in detail herein. However, it should be understood that the drawings and their detailed description are not intended to limit the claims to the specific forms disclosed. On the contrary, this application is intended to cover all modifications, equivalents, and alternatives within the spirit and scope of the disclosure of this application, as defined by the appended claims.
[0032] This disclosure includes references to "one embodiment", "a particular embodiment", "some embodiments", "various embodiments", or "an embodiment". The appearances of the phrases "in one embodiment", "in a particular embodiment", "in some embodiments", "in various embodiments", or "in an embodiment" do not necessarily mean the same embodiment. Specific features, structures, or characteristics may be combined in any suitable manner that is not inconsistent with this disclosure.
[0033] In the appended claims, when an element is "configured to" perform one or more tasks, it is specifically intended not to invoke 35 U.S.C. § 112(f) with respect to that claim element. Accordingly, none of the claims of the present application are intended to be construed as having means-plus-function elements. If the applicant intends to include § 112(f) during prosecution, the claim element is described using the "means for" construction for performing the function.
[0034] As used in this specification, the term "based on" is used to describe one or more elements that influence a determination. This term does not exclude the possibility that additional factors may influence the determination. That is, the determination may be based solely on a particular element or on a particular element and other unspecified elements. Consider the phrase "determine A based on B". This phrase defines that B is used to determine A or is a factor that affects the determination of A. This phrase does not presuppose that the determination of A can also be based on other factors such as C. This phrase is intended to include embodiments in which A is determined based only on B.
[0035] As used herein, the phrase "in response to" describes one or more factors that trigger an effect. This phrase does not exclude the possibility that additional factors may affect or otherwise cause the effect. That is, the effect may respond only to these factors or only to a particular factor and other unspecified factors.
[0036] As used in this specification, terms such as "first", "second", etc. are used as labels for the nouns that precede them and do not imply any kind of ordering (e.g., spatial, temporal, logical, etc.) unless otherwise specified. As used in this specification, the term "or" is used inclusively, exclusively, or not at all. For example, the phrase "at least one of x, y, or z" means any one of x, y, and z, as well as any combination thereof (e.g., x and y, but not z). In certain situations, the context of the use of the term "or" may indicate that it is being used in an exclusive sense. For example, "select one of x, y, or z" means that only one of x, y, and z is selected in that instance.
[0037] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the disclosed embodiments. However, one of ordinary skill in the art should recognize that aspects of the disclosed embodiments may be practiced without these specific details. In some instances, well-known structures, computer program instructions, and techniques have not been shown in detail to avoid obscuring the disclosed embodiments.
[0038] U.S. Patent Application No. 13 / 969,372, filed on August 16, 2013 (currently U.S. Patent No. 8,812,144), discusses techniques for generating music content based on one or more music attributes, which is hereby incorporated by reference in its entirety. To the extent that any interpretation is made based on a recognized conflict between the definition of the '372 application and the remainder of the present disclosure, the present disclosure is intended to apply. The music attributes may be input by a user or determined based on environmental information such as ambient noise, lighting, etc. The '372 disclosure discusses techniques for selecting stored loops and / or tracks or generating new loops / tracks, and techniques for layering the selected loops / tracks to generate output music content.
[0039] U.S. Patent Application No. 16 / 420,456, filed on May 23, 2019 (currently U.S. Patent No. 10,679,596), discusses techniques for generating music content, which is hereby incorporated by reference in its entirety. To the extent that any interpretation is made based on a recognized conflict between the definition of the '456 application and the remainder of the present disclosure, the present disclosure is intended to apply. The music may be generated based on user input or using computer-implemented methods. The '456 disclosure discusses various embodiments of music generators.
[0040] The present disclosure generally relates to systems for generating custom music content by selecting and combining audio tracks based on various parameters. In various embodiments, machine learning algorithms (including neural networks such as deep learning neural networks) are configured to generate and customize music content for a particular user. In some embodiments, a user can create their own control elements, and the computing system can be trained to generate output music content according to the intended function of the user-defined control elements. In some embodiments, the playback data of the music content generated by the techniques described herein may be recorded to record and track the use of various music content by different rights holders (e.g., copyright holders). The various techniques described below provide custom music related to different contexts, facilitate generating music according to a particular voice, enable the user to have more control over how the music is generated, generate music that achieves one or more specific goals, and generate music in real time along with other content, etc.
[0041] As used herein, the term "audio file" refers to audio information for music content. For example, the audio information can include data that describes music content as raw audio in a format such as wav, aiff, or FLAC. Properties of the music content may be included in the audio information. The properties may include, for example, quantifiable music properties such as instrument classification, pitch transcription, beat timing, tempo, file length, and audio amplitude in a plurality of frequency bins. In some embodiments, the audio file includes audio information over a particular time interval. In various embodiments, the audio file includes a loop. As used herein, the term "loop" refers to audio information for a single meter over a particular time interval. The various techniques discussed with reference to the audio file can also be performed using a loop that includes a single instrument. The audio file or loop may be played back repeatedly (e.g., a 30-second audio file may be played back 4 times in a row to produce 2 minutes of music content), but the audio file may also be played back once, for example, without being repeated.
[0042] In some embodiments, an image representation of a music file is generated and used to generate music content. The image representation of the audio file is generated based on the data of the audio file and the MIDI representation of the audio file. The image representation may be, for example, a two-dimensional image representation of pitch and rhythm determined from the MIDI representation of the audio file. Rules (e.g., composition rules) can be applied to the image representation to select an audio file to use to generate new music content. In various embodiments, a machine learning / neural network is implemented on the image representation to select an audio file for combination to generate new music content. In some embodiments, the image representation is a compressed (e.g., low resolution) version of the audio file. By compressing the image representation, the search speed for selected music content in the image representation can be improved.
[0043] In some embodiments, the music generator can generate new music content based on various parameter representations of a music file. For example, an audio file typically has an audio signal that can be represented as a graph of a signal (e.g., signal amplitude, frequency, or a combination thereof) over time. However, time-based representations are tempo-dependent. In various embodiments, an audio file can also be represented using a graph of the signal over beats (e.g., a signal graph). The signal graph is tempo-independent and allows for tempo-invariant changes to the audio parameters of the music content.
[0044] In some embodiments, the music generator enables the user to create and label user-defined controls. For example, the user can create controls that the music generator can then be trained to affect the music according to the user's preferences. In various embodiments, the user-defined control device is a high-level control device such as a control device that adjusts mood, intensity, or genre. Such controls are typically subjective measures based on the listener's personal preferences. In some embodiments, the user creates and labels controls for user-defined parameters. The music generator can then play various music files and change the music according to the parameters defined by the user. The music generator can learn and remember the user's preferences based on the user's adjustment of the user-defined parameters. Thus, during subsequent playback, the user-defined controls for the user-defined parameters can be adjusted by the user, and the music generator adjusts the music playback according to the user's preferences. In some embodiments, the music generator can also select music content according to the user's preferences set by the user-defined parameters.
[0045] In some embodiments, the music content generated by the music generator includes music with various stakeholders (e.g., rights holders or copyright holders). In commercial applications that continuously play the generated music content, it is difficult to provide rewards based on the playback of individual audio tracks (files). Accordingly, in various embodiments, techniques for recording playback data of continuous music content are implemented. The recorded playback data can include information related to the playback time of individual audio tracks within the continuous music content that matches the stakeholders for the individual audio tracks. Further, techniques for preventing the falsification of the playback data information may be implemented. For example, the playback data information can be stored in a publicly accessible fixed blockchain ledger.
[0046] This disclosure first describes, with reference to FIGS. 1 and 2, an example of a music generation module and an overall system configuration with multiple applications. Techniques for generating music content from an image representation are described with reference to FIGS. 3-7. Implementation methods for implementing control elements generated by a user are described with reference to FIGS. 8 and 10. Techniques for realizing audio technology are described with reference to FIGS. 11-17. Techniques for recording information about the generated music or elements in a blockchain or other cryptographic ledger are described with reference to FIGS. 18-19. FIGS. 20A-20B show an exemplary application interface.
[0047] Generally, the disclosed music generator includes an audio file, metadata (e.g., information describing the audio file), and a grammar for combining the audio files based on the metadata. The generator can generate a music experience using rules for identifying audio files based on the metadata of the music experience and target characteristics. This can be set up to expand the set of experiences that can be created by adding or changing rules, audio files, and / or metadata. The adjustments can be made manually (e.g., an artist adds new metadata), or the music generator can increase rules / audio files / metadata when monitoring the music experience and desired goals / characteristics within a given environment. For example, listener-defined controls can be implemented to obtain user feedback regarding the goals or features of the music.
[0048] Overview of a Representative Music Generator FIG. 1 is a diagram showing an exemplary music generator according to some embodiments. In the illustrated embodiment, the music generation module 160 receives various information from a plurality of different sources and generates output music content 140.
[0049] In the illustrated embodiment, module 160 accesses the stored audio files and the corresponding attributes 110 of the stored audio files, and combines the audio files to generate output music content 140. In some embodiments, music generation module 160 selects audio files based on those attributes and combines the audio files based on target music attributes 130. In some embodiments, the audio files may be selected based on environmental information 150 combined with target music attributes 130. In some embodiments, environmental information 150 is used indirectly to determine target music attributes 130. In some embodiments, target music attributes 130 are explicitly specified by the user, for example, by specifying a desired energy level, mood, a plurality of parameters, and the like. For example, the listener-defined control device described herein may be implemented to specify the preferences of the listener used as target music attributes. Examples of target music attributes 130 include energy, complexity, and diversity, but more specific attributes (e.g., corresponding to attributes of stored tracks) can also be specified. Generally, when a higher level of target music attributes is specified, lower level music attributes may be determined by the system before generating the output music content.
[0050] Complexity refers to the many audio files, loops, and / or instruments included in the composition. Energy may be related to other attributes or orthogonal to them. For example, changing the key or tempo can affect the energy. However, for a given tempo and key, the energy can be changed by adjusting the type of instrument (e.g., by adding a hi-hat or white noise), complexity, volume, etc. Diversity can represent the amount of temporal variation in the generated music. The variation may be generated with respect to a static set of other music attributes (e.g., by selecting different tracks for a given tempo and key), or by changing the music attributes over time (e.g., by changing the tempo and key more frequently if greater variation is desired). In some embodiments, the target music attributes are considered to exist in a multi-dimensional space, and the music generation module 160 can make course corrections and slowly move through that space based on environmental changes and / or user input, as needed.
[0051] In some embodiments, the attributes stored with the audio file include information about one or more audio files, including tempo, volume, energy, diversity, spectrum, envelope, modulation, periodicity, attack and decay times, noise, artist, instrument, theme, etc. Note that in some embodiments, the audio files are partitioned such that a set of one or more audio files is specific to a particular audio file type (e.g., one sound source or one sound source type).
[0052] In the illustrated embodiment, module 160 accesses the stored rule set 120. In some embodiments, the stored rule set 120 defines rules such as the number of audio files to overlay so that the audio files are played simultaneously (which can correspond to the complexity of the output music), the progression of the primary / secondary keys used when transitioning between audio files or music phrases (which instruments are used together (e.g., instruments having an affinity for each other)), etc., to achieve the targeted music attributes. In another way, the music generation module 160 uses the stored rule set 120 to achieve one or more declarative goals defined by the target music attributes (and / or target environmental information). In some embodiments, the music generation module 160 includes one or more pseudo-random number generators configured to introduce pseudo-random numbers to avoid repetitive output music.
[0053] In some examples, the environmental information 150 includes one or more of lighting information, ambient noise, user information (such as facial expression, body posture, activity level, movement, skin temperature, performance of a particular activity, type of clothing, etc.), temperature information, purchasing activities in the region, time, day of the week, season, number of people present, weather, etc. In some embodiments, the music generation module 160 does not receive / process the environmental information. In some embodiments, the environmental information 150 is received by another module that determines the target music attributes 130 based on the environmental information. The target music attributes 130 can also be derived based on other types of content, such as video data. In some embodiments, the environmental information is used, for example, to adjust one or more stored rule sets 120 to achieve one or more environmental goals. Similarly, the music generator may use the environmental information to adjust the stored attributes of one or more audio files, for example, to indicate the target music attributes or target audience characteristics to which these audio files are particularly relevant.
[0054] As used herein, the term "module" refers to a circuit configured to perform a specified operation, or a non-transitory physical computer-readable medium storing information (e.g., program instructions) for instructing another circuit (e.g., a processor) to perform the specified operation. A module can be implemented in multiple ways, including performing operations as a hardwired circuit or as a memory storing program instructions executable by one or more processors. The hardware circuit can include, for example, custom very large scale integrated circuits or gate arrays, logic chips, transistors, or other discrete components such as off-the-shelf semiconductors. A module can also be implemented in programmable hardware devices such as field programmable gate arrays, programmable array logic, programmable logic devices, etc. A module can also be in any suitable form of a non-transitory computer-readable medium storing program instructions executable to perform the specified operation.
[0055] As used herein, the phrase "music content" means both the music itself (the audible representation of the music) and the information that can be used to play the music. Thus, a piece of music recorded as a file on a storage medium (e.g., a compact disc, a flash drive, etc.) is an example of music content, and the sound generated by outputting this recorded file or other electronic representation (e.g., via a speaker) is also an example of music content.
[0056] The term "music" includes the well-understood meaning, including sounds generated by musical instruments and vocal sounds. Thus, music includes, for example, instrumental performances or recordings, a cappella performances or recordings, and performances or recordings that include both musical instruments and vocals. One of ordinary skill in the art will recognize that "music" does not encompass all vocal recordings. Works that do not include musical attributes such as rhythm, like speeches, news, audiobooks, etc., are not music.
[0057] One of the "contents" of music can be appropriately distinguished from other music contents. For example, a digital file corresponding to the first song can represent the music content of the first song, and a digital file corresponding to the second song can represent the music content of the second song. The phrase "music content" can also be used to distinguish specific intervals within a given musical work, such that different parts of the same song can be regarded as different music contents. Similarly, different tracks (e.g., piano track, guitar track) within a given musical work can also correspond to different music contents. In the context of a potentially infinite stream of generated music, the phrase "music content" can be used to refer to a portion of the stream (e.g., a few measures or minutes).
[0058] The music content generated by embodiments of the present disclosure may be "new music content" that is a combination of musical elements that have not been generated before. The related (but more expansive) concept, namely "original music content", will be described in more detail below. To facilitate the explanation of this term, the concept of a "control entity" for an instance of music content generation is described. Unlike "original music content", "new music content" does not refer to the concept of a control entity. Thus, new music content refers to music content that has not been generated by any entity or computer system.
[0059] Conceptually, the present disclosure refers to several “entities” as controlling particular instances of computer-generated music content. Such entities own (to the extent such rights actually exist) the legal rights (e.g., copyrights) corresponding to the content generated by the computer. In one embodiment, the controlling entity is the individual who creates (e.g., codes various software routines) the music generator implemented by the computer, or the individual who operates (e.g., supplies input to) a particular instance of computer-implemented music generation. In other embodiments, the computer-implemented music generator may be created by a legal entity (e.g., a corporation or other business organization), such as in the form of a software product, a computer system, or a computing device. In some cases, such a computer-implemented music generator can be deployed to many clients. Depending on the terms of the license associated with the distribution of this music generator, the controlling entity can be the creator, the distributor, or the client in various cases. In the absence of such an explicit legal agreement, the controlling entity of the computer-implemented music generator is the entity that facilitates (e.g., supplies input to and thereby operates) a particular computer-generated instance of music content.
[0060] In the context of the present disclosure, the computer generation of "original music content" by a control entity refers to: 1) a combination of music elements that have not been previously generated by either the control entity or others, and 2) a combination of music elements that have been generated by the control entity but were generated in a first instance. Content type 1) is herein referred to as "new music content" and is similar to the definition of "new music content", but the definition of "new music content" refers to the concept of a "control entity" while the definition of "new music content" does not. On the other hand, content type 2 is herein referred to as "proprietary music content". The term "proprietary" here does not imply any implicit legal rights in the content (even if such rights exist), but is merely used to indicate that the music content was originally generated by the control entity. Therefore, a control entity that "plays" music content that was previously and originally generated by the control entity constitutes the "generation of original music content" of the present disclosure. "Non-original music content" with respect to a particular control entity refers to music content that is not the "original music content" of that control entity.
[0061] A portion of the music content can include music components from one or more other music contents. Creating music content in this way is called "sampling" music content and is common in certain music works, particularly music works of a certain genre. Such music content is herein referred to as using terms such as "music content with sampled components", "derivative music content", or other similar terms. In contrast, music content that does not include sampled components is herein represented using terms such as "music content without sampled components", "non-derivative music content", or other similar terms.
[0062] Note that in the application of these terms, when a particular piece of music content is reduced to a sufficient level of granularity, it can be argued that this music content is derivative (substantially all music content is derivative). The terms "derivative" and "non-derivative" are not used in this sense in the present disclosure. With regard to the computer generation of music content, when the computer generation selects a component part from existing music content of an entity other than the control entity (for example, when a computer program selects a specific part of an audio file of a popular artist's work so as to include a part of the music content to be generated), such computer generation is said to be derivative (and results in derivative music content). On the other hand, the generation of music content by a computer is said to be non-derivative (and results in non-derivative music content) when the computer generation does not utilize components of such existing content. Note that a part of the "original music content" may be derivative music content, or may be non-derivative music content.
[0063] Note that the term "derivative" is intended to have a broader meaning in the present disclosure than the term "derivative" used in the U.S. Copyright Law. For example, derivative music content may or may not be a derivative work under the U.S. Copyright Law. The term "derivative" in the present disclosure is not intended to convey a negative meaning, but is merely used to indicate whether a particular fragment of music content "borrows" a part of the content from another work.
[0064] Furthermore, the phrases "new music content", "novel music content", and "original music content" do not encompass music content that is entirely different from combinations of existing music elements. For example, simply changing a few notes of an existing music work does not produce new, novel, or original music content as these phrases are used in this disclosure. Similarly, merely changing the key or tempo, or adjusting the relative intensities of frequencies of an existing music work, does not produce new, novel, or original music content. Further, the phrases new, novel, original music content are not intended to cover music content that is a borderline case between original and non - original content. Instead, these terms are intended to cover music content that is clearly original (herein referred to as "protectable" music content) without doubt, including music content that is subject to copyright protection. Further, as used herein, the term "available" music content refers to music content that does not infringe the copyrights of entities other than the control entity. New and / or original music content is often protected and available. This has advantages in preventing the copying of music content and / or paying royalties for music content.
[0065] The various embodiments discussed herein use rule - based engines, but any of a variety of other types of computer - implemented algorithms may be used for any of the computer learning and / or music generation techniques discussed herein. However, a rule - based approach may be particularly effective in a music context.
[0066] Overview of applications, memory elements, and data that may be used in a typical music system The music generation module can generate music content by interacting with multiple different applications, modules, memory elements, etc. For example, an end user can install one of multiple types of applications for different types of computing devices (e.g., mobile devices, desktop computers, DJ equipment, etc.). Similarly, another type of application may be provided to enterprise users. By interacting with the application during music content generation, the music generator can receive external information used to determine target music attributes in order to update one or more rule sets used to generate music content. In addition to interacting with one or more applications, the music generation module can interact with other modules to receive rule sets, updated rule sets, etc. Finally, the music generation module can access one or more rule sets, audio files, and / or the generated music content stored in one or more memory elements. In addition, the music generation module can store any of the above items in one or more memory elements, which may be accessed locally or via a network (e.g., cloud-based).
[0067] Figure 2 is a block diagram showing an exemplary overview of a system for generating output music content based on inputs from multiple different sources, according to some embodiments. In the illustrated embodiment, system 200 includes a rule module 210, a user application 220, a web application 230, an enterprise application 240, an artist application 250, an artist rule generation module 260, storage of generated music 270, and an external input 280.
[0068] In the illustrated embodiment, user application 220, web application 230, and enterprise application 240 receive external input 280. In some embodiments, external input 280 includes environmental input, target music attributes, user input, sensor input, and the like. In some embodiments, user application 220 is installed on a user's mobile device and includes a graphical user interface (GUI) that enables the user to interact / communicate with rule module 210. In some embodiments, web application 230 is not installed on the user device and is configured to operate within the browser of the user device and can be accessed via a website. In some embodiments, enterprise application 240 is an application used by a large entity to interact with the music generator. In some embodiments, application 240 is used in combination with user application 220 and / or web application 230. In some embodiments, application 240 communicates with one or more external hardware devices and / or sensors to collect information about the surrounding environment.
[0069] In the illustrated embodiment, rule module 210 communicates with user application 220, web application 230, and enterprise application 240 to generate output music content. In some embodiments, music generator 160 is included in rule module 210. Note that rule module 210 may be included in one of applications 220, 230, and 240, or may be installed on a server and accessed via a network. In some embodiments, applications 220, 230, and 240 receive the generated output music content from rule module 210 and play the content. In some embodiments, rule module 210 may request input from applications 220, 230, and 240 regarding, for example, target music attributes and environmental information, and use this data to generate music content.
[0070] In the illustrated embodiment, the stored rule set 120 is accessed by the rule module 210. In some embodiments, the rule module 210 modifies and / or updates the stored rule set 120 based on communication with applications 220, 230, and 240. In some embodiments, the rule module 210 accesses the stored rule set 120 to generate output music content. In the illustrated embodiment, the stored rule set 120 can include rules from the artist rule generation module 260, which will be described in more detail below.
[0071] In the illustrated embodiment, the artist application 250 communicates with the artist rule generation module 260 (which may be part of the same application or cloud-based, for example). In some embodiments, the artist application 250 enables an artist to create a rule set for a particular sound, for example, based on previous compositions. This functionality is further described in U.S. Patent No. 10,679,596. In some embodiments, the artist rule generation module 260 is configured to store the artist rule set generated for use by the rule module 210. A user can purchase rule sets from a particular artist and then use them to generate output music via a particular application. A rule set for a particular artist may sometimes be referred to as a signature pack.
[0072] In the illustrated embodiment, the stored audio files and corresponding attributes 110 are accessed by the module 210 when applying rules to select and combine tracks to generate output music content. In the illustrated embodiment, the rule module 210 stores the generated output music content 270 in a storage element.
[0073] In some embodiments, one or more of the elements of FIG. 2 are implemented on a server and accessed via a network, referred to as a cloud-based implementation. For example, the stored rule set 120, audio file / attributes 110, and generated music 270 can all be stored in the cloud and accessed by module 210. In another example, module 210 and / or module 260 may also be implemented within the cloud. In some embodiments, the generated music 270 is stored in the cloud and watermarked. This allows, for example, not only to generate large amounts of custom music content but also to detect copies of the generated music.
[0074] In some embodiments, one or more of the disclosed modules are configured to generate other types of content in addition to music content. For example, the system can be configured to generate output visual content based on target music attributes, determined environmental conditions, currently used rule sets, and the like. As another example, the system can search a database or the Internet based on the current attributes of the generated music, dynamically change as the music changes, and display a collage of images that match the attributes of the music.
[0075] Examples of machine learning approaches As described herein, the music generation module 160 shown in FIG. 1 can implement various artificial intelligence techniques (e.g., machine learning techniques) to generate output music content 140. In various embodiments, the AI techniques implemented include a combination of more traditional machine learning techniques and knowledge-based systems with deep neural networks (DNNs). This combination can match the strengths and weaknesses of each of these techniques to the challenges specific to music composition and personalization systems. Music content is composed of multiple levels. For example, a piece of music has sections, phrases, melodies, notes, and textures. DNNs are effective in analyzing and generating very high-level and very low-level details of music content. For example, a DNN can classify the texture of a sound as belonging to a clarinet or an electric guitar at a low level, or detect poetry and choruses at a high level. Intermediate levels of details of music content, such as melody construction, orchestration, etc., are more difficult. DNNs are typically excellent at capturing a wide range of styles in a single model, and thus, DNNs can be implemented as generation tools with many expressive ranges.
[0076] In some embodiments, the music generation module 160 utilizes expert knowledge by having human-created audio files (e.g., loops) as the basic units of music content used by the music generation module. For example, the social context of expert knowledge can be incorporated through the selection of rhythm, melody, and texture, and a discovery method can be recorded in a multi-level structure. Different from the separation of DNNs based on structural levels from traditional machine learning, expert knowledge can be applied to any area where the musicality can be improved without imposing overly strong restrictions on the trainability of the music generation module 160.
[0077] In some embodiments, the music generation module 160 uses a DNN to find patterns of how to combine, vertically by overlaying layers of audio on top of each other, and horizontally by concatenating audio files or loops into a sequence. For example, the music generation module 160 can implement a long short-term memory (LSTM) recurrent neural network trained on the Mel-frequency cepstral coefficient (MFCC) audio features of loops used in multitrack audio recordings. In some embodiments, the network is trained to predict and select the audio features of loops for the upcoming beats based on knowledge of the audio features of the previous beats. For example, the network can be trained to predict the audio features of loops for the next 8 beats based on knowledge of the audio features of the last 128 beats. Thus, the network is trained to utilize low-dimensional feature representations to predict the upcoming beats.
[0078] In certain embodiments, the music generation module 160 uses known machine learning algorithms to assemble a sequence of multitrack audio into a musical structure with dynamics of intensity and complexity. For example, the music generation module 160 can implement a hierarchical hidden Markov model that can act like a state machine that makes state transitions with probabilities determined by multiple levels of a hierarchical structure. As an example, a particular type of drop is likely to occur after a buildup section, but is less likely to occur if there is no drum at the end of that buildup. In various embodiments, in contrast to DNN training where what is learned is more opaque, the probabilities can be trained transparently.
[0079] The Markov model may handle larger temporal structures, and thus, presenting an exemplary track may not be easily trainable as the exemplary track may be too long. A feedback control element (such as approval / disapproval of the user interface) can be used to provide feedback to the music at any time. In certain embodiments, the feedback control element is implemented as one of the UI control elements 830 shown in FIG. 8. Next, using the correlation between the music structure and the feedback, a structure model used for composition, such as a transition table or a Markov model, can be updated. This feedback can also be collected directly from measurements of heart rate, sales, or other metrics for which the system can determine a clear classification. The above-mentioned expertise heuristics are as probabilistic as possible and are designed to be trained in the same way as the Markov model.
[0080] In certain embodiments, the training can be performed by a composer or a DJ. Such training may be separate from listener training. For example, training performed by a listener (such as a typical user) may be limited to identifying correct or incorrect classifications based on positive and negative model feedback, respectively. In the case of a composer or a DJ, the training includes hundreds of time steps and details of the layers and volume controls used to more clearly indicate elements that prompt changes in the music content. For example, the training performed by a composer and a DJ can include array prediction training similar to the global training of the above-mentioned DNN.
[0081] In various embodiments, the DNN is trained by incorporating multi-track audio and interface interactions to predict what a DJ or composer will do next. In some embodiments, these interactions can be recorded and used to develop new, more transparent discovery methods. In some embodiments, the DNN receives as input many previous musical measures and utilizes a low-dimensional feature representation along with additional features that describe the changes to the track applied by the DJ or composer as described above. For example, the DNN can receive as input the last 32 musical measures and utilize a low-dimensional feature representation along with additional features to describe the changes to the track applied by the DJ or composer. These changes may include adjustments to the gain of a particular track, applied filters, delays, etc. For example, a DJ can use the same drum loop that is repeated for 5 minutes during a performance, but can gradually increase the gain and delay on the track over time. Thus, the DNN can be trained to predict such gain and delay changes in addition to loop selection. If a loop is not played for a particular instrument (e.g., if the drum loop is not played), the feature set may be all zeros for that instrument, whereby the DNN can learn that predicting all zeros can be a successful strategy that leads to selective layering.
[0082] In some cases, the DJ or composer records a live performance using a mixer and devices such as TRAKTOR (Native Instruments GmbH). These recordings are typically captured at high resolution (e.g., 4-track recording or MIDI). In some embodiments, the system decomposes the recording into component loops and obtains information regarding not only the acoustic quality of each individual loop but also the combinations of loops within the components. By using this information to train a DNN (or other machine learning), the DNN can correlate both the loop configuration (e.g., sequencing, layering, loop timing, etc.) and acoustic quality to inform the music generation module 160 on how to create a music experience similar to the artist's performance without using the actual loops used in the artist's performance.
[0083] Exemplary music generator using an image representation of an audio file Popular music often has combinations of rhythm, texture, and pitch. When generating music note by note for each instrument in a song (as done by a music generator), rules can be implemented based on these combinations to create coherent music. Generally, the stricter the rules, the less room for creative variation and the higher the likelihood of creating a copy of existing music.
[0084] When creating music by combining songs that have already been recorded as audio, it is necessary to create combinations considering multiple invariant combinations of notes for each phrase. However, when drawing from a library of thousands of audio recordings, searching through all possible combinations can be computationally expensive. Additionally, it may be necessary to perform notable comparisons to check for harmonically inconsistent combinations, especially with respect to beats. A new rhythm created by combining multiple files can also be checked against rules regarding the rhythmic composition of the combined phrases.
[0085] It is not always possible to extract the functions necessary to create combinations from audio files. Even if it is possible, extracting the necessary functions from audio files can be computationally expensive. In various embodiments, symbolic audio representations are used for music composition to reduce computational costs. Symbolic audio representations can depend on the texture of the music composer's instrument, the stored rhythm, and the storage of pitch information. A common format for symbolic music representation is MIDI. MIDI includes accurate timing, pitch, and performance control information. In some embodiments, MIDI can be simplified and further compressed through a piano roll representation where notes are shown as bars on a discrete time / pitch graph, typically at a pitch of 8 octaves.
[0086] In some embodiments, a music generator is configured to generate output music content by generating an image representation of a music file and selecting a combination of music based on the analysis of the image representation. The image representation may be a further compressed representation from the piano roll representation. For example, the image representation may be a low-resolution representation generated based on the MIDI representation of the audio file. In various embodiments, composition rules are applied to the image representation to select music content from the music file and combine and generate the output music content. The composition rules can be applied, for example, using a rule-based method. In some embodiments, a machine learning algorithm or model (e.g., a deep learning neural network) is implemented to select and combine audio files to generate the output music content.
[0087] FIG. 3 is a block diagram showing an exemplary music generation system configured to output music content based on the analysis of an image representation of an audio file according to some embodiments. In the illustrated embodiment, system 300 includes an image representation generation module 310, a music selection module 320, and a music generation module 160.
[0088] In the illustrated embodiment, the image representation generation module 310 is configured to generate one or more image representations of an audio file. In a particular embodiment, the image representation generation module 310 receives audio file data 312 and MIDI representation data 314. The MIDI representation data 314 includes the MIDI representation of a particular audio file within the audio file data 312. For example, for a particular audio file within an audio file, the data 312 can have a corresponding MIDI representation within the MIDI representation data 314. In some embodiments having multiple audio files within the audio file data 312, each audio file within the audio file data 312 has a corresponding MIDI representation within the MIDI representation data 314. In the illustrated embodiment, the MIDI representation data 314 is provided to the image representation generation module 310 along with the audio file data 312. However, in some contemplated embodiments, the image representation generation module 310 can generate the MIDI representation data 314 itself from the audio file data 312.
[0089] As shown in FIG. 3, the image representation generation module 310 generates an image representation 316 from the audio file data 312 and the MIDI representation data 314. The MIDI representation data 314 includes pitch, time, and velocity data of the notes associated with the music file, while the music file data 312 includes data for the playback of the music itself. In a particular embodiment, the image representation generation module 310 generates an image representation of the audio file based on the pitch, time, and velocity data from the MIDI representation data 314. The image representation can be, for example, a two-dimensional image representation of the audio file. In a 2D image representation of an audio file, the x-axis represents time (rhythm), the y-axis represents pitch (similar to a piano roll representation), and the pixel value at each x-y coordinate represents velocity.
[0090] The 2D image representation of an audio file can have various image sizes, which are typically selected to correspond to the music structure. For example, in one possible embodiment, the 2D image display is 32 (x-axis) × 24 images (y-axis). An image representation with a width of 32 pixels enables each pixel to represent one quarter of a beat in the time dimension. Thus, an 8-beat music can be represented by an image representation with a width of 32 pixels. This representation may not have sufficient detail to capture the expressive details of the music within the audio file, but the expressive details are retained in the audio file itself, which is used in combination with the image representation by system 300 for generating the output music content. However, with a time resolution of 1 / 4 beat, it can cover a fairly wide range of common pitch and rhythm combination rules.
[0091] Figure 4 shows an example of an image representation 316 of an audio file. The image representation 316 has a width of 32 pixels (time) and a height of 24 pixels (pitch). Each pixel (square) 402 has a value representing the tempo at that time and the pitch within the audio file. In various embodiments, the image representation 316 may be a grayscale image representation of the audio file, where the pixel values are represented by varying the intensity of gray. The change in gray based on the pixel values is small and may be difficult for many people to notice. FIGS. 5A and 5B each show an example of a grayscale image for a melodic image feature representation and a drum beat image feature representation. However, other representations (e.g., color or numerical) may also be considered. In these representations, each pixel can have a plurality of different values corresponding to different musical attributes.
[0092] In certain embodiments, the image representation 316 is an 8-bit representation of an audio file. Thus, each pixel can have 256 possible values. The MIDI representation typically has 128 possible values for velocity. In various embodiments, the details of the velocity values are not as important as the task of selecting an audio file for the combination. Thus, in such embodiments, the pitch axis (y-axis) may be banded so that in each set, the 8-octave range of 4 octaves is covered by 2 sets of octaves. For example, the 8 octaves can be defined as follows: [Number]
[0093] Within these defined ranges of octaves, the rows and values of the pixels determine the octave and velocity of the note. For example, a pixel value of 10 in row 1 represents a note in octave 0 with a velocity of 10, and a pixel value of 74 in row 1 represents a note in octave 2 with a velocity of 10. As another example, a pixel value of 79 in row 13 represents a note in octave 3 with a velocity of 15, and a pixel value of 207 in row 13 represents a note in octave 7 with a velocity of 15. Thus, using the defined range of octaves described above, the first 12 rows (rows 0 - 11) determine the pixel values that represent one of the first 4 octaves (pixel values also determine the velocity of the note), representing the first set of 4 octaves (octaves 0, 2, 4, 6). Similarly, the second 12 rows (rows 12 - 23) represent the second set of 4 octaves (octaves 1, 3, 5, and 7), and the pixel values determine the pixel values that represent one of the second 4 octaves (pixel values also determine the velocity of the note).
[0094] As described above, by banding the pitch axis to cover an 8-octave range, the speed for each octave can be defined by 64 values instead of the 128 values of MIDI representation. Thus, a 2D image representation (e.g., image representation 316) can be more compressed (e.g., lower resolution) than the MIDI representation of the same audio file. In some embodiments, the 64 values may be larger than the values required by system 300 to select musical combinations, so further compression of the image representation may be allowed. For example, by representing the start of a note with an odd pixel value and the duration of the note with a pixel value, the speed resolution can be further reduced to enable compression in the temporal representation. By reducing the resolution in this way, two notes played continuously at the same speed can be distinguished from a single long note based on the odd or even value of the pixel.
[0095] As described above, the compactness of the image representation reduces the size of the file required for the musical representation (e.g., compared to the MIDI representation). Thus, by implementing an image representation of the audio file, the required disk storage can be reduced. Further, the compressed image representation may be stored in a high-speed memory that can quickly search through possible musical combinations. For example, an 8-bit image representation can be stored in the graphics memory on a computer device, thereby enabling large parallel searches to be performed together.
[0096] In various embodiments, the image representations generated for multiple audio files are combined into a single image representation. For example, the image representations of dozens, hundreds, or thousands of audio files can be combined into a single image representation. The single image representation may be a large searchable image that can be used for parallel searching of multiple audio files that make up a single image. For example, a single image can be searched in a similar way to a large texture in a video game that uses software such as MegaTextures (from id Software).
[0097] FIG. 6 is a block diagram showing an exemplary system configured to generate a single image representation according to some embodiments. In the illustrated embodiment, system 600 includes a single image representation generation module 610 and a texture feature extraction module 620. In certain embodiments, the single image representation generation module 610 and the texture feature extraction module 620 are disposed within the image representation generation module 310 shown in FIG. 3. However, the single image representation generation module 610 or the texture feature extraction module 620 may be disposed outside of the image representation generation module 310.
[0098] As shown in the exemplary embodiment of FIG. 6, a plurality of image representations 316A-N are generated. The image representations 316A-N may be the number of individual image representations for the number of N individual audio files. The single image representation generation module 610 can combine the individual image representations 316A-N into a single combined image representation 316. In some embodiments, the individual image representations combined by the single image representation generation module 610 include individual image representations for different devices. For example, different musical instruments in an orchestra can be represented by individual image representations, which are then combined into a single image representation for music search and selection.
[0099] In certain embodiments, the individual image representations 316A-N are combined with the individual image representations 316 arranged adjacent to each other without overlap. Thus, the single image representation 316 is a complete dataset representation of all the individual image representations 316A-N without data loss (e.g., without data from one image representation change data for another image representation). FIG. 7 shows an example of a single image representation 316 of a plurality of audio files. In the illustrated embodiment, the single image representation 316 is a combined image generated from the individual image representations 316A, 316B, 316C, and 316D.
[0100] In some embodiments, a single image representation 316 has texture features 622 added thereto. In the illustrated embodiment, the texture features 622 are added to a single image representation 316 as a single row. Referring to FIG. 6, the texture features 622 are determined by a texture feature extraction module 620. The texture features 622 may include, for example, the musical instrument texture of the music within an audio file. For example, the texture features may include features from different musical instruments such as drums, stringed instruments, and the like.
[0101] In certain embodiments, the texture feature extraction module 620 extracts texture features 622 from the audio file data 312. The texture feature extraction module 620 may implement, for example, a rule-based method, a machine learning algorithm or model, a neural network, or other feature extraction techniques to determine texture features from the audio file data 312. In some embodiments, the texture feature extraction module 620 may extract texture features 622 from an image representation 316 (e.g., a plurality of image representations or a single image representation). For example, the texture feature extraction module 620 may implement image-based analysis (such as an image-based machine learning algorithm or model) to extract texture features 622 from the image representation 316.
[0102] The addition of texture features 622 to the single-image representation 316 provides additional information to the single-image representation that is not typically available in the MIDI representation or piano roll representation of an audio file. In some embodiments, the rows having texture features 622 in the single-image representation 316 (shown in FIG. 7) may not need to be human-readable. For example, the texture features 622 may only need to be machine-readable for implementation in a music generation system. In certain embodiments, the texture features 622 are added to the single-image representation 316 and used for image-based analysis of the single-image representation. For example, the texture features 622 may be used by an image-based machine learning algorithm or model used for music selection, as described below. In some embodiments, the texture features 622 may be ignored during music selection, for example, in rule-based selection as described below.
[0103] Referring back to FIG. 3, in the illustrated embodiment, an image representation 316 (e.g., a plurality of image representations or a single image representation) is provided to the music selection module 320. The music selection module 320 can select an audio file or a portion of an audio file that is combined within the music generation module 160. In certain embodiments, the music selection module 320 applies a rule-based method to search for and select an audio file or a portion of an audio file for combination by the music generation module 160. As shown in FIG. 3, the music selection module 320 accesses rules for the rule-based method from the stored rule set 120. For example, the rules accessed by the music selection module 320 may include, but are not limited to, rules for search and selection such as composition rules and note combination rules. Applying the rules to the image representation 316 can be performed using graphics processing available on a computer device.
[0104] For example, in various embodiments, the note combination rules can be expressed as calculations of vectors and matrices. Graphics processing units are typically optimized for performing vector and matrix calculations. Note that, for example, the intervals between one pitch step can typically be inconsistent and may often be avoided. Such notes can be found by additionally searching for adjacent pixels in a layered image (or segment of a large image) based on the rules. Thus, in various embodiments, the disclosed module can call a kernel to execute all or part of the disclosed operations on the graphics processor of a computer device.
[0105] In some embodiments, the pitch banding in the above-described image representation enables the use of graphics processing for the embedding of high-pass or low-pass filtering of audio. Removing (e.g., filtering out) pixel values below a threshold can simulate high-pass filtering, and removing pixel values above a threshold can simulate low-pass filtering. For example, in the above example of banding, removing (filtering out) pixel values less than 64 with a filter can have a similar effect to applying a high-pass filter with a shelf at B1 by removing octaves 0 and 1 in the example. Thus, the use of filters on each audio file can be efficiently simulated by applying rules to the image representation of the audio file.
[0106] In various embodiments, when overlaying music files to create music, the pitch of a particular audio file can be changed. Changing the pitch can widen the range of successful combinations and potentially widen the search space for combinations. For example, each audio file can be tested with 12 different pitch shift keys. Optimized search for these combinations may be possible by offsetting the row order of the image representation when analyzing the image and adjusting the octave shift as needed.
[0107] In certain embodiments, the music selection module 320 implements a machine learning algorithm or model on the image representation 316 to search for and select an audio file or a portion of an audio file for combination by the music generation module 160. The machine learning algorithm / model may include, for example, a deep learning neural network or other machine learning algorithm that classifies images based on training of the algorithm. In such embodiments, the music selection module 320 includes one or more machine learning models that are trained based on combinations and sequences of audio files that provide desired music characteristics.
[0108] In some embodiments, the music selection module 320 includes a machine learning model that continuously learns during the selection of output music content. For example, the machine learning model can receive user input or other input that reflects characteristics of the output music content that can be used to adjust classification parameters executed by the machine learning model. Similar to rule-based methods, the machine learning model can be implemented using a graphics processing unit on a computer device.
[0109] In some embodiments, the music selection module 320 implements a combination of a rule-based method and a machine learning model. In one intended embodiment, the machine learning model is trained to find combinations of audio files and image representations to start the search for music content and combine if the search is performed using a rule-based method. In some embodiments, the music selection module 320 tests the harmonic and rhythm rule coherence in the music selected for combination by the music generation module 160. For example, the music selection module 320 can test the harmony and rhythm in the selected audio file 322 before providing the selected audio file to the music generation module 160, as described below.
[0110] In the exemplary embodiment of FIG. 3, as described above, the music selected by the music selection module 320 is provided to the music generation module 160 as the selected music file 322. The selected audio file 322 may include a complete or partial audio file that is combined by the music generation module 160 to generate the output music content 140, as described herein. In some embodiments, the music generation module 160 accesses the stored rule set 120 to search for rules applied to the selected audio file 322 to generate the output music content 140. The rules retrieved by the music generation module 160 may be different from the rules applied by the music selection module 320.
[0111] In some embodiments, the selected audio file 322 includes information for combining the selected audio files. For example, the machine learning model implemented by the music selection module 320 can provide instructions in the output that describe how to combine the music content in addition to the selection of music to combine. These instructions are then provided to the music generation module 160 and implemented by the music generation module to combine the selected audio files. In some embodiments, the music generation module 160 tests the harmony and rhythm rule coherence before finalizing the output music content 140. Such tests can be performed in addition to or instead of the tests performed by the music selection module 320.
[0112] Exemplary Control for Music Content Generation In various embodiments, as described herein, the music generation system is configured to automatically generate output music content by selecting and combining audio tracks based on various parameters. As described herein, a machine learning model (or other AI technology) is used to generate music content. In some embodiments, the AI technology is implemented to customize music content for a particular user. For example, the music generation system can implement various types of adaptive control for personalizing music generation. By personalizing music generation, in addition to content generation by the AI technology, content control by the composer or listener becomes possible. In some embodiments, a user creates their own control element, and this control element can be trained (e.g., using AI technology) so that the music generation system generates output music content according to the intended function of the control element created by the user. For example, a user can create a control element that trains the music generation system to affect the music according to the user's preferences.
[0113] In various embodiments, the control elements created by the user are high-level controls such as controls that adjust mood, intensity, or genre. Such user-generated control elements are typically subjective metrics based on the individual preferences of the listener. In some embodiments, the user labels the control elements created by the user and defines the parameters specified by the user. The music generation system can play various music contents, and the user can use the control elements to change the user-specified parameters within the music content. The music generation system can learn and remember how the user-defined parameters change the audio parameters within the music content. Thus, during subsequent playback, the control elements created by the user can be adjusted by the user, and the music generation system adjusts the audio parameters in music playback according to the adjustment level of the parameters specified by the user. In some contemplated embodiments, the music generation system can also select music content according to the user's preferences set by the parameters specified by the user.
[0114] FIG. 8 is a block diagram showing an exemplary system configured to implement user-generated controls in music content generation according to some embodiments. In the illustrated embodiment, system 800 includes a music generation module 160 and a user interface module 820. In various embodiments, the music generation module 160 implements the techniques described herein to generate output music content 140. For example, the music generation module 160 can access the stored audio file 810 and generate the output music content 140 based on the stored rule set 120.
[0115] In various embodiments, the music generation module 160 changes music content based on input from one or more UI control elements 830 implemented in the UI module 820. For example, the user can adjust the level of the control element 830 during interaction with the UI module 820. Examples of control elements include, but are not limited to, sliders, dials, buttons, or knobs. The level of the control element 830 sets the control element level 832 provided to the music generation module 160. The music generation module 160 can then change the output music content 140 based on the control element level 832. For example, the music generation module 160 can implement AI techniques to change the output music content 140 based on the control element level 830.
[0116] In certain embodiments, one or more of the control elements 830 are user-defined control elements. For example, the control element may be defined by a composer or listener. In such embodiments, the user can create and label a UI control element that specifies the parameters the user wishes to implement to control the output music content 140 (e.g., the user creates a control element to control the parameters specified by the user in the control output music content 140).
[0117] In various embodiments, the music generation module 160 can be trained or learned to affect the output music content 140 in a specific way based on input from control elements generated by the user. In some embodiments, the music generation module 160 is trained to change audio parameters within the output music content 140 based on the level of the user-generated control elements set by the user. Training the music generation module 160 can include, for example, determining the relationship between the audio parameters within the output music content 140 and the level of the control elements created by the user. Next, the relationship between the audio parameters within the output music content 140 and the level of the user-generated control elements can be utilized by the music generation module 160 to change the output music content 140 based on the input level of the user-generated control elements.
[0118] Figure 9 shows a flowchart of a method for training the music generation module 160 based on control elements created by a user, according to some embodiments. Method 900 begins with the user creating and labeling control elements in 910. For example, as described above, the user can create and label a UI control element for controlling user-specified parameters within the output music content 140 generated by the music generation module 160. In various embodiments, the label of the UI control element describes the user-specified parameter. For example, the user can label the control element "attitude" to specify that the user desires to control the (user-defined) attitude in the generated music content.
[0119] After creating the UI control element, method 900 continues the playback session 915. The playback session 915 can be used to train a system (e.g., music generation module 160) on how to change audio parameters based on the level of the UI control element created by the user. In the playback session 915, an audio track is played at 920. The audio track may be a loop or sample of music from an audio file stored on or accessed by the device.
[0120] At 930, the user provides input regarding his / her interpretation of user-specified parameters within the played audio track. For example, in certain embodiments, the user listens to the audio track and is requested to select the level of user-specified parameters that the user believes describe the music within the audio track. The level of the user-specified parameters can be selected, for example, using the control elements created by the user. This process can be repeated for multiple audio tracks within the playback session 915 to generate multiple data points for the level of the user-specified parameters.
[0121] In some contemplated embodiments, the user may be requested to listen to multiple audio tracks at once and comparatively evaluate the audio tracks based on user-defined parameters. For example, in an example of a control generated by a user that defines "attributes", the user can listen to multiple audio tracks and select which audio track has more "attributes" and / or which audio track has fewer "attributes". Each selection made by the user may be a data point for the level of the user-specified parameters.
[0122] After playback session 915 is completed, the levels of the audio parameters within the audio track from the playback session are evaluated at 940. Examples of audio parameters include, but are not limited to, volume, tone, bass, treble, reverb, etc. In some embodiments, the levels of the audio parameters within the audio track are evaluated when the audio track is played (e.g., during playback session 915). In some embodiments, the audio parameters are evaluated after playback session 915 has ended.
[0123] In various embodiments, the audio parameters within the audio track are evaluated from the metadata of the audio track. For example, an audio analysis algorithm can be used to generate the metadata of the audio track or symbolic music data (e.g., MIDI), which may be a short, pre-recorded music file. The metadata may include, for example, the pitch present in the recording, the start count per beat, the ratio for notes without pitch, the volume level, and other quantifiable characteristics of the sound.
[0124] At 950, the correlation between the level of the user-selected parameter for the user-specified parameter and the audio parameter is determined. Since the user-selected level for the user-specified parameter corresponds to the level of the control element, the correlation between the user-selected level for the user-specified parameter and the audio parameter can be utilized to define the relationship between the levels of one or more audio parameters and the levels of the control elements within 960. In various embodiments, the correlation relationship between the level of the parameter selected by the user and the audio parameter, and the relationship between the levels of one or more audio parameters and the levels of the control elements are determined using AI techniques (e.g., regression models or machine learning algorithms).
[0125] Returning to FIG. 8, the relationship between the levels of one or more audio parameters and the level of the control element may then be realized by the music generation module 160 to determine how to adjust the audio parameters within the output music content 140 based on the input of the control element level 832 received from the control element 830 generated by the user. In a particular embodiment, the music generation module 160 implements a machine learning algorithm to generate the output music content 140 based on the input and relationship of the control element level 832 received from the control element 830 created by the user. For example, the machine learning algorithm can analyze how the metadata description of the audio track changes through the recording. The machine learning algorithm may include, for example, a neural network, a Markov model, or a dynamic Bayesian network.
[0126] As described herein, the machine learning algorithm can be trained to predict the metadata of the upcoming music fragment when the metadata of the music up to that point is provided. The music generation module 160 can then implement the prediction algorithm by searching a pool of pre-recorded audio files for those having the properties closest to the predicted metadata of what is to come next. Selecting the closest matching audio file to play next helps create output music content in which the music properties progress sequentially, similar to the recordings of the examples on which the prediction algorithm was trained.
[0127] In some embodiments, parametric control of the music generation module 160 using a prediction algorithm may be included in the prediction algorithm itself. In such embodiments, as inputs to the algorithm, along with music metadata, several predefined parameters can be used, and the prediction varies based on these parameters. Alternatively, parametric control can be applied to the prediction to modify them. As an example, generative synthesis is performed by sequentially selecting the nearest music fragments predicted to come next by the prediction algorithm and appending the audio of the file from end to end. At some point, the listener can increase the control element level (such as the start control element for each beat), and the output of the prediction model is changed by increasing the start data field for each predicted beat. When selecting the next audio file to add to the composition, in this scenario, there is a higher likelihood of selecting an audio file with large per-beat properties.
[0128] In various embodiments, a generation system such as the music generation module 160 that utilizes metadata descriptions of music content can use hundreds or thousands of data fields within the metadata of each music fragment. To provide more diversity, multiple simultaneous tracks characterized by different sound sources and sound types may be used. In such cases, the prediction model has thousands of data fields representing the properties of the music, each having an apparent effect on the listening experience. In such cases, an interface for changing each data field of the output of the prediction model can be used for the listener to control the music, generating thousands of control elements. Alternatively, multiple data fields can be combined and exposed as a single control element. The more the music properties are affected by a single control element, the more abstract the control element becomes from specific music properties, and the more subjective the labeling of these controls becomes. In this way, primary control elements and sub-parameter control elements (described below) can be implemented for dynamic and individualized control of the output music content 140.
[0129] As described herein, a user can specify their control element and train the music generation module 160 with respect to a method of operating based on user adjustment of the control element. This process reduces bias and complexity, and the data fields may be completely hidden from the listener. For example, in some embodiments, the listener is provided with a user-generated control element on the user interface. The listener is then presented with a short music clip and asked to set the level of the control element that is thought to best represent the music they heard. By repeating this process, a plurality of data points can be generated that can be used to regressively model the desired effect of control over music. In some embodiments, these data points can be added as additional inputs in a prediction model. The prediction model can then attempt to predict the musical properties that generate a composition sequence similar to the trained sequence while also matching the expected behavior of the control element set at a particular level. Alternatively, a control element mapper in the form of a regression model may be used to map the prediction modifier to the control element without retraining the prediction model.
[0130] In some embodiments, training for a given control element may include both global training (e.g., based on feedback from multiple user accounts) and local training (e.g., based on feedback from the current user account). In some embodiments, a set of control elements specific to a subset of musical elements provided by the composer may be created. For example, a scenario may include an artist creating a loop pack and then training the music generation module 160 using examples of performance or synthesis previously created using these loops. Patterns in these examples can be modeled with regression or neural network models and used to create rules for the construction of new music with similar patterns. These rules can be parameterized and exposed as control elements for the composer to manually change offline before starting to use the music generation module 160, or for the listener to adjust while listening. Examples that the composer feels are opposite to the desired effect of the control can also be used for negative reinforcement.
[0131] In some embodiments, in addition to utilizing exemplary musical patterns, the music generation module 160 can generate musical patterns that correspond to input from the composer before the listener begins to hear the generated music. The composer can do this through direct feedback (described below), for example, tapping a pro control element for positive reinforcement of the pattern or a contra control element for negative reinforcement.
[0132] In various embodiments, the music generation module 160 can enable a composer to create their own sub-parameter control elements of the control elements learned by the music generation module, as described below. For example, a control element for "intensity" may be created as a primary control element from learned patterns regarding the number of expressed notes per beat and the texture quality of the instruments during performance. The composer can then create two sub-parameter control elements by selecting patterns related to the expression of notes, such as a "rhythm intensity" control element and a "texture intensity" control element for the texture pattern. Examples of sub-parameter control elements include control elements for vocals, intensity in a specific frequency range (e.g., bass), complexity, tempo, and the like. These sub-parameter control elements can be used together with more abstract control elements such as energy (e.g., primary control elements). These composer skill control elements may be trained for the music generation module 160 by the composer, similar to the user-created controls described herein.
[0133] As described herein, training the music generation module 160 to control audio parameters based on inputs from user-created control elements enables implementing individual control elements for different users. For example, one user can associate an increase in attitude with an increase in base content, and another user can associate an increase in attitude with a specific type of vocal or a specific tempo range. The music generation module 160 can change audio parameters for different specifications of attitude based on the training of the music generation module for a particular user. In some embodiments, the individualized controls can be used in combination with global rules or control elements that are implemented in the same way for many users. The combination of overall feedback or control and local feedback or control can provide high-quality music production that provides special controls for the individuals involved.
[0134] In various embodiments, as shown in FIG. 8, one or more UI control elements 830 are implemented within the UI module 820. As described above, a user can use the control elements 830 during interaction with the UI module 820 to adjust the control element level 832 and change the output music content 140. In certain embodiments, one or more of the control elements 830 are system-defined control elements. For example, the control elements may be defined as parameters controllable by the system 800. In such embodiments, a user can adjust the system-defined control elements to change the output music content 140 according to the parameters defined by the system.
[0135] In certain embodiments, system-defined UI control elements (e.g., knobs or sliders) enable a user to control the abstract parameters of the output music content 140 automatically generated by the music generation module 160. In various embodiments, the abstract parameters act as primary control element inputs. Examples of abstract parameters include, but are not limited to, intensity, complexity, mood, genre, and energy level. In some embodiments, the intensity control element may adjust the number of built-in low-frequency loops. The complexity control element can guide the number of overlaid tracks. Other adjustment elements, such as the mood adjustment element, vary among various attributes, from calm to happy, and can affect, for example, the musical key being played.
[0136] In various embodiments, system-defined UI control elements (e.g., knobs or sliders) enable a user to control the energy level of output music content 140 automatically generated by music generation module 160. In some embodiments, the label of the control element (e.g., "Energy") may vary in size, color, or other characteristics to reflect user input that adjusts the energy level. In some embodiments, when a user adjusts a control element, the current level of the control element may be output until the user releases the control element (e.g., releases a mouse click or removes a finger from a touch screen).
[0137] Energy may be an abstract parameter related to multiple more specific music attributes, as defined by the system. As an example, energy may be related to tempo in various embodiments. For example, a change in energy level may be associated with a tempo change of a selected number of beats per minute (e.g., ~6 beats per minute). In some embodiments, within a given range for a certain parameter (e.g., tempo), music generation module 160 can explore musical variations by changing other parameters. For example, music generation module 160 can generate buildups and drops, create tension, vary the number of tracks that are layered simultaneously, change keys, add or remove vocals, add or remove bass, play different melodies, etc.
[0138] In some embodiments, one or more sub-parameter control elements are implemented as control element 830. The sub-parameter control elements may enable more specific control of attributes incorporated into primary control elements such as energy control elements. For example, the energy control element can change the number of percussion layers used and the amount of vocals, but separate control elements can directly control these sub-parameters, so not all control elements necessarily act independently. In this way, the user can select the level of control specificity desired. In some embodiments, the sub-parameter control elements may be implemented for the control elements generated by the user as described above. For example, the user can create and label control elements that specify sub-parameters of other user-specified parameters.
[0139] In some embodiments, the user interface module 820 allows the user the option to expand the UI control element 830 and presents one or more sub-parameter user control elements. Additionally, a particular artist can provide attribute information used to guide music synthesis under user control of high-level control elements (e.g., an energy slider). For example, the artist can provide an "artist pack" with tracks from that artist and rules for music composition. The artist can use the artist interface to provide values for the user control elements of the sub-parameters. For example, a DJ has rhythms and drums as control elements exposed to the user, allowing the listener to incorporate some rhythms and drums. In some embodiments, as described herein, an artist or user can generate their own custom control elements.
[0140] In various embodiments, a human-in-the-loop generation system can be used to generate artifacts with the help of human intervention and control, potentially improving the quality and fitness of the generated music for personal purposes. In some embodiments of the music generation module 160, by controlling the generation process via the interface control element 830 implemented in the UI module 820, a listener can become a listener-composer. The design and implementation of these control elements can affect the balance between the roles of the individual listener and composer. For example, highly detailed and technical control elements reduce the influence of the generation algorithm and give more creative control to the user's hand, while requiring more practical interaction and technical skills for management.
[0141] Conversely, a higher level of control elements can reduce the required effort and interaction time while reducing creative control. For example, for individuals who desire a more listener-like role, primary control elements may be preferred as described herein. The primary control elements can be based on abstract parameters such as mood, intensity, or genre. These abstract music parameters are often subjective scales that are individually interpreted. For example, in many cases, the listening environment affects how a listener perceives music. Thus, music that a listener might call "relaxing" at a party may be too energetic and intense for a meditation session.
[0142] In some embodiments, one or more UI control elements 830 are implemented to receive user feedback regarding the output music content 140. The user feedback control elements can include, for example, star ratings, likes / dislikes, etc. In various embodiments, the user feedback may be used to train the system to the specific tastes and / or more global tastes of the users applied to multiple users. In embodiments having like / dislike (e.g., positive / negative) feedback, the feedback is binary. Binary feedback including strong positive and strong negative responses can be effective in providing positive and negative reinforcement for the function of the control element 830. In some contemplated embodiments, the input from the like / dislike control element can be used to control the output music content 140 (e.g., the like / dislike control element is used to control the output itself). For example, the like control element can be used to change the maximum number of repeats of the currently playing music content 140.
[0143] In some embodiments, the counter for each audio file tracks the number of times a section of that audio file (e.g., an 8 - beat segment) has been recently played. Once a file is used beyond a desired threshold, a bias can be applied to that selection. This bias may gradually return to zero over time. Along with rule - defined music sections (e.g., build - up, drop - down, breakdown, intro, sustain) that set desired functions of the music, this iteration counter and bias can be used to shape the music into segments with coherent themes. For example, the music generation module 160 can increment the counter on the press of the opposite, so that the audio content of the output music content 140 is encouraged to change earlier without impairing the music function of the section. Similarly, the music generation module 160 can decrement the counter by pressing the approval so that the audio content of the output music content 140 does not deviate from the repetition for a longer period. Before reaching the threshold and before the bias is applied, other machine - learning and rule - based mechanisms within the music generation module 160 can guide the selection of other audio content.
[0144] In some embodiments, the music generation module 160 is configured to determine various context information (e.g., the environmental information 150 shown in FIG. 1) before and after receiving user feedback. For example, upon receiving an "approval" instruction from the user, the music generation module 160 can determine the time, location, device speed, biometric data (e.g., heart rate), etc. from the environmental information 150. In some embodiments, this context information may be used to train a machine - learning model to generate music that the user prefers in various different contexts (e.g., the machine - learning model recognizes the context).
[0145] In various embodiments, the music generation module 160 determines the current type of environment and performs different operations for the same user adjustments in different environments. For example, the music generation module 160 can perform environmental measurements and listener biometric measurements when a listener trains an "Attitude" control element. During training, the music generation module 160 is trained to include these measurements as part of the control element. In this example, when the listener is performing high-intensity exercise in a gym, the "Attitude" control element may affect the intensity of the drum beats. When sitting at a computer, changing the "Attitude" control element may not affect the drum beats, but may increase the distortion of the baseline. In such embodiments, a single user control element can have different sets of rules, or different machine learning models trained differently, that are used differently, either alone or in combination, in different listening environments.
[0146] In contrast to context recognition, when the expected behavior of a control element is static, a large number of controls may be required or desired for all uses of the listening context music generation module 160. Thus, in some embodiments, the disclosed techniques can provide functionality for multiple environments with a single control element. Implementing a single control element for multiple environments can reduce the number of control elements, simplify the user interface, and enable faster search. In some embodiments, the behavior of the control element is made dynamic. The dynamism of the control element is obtained by utilizing environmental measurements such as the sound level recorded by a microphone, heart rate measurement, time, movement speed, etc. These measurements can be used as additional inputs to the training of the control element. Thus, the same listener interaction with the control element potentially has different music effects depending on the environmental situation in which the interaction occurs.
[0147] In some embodiments, the context recognition function described above is different from the concept of a generative music system that changes the generation process based on the environmental context. For example, these techniques can modify the effect of user control elements based on the environmental context, which can be used alone or in combination with the concept of generating music based on the environmental context and the output of user control.
[0148] In some embodiments, the music generation module 160 is configured to control the output music content 140 generated to achieve a predetermined goal. Examples of the described goals include, but are not limited to, sales goals, biometric goals such as heart rate or blood pressure, and ambient noise goals. The music generation module 160 can learn how to modify control elements generated manually (user-created) or algorithmically (system-defined) using the techniques described herein to generate the output music content 140 in order to achieve a predetermined goal.
[0149] The target state may be a measurable environment, and the listener may state what the listener wishes to achieve while listening to music using the music generation module 160 and with the help of the music generation module 160. These target states may be directly affected or mediated by psychological effects such as a focus that encourages a particular piece of music. That is, it may also be affected by changing the listener's acoustic experience of the space through music. As an example, a listener can set a goal to lower their heart rate while running. By recording the listener's heart rate under different states of available control elements, the music generation module 160 learns that the listener's heart rate typically decreases when a control element named "attitude" is set to a low level. Therefore, to help the listener achieve a low heart rate, the music generation module 160 can automate the "attitude" control to a low level.
[0150] By generating the type of music that a listener expects in a particular environment, the music generation module 160 can assist in creating a particular environment. Examples include heart rate, overall volume in the listener's physical space, sales in a store, and the like. Some environmental sensors and state data may not be suitable for the target state. For example, time may be an environmental metric used as an input to achieve a target state for sleep induction, but the music generation module 160 cannot control the time itself.
[0151] In various embodiments, the sensor input may be decoupled from the control element mapper while attempting to reach the state target, but the sensor may continue to record and instead provide measurements for comparing the actual target state with the target state. The difference between the target state and the actual environmental state can be formulated as a reward function for a machine learning algorithm that can adjust the mapping within the control element mapper in a mode attempting to achieve the target state. This algorithm can adjust the mapping to reduce the difference between the target and the actual environmental state.
[0152] Although music content has many physiological and psychological effects, creating the music content that a listener expects in a particular environment does not necessarily help create such an environment for the listener. In some cases, it may have no effect or even a negative impact on reaching the target state. In some embodiments, the music generation module 160 can adjust music properties based on past results while branching in another direction if the change does not meet the threshold. For example, if reducing the "attitude" control element does not lower the listener's heart rate, the music generation module 160 can transition to and develop a new strategy using other control elements, or generate new control elements using the actual state of the target variable as positive or negative reinforcement for a regression model or a neural network model.
[0153] In some embodiments, if it is found that the context affects the expected behavior of a control element for a particular listener, it may mean that a data point (e.g., an audio parameter) changed by the control element in a particular context is related to the context for that listener. Thus, these data points provide a good starting point for creating music that creates environmental changes. For example, if a listener always manually raises the "rhythmic" control element when going to a train station, the music generation module 160 can start automatically increasing this control element when it detects that the listener is at a train station.
[0154] In some embodiments, as described herein, the music generation module 160 is trained to implement control elements that match the user's expectations. If the music generation module 160 is trained end-to-end for each control element (e.g., from the control element level to the output music content 140), the complexity of the training for each control element will increase, which may slow down the training. Furthermore, it is difficult to establish the ideal combined effect of multiple control elements. However, for each control element, the music generation module 160 should ideally be trained to perform the music changes expected based on the control element. For example, the music generation module 160 may be trained by the listener for the "energy" control element to increase the rhythm density as the "energy" increases. Since the listener is exposed to not only the individual layers of the music content but also the final output music content 140, the music generation module 160 can be trained to affect the final output music content using the control element. However, this can be a multi-step problem, such as in a particular control setting, the music should sound like X, and to create music that sounds like X, a set of audio files Y should be used on each track.
[0155] In certain embodiments, a teacher / student framework is employed to address the above problems. FIG. 10 is a block diagram showing an exemplary teacher / student framework system according to some embodiments. In the illustrated embodiment, system 1000 includes a teacher model implementation module 1010 and a student model implementation module 1020.
[0156] In certain embodiments, the teacher model implementation module 1010 implements a trained teacher model. For example, the trained teacher model may be a model that learns how to predict how the final mix (e.g., stereo mix) should sound without considering at all the set of loops available in the final mix. In some embodiments, the learning process of the teacher model utilizes real-time analysis of the output music content 140 using the fast Fourier transform to calculate the distribution of sounds over different frequencies for a sequence of short time steps. The teacher model can explore the patterns of these sequences using a time sequence prediction model such as a recurrent neural network (RNN). In some embodiments, the teacher model in the teacher model implementation module 1010 may be trained offline on a stereo recording where individual loops or audio files are not available.
[0157] In the illustrated embodiment, the teacher model implementation module 1010 receives the output music content 140 and generates a compact description 1012 of the output music content. Using the trained teacher model, the teacher model implementation module 1010 can generate the compact description 1012 without considering at all the audio tracks or audio files within the output music content 140. The compact description 1012 may include an explanation X of how any output music content 140 should sound as determined by the teacher model implementation module 1010. The compact description 1012 is more compact than the output music content 140 itself.
[0158] The compact description 1012 may be provided to the student model implementation module 1020. The student model implementation module 1020 implements the trained student model. For example, the trained student model can be a model that learns a method of producing music that matches the compact description using an audio file or a loop Y (different from X). In the illustrated embodiment, the student model implementation module 1020 generates student output music content 1014 that substantially matches the output music content 140. As used herein, the phrase "substantially matches" indicates that the student output music content 1014 produces sound in the same way as the output music content 140. For example, a trained listener can consider that the student outputs the music content 1014 and the music content 140 is the same voice.
[0159] In many cases, control elements are expected to affect similar patterns in music. For example, control elements can affect both pitch relationships and rhythms. In some embodiments, the music generation module 160 is trained for a number of control elements according to one teacher model. By training the music generation module 160 for a number of control elements using a single teacher model, it may not be necessary to relearn the same basic pattern for each control element. In such embodiments, the student model of the teacher model learns how to change the selection of loops for each track in order to achieve the desired attributes in the final music mix. In some embodiments, the characteristics of the loops may be pre-calculated to reduce the learning challenge and baseline performance (although at the expense of potentially reducing the possibility of finding the optimal mapping of control elements).
[0160] Non-limiting examples of musical attributes pre-computed for each loop or audio file that can be used for student model training include: the ratio of low to high frequencies, the number of sound starts / second, the ratio of detected to undetected sounds, the spectral range, the average start intensity. In some embodiments, the student model is a simple regression model trained to select a loop for each track to obtain the musical properties closest in the final stereo mix. In various embodiments, the student / teacher model framework can have several advantages. For example, if a new property is added to the loop pre-computation routine, there is no need to re-train the entire end-to-end model, i.e., the student model.
[0161] As another example, since the characteristics of the final stereo mix that affect different controls are likely to be common to other control elements, training the music generation module 160 for each control element as an end-to-end model means that each model needs to learn the same thing to reach the best loop selection, making the training slower and more difficult than required. Since only the stereo output needs to be analyzed in real time and the output music content is generated in real time for the listener, the music generation module 160 can obtain a "free" signal computationally. Even FFT may already be applied for visualization and audio mixing purposes. In this way, the teacher model can be trained to predict the combined behavior of the control elements, and the music generation module 160 can be trained to find a way to adapt to other control elements while still generating the desired output music content. This may encourage the training of control elements to reduce control elements that have the effect of emphasizing the unique effects of specific control elements and reducing the influence of other control elements.
[0162] Exemplary low-resolution pitch detection system Robust pitch detection for polyphonic music content and diverse instrument types has traditionally been difficult to achieve. Tools that implement end-to-end music transcription can attempt to perform an audio recording and generate a score written in the MIDI format, or a symbolic music representation. Without knowledge of beat placement and tempo, these tools need to infer the music rhythm structure, instrumentation, and pitch. The results can vary, and a common problem is that too many short notes that do not exist in the audio file are detected, and the harmonics of the notes are detected as the fundamental pitch.
[0163] However, pitch detection is also useful in situations where end-to-end transcription is not required. For example, to create a harmonious combination of musical loops, it may be sufficient to know which pitches are audible on each beat without knowing the exact position of the notes. If the loop length and tempo are known, there is no need to estimate the temporal position of the beats from the audio.
[0164] In some embodiments, the pitch detection system is configured to detect which fundamental pitches (e.g., C, C#, … B) are present in a short music audio file of a known beat length. By reducing the problem scope and focusing on robustness to instrument textures, high-precision results can be achieved for beat-resolution pitch detection.
[0165] In some embodiments, the pitch detection system is trained on examples where the ground truth is known. In some embodiments, the audio data is created from the score data. MIDI and other symbolic music formats can be synthesized using a software audio synthesizer with random parameters for textures and effects. For each audio file, the system can generate a log spectrogram 2D representation with multiple frequency bins for each pitch class. This 2D representation is used as input to a neural network or other AI techniques, and a number of convolutional layers are used to create a feature representation of the frequency and time representation of the audio. The convolutional stride and padding can be varied according to the length of the audio file to generate a fixed model output shape with different tempo inputs. In some embodiments, the pitch detection system adds a regression layer to the convolutional layers to output a sequence that is temporarily dependent on the prediction. The cross-category entropy loss can be used to compare the logical output of the neural network to the binary representation of the score.
[0166] The design of convolutional layers combined with a regression layer is similar to and is changed from the task of transcribing speech to text. For example, transcribing speech to text typically needs to be sensitive to relative pitch changes but not to absolute pitch changes. Thus, the frequency range and resolution are typically small. Further, the text may need to be invariant to speed in a way that is not desirable for music at a static tempo. For example, the length of the output sequence is known in advance, reducing the complexity of training, so the connectionist temporal classification (CTC) loss calculation often used in the task of transcribing speech to text may not be needed.
[0167] The following representation has 12 pitch classes for each beat, and 1 represents the presence of its fundamental tone in the score used to synthesize audio. (C, C#... B) and each row represents a beat, and the subsequent rows represent the scores at different beats:
Number
[0168] In some embodiments, the neural network is trained on classical music and pseudo-randomly generated music scores of harmony and polyphony of 1 to 4 parts (or more). Data augmentation may help to enhance the robustness to effects such as filters and reverberation on music content. This can be a difficulty in pitch detection (for example, some of the basic tones may remain after the original note has ended). In some embodiments, the dataset can be biased, and loss weighting is used because it is much more likely that pitch classes do not play notes on each beat.
[0169] In some embodiments, the output format enables avoiding harmonic collisions at each beat while maximizing the range of harmonic contexts that the loop can use. For example, the base loop contains only F and can descend to E on the last beat of the loop. This loop sounds like it would be harmonically acceptable to most people in the key of F. Without given temporal resolution and only knowing that E and F are included in the audio, it could be a sustained E with a short F at the end. This would likely not be acceptable to most people in the context of the key of F. At higher resolution, as individual notes increase, harmonics, fretboard tones, slides may be detected and thus additional notes may be mis-identified. According to some embodiments, by developing a system with optimal resolution of temporal and pitch information to combine short audio recordings of musical instruments to create a musically mixed in a harmonically sound combination, the complexity of the pitch detection problem can be reduced and the robustness to short and insignificant pitch events can be increased.
[0170] In various embodiments of the music generation system described herein, the system enables a listener to select audio content that is used to create (generate) a pool from which the system constructs new music. This approach may be different from creating a playlist in that the user does not need to select individual tracks or order the selections. Additionally, content from multiple artists can be used simultaneously. In some embodiments, the music content is grouped into “packs” designed by a software provider or contributing artist. One pack contains multiple audio files with corresponding image features and feature metadata files. A single pack may contain, for example, 20 to 100 audio files that can be used by the music generation system to create music. In some embodiments, it is possible to select a single pack or a combination of multiple packs. During playback, packs can be added or removed without stopping the music.
[0171] Exemplary Audio Techniques for Music Content Generation In various embodiments, a software framework for managing audio generated in real time can benefit from supporting certain types of functionality. For example, audio processing software can follow the metaphor of a modular signal chain inherited from analog hardware, where different modules that provide audio generation and audio effects are chained together in an audio signal graph. Individual modules typically expose various continuous parameters that enable real-time changes to the module's signal processing. In the early days of electronic music, the parameters themselves were often analog signals, and thus the parameter processing chain and the signal processing chain coincided. Since the digital revolution, parameters have tended to be separate digital signals.
[0172] The embodiments disclosed herein recognize that a flexible control system that enables adjustment and combination of parameter manipulation can be advantageous for a real-time music generation system, whether the system interacts with a human performer or implements machine learning or other artificial intelligence techniques to generate music. Further, the present disclosure recognizes that it can also be advantageous for the effect of parameter changes to be invariant with respect to tempo changes.
[0173] In some embodiments, a music generation system generates new music content from reproduced music content based on different parameter representations of an audio signal. For example, an audio signal can be represented by both a graph of the signal over time (e.g., an audio signal graph) and a graph of the signal over beats (e.g., a signal graph). The signal graph is tempo-invariant and enables tempo-invariant changes based on the audio signal graph in addition to tempo-invariant changes of the audio parameters of the music content.
[0174] FIG. 11 is a block diagram showing an exemplary system configured to implement audio techniques in music content generation according to some embodiments. In the illustrated embodiment, system 1100 includes a graph generation module 1110 and an audio technique music generation module 1120. The audio technique music generation module 1120 may operate as a music generation module (e.g., the audio technique music generation module is music generation module 160 described herein), or the audio technique music generation module may be implemented as part of a music generation module (e.g., as part of music generation module 160).
[0175] In the illustrated embodiment, music content 1112 including music file data is accessed by a graph generation module 1110. The graph generation module 1110 can generate a first graph 1114 and a second graph 1116 for an audio signal within the accessed music content 1112. In a particular embodiment, the first graph 1114 is an audio signal graph that graphs the audio signal as a function of time. The audio signal may include, for example, amplitude, frequency, or a combination of both. In a particular embodiment, the second graph 1116 is a signal graph that graphs the audio signal as a function of beats.
[0176] In a particular embodiment, as shown in the exemplary embodiment of FIG. 11, the graph generation module 1110 is disposed within the system 1100 and generates the first graph 1114 and the second graph 1116. In such an embodiment, the graph generation module 1110 may be collocated with an audio technology music generation module 1120. However, other embodiments are conceivable where the graph generation module 1110 is disposed in a separate system and the audio technology music generation module 1120 accesses the graph from a separate system. For example, the graph may be generated and stored on a cloud-based server accessible by the audio technology music generation module 1120.
[0177] FIG. 12 shows an example of an audio signal graph (e.g., the first graph 1114). FIG. 13 shows an example of a signal graph (e.g., the second graph 1116). In the graphs of FIGS. 12 and 13, each change in the audio signal is represented as a node (e.g., the audio signal node 1202 in FIG. 12 and the signal node 1302 in FIG. 13). Thus, the parameters of a particular node determine (e.g., define) the change in the audio signal at the particular node. Since the first graph 1114 and the second graph 1116 are based on the same audio signal, the graphs may have a similar structure, and the difference between the graphs is the X-axis scale (time vs. beat). By having a similar structure within the graphs, a change in the parameters (described below) of a node (e.g., the node 1202 in the first graph 1114) in one graph corresponding to a node (e.g., the node 1302 in the second graph 1116) in the other graph can be determined by the parameters either downstream or upstream of the node in one of the graphs.
[0178] Returning to FIG. 11, the first graph 1114 and the second graph 1116 are received (or accessed) by the audio technology music generation module 1120. In a particular embodiment, the audio technology music generation module 1120 generates new music content 1122 from the playback music content 1118 based on the audio modifier parameters selected from the first graph 1114 and the audio modifier parameters selected from the second graph 1116. For example, the audio technology music generation module 1120 can modify the playback music content 1118 by any of the audio modifier parameters from the first graph 1114, the audio modifier parameters from the second graph 1116, or a combination thereof. The new music content 1122 is generated by modifying the playback music content 1118 based on the voice modification factor parameters.
[0179] In various embodiments, the audio technology music generation module 1120 can select audio modifier parameters based on whether tempo variant changes, tempo invariant changes, or combinations thereof are desired in the modification of the playback content 1118. For example, tempo variant changes may be performed based on audio modifier parameters selected or determined from the first graph 1114, and tempo invariant changes may be performed based on audio modifier parameters selected or determined from the second graph 1116. In embodiments where a combination of tempo variant changes and tempo invariant changes is desired, the audio modification parameters may be selected from both the first graph 1114 and the second graph 1116. In some embodiments, the audio modifier parameters from each individual graph are applied separately to different properties (e.g., amplitude or frequency) or different layers (e.g., different instrument layers) within the playback music content 1118. In some embodiments, the audio modifier parameters from each graph are combined into a single audio modifier parameter for application to a single property or layer within the music content 1118.
[0180] FIG. 14 shows an exemplary system for performing real-time modification of music content using a music technology music generation module 1420 according to some embodiments. In the illustrated embodiment, the audio technology music generation module 1420 includes a first node determination module 1410, a second node determination module 1420, an audio parameter determination module 1430, and an audio parameter modification module 1440. Collectively, the first node determination module 1410, the second node determination module 1420, the audio parameter determination module 1430, and the audio parameter modification module 1440 implement the system 1400.
[0181] In the illustrated embodiment, the Audio Technology Music Generation Module 1420 receives playback music content 1418 including an audio signal. The Audio Technology Music Generation Module 1420 can process the audio signal via a first graph 1414 (e.g., a time-based audio signal graph) and a second graph 1416 (e.g., a beat-based signal graph) within the first node determination module 1410. When the audio signal passes through the first graph 1414, the parameters of each node in the graph determine the change of the audio signal. In the illustrated embodiment, the second node determination module 1420 can receive the information on the first node 1412 and determine the information of the second node 1422. In a particular embodiment, the second node determination module 1420 reads the parameters within the second graph 1416 based on the position of the first node in the first node information 1412 within the audio signal passing through the first graph 1414. Thus, by way of example, an audio signal heading towards node 1202 in the first graph 1414 (shown in FIG. 12) can trigger the second node determination module 1420 to determine the corresponding (parallel) node 1302 in the second graph 1416 (shown in FIG. 13).
[0182] As shown in FIG. 14, the Audio Parameter Determination Module 1430 can receive the second node information 1422 and determine (e.g., select) specific audio parameters 1432 based on the second node information. For example, the Audio Parameter Determination Module 1430 can select the audio parameters based on a portion (e.g., x number of the next beat) of the next beat of the second graph 1416 following the position of the second node as identified by the second node information 1422. In some embodiments, a conversion from beats to real time may be implemented to determine the portion of the second graph 1416 from which the audio parameters are read. The specified audio parameters 1432 may be provided to the Audio Parameter Change Module 1440.
[0183] The voice parameter change module 1440 can control the change of music content to generate new music content. For example, the voice parameter change module 1440 can change the played music content 1418 to generate new music content 1122. In certain embodiments, the audio parameter change module 1440 changes the properties of the played music content 1418 by changing specific audio parameters 1432 (determined by the audio parameter determination module 1430) for the audio signals within the played music content. For example, changing the audio parameters 1432 specified for the audio signals within the played music content 1418 changes properties such as the amplitude, frequency, or a combination of both within the audio signals. In various embodiments, the voice parameter change module 1440 changes the properties of different voice signals within the played music content 1418. For example, the different audio signals in the played music content 1418 may correspond to different musical instruments represented in the played music content 1418.
[0184] In some embodiments, the voice parameter change module 1440 uses a machine learning algorithm or other AI technology to change the characteristics of the voice signals within the played music content 1418. In some embodiments, the voice parameter change module 1440 changes the properties of the played music content 1418 according to user input to the module, which may be provided via a user interface associated with the music generation system. Embodiments are also contemplated in which the voice parameter change module 1440 changes the properties of the played music content 1418 using a combination of AI technology and user input. Various embodiments for changing the characteristics of the played music content 1418 by the voice parameter change module 1440 enable real-time operation of the music content (e.g., operation during playback). As described above, real-time operation can include applying tempo change modifications, tempo-invariant combinations, or a combination of both to the audio signals of the played music content 1418.
[0185] In some embodiments, the audio technology music generation module 1420 implements a two-layer parameter system to change the characteristics of the reproduced music content 1418 by the voice parameter change module 1440. In the two-layer parameter system, there may be a distinction between "automation" that directly controls the voice parameter values (e.g., tasks automatically executed by the music generation system) and "modulation" that multiplicatively superimposes voice parameter changes on top of the automation, as described below. The two-layer parameter system allows different parts of the music generation system (e.g., different machine learning models in the system architecture) to consider different musical aspects separately. For example, one part of the music generation system can set the volume of a particular instrument according to the intended section type of the composition, and another part can superimpose periodic variations in volume for additional interest.
[0186] Exemplary Techniques for Real-Time Audio Effects in Music Content Generation Music technology software typically enables composers / producers to control various abstract envelopes by automation. In some embodiments, automation is a pre-programmed temporal manipulation of some audio processing parameters (such as volume or reverb amount). Automation is typically either a manually defined breakpoint envelope (e.g., a piecewise linear function) or a programmable function such as a sine wave (also known as a low-frequency oscillator (LFO)).
[0187] The disclosed music generation system may be different from typical music software. For example, most parameters are, in a sense, automated by default. The AI technology of the music generation system can control most or all audio parameters in various ways. At a base level, the neural network can predict appropriate settings for each audio parameter based on its training. However, it may be useful to provide high-level automation rules for the music generation system. For example, large-scale music structures may instruct a slow build of volume as an additional consideration in addition to lower-level settings that might otherwise be predicted.
[0188] This disclosure generally relates to an information architecture and a procedural approach for combining multiple parametric instructions issued simultaneously by different levels of a hierarchical generation system to create a continuous output that is musically coherent and evolving. The disclosed music generation system can generate long-form music experiences intended to be experienced continuously for hours. Long-form music experiences need to create a consistent musical experience journey for a more satisfying experience. To do this, the music generation system can reference itself on a long time scale. These references can vary from direct to summarized.
[0189] In certain embodiments, to facilitate larger-scale music rules, a music generation system (e.g., music generation module 160) exposes an automation API (Application Programming Interface). FIG. 15 shows a block diagram of an exemplary API module in a system for the automation of audio parameters, according to some embodiments. In the illustrated embodiment, system 1500 includes API module 1505. In one embodiment, API module 1505 includes automation module 1510. The music generation system can support both wave table style LFOs and any breakpoint envelope. Automation module 1510 can apply automation 1512 to any audio parameter 1520. In some embodiments, automation 1512 is applied recursively. For example, any programmable automation, such as a sine wave, has parameters of its own (frequency, amplitude, etc.) and automation can be applied to those parameters.
[0190] In various embodiments, automation 1512 includes a signal graph parallel to the audio signal graph, as described above. The signal graph can be treated similarly via a "pull" technique. In the "pull" technique, API module 1505 can request recalculation from automation module 1510 as needed, and automation 1512 can perform recalculation such that it recursively requests upstream automation on which it depends. In certain embodiments, the signal graph for automation is updated at a controlled rate. For example, the signal graph can be updated each time the execution engine update routine is executed, which can be aligned with the audio block rate (e.g., after the audio signal graph has rendered one block (e.g., one block is 512 samples)).
[0191] In some embodiments, it is desirable for the audio parameter 1520 itself to vary at the audio sample rate; otherwise, discontinuous parameter changes at the audio block boundaries can introduce audible artifacts. In certain embodiments, the music generation system manages this issue by treating the automation updates as parameter value targets. When the real-time audio thread renders an audio block, the audio thread smoothly ramps the specified parameter from its current value to the supplied target value over the course of that block.
[0192] The music generation system described herein (e.g., the music generation module 160 shown in FIG. 1) may have an architecture with a hierarchical nature. In some embodiments, different parts of the hierarchy can provide multiple cues for the value of a particular audio parameter. In certain embodiments, the music generation system provides two separate mechanisms for combining / resolving multiple cues, namely, modulation and overriding. In the illustrated embodiment of FIG. 15, modulation 1532 is implemented by a modulation module 1530 and overriding 1542 is implemented by an overriding module 1540.
[0193] In some embodiments, automation 1512 may be declared to be modulation 1532. Such a declaration may mean that instead of directly setting the value of the audio parameter, the automation 1512 should act multiplicatively on the current value of the audio parameter. Thus, a large music section can apply a long modulation 1532 to an audio parameter, and the value of the modulation is multiplied regardless of the value indicated by other parts of the music generation system.
[0194] In various embodiments, the API module 1505 includes an overwrite module 1540. The overwrite module 1540 may be, for example, an overwrite function for audio parameter automation. The overwrite module 1540 may be intended to be used by an external control interface (e.g., an artist control user interface). The overwrite module 1540 may control the audio parameters 1520 regardless of what the music generation system attempts to do with the music parameters 1520. If there are overwrite parameters 1520 that are overwritten by an overwrite 1542, the music generation system may generate a "shadow parameter" 1522 that tracks the audio parameters that could have been overwritten if they had not been overwritten (e.g., if the overwrite parameters are based on an automation 1512 or a modulation 1532). Thus, when the overwrite 1542 is "released" (e.g., removed by an artist), the audio parameters 1520 can snap back to where they would have been according to the automation 1512 or the modulation 1532.
[0195] In various embodiments, these two approaches can be combined. For example, the overwrite 1542 can be a modulation 1532. If the overwrite 1542 is a modulation 1532, the base value of the overwrite parameter 1520 is set by the music generation system, but may be multiplicatively modulated by the overwrite 1542. Each audio parameter 1520 may have at the same time not only one (or zero) of each overwrite 1542, but also one (or zero) automation 1512 and one (or zero) modulation 1532.
[0196] In various embodiments, the abstract class hierarchy is defined as follows (note that there is some multiple inheritance):
Number
[0197] Based on the hierarchy of abstract classes, a mono is considered either an Automation or Automatable. In some embodiments, any automation can be applied to anything that is automatable. Automations include LFOs, breakpoint envelopes, etc. All of these automations are tempo-locked. That is, they change over time according to the current beat.
[0198] An automation can itself have parameters that are automatable. For example, the frequency and amplitude of LFO automation are automatable. Thus, there is a signal graph of dependent automations and automation parameters that runs in parallel with the audio signal graph at the control rate rather than the audio rate. As described above, the signal graph uses the pull model. The music generation system holds a track of any automation 1512 applied to the audio parameter 1520 and updates these for each "game loop". Next, the automation 1512 recursively requests updates to its own automation audio parameters 1520. This recursive update logic may exist in the base class Beat-Dependent, which is expected to be called frequently (but not necessarily regularly). The update logic can have a prototype described as follows:
Number
[0199] In certain embodiments, the BeatDependent class maintains a list of its own dependencies (e.g., other BeatDependent instances) and recursively calls their update functions. The updateCounter can be passed through the chain so that the signal graph can cycle without double-updating. This can be important because an automation may be applied to several different targets of automation. In some embodiments, this is not a problem because the second update has the same currentBeat as the first, and these update routines are ineffective unless the beat changes.
[0200] In various embodiments, when automation is applied to automation, in each cycle of the "game loop", the music generation system can request (recursively) updated values from each automation and use it to set the values of the automation. In this case, "setting" may depend on a particular subclass and may also depend on whether the parameter is also being modulated and / or disabled.
[0201] In certain embodiments, modulation 1532 is automation 1512 that is applied multiplicatively rather than absolutely. For example, modulation 1532 can be applied to an already automated audio parameter 1520, and the effect is a percentage of the automated value. This can enable multiplicative, for example, ongoing vibrations around a moving means.
[0202] In some embodiments, the audio parameter 1520 is overwritable, meaning that any automation 1512 or modulation 1532 applied thereto, or other (less privileged) requests, are overwritten by the overwrite value of the overwrite 1542, as described above. This overwrite enables external control over some aspects of the music generation system, while the music generation system continues as it would otherwise. When the audio parameter 1520 is disabled, the music generation system keeps track of what the value was (e.g., tracks the applied automation / modulation and other requests). When the overwrite is removed, the music generation system snaps the parameter and checks where the parameter was.
[0203] To facilitate modulation 1532 and disable 1542, the music generation system may abstract the setValue method of the parameter. There is also a private method _setValue that actually sets the value. Examples of public methods are shown below:
Number
[0204] The public method can refer to a member variable of the Parameter class called _unmodified. This variable is an instance of the aforementioned ShadowParameter. Each audio parameter 1520 has a shadow parameter 1522 that tracks where it would be if not modulated. If the audio parameter 1520 is not currently modulated, both the audio parameter 1520 and its shadow parameter 1522 are updated with the requested value. Otherwise, the shadow parameter 1522 tracks the request and the actual audio parameter value 1520 is set elsewhere (e.g., in the updateModulations routine - where the modulation factor is multiplied by the shadow parameter value to give the actual parameter value).
[0205] In various embodiments, large-scale structures in a long-form music experience are achieved by various mechanisms. One broad approach is the long-term use of musical self-reference. For example, a very direct self-reference is to exactly repeat a previously played audio segment. In music theory, a repeated segment is called a theme (or motif). More typically, musical content uses themes and variations, where the theme is later repeated with some variations, giving a sense of coherence while maintaining a sense of progression. The music generation system disclosed herein can use themes and variations to create large-scale structures in several ways, including direct repetition or the use of an abstract envelope.
[0206] An abstract envelope is a value of audio parameters over time. Abstractly from the controlled audio parameters, the abstract envelope can be applied to other audio parameters. For example, a set of audio parameters can be automated cooperatively by a single control abstract envelope. This technique can "bind" different layers perceptually in a short period. The abstract envelope can be temporarily reused and applied to different audio parameters. In this way, the abstract envelope becomes an abstract musical theme, and this theme is repeated by applying this envelope to other audio parameters in a later listening experience. In this way, a sense of structure and long-term consistency is established while there are variations in this theme.
[0207] The abstract envelope is regarded as a musical theme and can abstract many musical features. Examples of musical features that can be abstracted include, but are not limited to, the following: · Built in tension (such as track volume, level of distortion, etc.). · Rhythm (volume adjustment and / or gate setting can give a rhythmic effect to pads, etc.). · Melody (pitch filtering can mimic the contour of a melody applied to pads, etc.).
[0208] Exemplary additional audio techniques for real-time music content generation Real-time music content generation can pose unique challenges. For example, due to strict real-time constraints, function calls and subroutines with unpredictable and unconstrained execution times should be avoided. Avoiding this problem may eliminate most uses of high-level programming languages and most uses of low-level languages such as C and C++. Things that allocate memory from the heap (e.g., malloc under the hood) can be excluded, as can things that can block, such as locking a mutex. This makes multithreaded programming particularly difficult in real-time music content generation. Most standard memory management approaches may also be infeasible, and as a result, dynamic data structures such as C++ STL containers are restricted in their use for real-time music content generation.
[0209] Another area of concern is the management of audio parameters that involve DSP (digital signal processing) functions (such as the cutoff frequency of a filter). For example, when changing an audio parameter dynamically, audible artifacts may occur unless the audio parameter is changed continuously. Therefore, communication between the real-time DSP audio thread and the user-facing interface or the programmatic interface may be required to change an audio parameter.
[0210] To address these constraints, various audio software can be implemented, and various approaches exist. For example, · Inter-thread communication may be handled with a lock-free message queue. · Functions are written in plain C and utilize function pointer callbacks. · Memory management may be implemented via custom "zones" or "regions". · The "2-speed" system may be implemented such that real-time audio thread calculations are executed at the audio rate and the audio thread is controlled to execute at the "control rate". The control audio thread can set the change target of the audio parameters for the real-time audio thread to smoothly ramp.
[0211] In some embodiments, synchronization between control rate audio parameter operations for use in actual DSP routines and real-time audio thread safe storage of audio parameter values may require some kind of thread-safe communication of audio parameter targets. Since most audio parameters of audio routines are continuous (not discrete), they are usually represented by floating-point data types. Various distortions to the data have historically been required due to the lack of lock-free atomic floating-point data types.
[0212] In certain embodiments, a simple lock-free atomic floating-point data type is implemented in the music generation system described herein. The lock-free atomic floating-point data type can be achieved by treating the floating-point type as a bit string and "tricking" the compiler into treating it as an atomic integer type of the same bit width. This approach can support atomic getting / setting suitable for the music generation system described herein. An implementation example of the lock-free atomic floating-point data type is shown below:
Number
[0213] In some embodiments, dynamic memory allocation from the heap is not executable for real-time code related to music content generation. For example, static stack-based allocation can make it difficult to use programming techniques such as dynamic memory containers and functional programming approaches. In certain embodiments, the music generation system described herein implements a "memory zone" for memory management in a real-time context. As used herein, a "memory zone" is an area of heap-allocated memory that is allocated upfront without real-time constraints (e.g., when real-time constraints do not yet exist or are paused). Memory storage objects are generated within the area of heap-allocated memory without the need to request more memory from the system, thereby making the memory safe in real-time. Garbage collection can include deallocating the entire memory zone. Also, the implementation of memory by the music generation system may be safe, real-time safe, and efficient multi-threaded.
[0214] FIG. 16 shows a block diagram of an exemplary memory zone 1600 according to some embodiments. In the illustrated embodiment, the memory zone 1600 includes a heap-allocated memory module 1610. In various embodiments, the heap-allocated memory module 1610 receives and stores a first graph 1114 (e.g., an audio signal graph), a second graph 1116 (e.g., a signal graph), and audio signal data 1602. Each of the stored items can be retrieved, for example, by an audio parameter change module 1440 (shown in FIG. 14).
[0215] An example implementation of a memory zone is shown below:
Number
[0216] In some embodiments, different audio threads of a music generation system need to communicate with each other. Typical thread-safety approaches (which may involve locking of "mutually exclusive" data structures) may not be usable in a real-time context. In certain embodiments, serialization of dynamic routing data to a pool of single-producer single-consumer circular buffers is implemented. A circular buffer is a type of FIFO (first-in first-out) queue data structure that does not require dynamic memory allocation after initialization. A single-producer single-consumer thread-safe circular buffer allows one audio thread to push data into the queue while another audio thread pulls the data out. For the music generation system described herein, the circular buffer can be extended to allow for multi-producer single-consumer audio threads. These buffers can be implemented by pre-assigning a static array of circular buffers and dynamically routing the serialized data to specific "channels" (e.g., specific circular buffers) according to an identifier added to the music content generated by the music generation system. The static array of circular buffers is accessible by a single user (e.g., a single consumer).
[0217] FIG. 17 shows a block diagram of an exemplary system for storing new music content, according to some embodiments. In the illustrated embodiment, system 1700 includes a circular buffer static array module 1710. The circular buffer static array module 1710 includes a plurality of circular buffers and enables storage of a plurality of producers, single-consumer audio threads according to thread identifiers. For example, the circular buffer static array module 1710 can receive new music content 1122 and store the new music content at 1712 for access by a user.
[0218] In various embodiments, abstract data structures such as dynamic containers (vectors, queues, lists) are typically implemented in a non-real-time safe manner. However, these abstract data structures are useful for audio programming. In certain embodiments, the music generation system described herein implements a custom list data structure (e.g., a singly linked list). Many functional programming techniques can be implemented from a custom list data structure. The implementation of the custom list data structure may use a "memory zone" (described above) for underlying memory management. In some embodiments, the custom list data structure is serializable, which can be made safe for real-time use and can communicate between audio threads using the multi-producer, single-consumer audio threads described above.
[0219] Exemplary blockchain ledger technology The disclosed system, in some embodiments, can utilize a secure recording technology such as a blockchain or other cryptographic ledger to record information about the generated music or elements thereof, such as loops or tracks. In some embodiments, the system combines a plurality of audio files (e.g., tracks or loops) to generate output music content. This combination can be performed by combining multiple layers of audio content so that they at least partially overlap in time. The output content may be individual pieces of music or continuous. Tracking the use of music elements can be difficult in the context of continuous music, for example, to provide royalties to relevant stakeholders. Thus, in some embodiments, the disclosed system records the identifiers and usage information (e.g., timestamps or play counts) of the audio files used in the synthesized music content. Further, the disclosed system can utilize various algorithms to track, for example, the playback time in the context of blended audio files.
[0220] As used herein, the term "blockchain" refers to a series of cryptographically linked records (referred to as blocks). For example, each block may include the cryptographic hash of the previous block, a timestamp, and transaction data. A blockchain can be used as a publicly distributed ledger and can be managed by a network of computing devices that use a protocol agreed upon for communication and verification of new blocks. Among the implementations of blockchains, some are immutable while others allow for the modification of subsequent blocks. Generally, a blockchain can record transactions in a verifiable and permanent manner. Although the blockchain ledger is discussed herein for illustrative purposes, it should be understood that the disclosed techniques can be used with other types of cryptographic ledgers in other embodiments.
[0221] FIG. 18 is a diagram showing an example of reproduction data according to some embodiments. In the illustrated embodiment, the database structure includes entries for a plurality of files. Each entry shown includes a file identifier, a start timestamp, and a total time. The file identifier can uniquely identify an audio file tracked by the system. The start timestamp can indicate the first inclusion of the audio file in the mixed audio content. This timestamp may be based on the local clock of the playback device or, for example, based on an Internet clock. The total time can indicate the length of the interval during which the audio file was incorporated. Note that this may be different from the length of the audio file. For example, when only a portion of the audio file is used, when the audio file is accelerated or decelerated within the mix, etc. In some embodiments, when an audio file is incorporated at multiple different times, each time results in an entry. In other embodiments, additional play for a file may result in an increase in the time field of an existing entry if an entry for the file already exists. In still other embodiments, the data structure may track the number of times each audio file is used rather than the length of the capture. Further, other encodings of time-based usage data are conceivable.
[0222] In various embodiments, different devices can determine, store, and use a ledger to record reproduction data. Exemplary scenarios and topologies are described later with reference to FIG. 19. The reproduction data may be temporarily stored in a computing device before being committed to the ledger. The stored reproduction data may be encrypted, for example, to reduce or avoid manipulation of entries or insertion of false entries.
[0223] FIG. 19 is a block diagram showing an exemplary composition system according to some embodiments. In the example shown, the system includes a playback device 1910, a computing system 1920, and a ledger 1930.
[0224] In the illustrated embodiment, the playback device 1910 receives a control signal from the computing system 1920 and transmits playback data to the computing system 1920. In this embodiment, the playback device 1910 includes a playback data recording module 1912 that can record playback data based on the audio mix played by the playback device 1910. The playback device 1910 also includes a playback data storage module 1914 configured to temporarily store, log, or both, the playback data. The playback device 1910 may report the playback data to the computing system 1920 periodically or in real time. The playback data may be stored for later reporting, for example, when the playback device 1910 is offline.
[0225] In the illustrated embodiment, the computing system 1920 receives the playback data and commits an entry reflecting the playback data to the ledger 1930. The computing system 1920 also sends a control signal to the playback device 1910. This control signaling may include various types of information in different embodiments. For example, the control signal transmission may include configuration data, mixing parameters, audio samples, machine learning updates, etc. for use by the playback device 1910 to compose music content. In other embodiments, the computing system 1920 can compose music content and stream the music content data to the playback device 1910 via a control signal. In these embodiments, the modules 1912 and 1914 may be included in the computing system 1920. Generally, the modules and functions described with reference to FIG. 19 may be distributed among multiple devices according to various topologies.
[0226] In some embodiments, the playback device 1910 is configured to commit entries directly to the ledger 1930. For example, a playback device such as a mobile phone can create music content, determine playback data, and store the playback data. In this scenario, the mobile device can report the playback data to a server such as the computing system 1920, or directly to a computing system (or a set of computing nodes) that maintains the ledger 1930.
[0227] In some embodiments, the system maintains a record of the rights holders, along with, for example, audio file identifiers or mappings to a set of audio files. The record of this entity can be maintained in the ledger 1930, or in a separate ledger, or in some other data structure. This allows, for example, the rights owner to remain anonymous if the ledger 1930 is public but contains non-identifying entity identifiers mapped to entities within some other data structure.
[0228] In some embodiments, the music synthesis algorithm can generate a new audio file from two or more existing audio files for inclusion in a mix. For example, the system can generate a new audio file C based on two audio files A and B. One implementation for such blending uses interpolation between the vector representations of the audio representations of files A and B and the generation of file C using an inverse transformation from the vector to the audio representation. In this example, the playback times of audio files A and B may both be incremented, or may be incremented by less than the actual playback time, for reasons such as being blended.
[0229] For example, when audio file C is incorporated into the mixed content for 20 seconds, audio file A has playback data indicating 15 seconds, and audio file B has playback data indicating 5 seconds (note that the total of the blended audio files may or may not match the usage period of the resulting file C). In some embodiments, the playback time of each original file is based on it, similar to the mixed file C. For example, in a vector embodiment, in the case of an n-dimensional vector representation, the interpolation vector a has the following distance d from the vector representations of audio files A and B:
Number
[0230] In these embodiments, the playback time i of each original file can be determined as follows:
Number
[0231] In some embodiments, the form of remuneration can be incorporated into a ledger structure. For example, a specific entity can include information associating an audio file with execution requirements such as displaying a link or including an advertisement. In these embodiments, the composition system can provide proof of performance of related operations (e.g., displaying an advertisement) when including the audio file in the mix. The proof of execution can be reported according to one of various suitable reporting templates that require specific fields to indicate how and when the operation was executed. The proof of execution includes time information and can utilize encryption to avoid false claims of execution. In these embodiments, the use of an audio file that also does not show evidence of performance of the related required operations may require some other form of remuneration such as royalty payment. Generally, different forms of remuneration can be registered when the entity transmitting the audio file is different.
[0232] As described above, the disclosed technology can provide a reliable record of audio file usage in a music mix, even if composed in real time. The publicity of the ledger can instill trust in the fairness of compensation. This can encourage the participation of artists and other collaborators, and potentially improve the diversity and quality of audio files available for automatic mixing.
[0233] In some embodiments, an artist pack may be created with elements used by a music engine to create a continuous soundscape. The artist pack may be a curated set of elements stored in one or more data structures related to an entity such as an artist or group, either professionally (or otherwise). Examples of these elements include, but are not limited to, loops, composition rules, discovery methods, and neural network vectors. Loops can be included in a database of music phrases. Each loop is typically a single instrument or related set of instruments that plays the progression of music over a certain period. These range from short loops (e.g., 4 bars) to long loops (e.g., 32 - 64 bars). Loops can be arranged into layers such as melody, harmony, drums, bass, top, FX, etc. Also, the loop database may be represented as a variational auto - encoder with encoded loop representations. In this case, the loops themselves are not necessary, and the NN is used to generate the audio encoded by the NN.
[0234] Heuristics refer to the parameters, rules, and data that guide a music engine. Parameters guide elements such as section length, use of effects, frequency of variation techniques, musical complexity, or any kind of parameter that can generally be used to enhance the decision - making of the music engine when it constructs and renders music.
[0235] The ledger records transactions related to the consumption of content by related right holders. This can be, for example, a loop, heuristics, or a neural network vector. The purpose of the ledger is to record these transactions and enable transparent accounting. The ledger is intended to capture transactions when they occur, including content consumption, use of parameters in the music engine guide, use of vectors on the neural network, etc. The ledger can record various transaction types, such as individual events (e.g., this loop was played at this point in time), this pack was played during this time period, or this machine learning module (e.g., neural network module) was used during this time period.
[0236] The ledger can associate multiple right holders with any artist pack, and more specifically, can associate them with specific loops or other elements of the artist pack. For example, a label, artist, and composer can have rights to a given artist pack. The ledger can be allowed to associate payment details of the pack that specify what percentage each party receives. For example, the artist can receive 25%, the label can receive 25%, and the composer can receive 50%. Using a blockchain to manage these transactions allows for small payments to be made to each right holder in real time or accumulated over an appropriate period.
[0237] As described above, in some implementations, the loop may be replaced by a VAE that is essentially an encoding of the loop within the machine learning module. In this case, the ledger can associate playback time with a specific artist pack that includes the machine learning module. For example, if an artist pack is played for 10% of the total playback time of all devices, this artist can receive 10% of the total revenue distribution.
[0238] In some embodiments, the system enables an artist to create an artist profile. The profile includes information related to the artist, such as biographics, a profile picture, bank details, and other data necessary to verify the artist's identity. Once the artist profile is created, the artist can upload and issue artist packs. These packs contain elements that the music engine uses to create a soundscape.
[0239] For each created artist pack, a rights holder can be defined and associated with the pack. Each rights holder can claim a percentage of the pack. Additionally, each rights holder creates a profile and associates a bank account and payment profile. The artist themselves can be the rights holder and own 100% of the rights related to their pack.
[0240] In addition to recording events used for revenue recognition in a ledger, the ledger can manage promotions related to artist packs. For example, an artist pack can conduct a free promotion for a month where the revenue is different from when no promotion is being run. The ledger automatically calculates these revenue inputs when calculating payments to rights holders.
[0241] This copyright management model enables an artist to sell the rights of a pack to one or more external rights holders. For example, at the launch of a new package, the artist can pre - fund the package by selling 50% of the share of the pack to fans or investors. In this case, the number of investors / rights holders can be arbitrarily large. For example, the artist can sell 50% to 100K users and receive 1 / 100K of the revenue. Since all accounting is managed in the ledger, investors will receive compensation directly in this scenario, eliminating the need for auditing the artist's account.
[0242] Exemplary User and Enterprise GUIs Figures 20A-20B are block diagrams showing graphical user interfaces according to some embodiments. In the illustrated embodiments, Figure 20A includes a GUI displayed by user application 2010, and Figure 20B includes a GUI displayed by enterprise application 2030. In some embodiments, the GUIs shown in Figures 20A and 20B are generated by a website rather than by an application. In various embodiments, any of a variety of suitable elements may be displayed, including one or more elements such as dials (e.g., for controlling volume, energy, etc.), buttons, knobs, display boxes (e.g., for providing updated information to the user), etc.
[0243] In Figure 20A, user application 2010 displays a GUI that includes section 2012 for selecting one or more artist packs. In some embodiments, pack 2014 can alternatively or additionally include a theme pack or pack for a particular occasion (e.g., wedding, birthday party, graduation, etc.). In some embodiments, the number of packs shown in section 2012 is greater than the number of packs that can be displayed in section 2012 at one time. Thus, in some embodiments, the user scrolls up and down in section 2012 to display one or more packs 2014. In some embodiments, the user can select an artist pack 2014 for which they want to listen to output music content. In some embodiments, the artist pack can be, for example, purchased and / or downloaded.
[0244] Selection element 2016 enables the user to adjust one or more music attributes (e.g., energy level) in the illustrated embodiment. In some embodiments, selection element 2016 enables the user to add / remove / change one or more target music attributes. In various embodiments, selection element 2016 may render one or more UI control elements (e.g., control element 830).
[0245] In the illustrated embodiment, the selection element 2020 enables a user to expose the environment to the device (e.g., a mobile device) in order to determine the target music attributes. In some embodiments, after the user selects the selection element 2020, the device uses one or more sensors (e.g., a camera, a microphone, a thermometer, etc.) to collect information about the environment. In some embodiments, the application 2010 also selects or proposes one or more artist packs based on the environmental information collected by the application when the user selects the element 2020.
[0246] The selection element 2022 enables a user to combine multiple artist packs to generate a new rule set in the illustrated embodiment. In some embodiments, the new rule set is based on the user selecting one or more packs for the same artist. In other embodiments, the new rule set is based on the user selecting one or more packs for different artists. The user can indicate the weights of different rule sets such that, for example, a high-weight rule set has a greater effect on the generated music than a low-weight rule set. The music generator can combine rule sets in a plurality of different ways, such as by switching between rules from different rule sets, taking the average value of rules from a plurality of different rule sets, etc.
[0247] In the illustrated embodiment, the selection element 2024 enables a user to manually adjust rules within one or more rule sets. For example, in some embodiments, the user may want to adjust the music content generated at a finer level by adjusting one or more rules within the rule set used to generate the music content. In some embodiments, this enables the user of the application 2010 to become their own disc jockey by using the controls shown in the GUI of FIG. 20A to adjust the rule set that the music generator uses to generate the output music content. These embodiments may also enable a finer control of the target music attributes.
[0248] In FIG. 20B, the enterprise application 2030 displays a GUI that includes an artist pack selection section 2012 and an artist pack 2014. In the illustrated embodiment, the enterprise GUI displayed by the application 2030 also includes elements 2016 for adjusting / adding / deleting one or more music attributes. In some embodiments, the GUI shown in FIG. 20B is used in a business or in-store setting to generate a particular environment (e.g., to optimize sales) by generating music content. In some embodiments, an employee uses the application 2030 to select one or more artist packs that have previously been shown to increase sales (e.g., the metadata of a given rule set can indicate the actual experimental results using the rule set in a real-world context).
[0249] In the illustrated embodiment, the input hardware 2040 transmits information to the application or website displaying the enterprise application 2030. In some embodiments, the input hardware 2040 is one of a cash register, a thermal sensor, an optical sensor, a clock, a noise sensor, etc. In some embodiments, the target music attributes and / or rule sets for generating output music content for a specific environment are adjusted using information transmitted from one or more of the above hardware devices. In the illustrated embodiment, the selection element 2038 enables the user of the application 2030 to select one or more hardware devices for receiving environmental inputs.
[0250] In the illustrated embodiment, the display 2034 displays environmental data to the user of the application 2030 based on information from the input hardware 2040. In the illustrated embodiment, the display 2032 indicates changes to the rule set based on the environmental data. In some embodiments, the display 2032 enables the user of the application 2030 to view the changes made based on the environmental data.
[0251] In some embodiments, the elements shown in FIGS. 20A and 20B are for theme packs and / or opportunity packs. That is, in some embodiments, users or businesses using the GUIs displayed by the applications 2010 and 2030 can select / adjust / change the rule set to generate music content for one or more opportunities and / or themes.
[0252] Detailed example of a music generation system Figures 21-23 show details regarding a particular embodiment of the music generation module 160. It should be noted that these particular examples are disclosed for purposes of illustration and are not intended to limit the scope of the present disclosure. In these embodiments, the construction of music from loops is performed by a client system such as a personal computer, a mobile device, a media device, etc. As used in the description of Figures 21-23, the term "loop" can be replaced with the term "audio file". Generally, loops are included in audio files as described herein. Loops can be specifically curated into loop packs, which can be called artist packs. Loops can be analyzed for music properties and those properties can be stored as loop metadata. The audio within the constructed track can be analyzed (e.g., in real time), filtered, mixed, and mastered for the output stream. Various feedback is sent to the server, including, for example, explicit feedback from the user's interaction with sliders or buttons, and implicit feedback generated by sensors based on, for example, volume changes, listening length, environmental information, etc. In some embodiments, control inputs have known effects (e.g., to directly or indirectly specify target music attributes) and are used by the configuration module.
[0253] The following discussion introduces various terms used with reference to Figures 21-23. In some embodiments, a loop library is a master library of loops that may be stored by a server. Each loop can include audio data and metadata that describes the audio data. In some embodiments, a loop package is a subset of the loop library. The loop package may be a pack for a particular artist, a particular mood, a particular type of event, etc. The client device can download a loop package, for example, for offline listening or downloading a portion of the loop pack in response to a request for online listening.
[0254] In some embodiments, the generated stream is data that specifies the music content that the user hears when the user uses the music generation system. It should be noted that the actual output audio signal can vary slightly for a given generated stream, for example, based on the capabilities of the audio output device.
[0255] In some embodiments, the composition module constructs a composition from loops available within a loop package. The configuration module can receive loops, loop metadata, and user input as parameters and can be executed by a client device. In some embodiments, the configuration module outputs an execution script that is sent to an execution module and one or more machine learning engines. In some embodiments, the execution script outlines which loops are played on each track of the generated stream and what effects are applied to the stream. The execution script can utilize beat - relative timing to represent when an event occurs. Also, the execution script can encode effect parameters (e.g., for effects such as reverb, delay, compression, equalization, etc.).
[0256] In some embodiments, the execution module receives the execution script as input and renders it to the generated stream. The execution module can generate a number of tracks specified by the execution script and mix the tracks into a stream (e.g., a stereo stream), although the stream can have various encodings including, in various embodiments, audio encoding, object - based audio encoding, multi - channel stereo, etc. In some embodiments, when a particular execution script is provided, the execution module always generates the same output.
[0257] In some embodiments, the analysis module is a server implementation module that receives feedback information and configures the configuration module (e.g., in real time, periodically, based on an administrator command, etc.). In some embodiments, the analysis module uses a combination of machine learning techniques to correlate user feedback with execution scripts and loop library metadata.
[0258] FIG. 21 is a block diagram showing an exemplary music generation system including an analysis module and a configuration module according to some embodiments. In some embodiments, the system of FIG. 21 is configured to generate a potentially infinite stream of music by allowing a user to directly control the mood and style of the music. In the illustrated embodiment, the system includes an analysis module 2110, a configuration module 2120, an execution module 2130, and an audio output device 2140. In some embodiments, the analysis module 2110 is implemented by a server, and the configuration module 2120 and the execution module 2130 are implemented by one or more client devices. In other embodiments, modules 2110, 2120, and 2130 may all be implemented on a client device, or may all be implemented on the server side.
[0259] In the illustrated embodiment, the analysis module 2110 stores one or more artist packs 2112 and implements a feature extraction module 2114, a client simulator module 2116, and a deep neural network 2118.
[0260] In some embodiments, the feature extraction module 2114 adds loops to the loop library after analyzing loop audio (note that some loops may be received with pre-generated metadata and may not require analysis). For example, raw audio in formats such as wav, aiff, or FLAC can be analyzed for quantifiable musical characteristics such as instrument classification, pitch transcription, beat timing, tempo, file length, and audio amplitude in multiple frequency bins. Also, the analysis module 2110 can store more abstract musical characteristics or mood descriptions for loops, for example, based on manual tagging by an artist or machine listening. For example, mood can be quantified using a plurality of discrete categories having a range of values for each category for a given loop.
[0261] Consider loop A, for example, which is analyzed to determine that notes G2, Bb2, and D2 are used, the first beat starts at 6 milliseconds into the file, the tempo is 122 bpm, the file is 6483 milliseconds long, and the loop has normalized amplitude values of 0.3, 0.5, 0.7, 0.3, and 0.2 across five frequency bins. The artist could label the loop as "funk genre" with the following mood values: [Table 1]
[0262] The analysis module 2110 can store this information in a database, and a client can download a subset of the information, for example, as a loop package. The artist pack 2112 is shown for illustration, but the analysis module 2110 can provide various types of loop packages to the composition module 2120.
[0263] In the illustrated embodiment, the client simulator module 2116 analyzes various types of feedback and provides feedback information in a format supported by the deep neural network 2118. In the illustrated embodiment, the deep neural network 2118 also receives, as input, an execution script generated by the configuration module. In some embodiments, the deep neural network configures the configuration module based on these inputs, for example, to improve the correlation between the type of generated music output and the desired feedback. For example, the deep neural network may periodically push updates to the client device implementing the configuration module 2120. Note that the deep neural network 2118 is shown for illustrative purposes and may provide powerful machine learning performance in the disclosed embodiments, but is not intended to limit the scope of the present disclosure. In various embodiments, various types of machine learning techniques may be implemented alone or in various combinations to perform similar functions. The machine learning module may be used, in some embodiments, to directly implement a rule set (e.g., an arrangement rule or technique), or, for example, in the illustrated embodiment, to control a module implementing another type of rule set using the deep neural network 2118.
[0264] In some embodiments, the analysis module 2110 generates composition parameters for the composition module 2120 to improve the correlation between the desired feedback and the use of specific parameters. For example, actual user feedback may be used to adjust the composition parameters, for example, in an attempt to reduce negative feedback.
[0265] As an example, consider a situation where module 2110 discovers a correlation between negative feedback (e.g., explicit low ranking, low volume listening, short listening time, etc.) and composition using multiple layers. In some embodiments, module 2110 uses techniques such as backpropagation to adjust probability parameters used to add more tracks in order to determine that this reduces the frequency of this problem. For example, module 2110 can predict that reducing the probability parameter by 50% will reduce negative feedback by 8%, decide to reduce it, and push the updated parameter to the composition module (note that the probability parameter, although discussed in detail below, can be adjusted similarly for any of the various parameters for the statistical model).
[0266] As another example, consider a situation where module 2110 discovers that negative feedback correlates with the user setting mood control to high tension. Also, a correlation may be found between loops with low tension tags and users who require high tension. In this case, module 2110 can increase the parameter so that the probability of selecting a loop with a high tension tag increases when the user requests high tension music. Thus, machine learning can be based on various information including composition output, feedback information, user control input, etc.
[0267] In the illustrated embodiment, composition module 2120 includes a section sequencer 2122, a section arranger 2124, a technical implementation module 2126, and a loop selection module 2128. In some embodiments, composition module 2120 arranges and constructs sections of a composition based on loop metadata and user control input (e.g., mood control).
[0268] In some embodiments, the section sequencer 2122 arranges different types of sections. In some embodiments, the section sequencer 2122 implements a finite state machine for continuously outputting the next type of section during operation. For example, the composition module 2120 may be configured to use different types of sections such as an intro, a build-up, a drop, a breakdown, and a bridge, which will be described in more detail below with reference to FIG. 23. Further, each section can include a plurality of subsections that define how the music changes throughout the section, for example, a transition-in subsection, a main content subsection, and a transition-out subsection.
[0269] The section arranger 2124, in some embodiments, constructs subsections according to arrangement rules. For example, one rule can be specified for transition-in by gradually adding tracks. In another rule, transition-in can be specified by gradually increasing the gain of a set of tracks. In another rule, it can be specified to chop vocal loops to create a melody. In some embodiments, the probability of a loop in the loop library added to a track is a function of user input parameters such as the current position within a section or subsection, loops that temporally overlap on another track, and mood variables. The function can be adjusted, for example, by adjusting coefficients based on machine learning.
[0270] In some embodiments, the technical implementation module 2120 is configured to facilitate section placement, for example, by adding rules specified by an artist or determined by analyzing the compositions of a particular artist. "Technique" can describe how a particular artist implements placement rules at the technical level. For example, in a placement rule that specifies a transition in by gradually adding tracks, one technique may indicate adding tracks in the order of drums, bass, pads, and then vocals, while another technique can indicate adding tracks in the order of bass, pads, vocals, and then drums. Similarly, in the case of an array rule that specifies chopping a vocal loop to create a melody, the technique may indicate chopping the vocal at a beats per second and repeating the chopped section of the loop 2 times before moving on to the next chopped section.
[0271] In the illustrated embodiment, the loop selection module 2128 selects loops according to placement rules and techniques for inclusion in the section by the section arranger 2124. When a section is complete, a corresponding execution script may be generated and sent to the execution module 2130. The execution module 2130 may receive execution script portions at various granularities. This includes, for example, the overall execution script for a performance of a certain length, the execution script for each section, the execution script for each sub-section, and the like. In some embodiments, placement rules, techniques, or loop selection are implemented statistically, for example, with different approaches using different time percentages.
[0272] In the illustrated embodiment, the execution module 2130 includes a filter module 2131, an effect module 2132, a mix module 2133, a master module 2134, and an execution module 2135. In some embodiments, these modules process an execution script and generate music data in a format supported by the audio output device 2140. The execution script can specify loops to be played, effects (e.g., per track or per subsection) to be applied by module 2132 when they should be played, filters to be applied by module 2131, and the like.
[0273] For example, the execution script can be specified to apply a low-pass filter ramp from 1000 to 20000 Hz to a specific track. As another example, the execution script can be specified to apply reverb with a 0.2 wet setting (5000 to 15000 milliseconds) to a specific track.
[0274] In some embodiments, the mix module 2133 is configured to perform automatic level control on the tracks to be combined. In some embodiments, the mix module 2133 also performs mixing by using frequency domain analysis of the combined tracks to measure frequencies with excessive or insufficient energy and applying gains to tracks in different frequency bands. In some embodiments, the master module 2134 is configured to perform multiband compression, equalization (EQ), or limiting procedures to generate data for final formatting by the execution module 2135. The embodiment of FIG. 21 can automatically generate various output music contents according to user input or other feedback information, while machine learning techniques can improve the user experience over time.
[0275] FIG. 22 is a diagram showing an example of an enhanced section of music content according to some embodiments. The system of FIG. 21 can configure such a section by applying placement rules and techniques. In the illustrated example, the build-up section includes three sub-sections and separate tracks for vocals, pads, drums, bass, and white noise.
[0276] In the illustrated example, the drum loop A is included in the sub-section transition, which is also repeated in the main content sub-section. The sub-section transition also includes the bass loop A. As shown, the gain of the section is low and increases linearly throughout the section (although non-linear increases and decreases are conceivable). In the illustrated example, the main content and the sub-sections of the transition out include various vocal, pad, drum, and bass loops. As described above, the disclosed techniques for automatically sequencing, arranging, and implementing sections can generate an almost infinite stream of output music content based on various user-adjustable parameters.
[0277] In some embodiments, the computer system displays an interface similar to FIG. 22 and enables an artist to specify the techniques used to compose a section. For example, an artist can create a structure as shown in FIG. 22 and parse it into code for the composition module.
[0278] FIG. 23 is a diagram showing an exemplary technique for placing sections of music content according to some embodiments. In the illustrated embodiment, the generated stream 2310 includes a plurality of sections 2320, each including a start sub-section 2322, a development sub-section 2324, and a transition sub-section 2326. In the example shown, a plurality of types of each section / sub-section are shown in a table connected by dotted lines. In the illustrated embodiment, the circular elements are examples of placing tools, which can be further implemented using specific techniques as discussed below. As shown, various compositional decisions can be performed pseudo-randomly according to statistical percentages. For example, the type of sub-section, the arrangement tool for a particular type or sub-section, or the technique used to implement the arrangement tool can be determined statistically.
[0279] In the example shown, a given section 2320 is one of five types: intro, build-up, drop, breakdown, and bridge, each having a different function to control the intensity of the entire section. In this example, the state sub-section is one of three types: slow build, sudden shift, or minimal, each type being different. The development sub-section in this example is any of three types: reduce, transform, augment. In this example, the transition sub-section is any of three types: collapse, ramp, hint. Different types of sections and sub-sections may be selected, for example, based on rules or pseudo-randomly.
[0280] In the illustrated example, the behavior of different sub-section types is implemented using one or more arrangement tools. In the case of a slow build, in this example, a low-pass filter is applied for 40% of the time and layers are added for 80% of the time. In the conversion progress sub-section, in this example, 25% of the time loop is chopped. Various additional arrangement tools are shown, such as one-shot, dropout beat, application of reverb, addition of pads, addition of themes, removal of layers, white noise, etc. These examples are included for illustrative purposes and are not intended to limit the scope of the present disclosure. Further, for ease of explanation, these examples may not be complete (e.g., actual arrangements may typically include a much larger number of arrangement rules).
[0281] In some embodiments, one or more arrangement tools may be implemented using a particular technique (which may be an artist specified or determined based on analysis of the artist's content). For example, a one-shot may be implemented using a sound effect or a vocal, loop chopping may be implemented using a stutter or half-chopping technique, layer removal may be implemented by vocal synthesis or removal, white noise may be implemented using a ramp or pulse function, etc. In some embodiments, the particular technique selected for a given arrangement tool may be selected according to a statistical function (e.g., 30% of layer removal may remove the composition and 70% of the time that may remove the vocal for a given artist). As described above, arrangement rules or techniques may be automatically determined, for example, by analyzing existing compositions using machine learning.
[0282] Method Example FIG. 24 is a flowchart method for using a ledger according to some embodiments. The method shown in FIG. 24 can be used in particular with any of the computer circuits, systems, devices, elements, or components disclosed herein. In various embodiments, some of the method elements shown can be implemented simultaneously in a different order than shown or can be omitted. Additional method elements can also be executed as desired.
[0283] In the illustrated embodiment, at 2410, the computing device determines playback data indicative of the characteristics of the playback of the music content mix. The mix may include a determined combination of a plurality of audio tracks (note that the combination of tracks may be determined in real time, for example, immediately prior to the output of the current portion of the music content mix that is a continuous stream of content). This determination may be based on the composition of the content mix (e.g., by a playback device such as a server or a mobile phone) or received from another device that determines which audio files to include in the mix. The playback data may be stored (e.g., in an offline mode) or encrypted. The playback data may be reported periodically or in response to a particular event (e.g., the restoration of connectivity to the server).
[0284] In the illustrated embodiment, the computing device records information specifying individual playback data for one or more of the plurality of audio tracks within the music content mix in an electronic blockchain ledger data structure. In the illustrated embodiment, the information specifying the individual playback data for an individual audio track includes usage data for the individual audio track and signature information associated with the individual audio track.
[0285] In one aspect, the characteristic information is an identifier for one or more entities. For example, the signature information may be a string or a unique identifier. In other embodiments, the signature information may be obfuscated by encryption or other means to avoid others from identifying the entity. In some embodiments, the usage data includes at least one of the time played for the music content mix or the number of times played for the music content mix.
[0286] In some embodiments, data identifying individual audio tracks within a music content mix is retrieved from a data store that also indicates actions to be performed in connection with including the one or more individual audio tracks. In these embodiments, the recording can include recording an indication of the performance of the indicated actions.
[0287] In some embodiments, the system determines rewards for a plurality of entities associated with a plurality of audio tracks based on information identifying individual playback data recorded in an electronic blockchain ledger.
[0288] In some embodiments, the system determines usage data for a first individual audio track that is not included in the music content mix in its original musical form. For example, the audio track can be modified and used to generate a new audio track, and the usage data can be adjusted to reflect this modification or use. In some embodiments, the system generates a new audio track based on interpolation between vector representations of audio in at least two of the plurality of audio tracks, and the usage data is determined based on the distance between the vector representation of the first individual audio track and the vector representation of the new audio track. In some embodiments, the usage data is based on the ratio of the interpolated vector representations and the Euclidean distance from the vectors in at least two of the plurality of audio tracks.
[0289] FIG. 25 is a flowchart of a method of using an image representation to combine audio files according to some embodiments. The method shown in FIG. 25 can be used in particular with any of the computer circuits, systems, devices, elements, or components disclosed herein. In various embodiments, some of the illustrated method elements can be performed simultaneously in a different order than shown, or can be omitted. Additional method elements can also be performed as desired.
[0290] In the illustrated embodiment, the computing device generates (2510) a plurality of image representations of a plurality of audio files, where an image representation for a particular audio file is generated based on data in the particular audio file and the MIDI representation of the particular audio file. In some embodiments, the pixel values in the image representation represent the tempo in the audio file where the image representation is compressed at the tempo resolution.
[0291] In some embodiments, the image representation is a two-dimensional representation of the audio file. In some embodiments, the pitch is represented by columns of the two-dimensional representation for time, and the pixel values of the two-dimensional representation represent tempo in rows of the two-dimensional representation. In some embodiments, the pitch is represented by columns of the two-dimensional representation for time, and the pixel values of the two-dimensional representation represent tempo in rows of the two-dimensional representation. In some embodiments, the pitch axis is banded into two sets of octaves in the range of 8 octaves, the first 12 rows of pixels represent the first 4 octaves that determine the pixel values of the pixels, the second 12 rows of pixels represent the second 4 octaves of the pixels, and the second 4 octaves that determine the pixel values of the pixels represent one of the second 4 octaves. In some embodiments, odd pixel values along the time axis represent the start of a note, and even pixel values along the time axis represent the continuation of a note. In some embodiments, each pixel represents a part of a beat in the time dimension.
[0292] In the illustrated embodiment, the arithmetic unit selects a plurality of audio files based on a plurality of image representations (2520).
[0293] In the illustrated embodiment, the arithmetic unit combines a plurality of music files to generate output music content (2530).
[0294] In some embodiments, one or more configuration rules are applied to select a plurality of audio files based on a plurality of image representations. In some embodiments, applying one or more composition rules includes removing pixel values in an image representation that exceed a first threshold and removing pixel values in an image representation that are below a second threshold.
[0295] In some embodiments, one or more machine learning algorithms are applied to an image representation to select and combine a plurality of audio files to generate output music content. In some embodiments, harmony and rhythm coherence are tested in the output music content.
[0296] In some embodiments, a single image representation is generated from a plurality of image representations, and a description of texture features is added to the single image representation from which texture features are extracted from a plurality of audio files. In some embodiments, the single image representation is stored together with a plurality of audio files. In some embodiments, a plurality of audio files are selected by applying one or more composition rules to the single image representation.
[0297] FIG. 26 is a flowchart of a method of implementing a control element generated by a user according to some embodiments. The method shown in FIG. 26 can be used, in particular, with any of the computer circuits, systems, devices, elements, or components disclosed herein. In various embodiments, some of the method elements shown can be implemented simultaneously in a different order than shown or omitted. Additional method elements can also be executed as desired.
[0298] In the illustrated embodiment, at 2610, the computing device accesses a plurality of audio files. In some embodiments, the audio files are accessed from the memory of the computer system, and the user has rights to the accessed audio files.
[0299] In the illustrated embodiment, at 2620, the computing device generates output music content by combining music content from two or more audio files using at least one trained machine learning algorithm. In some embodiments, the combination of music content is determined by at least one trained machine learning algorithm based on the music content within two or more audio files. In some embodiments, at least one trained machine learning algorithm combines music content by sequentially selecting music content from two or more audio files based on the music content within the two or more audio files.
[0300] In some embodiments, at least one trained machine learning algorithm is trained to select music content for beats coming after a specified time based on the metadata of the music content played up to the specified time. In some embodiments, at least one trained machine learning algorithm is further trained to select music content for beats coming after a specified time based on the level of a control element.
[0301] In the illustrated embodiment, the arithmetic unit implements, on the user interface, a control element generated by the user for a change in a user-specified parameter in the generated output music content, and the levels of one or more audio parameters in the generated output music content are determined based on the level of the control element, and the relationship between the levels of the one or more audio parameters and the level of the control element is determined based on user input during at least one music playback session. In some embodiments, the level of the user-specified parameter changes based on one or more environmental conditions.
[0302] In some embodiments, the relationship between the levels of one or more audio parameters and the level of the control element is determined by playing a plurality of audio tracks during at least one music playback session, the plurality of audio tracks having changing audio parameters, receiving, for each of the audio tracks, an input specifying the level of a parameter within the audio track selected by the user, evaluating, for each of the audio tracks, the level of one or more audio parameters within the audio track, and determining the relationship between the levels of the one or more audio parameters and the level of the control element based on the correlation between each level of the parameter selected by the user and each level of the one or more audio parameters.
[0303] In some embodiments, the relationship between the levels of one or more audio parameters and the levels of control elements is determined using one or more machine learning algorithms. In some embodiments, the relationship between the levels of one or more audio parameters and the levels of control elements is improved based on user variations in the levels of control elements during playback of the generated output music content. In some embodiments, the levels of one or more audio parameters within an audio track are evaluated using metadata from the audio track. In some embodiments, the relationship between the levels of one or more audio parameters and the levels of user-specified parameters is further based on additional user input during one or more additional music playback sessions.
[0304] In some embodiments, the computing device implements at least one additional control element generated by the user on the user interface to vary additional user-specified parameters within the generated output music content. Here, the additional user-specified parameters are sub-parameters of the user-specified parameters. In some embodiments, the generated output music content is modified based on user adjustment of the levels of the control elements. In some embodiments, a feedback control element is implemented on the user interface, and the feedback control element enables the user to provide positive or negative feedback on the generated output music content during playback. In some embodiments, at least one trained machine algorithm modifies the generation of subsequent generated output music content based on the feedback received during playback.
[0305] FIG. 27 is a flowchart of a method for generating music content by changing audio parameters according to some embodiments. The method shown in FIG. 27 can be used in particular with any of the computer circuits, systems, devices, elements, or components disclosed herein. In various embodiments, some of the method elements shown can be performed simultaneously in a different order than shown or can be omitted. Additional method elements can also be executed as desired.
[0306] In the illustrated embodiment, at 2710, the computing device accesses a set of music content. In some embodiments.
[0307] In the illustrated embodiment, the computing device generates a first graph of the audio signal of the music content, where the first graph is a graph of the audio parameters versus time.
[0308] In the illustrated embodiment, the computing device generates a second graph of the audio signal of the music content, where the second graph is a signal graph of the audio parameters versus beats. In some embodiments, the second graph of the audio signal has a similar structure to the first graph of the audio signal.
[0309] In the illustrated embodiment, the computing device generates new music content from the reproduced music content by changing the audio parameters in the reproduced music content, where the audio parameters are changed based on a combination of the first graph and the second graph.
[0310] In some embodiments, the audio parameters of the first graph and the second graph are defined by the nodes of the graph that determine the changes in the properties of the audio signal. In some embodiments, the step of generating new music content includes receiving the reproduced music content, determining a first node in the first graph corresponding to the audio signal in the reproduced music content, determining a second node in the second graph corresponding to the first node, determining one or more specific audio parameters based on the second node, and changing one or more properties of the audio signal in the reproduced music content by changing the specific audio parameters. In some embodiments, one or more additional specified audio parameters are determined based on the first node, and one or more properties of the additional audio signals in the reproduced music content are changed by changing the additional specified audio parameters.
[0311] In some embodiments, the step of determining one or more audio parameters includes determining a part of the second graph to implement the audio parameters based on the position of the second node in the second graph, and selecting one or more audio parameters as one or more audio specified parameters from the determined part of the second graph. In some embodiments, by changing one or more specified audio parameters, the part of the reproduced music content corresponding to the determined part of the second graph is changed. In some embodiments, the changed properties of the audio signal in the reproduced music content include signal amplitude, signal frequency, or a combination thereof.
[0312] In some embodiments, one or more automations are applied to audio parameters, and at least one of the at least one automation is a pre-programmed temporal manipulation of at least one audio parameter. In some embodiments, one or more modulations are applied to the audio parameters, wherein at least one of the at least one modulation multiplicatively modifies at least one audio parameter over at least one automation.
[0313] The following numbered items describe various non-limiting embodiments disclosed in the present specification. Set A (A1) A method comprising: generating, by a computer system, a plurality of graphical representations of a plurality of audio files, wherein a graphical representation of a particular audio file is generated based on data within the particular audio file and a MIDI representation of the particular audio file; selecting, based on the plurality of graphical representations, a plurality of the audio files of the audio files; combining the plurality of the audio files of the audio files to produce output music content; and a method comprising the steps. (A2) The method according to any of the preceding items in Set A, wherein the pixel values of the graphical representation represent the tempo of the audio file, and the graphical representation is compressed at a tempo resolution. (A3) The method according to any of the preceding items in Set A, wherein the graphical representation is a two-dimensional representation of the audio file. (A4) The method according to any of the preceding items in Set A, wherein pitch is represented by rows of the two-dimensional representation, time is represented by columns of the two-dimensional representation, and pixel values of the two-dimensional representation represent tempo. (A5) The two-dimensional representation is 32 pixels wide × 24 pixels high, and each pixel represents a small part of a beat in the time dimension, by any of the methods of the foregoing items in Set A. (A6) The pitch axis is banded into two octave sets in the range of 8 octaves, and the first 12 rows of pixels represent the first 4 octaves having pixel values of the pixels that determine which of the first 4 octaves is represented, and the second 12 rows of pixels represent the second 4 octaves having pixel values of the pixels that determine which of the second 4 octaves is represented, by any of the methods of the foregoing items in Set A. (A7) Odd pixel values along the time axis represent the start of a note, and even pixel values along the time axis represent the continuation of a note, by any of the methods of the foregoing items in Set A. (A8) A step of applying one or more pre-set composition rules to select a plurality of audio files of the audio file based on the plurality of image representations, by any of the methods of the foregoing items in Set A. (A9) The step of applying one or more composition rules includes removing pixel values in image representations that exceed a first threshold and removing pixel values in image representations that are below a second threshold, by any of the methods of the foregoing items in Set A. (A10) A step of selecting and combining the plurality of audio files among the audio files and applying one or more machine learning algorithms to the image representation to generate the output music content, by any of the methods of the foregoing items in Set A. (A11) A step of testing the harmony and rhythm coherence in the output music content, by any of the methods of the foregoing items in Set A. (A12) A step of generating a single image representation from the plurality of image representations; extracting one or more texture features from the plurality of audio files; adding a description of the extracted texture features to the single image representation; Any method of the foregoing items within set A, further comprising. (A13) A non-transitory computer-readable medium having stored instructions, executable by a computing device to perform operations including any combination of the operations performed by any method of the foregoing items within set A. (A14) An apparatus, one or more processors; one or more memories storing program instructions, comprising, the program instructions being executable by the one or more processors to perform any combination of the operations performed by any method of the foregoing items within set A. Set B (B1) A method, accessing, by a computer system, a set of music content; generating, by the computer system, a first graph of an audio signal of the music content, the first graph being a graph of an audio parameter versus time; generating, by the computer system, a second graph of the audio signal of the music content, the second graph being a signal graph of the audio parameter versus beats; generating, by the computer system, new music content from the played music content by changing the audio parameter in the played music content, the audio parameter being changed based on a combination of the first graph and the second graph; A method comprising. (B2) Any of the methods of the foregoing items within Set B, wherein the second graph of the audio signal has a structure similar to that of the first graph of the audio signal. (B3) Any of the methods of the foregoing items within Set B, wherein the audio parameters in the first graph and the second graph are defined by nodes in a graph that determines changes in the characteristics of the audio signal. (B4) The step of generating new music content includes the step of receiving the reproduced music content, the step of determining a first node in a first graph corresponding to an audio signal in the reproduced music content, and the step of determining a second node in a second graph corresponding to the first node, the step of determining one or more specific audio parameters based on the second node, and the step of changing one or more properties of the audio signal in the reproduced music content by changing the specific audio parameters, and Any of the methods of the foregoing items within Set B. (B5) the step of determining one or more additional specified audio parameters based on the first node, and the step of changing one or more properties of an additional audio signal in the reproduced music content by modifying the additional specified audio parameters, and Any of the methods according to any of the foregoing items within Set B. (B6) The step of determining the one or more audio parameters includes the step of determining a portion of the second graph that implements the audio parameter based on the position of the second node in the second graph, and the step of selecting the audio parameter as the one or more audio specified parameters from the determined portion of the second graph, and Any of the methods according to any of the foregoing items within Set B. (B7) The step of changing the one or more specified audio parameters is the method according to any of the foregoing items in set B that changes a portion of the reproduced music content corresponding to the determined portion of the second graph. (B8) The property of the audio signal in the reproduced music content that is changed is the method according to any of the foregoing items in set B, including signal amplitude, signal frequency, or a combination thereof. (B9) The step of applying one or more automations to the audio parameters, wherein at least one of the automations is a pre-programmed temporal manipulation of at least one audio parameter, the method according to any of the foregoing items in set B that further includes this step. (B10) The step of applying one or more modulations to the audio parameters, wherein at least one of the modulations multiplicatively changes at least one audio parameter on top of at least one automation, the method according to any of the foregoing items in set B that further includes this step. (B11) The method according to any of the foregoing items in set B that further includes the step of providing, by the computer system, a heap-allocated memory for storing one or more objects associated with the set of music content, wherein the one or more objects include at least one of the audio signal, the first graph, and the second graph. (B12) The method according to any of the foregoing items in set B that further includes the step of storing the one or more objects in a data structure list within the heap-allocated memory, wherein the data structure list includes a serialized list of links for the objects. (B13) A non-transitory computer-readable medium having stored instructions, executable by a computing device to perform operations including any combination of the operations performed by the method according to any of the foregoing items in set B. (B14) A device comprising: One or more processors; One or more memories storing program instructions; The program instructions being executable by the one or more processors to perform any combination of operations executed by any of the methods of the foregoing items in Set B. Set C (C1) A method comprising: Accessing a plurality of audio files by a computer system; Generating output music content by combining music content from two or more audio files using at least one trained machine learning algorithm; Implementing a control element created by a user to vary user-specified parameters in the generated output music content on a user interface associated with the computer system; And The level of one or more audio parameters in the generated output music content is determined based on the level of the control element, and the relationship between the level of the one or more audio parameters and the level of the control element is based on user input during at least one music playback session. (C2) The combining of the music content is determined by the at least one trained machine learning algorithm based on the music content in two or more audio files, according to any of the methods of the foregoing items in Set C. (C3) The at least one trained machine learning algorithm combines the music content by sequentially selecting music content from the two or more audio files based on the music content in the two or more audio files, according to any of the methods of the foregoing items in Set C. (C4) Any of the methods of the preceding items in set C, wherein the at least one trained machine learning algorithm is trained to select music content for beats coming after the specified time based on metadata of music content played up to the specified time. (C5) Any of the methods of the preceding items in set C, wherein the at least one trained machine learning algorithm is further trained to select music content for beats coming after the specified time based on the level of the control element. (C6) The relationship between the levels of the one or more audio parameters and the level of the control element is During the at least one music playback session, a plurality of audio tracks are played, the plurality of audio tracks having varying audio parameters, For each of the audio tracks, an input is received specifying the level selected by the user of the user-specified parameter within the audio track, For each of the audio tracks, the levels of one or more audio parameters within the audio track are evaluated, Based on the correlation between each of the levels selected by the user of the user-specified parameter and each of the evaluated levels of the one or more audio parameters, determining the relationship between the levels of the one or more audio parameters and the level of the control element, Any of the methods of the preceding items in set C, determined thereby. (C7) Any of the methods of the preceding items in set C, wherein the relationship between the levels of the one or more audio parameters and the level of the control element is determined using one or more machine learning algorithms. (C8) A method according to any of the preceding items in set C, further comprising the step of finely adjusting the relationship between the level of the one or more audio parameters and the level of the control element based on a user variation in the level of the control element during playback of the generated output music content. (C9) A method according to any of the preceding items in set C, wherein the level of the one or more audio parameters in the audio track is evaluated using metadata from the audio track. (C10) A method according to any of the preceding items in set C, further comprising the step of varying, by the computer system, the level of the user-specified parameter based on one or more environmental conditions. (C11) A method according to any of the preceding items in set C, further comprising implementing, on the user interface associated with the computer system, at least one additional control element generated by the user for variation of an additional user-specified parameter in the generated output music content, wherein the additional user-specified parameter is a sub-parameter of the user-specified parameter. (C12) A method according to any of the preceding items in set C, further comprising accessing the audio file from the memory of the computer system, wherein the user has rights to the accessed audio file. (C13) A non-transitory computer-readable medium having stored instructions, executable by a computing device to perform operations including any combination of the operations performed by a method according to any of the preceding items in set C. (C14) An apparatus, one or more processors; one or more memories storing program instructions; A device that includes and whose program instructions are executable by the one or more processors to perform any combination of operations executed by any of the methods of the foregoing items in set C. Set D (D1) A method comprising: A step of determining playback data of a music content mix by a computer system, wherein the playback data indicates characteristics of playback of the music content mix, and the music content mix includes a determined combination of a plurality of audio tracks; A step of recording, by the computer system, in an electronic blockchain ledger data structure, information specifying individual playback data for one or more of the plurality of audio tracks in the music content mix, wherein the information specifying individual playback data for an individual audio track includes usage data for the individual audio track and signature information associated with the individual audio track; A method comprising. (D2) The usage data includes at least one of the time the music content mix was played or the number of times the music content mix was played, according to any of the methods of the foregoing items in set D. (D3) Any of the methods of the foregoing items in set D, further comprising a step of determining the information specifying the individual playback data for the individual audio track based on the playback data of the music content mix and data identifying the individual audio tracks in the music content mix. (D4) The data identifying the individual audio tracks in the music content mix is retrieved from a data store that also indicates operations to be performed in connection with including one or more individual audio tracks, and the recording includes recording an indication of proof of execution of the indicated operations, according to any of the methods of the foregoing items in set D. (D5) A method according to any of the preceding items in set D, further comprising the step of identifying one or more entities associated with the individual audio tracks based on signature information for a plurality of entities associated with a plurality of audio tracks and data specifying the signature information. (D6) A method according to any of the preceding items in set D, further comprising the step of determining a reward for the plurality of entities associated with the plurality of audio tracks based on information identifying individual playback data recorded in the electronic blockchain ledger. (D7) A method according to any of the preceding items in set D, wherein the playback data includes usage data of one or more machine learning modules used to generate the music mix. (D8) A method according to any of the preceding items in set D, further comprising the step of storing, at least temporarily, the playback data of the music content mix by a playback device. (D9) A method according to any of the preceding items in D, wherein the stored playback data is communicated by the playback device to another computer system periodically or in response to an event. (D10) A method according to any of the preceding items in set D, further comprising the step of encrypting, by the playback device, the playback data of the music content mix. (D11) A method according to any of the preceding items in set D, wherein the blockchain ledger is publicly accessible and immutable. (D12) A method according to any of the preceding items in set D, further comprising the step of determining usage data of a first individual audio track not included in the music content mix in its original music format. (D13) A method according to any of the preceding items in set D, further comprising the step of generating a new audio track based on interpolation between vector representations of audio in at least two of the plurality of audio tracks. The step of determining the usage data is any of the methods of the foregoing items in set D based on the distance between the vector representation of the first individual audio track and the vector representation of the new audio track. (D14) A non-transitory computer-readable medium having stored instructions, executable by a computing device to perform operations including any combination of operations performed by any of the methods of the foregoing items in set D. (D15) An apparatus, one or more processors, one or more memories storing program instructions, comprising, the program instructions being executable by the one or more processors to perform any combination of operations performed by any of the methods of the foregoing items in set D. Set E (E1) A method, storing, by a computing system, data specifying a plurality of tracks of a plurality of different entities, the data including signature information for each track of the plurality of music tracks; layering, by the computing system, the plurality of tracks to generate output music content; recording, in an electronic blockchain ledger data structure, information specifying the tracks included in the layering and the signature information for each of the tracks; A method comprising. (E2) Any of the methods in set E, wherein the blockchain ledger is publicly accessible and immutable. (E3) detecting a use of one of the tracks by an entity that does not match the signature information; Any of the methods in set E further comprising. (E4) A method comprising: storing, by a computing system, data specifying a plurality of tracks for a plurality of different entities, the data including metadata linking to other content; selecting and layering, by the computing system, a plurality of tracks to generate output music content; searching, based on the metadata of the selected tracks, for content to be output in relation to the output music content; A method comprising the above steps. (E5) Any method in set E, wherein the other content includes visual advertisements. (E6) A method comprising: storing, by a computing system, data specifying a plurality of tracks; causing, by the computing system, output of a plurality of music examples; receiving, by the computing system, user input indicating the user's opinion on whether one or more of the music examples exhibit a first music parameter; storing, based on the receiving step, custom rule information for the user; receiving, by the computing system, user input indicating an adjustment to the first music parameter; selecting and layering, by the computing system, a plurality of tracks based on the custom rule information to generate output music content; A method comprising the above steps. (E7) Any method in set E, wherein the user specifies the name of the first music parameter. (E8) The user specifies one or more goals for one or more ranges of the first music parameter, and the selecting step and the layering step are any method of the items in set E based on past feedback information. (E9) In response to determining that the selecting step and the layering step based on the past feedback information did not meet one or more goals, adjusting a plurality of other music attributes and determining the effect of adjusting the plurality of other music attributes for one or more goals, any method of the items in set E that further includes. (E10) Any method of the items in set E, where the goal is the user's heart rate. (E11) Any method of the items in set E, where the selecting step and the layering step are performed by a machine learning engine. (E12) Any method of the items in set E that further includes a step of training the machine learning engine, including training a teacher model based on a plurality of types of user input and training a student model based on the received user input. (E13) A method, Storing data for specifying a plurality of tracks by a computing system, Determining by the computing system that the audio device is located in a first type of environment, Receiving user input for adjusting a first music parameter while the audio device is placed in the first type of environment, Selecting and layering a plurality of tracks by the computing system based on the adjusted first music parameter and the first type of environment to generate output music content via the audio device, The step of determining, by the computing system, that the audio device is located in a second type of environment; The step of receiving a user input for adjusting a first music parameter while the audio device is placed in the second type of environment; The step of selecting and layering a plurality of tracks based on the adjusted first music parameter and the second type of environment by the computing system, and outputting music content via the audio device; A method comprising. (E14) The step of selecting and layering adjusts different music attributes based on an adjustment to the first music parameter when the audio device is in the first type of environment, compared to when the audio device is in the second type of environment, any method of an item in set E. (E15) A method comprising: The step of storing, by a computing system, data specifying a plurality of tracks; The step of storing, by the computing system, data specifying music content; The step of determining one or more parameters of the stored music content; The step of selecting and layering a plurality of tracks by the computing system to generate output music content based on one or more parameters of the stored music content; The step of causing the output of the output music content and the stored music content to overlap in time; A method comprising. (E16) A method comprising: The step of storing, by a computing system, data specifying a plurality of tracks and corresponding track attributes; The step of storing, by the computing system, data specifying music content; Determining one or more parameters of the stored music content and determining one or more tracks included in the stored music content; Selecting, by the computing system, tracks among the stored tracks and tracks among the determined tracks, layering them, and generating output music content based on the one or more parameters; A method comprising. (E17) A method comprising: Storing, by a computing system, data specifying a plurality of tracks; Selecting and layering, by the computing system, a plurality of tracks to generate output music content; Including The selection is A neural network module configured to select tracks for future use based on most recently used tracks; One or more hierarchical hidden Markov models configured to select a structure for future output music content based on the selected structure and to constrain the neural network module; A method performed using both. (E18) Further comprising training one or more hierarchical hidden Markov models based on positive or negative user feedback regarding the output music content, a method according to any of the items in set E. (E19) Further comprising training the neural network model based on user selections provided in response to a plurality of examples of music content, a method according to any of the items in set E. (E20) Any method in set E, wherein the neural network module includes a plurality of layers including at least one globally trained layer and at least one layer specifically trained based on feedback from a specific user account. (E21) A method comprising: storing, by a computing system, data specifying a plurality of tracks; selecting and layering, by the computing system, a plurality of tracks to generate output music content; wherein the selection is performed using both a globally trained machine learning module and a locally trained machine learning module. (E22) Any method in set E, comprising training the globally trained machine learning module using feedback from a plurality of user accounts and training the locally trained machine learning module based on feedback from a single user account. (E23) Any method in set E, wherein both the globally trained machine learning module and the locally trained machine learning module provide output information based on user adjustment of music attributes. (E24) Any method in set E, wherein the globally trained machine learning module and the locally trained machine learning module are included in different layers of a neural network. (E25) A method comprising: receiving a first user input indicating a value of a high-level music composition parameter; Automatically selecting and combining audio tracks based on the values of the music composition parameters at the upper level, including using a plurality of values of one or more sub-parameters associated with the music composition parameters at the upper level during the output of music content, and steps; Receiving a second user input indicating one or more values of the one or more sub-parameters; Based on the second user input, the step of automatically selecting and combining; A method including. (E26) The music composition parameter at the upper level is an energy parameter, and the one or more sub-parameters include any of the items in set E including tempo, number of layers, vocal parameters, or bass parameters. Any method of the items in set E. (E27) A method, Accessing score information by a computing system; Synthesizing music content based on the score information by the computing system; Analyzing the music content to generate frequency information regarding a plurality of frequency bins at different time points; Training a machine learning engine, including inputting the frequency information and using the score information as a label for the training; A method including. (E28) Inputting music content into the trained machine learning engine; Generating score information of the music content using the machine learning engine; Further including, and the generated score information includes frequency information of a plurality of frequency bins at different time points. Any method of the items in set E. (E29) Any method of the items in set E where the different time points have a certain distance between them. (E30) Any method of the items in set E, wherein the different time points correspond to the beats of the music content. (E31) Any method of the items in set E, wherein the machine learning engine includes a convolutional layer and a recurrent layer. (E32) Any method of the items in set E, wherein the frequency information includes a binary indication of whether the music content contains content in a frequency bin at a certain time point. (E33) Any method of the items in set E, wherein the music content includes a plurality of different musical instruments.
[0314] Although specific embodiments have been described above, these embodiments are not intended to limit the scope of the present disclosure, even if only a single embodiment is described with respect to a particular feature. The examples of features provided in the present disclosure are intended to be illustrative rather than limiting, unless otherwise specified. As will be apparent to those skilled in the art who enjoy the benefits of the present disclosure, the foregoing description is intended to cover such alternatives, modifications, and equivalents.
[0315] The scope of the present disclosure includes any feature or combination of features (whether explicit or implicit) disclosed herein, or any generalization thereof, regardless of whether any or all of the problems solved in the present specification are alleviated. Accordingly, new claims may be formed during the examination of this application (or an application claiming priority based on this application) for any such combination of features. In particular, referring to the appended claims, the features of the dependent claims may be combined with the features of the independent claims, and the features of each independent claim may be combined in any suitable manner, rather than in a specific combination not listed in the scope of the appended claims.
Description of Reference Numerals
[0316] 110 Stored audio files and corresponding attributes 120 Stored rule set 130 Target music attributes 140 Output music content 150 Environmental information 160 Music generation module
Claims
1. 1. A method comprising: accessing, by a computer system, a plurality of audio files; generating output musical content by combining musical content from two or more audio files using at least one trained machine learning algorithm; implementing a user generated control element on a user interface associated with the computer system, the control element having a user generated label to describe a user defined parameter that controls a combination of a plurality of audio parameters in the generated output musical content, the audio parameter being different from the user defined parameter, a level of the control element on the user interface determining a level of the user defined parameter, the levels of the plurality of audio parameters in the generated output musical content being determined according to a relationship defining a level of each of the plurality of audio parameters based on the level of the user defined parameter, the relationship being determined from user input during a playback session of a plurality of audio tracks, the user input being a selection of a level of the user defined parameter in each of the plurality of audio tracks played during the playback session, the relationship being based on an evaluation of a correlation between the level of the audio parameter and a selected level of the user defined parameter for each of the plurality of audio tracks; The method includes:
2. The method of claim 1 , wherein the musical content combinations are determined by the at least one trained machine learning algorithm based on musical content in the two or more audio files.
3. 2. The method of claim 1, wherein the at least one trained machine learning algorithm combines musical content by sequentially selecting musical content from the two or more audio files based on musical content within the two or more audio files.
4. 2. The method of claim 1, wherein the at least one trained machine learning algorithm is trained to select musical content for a beat that comes after a specified time based on metadata of musical content played up to the specified time.
5. 5. The method of claim 4, wherein the at least one trained machine learning algorithm is further trained to select musical content for a beat that comes after the specified time based on a level of the control element.
6. 2. The method of claim 1 , wherein the relationship defining a level of each of the plurality of audio parameters based on a level of the user-defined parameter is determined by evaluating the level of the audio parameter and a selected level of the user-defined parameter for each of the plurality of audio tracks using one or more machine learning algorithms.
7. 2. The method of claim 1, wherein the user-defined parameters are parameters created by the user to enable custom control of a combination of the plurality of audio parameters, and the combination of the plurality of audio parameters that is custom controlled by the user-defined parameters is determined during the playback session.
8. 2. The method of claim 1, wherein adjusting the level of the control element adjusts the level of each of the plurality of audio parameters in the combination of the plurality of audio parameters according to the relationship that defines the level of each of the plurality of audio parameters based on the level of the user-defined parameter.
9. 2. The method of claim 1, further comprising: fine-tuning the relationship defining a level of each of the plurality of audio parameters based on a level of the user-defined parameter in accordance with a user variation of a level of the control element during playback of the generated output musical content.
10. The method of claim 1 , wherein the levels of the audio parameters in the multiple audio tracks are evaluated using metadata from the multiple audio tracks.
11. The method of claim 1 , further comprising the step of varying, by the computer system, levels of the user-defined parameters based on one or more environmental conditions.
12. implementing on the user interface associated with the computer system at least one additional control element created by the user for variation of an additional user-defined parameter within the generated output musical content; Further comprising: The method of claim 1 , wherein the additional user-defined parameter is a sub-parameter of the user-defined parameter, the additional user-defined parameter controlling another combination of a plurality of audio parameters.
13. The method of claim 1 , further comprising the step of accessing the audio file from a memory of the computer system, the user having rights to the accessed audio file.
14. A non-transitory computer-readable medium having instructions stored thereon, the instructions being executable by a computing device to perform operations, the operations including: accessing a plurality of audio files; generating output musical content by combining musical content from two or more audio files using at least one trained machine learning algorithm; implementing a user generated control element on a user interface, said control element having a user generated label to describe a user defined parameter that controls a combination of a plurality of audio parameters in the generated output musical content, said audio parameter being different from said user defined parameter, a level of said control element on said user interface determining a level of said user defined parameter, and levels of said plurality of audio parameters in the generated output musical content being determined according to a relationship defining a level of each of said plurality of audio parameters based on a level of said user defined parameter, said relationship being determined from a user input during a playback session of a plurality of audio tracks, said user input being a selection of a level of said user defined parameter in each of said plurality of audio tracks played during said playback session, said relationship being based on an evaluation of a correlation between a level of said audio parameter and a selected level of said user defined parameter for each of said plurality of audio tracks; A non-transitory computer readable medium comprising:
15. 15. The non-transitory computer-readable medium of claim 14, wherein the operations further comprise modifying the generated output musical content based on user adjustment of a level of the control element.
16. 15. The non-transitory computer-readable medium of claim 14, wherein the relationship defining a level of each of the plurality of audio parameters based on a level of the user-defined parameter is further based on additional user input during one or more additional music playback sessions.
17. 20. The non-transitory computer-readable medium of claim 16, wherein the operations further include implementing a feedback control element on the user interface, the feedback control element enabling the user to provide positive or negative feedback on the generated output musical content during playback.
18. 20. The non-transitory computer-readable medium of claim 17, wherein the at least one trained machine learning algorithm alters production of subsequent generated output musical content based on the feedback received during the playback.
19. An apparatus comprising: one or more processors; One or more memories having stored therein program instructions, the program instructions comprising: Access multiple audio files, generating output musical content by combining musical content from two or more audio files using at least one trained machine learning algorithm; implementing a user generated control element on a user interface, said control element having a user generated label to describe a user defined parameter that controls a combination of a plurality of audio parameters in the generated output musical content, said audio parameter being different from said user defined parameter, a level of said control element on said user interface determining a level of said user defined parameter, and levels of said plurality of audio parameters in the generated output musical content being determined according to a relationship defining a level of each of said plurality of audio parameters based on a level of said user defined parameter, said relationship being determined from a user input during a playback session of a plurality of audio tracks, said user input being a selection of a level of said user defined parameter in each of said plurality of audio tracks played during said playback session, said relationship being based on an evaluation of a correlation between a level of said audio parameter and a selected level of said user defined parameter for each of said plurality of audio tracks. The apparatus is executable by the one or more processors to perform the steps of:
20. 20. The device of claim 19, wherein the program instructions stored in the one or more memories are further executable to play the generated output musical content on the device and to enable the user to vary the level of the user-defined parameter by setting a level of the control element.
Citation Information
Patent Citations
Content generation device and content generation method
JP2006084749A
Device and program for creating music piece
JP2009020387A
Acoustic processing device
JP2015011146A
Musical tone information processing apparatus, and program
JP2015099358A