Music content generation
Through a combination of machine learning algorithms and deep learning neural networks, customized music content is generated, which solves the problem of users getting tired of the same songs in streaming services, and realizes personalized music generation and copyright protection.
Patent Information
- Application Number
- CN202510574406.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2020-08-21
- Filing Date
- 2021-02-11
- Publication Date
- 2025-08-26
AI Technical Summary
Existing streaming music services cannot adjust music according to users' tastes, environments, etc., resulting in users getting tired of hearing songs from the same genre, and song selection is limited by license agreements and quantity, making it impossible to achieve a personalized music experience.
Generate customized music content through machine learning algorithms, leverages a combination of deep learning neural networks and expert knowledge to generate music based on user input and environmental information, allowing users to create and train controls to adjust music preferences, and record playback data to prevent tampering.
It realizes personalized music generation based on user preferences and environment changes, improves the diversity and novelty of music selection, meets users' music needs, and records and protects copyrights.
Smart Images

Figure CN120544524A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with the application date of February 11, 2021, application number 202180013939.X, and name “Music Content Generation”. Technical Field
[0002] The present disclosure relates to audio engineering and, more particularly, to generating musical content. Background Art
[0003] Streaming music services typically provide songs to users via the Internet. Users can subscribe to these services and stream music through a web browser or application. Examples of such services include PANDORA, SPOTIFY, GROOVESHARK, etc. Users can often select music genres or specific artists to stream. Users can usually rate songs (e.g., using a star rating or like / dislike system), and some music services can customize which songs are streamed to users based on previous ratings. The cost of running a streaming service (which may include paying royalties for each streamed song) is typically covered by the cost of the user's subscription and / or advertisements that appear between songs.
[0004] Song selection may be limited by licensing agreements and the number of songs created for a particular genre. Users may become tired of hearing the same songs in a particular genre. Furthermore, these services may not be able to adapt the music to the user's taste, environment, behavior, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0005] Figure 1 is a schematic diagram illustrating an exemplary music generator.
[0006] Figure 2 is a block diagram illustrating an exemplary overview of a system for generating output music content based on input from multiple different sources, according to some embodiments.
[0007] Figure 3 is a block diagram illustrating an exemplary music generator system configured to output music content based on analysis of image representations of audio files, in accordance with some embodiments.
[0008] Figure 4 Depicts an example of an image representation of an audio file.
[0009] Figure 5A and Figure 5B Depicted are examples of grayscale images used for melody image feature representation and drum beat image feature representation, respectively.
[0010] Figure 6 is a block diagram illustrating an exemplary system configured to generate a single image representation in accordance with some embodiments.
[0011] Figure 7 Depicts an example of a single image representation of multiple audio files.
[0012] Figure 8 is a block diagram illustrating an exemplary system configured to implement user-created controls in music content generation, in accordance with some embodiments.
[0013] Figure 9 Depicted is a flow diagram of a method for training a music generator module based on user-created control elements, according to some embodiments.
[0014] Figure 10 is a block diagram illustrating an exemplary teacher / student framework system in accordance with some embodiments.
[0015] Figure 11 is a block diagram illustrating an exemplary system configured to implement audio techniques in music content generation, in accordance with some embodiments.
[0016] Figure 12 Depicts an example of an audio signal graph.
[0017] Figure 13 Depicts an example of an audio signal graph.
[0018] Figure 14 Depicted is an exemplary system for implementing real-time modification of musical content using an Audio Technology Music Generator module in accordance with some embodiments.
[0019] Figure 15 Depicted is a block diagram of an exemplary API module for audio parameter automation in a system according to some embodiments.
[0020] Figure 16 A block diagram of an exemplary memory area is depicted in accordance with some embodiments.
[0021] Figure 17 Depicted is a block diagram of an exemplary system for storing new music content in accordance with some embodiments.
[0022] Figure 18 is a schematic diagram illustrating example playback data according to some embodiments.
[0023] Figure 19 is a block diagram illustrating an example authoring system in accordance with some embodiments.
[0024] Figure 20A-Figure 20B is a block diagram illustrating a graphical user interface according to some embodiments.
[0025] Figure 21is a block diagram illustrating an example music generator system including analysis and composition modules according to some embodiments.
[0026] Figure 22 is a diagram illustrating example enhanced sections of music content according to some embodiments.
[0027] Figure 23 is a diagram illustrating an example technique for arranging chapters of music content according to some embodiments.
[0028] Figure 24 is a flowchart method for using a ledger according to some embodiments.
[0029] Figure 25 is a flowchart method for combining audio files using image representations according to some embodiments.
[0030] Figure 26 is a flowchart method for implementing a user-created control element according to some embodiments.
[0031] Figure 27 is a flowchart method for generating music content by modifying audio parameters according to some embodiments.
[0032] Although the embodiments disclosed herein are susceptible to various modifications and alternative forms, specific embodiments are shown by way of example in the drawings and described in detail herein. However, it should be understood that the drawings and detailed description thereof are not intended to limit the scope of the claims to the particular forms disclosed. On the contrary, this application is intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the disclosure of this application as defined by the appended claims.
[0033] This disclosure includes references to "one embodiment," "a specific embodiment," "some embodiments," "various embodiments," or "an embodiment." The appearances of the phrases "in one embodiment," "in a specific embodiment," "in some embodiments," "in various embodiments," or "in an embodiment" are not necessarily referring to the same embodiment. The particular features, structures, or characteristics may be combined in any suitable manner consistent with this disclosure.
[0034] Recitation in the appended claims of an element being "configured to" perform one or more tasks is expressly not intended to invoke 35 U.S.C. §112(f) with respect to that claim element. Therefore, none of the claims filed in this application are intended to be interpreted as having "means-plus-function" elements. If the applicant wishes to invoke §112(f) during patent application procedure, it will recite the claim element using the "means for [performing the function]" construction.
[0035] As used herein, the term "based on" is used to describe one or more factors that influence a determination. The term does not exclude the possibility that additional factors may influence the determination. That is, a determination may be based solely on the specified factors or on the specified factors as well as other unspecified factors. Consider the phrase "A is determined based on B." The phrase specifies that B is a factor used to determine A or that influences the determination of A. The phrase does not exclude that the determination of A may also be based on certain other factors, such as C. The phrase is also intended to encompass embodiments in which A is determined solely based on B. As used herein, the phrase "based on" is synonymous with the phrase "based at least in part on."
[0036] As used herein, the phrase "in response to" describes one or more factors that trigger an effect. The phrase does not exclude the possibility of additional factors that may influence or otherwise trigger the effect. That is, the effect may be responsive to only those factors, or it may be responsive to the specified factors as well as other unspecified factors.
[0037] As used herein, the terms "first," "second," and the like are used as labels for the nouns that precede them and do not imply any type of ordering (e.g., spatial, temporal, logical, etc.) unless otherwise indicated. As used herein, the term "or" is used as an inclusive or and not as an exclusive or. For example, the phrase "at least one of x, y, or z" means any one of x, y, and z, as well as any combination thereof (e.g., x and y, but not z). In certain cases, the context in which the term "or" is used may indicate that it is used in an exclusive sense, for example, where "select one of x, y, or z" means that only one of x, y, and z is selected in this example.
[0038] In the following description, numerous specific details are set forth to provide a thorough understanding of the disclosed embodiments. However, one of ordinary skill in the art will recognize that various aspects of the disclosed embodiments can be practiced without these specific details. In some instances, well-known structures, computer program instructions, and techniques are not shown in detail to avoid obscuring the disclosed embodiments. DETAILED DESCRIPTION
[0039] U.S. Patent Application No. 13 / 969,372, filed August 16, 2013 (now U.S. Patent No. 8,812,144), the entire contents of which are incorporated herein by reference, discusses techniques for generating musical content based on one or more musical attributes. To the extent any interpretation based on a perceived conflict between the definitions of the aforementioned application 372 and the remainder of this disclosure, the present disclosure is intended to govern. The musical attributes may be input by a user or may be determined based on environmental information such as ambient noise, lighting, etc. The aforementioned disclosure 372 discusses techniques for selecting stored loops and / or tracks or generating new loops / tracks and layering the selected loops / tracks to generate output musical content.
[0040] U.S. Patent Application No. 16 / 420,456, filed May 23, 2019 (now U.S. Patent No. 10,679,596), the entire contents of which are incorporated herein by reference, discusses techniques for generating music content. With respect to any interpretation based on a perceived conflict between the definitions of the aforementioned application '456 and the remainder of this disclosure, this disclosure is intended to govern. Music can be generated based on user input or using a computer-implemented method. The aforementioned disclosure '456 discusses various music generator embodiments.
[0041] The present disclosure generally relates to a system for generating customized music content by selecting and combining tracks based on various parameters. In various embodiments, a machine learning algorithm (including a neural network such as a deep learning neural network) is configured to generate and customize music content for a specific user. In some embodiments, a user can create their own control elements, and a computing system can be trained to generate output music content based on the user's expected functions of the user-defined control elements. In some embodiments, playback data of the music content generated by the technology described herein can be recorded to record and track the use of various music content by different rights holders (e.g., copyright holders). The various technologies discussed below can provide more relevant customized music for different contexts, promote the generation of music based on specific sounds, allow users to better control how music is generated, generate music that achieves one or more specific goals, generate music in real time along with other content, etc.
[0042] As used herein, the term "audio file" refers to sound information for musical content. For example, the sound information may include data describing the musical content as raw audio in formats such as wav, aiff, or FLAC. Attributes of the musical content may be included in the sound information. Attributes may include, for example, quantifiable musical attributes such as instrument classification, pitch transcription, beat timing, tempo, file length, and audio amplitude in multiple frequency segments. In some embodiments, an audio file includes sound information within a specific time interval. In various embodiments, an audio file includes a loop. As used herein, the term "loop" refers to the sound information of a single instrument within a specific time interval. The various techniques discussed with reference to audio files may also be performed using loops including a single instrument. An audio file or loop may be played in a repeated manner (for example, a 30-second audio file may be played four times in succession to generate 2 minutes of musical content), but an audio file may also be played once, for example without repeating.
[0043] In some embodiments, an image representation of an audio file is generated and used to generate music content. The image representation of the audio file can be generated based on the data in the audio file and the MIDI representation of the audio file. For example, the image representation can be a two-dimensional (2D) image representation of the pitch and tempo determined from the MIDI representation of the audio file. Rules (e.g., composition rules) can be applied to the image representation to select the audio file to be used to generate new music content. In various embodiments, machine learning / neural networks are implemented on the image representation to select audio files to be combined to generate new music content. In some embodiments, the image representation is a compressed (e.g., lower resolution) version of the audio file. The compressed image representation can increase the speed of searching for selected music content in the image representation.
[0044] In some embodiments, a music generator can generate new music content based on various parameter representations of an audio file. For example, an audio file typically has an audio signal that can be represented as a graph of a signal (e.g., signal amplitude, frequency, or a combination thereof) relative to time. However, the time-based representation depends on the tempo of the music content. In various embodiments, an audio file is also represented using a graph of the signal relative to the tempo (e.g., a signal graph). The signal graph is independent of tempo, allowing the audio parameters of the music content to be modified without changing the tempo.
[0045] In some embodiments, the music generator allows users to create and label user-defined controls. For example, a user can create a control, which the music generator can then train to influence the music based on the user's preferences. In various embodiments, user-defined controls are advanced controls, such as those that adjust mood, intensity, or genre. These controls are typically subjective measurements based on the listener's personal preferences. In some embodiments, the user creates and labels controls for user-defined parameters. The music generator can then play various music files and allow the user to modify the music based on the user-defined parameters. The music generator can learn and store the user's preferences based on the user's adjustments to the user-defined parameters. Therefore, during later playback, the user can adjust the user-defined controls for the user-defined parameters, and the music generator can adjust the music playback based on the user's preferences. In some embodiments, the music generator can also select music content based on the user's preferences set by the user-defined parameters.
[0046] In some embodiments, the music content generated by the music generator includes music with various stakeholder entities (e.g., rights holders or copyright holders). In commercial applications where the generated music content is played back continuously, compensation based on playback of a single track (file) may be difficult. Therefore, in various embodiments, technology is implemented for recording playback data of continuous music content. The recorded playback data may include information related to the playback time of individual tracks within the continuous music content that is matched to the stakeholders of each individual track. In addition, technology can be implemented to prevent tampering with playback data information. For example, playback data information can be stored in a publicly accessible immutable blockchain ledger.
[0047] This disclosure initially refers to Figure 1 and Figure 2 Describes the example music generator module and the overall system organization with various applications. Figures 3 to 7 Techniques for generating musical content from image representations are discussed. Figure 8 and Figure 10 Discusses techniques for implementing user-created control elements. Figures 11 to 17 Discusses techniques for generating and implementing audio technology. Figures 18 and 19 Techniques for recording information about generated music or elements in a blockchain or other encrypted ledger are discussed. FIG. 20A to FIG. 20B An exemplary application interface is shown.
[0048] In general, the disclosed music generator includes audio files, metadata (e.g., information describing the audio files), and a syntax for combining audio files based on the metadata. The generator can create music experiences using rules to identify audio files based on the metadata and target characteristics of the music experience. The generator can be configured to expand the set of experiences it can create by adding or modifying rules, audio files, and / or metadata. Adjustments can be made manually (e.g., an artist adds new metadata), or the music generator can add rules / audio files / metadata as it monitors the music experience within a given environment and the desired goals / characteristics. For example, listener-defined controls can be implemented to obtain user feedback on music goals or characteristics.
[0049] Overview of an Exemplary Music Generator
[0050] Figure 1 is a schematic diagram illustrating an exemplary music generator according to some embodiments. In the illustrated embodiment, the music generator module 160 receives various information from a number of different sources and generates output music content 140.
[0051] In the illustrated embodiment, module 160 accesses stored audio file(s) and corresponding attributes 110 for the stored audio file(s), and combines the audio files to generate output music content 140. In some embodiments, music generator module 160 selects audio files based on their attributes and combines the audio files based on target music attributes 130. In some embodiments, audio files can be selected based on environmental information 150 in conjunction with target music attributes 130. In some embodiments, environmental information 150 is indirectly used to determine target music attributes 130. In some embodiments, target music attributes 130 are explicitly specified by the user, for example, by specifying a desired energy level, mood, multiple parameters, etc. For example, the listener-defined controls described herein can be implemented to specify listener preferences as target music attributes. Examples of target music attributes 130 include energy, complexity, and diversity, but more specific attributes (e.g., attributes corresponding to stored tracks) can also be specified. Generally speaking, when higher-level target music attributes are specified, the system can determine lower-level specific music attributes before generating output music content.
[0052] Complexity can refer to the number of audio files, loops, and / or instruments included in a work. Energy may or may not be related to other attributes. For example, changing the mode or tempo may affect energy. However, for a given tempo and mode, energy can be changed by adjusting the instrument type (e.g., by adding high hats or white noise), complexity, volume, etc. Diversity can refer to the amount of change in the generated music over time. Diversity can be generated for a set of static other musical attributes (e.g., by selecting different tracks for a given tempo and mode), or can be generated by changing the musical attributes over time (e.g., by changing the tempo and mode more frequently when greater diversity is desired). In some embodiments, the target musical attributes can be considered to exist in a multidimensional space, and the music generator module 160 can slowly move through the space, for example, performing route corrections based on environmental changes and / or user input if necessary.
[0053] In some embodiments, the attributes stored with the audio files include information about one or more audio files, including: tempo, volume, energy, genre, spectrum, envelope, modulation, periodicity, rise and decay times, noise, artist, instrument, theme, etc. Note that in some embodiments, the audio files are partitioned such that a group of one or more audio files is specific to a particular audio file type (e.g., an instrument or an instrument type).
[0054] In the illustrated embodiment, module 160 accesses stored rule set(s) 120. In some embodiments, the stored rule set(s) 120 specify rules for how many audio files to cover so that they play simultaneously (which may correspond to the complexity of the output music), which major / minor progression to use when transitioning between audio files or phrases, which instruments to use together (e.g., instruments that have an affinity for each other), etc., to achieve the target music attributes. In other words, the music generator module 160 uses the stored rule set(s) 120 to achieve one or more declarative goals defined by the target music attributes (and / or target environment information). In some embodiments, the music generator module 160 includes one or more pseudo-random number generators that are configured to introduce pseudo-randomness to avoid repetitive output music.
[0055] In some embodiments, the environmental information 150 includes one or more of the following: lighting information, ambient noise, user information (facial expressions, body posture, activity level, movement, skin temperature, performance of specific activities, clothing type, etc.), temperature information, purchasing activity in the area, time of day, day of the week, time of year, number of people present, weather conditions, etc. In some embodiments, the music generator module 160 does not receive / process environmental information. In some embodiments, the environmental information 150 is received by another module that determines target music attributes 130 based on the environmental information. The target music attributes 130 can also be derived based on other types of content (e.g., video data). In some embodiments, the environmental information is used to adjust one or more stored rule sets 120, for example, to achieve one or more environmental goals. Similarly, the music generator can use the environmental information to adjust the stored properties of one or more audio files, for example, to indicate target music attributes or target audience characteristics that are particularly relevant to those audio files.
[0056] As used herein, the term "module" refers to a physical, non-transitory computer-readable medium that is configured to perform a specified operation or stores information (e.g., program instructions) that instructs other circuits (e.g., processors) to perform a specified operation. A module can be implemented in a variety of ways, including as a hard-wired circuit or as a memory in which program instructions are stored that can be executed by one or more processors to perform the operation. Hardware circuits can include, for example, customized very large scale integration (VLSI) circuits or gate arrays, off-the-shelf semiconductors such as logic chips, transistors, or other discrete components. A module can also be implemented in a programmable hardware device, such as a field programmable gate array, programmable array logic, a programmable logic device, or the like. A module can also be a non-transitory computer-readable medium in any suitable form that stores program instructions that can be executed to perform a specified operation.
[0057] As used herein, the phrase "music content" refers to both the music itself (the audible representation of the music) and the information that can be used to play the music. Thus, a song recorded as a file on a storage medium (such as, but not limited to, a compact disc, a flash drive, etc.) is an example of music content; the sound produced by outputting such a recorded file or other electronic representation (e.g., through a speaker) is also an example of music content.
[0058] The term "music" includes its commonly understood meaning, including sounds produced by musical instruments as well as the human voice. Thus, music includes, for example, instrumental performances or recordings, a cappella performances or recordings, and performances or recordings that include both musical instruments and voices. One of ordinary skill in the art will recognize that "music" does not include all sound recordings. Works that do not contain musical attributes such as rhythm or meter (e.g., speeches, news broadcasts, and audiobooks) are not music.
[0059] A piece of musical "content" can be distinguished from another piece of musical content in any suitable manner. For example, a digital file corresponding to a first song can represent the first piece of musical content, while a digital file corresponding to a second song can represent the second piece of musical content. The phrase "musical content" can also be used to distinguish specific intervals within a given musical work, so that different parts of the same song can be considered different pieces of musical content. Similarly, different tracks within a given musical work (e.g., a piano track, a guitar track) can also correspond to different pieces of musical content. In the context of a potentially endless stream of generated music, the phrase "musical content" can be used to refer to a specific portion of the stream (e.g., a few bars or a few minutes).
[0060] The music content generated by embodiments of the present disclosure may be "new music content"—a combination of music elements that has never been generated before. A related (but broader) concept—"original music content"—will be further described below. To facilitate explanation of this term, the concept of a "controlling entity" is described in connection with instances of music content generation. Unlike the phrase "original music content," the phrase "new music content" does not involve the concept of a controlling entity. Therefore, new music content refers to music content that has never been generated before by any entity or computer system.
[0061] Conceptually, the present disclosure refers to some "entities" as controlling specific instances of computer-generated music content. Such entities have any legal rights (e.g., copyright) that may correspond to the computer-generated content (to the extent that any such rights may actually exist). In one embodiment, the individual that creates (e.g., encodes various software routines) a computer-implemented music generator or operates (e.g., provides input to) a specific instance of computer-implemented music generation will be the controlling entity. In other embodiments, the computer-implemented music generator can be created by a legal entity (e.g., a company or other commercial organization), such as in the form of a software product, a computer system, or a computing device. In some instances, such a computer-implemented music generator can be deployed to many clients. In various instances, the controlling entity can be a creator, a distributor, or a client, according to the terms of the license associated with the distribution of the music generator. If there is no such clear legal agreement, the controlling entity of the computer-implemented music generator is the entity that promotes (e.g., provides input and thereby operates) the computer-generated specific instance of music content.
[0062] Within the meaning of the present disclosure, computer generation of "original musical content" by a controlling entity refers to 1) a combination of musical elements that has never been generated before by the controlling entity or anyone else, and 2) a combination of musical elements that have been generated before but were originally generated by the controlling entity. Content type 1) is referred to herein as "novel musical content", which is similar to the definition of "new musical content", except that the definition of "novel musical content" refers to the concept of a "controlling entity" while the definition of "new musical content" does not. On the other hand, content type 2) is referred to herein as "proprietary musical content". Note that the term "proprietary" in this context does not refer to any implied legal rights in the content (although such rights may exist), but is merely used to indicate that the musical content was originally generated by the controlling entity. Therefore, the "regeneration" of musical content previously and originally generated by the controlling entity by a controlling entity constitutes "generation of original musical content" within the meaning of the present disclosure. "Non-original musical content" for a particular controlling entity is musical content that is not "original musical content" of that controlling entity.
[0063] Some music content segments can include music components from one or more other music content segments. Creating music content in this way is called "sampled" music content, and is common in specific music works, particularly in specific music genres. Such music content is referred to as "music content with sampled components," "derivative music content," or other similar terms in this article. In contrast, music content that does not include sampled components is referred to as "music content without sampled components," "non-derivative music content," or other similar terms in this article.
[0064] In applying these terms, it should be noted that if any particular musical content is reduced to a sufficient level of granularity, then it can be argued that the musical content is derivative (effectively meaning that all musical content is derivative). In this disclosure, the terms "derivative" and "non-derivative" are not used in this sense. With respect to computer-generation of musical content, if the computer-generation selects portions of components from pre-existing musical content from an entity other than the controlling entity (e.g., a computer program selects specific portions of an audio file of a popular artist's work to include in a piece of musical content being generated), then such computer-generation is considered derivative (and produces derivative musical content). On the other hand, if the computer-generation does not utilize such components of such pre-existing content, then the computer-generation of musical content is considered non-derivative (and produces non-derivative musical content). Note that some segments of "original musical content" may be derivative musical content, while some segments may be non-derivative musical content.
[0065] Please note that the term "derivative" in this disclosure is intended to have a broader meaning than the term "derivative work" as used in U.S. copyright law. For example, derivative musical content may or may not be a derivative work under U.S. copyright law. The term "derivative" in this disclosure is not intended to convey a negative connotation; it is simply used to indicate whether a particular piece of musical content has "borrowed" portions of another work.
[0066] Furthermore, the phrases "new musical content," "novel musical content," and "original musical content" are not intended to encompass musical content that differs only slightly from a pre-existing combination of musical elements. For example, simply changing a few notes of a pre-existing musical composition does not produce new, novel, or original musical content, as those phrases are used in this disclosure. Similarly, simply changing the key or tempo of a pre-existing musical composition or adjusting the relative intensity of the frequencies (e.g., using an equalizer interface) does not produce new, novel, or original musical content. Furthermore, the phrases "new, novel, and original musical content" are not intended to encompass musical content that falls between original and non-original content; rather, these terms are intended to encompass musical content that is undoubtedly and provably original, including musical content that is eligible for copyright protection by a controlling entity (referred to herein as "protected" musical content). Furthermore, as used herein, the term "available" musical content refers to musical content that does not infringe the copyright of any entity other than the controlling entity. New and / or original musical content is often protected and available. This may be advantageous in preventing the copying of musical content and / or the payment of royalties for the musical content.
[0067] Although various embodiments discussed herein use rule-based engines, various other types of computer-implemented algorithms may be used with any of the computer learning and / or music generation techniques discussed herein. However, rule-based approaches may be particularly effective in a music context.
[0068] Overview of Applications, Storage Elements, and Data That Can Be Used in an Exemplary Music System
[0069] The music maker module can interact with a plurality of different applications, modules, storage elements etc. to generate music content. For example, a terminal user can install one of a plurality of types of applications for different types of computing devices (for example, mobile device, desktop computer, DJ equipment etc.). Similarly, another type of application can be provided to business users. When generating music content, interacting with the application can allow the music maker to receive external information, which can be used to determine the target music attribute and / or update one or more rule sets for generating music content. Except interacting with one or more applications, the music maker module can also interact with other modules to receive rule set, update rule set etc. Finally, the music maker module can access one or more rule sets, audio files and / or the music content of the generation stored in one or more storage elements. In addition, the music maker module can store any of the items listed above in one or more storage elements, which can be local or (for example, cloud-based) via network access.
[0070] Figure 2 2 is a block diagram illustrating an exemplary overview of a system for generating output music content based on input from multiple different sources, according to some embodiments. In the illustrated embodiment, system 200 includes a rules module 210, a user application 220, a web application 230, an enterprise application 240, an artist application 250, an artist rule generator module 260, a store of generated music 270, and external input 280.
[0071] In the illustrated embodiment, the user application 220, the web application 230, and the enterprise application 240 receive external input 280. In some embodiments, the external input 280 includes: environmental input, target music attributes, user input, sensor input, etc. In some embodiments, the user application 220 is installed on the user's mobile device and includes a graphical user interface (GUI) that allows the user to interact / communicate with the rule module 210. In some embodiments, the web application 230 is not installed on the user device, but is configured to run within the browser of the user device and can be accessed through a website. In some embodiments, the enterprise application 240 is an application used by a larger entity to interact with the music generator. In some embodiments, the application 240 is used in conjunction with the user application 220 and / or the web application 230. In some embodiments, the application 240 communicates with one or more external hardware devices and / or sensors to collect information about the surrounding environment.
[0072] In the illustrated embodiment, the rule module 210 communicates with the user application 220, the web application 230, and the enterprise application 240 to generate output music content. In some embodiments, the music generator 160 is included in the rule module 210. Please note that the rule module 210 can be included in one of the applications 220, 230, and 240, or can be installed on a server and accessed via a network. In some embodiments, the applications 220, 230, and 240 receive the generated output music content from the rule module 210 and cause the content to be played. In some embodiments, the rule module 210 requests input from the applications 220, 230, and 240, such as about target music attributes and environmental information, and can use this data to generate music content.
[0073] In the illustrated embodiment, the stored rule set(s) 120 are accessed by the rules module 210. In some embodiments, the rules module 210 modifies and / or updates the stored rule set 120 based on communications with the applications 220, 230, and 240. In some embodiments, the rules module 210 accesses the stored rule set(s) 120 to generate output music content. In the illustrated embodiment, the stored rule set(s) 120 may include rules from the artist rule generator module 260, discussed in further detail below.
[0074] In the illustrated embodiment, the artist application 250 communicates with an artist rule generator module 260 (e.g., which can be part of the same application or can be cloud-based). In some embodiments, the artist application 250 allows artists to create rule sets for their specific sounds, for example, based on previous works. U.S. Patent No. 10,679,596 further discusses this functionality. In some embodiments, the artist rule generator module 260 is configured to store the generated artist rule sets for use by the rule module 210. Users can purchase rule sets from a specific artist before using that artist to generate output music through their specific application. Rule sets for specific artists can be referred to as signature packs.
[0075] In the embodiment shown, the stored audio file(s) and corresponding attribute(s) 110 are accessed by the module 210 when applying the rules to select and combine tracks to generate output music content. In the embodiment shown, the rules module 210 stores the generated output music content 270 in a storage element.
[0076] In some embodiments, the system is implemented on a server and accessed via a network. Figure 2260, which can be referred to as a cloud-based implementation. For example, stored rule set(s) 120, audio file(s) / attribute(s) 110, and generated music 270 can all be stored on the cloud and accessed by module 210. In another example, module 210 and / or module 260 can also be implemented in the cloud. In some embodiments, generated music 270 is stored in the cloud and digitally watermarked. This can allow, for example, detection of duplicate generated music and generation of large amounts of customized music content.
[0077] In some embodiments, one or more of the disclosed modules are configured to generate other types of content in addition to music content. For example, the system can be configured to generate output visual content based on target music attributes, determined environmental conditions, currently used rule sets, etc. As another example, the system can search a database or the internet based on the current attributes of the music being generated and display an image collage that dynamically changes as the music changes and matches the attributes of the music.
[0078] Exemplary Machine Learning Methods
[0079] As stated in this article, Figure 1 The illustrated music generator module 160 can implement a variety of artificial intelligence (AI) techniques (e.g., machine learning techniques) to generate the output music content 140. In various embodiments, the implemented artificial intelligence techniques include a combination of deep neural networks (DNNs) with more traditional machine learning techniques and knowledge-based systems. This combination can combine the respective strengths and weaknesses of these techniques with the challenges inherent in musical compositions and personalization systems. Musical content has multiple levels of structure. For example, a song has sections, phrases, melodies, notes, and textures. DNNs can effectively analyze and generate music content at very high and very low levels of detail. For example, a DNN can be good at classifying the texture of a sound as belonging to a low-level clarinet or electric guitar, or detecting high-level verses and choruses. Intermediate levels of music content details, such as melody construction, orchestration, etc., may be more difficult. DNNs are generally good at collecting a wide range of styles in a single model, and therefore, DNNs can be implemented as generation tools with a large range of expressions.
[0080] In some embodiments, the music generator module 160 utilizes expert knowledge by using human-written audio files (e.g., loops) as the basic units of musical content used by the music generator module. For example, the social context of expert knowledge can be embedded through the selection of rhythm, melody, and texture to record heuristics in multi-level structures. Unlike the separation of DNNs and traditional machine learning based on structural levels, expert knowledge can be applied to any area where it can increase musicality without placing excessive restrictions on the trainability of the music generator module 160.
[0081] In some embodiments, the music generator module 160 uses a DNN to find patterns in how audio layers are combined, vertically by layering sounds on top of each other, and horizontally by combining audio files or loops into sequences. For example, the music generator module 160 can implement an LSTM (long-short-term memory) recurrent neural network that is trained on MFCC (Mel-frequency cepstral coefficients) audio features of loops used in multi-track audio recordings. In some embodiments, the network is trained to predict and select audio features of loops of upcoming beats based on knowledge of the audio features of previous beats. For example, the network can be trained to predict audio features of loops of the next 8 beats based on knowledge of the audio features of the last 128 beats. Thus, the network is trained to predict upcoming beats using low-dimensional feature representations.
[0082] In certain embodiments, the music generator module 160 uses known machine learning algorithms to assemble a sequence of multi-track audio into a musical structure with dynamic characteristics of intensity and complexity. For example, the music generator module 160 can implement a hierarchical hidden Markov model, which can perform state transitions like a state machine, with probabilities determined by a multi-level hierarchical structure. For example, a particular type of drop may be more likely to occur after a rising chapter, but less likely to occur if the rising chapter ends without drums. In various embodiments, probabilities can be trained transparently, in contrast to DNN training where the content being learned is more opaque.
[0083] Markov models can handle larger temporal structures and therefore may not be easily trained by presenting example tracks, as the examples may be too long. Feedback control elements (e.g., thumbs up / down on a user interface) can be used to provide feedback on the music at any time. In certain embodiments, the feedback control element is implemented as one of the (one or more) UI control elements 830, such as Figure 8As shown. The correlation between the musical structure and the feedback can then be used to update the structural model used for the composition, such as a transition table or a Markov model. This feedback can also be collected directly from measurements of heart rate, sales, or any other metric for which the system can determine a clear classification. As mentioned above, the expert knowledge heuristic is also designed to be as probabilistic as possible and is trained in the same manner as the Markov model.
[0084] In certain embodiments, training can be performed by a composer or DJ. Such training can be separate from audience training. For example, training performed by an audience (e.g., a typical user) may be limited to identifying correct or incorrect classifications based on positive and negative model feedback, respectively. For composers and DJs, training may include hundreds of time steps and include detailed information about the layers and volume controls used to provide more explicit details to understand the factors that drive changes in musical content. For example, training performed by composers and DJs can include sequence prediction training similar to the global training of the DNN described above.
[0085] In various embodiments, a DNN is trained using multiple tracks of audio and interface interactions to predict what a DJ or composer will do next. In some embodiments, these interactions can be recorded and used to develop new, more transparent heuristics. In some embodiments, the DNN receives as input a number of prior music metrics and utilizes the low-dimensional feature representation described above, along with additional features describing the modifications the DJ or composer has applied to the track. For example, the DNN may receive as input the last 32 bars of music and utilize the low-dimensional feature representation along with additional features describing the modifications the DJ or composer has applied to the track. These modifications may include adjusting the gain, applied filters, delay, and more for a particular track. For example, a DJ may use the same drum loop for five minutes during a performance, but may gradually increase the gain and delay over time. Therefore, in addition to loop selection, the DNN can be trained to predict these gain and delay changes. When no loop is playing for a particular instrument (e.g., no drum loop), the feature set for that instrument may be all zeros, which can allow the DNN to learn that predicting all zeros is likely a successful strategy, potentially leading to selective stratification.
[0086] In some instances, DJs or composers record live performances using a mixer and equipment such as TRAKTOR (Local Instruments GmbH). These recordings are typically captured in high resolution (e.g., 4-track recordings or MIDI). In some embodiments, the system decomposes the recordings into their constituent loops, generating information about the combination of loops in the composition and the sound quality of each individual loop. Training a DNN (or other machine learning) with this information provides the DNN with the ability to correlate the composition (e.g., sequencing, layering, timing of loops, etc.) and the sound quality of the loops to inform the music generator module 160 how to create a musical experience similar to the artist's performance without using the actual loops used by the artist in the performance.
[0087] Exemplary music generator using image representations of audio files
[0088] Popular music often features widely observed combinations of rhythms, textures, and pitches. When creating music note by note for each instrument in a piece (as can be done by a music generator), rules can be implemented based on these combinations to create coherent music. Generally speaking, the stricter the rules, the less room there is for creative variation, making it more likely that copies of existing music will be created.
[0089] When creating music by combining musical phrases that have already been performed and recorded as audio, it may be necessary to consider multiple, unchangeable combinations of notes within each phrase to create the composition. However, when drawing from a library of thousands of recordings, searching every possible combination can be computationally expensive. Furthermore, note-by-note comparisons may be necessary to check for harmonically dissonant combinations, especially across beats. New rhythms created by combining multiple files can also be checked against the rhythmic composition rules of the combined phrases.
[0090] It may not always be possible to extract the necessary features from the audio files for combination. Even when possible, extracting the required features from the audio files may be computationally expensive. In various embodiments, symbolic audio representations are used in music composition to reduce computational overhead. The symbolic audio representations may rely on the music composer's memory of the textures of the instruments and stored rhythm and pitch information. A common symbolic music representation format is MIDI. MIDI contains precise timing, pitch, and performance control information. In some embodiments, MIDI can be further simplified and compressed by a piano roll representation, in which notes are displayed as bars on a discrete time / pitch graph, typically with 8 octaves.
[0091] In some embodiments, the music generator is configured to generate output music content by generating an image representation of the audio file and selecting a combination of music based on an analysis of the image representation. The image representation can be a representation further compressed from a piano roll representation. For example, the image representation can be a low-resolution representation generated based on a MIDI representation of the audio file. In various embodiments, composition rules are applied to the image representation to select music content from the audio file to combine and generate the output music content. For example, a rule-based approach can be used to apply the composition rules. In some embodiments, a machine learning algorithm or model (e.g., a deep learning neural network) is implemented to select and combine audio files for generating the output music content.
[0092] Figure 3 is a block diagram illustrating an exemplary music generator system configured to output music content based on analysis of image representations of audio files according to some embodiments. In the illustrated embodiment, system 300 includes an image representation generation module 310, a music selection module 320, and music generator module 160.
[0093] In the illustrated embodiment, the image representation generation module 310 is configured to generate one or more image representations of audio files. In certain embodiments, the image representation generation module 310 receives audio file data 312 and MIDI representation data 314. The MIDI representation data 314 includes MIDI representation(s) of a specified audio file(s) in the audio file data 312. For example, for a specified audio file in the audio file data 312, there may be a corresponding MIDI representation in the MIDI representation data 314. In some embodiments where the audio file data 312 contains multiple audio files, each audio file in the audio file data 312 may have a corresponding MIDI representation in the MIDI representation data 314. In the illustrated embodiment, the MIDI representation data 314 is provided to the image representation generation module 310 along with the audio file data 312. However, in some contemplated embodiments, the image representation generation module 310 may itself generate the MIDI representation data 314 from the audio file data 312.
[0094] like Figure 3As shown, the image representation generation module 310 generates (one or more) image representations 316 from audio file data 312 and MIDI representation data 314. The MIDI representation data 314 may include the pitch, time (or tempo), and velocity (or note intensity) data of the notes in the music associated with the audio file, while the audio file data 312 includes data for playing back the music itself. In certain embodiments, the image representation generation module 310 generates an image representation for the audio file based on the pitch, time, and velocity data from the MIDI representation data 314. The image representation can be, for example, a two-dimensional (2D) image representation of the audio file. In the 2D image representation of the audio file, the x-axis represents time (tempo) and the y-axis represents pitch (similar to a piano roll representation), where the pixel value at each xy coordinate represents velocity.
[0095] The 2D image representation of the audio file can have a variety of image sizes, although the image size is typically selected to correspond to the musical structure. For example, in one contemplated embodiment, the 2D image representation is a 32 (x-axis) x 24 image (y-axis). A 32-pixel wide image representation allows each pixel to represent a quarter beat in the time dimension. Thus, 8 beats of music can be represented using a 32-pixel wide image representation. While this representation may not have enough detail to capture the expressive details of the music in the audio file, the expressive details are retained in the audio file itself, which is combined with the image representation of the system 300 to generate the output musical content. However, the quarter beat time resolution does allow for a large coverage of common pitch and rhythm combination rules.
[0096] Figure 4 An example of an image representation 316 of an audio file is depicted. The image representation 316 is 32 pixels wide (time) and 24 pixels high (pitch). Each pixel (square) 402 has a value representing the tempo and pitch of the audio file. In various embodiments, the image representation 316 can be a grayscale image representation of the audio file, where pixel values are represented by varying grayscale intensities. The grayscale variations based on pixel values may be subtle and imperceptible to many people. Figure 5A and Figure 5B Examples of grayscale images for melody image feature representation and drum beat image feature representation are depicted, respectively. However, other representations (e.g., color or digital) can also be considered. In these representations, each pixel can have multiple different values corresponding to different musical attributes.
[0097] In a particular embodiment, the image representation 316 is an 8-bit representation of the audio file. Thus, each pixel can have 256 possible values. MIDI representations typically have 128 possible velocity values. In various embodiments, the details of the velocity values may not be as important as the task of selecting the audio file to combine. Therefore, in such embodiments, the pitch axis (y-axis) can be banded into two groups of octaves over an 8-octave range, with each group having 4 octaves. For example, 8 octaves can be defined as follows:
[0098] Octave 0: rows 0-11, values 0-63;
[0099] Octave 1: rows 12-23, values 0-63;
[0100] Octave 2: rows 0-11, values 64-127;
[0101] Octave 3: lines 12-23, values 64-127;
[0102] Octave 4: rows 0-11, values 128-191;
[0103] Octave 5: Lines 12-23, Values 128-191
[0104] Octave 6: rows 0-11, values 192-255; and
[0105] Octave 7: Rows 12-23, values 192-255.
[0106] With these defined octave ranges, the rows and values of the pixels determine the octave and velocity of the note. For example, a pixel value of 10 in row 1 represents a note in octave 0 with a velocity of 10, while a pixel value of 74 in row 1 represents a note in octave 2 with a velocity of 10. As another example, a pixel value of 79 in row 13 represents a note in octave 3 with a velocity of 15, while a pixel value of 207 in row 13 represents a note in octave 7 with a velocity of 15. Thus, using the defined ranges of octaves above, the first 12 rows (rows 0-11) represent the first set of 4 octaves (octaves 0, 2, 4, and 6), with the pixel value determining which of the first 4 octaves is represented (the pixel value also determines the velocity of the note). Similarly, the second 12 rows (rows 12-23) represent a second set of 4 octaves (octaves 1, 3, 5, and 7), with the pixel value determining which of the second 4 octaves is represented (the pixel value also determines the velocity of the note).
[0107] As described above, by banding the pitch axis to cover an 8-octave range, the velocity of each octave can be defined by 64 values rather than the 128 values of the MIDI representation. Thus, a 2D image representation (e.g., image representation 316) can be compressed (e.g., have a lower resolution) than a MIDI representation of the same audio file. In some embodiments, further compression of the image representation may be allowed because 64 values may be more than required for the system 300 to select a musical combination. For example, the velocity resolution can be further reduced to allow compression in the time representation by having odd pixel values represent note onsets and even pixel values represent note sustains. Reducing the resolution in this way allows two notes of the same velocity played in rapid succession to be distinguished from one longer note based on odd or even pixel values.
[0108] As mentioned above, the compactness of the image representation reduces the file size required to represent the music (e.g., compared to a MIDI representation). Thus, implementing an image representation of an audio file reduces the amount of disk storage required. Furthermore, the compressed image representation can be stored in high-speed memory, allowing for rapid searches of possible music combinations. For example, an 8-bit image representation can be stored in graphics memory on a computer device, allowing for large-scale parallel searches.
[0109] In various embodiments, image representations generated for multiple audio files are combined into a single image representation. For example, image representations for tens, hundreds, or thousands of audio files can be combined into a single image representation. The single image representation can be a large, searchable image that can be used to search the multiple audio files that make up the single image in parallel. For example, software such as MegaTextures (from id Software) can be used to search the single image in a manner similar to large textures in video games.
[0110] Figure 6 is a block diagram illustrating an exemplary system configured to generate a single image representation according to some embodiments. In the illustrated embodiment, the system 600 includes a single image representation generation module 610 and a texture feature extraction module 620. In certain embodiments, the single image representation generation module 610 and the texture feature extraction module 620 are located in the image representation generation module 310, as shown in FIG. Figure 3 However, a single image representation generation module 610 or a texture feature extraction module 620 may be located outside the image representation generation module 310 .
[0111] like Figure 6As shown in the illustrated embodiment, multiple image representations 316A-N are generated. The image representations 316A-N can be N separate image representations of N separate audio files. The single image representation generation module 610 can combine the separate image representations 316A-N into a single combined image representation 316. In some embodiments, the separate image representations combined by the single image representation generation module 610 include separate image representations for different musical instruments. For example, different instruments in an orchestra can be represented by separate image representations, which are then combined into a single image representation for use in searching and selecting music.
[0112] In certain embodiments, the individual image representations 316A-N are combined into a single image representation 316, where the individual image representations are placed adjacent to each other without overlapping. Thus, the single image representation 316 is a complete dataset representation of all the individual image representations 316A-N, with no data loss (e.g., no data from one image representation modifies data for another image representation). Figure 7 Depicted is an example of a single image representation 316 of multiple audio files. In the illustrated embodiment, the single image representation 316 is a combined image generated from individual image representations 316A, 316B, 316C, and 316D.
[0113] In some embodiments, a single image representation 316 is appended with textural features 622. In the embodiment shown, the textural features 622 are appended to the single image representation 316 as a single row. Figure 6 The texture features 622 are determined by the texture feature extraction module 620. The texture features 622 may include, for example, the texture of the instrument of the music in the audio file. For example, the texture features may include features from different instruments (e.g., drums, string instruments, etc.).
[0114] In certain embodiments, the texture feature extraction module 620 extracts texture features 622 from the audio file data 312. The texture feature extraction module 620 can implement, for example, a rule-based approach, a machine learning algorithm or model, a neural network, or other feature extraction techniques to determine texture features from the audio file data 312. In some embodiments, the texture feature extraction module 620 can extract texture features 622 from the image representation(s) 316 (e.g., multiple image representations or a single image representation). For example, the texture feature extraction module 620 can implement image-based analysis (e.g., an image-based machine learning algorithm or model) to extract texture features 622 from the image representation(s) 316.
[0115] Adding texture features 622 to the single image representation 316 provides the single image representation with additional information that is not typically available in a MIDI representation or piano roll representation of an audio file. In some embodiments, the single image representation 316 ( Figure 7 316) may not need to be human-readable. For example, the texture features 622 may only need to be machine-readable for implementation in a music generation system. In certain embodiments, the texture features 622 are attached to the single image representation 316 for image-based analysis of the single image representation. For example, the texture features 622 can be used by an image-based machine learning algorithm or model used in music selection, as described below. In some embodiments, the texture features 622 can be ignored during music selection, such as in rule-based selection, as described below.
[0116] Back to Figure 3 In the illustrated embodiment, the image representation(s) 316 (e.g., multiple image representations or a single image representation) are provided to the music selection module 320. The music selection module 320 can select audio files or portions of audio files to combine in the music generator module 160. In certain embodiments, the music selection module 320 applies a rule-based approach to search for and select audio files or portions of audio files for the music generator module 160 to combine. Figure 3 As shown, the music selection module 320 accesses rules of the rule-based method from the stored (one or more) rule sets 120. For example, the rules accessed by the music selection module 320 may include rules for searching and selecting, such as, but not limited to, composition rules and note combination rules. Applying the rules to the (one or more) image representations 316 may be implemented using graphics processing available on a computer device.
[0117] For example, in various embodiments, note combination rules can be expressed as vector and matrix calculations. Graphics processing units are typically optimized for vector and matrix calculations. For example, notes that are one pitch apart are generally dissonant and are often avoided. Notes such as these can be found by searching adjacent pixels in an overlaid layered image (or a fragment of a larger image) based on a rule. Thus, in various embodiments, the disclosed modules can call a kernel to perform all or part of the disclosed operations on a graphics processor of a computing device.
[0118] In some embodiments, the pitch grading in the image representation as described above allows for the use of graphical processing to implant high-pass or low-pass filtering of the audio. Removing (e.g., filtering out) pixel values below a threshold can simulate high-pass filtering, while removing pixel values above a threshold can simulate low-pass filtering. For example, filtering out (removing) pixel values below 64 in the graded example above may have an effect similar to applying a high-pass filter with a shelf at B1 by removing octaves 0 and 1 in the example. Thus, the use of a filter for each audio file can be effectively simulated by applying rules to the image representation of the audio file.
[0119] In various embodiments, when layering audio files together to create music, the pitch of a given audio file can be altered. This alteration in pitch can potentially open up a wider range of possible successful combinations and combinatorial search space. For example, each audio file can be tested in 12 different pitch-shifted modes. Shifting the order of rows in the image representation while parsing the image, and adjusting octave shifts when necessary, can allow for a refined search through these combinations.
[0120] In certain embodiments, the music selection module 320 implements a machine learning algorithm or model on the image representation(s) 316 to search for and select audio files or portions of audio files for combination by the music generator module 160. The machine learning algorithm / model may include, for example, a deep learning neural network or other machine learning algorithm that classifies images based on algorithmic training. In such embodiments, the music selection module 320 includes one or more machine learning models that are trained based on combinations and sequences of audio files that provide desired musical attributes.
[0121] In some embodiments, the music selection module 320 includes a machine learning model that continuously learns during the selection of output music content. For example, the machine learning model can receive user input or other input reflecting attributes of the output music content, which can be used to adjust classification parameters implemented by the machine learning model. Similar to rule-based methods, the machine learning model can be implemented using a graphics processing unit on a computer device.
[0122] In some embodiments, the music selection module 320 implements a combination of a rule-based approach and a machine learning model. In one contemplated embodiment, the machine learning model is trained to find combinations of audio files and image representations for initiating a search for music content to combine, wherein the search is implemented using a rule-based approach. In some embodiments, the music selection module 320 tests the coherence of harmonic and rhythmic rules in the music selected for combination by the music generator module 160. For example, the music selection module 320 can test the harmony and rhythm in the selected audio file 322 before providing the selected audio file to the music generator module 160, as described below.
[0123] exist Figure 3 In the illustrated embodiment, as described above, the music selected by the music selection module 320 is provided to the music generator module 160 as a selected audio file 322. The selected audio file 322 may include a complete or partial audio file that is combined by the music generator module 160 to generate the output music content 140, as described herein. In some embodiments, the music generator module 160 accesses the stored rule set(s) 120 to retrieve rules that are applied to the selected audio file 322 for use in generating the output music content 140. The rules retrieved by the music generator module 160 may be different from the rules applied by the music selection module 320.
[0124] In some embodiments, the selected audio files 322 include information for combining the selected audio files. For example, the machine learning model implemented by the music selection module 320 can provide instructions to the output, which describe how to combine the music content in addition to the selection of the music to be combined. These instructions can then be provided to the music generator module 160 and implemented by the music generator module to combine the selected audio files. In some embodiments, the music generator module 160 tests the consistency of harmony and rhythmic rules before finally determining the output music content 140. Such tests can supplement or replace the tests implemented by the music selection module 320.
[0125] Example controls for music content generation
[0126] In various embodiments, as described herein, the music generator system is configured to automatically generate output music content by selecting and combining tracks based on various parameters. As described herein, a machine learning model (or other AI technology) is used to generate music content. In some embodiments, AI technology is implemented to customize music content for a specific user. For example, the music generator system can implement various types of adaptive controls for personalized music generation. In addition to generating content through AI technology, personalized music generation allows composers or listeners to control content. In some embodiments, users create their own control elements, and the music generator system can train them (e.g., using AI technology) to generate output music content based on the user's expected functions of the user-created control elements. For example, a user can create a control element, and the music generator system trains the control element to affect music according to the user's preferences.
[0127] In various embodiments, user-created control elements are advanced controls, such as controls that adjust mood, intensity, or genre. Such user-created control elements are typically subjective measures based on the listener's personal preferences. In some embodiments, the user labels the user-created control elements to define user-specified parameters. The music generator system can play a variety of music content and allow the user to modify the user-specified parameters in the music content using the control elements. The music generator system can learn and store how the user-defined parameters change the audio parameters in the music content. Therefore, during later playback, the user can adjust the user-created control elements, and the music generator system adjusts the audio parameters in the music playback according to the adjustment level of the user-specified parameters. In some contemplated embodiments, the music generator system can also select music content based on user preferences set by the user-specified parameters.
[0128] Figure 8 8 is a block diagram illustrating an exemplary system configured to implement user-created controls in music content generation according to some embodiments. In the illustrated embodiment, system 800 includes a music generator module 160 and a user interface (UI) module 820. In various embodiments, music generator module 160 implements the techniques described herein for generating output music content 140. For example, music generator module 160 can access stored audio file(s) 810 and generate output music content 140 based on stored rule set(s) 120.
[0129] In various embodiments, the music generator module 160 modifies the music content based on input from one or more UI control elements 830 implemented in the UI module 820. For example, during interaction with the UI module 820, a user can adjust the level of (one or more) control elements 830. Examples of control elements include, but are not limited to, sliders, dials, buttons, or knobs. The level of (one or more) control elements 830 then sets (one or more) control element levels 832, which are provided to the music generator module 160. The music generator module 160 can then modify the output music content 140 based on the (one or more) control element levels 832. For example, the music generator module 160 can implement AI techniques to modify the output music content 140 based on the (one or more) control element levels 830.
[0130] In certain embodiments, one or more control elements 830 are user-defined control elements. For example, a control element may be defined by a composer or an audience. In such embodiments, a user may create and label a UI control element that specifies a parameter that the user wants to implement to control the output music content 140 (e.g., a user creates a control element for controlling a user-specified parameter in controlling the output music content 140).
[0131] In various embodiments, the music generator module 160 can learn or be trained to affect the output music content 140 in a specified manner based on input from a user-created control element. In some embodiments, the music generator module 160 is trained to modify audio parameters in the output music content 140 based on the level of the user-created control element set by the user. Training the music generator module 160 can include, for example, determining the relationship between the audio parameters in the output music content 140 and the level of the user-created control element. The music generator module 160 can then use the relationship between the audio parameters in the output music content 140 and the level of the user-created control element to modify the output music content 140 based on the input level of the user-created control element.
[0132] Figure 9 A flow chart depicts a method for training the music generator module 160 based on user-created control elements, according to some embodiments. Method 900 begins with a user creating and labeling a control element at 910. For example, as described above, a user can create and label a UI control element for controlling a user-specified parameter in the output music content 140 generated by the music generator module 160. In various embodiments, the label of the UI control element describes the user-specified parameter. For example, a user can label a control element as "attitude" to specify that the user wants to control an attitude (defined by the user) in the generated music content.
[0133] After the UI control element is created, the method 900 continues with a playback session 915. The playback session 915 can be used to train the system (e.g., the music generator module 160) on how to modify audio parameters based on the level of the UI control element created by the user. In the playback session 915, an audio track is played at 920. The audio track can be a loop or sample of music from an audio file stored on the device or accessed by the device.
[0134] At 930, the user provides input regarding his / her interpretation of the user-specified parameter in the audio track being played. For example, in certain embodiments, the user is asked to listen to the audio track and select the level of the user-specified parameter that the user believes describes the music in the audio track. For example, the level of the user-specified parameter can be selected using a user-created control element. This process can be repeated for multiple audio tracks in the playback session 915 to generate multiple data points for the level of the user-specified parameter.
[0135] In some contemplated embodiments, a user may be asked to listen to multiple tracks at once and compare and rate the tracks based on user-defined parameters. For example, in the example of a user-created control defining "attitude," the user may listen to multiple tracks and select which tracks have more "attitude" and / or which tracks have less "attitude." Each selection made by the user may be a data point at the level of the user-specified parameter.
[0136] After the playback session 915 is complete, the levels of audio parameters in the audio track from the playback session are evaluated at 940. Examples of audio parameters include, but are not limited to, volume, pitch, bass, treble, reverb, etc. In some embodiments, the levels of audio parameters in the audio track are evaluated while the audio track is playing (e.g., during the playback session 915). In some embodiments, the audio parameters are evaluated after the playback session 915 ends.
[0137] In various embodiments, audio parameters in an audio track are evaluated from the track's metadata. For example, an audio analysis algorithm can be used to generate metadata or symbolic music data (e.g., MIDI) for an audio track (which may be a short, pre-recorded music file). The metadata can include, for example, the pitch of the notes present in the recording, the onset of each beat, the ratio of pitched to non-pitch sounds, the volume, and other quantifiable properties of the sound.
[0138] At 950, a correlation between the user-selected level of the user-specified parameter and the audio parameter is determined. Because the user-selected level of the user-specified parameter corresponds to the level of the control element, the correlation between the user-selected level of the user-specified parameter and the audio parameter can be used to define a relationship between the level of one or more audio parameters and the level of the control element at 960. In various embodiments, AI techniques (e.g., a regression model or a machine learning algorithm) are used to determine the correlation between the user-selected level of the user-specified parameter and the audio parameter, as well as the relationship between the level of one or more audio parameters and the level of the control element.
[0139] Back to Figure 8 , the relationship between the levels of one or more audio parameters and the levels of the control elements can then be implemented by the music generator module 160 to determine how to adjust the audio parameters in the output music content 140 based on the input of the control element levels 832 received from the user-created control elements 830. In certain embodiments, the music generator module 160 implements a machine learning algorithm to generate the output music content 140 and the relationship based on the input of the control element levels 832 received from the user-created control elements 830. For example, the machine learning algorithm can analyze how the metadata description of the audio track changes throughout the recording. The machine learning algorithm can include, for example, a neural network, a Markov model, or a dynamic Bayesian network.
[0140] As described herein, a machine learning algorithm can be trained to predict metadata for an upcoming music segment when provided with music metadata up to that point. The music generator module 160 can implement the prediction algorithm by searching a pool of pre-recorded audio files for those files that have attributes that most closely match the metadata predicted to appear next. Selecting the closest matching audio file to play next helps create output music content with a sequential progression of musical attributes that is similar to the example recordings used to train the prediction algorithm.
[0141] In some embodiments, parameter control of the music generator module 160 using the prediction algorithm can be included in the prediction algorithm itself. In such embodiments, some predefined parameters can be used as the input of the algorithm together with the music metadata, and the prediction changes based on the parameters. Alternatively, parameter control can be applied to the prediction to modify these predictions. As an example, the closest music fragment that will appear next is predicted by the sequential selection prediction algorithm and the audio of the file is connected end to end to make a generated work. At some point, the audience can increase the control element level (for example, the control element of each beat starting), and modify the output of the prediction model by increasing the "each beat starting" data field of the prediction. When selecting the next audio file to be attached to the work, it is more likely to select an audio file with a higher every beat starting attribute in this scenario.
[0142] In various embodiments, a generation system that utilizes metadata descriptions of musical content, such as the music generator module 160, may use hundreds or thousands of data fields in the metadata for each piece of music. To provide more variability, multiple concurrent tracks can be used, each featuring different instruments and vocal types. In these instances, the predictive model may have tens of thousands of data fields representing musical attributes, each with a significant impact on the listening experience. To allow listeners to control the music in such instances, an interface for modifying each data field output by the predictive model can be used, thereby creating thousands of control elements. Alternatively, multiple data fields can be combined and exposed as a single control element. As a single control element influences more musical attributes, the control element becomes more abstract for specific musical attributes, and the labeling of these controls becomes subjective. In this way, primary control elements and sub-parameter control elements (described below) can be implemented for dynamic and personalized control of the output musical content 140.
[0143] As described herein, users can specify their own control elements and train the music generator module 160 on how to behave based on the user's adjustments to the control elements. This process can reduce bias and complexity, and the data fields can be completely hidden from the audience. For example, in some embodiments, the user-created control elements are presented to the audience on a user interface. The audience is then presented with a short musical clip and asked to set the level of the control element that they think best describes the music they are hearing. By repeating this process, multiple data points can be created that can be used to perform regression modeling on the expected effect of the control on the music. In some embodiments, these data points can be added as additional inputs to the predictive model. The predictive model can then attempt to predict a musical sequence that will produce a composition similar to the sequence it has been trained on while also matching the musical properties of the expected behavior of the control element set to a specific level. Alternatively, a control element mapper in the form of a regression model can be used to map predicted modifiers to control elements without retraining the predictive model.
[0144] In some embodiments, training for a given control element can include global training (e.g., training based on feedback from multiple user accounts) and local training (e.g., training based on feedback from the current user account). In some embodiments, a set of control elements can be created that are specific to a subset of musical elements provided by a composer. For example, a scenario might include an artist creating a loop package and then training the music generator module 160 using examples of performances or works they have previously created using these loops. The patterns in these examples can be modeled using a regression or neural network model and used to create rules for creating new music with similar patterns. These rules can be parameterized and exposed as control elements for composers to manually modify offline before listeners begin using the music generator module 160, or for listeners to adjust while listening. Examples that the composer feels are opposite to the desired effect of the control can also be used for negative reinforcement.
[0145] In some embodiments, in addition to utilizing patterns in sample music, the music generator module 160 can find patterns in the music it creates that correspond to composer input before listeners begin listening to the generated music. The composer can do this through direct feedback (described below), such as tapping a thumbs-up control element for positive reinforcement of a pattern, or tapping a thumbs-down control element for negative reinforcement.
[0146] In various embodiments, the music generator module 160 can allow composers to create their own sub-parameter control elements, such as the control elements learned by the music generator module as described below. For example, a control element for "intensity" may have been created as a primary control element from a learning pattern related to the number of note onsets per beat and the texture quality of the instrument played. The composer can then create two sub-parameter control elements by selecting a pattern related to note onsets, such as a "rhythm intensity" control element and a "texture intensity" control element for a texture pattern. Examples of sub-parameter control elements include a vocal control element, the intensity of a specific frequency range (e.g., bass), complexity, tempo, etc. These sub-parameter control elements can be combined with more abstract control elements (e.g., primary control elements) such as injected energy. These composer skill control elements can be trained for the music generator module 160 by the composer in a manner similar to the user-created controls described herein.
[0147] As described herein, training the music generator module 160 to control audio parameters based on input from user-created control elements allows for the implementation of separate control elements for different users. For example, one user may associate increased attitude with increased bass content, while another user may associate increased attitude with a certain type of vocals or a certain beat count range. The music generator module 160 can modify the audio parameters for different attitude descriptions based on the training of the music generator module for a specific user. In some embodiments, personalized controls can be used in conjunction with global rules or control elements that are implemented in the same manner for many users. The combination of global and local feedback or control can provide high-quality music production with specialized controls for the individuals involved.
[0148] In various embodiments, such as Figure 8 As shown, one or more UI control elements 830 are implemented in the UI module 820. As described above, during interaction with the UI module 820, the user can use (one or more) control elements 830 to adjust (one or more) control element levels 832 to modify the output music content 140. In certain embodiments, the one or more control elements 830 are system-defined control elements. For example, the control elements can be defined as controllable parameters by the system 800. In such embodiments, the user can adjust the system-defined control elements to modify the output music content 140 according to the system-defined parameters.
[0149] In certain embodiments, system-defined UI control elements (e.g., knobs or sliders) allow the user to control abstract parameters of the output music content 140 automatically generated by the music generator module 160. In various embodiments, the abstract parameters serve as the primary control element input. Examples of abstract parameters include, but are not limited to, intensity, complexity, mood, genre, and energy level. In some embodiments, an intensity control element can adjust the number of low-frequency loops incorporated. A complexity control element can direct the number of tracks covered. Other control elements, such as a mood control element, can range from sober to happy and affect, for example, the key of the music being played, as well as other attributes.
[0150] In various embodiments, a system-defined UI control element (e.g., a knob or slider) allows a user to control the energy level of the output music content 140 automatically generated by the music generator module 160. In some embodiments, the label of the control element (e.g., "Energy") can change size, color, or other attributes to reflect the user input of adjusting the energy level. In some embodiments, as the user adjusts the control element, the current level of the control element can be output until the user releases the control element (e.g., releases a mouse click or removes a finger from a touch screen).
[0151] The system-defined energy can be an abstract parameter that is related to a number of more specific musical attributes. For example, in various embodiments, energy can be related to tempo. For example, changes in energy levels can be associated with tempo changes of a selected number of beats per minute (e.g., ~6 beats per minute). In some embodiments, within a given range of one parameter (e.g., tempo), the music generator module 160 can explore musical changes by changing other parameters. For example, the music generator module 160 can create rises and falls, create tension, change the number of tracks layered simultaneously, change the key, add or remove vocals, add or remove bass, play different melodies, etc.
[0152] In some embodiments, one or more sub-parameter control elements are implemented as (one or more) control elements 830. Sub-parameter control elements can allow for more specific control of properties that are incorporated into main control elements such as energy control elements. For example, the energy control element can modify the number of layers of percussive sounds and the amount of vocals used, but separate control elements allow for direct control of these sub-parameters so that all control elements are not necessarily independent. This way, the user can select the level of control specificity they wish to use. In some embodiments, sub-parameter control elements can be implemented for user-created control elements, as described above. For example, a user can create and label a control element that specifies a sub-parameter of another user-specified parameter.
[0153] In some embodiments, user interface module 820 allows the user to expand UI control element 830 to display the option of one or more sub-parameter user control elements. In addition, a specific artist can provide attribute information for guiding music creation under the user control of a high-level control element (e.g., an energy slider). For example, an artist can provide an "artist package" that has the artist's tracks and music creation rules. Artists can use the artist interface to provide values for sub-parameter user control elements. For example, a DJ may use rhythm and drums as control elements disclosed to the user to allow the audience to more or less combine rhythm and drums. In some embodiments, as described herein, artists or users can generate their own custom control elements.
[0154] In various embodiments, a human-in-the-loop generation system can be used to generate artifacts with the help of human intervention and control, potentially improving the quality and suitability of the generated music for personal purposes. For some embodiments of the music generator module 160, listeners can become listener-composers by controlling the generation process through interface control elements 830 implemented in the UI module 820. The design and implementation of these control elements may affect the balance between an individual's listener and composer roles. For example, highly detailed and technical control elements may reduce the impact of the generation algorithm and hand more creative control to the user, while requiring more hands-on interaction and technical skills to manage.
[0155] Conversely, higher-level control elements may reduce the required effort and interaction time while also reducing creative control. For example, for individuals who desire a more listener-like role, primary control elements, as described herein, may be advantageous. For example, primary control elements may be based on abstract parameters such as mood, intensity, or genre. These abstract parameters of music may be subjective metrics that are often interpreted individually. For example, in many cases, the listening environment influences how listeners describe music. Thus, music that listeners might describe as "relaxing" at a party may be too energetic and intense for a meditation session.
[0156] In some embodiments, one or more UI control elements 830 are implemented to receive user feedback about output music content 140. User feedback control elements can include, for example, star ratings, thumbs up / thumbs down, etc. In various embodiments, user feedback can be used to train the system to adapt to the specific taste of the user and / or apply to more global tastes of multiple users. In an embodiment with thumbs up / thumbs down (e.g., positive / negative) feedback, the feedback is binary. Binary feedback including strong positive and strong negative responses can effectively provide positive and negative reinforcement for the function of (one or more) control elements 830. In some contemplated embodiments, the input from the thumbs up / thumbs down control element can be used to control the output music content 140 (e.g., the thumbs up / thumbs down control element is used to control the output itself). For example, the thumbs up control element can be used to modify the maximum number of repetitions of the output music content 140 currently being played.
[0157] In some embodiments, a counter for each audio file tracks how many times a section of the audio file (e.g., an 8-beat section) has been played recently. Once a file is used above a desired threshold, a bias can be applied to its selection. Over time, this bias may gradually return to zero. Together with the music sections defined by rules that set the desired function of the music (e.g., build-up, drop-off, breakdown, introduction, sustain), this repetition counter and bias can be used to shape the music into sections with coherent themes. For example, the music generator module 160 can increase the count when the thumb is pressed down, so that the audio content of the output music content 140 is encouraged to change more quickly without destroying the musical function of the section. Similarly, the music generator module 160 can decrease the count when the thumb is pressed up, so that the audio content of the output music content 140 does not deviate from repetition over a longer period of time. Before the threshold is reached and the bias is applied, other machine learning and rule-based mechanisms in the music generator module 160 may still result in the selection of other audio content.
[0158] In some embodiments, the music generator module 160 is configured to determine various contextual information (e.g., Figure 1 Contextual information 150 is shown. For example, in conjunction with receiving a "thumbs-up" indication from the user, music generator module 160 can determine the time of day, location, device speed, biometric data (e.g., heart rate), etc. from contextual information 150. In some embodiments, this contextual information can be used to train a machine learning model to generate music that the user enjoys in a variety of different contexts (e.g., the machine learning model is context-aware).
[0159] In various embodiments, the music generator module 160 determines the current environment type and takes different actions for the same user adjustment in different environments. For example, when a listener trains an "attitude" control element, the music generator module 160 can take environmental measurements and listener biometrics. During training, the music generator module 160 is trained to include these measurements as part of the control element. In this example, when the listener is doing a high-intensity workout at the gym, the "attitude" control element may affect the intensity of the drum beat. When sitting at a computer, changing the "attitude" control element may not affect the drum beat, but may increase the distortion of the bass line. In such embodiments, a single user control element can have different rule sets or differently trained machine learning models that are used individually or in combination in different ways in different listening environments.
[0160] In contrast to context awareness, if the expected behavior of the control element is static, it is likely that many controls may become necessary or desirable for each listening context music generator module 160 used. Therefore, in some embodiments, the disclosed technology can provide functionality for multiple environments with a single control element. Implementing a single control element for multiple environments can reduce the number of control elements, making the user interface simpler and the search faster. In some embodiments, the control element behavior is dynamic. The power of the control element can come from utilizing environmental measurements, such as: sound levels recorded by a microphone, heart rate measurements, time of day, and movement rate. These measurements can be used as additional input for control element training. Therefore, the same listener's interaction with the control element may produce different musical effects, depending on the environmental context in which the interaction occurs.
[0161] In some embodiments, the context-aware functionality described above is distinct from the concept of a generative music system that modifies the generation process based on the environmental context. For example, these techniques can modify the effects of user-controlled elements based on the environmental context, which can be used alone or in conjunction with the concept of generating music based on the environmental context and the output of user controls.
[0162] In some embodiments, the music generator module 160 is configured to control the generated output music content 140 to achieve a given goal. Examples of given goals include, but are not limited to, sales goals, biometric goals such as heart rate or blood pressure, and environmental noise goals. The music generator module 160 can learn how to use the techniques described herein to modify manually (user-created) or algorithmically (system-defined) generated control elements to generate output music content 140 to meet the given goal.
[0163] Target states can be measurable environmental and listener states that a listener desires to experience while listening to music using and with the aid of the music generator module 160. These target states may be influenced directly—by changing the acoustic experience of the space in which the listener is located through the music—or may be mediated through psychological effects, such as particular music encouraging focus. As an example, a listener may set a goal to have a lower heart rate during a run. By recording the listener's heart rate under different states of the available controls, the music generator module 160 learns that the listener's heart rate typically decreases when a control named "attitude" is set to a low level. Therefore, to help the listener achieve a low heart rate, the music generator module 160 may automate the "attitude" control to a low level.
[0164] The music generator module 160 can help create a specific environment by creating the kind of music that the listener desires in that specific environment. Examples include heart rate, the overall volume in the listener's physical space, store sales, etc. Some environmental sensors and state data may not be suitable for the target state. For example, the time of day can be an environmental metric used as input to achieve the target state of inducing sleep, but the music generator module 160 itself cannot control the time of day.
[0165] In various embodiments, while sensor inputs may be disconnected from the control element mapper while attempting to reach a state target, the sensors may continue to record and instead provide measurements for comparing the actual state to the target state. The difference between the target and actual environmental states may be formulated as a reward function for a machine learning algorithm, which may adjust the mappings in the control element mapper in a mode that attempts to achieve the target state. The algorithm may adjust the mappings to reduce the difference between the target and actual environmental states.
[0166] While musical content has many physiological and psychological effects, creating the musical content that a listener desires in a particular environment may not always help create that environment for the listener. In certain situations, there may be no effect or a negative impact on achieving the target state. In some embodiments, if the change does not meet a threshold, the music generator module 160 can adjust the properties of the music based on past results while branching in other directions. For example, if lowering the "attitude" control element does not result in a decrease in the listener's heart rate, the music generator module 160 can switch and develop new strategies using other control elements, or generate new control elements using the actual state of the target variable as positive reinforcement or negative reinforcement for a regression or neural network model.
[0167] In some embodiments, if context is found to influence the expected behavior of a control element for a particular listener, it may be suggested that the data points (e.g., audio parameters) modified by the control element in a particular context are relevant to the listener's context. Therefore, these data points can provide a good starting point for attempting to generate music that produces environmental changes. For example, if a listener always manually increases the "tempo" control element when going to a train station, the music generator module 160 can begin automatically increasing this control element when it detects that the listener is at the train station.
[0168] In some embodiments, as described herein, the music generator module 160 is trained to implement control elements that match user expectations. If the music generator module 160 were trained end-to-end for each control element (e.g., from the control element level to the output music content 140), the complexity of training each control element could be high, which could slow down training. Furthermore, establishing the ideal combined effect of multiple control elements could be difficult. However, ideally, for each control element, the music generator module 160 should be trained to implement the desired musical changes based on the control element. For example, the music generator module 160 could be trained by listeners on the "Energy" control element so that rhythmic density increases with increasing "Energy." Because listeners are exposed to the final output music content 140, not just the individual layers of the music content, the music generator module 160 can be trained to use control elements to influence the final output music content. However, this can become a multi-step problem, for example, if the music should sound like X for a specific control setting, and to create music that sounds like X, a set of Y audio files should be used for each track.
[0169] In certain embodiments, a teacher / student framework is employed to address the above-mentioned issues. Figure 10 is a block diagram illustrating an exemplary teacher / student framework system according to some embodiments. In the illustrated embodiment, the system 1000 includes a teacher model implementation module 1010 and a student model implementation module 1020.
[0170] In certain embodiments, the teacher model implementation module 1010 implements a trained teacher model. For example, the trained teacher model can be a model that learns how to predict how a final mix (e.g., a stereo mix) should sound without considering the set of loops available in the final mix. In some embodiments, the learning process of the teacher model utilizes real-time analysis of the output music content 140 using a fast Fourier transform (FFT) to calculate the distribution of sounds at different frequencies for a sequence of short time steps. The teacher model can utilize a time series prediction model such as a recurrent neural network (RNN) to search for patterns in these sequences. In some embodiments, the teacher model in the teacher model implementation module 1010 can be trained offline on stereo recordings for which no separate loops or audio files are available.
[0171] In the illustrated embodiment, the teacher model implementation module 1010 receives the output music content 140 and generates a compact description 1012 of the output music content. Using a trained teacher model, the teacher model implementation module 1010 can generate the compact description 1012 without regard to the tracks or audio files in the output music content 140. The compact description 1012 can include a description X of what the output music content 140 should sound like, as determined by the teacher model implementation module 1010. The compact description 1012 is more compact than the output music content 140 itself.
[0172] The compact description 1012 can be provided to a student model implementation module 1020. The student model implementation module 1020 implements a trained student model. For example, the trained student model may be a model that learns how to use audio file or loop Y (as opposed to X) to produce music that matches the compact description. In the illustrated embodiment, the student model implementation module 1020 generates student output musical content 1014 that substantially matches the output musical content 140. As used herein, the phrase "substantially matches" means that the student output musical content 1014 sounds similar to the output musical content 140. For example, a trained listener may perceive the student output musical content 1014 and the output musical content 140 to sound the same.
[0173] In many instances, control elements are expected to affect similar patterns in music. For example, control elements may affect pitch relationships and rhythm. In some embodiments, the music generator module 160 is trained for a large number of control elements based on a teacher model. By using a single teacher model to train the music generator module 160 for a large number of control elements, it may not be necessary to relearn similar basic patterns for each control element. In such embodiments, the student model of the teacher model then learns how to change the loop selection of each track to achieve the desired properties in the final music mix. In some embodiments, the properties of the loop can be pre-calculated to reduce the learning challenge and baseline performance (although it may be at the expense of potentially reducing the possibility of finding the best mapping for the control element).
[0174] Non-limiting examples of musical attributes that can be pre-computed for each loop or audio file that can be used for student model training include the following: ratio of bass to treble frequencies, number of note onsets per second, ratio of detected tonal to atonal sounds, spectral range, average onset intensity. In some embodiments, the student model is a simple regression model that is trained to select loops for each track to obtain the closest musical attribute in the final stereo mix. In various embodiments, the student / teacher model framework may have several advantages. For example, if new attributes are added to the loop's pre-computed routine, there is no need to retrain the entire end-to-end model, only the student model.
[0175] As another example, since properties that affect the final stereo mix for different controls may be common to other control elements, training the music generator module 160 for each control element as an end-to-end model would mean that each model would need to learn the same thing (stereo mix musical features) to obtain optimal loop selection, making training slower and more difficult than it might need to be. Only the stereo output needs to be analyzed in real time, and since the output musical content is generated in real time for the listener, the music generator module 160 can obtain a "free" signal through computation. Even the FFT might have been applied for visualization and audio mixing purposes. In this way, the teacher model can be trained to predict the combined behavior of the control elements, and the music generator module 160 is trained to find ways to adapt to the other control elements while still producing the desired output musical content. This can encourage training of control elements to emphasize the unique effects of specific control elements and reduce control elements that have the effect of weakening the influence of other control elements.
[0176] Exemplary Low-Resolution Pitch Detection System
[0177] Pitch detection that is robust to polyphonic musical content and multiple instrument types can traditionally be difficult to achieve. Tools implementing end-to-end music transcription may record audio and attempt to generate a symbolic musical representation in the form of written sheet music or MIDI. Without knowledge of beat positions or tempo, these tools may need to infer musical rhythmic structure, instrumentation, and pitch. Results can vary, with a common problem being the detection of too many short, non-existent notes in an audio file and the detection of harmonics of notes as the fundamental pitch.
[0178] However, pitch detection can also be useful in situations where end-to-end transcription is not required. For example, to produce harmonically plausible compositions of musical loops, it is sufficient to know which pitches can be heard on each beat, without knowing the exact positions of the notes. If the length and tempo of the loop are known, then inferring the temporal positions of the beats from the audio may not be necessary.
[0179] In some embodiments, the pitch detection system is configured to detect which fundamental pitches (e.g., C, C#, ..., B) exist in short music audio files of known beat length. By reducing the problem scope and focusing on robustness to instrument texture, high-precision results for beat-resolution pitch detection can be achieved.
[0180] In some embodiments, the pitch detection system is trained on examples with known ground truth. In some embodiments, the audio data is created from sheet music data. MIDI and other notated music formats can be synthesized using software audio synthesizers with random parameters to obtain texture and effects. For each audio file, the system can generate a logarithmic spectrum with multiple frequency bins for each pitch level. Figure 2 D representation. This 2D representation is used as input to a neural network or other AI technique, where multiple convolutional layers are used to create feature representations of the audio frequency and time representations. The convolution stride and padding can be varied based on the audio file length to produce a constant model output shape with varying tempo inputs. In some embodiments, the pitch detection system appends a recurrent layer to the convolutional layer to output a time-dependent prediction sequence. A categorical cross-entropy loss can be used to compare the logical output of the neural network with the binary representation of the musical score.
[0181] The design of a combination of convolutional layers and recurrent layers may be similar to that used in speech-to-text work, but with some modifications. For example, speech-to-text often needs to be sensitive to relative pitch changes rather than absolute pitch. Therefore, the frequency range and resolution are often small. In addition, the text may need to remain unchanged so that it can be sped up in ways that are undesirable in music with static beats. The computation of the connectionist temporal classification (CTC) loss often used in speech-to-text tasks may not be needed, for example, because the length of the output sequence is known in advance, which reduces the complexity of training.
[0182] The following representation shows that there are 12 pitch levels per beat, where 1 indicates that the fundamental note is present in the musical score used to synthesize the audio (C, C#....B) and each row represents a beat, e.g., subsequent rows represent musical scores for different beats:
[0183]
[0184] In some embodiments, the neural network is trained on pseudo-randomly generated scores of classical music and 1-4 part (or more) harmony and polyphony. Data augmentation can help enhance the robustness of the musical content through filters and effects such as reverberation, which can be a challenge for pitch detection (e.g., because parts of the fundamental pitch may still exist after the original note ends). In some embodiments, the dataset can be biased and loss weights used because the pitch class is more likely to not play a note on every beat.
[0185] In some embodiments, the output format allows harmonic conflicts to be avoided on every beat while maximizing the range of harmonic contexts that the loop can use. For example, a bass loop may consist of only F and move down to E on the last beat of the loop. To most people in the key of F, this loop would sound harmonically fine. If no time resolution was provided, and only the knowledge was that there was an E and an F in the audio, it would probably end up being a sustained E with a short F, which would sound unacceptable to most people in the context of the key of F. As the resolution increases, the chances of harmonics, fretboard sounds, and glissando being detected as a single note increase, and thus additional notes may be mistakenly identified. According to some embodiments, the complexity of the pitch detection problem can be reduced and robustness to short, less significant pitch events can be increased by developing a system with optimal resolution of time and pitch information for combining short audio recordings of instruments to create musical mixes with combinations of harmonic sounds.
[0186] In various embodiments of the music generator system described herein, the system may allow listeners to select audio content that is used to create a pool from which the system constructs (generates) new music. This approach may be different from creating a playlist because the user does not need to select individual tracks or organize the selections in a sequential order. In addition, content from multiple artists can be used together at the same time. In some embodiments, the music content is grouped into "packages" designed by the software provider or contributing artists. A package contains multiple audio files with corresponding image features and feature metadata files. For example, a single package may contain 20 to 100 audio files that can be used by the music generator system to create music. In some embodiments, a single package can be selected or multiple packages can be selected in combination. During playback, packages can be added or deleted without stopping the music.
[0187] Exemplary Audio Techniques for Music Content Generation
[0188] In various embodiments, software frameworks for managing real-time audio generation can benefit from supporting specific types of functionality. For example, audio processing software might follow a modular signal chain metaphor inherited from analog hardware, where different modules providing audio generation and audio effects are linked together to form an audio signal graph. Individual modules often expose various continuous parameters, allowing the module's signal processing to be modified in real time. In the early days of electronic music, parameters themselves were often analog signals, so the parameter processing chain and the signal processing chain coincided. Since the digital revolution, parameters are often a separate digital signal.
[0189] The embodiments disclosed herein recognize that for real-time music generation systems—whether the system interacts live with a human performer or the system implements machine learning or other artificial intelligence (AI) techniques to generate music—a flexible control system that allows for the coordination and combination of parameter manipulations may be advantageous. Additionally, the present disclosure recognizes that it may also be advantageous for the effects of parameter changes to be invariant to changes in tempo.
[0190] In some embodiments, a music generator system generates new music content from replayed music content based on different parameter representations of an audio signal. For example, an audio signal can be represented by both a signal graph relative to time (e.g., an audio signal graph) and a signal graph relative to tempo (e.g., a signal graph). The signal graph is invariant to tempo, which allows tempo-invariant modification of audio parameters of the music content in addition to tempo-varying modifications based on the audio signal graph.
[0191] Figure 11 1 is a block diagram illustrating an exemplary system configured to implement audio technology in music content generation according to some embodiments. In the illustrated embodiment, system 1100 includes a graphics generation module 1110 and an audio technology music generator module 1120. Audio technology music generator module 1120 can operate as a music generator module (e.g., the audio technology music generator module is music generator module 160, as described herein) or the audio technology music generator module can be implemented as part of a music generator module (e.g., as part of music generator module 160).
[0192] In the illustrated embodiment, music content 1112 comprising audio file data is accessed by a graph generation module 1110. Graph generation module 1110 may generate a first graph 1114 and a second graph 1116 for the audio signal in the accessed music content 1112. In certain embodiments, first graph 1114 is an audio signal graph that plots the audio signal as a function of time. The audio signal may include, for example, amplitude, frequency, or a combination of both. In certain embodiments, second graph 1116 is a signal graph that plots the audio signal as a function of tempo.
[0193] In certain embodiments, Figure 11 As shown in the illustrated embodiment of FIG, a graph generation module 1110 is located in the system 1100 to generate a first graph 1114 and a second graph 1116. In such an embodiment, the graph generation module 1110 can be collocated with the audio-technical music generator module 1120. However, other embodiments are contemplated in which the graph generation module 1110 is located in a separate system and the audio-technical music generator module 1120 accesses the graph from the separate system. For example, the graph can be generated and stored on a cloud-based server accessible to the audio-technical music generator module 1120.
[0194] Figure 12 An example of an audio signal graph (eg, first graph 1114 ) is depicted. Figure 13 An example of a signal graph (eg, second graph 1116) is depicted. Figure 12 and Figure 13 In the diagram shown, each change in the audio signal is represented as a node (e.g., Figure 12 The audio signal node 1202 and Figure 13 1302 in the signal node). Thus, the parameters of a given node determine (e.g., define) the change in the audio signal at the given node. Because the first graph 1114 and the second graph 1116 are based on the same audio signal, the graphs may have similar structures, with the change between the graphs being the x-axis scale (time vs. beats). Having similar structures in the graphs allows for modification of parameters of a node in one graph (e.g., node 1302 in the second graph 1116) corresponding to a node in another graph (e.g., node 1202 in the first graph 1114) that is determined by parameters downstream or upstream of the node in one graph (e.g., node 1202 in the first graph 1114) (as described below).
[0195] Back to Figure 11 , first map 1114 and second map 1116 are received (or accessed) by audio-techno music generator module 1120. In certain embodiments, audio-techno music generator module 1120 generates new music content 1122 from playback music content 1118 based on the audio modifier parameters selected from first map 1114 and the audio modifier parameters selected from second map 1116. For example, audio-techno music generator module 1120 may modify playback music content 1118 using the audio modifier parameters from first map 1114, the audio modifier parameters from second map 1116, or a combination thereof. New music content 1122 is generated by modifying playback music content 1118 based on the audio modifier parameters.
[0196] In various embodiments, the audio-technique music generator module 1120 can select audio modifier parameters to implement in modifying the playback content 1118 based on whether a beat-change modification, a beat-invariant modification, or a combination thereof is desired. For example, a beat-change modification can be performed based on audio modifier parameters selected or determined from the first graph 1114, while a beat-invariant modification can be performed based on audio modifier parameters selected or determined from the second graph 1116. In embodiments where a combination of beat-change and beat-invariant modifications is desired, audio modifier parameters can be selected from both the first graph 1114 and the second graph 1116. In some embodiments, the audio modifier parameters from each separate graph are applied separately to different attributes (e.g., amplitude or frequency) or different layers (e.g., different instrument layers) in the playback music content 1118. In some embodiments, the audio modifier parameters from each graph are combined into a single audio modifier parameter to be applied to a single attribute or layer in the playback music content 1118.
[0197] Figure 14 An exemplary system for implementing real-time modification of music content using an audio-technical music generator module 1420 is depicted in accordance with some embodiments. In the illustrated embodiment, the audio-technical music generator module 1420 includes a first node determination module 1410, a second node determination module 1420, an audio parameter determination module 1430, and an audio parameter modification module 1440. The first node determination module 1410, the second node determination module 1420, the audio parameter determination module 1430, and the audio parameter modification module 1440 together implement the system 1400.
[0198] In the illustrated embodiment, the audio-technical music generator module 1420 receives playback music content 1418 including an audio signal. The audio-technical music generator module 1420 may process the audio signal in the first node determination module 1410 through a first graph 1414 (e.g., a time-based audio signal graph) and a second graph 1416 (e.g., a beat-based signal graph). As the audio signal passes through the first graph 1414, the parameters of each node in the graph determine the change in the audio signal. In the illustrated embodiment, the second node determination module 1420 may receive information about the first node 1412 and determine information about the second node 1422. In a specific embodiment, the second node determination module 1420 reads the parameters in the second graph 1416 based on the position of the first node found in the first node information 1412 in the audio signal passing through the first graph 1414. Thus, as an example, the first node determined by the first node determination module 1410 to go to the first graph 1414 ( Figure 12 The audio signal of the node 1202 in the second graph 1416 (shown) can trigger the second node determination module 1420 to determine Figure 13 ) in the corresponding (parallel) node 1302.
[0199] like Figure 14 As shown, the audio parameter determination module 1430 can receive the second node information 1422 and determine (e.g., select) a specified audio parameter 1432 based on the second node information. For example, the audio parameter determination module 1430 can select the audio parameter based on a portion of the next beat (e.g., x next beats) in the second graph 1416 after the position of the second node identified in the second node information 1422. In some embodiments, a beat-to-real-time conversion can be implemented to determine the portion of the second graph 1416 from which the audio parameter can be read. The specified audio parameter 1432 can be provided to the audio parameter modification module 1440.
[0200] The audio parameter modification module 1440 can control the modification of musical content to generate new musical content. For example, the audio parameter modification module 1440 can modify the playback musical content 1418 to generate new musical content 1122. In certain embodiments, the audio parameter modification module 1440 modifies the properties of the playback musical content 1418 by modifying specified audio parameters 1432 (determined by the audio parameter determination module 1430) of the audio signal in the playback musical content. For example, modifying the specified audio parameters 1432 of the audio signal in the playback musical content 1418 modifies a property of the audio signal such as amplitude, frequency, or a combination thereof. In various embodiments, the audio parameter modification module 1440 modifies the properties of different audio signals in the playback musical content 1418. For example, different audio signals in the playback musical content 1418 may correspond to different instruments represented in the playback musical content 1418.
[0201] In some embodiments, the audio parameter modification module 1440 uses a machine learning algorithm or other AI technology to modify the properties of the audio signal in the playback music content 1418. In some embodiments, the audio parameter modification module 1440 modifies the properties of the playback music content 1418 based on user input to the module, which user input can be provided through a user interface associated with the music generation system. It is also conceivable that the audio parameter modification module 1440 uses a combination of AI technology and user input to modify the properties of the playback music content 1418. Various embodiments for modifying the properties of the playback music content 1418 by the audio parameter modification module 1440 allow real-time manipulation of the music content (e.g., manipulation during playback). As described above, real-time manipulation can include applying a beat change modification, a beat-invariant combination, or a combination of the two to the audio signal in the playback music content 1418.
[0202] In some embodiments, the audio-technical music generator module 1420 implements a 2-tier parameter system for modifying properties of the played back music content 1418 via the audio parameter modification module 1440. In the 2-tier parameter system, a distinction may be made between "automation" (e.g., tasks performed automatically by the music generation system) that directly controls the audio parameter values and "modulation" that multiply audio parameter modifications over the automation, as described below. The 2-tier parameter system may allow different parts of the music generation system (e.g., different machine learning models in the system architecture) to consider different aspects of the music. For example, one part of the music generation system may set the volume of a specific instrument based on the type of intended section of the piece, while another part may overlay periodic variations in volume for added interest.
[0203] Exemplary techniques for real-time audio effects in music content generation
[0204] Music technology software often allows composers / producers to control various abstract envelopes through automation. In some embodiments, automation is a pre-programmed temporal manipulation of some audio processing parameter (e.g., volume or reverb amount). Automation is typically a manually defined breakpoint envelope (e.g., a piecewise linear function) or a programmed function such as a sine wave (also known as a low-frequency oscillator (LFO)).
[0205] The disclosed music generator system may differ from typical music software. For example, most parameters are automated by default. The AI technology in the music generator system can control most or all audio parameters in various ways. At a basic level, a neural network can predict the appropriate setting for each audio parameter based on its training. However, it may be helpful to provide the music generator system with some higher-level automation rules. For example, a large musical structure may require a slow build-up of volume as an additional consideration, beyond the low-level settings that can be predicted.
[0206] The present disclosure generally relates to an information architecture and procedural methods for combining multiple parameter commands issued simultaneously by different levels of a hierarchical generative system to create a continuous output that is musically coherent and varied. The disclosed music generator system can create long-form musical experiences designed to be experienced continuously for hours. Long-form musical experiences require creating a coherent musical journey for a more satisfying experience. To this end, the music generator system can reference itself over a wide range of timescales. These references can range from direct to abstract.
[0207] In certain embodiments, to facilitate larger-scale music rules, the music generator system (eg, music generator module 160) exposes an automation API (application programming interface). Figure 15 A block diagram depicts an exemplary API module in a system for audio parameter automation, according to some embodiments. In the illustrated embodiment, system 1500 includes API module 1505. In certain embodiments, API module 1505 includes automation module 1510. The music generator system can support wavetable LFOs and arbitrary breakpoint envelopes. Automation module 1510 can apply automation 1512 to any audio parameter 1520. In some embodiments, automation 1512 is applied recursively. For example, any program automation (such as a sine wave) that has parameters (frequency, amplitude, etc.) can have automation applied to those parameters.
[0208] In various embodiments, automation 1512 includes a signal graph that is parallel to the audio signal graph, as described above. The signal graph can be processed similarly: through a "pull" technique. In a "pull" technique, API module 1505 can request automation module 1510 to recalculate as needed, and the recalculation is performed so that automation 1512 recursively requests the upstream automation it depends on to do the same. In certain embodiments, the signal graph for automation is updated at a controlled rate. For example, the signal graph can be updated once per run of the performance engine update routine, which can be aligned with the block rate of the audio (e.g., after the audio signal graph renders a block (e.g., a block is 512 samples)).
[0209] In some embodiments, it may be desirable for the audio parameters 1520 themselves to change at the audio sample rate, as otherwise discontinuous parameter changes at audio block boundaries can result in audible artifacts. In certain embodiments, the music generator system manages this issue by treating automation updates as parameter value targets. When the real-time audio thread renders an audio block, the audio thread will smoothly fade the given parameter from its current value to the provided target value over the course of the block.
[0210] The music generator system described herein (e.g., Figure 1 The music generator module 160 shown in FIG. 1 may have an architecture with a hierarchical nature. In some embodiments, different parts of the hierarchy may provide multiple suggestions for values of a particular audio parameter. In certain embodiments, the music generator system provides two independent mechanisms for combining / resolving multiple suggestions: modulation and overlay. Figure 15 In the illustrated embodiment, modulation 1532 is implemented by modulation module 1530 and overlay 1542 is implemented by overlay module 1540.
[0211] In some embodiments, automation 1512 can be declared as a modulation 1532. Such a declaration can mean that rather than directly setting the value of an audio parameter, automation 1512 should act on the current value of the audio parameter in a multiplicative manner. Thus, a large musical section can apply a long modulation 1532 to an audio parameter (e.g., a slow fader on a volume fader), and the value of the modulation will be multiplied by whatever value the rest of the music generator system may dictate.
[0212] In various embodiments, the API module 1505 includes an overlay module 1540. The overlay module 1540 can be, for example, an overlay tool for audio parameter automation. The overlay module 1540 can be intended to be used by an external control interface (e.g., an artist control user interface). The overlay module 1540 can control the audio parameter 1520 regardless of what the music generator system attempts to do with it. When an overlay 1542 overwrites an audio parameter 1520, the music generator system can create a "shadow parameter" 1522 that tracks where the audio parameter would be if it were not overwritten (e.g., where the audio parameter would be based on automation 1512 or modulation 1532). Therefore, when the overlay 1542 is "released" (e.g., removed by the artist), the audio parameter 1520 can quickly return to where it should have been based on automation 1512 or modulation 1532.
[0213] In various embodiments, these two approaches can be combined. For example, the override 1542 can be a modulation 1532. When the override 1542 is a modulation 1532, the base value of the audio parameter 1520 can still be set by the music generator system, but then modulated in a multiplicative manner by the override 1542 (overriding any other modulation). Each audio parameter 1520 can simultaneously have one (or zero) automation 1512 and one (or zero) modulation 1532, as well as one (or zero) of each override 1542.
[0214] In various embodiments, the abstract class hierarchy is defined as follows (note that there is some multiple inheritance):
[0215] Automatable
[0216] Automation Parameters
[0217] parameter
[0218] Audio Node Parameters
[0219] Macro parameters
[0220] Shadow Parameters
[0221] Beat Dependence
[0222] automation
[0223] Envelope
[0224] Parameter Follower
[0225] regular
[0226] Transformation Automation
[0227] Hyperautomation
[0228] Macro parameters
[0229] Shadow Parameters
[0230] Based on an abstract class hierarchy, things can be considered automated, or automatable. In some embodiments, any automation can be applied to anything automatable. Automations include things like LFOs and breakpoint envelopes. These automations are beat-locked, meaning they change over time based on the current beat.
[0231] Automations themselves may have automatable parameters. For example, the frequency and amplitude of an LFO automation are automatable. Therefore, there is a signal graph of dependent automations and automation parameters that runs in parallel with the audio signal graph, but at a control rate rather than the audio rate. As described above, the signal graph uses a pull model. The music generator system keeps track of any automations 1512 applied to audio parameters 1520 and updates these once per "game loop". Automations 1512 in turn recursively request updates to their own automated audio parameters 1520. This recursive update logic may reside in a base class BeatDependency, which is expected to be called frequently (but not necessarily regularly). The update logic may have a prototype described as follows:
[0232] BeatDependency::update(double currentBeat, int updateCounter, boolean override)
[0233] In certain embodiments, the Beat Dependency class maintains a list of its own dependencies (e.g., other Beat Dependency instances) and recursively calls their update functions. An update counter can be passed up so that the signal graph can have loops without double updating. This can be important because automation can apply to several different automatables. In some embodiments, this may not matter because the second update will have the same current beat as the first, and unless the beat changes, these update routines should have no effect.
[0234] In various embodiments, as automation is applied to each cycle of an automatable "game loop," the music generator system can request an updated value from each automation (recursively) and use it to set the automatable value. In this case, the "setting" may depend on the specific subclass, and also on whether the parameter is also being modulated and / or overridden.
[0235] In certain embodiments, modulation 1532 is automation 1512 applied multiplicatively rather than absolutely. For example, modulation 1532 can be applied to an already automated audio parameter 1520, where the effect will be a percentage of the automation value. This multiplicative approach might allow for continuous oscillation around a moving average, for example.
[0236] In some embodiments, audio parameters 1520 can be overridden, as described above, meaning that any automation 1512 or modulation 1532 applied to them, or other (less privileged) requests are overridden by the overridden value in overlay 1542. Such overriding can allow external control of specific aspects of the music generator system while the music generator system continues to operate. When an audio parameter 1520 is overridden, the music generator system keeps track of what the value will be (e.g., to keep track of applied automation / modulation and other requests). When the overlay is released, the music generator system snaps the parameter back to where it should have been.
[0237] To facilitate modulation 1532 and overriding 1542, the music generator system can abstract the setValue method of Parameter. There may also be a private method _setValue that actually sets the value. An example of a public method is as follows:
[0238]
[0239] The public methods can reference a member variable of the Parameter class named _unmodulated. This variable is an instance of ShadowParameter, as described above. Each audio parameter 1520 has a shadow parameter 1522 that tracks where it would be if it were not modulated. If the audio parameter 1520 is not currently modulated, both the audio parameter 1520 and its shadow parameter 1522 are updated with the requested value. Otherwise, the shadow parameter 1522 tracks the request and the actual audio parameter value 1520 is set elsewhere (e.g., in the updateModulations routine—where the modulation factor is multiplied by the shadow parameter value to give the actual parameter value).
[0240] In various embodiments, large-scale structure in long-form musical experiences is achieved through various mechanisms. One broad approach might be to use musical self-reference over time. For example, a very straightforward self-reference would be to repeat exactly a particular audio segment that was previously played. In music theory, a repeated segment may be called a theme (or motif). More typically, musical content uses a theme and variations, whereby a theme is repeated at a later time with some variations to provide a sense of continuity but maintain a sense of progression. The music generator system disclosed herein can use the theme and variations to create large-scale structure in a variety of ways, including direct repetition or through the use of abstract envelopes.
[0241] An abstract envelope is the time-varying value of an audio parameter. Abstracted from the audio parameter it controls, an abstract envelope can be applied to any other audio parameter. For example, a collection of audio parameters can be consistently automated using a single, controlled abstract envelope. This technique can perceptually "glue" different layers together in the short term. An abstract envelope can also be reused and applied to different audio parameters temporarily. In this way, an abstract envelope becomes an abstract musical theme, which can be repeated by applying the envelope to different audio parameters later in the listening experience. Thus, variations on the theme exist while building a sense of structure and long-term coherence.
[0242] As a musical theme, the abstract envelope can abstract many musical features. Examples of musical features that can be abstracted include, but are not limited to:
[0243] Build tension (volume, distortion level, etc.) in any track.
[0244] Rhythm (volume adjustment and / or gate control to create rhythmic effects applied to pads, etc.)
[0245] Melodic (pitch filtering can emulate a melodic contour applied to pads, etc.)
[0246] Exemplary Additional Audio Techniques for Real-Time Music Content Generation
[0247] Real-time music content generation can present unique challenges. For example, due to strict real-time constraints, function calls or subroutines with unpredictable and potentially infinite execution times should be avoided. Avoiding this issue can preclude the use of most high-level programming languages, as well as most low-level languages (such as C and C++). Anything that allocates memory from the heap (e.g., via the malloc function under the hood) is likely to be ruled out, as is anything that might block, such as locking a mutex. This can make multithreaded programming particularly difficult for real-time music content generation. Most standard memory management methods may also be impractical, so dynamic data structures (e.g., C++, STL containers) are of limited use for real-time music content generation.
[0248] Another challenging area can be managing audio parameters involved in DSP (digital signal processing) functions (e.g., filter cutoff frequencies). For example, when dynamically changing audio parameters, audible artifacts may occur unless the audio parameters change continuously. Therefore, communication between the real-time DSP audio thread and the user-facing or programming interface may be required to change the audio parameters.
[0249] Various audio software can be implemented to deal with these limitations, and there are various approaches. For example:
[0250] Inter-thread communication can be handled using lock-free message queues.
[0251] Functions can be written in pure C and use function pointer callbacks.
[0252] Memory management can be implemented through custom "regions" or "arenas"
[0253] A "dual-speed" system can be implemented with a real-time audio thread running at the audio rate and a control audio thread running at a "control rate." The control audio thread can set a target for audio parameter changes, and the real-time audio thread smoothly ramps up to that target.
[0254] In some embodiments, synchronization between the control rate audio parameter manipulation and the real-time audio thread-safe storage of the audio parameter values for the actual DSP routines may require some thread-safe communication of the audio parameter targets. Most audio parameters of the audio routines are continuous (rather than discrete) and therefore typically represented by floating-point data types. Due to the lack of lock-free atomic floating-point data types, various contortions of the data have historically been required.
[0255] In a specific embodiment, a simple lock-free atomic floating-point data type is implemented in the music generator system described herein. This can be achieved by treating the floating-point type as a sequence of bits and "tricking" the compiler into treating it as an atomic integer type with the same bit width. This approach supports atomic get / sets, which is suitable for the music generator system described herein. An example implementation of the lock-free atomic floating-point data type is described below:
[0256] / / atomic float
[0257] class af32{
[0258] public:
[0259] af32(){}
[0260] af32(float x){operator()(x);}
[0261] ~af32(){}
[0262] af32(const af32&x):valueStore(x()){}
[0263] af32&operator=(const af32&x){this->operator()(x()); return*this;}
[0264] float operator()()const{uint32_t voodoo=atomic_load(&valueStore);
[0265] return ((float) &voodoo);}
[0266] void operator()(float value){
[0267] uint32_t voodoo=((uint32_t)&value); atomic_store(&_valueStore,voodoo);
[0268] }
[0269] private:
[0270] std::atomic_uint32_t_valueStore{0};
[0271] };
[0272] In some embodiments, dynamic memory allocation from the heap is not feasible for real-time code associated with music content generation. For example, static stack-based allocation may make it difficult to use programming techniques such as dynamic storage containers and functional programming methods. In a specific embodiment, the music generator system described herein implements "memory areas" for memory management in a real-time context. As used herein, a "memory area" is an area of heap-allocated memory that is pre-allocated in the absence of real-time constraints (e.g., when real-time constraints do not yet exist or are suspended). Memory storage objects can then be created in the heap-allocated memory area without requesting more memory from the system, thereby making the memory real-time safe. Garbage collection may include deallocation of memory areas as a whole. The memory implementation of the music generator system can also be multi-threaded safe, real-time safe, and efficient.
[0273] Figure 16 A block diagram of an exemplary memory area 1600 is depicted in accordance with some embodiments. In the illustrated embodiment, the memory area 1600 includes a heap allocated memory module 1610. In various embodiments, the heap allocated memory module 1610 receives and stores a first map 1114 (e.g., an audio signal map), a second map 1116 (e.g., a signal map), and audio signal data 1602. For example, the audio signal may be modified by the audio parameter modification module 1440 ( Figure 14 ) to retrieve each stored item.
[0274] An example implementation of a memory region is described below:
[0275] / / Memory pool class memory area
[0276] {public:
[0277] MemoryZone(uint64_t sz):sz(sz),zone((char*)malloc(sz)){}
[0278] ~MemoryZone(){free(zone);}
[0279] void* bags(size_t obj_size, size_t alignment) {
[0280] uint64_t p = atomic_load(&p);
[0281] uint64_t q = p % uint64_t(alignment);
[0282] if (p + q > sz) return nullptr;
[0283] uint64_t pp = atomic_fetch_add(&p, uint64_t(obj_size) + q);
[0284] [[ID=二十八]]if (pp == p) { return zone_ + p + q;}
[0285] else { return bags(obj_size, alignment);}
[0286] } uint64_t used() { return atomic_load(&p);}
[0287] uint64_t available() { return int64_t(sz) - int64_t(atomic_load(&p));}
[0288] void hose() { atomic_store(&p, 0ULL);}
[0289] private:
[0290] char* zone;
[0291] uint64_t sz;
[0292] It should be noted that there are some unclear or potentially incorrect parts in the original code (such as the variable `zone_` in line 29 which is not defined before), and the translation is done as accurately as possible based on the existing content.std::atomic_uint64_t p_{0};
[0293] };
[0294] In some embodiments, different audio threads of the music generator system need to communicate with each other. Typical thread-safe methods (which may involve locking "mutex" data structures) may not be usable in a real-time context. In certain embodiments, dynamically routed data serialization to a single-producer, single-consumer circular buffer pool is implemented. A circular buffer is a FIFO (first-in, first-out) queue data structure that generally does not require dynamic memory allocation after initialization. A single-producer, single-consumer thread-safe circular buffer may allow one audio thread to push data into the queue while another audio thread pulls data out. For the music generator system described herein, the circular buffer can be expanded to allow multiple producer, single-consumer audio threads. These buffers can be implemented by pre-allocating a static array of circular buffers and dynamically routing the serialized data to a specific "channel" (e.g., a specific circular buffer) based on an identifier added to the music content produced by the music generator system. A single user (e.g., a single consumer) can access the static array of circular buffers.
[0295] Figure 17 A block diagram depicts an exemplary system for storing new music content, according to some embodiments. In the illustrated embodiment, system 1700 includes a circular buffer static array module 1710. Circular buffer static array module 1710 can include multiple circular buffers that allow for storage of multi-producer, single-consumer audio threads based on thread identifiers. For example, circular buffer static array module 1710 can receive new music content 1122 and store the new music content in 1712 for access by a user.
[0296] In various embodiments, abstract data structures such as dynamic containers (vectors, queues, lists) are typically implemented in a non-real-time safe manner. However, these abstract data structures may be very useful for audio programming. In a specific embodiment, the music generator system described herein implements a custom list data structure (e.g., a singly linked list). Many functional programming techniques can be implemented from the custom list data structure. The custom list data structure implementation can use "memory areas" (as described above) for underlying memory management. In some embodiments, the custom list data structure is serializable, which can make it safe for real-time use and enable communication between audio threads using the above-mentioned multi-producer, single-consumer audio thread.
[0297] Exemplary blockchain ledger technology
[0298] In some embodiments, the disclosed system may utilize secure recording technologies such as blockchain or other cryptographic ledgers to record information about the generated music or its elements, such as loops or tracks. In some embodiments, the system combines multiple audio files (e.g., tracks or loops) to generate output music content. The combination may be performed by combining multiple audio content layers so that they at least partially overlap in time. The output content may be discrete pieces of music or continuous. In the context of continuous music, tracking the use of musical elements may be challenging, for example, in order to provide royalties to relevant stakeholders. Therefore, in some embodiments, the disclosed system records identifiers and usage information (e.g., timestamps or play counts) of audio files used in the composed music content. In addition, for example, the disclosed system may utilize various algorithms to track playback time in the context of mixed audio files.
[0299] As used herein, the term "blockchain" refers to a set of records (called blocks) that are cryptographically linked. For example, each block may include a cryptographic hash of the previous block, a timestamp, and transaction data. A blockchain can serve as a public distributed ledger and can be managed by a network of computing devices that communicate and verify new blocks using an agreed-upon protocol. Some blockchain implementations may be immutable, while others may allow blocks to be subsequently altered. In general, a blockchain can record transactions in a verifiable and permanent manner. Although a blockchain ledger is discussed herein for illustrative purposes, it should be understood that in other embodiments, the disclosed technology can be used with other types of cryptographic ledgers.
[0300] Figure 18 is a diagram illustrating example playback data according to some embodiments. In the illustrated embodiment, a database structure includes entries for multiple files. Each illustrated entry includes a file identifier, a start timestamp, and a total time. The file identifier can uniquely identify an audio file tracked by the system. The start timestamp can indicate the first time the audio file was included in the mixed audio content. For example, the timestamp can be based on the local clock of the playback device or on an internet clock. The total time can indicate the length of the interval in which the audio files were merged. Note that this may be different from the length of the audio file, for example, if only a portion of the audio file is used, if the audio file is sped up or slowed down in the mix, etc. In some embodiments, when audio files are merged at multiple different times, an entry is generated each time. In other embodiments, if an entry for a file already exists, the additional playback of the file may result in the time field of the existing entry being incremented. In yet other embodiments, the data structure can track the number of times each audio file is used rather than the length of the merge. Additionally, other encodings of time-based usage data are contemplated.
[0301] In various embodiments, different devices may determine, store, and use the ledger to record playback data. Figure 19 Discuss example scenarios and topologies. Replay data can be temporarily stored on a computing device before being submitted to the ledger. The stored replay data can be encrypted, for example, to reduce or prevent manipulation of entries or insertion of erroneous entries.
[0302] Figure 19 19 is a block diagram illustrating an example composition system according to some embodiments. In the illustrated example, the system includes a playback device 1910, a computing system 1920, and a ledger 1930.
[0303] In the illustrated embodiment, playback device 1910 receives control signaling from computing system 1920 and sends playback data to computing system 1920. In this embodiment, playback device 1910 includes a playback data recording module 1912 that can record playback data based on the audio mix played by playback device 1910. Playback device 1910 also includes a playback data storage module 1914 that is configured to temporarily store playback data in a ledger, or both. Playback device 1910 can periodically report playback data to computing system 1920 or can report playback data in real time. For example, when playback device 1910 is offline, playback data can be stored for later reporting.
[0304] In the illustrated embodiment, computing system 1920 receives playback data and submits entries reflecting the playback data to ledger 1930. Computing system 1920 also sends control signaling to playback device 1910. In different embodiments, the control signaling may include various types of information. For example, the control signaling may include configuration data, mixing parameters, audio samples, machine learning updates, etc. for playback device 1910 to use in composing music content. In other embodiments, computing system 1920 may compose music content and stream the music content data to playback device 1910 via control signaling. In these embodiments, modules 1912 and 1914 may be included in computing system 1920. In general, reference Figure 19 The modules and functions discussed can be distributed across multiple devices according to various topologies.
[0305] In some embodiments, playback device 1910 is configured to submit entries directly to ledger 1930. For example, a playback device such as a mobile phone can compose music content, determine playback data, and store the playback data. In this case, the mobile device can report the playback data to a server, such as computing system 1920, or directly to the computing system (or a group of computing nodes) that maintains ledger 1930.
[0306] In some embodiments, the system maintains a record of rights holders, for example, with a mapping to an audio file identifier or to a collection of audio files. This entity record can be maintained in the ledger 1930 or in a separate ledger or in some other data structure. This can allow rights holders to remain anonymous, for example, when the ledger 1930 is public but includes non-identifying entity identifiers that are mapped to entities in some other data structure.
[0307] In some embodiments, a music composition algorithm can generate a new audio file from two or more existing audio files for inclusion in a mix. For example, the system can generate a new audio file C based on two audio files A and B. One technique for this type of mix uses interpolation between vector representations of the audio from files A and B and an inverse transform from vectors to audio representations to generate file C. In this example, the playtime of both audio files A and B can be increased, but the amount of increase may be less than their actual playtime, for example, because they are being mixed.
[0308] For example, if audio file C is incorporated into the mixed content for 20 seconds, audio file A may have playback data indicating 15 seconds, while audio file B may have playback data indicating 5 seconds (and note that the sum of the mixed audio files may or may not match the used length of the resulting file C). In some embodiments, the playback time of each original file is based on its similarity to the mixed file C. For example, in a vector embodiment, for an n-dimensional vector representation, the interpolated vector a has the following distance d from the vector representations of audio files A and B:
[0309] d(a,c)=((a1-c1) 2 +(a2-c2) 2 +…+(an-cn) 2 ) 1 / 2
[0310] d(b,c)=((b1-c1) 2 +(b2-c2) 2 +…+(bn-cn) 2 ) 1 / 2
[0311] In these embodiments, the playback time i of each original file can be determined as:
[0312]
[0313] Where t represents the playback time of file C.
[0314] In some embodiments, a form of compensation can be incorporated into the ledger structure. For example, a particular entity can include information associating an audio file with a performance requirement, such as displaying a link or including an advertisement. In these embodiments, when an audio file is included in a mix, the synthesis system can provide proof of performance of the associated operation (e.g., displaying an advertisement). The proof of performance can be reported in one of a variety of appropriate reporting templates that require specific fields to show how and when the operation was performed. The proof of performance can include time information and utilize cryptography to avoid erroneous performance assertions. In these embodiments, the use of audio files that also do not show proof of performance of the associated required operation may require some other form of compensation, such as a royalty payment. Typically, different entities that submit audio files can register for different forms of compensation.
[0315] As described above, the disclosed technology can provide a reliable record of audio files used in music mixing, even when synthesizing in real time. The public nature of the ledger may provide confidence in the fairness of compensation. This in turn may encourage participation from artists and other collaborators, potentially increasing the variety and quality of audio files available for automatic mixing.
[0316] In some embodiments, an artist pack can be made with the elements that a music engine uses to create a continuous soundscape. An artist pack can be a professionally (or otherwise) curated set of elements that are stored in one or more data structures associated with an entity such as an artist or group. Examples of these elements include, but are not limited to, loops, composition rules, heuristics, and neural network vectors. Loops can be contained in a database of musical phrases. Each loop is typically a single instrument or a group of related instruments playing a musical progression over a period of time. These can range from short loops (e.g., 4 bars) to longer loops (e.g., 32 to 64 bars), etc. Loops can be organized into layers, such as melody, harmony, drums, bass, top, FX, etc. The loop database can also be represented as a variational autoencoder with an encoded loop representation. In this case, the loops themselves are not needed, but rather the NN is used to generate the sounds encoded in the NN.
[0317] Heuristics are parameters, rules, or data that guide a music engine in creating music. Parameters guide elements such as part length, the use of effects, the frequency of variational techniques, the complexity of the music, or generally speaking, any type of parameter that can be used to enhance the decision-making of a music engine when creating and rendering music.
[0318] The ledger records transactions related to the consumption of content associated with rights holders. This could be, for example, loops, heuristics, or neural network vectors. The goal of the ledger is to record these transactions and make them available for transparent accounting. The ledger is designed to capture transactions that occurred, which might include the consumption of content, the use of parameters to guide a music engine, and the use of vectors on a neural network. The ledger can record a variety of transaction types, including discrete events (e.g., this loop was played at this time), this package was played for this amount of time, or this machine learning module (e.g., a neural network module) was used for this amount of time.
[0319] The ledger can associate multiple rights holders with any given artist package, or more specifically, with specific loops or other elements of an artist package. For example, a label, artist, and composer may hold rights to a given artist package. The ledger can allow them to associate payment details for the package, specifying the percentage each party will receive. For example, the artist could receive 25%, the record label 25%, and the composer 50%. Using a blockchain to manage these transactions can allow for micropayments to be made to each rights holder in real time, or accumulated over an appropriate period of time.
[0320] As mentioned above, in some implementations, the loop may be replaced with a VAE that is essentially the encoding of the loop in the machine learning module. In this case, the ledger can associate play time with a specific artist package that contains the machine learning module. For example, if an artist package accounts for 10% of the total play time across all devices, then that artist can receive 10% of the total revenue distribution.
[0321] In some embodiments, the system allows artists to create an artist profile. This profile includes information about the artist, including a biography, a profile picture, bank details, and other data necessary to verify the artist's identity. After creating an artist profile, artists can upload and publish artist packages. These packages contain the components that the music engine uses to create soundscapes.
[0322] For each artist package created, rights holders can be defined and associated with the package. Each rights holder can claim a percentage of the package. Additionally, each rights holder creates a profile and associates a bank account with their profile for payment. Artists themselves are rights holders and may own 100% of the rights associated with their work's package.
[0323] In addition to recording events in the ledger for revenue recognition, the ledger can also manage promotions associated with artist packages. For example, an artist package might have a free month promotion, where the revenue generated would be different from the revenue generated if the promotion was not running. The ledger automatically accounts for these revenue inputs when calculating payments to rights holders.
[0324] The same rights management model can allow artists to sell the rights to their packages to one or more external rights holders. For example, when launching a new package, an artist can pre-fund their package by selling a 50% stake in the package to followers or investors. In this scenario, the number of investors / rights holders can be arbitrarily large. For example, an artist could sell their 50% stake to 100,000 users, who would receive 1 / 100,000 of the revenue generated by the package. Because all accounting is managed by the ledger, in this scenario, investors would be paid directly, without the need to audit the artist's accounts.
[0325] Exemplary User and Enterprise GUIs
[0326] FIG. 20A to FIG. 20B is a block diagram illustrating a graphical user interface according to some embodiments. In the illustrated embodiment, Figure 20A Contains the GUI displayed by the user application 2010, and Figure 20B Contains a GUI displayed by enterprise application 2030. In some embodiments, Figure 20A and Figure 20B The displayed GUI is generated by the website rather than the application. In various embodiments, any of a variety of suitable elements may be displayed, including one or more of the following elements: a dial (e.g., to control volume, energy, etc.), a button, a knob, a display box (e.g., to provide updated information to the user), etc.
[0327] exist Figure 20A , user application 2010 displays a GUI including a section 2012 for selecting one or more artist packs. In some embodiments, packs 2014 may alternatively or additionally include themed packs or packs for specific occasions (e.g., weddings, birthday parties, graduations, etc.). In some embodiments, the number of packs displayed in section 2012 is greater than the number that can be displayed in section 2012 at one time. Therefore, in some embodiments, the user scrolls up and / or down in section 2012 to view one or more packs 214. In some embodiments, the user can select an artist pack 2014 based on the output music content that the user wants to hear. In some embodiments, for example, the artist packs can be purchased and / or downloaded.
[0328] In the illustrated embodiment, selection element 2016 allows the user to adjust one or more music attributes (e.g., energy level). In some embodiments, selection element 2016 allows the user to add / delete / modify one or more target music attributes. In various embodiments, selection element 2016 may present one or more UI control elements (e.g., control element 830).
[0329] In the illustrated embodiment, selection element 2020 allows a user to have a device (e.g., a mobile device) listen to the environment to determine target music attributes. In some embodiments, the device uses one or more sensors (e.g., a camera, microphone, thermometer, etc.) to collect information about the environment after the user selects selection element 2020. In some embodiments, application 2010 also selects or suggests one or more artist packs based on the environmental information collected by the application when the user selects element 220.
[0330] In the illustrated embodiment, selection element 2022 allows a user to combine multiple artist packages to generate a new rule set. In some embodiments, the new rule set is based on the user selecting one or more packages for the same artist. In other embodiments, the new rule set is based on the user selecting one or more packages for different artists. The user can indicate weights for different rule sets, for example, so that highly weighted rule sets have a greater influence on the generated music than lower-weighted rule sets. The music generator can combine rule sets in a variety of different ways, for example, by switching between rules from different rule sets, averaging the values of rules from multiple different rule sets, etc.
[0331] In the illustrated embodiment, the selection element 2024 allows the user to manually adjust the rules in one or more rule sets. For example, in some embodiments, the user wishes to adjust the music content being generated at a more granular level by adjusting one or more rules in the rule set used to generate the music content. In some embodiments, this allows the user of the application 2010 to manually adjust the rules in one or more rule sets. Figure 2 The controls displayed in the GUI in the
[0066] become their own disc jockeys (DJs) to adjust the rule set used by the music generator to generate output music content. These embodiments can also allow for more refined control over target music attributes.
[0332] exist Figure 20B In the embodiment shown, the enterprise application 2030 displays a GUI that further includes an artist package selection portion 212 having artist packages 214. In the embodiment shown, the enterprise GUI displayed by the application 2030 also includes an element 216 for adjusting / adding / deleting one or more music attributes. In some embodiments, the GUI displayed in Figure 20BThe GUI in the rule set is used in a business or storefront to create a specific environment by generating music content (e.g., for optimizing sales). In some embodiments, an employee uses application 2030 to select one or more artist packages that have been previously shown to increase sales (e.g., metadata for a given rule set can indicate actual experimental results of using the rule set in a real-world context).
[0333] In the illustrated embodiment, input hardware 2040 sends information to the application or website displaying enterprise application 2030. In some embodiments, input hardware 2040 is one of the following: a cash register, a heat sensor, a light sensor, a clock, a noise sensor, etc. In some embodiments, information sent from one or more of the above-listed hardware devices is used to adjust target music attributes and / or a rule set for generating output music content for a particular environment. In the illustrated embodiment, selection element 2038 allows a user of application 2030 to select one or more hardware devices from which to receive environmental input.
[0334] In the illustrated embodiment, display 2034 displays environmental data to a user of application 2030 based on information from input hardware 2040. In the illustrated embodiment, display 2032 shows changes to a rule set based on the environmental data. In some embodiments, display 2032 allows a user of application 2030 to see changes made based on the environmental data.
[0335] In some embodiments, Figure 20A and Figure 20B The elements shown in are used for theme packages and / or occasion packages. That is, in some embodiments, users or businesses using the GUI displayed by application 2010 and application 2030 can select / adjust / modify rule sets to generate music content for one or more occasions and / or themes.
[0336] Detailed example music generator system
[0337] Figures 21 to 23 Details regarding specific embodiments of the music generator module 160 are shown. Note that while these specific examples are disclosed for illustrative purposes, they are not intended to limit the scope of the present disclosure. In these embodiments, building music from loops is performed by a client system such as a personal computer, mobile device, media device, etc. Figures 21 to 23As used in the discussion herein, the term "loop" may be interchangeable with the term "audio file". Typically, as described herein, loops are included in audio files. Loops may be grouped into professionally curated packages of loops, which may be referred to as artist packages. Loops may be analyzed for musical properties, and the properties may be stored as loop metadata. The audio in a constructed track may be analyzed (e.g., in real time) and filtered to mix and control output streaming media. Various feedback may be sent to a server, including explicit feedback such as from user interaction with a slider or button, and implicit feedback (e.g., generated by sensors based on volume changes, based on listening length, environmental information, etc.). In some embodiments, the control input has a known effect (e.g., directly or indirectly specifies a target musical property) and is used by the combination module.
[0338] The following discussion introduces the reference Figures 21 to 23 Various terms are used. In some embodiments, a loop library is a master library of loops that can be stored by a server. Each loop can include audio data and metadata describing the audio data. In some embodiments, a loop package is a subset of the loop library. A loop package can be a package for a specific artist, a specific mood, a specific event type, etc. A client device can download a loop package for offline listening or download portions of a loop package on demand, such as for online listening.
[0339] In some embodiments, the generated streaming media is data specifying the music content that the user hears when using the music generator system. Note that for a given generated streaming media, the actual output audio signal may be slightly different, for example, based on the capabilities of the audio output device.
[0340] In some embodiments, the combination module builds a combination from the loops available in the loop package. The combination module can receive loops, loop metadata, and user input as parameters and can be executed by the client device. In some embodiments, the combination module outputs a performance script that is sent to the performance module and one or more machine learning engines. In some embodiments, the performance script outlines which loops will be played on each track of the generated streaming media and what effects will be applied to the streaming media. The performance script can use beat-related time to represent the time when events occur. The performance script can also encode effect parameters (for example, for effects such as reverb, delay, compression, equalization, etc.).
[0341] In some embodiments, the performance module receives a performance script as input and presents it as a generated streaming media. The performance module can generate multiple tracks specified by the performance script and mix these tracks into a streaming media (e.g., a stereo streaming media, although the streaming media can have various encodings in various embodiments, including surround encoding, object-based audio encoding, multi-channel stereo, etc.). In some embodiments, when a specific performance script is provided, the performance module will always generate the same output.
[0342] In some embodiments, the analysis module is a server-implemented module that receives feedback information and configures the composition module (e.g., in real time, periodically, based on administrator commands, etc.) In some embodiments, the analysis module uses a combination of machine learning techniques to correlate user feedback with performance script and loop library metadata.
[0343] Figure 21 is a block diagram illustrating an example music generator system including analysis and combination modules according to some embodiments. In some embodiments, Figure 21 The system is configured to generate potentially unlimited music streaming, allowing users to directly control the mood and style of the music. In the illustrated embodiment, the system includes an analysis module 2110, a combination module 2120, a performance module 2130, and an audio output device 2140. In some embodiments, the analysis module 2110 is implemented by a server, and the combination module 2120 and the performance module 2130 are implemented by one or more client devices. In other embodiments, modules 2110, 2120, and 2130 can all be implemented on the client device or can all be implemented on the server.
[0344] In the illustrated embodiment, the analysis module 2110 stores one or more artist packages 2112 and implements a feature extraction module 2114 , a client simulator module 2116 , and a deep neural network 2118 .
[0345] In some embodiments, the feature extraction module 2114 adds the loop to the loop library after analyzing the loop audio (although note that some loops may be received with already generated metadata and may not require analysis). For example, raw audio in a format such as wav, aiff, or FLAC can be analyzed to obtain quantifiable musical attributes such as instrument classification, pitch transcription, beat timing, tempo, file length, and audio amplitude in multiple frequency bins. The analysis module 2110 can also store more abstract musical attributes or emotional descriptions of the loop, for example, based on manual labeling by the artist or machine listening. For example, multiple discrete categories can be used to quantify emotion, where a range of values for each category is used for a given loop.
[0346] For example, consider loop A, which was analyzed to determine that the notes G2, Bb2, and D2 are used, the first beat starts at 6 milliseconds in the file, the beat count is 122 bpm, the file length is 6483 milliseconds, and the loop has normalized amplitude values of 0.3, 0.5, 0.7, 0.3, and 0.2 across five frequency bins. The artist may label the loop as "Funk" with the following mood values:
[0347] transcendence Peace strength joy sad nervous high high Low middle none Low
[0348] The analysis module 2110 can store this information in a database and the client can download sub-portions of this information, for example, as loop packages. Although artist packages 2112 are shown for illustrative purposes, the analysis module 2110 can provide various types of loop packages to the combination module 2120.
[0349] In the illustrated embodiment, the client simulator module 2116 analyzes various types of feedback to provide feedback information in a format supported by the deep neural network 2118. In the illustrated embodiment, the deep neural network 2118 also receives as input a performance script generated by the combination module. In some embodiments, the deep neural network configures the combination module based on these inputs, for example, to improve the correlation between the type of music output generated and the desired feedback. For example, the deep neural network can periodically push updates to the client device implementing the combination module 2120. Please note that the deep neural network 2118 is shown for illustrative purposes and can provide powerful machine learning performance in the disclosed embodiments, but is not intended to limit the scope of this disclosure. In various embodiments, various types of machine learning techniques can be implemented alone or in various combinations to perform similar functions. Please note that the machine learning module can be used to directly implement a rule set (e.g., a permutation rule or technique) in some embodiments, or can be used to control a module that implements other types of rule sets, for example, using the deep neural network 2118 in the illustrated embodiment.
[0350] In some embodiments, the analysis module 2110 generates combination parameters for the combination module 2120 to improve the correlation between desired feedback and the use of specific parameters. For example, actual user feedback can be used to adjust the combination parameters, for example, to try to reduce negative feedback.
[0351] As an example, consider a case where module 2110 discovers a correlation between negative feedback (e.g., clearly low rankings, low volume listening, short listening times, etc.) and works that use a large number of layers. In some embodiments, module 2110 uses techniques such as backpropagation to determine that adjusting the probability parameter for adding more tracks will reduce the frequency of this problem. For example, module 2110 may predict that reducing the probability parameter by 50% will reduce negative feedback by 8%, and may determine to perform the reduction and push the updated parameter to the combination module (note that the probability parameter is discussed in detail below, but any of the various parameters of the statistical model can be adjusted similarly).
[0352] As another example, consider a scenario where module 2110 discovers that negative feedback correlates with the user setting the mood control to high tension. A correlation may also be discovered between loops labeled low tension and users requesting high tension. In this case, module 2110 may add a parameter that increases the probability of selecting loops labeled high tension when the user requests high tension music. Thus, machine learning can be based on a variety of information, including combined output, feedback information, user control input, and so on.
[0353] In the illustrated embodiment, the combination module 2120 includes a section sequencer 2122, a section arranger 2124, a technique implementation module 2126, and a loop selection module 2128. In some embodiments, the combination module 2120 organizes and constructs the sections of the combination based on loop metadata and user control input (e.g., mood control).
[0354] In some embodiments, the section sequencer 2122 sequences sections of different types. In some embodiments, the section sequencer 2122 implements a finite state machine to sequentially output sections of the next type during operation. For example, the combination module 2120 can be configured to use sections of different types, such as intros, rises, falls, breakdowns, and bridges, as described below with reference to Figure 23 Further discussed in detail. In addition, each section may include multiple subsections that define how the music changes throughout the section, for example, including a transition subsection, a main content subsection, and a transition subsection.
[0355] In some embodiments, the part arranger 2124 constructs subparts according to arrangement rules. For example, one rule can specify to transfer in by gradually adding tracks. Another rule can specify to transfer in by gradually increasing the gain of a group of tracks. Another rule can specify to create a melody by syncopating vocal loops. In some embodiments, the probability that a loop in the loop library is attached to a track is a function of the current position in the part or subpart, the loops that overlap in time on another track, and user input parameters such as emotional variables (which can be used to determine the target attributes of the generated music content). For example, the function can be adjusted by adjusting the coefficient based on machine learning.
[0356] In some embodiments, the technique implementation module 2120 is configured to facilitate the arrangement of sections by adding rules, such as those specified by the artist or determined by analyzing the work of a particular artist. A "technique" can describe how a particular artist implements the arrangement rule at the technical level. For example, for an arrangement rule that specifies transitioning in by gradually adding tracks, one technique can indicate adding tracks in the order of drums, bass, pads, and then vocals, while another technique can indicate adding tracks in the order of bass, pads, vocals, and then drums. Similarly, for an arrangement rule that specifies looping a vocal to create a melody, one technique can indicate looping the vocal on every second beat and repeating the looped loop twice before moving on to the next looped section.
[0357] In the illustrated embodiment, loop selection module 2128 selects loops for inclusion in a segment of segment arranger 2124 based on permutation rules and techniques. Once a segment is complete, a corresponding performance script can be generated and sent to performance module 2130. Performance module 2130 can receive performance script segments at various granularity levels. This can include, for example, the entire performance script for a particular length of performance, a performance script for each segment, a performance script for each sub-segment, and so on. In some embodiments, the permutation rules, techniques, or loop selection are statistically implemented, e.g., different methods use different percentages of the time.
[0358] In the illustrated embodiment, the performance module 2130 includes a filter module 2131, an effects module 2132, a mixing module 2133, a master module 2134, and an execution module 2135. In some embodiments, these modules process a performance script and generate music data in a format supported by the audio output device 2140. The performance script may specify which loops to play, when they should play, what effects module 2132 should apply (e.g., on a per-track or per-subpart basis), what filters module 2131 should apply, etc.
[0359] For example, a performance script may specify that a low pass filter from 1000 to 20000 Hz be applied to a particular track for a duration of 1000 to 5000 milliseconds. As another example, a performance script may specify that a reverb with a 0.2 wet setting be applied to a particular track for a duration of 5000 to 15000 milliseconds.
[0360] In some embodiments, the mixing module 2133 is configured to perform automatic level control on the combined tracks. In some embodiments, the mixing module 2133 uses frequency domain analysis of the combined tracks to measure frequencies with too much or too little energy and applies gain to the tracks in different frequency bands to achieve a uniform mix. In some embodiments, the main module 2134 is configured to perform multi-band compression, equalization (EQ), or limiting procedures to generate data for final formatting by the execution module 2135. Figure 21 Embodiments may automatically generate various output music content based on user input or other feedback information, while machine learning techniques may allow for improvements in user experience over time.
[0361] Figure 22 is a diagram illustrating an example enhancement portion of music content according to some embodiments. Figure 21 The system can compose such a part by applying arrangement rules and techniques. In the example shown, the enhancement part includes three sub-parts and separate tracks for vocals, pads, drums, bass and white noise.
[0362] In the example shown, the transitions in the subsections include drum loop A, which is also repeated for the main content subsection. The transitions in the subsections also include bass loop A. As shown, the gain of the section starts low and increases linearly throughout the section (although nonlinear increases or decreases are contemplated). In the example shown, the main content and rollout subsections include various vocals, pads, drums, and bass loops. As described above, the disclosed techniques for automatically sequencing sections, arranging sections, and implementing techniques can generate a nearly unlimited stream of output music content based on a variety of user-adjustable parameters.
[0363] In some embodiments, the computer system displays a display similar to Figure 22 The interface allows the artist to specify the technology used for the component. For example, the artist can create a structure such as Figure 22 As shown, the code can be parsed into a combined module.
[0364] Figure 23 23 is a diagram illustrating an example technique for arranging parts of music content according to some embodiments. In the illustrated embodiment, the generated streaming media 2310 includes a plurality of parts 2320, each of which includes a start sub-part 2322, a development sub-part 2324, and a transition sub-part 2326. In the illustrated example, the various types of each part / sub-part are shown in a table connected by dotted lines. In the illustrated embodiment, the circular element is an example of an arrangement tool, which can be further implemented using the specific technology described below. As shown, various combination decisions can be performed pseudo-randomly according to statistical percentages. For example, the type of sub-part, the arrangement tool of a specific type or sub-part, or the technology for implementing the arrangement tool can be statistically determined.
[0365] In the example shown, a given section 2320 is one of five types: intro, buildup, decline, breakdown, and bridge, each with a different function for controlling the intensity of the section. In this example, the state subsection is one of three types: slow build, sudden transition, or minimal, each with a different behavior. In this example, the development subsection is one of three types: decrease, transition, or increase. In this example, the transition subsection is one of three types: collapse, gradient, or hint. For example, the different types of sections and subsections can be selected based on rules or pseudo-randomly.
[0366] In the examples shown, the behavior of the different subsection types is implemented using one or more arrangement tools. For the slow build, in this example, a low pass filter is applied 40% of the time and layers are added 80% of the time. For the transform development subsection, in this example, the loop is sliced 25% of the time. Various additional arrangement tools are shown, including one-shots, differential beats, applying reverb, adding pads, adding themes, removing layers, and white noise. These examples are included for illustrative purposes and are not intended to limit the scope of this disclosure. Furthermore, for ease of illustration, these examples may not be complete (e.g., actual arrangements may typically involve a greater number of arrangement rules).
[0367] In some embodiments, one or more arrangement tools may be implemented using specific techniques (which may be artist-specified or determined based on analysis of the artist's content). For example, a single shot may be implemented using sound effects or vocals, loop segmentation may be implemented using stuttering or half-cutting techniques, layer removal may be implemented by removing synthesizers or removing vocals, white noise may be implemented using a gradient or pulse function, etc. In some embodiments, the specific technique selected for a given arrangement tool may be selected based on a statistical function (e.g., layer removal may remove synthesizers 30% of the time, and vocals may be removed 70% of the time for a given artist). As described above, arrangement rules or techniques may be automatically determined by analyzing existing works, such as using machine learning.
[0368] Example Method
[0369] Figure 24 is a flowchart method for using a ledger according to some embodiments. Figure 24 The methods shown can be used in conjunction with any computer circuit, system, device, element, or assembly disclosed herein. In various embodiments, some of the method elements shown can be performed simultaneously in a different order than shown, or can be omitted. Additional method elements can also be performed as needed.
[0370] At 2410, in the illustrated embodiment, the computing device determines playback data indicating playback characteristics of a music content mix. The mix may include a determined combination of multiple tracks (note that the combination of tracks may be determined in real time, e.g., just prior to outputting the current portion of the music content mix, which may be a continuous stream of content). This determination may be based on a composition of the content mix (e.g., by a server or playback device such as a mobile phone) or may be received from another device that determines which audio files are included in the mix. The playback data may be stored (e.g., in an offline mode) and may be encrypted. The playback data may be reported periodically or in response to a specific event (e.g., re-establishing connection to the server).
[0371] At 2420, in the illustrated embodiment, the computing device records information specifying individual playback data for one or more of the plurality of tracks in the music content mix in an electronic blockchain ledger data structure. In the illustrated embodiment, the information specifying the individual playback data for the individual tracks includes usage data for the individual tracks and signature information associated with the individual tracks.
[0372] In some embodiments, the signature information is an identifier of one or more entities. For example, the signature information can be a string of characters or a unique identifier. In other embodiments, the signature information can be encrypted or otherwise obfuscated to prevent others from identifying the entity. In some embodiments, the usage data includes at least one of the following: the play time of the music content mix or the number of times the music content mix was played.
[0373] In some embodiments, data identifying individual tracks in a music content mix is retrieved from a data store that also indicates operations to be performed in association with the inclusion of one or more individual tracks. In these embodiments, the record may include an indication of proof of performance of the recorded indicated operations.
[0374] In some embodiments, the system determines compensation for multiple entities associated with multiple audio tracks based on information specifying individual playback data recorded in an electronic blockchain ledger.
[0375] In some embodiments, the system determines usage data for a first individual track that is not included in a mix of music content in its original musical form. For example, the track may be modified, used to generate a new track, etc., and the usage data may be adjusted to reflect the modification or usage. In some embodiments, the system generates the new track based on interpolating between vector representations of audio in at least two of the plurality of tracks, and the usage data is determined based on a distance between the vector representation of the first individual track and the vector representation of the new track. In some embodiments, the usage data is based on a ratio of the Euclidean distance to the interpolated vector representation to the vector representation of the at least two of the plurality of tracks.
[0376] Figure 25 is a flowchart method for combining audio files using image representations according to some embodiments. Figure 25 The methods shown can be used in conjunction with any computer circuit, system, device, element, or assembly disclosed herein. In various embodiments, some of the method elements shown can be performed simultaneously in a different order than shown, or can be omitted. Additional method elements can also be performed as needed.
[0377] At 2510, in the illustrated embodiment, a computing device generates a plurality of image representations of a plurality of audio files, wherein the image representations of the specified audio files are generated based on data in the specified audio files and a MIDI representation of the specified audio files. In some embodiments, pixel values in the image representations represent tempos in the audio files, wherein the image representations are compressed at a tempo resolution.
[0378] In some embodiments, the image representation is a two-dimensional representation of the audio file. In some embodiments, pitch is represented by rows in the two-dimensional representation, time is represented by columns in the two-dimensional representation, and pixel values in the two-dimensional representation represent tempo. In some embodiments, pitch is represented by rows in the two-dimensional representation, time is represented by columns in the two-dimensional representation, and pixel values in the two-dimensional representation represent tempo. In some embodiments, the pitch axis is scaled into two groups of octaves within an 8-octave range, where the first 12 rows of pixels represent the first 4 octaves, with the pixel values of the pixels determining which of the first 4 octaves is represented, and the second 12 rows of pixels represent the second 4 octaves, with the pixel values of the pixels determining which of the second 4 octaves is represented. In some embodiments, odd-numbered pixel values along the time axis represent note onsets, and even-numbered pixel values along the time axis represent note durations. In some embodiments, each pixel represents a portion of a beat in the time dimension.
[0379] At 2520 , in the illustrated embodiment, the computing device selects a plurality of audio files based on the plurality of image representations.
[0380] At 2530 , in the illustrated embodiment, the computing device combines the plurality of audio files to generate output music content.
[0381] In some embodiments, one or more composition rules are applied to select a plurality of audio files based on the plurality of image representations. In some embodiments, applying the one or more composition rules includes removing pixel values in the image representations that are above a first threshold and removing pixel values in the image representations that are below a second threshold.
[0382] In some embodiments, one or more machine learning algorithms are applied to the image representation for selecting and combining multiple audio files and generating output music content.In some embodiments, harmony and rhythmic coherence are tested in the output music content.
[0383] In some embodiments, a single image representation is generated from the plurality of image representations and a description of the texture features is appended to the single image representation from which the texture features are extracted from the plurality of audio files. In some embodiments, the single image representation is stored with the plurality of audio files. In some embodiments, the plurality of audio files are selected by applying one or more authoring rules to the single image representation.
[0384] Figure 26 is a flowchart method for implementing a user-created control element according to some embodiments. Figure 26 The methods shown can be used in conjunction with any computer circuit, system, device, element, or assembly disclosed herein. In various embodiments, some of the method elements shown can be performed simultaneously in a different order than shown, or can be omitted. Additional method elements can also be performed as needed.
[0385] In the illustrated embodiment, a computing device accesses a plurality of audio files at 2610. In some embodiments, the audio files are accessed from a memory of the computer system, where a user has rights to the accessed audio files.
[0386] At 2620, in the illustrated embodiment, the computing device generates output music content by combining music content from two or more audio files using at least one trained machine learning algorithm. In some embodiments, the combination of the music content is determined by the at least one trained machine learning algorithm based on the music content within the two or more audio files. In some embodiments, the at least one trained machine learning algorithm combines the music content by sequentially selecting music content from the two or more audio files based on the music content within the two or more audio files.
[0387] In some embodiments, the at least one trained machine learning algorithm has been trained to select music content for an upcoming beat after a specified time based on metadata of the music content played up to the specified time. In some embodiments, the at least one trained machine learning algorithm has been further trained to select music content for an upcoming beat after the specified time based on a level of a control element.
[0388] In the illustrated embodiment, at 2630, the computing device implements a control element created by a user on a user interface for changing a user-specified parameter in the generated output music content, wherein the level of one or more audio parameters in the generated output music content is determined based on the level of the control element, and wherein the relationship between the level of the one or more audio parameters and the level of the control element is based on user input during at least one music playback session. In some embodiments, the level of the user-specified parameter varies based on one or more environmental conditions.
[0389] In some embodiments, the relationship between the levels of one or more audio parameters and the level of a control element is determined by: playing multiple audio tracks during at least one music playback session, wherein the multiple audio tracks have varying audio parameters; for each audio track, receiving input of a user-selected level of a user-specified parameter in the specified audio track; evaluating the levels of one or more audio parameters in the audio track for each audio track; and determining the relationship between the levels of the one or more audio parameters and the level of the control element based on a correlation between each user-selected level of the user-specified parameter and each evaluated level of the one or more audio parameters.
[0390] In some embodiments, one or more machine learning algorithms are used to determine a relationship between the levels of one or more audio parameters and the levels of control elements. In some embodiments, the relationship between the levels of the one or more audio parameters and the levels of control elements is refined based on user changes to the levels of the control elements during playback of the generated output music content. In some embodiments, metadata from the audio track is used to evaluate the levels of the one or more audio parameters in the audio track. In some embodiments, the relationship between the levels of the one or more audio parameters and the levels of the user-specified parameters is further based on additional user input during one or more additional music playback sessions.
[0391] In some embodiments, the computing device implements at least one additional control element created by a user on a user interface for changing an additional user-specified parameter in the generated output musical content, wherein the additional user-specified parameter is a subparameter of the user-specified parameter. In some embodiments, the generated output musical content is modified based on the user's adjustment of the level of the control element. In some embodiments, a feedback control element is implemented on the user interface, wherein the feedback control element allows the user to provide positive or negative feedback on the generated output musical content during playback. In some embodiments, at least one trained machine algorithm modifies the generation of subsequently generated output musical content based on the feedback received during playback.
[0392] Figure 27 is a flowchart method for generating music content by modifying audio parameters according to some embodiments. Figure 27The methods shown can be used in conjunction with any computer circuit, system, device, element, or assembly disclosed herein. In various embodiments, some of the method elements shown can be performed simultaneously in a different order than shown, or can be omitted. Additional method elements can also be performed as needed.
[0393] At 2710, in the illustrated embodiment, a computing device accesses a set of music content. In some embodiments.
[0394] At 2720 , in the illustrated embodiment, the computing device generates a first graph of an audio signal of the music content, where the first graph is a graph of an audio parameter with respect to time.
[0395] In 2730, in the illustrated embodiment, the computing device generates a second graph of the audio signal of the music content, wherein the second graph is a signal graph of audio parameters relative to tempo. In some embodiments, the second graph of the audio signal has a similar structure to the first graph of the audio signal.
[0396] At 2740 , in the illustrated embodiment, the computing device generates new music content from the played back music content by modifying audio parameters in the played back music content, wherein the audio parameters are modified based on a combination of the first graph and the second graph.
[0397] In some embodiments, the audio parameters in the first and second graphs are defined by nodes in the graphs that determine changes in properties of the audio signal. In some embodiments, generating new music content includes: receiving playback music content; determining a first node in the first graph that corresponds to an audio signal in the playback music content; determining a second node in the second graph that corresponds to the first node; determining one or more specified audio parameters based on the second node; and modifying one or more properties of the audio signal in the playback music content by modifying the specified audio parameters. In some embodiments, one or more additional specified audio parameters are determined based on the first node, and one or more properties of the additional audio signal in the playback music content are modified by modifying the additional specified audio parameters.
[0398] In some embodiments, determining the one or more audio parameters includes determining a portion of the second graph to implement the audio parameters based on a position of the second node in the second graph, and selecting audio parameters from the determined portion of the second graph as the one or more audio-designated parameters. In some embodiments, modifying the one or more designated audio parameters modifies a portion of the played back music content corresponding to the determined portion of the second graph. In some embodiments, the modified attribute of the audio signal in the played back music content includes signal amplitude, signal frequency, or a combination thereof.
[0399] In some embodiments, one or more automations are applied to the audio parameters, wherein at least one of the automations is a pre-programmed temporal manipulation of the at least one audio parameter. In some embodiments, one or more modulations are applied to the audio parameters, wherein at least one of the modulations multiply modifies the at least one audio parameter over the at least one automation.
[0400] The following numbered clauses list various non-limiting embodiments disclosed herein:
[0401] Set A
[0402] A1. A method comprising:
[0403] The computer system generates a plurality of image representations of the plurality of audio files, wherein the image representation of the specified audio file is generated based on data in the specified audio file and a MIDI representation of the specified audio file;
[0404] selecting a plurality of audio files based on the plurality of image representations; and
[0405] Combine multiple audio files to generate output music content.
[0406] A2. The method of any preceding clause in Set A, wherein pixel values in the image representation represent a tempo in the audio file, and wherein the image representation is compressed at a tempo resolution.
[0407] A3. A method according to any preceding clause in Set A, wherein the image representation is a two-dimensional representation of the audio file.
[0408] A4. A method according to any preceding clause in set A, wherein pitch is represented by rows in the two-dimensional representation, wherein time is represented by columns in the two-dimensional representation, and wherein pixel values in the two-dimensional representation represent velocity.
[0409] A5. A method according to any preceding clause in Set A, wherein the two-dimensional representation is 32 pixels wide by 24 pixels high, and wherein each pixel represents a portion of a beat in the time dimension.
[0410] A6. A method according to any of the preceding clauses in set A, wherein the pitch axis is brought into two sets of octaves within an 8-octave range, wherein the first 12 rows of pixels represent the first 4 octaves, wherein the pixel values of the pixels determine which of the first 4 octaves is represented, and wherein the second 12 rows of pixels represent the second 4 octaves, wherein the pixel values of the pixels determine which of the second 4 octaves is represented.
[0411] A7. A method according to any preceding clause in Set A, wherein odd pixel values along the time axis represent note onsets and even pixel values along the time axis represent note durations.
[0412] A8. The method of any preceding clause in Set A, further comprising applying one or more authoring rules to select a plurality of audio files based on the plurality of image representations.
[0413] A9. A method according to any preceding clause in set A, wherein applying one or more authoring rules comprises removing pixel values in the image representation above a first threshold and removing pixel values in the image representation below a second threshold.
[0414] A10. A method according to any preceding clause in Set A, further comprising applying one or more machine learning algorithms to the image representation for selecting and combining multiple audio files and generating output music content.
[0415] A11. A method according to any preceding clause in Set A, further comprising testing harmonic and rhythmic coherence in the output musical content.
[0416] A12. A method according to any of the preceding clauses in Set A, further comprising:
[0417] generating a single image representation from multiple image representations;
[0418] extracting one or more texture features from the plurality of audio files; and
[0419] The descriptions of the extracted texture features are appended to a single image representation.
[0420] A13. A non-transitory computer-readable medium having instructions stored thereon, the instructions being executable by a computing device to perform operations comprising:
[0421] Any combination of operations performed by the methods of any of the preceding clauses in Set A.
[0422] A14. An apparatus comprising:
[0423] one or more processors; and
[0424] One or more memories having program instructions stored thereon, the program instructions being executable by the one or more processors to:
[0425] Perform any combination of the operations performed by the methods according to any of the preceding clauses in Set A.
[0426] Set B
[0427] B1. A method comprising:
[0428] A computer system accesses a set of music content;
[0429] The computer system generates a first map of an audio signal of the musical content, wherein the first map is a map of an audio parameter with respect to time;
[0430] The computer system generates a second graph of the audio signal of the musical content, wherein the second graph is a graph of the signal of the audio parameter relative to tempo; and
[0431] The computer system generates new music content from the played back music content by modifying audio parameters in the played back music content, wherein the audio parameters are modified based on a combination of the first map and the second map.
[0432] B2. The method of any preceding clause in Set B, wherein the second map of the audio signal has a similar structure as the first map of the audio signal.
[0433] B3. A method according to any preceding clause in Set B, wherein the audio parameters in the first and second graphs are defined by nodes in the graphs that determine changes in properties of the audio signal.
[0434] B4. The method according to any preceding clause in Set B, wherein generating the new music content comprises:
[0435] receiving and playing back music content;
[0436] Determining a first node in the first graph corresponding to an audio signal in the played back music content;
[0437] determining a second node in the second graph corresponding to the first node;
[0438] determining one or more specified audio parameters based on the second node; and
[0439] Modify one or more properties of an audio signal in the playback music content by modifying specified audio parameters.
[0440] B5. A method according to any of the preceding clauses in Set B, further comprising:
[0441] determining one or more additional specified audio parameters based on the first node; and
[0442] One or more properties of the additional audio signal in the played back music content are modified by modifying the additional specified audio parameters.
[0443] B6. The method of any preceding clause in Set B, wherein determining one or more audio parameters comprises:
[0444] determining, based on the position of the second node in the second graph, a portion of the second graph to implement the audio parameter; and
[0445] An audio parameter is selected from the determination portion of the second diagram as the one or more audio-specific parameters.
[0446] B7. A method according to any preceding clause in set B, wherein modifying one or more specified audio parameters modifies a portion of the played back music content corresponding to a determined portion of the second graph.
[0447] B8. A method according to any preceding clause in Set B, wherein the modified property of the audio signal in the playback music content comprises signal amplitude, signal frequency, or a combination thereof.
[0448] B9. The method of any preceding clause in Set B, further comprising applying one or more automations to audio parameters, wherein at least one of the automations is a pre-programmed temporal manipulation of at least one audio parameter.
[0449] B10. The method of any preceding clause in Set B, further comprising applying one or more modulations to the audio parameters, wherein at least one of the modulations modifies the at least one audio parameter multiplicatively over the at least one automation.
[0450] B11. The method according to any of the preceding clauses in set B further includes the computer system providing a heap-allocated memory for storing one or more objects associated with the group of music content, the one or more objects comprising at least one of the following: an audio signal, a first image, and a second image.
[0451] B12. The method of any preceding clause in Set B, further comprising storing the one or more objects in a data structure list in heap-allocated memory, wherein the data structure list comprises a serialized list of links for the objects.
[0452] B13. A non-transitory computer-readable medium having instructions stored thereon, the instructions being executable by a computing device to perform operations comprising:
[0453] Any combination of operations performed by the methods of any of the preceding clauses in Set B.
[0454] B14. An apparatus comprising:
[0455] one or more processors; and
[0456] One or more memories having program instructions stored thereon, the program instructions being executable by the one or more processors to:
[0457] Perform any combination of the operations performed by the methods according to any of the preceding clauses in Set B.
[0458] Set C
[0459] C1. A method comprising:
[0460] A computer system accesses multiple audio files;
[0461] generating output music content by combining music content from two or more audio files using at least one trained machine learning algorithm; and
[0462] A control element created by a user for changing a user-specified parameter in generated output music content is implemented on a user interface associated with a computer system, wherein the level of one or more audio parameters in the generated output music content is determined based on the level of the control element, and wherein the relationship between the level of the one or more audio parameters and the level of the control element is based on user input during at least one music playback session.
[0463] C2. A method according to any preceding clause in Set C, wherein the combination of musical content is determined by at least one trained machine learning algorithm based on the musical content within two or more audio files.
[0464] C3. A method according to any preceding clause in Set C, wherein at least one trained machine learning algorithm combines the musical content by sequentially selecting musical content from two or more audio files based on the musical content within the two or more audio files.
[0465] C4. A method according to any preceding clause in set C, wherein at least one trained machine learning algorithm has been trained to select music content for an upcoming beat after a specified time based on metadata of the music content played to the specified time.
[0466] C5. A method according to any preceding clause in set C, wherein at least one trained machine learning algorithm has been further trained to select musical content for an upcoming beat after a specified time based on the level of the control element.
[0467] C6. A method according to any preceding clause in Set C, wherein the relationship between the level of the one or more audio parameters and the level of the control element is determined by:
[0468] playing a plurality of audio tracks during at least one music playback session, wherein the plurality of audio tracks have varying audio parameters;
[0469] For each audio track, receiving input of a user-selected level of a user-specified parameter in the specified audio track;
[0470] evaluating, for each audio track, a level of one or more audio parameters in the audio track; and
[0471] Based on the correlation between each user-selected level of the user-specified parameter and each evaluated level of the one or more audio parameters, a relationship between the levels of the one or more audio parameters and the level of the control element is determined.
[0472] C7. A method according to any preceding clause in set C, wherein one or more machine learning algorithms are used to determine the relationship between the level of one or more audio parameters and the level of the control element.
[0473] C8. A method according to any preceding clause in set C, further comprising refining the relationship between the levels of one or more audio parameters and the levels of control elements based on user changes in the levels of the control elements during playback of the generated output music content.
[0474] C9. A method according to any preceding clause in set C, wherein the level of one or more audio parameters in the audio track is evaluated using metadata from the audio track.
[0475] C10. The method of any preceding clause in Set C, further comprising the computer system changing the level of the user-specified parameter based on one or more environmental conditions.
[0476] C11. The method according to any of the preceding clauses in set C further includes implementing at least one additional control element created by a user on a user interface associated with the computer system for changing additional user-specified parameters in the generated output music content, wherein the additional user-specified parameters are sub-parameters of the user-specified parameters.
[0477] C12. The method of any preceding clause in Set C, further comprising accessing an audio file from a memory of the computer system, wherein the user has rights to the accessed audio file.
[0478] C13. A non-transitory computer-readable medium having instructions stored thereon, the instructions executable by a computing device to perform operations comprising:
[0479] Any combination of operations performed by the methods of any of the preceding clauses in Set C.
[0480] C14. An apparatus comprising:
[0481] one or more processors; and
[0482] One or more memories having program instructions stored thereon, the program instructions being executable by the one or more processors to:
[0483] Perform any combination of the operations performed by the methods according to any of the preceding clauses in Set C.
[0484] Set D
[0485] D1. A method comprising:
[0486] The computer system determines playback data for a music content mix, wherein the playback data indicates playback characteristics of the music content mix, wherein the music content mix includes a determined combination of a plurality of audio tracks; and
[0487] The computer system records information specifying individual playback data for one or more of a plurality of tracks in a music content mix in an electronic blockchain ledger data structure, wherein the information specifying the individual playback data for the individual tracks includes usage data for the individual tracks and signature information associated with the individual tracks.
[0488] D2. A method according to any preceding clause in set D, wherein the usage data comprises at least one of: a play time of the music content mix or a number of play times of the music content mix.
[0489] D3. A method according to any preceding clause in set D, further comprising determining information specifying individual playback data of the individual tracks based on the playback data of the music content mix and data identifying the individual tracks in the music content mix.
[0490] D4. A method according to any preceding clause in set D, wherein data identifying individual tracks in a mix of musical content is retrieved from a data store, the data store also indicating operations to be performed associated with the inclusion of one or more individual tracks, wherein the recording includes an indication of recording proof of performance of the indicated operations.
[0491] D5. A method according to any preceding clause in set D, further comprising identifying one or more entities associated with an individual audio track based on the signature information and data specifying the signature information of a plurality of entities associated with the plurality of audio tracks.
[0492] D6. A method according to any of the preceding clauses in Set D, further comprising determining compensation to multiple entities associated with multiple audio tracks based on information specifying individual playback data recorded in the electronic blockchain ledger.
[0493] D7. A method according to any preceding clause in set D, wherein the playback data includes usage data of one or more machine learning modules for generating the music mix.
[0494] D8. A method according to any preceding clause in set D, further comprising the playback device at least temporarily storing playback data of the music content mix.
[0495] D9. A method according to any preceding clause in set D, wherein the stored playback data is transmitted by the playback device to another computer system periodically or in response to an event.
[0496] D10. A method according to any of the preceding clauses in set D, further comprising the playback device encrypting playback data of the music content mix.
[0497] D11. A method according to any of the preceding clauses in Set D, wherein the blockchain ledger is publicly accessible and immutable.
[0498] D12. A method according to any of the preceding clauses in Set D, further comprising:
[0499] Usage data of a first individual audio track not included in the combination of music contents in its original musical form is determined.
[0500] D13. A method according to any of the preceding clauses in Set D, further comprising:
[0501] generating a new audio track based on interpolating between vector representations of audio in at least two of the plurality of audio tracks;
[0502] The determining of the usage data is based on a distance between a vector representation of the first individual audio track and a vector representation of the new audio track.
[0503] D14. A non-transitory computer-readable medium having instructions stored thereon, the instructions being executable by a computing device to perform operations comprising:
[0504] Any combination of operations performed according to the methods of any of the preceding clauses in Set D.
[0505] D15. A device comprising:
[0506] one or more processors; and
[0507] One or more memories having program instructions stored thereon, the program instructions being executable by the one or more processors to:
[0508] Perform any combination of the operations performed by the methods according to any of the preceding clauses in Set D.
[0509] Set E
[0510] E1. A method comprising:
[0511] A computing system stores data specifying a plurality of tracks of a plurality of different entities, wherein the data includes signature information for respective tracks of the plurality of music tracks;
[0512] The computing system layers the plurality of tracks to generate output music content;
[0513] Information of the tracks included in a specified layer and individual signature information of those tracks are recorded in an electronic blockchain ledger data structure.
[0514] E2. A method according to any clause in Set E, wherein the blockchain ledger is publicly accessible and immutable.
[0515] E3. A method according to any of the clauses in Set E, further comprising:
[0516] Detects use of one of the tracks by an entity that does not match the signature information.
[0517] E4. A method comprising:
[0518] A computing system stores data specifying a plurality of tracks for a plurality of different entities, wherein the data includes metadata linking to other content;
[0519] A computing system selects and layers the plurality of tracks to generate output music content;
[0520] Content to be output together with the output music content is retrieved based on metadata of the selected track.
[0521] E5. A method according to any clause in Set E, wherein the other content includes visual advertising.
[0522] E6. A method comprising:
[0523] The computing system stores data specifying a plurality of tracks;
[0524] The computing system causes output of a plurality of music examples;
[0525] The computing system receives user input indicating a user opinion regarding whether one or more of the music examples exhibit a first music parameter;
[0526] Based on receiving and storing user-defined rule information;
[0527] The computing system receives user input indicating an adjustment to a first musical parameter;
[0528] The computing system selects and layers the plurality of tracks based on the custom rule information to generate output music content.
[0529] E7. A method according to any clause in set E, wherein the user specifies a name for the first musical parameter.
[0530] E8. A method according to any of the clauses in set E, wherein the user specifies one or more target objectives for one or more ranges of the first musical parameter, and wherein the selection and stratification are based on past feedback information.
[0531] E9. The method according to any clause in set E further includes: in response to determining that the selection and stratification based on past feedback information does not meet one or more target objectives, adjusting multiple other musical attributes and determining the effect of adjusting the multiple other musical attributes on the one or more target objectives.
[0532] E10. A method according to any clause in set E, wherein the target objective is the user's heart rate.
[0533] E11. A method according to any of the clauses in set E, wherein the selecting and stratifying are performed by a machine learning engine.
[0534] E12. A method according to any of the clauses in set E, further comprising training a machine learning engine, comprising training a teacher model based on multiple types of user input and training a student model based on received user input.
[0535] E13. A method comprising:
[0536] The computing system stores data specifying a plurality of tracks;
[0537] The computing system determines that the audio device is located in a first type of environment;
[0538] receiving user input for adjusting a first music parameter while the audio device is in a first type of environment;
[0539] The computing system selects and layers the plurality of tracks based on the adjusted first music parameter and the first type of environment to generate output music content through the audio device;
[0540] The computing system determines that the audio device is located in a second type of environment;
[0541] receiving user input adjusting a first music parameter when the audio device is located in a second type of environment;
[0542] The computing system selects and layers a plurality of tracks based on the adjusted first music parameter and the second type of environment to generate output music content through the audio device.
[0543] E14. A method according to any clause in set E, wherein different musical attributes are selected and layered based on the adjustment of the first musical parameter when the audio device is in a first type of environment compared to when the audio device is in a second type of environment.
[0544] E15. A method comprising:
[0545] The computing system stores data specifying a plurality of tracks;
[0546] The computing system stores data specifying musical content;
[0547] determining one or more parameters of the stored musical content;
[0548] The computing system selects and layers the plurality of tracks based on one or more parameters of the stored music content to generate output music content; and
[0549] The output of the music content and the output of the stored music content are caused so that they overlap in time.
[0550] E16. A method comprising:
[0551] The computing system stores data specifying a plurality of tracks and corresponding track attributes;
[0552] The computing system stores data specifying musical content;
[0553] determining one or more parameters of the stored music content and determining one or more tracks included in the stored music content;
[0554] Based on the one or more parameters, the computing system selects and layers the stored tracks and the determined tracks to generate the output music content.
[0555] E17. A method comprising:
[0556] The computing system stores data specifying a plurality of tracks;
[0557] A computing system selects and layers the plurality of tracks to generate output music content;
[0558] Where selection is performed using both:
[0559] a neural network module configured to select tracks for future use based on recently used tracks; and
[0560] One or more hierarchical hidden Markov models are configured to select a structure for the music content to be output and constrain the neural network module based on the selected structure.
[0561] E18. A method according to any clause in Set E, further comprising:
[0562] One or more hierarchical hidden Markov models are trained based on positive or negative user feedback regarding the output music content.
[0563] E19. A method according to any clause in Set E, further comprising:
[0564] A neural network model is trained based on user selections provided in response to a plurality of music content samples.
[0565] E20. A method according to any clause in set E, wherein the neural network module comprises a plurality of layers, including at least one globally trained layer and at least one layer specifically trained based on feedback from a particular user account.
[0566] E21. A method comprising:
[0567] The computing system stores data specifying a plurality of tracks;
[0568] A computing system selects and layers multiple tracks to generate output music content, where the selection is performed using both a globally trained machine learning module and a locally trained machine learning module.
[0569] E22. A method according to any of the clauses in set E, further comprising training a globally trained machine learning module using feedback from multiple user accounts, and training a locally trained machine learning module based on feedback from a single user account.
[0570] E23. A method according to any of the clauses in set E, wherein both the globally trained machine learning module and the locally trained machine learning module provide output information based on user adjustments to the music attributes.
[0571] E24. A method according to any clause in set E, wherein the globally trained machine learning module and the locally trained machine learning module are included in different layers of the neural network.
[0572] E25. A method comprising:
[0573] receiving a first user input indicating a value of a higher-level music composition parameter;
[0574] automatically selecting and combining audio tracks based on the value of a higher-level music composition parameter, including using multiple values of one or more sub-parameters associated with the higher-level music composition parameter when outputting the musical content;
[0575] receiving a second user input indicating one or more values for the one or more sub-parameters; and
[0576] Constraints are automatically selected and combined based on second user input.
[0577] E26. A method according to any clause in set E, wherein the higher-level music composition parameter is an energy parameter, and wherein the one or more sub-parameters include speed, number of layers, vocal parameters, or bass parameters.
[0578] E27. A method comprising:
[0579] The computing system accesses the musical score information;
[0580] The computing system synthesizes musical content based on the musical score information;
[0581] analyzing the music content to generate frequency information for a plurality of frequency bins at different time points;
[0582] Training the machine learning engine involves inputting frequency information and using musical score information as training labels.
[0583] E28. A method according to any clause in Set E, further comprising:
[0584] Inputting musical content into a trained machine learning engine; and
[0585] A machine learning engine is used to generate musical score information for the music content, wherein the generated musical score information includes frequency information of multiple frequency bins at different time points.
[0586] E29. A method according to any clause in set E, wherein different time points have a fixed distance between them.
[0587] E30. A method according to any clause in set E, wherein the different time points correspond to beats of the musical content.
[0588] E31. A method according to any clause in set E, wherein the machine learning engine comprises convolutional layers and recurrent layers.
[0589] E32. A method according to any clause in set E, wherein the frequency information comprises a binary indication of whether the music content comprises content in the frequency bin at a point in time.
[0590] E33. A method according to any clause in set E, wherein the musical content comprises a plurality of different instruments.
[0591] Although specific embodiments have been described above, these embodiments are not intended to limit the scope of the present disclosure, even if only a single embodiment is described with respect to a particular feature. Unless otherwise stated, the feature examples provided in this disclosure are intended to be illustrative rather than restrictive. The above description is intended to cover such alternatives, modifications, and equivalents as would be apparent to one skilled in the art having the benefit of this disclosure.
[0592] The scope of the present disclosure includes any feature or combination of features disclosed herein (explicit or implicit), or any generalization thereof, whether or not it mitigates any or all of the problems addressed herein. Accordingly, new claims may be made during prosecution of the present application (or an application claiming priority thereto) to any such combination of features. In particular, features from dependent claims may be combined with features of the independent claims, with reference to the appended claims, and features from individual independent claims may be combined in any appropriate manner, and not merely in the specific combinations recited in the appended claims.
Claims
1. A method comprising: Accessing a set of music content by a computer system; generating, by the computer system, a first graph of an audio signal of the music content, wherein the first graph is a graph of an audio parameter with respect to time; generating, by the computer system, a second graph of the audio signal of the music content, wherein the second graph is a signal graph of the audio parameter relative to tempo; and New music content is generated from the played back music content by a computer system by modifying audio parameters in the played back music content, wherein the audio parameters are modified based on a combination of the first map and the second map.
2. The method according to claim 1, wherein The audio parameters in the first and second graphs are defined by nodes in the graphs that determine changes in properties of audio signals.
3. The method according to claim 2, wherein: Generating the new music content includes: receiving the playback music content; Determining a first node in the first graph corresponding to the audio signal in the played back music content; determining a second node in the second graph corresponding to the first node; determining one or more specified audio parameters based on the second node; and One or more properties of the audio signal in the played back music content are modified by modifying the specified audio parameters.
4. The method according to claim 3, further comprising: determining one or more additional specified audio parameters based on the first node; as well as One or more properties of the additional audio signal in the played back music content are modified by modifying the additional specified audio parameters.
5. The method according to claim 3, wherein Determining the one or more specified audio parameters includes: determining a portion of the second graph to implement the audio parameter based on a position of a second node in the second graph; and An audio parameter is selected from the determined portion of the second map as the one or more specified audio parameters.
6. The method according to claim 5, wherein: Modifying the one or more specified audio parameters modifies a portion of the played back music content corresponding to the determined portion of the second graph.
7. The method according to claim 3, wherein: The modified attribute of the audio signal in the played back music content includes signal amplitude, signal frequency or a combination thereof.
8. The method of claim 1 , further comprising applying one or more automations to the audio parameters, wherein At least one of the automations is a pre-programmed temporal manipulation of at least one of the audio parameters.
9. The method of claim 8, further comprising applying one or more modulations to the audio parameters, wherein At least one of the modulations modifies at least one of the audio parameters multiplicatively over at least one automation.
10. The method of claim 1, further comprising providing, by the computer system, heap-allocated memory for storing one or more objects associated with the set of music content, the one or more objects comprising at least one of: the audio signal, the first graph, and the second graph.
11. The method of claim 10, further comprising storing the one or more objects in a list of data structures in the heap-allocated memory, wherein The list of data structures includes a serialized list of links for the objects.
12. A non-transitory computer-readable medium having instructions stored thereon, the instructions being executable by a computing device to perform operations comprising: access a collection of music content; generating a first map of an audio signal of the music content, wherein the first map is a map of an audio parameter with respect to time; generating a second map of the audio signal of the music content, wherein the second map is a signal map of the audio parameter with respect to tempo; and New music content is generated from the played back music content by modifying audio parameters in the played back music content, wherein the audio parameters are modified based on a combination of the first map and the second map.
13. The non-transitory computer readable medium of claim 12, wherein: The audio parameters in the first graph and the second graph are defined by nodes in the graph that determine changes in properties of audio signals, and wherein generating the new music content comprises: receiving the playback music content; Determining a first node in the first graph corresponding to the audio signal in the played back music content; determining a second node in the second graph corresponding to the first node; determining one or more specified audio parameters based on the second node; and One or more properties of the audio signal in the played back music content are modified by modifying the specified audio parameters.
14. The non-transitory computer readable medium of claim 13, wherein: The modified attribute of the audio signal in the played back music content includes signal amplitude, signal frequency or a combination thereof.
15. The non-transitory computer-readable medium of claim 12, further comprising storing one or more objects associated with the set of music content in heap-allocated memory, wherein The one or more objects include at least one of: the audio signal, the first graph, and the second graph.
16. The non-transitory computer-readable medium of claim 12, further comprising adding an identifier to the new music content and storing the new music content in at least one circular buffer in a static array of circular buffers.
17. The non-transitory computer readable medium of claim 16, wherein: The static array of the circular buffer can be accessed by a single user.
18. An apparatus comprising: one or more processors; as well as One or more memories having stored thereon program instructions executable by the one or more processors for: access a collection of music content; generating a first map of an audio signal of the music content, wherein the first map is a map of an audio parameter with respect to time; generating a second map of the audio signal of the music content, wherein the second map is a signal map of the audio parameter with respect to tempo; and New music content is generated from the played back music content by modifying audio parameters in the played back music content, wherein the audio parameters are modified based on a combination of the first map and the second map.
19. The apparatus of claim 18, further comprising an application programming interface, wherein The application programming interface implements one or more automated applications of the audio parameters.
Citation Information
Patent Citations
Music generator
US10679596B2
Music generator
US20190362696A1
Music generator
US8812144B2