Dynamic Control of Generative Music Pieces
User interfaces for generative music systems allow real-time feedback and customization, addressing the challenges of user engagement and personalization of musical content.
Patent Information
- Application Number
- JP2025521007
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-14
- Filing Date
- 2023-10-13
- Publication Date
- 2025-10-28
AI Technical Summary
Existing generative music systems struggle to effectively convey musical aspects to users and determine their precise preferences, lacking user-accessible feedback mechanisms during real-time music generation.
Implement user interfaces that allow graphical viewing and real-time feedback integration of musical decisions, enabling users to influence generative music composition through interfaces that display internal engine processes and allow immediate adjustment of musical parameters.
Enables users to dynamically control and customize generative music by providing immediate feedback, enhancing user engagement and personalization of musical content.
Smart Images

Figure 2025535756000001_ABST
Abstract
Description
[Technical Field]
[0001] TECHNICAL FIELD This disclosure relates to audio engineering, and more particularly to computer-composed musical content. [Background technology]
[0002] A generative music system may use a computing system to compose musical content. For example, the AiMi platform allows a user to select a musical genre and listen to dynamically generated songs in that genre. The user may also provide feedback about the songs, and the system may adjust song parameters based on the user feedback.
[0003] In the context of generative music, it can be difficult to convey aspects of a piece of music to a user or to determine the precise preferences of a given user. [Brief explanation of the drawings]
[0004] [Figure 1] FIG. 1 illustrates an exemplary music generator module that generates musical content based on multiple different types of inputs, according to some embodiments.
[0005] [Figure 2] 10A-10C illustrate examples of interfaces showing multiple mixed musical phrases, according to some embodiments.
[0006] [Figure 3] 10A-10C illustrate examples of user resizing of a displayed musical phrase, according to some embodiments.
[0007] [Figure 4] 10A-10C illustrate examples of user musical phrase feedback and separate listening functionality for musical phrases, according to some embodiments.
[0008] [Figure 5] 10A-10C illustrate examples of user phrase feedback, according to some embodiments.
[0009] [Figure 6] FIG. 10 illustrates an example of user section feedback, according to some embodiments.
[0010] [Figure 7A] 10 is a screenshot illustrating an example interface, according to some embodiments. [Figure 7B] 10 is a screenshot illustrating an example interface, according to some embodiments. [Figure 7C] 10 is a screenshot illustrating an example interface, according to some embodiments. [Figure 8A] 10 is a screenshot illustrating an example interface, according to some embodiments. [Figure 8B] 10 is a screenshot illustrating an example interface, according to some embodiments. [Figure 8C] 10 is a screenshot illustrating an example interface, according to some embodiments. [Figure 9A] 10 is a screenshot illustrating an example interface, according to some embodiments. [Figure 9B] 10 is a screenshot illustrating an example interface, according to some embodiments. [Figure 9C] 10 is a screenshot illustrating an example interface, according to some embodiments. [Figure 10A] 10 is a screenshot illustrating an example interface, according to some embodiments. [Figure 10B] 10 is a screenshot illustrating an example interface, according to some embodiments.
[0011] [Figure 11A] FIG. 10 is a block diagram illustrating an example of an interface displaying currently mixed musical phrases within a generative music section.
[0012] [Figure 11B] 10A-10C illustrate examples of interfaces for modifying various characteristics of a selected musical phrase within a generative music section, according to some embodiments.
[0013] [Figure 11C] 10A-10C illustrate examples of interfaces for playing selected musical phrases within a generative music section, according to some embodiments.
[0014] [Figure 12] 1 is a flow diagram illustrating an example method, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0015] In general, compared to musical content that has already been recorded and resides in a library, generative music that is generated "on the fly" can be more difficult to allow users to control. Typical generative music software generates its music in real time, does not offer user-accessible options for providing feedback as the music is being generated, and traditional user interfaces are not particularly effective for user customization of generative music. On the other hand, while music players for music that resides statically in a library offer various features that allow music to be rated and evaluated, these features are not necessarily available (or even applicable) to dynamically created generative music.
[0016] The disclosed system may implement various techniques for immediately incorporating and incorporating user feedback into the generative music. User interfaces, described below, may allow a user to graphically view internal engine musical decisions (e.g., musical phrase selections, mix decisions, attributes, parameters, etc.), modify the results of these decisions, and send feedback to the system to be incorporated into the musical composition. In some cases, this feedback may adjust decisions in real time and be used for the musical composition on the fly. As described in more detail below, the disclosed user interface techniques may provide a user with a view of what is currently happening in the generative music composition while also providing the user with the ability to influence subsequent compositions.
[0017] Music Generator Example Overview Generally speaking, the disclosed music generator includes loop data, metadata (e.g., information describing the loops), and a grammar for combining loops based on the metadata. The generator may create a musical experience using rules that identify loops based on the metadata and target characteristics of the musical experience. The generator may be configured to expand the set of experiences that can be created by adding or modifying rules, loops, and / or metadata. Adjustments may be performed manually (e.g., an artist adds new metadata), or the music generator may enhance the rules / loops / metadata as it monitors the musical experience within a given environment and desired goals / characteristics. For example, if the music generator observes a crowd and sees people smiling, it can enhance its rules and / or metadata to note that a particular loop combination makes people smile. Similarly, if cash register sales increase, the rule generator can use that feedback to increase the rules / metadata for the associated loops that correlate with increased sales.
[0018] As used herein, the term "loop" refers to sound information of a single instrument over a particular time interval. While a loop may be played repeatedly (e.g., a 30-second loop may be played four times in succession to generate two minutes of musical content), a loop may also be played once, for example, without repeating. The various techniques described with respect to loops may also be implemented using audio files containing multiple instruments. The term "track" encompasses loops and audio files containing sounds from multiple instruments. Furthermore, a "track" may be a recorded or computer-generated musical phrase (e.g., synthesized entirely from scratch or generated by combining previously recorded or generated sounds). In general, the term "musical phrase" refers to information specifying a sequence of sounds over a time interval. A musical phrase may be a loop or a track and may include a single instrument or multiple instruments. It should be understood that the various techniques described with respect to one of a loop, a track, or a musical phrase may apply to all three or any of them in various embodiments.
[0019] 1 is a diagram illustrating an exemplary music generator according to some embodiments. In the illustrated embodiment, a music generator module 160 receives various information from multiple different sources and generates output music content 140.
[0020] In the illustrated embodiment, module 160 accesses stored loops and corresponding attributes 110 for the stored loops and combines the loops to generate output musical content 140. In particular, music generator module 160 selects loops based on their attributes and combines the loops based on target musical attributes 130 and / or environmental information 150. In some embodiments, environmental information is used indirectly to determine target musical attributes 130. In some embodiments, target musical attributes 130 are explicitly specified by a user, for example, by specifying a desired energy level, mood, multiple parameters, etc. Examples of target musical attributes 130 include, for example, energy, complexity, and variety, although more specific attributes (e.g., corresponding to attributes of stored tracks) may also be specified. Music attributes may be input by a user or may be determined based on environmental information such as ambient noise, lighting, etc. Generally speaking, when higher-level target musical attributes are specified, lower-level specific musical attributes may be determined by the system before generating the output musical content. Examples of techniques for generating musical content based on one or more musical attributes are described in U.S. patent application Ser. No. 13 / 969,372 (now U.S. Pat. No. 8,812,144), entitled "Music Generator," filed Aug. 16, 2013, and incorporated herein by reference in its entirety. The disclosure of Ser. No. 13 / 969,372 describes techniques such as selecting stored loops and / or tracks, generating new loops / tracks, and layering selected loops / tracks to generate output musical content. To the extent any interpretation is made based on a perceived conflict between any definition in the incorporated application and the remainder of this disclosure, the present disclosure is intended to control.
[0021] Complexity may refer to the number of loops and / or instruments included in a piece of music. Energy may be related to other attributes or may not affect other attributes. For example, changing the key or tempo may affect energy. However, for a given tempo and key, energy may be changed by adjusting the type of instrument (e.g., by adding hi-hats or white noise), complexity, volume, etc. Diversity may refer to the amount of musical variation produced over time. Diversity may be generated relative to a static set of other musical attributes (e.g., by selecting different tracks for a given tempo and key), or by changing musical attributes over time (e.g., by changing tempo and key more frequently when greater variety is desired). In some embodiments, the target musical attributes are considered to exist in a multi-dimensional space, and music generator module 160 may slowly move through that space, using course corrections as needed, for example, based on environmental changes and / or user input.
[0022] In some embodiments, the attributes stored with the loops include information about one or more loops including tempo, volume, energy, diversity, spectrum, envelope, modulation, periodicity, attack and decay times, noise, artist, instrument, theme, gain, etc. Note that in some embodiments, loops are divided so that a set of one or more loops is specific to a particular loop type (e.g., one instrument or one type of instrument).
[0023] In the illustrated embodiment, module 160 accesses a stored rule set 120. The stored rule set 120, in some embodiments, specifies rules such as how many loops to overlay to be played simultaneously (which may correspond to the complexity of the output music), which major / minor key progressions to use when transitioning between loops or musical phrases, which instruments to use together (e.g., instruments that have an affinity with each other), etc., to achieve target musical attributes. In other words, music generator module 160 uses stored rule sets 120 to achieve one or more declarative goals defined by target musical attributes (and / or target environmental information). In some embodiments, music generator module 160 includes one or more pseudo-random number generators configured to introduce pseudo-randomness to avoid repetitive output music.
[0024] In some embodiments, environmental information 150 includes one or more of lighting information, ambient noise, user information (e.g., facial expressions, body posture, activity level, movement, skin temperature, performance of a particular activity, type of clothing, etc.), temperature information, purchasing activity in a certain area, time of day, day of the week, season, number of people present, weather conditions, etc. In some embodiments, music generator module 160 does not receive / process environmental information. In some embodiments, environmental information 150 is received by another module that determines target musical attributes 130 based on the environmental information. Target musical attributes 130 may also be derived based on other types of content, e.g., video data. In some embodiments, environmental information is used to adjust one or more stored rule sets 120, for example, to achieve one or more environmental goals. Similarly, the music generator may use environmental information to adjust stored attributes for one or more loops, for example, to indicate target musical attributes or target audience characteristics to which those loops are particularly relevant.
[0025] As used herein, the term “module” refers to a physical, non-transitory, computer-readable medium that is configured to perform specified operations or that stores information (e.g., program instructions) that instruct other circuitry (e.g., processors) to perform specified operations. A module may be implemented in multiple ways, including as hardwired circuitry or as memory that stores program instructions executable by one or more processors to perform operations. Hardware circuitry may include, for example, custom very large-scale integrated circuit (VLSI) circuits or off-the-shelf semiconductors such as gate arrays, logic chips, transistors, or other discrete components. A module may also be implemented in programmable hardware devices such as field programmable gate arrays, programmable array logic, programmable logic devices, etc. A module may be any suitable form of non-transitory computer-readable medium that stores program instructions executable to perform specified operations.
[0026] As used herein, "musical content" refers to both the music itself (the audible representation of the music) and information that can be used to play the music. Thus, a song recorded as a file on a storage medium (such as, but not limited to, a compact disc, a flash drive, etc.) is an example of musical content, and the sound produced by playing this recorded file or other electronic representation (e.g., through a speaker) is also an example of musical content.
[0027] The term "music" includes its well-understood meaning, which includes sounds produced by instruments and vocal sounds. Thus, music includes, for example, instrumental performances or recordings, a cappella performances or recordings, and performances or recordings that include both instruments and voices. Those skilled in the art will recognize that "music" does not encompass all vocal recordings. For example, works that do not contain musical attributes such as rhythm or rhyme, such as speeches, news broadcasts, audiobooks, etc., are not music.
[0028] One musical "content" may be distinguished from another musical content in any suitable manner. For example, a digital file corresponding to a first song may represent the first musical content, and a digital file corresponding to a second song may represent the second musical content. The phrase "musical content" may be used to distinguish particular sections within a given musical work, such that different parts of the same song may be considered different parts of different musical content. Similarly, different tracks (e.g., piano track, guitar track) within a given musical work may correspond to different musical content. In the context of a potentially infinite stream of generated music, the phrase "musical content" may be used to refer to the same portion of the stream (e.g., a few bars or minutes).
[0029] Musical content generated by embodiments of the present disclosure may be "new musical content," i.e., a combination of musical elements that has not been previously generated. A related (but more expansive) concept, "original musical content," is described further below. To facilitate the explanation of this term, the concept of a "controlling entity" is described for an instance of musical content generation. Unlike the phrase "original musical content," the phrase "new musical content" does not refer to the concept of a controlling entity. Thus, new musical content refers to musical content that has not previously been generated by any entity or computing system.
[0030] Conceptually, this disclosure refers to some “entity” as controlling a particular instance of computer-generated musical content. Such an entity owns any legal rights (e.g., copyright) that may correspond to the computer-generated content (to the extent such rights may actually exist). In one embodiment, the controlling entity would be the individual who creates the computer-implemented music generator (e.g., codes the various software routines) or operates (e.g., provides inputs to) a particular instance of computer-implemented music generation. In other embodiments, a computer-implemented music generator may be created by a legal entity (e.g., a corporation or other business organization) in the form of a software product, computer system, computing device, or the like. In some instances, such a computer-implemented music generator may be deployed to many clients. Depending on the terms of the license associated with the distribution of this music generator, the controlling entity may, in various instances, be the creator, distributor, or client. In the absence of such an explicit legal agreement, the controlling entity for a computer-implemented music generator is the entity that facilitates (e.g., provides inputs to and operates on) a particular instance of computer-generated musical content.
[0031] Within the meaning of this disclosure, computer-generated "original musical content" by a controlling entity refers to 1) a combination of musical elements not previously generated by the controlling entity or anyone else, and 2) a combination of musical elements previously generated but generated in a first instance by the controlling entity. Content type 1) is referred to herein as "new musical content," similar to the definition of "new musical content," except that the definition of "new musical content" refers to the concept of a "controlling entity," whereas the definition of "new musical content" does not. Content type 2) is, on the other hand, referred to herein as "proprietary musical content." Note that the term "proprietary" in this context does not refer to any implied legal rights in the content (although such rights may exist), but is used merely to indicate that the musical content was originally generated by the controlling entity. Thus, a controlling entity's "regeneration" of musical content previously originally generated by the controlling entity constitutes "generation of original musical content" within this disclosure. "Non-original musical content" with respect to a particular controlling entity is musical content that is not that controlling entity's "original musical content."
[0032] Some musical content may include musical components from one or more other musical content. Creating musical content in this manner is called "sampling" the musical content and is common in certain musical works, especially in certain musical genres. Such musical content is referred to herein as "musical content with sampled components," "derivative musical content," or other similar terms. In contrast, musical content that does not include sampled components is referred to herein as "musical content without sampled components," "non-derivative musical content," or other similar terms.
[0033] In applying these terms, it is noted that if any particular musical content is reduced to a sufficient level of granularity, a claim may be made that the musical content is derivative (meaning, in effect, that all musical content is derivative). The terms "derivative" and "non-derivative" are not used in this sense in this disclosure. With respect to computer-generation of musical content, if the computer-generation selects some of the components from pre-existing musical content of an entity other than the controlling entity (e.g., if a computer program selects particular portions of an audio file of a popular artist's work for inclusion in the musical content being generated), such computer-generation is said to be derivative (resulting in derivative musical content). On the other hand, computer-generation of musical content is said to be non-derivative (resulting in non-derivative musical content) if the computer-generation does not utilize such components of such pre-existing content. Some "original musical content" may be derivative musical content, and some content may be non-derivative musical content.
[0034] It should be noted that the term "derivative" is intended in this disclosure to have a broader meaning than the term "derivative work" as used in U.S. copyright law. For example, derivative musical content may or may not be a derivative work under U.S. copyright. The term "derivative" in this disclosure is not intended to convey a negative connotation, but is merely used to imply whether particular musical content "borrows" a portion of content from another work.
[0035] Furthermore, the phrases “new musical content,” “novel musical content,” and “original musical content” are not intended to encompass musical content that differs only insignificantly from a combination of existing musical elements. For example, simply changing a few notes in an existing musical composition does not result in new, novel, or original musical content as those phrases are used in this disclosure. Similarly, simply changing the key or tempo of an existing musical composition or adjusting the relative intensities of frequencies (e.g., using an equalizer interface) does not create new, novel, or original musical content. Furthermore, the phrases “new, novel, and original musical content” are not intended to cover musical content that is borderline between original and non-original content; instead, these terms are intended to cover musical content that is undoubtedly and demonstrably original, including musical content that would be eligible for copyright protection from a controlling entity (referred to in the specification as “protectable” musical content). Furthermore, as used herein, the term “available” musical content refers to musical content that does not infringe the copyright of any entity other than the controlling entity. New and / or original music content is often protectable and available, which may be advantageous in preventing copying of the music content and / or paying royalties for the music content.
[0036] Although various embodiments described herein use a rule-based engine, various other types of computer-implemented algorithms may be used for any of the computer learning and / or music generation techniques described herein. Rule-based techniques may or may not include statistical and machine learning techniques. Exemplary techniques for generating musical content using statistical rules and machine learning engines are described in more detail in U.S. patent application Ser. No. 16 / 420,456 (now U.S. Pat. No. 10,679,596), filed May 23, 2019, entitled "Music Generator," which is incorporated herein by reference in its entirety. Further examples of techniques for machine learning training and processing audio data are described in U.S. patent application Ser. No. 17 / 174,052, filed February 11, 2021, entitled "Music Content Generation Using Image Representations of Audio Files," which is incorporated herein by reference in its entirety.
[0037] Example of a user interface with a bubble element 2 is a diagram illustrating an exemplary user interface for a generative music application, according to some embodiments. In the illustrated example, the interface includes a circular bubble corresponding to the current section of the generative music, which includes multiple smaller bubbles indicating the musical phrase currently being mixed. In this example, the musical phrase includes beats, vocals, effects (FX), effects 2, top, melody, and pads.
[0038] As the song plays, tracks may be removed from or inserted into the mix, and track volumes may be dynamically adjusted. In some embodiments, the illustrated interface elements reflect song changes in real time, for example, by adding / removing bubbles, changing the size of bubbles to reflect the current volume of a given track, etc. This may provide the user with quick, detailed information about the current song state.
[0039] As described in more detail below, the user may interact with the interface to adjust subsequent songs. User-initiated changes may begin occurring immediately in the current song, or may be delayed for certain types of input (e.g., to incorporate changes while continuously playing pleasant musical content). For example, changes to the volume of a musical phrase may be reflected immediately in the mix, and feedback regarding a given musical phrase or section may be reflected in subsequent song decisions (e.g., by adjusting the number of times a musical phrase is included, the amount of time it is played, etc.).
[0040] For purposes of illustration, circular bubbles are described herein, however, in other embodiments, a variety of other shapes may be displayed, including, but not limited to, oval, elliptical, polygonal, three-dimensional shapes, etc.
[0041] 3 is a diagram illustrating an example of a user adjusting bubble size, according to some embodiments. In the illustrated example, a user selects an FX bubble (dotted line 310 indicates its initial position) and inputs a change in the size of the bubble (e.g., by dragging a finger on a touchscreen or moving a cursor using a mouse), with the resized FX bubble shown using a thick line 320. In some embodiments, this immediately increases the volume of the FX musical phrase in the output mix.
[0042] Note that the example inputs of Figure 3 may also affect future compositions; for example, the music engine may mix that musical phrase (or similar musical phrases) in a subsequent mix using a higher volume than would have been used prior to the user feedback. Additionally, while the example inputs of Figure 3 may have an immediate effect, the music engine may dynamically adjust the volume of the adjusted musical phrase shortly thereafter (although the adjustment may be scaled based on user input). Thus, the user-specified volume may not be a static change, but may generally affect the composition in a desired direction.
[0043] 4 is a diagram illustrating example interface elements for a particular musical phrase or channel, according to some embodiments. In the illustrated example, a user has selected a vocal musical phrase. The interface provides a selectable "listen in isolation" element to play that musical phrase without other musical phrases being mixed in. This allows the user to determine whether the selected musical phrase is indeed the musical phrase they want to adjust.
[0044] The interface also provides a like element 402 and a dislike element 404 that allow the user to provide feedback regarding the musical phrase. A like may indicate positive feedback, and a dislike may indicate negative feedback. The music engine may use this feedback to adjust parameters associated with future mixes of that musical phrase, such as including it more or less frequently, at shorter or longer time intervals when included, at a lower or higher volume, combined more or less frequently with other musical phrases being mixed at the same time, or combined more or less frequently with other environmental information relative to the user at the time the feedback is provided. While thumbs up and thumbs down are included as an example, it is noted that various other feedback interfaces (e.g., star rating methods, number scales, up / down arrows, like / dislike, etc.) may be implemented to allow the user to provide a rating or otherwise respond to a given musical phrase. In some embodiments, a user's pressing of the like element 402 and the dislike element 404 modifies the user's profile information to indicate the user's musical tastes, for example, for use in subsequent songs.
[0045] Thus, in some embodiments, feedback related to musical phrases (or feedback related to loops / phrases used to construct musical phrases) is used to adjust future musical phrase selection. For example, in some embodiments, the system is configured to encode musical phrases into vectors and may use user feedback to select musical phrases with vectors similar to the current musical phrase more or less frequently. For example, the system may search within the vector space to find similar sound loops in response to a "like" for future inclusion. Similarly, the system may reduce the likelihood of playing loops (or prevent inclusion of loops altogether) that are within a specified distance in the vector space from the currently playing loop.
[0046] The computing system may create the vectors using a self-supervised neural network (which may also be called "unsupervised") trained to generate vectors such that the Euclidean distance between vectors corresponds to the audio similarity of the musical phrases. The system may generate training musical phrases or loops by applying various audio transformations, such as rate shifting, adding noise and reverb, to existing musical phrases. The training may reward the model generating the vectors such that greater transformations of sound (when generating the training musical phrases) produce more diverse vectors from the neural network. The training may also reward the model generating vectors such that the variation between vectors of modified musical phrases (generated by modifying the same original musical phrase) is less than the variation between vectors of distinct, unmodified musical phrases. The training may also reward systems that can successfully classify musical instrument types, which may promote clustering of musical phrases in the vector space by instrument.
[0047] 5 is a diagram illustrating an example of phrase-level interaction, according to some embodiments. In the illustrated example, a user has selected a vocal track, and the illustrated interface includes elements 510, 520, 530, and 540 for the four most recent phrases of the track. These phrases may be temporal sections of the track (or may be looped). The selected phrase may be played in isolation to the user, who may adjust the volume via element 550 or provide other feedback (e.g., like or dislike, a rating, etc., as described above) via element 560. Adjusting the volume may affect the selected phrase relative to the overall volume of the track (which may remain at the same or similar volume relative to other phrases on the track). Providing feedback may affect parameters for the phrase's inclusion in future mixes (e.g., likelihood of inclusion, number of loops if included, adjustments to the phrase for inclusion, etc.).
[0048] Note that this interface may display a variety of information about the selected phrase, such as a graph of amplitude over time, frequency information, tempo, key, loop count, etc.
[0049] 6 is a diagram illustrating an example of mix-level interaction, according to some embodiments. In the illustrated example, the user selects the outer bubble corresponding to the entire mix. The interface shows the current section being composed (in this example, "intensity buildup"). The interface also provides a like factor 602 and a dislike factor 604 for the user to provide feedback on the current section (although various feedback interfaces may be implemented). User feedback on a song section may be used to adjust parameters for the section's inclusion (e.g., based on the previous section up to the time the user provided feedback), such as the number of times it is included, its play length when included, its volume when included, its inclusion with other sections, etc.
[0050] In some embodiments, the interface is configured to display one or more tracks that are not currently mixed. These tracks may be indicated using a different color, line type, or other visual distinction from the currently mixed tracks. These tracks may also be displayed separately from the mixed tracks (tracks included in the mix are shown as touching or overlapping bubbles in some embodiments). The displayed unmixed tracks may be determined by the song machine learning engine to be appropriate for the current mix, but may not currently be included due to, for example, statistical rules or thresholds for tracks included in a particular type of section currently being played.
[0051] The user may, for example, select and drag bubbles for these tracks to include them in the mix. Similarly, the user may listen to tracks that are not currently mixed separately and provide feedback on those tracks, which may be used to control the future inclusion of such tracks.
[0052] It is noted that various functions described herein may be performed on the client side (e.g., on the user device), the server side, or both. As an example of a split function, the client device may perform certain disclosed functions, such as displaying the disclosed interface, and the server may implement adjustments to song parameters based on user input. Example Interface Screenshots
[0053] 7A-10B, described in detail below, show example screenshots from one example of a generative music application embodiment. In the illustrated example, a version of the AiMi application displays interface elements related to "ambient" and "chill" music experiences.
[0054] FIG. 7A shows an example interface showing the currently mixed tracks. FIG. 7B shows user selection of the overall mix and like / dislike factors for receiving user input. The interface also shows the current song section ("build"). FIG. 7B shows user selection of the current song section ("sustain") and shows the preceding and succeeding sections (build and drop). In this example, the user can provide feedback via the like / dislike interface for the individual sections themselves. FIGS. 8A-8C show examples of subsequent interfaces after the initial user feedback.
[0055] In FIG. 8A, the user has selected the bad option for the current section ("Jam Drop"). The interface provides additional elements for more granular feedback. In this example, the user can choose to play this song section for a shorter time interval or less frequently in future songs. In FIG. 8B, the interface provides an indication that the Jam Drop section will be extended for an additional 10 seconds. In FIG. 8C, the interface provides an indication that the Jam Drop section will be played less frequently. Generally speaking, the interface may provide various indications of actions taken based on user feedback regarding a mix, track, section, etc.
[0056] Figure 9A shows an interface displayed in response to a selection to use a beat track. In this example, the beat track and other tracks have been resized and moved relative to Figure 7A based on the user's selection of this track. Figure 9B shows a user's selection of a dislike element for the beat track and an indication that AiMi will play such tracks less (called musical "ideas" in the interface). Figure 9C shows a user's selection of a like element and an indication that such musical ideas will be played more in future songs.
[0057] 10A and 10B show an example interface with beat track and overall mix selection for a "Chill" song, respectively. As shown, the interface may use different colors to match different songs. An example interface with more user-adjustable parameters.
[0058] 11A-11C, described in more detail below, illustrate examples of fine track adjustment, according to some embodiments. Note that in these embodiments, each bubble has a fixed size in the illustrated interface, but one or more inner bubbles indicate the current value of a parameter (e.g., gain), a user-specified value of a parameter (e.g., maximum gain), or some combination thereof. In other embodiments, the inner bubbles may represent multiple different parameters.
[0059] As with the previous example, the user interface of the generative music application displays the track of the current section of the generative music (in this example, the "Jam Drop" section). In some embodiments, the interface also displays (and potentially allows for modification of) the next section to be played. One or more attributes of a particular track (in this example, the gain of the "Beat" track) may be modified using various interface elements.
[0060] FIG. 11A is a block diagram illustrating an example of a currently mixed interface within a generative music section. As shown, each track (represented by a bubble, e.g., circle 1110) includes two concentric circles. In this example, solid circle 1112 reflects the track's current gain, and dotted circle 1114 reflects the track's maximum gain (e.g., based on a default maximum value or previous user input). As described with reference to FIG. 11C, a user may define the maximum gain using interface elements. Note that the dotted lines in FIG. 11A differ from those depicted in FIG. 3. That is, the dotted circles in FIG. 11C indicate the track's maximum gain (although they may actually be shown in the user interface), while the dotted circles in FIG. 3 illustrate the initial volume of each track before modification (although they may not be shown after modification).
[0061] Figure 11B depicts an example of an interface for modifying various characteristics of a selected track within the generated music section. In particular, Figure 11B includes like button element 1120 and dislike button element 1122, slider 1124, and circles 1126 and 1128. The interface of Figure 11B may include more or fewer interface elements than those depicted.
[0062] The user may modify the track's maximum gain (represented by dashed circle 1126). In some embodiments, the user drags slider 1124 (e.g., using a touchscreen or mouse cursor), thereby updating the maximum gain value (shown at 80%) and the radius of dotted circle 1126. In some embodiments, the maximum gain value is used to limit changes the system may make to each track, for example, when varying gain based on a music algorithm. The maximum gain value may also limit the effect of other user changes (e.g., adjustments to another parameter may be constrained so that the gain of the beat does not exceed a specified maximum). The user may use buttons 1120 and 1122 to provide feedback on, and potentially modify, the selected track, as described with respect to FIG. 4.
[0063] While maximum gain is the attribute being modified in the depicted illustration, other attributes of the track (e.g., current gain, tempo, intensity, effects, syncopation, volume, instrument, etc.) may be modified in a similar manner using similar or different interface elements. For example, when a user taps slider 1124, multiple sub-sliders may appear, each capable of adjusting a particular attribute of the track (potentially including gain). Additionally, interactions with the UI may affect other attributes beyond those shown in FIG. 11G, e.g., related attributes in other tracks and / or loops.
[0064] In some embodiments (not explicitly shown), a given bubble / circle at one interface level may correspond to a group of tracks or instruments (which may be called channels). In these embodiments, the user may select a group (e.g., by double-tapping), and the interface may show different levels with an expanded view of the channels within that group. The user may then provide independent input for the channels. For example, a melodic group may contain two melodic channels. The user may select a melodic group and then adjust the gain (or other sub-slider parameter) of one of the melodies.
[0065] 11C depicts an example interface for playing a selected track within the generated music section. As shown, the "Beats" track is selected, causing the system to play only the audio for that track. As shown, this selection obscures the other bubbles. Additional control buttons (e.g., play, pause, share, or record buttons) may be included in the user interface to control the track being played.
[0066] 12 is a flow diagram illustrating an example method, according to some embodiments. The method shown in FIG. 12 may be used in conjunction with, among other things, any of the computer circuitry, systems, devices, elements, or components disclosed herein. In various embodiments, some of the elements of the method shown may be performed simultaneously, in a different order than shown, or may be omitted. Additional method elements may be performed as desired.
[0067] At 1210, in the illustrated embodiment, the computing system selects a set of tracks to include in the generated musical content.
[0068] At 1210, in the illustrated embodiment, the computing system determines a gain value for each of a plurality of selected tracks.
[0069] At 1230, in the illustrated embodiment, the computing system mixes the selected tracks based on the determined gain values to generate the output musical content.
[0070] At 1240, in the illustrated embodiment, the computing system may cause a display of an interface visually indicating the selected tracks and their determined gain values relative to other selected tracks. In some embodiments, the displayed visual representation of the selected tracks includes a bubble element for each track.
[0071] At 1250, in the illustrated embodiment, the computing system receives, via the interface, a user gain input directing an adjustment to the gain value of one of the selected tracks. In some embodiments, the user gain input adjusts the size of the displayed visual track representation.
[0072] At 1260, in the illustrated embodiment, the computing system adjusts the mix of the selected tracks based on user input.
[0073] In some embodiments, the computing system, in response to user input selecting the displayed visual track representation, causes the display of an interface option for isolating and playing the selected track. In some embodiments, the interface further visually indicates, for a given track, both the track's current gain value and a user-directed gain level for the track. In some embodiments, the user-directed gain level is a maximum gain level for the track. In some embodiments, the interface includes, for the user-directed track, multiple user interface elements (e.g., groups), which are user-adjustable to modify one or more additional musical parameters for the user-directed track. Adjustments to the mix may be further based on the one or more modified additional musical parameters for the user-directed track.
[0074] In some embodiments, the computing system causes the display of a feedback interface in response to user input selecting the displayed visual track representation, receives user feedback input providing feedback on the selected track, and adjusts parameters for selecting future tracks for the mix based on the user feedback input. The computing system may determine vectors for the tracks, the vectors generated by an unsupervised machine learning model, and adjusting the parameters for selecting future tracks includes increasing or decreasing selection of tracks within a threshold distance in a vector space of the selected track.
[0075] In some embodiments, in response to a user input selecting the displayed visual track representation, the computing system causes the display of a phrase interface configured to receive user phrase input regarding a phrase included in the track, and may adjust the inclusion of the phrase in future versions of the track based on the user phrase input.
[0076] In some embodiments, the computing system generates a display of the currently playing song section in response to user input selecting a mix of the displayed tracks. The interface may visually indicate one or more additional tracks not currently mixed but suitable for mixing with the mixed and selected tracks. The computing system may adjust one or more parameters for future inclusion of the song section based on user feedback input corresponding to the song section.
[0077] The various techniques described herein may be implemented by one or more computer programs. The term "program" should be interpreted broadly to cover a set of instructions in a programming language that can be executed by a computing device. These programs may be written in any suitable computer language, including low-level languages such as assembly and high-level languages such as Python®. Programs may also be written in compiled languages such as C or C++, or interpreted languages such as JavaScript®.
[0078] Program instructions may be stored on a “computer-readable storage medium” or “computer-readable medium” to facilitate execution of the program instructions by a computing system. Generally speaking, these phrases include any tangible or non-transitory storage or memory medium. The terms “tangible” and “non-transitory” are intended to exclude propagating electromagnetic signals, but are not intended to otherwise limit the type of storage medium. Thus, the phrase “computer-readable storage medium” or “computer-readable medium” is intended to cover types of storage devices that do not necessarily store information permanently (e.g., random access memory (RAM)). Thus, the term “non-transitory” is a limitation on the nature of the medium itself (i.e., the medium must not be a signal), as opposed to a limitation on the data storage permanence of the medium (e.g., RAM vs. ROM).
[0079] The phrases "computer-readable storage medium" and "computer-readable medium" are intended to refer to both storage media within a computer system and removable media such as CD-ROMs, memory sticks, and portable hard disks. These phrases cover any type of volatile memory within a computer system, including DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc., as well as non-volatile memory such as magnetic media, e.g., hard drives, or optical storage devices. These phrases are expressly intended to cover the memory of a server facilitating the download of program instructions, the memory in intermediate computer systems involved in the download, and the memory of any destination computing device. Furthermore, these phrases are intended to cover combinations of different types of memory.
[0080] Furthermore, the computer-readable medium or storage medium may be located on a first set of one or more computer systems on which the program executes, as well as on a second set of one or more computer systems connected to the first set via a network. In the latter instance, the second set of computer systems may provide the program instructions to the first set of computer systems for execution. In short, the phrases "computer-readable storage medium" and "computer-readable medium" may include two or more media that may reside in different locations, for example, on different computers connected via a network.
[0081] The present disclosure includes references to "an embodiment" or groupings of "embodiments" (e.g., "some embodiments" or "various embodiments"). Embodiments are different implementations or instances of the disclosed concepts. References to "an embodiment," "one embodiment," "particular embodiment," etc. do not necessarily refer to the same embodiment. Numerous possible embodiments are contemplated, including those specifically disclosed, as well as modifications or alternatives that are within the spirit or scope of the present disclosure.
[0082] This disclosure may describe potential advantages that may result from the disclosed embodiments. Not all implementations of these embodiments necessarily reveal any or all potential advantages. Whether advantages are realized for a particular implementation depends on many factors, some of which are outside the scope of this disclosure. Indeed, there are several reasons why an implementation falling within the scope of the claims may not exhibit some or all of any disclosed advantages. For example, a particular implementation may include other circuitry outside the scope of this disclosure that, in combination with one of the disclosed embodiments, negates or reduces one or more of the disclosed advantages. Furthermore, suboptimal design implementation of a particular implementation (e.g., implementation techniques or tools) may also negate or reduce a disclosed advantage. Even assuming skilled implementation, realization of advantages may still depend on other factors, such as the environmental conditions in which the implementation is deployed. For example, inputs provided to a particular implementation may prevent one or more problems addressed in this disclosure from occurring on a particular occasion, resulting in the benefits of that solution not being realized. Given the existence of possible factors external to this disclosure, it is expressly intended that any potential advantages described herein should not be construed as claim limitations that must be met in order to demonstrate infringement. Rather, the identification of such potential advantages is intended to illustrate the types of improvements available to a designer having the benefit of this disclosure. The liberal description of such advantages (e.g., stating that a particular advantage "may result") is not intended to convey doubt as to whether such advantages are actually realizable, but rather to recognize the technological reality that realization of such advantages often depends on additional factors.
[0083] Unless otherwise expressly stated, the embodiments are non-limiting. That is, the disclosed embodiments are not intended to limit the scope of claims drafted based on this disclosure, even if only a single example of a particular feature is described. The disclosed embodiments are intended to be illustrative, not limiting, unless stated to the contrary in this disclosure. Accordingly, the application is intended to permit claims covering the disclosed embodiments, as well as alternatives, modifications, and equivalents that will be apparent to those skilled in the art having the benefit of this disclosure.
[0084] For example, features in the present application may be combined in any suitable manner. Accordingly, during prosecution of this application (or an application claiming priority thereto), new claims may be formulated to such combinations of features. In particular, with respect to the appended claims, features from dependent claims may be combined, where appropriate, with features of other dependent claims, including claims that are dependent on other independent claims. Similarly, features from each independent claim may be combined, where appropriate.
[0085] Thus, the accompanying dependent claims may each be drafted so as to depend on a single other dependent claim, although additional dependencies are also contemplated. Any combination of features in the dependent claims consistent with this disclosure is contemplated and may be claimed in this or another application. In short, combinations are not limited to those specifically recited in the accompanying claims.
[0086] It is also contemplated that, where appropriate, a claim drafted in one form or statutory type (e.g., apparatus) is intended to support a corresponding claim in another form or statutory type (e.g., method).
[0087] Because this disclosure is a legal document, various terms and phrases may be subject to administrative and judicial interpretation. It is hereby declared that the definitions provided in the following paragraphs, and throughout this disclosure, are to be used in determining how to interpret any claims drafted based on this disclosure.
[0088] Reference to the singular form of an item (i.e., a noun or noun phrase preceded by "a," "an," or "the") is intended to mean "one or more" unless the context clearly dictates otherwise. Thus, a reference to "an item" in a claim does not exclude additional instances of that item without attendant context. A "plurality" of items refers to a set of two or more items.
[0089] The word "may" is used herein in a permissive sense (ie, having the possibility, being able to) and not in a mandatory sense (ie, must).
[0090] The terms "comprising" and "including" and their forms are open-ended and mean "including, but not limited to."
[0091] In this disclosure, when the term "or" is used in reference to a list of alternatives, it will generally be understood to be used in an inclusive sense unless the context provides otherwise. Thus, a list of "x or y" is equivalent to "x or y, or both," and thus covers 1) x but not y, 2) y but not x, and 3) both x and y. On the other hand, a phrase such as "either x or y, but not both" makes it clear that "or" is used in an exclusive sense.
[0092] The enumeration of "w, x, y, or z, or any combination thereof" or "at least one of w, x, y, and z" is intended to cover all possibilities involving a single element up to the total number of elements in the set. For example, given the set [w, x, y, z], these phrases cover any single element of the set (e.g., w but not x, y, or z), any two elements (e.g., w and x but not y or z), any three elements (e.g., w, x, and y but not z), and all four elements. Thus, the phrase "at least one of w, x, y, and z" refers to at least one element of the set [w, x, y, z], thereby covering all possible combinations of this list of elements. This phrase should not be interpreted as requiring that there be at least one instance of w, at least one instance of x, at least one instance of y, and at least one instance of z.
[0093] In this disclosure, various "labels" may precede nouns or noun phrases. Unless the context provides otherwise, different labels used for a feature (e.g., "first circuit," "second circuit," "particular circuit," "given circuit," etc.) refer to different instances of the feature. Additionally, the labels "first," "second," and "third," when applied to features, do not imply any type of ordering (e.g., spatial, temporal, logical, etc.) unless otherwise specified.
[0094] The phrase "based on" or "based on" is used to describe one or more factors that influence the determination. This term does not exclude that additional factors may influence the determination. That is, the determination may be based only on the specified factors, or on the specified factors and other unspecified factors. Consider the phrase "determining A based on B." This phrase specifies that B is used to determine A or is a factor that influences the determination of A. This phrase does not exclude that the determination of A may also be based on some other factor, such as C. This phrase is also intended to cover embodiments in which A is determined only based on B. As used herein, the phrase "based on" is synonymous with the phrase "based at least in part on."
[0095] The phrases "in response to" and "responsive to" describe one or more factors that trigger an effect. This phrase does not exclude that additional factors may influence or otherwise trigger the effect, either in conjunction with or independently of the specified factors. That is, the effect may be responsive only to those factors, or to the specified factors and other unspecified factors. Consider the phrase "performing A in response to B." This phrase specifies that B is the factor that triggers the execution of A or triggers a particular outcome for A. This phrase does not exclude that performing A may also be responsive to some other factor, such as C. This phrase does not exclude that performing A may be responsive to B and C jointly. This phrase is also intended to cover embodiments in which A is performed in response to B only. As used herein, the phrase "responsive to" is synonymous with the phrase "at least partially responsive to." Similarly, the phrase "in response to" is synonymous with the phrase "at least partially responsive to."
[0096] Within this disclosure, different entities (which may be variously referred to as "units," "circuits," other components, etc.) are described or claimed as being "configured" to perform one or more tasks or operations. This formulation, i.e., an entity configured to perform one or more tasks, is used herein to refer to a structure (i.e., something physical). More specifically, this formulation is used to indicate that the structure is arranged to perform one or more tasks during operation. A structure can be said to be "configured" to perform some task even if the structure is not currently operating. Thus, an entity described or listed as being "configured" to perform some task refers to something physical, such as a device, a circuit, a system having a processor unit and a memory that stores executable program instructions to implement the task. This phrase is not used herein to refer to something intangible.
[0097] In some cases, various units / circuits / components may be described herein as performing a set of tasks or operations, and it will be understood that these entities are "configured to" perform those tasks / operations, even if not specifically stated otherwise.
[0098] The term "configured" is not intended to mean "configurable." For example, an unprogrammed FPGA is not considered to be "configured" to perform a particular function. However, this unprogrammed FPGA may be "configurable" to perform that function. After appropriate programming, the FPGA may be said to be "configured" to perform a particular function.
[0099] For purposes of U.S. patent applications based on this disclosure, reciting in a claim that a structure is "configured to" perform one or more tasks is expressly intended not to invoke 35 U.S.C. §112(f) with respect to that claim element. If applicant wishes to invoke Section 112(f) during prosecution of a U.S. patent application based on this disclosure, applicant will recite claim elements using "means for" [performing] [function] constructs.
[0100] Different "circuits" may be described in this disclosure. These circuits or "circuitry" comprise hardware that includes various types of circuit elements, such as combinational logic, clocked storage (e.g., flip-flops, registers, latches, etc.), finite state machines, memories (e.g., random access memory, embedded dynamic random access memory), programmable logic arrays, etc. Circuitry may be custom designed or obtained from standard libraries. In various implementations, circuitry may include digital components, analog components, or a combination of both, as appropriate. Particular types of circuitry may be generally referred to as "units" (e.g., decode units, arithmetic logic units (ALUs), functional units, memory management units (MMUs), etc.). Such units may also be referred to as circuits or circuitry.
[0101] Thus, the disclosed circuits / units / components and other elements illustrated in the drawings and described herein include hardware elements such as those described in the preceding paragraphs. In many instances, the internal arrangement of hardware elements within a particular circuit may be specified by describing the function of that circuit. For example, a particular "decode unit" may be described as performing the function of "processing the opcode of an instruction and routing the instruction to one or more of a plurality of functional units," meaning that the decode unit is "configured" to perform this function. This specification of this function is sufficient to suggest a set of possible structures for the circuit to one of ordinary skill in the computer arts.
[0102] In various embodiments, as described in the preceding paragraphs, circuits, units, and other elements may be defined by the functions or operations they are configured to implement. The arrangement and placement of such circuits / units / components relative to one another and the manner in which they interact form a microarchitecture definition of hardware that will ultimately be fabricated in an integrated circuit or programmed into an FPGA to form a physical implementation of the microarchitecture definition. Thus, a microarchitecture definition will be recognized by those skilled in the art as a structure from which many physical implementations may be derived, all of which fall within the broader structure described by the microarchitecture definition. That is, those skilled in the art presented with a microarchitecture definition provided in accordance with this specification may, without undue experimentation, apply ordinary skill to implement the structure by coding a description of the circuits / units / components in a hardware description language (HDL), such as Verilog or VHDL. HDL descriptions are often expressed in a manner that may appear functional. However, to those skilled in the art, this HDL description is the method used to translate the structure of a circuit, unit, or component into the next level of implementation detail. Such HDL descriptions may take the form of behavioral code (which is typically not synthesizable), register transfer language (RTL) code (which, in contrast to behavioral code, is typically synthesizable), or structural code (e.g., a netlist specifying logic gates and their connectivity). The HDL descriptions may subsequently be synthesized against a library of cells designed for a given integrated circuit manufacturing technology and may be modified for timing, power, and other reasons to produce a final design database that is sent to a foundry to generate masks and ultimately produce the integrated circuit. Some hardware circuits, or portions thereof, may also be custom designed in a schematic editor and captured in the integrated circuit design along with the synthesized circuit. An integrated circuit may include transistors and other circuit elements (e.g., passive elements such as capacitors, resistors, inductors, etc.), and interconnections between the transistors and circuit elements.Some embodiments may implement multiple integrated circuits coupled together to implement a hardware circuit, and / or in some embodiments, discrete elements may be used. Alternatively, an HDL design may be synthesized into a programmable logic array, such as a field programmable gate array (FPGA), and implemented within an FPGA. This separation between the design of a group of circuits and the subsequent low-level implementation of those circuits generally results in a scenario in which a circuit or logic designer never specifies a particular set of structures for the low-level implementation beyond describing what the circuit is configured to do, as this process is performed at different stages of the circuit implementation process.
[0103] The fact that many different low-level combinations of circuit elements may be used to implement the same specification for a circuit results in numerous equivalent structures for that circuit. As noted above, these low-level circuit implementations may vary according to changes in manufacturing technology, the foundry selected to manufacture the integrated circuit, the library of cells provided for a particular project, etc. In many cases, the choices made by different design tools or methodologies to generate these different implementations may be arbitrary.
[0104] Furthermore, it is common for a single implementation of a particular functional specification of a circuit to include a large number of devices (e.g., millions of transistors) for a given embodiment. Thus, the sheer volume of this information makes it impractical to provide a complete enumeration of the low-level structures used to implement a single embodiment, much less a vast array of equivalent possible implementations. For this reason, this disclosure describes the structure of a circuit using functional abbreviations commonly used in the industry.
[0105] In this disclosure, various “modules” operable to perform specified functions are illustrated in the figures and described in detail. As used herein, a “module” refers to software or hardware operable to perform a specified set of operations. A module may refer to a set of software instructions executable by a computer system to perform a set of operations. A module may also refer to hardware configured to perform a set of operations. Hardware modules may comprise general-purpose hardware, non-transitory computer-readable media that store program instructions, or specialized hardware such as customized ASICs. Thus, a module described as “executable” to perform an operation refers to a software module, while a module described as “configured” to perform an operation refers to a hardware module. A module described as “operable” to perform an operation refers to a software module, a hardware module, or some combination thereof. Furthermore, for any description herein referring to a module being “executable” to perform particular operations, it should be understood that in other embodiments, these operations may be implemented by hardware modules “configured” to perform the operations, and vice versa.
Claims
1. 1. A method comprising: selecting, by a computing system, a set of musical phrases to be included in the generated musical content; determining, by said computing system, a gain value for each of a plurality of selected musical phrases; mixing, by the computing system, the selected musical phrases based on the determined gain values to generate output musical content; causing, by said computing system, a display of an interface visually indicating the selected musical phrase and their determined gain values relative to other selected musical phrases; receiving, by the computing system, via the interface, a user gain input indicating a gain value of one of the selected musical phrases to be adjusted; and adjusting, by the computing system, a mix of the selected musical phrase based on the user gain input.
2. 10. The method of claim 1, further comprising: in response to user input selecting a displayed visual musical phrase representation, the computing system causing the display of an interface option for isolating and playing the selected musical phrase.
3. in response to a user input selecting a displayed visual musical phrase representation, the computing system causes a display of a feedback interface and receives user feedback input providing feedback on the selected musical phrase; The method of claim 1 , further comprising: adjusting parameters for selecting future musical phrases to mix based on the user feedback input.
4. determining vectors for the musical phrases, the vectors being generated by an unsupervised machine learning model; 4. The method of claim 3, wherein adjusting parameters for selecting future musical phrases further comprises increasing or decreasing selection of musical phrases within a threshold distance in a vector space of the selected musical phrases.
5. 10. The method of claim 1, further comprising, in response to a user input selecting the displayed visual musical phrase representation, causing the display of a phrase interface configured to receive user phrase input regarding a phrase included in the musical phrase.
6. The method of claim 5 , further comprising adjusting inclusion of phrases in future versions of the musical phrase based on the user phrase input.
7. The method of claim 1 , wherein the user gain input adjusts the size of a displayed visual musical phrase representation.
8. The method of claim 1 , wherein the displayed visual representation of the selected musical phrases includes a bubble element for each musical phrase.
9. 10. The method of claim 1, further comprising: in response to a user input selecting the displayed mix of musical phrases, the computing system causing a display of a currently playing song section.
10. The method of claim 9 , further comprising adjusting one or more parameters for future inclusion of the song section based on user feedback input corresponding to the song section.
11. 10. The method of claim 9, wherein the interface visually indicates one or more additional musical phrases that are not currently mixed but are suitable for mixing with the mixed selected musical phrase.
12. The method of claim 1 , wherein the interface further visually indicates, for a given musical phrase, both the current gain value for the musical phrase and the user-directed gain level for the musical phrase.
13. The method of claim 12 , wherein the user-indicated gain level is a maximum gain level for the musical phrase.
14. the interface includes a plurality of user interface elements for a user-directed musical phrase, the plurality of user interface elements being user-adjustable to modify one or more additional musical parameters for the user-directed musical phrase; The method of claim 1 , wherein the adjustment is further based on one or more modified additional musical parameters for the user-indicated musical phrase.
15. The method of claim 1 , wherein the interface includes groups of the selected musical phrases, the groups being user selectable to specify user gain inputs for individual musical phrases within the groups.
16. one or more memories; Executing program instructions stored in said one or more memories; selecting a set of musical phrases to be included in the generated musical content; determining a gain value for each of a plurality of selected musical phrases; mixing the selected musical phrases based on the determined gain values to generate output musical content; generating a display of an interface visually indicating the selected musical phrase and their determined gain values relative to other selected musical phrases; receiving a user gain input via the interface indicating a gain value of one of the selected musical phrases to be adjusted; and one or more processors configured to: adjust a mix of the selected musical phrase based on the user gain input.
17. A non-transitory computer-readable medium having stored thereon instructions executable by a computing device to perform operations, the operations comprising: selecting a set of musical phrases to be included in the generated musical content; determining a gain value for each of a plurality of selected musical phrases; mixing the selected musical phrases based on the determined gain values to generate output musical content; generating a display of an interface visually indicating the selected musical phrase and their determined gain values relative to other selected musical phrases; receiving a user gain input via the interface indicating a gain value of one of the selected musical phrases to be adjusted; and adjusting a mix of the selected musical phrase based on the user gain input.
18. 20. The non-transitory computer-readable medium of claim 17, wherein the operations further include, in response to user input selecting a displayed visual musical phrase representation, causing the display of an interface option for isolating and playing the selected musical phrase.
19. The operation is causing display of a feedback interface in response to user input selecting a displayed visual musical phrase representation and receiving user feedback input providing feedback on the selected musical phrase; 20. The non-transitory computer-readable medium of claim 17, further comprising: adjusting parameters for selecting future musical phrases to mix based on the user feedback input.
20. 20. The non-transitory computer-readable medium of claim 17, wherein the interface visually indicates, for a given musical phrase, both a current gain value for the musical phrase and a user-directed gain level for the musical phrase.