Dynamic control of generative musical works
By designing a graphical user interface, allowing users to control and feedback the composition decisions of the generative music system in real time, it solves the problem that users have difficulty in fine-grained control of real-time generated music, and achieves higher personalization and flexibility.
Patent Information
- Application Number
- CN202380072488.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-14
- Filing Date
- 2023-10-13
- Publication Date
- 2025-05-23
AI Technical Summary
In a generative music system, it is difficult to achieve fine-grained control and feedback on real-time generated music by users, and traditional user interfaces are not suitable for dynamically created generative music.
A user interface is designed that allows users to graphically view and modify music generator composition decisions and reflect these feedbacks in real time during the music generation process. This interface represents phrases and tracks through bubbles, and users can adjust the volume and other attributes of the phrases to provide feedback to influence subsequent compositions.
It realizes real-time control and feedback of generated music by users, enhances users' understanding and participation in the music generation process, and improves the personalization and flexibility of music generation.
Smart Images

Figure CN120035855A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to audio engineering and, more particularly, to computer-generated musical content. Background Art
[0002] Generative music systems can use computer systems to create music content. For example, the AiMi platform allows users to select a music genre and listen to dynamically generated compositions in that genre. Users can also provide feedback on the composition, and the system can adjust the composition parameters based on user feedback.
[0003] In the context of generative music, it can be challenging to communicate various aspects of a composition to a user or to determine the fine-grained preferences of a specific user. BRIEF DESCRIPTION OF THE DRAWINGS
[0004] Figure 1 is a diagram illustrating an exemplary music generator module that generates music content based on a variety of different types of inputs in accordance with some embodiments.
[0005] Figure 2 is a diagram illustrating an example interface showing a plurality of mixed phrases in accordance with some embodiments.
[0006] Figure 3 is a diagram showing an example user adjusting the size of a displayed phrase in accordance with some embodiments.
[0007] Figure 4 is a diagram showing example user phrase feedback and individual listening functionality of phrases in accordance with some embodiments.
[0008] Figure 5 is a diagram showing example user phrase feedback in accordance with some embodiments.
[0009] Figure 6 is a diagram showing example user snippet feedback in accordance with some embodiments.
[0010] Figures 7A-10B is a screenshot illustrating an example interface according to some embodiments.
[0011] Fig.11A is a block diagram illustrating an example interface that displays phrases currently mixed in a generative music clip.
[0012] Fig. 11B is a diagram illustrating an example interface for changing various properties of a selected phrase in a generative music piece in accordance with some embodiments.
[0013] Fig. 11Cis a diagram illustrating an example interface for playing a selected phrase in a generative music piece in accordance with some embodiments.
[0014] Fig.12 is a flow chart illustrating an example method according to some embodiments. DETAILED DESCRIPTION
[0015] Typically, generative music that is generated "on the fly" can be more difficult for a user to control than music content that has already been recorded and exists in a library. Typical generative music software generates its music in real time, and does not provide user-accessible options for providing feedback during the music generation process, and traditional user interfaces are not particularly effective for user customization of generative music. On the other hand, music players for music that exists statically in a library provide various functions that allow music to be scored and evaluated, but these functions are not necessarily applicable to (or even incompatible with) dynamically created generative music.
[0016] The disclosed system can implement various techniques for incorporating and reflecting user feedback on generative music in real time. The user interface discussed below can allow a user to graphically view internal engine composition decisions (e.g., musical phrase selection, mixing decisions, attributes, parameters, etc.), modify the results of these decisions, and send feedback to the system, which will be reflected in the composition. In some cases, this feedback adjusts the decisions in real time and can be used immediately in the composition. As discussed in detail below, the disclosed user interface technology can therefore simultaneously achieve: providing a user with a view of what is currently happening in a generative music work, and giving the user the ability to influence subsequent composition.
[0017] Overview of the Sample Music Generator
[0018] In general, the disclosed music generator includes loop data, metadata (e.g., information describing the loop), and a grammar for combining loops based on the metadata. The generator can use rules to create a music experience to identify loops based on metadata and target characteristics of the music experience. It can be configured to expand the set of experiences it can create by adding or modifying rules, loops, and / or metadata. Adjustments can be performed manually (e.g., an artist adds new metadata), or the music generator can expand rules / loops / metadata as it monitors the music experience and desired goals / characteristics within a given environment. For example, if the music generator observes a crowd and sees that people are smiling, it can expand its rules and / or metadata to note that certain loop combinations cause people to smile. Similarly, if cash register sales increase, the rule generator can use this feedback to expand the rules / metadata of associated loops related to sales growth.
[0019] As used herein, the term "loop" refers to the sound information of a single instrument within a specific time interval. The loop can be played in a repeated manner (for example, a 30-second loop can be played four times continuously to generate 2 minutes of music content), but the loop can also be played once, for example, without repetition. The various techniques discussed with reference to the loop can also be performed using an audio file comprising multiple instruments. The term "track" includes loops and audio files containing multiple instrument sounds. In addition, a "track" can be recorded or a computer-generated phrase (for example, completely synthesized from scratch or generated by combining previously recorded or previously generated sounds). In general, the term "phrase" refers to the information of a sound sequence within a specified time interval. A phrase can be a loop or a track, and can include a single instrument or multiple instruments. It should be understood that the various techniques discussed with reference to one of the loops, tracks or phrases can be applied to all three or any one of the three in various embodiments.
[0020] Figure 1 is a diagram illustrating an exemplary music generator according to some embodiments. In the illustrated embodiment, the music generator module 160 receives various information from a number of different sources and generates output music content 140.
[0021] In the illustrated embodiment, the module 160 accesses the stored (one or more) loops and the (one or more) corresponding attributes 110 of the stored (one or more) loops, and combines these loops to generate the output music content 140. Specifically, the music generator module 160 selects loops based on the attributes of the loops, and combines the loops based on the target music attributes 130 and / or the environmental information 150. In some embodiments, the environmental information is indirectly used to determine the target music attributes 130. In some embodiments, the target music attributes 130 are explicitly specified by the user, for example, by specifying the desired energy level, mood, multiple parameters, etc. Examples of target music attributes 130 include energy, complexity, and diversity, but more specific attributes (e.g., corresponding to the attributes of the stored tracks) may also be specified. The music attributes may be input by the user or determined based on environmental information (e.g., ambient noise, lighting, etc.). In general, when a higher level of target music attributes is specified, the system may determine a lower level of specific music attributes before generating the output music content. Example techniques for generating music content based on one or more musical attributes are described in detail in U.S. Patent Application No. 13 / 969,372, filed on August 16, 2013 (now U.S. Patent No. 8,812,144), entitled "Music Generator," which is incorporated herein by reference in its entirety. The '372 publication discusses techniques such as selecting stored loops and / or tracks or generating new loops / tracks, and layering selected loops / tracks to generate output music content. To the extent any possible conflict of interpretation between any definition in the incorporated application and the remainder of this disclosure exists, the language of this disclosure shall control.
[0022] Complexity can refer to many loops and / or instruments included in the composition. Energy may be related to other attributes, or may be orthogonal to other attributes. For example, changing the pitch or rhythm can affect energy. However, for a given rhythm and pitch, the energy can be changed by adjusting the instrument type (for example, by adding hi-hats or white noise), complexity, volume, etc. Diversity can refer to the amount of change of the generated music over time. Diversity can be generated for a set of static other music attributes (for example, by selecting different tracks for a given rhythm and pitch), or diversity can be generated by changing the music attributes over time (for example, by changing the rhythm and pitch more often when greater diversity is desired). In some embodiments, the target music attribute can be considered to exist in a multidimensional space, and the music generator module 160 can be based on environmental changes and / or user input, for example, slowly moving through the space with course correction (if necessary).
[0023] In some embodiments, the properties stored with the loops include information about one or more loops including: tempo, volume, energy, diversity, spectrum, envelope, modulation, periodicity, rise and decay times, noise, artist, instrument, theme, gain, etc. Note that in some embodiments, the loops are partitioned so that a set of one or more loops is specific to a particular loop type (e.g., one instrument or one type of instrument).
[0024] In the illustrated embodiment, module 160 accesses (one or more) stored rule sets 120. In some embodiments, (one or more) stored rule sets 120 specify rules for how many loops overlap so that they are played simultaneously (which can correspond to the complexity of the output music), which major / minor progressions to use when switching between loops or phrases, which instruments to use together (e.g., instruments that have an affinity for each other), etc. to achieve the target music attributes. In other words, the music generator module 160 uses the stored one or more rule sets 120 to achieve one or more declarative goals defined by the target music attributes (and / or target environment information). In some embodiments, the music generator module 160 includes one or more pseudo-random number generators that are configured to introduce pseudo-randomness to avoid repetitive output music.
[0025] In some embodiments, the environmental information 150 includes one or more of the following: lighting information, ambient noise, user information (facial expressions, body posture, activity level, movement, skin temperature, performance of certain activities, clothing type, etc.), temperature information, purchasing activity in the area, time of day, day of the week, time of year, number of people present, weather conditions, etc. In some embodiments, the music generator module 160 does not receive / process environmental information. In some embodiments, the environmental information 150 is received by another module that determines the target music attributes 130 based on the environmental information. The target music attributes 130 can also be derived based on other types of content (e.g., video data). In some embodiments, the environmental information is used to adjust one or more stored rule sets 120, for example to achieve one or more environmental goals. Similarly, the music generator can use environmental information to adjust the properties of one or more stored loops, for example, to indicate target music attributes or target audience characteristics that are particularly relevant to these loops.
[0026] As used herein, the term "module" refers to a circuit configured to perform a specified operation, or refers to a physical non-transient computer-readable medium that stores information (e.g., program instructions) that instructs other circuits (e.g., processors) to perform specified operations. Modules can be implemented in a variety of ways, including as hard-wired circuits or as memories in which program instructions are stored, which can be executed by one or more processors to perform operations. Hardware circuits can, for example, include customized very large scale integration (VLSI) circuits or gate arrays, off-the-shelf semiconductors (such as logic chips, transistors), or other discrete components. Modules can also be implemented in programmable hardware devices such as field programmable gate arrays, programmable array logic, programmable logic devices, etc. Modules can also be any suitable form of non-transient computer-readable media that stores executable program instructions to perform specified operations.
[0027] As used herein, the phrase "music content" refers to both the music itself (the audible representation of the music) and the information that can be used to play the music. Thus, a song recorded as a file on a storage medium (such as, but not limited to, a compact disk, a flash drive, etc.) is an example of music content; the sound produced by outputting the recorded file or other electronic representation (e.g., through a speaker) is also an example of music content.
[0028] The term "music" includes its well-understood meaning, including sounds produced by musical instruments as well as human voices. Thus, music includes, for example: instrumental performances or recordings, a cappella performances or recordings, and performances or recordings that include both musical instruments and human voices. One of ordinary skill in the art will recognize that "music" does not include all recordings of human voices. Works that do not include musical attributes such as rhythm or cadence (e.g., speech, news broadcasts, and audiobooks) are not music.
[0029] A piece of musical "content" may be distinguished from another piece of musical content in any suitable manner. For example, a digital file corresponding to a first song may represent the first piece of musical content, while a digital file corresponding to a second song may represent the second piece of musical content. The phrase "musical content" may also be used to distinguish specific intervals within a given musical work, so that different portions of the same song may be considered different musical content. Similarly, different tracks (e.g., piano track, guitar track) within a given musical work may also correspond to different musical content. In the context of a potentially endless stream of generated music, the phrase "musical content" may be used to refer to a portion of that stream (e.g., a few bars or minutes).
[0030] The music content generated by the embodiments of the present disclosure may be "new music content" - a combination of music elements that has never been generated before. A related (but broader) concept - "original music content" is further described below. To facilitate the explanation of this term, the concept of a "controlling entity" relative to a music content generation instance is described. Unlike the phrase "original music content", the phrase "new music content" does not involve the concept of a controlling entity. Therefore, new music content refers to music content that has never been generated by any entity or computer system before.
[0031] Conceptually, the present disclosure refers to a certain "entity" as a specific instance of controlling computer-generated music content. Such an entity has any legal rights (e.g., copyright) that may correspond to the computer-generated content (to the extent that any such rights may actually exist). In one embodiment, the individual who creates (e.g., encodes various software routines) a computer-implemented music generator or operates (e.g., provides input to) a specific instance of computer-implemented music generation will be the controlling entity. In other embodiments, the computer-implemented music generator can be created by a legal entity (e.g., a company or other commercial organization) such as in the form of a software product, a computer system, or a computing device. In some cases, such a computer-implemented music generator can be deployed to many clients. According to the licensing terms associated with the distribution of the music generator, in various cases, the controlling entity can be a creator, a distributor, or a client. If there is no such clear legal agreement, the controlling entity of the computer-implemented music generator is an entity that promotes (e.g., provides input and thereby operates) a specific instance of computer-generated music content.
[0032] Within the meaning of the present disclosure, computer generation of "original musical content" by a controlling entity refers to: 1) a combination of musical elements that have never been generated before by the controlling entity or anyone else, and 2) a combination of musical elements that have been generated before but were originally generated by the controlling entity. Content type 1) is referred to herein as "novel musical content" and is similar to the definition of "new musical content", except that the definition of "novel musical content" involves the concept of "controlling entity", while the definition of "new musical content" does not involve the concept of "controlling entity". On the other hand, content type 2) is referred to herein as "proprietary musical content". Note that the term "proprietary" in this context does not refer to any implied legal rights in the content (although such rights may exist), but is only used to indicate that the musical content was originally generated by the controlling entity. Therefore, the "regeneration" of musical content previously originally generated by the controlling entity by a controlling entity constitutes "generation of original musical content" within the present disclosure. "Non-original musical content" with respect to a particular controlling entity is musical content that is not the "original musical content" of the controlling entity.
[0033] Some music contents may include music components from one or more other music contents. Creating music contents in this way is called "sampling" music contents, and is common in some music works, especially in some music genres. Such music contents are referred to as "music contents with sampling components", "derivative music contents" or other similar terms are used in this article. By contrast, music contents that do not include sampling components are referred to as "music contents without sampling components", "non-derivative music contents" or other similar terms are used in this article.
[0034] In applying these terms, note that if any particular musical content is reduced to a sufficient level of granularity, an argument can be made that that musical content is derivative (effectively meaning that all musical content is derivative). The terms "derivative" and "non-derivative" are not used in this sense in this disclosure. With respect to computer generation of musical content, if the computer generation selects portions of components from pre-existing musical content of an entity other than the controlling entity (e.g., a computer program selects specific portions of an audio file of a work by a popular artist to include in a piece of musical content being generated), then such computer generation is considered derivative (and produces derivative musical content). On the other hand, if the computer generation does not utilize such components of such pre-existing content, then the computer generation of musical content is considered non-derivative (and produces non-derivative musical content). Note that some "original musical content" may be derivative musical content, while some may be non-derivative musical content.
[0035] Note that the term "derivative" is intended to have a broader meaning within this disclosure than the term "derivative work" as used in U.S. copyright law. For example, derivative musical content may or may not be a derivative work under U.S. copyright law. The term "derivative" in this disclosure is not intended to convey a negative connotation; it is only used to indicate whether particular musical content "borrows" portions of content from another work.
[0036] In addition, the phrases "new musical content", "novel musical content" and "original musical content" are not intended to include musical content that is only slightly different from the pre-existing combination of musical elements. For example, only changing a few notes of a pre-existing musical work will not result in new, novel or original musical content, as those phrases are used in the present disclosure. Similarly, only changing the pitch or rhythm of a pre-existing musical work or adjusting the relative strength of its frequency (e.g., using an equalizer interface) will not produce new, novel or original musical content. In addition, the phrases "new musical content, novel musical content and original musical content" are not intended to cover those musical contents that are borderline cases between original content and non-original content; on the contrary, these terms are intended to cover musical content that is undoubtedly and provably original, including musical content that will be eligible for copyright protection for the controlling entity (referred to as "protectable" musical content in this article). In addition, as used herein, the term "available" musical content refers to musical content that does not violate the copyright of any entity other than the controlling entity. New and / or original musical content is often protectable and available. This may be advantageous in preventing the copying of musical content and / or paying royalties for musical content.
[0037] Although the various embodiments discussed herein use rule-based engines, various other types of computer-implemented algorithms may be used for any of the computer learning and / or music generation techniques discussed herein. Rule-based techniques may or may not include statistical and machine learning techniques. Example techniques for generating music content using statistical rules and machine learning engines are described in detail in U.S. patent application number 16 / 420,456 (now U.S. Patent No. 10,679,596), filed on May 23, 2019, and entitled “Music Generator,” the entire contents of which are incorporated herein by reference. Additional example techniques for machine learning training and processing audio data are described in U.S. patent application number 17 / 174,052, filed on February 11, 2021, and entitled “Music Content Generation Using Image Representations of Audio Files,” the entire contents of which are incorporated herein by reference.
[0038] Example user interface with bubble elements
[0039] Figure 2 is a diagram showing an example user interface for a generative music application according to some embodiments. In the example shown, the interface includes a circular bubble corresponding to the current segment of generative music, which in turn includes multiple smaller bubbles indicating the phrase currently being mixed. In this example, the phrase includes a beat, vocal, effect (FX), effect 2, top, melody, and pad.
[0040] As the composition plays, tracks can be removed from or inserted into the mix, and the volume of the tracks can be adjusted dynamically. In some embodiments, the interface elements shown reflect composition changes in real time, such as by adding / removing bubbles, changing bubble size to reflect the current volume of a given track, etc. This can quickly provide the user with detailed information about the current composition status.
[0041] As discussed in detail below, the user may also interact with the interface to adjust subsequent compositions. User-initiated changes may begin to occur immediately in the current composition, or may be delayed for certain types of input (e.g., in order to incorporate these changes while continuing to play the enjoyable musical content). For example, changes to the volume of a phrase may be immediately reflected in the mix, while feedback about a given phrase or musical segment may be reflected in subsequent compositional decisions (e.g., by adjusting the number of times a phrase is included, the amount of time it is played, etc.).
[0042] Although circular bubbles are discussed herein for purposes of explanation, various other shapes may be displayed in other embodiments including, but not limited to, ovals, ellipses, polygons, three-dimensional shapes, and the like.
[0043] Figure 3 is a diagram showing an example user adjustment to bubble size according to some embodiments. In the example shown, the user selects the FX bubble (with dashed line 310 indicating its initial position), and enters a change to the bubble size (e.g., by dragging a finger on a touch screen or moving a cursor using a mouse), where the resized FX bubble is shown using a bold line 320. In some embodiments, this immediately increases the volume of the FX phrase in the output mix.
[0044] It should be noted that Figure 3 The example input can also affect future compositions, for example, the music engine can mix the phrase (or similar phrases) at a higher volume in subsequent mixes than it would have used before the user feedback. Figure 3 An example input of may have an immediate effect, but the music engine can then dynamically adjust the volume of the adjusted phrase (although that adjustment may be scaled based on the user input). Thus, a user-specified volume may not be a static change, but can generally affect the composition in the desired direction.
[0045] Figure 4 is a diagram showing example interface elements for a particular phrase or channel according to some embodiments. In the example shown, the user has selected a vocal phrase. The interface provides a "listen alone" element that can be selected to play the phrase without mixing in other phrases. This allows the user to determine if the selected phrase is one they really want to adjust.
[0046] The interface also provides thumbs up and thumbs down elements 402 and 404 that allow users to provide feedback about a phrase. Thumbs up can represent positive feedback, while thumbs down can represent negative feedback. The music engine can use this feedback to adjust the parameters related to the future mixing of the phrase, for example, more frequently (more frequently) or less frequently (less frequently) include it, when included, the time interval is shorter or longer, the volume is smaller or larger, more often (more often) or less often (less often) is combined with other phrases that are also mixed at the same time, more often or less often (less often) is combined with other environmental information when the user provides feedback, etc. It should be noted that thumbs up and thumbs down are included as an example, and various other feedback interfaces (for example, star rating method, digital scale, arrow up / down, like / dislike, etc.) can be implemented to allow the user to rate a given phrase or respond in other ways. In some embodiments, the user presses thumbs up / down elements 402 and 404 to modify the user's profile information, for example, to indicate their music preferences in subsequent compositions.
[0047] Thus, in some embodiments, feedback related to phrases (or feedback related to loops / phrases used to construct phrases) is used to adjust future phrase selections. For example, in some embodiments, the system is configured to encode phrases into vectors, and user feedback can be used to select phrases with vectors similar to the current phrase more often or less often. For example, the system can search in vector space in response to a "thumbs up" to find similar sound loops for future inclusion. Similarly, the system can reduce the likelihood of playing loops (or completely prevent inclusion of loops) within a specified distance from the currently playing loop in vector space.
[0048] The computing system can create vectors using a self-supervised neural network (which may also be referred to as "unsupervised") that is trained to produce vectors such that the Euclidean distance between vectors corresponds to the audio similarity of the phrases. The system can generate training phrases or loops by applying various audio transformations (e.g., rate shifts, adding noise and reverb, etc.) to existing phrases. Training can reward the model used to generate vectors so that greater transformations of the sound (when generating training phrases) produce more diverse vectors from the neural network. Training can also reward the model used to generate vectors so that the variation between vectors of modified phrases (which are generated by modifying the same original phrases) is less than the variation between vectors of different unmodified phrases. Training can also reward the system for being able to successfully classify instrument types, which can encourage the clustering of phrases by instrument in the vector space.
[0049] Figure 5 is an illustration showing an example short - phrase - level interaction according to some embodiments. In the example shown, the user has selected a vocal track, and the interface shown includes elements for the four most recent short phrases 510, 520, 530, and 540 of that track. These short phrases can be temporal segments of the track (which can be loops). The selected short phrase can be played for the user individually, and the user can adjust the volume via element 550 or provide other feedback (such as thumbs up or down, rating, etc. as discussed above) via element 560. Adjusting the volume can affect the selected short phrase relative to the overall volume of the track (the overall volume of the track can remain the same or similar for other short phrases of the track). Providing feedback can affect parameters regarding the inclusion of the short phrase in future mixes (such as the likelihood of inclusion, the number of times to loop when included, adjustments to the short phrase for inclusion, etc.).
[0050] Note that the interface can display various information about the selected short phrase, such as a graph of amplitude over time, frequency information, rhythm, pitch, number of loops, etc.
[0051] Figure 6 is an illustration showing an example mix - level interaction according to some embodiments. In the example shown, the user has selected an outer bubble corresponding to the entire mix. The interface shows the current segment being composed (in this example, an "intensity buildup"). The interface also provides thumbs - up and thumbs - down elements 602 and 604 for the user to provide feedback on the current segment (although various feedback interfaces can be implemented). The user's feedback on the composition segment can be used to adjust parameters regarding the inclusion of the segment, such as the number of times it is included, the length it is played when included, the volume when included, and its inclusion with other segments (e.g., based on segments for which the user has previously provided feedback).
[0052] In some embodiments, the interface is configured to display one or more tracks that are not currently being mixed. These tracks can be represented using a different color, line type, or other visual differences from the tracks currently being mixed. These tracks can also be displayed separately from the tracks being mixed (whereas in some embodiments, tracks included in the mix are shown as touching or overlapping bubbles). The unmixed tracks shown can be determined by the composition machine - learning engine to be suitable for the current mix but may not currently be included due to, for example, statistical rules or thresholds regarding the tracks included in the particular type of segment currently being played.
[0053] The user can select and drag the bubbles of these tracks, for example, to include those tracks in the mix. Likewise, the user can listen to the tracks that are not currently being mixed separately and provide feedback on these tracks, which can be used to control the future inclusion of such tracks.
[0054] It should be noted that the various functions described herein can be performed on the client side (e.g., on a user device), the server side, or both. As an example of splitting functions, a client device can perform certain public functions, such as displaying a public interface, while a server can implement adjustment of composition parameters based on user input.
[0055] Example interface screenshots
[0056] This will be discussed in detail below Figures 7A-10B , Figures 7A-10B Example screenshots from an example generative music application embodiment are shown. In the example shown, a version of the AiMi application displays interface elements associated with example "ambient" and "chill" music experiences.
[0057] Fig. 7A An example interface is shown showing the audio tracks currently being mixed. Figure 7B The user's selection of the overall mix and a thumbs up / down element for receiving user input are shown. The interface also shows the current composition segment ("build"). Figure 7B The user's selection of the current composition segment ("sustain") is shown and the preceding and following segments (progressive segment and climax segment (drop)) are shown. In this example, the user can provide feedback on the individual segments themselves through a thumbs up / down interface. Figures 8A-8C An example subsequent interface after initial user feedback is shown.
[0058] exist Fig. 8A In the example, the user selected the thumbs down option for the current segment ("Jam Drop"). The interface provides additional elements for more granular feedback. In this example, the user can choose to play the composition segment at shorter intervals or less often in future compositions. Figure 8B In the interface, the following instructions are provided: The improvisation climax will be extended by an additional 10 seconds. Figure 8C In the example, the interface provides an indication that the climax improvisation will be played less often. In general, the interface can provide various indications of actions that have been taken based on user feedback about a mix, track, segment, etc.
[0059] Fig.9A An interface is shown in response to a user's selection of a click track. In this example, the click track and other tracks are displayed relative to each other according to the user's selection of the track. Fig. 7A Resized and moved. Fig. 9B A user's selection of a thumbs down element for a click track is shown, along with an indication that AiMi will play fewer tracks like this one (referred to in this interface as musical "ideas"). Fig. 9C A user selection of a thumbs-up element is shown along with an indication that more musical ideas like this one will be played in future compositions.
[0060] Fig. 10A and 10B Example interface showing selection of the click track and overall mix for the composition "Soothing". As shown, the interface can use different colors to match different compositions. Example interface with more granular user-adjustable parameters
[0061] This will be discussed in detail below Fig.11A -C, Fig.11A -C shows example fine track adjustments according to some embodiments. Note that in these embodiments, each bubble has a fixed size in the illustrated interface, but one or more inner bubbles display the current value of a parameter (e.g., gain), a user-indicated parameter value (e.g., maximum gain), or some combination of the two. In other embodiments, the inner bubbles may represent multiple different parameters.
[0062] Similar to the previous example, the user interface of the generative music application displays the track of the current segment of generative music (in this example, the "high riff segment"). In some embodiments, the interface also displays (and may allow modification of) the next segment to be played. Various interface elements can be used to modify one or more properties of a particular track (in this example, the gain of the "beat" track).
[0063] Fig.11A 1 is a block diagram illustrating an example interface that displays the currently mixed tracks in a generative music clip. As shown, each track (represented by a bubble, e.g., circle 1110) includes two concentric circles. In this example, the solid circle 1112 reflects the current gain of the track, while the dashed circle 1114 reflects the maximum gain of the track (e.g., based on a default maximum value or based on previous user input). As will be seen in reference to Fig. 11C As described above, the user can define the maximum gain using interface elements. Note that Fig.11A The dotted line in Figure 3 The dashed lines are different from those shown in: Fig. 11CThe dashed circle shows the maximum gain of the track (and may actually be displayed in the user interface), while Figure 3 The dashed circle shows the initial volume of its corresponding track before the change (and may not be shown after the change).
[0064] Fig. 11B An example interface for changing various properties of a selected track in a generative music clip is shown. In particular, Fig. 11B Included are thumbs up and down button elements 1120 and 1122 , a slider 1124 , and circles 1126 and 1128 . Fig. 11B The interface may include more or fewer interface elements than shown.
[0065] The user can modify the maximum gain of the track (represented by dashed circle 1126). In some embodiments, the user drags slider 1124 (e.g., using a touch screen or mouse cursor), which updates the maximum gain value (displayed as 80%) and the radius of dashed circle 1126. In some embodiments, the maximum gain value is used to limit changes that the system may make to the corresponding track, for example, when the gain is changed based on a composition algorithm. The maximum gain value can also limit the impact of other changes made by the user (e.g., an adjustment to another parameter may be limited so that it does not cause the gain of the beat to exceed the indicated maximum value). The user can also use buttons 1120 and 1122 to provide feedback related to the selected track (and potentially modify the selected track), as shown in Figure 1. Figure 4 As described.
[0066] Although the property modified in the illustrated figure is maximum gain, other properties of the track (e.g., current gain, tempo, intensity, effects, syncopation, volume, instruments, etc.) can be similarly modified using similar or different interface elements. For example, a user clicking on slider 1124 can display multiple sub-sliders, each of which can adjust a specific property of the track (possibly including gain). In addition, interaction with the UI can affect other properties besides the properties shown in FIG. 11G, such as related properties in other tracks and / or loops.
[0067] In some embodiments (not explicitly shown), a given bubble / circle at one interface level may correspond to a group of tracks or instruments (which may be referred to as channels). In these embodiments, a user may select a group (e.g., by double-clicking), and the interface may display a different level with an expanded view of the channels in that group. The user may then provide independent input for the channels. For example, a melody group may include two melody channels. The user may select a melody group and then adjust the gain (or other sub-slider parameter) of one of the melodies.
[0068] Fig. 11CDepicted is an example interface for playing a selected track in a generative music clip. As shown, the "Beat" track is selected, causing the system to play only the audio of that track. As shown, this selection reduces the visibility of other bubbles. The user interface can include additional control buttons (e.g., play, pause, share, or record buttons) to control the track being played.
[0069] Fig.12 is a flow chart illustrating an example method according to some embodiments. Fig.12 The method shown can be used in conjunction with any computer circuit, system, device, element or component disclosed herein. In various embodiments, some of the method elements shown can be performed in parallel, in a different order than shown, or can be omitted. Additional method elements can be performed as needed.
[0070] At 1210, in the illustrated embodiment, the computing system selects a set of audio tracks to include in the generative music content.
[0071] At 1220, in the illustrated embodiment, the computing system determines gain values for each of the plurality of selected audio tracks.
[0072] At 1230, in the illustrated embodiment, the computing system mixes the plurality of selected audio tracks based on the determined gain values to generate output music content.
[0073] At 1240, in the illustrated embodiment, the computing system causes an interface to be displayed that visually indicates the selected audio track and its determined gain value relative to the other selected audio tracks. In some embodiments, the displayed visual indication of the selected audio track includes a bubble element for each audio track.
[0074] At 1250, in the illustrated embodiment, the computing system receives via the interface a user gain input indicating an adjustment of a gain value of one of the selected tracks. In some embodiments, the user gain input adjusts the size of the displayed visual track indication.
[0075] At 1260, in the illustrated embodiment, the computing system adjusts the mix of the selected audio track based on the user input.
[0076] In some embodiments, the computing system is responsive to user input selecting the displayed visual track indication so that an interface option for playing the selected track alone is displayed. In some embodiments, the interface also visually indicates for a given track: both the current gain value of the track and the user-indicated gain level of the track. In some embodiments, the user-indicated gain level is the maximum gain level of the track. In some embodiments, the interface includes a plurality of user interface elements (e.g., a set) for the user-indicated track that can be adjusted by the user to modify one or more additional composition parameters for the user-indicated track. Adjustments to the mix can also be based on one or more modified additional composition parameters for the user-indicated track.
[0077] In some embodiments, the computing system, in response to user input selecting the displayed visual track indication, causes a feedback interface to be displayed and user feedback input to be received (the user feedback input providing user feedback on the selected track), and adjusts parameters for selecting future tracks for mixing based on the user feedback input. The computing system may determine a vector of the tracks, wherein the vector is generated by an unsupervised machine learning model, and adjusting the parameters for selecting future tracks includes increasing or decreasing the selection of the tracks when the selected tracks are within a threshold distance in the vector space.
[0078] In some embodiments, the computing system causes a phrase interface to be displayed in response to user input selecting the displayed visual track indication, the phrase interface being configured to receive user phrase input regarding a phrase included in the track. The computing system can adjust the inclusion of the phrase in future versions of the track based on the user phrase input.
[0079] In some embodiments, the computing system causes an indication of the currently playing composition segment to be displayed in response to user input selecting the displayed track mix. The interface can visually indicate one or more additional tracks that are not currently mixed but are suitable for mixing with the selected tracks that are already mixed. The computing system can adjust one or more parameters for future inclusion of the composition segment based on user feedback input corresponding to the composition segment.
[0080] The various techniques described herein may be performed by one or more computer programs. The term "program" should be broadly interpreted to encompass a sequence of instructions in a programming language executable by a computing device. These programs may be written in any suitable computer language, including low-level languages (such as assembly language) and high-level languages (such as Python). Programs may be written in compiled languages (such as C or C++) or in interpreted languages (such as JavaScript).
[0081] Program instructions may be stored on a "computer-readable storage medium" or "computer-readable medium" to facilitate the execution of the program instructions by a computer system. In general, these phrases include any tangible or non-transitory storage or memory medium. The terms "tangible" and "non-transitory" are intended to exclude propagating electromagnetic signals, but do not otherwise limit the type of storage medium. Thus, the terms "computer-readable storage medium" or "computer-readable medium" are intended to cover types of storage devices that do not necessarily store information permanently (such as random access memory (RAM)). Accordingly, the term "non-transitory" is a limitation on the nature of the medium itself (i.e., the medium cannot be a signal), rather than a limitation on the persistence of the medium's data storage (such as RAM vs. ROM).
[0082] The terms "computer-readable storage medium" or "computer-readable medium" are intended to refer to storage media internal to a computer system as well as removable media such as a CD-ROM, memory stick, or portable hard drive. These phrases encompass any type of volatile memory in a computer system, including DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc., as well as non-volatile memory such as magnetic media (e.g., hard drives) or optical storage devices. These phrases are intended to explicitly encompass the memory of a server that facilitates the download of program instructions, the memory of any intermediate computer systems involved in the download process, and the memory of all target computing devices. Furthermore, these phrases are intended to encompass combinations of different types of memory.
[0083] In addition, the computer-readable medium or storage medium may be located in a first group of one or more computer systems executing the program, and may also be located in a second group of one or more computer systems connected to the first group via a network. In the latter case, the second group of computer systems may provide program instructions to the first group of computer systems for execution. In short, the phrase "computer-readable storage medium" or "computer-readable medium" may include two or more media that may be located in different locations, such as in different computers connected via a network.
[0084] The present disclosure includes references to "one embodiment" or groups of "embodiments" (e.g., "some embodiments" or "various embodiments"). Embodiments are different implementations or instances of the disclosed concepts. References to "one embodiment," "an embodiment," "a particular embodiment," etc. do not necessarily refer to the same embodiment. A large number of possible embodiments are contemplated, including those specifically disclosed, as well as modifications or substitutions that fall within the spirit or scope of the present disclosure.
[0085] The present disclosure may discuss potential advantages that may arise from the disclosed embodiments. Not all implementations of these embodiments necessarily exhibit any or all potential advantages. Whether an advantage is achieved for a particular implementation depends on many factors, some of which are beyond the scope of the present disclosure. In fact, there are many reasons why an implementation that falls within the scope of the claims may not exhibit some or all of any disclosed advantages. For example, a particular implementation may include other circuits outside the scope of the present disclosure that, in combination with one of the disclosed embodiments, negate or reduce one or more of the disclosed advantages. In addition, suboptimal design execution (e.g., implementation techniques or tools) of a particular implementation may also negate or reduce the disclosed advantages. Even assuming a skilled implementation, the realization of the advantages may still depend on other factors, such as the environmental conditions in which the implementation is deployed. For example, the input provided to a particular implementation may prevent one or more problems solved in the present disclosure from occurring in a particular occasion, and as a result, the benefits of its solution may not be achieved. In view of the possible factors outside the present disclosure, it is expressly stated that any potential advantages described herein should not be interpreted as claim limitations that must be met to prove infringement. Instead, the identification of such potential advantages is intended to illustrate the various types of improvements available to designers who benefit from the present disclosure. Describing these advantages generously (e.g., stating that a particular advantage “may occur”) is not intended to express doubt that these advantages can actually be achieved, but rather to recognize that achieving these advantages often depends on the technical realities of additional factors.
[0086] Unless otherwise stated, the embodiments are non-limiting. That is, the disclosed embodiments are not intended to limit the scope of claims drafted based on the present disclosure, even if only a single example is described therein with respect to a particular feature. The disclosed embodiments are intended to be illustrative rather than restrictive, unless there is any contrary statement in the present disclosure. Therefore, the present application is intended to allow claims covering the disclosed embodiments and such alternatives, modifications, and equivalents that are obvious to those skilled in the art having the benefit of the present disclosure.
[0087] For example, the features in this application may be combined in any suitable manner. Therefore, during the examination of this application (or an application claiming priority thereto), new claims may be formulated for any such feature combinations. In particular, with reference to the attached claims, features from dependent claims may be combined with features of other dependent claims, including claims that are dependent on other independent claims, where appropriate. Similarly, features from individual independent claims may be combined as appropriate.
[0088] Thus, while the attached dependent claims may be drafted so that each is dependent on a single other claim, additional dependencies are contemplated. Combinations of features in any dependent claims consistent with the present disclosure are contemplated and may be claimed in this or another application. In short, the combinations are not limited to those specifically listed in the attached claims.
[0089] It is also contemplated that claims drafted in one format or legal type (eg, apparatus) are intended to support corresponding claims in another format or legal type (eg, method), where appropriate.
[0090] Because this disclosure is a legal document, various terms and phrases may be subject to administrative and judicial interpretation. It is hereby notified that the definitions provided in the following paragraphs as well as throughout the disclosure will be used to determine how to interpret claims drafted based on this disclosure.
[0091] Reference to an item in the singular (i.e., a noun or noun phrase preceded by "a," "an," or "the") is intended to mean "one or more" unless the context clearly dictates otherwise. Thus, reference to "an item" in a claim does not exclude other instances of that item in the absence of accompanying context. A "plurality" of an item refers to a group of two or more items.
[0092] The word "may" is used herein in a permissive sense (ie, having the potential to, being able to), rather than the mandatory sense (ie, must).
[0093] The terms "include" and "including" and their various forms are open ended and mean "including, but not limited to."
[0094] When the term "or" is used in this disclosure with respect to a list of options, it will generally be understood to be used in an inclusive sense unless the context dictates otherwise. Thus, the recitation of "x or y" is equivalent to "x or y, or both," thus covering 1) x but not y, 2) y but not x, and 3) x and y. On the other hand, phrases such as "either x or y, but not both" clearly indicate that the use of "or" is exclusive.
[0095] The description "w, x, y, or z, or any combination thereof" or "at least one of w, x, y, and z ..." is intended to cover all possibilities involving individual elements, up to the total number of elements in the set. For example, given the set [w, x, y, z], these phrases cover any individual element in the set (e.g., w but not x, y, or z), any two elements (e.g., w and x but not y or z), any three elements (e.g., w, x, and y but not z), and all four elements. The phrase "at least one of w, x, y, and z ..." thus refers to at least one element in the set [w, x, y, z], thereby covering all possible combinations in that list of elements. The phrase should not be interpreted as requiring the presence of at least one instance of w, at least one instance of x, at least one instance of y, and at least one instance of z.
[0096] In this disclosure, various "labels" may precede a noun or noun phrase. Unless the context dictates otherwise, different labels (e.g., "first circuit," "second circuit," "particular circuit," "given circuit," etc.) used for a feature refer to different instances of that feature. Furthermore, the labels "first," "second," and "third" do not imply any type of ordering (e.g., spatial, temporal, logical, etc.) when applied to features, unless otherwise noted.
[0097] The phrase "based on" is used to describe one or more factors that influence a determination. The term does not exclude the possibility that other factors may influence the determination. That is, a determination may be based only on the specified factors or on the specified factors as well as other unspecified factors. Consider the phrase "A is determined based on B." The phrase specifies that B is a factor used to determine A or influences the determination of A. The phrase does not exclude that the determination of A may also be based on some other factor, such as C. The phrase is also intended to cover embodiments in which A is determined based only on B. As used herein, the phrase "based on" is synonymous with the phrase "based at least in part on."
[0098] The phrases "in response to" and "in response to" describe one or more factors that trigger an effect. The phrase does not exclude the possibility that other factors may influence or otherwise trigger the effect, either together with or independently of the specified factors. That is, the effect may be in response to these factors alone, or it may be in response to the specified factors and other unspecified factors. Consider the phrase "in response to B, A is performed." This phrase specifies that B is the factor that triggers the performance of A, or the factor that triggers a specific result of A. This phrase does not exclude that the performance of A may also be in response to some other factor, such as C. This phrase also does not exclude that the performance of A may be jointly in response to B and C. The phrase is also intended to cover embodiments in which A is performed in response to B alone. As used herein, the phrase "in response to" is synonymous with the phrase "at least in response to." Similarly, the phrase "in response to" is synonymous with the phrase "at least in response to."
[0099] In the present disclosure, different entities (which may be variously referred to as "units", "circuits", other components, etc.) may be described or claimed as being "configured" to perform one or more tasks or operations. This expression - [the entity] is configured to [perform one or more tasks] - is used herein to refer to a structure (i.e., something physical). More specifically, the expression is used to indicate that the structure is arranged to perform one or more tasks during operation. A structure can be said to be "configured to" perform some tasks, even if the structure is not currently being operated. Therefore, an entity described or recited as being "configured to" perform certain tasks refers to something physical, such as a device, a circuit, a system having a processor unit and a memory storing program instructions that can be executed to perform tasks, etc. The phrase is not used herein to refer to something intangible.
[0100] In some cases, various units / circuits / components may be described herein as performing a set of tasks or operations. It should be understood that even if not specifically stated, these entities are also "configured to" perform these tasks / operations.
[0101] The term "configured to" does not mean "configurable to". For example, an unprogrammed FPGA would not be considered "configured to" perform a specific function. However, such an unprogrammed FPGA may be "configurable to" perform that function. After being properly programmed, the FPGA can be said to be "configured to" perform that specific function.
[0102] For purposes of a U.S. patent application based on the present disclosure, reciting in a claim that a structure is “configured to” perform one or more tasks expressly indicates that no 35 U.S.C. §112(f) is invoked with respect to that claim element. If an applicant wishes to invoke section 112(f) during prosecution of a U.S. patent application based on the present disclosure, it would recite the claim element using “means for [performing the function].”
[0103] Different "circuits" may be described in the present disclosure. These circuits or "circuitry systems" constitute hardware and include various types of circuit elements, such as combinational logic, clock storage devices (e.g., flip-flops, registers, latches, etc.), finite state machines, memories (e.g., random access memory, embedded dynamic random access memory), programmable logic arrays, etc. The circuits can be custom designed or obtained from standard libraries. In various embodiments, the circuits may appropriately include digital components, analog components, or a combination of both. Specific types of circuits are often referred to as "units" (e.g., decoding units, arithmetic logic units (ALUs), functional units, memory management units (MMUs), etc.). Such units are also referred to as circuits or circuit systems.
[0104] Thus, the disclosed circuits / units / components and other elements shown in the figures and described herein include hardware elements, as described in the previous paragraph. In many instances, the internal arrangement of hardware elements within a particular circuit can be specified by describing the functionality of the circuit. For example, a particular "decode unit" may be described as performing the function of "processing an opcode for an instruction and routing the instruction to one or more of a plurality of functional units," meaning that the decode unit is "configured to" perform that function. For a person skilled in the computer arts, such a functional specification is sufficient to suggest a set of possible structures for the circuit.
[0105] In various embodiments, as described in the previous paragraph, circuits, units, and other elements may be defined by the functions or operations they are configured to implement. The arrangement and the relationship between such circuits / units / components and the way they interact form a microarchitecture definition of hardware, which is ultimately manufactured in an integrated circuit or programmed into an FPGA to form a physical implementation of the microarchitecture definition. Therefore, those skilled in the art recognize that the microarchitecture definition is a structure from which many physical implementations can be derived, all of which belong to the broader structure described by the microarchitecture definition. That is, a technician presented by the microarchitecture definition provided by the present disclosure can implement the structure by encoding the description of the circuit / unit / component in a hardware description language (HDL) (e.g., Verilog or VHDL) without excessive experimentation and under the application of ordinary technicians. HDL descriptions are often expressed in a way that may appear functional. However, for those skilled in the art, such HDL descriptions are a way of converting the structure of a circuit, unit, or component into the next level of implementation details. Such HDL descriptions can take the form of behavioral code (usually not synthesizable), register transfer language (RTL) code (unlike behavioral code, RTL code is usually synthesizable) or structural code (e.g., a netlist specifying logic gates and their connections). The HDL description can then be synthesized against a library of cells designed for a given integrated circuit manufacturing technology, and can be modified for timing, power, and other reasons to produce a final design database that is transmitted to a foundry to generate masks and ultimately produce integrated circuits. Some hardware circuits or portions thereof may also be custom designed in a schematic editor and captured into the integrated circuit design along with the synthesized circuits. An integrated circuit may include transistors and other circuit elements (e.g., passive elements such as capacitors, resistors, inductors, etc.) and interconnections between transistors and circuit elements. Some embodiments may implement multiple integrated circuits coupled together to implement hardware circuits, and / or discrete elements may be used in some embodiments. Alternatively, the HDL design may be synthesized into a programmable logic array, such as a field programmable gate array (FPGA), and may be implemented in an FPGA. This decoupling between the design of a set of circuits and the subsequent low-level implementation of these circuits often results in a situation where the circuit or logic designer never specifies a specific set of structures for the low-level implementation beyond the description of what the circuit is configured to do, because the process is performed at different stages of the circuit implementation process.
[0106] The fact that many different low-level combinations of circuit elements can be used to achieve the same specification of a circuit results in a large number of equivalent structures for that circuit. As mentioned above, these low-level circuit implementations can vary depending on variations in manufacturing technology, the foundry chosen to fabricate the integrated circuit, the cell libraries provided for a particular project, etc. In many cases, the choices made by different design tools or methodologies can be arbitrary.
[0107] Furthermore, for a given embodiment, a single implementation of a circuit's specific functional specifications typically includes a large number of devices (e.g., millions of transistors). Thus, the sheer volume of this information makes it impractical to provide a complete description of the low-level structure for implementing a single embodiment, let alone the large number of equivalent possible implementations. Therefore, this disclosure uses functional shorthand commonly used in the industry to describe the structure of a circuit.
[0108] In the present disclosure, various "modules" operable to perform specified functions are shown in the figures and described in detail. As used herein, "module" refers to software or hardware operable to perform a set of specific operations. A module may refer to a set of software instructions that can be executed by a computer system to perform the set of operations. A module may also refer to hardware configured to perform the set of operations. Hardware modules may constitute general-purpose hardware and non-volatile computer-readable media storing program instructions, or special-purpose hardware such as custom ASICs. Therefore, a module described as "executable" to perform an operation refers to a software module, and a module described as "configured" to perform an operation refers to a hardware module. A module described as "operable" to perform an operation refers to a software module, a hardware module, or some combination of the two. In addition, for any discussion of modules mentioned herein that are "executable" to perform specific operations, it should be understood that in other embodiments, these operations may also be implemented by hardware modules "configured" to perform operations, and vice versa.
Claims
1. A method, include: selecting, by a computing system, a set of musical phrases for inclusion in the generative musical content; Determining, by the computing system, a gain value for each of a plurality of selected musical phrases; mixing, by the computing system, the selected phrases based on the determined gain values to generate output music content; causing, by the computing system, to display an interface that visually indicates the selected phrase and the determined gain value of the selected phrase relative to other selected phrases; receiving, by the computing system via the interface, a user gain input indicating adjustment of a gain value of one of the selected phrases; as well as A mix of the selected phrase is adjusted by the computing system based on the user input.
2. The method according to claim 1, further comprising: include: In response to user input selecting a displayed visual phrase indication, the computing system causes display of an interface option for soloing the selected phrase.
3. The method according to claim 1, further comprising: include: In response to user input selecting the displayed visual phrase indication, the computing system causes display of a feedback interface and receives user feedback input providing feedback on the selected phrase; as well as Based on the user feedback input, parameters for selecting future phrases for mixing are adjusted.
4. The method according to claim 3, further comprising: include: determining a vector of a musical phrase, wherein the vector is generated by an unsupervised machine learning model; Wherein, adjusting the parameters for selecting future phrases includes: increasing or decreasing the selection of phrases within a threshold distance of the selected phrases in the vector space.
5. The method according to claim 1, further comprising: include: In response to user input selecting the displayed visual phrase indication, a phrase interface is displayed, the phrase interface being configured to receive user phrase input regarding a phrase contained in the phrase.
6. The method according to claim 5, further comprising: include: Based on the user phrase input, inclusion of the phrase in future versions of the phrase is adjusted.
7. The method according to claim 1, in, The user gain input adjusts the size of the displayed visual phrase indication.
8. The method according to claim 1, in, The displayed visual indication of the selected phrases includes a bubble element for each phrase.
9. The method according to claim 1, further comprising: include: In response to user input selecting the displayed phrase mixture, the computing system causes display of an indication of the composition section currently being played.
10. The method according to claim 9, further comprising: include: Based on user feedback input corresponding to the composition segment, one or more parameters for future inclusion of the composition segment are adjusted.
11. The method according to claim 9, in, The interface visually indicates one or more additional phrases that are not currently mixed but are suitable for mixing with the already mixed selected phrases.
12. The method according to claim 1, in, The interface also visually indicates, for a given phrase, both a current gain value for the phrase and a user-indicated gain level for the phrase.
13. The method according to claim 12, in, The user-indicated gain level is the maximum gain level for the phrase.
14. The method according to claim 1, in: The interface includes a plurality of user interface elements for the user-indicated musical phrase, the plurality of user interface elements being adjustable by the user to modify one or more additional composition parameters for the user-indicated musical phrase; and The adjustments are also based on one or more modified additional composition parameters for the user-indicated phrase.
15. The method according to claim 1, in, The interface includes a group of the selected phrases, the group being selectable by a user to specify a user gain input for each phrase in the group.
16. A device, include: one or more memories; as well as one or more processors configured to execute program instructions stored on the one or more memories to: selecting a set of musical phrases to include in the generative musical content; determining gain values for each of a plurality of selected phrases; mixing the selected phrases based on the determined gain values to generate output music content; causing display of an interface that visually indicates the selected phrase and the determined gain value of the selected phrase relative to other selected phrases; receiving, via the interface, a user gain input indicating adjustment of a gain value of one of the selected phrases; as well as A mix of the selected phrase is adjusted based on the user input.
17. A non-transitory computer readable medium having stored thereon instructions executable by a computing device to perform operations, the operations include: selecting a set of musical phrases to include in the generative musical content; determining gain values for each of a plurality of selected phrases; mixing the selected phrases based on the determined gain values to generate output music content; causing display of an interface that visually indicates the selected phrase and the determined gain value of the selected phrase relative to other selected phrases; receiving, via the interface, a user gain input indicating adjustment of a gain value of one of the selected phrases; as well as A mix of the selected phrase is adjusted based on the user input.
18. The non-transitory computer readable medium of claim 17, in, The operations also include causing display of an interface option for soloing the selected phrase in response to user input selecting the displayed visual phrase indication.
19. The non-transitory computer readable medium of claim 17, in, The operations also include: In response to user input selecting the displayed visual phrase indication, causing a feedback interface to be displayed and receiving user feedback input providing feedback on the selected phrase; and Parameters for selecting future phrases for mixing are adjusted based on the user feedback input.
20. The non-transitory computer readable medium of claim 17, in, The interface also visually indicates, for a given phrase, both a current gain value for the phrase and a user-indicated gain level for the phrase.
Citation Information
Patent Citations
Music generator
US10679596B2
Music generator
US20190362696A1
Music Content Generation Using Image Representations of Audio Files
US20210248983A1
Music generator
US8812144B2