Generating audio-based and / or audiovisual-based music content using generative models
Through the unified user interface and generative model processing multimodal input, synchronized lyrics and music creation content is generated, and resource waste and matching problems caused by multi-model interaction are solved, and efficient music content generation is achieved.
Patent Information
- Application Number
- CN202510493223.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-19
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, users need to interact with multiple generative models to generate music content, resulting in wasted computing and network resources, and the generated lyrics and music creation content often cannot be objectively matched, requiring additional synchronization processing.
Using a single or multiple generative models, multimodal input is processed through a unified user interface, lyric content and music creation content are generated, and the seed and synchronization verification engines are used to ensure content synchronization, reducing user interaction and post-processing needs.
It realizes the rapid and efficient generation of synchronized audio and audio-visual music content on limited hardware devices, reduces user input, improves interaction efficiency, and ensures that the content matches user preferences.
Smart Images

Figure CN120340446A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to using generative models to generate audio-based music content and / or audiovisual-based music content. Background Art
[0002] Various generative models (GMs) have been proposed, which can be used to process user inputs to generate outputs that reflect generative content in response to the user inputs. For example, large language models (LLMs) have been developed, which can be used to process user inputs to generate LLM outputs that reflect text-based generative content in response to the user inputs. Further, music generative models have been developed, which can be used to process user inputs to generate music generative outputs that reflect audio-captured music in response to the user inputs. Additionally, image and video generative models have been developed, which can be used to process user inputs to generate image and / or video generative outputs that reflect image- and / or video-based generative content in response to the user inputs.
[0003] However, in many cases, users have to interact with various different GMs to obtain the generative content. For example, assume that a user wants to generate music content. In this example, the user can interact with an LLM to generate lyric content for the music content and a music generative model for music composition content. However, users interacting with these different GMs typically need to perform different interactions with these different GMs to obtain the desired music content, which wastes computing resources by requiring these different interactions and also wastes network resources because these different GMs are typically executed at remote servers due to their size. Further, the lyric content and music composition generated by these different GMs may be meaningless because the lyric content may not objectively match the music composition content, rendering the lyric content and music composition content unusable, thus wasting computing and / or network resources during these different interactions. Additionally, even if the lyric content generated using an LLM does objectively match the music composition generated using a music generative model, additional processing may be required to properly synchronize the lyric content and the music composition content. Summary of the Invention
[0004] The implementations described herein involve using a generative model (GM) to generate audio-based music content that includes at least lyric content and music composition content. In some implementations, the audio-based music content can further include visual multimedia content (e.g., generative or non-generative visual multimedia content), resulting in audio-visual based music content. The processor of the system can: receive user input associated with the user's client device, the user input including a request for music content; generate the music content; and cause the music content to be rendered at the client device. In some implementations, the processor can cause a single GM to process GM input (including at least the user input) to generate a GM output, and can determine the lyric content and the music composition content based on the GM output. In other implementations, the processor can cause multiple GMs to process corresponding GM inputs (each GM input including at least the user input) to generate corresponding GM outputs, and can determine the lyric content and the music composition content based on the corresponding GM outputs. In various implementations, the processor can receive additional user input associated with the user's client device, the additional user input including a request to modify the music content.
[0005] In implementations where a single GM is used to process GM input to generate a GM output, the single GM can be a multimodal GM that is fine-tuned to receive multimodal inputs such as text-based user input, audio-based user input, and / or visual-based user input, and is fine-tuned to generate multimodal outputs such as text-based output, audio-based output, and / or visual-based output. Some examples of multimodal GMs that can receive multimodal inputs and generate multimodal outputs are Bard, Gemini, GPT, etc. Thus, when using a single GM to process GM input (including user input and optionally including other context, prompts, etc.), the GM output can include various probability distributions over sequences of tokens. For example, when determining the lyric content, the processor can employ various decoding techniques from a sequence of words or word units (e.g., text-based output) or from a sequence of phonemes or articulatory units (e.g., audio-based output), and determine the lyric content based on the probability distribution over the sequence of words or word units or over the sequence of phonemes or articulatory units. Further, when determining the lyric content, the processor can employ various decoding techniques from a sequence of notes or note units and determine the music composition content based on the probability distribution over the sequence of notes or note units.
[0006] One or more technical advantages can be achieved by leveraging a single GM as described herein to generate audio-based music content and / or audiovisual-based music content. As a non-limiting example, a single unified user interface is utilized to enable a user to provide simplified user input to generate audio-based music content and / or audiovisual-based music content. Thus, the user does not need to interact with multiple GMs to generate audio-based music content and / or audiovisual-based music content. These techniques are particularly advantageous given the hardware constraints of some client devices. For example, assume the user's client device is a mobile device of the user with a limited display size (e.g., relative to the display of a laptop or desktop computer). In this case, the single unified user interface enables the user to provide simplified user input to generate audio-based music content and / or audiovisual-based music content without the user having to switch between GM applications, between tabs of a web browser application, etc. to generate audio-based music content and / or audiovisual-based music content, thereby reducing the amount of user input received at the mobile device and enabling more rapid and efficient interaction between the user and the mobile device. As another non-limiting example, and as a result of the single GM being fine-tuned to generate audio-based music content and / or audiovisual-based music content, the need for post-processing audio-based music content and / or audiovisual-based music content to ensure its synchronization is eliminated.
[0007] In an implementation where multiple GMs are utilized to process respective GM inputs to generate respective GM outputs, each of the multiple GMs can be a unimodal GM and / or a multimodal GM, the unimodal GM and / or multimodal GM being jointly fine-tuned to receive respective unimodal or multimodal inputs and being jointly fine-tuned to generate respective outputs. As described above, some examples of multimodal GMs that are capable of receiving multimodal inputs and generating multimodal outputs are Bard, Gemini, GPT, etc. Further, an example of a unimodal GM that is capable of receiving unimodal input and generating an audio-based output is AudioLM; some examples of unimodal GMs that are capable of receiving unimodal input and generating text-based outputs are PaLM, LaMDA, etc.; and some examples of unimodal GMs that are capable of receiving unimodal input and generating vision-based outputs are Imagen, Dall-E, Sora, etc. Thus, the respective GM inputs (including user input and optionally including other contexts, prompts, etc.) can be customized for the respective multiple GMs to generate respective GM outputs, and each of the respective GM outputs can include a respective probability distribution over a sequence of tokens in the same or a similar manner as described above.
[0008] One or more technical advantages can be achieved by leveraging multiple GMs as described herein to generate audio-based music content and / or audiovisual-based music content. As a non-limiting example, a single unified user interface is utilized to enable a user to provide simplified user input to generate audio-based music content and / or audiovisual-based music content. Even though the multiple GMs are different GMs in these implementations, the user only needs to provide a single user input to invoke calls to each of the multiple different GMs, such that the user may not even be aware that multiple GMs are being utilized to generate audio-based music content and / or audiovisual-based music content. These techniques are particularly advantageous given the hardware constraints of some client devices. For example, assume the user's client device is the user's mobile device with a limited display size (e.g., relative to the display of a laptop or desktop computer). In this case, the single unified user interface enables the user to provide simplified user input to generate audio-based music content and / or audiovisual-based music content without the user having to switch between GM applications, between tabs of a web browser application, etc. to generate audio-based music content and / or audiovisual-based music content, thereby reducing the amount of user input received at the mobile device and enabling a faster and more efficient interaction between the user and the mobile device. As another non-limiting example, and as a result of the multiple GMs being jointly fine-tuned to generate audio-based music content and / or audiovisual-based music content, the need for post-processing audio-based music content and / or audiovisual-based music content to ensure its synchronization is eliminated.
[0009] In implementations where additional user input is received that includes a request to modify audio-based music content and / or audiovisual-based music content associated with the user's client device, the processor can determine a seed to be utilized when processing the additional user input based on the additional user input and based on previously rendered music content. The seed can be a corresponding lower-level representation of the previously rendered lyric content and / or music composition content. For example, the corresponding lower-level representation of the lyric content and / or music composition content can be a corresponding embedding in a corresponding embedding space. Thus, if the additional user input requests modification of the lyric content, but the music composition content remains the same, when processing the additional user input and the seed, the seed will ensure that the lyric content is modified (e.g., as requested by the user), but the music composition content will not be modified. Similarly, if the additional user input requests modification of the music composition content, but the lyric content remains the same, when processing the additional user input and the seed, the seed will ensure that the music composition content is modified (e.g., as requested by the user), but the lyric content will not be modified. It should be understood that the seed determined by the processor will be based on how the additional user input requests modification of the music content.
[0010] One or more technical advantages can be achieved by modifying audio-based music content and / or audiovisual-based music content by utilizing seeds as described herein. As a non-limiting example, a seed can constrain the extent to which audio-based music content and / or audiovisual-based music content is modified based on additional user input. As a result, the seed enables a user to quickly and efficiently modify audio-based music content and / or audiovisual-based music content without the user having to reprompt these GMs with detailed instructions about what they like and / or dislike about the audio-based music content and / or audiovisual-based music content. As a result, the length of any additional user input that is processed to modify the audio-based music content and / or audiovisual-based music content is reduced because the additional user input and the determined seed (which can be a lower-level representation of the music content) automatically embed that information, thereby saving computational resources and network resources when modifying the audio-based music content and / or audiovisual-based music content. Further, in the absence of using a seed as described herein when modifying the audio-based music content and / or audiovisual-based music content, any resulting music content subsequently generated may be significantly different from the original music content that was rendered for presentation to the user.
[0011] In an implementation where the lyric content is audibly rendered, the processor can optionally cause the lyric content to be audibly rendered in the voice of the user providing the user input. For example, the lyric content can correspond to text determined based on the GM output. Thus, when synthesizing the audio data that captures the lyric content, the processor can utilize the voice embedding of the user (e.g., stored in the user profile database or obtained by requesting the user to say a few words during the interaction) and / or a set of one or more prosody attributes associated with the user (e.g., stored in the user profile database or obtained by requesting the user to say a few words during the interaction) to synthesize the audio data such that the audio data is audibly perceived as being spoken or sung by the user providing the user input. As another example, the lyric content can correspond to audio data determined based on the GM output. Thus, instead of synthesizing the audio data that captures the lyric content, the system can use the voice embedding of the user and / or a set of one or more prosody attributes associated with the user to adapt the lyric content such that the lyric content is audibly perceived as being spoken or sung by the user providing the user input.
[0012] One or more technical advantages can be achieved by causing the lyrics content to be audibly rendered with the voice of the user providing the user input as described herein. As a non-limiting example, the lyrics content may better resonate with the user or an additional user (e.g., the user's child, the user's spouse, the user's friend, etc.). While what resonates with the user consuming the lyrics content will depend on the user's subjective preferences and goals, it will make the resulting lyrics content more objectively and conveniently more relevant to the user's subjective preferences.
[0013] In some implementations where the audio-based music content further includes visual multimedia content (e.g., to produce audio-visual based music content), the processor may generate generative visual multimedia content (e.g., generative images, generative videos, etc.). In these implementations, a single GM may be used or separate image / video generative models may be used to generate the generative visual multimedia content. In additional or alternative implementations where the audio-based music content further includes visual multimedia content (e.g., to produce audio-visual based music content), the processor may obtain non-generative visual multimedia content (e.g., non-generative images, non-generative videos, etc.). In these implementations, the non-generative visual multimedia content may be obtained from, for example, an image / video search system, the photo / video album of the user providing the user input, etc.
[0014] One or more technical advantages can be achieved by including visual multimedia content as described herein. As a non-limiting example, a single unified user interface is utilized to enable a user to provide simplified user input to generate audiovisual-based music content. Whether using a single GM or multiple GMs, the user only needs to provide a single user input to cause the generation of audiovisual-based music content. These techniques are particularly advantageous given the hardware constraints of some client devices. For example, assume that the user's client device is the user's mobile device with a limited display size (e.g., relative to the display of a laptop computer or a desktop computer). In this case, the single unified user interface enables the user to provide simplified user input to generate audiovisual-based music content without the user having to switch between GM applications, between tabs of a web browser application, etc. to generate audiovisual-based music content, thereby reducing the amount of user input received at the mobile device and enabling more rapid and efficient interaction between the user and the mobile device. As another non-limiting example, and as a result of a single GM being fine-tuned and / or multiple GMs being jointly fine-tuned to generate audiovisual-based music content, the need for post-processing of the audiovisual-based music content to ensure its synchronization is eliminated. As another non-limiting example, the visual multimedia content can better resonate with the user or an additional user (e.g., the user's child, the user's spouse, the user's friend, etc.). While what resonates with the user consuming the visual multimedia content will depend on the user's subjective preferences and goals, it will make the resulting visual multimedia content more objectively and conveniently relevant to the user's subjective preferences.
[0015] The above description is provided as an overview of some implementations of the present disclosure. Those implementations and further descriptions of other implementations will be described in more detail below. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 A block diagram depicting an example environment that illustrates various aspects of the present disclosure and in which some implementations disclosed herein can be implemented.
[0017] Figure 2 Depicts a processing flow for utilizing various components of an example environment from Figure 1 in accordance with various implementations.
[0018] Figure 3 A flowchart depicting an example method of generating music content using a single generative model (GM) in accordance with various implementations.
[0019] Figure 4 A flowchart depicting an example method of generating music content using multiple generative models (GMs) in accordance with various implementations.
[0020] Figure 5A , Figure 5B , Figure 5C , Figure 5D , Figure 5E and Figure 5F depict various non - limiting examples of generating music content according to various implementations.
[0021] Figure 6 depicts an example architecture of a computing device according to various implementations. DETAILED DESCRIPTION
[0022] Turning now to Figure 1 , a block diagram of an example environment is depicted that illustrates aspects of the present disclosure and in which the implementations disclosed herein can be implemented. The example environment includes a client device 110 and a generative content system 120. In some implementations, all or some aspects of the generative content system 120 can be implemented locally at the client device 110. In additional or alternative implementations, all or some aspects of the generative content system 120 can be implemented remotely (e.g., at a remote server) from the client device 110 as depicted in Figure 1 . In those implementations, the client device 110 and the generative content system 120 can be communicatively coupled to each other via one or more networks 199 (such as one or more wired or wireless local area networks (“LANs”, including mesh networks, near - field communication, etc.) or wide area networks (“WANs”, including the Internet)).
[0023] The client device 110 can be, for example, one or more of the following: a desktop computer, a laptop computer, a tablet computer, a mobile phone, a computing device of a vehicle (e.g., an in - vehicle communication system, an in - vehicle entertainment system, an in - vehicle navigation system), a stand - alone interactive speaker (optionally having a display), a smart appliance (such as a smart TV) and / or a wearable device of a user that includes a computing device (e.g., a watch of the user that has a computing device, glasses of the user that have a computing device, a virtual or augmented reality computing device). Additional and / or alternative client devices can be provided.
[0024] The client device 110 may execute one or more software applications via the application engine 114, through which touch inputs and / or other user inputs may be submitted, and / or content (e.g., text) responsive to the touch inputs and / or other user inputs may be rendered (e.g., audibly and / or visually). The application engine 114 may execute one or more software applications that are separate from the operating system of the client device 110 (e.g., installed “on top of” the operating system), or alternatively may be directly implemented by the operating system of the client device 110. For example, the application engine 114 may execute a web browser, a generative music creator, or an automated assistant installed on top of the operating system of the client device 110. As another example, the application engine 114 may execute a web browser software application, a generative music creator software application, or an automated assistant software application that is integrated as part of the operating system of the client device 110. The application engine 114 (and the one or more software applications executed by the application engine 114) may interact with or otherwise provide access to the generative content system 120 (e.g., as a front end).
[0025] In various implementations, the client device 110 may include a user input engine 111 configured to detect user input provided by a user of the client device 110 using one or more user interface input devices. For example, the client device 110 may be equipped with one or more microphones that capture audio data, such as audio data corresponding to the user's spoken words or other sounds in the environment of the client device 110. Additionally or alternatively, the client device 110 may be equipped with one or more visual components configured to capture visual data corresponding to images and / or movement (e.g., gestures) detected in the field of view of one or more of the visual components. Additionally or alternatively, the client device 110 may be equipped with one or more touch-sensitive components (e.g., a keyboard and mouse, a stylus, a touch screen, a touch panel, one or more hardware buttons, etc.) configured to capture signals corresponding to typed and / or touch inputs directed at the client device 110.
[0026] In some versions of those implementations, the client device 110 may utilize one or more ML models stored in the machine learning (ML) model database 180 to process user input. For example, the user input received at the client device 110 may be an oral utterance. In these examples, the user input engine 111 may use an automatic speech recognition (ASR) model (e.g., a recurrent neural network (RNN) model, a Transformer model, and / or any other type of ML model capable of performing ASR) stored in the ML model database 180 to process the audio data captured by the oral utterance and generated by the microphone of the client device 110 to generate an ASR output. The ASR output may include, for example, speech hypotheses (e.g., word hypotheses and / or transcription hypotheses) predicted to correspond to the oral utterance captured in the audio data, one or more corresponding predicted values (e.g., probabilities, log-likelihoods, and / or other values) for each speech hypothesis in the speech hypotheses, multiple phonemes predicted to correspond to the oral utterance captured in the audio data, one or more corresponding predicted values (e.g., probabilities, log-likelihoods, and / or other values) for each phoneme in the multiple phonemes, and / or other ASR outputs. In these implementations, the user input engine 111 may select one or more speech hypotheses in the speech hypotheses as the recognized text corresponding to the oral utterance (e.g., based on the corresponding predicted values for each speech hypothesis in the speech hypotheses), such as when the user input engine 111 utilizes an end-to-end ASR model. In other implementations, the user input engine 111 may select one or more phonemes in the predicted phonemes (e.g., based on the corresponding predicted values for each phoneme in the predicted phonemes), and determine the recognized text corresponding to the oral utterance based on the selected one or more predicted phonemes, such as when the user input engine 111 utilizes a non-end-to-end ASR model. In these implementations, the user input engine 111 may optionally employ an additional mechanism (e.g., a directed acyclic graph) to determine the recognized text corresponding to the oral utterance based on the selected one or more predicted phonemes.
[0027] In various implementations, the client device 110 may include a rendering engine 112 configured to render content for audible and / or visual presentation to a user of the client device 110 using one or more user interface output devices. For example, the client device 110 may be equipped with speakers that enable content to be rendered as audible content via the client device 110. Additionally or alternatively, the client device 110 may be equipped with a display or a projector that enables content to be rendered as text content and optionally rendered via the client device 110 together with other visual content (e.g., images, videos, etc.).
[0028] In some versions of those implementations, the client device 110 may utilize one or more ML models stored in the ML model database 180 to process the content described herein. For example, and as described above, the content may be audibly rendered as audible content via the speaker of the client device 110. In these examples, the rendering engine 112 may use a text-to-speech (TTS) model stored in the ML model database 180 to process the content (e.g., the lyric content generated using the generative content system 120) to generate synthetic speech audio data including computer-generated synthetic speech capturing the lyric content. In implementations where the rendering engine 112 utilizes the TTS model to process the content, the rendering engine 112 may use a particular set of one or more prosody attributes (e.g., defining the tone, pitch rhythm, speed, etc. of the computer-generated synthetic speech) and / or use a particular voice embedding to generate the synthetic speech to reflect different personas and / or speaking styles, such as a particular set of one or more prosody attributes associated with the user of the client device 110 and / or a voice embedding associated with the user of the client device 110.
[0029] Notably, although the ML models stored in the ML model database 180 were described above as being implemented locally by the client device 110, it should be understood that this is for illustrative purposes and not meant to be limiting. For example, the audio data capturing the spoken utterance may additionally or alternatively be streamed to the generative content system 120, and the generative content system 120 may utilize the ASR model stored in the ML model database 180 (or a separate cloud-based ASR model) to generate the ASR output. Additionally, for example, a summary of the content may be additionally or alternatively processed by the generative content system 120 using the TTS model stored in the ML model database 180 (or a separate cloud-based TTS model) to generate synthetic speech audio data, and the synthetic speech audio data may be streamed to the client device 110 (or an additional client device of the user) to cause the synthetic speech audio data to be audibly rendered for presentation to the user of the client device 110.
[0030] In various implementations, the client device 110 may include a context engine 113 that is configured to determine a client device context (e.g., a current or recent context) of the client device 110 and / or a user context of a user of the client device 110 (or, when the client device 110 is associated with multiple users, the active user of the client device 110). In some of those implementations, the context engine 113 may determine the context based on data stored in the user profile database 110A. The data stored in the user profile database 110A may include, for example, user interaction data characterizing the current or recent interactions of the client device 110 and / or the user of the client device 110, location data characterizing the current or recent location of the client device 110 and / or the geographic region associated with the user of the client device 110, user attribute data characterizing one or more attributes of the user of the client device 110, user preference data characterizing one or more preferences of the user of the client device 110, and / or any other data that the context engine 113 may access via the user profile database 110A or otherwise.
[0031] For example, the context engine 113 may determine the current context based on the current state of a conversation session (e.g., considering one or more recent user inputs provided by the user during the conversation session) and / or the current location of the client device 110. For example, the context engine 113 may determine the current context of "visitor looking for upcoming events in Louisville, Kentucky" based on a recently issued query and the expected future location of the client device 110 (e.g., based on a recently booked hotel accommodation). As another example, the context engine 113 may determine the current context based on which software application is active in the foreground of the client device 110, the current or recent state of the active software application, and / or the content currently or recently rendered by the active software application. The context determined by the context engine 113 may be utilized, for example, to supplement or rewrite a user input received at the client device 110, generate an implicit user input (e.g., an implicit query or prompt formed independent of any explicit user input provided by the user of the client device 110), and / or determine to submit an implicit user input and / or render results (e.g., content) for the implicit user input.
[0032] Further, the client device 110 and / or the generative content system 120 may include one or more memories for storing data and / or software applications, one or more processors for accessing data and executing software applications, and / or other components that facilitate communication over one or more networks in the network 199. In some implementations, one or more of the software applications may be installed locally at the client device 110, while in other implementations, one or more of the software applications may be remotely hosted (e.g., by one or more servers) and may be accessible by the client device 110 over one or more networks in the network 199.
[0033] Although Figure 1 aspects are shown or described with respect to a single client device with a single user, it should be understood that this is for illustrative purposes only and is not meant to be limiting. For example, one or more additional client devices of the user and / or additional users may also implement the techniques described herein. For example, the client device 110, one or more additional client devices, and / or any other computing device of the user may form a device ecosystem that may employ the techniques described herein. These additional client devices and / or computing devices may communicate with the client device 110 (e.g., via the network 199). As another example, a given client device may be utilized by multiple users in a shared setting (e.g., a group of users, a household, a workplace, a hotel, etc.).
[0034] The generative content system 120 is shown in Figure 1 as including a generative model (GM) training engine 130, a GM inference engine 140, a visual multimedia content engine 150, a synchronization verification engine 160, and a modification engine 170. Some of these engines may be combined and / or omitted in various implementations. Further, these engines may include various sub-engines. For example, the GM training engine 130 is shown in Figure 1 as including a GM fine-tuning instance engine 131 and a GM fine-tuning engine 132. Further, the GM inference engine 140 is shown in Figure 1 as including a GM input engine 141, a GM processing engine 142, and a GM output engine 143. Additionally, the modification engine 170 is shown in Figure 1 as including a lyrics seed engine 171, a music composition seed engine 172, and a visual multimedia content seed engine 173. Similarly, some of these sub-engines may be combined and / or omitted in various implementations. Thus, it should be understood that Figure 1 the various engines and sub-engines of the generative content system 120 shown are not meant to be limiting.
[0035] Further, the generative content system 120 is in Figure 1is shown as being docked with various databases such as GM database 120A and fine-tuning data database 130A. Although specific engines and / or sub-engines are depicted as having access to specific databases, it should be understood that this is for illustrative purposes only and is not meant to be limiting. For example, in some implementations, each of the various engines and / or sub-engines of the generative content system 120 can access each of the various databases. Further, some of these databases may be combined and / or omitted in various implementations. Thus, it should be understood that the various databases Figure 1 docked with the generative content system 120 shown are not meant to be limiting.
[0036] In addition, the generative content system 120 is shown in Figure 1 as being docked with other systems such as external system 190. The external system may include, for example, a search system (e.g., a text-based search system, an image-based search system, a video-based search system, etc.) and / or other generative systems (other text-based generative systems, other image-based generative systems, other video-based generative systems, other audio-based generative systems, etc.). In some implementations, the external system 190 is a first-party system, while in other implementations, the external system 190 is a third-party system. As used herein, the term "first party" or "first-party entity" refers to the entity that controls, develops, and / or maintains the generative content system 120, while the term "third party" or "third-party entity" refers to an entity different from the entity that controls, develops, and / or maintains the generative content system 120.
[0037] As described in more detail herein (e.g., with respect to Figure 2 , Figure 3 , Figure 4 and Figures 5A to 5F ), the generative content system 120 can be utilized to generate music content to be rendered for presentation to a user of the client device 110 in response to receiving user input requesting music content. The music content may include, for example, lyric content (e.g., words or phrases for the music content), music composition content (e.g., notes of the music content or a piece of music played by one or more musical instruments), and visual multimedia content related to the lyric content and / or music composition content (e.g., images and / or videos). In some implementations, (e.g., as described with respect to Figure 3 ) a single call to a single GM can be used to generate the music content. In these implementations, the single GM can be fine-tuned to generate the music content. In additional or alternative implementations, (e.g., as described with respect to Figure 4As described, corresponding calls to multiple GMs can be used to generate music content. In these implementations, each of the multiple GMs can be jointly fine-tuned in an end-to-end manner to generate a corresponding portion of the music content. In various implementations, the generative content system 120 can be utilized to refine the music content to be rendered for presentation to the user of the client device 110 in response to receiving additional user input requesting a modification of the music content (e.g., as described with respect to Figures 5A to 5F As described). In these implementations, and when modifying the music content, the generative content system 120 can determine the seed of one or more portions of the music content initially generated based on the user input, and utilize the seed and the additional user input for further processing by the GM to generate a modified version of the music content. By using a seed as described herein, the music content can be efficiently modified as specified by the additional user input while maintaining certain aspects of the music content.
[0038] As described above, in implementations where a single call to a single GM is used to generate music content, the single GM can be fine-tuned to generate the music content. The single GM can be stored in the GM model database 120A and can include any GM (e.g., Bard, Gemini, GPT, and / or any other GM, such as any other GM that is encoder-only, decoder-only, sequence-to-sequence, and optionally includes an attention mechanism or other memory). Notably, the GM stored in the GM database 120A can include billions of weights and / or parameters learned by initially training the GM on a large amount of diverse data. This enables these GMs to generate GM outputs as a probability distribution over a sequence of tokens as described herein. Further, in implementations where a single call to a single GM is used to generate music content, the single GM can be a multimodal GM that is fine-tuned to be capable of processing text-based user input (e.g., typed user input provided by the user of the client device 110), audio-based user input (e.g., spoken user input provided by the user of the client device 110), and / or visual-based user input (e.g., images and / or videos provided by the user of the client device 110) to generate text-based content (e.g., text corresponding to lyric content as described herein and / or text corresponding to music composition content such as notes as described herein), audio-based content (e.g., audio data corresponding to lyric content as described herein and / or audio data corresponding to music composition content as described herein), and / or visual-based content (e.g., images and / or videos associated with the music content as described herein). Further, by fine-tuning a single GM to generate music content, any resulting music content is synchronized without the need for any additional post-processing of the music content.
[0039] However, in various implementations, the synchronization verification engine 160 can be utilized to verify that music content is actually synchronized when played for presentation to a user of the client device 110. For example, the synchronization verification engine 160 can simulate the playback of music content without the need to render the music content for presentation to the user. During the simulated playback of the music content, the synchronization verification engine 160 can verify, for example, that lyric content 205 is logically arranged with respect to the playback of music composition content 206, and correct any potential errors by inserting a delay for the lyric content 205, removing a delay for the music composition content 206, adjusting the tempo or beat of the playback of the lyric content 205, and so on. In some versions of those implementations, the synchronization verification engine 160 can accelerate the playback of the lyric content 205 and the music composition content 206 to reduce the latency that causes the music content to be rendered for presentation to the user.
[0040] When fine-tuning a single GM, the GM fine-tuning instance engine 131 can access the fine-tuning data database 130A to obtain a plurality of fine-tuning instances. Each fine-tuning instance among the plurality of fine-tuning instances can include corresponding fine-tuning user input, corresponding fine-tuning lyric content, and corresponding fine-tuning music composition content (and optionally include corresponding fine-tuning visual multimedia content). Further, when fine-tuning a single GM based on a given fine-tuning instance among the plurality of fine-tuning instances, the GM fine-tuning engine 132 can process the corresponding user input to generate predicted lyric content and predicted music composition content (and optionally generate predicted visual multimedia content). In some implementations, the GM fine-tuning engine 132 can compare the predicted lyric content with the corresponding fine-tuning lyric content of the given fine-tuning instance, and compare the predicted music composition content with the corresponding fine-tuning music composition content of the given fine-tuning instance to generate one or more losses (and optionally can compare the predicted visual multimedia content with the corresponding fine-tuning visual multimedia content of the given fine-tuning instance). Additionally, the GM fine-tuning engine 132 can update the single GM based on the one or more losses. Although specific learning techniques for fine-tuning a single GM are described above (e.g., supervised fine-tuning (SFT) techniques), it should be understood that this is for illustrative purposes and does not imply a limitation.
[0041] For example, the GM fine-tuning engine 132 may additionally or alternatively utilize reinforcement learning from human feedback (RLHF) techniques, where the predicted lyric content and the predicted music composition content (and optionally, the predicted visual multimedia content) are provided for presentation to developers associated with the generative content system 120, and given the corresponding fine-tuning user input processed using a single GM, the developers may obtain feedback on the predicted lyric content and the predicted music composition content. However, it should be noted that techniques that require the participation of developers (or other users, such as MechanicalTurks) consume additional computational and financial resources.
[0042] In addition, for example, the GM fine-tuning instance engine 131 may access the fine-tuning data database 130A to obtain a plurality of first fine-tuning instances and a plurality of second fine-tuning instances. Each first fine-tuning instance among the plurality of first fine-tuning instances may include a corresponding fine-tuning user input and corresponding fine-tuning lyric content, and each second fine-tuning instance among the plurality of second fine-tuning instances may include a corresponding fine-tuning user input and corresponding fine-tuning music composition content. Thus, in this case, for each first fine-tuning instance among the plurality of first fine-tuning instances that includes the corresponding fine-tuning lyric content, there is a corresponding second fine-tuning instance among the plurality of second fine-tuning instances that includes the corresponding fine-tuning music composition content for the corresponding fine-tuning lyric content. The GM fine-tuning engine 132 may process the corresponding user input in the same or a similar manner as described above to generate the predicted lyric content and the predicted music composition content.
[0043] Similarly, as described above, in an implementation of using corresponding calls to multiple GMs to generate music content, each of the multiple GMs can be jointly fine-tuned in an end-to-end manner to generate a corresponding part of the music content. Each of the multiple GMs can be stored in the GM model database 120A and can include any GM (e.g., Bard, Gemini, GPT, and / or any other GM, such as any other GM that is encoder-only, decoder-only, sequence-to-sequence, and optionally includes an attention mechanism or other memory). Further, in an implementation of using corresponding calls to multiple GMs to generate music content, each of the GMs can have a corresponding modality. For example, a first GM can be fine-tuned to be capable of processing text-based user input (e.g., typed user input provided by a user of the client device 110), audio-based user input (e.g., spoken user input provided by a user of the client device 110), and / or vision-based user input (e.g., images and / or videos provided by a user of the client device 110) to generate text-based content (e.g., text corresponding to lyric content as described herein and / or text corresponding to music composition content such as musical notes as described herein). Further, a second GM can be fine-tuned to be capable of processing text-based user input, audio-based user input, and / or vision-based user input to generate audio-based content (e.g., audio data corresponding to lyric content as described herein and / or audio data corresponding to music composition content as described herein). Additionally, a third GM can be fine-tuned to be capable of processing text-based user input, audio-based user input, and / or vision-based user input to generate vision-based content (e.g., images and / or videos associated with the music content as described herein). Further, by jointly fine-tuning these multiple GMs in an end-to-end manner to generate music content, any resulting music content is synchronized without the need for any additional post-processing of the music content. However, the synchronization verification engine 160 can be utilized to verify that the music content is actually synchronized as described above.
[0044] When jointly fine-tuning multiple GMs in an end-to-end manner, the GM fine-tuning instance engine 131 can access the fine-tuning data database 130A to obtain multiple corresponding fine-tuning instances for each of the multiple GMs. For example, each of the multiple first fine-tuning instances to be utilized when fine-tuning the first GM to generate lyric content may include a corresponding fine-tuning user input and corresponding fine-tuning lyric content. Further, each of the multiple second fine-tuning instances to be utilized when fine-tuning the second GM to generate music composition content may include a corresponding fine-tuning user input and corresponding fine-tuning music composition content. Additionally, each of the multiple third fine-tuning instances to be utilized when fine-tuning the third GM to generate visual multimedia content associated with music content may include a corresponding fine-tuning user input and corresponding fine-tuning visual multimedia content. Thus, in this case, for each of the multiple first fine-tuning instances that includes corresponding fine-tuning lyric content, there is a corresponding one of the multiple second fine-tuning instances that includes corresponding fine-tuning music composition content for the corresponding fine-tuning lyric content, and there is a corresponding one of the multiple third fine-tuning instances that includes corresponding fine-tuning visual multimedia content for the corresponding fine-tuning lyric content and corresponding fine-tuning music composition content. The GM fine-tuning engine 132 can cause each of the multiple GMs to process the corresponding user input to separately generate predicted lyric content, predicted music composition content, and predicted visual multimedia content in the same or a similar manner as described above. However, when jointly fine-tuning multiple GMs in an end-to-end manner, one or more of the losses can be shared across the multiple GMs to ensure that the music content generated using the multiple GMs is synchronized when played to be presented to the user of the client device 110. However, the synchronization verification engine 160 can be utilized to verify that the music content is actually synchronized as described above when played to be presented to the user of the client device 110.
[0045] Now turning to Figure 2 , depicts a diagram for utilizing from Figure 1The process flow of the various components of an example environment. For purposes of illustration, assume that a user of client device 110 provides user input 201, and that this user input 201 is detected via user input engine 111. For example, assume that user input 201 is "write me a song about patent law". In this example, GM input engine 141 may process user input 201 to generate GM input 203. Notably, in generating GM input 203, GM input engine 141 may utilize explicit GM (e.g., stored in GM database 140A). Explicit GM can be a form of GM that processes user input 201 (and optionally processes context 202 determined by context engine 113 of client device 110) to generate GM input 203. GM input 203 may then be provided to GM processing engine 142 to generate GM output 204. In other words, GM input engine 141 may utilize explicit GM to process the original user input 201 and place the original user input in a structured form more suitable for processing by GM processing engine 142. Further, GM input engine 141 may utilize explicit GM to incorporate context 202 into the GM input and optionally utilize any other dynamic cues to assist GM processing engine 142 in generating GM output 204. For example, and based on user input 201 being "write me a song about patent law", context 202 may include recent news about patent law or search results for patent news (e.g., obtained via a call to one of external systems 190 such as the Internet), an indication that the user's occupation is "patent attorney" based on user profile data stored in user profile database 110A, and / or other context. Further, and based on user input 201 being "write me a song about patent law", dynamic cues may include, for example, "write a song about patent law for a patent attorney, be specific in the lyrics and mention pertinent statutes and regulations for patent law" etc.
[0046] In an implementation of using a single GM to generate music content, the GM input 203 can include only a single GM input. Further, in these implementations, the GM processing engine 142 can use the single GM to process the GM input 203 to generate the GM output 204. Additionally, in these implementations, the GM output 204 can include a probability distribution over a sequence of tokens. For example, when determining the lyric content 205, the GM output engine 143 can employ various decoding techniques from a sequence of words or word units (e.g., text-based output) or from a sequence of phonemes or articulatory units (e.g., audio-based output), and determine the lyric content 205 based on the probability distribution over the sequence of words or word units or over the sequence of phonemes or articulatory units. Further, when determining the music composition content 206, the GM output engine 143 can employ various decoding techniques from a sequence of notes or note units and determine the music composition content 206 based on the probability distribution over the sequence of notes or note units.
[0047] In an implementation of using multiple GMs to generate music content, the GM input 203 can include a respective GM input for each of the multiple GMs, where each GM input in the respective GM inputs can vary because the context 202 or the dynamic cue can vary for each of the GMs. Further, in these implementations, the GM processing engine 142 can use each of the multiple GMs to process a respective one of the GM inputs in the GM input 203 to generate the GM output 204. Additionally, in these implementations, the GM output 204 can include a respective probability distribution over a respective sequence of tokens. For example, when determining the lyric content 205, the GM output engine 143 can employ various decoding techniques from a sequence of words or word units (e.g., text-based output) or from a sequence of phonemes or articulatory units (e.g., audio-based output), and determine the lyric content 205 based on the probability distribution over the sequence of words or word units or over the sequence of phonemes or articulatory units. It is noted that a first GM can be used to determine the probability distribution over the sequence of words or word units or over the sequence of phonemes or articulatory units. Further, when determining the music composition content 206, the GM output engine 143 can employ various decoding techniques from a sequence of notes or note units and determine the music composition content 206 based on the probability distribution over the sequence of notes or note units. It is noted that a second GM different from the first GM can be used to determine the probability distribution over the sequence of notes or note units.
[0048] Further, the rendering engine 112 may cause the lyrics content 205 and / or the music composition content 206 to be rendered as music content at the user's client device 110 and in response to the user input 201. In various implementations, the visual multimedia content engine 150 may determine the visual multimedia content 207 to be rendered together with the music content. In some versions of those implementations, the visual multimedia content 207 may be a generated visual multimedia content (e.g., a generated image, a generated video, a generated animation, or a gif, etc.). In an implementation in which a single GM is used to generate music content, the visual multimedia content engine 150 may determine the visual multimedia content 207 based on the GM output. In an implementation in which multiple GMs are used to generate music content, a separate image generation GM may be used to generate the visual multimedia content 207. In other versions of those implementations, the visual multimedia content 207 may be a non-generated visual multimedia content (e.g., a non-generated image, a non-generated video, a non-generated animation, or a gif, etc.). In these implementations, the visual multimedia content engine 150 is non-generative visual multimedia content, and the visual multimedia content engine 150 can obtain the non-generative visual multimedia content from one or more databases (for example, an image / video album of a user of the client device 110, an image / video of the user of the client device 110 obtained via a call to one of the external systems 190 such as the Internet).
[0049] In various implementations, and as indicated at block 208, the generative content system 120 may receive additional user input to modify the music content originally rendered for presentation to the user. If no additional user input is received, the generative content system 120 may wait to receive the additional user input at block 208. However, if additional user input is received, the modification engine 170 may determine a seed 209 to be utilized in generating a modified version of the music content. Continuing with the example above, where the user input 201 is "write me a song about patent law," further assume that the user of the client device 110 provides additional user input of "the lyrics sound great, can you include some additional lyrics about the current state of 103 and obviousness rationales." In this example, the additional user input indicates that the user of the client device 110 is satisfied with the originally rendered lyrics content 205 and the music composition content 206, but indicates a desire to add additional lyrics.
[0050] Accordingly, in this example, the modification engine 170 (and more specifically, the lyric seed engine 171) may determine the seed of the lyric content 205, and the modification engine 170 (and more specifically, the music composition seed engine 173) may determine the seed of the music composition content 206. In an implementation where visual multimedia content 207 is included and includes generative visual multimedia content, the modification engine 170 (and more specifically, the visual multimedia content seed engine 173) may determine the seed of the visual multimedia content 207. The seed 209 may be a corresponding lower-level representation of the lyric content 205 and / or the music composition content 206. For example, the corresponding lower-level representation of the lyric content and / or the music composition content may be a corresponding embedding in a corresponding embedding space. Thus, the GM input engine 141 may cause the explicit GM to include the seed in the processing of the additional GM input to generate a modified version of the lyric content 205 to include additional details about "the current state of 103andobviousness rationales" as requested by the user via the additional user input. Further, the rendering engine 112 may cause the modified version of the lyric content 205 and / or the music composition content 206 to be rendered as music content at the client device 110 of the user and in response to the additional user input. The user may continue to interact with the generative content system 120 in this manner to continue modifying the music content. Optionally, one or more selectable elements may be provided to the user of the client device 110 to share the music content generated via the generative content system 120.
[0051] Now turning to Figure 3 , a flowchart of an example method 300 is depicted that illustrates using a single generative model (GM) to generate music content. For convenience, the operations of method 300 are described with reference to the system that performs these operations. The system of method 300 includes one or more processors, memories, and / or other components of a computing device (e.g., Figure 1 the client device 110 of Figure 1 the generative content system 120 of Figure 6 the computing device 610, one or more servers, and / or other computing devices). Additionally, although the operations of method 300 are shown in a particular order, this is not intended to be limiting. One or more operations may be reordered, omitted, and / or added.
[0052] At block 352, the system receives user input associated with a client device, the user input including a request for music content, and the music content includes at least lyric content and music composition content. The user input may be received via typed input, verbal input, touch input, etc.
[0053] At block 354, the system processes the GM input using a generative model (GM) to generate a GM output, where the GM input includes at least the user input. For example, the system may generate the GM input (e.g., as described regarding the GM input processing engine 141 with respect to Figure 1 and Figure 2 ), and may use the GM to process the GM input to generate the GM output (e.g., as described regarding the GM processing engine 142 with respect to Figure 1 and Figure 2 ).
[0054] At block 356, the system determines the lyric content and the music composition content based on the GM output. For example, the system may determine the lyric content and the music composition content based on one or more probability distributions over one or more sequences of tokens (e.g., as described regarding the GM output engine 143 with respect to Figure 1 and Figure 2 ). In some implementations, block 356 may further include sub-block 356A. In these implementations, at sub-block 356A, the system determines the visual multimedia content to be rendered at the client device when the music content is being rendered at the client device. The visual multimedia content may include, for example, generative visual multimedia content and / or non-generative visual multimedia content. In implementations where the visual multimedia content includes generative visual multimedia content, the system may determine the generative visual multimedia content based on one or more of the probability distributions over one or more sequences of tokens (e.g., as described regarding the GM output engine 143 with respect to Figure 1 and Figure 2 ). In additional or alternative implementations, in cases where the visual multimedia content includes generative visual multimedia content, the system may utilize a separate image generation GM or a separate video generation GM that is separate from the GM used to process the GM input. In implementations where the visual multimedia content includes non-generative visual multimedia content, the system may obtain the non-generative visual multimedia content from one or more databases that are personal to the user providing the user input and based on one or more entities mentioned in the user input.
[0055] At block 358, the system causes music content to be rendered at the client device. In some implementations, the system may cause lyric content to be visually rendered at a display of the client device. In additional or alternative implementations, the system may cause lyric content to be audibly rendered via a speaker of the client device. In some implementations, the system may cause music composition content to be audibly rendered via a speaker of the client device and optionally rendered together with the lyric content. In additional or alternative implementations, the system may cause selectable elements or links to be rendered via a display of the client device and, when selected, cause music composition content to be audibly rendered via a speaker of the client device.
[0056] In some implementations, block 358 may further include sub-block 358A. In these implementations, at sub-block 358A, the system causes lyric content to be rendered in the voice of a user of the client device. For example, the lyric content may correspond to text determined based on the GM output. Thus, when synthesizing the audio data capturing the lyric content, the system may utilize the voice embedding of the user (e.g., stored in the user profile database 110A or obtained by requesting the user to say a few words during an interaction) and / or a set of one or more prosodic attributes associated with the user (e.g., stored in the user profile database 110A or obtained by requesting the user to say a few words during an interaction) to synthesize the audio data such that the audio data is audibly perceived as being spoken or sung by the user providing the user input. As another example, the lyric content may correspond to audio data determined based on the GM output. Thus, instead of synthesizing the audio data capturing the lyric content, the system may use the voice embedding of the user and / or a set of one or more prosodic attributes associated with the user to adapt the lyric content such that the lyric content is audibly perceived as being spoken or sung by the user providing the user input.
[0057] At block 360, the system determines whether additional user input has been received. The additional user input may be received via typed input, verbal input, touch input, etc. If, at an iteration of block 360, the system determines that no additional user input has been received, the system may continue to monitor for additional user input at block 360.
[0058] If, at an iteration of block 360, the system determines that additional user input has been received, the system proceeds to block 362. At block 362, the system determines whether additional user input has been provided to modify the music content. If, at an iteration of block 362, the system determines that no additional user input has been provided to modify the music content, the system returns to block 360. However, it should be noted that if no additional user input is provided to modify the music content, the system may still respond to the user. Nevertheless, the system may still continue to monitor for additional user input provided to modify the music content such that the monitoring persists across a conversation session between the user and the system and such that the monitoring persists across multiple conversation sessions between the user and the system.
[0059] If, at an iteration of block 362, the system determines that additional user input has been provided to modify the music content, the system proceeds to block 364. At block 364, the system determines one or more seeds for the lyric content and / or music composition content. For example, the system may determine seeds for the lyric content and music composition content (e.g., as described with respect to Figure 1 and Figure 2 lyric seed engine 171 and music composition seed engine 172). Additionally or alternatively, the system may determine corresponding seeds for the lyric content and additional corresponding seeds for the music composition content (e.g., as described with respect to Figure 1 and Figure 2 lyric seed engine 171 and music composition seed engine 172). It should be understood that the seeds determined at block 364 may vary based on how the additional user input requests modification of the music content. The system returns to block 354 and continues method 300.
[0060] However, when returning to block 354 and continuing method 300, the system may process additional GM input to generate additional GM output. The additional GM input includes at least the seeds determined at block 364 and the additional user input. Thus, by continuing iterations of method 300 and by leveraging the seeds, the modified version of the music content should retain aspects of the originally rendered music content but also include modifications based on how the additional user input requests modification of the music content.
[0061] Although Figure 3 method 300 is described with respect to using a single GM to generate music content, it should be understood that this is one technique contemplated herein and is not intended to be limiting.
[0062] Now turning to Figure 4, depicts a flowchart of an example method 400 showing the use of multiple generative models (GMs) to generate music content. For convenience, the operations of method 400 are described with reference to the system performing the operation. The system of method 400 includes one or more processors, a memory, and / or other components (e.g., Figure 1 client device 110 of Figure 1 generative content system 120 of Figure 6 computing device 610, one or more servers, and / or other computing devices). Additionally, although the operations of method 400 are shown in a particular order, this is not intended to be limiting. One or more operations may be reordered, omitted, and / or added.
[0063] At block 452, the system receives user input associated with the client device, the user input including a request for music content, and the music content includes at least lyric content and music composition content. The user input may be received via typed input, verbal input, touch input, etc.
[0064] At block 454, the system processes a first GM input using a first generative model (GM) to generate a first GM output, the first GM input including at least the user input. For example, a first GM may be utilized to generate lyric content. Further, the first GM input may be customized for the first GM for generating lyric content by including, for example, appropriate context information along with the user input for generating lyric content, appropriate dynamic cues along with the user input for generating lyric content, etc. The first GM may be, for example, a large language model (LLM), an audio generative model, etc. The first GM input and the first GM's processing of the first GM input are described in more detail herein (e.g., as described with respect to Figure 1 and Figure 2 GM input engine 141 and GM processing engine 142).
[0065] At block 456, the system determines the lyric content based on the first GM output. The first GM output may be, for example, a probability distribution over a sequence of words or word units (e.g., when the first GM is an LLM) or a sequence of phonemes or pronunciation units (e.g., when the first GM is an audio generative model). Based on the probability distribution, the system may determine the lyric content from the sequence of words or word units or the sequence of phonemes or phoneme units (e.g., as described with respect to Figure 1 and Figure 2 GM output engine 143).
[0066] At block 458, the system uses a second GM to process a second GM input to generate a second GM output, where the second GM input includes at least user input. For example, the second GM can be utilized to generate music composition content. Further, the second GM input can be customized for the second GM for generating music composition content by including, for example, appropriate context information along with the user input for generating music composition content, appropriate dynamic cues along with the user input for generating music composition content, etc. The second GM can be, for example, another large language model (LLM) and / or another audio generative model (e.g., AudioLM), etc. The second GM input and the second GM's processing of the second GM input are described in more detail herein (e.g., as described regarding Figure 1 and Figure 2 the GM input engine 141 and GM processing engine 142).
[0067] At block 460, the system determines music composition content based on the second GM output. The second GM output can be, for example, a probability distribution over a sequence of notes or note units. Based on the probability distribution, the system can determine music composition content from the sequence of notes or note units (e.g., as described regarding Figure 1 and Figure 2 the GM output engine 143).
[0068] At block 462, the system causes the music content to be rendered at the client device. In some implementations, the system can cause the lyric content to be visually rendered at the display of the client device. In additional or alternative implementations, the system can cause the lyric content to be audibly rendered via the speaker of the client device. In some implementations, the system can cause the music composition content to be audibly rendered via the speaker of the client device and, optionally, rendered together with the lyric content. In additional or alternative implementations, the system can cause selectable elements or links to be rendered via the display of the client device and, when selected, cause the music composition content to be audibly rendered via the speaker of the client device.
[0069] At block 464, the system determines whether additional user input has been received. Additional user input can be received via typed input, verbal input, touch input, etc. If, at an iteration of block 464, the system determines that no additional user input has been received, the system can continue to monitor for additional user input at block 464.
[0070] If, at an iteration of block 464, the system determines that no additional user input has been received, the system advances to block 466. At block 466, the system determines whether additional user input has been provided to modify the music content.
[0071] If, at an iteration of block 466, the system determines that no additional user input has been provided to modify the music content, the system returns to block 464. However, it should be noted that if no additional user input is provided to modify the music content, the system may still respond to the user. Nevertheless, the system may still continue to monitor for additional user input provided to modify the music content such that the monitoring persists across a dialogue session between the user and the system and such that the monitoring persists across multiple dialogue sessions between the user and the system.
[0072] If, at an iteration of block 466, the system determines that additional user input has been provided to modify the music content, the system advances to block 468. At block 468, the system determines one or more seeds for the lyric content and / or music composition content. For example, the system may determine seeds for the lyric content and music composition content (e.g., as described with respect to Figure 1 and Figure 2 lyric seed engine 171 and music composition seed engine 172). Additionally or alternatively, the system may determine corresponding seeds for the lyric content and additional corresponding seeds for the music composition content (e.g., as described with respect to Figure 1 and Figure 2 lyric seed engine 171 and music composition seed engine 172). It should be understood that the seeds determined at block 468 may vary based on how the additional user input requests modification of the music content. The system returns to block 454 and continues method 400.
[0073] However, upon returning to block 454 and continuing method 400, the system may process additional GM input to generate additional GM output. The additional GM input includes at least the seeds determined at block 468 and the additional user input. Thus, by continuing iterations of method 400 and by leveraging the seeds, the modified version of the music content should retain aspects of the original rendered music content but also include modifications based on how the additional user input requests modification of the music content.
[0074] Although Figure 4 method 400 does not include corresponding portions of the respective sub - blocks described for Figure 3 method 300, it should be understood that this is for the sake of brevity and is not intended to be limiting. For example, in some implementations, and although not shown in Figure 4depicted in method 400, but the system can determine visual multimedia content to be rendered at the client device while music content is being rendered at the client device. The visual multimedia content can include, for example, generative visual multimedia content and / or non-generative visual multimedia content. In an implementation where the visual multimedia content includes generative visual multimedia content, the system can determine the generative visual multimedia content based on one or more probability distributions in a probability distribution over one or more token sequences from the first GM or the second GM (e.g., as described regarding Figure 1 and Figure 2 the GM output engine 143). In an additional or alternative implementation, in a case where the visual multimedia content includes generative visual multimedia content, the system can utilize a separate image generation GM or a separate video generation GM that is separate from both the first GM and the second GM. In an implementation where the visual multimedia content includes non-generative visual multimedia content, the system can obtain the non-generative visual multimedia content from one or more databases that are personal to the user providing the user input and based on one or more entities mentioned in the user input.
[0075] As another example, in some implementations, and although not depicted in Figure 4 method 400, the system can cause the lyrics content to be rendered in the voice of the user of the client device. For example, the lyrics content can correspond to text determined based on the output of the first GM (e.g., when the first GM is an LLM). Thus, when synthesizing the audio data capturing the lyrics content, the system can utilize the voice embedding of the user (e.g., stored in the user profile database 110A or obtained by requesting the user to say a few words during the interaction) and / or a set of one or more prosodic attributes associated with the user (e.g., stored in the user profile database 110A or obtained by requesting the user to say a few words during the interaction) to synthesize the audio data such that the audio data is audibly perceived as being spoken or sung by the user providing the user input. Additionally, for example, the lyrics content can correspond to the audio data determined based on the output of the first GM (e.g., when the first GM is an audio generative model). Thus, instead of synthesizing the audio data capturing the lyrics content, the system can use the voice embedding of the user and / or a set of one or more prosodic attributes associated with the user to adapt the lyrics content such that the lyrics content is audibly perceived as being spoken or sung by the user providing the user input.
[0076] Now turning to Figure 5A 、 Figure 5B 、 Figure 5C 、 Figure 5D 、 Figure 5E and Figure 5F, depicting various non-limiting examples of generating music content. The client device 110 (e.g., from Figure 1 The client device 110 of the embodiment of the present invention may include various user interface components, including, for example, a microphone for generating audio data based on spoken words and / or other audible input, a speaker for audibly rendering synthesized speech and / or other audible output, and / or a display 191 for visually rendering visual output. Further, the display 191 of the client device 110 may include various system interface elements 192, 193, and 194 (e.g., hardware and / or software interface elements) that can be interacted with by a user of the client device 110 to cause the client device 110 to perform one or more actions. Display 191 of client device 110 enables a user to interact with content rendered on display 191 via touch input (e.g., by directing user input to display 191 or portions thereof (e.g., to text entry box 195, to a keyboard (not depicted), or to other portions of display 191)) and / or via verbal input (e.g., at client device 110, by selecting microphone interface element 196—or simply by speaking without selecting microphone interface element 196 (i.e., the automated assistant may monitor one or more terms or phrases, gestures, gazes, mouth movements, lip movements, and / or other conditions that activate verbal input)). Although Figure 5A , Figure 5B , Figure 5C , Figure 5D , Figure 5E and Figure 5F The client device 110 is depicted as a mobile phone, but it should be understood that this is for the purpose of example and is not intended to be limiting. For example, the client device 110 may be a stand-alone speaker with a display, a stand-alone speaker without a display, a home automation device, an in-vehicle system, a laptop computer, a desktop computer, and / or any other device capable of executing an automated assistant to participate in a human-computer conversation session with a user of the client device 110.
[0077] Specifically refer to Figure 5A , assuming that a user of client device 110 accesses a generative content application via client device 110, the generative content application enables the user to interact with a generative content system (e.g., Figure 1 Assume further that the user provides user input 552A “Help me write a lullaby for Frank that reminds him daddy loves him”. In response to receiving user input 552A, the generative content system may generate music content (e.g., Figure 3method 300 or Figure 4 described by method 400). The music content may include lyric content as indicated at 554A1 and music composition content as indicated at 554A2. It should be noted that the lyric content may be rendered visually and / or audibly via the speaker of the client device 110. Further, the music composition content may be rendered audibly via the speaker of the client device 110. However, in various implementations, the lyric content and / or music composition content may not be rendered audibly until additional user input is provided to cause the lyric content and / or music composition to be rendered audibly.
[0078] Specific reference Figure 5B and continue Figure 5A to the example of, and further assume that the user provides additional user input 552B of "Frank also loves planes and cats, can you include these in the lyrics as well?". In response to receiving the additional user input 552B, the generative content system may determine the seed of the lyric content and / or music composition content originally rendered in Figure 5A and as indicated at 554B1. It should be noted that in the example of Figure 5B , the additional user input 552B does not request modification of any music composition content. However, the additional user input 552B requests modification of the lyric content. Specifically, the additional user input 552B requests supplementing the lyric content to include lyric content related to planes and cats. Thus, when subsequently processing the seed along with the additional user input, the generative content system ensures that the lyric content and music composition content are substantially similar to the lyric content and music composition content originally rendered in Figure 5A , but in Figure 5B , the lyric content as indicated by 554B2 includes additional lyrics related to planes and cats, and the music composition content as indicated by 554A2 does not change.
[0079] Specific reference Figure 5C and continue Figure 5A and Figure 5B to the example of, and further assume that the user provides additional user input 552C of "The lyrics are perfect, can you make the music a bit more rocky?". In response to receiving the additional user input 552C, the generative content system may determine the seed in Figure 5BThe seeds of the lyric content and / or musical composition content that are subsequently rendered therein and as indicated at 554C1. It should be noted that in Figure 5C 's example, the additional user input 552C does not request to modify any lyric content. However, the additional user input 552C requests to modify the musical composition. Specifically, the additional user input 552C requests that the musical composition content reflect the genre of rock music. Thus, when subsequently processing the seeds along with the additional user input, the generative content system ensures that the lyric content and the musical composition content are substantially similar to the lyric content and the musical composition content that are subsequently rendered in Figure 5B However, in Figure 5C as indicated at 554C3, the musical composition content includes more elements represented by the rock music genre, and the lyric content as indicated at 554B2 remains unchanged.
[0080] Specifically refer to Figure 5D and continue with Figures 5A to 5C 's example. Further assume that the user provides the additional user input 552D of "Can you add some images of planes and cats to go along with the lullaby as well?". In response to receiving the additional user input 552D, as indicated at 554D1, the generative content system may invoke an image generative model (e.g., the same model that is utilized to generate the lyric content and / or the musical composition content, and / or a separate image generative model). For example, the generative content system may generate a prompt that, when processed by the image generative model, causes generative images of planes and cats to be rendered at the client device 110, as indicated at 554D2 and 554D3 respectively. Although Figure 5D 's example is described with respect to a generative content system that causes generative images of planes and cats to be generated and rendered, it should be understood that this is for purposes of illustration and not meant to be limiting. Instead, it should be understood that non-generative images of planes and / or cats may be obtained (e.g., by invoking an image search system, etc. via the external system 190). It should be noted that these generative images may be rendered when the music content is played at the client device 110.
[0081] Specifically refer to Figure 5E and continue with Figures 5A to 5D, further assuming that the user provides additional user input 552E of “Can you addsome images of me and Frank as well?” In response to receiving the additional user input 552E, the generative content system can access and obtain images of the user and Frank, as indicated at 554E1. For example, the generative content system can access the user's photo album (or cause the automated assistant to access the user's photo album) to obtain, for example, a first image of the user and Frank as indicated at 554E2, a second image of the user and Frank as indicated at 554E3, and / or other images and / or videos of the user and Frank. Notably, these non-generative images can be rendered when the music content is played at the client device 110.
[0082] Specific reference Figure 5F , and continue Figures 5A to 5E , further assume that the user provides additional user input 552F of “Perfect, let's hear the lullaby and use my voice to sing the lyrics”. In response to receiving additional user input 552F, the generative content system may synthesize and / or adapt the lyric content to be audibly rendered in the voice of the user of client device 110. Thus, when playing the lullaby for presentation to the user of client device 110 as indicated at 554F1, the generative content system may audibly render the lyric content from Figure 5B A modified version of the lyrics content and as in Figure 5F Further, the generative content system may audibly render the user's voice from Figure 5C A modified version of the music composition content, as indicated by 554F3. In addition, the generative content system can visually render Figure 5D and Figure 5E The generated and non-generated images requested in , as indicated by 554F4. The user can continue to interact with the generated content as desired.
[0083] although Figures 5A to 5F Although described with respect to specific examples, it should be understood that these examples are provided for the purpose of illustrating the technology contemplated herein and are not intended to be limiting. Figures 5A to 5F The interactive script is depicted in , but it should be understood that this is for illustrative purposes and is not intended to be limiting.
[0084] Now turn to Figure 6, which depicts a block diagram of an example computing device 610 that can optionally be utilized to perform one or more aspects of the techniques described herein. In some implementations, one or more of a client device, a multimodal response system component, or other cloud-based software application components and / or other components can include one or more components of the example computing device 610.
[0085] The computing device 610 generally includes at least one processor 614 that communicates with a plurality of peripheral devices via a bus subsystem 612. These peripheral devices can include a storage subsystem 624 (which includes, for example, a memory subsystem 625 and a file storage subsystem 626), a user interface output device 620, a user interface input device 622, and a network interface subsystem 616. The input and output devices allow user interaction with the computing device 610. The network interface subsystem 616 provides an interface to an external network and is coupled to corresponding interface devices in other computing devices.
[0086] The user interface input device 622 can include a keyboard, a pointing device (such as a mouse, trackball, touchpad, or graphics tablet), a scanner, a touchscreen integrated into a display, an audio input device (such as a voice recognition system, a microphone), and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and ways of inputting information into the computing device 610 or onto a communication network.
[0087] The user interface output device 620 can include a display subsystem, a printer, a fax machine, or a non-visual display, such as an audio output device. The display subsystem can include a cathode ray tube (CRT), a flat panel device (such as a liquid crystal display (LCD)), a projection device, or some other mechanism for creating a visible image. The display subsystem can also provide a non-visual display, such as via an audio output device. In general, the use of the term "output device" is intended to include all possible types of devices and ways of outputting information from the computing device 610 to a user or another machine or computing device.
[0088] The storage subsystem 624 stores programming and data constructs that provide the functionality of some or all of the modules described herein. For example, the storage subsystem 624 can include aspects for performing selected aspects of the methods disclosed herein and logic for implementing Figure 1 and Figure 2 the various components depicted.
[0089] These software modules are typically executed by the processor 614, either alone or in conjunction with other processors. The memory 625 used in the storage subsystem 624 may include multiple memories, including a main random access memory (RAM) 630 for storing instructions and data during program execution and a read-only memory (ROM) 632 for storing fixed instructions. The file storage subsystem 626 may provide persistent storage for program and data files and may include a hard disk drive, a floppy disk drive and associated removable media, a CD-ROM drive, an optical disk drive, or a removable media cartridge. Modules implementing the functionality of certain implementations may be stored by the file storage subsystem 626 in the storage subsystem 624 or in other machines accessible to the processor 614.
[0090] The bus subsystem 612 provides the mechanism for enabling the various components and subsystems of the computing device 610 to communicate with each other as intended. Although the bus subsystem 612 is schematically shown as a single bus, alternative implementations of the bus subsystem 612 may use multiple buses.
[0091] The computing device 610 may have different types, including workstations, servers, computing clusters, blade servers, server farms, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of the computing device 610 depicted in Figure 6 is only intended as a specific example for illustrating some implementations. Many other configurations of the computing device 610 are possible, which have more or fewer components compared to the computing device depicted in Figure 6
[0092] In cases where the systems described herein collect or otherwise monitor personal information about a user or information that can utilize personal information and / or monitoring, the user may be given the opportunity to control whether the program or feature collects user information (e.g., information about the user's social network, social actions or activities, occupation, user preferences, or the user's current geographical location), or to control whether and / or how to receive content that may be more relevant to the user from the content server. Additionally, certain data may be processed in one or more ways before it is stored or used so that personally identifiable information is removed. For example, the user's identity may be processed so that the user's personally identifiable information cannot be determined, or in the case of obtaining geographical location information, the user's geographical location may be generalized (such as generalized to the city, zip code, or state level) so that the user's specific geographical location cannot be determined. Thus, the user can control how information about the user is collected and / or used.
[0093] In some implementations, a method implemented by one or more processors is provided, and the method includes: receiving user input associated with a client device of a user, the user input including a request for music content, and the music content including lyric content and music composition content; and generating the music content in response to the user input. Generating the music content in response to the user input includes: processing GM input using a generative model (GM) to generate a GM output, the GM input including at least the user input; and determining the lyric content and the music composition content based on the GM output. The method further includes: causing the music content to be audibly rendered at the client device.
[0094] These and other implementations of the techniques disclosed herein may optionally include one or more of the following features.
[0095] In some implementations, the method further includes: receiving additional user input associated with the client device of the user, the additional user input including a request to modify the lyric content and / or the music composition content; generating a modified version of the music content in response to the additional user input, the modified version of the music content including a modified version of the lyric content and / or a modified version of the music composition content; and causing the modified version of the music content to be audibly rendered at the client device.
[0096] In some versions of those implementations, generating the modified version of the music content in response to the additional user input may include: processing additional GM input using the GM to generate an additional GM output, the additional GM input including at least the additional user input and one or more seeds associated with the music content; and determining the modified version of the lyric content and / or the modified version of the music composition content based on the additional GM output.
[0097] In some further versions of those implementations, each of the one or more seeds associated with the music content may be a corresponding lower-level representation of the lyric content and / or the music composition content.
[0098] In some yet further versions of those implementations, the corresponding lower-level representation of the lyric content and / or the music composition content may be a corresponding embedding in an embedding space.
[0099] In some implementations, the method may further include: determining visual multimedia content to be visually rendered at the client device while the music content is being audibly rendered at the client device; and causing the visual multimedia content to be visually rendered via a display of the client device while the music content is being audibly rendered at the client device.
[0100] In some versions of those implementations, the visual multimedia content can be generative visual multimedia content.
[0101] In some further versions of those implementations, determining the visual multimedia content to be visually rendered at the client device when the music content is being audibly rendered at the client device can include: determining the generative visual multimedia content based on the GM output.
[0102] In some even further versions of those implementations, the generative visual multimedia content can be synchronized with the music content.
[0103] In additional or alternative further versions of those implementations, determining the visual multimedia content to be visually rendered at the client device when the music content is being audibly rendered at the client device can include: using image GM to process an image GM input to generate an image GM output, the image GM input including at least the user input; and determining the generative visual multimedia content based on the image GM output.
[0104] In some even further versions of those implementations, the method can further include: before causing the visual multimedia content to be visually rendered at the client device while the music content is being audibly rendered at the client device: synchronizing the generative visual multimedia content with the music content.
[0105] In additional or alternative further versions of those implementations, the image GM input can further include the lyric content and / or the music composition content.
[0106] In additional or alternative versions of those implementations, the visual multimedia content can be non - generative visual multimedia content.
[0107] In some further versions of those implementations, determining the visual multimedia content to be visually rendered at the client device when the music content is being audibly rendered at the client device can include: identifying one or more entities included in the request for the music content; and causing the non - generative visual multimedia content to be obtained based on one or more of the entities included in the request for the music content.
[0108] In some even further versions of those implementations, the method can further include: before causing the visual multimedia content to be visually rendered at the client device while the music content is being audibly rendered at the client device: synchronizing the non - generative visual multimedia content with the music content.
[0109] In even some further versions of those implementations, the non-generative visual multimedia content can be obtained from a visual multimedia content database that is personal to the user of the client device.
[0110] In some implementations, causing the music content to be aurally rendered at the client device can include: causing the lyric content to be aurally rendered via one or more speakers of the client device; and simultaneously causing the music composition content to be aurally rendered via one or more speakers of the client device.
[0111] In some versions of those implementations, the lyric content is aurally rendered in the voice of the user of the client device.
[0112] In some further versions of those implementations, the lyric content determined based on the GM output can be text corresponding to the lyric content, and the lyric content aurally rendered via the one or more speakers of the client device can include synthetic speech audio data capturing the lyric content.
[0113] In some even further versions of those implementations, the method can further include: using a text-to-speech (TTS) model to cause a representation of the lyric content and the voice of the user of the client device to be processed to generate the lyric content in the form of the voice of the user of the client device.
[0114] In some even further versions of those implementations, the representation of the voice of the user of the client device can include one or more of the following: a voice embedding of the user of the client device, or prosodic attributes of the voice of the user of the client device.
[0115] In some implementations, the GM output can include at least a first probability distribution over a first sequence of tokens and a second probability distribution over a second sequence of tokens.
[0116] In some versions of those implementations, the first sequence of tokens can be a sequence of words or word units or a sequence of phonemes or articulatory units, and determining the lyric content based on the GM output can include: determining the lyric content from the sequence of words or word units or the sequence of phonemes or articulatory units based on the first probability distribution.
[0117] In additional or alternative versions of those implementations, the second sequence of tokens can be a sequence of notes or note units, and determining the music composition content based on the GM output can include: determining the music composition content from the sequence of notes or note units based on the second probability distribution.
[0118] In some implementations, the method may further include: before receiving the user input associated with the client device of the user: fine-tuning the GM to generate the GM output, and determining the lyric content and the music composition content based on the GM output.
[0119] In some versions of those implementations, fine-tuning the GM may include: obtaining a plurality of first fine-tuning instances, each first fine-tuning instance in the plurality of first fine-tuning instances including a corresponding fine-tuning user input and a corresponding fine-tuning lyric content; and fine-tuning the GM based on the plurality of first fine-tuning instances.
[0120] In some versions of those implementations, fine-tuning the GM may further include: obtaining a plurality of second fine-tuning instances, each second fine-tuning instance in the plurality of second fine-tuning instances including a corresponding fine-tuning user input and a corresponding fine-tuning music composition content; and fine-tuning the GM based on the plurality of second fine-tuning instances.
[0121] In some implementations, the method may further include: before causing the music content to be audibly rendered at the client device: causing the lyric content to be visually rendered via a display of the client device; and determining whether a user confirmation to cause the music content to be audibly rendered at the client device is received. Causing the music content to be audibly rendered at the client device is in response to determining that the user confirmation to cause the music content to be audibly rendered at the client device has been received.
[0122] In some versions of those implementations, the method may further include: in response to determining that the user confirmation to cause the music content to be audibly rendered at the client device has not been received: refraining from causing the music content to be audibly rendered at the client device.
[0123] In some implementations, a method implemented by one or more processors is provided, and the method includes: receiving a user input associated with a client device of a user, the user input including a request for music content, the music content including lyric content and music composition content; and generating the music content in response to the user input. Generating the music content in response to the user input includes: processing a first GM input using a first generative model (GM) to generate a first GM output, the first GM input including at least the user input; determining the lyric content based on the first GM output; processing a second GM input using a second generative model (GM) to generate a second GM output, the second GM input including at least the user input; and determining the music composition content based on the second GM output. The method further includes: causing the lyric content and the music composition content to be audibly rendered at the client device.
[0124] These and other implementations of the technology disclosed herein may optionally include one or more of the following features.
[0125] In some implementations, the method may further include: receiving additional user input associated with the client device of the user, the additional user input including a request to modify the lyric content and / or the music composition content; generating a modified version of the music content in response to the additional user input, the modified version of the music content including a modified version of the lyric content and / or a modified version of the music composition content; and causing the modified version of the music content to be audibly rendered at the client device.
[0126] In some versions of those implementations, generating the modified version of the music content in response to the additional user input may include: determining that the request included in the additional user input is a request to modify the lyric content and / or a request to modify the music composition content; in response to determining that the request included in the additional user input is a request to modify the lyric content: processing additional first GM input using the first GM to generate additional first GM output, the additional first GM input including at least the additional user input and one or more seeds associated with the lyric content, and determining the modified version of the lyric content based on the additional first GM output; and in response to determining that the request included in the additional user input is a request to modify the music composition content: processing additional second GM input using the second GM to generate additional second GM output, the additional second GM input including at least the additional user input and one or more seeds associated with the music composition content; and determining the modified version of the music composition content based on the additional second GM output.
[0127] In some further versions of those implementations, the method further includes: in response to determining that the request included in the additional user input is not a request to modify the lyric content: avoiding any further processing using the first GM; and in response to determining that the request included in the additional user input is not a request to modify the music composition content, avoiding any further processing using the second GM.
[0128] In additional or alternative further versions of those implementations, each of the one or more seeds associated with the lyric content may be a corresponding lower-level representation of the lyric content.
[0129] In additional or alternative further versions of those implementations, each of the one or more seeds associated with the music composition content may be a corresponding lower-level representation of the music composition content.
[0130] In some implementations, the method may further include: determining visual multimedia content to be visually rendered at the client device while the music content is being audibly rendered at the client device; and causing the visual multimedia content to be visually rendered via a display of the client device while the music content is being audibly rendered at the client device.
[0131] In some versions of those implementations, the visual multimedia content may be generative visual multimedia content.
[0132] In some further versions of those implementations, determining the visual multimedia content to be visually rendered at the client device while the music content is being audibly rendered at the client device may include: determining the generative visual multimedia content based on a first GM output or a second GM output.
[0133] In some yet further versions of those implementations, the generative visual multimedia content may be synchronized with the music content.
[0134] In additional or alternative further versions of those implementations, determining the visual multimedia content to be visually rendered at the client device while the music content is being audibly rendered at the client device may include: using an image GM other than the first GM and other than the second GM to process an image GM input to generate an image GM output, the image GM input including at least the user input; and determining the generative visual multimedia content based on the image GM output.
[0135] In some yet further versions of those implementations, the method may further include: before causing the visual multimedia content to be visually rendered at the client device while the music content is being audibly rendered at the client device: synchronizing the generative visual multimedia content with the music content.
[0136] In additional or alternative further versions of those implementations, the image GM input may further include the lyric content and / or the music composition content.
[0137] In additional or alternative versions of those implementations, the visual multimedia content may be non - generative visual multimedia content.
[0138] In some further versions of those implementations, determining the visual multimedia content to be visually rendered at the client device when the music content is audibly rendered at the client device may include: identifying one or more entities included in the request for the music content; and causing the non-generative visual multimedia content to be obtained based on one or more of the entities included in the request for the music content.
[0139] In some even further versions of those implementations, the method may further include: before causing the visual multimedia content to be visually rendered at the client device while the music content is being audibly rendered at the client device: synchronizing the non-generative visual multimedia content with the music content.
[0140] In additional or alternative even further versions of those implementations, the non-generative visual multimedia content may be obtained from a visual multimedia content database that is personal to the user of the client device.
[0141] In some implementations, causing the music content to be audibly rendered at the client device may include: causing the lyric content to be audibly rendered via one or more speakers of the client device; and simultaneously causing the music composition content to be audibly rendered via one or more speakers of the client device.
[0142] In some even further versions of those implementations, the lyric content is audibly rendered in the voice of the user of the client device.
[0143] In some even further versions of those implementations, the lyric content determined based on the first GM output may be text corresponding to the lyric content, and the lyric content audibly rendered via the one or more speakers of the client device may include synthetic speech audio data capturing the lyric content.
[0144] In some even further versions of those implementations, the method may further include: using a text-to-speech (TTS) model to cause the lyric content and a representation of the voice of the user of the client device to be processed to generate the lyric content in the form of the voice of the user of the client device.
[0145] In some even further versions of those implementations, the representation of the voice of the user of the client device may include one or more of the following: a voice embedding of the user of the client device, or prosodic attributes of the voice of the user of the client device.
[0146] In some implementations, the first GM output may include at least a first probability distribution over a first sequence of tokens, and the second GM output may include a second probability distribution over a second sequence of tokens.
[0147] In some versions of those implementations, the first sequence of tokens may be a sequence of words or word units, and determining the lyric content based on the first GM output may include: determining the lyric content from the sequence of words or word units based on the first probability distribution.
[0148] In additional or alternative versions of those implementations, the second sequence of tokens may be a sequence of notes or note units, and determining the music composition content based on the second GM output may include: determining the music composition content from the sequence of notes or note units based on the second probability distribution.
[0149] In some implementations, the method may further include: before receiving the user input associated with the client device of the user: fine-tuning the first GM to generate the first GM output, determining the lyric content based on the first GM output; and fine-tuning the second GM to generate the second GM output, determining the music composition content based on the second GM output.
[0150] In some further versions of those implementations, fine-tuning the first GM may include: obtaining a plurality of first fine-tuning instances, each first fine-tuning instance of the plurality of first fine-tuning instances including a corresponding fine-tuning user input and a corresponding fine-tuning lyric content; and fine-tuning the first GM based on the plurality of first fine-tuning instances.
[0151] In additional or alternative versions of those implementations, fine-tuning the second GM includes: obtaining a plurality of second fine-tuning instances, each second fine-tuning instance of the plurality of second fine-tuning instances including a corresponding fine-tuning user input and a corresponding fine-tuning music composition content; and fine-tuning the second GM based on the plurality of second fine-tuning instances.
[0152] In some implementations, the method may further include: before causing the music content to be audibly rendered at the client device: causing the lyric content to be visually rendered via a display of the client device; and determining whether a user confirmation to cause the music content to be audibly rendered at the client device is received. Causing the music content to be audibly rendered at the client device may be in response to determining that the user confirmation to cause the music content to be audibly rendered at the client device has been received.
[0153] In some further versions of those implementations, the method may further include: in response to determining that the user confirmation that causes the music content to be audibly rendered at the client device has not been received: avoiding causing the music content to be audibly rendered at the client device.
[0154] Additionally, some implementations include one or more processors of one or more computing devices (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or a tensor processing unit (TPU)), where the one or more processors are operable to execute instructions stored in an associated memory, and where the instructions are configured to cause the execution of any of the methods described above. Some implementations further include one or more non-transitory computer-readable storage media that store computer instructions that can be executed by the one or more processors to execute any of the methods described above. Some implementations further include a computer program product that includes instructions that can be executed by the one or more processors to execute any of the methods described above.
[0155] It should be understood that all combinations of the foregoing concepts and additional concepts described in more detail herein are considered to be part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter that appear at the end of this disclosure are considered to be part of the subject matter disclosed herein.
Claims
1. A method implemented by one or more processors, the method comprising: Receiving user input associated with a client device of a user, the user input including a request for music content, and the music content including lyric content and music composition content; Generating the music content in response to the user input, wherein generating the music content in response to the user input includes: Processing GM input using a generative model (GM) to generate a GM output, the GM input including at least the user input; and Determining the lyric content and the music composition content based on the GM output; and Causing the music content to be audibly rendered at the client device.
2. The method according to claim 1, further comprising: Receiving additional user input associated with the client device of the user, the additional user input including a request to modify the lyric content and / or the music composition content; Generating a modified version of the music content in response to the additional user input, the modified version of the music content including a modified version of the lyric content and / or a modified version of the music composition content; And Causing the modified version of the music content to be audibly rendered at the client device.
3. The method according to claim 2, wherein generating the modified version of the music content in response to the additional user input includes: Processing additional GM input using the generative model to generate an additional GM output, the additional GM input including at least the additional user input and one or more seeds associated with the music content; And Determining the modified version of the lyric content and / or the modified version of the music composition content based on the additional GM output.
4. The method according to claim 3, wherein each of the one or more seeds associated with the music content is a corresponding lower-level representation of the lyric content and / or the music composition content.
5. The method according to claim 4, wherein the corresponding lower-level representation of the lyric content and / or the music composition content is a corresponding embedding in an embedding space.
6. The method according to claim 1, further comprising: Determining visual multimedia content to be visually rendered at the client device when the music content is being audibly rendered at the client device; And Causing the visual multimedia content to be visually rendered via a display of the client device when the music content is being audibly rendered at the client device.
7. The method according to claim 6, wherein the visual multimedia content is generative visual multimedia content.
8. The method according to claim 7, wherein determining the visual multimedia content to be visually rendered at the client device when the music content is being audibly rendered at the client device includes: Determining the generative visual multimedia content based on the GM output.
9. The method according to claim 8, wherein the generative visual multimedia content is synchronized with the music content.
10. The method according to claim 7, wherein determining the visual multimedia content to be visually rendered at the client device when the music content is audibly rendered at the client device comprises: Processing an image GM input using an image generative model to generate an image GM output, the image GM input comprising at least the user input; And Determining the generative visual multimedia content based on the image GM output.
11. The method according to claim 10, wherein the image GM input further comprises the lyric content and / or the music composition content.
12. The method according to claim 6, wherein the visual multimedia content is non-generative visual multimedia content.
13. The method according to claim 12, wherein determining the visual multimedia content to be visually rendered at the client device when the music content is audibly rendered at the client device comprises: Identifying one or more entities included in the request for the music content; And Causing the non-generative visual multimedia content to be obtained based on at least one of the one or more entities included in the request for the music content.
14. The method according to claim 13, further comprising: Before causing the visual multimedia content to be visually rendered at the client device when the music content is being audibly rendered at the client device: Synchronizing the non-generative visual multimedia content with the music content.
15. The method according to claim 13, wherein the non-generative visual multimedia content is obtained from a visual multimedia content database that is personal to the user of the client device.
16. The method according to claim 1, wherein causing the music content to be audibly rendered at the client device comprises: Causing the lyric content to be audibly rendered via one or more speakers of the client device; And Simultaneously causing the music composition content to be audibly rendered via the one or more speakers of the client device.
17. The method according to claim 16, wherein the lyric content is audibly rendered in the voice of the user of the client device.
18. A system, comprising: At least one processor; And A memory storing instructions that, when executed by the at least one processor, cause the at least one processor to be operable to: Receive a user input associated with a client device of a user, the user input comprising a request for music content, and the music content comprising lyric content and music composition content; Generate the music content in response to the user input, wherein generating the music content in response to the user input comprises: Processing a GM input using a generative model (GM) to generate a GM output, the GM input comprising at least the user input; and Determine the lyric content and the music composition content based on the GM output; and Cause the music content to be audibly rendered at the client device.
19. A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to be operable to perform operations, the operations including: Receiving user input associated with a client device of a user, the user input including a request for music content, and the music content including lyric content and music composition content; Generating the music content in response to the user input, wherein generating the music content in response to the user input includes: Processing a GM input using a generative model (GM) to generate a GM output, the GM input including at least the user input; and Determining the lyric content and the music composition content based on the GM output; and Causing the music content to be audibly rendered at the client device.