Generative alarm sounds

The system generates personalized alarm sounds using machine learning models, addressing the lack of customization in pre-loaded alarm sounds by incorporating text prompts and contextual signals, thereby reducing user input and enhancing the alarm setting experience.

WO2025221483A1PCT designated stage Publication Date: 2025-10-23GOOGLE LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/023149
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-16
Filing Date
2025-04-04
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Existing alarm sounds are pre-loaded and lack personalization, requiring extensive user input for customization.

Method used

A system that generates original audio clips based on text prompts and contextual signals using machine learning models, allowing users to customize alarm sounds by inputting a simple text prompt and obtaining contextual signals such as calendar events and environmental factors.

Benefits of technology

Enables personalized alarm sounds with reduced user effort by generating audio clips tailored to user preferences and contextual information, enhancing the alarm setting experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025023149_23102025_PF_FP_ABST
    Figure US2025023149_23102025_PF_FP_ABST
Patent Text Reader

Abstract

A system that includes at least one processor and a storage device that stores instructions executable by the at least one processor to obtain a text prompt. The storage device further stores instructions executable by the at least one processor to obtain one or more contextual signals. The storage device further stores instructions executable by the at least one processor to generate, by one or more generative models, audio data for an audio clip based on the text prompt and the one or more contextual signals. The storage device further stores instructions executable by the at least one processor to output audio data for the audio clip.
Need to check novelty before this filing date? Find Prior Art

Description

GENERATIVE ALARM SOUNDS

[0001] This application claims benefit of U.S. Provisional Application No. 63 / 634,647 filed April 16, 2024, the entire content of which is hereby incorporated by reference.BACKGROUND

[0002] Users may interact with a device application to set an alarm and a corresponding alarm sound to be output by a user device at a user-specified time (e.g., an alarm for waking up in the morning, an alarm to remind a user at certain moments during the day, etc.). Alarm sounds output by the user device are currently pre-loaded or include sounds and / or music that has already been recorded.SUMMARY

[0003] In general, described herein are techniques for generating original audio clips (e.g., alarm sounds) based on text prompts and contextual signals. An original sound generator may include one or more machine learning models (e.g., large language models, diffusion models) trained to generate music (e.g., a melody and / or harmony) based on a text prompt and contextual signals. The original sound generator may provide the one or more machine learning models a text prompt input by a user operating a user computing device. The text prompt may include an open-ended phrase conveying customizations and / or preferences for original audio clips that are oftentimes input by the user operating the user computing device. The original sound generator may additionally provide the one or more machine learning models contextual signals associated with the user computing device. Contextual signals may specify personalization and / or customization preferences for original audio clips (e.g., original alarm sounds) that are not explicitly conveyed in a text prompt. Contextual signals may include events based data of the user computing device (e.g., data representing a calendar or special occasions saved to a user device, data representing interests of a user operating the user device, etc.) and / or environmental factors associated with the user computing device (e.g., time of day, day of the week, weather, etc. at a location where the user device is positioned). The original sound generator may apply one or more machine learning models to generate an original audio clip based on a text prompt and contextual signals. The original sound generator may output the original audio clip to a software application of the user device. For example, the original sound generator may output the original audio clip to an alarmclock application of a user device, as well as instructions for how and / or when the alarm clock application will output the original audio clip via an output device.

[0004] In one example, a method includes obtaining, by one or more processors, a text prompt. The method may further include obtaining, by the one or more processors, one or more contextual signals. The method may further include generating, by one or more generative models executing on the one or more processors, audio data for an audio clip based on the text prompt and the one or more contextual signals. The method may further include outputting, by the one or more processors, audio data for the audio clip.

[0005] In another example, a system includes at least one processor and a storage device that stores instructions executable by the at least one processor to obtain a text prompt. The storage device further stores instructions executable by the at least one processor to obtain one or more contextual signals. The storage device further stores instructions executable by the at least one processor to generate, by one or more generative models, audio data for an audio clip based on the text prompt and the one or more contextual signals. The storage device further stores instructions executable by the at least one processor to output audio data for the audio clip.

[0006] In another example, a device comprises at least one processor and a storage device that stores instructions executable by the at least one processor. The instructions executable by the at least one processor may be configured to obtain a text prompt. The instructions executable by the at least one processor may further be configured to obtain one or more contextual signals. The instructions executable by the at least one processor may further be configured to generate, by one or more generative models, audio data for an audio clip based on the text prompt and the one or more contextual signals. The instructions executable by the at least one processor may further be configured to output audio data for the audio clip.

[0007] In another example, a computer-readable storage medium encoded with instructions, that when executed, cause at least one processor of a computing device to obtain a text prompt. The instructions may further cause the at least one processor to obtain one or more contextual signals. The instructions may further cause the at least one processor to generate, by one or more generative models, audio data for an audio clip based on the text prompt and the one or more contextual signals. The instructions may further cause the at least one processor to output audio data for the audio clip.

[0008] The details of one or more examples of the disclosure are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description and drawings, and from the claims.BRIEF DESCRIPTION OF DRAWINGS

[0009] FIG. l is a conceptual diagram illustrating an example computing environment for generating audio data for original audio clips as alarm sounds, in accordance with one or more techniques of this disclosure.

[0010] FIG. 2 is a block diagram illustrating a computing device for outputting audio data for original audio clips, in accordance with one or more techniques of this disclosure.

[0011] FIG. 3 is a flowchart illustrating an example operation for outputting audio data for original audio clips generated based on a text prompt and one or more contextual signals, in accordance with one or more techniques of this disclosure.DETAILED DESCRIPTION

[0012] FIG. 1 is a conceptual diagram illustrating example computing environment 101 for generating audio data for original audio clips as alarm sounds, in accordance with one or more techniques of this disclosure. As shown in the example of FIG. 1, computing environment 101 includes computing device 100 and computing system 150. Computing system 150 may include, but is not limited to, remote computing systems, such as one or more desktop computers, laptop computers, mainframes, servers, cloud computing systems, etc. capable of sending information to and receiving information from computing device 100 via a network or a wired connection. Computing device 100 may include, but is not limited to, portable, mobile, or other devices, such as mobile phones (including smartphones), wearable computing devices (e.g., smart watches, smart glasses, etc.), laptop computers, desktop computers, tablet computers, smart speakers, smart television platforms, server computers, mainframes, infotainment systems (e.g., vehicle head units), etc. In some examples, computing device 100 may represent a cloud computing system that provides one or more services via a network. That is, in some examples, computing device 100 may be a distributed computing system.

[0013] Computing device 100, in the example of FIG. 1, may include user interface device 102 (“UI device 102). UI device 102 of computing device 100 may be configured to function as an input device and / or an output device for computing device 100. UI device 102 may be implemented using various technologies. For instance, UI device 102may be configured to receive input from a user through tactile, audio, and / or video feedback. Examples of input devices include a presence-sensitive display, a presencesensitive or touch-sensitive input device, a mouse, a keyboard, a voice responsive system, video camera, microphone or any other type of device for detecting a command from a user. In some examples, a presence-sensitive display includes a touch-sensitive or presence-sensitive input screen, such as a resistive touchscreen, a surface acoustic wave touchscreen, a capacitive touchscreen, a projective capacitance touchscreen, a pressure sensitive screen, an acoustic pulse recognition touchscreen, or another presence-sensitive technology. That is, UI device 102 of computing device 100 may include a presencesensitive device that may receive tactile input from a user of computing device 100. UI device 102 may receive indications of the tactile input by detecting one or more gestures from the user (e.g., when the user touches or points to one or more locations of UI device 102 with a finger or a stylus pen).

[0014] UI device 102 may additionally or alternatively be configured to function as an output device by providing output to a user using tactile, audio, or video stimuli. Examples of output devices include a sound card, a video graphics adapter card, or any of one or more display devices, such as a liquid crystal display (LCD), dot matrix display, light emitting diode (LED) display, miniLED, microLED, organic light-emitting diode (OLED) display, e-ink, or similar monochrome or color display capable of outputting visible information to a user of computing device 100. Additional examples of an output device include a speaker, a haptic device, or other device that can generate intelligible output to a user. For instance, UI device 102 may present output to a user of computing device 100 as a graphical user interface that may be associated with functionality provided by computing device 100. In this way, UI device 102 may present various user interfaces of applications executing at or accessible by computing device 100 (e.g., an electronic message application, an Internet browser application, etc.). A user of computing device 100 may interact with a respective user interface of an application to cause computing device 100 to perform operations relating to a function.

[0015] In some examples, UI device 102 of computing device 100 may detect two- dimensional and / or three-dimensional gestures as input from a user of computing device 100. For instance, a sensor of UI device 102 may detect the user's movement (e.g., moving a hand, an arm, a pen, a stylus, etc.) within a threshold distance of the sensor of UI device 102. UI device 102 may determine a two- or three-dimensional vector representation of the movement and correlate the vector representation to a gesture input(e.g., a hand-wave, a pinch, a clap, a pen stroke, etc.) that has multiple dimensions. In other words, UI device 102 may, in some examples, detect a multidimensional gesture without requiring the user to gesture at or near a screen or surface at which UI device 102 outputs information for display. Instead, UI device 102 may detect a multi-dimensional gesture performed at or near a sensor which may or may not be located near the screen or surface at which UI device 102 outputs information for display.

[0016] In the example of FIG. 1, computing device 100 may include user interface module 104 (“UI module 104”), application modules 106A-106N (collectively “application modules 106”), and alarm sound generator client module 110. Modules 104, 106, and 110 may perform operations described herein using hardware, software, firmware, or a mixture thereof residing in and / or executing at computing device 100. Computing device 100 may execute modules 104, 106, and 110 with one processor or with multiple processors. In some examples, computing device 100 may execute modules 104, 106, and 110 as a virtual machine executing on underlying hardware. Modules 104, 106, and 110 may execute as one or more services of an operating system or computing platform or may execute as one or more executable programs at an application layer of a computing platform.

[0017] UI module 104, as shown in the example of FIG. 1, may be operable by computing device 100 to perform one or more functions, such as receive input and send indications of such input to other components associated with computing device 100, such as alarm sound generator client module 110 and / or application modules 106. UI module 104 may also receive data from components associated with computing device 100 such as alarm sound generator client module 110 and / or application modules 106. Using the data received, UI module 104 may cause other components associated with computing device 100, such as UI device 102, to provide output based on the data. For instance, UI module 104 may receive data from alarm sound generator client module 110 and / or one of application modules 106 to display a graphical user interface (GUI).

[0018] Application modules 106, as shown in the example of FIG. 1, may include functionality to perform any variety of operations on computing device 100. For instance, application modules 106 may include an alarm clock application, a word processor, a text application, a web browser, a multimedia player, a calendar application, an operating system, a distributed computing application, a graphic design application, a video editing application, a web development application, or any other application.Alarm sound generator client module 110 may include an application module similar toone of application modules 106. For example, alarm sound generator client module 110 may include a clock application and / or an alarm clock application. Alarm sound generator client module 110, for example, may include an alarm clock application with functionality to set alarms that play alarm sounds at particular times. Alarm sound generator client module 110, in various examples, may provide data to UI module 104 causing UI device 102 to display a GUI allowing a user to set alarms that output an original alarm sound generated according to the techniques described herein.

[0019] In accordance with the techniques described herein, computing device 100 may output audio data for an original audio clip generated based on a text prompt and contextual signals. A user operating computing device 100 may customize and personalize the original audio clip by inputting a simple text prompt. Alarm sound generator client module 110 of computing device 100, for example, may include an alarm clock application with functionality to initiate the generation of audio data for an original audio clip that may be used as an original alarm sound. Alarm sound generator client module 110 may provide data to UI module 104 causing UI device 102 to display a GUI that includes a field for a user operating computing device 100 to input a text prompt. In one example, UI device 102 may detect an audio input from a user operating computing device 100 representing speech of the user speaking a text prompt. UI device 102 may send the audio input to alarm sound generator client module 110 and / or alarm sound generator 108. Alarm sound generator client module 110 and / or alarm sound generator 108 may obtain the text prompt based on the input audio by at least applying speech-to- text techniques (e.g., acoustic modeling, language modeling, Hidden Markov Models, deep learning, etc.) to convert the audio input into a string representing the text prompt.

[0020] In another example, UI device 102 may provide the user operating computing device 100 the ability to provide inputs (e.g., to select letters, emojis, etc.) at a graphical keyboard, for example, to compose a text input that specifies preferences for generating audio data for the original audio clip. UI device 102 may detect input signals by the user operating computing device 100 representing the text prompt, and send the input signals to alarm sound generator client module 110 via UI module 104. For example, alarm sound generator client module 110 may obtain a text prompt including the string of “set wakeup alarm at 6AM on weekdays to play chill reggae music that changes every day” as a result of a user typing the text prompt with UI device 102. In some examples, alarm sound generator client module 110 may generate data to output a GUI to UI device 102, via UI module 104, that prompts a user to select a preferred mood to be applied to anoriginal audio clip that may be used as an alarm sound. Alarm sound generator client module 1110 may the selected mood in the text prompt, such that alarm sound generator 108 may generate audio data for the original audio clip based on the selected mood.

[0021] Alarm sound generator client module 110 may send the text prompt with the request to generate audio data for an original audio clip to alarm sound generator 108 of computing system 150 via a network, for example. Alarm sound generator 108 may include one or more machine learning models (e.g., language models) trained to generate audio data for original audio clips (e.g., original alarm sounds) based on a text prompt and / or contextual signals of preferences not explicitly included in the text prompt. Although illustrated as stored at computing system 150, some or all of the functionality of alarm sound generator 108 may be stored at computing device 100, as discussed in more detail with respect to FIG. 2. For example, some or all of the functionality of alarm sound generator 108 may be stored in an application module of computing device 100 (e.g., alarm sound generator client module 110).

[0022] Alarm sound generator 108 and / or alarm sound generator client module 110 may additionally or alternatively obtain contextual signals, after obtaining explicit permission from a user operating computing device 100, to further personalize an original audio clip (e.g., alarm sound) requested by a user. Alarm sound generator 108 and / or alarm sound generator client module 110 may obtain contextual signals to further personalize an original audio clip for a user based on information not explicitly conveyed in the text prompt. Alarm sound generator 108 and / or alarm sound generator client module 110 may obtain contextual signals that include data representing preferences for an original audio clip that may be used as an alarm sound. For example, alarm sound generator 108 may obtain, after receiving explicit consent from a user operating computing device 100, a contextual signal including an event (e.g., calendar event, meeting, special occasions) specified by data stored at an application module of application modules 106 (e.g., a calendar application). In another example, alarm sound generator 108 may alternatively or additionally obtain, after receiving explicit consent from a user operating computing device 100, a contextual signal including environmental factors (e.g., time of day, day of the week, weather, location, etc.) specified by data stored at an application module of application module 106 (e.g., a weather application, a location based service, a calendar application, etc.).

[0023] In some instances, alarm sound generator 108 may include a machine learning model trained to parse text to identify additional information that may improvepersonalization of an original audio clip that may be used as an alarm sound. For example, alarm sound generator 108 may apply the machine learning model to parse the text prompt of “set wakeup alarm at 6AM on weekdays to play chill reggae music that changes every day” to determine contextual intents of alarm sound customization and personalization based on context signals identified in the text prompt. In this example, alarm sound generator 108 may apply the machine learning model to determine a mood of “chill” specified in the text prompt. Additionally, or alternatively, alarm sound generator 108 may apply the machine learning model to determine additional information for contextual signals representing a time period to generate audio data for a new, original audio clip of “every day,” and / or an intent to output audio data for the original audio clip at “6AM on weekdays.” Alarm sound generator 108 may request the additional information from alarm sound generator client module 110, via a network. Alarm sound generator client module 110 may obtain the additional information as contextual signals from one or more application modules of application modules 106, after receiving explicit consent from a user operating computing device 100. Alarm sound generator client module 110 may send the additional information as contextual signals to alarm sound generator 108 via a network, for example.

[0024] In some instances, alarm sound generator client module 110 may obtain a contextual signal based on a text prompt. In one example, alarm sound generator client module 110 may include a machine learning model (e.g., a deep neural network) trained to identify semantic context (e.g., meaning or context) of text prompts. Alarm sound generator client module 110 may apply the machine learning model to parse a text prompt to identify and obtain contextual signals based on the text prompt. In some examples, alarm sound generator client module 110 may additionally or alternatively identify additional information that may be collected as contextual signals based on a contextual intent of the text prompt. Alarm sound generator client module 110 may request additional information (e.g., calendar information, time of day, day of the week, weather, location, etc.), after explicit consent from a user, from application modules 106 to improve personalization of an original audio clip that may be used as an alarm sound. Alarm sound generator client module 110 may obtain the additional information as contextual signals based on whether the machine learning model of alarm sound generator client module 110 determines the additional information is relevant to a contextual intent of the text prompt. For example, alarm sound generator client module 110 may obtain a text prompt of “nature sounds for my morning meeting,” and determine additionalinformation may be relevant to generating audio data for an original audio clip based on the text prompt. In this example, alarm sound generator client module 110 may determine a calendar associated with the user may be relevant (e.g., to understand when the morning meeting occurs) and / or a location of the user may be relevant (e.g., to generate audio data for original audio clips of nature sounds that are localized based on an area where computing device 100 is located). Alarm sound generator client module 110 may, after receiving explicit consent from a user operating computing device 100, obtain the calendar from a calendar application (e.g., application module 106B) and a location of the user from a Global Positioning System (GPS) receiver that may be included as part of computing device 100. Alarm sound generator client module 110 may generate contextual signals that include data representing the relevant calendar entry (e.g., “morning meeting”) and the location. Alarm sound generator client module 110 may send the contextual signals to alarm sound generator 108 via a network.

[0025] Throughout the disclosure, examples are described where a computing device (e.g., computing device 100) and / or a computing system (e.g., computing system 150) analyzes information (e.g., wireless ID tags and respective information, locations, context, motion, etc.) associated with a computing device and a user of the computing device, only if the computing device receives permission from the user of the computing device to analyze the information. For example, in situations discussed above and below, before computing device 100 or computing system 150 can collect or may make use of information associated with a user, the user may be provided with an opportunity to provide input to control whether programs or features of the computing device and / or computing system can collect and make use of user information (e.g., information about a user’s or user device’s current location, such as by GPS or wireless ID tag, etc.), or to dictate whether and / or how to the device and / or system may receive content that may be relevant to the user. In addition, certain data may be treated in one or more ways before it is stored or used by computing device 100 and / or computing system 150, so that personally identifiable information is removed. For example, a user’s identity and image may be treated so that no personally identifiable information can be determined about the user, or a user’s geographic location may be generalized where location information is obtained (such as to a city, ZIP code, or state level), so that a particular location of a user cannot be determined. Thus, the user may have control over how information is collected about the user and used by the computing device and computing system.

[0026] In some instances, alarm sound generator 108 and / or alarm sound generator client module 110 may generate a modified text prompt based on the text prompt and the one or more contextual signals. Alarm sound generator 108 may generate a modified text prompt that includes information associated with both the original text prompt input by a user operating computing device 100 and information associated with preferences specified in the obtained contextual signals. Alarm sound generator 108 may generate a contextual string for each contextual signal that conveys a preference represented in a corresponding contextual signal. Alarm sound generator 108 may append, prepend, or otherwise amend a text prompt input by a user operating computing device 100 to include the one or more contextual strings, each associated with corresponding one or more contextual signals. For example, alarm sound generator 108 may generate a modified text prompt of “set wakeup alarm at 6AM on weekdays to play chill reggae music that changes every day (mood: chill) (refresh time period: every day) (time of day: 6AM) (days of the week: weekdays)” based on determined information included in the example text prompt and example contextual signals previously discussed. In instances where alarm sound generator client module 110 modifies a text prompt with obtained contextual signals, alarm sound generator client module 110 may send the modified text prompt to alarm sound generator 108 of computing system 150 via a network.

[0027] Alarm sound generator 108 may apply one or more machine learning models to generate audio data for an audio clip based on a text prompt and one or more contextual signals. Alarm sound generator 108 may include one or more machine learning models (e.g., one or more language models) trained to generate music (e.g., melody and / or harmony) or other audio that may be used as an alarm sound output by computing device 100. In some examples, alarm sound generator 108 may provide the one or more machine learning models the text prompt and the one or more contextual signals. Alarm sound generator 108 may train and apply the machine learning model to generate original audio based on the text prompt and data included in the one or more contextual signals.

[0028] In some examples, alarm sound generator 108 may provide the one or more machine learning models the modified text prompt that includes a text prompt input by a user operating computing device 100 and obtained contextual strings. Alarm sound generator 108 may apply the one or more machine learning models to generate audio data for an audio clip based on the modified text prompt. For example, alarm sound generator 108 may provide the modified text prompt to a generative model that includes three independently, pre-trained models for conditional autoregressive music generation. Thegenerative model of alarm sound generator 108 may generate tokens based on the modified text prompt and predict audio tokens of an original audio clip based on the text prompt tokens. For example, alarm sound generator 108 may include a generative model pre-trained to tokenize a text prompt (e.g., the modified text prompt) and / or contextual signals by generating prompt tokens that include vector representations of the text prompt and / or contextual signals in a high-dimensional space. The generative model of alarm sound generator 108 may process the prompt tokens through layers of one or more machine learning models (e.g., neural networks) to determine a context of the text prompt and / or contextual signals. The generative model may predict audio tokens in an autoregressive manner based on the determined context with a token generation strategy (e.g., simple sampling, nucleus sampling, beam search, etc.). The generative model may predict audio tokens as short audio sounds that may represent one or more frames of the original audio clip.

[0029] In some instances, alarm sound generator 108 may generate audio data for a plurality of original audio clips based on a text prompt and one or more contextual signals. Alarm sound generator 108 may determine a set of candidate audio clips based on audio data of the plurality of generated audio clips. Alarm sound generator 108 may output audio data for the set of candidate audio clips to alarm sound generator client module 110 of computing device 100 via a network. Alarm sound generator client module 110 may send data to UI device 102, via UI module 104, to output audio data for the set of candidate audio clips. Alarm sound generator client module 110 may send data to UI device 102 that includes audio data for the set of candidate audio clips and an option for a user to select a candidate audio clip of the set of the candidate audio clips. UI device 102 may detect a signal specifying an indication of user feedback as a selection of a candidate audio clip. UI device 102 may send the signal specifying the indication to alarm sound generator client module 110. Alarm sound generator client module 110 may process the signal and send the indication of the selected candidate audio clip to alarm sound generator 108 via a network. Alarm sound generator 108 may tune the one or more machine learning models based on the indication of the selected candidate audio clip. In this way, alarm sound generator 108 may tune the one or more machine learning models to generate audio data for original audio clips that are personalized based on user feedback.

[0030] In some instances, alarm sound generator 108 may generate alarm instructions in addition to audio data for the original audio clip. Alarm sound generator 108 maygenerate alarm instructions based on the text prompt and one or more contextual signals. Alarm sound generator 108 may apply a generative model to process the text prompt and one or more contextual signals to identify customization and personalization preferences. For example, alarm sound generator 108 may apply the generative model to identify that a text prompt requested an original audio clip for an alarm sound that changes volume over time. In some instances, alarm sound generator 108 may generate audio data for an original audio clip that includes audio waveforms with changing attributes, such as a change in amplitude, intensity, or speed of the audio waveforms over time. In other instances, alarm sound generator 108 may generate alarm instructions for computing device 100 to change the attributes of audio data of the audio clip output by computing device 100 when playing back original audio clips based on audio data for the audio clip. In another example, alarm sound generator 108 may apply the generative model to identify that a contextual signal (e.g., a contextual signal obtained from a calendar application after explicit consent from a user) specifies that a user of computing device 100 is on vacation during a week. Alarm sound generator 108 may apply the generative model to generate alarm instructions to output a request to the user to temporarily disable, or otherwise adjust, one or more alarms during the week of vacation, in order to configure less intrusive alarms, for example.

[0031] In the example previously discussed, alarm sound generator 108 may generate and output audio data for an audio clip that includes a melody and vocal sounds corresponding to “chill reggae music” based on the text prompt of “set wakeup alarm at 6AM on weekdays to play chill reggae music that changes every day” and any obtained contextual signals. Alarm sound generator 108 may generate and output audio data for a new audio clip that includes a different melody and different vocal sounds corresponding to “chill reggae music” daily, according to the text prompt. In some instances, alarm sound generator 108 may generate and output audio data for the audio clip with alarm instructions on when and / or how computing device 100 should output the generated audio data for audio clip. For example, alarm sound generator 108 may generate audio data for an audio clip with alarm instructions for alarm sound generator client module 110 to send data to UI device 102, via UI module 104, to output the generated audio data for audio clip at 6:00 A.M. on weekdays. Alarm sound generator 108 may output audio data for audio clip and alarm instructions to alarm sound generator client module 110 via a network.

[0032] Alarm sound generator 108 may output audio data for the original audio clip to computing device 100 via a network. In some instances, alarm sound generator 108 may output audio data for the original audio clip and alarm instructions. Computing device 100 may store audio data for the original audio clip and any alarm instructions generated by alarm sound generator 108. For example, alarm sound generator client module 110 may determine when to output audio data for the audio clip based on the alarm instructions specifying an alarm time corresponding to the audio clip. When alarm sound generator client module 110 determines the time maintained by an operating system of computing device 100 is equal to the alarm time, alarm sound generator client module 110 may send data to UI device 102, via UI module 104, to output audio data for the original audio clip. Alarm sound generator client module 110 may send data to UI device 102 that includes audio data for the original audio clip and alarm instructions for outputting audio data for the original audio clip, such as increasing the volume of UI device 102 over time when outputting the original audio clip.

[0033] The techniques may provide one or more technical advantages that realize one or more practical applications. For example, computing device 100 may output audio data for an original audio clip that a user operating computing device 100 may personalize and customize. Computing device 100, or more specifically alarm sound generator client module 110, may include an alarm clock application that improves a user’s experience with setting alarms. Alarm sound generator client module 110 may conventionally require a user operating computing device 100 to input a limited number of selections (e.g., set time, select pre-loaded alarm sound, etc.), which may result in many user inputs and extensive user effort when trying to customize or personalize an alarm. By generating audio data for original audio clips based on a text prompt and contextual signals, alarm sound generator 108 may generate alarm clock sounds and instructions for outputting generated alarm clock sound according to a user’s preference. In this way, computing device 100 may implement personalized clock alarms based on a simple text prompt that a user operating computing device 100 may customize.

[0034] FIG. 2 is a block diagram illustrating example computing device 200 for setting alarms that output audio data for generated audio clips, in accordance with one or more techniques of this disclosure. Computing device 200, alarm sound generator client module 210, and alarm sound generator 208 may be one example of computing device 100, alarm sound generator client module 110, and alarm sound generator 108, respectively, in accordance with one or more techniques of this disclosure. FIG. 2illustrates only one particular example of computing device 200, and many other examples of computing device 200 may be used in other instances and may include a subset of components included in example computing device 200 or may include additional components not shown in FIG. 2.

[0035] As shown in FIG. 2, computing device 200 may include one or more user interface devices 202 (“UI device 202” or “display 202”), one or more processors 224 (“processor 224”), one or more storage devices 228 (“storage device 228”), one or more communication units 226 (“communication unit 226”). Also shown in FIG. 2, UI device 202 may include one or more input devices 234 (“input device 234”) and one or more output devices (“output devices 236”). As also shown in FIG. 2, storage device 228 may include user interface module 204 (“UI module 204”), one or more application modules 206 (“application module 206”), alarm sound generator client module 210, alarm sound generator 208, operating system 230 (“OS 230”), and database 232.

[0036] Communication channels 250 (“COMM channel 250”) may interconnect each of the components 202, 224, 226, and 228 for inter-component communications (physically, communicatively, and / or operatively). In some examples, communication channel 250 may include a system bus, a network connection, an inter-process communication data structure, or any other method for communicating data.

[0037] Communication unit 226 of computing device 200 may communicate with one or more external devices via one or more wired and / or wireless networks by transmitting and / or receiving network signals on the one or more networks. Examples of communication units 226 include a network interface card (e.g., such as an Ethernet card), an optical transceiver, a radio frequency transceiver, a GNSS receiver, or any other type of device that can send and / or receive information. Other examples of communication unit 226 may include short wave radios, cellular data radios (for terrestrial and / or satellite cellular networks), wireless network radios, as well as universal serial bus (USB) controllers.

[0038] In some examples, UI device 202 may be a presence-sensitive display configured to detect input (e.g., touch and non-touch input) from a user of respective computing device 200. UI device 202 may output information to a user in the form of a UI, which may be associated with functionality provided by computing device 200. Such UIs may be associated with computing platforms, operating systems, applications, and / or services executing at or accessible from computing device 200 (e.g., alarm applications, assistant applications, electronic message applications, chat applications, Internet browserapplications, mobile or desktop operating systems, social media applications, electronic games, menus, and other types of applications). Computing device 200 may also output, via UI device 202, an option for a user operating computing device 200 to grant computing device 200 explicit consent to provide information to alarm sound generator 208, such as a text prompt or one or more contextual signals.

[0039] Input device 234 of computing device 200 may receive input. Examples of input are tactile, audio, and video input. Input device 234 of computing device 200, in one example, includes a presence-sensitive display, a fingerprint sensor, touch-sensitive screen, mouse, keyboard, voice responsive system, video camera, microphone or any other type of device for detecting input from a human or machine.

[0040] Input devices 234 may include one or more sensors. Numerous examples of sensors exist and include any input component configured to obtain environmental information about the circumstances surrounding computing device 200 and / or physiological information that defines the activity state and / or physical well-being of a user of computing device 200. In some examples, a sensor may be an input component that obtains physical position, movement, and / or location information of computing device 200. For instance, sensors may include one or more location sensors (e.g., GNSS components, Wi-Fi components, cellular components), one or more temperature sensors, one or more motion sensors (e.g., multi-axial accelerometers, gyros), one or more pressure sensors (e.g., barometer), one or more ambient light sensors, and one or more other sensors (e.g., microphone, camera, infrared proximity sensor, hygrometer, and the like). Other sensors may include a heart rate sensor, magnetometer, glucose sensor, hygrometer sensor, olfactory sensor, compass sensor, step counter sensor, to name a few other non-limiting examples.

[0041] Output device 236 of computing device 200 may generate one or more outputs. Examples of outputs are tactile, audio, and video output. Output device 236 of computing device 200, in one example, includes a presence-sensitive display, sound card, video graphics adapter card, speaker, liquid crystal display (LCD), or any other type of device for generating output to a human or machine.

[0042] Processor 224 may implement functionality and / or execute instructions within computing device 200. For example, processor 224 may receive and execute instructions that provide the functionality of modules 204-208 and OS 230. These instructions executed by processor 224 may cause computing device 200 to store and / or modify information within storage device 228 or processor 224 during program execution.Processor 224 may execute instructions of modules 204-208 and OS 230 to perform one or more operations. That is modules 204-208 and OS 230 may be operable by processor 224 to perform various functions described herein.

[0043] Storage device 228 within computing device 200 may store information for processing during operation of computing device 200 (e.g., computing device 200 may store data accessed by modules 204-208 and OS 230 during execution at computing device 200). In some examples, storage device 228 may be a temporary memory, meaning that a primary purpose of storage device 228 is not long-term storage. Storage device 228 on computing device 200 may be configured for short-term storage of information as volatile memory and therefore not retain stored contents if powered off. Examples of volatile memories include random access memories (RAM), dynamic random access memories (DRAM), static random access memories (SRAM), and other forms of volatile memories known in the art.

[0044] Storage device 228 may include one or more computer-readable storage media. Storage device 228 may be configured to store larger amounts of information than volatile memory. Storage device 228 may further be configured for long-term storage of information as non-volatile memory space and retain information after power on / off cycles. Examples of non-volatile memories include magnetic hard discs, optical discs, floppy discs, flash memories, or forms of electrically programmable memories (EPROM) or electrically erasable and programmable (EEPROM) memories. Storage device 228 may store program instructions and / or information (e.g., within database 232) associated with modules 204-208 and OS 230.

[0045] Computing device 200 may include OS 230. OS 230 may control the operation of components of computing device 200. For example, OS 230 may facilitate the communication of modules 204-208 with processor 224, storage device 228, and communication units 226. In some examples, OS 230 may manage interactions between software applications (e.g., application module 206) and a user of computing device 200. OS 230 may have a kernel that facilitates interactions with underlying hardware of computing device 200 and provides a fully formed application space capable of executing a wide variety of software applications having secure partitions in which each of the software applications executes to perform various operations. In some examples, UI module 204 may be considered a component of OS 230.

[0046] In accordance with the techniques described herein, alarm sound generator 208 may generate audio data for original audio clips that may be used as alarm sounds. Someor all of the functionality of alarm sound generator 208 may be stored at an application module of computing device 200 (e.g., alarm sound generator client module 210). In the example of FIG. 2, alarm sound generator 208 includes prompt module 240 and generative models 242. Prompt module 240 may include text prompt 244, contextual signals 246, and machine learning model 248. Text prompt 244 may include a string specifying preferences for an original audio clip that may be used as an alarm sound. Prompt module 240 may obtain text prompt 244 as an input by a user interacting with computing device 200. Alarm sound generator client module 210 , for example, may include an alarm application configured to generate data for a graphical user interface that prompts a user to input text prompt 244 to set a new alarm with originally generated audio data for audio clips. Alarm sound generator client module 210 may output the data for the graphical user interface to output device 236, via UI module 204. Input device 234 may detect signals representing a user inputting words that make up text prompt 244. Input device 234 may send the signal data to prompt module 240, via UI module 204. Prompt module 240 may store the signal data as text prompt 244. For example, prompt module 240 may store text prompt 244 that includes a string of “calm summer forest sounds that start out quiet and rise in volume over time.”

[0047] In some instances, prompt module 240 may randomly generate text prompt 244. Alarm sound generator client module 210 may generate data for a graphical user interface to allow a user to select a “fun” or “random” mode for an alarm sound. Alarm sound generator client module 210 may send the data, via UI module 104, to output devices 236 to present the “fun” or “random” mode to the user operating computing device 200. Input devices 234 may detect a signal indicating a user has selected the “fun” or “random” mode for an alarm. Prompt module 240 may obtain, via UI module 204, the indication of the user selecting the “fun” or “random” mode. Prompt module 240 may apply machine learning model 248 to randomly generate a text prompt from an alarm sound as text prompt 244. In some examples, prompt module 240 may send the randomly generated text prompt to alarm sound generator client module 210 . Alarm sound generator client module 210 may generate data for a graphical user interface that includes the randomly generated text prompt and an option for a user operating computing device 200 to select whether to accept the randomly generated text prompt. Responsive to a user selecting to reject the randomly generated text prompt, prompt module 240 may randomly generate another text prompt until a user accepts the randomly generated text prompt. By receiving indications of a user approving and rejecting randomly generated text prompts,prompt module 240 may tune machine learning model 248 to generate random text prompts that are personalized to the user.

[0048] Contextual signals 246 may include data representing preferences for an original audio clip that may not be explicitly conveyed in text prompt 244. Prompt module 240 may obtain, after receiving explicit consent from a user operating computing device 200, contextual signals 246 from application module 206 and / or other data that may be stored at storage devices 228. Contextual signals 246 may include events based on data of computing device 200 and / or environmental factors associated with computing device 200. For example, contextual signals 246 may include calendar data from a calendar application of application module 206, weather data from a weather application of application module 206, or the like.

[0049] In some instances, prompt module 240 may obtain contextual signals 246 based on a contextual intent of text prompt 244. Prompt module 240 may apply machine learning model 248 to determine a contextual intent of text prompt 244 and output one or more contextual signals to be stored as contextual signals 246. Machine learning model 248 may include a machine learning model (e.g., neural network, language model, etc.) trained to determine additional information that may be included as part of contextual signals 246. For example, machine learning model 248 may parse text prompt 244 to determine additional information that may be relevant to generation of audio data for an original audio clip. Prompt module 240 may obtain the additional information (e.g., from application module 206) based on the indications output by machine learning model 248.

[0050] Prompt module 240 may provide text prompt 244 and contextual signals 246 to generative models 242. Generative models 242 may include generative machine learning models trained to generate sounds based on a text prompt. For example, generative models 242 may include a diffusion machine learning model and / or language models trained for conditional autoregressive music generation. In examples where generative models 242 include diffusion machine learning models, generative models 242 may generate audio data for an audio clip by at least noising and denoising an audio signal based at least on a modified text prompt generated based on text prompt 244 and contextual signals 246. Generative models 242, in the example of FIG. 2, may include generative model 242 A and generative model 242B. Generative model 242 A may include a language model trained to generate melodies based on text prompt 244 and contextual signals 246. Generative model 242A may generate text prompt tokens by processing text prompt 244 and contextual signals 245 that include vector representations of text prompt244 and / or contextual signals 246 in a high-dimensional space. Generative model 242A may process the prompt tokens through layers of one or more machine learning models (e.g., neural networks) to determine a context of text prompt and / or contextual signals 246. In examples where generative models 242 include autoregressive machine learning models, generative model 242A may autoregressively predict audio tokens based on the determined context. Generative model 242A may predict an audio token as a frame of a melody to be included in audio data for an original audio clip. Generative model 242A may apply the text prompt tokens as conditioning signals for an autoencoder of generative model 242Ato convert audio tokens to waveforms of the melody to be included audio data for in an original audio clip.

[0051] Generative model 242B may include a language model trained to generate vocal sounds based on text prompt 244 and contextual signals 246. Generative model 242B may generate text prompt tokens by processing text prompt 244 and contextual signals245 that include vector representations of text prompt 244 and / or contextual signals 246 in a high-dimensional space. Generative model 242B may process the prompt tokens through layers of one or more machine learning models (e.g., neural networks) to determine a context of text prompt and / or contextual signals 246. In examples where generative models 242 include autoregressive machine learning models, generative model 242B may autoregressively predict audio tokens based on the determined context. Generative model 242B may predict an audio token as a frame of vocal sounds to be included in audio data for an original audio clip. Generative model 242A may apply the text prompt tokens as conditioning signals for an autoencoder of generative model 242A to convert audio tokens to waveforms of the vocal sounds to be included in audio data for an original audio clip.

[0052] Generative models 242 may combine the waveforms of the melody generated by generative model 242A and the waveforms of the vocal sounds generated by generative model 242B to generate audio data for an original audio clip. For example, generative models 242 may concatenate the waveforms of a melody, generated by generative model 242A, that includes a short, noisy song clip and the waveforms of vocal sounds, generated by generative model 242B, that includes speech of a drill sergeant yelling a command based on text prompt 244 specifying to create an alarm that starts with noisy clips and ends with a drill sergeant yelling the command. Generative models 242 may generate audio data for original audio clips in periodic intervals. For example, generative models 242 may generate audio data for a new original audio clip every day as a new alarmsound based on text prompt 244 and contextual signals 246. Generative models 242 may output audio data for the original audio clip to alarm sound generator client module 210 that may generate data to output audio data for the original audio clip as an alarm sound to output devices 236, via UI module 204.

[0053] Generative models 242 may be tuned to generate audio data for original audio clips for alarm sounds based on sample audio data. For example, database 232 may include sample audio data input by a user of computing device 200, such as music samples, a plurality of alarm tones, and / or user feedback (e.g., user feedback based on selections of candidate original audio clips to use as an alarm sound). In one example, text prompt 244 may specify a preference to generate audio data for an original audio clip based on a music sample. Generative models 242 may obtain a music sample input by a user operating computing device 200 and tune each of generative models 242 based on the music sample.

[0054] Generative models 242 may generate audio data for original audio clips based on text prompt 244 and contextual signals 246. In one example, generative models 242 may generate audio data for original audio clips that change attributes (e.g., amplitude, intensity, speed, etc.) of audio in the original audio clip. For example, text prompt 244 may include the string of “jazz sounds that start out slow and increase speed over time.” Generative models 242 may determine an indication to change attributes of audio data for the audio clip as “starts out slow and increases speed over time.” Generative models 242 may generate, based on the indication to change attributes, audio data for an original audio clip with audio that increases speed over time, such that audio data for the original audio clip starts out with slow audio (e.g., audio with a playback speed of 0.5) that gets faster over time (e.g., increasing the audio playback speed to 1.5). In another example, generative models 242 may determine a mood of “energized,” based on an example text prompt, and generate audio data for an original audio clip that is upbeat. In another example, generative models 242 may obtain a contextual signal of contextual signals 246 specifying a time of day audio data for an originally generated audio clip may be output via output device 236. For example, generative models 242 may generate audio data for a soothing audio clip based on contextual signals 246 specifying audio data for the audio clip will be output in the morning, or may generate audio data for a louder, more energetic, audio clip based on contextual signals 246 specifying audio data for the audio clip will be output in the middle of the day.

[0055] In another example, generative models 242 may generate audio data for original audio clips based on contextual signals 246 specifying a day of the week or a user’s schedule for a given day (e.g., generative models 242 generates a loud alarm for urgent events on a user’s schedule or generative models 242 generates a soothing alarm for weekends). Generative models 242 may determine an event based on the one or more contextual signals. Generative models 242 may generate audio data for a new audio clip based on the event. For example, generative models 242 may obtain a text prompt specifying a user’s preference to create alarm sounds that include audio data for audio clips of nature sounds that correspond to the current season where the user is located, as well as contextual signals representing a user’s location, weather information, date and time information, or the like. Generative models 242 may determine an event, with respect to the text prompt, based on the contextual signals indicating the user is at a location that has a change in season compared to a previous day, for example. For example, generative models 242 may generate audio data for a new audio clip based on the event of changing seasons.

[0056] In some instances, generative models 242 may generate audio data for a plurality of original audio clips. Generative models 242 may assign an audio clip value between 0 and 1 to each of the plurality of original audio clips based on a confidence in an original audio clip being of interest to a user. Generative models 242 may determine a set of candidate original audio clips based on the audio clip values of each of the plurality of original audio clips. For example, generative models 242 may determine the set of candidate original audio clips as audio clips of the plurality of audio clips that are assigned an audio clip value satisfying a threshold. Generative models 242 may output audio data for the set of candidate original audio clips to alarm sound generator client module 210 . Alarm sound generator client module 210 may generate data for a graphical user interface to present the set of candidate original audio clips and allow the user to select an original audio clip of the set of candidate original audio clip as an alarm sound. In some examples, alarm sound generator client module 210 may generate data for a graphical user interface that allows a user to rank or otherwise provide feedback for the candidate original audio clips. Alarm sound generator client module 210 may provide the user feedback specifying the user ranking or selections to generative models 242. Responsive to a user rejecting each candidate original audio clip, generative models 242 may generate a new plurality of original audio clips, determine a new set of candidate original audio clips, and output audio data for the new set of candidate original audioclips for a user to select and / or provide feedback. Generative models 242 may tune each language model based on the user feedback. In this way, generative models 242 become fine-tuned to generate original audio clips that are personalized to a user’s preferences.

[0057] In some instances, alarm instruction generator 238 may generate alarm instructions for alarm sound generator client module 210 to output audio data for the original audio clip via output device 236. Alarm instruction generator 238 may generate alarm instructions based on text prompt 244 and / or contextual signals 246. Alarm instructions may include instructions for changing attributes of audio data for a generated audio clip based on customization and / or personalization preferences specified in text prompt 244 and / or contextual signals 246. For example, text prompt 244 may include the string of “calm summer forest sounds that start out quiet and rise in volume over time.” Alarm instruction generator 238 may identify an indication to change an attribute (e.g., volume, intensity, speed, etc.) of audio data for an original audio clip over time based on text prompt 244 and / or contextual signals 246. For example, alarm instruction generator 238 may generate alarm instructions for audio of output device 236 to progressively increase in intensity when outputting audio data for the original audio clip associated with text prompt 244. In another example, contextual signals 246 may include a preference to increase the volume of an original audio clip during successive snooze cycles. Alarm instruction generator 238 may generate alarm instructions that specify when a user selects “snooze” in response to output device 236 outputting audio data for the original audio clip, output device 236 will increase the volume of audio of output device 236 when outputting audio data for the original audio clip during the snooze cycle. Alarm instruction generator 238 may output alarm instructions to alarm sound generator client module 210. Alarm sound generator client module 210 may implement the alarm instructions when generating data to output audio data for an original audio clip generated by generative models 242 via output device 236.

[0058] FIG. 3 is a flowchart illustrating an example operation for outputting audio data for original audio clips generated based on a text prompt and one or more contextual signals, in accordance with one or more techniques of this disclosure. FIG. 3 may be discussed with respect to FIG. 2 for example purposes only.

[0059] Computing device 200 may obtain a text prompt (302). For example, computing device 200 may obtain text prompt 244 that includes a simple phrase of preferences a user operating computing device 200 wants in an original audio clip that may be output as an alarm sound. Computing device 200, or more specifically alarm sound generator 208,may obtain one or more contextual signals (304). Alarm sound generator 208 may obtain one or more contextual signals that include customization and personalization preferences that may provide further context for generating audio data for an original audio clip. For example, alarm sound generator 208 may obtain, after explicit consent from a user operating computing device 200, contextual signals such as event data or environmental factor data that may be stored at computing device 200 (e.g., stored at any of application modules 206).

[0060] Computing device 200 may generate, by one or more generative models, audio data for an audio clip based on the text prompt and contextual signals (306). Alarm sound generator 208 of computing device 200 may apply generative models 242 to generate audio data for an original audio clip based on the text prompt and contextual signals. In some instances, generative models 242 may include language models that are pre-trained to generate sounds based on text prompt and contextual signal data. In some examples, alarm sound generator 208 may modify text prompt 244 to include data of contextual signals 246, and provide generative models 242 the modified text prompt. Generative models 242 may generate audio data for one or more audio clips based on a single text prompt that may be continuously modified based on updated contextual signals.

[0061] Computing device 200 may output audio data for the audio clip (308). Computing device 200 may output audio data for the audio clip as an alarm sound via output device 236, for example. Alarm sound generator client module 210 may include an alarm clock application that generates data corresponding to when and how audio data for an original audio clip generated by alarm sound generator 208 may be output. Alarm sound generator 208 may provide alarm sound generator client module 210 alarm instructions that specify, for example, a date and time audio data for an original audio clip is output as an alarm sound and / or additional alarm instructions such as whether to increase the volume of computing device 200 over a period of time.

[0062] Example 1 : A method includes obtaining, by one or more processors, a text prompt; obtaining, by the one or more processors, one or more contextual signals; generating, by one or more generative models executing on the one or more processors, audio data for an audio clip based on the text prompt and the one or more contextual signals; and outputting, by the one or more processors, audio data for the audio clip.

[0063] Example 2: The method of example 1, wherein the one or more contextual signals includes data representing preferences for the audio clip associated with at least one of an event based on data associated with a user device and one or more environmental factors.

[0064] Example 3 : The method of any of examples 1 and 2, wherein generating audio data for the audio clip comprises: determining, by the one or more generative models, a mood based on the text prompt; and generating audio data for the audio clip based at least on the mood.

[0065] Example 4: The method of any of examples 1 through 3, wherein generating audio data for the audio clip comprises: generating a modified text prompt based on the text prompt and the one or more contextual signals; generating a plurality of prompt tokens based on the modified text prompt; and predicting a plurality of audio tokens based on the plurality of prompt tokens.

[0066] Example 5: The method of any of examples 1 through 4, wherein generating audio data for the audio clip comprises: generating a modified text prompt based on the text prompt and the one or more contextual signals; and noising and denoising an audio signal based at least on the modified text prompt.

[0067] Example 6: The method of any of examples 1 through 5, wherein obtaining the one or more contextual signals comprises: requesting additional information from one or more application modules of a user device.

[0068] Example 7: The method of any of examples 1 through 6, further includes generating, based on the text prompt and the one or more contextual signals, audio data for a new audio clip in periodic intervals.

[0069] Example 8: The method of any of examples 1 through 7, further includes determining, based on the one or more contextual signals, an event; and generating audio data for a new audio clip based on the event.

[0070] Example 9: The method of any of examples 1 through 8, further includes providing sample audio data to the one or more generative models, wherein the sample audio data includes at least one of a plurality of music samples, a plurality of alarm tones, and user feedback; and tuning the one or more generative models based on the sample audio data.

[0071] Example 10: The method of any of examples 1 through 9, further includes identifying an indication to change attributes of the audio clip over time; and instructing a user device to output audio data for the audio clip based on the indication.

[0072] Example 11 : The method of any of examples 1 through 10, wherein generating audio data for the audio clip includes: determining, based at least on the text prompt, an indication to change attributes of the audio clip; and generating audio data for the audio clip based on the indication.

[0073] Example 12: The method of any of examples 1 through 11, further includes generating audio data for a plurality of audio clips based on the text prompt and the one or more contextual signals; determining a set of candidate audio clips based on the plurality of audio clips; outputting audio data for the set of candidate audio clips to a user device; obtaining, from the user device, user feedback for the set of candidate audio clips; and tuning the one or more generative models based on the user feedback.

[0074] Example 13: The method of any of examples 1 through 12, further includes randomly generating the text prompt.

[0075] Example 14: A system includes at least one processor; and a storage device that stores instructions executable by the at least one processor to: obtain a text prompt; obtain one or more contextual signals; generate, by one or more generative models, audio data for an audio clip based on the text prompt and the one or more contextual signals; and output audio data for the audio clip.

[0076] Example 15: The system of example 14, wherein the one or more contextual signals includes data representing preferences for the audio clip associated with at least one of an event based on data associated with a user device and one or more environmental factors.

[0077] Example 16: The system of any of examples 14 and 15, wherein to generate audio data for the audio clip, the storage device stores instructions executable by the at least one processor to: determine, by the one or more generative models, a mood based on the text prompt; and generate audio data for the audio clip based at least on the mood.

[0078] Example 17: The system of any of examples 14 through 16, wherein to generate audio data for the audio clip, the storage device stores instructions executable by the at least one processor to: generate a modified text prompt based on the text prompt and the one or more contextual signals; generate a plurality of prompt tokens based on the modified text prompt; and predict a plurality of audio tokens based on the plurality of prompt tokens.

[0079] Example 18: The system of any of examples 14 through 17, wherein to generate audio data for the audio clip, the storage device stores instructions executable by the at least one processor to: generate a modified text prompt based on the text prompt and the one or more contextual signals; and noise and denoise an audio signal based at least on the modified text prompt.

[0080] Example 19: The system of any of examples 14 through 18, wherein to obtain the one or more contextual signals, the storage device stores instructions executable by the atleast one processor to: request additional information from one or more application modules of a user device.

[0081] Example 20: The system of any of examples 14 through 19, wherein the storage device further stores instructions executable by the at least one processor to: generate, based on the text prompt and the one or more contextual signals, audio data for a new audio clip in periodic intervals.

[0082] Example 21 : The system of any of examples 14 through 20, wherein the storage device further stores instructions executable by the at least one processor to: determine, based on the one or more contextual signals, an event; and generate audio data for a new audio clip based on the event.

[0083] Example 22: The system of any of examples 14 through 21, wherein the storage device further stores instructions executable by the at least one processor to: provide sample audio data to the one or more generative models, wherein the sample audio data includes at least one of a plurality of music samples, a plurality of alarm tones, and user feedback; and tune the one or more generative models based on the sample audio data.

[0084] Example 23: The system of any of examples 14 through 22, wherein the storage device further stores instructions executable by the at least one processor to: identify an indication to change attributes of the audio clip over time; and instruct a user device to output audio data for the audio clip based on the indication.

[0085] Example 24: The system of any of examples 14 through 23, wherein to generate audio data for the audio clip, the storage device further stores instructions executable by the at least one processor to: determine, based at least on the text prompt, an indication to change attributes of the audio clip; and generate audio data for the audio clip based on the indication.

[0086] Example 25: The system of any of examples 14 through 24, wherein the storage device further stores instructions executable by the at least one processor to: generate audio data for a plurality of audio clips based on the text prompt and the one or more contextual signals; determine a set of candidate audio clips based on the plurality of audio clips; output audio data for the set of candidate audio clips to a user device; obtain, from the user device, user feedback for the set of candidate audio clips; and tune the one or more generative models based on the user feedback.

[0087] Example 26: The system of any of examples 14 through 25, wherein the storage device further stores instructions executable by the at least one processor to: randomly generate the text prompt.

[0088] Example 27: Computer-readable storage medium encoded with instructions that, when executed, cause at least one processor of a computing system to: obtain a text prompt; obtain one or more contextual signals; generate, by one or more generative models, audio data for an audio clip based on the text prompt and the one or more contextual signals; and output audio data for the audio clip.

[0089] Example 28: The computer-readable storage medium of example 27, wherein the one or more contextual signals includes data representing preferences for the audio clip associated with at least one of an event based on data associated with a user device and one or more environmental factors.

[0090] Example 29: The computer-readable storage medium of any of examples 27 and 28, wherein to generate audio data for the audio clip, the instructions cause the at least one processor of the computing system to: determine, by the one or more generative models, a mood based on the text prompt; and generate audio data for the audio clip based at least on the mood.

[0091] Example 30: The computer-readable storage medium of any of examples 27 through 29, wherein to generate audio data for the audio clip, the instructions cause the at least one processor of the computing system to: generate a modified text prompt based on the text prompt and the one or more contextual signals; generate a plurality of prompt tokens based on the modified text prompt; and predict a plurality of audio tokens based on the plurality of prompt tokens.

[0092] Example 31 : The computer-readable storage medium of any of examples 27 through 30, wherein to generate audio data for the audio clip, the instructions cause the at least one processor of the computing system to: generate a modified text prompt based on the text prompt and the one or more contextual signals; and noise and denoise an audio signal based at least on the modified text prompt.

[0093] Example 32: The computer-readable storage medium of any of examples 27 through 31, wherein to obtain the one or more contextual signals, the instructions cause the at least one processor of the computing system to: request additional information from one or more application modules of a user device.

[0094] Example 33: The computer-readable storage medium of any of examples 27 through 32, wherein the instructions further cause the at least one processor of the computing system to: generate, based on the text prompt and the one or more contextual signals, audio data for a new audio clip in periodic intervals.

[0095] Example 34: The computer-readable storage medium of any of examples 27 through 33, wherein the instructions further cause the at least one processor of the computing system to: determine, based on the one or more contextual signals, an event; and generate audio data for a new audio clip based on the event.

[0096] Example 35: The computer-readable storage medium of any of examples 27 through 34, wherein the instructions further cause the at least one processor of the computing system to: provide sample audio data to the one or more generative models, wherein the sample audio data includes at least one of a plurality of music samples, a plurality of alarm tones, and user feedback; and tune the one or more generative models based on the sample audio data.

[0097] Example 36: The computer-readable storage medium of any of examples 27 through 35, wherein the instructions further cause the at least one processor of the computing system to: identify an indication to change attributes of the audio clip over time; and instruct a user device to output audio data for the audio clip based on the indication.

[0098] Example 37: The computer-readable storage medium of any of examples 27 through 36, wherein to generate audio data for the audio clip, the instructions further cause the at least one processor of the computing system to: determine, based at least on the text prompt, an indication to change attributes of the audio clip; and generate audio data for the audio clip based on the indication.

[0099] Example 38: The computer-readable storage medium of any of examples 27 through 37, wherein the instructions further cause the at least one processor of the computing system to: generate audio data for a plurality of audio clips based on the text prompt and the one or more contextual signals; determine a set of candidate audio clips based on the plurality of audio clips; output audio data for the set of candidate audio clips to a user device; obtain, from the user device, user feedback for the set of candidate audio clips; and tune the one or more generative models based on the user feedback.

[0100] Example 39: The computer-readable storage medium of any of examples 27 through 38, wherein the instructions further cause the at least one processor of the computing system to: randomly generate the text prompt.

[0101] Example 40: A device includes at least one processor; and a storage device that stores instructions executable by the at least one processor to: obtain a text prompt; obtain one or more contextual signals; generate, by one or more generative models, audio datafor an audio clip based on the text prompt and the one or more contextual signals; and output audio data for the audio clip.

[0102] Example 41 : The device of example 40, wherein the one or more contextual signals includes data representing preferences for the audio clip associated with at least one of an event based on data associated with a user device and one or more environmental factors.

[0103] Example 42: The device of any of examples 40 through 41, wherein to generate audio data for the audio clip, the storage device stores instructions executable by the at least one processor to: determine, by the one or more generative models, a mood based on the text prompt; and generate audio data for the audio clip based at least on the mood.

[0104] Example 43: The device of any of examples 40 through 42, wherein to generate audio data for the audio clip, the storage device stores instructions executable by the at least one processor to: generate a modified text prompt based on the text prompt and the one or more contextual signals; generate a plurality of prompt tokens based on the modified text prompt; and predict a plurality of audio tokens based on the plurality of prompt tokens.

[0105] Example 44: The device of any of examples 40 through 43, wherein to generate audio data for the audio clip, the storage device stores instructions executable by the at least one processor to: generate a modified text prompt based on the text prompt and the one or more contextual signals; and noise and denoise an audio signal based on the modified text prompt.

[0106] Example 45: The device of any of examples 40 through 44, wherein to obtain the one or more contextual signals, the storage device stores instructions executable by the at least one processor to: request additional information from one or more application modules of a user device.

[0107] Example 46: The device of any of examples 40 through 45, wherein the storage device further stores instructions executable by the at least one processor to: generate, based on the text prompt and the one or more contextual signals, audio data for a new audio clip in periodic intervals.

[0108] Example 47: The device of any of examples 40 through 46, wherein the storage device further stores instructions executable by the at least one processor to: determine, based on the one or more contextual signals, an event; and generate audio data for a new audio clip based on the event.

[0109] Example 48: The device of any of examples 40 through 47, wherein the storage device further stores instructions executable by the at least one processor to: provide sample audio data to the one or more generative models, wherein the sample audio data includes at least one of a plurality of music samples, a plurality of alarm tones, and user feedback; and tune the one or more generative models based on the sample audio data.

[0110] Example 49: The device of any of examples 40 through 48, wherein the storage device further stores instructions executable by the at least one processor to: identify an indication to change attributes of the audio clip over time; and instruct a user device to output audio data for the audio clip based on the indication.

[0111] Example 50: The device of any of examples 40 through 49, wherein to generate audio data for the audio clip, the storage device further stores instructions executable by the at least one processor to: determine, based at least on the text prompt, an indication to change attributes of the audio clip; and generate audio data for the audio clip based on the indication.

[0112] Example 51 : The device of any of examples 40 through 50, wherein the storage device further stores instructions executable by the at least one processor to: generate audio data for a plurality of audio clips based on the text prompt and the one or more contextual signals; determine a set of candidate audio clips based on the plurality of audio clips; output audio data for the set of candidate audio clips to a user device; obtain, from the user device, user feedback for the set of candidate audio clips; and tune the one or more generative models based on the user feedback.

[0113] Example 52: The device of any of examples 40 through 51, wherein the storage device further stores instructions executable by the at least one processor to: randomly generate the text prompt.

[0114] Example 53: A computing system comprising means for performing any combination of examples 1-52.

[0115] Example 54: A computing device comprising means for performing any combination of examples 1-52.

[0116] Example 55: Anon-transitory computer-readable storage medium encoded with instructions that, when executed by one or more processors, cause the one or more processors to perform any combination of examples 1-52.

[0117] By way of example, and not limitation, such computer-readable storage media can comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other storage mediumthat can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. It should be understood, however, that computer-readable storage mediums and media and data storage media do not include connections, carrier waves, signals, or other transient media, but are instead directed to non-transient, tangible storage media. Disk and disc, as used herein, includes compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and Blu-ray disc, where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable medium.

[0118] The techniques described in this disclosure may be implemented, at least in part, in hardware, software, firmware, or any combination thereof. For example, various aspects of the described techniques may be implemented within one or more processors, including one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or any other equivalent integrated or discrete logic circuitry, as well as any combinations of such components. The term “processor” or “processing circuitry” may generally refer to any of the foregoing logic circuitry, alone or in combination with other logic circuitry, or any other equivalent circuitry. A control unit including hardware may also perform one or more of the techniques of this disclosure.

[0119] Such hardware, software, and firmware may be implemented within the same device or within separate devices to support the various techniques described in this disclosure. In addition, any of the described units, modules or components may be implemented together or separately as discrete but interoperable logic devices. Depiction of different features as modules or units is intended to highlight different functional aspects and does not necessarily imply that such modules or units must be realized by separate hardware, firmware, or software components. Rather, functionality associated with one or more modules or units may be performed by separate hardware, firmware, or software components, or integrated within common or separate hardware, firmware, or software components.

[0120] Various examples of the invention have been described. These and other examples are within the scope of the following claims.

Claims

WHAT IS CLAIMED IS:

1. A method comprising: obtaining, by one or more processors, a text prompt; obtaining, by the one or more processors, one or more contextual signals; generating, by one or more generative models executing on the one or more processors, audio data for an audio clip based on the text prompt and the one or more contextual signals; and outputting, by the one or more processors, audio data for the audio clip.

2. The method of claim 1, wherein the one or more contextual signals includes data representing preferences for the audio clip associated with at least one of an event based on data associated with a user device and one or more environmental factors.

3. The method of claim 1, wherein generating audio data for the audio clip comprises: determining, by the one or more generative models, a mood based on the text prompt; and generating audio data for the audio clip based at least on the mood.

4. The method of claim 1, wherein generating audio data for the audio clip comprises: generating a modified text prompt based on the text prompt and the one or more contextual signals; generating a plurality of prompt tokens based on the modified text prompt; and predicting a plurality of audio tokens based on the plurality of prompt tokens.

5. The method of claim 1, wherein generating audio data for the audio clip comprises: generating a modified text prompt based on the text prompt and the one or more contextual signals; and noising and denoising an audio signal based at least on the modified text prompt.

6. The method of claim 1, wherein obtaining the one or more contextual signals comprises: requesting additional information from one or more application modules of a user device.

7. The method of claim 1, further comprising: generating, based on the text prompt and the one or more contextual signals, audio data for a new audio clip in periodic intervals.

8. The method of claim 1, further comprising: determining, based on the one or more contextual signals, an event; and generating audio data for a new audio clip based on the event.

9. The method of claim 1, further comprising: providing sample audio data to the one or more generative models, wherein the sample audio data includes at least one of a plurality of music samples, a plurality of alarm tones, and user feedback; and tuning the one or more generative models based on the sample audio data.

10. The method of claim 1, further comprising: identifying an indication to change attributes of the audio clip over time; and instructing a user device to output audio data for the audio clip based on the indication.

11. The method of claim 1, wherein generating audio data for the audio clip comprises: determining, based at least on the text prompt, an indication to change attributes of the audio clip; and generating audio data for the audio clip based on the indication.

12. The method of claim 1, further comprising: generating audio data for a plurality of audio clips based on the text prompt and the one or more contextual signals; determining a set of candidate audio clips based on the plurality of audio clips; outputting audio data for the set of candidate audio clips to a user device;obtaining, from the user device, user feedback for the set of candidate audio clips; and tuning the one or more generative models based on the user feedback.

13. The method of claim 1, further comprising: randomly generating the text prompt.

14. A system comprising means for performing any of the method of claims 1-13.

15. A computer program product comprising at least one non-transitory computer- readable media including one or more instructions that, when executed by at least one processor, cause the at least one processor to perform any of the method of claims 1-13.

Citation Information

Patent Citations

  • Method and device for generating multimedia resources based on large model and storage medium

    CN117789680A

  • Generating multi-modal response(s) through utilization of large language model(s)

    US11907674B1

  • Context-sensitive generation of conversational responses

    WO2016195912A1

  • US202463634647P