Generating new musical expression based on musical analysis and user input

The computing system addresses the challenge of creating engaging audiovisual content by using machine learning to analyze and transform user input, enabling users to generate high-quality content with minimal effort, suitable for social media platforms.

WO2026024623A1PCT designated stage Publication Date: 2026-01-29GHOST NOTES INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/038484
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-23
Filing Date
2025-07-21
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Traditional methods of creating engaging audiovisual content require significant time, expertise, and specialized software, posing barriers for average users.

Method used

A computing system that applies machine learning models to analyze musical and visual data, along with user input, to generate new musical and visual expressions dynamically, allowing users to create high-quality content with minimal skill.

Benefits of technology

Enables users to effortlessly produce captivating audiovisual content that is stylistically faithful to the input, without requiring advanced knowledge or tools, facilitating easy creation and sharing on social media platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025038484_29012026_PF_FP_ABST
    Figure US2025038484_29012026_PF_FP_ABST
Patent Text Reader

Abstract

Techniques of this disclosure may include a computing system that receives audio data having a first set of musical style characteristics and applies a machine learning model to the audio data to generate structured data including one or more data values. Each data value may correspond to a respective musical style characteristic from the first set of musical style characteristics. The computing system may receive input, and generate, based on the structured data and the input, one or more of visual data having a set of visual style characteristics associated with the first set of musical style characteristics and audio data having a second set of musical style characteristics associated with the first set of musical style characteristics. The computing system may output one or more of the visual data having the set of visual style characteristics and the audio data having the second set of musical style characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

GENERATING NEW MUSICAL EXPRESSION BASED ON MUSICAL ANALYSIS AND USER INPUTCROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 674,752, filed July 23, 2024, which is incorporated by reference herein in its entirety.BACKGROUND

[0002] The proliferation of social media platforms has driven an increasing demand for engaging audiovisual content. Users often seek ways to create and share visually and audibly appealing media that captures attention and enhances online interactions. Traditional methods of content creation, however, often require significant time, expertise, and specialized software, which may pose barriers to the average user.SUMMARY

[0003] In general, this disclosure is directed to techniques for dynamically generating new musical and / or visual expression based on analysis of musical data, visual data, and / or user input. An example computing system may receive musical data (e.g., from an input device), such as audio data considered to have a “musical style”, or a first set of musical style characteristics, e.g., sets of patterns and structures that are characteristic of a particular composer, period, genre, etc. For example, the audio data may include particular note pitches, melodies, chord progressions, key signatures, harmonies, rhythms, textures, forms, etc. that together create a recognizable and distinct musical style or identity. In some examples, the audio data may be in the form of pre-existing music from a library, a recording of live music, a voice recording, and the like. The computing system may apply a machine learning model to the audio data to generate structured data (e.g., tone row data) including data values corresponding to a respective musical style characteristic from the first set of musical style characteristics. In some examples, the computing system may additionally or alternatively receive input visual data (e.g., images, videos, etc.), input indicative of a tactile event, input indicative of motion (e.g., captured gestures), user generated content, and / or biometric data. For example, the computing system may receive input from an application executing at a computing device when a user interacts with agraphical user interface (GUI) of an application and / or an input device of the computing device (e.g., a camera) detects motion. In some examples, the computing system may map received input to at least one data value included in the structured data to generate audio data having a second set of musical style characteristics associated with the first set of musical style characteristics. That is, the generated audio data may be considered “new” or “altered” music that is “stylistically faithful” to the audio data initially received, i.e., the computing system may transform input audio data based on user input, in which the transformed audio data is different from the input audio data, but includes musical style characteristics that have a threshold similarity to the musical style characteristics from which they were derived. In some examples, the computing system may additionally or alternatively generate instructions for generating visual content, such as GUIs or other visual data associated with the transformed audio data and any / or any input data. For example, a particular gesture from a user may result in the computing system generating one or more visuals or visual elements, such as a GUI, images, visual effects, colors, etc. In some examples, the computing system may output at least a portion of the audio data having the second set of musical style characteristics with the instructions. As such, the computing system may generate new musical expression (e.g., in the form of new audiovisual content) from input audio data, input visual data, and / or user input.

[0004] In one example, the disclosure is directed toward a method that includes receiving, by a computing system, first audio data having a first set of musical style characteristics, and applying, by the computing system, a machine learning model to the first audio data to generate structured data including one or more data values, wherein each of the one or more data values correspond to a respective musical style characteristic from the first set of musical style characteristics. The method further includes receiving, by the computing system, at least one input, and generating, by the computing system, based on the structured data and the at least one input, one or more of visual data having a set of visual style characteristics associated with the first set of musical style characteristics and second audio data having a second set of musical style characteristics associated with the first set of musical style characteristics. The method further includes outputting, by the computing system, one or more of at least a portion of the visual data having the set of visual style characteristics and at least a portion of the second audio data having the second set of musical style characteristics.

[0005] In another example, the disclosure is directed toward a computing system that includes one or more processors, and one or more storage devices that store instructions. The instructions, when executed by the one or more processors, cause the one or more processors to receive first audio data having a first set of musical style characteristics, and apply a machine learning model to the first audio data to generate structured data including one or more data values, wherein each of the one or more data values correspond to a respective musical style characteristic from the first set of musical style characteristics. The instructions further cause the one or more processors to receive at least one input, and generate, based on the structured data and the at least one input, one or more of visual data having a set of visual style characteristics associated with the first set of musical style characteristics and second audio data having a second set of musical style characteristics associated with the first set of musical style characteristics. The instructions further cause the one or more processors to output one or more of at least a portion of the visual data having the set of visual style characteristics and at least a portion of the second audio data having the second set of musical style characteristics.

[0006] In another example, the disclosure is directed toward a non-transitory computer-readable storage medium encoded with instructions that, when executed by one or more processors, cause one or more processors to receive first audio data having a first set of musical style characteristics, and apply a machine learning model to the first audio data to generate structured data including one or more data values, wherein each of the one or more data values correspond to a respective musical style characteristic from the first set of musical style characteristics. The instructions further cause the one or more processors to receive at least one input, and generate, based on the structured data and the at least one input, one or more of visual data having a set of visual style characteristics associated with the first set of musical style characteristics and second audio data having a second set of musical style characteristics associated with the first set of musical style characteristics. The instructions further cause the one or more processors to output one or more of at least a portion of the visual data having the set of visual style characteristics and at least a portion of the second audio data having the second set of musical style characteristics.

[0007] The details of one or more examples of the disclosure are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the disclosure will be apparent from the description and drawings, and from the claims.BRIEF DESCRIPTION OF THE FIGURES

[0008] FIG. 1 A is a conceptual diagram illustrating an example computing system for dynamically generating new musical expression based on musical analysis and user input, in accordance with one or more techniques of this disclosure.

[0009] FIG. IB is a conceptual diagram illustrating another example of the computing system of FIG. 1 A configured to dynamically generate new musical expression based on musical analysis and detected motion, in accordance with one or more techniques of this disclosure.

[0010] FIG. 1C is a conceptual diagram illustrating another example of the computing system of FIG. 1 A configured to dynamically generate new musical expression based on musical analysis and visual input, in accordance with one or more techniques of this disclosure.

[0011] FIG. ID is a conceptual diagram illustrating an example of the computing system of FIG. 1 A configured to dynamically generate recommendations based on analysis of user generated content, in accordance with one or more techniques of this disclosure.

[0012] FIG. IE is a conceptual diagram illustrating an example of the computing system of FIG. 1 A configured to dynamically generate recommendations based on biometric data analysis, in accordance with one or more techniques of this disclosure.

[0013] FIG. 2 is a block diagram illustrating another example computing system for dynamically generating new musical expression based on analysis of audio, visual, or user input, in accordance with one or more techniques of this disclosure.

[0014] FIG. 3 is a conceptual diagram illustrating an audiovisual content generation module configured to transform structured data based on analysis of audio, visual, or user input, in accordance with one or more techniques of this disclosure.

[0015] FIG. 4 is a conceptual diagram illustrating a gesture analysis module configured to map gestures to dynamically generate new musical expression, in accordance with one or more techniques of this disclosure.

[0016] FIG. 5 is a conceptual diagram illustrating the computing system of FIG. IB configured to map gestures to a grid to dynamically generate new musical expression, in accordance with one or more techniques of this disclosure.

[0017] FIG. 6 is a conceptual diagram illustrating the computing system of FIG. 1C configured to dynamically generate new musical expression based on musical analysis and visual input, in accordance with one or more techniques of this disclosure.

[0018] FIG. 7 is a conceptual diagram illustrating a biometric data analysis module configured to perform analysis of biometric data for determining recommendations, in accordance with one or more techniques of this disclosure.

[0019] FIG. 8 is a conceptual diagram illustrating the computing system of FIG. ID configured to generate recommendations based on analysis of user generated content, in accordance with one or more techniques of this disclosure.

[0020] FIG. 9 is a flowchart illustrating an example operation of a computing system for dynamically generating new musical expression based on musical analysis and user input, in accordance with one or more techniques of this disclosure.DETAILED DESCRIPTION

[0021] FIG. 1 A is a conceptual diagram illustrating an example computing system for dynamically generating new musical expression based on analysis of input audio data, input visual data, and / or user input, in accordance with one or more techniques of this disclosure. In the example of FIG. 1, a user 107 interacts with computing device 112 that is in communication with computing system 100 via network 101. In some examples, some or all of the components and / or functionality attributed to computing system 100 may be implemented and / or performed by computing device 112, and vice versa. In some examples, analysis module 106 and / or other modules described herein may be employed as a cloud-based service (e.g., by computing system 100). In other examples, analysis module 106 and / or other modules described herein may be employed “on-device,” e.g., as an on-premise or a local service (e.g., by a user computing device, such as computing device 112).

[0022] Computing system 100 may be implemented on a plurality of computing devices, such as computing device 112, or may be a computing device. Example computing devices include, but are not limited to, portable devices, mobile devices, such as mobile phones (including smartphones), laptop computers, desktop computers, tablet computers, smart television platforms, server computers, mainframes, etc. In some examples, computing system 100 mayrepresent a cloud computing system that provides one or more services via network 101. That is, in some examples, computing system 100 may be a distributed computing system.

[0023] Computing system 100 may be in communication with computing device 112 via network 101. That is, computing system 100 may receive input from and send output to computing device 112 via network 101. Network 101 may include any public or private communication network, such as a cellular network, Wi-Fi network, a packet switched network such as the Internet, or other type of network for transmitting data between computing system 100 and computing device 112. Computing device 112 may send and receive data to and from computing system 100 across network 101 using any suitable communication techniques. For example, computing system 100 and computing device 112 may each be operatively coupled to network 101 using respective network links. Network 101 may include network hubs, network switches, network routers, etc., that are operatively inter-coupled, thereby providing for the exchange of information between computing device 112 and computing system 100. In some examples, network links of network 101 may be Ethernet or other network connections. Such connections may include wireless and / or wired connections.

[0024] As shown in the example of FIG. 1, computing device 112 includes one or more user interface (UI) components (“UI components 102”), which may be configured to function as input devices and / or output devices for computing device 112. UI components 102 may be implemented using various technologies. For instance, UI components 102 may be configured to receive input from user 120 through tactile, audio, and / or video feedback, and / or traditional input peripherals (e.g., via peripheral devices, user interfaces, or the like). Examples of input devices include a presence-sensitive display, a presence-sensitive or touch-sensitive input device (such as that shown in FIG. 1), a mouse, a keyboard, a voice responsive system, video camera, microphone or any other type of device for detecting input from user 107. In some examples, a presence-sensitive display includes a touch-sensitive or presence-sensitive input screen, such as a resistive touchscreen, a surface acoustic wave touchscreen, a capacitive touchscreen, a projective capacitive touchscreen, a pressure sensitive screen, an acoustic pulse recognition touch screen, or another presence-sensitive technology. That is, UI components 102 of computing device 112 may include a presence-sensitive device that may receive tactile input 105 from user 107. UI components 102 may receive indications of tactile input 105 by detecting one or more gesturesfrom user 107 (e.g., when user 107 touches or points to one or more locations of UI components 102 with a finger or a stylus pen).

[0025] One of or more of UI components 102 may additionally or alternatively be configured to function as an output device by providing output to user 107 using tactile, audio, or video stimuli (e.g., providing output in the form of physical sensations, sounds, tones, spoken words, music, other auditory cues, visual information such as images, graphics, animations, videos, other visual cues, etc.). Examples of output devices include a sound card, a video graphics adapter card, or any of one or more display devices, such as a liquid crystal display (LCD), dot matrix display, light emitting diode (LED) display, microLED, miniLED, organic light-emitting diode (OLED) display, e-ink, or similar monochrome or color display capable of outputting visible information to user 107. Additional examples of an output device include a speaker, a haptic device, or other device that can generate intelligible output to a user. For instance, UI components 102 may present output to user 107 as a graphical user interface that may be associated with functionality provided by computing device 112. In this way, UI components 102 may present various user interfaces of applications executing at or accessible by computing device 112 (e.g., an electronic message application, an Internet browser application, etc.). User 107 may interact with a respective user interface of an application to cause computing device 112 to perform operations relating to a function provided by the application. For example, as shown in the example of FIG. 1, user 107 may interact with generated user interface (GUI) 103 of an application to cause computing device to output data such as audio data 148, which, may be generated by computing system 100 responsive to receiving input indicative of tactile event 105. In some examples, computing device 112 may be configured to store data output by computing system 100, such as audio data 148 and / or any other types of data output by computing system 100.

[0026] In some examples, UI components 102 of computing device 112 may detect two- dimensional and / or three-dimensional gestures as input from user 107. For example, UI components 102 may include one or more sensors configured to detect movement (e.g., movement of a hand, an arm, a pen, a stylus, etc.) within a threshold distance of the sensor of UI components 102. In some examples, UI components 102 may determine a two- or three- dimensional vector representation of the movement and correlate the vector representation to a gesture input (e.g., a hand-wave, a pinch, a clap, a pen stroke, etc.) that has multiple dimensions. As such, in some examples, UI components 102 may detect a multidimensional gesture withoutrequiring the user to gesture at, or near, a screen or surface at which UI components 102 output information for display. Instead, UI components 102 may detect a multi-dimensional gesture performed at or near a sensor which may or may not be located near the screen or surface at which UI components 102 output information for display.

[0027] In the example of FIG. 1, computing system 100 includes user interface (UI) module 104. UI module 104 may perform operations described herein using hardware, software, firmware, or a mixture thereof residing in and / or executing at computing system 100. Computing system 100 may execute UI module 104 with one processor or with multiple processors. In some examples, computing system 100 may execute UI module 104 as a virtual machine executing on underlying hardware. UI module 104 may execute as one or more services of an operating system or computing platform or may execute as one or more executable programs at an application layer of a computing platform.

[0028] UI module 104, as shown in the example of FIG. 1, may be operable by computing system 100 to perform one or more functions, such as receive input and send indications of such input to other components associated with computing system 100. UI module 104 may also receive data from components associated with computing system 100. Using the data received, UI module 104 may cause other components associated with computing system 100, such as one or more UI components, to provide output based on the data. In some examples, computing system 100 may be a computing device, such as a user computing device, configured to implement some or all of the techniques described herein “on-device.” For example, in some examples, UI module 104 may cause one or more UI components of computing system 100 to display a GUI (e.g., GUI 103), one or more GUI elements, other visual data, and / or play aloud audio data 148. In some other examples, UI module 104 may send data to one or more UI components 102 of computing device 112 to display a GUI (e.g., GUI 103), one or more GUI elements, other visual data, and / or play aloud audio data 148. As such, the functionality described herein with respect to any modules and components included in computing system 100 may be performed and / or implemented by computing device 112, and the functionality described herein with respect to any modules and components included in computing device 112 may be performed and / or implemented by computing system 100.

[0029] GUI 103 may be an example GUI of an application executing at computing device 112 and / or computing system 100, such as a GUI of a social media application. In general, thetechniques described herein may provide a user of an application various functionalities for creating audio and / or visual (collectively referred to herein as “audiovisual”) content. For example, a user may interact with GUI 103 to produce new audiovisual content and / or alter existing audiovisual content. In some examples, GUI 103 may be an interface or screen for capturing and / or recording visual data, such as videos and images. In some examples, GUI 103 may include one or more user interface elements that a user may interact with to produce various output. More specifically, in some examples, computing system 100 may receive an indication of user input from computing device 112 in response to a gesture detected at a location of a presence-sensitive display of computing device 112 that corresponds to a GUI component associated with an application executing at computing device 112. In some examples, the gesture may be provided by a user tapping on a screen (e.g., tactile event 105), and / or by a user moving at or near a sensor or other input device. In some examples, the indication of a gesture may be an audible input. In some examples, the indication of the gesture may be provided by user 107 by using gesture control, such as by providing the gestures described above (e.g., a hand-wave, a pinch, a clap, a pen stroke, etc.) or by tapping the screen in a certain manner. As such, the techniques described herein may be executed by computing system 100 in response to an indication of a variety of gestures. In some examples, the one or more gestures may be determined and analyzed by analysis module 106 to generate a variety of output data.

[0030] As an example, GUI 103 may include a virtual piano that user 107 can interact with by providing input indicative of one or more tactile events 105, in which computing system 100 may output audio data 148 including notes pitches, note sequences, etc. that are based on one or more tactile events 105 (e.g., each finger tap from user 107 may correspond to a key on the virtual piano). In some examples, GUI 103 may be a “feed,” or a GUI displaying various types of shared content (e.g., user generated content) hosted on an application. As such, in general, GUI 103 may be one GUI from a plurality of GUIs included in application for creating audiovisual content. GUI 103 may include various user interface elements associated with various types of output generated by computing system 100, e.g., generated audio data, generated visual data, recommended user generated content, recommended audio data, recommended visual data, input data overlaid with output data, and the like.

[0031] In general, user 107 may be provided with an opportunity to provide input to control whether programs or features of computing device 112 and / or computing system 100 can collectand make use of user information (e.g., user 107’s personal data, biometric data, etc.), or to dictate whether and / or how computing device 112 and / or computing system 100 may receive content that may be relevant to user 107. Other user information may include data that includes the context of user usage (e.g., breadth of share (sharing publicly, or with a large group, or privately, or a specific person), context of share, etc.), either obtained from an application itself or from other sources. Other user information may include data for the user’s device, such as the applications running on the device, location data, etc. Furthermore, certain data may be treated in one or more ways before it is stored or used by computing device 112 and / or computing system 100 so that personally identifiable information (e.g., sensitive data, particular location data, etc.) is removed. Thus, user 107 may have control over how information is collected about them and used by computing device 112 and / or computing system 100. For example, user 120 may be prompted by computing device 112 to provide explicit consent for computing device 112 and / or computing system 100 to retrieve and / or store any or all of user 107’s data. In this way, some or all of the techniques described herein may only be implemented with explicit consent from user 107.

[0032] In general, computing system 100 may receive musical data (which may be referred to herein as “audio,” an “audio signal,” and / or “audio data”) considered to have a “musical style.” As used herein, the term “musical style” may be defined as identifiable patterns in musical data that may be indicative of a composer, period, genre, or culture. For example, the audio data may include particular note pitches, melodies, chord progressions, key signatures, harmonies, rhythms, textures, forms, etc. that together create a recognizable and distinct musical style or identity. These patterns may exist at various levels of the musical structure. For example, surface-level features and / or patterns may include specific melodic or rhythmic motifs. Mid-level features and / or patterns may include, for example, harmonic progressions, phrase structures, and formal elements. Deep-level abstractions and / or patterns may include, for example, overall tonal plans, large-scale formal structures, and conceptual approaches to composition. Furthermore, while musical style may be based on the presence of certain musical elements, musical style may also be based on the frequency, combination, and context of certain musical elements within a musical piece or body of work. As such, in some examples, “musical style” may be based on statistical phenomena, in which a particular musical style may be defined by a consistent recurrence of certain patterns and their relationships. In some examples, the “audio,” “audiosignal,” and / or “audio data” described herein may be considered “signatures” that include distinctive patterns or combinations of musical elements which are characteristic of a particular composer or style. Thus, in general, “audio,” “audio signal,” and / or “audio data,” as signatures, may be used to analyze and / or alter existing music and / or generate new music in the same particular style. For example, in some examples, computing system 100 may employ analysis module 106 to analyze and process input data, such as input audio data. In some examples, analysis module 106 may apply one or more machine learning models to the input audio data to identify patterns and / or combinations of musical elements present in the input audio data. As used herein, “artificial intelligence” (Al) encompasses machine learning techniques, including supervised, unsupervised, and reinforcement learning, as well as other computational inference techniques. Unless otherwise indicated, references to “machine learning” may be interchangeable with references to “Al.” In general, the one or more machine learning models may perform a variety of tasks, such as classification, regression, clustering, etc. In some examples, the one or more machine learning models may determine, based on identified patterns and / or combinations of musical elements present in the input audio data, a particular musical style for the input audio data, e.g., a composer, period, genre, or culture associated with the input audio data. In some examples described herein, similarity between different audios, audio signals, input and output audio data, etc. may be based on whether the one or more machine learning models determine the different audios, audio signals, input and output audio data, etc. to be associated with the same musical style, and / or each include musical style characteristics that have a threshold similarity.

[0033] That is, in some examples, computing system 100 may intelligently generate input vectors (e.g., vectors generated by Al). That is, analysis module 106 may ingest and transform raw input audio data into a vector representation, and may apply one or more machine learning models to that vector, thereby automatically detecting patterns and / or combinations of musical elements (e.g., timbre, harmony, meter, genre-specific motifs, etc.). The models may perform classification, regression, clustering, and / or other inference tasks, and may generate output vector(s) that identifies, for instance, a composer, historical period, genre, or cultural style exhibited by the audio. In some examples, two audio signals may be deemed similar when the model(s) map them to the same or sufficiently close output vectors that encode musical style characteristics above a similarity threshold.

[0034] As used throughout this disclosure, “original,” “initial,” or “input” audio data may refer to audio data received and / or used by computing system 100 to generate new musical expression, output such as audio data 148 or other types of output data described herein (e.g., visual data output). In some examples, the input audio data may include audio data that is captured in real time or near real time. In general, any input data described herein may be processed by computing system 100 and / or computing device 112, in real time and / or asynchronously. In some examples, the input audio data described herein may include Over the Air (OTA) real-time capture of non-copyrightable chord progressions, tone rows, and / or beat map data. In some examples, the input audio data and / or any other type of input data described herein may be captured and / or processed in real-time by computing system 100 and / or computing device 112. In some examples, data pre-processed and / or pre-computed by computing system 100 may be sent to computing device 112, such that computing device 112 may process, analyze, transform, etc. input data that is captured and / or detected by one or more of UI components 102 (e.g., sensors) in real-time. In some examples, for example, computing system 100 may, based on input data, build or refine a model for generating output, in which computing system 100 may send data indicative of the model to computing device 112, such that computing device 112 may apply the model to captured and / or detected data to generate similar output in real-time or near real-time. In some examples, the input audio data may include audio data that is stored in the public domain, on a user computing device, in computing system 100, in a collaborative database, in another database, etc. In some examples, the input audio data may include preexisting music from a library, a recording of live music, a voice recording, an instrument recording, etc.

[0035] In some examples, sampling of input audio data may be performed in a manner that is time-limited, e.g., the input audio data may only include a “snippet” of a longer audio file or musical performance, such that the data analyzed by computing system 100 includes data forms that are “most interesting” to a user. That is, in the example of the input audio data including audio recorded or captured by a user attending a live concert, the user may only record or capture a sample of signal data representing rhythmic or harmonic patterns, beats, instruments, or other musical style characteristics that a user considers to be preferential, most valuable, most similar, and most aligned, etc. to the user’s music taste. As such, computing system 100 may better identify user preferences for audio and / or visual content.

[0036] In general, computing system 100 may be a computing system configured to generate audiovisual content by transforming input audio data and / or input visual data. In some examples, the input audio data and / or visual data may be transformed based on user input. For examples, computing system 100 may receive (e.g., from one or more user interface (UI) components 102 of computing device 112) musical data such as audio data having a first set of musical style characteristics, e.g., particular note pitches, melodies, chord progressions, key signatures, etc. In some examples, the input audio data may be pre-existing music from a library, a recording of live music, a voice recording, etc. Computing system 100 may apply analysis module 106 of musical expression generation module 108 to the input audio data to generate structured data (e.g., tone row data) including data values corresponding to a respective musical style characteristic from the first set of musical style characteristics. In general, computing system 100 may covert structured or unstructured input data into a feature vector representation suitable for machine-learning inference. As used herein, “structured data” may include feature vectors computed by analysis module 106 before inference, and / or intermediate embeddings that one or more machine learning models derive from unstructured data, such as raw waveforms and / or image frames. In some examples, the one or more data values are indicative of one or more of a frequency spectrum, an amplitude, a timbre, a note pitch, a note interval, a beat, a fill, a loop, an ensemble, a duration, a riff, a melody contour, a melody motif, a chord progression, a key signature, a musical scale mode, a tonality, a tempo, a meter, a rhythmic pattern, an instrumentation, a density, a polyphony, a volume, a note articulation, a note expression, a structure, a section, a genre, historical context data, attack / decay / sustain / release biases, etc. As such, in some examples, analysis module 106 may determine a range of data values for each of the aforementioned musical style characteristics, i.e., analysis module 106 may determine a range in which each of the aforementioned musical style characteristics can vary.

[0037] As shown in the example of FIG. 1 A, UI components 102 of computing device 112 may generate GUI 103 associated with an application executing on computing device 112. As further shown in FIG. 1 A, user 107 may provide input indicative of tactile event 105. That is, computing system 100 may receive at least one input indicative of at least one tactile event 105, in which computing system 100 receives the at least one input in response to the at least one tactile event 105 being detected at a location of a presence-sensitive display that corresponds to at least oneportion of GUI 103 associated with the application. For example, user 107 may provide input indicative of tactile event 105 by tapping a location of the presence-sensitive display.

[0038] In general, audiovisual content generation module 110 of musical expression generation module 108 may be configured to generate new musical expression (e.g., music, visual effects associated with music, etc.) based on input audio data, input visual data, user actions and motions, cultural considerations, style considerations, etc. In general, audiovisual content generation module 110 may dynamically generate various output data responsive to computing system 200 receiving various input data. In some examples, audiovisual content generation module 110 may generate multiple outputs responsive to a single input, e.g., responsive to a single tactile event 105 or a single motion, audiovisual content generation module 110 may generate a sequence of pitches or a string of musical events, including polyphonies, melody, harmony, chord progressions, beats, fills, riffs, and other compound forms.

[0039] In some examples, the output data may be consistent with the input data, e.g., initial or original audio data may be used to generate new or altered audio data that has musical style characteristics that are consistent, or have a threshold similarity to, the musical style characteristics from which they were derived. Furthermore, the output data may be consistent with other constraints, such as visual style characteristics, cultural considerations (e.g., audio and / or images that are culturally appropriate, as determined by analysis module 106), user preferences, etc. In general, the output of computing system 100 may include a file including audio and / or video components that are correlated with the outcome of any analyses described herein, e.g., the style of the output is correlated to and / or has a threshold similarity to the style of the input. For example, GUI 103 may be a GUI of an application that provides users the capability to customize and create audiovisual content in an interactive media environment. Computing system 200 may fine-tune and / or generate such audiovisual content (including, for example, music, instrumentation, visual aesthetics, or chyron layers, etc.) responsive to determining environmental and / or behavioral cues, such as brand icons, international flags, cultural style of dance, etc. As such, in general, audiovisual content generation module 110 may generate audiovisual content considered to be compositionally accurate, stylistically faithful, and culturally considerate to the various input data from which it was derived or based on.

[0040] In the example of FIG. 1 A, audiovisual content generation module 110 of musical expression generation module 108 may map input indicative of tactile event 105 to at least onedata value included in the structured data generated by analysis module 106. Specifically, audiovisual content generation module 110 may generate audio data having a second set of musical style characteristics associated with the first set of musical style characteristics. That is, the generated audio data may be considered “new” music that is “stylistically faithful” to the input audio data, i.e., audiovisual content generation module 110 may alter the input audio data based on the input indicative of tactile event 105, in which the altered audio data is different from the input audio data, but includes a second set of musical style characteristics that have a threshold similarity to the first set of musical style characteristics. For example, the new or second set of musical style characteristics may correspond to data values that are within the data value ranges determined by analysis module 106.

[0041] As an example, using an application hosted on computing device 112, user 107 may select an original song from a library, in which computing system 100 may receive input indicative of the selection and generate structured data including data values corresponding to the musical style characteristics of the original song. To manipulate the original song, or create new musical expression based on the original song, for example, user 107 may provide multiple tactile events 105 by quickly tapping their finger at a location of a presence-sensitive display that corresponds to at least one portion of GUI 103 associated with the application. UI components 102 may detect the multiple tactile events 105, and send, to computing system 100, input indicative of the multiple tactile events 105. In some examples, computing system 100 may determine, based on the input, at least one gesture, and map, based on a grid, the at least one gesture to at least one data value from the one or more data values included in the structured data. For example, each tactile event may be associated with a timestamp, in which analysis module 106 may determine that the sequence of tactile events 105 is fast paced. In this example, analysis module 106 may map each tactile event to a data value corresponding to a single note in the original song, but may generate audio data 148 in which the sequence of notes corresponding to the sequence of tactile events 105 are played at a faster pace than that of the original song. In other words, user 107 may manipulate, e.g., the tempo of the original song by providing faster paced or slower paced touch input. However, while audio data 148 output by computing system 100 may have a different tempo than the original song from which it was derived, audio data 148 may still have a threshold similarity to the original song, i.e., audio data 148 may be easily recognized as a derivation from the original song, or can be considered “stylistically faithful” tothe original song. In some examples, audio data 148 may be output with visual data generated based on the input received by computing system 100, and / or audio data 148.

[0042] In this way, computing system 100 may generate new musical expression (e.g., in the form of new audiovisual content) from input audio data, visual data, and / or user input, which may be as simple as tactile event 105. In other words, to create new, engaging, high-quality, and / or audibly and visually pleasing content, users may not have to possess an advanced skill set typically required for creating such content. Instead, a user may effortlessly provide a touch input and generate musical expression having a high level of artistry. In this way, users who do not have musical backgrounds or experience with creating high-quality audiovisual content may still create captivating media or music, e.g., for social media platforms where such content is in high demand.

[0043] FIG. IB is a conceptual diagram illustrating another example of the computing system of FIG. 1 A configured to dynamically generate new musical expression based on musical analysis and detected motion, in accordance with one or more techniques of this disclosure. In some examples, some or all of the components and / or functionality attributed to computing system 100 may be implemented or performed by computing device 112, and vice versa. That is, some or all of the techniques described with respect to FIG. IB, and / or other examples in this disclosure, may be implemented or performed “on-device,” e.g., by computing device 112, and / or by computing system 100, which may represent a distributed computing system or a cloud computing system that provides one or more services via network 101. As shown in the example of FIG. IB, in some examples, computing system 100 may receive input indicative of motion. For example, user 107 may provide input indicative of motion 109 (which may be represented by the left and right arrows included in FIG. IB), which may be defined as when user 107 moves from one position to another position, such as in the direction of the left and / or right arrow included in FIG. IB. That is, computing system 100 may receive at least one input indicative of at least one motion 109, in which computing system 100 receives the at least one input in response to the at least one motion 109 being detected at one or more of UI components 102, such as a camera. In some examples, in addition to or alternative to receiving input indicative of motion 109, computing system 100 may receive at least one input indicative of one or more of position, axis, direction, velocity, shape, signal, or a combination thereof. In some examples, audiovisual content generation module 110 may alter the input audio data based on at least oneinput indicative of at least one motion 109. Specifically, in some examples, analysis module 106 may determine, based on the input indicative of motion 109, at least one gesture. In some examples, analysis module 106 may determine one or more gestural patterns 117 from at least one motion 109. That is, in some examples, analysis module 109 may be configured to determine a sequence or combination of gestures by analyzing a plurality of captured or detected motions. For example, as shown in FIG. IB, analysis module 106 may determine gestural pattern 117 indicative of user 107 moving a certain distance in a certain x-direction over a certain period of time. As further described below, gestures and / or gestural patterns 117 may be used to predict future gestures, determine styles of dance, alter input audio, etc. For example, to alter input audio data, audiovisual content generation module 110 may map the at least one gesture or gestural pattern 117 to at least one data value included in the structured data generated by analysis module 106. In some examples, the altered audio data may be different from the input audio data, but may include musical style characteristics that have at least a threshold similarity to the musical style characteristics of the input audio data.

[0044] FIG. 1C is a conceptual diagram illustrating another example of the computing system of FIG. 1 A configured to dynamically generate new musical expression based on musical analysis and visual input, in accordance with one or more techniques of this disclosure. In some examples, computing system 100 may receive one or more visual data inputs 111, which may include visual data sourced from a device camera. In some examples, one or more visual data inputs 111 may include visual data sourced from a user’s feed in a social media application. That is, in some examples, computing system 100 may receive visual data (e.g., images and / or videos) that are hosted on an application executing at computing device 112 and / or visual data captured by a device camera. In some examples, computing system 100 may receive visual data stored on computing device 112, such as images and / or videos stored in a camera roll. In some examples, analysis module 106 may process one or more visual data inputs 111 to determine one or more visual data values, such as one or more of brightness, color, warmth, transparency, pixelation, shading, texture, hue, saturation, a visual object, foreground, background, etc.

[0045] As described with respect to FIG. 1 A and FIG. IB, computing system 100 may also receive input audio data having a first set of musical style characteristics. In some examples, analysis module 106 may apply a machine learning model to the input audio data to generate structured data including one or more data values, in which each of the one or more data valuescorrespond to a respective musical style characteristic from the first set of musical style characteristics. In the example of FIG. 1C, audiovisual content generation module 110 may generate, based on the structured data and one or more visual data inputs 111, audio data having a second set of musical style characteristics associated with the first set of musical style characteristics. Computing system 100 may also generate, based on the structured data and one or more visual data inputs 111, one or more visual data outputs (e.g., generated pictures, videos, etc.), in which computing system 100 may output the one or more visual data outputs with at least a portion of the audio data having the second set of musical style characteristics. That is, in some examples, computing system 100 may generate new musical expression based on input visual content and input audio data. In some examples, computing system 100 may generate new visual content based on the input visual content and input audio data. In some examples, computing system 100 may generate visual content based on input audio data, generate audio data based on input visual content, and / or generate visual content based on input visual content. In general, computing system 100 may determine one or more associations between visual data and audio data to generate new visual and / or audio data.

[0046] For instance, analysis module 106 may represent an intelligently generated input vector circuit that may receive audio having a first set of style characteristics and may generate a structured data vector that quantifies those characteristics. Audiovisual content generation module 110 of musical expression generation module 108 may represent an intelligently generated output vector circuit that may receive the structured data vector as input (and in some examples, additionally or alternatively one or more visual data inputs 111) to synthesize audio exhibiting a second, but stylistically related, set of musical characteristics. In some examples, audiovisual content generation module 110 may additionally generate corresponding visual outputs (e.g., synthesized images or video clips). That is, by translating between input and output vectors that encode learned associations between audio and visual data, musical expression generation module 108 may generate new musical expression from visual input and audio input, may generate new visual expression from audio input, and / or may generate combined audiovisual output from visual input and audio input.

[0047] As an example, analysis module 106 may determine visual data input 111 to include darker colored images, in which audiovisual content generation module 110 may generate audio data including lower note pitches. In another example, analysis module 106 may receive a stillimage from a concert (such as an image of a singer in front of a microphone, as shown in FIG. 1C), receive at least a portion of input audio data detected at a device microphone during the concert, and then generate a concert-like display based on the still image and input audio data. As such, computing system 100 may “visualize” content, and / or generate more complex or comprehensive audiovisual content from smaller samples of data (e.g., generate a concert video from a still concert image).

[0048] FIG. ID is a conceptual diagram illustrating another example of the computing system of FIG. 1 A configured to dynamically generate recommendations based on analysis of user generated content, in accordance with one or more techniques of this disclosure. In some examples, computing system 100 may receive data indicative of user generated content (UGC) 113 including one or more of input audio data having a first set of musical style characteristics and visual data having a first set of visual style characteristics. For example, UGC 113 may include a user generated video including a user performing a specific cultural dance to a specific song. Analysis module 106 may apply a machine learning model to the data indicative of UGC 113 to determine one or more user preferences. That is, for example, analysis module 106 may determine a specific culture associated with the dance and / or the song, and determine one or more user preferences, such as a preference for content associated with the identified culture. For example, UGC 113 may depict a user performing a culturally specific dance to a particular song. Analysis module 106 may convert the UGC into feature vectors capturing movement, rhythm, audio descriptors, etc. By applying machine learning inference to the input vectors, analysis module 106 may identify the cultural context of the dance and song and may derive vectors that encode the user’s affinity for the musical and / or visual characteristics associated with the specific dance and song. These vectors may be fed forward as output vectors to drive downstream recommendation and / or generation tasks.

[0049] Musical expression generation module 108 may generate, based on the one or more user preferences, one or more recommendations including one or more of audio data having a second set of musical style characteristics and visual data having a second set of visual style characteristics. That is, in some examples, musical expression generation module 108 may generate recommended audio data that is stylistically similar to the input audio data included in UGC 113. For example, if UGC 113 includes a particular cultural song, musical expression generation module 108 may generate one or more other song recommendations identified asbeing associated with the same culture. In some examples, musical expression generation module 108 may generate an altered version of the song included in UGC 113. In some examples, the visual data recommended by computing system 100 may include, for example, videos or images associated with the visual style of UGC 113. For example, based on UGC 113 including a particular cultural dance, computing system 100 may generate a recommended video including another dance identified as being associated with the same culture.

[0050] The one or more recommendations may be provided to a user of computing device 112 via one or more graphical user interfaces. For example, GUI 103 may be a user profile GUI of an application, e.g., a social media platform, in which UGC 113 is displayed. In some examples, GUI 103 may be a feed GUI where users of the application may interact. For example, UCG 113 may be presented to other users of the application as a recommendation. As such, in some examples, the one or more recommendations presented to a user of an application may be user generated content from other users of the application that is determined by computing system 100 to be of a similar style to UGC 113. In some other examples, however, the one or more recommendations generated by computing system 100 may be considered “outliers” or audio and / or visual content that is stylistically different from UGC 113, such that a user is provided recommendations that typically may not be recommended to them.

[0051] FIG. IE is a conceptual diagram illustrating another example of the computing system of FIG. 1 A configured to dynamically generate recommendations based on biometric data analysis, in accordance with one or more techniques of this disclosure. In some examples, computing system 100 may send, to computing device 112, one or more recommendations including one or more of audio data having a set of musical style characteristics and visual data having a set of visual style characteristics. For example, computing system 100 may send recommended content including an audio and a style of dance. Responsive to outputting the one or more recommendations (e.g., one or more of UI components 102 displaying the one or more recommendations to user 107 via an application GUI), sensor(s) 114 may measure or detect biometric data of user 107, such as biometric data including data indicative of one or more of neuroactivity, eye movement, heart rate, respiratory rate, blood pressure, heart rate variability, sleep duration, skin temperature, and blood oxygen. Computing system 100 may receive the biometric data, in which analysis module 106 may interpret the biometric data. In some examples, analysis module 106 may determine one or more user preferences by at least applyinga machine learning model to the biometric data. Specifically, analysis module 106 may assign a score to each musical style characteristic and / or visual style characteristic to determine whether a user is satisfied or dissatisfied with a respective characteristic. That is, when computing system 100 receives biometric measurements (e.g., heart rate, galvanic skin response, facial expression data), analysis module 106 may parse those measurements into physiological state vectors. Analysis module 106 may apply a learning model that scores musical style and / or visual style characteristics against the user’s state vectors to infer satisfaction and / or dissatisfaction levels. Based on the user preferences (e.g., a preference for lighter imagery, a preference for slower tempos, a preference for dances associated with a specific culture, a preference for alternative rock music, etc.), audiovisual content generation module 210 may generate or more updated recommendations including one or more of updated audio data and updated visual data, and computing system 100 may send, to computing device 112, the one or more updated recommendations. That is, audiovisual content generation module 110 may interpret the resulting output vectors to refine recommended content, e.g., suggesting imagery with lighter palettes, music with slower tempos when the user’s biometric vectors indicate overstimulation, etc. As such, in some examples, computing system 100 may be configured to recommend audiovisual content by determining whether a user’s biometric data indicates the user is satisfied with or enjoys the audiovisual content being presented to them. For example, analysis module 106 may determine that recommended audiovisual content including certain images or sounds may elevate user 107’s heart rate. As such, audiovisual content generation module 110 may generate one or more updated recommendations including audiovisual content that has a threshold level of dissimilarity to the initially recommended audiovisual content. Thus, by iteratively updating these output vectors, computing system 100 may fine-tune the content generated and / or recommended by audiovisual content generation module 110, such that user 107 is presented content that is audibly and visually pleasing to them.

[0052] FIG. 2 is a block diagram illustrating another example computing system for dynamically generating new musical expression based on analysis of audio, visual, or user input, in accordance with one or more techniques of this disclosure. Computing system 200 may be similar to computing system 100 of FIG. 1A, FIG. IB, FIG. 1C, FIG. ID, and / or FIG. IE. As shown in FIG. 2, computing system 200 includes UI components 227 including one or more communication channels 225, one or more input / output (I / O) devices 228, one or moreprocessor(s) 224, communication units 226, and one or more storage devices 230. Storage device 230 further includes UI module 204 and musical expression generation module 208 including analysis module 206 and audiovisual content generation module 210, which may be similar to UI module 104 and musical expression generation module 108 including analysis module 106 and audiovisual content generation module 110, respectively, of FIG. 1A, FIG. IB, FIG. 1C, FIG. ID, and / or FIG. IE. As shown in the example of FIG. 2, computing system 200 may include machine learning module 218. Analysis module 206 may further include gesture analysis module 220, visual data analysis module 221, biometric data analysis module 222, and user generated content (UGC) analysis module 223 (collectively referred to herein as “modules 221-223”). Storage device 230 further includes API module 215 and information storage unit 216. In some examples, computing system 200 may include additional units and modules not shown in the example of FIG. 2. Some or all of the components and / or functionality attributed to computing system 200 may be implemented or performed by a computing device in communication with computing system 200.

[0053] Computing system 200 may include one or more communication units 226 configured to communicate with external devices by transmitting and / or receiving data at computing system 200, such as to and from remote computer systems or computing devices. Example communication units 226 include a network interface card, an optical transceiver, a radio frequency transceiver, or any other type of device that can send and / or receive data. Other examples of communication units 226 may be devices configured to transmit and receive Ultrawideband®, Bluetooth®, GPS, 3G, 4G, and Wi-Fi®, etc. that may be found in computing devices, such as mobile devices and the like.

[0054] Communication channels 225 may interconnect each of the components shown in FIG. 2 for inter-component communications (physically, communicatively, and / or operatively). In some examples, communication channels 225 may include a system bus, a network connection (e.g., to a wireless connection), one or more inter-process communication data structures, or any other components for communicating data between hardware and / or software locally or remotely.

[0055] UI components 227 may include one or more I / O devices 228. I / O device 228 may receive various inputs and generate various outputs. Examples of inputs may be tactile, audio, kinetic, optical input, etc. Input devices of I / O devices 228, in one example, may include a touchscreen, a touchpad, a mouse, a keyboard, a voice responsive system, a video camera,buttons, a control pad, a microphone or any other type of device for detecting input from a human or machine. Output devices of I / O devices 228, may include a sound card, a video graphics adapter card, a speaker, a display, or any other type of device for generating output to a human or machine.

[0056] UI module 204, API module 215, information storage unit 216, musical expression generation module 208, analysis module 206, audiovisual content generation module 210, machine learning module 218, gesture analysis module 220, visual data analysis module 221, biometric data analysis module 222, and user generated content (UGC) analysis module 223 (hereinafter “modules 204-223”) may perform operations described herein using software, hardware, firmware, or a mixture of hardware, software, and firmware residing in and executing on computing system 200 or at one or more other computing devices (e.g., a cloud-based application not shown in FIG. 2). For example, some or all of modules 204-223 may be included in and executable on a local computing device, such as computing device 112 of FIG. 1. As such, the techniques described herein may all be implemented locally on a computing device.

[0057] Computing system 200 may execute one or more of modules 204-223 with one or more processors 224 or may execute any or part of one or more of modules 204-223 as or within a virtual machine executing on underlying hardware. One or more of modules 204-223 may be implemented in various ways, for example, as a downloadable or pre-installed application, remotely as a cloud application, or as part of the operating system of computing system 200. Other examples of computing system 200 that implement techniques of this disclosure may include additional components not shown in FIG. 2.

[0058] In the example of FIG. 2, one or more processors 224 may implement functionality and / or execute instructions within computing system 200. For example, one or more processors 224 may receive and execute instructions that provide the functionality of UI components 227, communication units 226, one or more storage devices 230 and an operating system to perform one or more operations described herein. For example, one or more processors 224 may receive and execute instructions that provide the functionality of some or all of modules 204-223 to perform one or more operations and various functions described herein. The one or more processors 224 may include a central processing unit (CPU). Examples of CPUs include, but are not limited to, a digital signal processor (DSP), a general-purpose microprocessor, a tensor processing unit (TPU); a neural processing unit (NPU); a neural processing engine; a core of aCPU, VPU, GPU, TPU, NPU or another processing device, an application specific integrated circuit (ASIC), a field programmable logic array (FPGA), or other equivalent integrated or discrete logic circuitry, or other equivalent integrated or discrete logic circuitry.

[0059] One or more storage devices 230 within computing system 200 may store information, such as information retrieved from a user computing device, or other data discussed herein, for processing during the operation of computing system 200. In some examples, one or more storage devices of storage devices 230 may be a volatile or temporary memory. Examples of volatile memories include random access memories (RAM), dynamic random-access memories (DRAM), static random-access memories (SRAM), and other forms of volatile memories known in the art. Storage devices 230, in some examples, may also include one or more computer- readable storage media. Storage devices 230 may be configured to store larger amounts of information for longer terms in non-volatile memory than volatile memory. Examples of nonvolatile memories include magnetic hard disks, optical discs, floppy discs, flash memories, or forms of electrically programmable memories (EPROM) or electrically erasable and programmable (EEPROM) memories. Storage devices 230 may store program instructions and / or data associated with the modules 204-223 of FIG. 2.

[0060] UI module 204 may receive information and instructions from one or more associated platforms, operating systems, applications, and / or services executing at a computing device and / or computing system 200. In addition, UI module 204 may act as an intermediary between the one or more associated platforms, operating systems, applications, and / or services executing at the computing device and / or computing system 200 and various output devices of the computing device and / or computing system 200 (e.g., speakers, LED indicators, vibrators, etc.) to produce output (e.g., graphical, audible, tactile, etc.) with a computing device and / or computing system 200.

[0061] Musical expression generation module 208 may be implemented on a computing device in various ways. For example, musical expression generation module 208 may be implemented as a downloadable or pre-installed application or “app.” In another example, musical expression generation module 208 may be implemented as part of an operating system of a computing device.

[0062] Information storage unit 216 may be a storage repository for various data received by computing system 200. In some examples, information storage unit 216 may operate, at least inpart, as a cache for data received by computing system 200. In general, information storage unit 216 may be configured as a database, flat file, table, or other data structure stored within storage device 230. In some examples, information storage unit 216 may be shared between various modules executing at computing system 200 (e.g., between one or more of modules 204-223 or other modules not shown in FIG. 2). In other examples, a different data repository may be configured for a module executing at computing system 200 that requires a data repository. Each data repository may be configured and managed by different modules and may store data in a different manner. In some examples, computing system 200 may receive and store information in information storage unit 216 over a specified period of time.

[0063] In some examples, computing system 200 may retrieve data using API module 215. In some examples, API module 215 may be configured to enable the exchanging of data in a standardized format. For example, API module 215 may support REST (Representational State Transfer), which is a widely-used architectural style for building APIs that use HTTP (Hypertext Transfer Protocol) to exchange data between applications. In some examples, computing system 200 may host one or more applications that may call or use an operating system of computing system 200 that includes a central intelligence layer, in which each application may communicate with the central intelligence layer using API module 215. API module 215 may include an API that is common or public across all applications. In some examples, the central intelligence layer may communicate with a central device data layer, which may be a centralized repository of data for computing system 200. The central device data layer may communicate with various components of computing system 200, e.g., using a private API included in API module 215.

[0064] In general, analysis module 206 may be configured to perform various techniques for analyzing various input data, such as input audio data, i.e., musical data. For example, analysis module 206 may be configured to perform pattern recognition, e.g., identify patterns in musical style data. In some examples, analysis module 206 may algorithmically interpret musical style data. In some examples, analysis module 206 may receive small, representative data samples (e.g., “thin slice data”) and, through iterative processes, develop detailed and comprehensive models of information. As such, analysis module 206 may be configured to use specific data samples to infer broader patterns and relationships.

[0065] In general, analysis module 206 may analyze various types of input data, such as input audio data, to generate new musical expression in the form of altered or generated audio data and / or visual data. In general, analysis module 206 may be configured to classify and / or characterize input data. As an example, based on at least a portion of input audio data, analysis module 206 may detect different instruments in an ensemble, separate those instruments or groups thereof into stems, and identify the unique waveform characteristics of each stem or instrument. In one example, waveform profiles determined by analysis module 206 may be run as executable search queries across a library of candidate audio files. In another example, the waveform profiles may be used to locate, rank, and / or automatically select most compatible candidate instruments to substitute as alternate instruments for an initial audio file. In another example, the waveform profiles may be used for style transfer onto a new or different stem. In general, analysis module 206 may automate the means by which instruments, music beds, loops, and other audio are assessed for compatibility. In some examples, analysis module 206 may perform hierarchical pairing of music and instrument assignments (e.g., based on suitability of characteristics such as waveform, spectral, class, type, cultural context, etc.). As such, analysis module 206 may be configured to identify and / or generate audio that is stylistically similar to the input audio, but may involve different tempos, pitches, instruments, or other characteristics, such as those described above.

[0066] Analysis module 206 may be configured to manipulate input data in various ways. For example, analysis module 206 may apply an algorithm for suppression of detected major 3rds in bass lines. In some examples, analysis module 206 may transform chords detected by one or more neural networks into tone row data. Analysis module 206 may be configured to perform object recognition on the basis of shapes, objects, elements, graphic values, movement, or combinations thereof detected within an input device’s (e.g., a camera) field of view. Analysis module 206 may be configured to process user-provided media (e.g., user generated content), and incorporate and / or transform the user-provided media into audibly and visually interactive play states (e.g., interactive experiences or environments where both audio and visual elements respond dynamically to user inputs). These play states may be used in gaming, virtual reality (VR), augmented reality (AR), interactive installations, multimedia art, etc.

[0067] In some examples, data received by computing system 200 may be preprocessed. Preprocessing techniques may include extracting one or more additional features from raw data.For example, feature extraction techniques may be applied to the user input or retrieved instructions to generate one or more new, additional features.

[0068] In some examples, cloud-based and / or device-based techniques may be used to acquire and process data. In some examples, computing system 200 may store musical style data, audio data, audio signal data, etc. from files, buffered streams, or extract musical style data, audio data, audio signal data, etc. from over the air (OTA).

[0069] In some examples, computing system 200 may exchange data with one or more external computing systems in communication with computing system 200 (e.g., via a network), such as a cloud-hosted service operated by a third-party vendor, a licensee system, client device, and the like. In some examples, the external computing system may exchange data with computing system 200 via a secure application programming interface (API).

[0070] For example, in some examples, computing system 200 may send, to an external computing system, raw data, intermediary output data, etc. Computing system 200 may receive data from the external computing system, which may be used to generate user-facing output (e.g., visual data, audio data having the second set of musical style characteristics, etc.). As one example, computing system 200 may send data including raw data and / or derived input vectors to the external computing system for further processing and / or analysis. In some examples, the data may include metadata that identifies the type of content (audio, visual, biometric, contextual) and / or any processing requests (e.g., a request for a particular model, handling of data, etc.). For example, in some examples, the external computing system may execute one or more proprietary or partner-supplied machine learning models to perform various tasks. The external computing system may return data to computing system 200 that can be processed by analysis module 206 and / or machine learning module 218, such as data indicative of output vectors (e.g., labelled feature vectors, similarity scores, etc.). In some examples, data received from the external computing system may be used by analysis module 206 and / or machine learning module 218 to interpret other data received by computing system 200. In some examples, the data received from the external computing system may be provided directly to audiovisual content generation module 210 to generate output. As such, in some examples, computing system 200 may outsource data processing and / or analysis steps to external computing systems or platforms (e.g., external systems or platforms that host specialized third- party models), thereby conserving memory and computational costs.

[0071] In some examples, an entity (e.g., a business entity) associated with computing system 200 may have a license agreement with another entity associated with another computing system (e.g., an external computing system or platform, client device, and the like). In some examples, computing system 200 may be a backend-as-a-service (BaaS) system. In some examples, computing system 200 (e.g., a licensor system) may receive, from an external computing system (e.g., a licensee system), raw data, intermediary output data, etc. In some examples, computing system 200 may receive audio data having a first set of musical style characteristics from an external computing system. As one example, computing system 200 may receive data including raw data and / or derived input vectors from the external computing system for further processing and / or analysis by analysis module 206 and / or machine learning module 218. Computing system 200 may send data to the external computing system, which may be user-facing output or data used to generate user-facing output. For example, computing system 200 may return data indicative of output vectors (e.g., labelled feature vectors, similarity scores, etc.), visual data, audio data having the second set of musical style characteristics, final output, etc. to the external computing system (e.g., licensee system).

[0072] Furthermore, in some examples, computing system 200 may receive input data from one or more external computing systems. For example, computing system 200 may receive training data, vendor-curated datasets, etc., which may be used by analysis module 206 and / or machine learning module 218 to improve output quality (e.g., input data provided by an external computing system may help to improve stylistic fidelity, cultural relevance, etc.).

[0073] In some examples, computing system 200 and one or more external computing systems may engage in a progressive co-processing loop, in which data may be shuttled back and forth in multiple passes before any audiovisual content is output to the user. For example, during a first pass, computing system 200 may transmit raw data to an external computing system, in which the external computing system may process the raw data to output feature vector(s). For example, the external computing system may output feature vector(s) indicative of a first set of musical style characteristics. Using the feature vector(s), computing system 200 may generate intermediary output, and may return the intermediary output and / or other information to the external computing system in a second pass. For example, computing system 200 may generate, using the feature vector(s), intermediary audio data having another set of musical style characteristics associated with the first set of musical style characteristics. Computing systemmay send the intermediary audio data and other information (e.g., beat-onset timestamps, candidate chord substitutions, user preferences, visual data, user biometric data, etc.) back to the external computing system. The external computing system may then apply a higher capacity generative model to the intermediary output and / or other information to further refine the output. In some examples, the external computing system may compare another set of musical style characteristics with the first set of musical style characteristics to determine a similarity score, and may refine the output audio data until a threshold similarity is met. Furthermore, in some examples, each iteration may involve passing only incremental information needed for a next refinement step, thereby conserving bandwidth. In general, computing system 200 and one or more external computing systems may engage in the progressive co-processing loop to generate output for any number of iterations until the generated output is determined to have a threshold level of similarity, certainty, etc. In this way, computing system 200 may conserve memory and computational costs while converging on final, high-fidelity audiovisual output that better aligns with the user’s stylistic preferences.

[0074] In general, analysis module 206 may employ machine learning module 218 to apply one or more machine learning models to data received and / or stored by computing system 200, which may be referred herein as “input data.” Machine learning module 218 may apply the one or more machine learning models to input audio data to generate the structured data including one or more data values, in which each of the one or more data values correspond to a respective musical style characteristic from the first set of musical style characteristics. In some examples, the structured data may include one or more data values indicative of one or more of a frequency spectrum, an amplitude, a timbre, a note pitch, a note interval, a beat, a fill, a riff, a melody contour, a melody motif, a chord progression, a key signature, a musical scale mode, a tonality, a tempo, a meter, a rhythmic pattern, an instrumentation, a density, a polyphony, a volume, a note articulation, a note expression, a structure, a section, a genre, and historical context data. That is, analysis module 206 may invoke machine learning module 218 as an Al-input-vector circuit. For example, analysis module 206 may convert raw audio into a structured feature vector whose elements may represent any combination of frequency spectrum, amplitude, timbre, pitch, interval, beat, riff, melody contour, chord progression, key signature, scale mode, tonality, tempo, meter, rhythmic pattern, instrumentation, density, polyphony, volume, articulation, expression, structural section, genre, or historical-context metadata. This structured vector maybe stored in text, binary, or other formats suitable for processing by machine learning module 218.

[0075] In some examples, one or more neural networks may be applied to input audio data to generate structured data that may either be time-shifted or in real time. In some examples, machine learning module 218 may employ one or more cloud-based and / or on-device machine learning models to generate the structured data. In some examples, machine learning module 218 may employ the MAD MOM model (accessible at https: / / github.com / CPJKU / madmom) which is an audio signal processing library configured to perform music information retrieval tasks. In general, machine learning module 218 may use algorithmic composition techniques to generate tone row data or other musical data. Tone row data may include all twelve tones used without repetition, as required by twelve-tone serialism. In some examples, machine learning module 218 may generate, based on input audio data having a first set of musical style characteristics (e.g., initial or original audio, or stored or extracted audio), tone row data including sequences of numbers from 1 to 12 in a randomized order, in which each number (representing a pitch class) appears only once in each generated row. Various other methods may be used to manage or manipulate the list of pitch classes, ensure proper randomization, etc. For example, the one or more data values described herein may be indicative of one or more note pitches, in which each of the one or more notes pitches is statistically weighted. In some examples, the tone row data may be included in an XML file. In some examples, the structured data (e.g., the tone row data) may be stored at least temporarily in information storage unit 216.

[0076] One or more modules of modules 221-223 may also employ machine learning module 218 to perform various tasks. As such, machine learning module 218 may involve machine learning techniques. In some examples, machine learning module 218 may perform various types of natural language processing (NLP) based on input data. In some examples, machine learning module 218 may perform classification, regression, clustering, anomaly detection, recommendation generation, translation, summarization, organization tasks, and / or other tasks. In some examples, machine learning model 218 may use recurrent neural networks (RNNs) and / or transformer models (self-attention models), such as GPT-3, BERT, and T5.

[0077] In some examples, machine learning module 218 may analyze and process input data, such as input audio data and / or input visual data. In some examples, machine learning module 218 may identify patterns and / or combinations of musical and / or visual elements present in theinput data. In some examples, machine learning module 218 may determine, based on identified patterns and / or combinations of musical and / or visual elements present in the input data, a particular musical and / or visual style for the input data, e.g., a composer, period, genre, culture, etc. associated with the input data. In some examples, machine learning module 218 may determine similarity between input and output data, e.g., by determining whether the input and output data is associated with the same musical and / or visual style, and / or each include musical and / or visual style characteristics that have a threshold similarity.

[0078] In some examples, machine learning module 218 may perform various types of classification to classify the input data into one or more classes or categories, such as binary classification, multiclass classification, or discrete categorical classification. The classifications may be single-label or multi-label. In some examples involving classification techniques, machine learning module 218 may be trained using supervised learning techniques.

[0079] In some examples, machine learning module 218 may perform various types of regression to provide output data in the form of a continuous numeric value, such as linear regression, polynomial regression, nonlinear regression, simple regression or multiple regression. The continuous numeric value may correspond to any number of different metrics or numeric representations.

[0080] In some examples, machine learning module 218 may perform various types of clustering to identify one or more clusters to which the input data most likely corresponds. Machine learning module 218 may identify one or more clusters within the input data. That is, in instances in which the input data includes multiple objects or other entities, machine learning module 218 may sort the multiple entities included in the input data into a number of clusters. In some examples involving clustering techniques, machine learning module 218 may be trained using unsupervised learning techniques.

[0081] In some examples, machine learning module 218 may include a parametric model, a nonparametric model, a linear model, or a non-linear model. In general, machine learning module 218 may be or include one or more various different types of machine-learned models, such as classifier models (e.g., linear classification models, quadratic classification models, etc.), regression models (e.g., simple linear regression models, multiple linear regression models, logistic regression models, stepwise regression models, multivariate adaptive regression splines, locally estimated scatterplot smoothing models, etc.), and neural networks (e.g., deep neuralnetworks, generative neural networks such as generative adversarial networks, transformer-based neural networks, feed forward neural networks, recurrent neural networks, convolutional neural networks, autoencoders, etc.). Machine learning model 218 may combine (e.g., stack) multiple various neural networks to form more complex networks. In some examples, machine learning model 218 may employ at least one neural network to process sequential data, such as timeseries data (e.g., audio data including time stamped data, sensor data versus time, visual data captured at various times, etc.), notes in a musical composition, sequential gestures, etc.

[0082] In some examples, the input data may or may not include feature embeddings. In some examples, one or more neural networks may be employed to provide feature embeddings based on the input data.

[0083] In some examples, machine learning module 218 may use reinforcement learning techniques, such as Markov decision processes, dynamic programming, Q functions or Q- learning, value function approaches, deep Q-networks, differentiable neural computers, asynchronous advantage actor-critics, deterministic policy gradient, etc.

[0084] In some implementations, machine learning module 218 may be an autoregressive model. In some implementations, machine learning module 218 may include or form part of a multiple model ensemble. Example ensemble techniques include a random forest, stacking, boosting, etc.

[0085] In some examples, machine learning module 218 may employ multiple machine-learned models that may or may not be linked and / or trained jointly.

[0086] In some examples, machine learning module 218 may be used to preprocess the input data for subsequent input into another model. Example preprocessing techniques may include dimensionality reduction techniques and embeddings (e.g., matrix factorization, principal components analysis, singular value decomposition, etc.), clustering, classification, regression, etc. In some examples, during training, the input data may be manipulated in various ways, such as by adding noise, changing color, shade, or hue, magnification, segmentation, amplification, etc.

[0087] In some examples, machine learning module 218 may provide output data responsive to receiving the input data. In some examples, the output data may include content (e.g., audiovisual content) either stored locally on the user device or in the cloud, that is relevantly shareable along with the initial content selection. In some examples, the output data may influence downstream processes or decision-making. For example, the output data from machinelearning module 218 may be used by audiovisual content generation unit 210 to generate audio data and / or visual data.

[0088] Machine learning module 218 described herein may be trained using various training types or techniques. For example, in some examples, machine learning module 218 may be trained using supervised learning, backward propagation of errors in conjunction with an optimization technique (e.g., gradient-based techniques), unsupervised learning techniques, semi-supervised techniques, reinforcement learning, etc. In some examples, generalization techniques may be performed during training to improve the generalization of machine learning module 218. In some examples, transfer learning techniques may be used to provide an initial model from which to begin training of machine learning module 218. In some examples, machine learning module 218 may continuously or periodically train one or more machine learning models, e.g., by establishing a feedback loop.

[0089] In some examples, machine learning module 218 may be included in different portions of computer-readable code on a computing device, e.g., machine learning module 218 may be included in an application or program, in an operating system of a computing device, etc.

[0090] The machine learning techniques described herein are readily interchangeable and combinable. Although certain example techniques have been described, many others may exist and may be used in conjunction with aspects of the present disclosure.

[0091] Thus, in general, machine learning module 218 may apply one or more machine learning techniques described above to various types of input data to generate various types of output data that may be used by one or more of modules 220-223 and / or audiovisual content generation module 210.

[0092] In general, audiovisual content generation module 210 may be configured to generate one or more of output audio data and output visual data. In general, computing system 200 may be configured to alter various input audio and / or visual data, which may be received from an input device or a library (e.g., information storage unit 216), in which the various audio and / or visual data may be altered based on one or more various user inputs, such as detected motion or gestures, tactile events, biometric data, user generated content, and the like. In some examples, audiovisual content generation module 210 may be configured to generate audio data based on visual data, and / or generate visual data based on audio data. Furthermore, when generating audio or visual content, audiovisual content generation unit 210 may be configured to only generateoutput that is stylistically faithful or has a threshold similarity to the content from which it was derived. For example, audiovisual content generation unit 210 may generate new audio that is “within key” of the initial audio from which it was derived, such that a user may create high quality musical expression without any “bad” notes, or create high quality musical expression that sounds harmonically consistent and pleasing. Furthermore, audiovisual content generation unit 210 may generate new visual output that is, for example, aesthetically similar to the visual content from which it was derived, and / or is determined to be associated with particular musical style characteristics. As such, audiovisual content generation module 210 may provide users the capability to create “synthesized” audio and visual content that is visually and audibly pleasing, without requiring users to have high levels of skill and experience that are typically required for creating such content.

[0093] In general, computing system 200 may receive various types of inputs to generate musical expression, and / or audiovisual content. For example, computing system 200 may receive input indicative of at least one tactile event (such as tactile event 105 of FIG. 1 A), input indicative of at least one motion (such as motion 109 of FIG. IB), visual data (such as images and / or videos, e.g., as described with respect to FIG. 1C), user generated content (such as images, videos, audio data, and / or user data associated with user generated content, e.g., as described with respect to FIG. ID), and / or biometric data associated with a user (e.g., biometric data, e.g., as described with respect to FIG. IE).

[0094] The modules of musical expression generation module 208 may each be configured to process the various types of input data and / or generate output data. For example, gesture analysis module 220 may be configured to process input indicative of at least one tactile event, input indicative of at least one motion, and / or other input indicative of at least one gesture. In some examples, gesture analysis module 220 may be configured to identify one or more gestural patterns based on input. In general, gestural analysis module 220 may be configured to translate gestural input to musical expression output. In some examples, gestural analysis module 220 may interpret behavior patterns captured in gestural input as musical expression (e.g., an orchestra conductor moving their hands in various positions). In some examples, gestural analysis module 220 may be configured to process gestural input for generating real-time or near real-time performance of multiple instruments or vocals by a single user, in which the multiple instruments or vocals may be controlled, manipulated, or “played” by different aspects of theuser’s movement (e.g., detectable body points). For example, one example may include a user using two hands to control different voices in an ensemble simultaneously. Another example may include a user using one hand to control an instrument, and the other hand to control dynamic functionality of the instrument (e.g., volume, signal processing, effects, etc.). Another example may include a single motion playing one voice, and a computer assisted voice playing accompaniment simultaneously in response. It should be noted that the various forms of musical expression described herein may be created virtually. That is, for example, the “instruments” referred to herein may not be physical instruments, but rather virtual instruments or audio data including sound produced by various instruments. In some examples, input data may be associated with a physical instrument, but may only include a visual representation of the instrument. For example, input visual data may be indicative of gestures and / or gestural patterns associated with, e.g., playing an air guitar. In some other examples, however, the input data may include audio data from a physical instrument and / or visual data including a physical instrument. As such, in general, the various types of input data to the computing system may correspond to physical entities (e.g., real instruments, a user providing live vocals, a user dancing, etc.), and / or may correspond to visual representations of entities (e.g., one or more gestures indicative of an instrument, conductor batons, etc.).

[0095] In some examples, gesture analysis module 220 may scale an axis of a grid with one or more data values included in the structured data generated by machine learning module 218, in which each of the one or more data values correspond to a respective musical style characteristic from a set of musical style characteristics. For example, machine learning module 218 may generate, based on initial audio data having a first set of musical style characteristics, structured data that includes a range of note pitches that are associated with the initial audio data. Gesture analysis module 220 may scale an axis of a grid (e.g., a grid associated with a touch-sensitive input device, a frame in which an input device (e.g., a camera) captures visual data (e.g., motion or movement), etc.) with the range of note pitches. In some examples, computing system 200 may capture or detect (e.g., via one or more of I / O devices 228, such as a sensor, camera, accelerometer, etc.) motion and / or visual data indicative of one or more user positions across the grid in one or more of an x-direction, y-direction, and z-direction, in which each user position from the one or more user positions is associated with one or more of a timestamp and a velocity.In some examples, capturing the visual data indicative of the one or more user positions is responsive to at least one gesture being detected at one or more input devices.

[0096] For example, computing system 200 and / or a computing device in communication with computing system 200 may capture visual data when an input device detects a user’s hand moving across a frame of the input device in a y-direction, in which the user’s hand movement may be indicative of one or more positions that are each associated with a timestamp and velocity (i.e., speed and direction). Gesture analysis module 220 may receive the captured visual data indicative of the user’s movement and map each user position from the one or more user positions to at least one axis value of the grid. Audiovisual content generation module 210 may generate, based on at least one mapped user position, audio data. In some examples, the generated audio data may be new or altered audio data that is different from the initial audio data from which the structured data was generated. That is, gesture analysis module 220 may map, for example, the range of note pitches derived from initial audio data to an axis of a grid, in which the mapped user positions may be used by audio generation unit 210 to generate a new sequence of note pitches that is different from the sequence of note pitches in the initial audio data. In some examples, however, the new or altered audio data may include new musical style characteristics that have a threshold similarity to the initial musical style characteristics of the initial audio. As such, while a user may alter or manipulate the initial audio by performing one or more gestures, the new or altered audio output by audiovisual content generation module 210 may remain stylistically faithful to the initial audio.

[0097] In some examples, gesture analysis module 220 may be configured to perform gesture prediction. That is, in some examples, gesture analysis module 220 may predict (e.g., by employing machine learning module 218) one or more gestures and / or gestural patterns based on the captured visual data. For example, continuing the example above, based on data indicative of the user’s hand repeatedly alternating between moving in a positive y-direction and moving in a negative y-direction (i.e., up and down), gesture analysis module 220 may predict the user’s hand to continue alternating between moving in a positive y-direction and moving in a negative y- direction. As such, while the captured data indicative of one or more user positions across the grid may map to axis values resulting in a sequence of, for example, note pitches, the gestures predicted by gesture analysis module 220 may result in an additional sequence of note pitches. Inthese examples, the audio generated by audiovisual content generation module 210 may be based on both sequences of note pitches.

[0098] In some examples, analysis module 206 may employ visual data analysis module 221 to process visual data, such as images and videos. In some examples, computing system 200 may receive one or more visual data inputs such as input images or videos, which may be visual data captured by an input device and / or visual data associated with user generated content. In some examples, the one or more visual inputs may be associated with a still image and / or a video (e.g., a video of a concert, a video of a user dancing, etc.). In some examples, computing system 200 may receive audio data, in which the audio data may be audio data stored in information storage unit 216. In some examples, the audio data may be audio detected at an input device, such as a voice recording, audio from a live musical performance, etc. In general, machine learning module 218 may apply a machine learning model to the audio data to generate structured data including one or more data values corresponding to musical style characteristics. In some examples, the one or more data values may be weighted data values. That is, in some examples, machine learning module 218 may assign a weight to each data value based on various factors, e.g., a frequency of a corresponding musical style characteristic, a determined importance ranking of a corresponding musical style characteristic, etc.

[0099] Visual data analysis module 221 may be configured to generate one or more visual outputs by applying an algorithm to the one or more weighted data values to generate one or more visual representations. For example, the one or more weighted data values may be indicative of one or more statistically weighted note pitches. Visual data analysis module 221 may apply an algorithm including one or more of a semantic rules table and a machine learning model to the weighted data values to generate one or more equally weighted visual data values.

[0100] For example, using a semantic rules table, visual data analysis module 221 may translate the weighted values and / or quality associated with at least a portion of audio data into equally weighted visual interpretation, i.e., musical style characteristics and their corresponding data values are imputed as contextually weighted graphics inputs in accordance with a set of rules. The visual data values may correspond to one or more visual representations and may each be indicative of one or more of brightness, color, warmth, transparency, pixelation, shading, texture, hue, saturation, a visual object, foreground, and background. In general, audiovisual content generation module 210 may generate one or more of audio and visual output, such as visualrepresentations corresponding to the weighted visual data values determined by visual data analysis module 221. For example, based on a note pitch with a higher statistical weighting, audiovisual content generation module 210 may generate a corresponding visual representation that includes a higher level of brightness, a lower level transparency, etc. The visual representation may be an image, a video, a visual effect, user interface element, and the like.

[0101] More specifically, musical style characteristics of the input audio (e.g., the harmonic and / or rhythmic accompaniment to a musical work) may inform and constrain interchangeable layers of customizable graphical filters, vertex shaders, effects, and motifs. The rendering, presentation, modification, weight, animation, behavior, and interplay of the visual representations described above may be based on the musical style characteristics, which may be influenced, generated, or altered by user input, such as input indicative of tactile events, motion, gestures, etc. Additionally, the visual representations may include visual style characteristics inherent to a selected instrument (e.g., a virtual instrument selected by a user of an application), an audio sample, a modality setting, ambient sounds, player vocalization, manner of performance, etc.

[0102] Furthermore, in some examples, various constraints may be applied to any data generated by analysis module 206 such that the audiovisual content output by computing system 200 remains stylistically faithful or maintains a threshold similarity to the audio and / or visual content from which it was derived. For example, the visual representations described above may be constrained by data models, such as the data models for generating tone row data. In some examples, the visual representations may be generated by applying a machine learning model trained on various types of audiovisual data, in which the machine learning model may employ one or more of the machine learning techniques described above with respect to machine learning module 218. That is, in some examples, machine learning module 218 may represent an Al system or agent that utilizes various machine learning models, rules, and data processing techniques to generate output.

[0103] In general, visual data analysis module 221 may be configured to generate, based on audio data and / or visual data (including, for example, captured data within a camera field of view, visual data sourced from an application, such as a social media application, visual data sourced from a library, etc.), visual content including visual representations, i.e., interpretations, of the audio data and / or visual data. In some examples, audiovisual content generation module210 may output the visual content with at least a portion of new or altered audio generated by audiovisual content generation module 210.

[0104] In some examples, with explicit user consent, analysis module 206 may employ biometric data analysis module 222 to process biometric data associated with a user. In some examples, the audio and / or visual content output by computing system 200 may be recommended audio and / or recommended visual content for a user. In some examples, the recommended audiovisual data includes one or more data values indicative of one or more of a dance posture, a dance motion, a dance sequence, note pitches, note intervals, beats, fills, riffs, melody contour, melody motifs, chord progressions, key signatures, musical scale modes, tonality, tempo, meter, rhythmic patterns, instrumentation, density, polyphony, volume, note articulation, note expression, structure, sections, genre, and historical context data.

[0105] In some examples, with explicit user consent, biometric data analysis module 222 may receive biometric data responsive to an input device detecting one or more physiological characteristics from a user when the recommended audiovisual content is presented to the user. In general, biometric data analysis module 222 may perform biometric processing with explicit user consent and under an explicit, revocable consent model that complies with applicable data protection statutes. For example, biometric feature vectors may be stored ephem erally in encrypted memory and / or may be irreversibly hashed to prevent reconstruction of user physiological data. Based on the biometric data, biometric data analysis module 222 may assign a score to each data value described above (e.g., a heart rate above a threshold value may correspond to a lower score, which may indicate a lower user satisfaction level), i.e., biometric data analysis module 222 may be configured to determine a user’s preferences for a particular musical style characteristic and / or visual style characteristic. Biometric data analysis module 222 may then generate, based on each assigned score, one or more updated recommendations. For example, biometric data analysis module 222 may determine that audiovisual content including lower note pitches and darker imagery results in a high heart rate for a user (e.g., by increasing anxiety in the user, etc.). As such, biometric data analysis module 222 may assign a lower score to the data values corresponding to the lower note pitches, a higher score to the data values corresponding to the higher note pitches, a lower score to the data values corresponding to dark colors, and a higher score to the data values corresponding to lighter colors. As such, the updated recommended audiovisual content generated by audiovisual content generation module 210 mayinclude higher pitched sounds and lighter imagery. In this way, biometric data analysis module 222 may be employed by computing system 200 to fine-tune or personalize the audiovisual content that is output to a user, such that the audiovisual content is more aligned to specific user preferences.

[0106] In some examples, analysis module 206 may employ UGC analysis module 223 to process user generated content, such as images, videos, captured movement, audio data, and / or user data associated with user generated content. In some examples, analysis module 206 may determine information indicative of user behavior, and / or may determine context information for a user. In some examples, UGC analysis module 223 may apply a machine learning model to the data indicative of user generated content to determine one or more user preferences. Specifically, in some examples, UGC analysis module 223 may identify one or more of at least one musical style characteristic from a first set of musical style characteristics included in input audio data and at least one visual style characteristic from a first set of visual style characteristics included in input visual data. For example, UGC analysis module 223 may identify one or more of a dance posture, a dance motion, a dance sequence, frequency spectrum, an amplitude, a timbre, a note pitch, a note interval, a beat, a fill, a riff, a melody contour, a melody motif, a chord progression, a key signature, a musical scale mode, a tonality, a tempo, a meter, a rhythmic pattern, an instrumentation, a density, a polyphony, a volume, a note articulation, a note expression, a structure, a section, a genre, historical context data, brightness, color, warmth, transparency, pixelation, shading, texture, hue, saturation, a visual object, foreground, and background. Continuing the example from FIG. ID, UGC analysis module 223 may receive UGC data from a user account and identify, for example, one or more dance postures, dance sequences, chord progressions, etc. as being indicative of a particular cultural song and dance. In some examples, UGC analysis module 223 may identify cultural icons, such as flags, traditional clothing, etc. As such, UGC analysis module 223 may identify a particular “style” of user generated audiovisual content (e.g., audio data having a first set of musical style characteristics and / or visual data having a first set of visual style characteristics), in which audiovisual content generation module 210 may generate recommended audiovisual content that has a threshold similarity to the user generated content. For example, audiovisual content generation module 210 may generate recommended audio data having a second set of musical style characteristicsassociated with the first set of musical style characteristics and / or visual data having a second set of visual style characteristics associated with the first set of visual style characteristics.

[0107] In some examples, UI module 204 of computing system 200 may output at least a portion of the audio data having the second set of musical style characteristics and / or instructions for generating one or more graphical user interfaces including visual data associated with input received by computing system 200. For example, UI module 204 may send, to a computing device in communication with computing system 200, and / or one or more of UI components 227, the new audio data having the second set of musical style characteristics and / or the instructions for generating one or more graphical user interfaces including the visual data.

[0108] FIG. 3 is a conceptual diagram illustrating an audiovisual content generation module configured to transform structured data based on analysis of audio, visual, or user input, in accordance with one or more techniques of this disclosure. In general, mapping module 332 may map input received by the computing system to one or more data values included in the structured data to generate new audio data, or audio data having a second set of musical style characteristics that is based on the first set of musical style characteristics of the input audio data, but is different from the first set of musical style characteristics of the input audio data.

[0109] In examples in which the computing system described herein receives at least one input indicative of at least one tactile event and / or motion, mapping module 332 may map, based on a visual grid (e.g., a grid associated with a touch-sensitive input device, a frame in which an input device (e.g., a camera) captures visual data (e.g., motion or movement), etc.), the at least one input to at least one data value from the one or more data values included in the structured data. As an example, a GUI associated with the application may include a virtual instrument, which a user may interact with to play one or more note pitches. In some examples, the input data may be associated with a timestamp and / or a velocity. In some examples, mapping module 332 may match or align, using timestamps, sequences of mapped data values with the data values associated with the initial audio. For example, in some examples, the velocities and timestamps associated with a sequence of user positions may result in altered audio with a sequence of note pitches that is faster than the corresponding sequence of note pitches in the initial audio.Mapping module 332 may determine a mapping between the faster sequence and the initial sequence, such that similarity determination module 334 may better determine whether the faster sequence achieves a threshold of similarity to the initial sequence.

[0110] In general, similarity determination module 334 may be configured to determine whether different audios, audio signals, input and output audio data, etc. are associated with the same musical style, and / or each include musical style characteristics that have a threshold similarity. In some examples, similarity determination module 334 may employ one or more machine learning models to determine, based on identified patterns and / or combinations of musical elements present in audio data, a particular musical style for the audio data, e.g., a composer, period, genre, or culture associated with the audio data. Similarity determination module 334 may employ one or more machine learning models to determine a similarity score or percentage between different audios, audio signals, input and output audio data, etc. For example, the similarity score or percentage may indicate a level of similarity between one or more musical style characteristics, data values, etc. associated with a respective audio.[OHl] In some examples, mapping module 332 may map various outputs to various inputs, and vice versa. Furthermore, in general, the output of the computing system described herein may include a file including audio and / or video components that are correlated with the outcome of any analyses described above, e.g., the style of the output is correlated to and / or is within a threshold similarity to the style of the input. As an example, mapping module 332 may overlay at least a portion of the altered or new audio data with visual data on the basis of one or more determined associations between the audio and visual data. In the example in which the computing system captures visual data indicative of a user moving across the grid (e.g., a user performs a sequence of dance movements), the visual data may include a video recording of the user’s movements. In this example, mapping module 332 may map at least a portion of the altered or new audio data to the visual data based on, for example, timestamps associated with the visual data. In another example, mapping module 332 may map at least a portion of the altered or new audio data to other data, such as one or more visual representations. For example, the one or more visual representations may correspond to one or more visual data values each indicative of one or more of brightness, color, warmth, transparency, pixelation, shading, texture, hue, saturation, a visual object, foreground, and background. As an example, a low note pitch produced by a user tapping their screen may be mapped by mapping module 332 to a darker color, in which the computing system may output instructions for generating a GUI element associated with the darker color when the note pitch is played.

[0112] In some examples, mapping module 332 may generate intermediary audio data having a third set of musical style characteristics associated with the first set of musical style characteristics. Specifically, in some examples, mapping module 332 may generate intermediary audio data that is not provided as final output by the computing system. In these examples, similarity determination module 334 may compare the third set of musical style characteristics with the first set of musical style characteristics to determine a similarity score. That is, while the third set of musical style characteristics may correspond to data values that are within the ranges determined by analysis module 206 of FIG. 2, in some examples, similarity determination module 334 may determine that the third set of musical style characteristics results in new or altered audio data that is too dissimilar to the initial or original audio data. As such, similarity determination module 334 may determine whether the similarity score is less than a threshold similarity score. In some examples, similarity determination module 334 may employ machine learning methods, such as those described herein, to determine similarity scores and / or percentages. Example threshold similarity score percentages may include, but are not limited to, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, and 100%.

[0113] Responsive to similarity determination module 334 determining the similarity score is less than the threshold similarity score, transformation module 336 may apply one or more transformations to the intermediary audio data having the third set of musical style characteristics to generate the audio data having the second set of musical style characteristics. That is, transformation module 336 may further apply one or more range constraints to one or more mapped data values, such that the new or altered set of musical style characteristics may vary less from the initial or original set of musical style characteristics. As such, the “audio data having the second set of musical style characteristics” described herein may be considered new or altered audio that is determined by similarity determination module 334 to have a similarity score that satisfies the threshold similarity score. In some examples, similarity determination module 334 may determine a similarity score for the intermediary audio data having the third set of musical style characteristics that satisfies the threshold similarity score. In these examples, the computing system may output, based on the mapping, at least a portion of the audio data having the third set of musical style characteristics, in which the audio data having the third set of musical style characteristics is the same as the audio data having the second set of musical style characteristics (i.e., in which no transformations are applied to the audio data by transformationmodule 336). As such, by determining a threshold of similarity between the new audio data and the original audio data is achieved, the computing system described herein may only output new audio data that is stylistically faithful to the original audio from which it was derived.

[0114] FIG. 4 is a conceptual diagram illustrating a gesture analysis module configured to map gestures to dynamically generate new musical expression, in accordance with one or more techniques of this disclosure. Gesture analysis module 420 may be similar to gesture analysis module 220 of FIG. 2. As shown in the example of FIG. 4, gesture analysis module 420 includes scaling module 439, gesture determination module 440, grid mapping module 441, and gesture prediction module 442. In general, gesture analysis module 420 may be configured to process input data indicative of one or more gestures and map the input data to one or more data values that result in new musical expression.

[0115] In general, audio data received and / or stored by the computing system may include a set of musical style characteristics that contribute to the sound or “style” of the audio. The computing system described herein may apply an algorithm or machine learning model to the audio data to generate structured data including one or more data values corresponding to a respective musical style characteristic from a set of musical style characteristics. As an example, the structured data may be tone row data corresponding to note pitches included in the audio data. In some examples, the structured data may include a range of data values for each musical style characteristic. Scaling module 439 may scale an axis of a grid with the data values included in the structured data. As an example, consider audio data having a particular set of note pitches, such as C and D. The structured data generated based on the audio data may include a range of data values for note pitch C, such as a range of data values corresponding to tone row C, Db, E, F, G#, A, B, C$, D, Eb, F$, G, or tone row C, D, Eb, F$, G#, Bb, A, C$, F, B, E, G, in which both example ranges start with a data value corresponding to note pitch C. The structured data may further include a range of data values for note pitch D, such as a range of data values corresponding to tone row D, E, F, G, Ab, B, C, D$, F$, G#, Bb, C$, or tone row D, F, G#, Eb, A, C$, F$, Bb, G, E, C, Ab, in which both example ranges start with a data value corresponding to note pitch D. In this example, scaling module 439 may scale an x-axis of a grid with a range of data values for note pitch C, and may scale a y-axis of a grid with a range of data values for note pitch D. In some examples, scaling module 439 may scale an axis of a grid with more than one range of data values. For example, if note pitch C and note pitch D are played at the same time(e.g., by different instruments) in the audio, scaling module 439 may scale a y-axis of a grid with both ranges of data values for note pitch C and D. As such, grid mapping module 441 may map one or more user positions associated with a particular timestamp to more than one axis value. As an example, a user’s hand position at a specific timestamp may be mapped to a data value included in the range of data values for note pitch C and a data value included in the range of data values for note pitch D. In this example, the audio generated by the computing system based on this mapping may include, for example, note pitch Db and note pitch E in place of note pitch C and note pitch D, respectively, in which note pitch Db and note pitch E are played at the same time.

[0116] Grid mapping module 441 may map each user position from the one or more user positions to at least one axis value. In some examples, grid mapping module 441 may map one or more user positions associated with a particular timestamp to more than one axis value. For example, continuing the example above, grid mapping module 441 may map a user’s single hand position at a specific timestamp to a data value included in the range of data values for note pitch C, such as a data value corresponding to Db, and a data value included in the range of data values for note pitch D, such as a data value corresponding to note pitch E. In another example, grid mapping module 441 may map a user’s left hand position at a specific timestamp to a data value included in the range of data values for note pitch C, such as a data value corresponding to Db, and map the user’s right hand position at the specific timestamp to a data value included in the range of data values for note pitch D, such as a data value corresponding to note pitch E. As such, grid mapping module 441 may map various user positions corresponding to multiple points (e.g., a left hand, a right hand, a torso, a head, an object, etc.) within the grid to multiple axes scaled with multiple ranges of data values corresponding to multiple musical style characteristics. In this way, a user may alter various musical style characteristics included in an audio by moving in various directions, speeds, etc. throughout the grid.

[0117] In some examples, gesture analysis module 420 may process input data indicative of at least one tactile event and / or at least one motion. In some examples, gesture determination module 440 may determine one or more gestures and / or one or more gestural patterns from input data. In some examples, gesture determination module 440 may determine additional information for one or more gestures and / or gestural patterns. In some examples, gesture determination module 440 may apply one or more machine learning models to determine a styleof dance, or one or more of a dance posture, a dance pose, a dance motion, a dance sequence, and the like. In some examples, the machine learning model employed may be trained using data indicative of one or more of a dance posture, a dance pose, a dance motion, a dance sequence, and the like, such as user positions corresponding to a particular style of dance or a particular choreography.

[0118] In some examples, gesture analysis module 420 may be configured to perform gesture prediction. Specifically, in some examples, gesture prediction module 442 may predict (e.g., by employing one or more machine learning techniques) one or more gestures based on captured visual data. In some examples, gesture prediction module 442 may apply a machine learning model to at least one mapped user position to generate at least one predicted user position. For example, gesture prediction module 442 may apply a machine learning model to one or more user positions mapped by grid mapping module 441 and predict one or more user positions indicative of a user’s hand moving in a positive y-direction, a user’s foot moving in a negative x- direction, etc. In some examples, the computing system may generate, based on the at least one mapped user position and the at least one predicted user position, the audio data. In some examples, the machine learning model employed may be configured to identify, based on at least one mapped user position, one or more of the dance posture, the dance pose, the dance motion, the dance sequence, etc. In some examples, the at least one predicted user position is based on one or more of an identified dance posture, an identified dance motion, and an identified dance sequence. As such, in some examples, the new or altered audio data generated by the computing system may be generated based on mapped user positions and / or predicted user positions, which may be based on particular styles of dance identified by the computing system. In some examples, the computing system may output visual data based on mapped user positions and / or predicted user positions, which may be based on particular styles of dance identified by the computing system.

[0119] FIG. 5 is a conceptual diagram illustrating the computing system of FIG. IB configured to map gestures to a grid to dynamically generate new musical expression, in accordance with one or more techniques of this disclosure. Computing system 500, UI module 504, musical expression generation module 508, analysis module 505, audiovisual content generation module 510, network 501, computing device 512, UI components 502, GUI 503, and user 507 may be similar to computing system 100, UI module 104, musical expression generation module 108,analysis module 106, audiovisual content generation module 110, network 101, computing device 112, UI components 102, GUI 103, and user 107 of FIG. IB, respectively. In some examples, analysis module 505 may scale axis 550 of grid 549 with one or more data values included in structured data, in which each of the one or more data values correspond to a respective musical style characteristic from a set of musical style characteristics. For example, data value 547 may correspond to note pitch C, and data value 546 may correspond to note pitch D. As shown in the example of FIG. 5, user 507 has user position 545 corresponding to the user’s head and user position 544 corresponding to the user’s hand. Responsive to one or more of UI components 502 determining at least one gesture, e.g., user 507 moving within grid frame 549 in any direction, computing device 512 may capture (e.g., via one or more of UI components 502, such as a camera) visual data indicative of user position 544 and user position 545 across grid 549 in one or more of an x-direction, y-direction, and z-direction, in which each of user position 544 and user position 545 is associated with one or more of a timestamp and a velocity. Computing system 500 may receive, from computing device 512 via network 501, the visual data.

[0120] In some examples, grid 549 may be part of application GUI 503, in which GUI 503 provides a means for interactive musical and audiovisual play. In some examples, grid 549 may represent a camera frame, and act as a user interface with which users may interact to control various input and output data. In some examples, GUI 503, with or without grid 549, may preview and / or replay audiovisual content output by audiovisual content generation module 510. In some examples, GUI 503 may be a GUI for editing content. In some examples, GUI 503 may provide a user capabilities to edit or customize audio settings, scaling of grid 549, instrument switching, vocal assignments, which points may correspond to various musical style characteristics, sound effects, visual effects, audio signal processing, etc. As such, in general, a user may have the ability to edit output provided by computing system 500, and / or provide additional input for computing system 500 to further edit or update output.

[0121] Analysis module 505 may map user position 545 to at least one axis value, such as data value 547. Analysis module 505 may map user position 544 to at least one axis value, such as data value 546. Based on the mapped user positions, audiovisual content generation module 510 may generate audio data 548 having a set of musical style characteristics corresponding to data values 546 and 547. In some examples, mapped data values 546 and 547 may result in asequence of corresponding note pitches. For example, audio data 548 may include a sequence of note pitches including note pitch C and note pitch D. In some examples, multiple musical style characteristics may be associated with the same timestamp. For example, mapped user position 544 and mapped user position 545 may be associated with a same timestamp, such that mapped data values 546 and 547 may result in simultaneous note pitches, such as simultaneous note pitch C and note pitch D. As such, audio data 548 may include multiple musical style characteristics, e.g., one or more of a frequency spectrum, an amplitude, a timbre, a note pitch, a note interval, a beat, a fill, a riff, a melody contour, a melody motif, a chord progression, a key signature, a musical scale mode, a tonality, a tempo, a meter, a rhythmic pattern, an instrumentation, a density, a polyphony, a volume, a note articulation, a note expression, a structure, a section, a genre, etc. that may be determined by multiple user positions at multiple timepoints that are mapped to multiple data values on multiple axes.

[0122] In some examples, audiovisual content generation module 510 may output audio data, such as audio data 548, based on data processed or determined by analysis module 505. As shown in the example of FIG. 5, computing system 500 may output audio data 548 to computing device 512, in which the audio included in audio data 548 may be played aloud to a user via one or more of UI components 502, such as a speaker. In some examples, computing system 500 may output visual data indicative of at least one mapped user position with at least a portion of the audio data. In some examples, computing system 500 may generate one or more graphical user interfaces, in which the one or more graphical user interfaces are associated with an application (e.g., an application executing at computing device 512). The one or more graphical user interfaces may include the visual data indicative of the at least one mapped user position. As an example, a user may perform a dance including multiple user positions (e.g., user positions 544 and 545) across grid frame 549, in which one or more UI components 503 may record and / or send the visual data associated with the dance to computing system 500 to generate the one or more graphical user interfaces. For example, computing system 500 may generate GUI 503, which may include a recorded video of user 507 performing the dance and / or other visual effects. In some examples, the recorded video and / or visual effects may be overlayed with audio data 548 based on user position timestamps. As such, computing system 500 may sync captured user movement or interaction with audio data to generate audiovisual content.

[0123] FIG. 6 is a conceptual diagram illustrating the computing system of FIG. 1C configured to dynamically generate new musical expression based on musical analysis and visual input, in accordance with one or more techniques of this disclosure. Computing system 600, UI module 604, musical expression module 608 including analysis module 606 and audiovisual content generation module 610, network 601, computing device 612, UI components 602, and GUI 603 may be similar to computing system 100, UI module 104, musical expression generation module 108 including analysis module 106 and audiovisual content generation module 110, network 101, computing device 112, UI components 102, and GUI 103, respectively, of FIG. 1C.

[0124] In the example of FIG. 6, computing system 600 may generate, based on one or more visual data inputs and structured data derived from input audio having a first set of musical style characteristics, one or more visual data outputs 655, in which computing system 600 may output the one or more visual data outputs 655 with at least a portion of the audio data 648 having a second set of musical style characteristics. As shown in the example of FIG. 6, one or more visual data outputs 655 may be a video generated from the still image of FIG. 1C, which may be a still image of a musical performance (e.g., a person singing with a microphone). In some examples, analysis module 606 may process one or more visual data inputs, such as a still image, to determine one or more visual data values, such as one or more of brightness, color, warmth, transparency, pixelation, shading, texture, hue, saturation, a visual object, foreground, background, greyscale, etc. Audiovisual content generation module 610 may generate one or more visual data outputs 655 (e.g., a generated video, etc.) based on analysis module 606 determining one or more associations between the input visual data and input audio data. In some examples, analysis module 606 may apply one or more of the machine learning techniques described above. In some examples, analysis module 606 may apply a visual physics machine learning engine that is trained on sentiment and semantic data (e.g., including musical data as non arbitrary graphics filters, data driven image processing values, etc.) to visualize data.

[0125] In some examples, to determine one or more associations between the input visual data and input audio data, analysis module 606 may apply a semantic rules table. As an example, an application executing at computing device 612 may employ a UI component 602, such as a microphone, in which the microphone may detect audio from the environment of computing device 612. Computing system 600 may receive the detected audio, in which analysis module 606 may parse, classify, translate, etc. the set of musical style characteristics present in the audio.As an example, analysis module 606 may interpret one or more data values for each style characteristic, in which the data values may be interpreted as a proxy for graphic control signal on a weighted and qualitative basis. The “quality”, “significance”, or “dominance” of an audio segment, either detected or sourced (e.g., from a library), may be defined by a number of factors, such as pitch, duration, clear versus distorted signals, harmonic versus rhythmic, warm versus cool (or major versus minor), dense versus spatial (or busy versus airy), gradual versus rapid attack (e.g., bowed violin versus plucked violin), mood, semantic descriptors (uplifting, suspenseful, angry, melancholy, etc.), genre, cultural cues, patterns, themes, signatures, and other style elements from a given musical work, corpus of works, or composer, etc.

[0126] For example, the weight of an audio signal may be a measure of the loudness or amplitude of the audio signal. Analysis module 606 may use these weights when applying graphic filtering or image processing, such that the loudness or amplitude may be equally represented in terms of brightness or transparency. As such, in the example of FIG. 6, the brightness of visual data output 655 may increase or lessen based on the changes in loudness or amplitude in the input audio. In other words, lesser and more distant sounds may produce less bright or more transparent values.

[0127] In another example, analysis module 606 may determine the input audio or other input information (e.g., cultural cues in input visual data, such as flags, icons, etc.) to be associated with a specific culture or location (e.g., a user’s current location, a location associated with an image, video, or audio, etc.). In this example, visual data output 655 may include visual effects determined by analysis module 606 to be associated with the culture or location, e.g., visual data output 655 may include red, white, and blue color combinations for a particular country. In another example, analysis module 606 may determine input audio to be associated with images or videos, or vice versa, based on semantic information and / or other contextual information. For example, a song included in a library may be tagged with semantic information such as genre, descriptions of the song (e.g., “uplifting” or “melancholy”), etc. Various images and / or videos hosted on a social media application, in a public library, etc. may also be tagged with similar semantic information. As such, in these examples, analysis module 606 may associate input audio data with visual data, or vice versa, based on having contextual data or semantic information with a threshold similarity. In some examples, additional weight may be afforded to semantic tags occurring within the duration of a given musical work, in which such tags may bemapped to particular moments and / or sequences, such as to more definitively associate a given semantic value with the specific musical bar or measure in which it occurs. In some examples, analysis module 606 may be configured to collate musical works to images and / or videos, and then determine visual representations for a determined “emotional quality” or descriptive label such as “uplifting”. In general, analysis module 606 may source visual and / or audio data in a variety of ways, e.g., by receiving visual and / or audio data from computing device 612, by retrieving visual and / or audio data from a database included in computing system 600, by determining information about received visual and / or audio data and using the information to run executable search queries across libraries of candidate audio files and / or visual data files, using APIs to access or interact with open source software and services, etc.

[0128] In general, visual data output 655 may be rendered contextually, i.e., in accordance with the rules basis described above, either in real-time or time shifted. In some examples, the nature, presentation, behavior, or elements of visual data output 655 may be iteratively informed by a conduit of audience behavior data and / or individual user session data retrieved from an application that hosts visual data output 655. In some examples, computing system 600 may implement interactive rendering. That is, in some examples, analysis module 606 may be configured to adapt or change the settings of, e.g., vertex and fragment shaders, pixelation, sketch, chromatic aberration, warping, and other visual representation, etc., responsive to user input. For example, visual output data 655, as shown in the example of FIG. 6, may include a video of a person singing in front of a microphone. Additionally, the video may include various visual representations and / or visual effects, such as lighting effects, color effects, etc. In some examples, a user may provide additional input to render visual data output 655 in a different way. For example, the user may want to alter the visual effects included in visual data output 655. In these examples, the user may provide additional input to computing system 600 by, for example, providing input to a prompt, input entry field, etc. associated with an application executing at computing device 612 that may host content such as visual data output 655 and / or provide a means for sharing such content (e.g., a social media application).

[0129] As such, computing system 600 may be configured to perform automated dynamic visualization of images or other content from computing device 612, such as images or content hosted on a social media application executing at computing device 612. In this way, users maybe provided the capability to create unique, engaging, and creative audiovisual content that is audibly and visually harmonic.

[0130] FIG. 7 is a conceptual diagram illustrating a biometric data analysis module configured to perform analysis of biometric data for generating recommendations, in accordance with one or more techniques of this disclosure. As shown in the example of FIG. 7, biometric data analysis module 722 includes data preprocessing module 751, motion detection module 739, neuroactivity detection module 573, heart rate detection module 754, and voice detection module 752. In some examples, biometric data analysis module 722 may include other modules not shown in FIG. 7 for detecting other types of physiological information. Biometric data analysis module 722 may be similar to biometric data analysis module 222 of FIG. 2. In general, with explicit user consent, the biometric data described herein may be received from a user and / or stored by the computing system, e.g., in a user profile. In some examples, one or more of the modules included in biometric data analysis module 722 may be implemented while a user is wearing a computing device.

[0131] In some examples, biometric data analysis module 722 may process biometric data associated with a user. In some examples, the audio and / or visual content output by the computing system may be recommended audio and / or recommended visual content for a user. In some examples, the recommended audiovisual data includes one or more data values indicative of one or more of a dance posture, a dance motion, a dance sequence, note pitches, note intervals, beats, fills, riffs, melody contour, melody motifs, chord progressions, key signatures, musical scale modes, tonality, tempo, meter, rhythmic patterns, instrumentation, density, polyphony, volume, note articulation, note expression, structure, sections, genre, and historical context data.

[0132] Responsive to an input device detecting one or more physiological characteristics from a user when recommended audiovisual content is presented to a user, biometric data analysis module 722 may receive information pertaining to the one or physiological characteristics. Data preprocessing module 751 may be configured to preprocess the information pertaining to the one or physiological characteristics and / or preprocess any data output by one or more of motion detection module 739, neuroactivity detection module 573, heart rate detection module 754, and voice detection module 752. For example, information indicating a user’s heart rate determined by heart rate detection module 754 may be sent to data preprocessing module 751, in which data preprocessing module 751 processes the information and performs steps to transform theinformation into data that can be used by a machine learning model or other components of the computing system.

[0133] Motion detection module 739 may be configured to receive data from motion sensors integrated within a computing device or system, such as accelerometers, gyroscopes, or magnetometers. For example, the sensors may capture changes in motion, orientation, and position of a computing device, which may be used by motion detection module 739 to recognize or determine one or more gestures. In some examples, motion and / or other sensors may be integrated within headphones and headsets. In these examples, for example, motion detection module 739 may receive data indicative of one or more motions for nodding a head. In this example, motion detection module 739 may determine, from the data, one or more gestures, such as a user nodding their head, and / or generate metadata including information about the detected motion. Motion detection module 739 may determine a user satisfaction level (e.g., a user nodding their head may correspond to a high user satisfaction level), and based on the user satisfaction level, assign a score to one or more data values corresponding to the visual and / or musical style characteristics included in the recommended content. In this way, the computing system may tailor recommended content based on the user’s determined preferences. In some examples, motion detection module 739 may compare motion data with the occurrence and characteristics of a user’s historical frequent motions, such as to better determine whether a particular motion or sequence of motions indicates a threshold satisfaction level.

[0134] Neuroactivity detection module 753 may be configured to receive data from neuroactivity sensors integrated within a computing device or system. For example, the neuroactivity sensors may monitor a user’s brain activity in various regions of the brain. Neuroactivity detection module 753 may interpret the sensor data and determine a user satisfaction level (e.g., a level of neuroactivity in a particular region of the brain above a threshold value may correspond to a high user satisfaction level), and based on the user satisfaction level, assign a score to one or more data values corresponding to the visual and / or musical style characteristics included in the recommended content. For example, neuroactivity detection module 753 may associate, based on a time relation, specific compositional sequences, elements, and qualities present in a recommended audio to elevated neurological activity. In some examples, neuroactivity detection module 753 may rank the scores assigned to one or more data values corresponding to the visual and / or musical style characteristics. In some examples, neuroactivity detection module 753 maycompare the detected neuroactivity with the user’s historical neuroactivity, such as to better determine whether particular neuroactivity indicates a threshold satisfaction level.

[0135] Heart rate detection module 754 may be configured to receive data from sensors configured to measure and monitor a user's heart rate, such as optical heart rate sensors or electrodes. These sensors may capture changes in blood flow and heartbeat patterns to accurately determine the user's heart rate. Heart rate detection module 754 may interpret the sensor data and determine a user satisfaction level (e.g., a heart rate above a threshold value may correspond to a low user satisfaction level), and based on the user satisfaction level, assign a score to one or more data values corresponding to the visual and / or musical style characteristics included in the recommended content. In some examples, heart rate detection module 754 may compare the detected heart rate data with the user’s historical heart rate data, such as to better determine whether a particular heart rate indicates a threshold satisfaction level.

[0136] Voice detection module 752 may receive data from input devices configured to detect audio input, such as a microphone. Voice detection module 752 may detect human voice patterns within the captured audio data and / or use voice recognition technology. Voice detection module 752 may interpret the audio data and determine a user satisfaction level (e.g., a user singing along to a song may correspond to a high user satisfaction level), and based on the user satisfaction level, assign a score to one or more data values corresponding to the visual and / or musical style characteristics included in the recommended content. In some examples, voice detection module 752 may compare the detected audio data with the user’s historical recorded audio data, such as to better determine whether a particular audio indicates a threshold satisfaction level.

[0137] As such, in general, based on sensor data indicative of physiological characteristics, biometric data analysis module 722 may determine user satisfaction levels for musical style characteristics and / or visual style characteristics included in recommended content, and update the recommended content accordingly. For example, heart rate detection module 754 may determine that audiovisual content including lower note pitches and darker imagery results in a high heart rate for a user (e.g., by increasing anxiety in the user, etc.). As such, heart rate detection module 754 may assign a lower score to the data values corresponding to the lower note pitches, a higher score to the data values corresponding to the higher note pitches, a lower score to the data values corresponding to dark colors, and a higher score to the data valuescorresponding to lighter colors. As such, the updated recommended audiovisual content generated by the computing system may include higher pitched sounds and lighter imagery. In some examples, the “recommended” audiovisual content may be generated content, e.g., including original or new audiovisual content, and / or existing audiovisual content. In some examples, the recommended audiovisual content may be time shifted or generated in real time. The audiovisual content recommended to a user may be optimized, fine-tuned or personalized (e.g., using feedback loops, reinforcement learning, etc.) for user or general audience enjoyment. Furthermore, the biometric data received from a user may determine preferences for a user profile, such that a user may receive content that is more audibly and visually pleasing to them, or more therapeutic.

[0138] FIG. 8 is a conceptual diagram illustrating the computing system of FIG. ID configured to generate recommendations based on analysis of user generated content, in accordance with one or more techniques of this disclosure. Computing system 800, UI module 804, musical expression module 808 including analysis module 806 and audiovisual content generation module 810, network 801, computing device 812, UI components 802, and GUI 803 may be similar to computing system 100, UI module 104, musical expression generation module 108 including analysis module 106 and audiovisual content generation module 110, network 101, computing device 112, UI components 102, and GUI 103, respectively, of FIG. ID

[0139] In the example of FIG. 8, computing system 800 generates, based on received data indicative of user generated content including audio data having a first set of musical style characteristics and / or visual data having a first set of visual style characteristics, one or more recommendations. As shown in the examples of FIG. 8, the one or more recommendations may include audio data 859 having a second set of musical style characteristics and visual data 857 having a second set of visual style characteristics. In some examples, analysis module 806 may apply a machine learning model to the received data indicative of the user generated content to determine one or more user preferences. For example, analysis module 806 may receive user generated content that includes captured motion (e.g., a user dancing) and / or audio data, in which analysis module 806 may determine dance style, music taste, etc. In some examples, analysis module 806 may receive other information pertaining to a user account, such as data indicative of a user’s engagement or interaction with an application, engagement or interactions with content hosted on the application, engagement or interactions with other user accounts, etc.,and / or information pertaining to the application, such as Key Performance Indicators (KPIs), trending audio, trending visual aesthetics, other trending content, etc. In some examples, analysis module 806 may be configured to query data stored by computing system 800 (e.g., data from computing device 812) to generate one or more recommendations.

[0140] As an example, if user generated content from a first user account includes the user dancing to a particular cultural song, analysis module 806 may determine the specific culture associated with the dance and / or the song, and determine a user preference for content associated with the identified culture. As such, in the example of FIG. 8, recommended visual data 857 and recommended audio 859 may be associated with the identified culture. In this example, recommended visual data 857 may be a video associated with a second user account for user 858, in which visual data 857 includes user 858 dancing to recommended audio 859. Therefore, a first user may receive recommended user generated content from other users, e.g., on GUI 803, which may be a feed for a social media application, that is stylistically similar or has a threshold similarity to the content that the first user has created. In this way, users of the application may discover users creating similar content, such that the users may potentially collaborate on content. In other examples, however, the recommendations may be stylistically different or have a threshold dissimilarity to the content that the user has created, such that the user may be presented with audiovisual content associated with other cultures, musical genres, styles, etc.

[0141] FIG. 9 is a flowchart illustrating an example operation of a computing system for dynamically generating new musical expression based on musical analysis and user input, in accordance with one or more techniques of this disclosure. For clarity purposes, FIG. 9 is described with respect to FIGS. 1A-8.

[0142] Computing system 100 receives first audio data having a first set of musical style characteristics (960). Analysis module 206 applies machine learning module 218 to the first audio data to generate structured data including one or more data values, in which each of the one or more data values correspond to a respective musical style characteristic from the first set of musical style characteristics (961). In some examples, machine learning module 218 includes at least one neural network. In some examples, each of the one or more data values are indicative of one or more of a frequency spectrum, an amplitude, a timbre, a note pitch, a note interval, a beat, a fill, a riff, a melody contour, a melody motif, a chord progression, a key signature, a musical scale mode, a tonality, a tempo, a meter, a rhythmic pattern, an instrumentation, adensity, a polyphony, a volume, a note articulation, a note expression, a structure, a section, a genre, and historical context data. In some examples, the one or more data values are indicative of one or more note pitches, in which each of the one or more notes pitches is statistically weighted.

[0143] Computing system 100 receives at least one input (962). In some examples, the at least one input is indicative of one or more of at least one tactile event 105 and at least one motion 109. In some examples, the at least one input is indicative of visual data associated with one or more of at least one user position, at least one object, at least one image, at least one video, at least one icon, and at least one visual representation. In some examples, computing system 100 receives one or more visual data inputs, such as one or more visual data inputs 111. In some examples, one or more of the first audio data having the first set of musical style characteristics and the at least one input is user generated content. In some examples, computing system 100 receives data indicative of user generated content 113 including one or more of the first audio data having a first set of musical style characteristics and visual data having a set of visual style characteristics. In some examples, computing system 100 outputs recommended visual data having a set of visual style characteristics, and recommended audio data having a set of musical style characteristics. In some examples, responsive to computing system 100 outputting one or more of recommended visual data and recommended audio data, computing system 100 receives, biometric data indicative of one or more physiological characteristics, such as neuroactivity, eye movement, heart rate, respiratory rate, blood pressure, heart rate variability, sleep duration, skin temperature, and blood oxygen.

[0144] Audiovisual content generation module 210 generates, based on the structured data and the at least one input, one or more of visual data having a set of visual style characteristics associated with the first set of musical style characteristics and second audio data having a second set of musical style characteristics associated with the first set of musical style characteristics (963). Computing system 100 outputs one or more of at least a portion of the visual data having the set of visual style characteristics and at least a portion of the second audio data having the second set of musical style characteristics (964). In some examples, based on the structured data and one or more visual data inputs, audiovisual content generation module 210 generates one or more visual data outputs 655. In some examples, computing system 100 outputs one or more visual data outputs 655 with at least a portion of second audio data 648 having thesecond set of musical style characteristics. In some examples, audiovisual content generation unit 210 generates instructions for generating one or more graphical user interfaces, such as GUI 603, including at least the portion of the visual data having the set of visual style characteristics associated with the first set of musical style characteristics.

[0145] In examples in which computing system 100 receives at least one input indicative of at least one tactile event 105, the at least one input may be received in response to the at least one tactile event 105 being detected at a location of a presence-sensitive display that corresponds to at least one portion of GUI 103 associated with an application. In examples in which computing system 100 receives at least one input indicative of at least one motion 109, the at least one input may be received in response to the at least one motion 109 being detected at one or more input devices.

[0146] In some examples, gesture determination module 440 determines, based on the at least one input indicative of at least one tactile event 105 and / or the at least one input indicative of the at least one motion 109, at least one gesture.

[0147] In some examples, scaling module 439 scales an axis of grid 539 with or more data values included in the structured data. In some examples, computing system 100 may receive at least one input indicative of at least one motion 109, which may include captured visual data indicative of one or more user positions across grid 539 in one or more of an x-direction, y- direction, and z-direction. In some examples, each user position from the one or more user positions is associated with one or more of a timestamp and a velocity. In some examples, grid mapping module 441 maps each user position from the one or more user positions to at least one axis value.

[0148] In some examples, the one or more user positions include at least two user positions 544 and 545 associated with the same timestamp, in which grid mapping module 441 maps each of user positions 544 and 545 associated with the same timestamp to at least one axis value. In some examples, each of the at least two mapped user positions associated with the same timestamp correspond to a respective musical style characteristic from the set of musical style characteristics. In some examples, audiovisual content generation module 210 generates, based on at least two mapped user positions associated with the same timestamp, at least a portion of audio data 548. In some examples, gesture prediction module 442 applies a machine learning model to at least one mapped user position to generate at least one predicted user position. Insome examples, computing system 100 trains the machine learning model using data indicative of one or more of a dance posture, a dance pose, a dance motion, and a dance sequence. In some examples, gesture analysis module 420 identifies, based on the at least one mapped user position, one or more of the dance posture, the dance pose, the dance motion, and the dance sequence, and wherein the at least one predicted user position is based on one or more of an identified dance posture, an identified dance motion, and an identified dance sequence. In these examples, audiovisual content generation module 210 generates, based on the at least one mapped user position and the at least one predicted user position, audio data 548.

[0149] In some examples, grid mapping module 441 maps, based on grid 549, the at least one gesture to at least one data value from the one or more data values included in the structured data to generate intermediary audio data having a third set of musical style characteristics associated with the first set of musical style characteristics. In some examples, similarity determination module 334 compares the third set of musical style characteristics with the first set of musical style characteristics to determine a similarity score, and determines whether the similarity score is less than a threshold similarity score. In some examples, responsive to similarity determination module 334 determining the similarity score is less than the threshold similarity score, transformation module 336 applies one or more transformations to the intermediary audio data having the third set of musical style characteristics to generate second audio data 148 having the second set of musical style characteristics. In some examples, transformation module 336 applies one or more range constraints to mapped data values. For example, example transformations may include transposing all note pitches by < ±3 semitones, scaling tempo by a factor between 0.85 and 1.15, substituting timbres from a curated compatibility table, reharmonizing at most two chords per eight-bar segment while preserving key center, etc. In general, the constraints may help ensure that the second set of musical style characteristics maintain a threshold level of similarity.

[0150] In some other examples, responsive to similarity determination module 334 determining the similarity score satisfies the threshold similarity score, computing system 100 outputs, based on the mapping, at least a portion of the intermediary audio data having the third set of musical style characteristics with the instructions, in which the intermediary audio data having the third set of musical style characteristics is the same as the second audio data 148 having the second set of musical style characteristics. In examples in which audiovisual content generation unit 210generates instructions for generating one or more graphical user interfaces including the visual data having the set of visual style characteristics associated with the first set of musical style characteristics, computing system 100 may output, based on the mapping, at least a portion of second audio data 148 having the second set of musical style characteristics with the instructions. In examples in which computing system 100 employs gesture analysis module 420, audiovisual content generation unit 210 may generate, based on at least one mapped user position, audio data 548, in which computing system 100 may output visual data indicative of the at least one mapped user position with at least a portion of audio data 548.

[0151] In examples in which computing system 100 receives one or more visual data inputs, such as one or more visual data inputs 111, one or more of the audio data and the one or more visual inputs may be associated with one or more of stored audio, detected audio, a still image, and a video. In some examples, the structured data includes one or more weighted data values, in which the one or more weighted data values are indicative of one or more statistically weighted note pitches, and visual data analysis module 221 applies an algorithm to the one or more weighted data values to generate one or more visual representations. In some examples, visual data analysis module 221 applies an algorithm including one or more of a semantic rules table and a machine learning model. In some examples, the one or more visual representations correspond to one or more visual data values each indicative of one or more of brightness, color, warmth, transparency, pixelation, shading, texture, hue, saturation, a visual object, foreground, and background.

[0152] In examples in which computing system 100 receives data indicative of user generated content 113 including one or more of first audio data having a first set of musical style characteristics and first visual data having a first set of visual style characteristics, UGC analysis module 223 applies a machine learning model to the data indicative of user generated content 113 to determine one or more user preferences. In some examples, UGC analysis module 223 identifies, based on user generated content 113, one or more of a dance posture, a dance motion, a dance sequence, frequency spectrum, an amplitude, a timbre, a note pitch, a note interval, a beat, a fill, a riff, a melody contour, a melody motif, a chord progression, a key signature, a musical scale mode, a tonality, a tempo, a meter, a rhythmic pattern, an instrumentation, a density, a polyphony, a volume, a note articulation, a note expression, a structure, a section, a genre, historical context data, brightness, color, warmth, transparency, pixelation, shading,texture, hue, saturation, a visual object, foreground, and background. In some examples, audiovisual content generation module 210 generates, based on the one or more user preferences, one or more recommendations including one or more of second audio data 859 having a second set of musical style characteristics and second visual data 857 having a second set of visual style characteristics. In some examples, computing system 100 outputs the one or more recommendations. In some examples, the data indicative of user generated content 113 is associated with a first user account, and the one or more recommendations include user generated content associated with a second user account.

[0153] In examples in which computing system 100 receives biometric data indicative of one or more physiological characteristics, responsive to outputting one or more recommendations, biometric data analysis module 722 determines a user satisfaction level by at least applying a machine learning model to the biometric data. In some examples, the one or more recommendations includes one or more data values indicative of one or more of a dance posture, a dance motion, a dance sequence, note pitches, note intervals, beats, fills, riffs, melody contour, melody motifs, chord progressions, key signatures, musical scale modes, tonality, tempo, meter, rhythmic patterns, instrumentation, density, polyphony, volume, note articulation, note expression, structure, sections, genre, and historical context data. In some examples, biometric data analysis module 722 assigns, based on the user satisfaction level, a score to each data value from the one or more data values. In some examples, audiovisual content generation module 210 generates, based on the assigned scores, one or more updated recommendations. Computing system 100 may then output the one or more updated recommendations.

[0154] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted over, as one or more instructions or code, a computer-readable medium and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media, which corresponds to a tangible medium such as data storage media, or communication media including any medium that facilitates transfer of a computer program from one place to another, e.g., according to a communication protocol. In this manner, computer-readable media generally may correspond to (1) tangible computer- readable storage media, which is non-transitory or (2) a communication medium such as a signal or carrier wave. Data storage media may be any available media that may be accessed by one ormore computers or one or more processors to retrieve instructions, code and / or data structures for implementation of the techniques described in this disclosure. A computer program product may include a computer-readable medium.

[0155] By way of example, and not limitation, such computer-readable storage media may comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other storage medium that may be used to store desired program code in the form of instructions or data structures and that may be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. It should be understood, however, that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but are instead directed to non-transient, tangible storage media. Disk and disc, as used herein, includes compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and Blu-ray disc, where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0156] Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Accordingly, the term “processor,” as used herein may refer to any of the foregoing structures or any other structure suitable for implementation of the techniques described herein. In addition, in some aspects, the functionality described herein may be provided within dedicated hardware and / or software modules. Also, the techniques could be fully implemented in one or more circuits or logic elements.

[0157] The techniques of this disclosure may be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC) or a set of ICs (e.g., a chip set). Various components, modules, or units are described in this disclosure to emphasize functional aspects of devices configured to perform the disclosed techniques, but do notnecessarily require realization by different hardware units. Rather, in some examples, various units may be combined in a hardware unit or provided by a collection of intraoperative hardware units, including one or more processors, in conjunction with suitable software and / or firmware.

[0158] It is to be recognized that, depending on the example, certain acts or events of any of the techniques described herein may be performed in a different sequence, may be added, merged, or left out altogether (e.g., not all described acts or events are necessary for the practice of the techniques). Moreover, in certain examples, acts or events may be performed concurrently, e.g., through multi-threaded processing, interrupt processing, or multiple processors, rather than sequentially.

[0159] In some examples, a computer-readable storage medium comprises a non-transitory medium. The term “non-transitory” indicates that the storage medium is not embodied in a carrier wave or a propagated signal. In certain examples, a non-transitory storage medium may store data that can, over time, change (e.g., in RAM or cache).

[0160] Example 1 : A method includes receiving, by a computing system, first audio data having a first set of musical style characteristics; applying, by the computing system, a machine learning model to the first audio data to generate structured data including one or more data values, wherein each of the one or more data values correspond to a respective musical style characteristic from the first set of musical style characteristics; receiving, by the computing system, at least one input; generating, by the computing system, based on the structured data and the at least one input, one or more of visual data having a set of visual style characteristics associated with the first set of musical style characteristics and second audio data having a second set of musical style characteristics associated with the first set of musical style characteristics; and outputting, by the computing system, one or more of at least a portion of the visual data having the set of visual style characteristics and at least a portion of the second audio data having the second set of musical style characteristics.

[0161] Example 2: The method of example 1, wherein each of the one or more data values are indicative of one or more of a frequency spectrum, an amplitude, a timbre, a note pitch, a note interval, a beat, a fill, a riff, a melody contour, a melody motif, a chord progression, a key signature, a musical scale mode, a tonality, a tempo, a meter, a rhythmic pattern, an instrumentation, a density, a polyphony, a volume, a note articulation, a note expression, a structure, a section, a genre, and historical context data.

[0162] Example 3: The method of example 2, wherein the one or more data values are indicative of one or more note pitches, and wherein each of the one or more notes pitches is statistically weighted.

[0163] Example 4: The method of any of examples 1 through 3, wherein the at least one input is indicative of at least one tactile event, wherein the at least one input is received in response to the at least one tactile event being detected at a location of a presence-sensitive display that corresponds to at least one portion of a graphical user interface associated with an application.

[0164] Example 5: The method of any of examples 1 through 4, wherein the at least one input is indicative of visual data associated with one or more of at least one user position, at least one object, at least one image, at least one video, at least one icon, and at least one visual representation.

[0165] Example 6: The method of any of examples 1-5, wherein the at least one input is indicative of the at least one motion, wherein the at least one input is received in response to the at least one motion being detected at one or more input devices.

[0166] Example 7: The method of any of examples 1-6, the method further comprising: generating, by the computing system, instructions for generating one or more graphical user interfaces including at least the portion of the visual data having the set of visual style characteristics associated with the first set of musical style characteristics.

[0167] Example 8: The method of example 7, further comprising: determining, by the computing system, based on the at least one input, at least one gesture; mapping, by the computing system, based on a grid, the at least one gesture to at least one data value from the one or more data values included in the structured data to generate intermediary audio data having a third set of musical style characteristics associated with the first set of musical style characteristics; comparing, by the computing system, the third set of musical style characteristics with the first set of musical style characteristics to determine a similarity score; determining, by the computing system, whether the similarity score is less than a threshold similarity score; responsive to determining the similarity score is less than the threshold similarity score, applying, by the computing system, one or more transformations to the intermediary audio data having the third set of musical style characteristics to generate the second audio data having the second set of musical style characteristics; and outputting, by the computing system, based on the mapping, at least a portion of the second audio data having the second set of musical style characteristicswith the instructions for generating the one or more graphical user interfaces including at least the portion of the visual data having the set of visual style characteristics associated with the first set of musical style characteristics.

[0168] Example 9: The method of example 8, further comprising: responsive to determining the similarity score satisfies the threshold similarity score, outputting, by the computing system, based on the mapping, at least a portion of the intermediary audio data having the third set of musical style characteristics with the instructions, wherein the intermediary audio data having the third set of musical style characteristics is the same as the second audio data having the second set of musical style characteristics.

[0169] Example 10: The method of example 8, wherein applying the one or more transformations to the intermediary audio data having the third set of musical style characteristics further comprises: applying, by the computing system, one or more range constraints to mapped data values.

[0170] Example 11 : The method of any of examples 1-10, wherein one or more of the first audio data having the first set of musical style characteristics and the at least one input is user generated content.

[0171] Example 12: The method of any of examples 1-11, wherein the visual data having the set of visual style characteristics is recommended visual data, and wherein the second audio data having the second set of musical style characteristics is recommended audio data.

[0172] Example 13: The method of example 12, wherein each visual style characteristic from the set of visual style characteristics and each musical style characteristic from the second set of musical style characteristics corresponds to a respective data value, the method further comprising: responsive to outputting one or more of at least a portion of the recommended visual data and at least a portion of the recommended audio data, receiving, by the computing system, biometric data including data indicative of one or more of neuroactivity, eye movement, heart rate, respiratory rate, blood pressure, heart rate variability, sleep duration, skin temperature, and blood oxygen; assigning, by the computing system, based on the biometric data, a score to each respective data value; determining, by the computing system, based on each assigned score, one or more of recommended visual data having a set of updated visual style characteristics and recommended audio data having an updated set of musical style characteristics; and outputting, by the computing system, one or more of at least a portion of the recommended visual datahaving the set of updated visual style characteristics and at least a portion of the recommended audio data having the updated set of musical style characteristics.

[0173] Example 14: The method of any of examples 1-13, wherein the machine learning model includes at least one neural network.

[0174] Example 15: The method of any of examples 1-14, wherein the first audio data having the first set of musical style characteristics and the at least one input are received from an external computing system, and wherein the method further comprises: sending, by the computing system, and to the external computing system, one or more of at least the portion of the visual data having the set of visual style characteristics and at least the portion of the second audio data having the second set of musical style characteristics.

[0175] Example 16: A computing system includes one or more processors; and one or more storage devices that store instructions, wherein the instructions, when executed by the one or more processors, cause the one or more processors to: receive first audio data having a first set of musical style characteristics; apply a machine learning model to the first audio data to generate structured data including one or more data values, wherein each of the one or more data values correspond to a respective musical style characteristic from the first set of musical style characteristics; receive at least one input; generate, based on the structured data and the at least one input, one or more of visual data having a set of visual style characteristics associated with the first set of musical style characteristics and second audio data having a second set of musical style characteristics associated with the first set of musical style characteristics; and output one or more of at least a portion of the visual data having the set of visual style characteristics and at least a portion of the second audio data having the second set of musical style characteristics.

[0176] Example 17: The computing system of example 16, wherein each of the one or more data values are indicative of one or more of a frequency spectrum, an amplitude, a timbre, a note pitch, a note interval, a beat, a fill, a riff, a melody contour, a melody motif, a chord progression, a key signature, a musical scale mode, a tonality, a tempo, a meter, a rhythmic pattern, an instrumentation, a density, a polyphony, a volume, a note articulation, a note expression, a structure, a section, a genre, and historical context data.

[0177] Example 18: The computing system of example 17, wherein the one or more data values are indicative of one or more note pitches, and wherein each of the one or more notes pitches is statistically weighted.

[0178] Example 19: The computing system of any of examples 16-18, wherein the at least one input is indicative of at least one tactile event, wherein the at least one input is received in response to the at least one tactile event being detected at a location of a presence-sensitive display that corresponds to at least one portion of a graphical user interface associated with an application.

[0179] Example 20: The computing system of any of examples 16-19, wherein the at least one input is indicative of visual data associated with one or more of at least one user position, at least one object, at least one image, at least one video, at least one icon, and at least one visual representation.

[0180] Example 21 : The computing system of any of examples 16-20, wherein the at least one input is indicative of at least one motion, and wherein the at least one input is received in response to the at least one motion being detected at one or more input devices.

[0181] Example 22: The computing system of any of examples 16-21, wherein the instructions further cause the one or more processors to: generate instructions for generating one or more graphical user interfaces including at least the portion of the visual data having the set of visual style characteristics associated with the first set of musical style characteristics.

[0182] Example 23: The computing system of example 22, wherein the instructions further cause the one or more processors to: determine, based on the at least one input, at least one gesture; map, based on a grid, the at least one gesture to at least one data value from the one or more data values included in the structured data to generate intermediary audio data having a third set of musical style characteristics associated with the first set of musical style characteristics; compare the third set of musical style characteristics with the first set of musical style characteristics to determine a similarity score; determine whether the similarity score is less than a threshold similarity score; responsive to determining the similarity score is less than the threshold similarity score, apply one or more transformations to the intermediary audio data having the third set of musical style characteristics to generate the second audio data having the second set of musical style characteristics; and output, based on the mapping, at least a portion of the second audio data having the second set of musical style characteristics with the instructions for generating the one or more graphical user interfaces including at least the portion of the visual data having the set of visual style characteristics associated with the first set of musical style characteristics.

[0183] Example 24: The computing system of example 23, wherein the instructions further cause the one or more processors to: responsive to determining the similarity score satisfies the threshold similarity score, output, based on the mapping, at least a portion of the intermediary audio data having the third set of musical style characteristics with the instructions, wherein the intermediary audio data having the third set of musical style characteristics is the same as the second audio data having the second set of musical style characteristics.

[0184] Example 25: The computing system of example 23, wherein to apply the one or more transformations to the intermediary audio data having the third set of musical style characteristics, the instructions further cause the one or more processors to: apply one or more range constraints to mapped data values.

[0185] Example 26: The computing system of any of examples 16-25, wherein one or more of the first audio data having the first set of musical style characteristics and the at least one input is user generated content.

[0186] Example 27: The computing system of any of examples 16-26, wherein the visual data having the set of visual style characteristics is recommended visual data, and wherein the second audio data having the second set of musical style characteristics is recommended audio data.

[0187] Example 28: The computing system of example 27, wherein each visual style characteristic from the set of visual style characteristics and each musical style characteristic from the second set of musical style characteristics corresponds to a respective data value, and wherein the instructions further cause the one or more processors to: responsive to outputting one or more of at least a portion of the recommended visual data and at least a portion of the recommended audio data, receive biometric data including data indicative of one or more of neuroactivity, eye movement, heart rate, respiratory rate, blood pressure, heart rate variability, sleep duration, skin temperature, and blood oxygen; assign, based on the biometric data, a score to each respective data value; determine, based on each assigned score, one or more of recommended visual data having a set of updated visual style characteristics and recommended audio data having an updated set of musical style characteristics; and output one or more of at least a portion of the recommended visual data having the set of updated visual style characteristics and at least a portion of the recommended audio data having the updated set of musical style characteristics.

[0188] Example 29: The computing system of any of examples 16-28, wherein the machinelearning model includes at least one neural network.

[0189] Example 30: The computing system of any of examples 16-29, wherein the first audio data having the first set of musical style characteristics and the at least one input are received from an external computing system, and wherein the instructions further cause the one or more processors to: send, to the external computing system, one or more of at least the portion of the visual data having the set of visual style characteristics and at least the portion of the second audio data having the second set of musical style characteristics.

[0190] Example 31 : A non-transitory computer-readable storage medium encoded with instructions that, when executed by one or more processors, cause one or more processors to: receive first audio data having a first set of musical style characteristics; apply a machine learning model to the first audio data to generate structured data including one or more data values, wherein each of the one or more data values correspond to a respective musical style characteristic from the first set of musical style characteristics; receive at least one input; generate, based on the structured data and the at least one input, one or more of visual data having a set of visual style characteristics associated with the first set of musical style characteristics and second audio data having a second set of musical style characteristics associated with the first set of musical style characteristics; and output one or more of at least a portion of the visual data having the set of visual style characteristics and at least a portion of the second audio data having the second set of musical style characteristics.

[0191] Example 32: The non-transitory computer-readable storage medium of example 31, wherein each of the one or more data values are indicative of one or more of a frequency spectrum, an amplitude, a timbre, a note pitch, a note interval, a beat, a fill, a riff, a melody contour, a melody motif, a chord progression, a key signature, a musical scale mode, a tonality, a tempo, a meter, a rhythmic pattern, an instrumentation, a density, a polyphony, a volume, a note articulation, a note expression, a structure, a section, a genre, and historical context data.

[0192] Example 33: The non-transitory computer-readable storage medium of example 32, wherein the one or more data values are indicative of one or more note pitches, and wherein each of the one or more notes pitches is statistically weighted.

[0193] Example 34: The non-transitory computer-readable storage medium of any of examples 31-33, wherein the at least one input is indicative of at least one tactile event, wherein the at least one input is received in response to the at least one tactile event being detected at a location of apresence-sensitive display that corresponds to at least one portion of a graphical user interface associated with an application.

[0194] Example 35: The non-transitory computer-readable storage medium of any of examples 31-34, wherein the at least one input is indicative of visual data associated with one or more of at least one user position, at least one object, at least one image, at least one video, at least one icon, and at least one visual representation.

[0195] Example 36: The non-transitory computer-readable storage medium of any of examples 31-35, wherein the at least one input is indicative of at least one motion, and wherein the at least one input is received in response to the at least one motion being detected at one or more input devices.

[0196] Example 37: The non-transitory computer-readable storage medium of any of examples 31-36, wherein the instructions further cause the one or more processors to: generate instructions for generating one or more graphical user interfaces including at least the portion of the visual data having the set of visual style characteristics associated with the first set of musical style characteristics.

[0197] Example 38: The non-transitory computer-readable storage medium of example 37, wherein the instructions further cause the one or more processors to: determine, based on the at least one input, at least one gesture; map, based on a grid, the at least one gesture to at least one data value from the one or more data values included in the structured data to generate intermediary audio data having a third set of musical style characteristics associated with the first set of musical style characteristics; compare the third set of musical style characteristics with the first set of musical style characteristics to determine a similarity score; determine whether the similarity score is less than a threshold similarity score; responsive to determining the similarity score is less than the threshold similarity score, apply one or more transformations to the intermediary audio data having the third set of musical style characteristics to generate the second audio data having the second set of musical style characteristics; and output, based on the mapping, at least a portion of the second audio data having the second set of musical style characteristics with the instructions for generating the one or more graphical user interfaces including at least the portion of the visual data having the set of visual style characteristics associated with the first set of musical style characteristics.

[0198] Example 39: The non-transitory computer-readable storage medium of example 38,wherein the instructions further cause the one or more processors to: responsive to determining the similarity score satisfies the threshold similarity score, output, based on the mapping, at least a portion of the intermediary audio data having the third set of musical style characteristics with the instructions, wherein the intermediary audio data having the third set of musical style characteristics is the same as the second audio data having the second set of musical style characteristics.

[0199] Example 40: The non-transitory computer-readable storage medium of example 38, wherein to apply the one or more transformations to the intermediary audio data having the third set of musical style characteristics, the instructions further cause the one or more processors to: apply one or more range constraints to mapped data values.

[0200] Example 41 : The non-transitory computer-readable storage medium of any of examples 31-40, wherein one or more of the first audio data having the first set of musical style characteristics and the at least one input is user generated content

[0201] Example 42: The non-transitory computer-readable storage medium of any of examples 31-41, wherein the visual data having the set of visual style characteristics is recommended visual data, and wherein the second audio data having the second set of musical style characteristics is recommended audio data.

[0202] Example 43 : The non-transitory computer-readable storage medium of example 42, wherein each visual style characteristic from the set of visual style characteristics and each musical style characteristic from the second set of musical style characteristics corresponds to a respective data value, and wherein the instructions further cause the one or more processors to: responsive to outputting one or more of at least a portion of the recommended visual data and at least a portion of the recommended audio data, receive biometric data including data indicative of one or more of neuroactivity, eye movement, heart rate, respiratory rate, blood pressure, heart rate variability, sleep duration, skin temperature, and blood oxygen; assign, based on the biometric data, a score to each respective data value; determine, based on each assigned score, one or more of recommended visual data having a set of updated visual style characteristics and recommended audio data having an updated set of musical style characteristics; and output one or more of at least a portion of the recommended visual data having the set of updated visual style characteristics and at least a portion of the recommended audio data having the updated set of musical style characteristics.

[0203] Example 44: The non-transitory computer-readable storage medium of any of examples 31-43, wherein the machine learning model includes at least one neural network.

[0204] Example 45: The non-transitory computer-readable storage medium of any of examples 31-44, wherein the first audio data having the first set of musical style characteristics and the at least one input are received from an external computing system, and wherein the instructions further cause the one or more processors to: send, to the external computing system, one or more of at least the portion of the visual data having the set of visual style characteristics and at least the portion of the second audio data having the second set of musical style characteristics.

[0205] Example 46: A computing device comprising means for performing any combination of the methods of examples 1-15.

[0206] Various examples have been described. These and other examples are within the scope of the following claims.

Claims

WHAT IS CLAIMED IS:

1. A method comprising: receiving, by a computing system, first audio data having a first set of musical style characteristics; applying, by the computing system, a machine learning model to the first audio data to generate structured data including one or more data values, wherein each of the one or more data values correspond to a respective musical style characteristic from the first set of musical style characteristics; receiving, by the computing system, at least one input; generating, by the computing system, based on the structured data and the at least one input, one or more of visual data having a set of visual style characteristics associated with the first set of musical style characteristics and second audio data having a second set of musical style characteristics associated with the first set of musical style characteristics; and outputting, by the computing system, one or more of at least a portion of the visual data having the set of visual style characteristics and at least a portion of the second audio data having the second set of musical style characteristics.

2. The method of claim 1, wherein each of the one or more data values are indicative of one or more of a frequency spectrum, an amplitude, a timbre, a note pitch, a note interval, a beat, a fill, a riff, a melody contour, a melody motif, a chord progression, a key signature, a musical scale mode, a tonality, a tempo, a meter, a rhythmic pattern, an instrumentation, a density, a polyphony, a volume, a note articulation, a note expression, a structure, a section, a genre, and historical context data.

3. The method of claim 2, wherein the one or more data values are indicative of one or more note pitches, and wherein each of the one or more notes pitches is statistically weighted.

4. The method of any of claims 1 through 3, wherein the at least one input is indicative of visual data associated with one or more of at least one user position, at least one object, at least one image, at least one video, at least one icon, at least one visual representation, and at least onemotion detected at one or more input devices.

5. The method of any of claims 1 through 4, wherein the first audio data having the first set of musical style characteristics and the at least one input are received from an external computing system, and wherein the method further comprises: sending, by the computing system, and to the external computing system, one or more of at least the portion of the visual data having the set of visual style characteristics and at least the portion of the second audio data having the second set of musical style characteristics.

6. The method of any of claims 1 through 5, the method further comprising: generating, by the computing system, instructions for generating one or more graphical user interfaces including at least the portion of the visual data having the set of visual style characteristics associated with the first set of musical style characteristics.

7. The method of claim 6, further comprising: determining, by the computing system, based on the at least one input, at least one gesture; mapping, by the computing system, based on a grid, the at least one gesture to at least one data value from the one or more data values included in the structured data to generate intermediary audio data having a third set of musical style characteristics associated with the first set of musical style characteristics; comparing, by the computing system, the third set of musical style characteristics with the first set of musical style characteristics to determine a similarity score; determining, by the computing system, whether the similarity score is less than a threshold similarity score; responsive to determining the similarity score is less than the threshold similarity score, applying, by the computing system, one or more transformations to the intermediary audio data having the third set of musical style characteristics to generate the second audio data having the second set of musical style characteristics; and outputting, by the computing system, based on the mapping, at least a portion of the second audio data having the second set of musical style characteristics with the instructions forgenerating the one or more graphical user interfaces including at least the portion of the visual data having the set of visual style characteristics associated with the first set of musical style characteristics.

8. The method of claim 7, further comprising: responsive to determining the similarity score satisfies the threshold similarity score, outputting, by the computing system, based on the mapping, at least a portion of the intermediary audio data having the third set of musical style characteristics with the instructions, wherein the intermediary audio data having the third set of musical style characteristics is the same as the second audio data having the second set of musical style characteristics.

9. The method of claim 7, wherein applying the one or more transformations to the intermediary audio data having the third set of musical style characteristics further comprises: applying, by the computing system, one or more range constraints to mapped data values.

10. The method of any of claims 1 through 9, wherein one or more of the first audio data having the first set of musical style characteristics and the at least one input is user generated content.

11. The method of any of claims 1 through 10, wherein the visual data having the set of visual style characteristics is recommended visual data, and wherein the second audio data having the second set of musical style characteristics is recommended audio data.

12. The method of claim 11, wherein each visual style characteristic from the set of visual style characteristics and each musical style characteristic from the second set of musical style characteristics corresponds to a respective data value, the method further comprising: responsive to outputting one or more of at least a portion of the recommended visual data and at least a portion of the recommended audio data, receiving, by the computing system, biometric data including data indicative of one or more of neuroactivity, eye movement, heart rate, respiratory rate, blood pressure, heart rate variability, sleep duration, skin temperature, and blood oxygen; assigning, by the computing system, based on the biometric data, a score to eachrespective data value; determining, by the computing system, based on each assigned score, one or more of recommended visual data having a set of updated visual style characteristics and recommended audio data having an updated set of musical style characteristics; and outputting, by the computing system, one or more of at least a portion of the recommended visual data having the set of updated visual style characteristics and at least a portion of the recommended audio data having the updated set of musical style characteristics.

13. A computing system comprising: one or more processors; and one or more storage devices that store instructions, wherein the instructions, when executed by the one or more processors, cause the one or more processors to: receive first audio data having a first set of musical style characteristics; apply a machine learning model to the first audio data to generate structured data including one or more data values, wherein each of the one or more data values correspond to a respective musical style characteristic from the first set of musical style characteristics; receive at least one input; generate, based on the structured data and the at least one input, one or more of visual data having a set of visual style characteristics associated with the first set of musical style characteristics and second audio data having a second set of musical style characteristics associated with the first set of musical style characteristics; and output one or more of at least a portion of the visual data having the set of visual style characteristics and at least a portion of the second audio data having the second set of musical style characteristics.

14. A non-transitory computer-readable storage medium encoded with instructions that, when executed by one or more processors, cause one or more processors to: receive first audio data having a first set of musical style characteristics; apply a machine learning model to the first audio data to generate structured data including one or more data values, wherein each of the one or more data values correspond to a respective musical style characteristic from the first set of musical style characteristics;receive at least one input; generate, based on the structured data and the at least one input, one or more of visual data having a set of visual style characteristics associated with the first set of musical style characteristics and second audio data having a second set of musical style characteristics associated with the first set of musical style characteristics; and output one or more of at least a portion of the visual data having the set of visual style characteristics and at least a portion of the second audio data having the second set of musical style characteristics.

15. A computing device comprising means for performing any combination of the methods of claims 1-12.

Citation Information

Patent Citations

  • System And Method Generating Synchronized Reactive Video Stream From Auditory Input

    US20210390937A1

  • Systems and methods for an immersive audio experience

    US20220343923A1

  • Comparison training for music generator

    WO2022040410A1

  • Multimedia music creation using visual input

    WO2022221716A1