Methods, electronic devices, and machine-readable media for training a machine learning system

By training a neural network system with adversarial components, meaningful music features are learned, which solves the problem of insufficient music feature learning in existing technologies, improves the performance of music recognition, generation and retrieval, and is suitable for multi-task applications.

CN113939870BActive Publication Date: 2026-03-27SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-06-01
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively learn and utilize the latent space of musical features, leading to a decline in the performance of music recognition, generation, and retrieval tasks, especially in interactive applications where they fail to meet human expectations for response.

Method used

By training a machine learning system with first and second neural network components, adversarial components are applied to impose feature distributions in the latent space, and backpropagation optimization techniques are combined to learn meaningful musical features, thereby constraining the distribution of the latent space using contextual constraints.

Benefits of technology

It enables the effective embedding of musical features in the latent space, improves the performance of music recognition, generation and retrieval, meets the requirements of human perception, and is suitable for multi-task applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113939870B_ABST
    Figure CN113939870B_ABST
Patent Text Reader

Abstract

A method includes receiving non-verbal input associated with input music content. The method also includes identifying one or more embeddings based on the input music content using a model that embeds a plurality of music features describing different music content and relationships between the different music content in a latent space. The method further includes at least one of (i) identifying stored music content based on the one or more identified embeddings or (ii) generating derived music content based on the one or more identified embeddings. In addition, the method includes presenting at least one of the stored music content or the derived music content. The model is generated by training a machine learning system that includes one or more first neural network components and one or more second neural network components such that the embeddings of the music features in the latent space have a predefined distribution.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates generally to machine learning systems. More specifically, the present disclosure relates to techniques for learning effective musical features for generation and retrieval based applications. BACKGROUND

[0002] Music is inherently complex, and a single musical theme or style can often be described along multiple dimensions. Some dimensions can broadly describe music and capture attributes that provide a more aggregated representation of the music. These dimensions can include musical features such as tonality, note density, complexity, and instruments. Other dimensions can describe music by considering sequential properties and temporal aspects of the music. These dimensions can include musical features such as cut time, harmonic progression, pitch contour, and repetition.

[0003] In recent years, neural networks have been used to learn a low-dimensional latent “music space” that encapsulates these types of musical features. Different musical passages can be associated with or represented by different embeddings in the space, such as different vectors in the space. The distance between two embeddings in the space can be used as a measure of similarity between two musical passages. Musical passages that are more similar to each other can be represented by embeddings that are separated by a smaller distance. Musical passages that are less similar to each other can be represented by embeddings that are separated by a larger distance. SUMMARY

[0004]

TECHNICAL SOLUTION

[0005] The present disclosure provides techniques for learning effective musical features for generation and retrieval based applications.

[0006] In a first embodiment, a method includes receiving non-verbal input associated with input musical content. The method also includes identifying one or more embeddings based on the input musical content using a model that embeds a plurality of musical features describing different musical content and relationships between different musical content in a latent space. The method further includes at least one of: (i) identifying stored musical content based on the one or more identified embeddings, or (ii) generating derived musical content based on the one or more identified embeddings. In addition, the method includes presenting at least one of the stored musical content or the derived musical content. The model is generated by training a machine learning system having one or more first neural network components and one or more second neural network components such that the embeddings of the musical features in the latent space have a predefined distribution.

[0007] In a second embodiment, an electronic device includes at least one memory, at least one speaker, and at least one processor operatively coupled to the at least one memory and the at least one speaker. The at least one processor is configured to receive non-verbal input associated with input music content. The at least one processor is also configured to identify one or more embeddings based on the input music content using a model that embeds a plurality of music features that describe different music content and relationships between the different music content in a latent space. The at least one processor is also configured to perform at least one of: (i) identify stored music content based on the one or more identified embeddings, or (ii) generate derived music content based on the one or more identified embeddings. In addition, the at least one processor is configured to present at least one of the stored music content or the derived music content via the at least one speaker. The model is generated by training a machine learning system having one or more first neural network components and one or more second neural network components such that the embeddings of the music features in the latent space have a predefined distribution.

[0008] In a third embodiment, a non-transitory machine-readable medium contains instructions that, when executed, cause at least one processor of an electronic device to receive non-verbal input associated with input music content. The medium also contains instructions that, when executed, cause the at least one processor to identify one or more embeddings based on the input music content using a model that embeds a plurality of music features that describe different music content and relationships between the different music content in a latent space. The medium also contains instructions that, when executed, cause the at least one processor to perform at least one of: (i) identify stored music content based on the one or more identified embeddings, or (ii) generate derived music content based on the one or more identified embeddings. In addition, the medium contains instructions that, when executed, cause the at least one processor to present at least one of the stored music content or the derived music content. The model is generated by training a machine learning system having one or more first neural network components and one or more second neural network components such that the embeddings of the music features in the latent space have a predefined distribution.

[0009] In a fourth embodiment, a method includes receiving reference music content, positive music content similar to the reference music content, and negative music content dissimilar to the reference music content. The method also includes generating a model that embeds a plurality of music features that describe the reference music content, the positive music content, and the negative music content and relationships between the reference music content, the positive music content, and the negative music content in a latent space. Generating the model includes training a machine learning system having one or more first neural network components and one or more second neural network components such that the embeddings of the music features in the latent space have a predefined distribution.

[0010] In another embodiment, the training of the machine learning system is performed such that: the one or more first neural network components receive the reference music content, the positive music content, and the negative music content, and are trained to generate the embeddings in the latent space that minimize a specified loss function; and the one or more second neural network components adversarially train the one or more first neural network components to generate the embeddings in the latent space having the predetermined distribution.

[0011] In another embodiment, the machine learning system is trained in two iterative stages; a first iterative stage trains the one or more first neural network components; and a second iterative stage adversarially trains the one or more first neural network components using the one or more second neural network components.

[0012] In another embodiment, the predetermined distribution is Gaussian and continuous.

[0013] In another embodiment, the machine learning system is trained such that: a classifier comprising the adjacency discriminator is trained to classify the embeddings of the reference music content generated by the one or more first neural network components as (i) related to the embeddings of the positive music content generated by the one or more first neural network components and (ii) unrelated to the embeddings of the negative music content generated by the one or more first neural network components; and the one or more second neural network components adversarially train the one or more first neural network components to generate the embeddings in the latent space having the predetermined distribution.

[0014] In another embodiment, the classifier is trained to determine whether two music contents in a musical piece are contiguous.

[0015] In another embodiment, the machine learning system is trained in two iterative stages; a first iterative stage trains the classifier to distinguish between related and unrelated music contents, wherein both the one or more first neural network components and the adjacency discriminator are updated; and a second iterative stage adversarially trains the one or more first neural network components using the one or more second neural network components.

[0016] In another embodiment, during the training: the classifier concatenates the embeddings of the reference music content and the embeddings of the positive music content to produce a first combined embedding; the classifier concatenates the embeddings of the reference music content and the embeddings of the negative music content to produce a second combined embedding; and the classifier is trained to classify the related and unrelated music contents using the first and second combined embeddings.

[0017] In another embodiment, the one or more first neural network components comprise a convolutional neural network; and the one or more second neural network components comprise an adversarial discriminator.

[0018] Other technical features can be readily apparent to one skilled in the art from the following figures, descriptions, and claims.

[0019] Before undertaking a detailed description of the embodiments, a thorough overview of some of the terminology used throughout this patent document will be presented. The terms "transmit," "receive," and "communicate," as well as derivatives thereof, encompass both direct and indirect communication. The terms "include" and "comprise," as well as derivatives thereof, mean inclusion without limitation. The term "or" is inclusive, meaning and / or. The phrase "associated with," as well as derivatives thereof, means to include, be included within, interconnect with, contain, be contained within, connect to or with, couple to or with, be communicable with, cooperate with, interleave, be proximate to, be bound to or with, have a property of, have relations with, or have

[0020] Moreover, various functions described below can be implemented or supported by one or more computer programs, each of which is formed from computer readable program code and embodied in a computer readable medium. The terms "application" and "program" refer to one or more computer programs, software components, sets of instructions, procedures, functions, objects, classes, instances, related data, or a portion thereof that are adapted for implementation in a suitable computer readable program code. The phrase "computer readable program code" includes any type of computer code, including source code, object code, and executable code. The phrase "computer readable medium" includes any type of medium capable of being accessed by a computer, such as read only memory (ROM), random access memory (RAM), a hard disk drive, a compact disc (CD), a digital video disc (DVD), or any other type of memory. A "non-transitory" computer readable medium excludes wired, wireless, optical, or other communication links. The non-transitory computer readable medium includes media where the data is permanent and does not change. The non-transitory computer readable medium includes media where the data is removable and can change, such as rewritable optical discs or erasable memory components.

[0021] As used herein, terms and phrases such as “have,” “has,” “can,” “having,” “include,” “including,” or “may” followed by a listing of items (such as digital, functions, operations, or components) are intended to mean that the item(s) are present, and there can be additional item(s) not expressly listed. Also, as used herein, the phrase “one or more of’ A or B, “at least one of’ A and / or B, or “one or more of’ A and B, can refer to all possible combinations of A and B. For example, “A or B,” “at least one of A and B,” and “one or more of A or B” can refer to all of the following: (1) including at least one A, (2) including at least one B, or (3) including at least one A and at least one B. Also, as used herein, the terms “first,” “second,” and the like, can modify various components, regardless of importance or significance of the components. These terms are merely used to distinguish a different component from another. For example, a first user device and a second user device can indicate different user devices than each other, regardless of the order or the importance of the devices. A first component can be denoted as a second component, and vice versa, without departing from the scope of the disclosure.

[0022] It should be understood that when an element (for example, a first element) is referred to as being “coupled / connected with / to” or “connected with / to” another element (for example, a second element), it can be directly coupled / connected with / to the other element or be indirectly coupled / connected with / to the other element through a third element. In contrast, it should be understood that when an element (for example, a first element) is referred to as being “directly coupled / connected with / to” or “directly connected / coupled with / to” another element (for example, a second element), there are no other elements (for example, a third element) between them.

[0023] As used herein, the phrase “configured (or set) to” can be interchangeably used with phrases “adapted to,” “capable of,” “designed to,” “suitable for,” “made to,” or “to be able to” depending on circumstances. The phrase “configured (or set) to” does not literally mean “specifically designed in hardware to” in nature. Instead, the phrase “configured (or set) to” can mean that a device can perform an operation with another device or component. For example, the phrase “a processor configured (or set) to perform A, B, and C” can mean a general-purpose processor (for example, a CPU or an application processor) that can perform the operations by executing one or more software programs stored in a memory device, or a dedicated processor (for example, an embedded processor) for performing the operations.

[0024] The terminology and phraseology used here is solely used for the purpose of describing some embodiments of the present disclosure and is not intended to limit the scope of other embodiments of the present disclosure. It should be understood that the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. All terminology and phrases used here, including technical and scientific terminology and phrases, have the same meaning as that ordinarily understood by a person of ordinary skill in the art to which embodiments of the present disclosure belong. It should also be understood that terms and phrases, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and should not be interpreted in an idealized or overly formal sense unless expressly so defined herein. In some cases, terms and phrases defined herein can be interpreted to exclude embodiments of the present disclosure.

[0025] Examples of the "electronic device" according to an embodiment of the present disclosure can include at least one of a smartphone, a tablet personal computer (PC), a mobile phone, a video phone, an e-book reader, a desktop PC, a laptop computer, a netbook computer, a workstation, a personal digital assistant (PDA), a portable multimedia player (PMP), an MP3 player, a mobile medical device, a camera, or a wearable device (e.g., smart glasses, a head-mounted device (HMD), electronic clothes, an electronic bracelet, an electronic necklace, an electronic appcessory, an electronic tattoo, a smart mirror, or a smart watch). Other examples of the electronic device include a smart home appliance. Examples of the smart home appliance can include at least one of a television, a digital video disc (DVD) player, an audio player, a refrigerator, an air conditioner, a cleaner, an oven, a microwave oven, a washing machine, a dryer, an air cleaner, a set-top box, a home automation control panel, a security control panel, a TV box (e.g., SAMSUNG HOMESYNC, APPLE TV, or GOOGLE TV), a smart speaker or a speaker with an integrated digital assistant (e.g., SAMSUNG GALAXY HOME, APPLE HOMEPOD, or AMAZON ECHO), a game console (e.g., XBOX, PLAYSTATION, or NINTENDO), an electronic dictionary, an electronic key, a camcorder, or an electronic frame. Other examples of the electronic device include at least one of various medical devices (e.g., various portable medical measuring devices (e.g., a blood glucose measuring device, a heart rate measuring device, or a body temperature measuring device), a magnetic resource angiography (MRA) device, a magnetic resource imaging (MRI) device, a computed tomography (CT) device, an imaging device, or an ultrasonic device), a navigation device, a global positioning system (GPS) receiver, an event data recorder (EDR), a flight data recorder (FDR), a car infotainment device, marine electronic equipment (e.g., marine navigation equipment or a gyro compass), avionics, security equipment, a vehicle head unit, an industrial or home robot, an automatic teller machine (ATM), a point of sale (POS) device, or an Internet of Things (IoT) device (e.g., a light bulb, various sensors, a gas or water meter, a sprinkler, a fire alarm, a thermostat, a street light, a toaster, a fitness device, a hot water tank, a heater, or a boiler). Other examples of the electronic device include at least one of furniture or a building / structure having an electronic board, an electronic signature receiving device, a projector, or various measuring devices (e.g., a device for measuring water, electricity, gas, or electromagnetic waves). Note that the electronic device according to various embodiments of the present disclosure can be one or a combination of the above-described devices. According to some embodiments of the present disclosure, the electronic device can be a flexible electronic device. The electronic device disclosed herein is not limited to the above-listed devices, and can include new electronic devices according to technological development.

[0026] In the following description, in accordance with various embodiments of the present disclosure, an electronic device is described with reference to the accompanying drawings. As used herein, the term "user" can represent a person or another device (e.g., an artificial intelligence electronic device) using the electronic device.

[0027] Throughout this patent document, other specific wordings and phrases can be provided definitions. Those of ordinary skill in the art will understand that, in many, if not most instances, such definitions apply to both preceding and like-performing uses of a word and phrase so defined.

[0028] No description in the present application should be interpreted as implying any particular element, step, or function is an essential element that must be included in the claim scope. The scope of the patent subject matter is defined only by the claims. The use of any other term in the claims, including but not limited to "mechanism," "module," "device," "unit," "component," "element," "member," "means," "machine," "system," "processor," or "controller," is understood by the Applicant to refer to a structure known to one of ordinary skill in the relevant art. BRIEF DESCRIPTION OF DRAWINGS

[0029] For a more complete understanding of the present disclosure and its advantages, reference is now made to the following description taken in conjunction with the accompanying drawings in which like reference numerals represent like parts:

[0030] Figure 1 An example network configuration including an electronic device according to the present disclosure is shown;

[0031] Figure 2 A first example system for machine learning to identify effective music features according to the present disclosure is shown;

[0032] Figure 3 A second example system for machine learning to identify effective music features according to the present disclosure is shown;

[0033] Figure 4 And Figure 5 A first example application of machine learning trained to identify effective music features according to the present disclosure is shown;

[0034] Figure 6 And 7 A second example application of machine learning trained to identify effective music features according to the present disclosure is shown;

[0035] Figure 8 , 9 And 10 a third example application of machine learning trained to identify effective music features according to the present disclosure is shown;

[0036] Figure 11 And 12An example method for machine learning to identify effective music features according to the present disclosure is shown; and

[0037] Figure 13 An example method using machine learning trained to identify effective music features according to the present disclosure is shown. DETAILED DESCRIPTION

[0038] The discussion below Figures 1 to 13 Various embodiments of the present disclosure are described with reference to the accompanying drawings. However, it should be understood that the present disclosure is not limited to these embodiments, and all variations and / or equivalents or alternatives thereof also fall within the scope of the present disclosure. Throughout the specification and drawings, identical or similar reference signs can be used to refer to identical or similar elements.

[0039] As mentioned above, music is inherently complex and can generally be described along multiple dimensions. Some dimensions can broadly describe music and capture properties that provide a more aggregated representation of the music, such as tonality, note density, complexity, and instruments. Other dimensions can describe music by considering sequential properties and temporal aspects of the music, such as cut-up, harmonic progression, pitch contour, and repetition. Neural networks have been used to learn a low-dimensional latent “music space” that encapsulates such music features, where different passages of music can be associated with or represented by different vectors or other embeddings in the space. The distance between two embeddings in the space can be used as a measure of similarity between two passages of music. This means, at a high level, that embeddings of similar music content should be geometrically closer in the latent space than different music content.

[0040] For certain tasks, it can be important to effectively learn such a latent space in order to help ensure that music is recognized, selected, generated, or used in a way that is consistent with human expectations. This is particularly true in interactive applications where the response of the machine is often dependent on a human operator. Thus, effective embeddings of music features can be used to interpret music in a way that is relevant to human perception. This type of embedding captures features that are useful for downstream tasks and are consistent with a distribution that is suitable for sampling and meaningful insertion. Unfortunately, learning useful music features is often at the expense of being able to effectively generate or decode from the learned music features (and vice versa).

[0041] The present disclosure provides techniques for learning effective musical features for generation and retrieval based applications. These techniques have the ability to learn meaningful musical features and conform to the useful distribution of those musical features embedded in a latent musical space. These techniques exploit context while simultaneously imposing a shape on the feature distribution in the latent space, such as by backpropagation, using an adversarial component. This allows for the joint optimization of the desired features by exploiting context, which improves the features, and constraining the distribution in the latent space, which makes it possible to generate samples. Instead of explicitly labeled data, neural network components or other machine learning algorithms can be trained under the assumption that two adjacent units of musical content, such as two adjacent sections or parts in the same musical piece, are related. In other words, the distance between the embeddings of two adjacent units of the same musical content in the latent space should be smaller than the distance between the embeddings of two randomly uncorrelated units of musical content in the latent space.

[0042] where the musical content can be analyzed and its features projected into a continuous low-dimensional space that is relevant to a human listener while maintaining a desired distribution in the low-dimensional space. Thus, these techniques can be used to effectively learn a feature space by embedding a large number of relevant musical features into the feature space. Furthermore, these approaches allow for training a single machine learning model and using it for a variety of downstream tasks. It is often the case that a unique model is trained for each particular task, as performance between tasks typically degrades when a single model is trained for multiple tasks. Each model trained using the techniques described in the present disclosure can be used to perform a variety of functions, such as searching for particular musical content based on an auditory nonverbal input, ranking musical content that most closely resembles an auditory nonverbal input, selecting particular musical content for playback based on an auditory nonverbal input, and autonomously generating music based on an auditory nonverbal input. Furthermore, the described techniques can jointly optimize multiple loss functions for embedding context, self-reconstruction, and constraining the distribution (such as by using backpropagation). As a single model can be used for multiple tasks, multiple loss functions can be simultaneously optimized and the feature distribution in the latent space can be constrained such that the distribution conforms to a particular subspace. This distribution allows for the learning of effective features using an additional loss function that exploits context. Furthermore, the trained machine learning model can be used to achieve improved performance in one or more downstream tasks.

[0043] Figure 1 An example network configuration 100 including an electronic device according to the present disclosure is shown. Figure 1 The illustrated embodiment of the network configuration 100 is for illustration only. Other embodiments of the network configuration 100 can be used without departing from the scope of the present disclosure.

[0044] According to an embodiment of the disclosure, the electronic device 101 is included in the network configuration 100. The electronic device 101 can include at least one of a bus 110, a processor 120, a memory 130, an input / output (I / O) interface 150, a display 160, a communication interface 170, a sensor 180, or a speaker 190. In some embodiments, the electronic device 101 can exclude at least one of the components, or can add at least one other component. The bus 110 includes a circuit for connecting the components 120-190 to each other and for transmitting communications (e.g., control messages and / or data) between the components.

[0045] The processor 120 includes one or more of a central processing unit (CPU), a graphics processing unit (GPU), an application processor (AP), or a communication processor (CP). The processor 120 can perform control over at least one other component of the electronic device 101 and / or perform operations or data processing related to communication. For example, the processor 120 can be used for training in order to learn effective music features, e.g., by embedding a large number of different music contents into a latent space of a desired distribution. The processor 120 can also or instead use a trained machine learning model for one or more generation and retrieval based applications, such as searching, ranking, playing, or generating music content.

[0046] The memory 130 can include a volatile and / or nonvolatile memory. For example, the memory 130 can store commands or data related to at least one other component of the electronic device 101. According to an embodiment of the disclosure, the memory 130 can store software and / or a program 140. The program 140 includes, for example, a kernel 141, middleware 143, an application programming interface (API) 145, and / or an application program (or "application") 147. At least a portion of the kernel 141, the middleware 143, or the application programming interface 145 can be denoted as an operating system (OS).

[0047] The kernel 141 can control or manage system resources (e.g., the bus 110, the processor 120, or the memory 130) used to execute operations or functions implemented in other programs (e.g., the middleware 143, the API 145, or the application programs 147). The kernel 141 provides an interface that allows the middleware 143, the API 145, or the application programs 147 to access the individual components of the electronic device 101 to control or manage the system resources. The application programs 147 can include one or more application programs for machine learning and / or training machine learning models, as described below. These functions can be performed by a single application or by a plurality of applications, each of which performs one or more of these functions. The middleware 143 can serve as a relay to allow, for example, the API 145 or the application 147 to communicate data with the kernel 141. A plurality of application programs 147 can be provided. The middleware 143 is capable of controlling work requests received from the application programs 147, for example, by assigning priorities for using the system resources (like the bus 110, the processor 120, or the memory 130) of the electronic device 101 to at least one of the plurality of application programs 147. The API 145 is an interface that allows the application programs 147 to control functions provided from the kernel 141 or the middleware 143. For example, the API 145 includes at least one interface or function (e.g., a command) for file control, window control, image processing, or text control.

[0048] The I / O interface 150 serves as an interface that can, for example, transmit a command or data input from a user or other external devices to other components of the electronic device 101. The I / O interface 150 can also output, to the user or other external devices, a command or data received from other components of the electronic device 101.

[0049] The display 160 includes, for example, a liquid crystal display (LCD), a light emitting diode (LED) display, an organic light emitting diode (OLED) display, a quantum dot light emitting diode (QLED) display, a microelectromechanical systems (MEMS) display, or an electronic paper display. The display 160 can also be a depth perception display, such as a multi-focal display. The display 160 is capable of displaying, for example, various contents (e.g., text, images, videos, icons, or symbols) to a user. The display 160 can include a touch screen and can receive, for example, touch, gesture, proximity, or hovering input using an electronic pen or a user's body part.

[0050] The communication interface 170 is, for example, capable of establishing communication between the electronic device 101 and external electronic devices (e.g., the first electronic device 102, the second electronic device 104, or the server 106). For example, the communication interface 170 can be connected with the network 162 or 164 through wireless or wired communication to communicate with external electronic devices. The communication interface 170 can be a wired or wireless transceiver or any other component for transmitting and receiving signals.

[0051] The electronic device 101 also includes one or more sensors 180, which can measure a physical quantity or detect an activation state of the electronic device 101 and convert the measured or detected information into an electrical signal. For example, the one or more sensors 180 can include one or more microphones, which can be used to capture non-verbal auditory input (e.g., a musical performance) from one or more users. The sensors 180 can also include one or more buttons for touch input, one or more cameras, gesture sensors, gyroscopes or gyro sensors, barometric sensors, magnetic sensors or magnetometers, acceleration sensors or accelerometers, grip sensors, proximity sensors, color sensors (e.g., red, green, blue (RGB) sensors), biophysical sensors, temperature sensors, humidity sensors, illumination sensors, ultraviolet (UV) sensors, electromyography (EMG) sensors, electroencephalogram (EEG) sensors, electrocardiogram (ECG) sensors, infrared (IR) sensors, ultrasonic sensors, iris sensors, or fingerprint sensors. The sensors 180 can also include an inertial measurement unit, which can include one or more accelerometers, gyroscopes, and other components. In addition, the sensors 180 can include a control circuit for controlling at least one of the sensors included herein. Any of these sensors 180 can be located within the electronic device 101.

[0052] In addition, the electronic device 101 includes one or more speakers 190, which can convert an electrical signal into audible sound. As described below, the one or more speakers 190 can be used to play musical content to at least one user. The musical content played through the one or more speakers 190 can include musical content that accompanies a musical performance by the user, musical content that is related to input provided by the user, or musical content that is generated based on input provided by the user.

[0053] The first external electronic device 102 or the second external electronic device 104 can be a wearable device or a wearable device (e.g., an HMD) that can be mounted to an electronic device. When the electronic device 101 is mounted in the electronic device 102 (e.g., an HMD), the electronic device 101 can communicate with the electronic device 102 through the communication interface 170. The electronic device 101 can be directly connected with the electronic device 102 to communicate with the electronic device 102 without involving a separate network. The electronic device 101 can also be an augmented reality wearable device, such as glasses, that includes one or more cameras.

[0054] The wireless communication can use at least one of, for example, long term evolution (LTE), long term evolution-advanced (LTE-A), a 5th generation wireless system (5G), a millimeter wave or 60 GHz wireless communication, wireless USB, code division multiple access (CDMA), wideband code division multiple access (WCDMA), a universal mobile telecommunications system (UMTS), wireless broadband (WiBro), or a global system for mobile communications (GSM) as a cellular communication protocol. The wired connection can include at least one of, for example, a universal serial bus (USB), a high definition multimedia interface (HDMI), recommended standard 232 (RS-232), or a plain old telephone service (POTS). The network 162 or 164 includes at least one communication network, for example, a computer network (like a local area network (LAN) or a wide area network (WAN)), the Internet, or a telephone network.

[0055] The first and second external electronic devices 102 and 104 and the server 106 can each be devices of the same or different types from the electronic device 101. According to certain embodiments of the present disclosure, the server 106 includes a group of one or more servers. Also, according to certain embodiments of the present disclosure, all or some of the operations performed on the electronic device 101 can be performed on another or multiple other electronic devices (e.g., the electronic devices 102 and 104 or the server 106). Furthermore, according to certain embodiments of the present disclosure, when the electronic device 101 is to automatically perform or request to perform certain functions or services, the electronic device 101 can request another device (such as the electronic devices 102 and 104 or the server 106) to perform at least some of the functions associated therewith, instead of or in addition to performing the functions or services by itself. The other electronic device (e.g., the electronic devices 102 and 104 or the server 106) is capable of performing the requested functions or additional functions and transferring a result of the performance to the electronic device 101. The electronic device 101 can provide the requested functions or services by processing the received result as is or additionally. To that end, for example, a cloud computing, distributed computing, or client-server computing technology can be used. Although Figure 1 It is illustrated that the electronic device 101 includes the communication interface 170 to communicate with the external electronic device 104 or the server 106 via the network 162 or 164, but the electronic device 101 can operate independently without a separate communication function according to some embodiments of the present disclosure.

[0056] The server 106 can optionally support the electronic device 101 by performing or supporting at least one operation (or function) implemented in the electronic device 101. For example, the server 106 can include a processing module or a processor that can support the processor 120 implemented in the electronic device 101.

[0057] Although Figure 1One example of a network configuration 100 including electronic device 101 is shown, but various changes can be made Figure 1 For example, network configuration 100 can include any number of each component in any suitable arrangement. In general, computing and communication systems have a wide variety of configurations, and Figure 1 The scope of the present disclosure is not limited to any particular configuration of computing and communication systems. Figure 1 One operational environment in which various features disclosed in this patent document can be used is shown, but these features can be used in any other suitable system.

[0058] Figure 2 A first example system 200 for machine learning to identify effective music features according to the present disclosure is shown. In particular, Figure 2 The illustrated system 200 can represent one model of machine learning that can be trained to learn a latent music space by embedding relevant music features into the space. Figure 2 The illustrated system 200 can be used in network configuration 100 of Figure 1 for example when system 200 of Figure 2 is implemented using or performed by server 106 in network configuration 100 of Figure 1 However, note that system 200 can be implemented using any other suitable device and in any other suitable environment.

[0059] In example embodiments of Figure 2 System 200 is implemented using a modified form of a deep structured semantic model (DSSM). More specifically, a contrastive DSSM is used to implement system 200. As Figure 2 illustrated, system 200 includes an embedding generator 202. Note that while Figure 2 embedding generator 202 and two embedding generators 202' and 202" are identified, embedding generators 202' and 202" can simply represent the same embedding generator 202 used to process different information. Of course, more than one embedding generator can be used here as well. Embedding generator 202 in this particular example includes a plurality of operational layers 208. Operational layers 208 in embedding generator 202 generally operate to perform various operations to convert non-verbal audible input data 210a-210c into output embeddings 212a-212c, respectively. Each output embedding 212a-212c represents music features of the associated non-verbal auditory input data 210a-210c in a latent music space.

[0060] The operational layers 208 in the embedding generator 202 can perform any suitable operations to generate the output embeddings 212a-c based on the non-verbal audible input data 210a-c. In some embodiments, the embedding generator 202 represents a convolutional neural network that includes the operational layers 208, such as one or more pooling layers, one or more normalization layers, one or more connection layers, and / or one or more convolutional layers. Each pooling layer can select or combine outputs from a previous layer for input to a next layer. For example, a pooling layer using max-pooling identifies a maximum output from a cluster in a previous layer for input to a next layer, while a pooling layer using average-pooling identifies an average of outputs from a cluster in a previous layer for input to a next layer. Each normalization layer can normalize outputs from a previous layer for input to a next layer. Each connection layer can form a connection for routing information between layers. Each convolutional layer can apply a convolution operation to an input in order to generate a result for output to a next layer. Optional connections between non-adjacent operational layers 208 can be represented here using lines 214, meaning that residuals or other data generated by one layer 208 can be provided to a non-adjacent layer 208. In particular embodiments, the embedding generator 202 can represent a fully connected convolutional neural network. However, note that the particular types of machine learning algorithms and connections between layers used here for the embedding generator 202 can vary as needed or desired, and other types of machine learning algorithms can be used here so long as the machine learning algorithms can generate embeddings of musical content in a latent space.

[0061] The embedding generator 202 here is used to process different reference input data 210a and generate different embeddings 212a, process different positive input data 210b and generate different embeddings 212b, and process negative input data 210c and generate different embeddings 212c. The positive input data 210b is known to be similar to the reference input data 210a (at least in terms of the musical features represented by the embeddings). In some cases, the reference input data 210a and the positive input data 210b can represent adjacent sections or portions in the same musical piece, helping to ensure similarity between the two. The negative input data 210c is known to be different from the reference input data 210a (at least in terms of the musical features represented by the embeddings). In some cases, the reference input data 210a and the negative input data 210c can represent different sections or portions in different musical pieces (e.g., different genres), helping to ensure non-similarity between the two.

[0062] For a DSSM, if q(z) represents the aggregated posterior distribution of all embeddings of length d generated by the DSSM function f(x) for x e X, one goal of training the DSSM is to match q(z) to a predefined desired distribution p(z), which can be defined as z, ~ Nd(μ, σ2), where μ = 0, σ2= 1. This can be achieved by connecting an adversarial discriminator 216 to the last layer of the embedding generator 202. Note that while Figure 2 While the adversarial discriminator 216 and two adversarial discriminators 216' and 216" are identified, the adversarial discriminators 216' and 216" can simply represent the same adversarial discriminator 216 used to process different information. Of course, more than one adversarial discriminator can be used here as well. The adversarial discriminator 216 is adversarially trained in conjunction with the embedding generator 202, which itself is a DSSM. The adversarial discriminator 216 is generally used to distinguish between the generated embeddings 212a-212c and vectors sampled from q(z). The adversarial discriminator 216 thereby helps to ensure that the aggregated posterior distribution of the embeddings 212a-212c from the embedding generator 202 conforms to the predetermined distribution. In some embodiments, the predefined distribution can be Gaussian and continuous, although other predefined distributions (e.g., uniform) can be used as well.

[0063] As can be seen from Figure 2 The adversarial discriminator 216 includes several operational layers 222 that can be used to perform the required functionality to support the adversarial discriminator 216. By performing the actual discrimination using only a few layers 222, most of the good features for music content classification will need to be learned by the embedding generator 202. This forces the system 200 to embed the input data 210a-210c in such a way that not only does it encode itself effectively, but it also effectively distinguishes itself from unrelated inputs. While three operational layers 222 are shown here, the adversarial discriminator 216 can include any suitable number of operational layers 222.

[0064] A loss function 224 is used during training of the system 200 to help establish appropriate parameters for the neural network or other machine learning algorithm that forms the embedding generator 202. For example, the loss function 224 can be used during training to minimize using methods such as stochastic gradient descent.

[0065] Standard DSSM training is well suited for metric learning because it explicitly trains the parameters of the DSSM to produce embeddings that are closer together for relevant items (according to a distance metric), while pushing irrelevant items further apart. However, the number of negative samples and the ratio of easy to hard samples is often heavily biased towards the easy end. This often results in poor performance because many samples can satisfy the constraints with very little loss, and do not provide truly meaningful updates during backpropagation. This often results in high inter-class variance and low intra-class variance, making fine-grained classification or meaningful similarity measures (important for music) challenging or impossible. To address this issue, a bootstrap approach has been used in the past, where particularly difficult samples are manually mined from the dataset and used at different stages of training. Unfortunately, this requires human intervention during the training process.

[0066] In some embodiments, the use of the adversarial discriminator 216 naturally helps to alleviate this problem by enforcing a predefined distribution on the embeddings produced by the embedding generator 202. During training, the parameters of the embedding generator 202 can be modified in order to find a way to achieve the desired similarity measure while respecting the predefined distribution of the learning space that does not allow for a learning space in which most of the samples can easily satisfy the similarity constraints.

[0067] In some embodiments, Figure 2 The illustrated system 200 can be trained in two phases. In the first phase, the embedding generator 202 can be trained using standard or other suitable DSSM training techniques, and negative samples can be used to compute the loss (e.g., soft-max loss) during this phase. In the second phase, the embedding generator 202 and the adversarial discriminator 216 are trained such that the adversarial discriminator 216 causes the embedding generator 202 to produce embeddings that look as if they have been sampled from the predefined distribution p(z). Thus, the parameters of the system 200 are being optimized according to two different losses, one learning the similarity measure and the other learning to describe the data such that the aggregated posterior distribution of the embeddings satisfies the predefined distribution (which can be Gaussian and continuous in some embodiments).

[0068] For the first phase of training, in some embodiments, the embedding generator 202 can be trained using Euclidean similarity. The Euclidean similarity between two embeddings can be represented as follows:

[0069] (1)

[0070] where and denote two embeddings, sim( , ) denotes the Euclidean similarity between the two embeddings, D( , represents the Euclidean distance metric between two embeddings in the latent feature space. In this example, the distance metric is represented with the Euclidean distance, although other distance terms (e.g., cosine distance) can be used as the distance metric. In some embodiments, negative samples can be included in the softmax function to compute P( | ), where represents the reconstructed vector or other embedding, and represents the input vector or other embedding. This can be expressed as follows:

[0071] (2)

[0072] The system 200 trains the embedding generator 202 to learn its parameters by minimizing the loss function 224, for example, by using stochastic gradient descent in some embodiments. This can be expressed as follows:

[0073] (3)

[0074] For the second phase of training, the embedding generator 202 and the adversarial discriminator 216 can be trained using a generative adversarial network (GAN) training process in some embodiments. In the GAN training process, the adversarial discriminator 216 is first trained to distinguish between the generated embeddings 212a-212c and vectors or other embeddings sampled from q(z). The embedding generator 202 is then trained to fool the associated adversarial discriminator 216. Some embodiments use a deterministic version of the GAN training process, where the randomness comes only from the data distribution and no additional randomness needs to be introduced.

[0075] Training alternates between the first phase using the DSSM process and the second phase using the GAN process until equation (3) converges. The higher learning rate of the GAN process (particularly for updating the embedding generator 202) relative to the DSSM loss can help to achieve the desired results. Otherwise, the GAN-based updates can have little or no impact, resulting in a model that behaves very similarly to a standard DSSM without any adversarial component.

[0076] Figure 3 A second example system 300 for machine learning to identify effective music features according to the present disclosure is shown. In particular, Figure 3 The illustrated system 300 can represent another model of machine learning that can be trained to learn a latent music space by embedding relevant music features into the space. Figure 3 The illustrated system 300 can be used in the network configuration 100 of Figure 1 for example, when the system 300 of Figure 3 usesFigure 1 The network configuration 100 is implemented by or executed by server 106. However, note that system 300 can be implemented using any other suitable device and in any other suitable environment.

[0077] exist Figure 3 In an example embodiment, system 300 uses an adversarial adjacency model, which can be said to represent a Siamese network paradigm. For example... Figure 3 As shown, system 300 includes an embedding generator 302. Note that, although Figure 3 Embedding generator 302 and embedding generator 302' are identified, but embedding generator 302' can simply refer to the same embedding generator 302 used to process different information. Of course, more than one embedding generator can also be used here. The embedding generator 302 in this particular example includes multiple operation layers 308, which are generally operated to perform various operations to transform the non-verbal audible input data 310a-310b into output embeddings 312a-312b, respectively. Each output embedding 312a-312b represents a musical feature of the associated non-verbal auditory input data 310a-310b in a defined latent music space.

[0078] The operational layer 308 in the embedding generator 302 can perform any suitable operation to generate output embeddings 312a-312b based on the non-verbal audible input data 310a-310b. In some embodiments, the embedding generator 302 represents a convolutional neural network including operational layers 308, such as one or more pooling layers, one or more normalization layers, one or more connection layers, and / or one or more convolutional layers. Line 314 can be used to represent optional connections between non-adjacent operational layers 308, meaning that residuals or other data generated by one layer 308 can be provided to non-adjacent layers 308. In a particular embodiment, the embedding generator 302 can represent a fully connected convolutional neural network. However, note that the specific type of machine learning algorithm used for the embedding generator 302 and the connections between layers can vary as needed or desired, and other types of machine learning algorithms can be used here, as long as the machine learning algorithm can generate embeddings of musical content in the latent space.

[0079] Here, the embedding generator 302 is used to process different reference input data 310a and generate different embeddings 312a, and to process different positive or negative input data 310b and generate different embeddings 312b. The positive input data 310b is known to be similar to the reference input data 310a, and the negative input data 310b is known to be dissimilar to the reference input data 310a (at least in terms of the musical features represented by the embeddings). In some cases, the reference input data 310a and the positive input data 310b can represent adjacent sections or portions of the same musical piece, helping to ensure similarity between the two. Further, in some cases, the reference input data 310a and the negative input data 310b can represent different sections or portions of different musical pieces (e.g., different genres), helping to ensure dissimilarity between the two.

[0080] Unlike the DSSM used in Figure 2 Figure 3 The embeddings 312a-312b in are not directly optimized for a desired metric. Rather, a classifier 324 implemented using a Siamese discriminator can be trained to determine whether two input units (embedding 312a and embedding 312b) are related. Because Siamese is used as a self-supervised replacement for manually designed similarity labels, the classifier 324 can be trained here to determine whether two embeddings 312a and 312b represent consecutive musical content in a musical piece.

[0081] In some embodiments, one version of the classifier 324 generates combined embeddings, where each combined embedding is formed by concatenating one embedding 312a with one embedding 312b to form a single classifier input. Thus, the combined embeddings can represent concatenated vectors. The combined embeddings are used to train the classifier 324 to produce a binary classification from each combined embedding. The binary classification can identify whether the embedding 312a concatenated with the embedding 312b is related (when the embedding 312b is used for positive input data 310b) or unrelated (when the embedding 312b is used for negative input data 310b). However, one goal here can include being able to embed a single unit (either a single embedding 312a or 312b). Thus, some embodiments of the system 300 use tied weights, where the lower layers of the embedding generator 302 are the same, and the embeddings are not concatenated until deep into the network. In other words, the two inputs (input data 310a and input data 310b) are independently embedded, but the same parameters are used to perform the embedding.

[0082] ​Classifier 324 is configured to use the connection embeddings 312a-312b to distinguish between relevant and irrelevant inputs. Again, by performing the actual discrimination in classifier 324 using only a few layers, most of the good features for music content classification will need to be learned by embedding generator 302. This forces system 300 to embed input data 310a-310b in such a way that not only encodes itself effectively, but also enables itself to be effectively distinguished from irrelevant inputs. System 300 can achieve this by embedding relevant units closer together. Note that, for ease of comparison, portions of the architecture of system 300 can use the same DSSM-type network as shown in Figure 2

[0083] Again, adversarial discriminator 316 can be connected to the last layer of embedding generator 302. Note that, while Figure 3 While adversarial discriminator 316 and adversarial discriminator 316' are identified, adversarial discriminator 316' can simply represent the same adversarial discriminator 316 used to process different information. Of course, more than one adversarial discriminator can be used here as well. Adversarial discriminator 316 can operate in the same or similar manner as adversarial discriminator 216. Again, one goal here is that the aggregated posterior distribution of embeddings 312a-312b from embedding generator 302 conforms to a predetermined distribution. In some embodiments, the predetermined distribution can be Gaussian and continuous, although other predetermined distributions (e.g., uniform) can be used as well.

[0084] In some embodiments, Figure 3 System 300 as shown can be trained in two stages. In a first stage, classifier 324 is trained to distinguish between relevant and irrelevant inputs. In a second stage, embedding generator 302 and adversarial discriminator 316 are trained such that adversarial discriminator 316 causes embedding generator 302 to produce embeddings that appear as if they have been sampled from a predefined distribution p(z). Thus, the embedding portion of the model can be updated in both stages of the training process.

[0085] For the first stage of training, in some embodiments, classifier 324 can be trained using cross-entropy with two classes (relevant and irrelevant). This can be represented as:

[0086] (4)

[0087] where M represents the number of classes (two in this example), y' represents the predicted probabilities, and y represents the ground truth. For the second stage of training, the GAN training process described above can be used. Depending on the implementation, compared to the method shown in Figure 2 Figure 3 ​​The illustrated methods can inherently be more stable and involve fewer adjustments to the learning rate between the two losses.

[0088] It should be noted here that, Figure 2 and 3 The two methods illustrated in and can both learn a latent music space by embedding a large number of music passages in the space, and the resulting model in either method can be used to perform a variety of functions. For example, a user input can be projected into the latent space in order to identify the embedding closest to the user input. The closest embedding can then be used to identify, rank, play, or generate music content for the user. It should also be noted here that the execution or use of these methods can be independent of the input. That is, Figure 2 and Figure 3 The methods illustrated in and can operate successfully regardless of how the music content and other non-verbal auditory data are represented. Thus, unlike other methods in which the performance of a music-based model can be determined by the input representation of the data, the methods here can train a machine learning model using any suitable music representation. These representations can include symbolic representations of music and raw audio representations of music, such as mel-frequency cepstral coefficient (MFCC) sequences, spectrograms, amplitude spectra, or chromagrams.

[0089] Although Figure 2 and 3 Two examples of systems 200 and 300 for machine learning to identify effective music features are illustrated, various changes can be made to Figure 2 and 3 For example, Figure 2 and Figure 3 The specific machine learning algorithms implemented in and to generate the embeddings can differ from the algorithms described above. Furthermore, the number of operational layers illustrated in the various embedding generators, adversarial discriminators, and classifiers can vary as desired or as needed.

[0090] Figure 4 and 5 A first example application 400 for machine learning according to the present disclosure is illustrated that is trained to identify effective music features. In particular, Figure 4 and Figure 5 A first example way in which a machine learning model (which can be trained as described above) can be used to perform a particular end-user application is illustrated. Figure 4 and Figure 5 The application 400 illustrated in can be used in the network configuration 100 of Figure 1 for example when using Figure 1at least one server 106 and at least one electronic device 101, 102, 104 in the network configuration 100 to execute the application 400. Note, however, that the application 400 can be executed using any other suitable devices and in any other suitable environment.

[0091] As shown in Figure 4 and Figure 5 An input utterance 402 is received, which in this example represents a request to generate music content to accompany a musical performance by at least one user. In some embodiments, the input utterance 402 can be received at an electronic device 101, 102, 104 of the user and sensed by at least one sensor 180 (e.g., a microphone) of the electronic device 101, 102, 104. Here, the input utterance 402 is digitized and communicated from the electronic device 101, 102, 104 to a cloud-based platform, which can be implemented using one or more servers 106.

[0092] An automatic speech recognition (ASR) and type classifier function 404 of the cloud-based platform analyzes the digitized version of the input utterance 402 in order to understand the input utterance 402 and identify a type of action that occurs in response to the input utterance 402. For example, the ASR and type classifier function 404 can perform natural language understanding (NLU) in order to derive a meaning of the input utterance 402. The ASR and type classifier function 404 can use the derived meaning of the input utterance 402 in order to determine whether a static function 406 or a continuous function 408 should be used to generate a response to the input utterance 402. The ASR and type classifier function 404 supports any suitable logic to perform speech recognition and select a type of response to provide.

[0093] If selected, the static function 406 can analyze the input utterance 402 or its derived meaning and generate a standard response 410. The standard response 410 can be provided to the electronic device 101, 102, 104 for presentation to the at least one user. The static function 406 is typically characterized by the fact that once the standard response 410 is provided, processing of the input utterance 402 can be complete. In contrast, the continuous function 408 can analyze the input utterance 402 or its derived meaning and interact with the electronic device 101, 102, 104 in order to provide a more continuous response to the user request. In this example, since the request is to generate music content to accompany a musical performance, the continuous function 408 can cause the electronic device 101, 102, 104 to generate and play music content that accompanies the musical performance.

[0094] To fulfill a user request here, non-verbal user input 412 is provided from at least one user and processed by one or more analysis functions 414 of the electronic device 101, 102, 104. Here, the non-verbal user input 412 represents a musical performance by at least one user. For example, the non-verbal user input 412 can be generated by one or more users playing one or more musical instruments. The user input 412 is captured by the electronic device 101, 102, 104, e.g., with a microphone of the electronic device. An analog-to-digital function 502 of the analysis functions 414 can be used to convert the captured user input 412 into corresponding digital data, which is used by the analysis functions 414 to generate one or more sets of input data 504 for a trained machine learning model, e.g., the system 200 of Figure 2 or the system 300 of Figure 3 The trained machine learning model uses the one or more sets of input data 504 to produce one or more embeddings 506, which represent the user input 412 projected in a learned latent space. Note that while not shown here, various pre-processing operations can be performed on the digitized user input 412 prior to generating the one or more embeddings 506. Any suitable pre-processing operations can be performed here, e.g., pitch detection.

[0095] The one or more embeddings 506 are used to determine one or more assistive actions 416, which in this example include playing musical content that accompanies the musical performance (e.g., through a speaker 190 of the electronic device 101, 102, 104). For example, the one or more embeddings 506 can be perturbed to generate one or more modified embeddings 508. The perturbation of the one or more embeddings 506 can occur in any suitable manner, e.g., by modifying values contained in the one or more embeddings 506 according to some specified criteria.

[0096] The one or more embeddings 506 and / or the one or more modified embeddings 508 can be used to select or generate music content to play to the user. For example, as part of the retrieving operation 510, the one or more embeddings 506 and / or the one or more modified embeddings 508 can be used to identify one or more similar embeddings in the latent space. Here, the one or more similar embeddings in the latent space are associated with music content that is similar to the music performance, and thus the electronic device 101, 102, 104 can retrieve and play to the user the music content associated with the one or more similar embeddings. As another example, as part of the generating operation 512, the one or more modified embeddings 508 can be decoded and used to generate derived music content, and the electronic device 101, 102, 104 can play to the user the derived music content. This process can be repeated as more non-verbal user input 412 is received and additional music content (whether retrieved or generated) is played to the user.

[0097] Note that while a single electronic device 101, 102, 104 is described here as being used by at least one user, the application 400 shown in Figure 4 and 5 The application 400 shown can be executed using multiple electronic devices 101, 102, 104. For example, the input utterance 402 can be received via a first electronic device 101, 102, 104, and the music content can be played via a second electronic device 101, 102, 104. The second electronic device 101, 102, 104 can be identified in any suitable manner, such as based on a prior configuration or based on the input utterance 402.

[0098] Figure 6 and 7 A second example application 600 is shown for machine learning that is trained to identify valid music features in accordance with the present disclosure. In particular, Figure 6 and 7 A second example way is shown in which a machine learning model (which can be trained as described above) can be used to perform a particular end-user application. Figure 6 and 7 The application 600 shown in Figure 1 may be used in the network configuration 100 of Figure 1 for example when the application 600 is executed using at least one server 106 and at least one electronic device 101, 102, 104 in the network configuration 100 of

[0099] As Figure 6As shown, a user input 602 is received, which in this example represents a request to compose music content. The user input 602 here can take various forms, several examples of which are shown in Figure 6 FIG. 6. For example, the user input 602 can request a composition of music content similar to pre-existing music content, or the user input 602 can request a composition of music content similar to non-verbal input provided by the user (e.g., by humming or playing an instrument). The duration of any non-verbal input provided by the user here can be relatively short, such as between about three seconds and about ten seconds. Although not shown here, the user input 602 can follow a similar path as the input utterance 402 in Figure 4 FIG. 4. That is, the user input 602 can be provided to a cloud-based platform and processed by the ASR and type classifier functionality 404, which can determine that the continuous functionality 408 should be used to generate a response to the user input 602.

[0100] The generation functionality 604 uses some specified music content as a starting seed and generates derived music content to be played back to the user by the presentation functionality 606. For example, if the user requests a composition of a piece of music similar to pre-existing music content, the user’s electronic device 101, 102, 104 can identify (or generate) one or more embeddings of the pre-existing music content in latent space and use the one or more embeddings to produce derived music content. If the user requests a composition of a piece of music similar to music input provided by the user, the user’s electronic device 101, 102, 104 can generate one or more embeddings of the user’s music input in latent space and use the one or more embeddings to produce derived music content. As a particular example, the one or more embeddings from the user’s music input can be used to select a pre-existing piece of music whose embeddings are similar to the embeddings from the user’s music input, and the pre-existing piece of music can be used as a seed.

[0101] Figure 7 One example implementation of the generation function 604 is shown in FIG. 7, which shows the generation function 604 receiving multiple sets of input data 702a-702n. Here, the input data 702a-702n represents seeds for generating derived music content. In some embodiments, the input data 702a-702n can represent different portions (e.g., five to ten second segments) of a pre-existing piece of music. The pre-existing piece of music can represent pre-existing music content specifically identified by a user, or the pre-existing piece of music can represent pre-existing music content selected based on music input by the user. Using a trained machine learning model (e.g., the system 200 of Figure 2 FIG. 2 or the system 300 of Figure 3The system 300) converts the input data 702a-702n into different embeddings 704. The embeddings 704 are provided to at least one recurrent neural network (RNN) 706, which processes the embeddings 704 to produce derived embeddings 708. The derived embeddings 708 represent embeddings in a latent space generated based on the music seeds represented by the input data 702a-702n. The derived embeddings 708 can then be decoded (similar to the generation operation 512 described above) to produce derived music content, which can be played to one or more users.

[0102] As shown here, the output generated by the at least one recurrent neural network 706 for at least one set of input data can be provided in a feed-forward manner for processing additional sets of input data. This can help the at least one recurrent neural network 706 generate different portions of derived music content that are generally consistent with each other (rather than being significantly different). Indeed, the embeddings 704 produced from the input data 702a-702n can be used to train the at least one recurrent neural network 706. It should be noted that while the use of at least one recurrent neural network 706 is shown here, any other suitable generative machine learning model can be used.

[0103] Again, it is noted that while a single electronic device 101, 102, 104 is described here as being used by at least one user, the application 600 shown in Figure 6 and 7 The application 600 shown in FIGS. 6A-6B can be executed using multiple electronic devices 101, 102, 104. For example, the user input 602 can be received via a first electronic device 101, 102, 104, and the music content can be played via a second electronic device 101, 102, 104. The second electronic device 101, 102, 104 can be identified in any suitable manner, such as based on a prior configuration or based on the user input 602.

[0104] Figure 8 , 9 and 10 show a third example application 800 for machine learning that is trained to identify valid music features in accordance with the present disclosure. In particular, Figure 8 , 9 and 10 show a third example manner in which a machine learning model (which can be trained as described above) can be used to perform a particular end-user application. Figure 8 , 9 The application 800 shown in FIGS. 8A-8B can be used in the network configuration 100 of Figure 1 , such as when using the application 600 shown in FIGS. 6A-6B. Figure 1at least one server 106 and at least one electronic device 101, 102, 104 in the network configuration 100 when executing the application 800. Note, however, that the application 800 can be executed using any other suitable devices and in any other suitable environment.

[0105] As shown in Figure 8 the user can provide an initial input utterance 802, which in this example requests playback of a particular type of music (e.g., classical music). The input utterance 802 is provided to the cloud-based platform and processed by the ASR and type classifier functionality 404, which can determine that the static functionality 406 should be used to generate a response to the input utterance 802. This results in a standard response 410, such as playback of some form of classical music or other requested music content.

[0106] As shown in Figure 9 the user can be dissatisfied with the standard response 410 and can provide a subsequent input utterance 902, which in this example requests playback of music content based on user-provided auditory data. The input utterance 902 is provided to the cloud-based platform and processed by the ASR and type classifier functionality 404, which can determine that the continuous functionality 408 should be used to generate a response to the input utterance 902.

[0107] The non-verbal user input 912 is provided to the user’s electronic device 101, 102, 104, such as in the form of a musical performance or other non-verbal input sound. The non-verbal user input 912 here need not represent a particular pre-existing song, but can be improvised in a particular style that the user wishes to hear. One or more analysis functions 914 of the electronic device 101, 102, 104 can convert the non-verbal user input 912 into one or more embeddings, such as in the same or similar manner as shown in Figure 5 and described above. The one or more embeddings are used to determine one or more secondary actions 916, which in this example include playing classical music or other related music content similar to the user input 912 (e.g., through a speaker 190 of the electronic device 101, 102, 104). For example, the one or more embeddings of the user input 912 can be used to identify one or more embeddings in a latent feature space that are closest to the embedding of the user input 912, and music content associated with the one or more identified embeddings in the latent space can be identified and played to the user.

[0108] Figure 10 One example of how the analysis functions 914 and secondary actions 916 can occur is shown. As Figure 10As shown, user input 912 is provided to trained machine learning model 1002, which can represent Figure 2 system 200 or Figure 3 system 300. Trained machine learning model 1002 uses user input 912 to generate at least one embedding 1004, which can occur in the manner described above. The at least one embedding 1004 is used to search learned latent space 1006, which contains embeddings 1008 of other music content. The search here can, for example, find one or more embeddings 1008’ closest to embedding 1004 (according to Euclidean, cosine, or other distance). The one or more identified embeddings 1008’ can be used to retrieve or generate music content 1010, which is played to the user via presentation function 1012. In this way, music content similar (based on perceptually relevant music features) to user input 912 can be identified, allowing music content played to the user to be based on and similar to user input 912.

[0109] Again, note that while a single electronic device 101, 102, 104 is described here as being used by at least one user, the applications 800 shown in Figure 8 , 9 and 10 can be performed using multiple electronic devices 101, 102, 104. For example, input utterance 802, 902 can be received via a first electronic device 101, 102, 104, and music content can be played via a second electronic device 101, 102, 104. The second electronic device 101, 102, 104 can be identified in any suitable manner, such as based on prior configuration or based on input utterance 802, 902.

[0110] Although Figure 4 , 5 , 6, 7, 8, 9, and 10 show example applications of machine learning trained to identify valid music features, various changes can be made to these figures. For example, machine learning that has been trained to identify valid music features can be used in any other suitable manner without departing from the scope of the present disclosure. The present disclosure is not limited to the particular end-user applications presented in these figures.

[0111] Figure 11 and 12 show example methods 1100 and 1200 for machine learning to identify valid music features according to the present disclosure. In particular, Figure 11 example method 1100 for training Figure 2 a machine learning model as shown in Figure 12 example method 1200 for training Figure 3 a machine learning model as shown in Figure 11 and12 Each of the methods 1100 and 1200 shown can be used in Figure 1 Executed in network configuration 100, for example when method 1100 or 1200 is... Figure 1 When server 106 in network configuration 100 is executed. However, note that each of methods 1100 and 1200 can be executed using any other suitable device and in any other suitable environment.

[0112] like Figure 11 As shown, the training of the machine learning model occurs in multiple stages 1102 and 1104. In the first stage 1102, in step 1106, the embedding generator is trained based on a similarity metric of the music content. This may include, for example, the processor 120 of server 106 training the embedding generator 202 using a standard or other suitable DSSM training technique. During this stage 1102, the embedding generator 202 may be trained based on Euclidean distance, cosine distance, or other distance metrics. Furthermore, negative samples may also be included in the softmax function. Overall, during this stage 1102, the embedding generator 202 can be trained to learn its parameters by minimizing the loss function 224.

[0113] In the second stage 1104, the machine learning model is trained adversarially by training an adversarial discriminator in step 1108 to distinguish between generated and sampled embeddings, and by training an embedding generator in step 1110 to attempt to deceive the adversarial discriminator. This may include, for example, the processor 120 of server 106 using a GAN training process. Here, the adversarial discriminator 216 is trained to distinguish between generated embeddings 212a-212c from embedding generator 202 and embeddings sampled from q(z). Furthermore, embedding generator 202 is trained to deceive adversarial discriminator 216. As a result, in step 1112, the adversarial discriminator is used to force the embeddings generated by the embedding generator to have a predetermined distribution. This may include, for example, embedding generator 202 and adversarial discriminator 216 being trained such that adversarial discriminator 216 causes embedding generator 202 to generate embeddings that appear to have been sampled from a predefined distribution p(z).

[0114] In step 1114, it is determined whether to repeat the training phase. This may include, for example, the processor 120 of server 106 determining whether the loss in equation (3) above has converged. As a specific example, this may include the processor 120 of server 106 determining whether the loss value calculated in equation (3) above remains within a threshold amount or a threshold percentage of each other in one or more iterations through phases 1102 and 1104. If not, the process returns to the first training phase 1102. Otherwise, in step 1116, the trained machine learning model is generated and output. At this point, the trained machine learning model can be put into use, for example, for one or more end-user applications, such as music content recognition, music content ranking, music content retrieval, and / or music content generation.

[0115] like Figure 12 As shown, the training of the machine learning model occurs in multiple stages 1202 and 1204. In the first stage 1202, at step 1206, training includes a classifier with an adjacency discriminator to distinguish between relevant and irrelevant content. This may include, for example, the processor 120 of server 106 training classifier 324 to identify whether embeddings 312a-312b are relevant (when embedding 312b is used for positive input data 310b) or irrelevant (when embedding 312b is used for negative input data 310b). As described above, the adjacency discriminator of classifier 324 can handle combined embeddings, for example, when each combined embedding represents embedding 312a and connecting embedding 312b. As a specific example, in some embodiments, classifier 324 may be trained using cross-entropy with two classes.

[0116] In the second stage 1204, the machine learning model is trained adversarially by training an adversarial discriminator in step 1208 to distinguish between generated and sampled embeddings, and by training an embedding generator in step 1210 to attempt to deceive the adversarial discriminator. This may include, for example, the processor 120 of server 106 using a GAN training process. Here, the adversarial discriminator 316 is trained to distinguish between generated embeddings 312a-312b from embedding generator 302 and embeddings sampled from q(z). Furthermore, embedding generator 302 is trained to deceive adversarial discriminator 316. As a result, in step 1212, the adversarial discriminator is used to force the embeddings generated by the embedding generator to have a predetermined distribution. This may include, for example, embedding generator 302 and adversarial discriminator 316 being trained such that adversarial discriminator 316 causes embedding generator 302 to generate embeddings that appear to have been sampled from a predefined distribution p(z).

[0117] At step 1214, it is determined whether to repeat the training phase. If not, the process returns to the first training phase 1202. Otherwise, at step 1216, the trained machine learning model has been generated and output. At this point, the trained machine learning model can be put into use, for example, for one or more end-user applications such as music content identification, music content ranking, music content retrieval, and / or music content generation.

[0118] Although Figure 11 and 12 Examples of methods 1100 and 1200 for machine learning to identify effective music features are shown, various changes can be made to Figure 11 and 12 For example, while shown as a series of steps, various steps in each figure can overlap, occur in parallel, occur in a different order, or occur any number of times. Also, any other suitable technique can be used to train a machine learning model designed in accordance with the teachings of this disclosure.

[0119] Figure 13 An example method 1300 is shown that uses machine learning trained to identify effective music features in accordance with this disclosure. In particular, Figure 13 An example method 1300 is shown that uses a machine learning model (which can be trained as described above) to support at least one end-user application. The machine learning model used here is generated by training a machine learning system having one or more neural networks and one or more adversarial discriminators such that a plurality of embeddings of music features in a latent space have a predetermined distribution. Figure 13 The method 1300 shown can be performed in the network configuration 100 of Figure 1 for example, when the method 1300 is performed by at least one electronic device 101, 102, 104 (possibly in conjunction with at least one server 106) in the network configuration 100 of Figure 1 However, note that the method 1300 can be performed using any other suitable devices and in any other suitable environment.

[0120] As Figure 13As shown, in step 1302, nonverbal input associated with the input musical content is obtained. This may include, for example, the processor 120 of electronic devices 101, 102, 104 receiving nonverbal input from at least one user via a microphone. In these cases, the nonverbal input may represent a musical performance or sound associated with a user's request to identify, accompany, compose, or play music. This may also or alternatively include the processor 120 of electronic devices 101, 102, 104 receiving a request to compose music similar to existing musical content. In those cases, the nonverbal input may represent one or more embeddings of existing musical content. If necessary, in step 1304, one or more embeddings of the nonverbal input are generated using a trained machine learning model. This may include, for example, the processor 120 of electronic devices 101, 102, 104 projecting a digital version of the nonverbal input into a latent musical space using a trained machine learning model.

[0121] In step 1306, one or more embeddings associated with the embeddings related to the input music content are identified. This may include, for example, the processor 120 of the electronic devices 101, 102, 104 identifying one or more nearest neighboring embeddings representing the embeddings associated with the input music content, for example, by using a trained embedding generator 202 or 302. As mentioned above, the distance between embeddings can be determined using various metrics, such as Euclidean, cosine, or other distance metrics.

[0122] One or more embedded elements of input music content and / or one or more identified embedded elements are used to perform desired user functions. In this example, this includes identifying stored music content associated with one or more identified embedded elements and / or generating exported music content in step 1308. This may include, for example, the processor 120 of electronic devices 101, 102, 104 identifying existing music content associated with one or more identified embedded elements, or creating exported music content based on one or more identified embedded elements. The stored and / or exported music content is presented in step 1310. This may include, for example, the processor 120 of electronic devices 101, 102, 104 playing the stored and / or exported music content through at least one speaker 190.

[0123] although Figure 13 An example of a machine learning method 1300 trained to identify valid musical features is shown, but it is possible to... Figure 13 Make various changes. For example, although it is displayed as a series of steps, Figure 13The various steps in the above-described processes can overlap, occur in parallel, occur in a different order, or occur any number of times. Moreover, machine learning models designed in accordance with the teachings of this disclosure can be used in any other suitable manner. As noted above, for example, the models can be used for music content identification, music content ranking, music content retrieval, and / or music content generation. In some cases, the same model can be used for at least two of these functions.

[0124] While the present disclosure has been described with reference to various example embodiments, it will be understood that various changes and modifications can be suggested by those skilled in the art. It is intended that the present disclosure encompass such changes and modifications as fall within the scope of the appended claims.

Claims

1. A method of training a machine learning system, comprising: receiving non-verbal inputs associated with musical passages; alternately: training one or more first neural network components configured to generate a generator in a generative adversarial network using the received non-verbal inputs associated with musical passages to generate a deep structured semantic model through the one or more first neural network components that generates a plurality of embeddings in a latent space that describe different musical passages and relationships between different musical passages, the training causing the generated embeddings to seek to minimize a loss function; adversarially training one or more second neural network components configured to generate a discriminator in the generative adversarial network in conjunction with the one or more first neural network components, the discriminator connected to a last layer of the one or more first neural network components, causing the one or more first neural network components to be trained to generate the deep structured semantic model through the one or more first neural network components such that the embeddings in the latent space have a predefined distribution; until the loss function converges; and generating the learned deep structured semantic model.

2. The method of claim 1, wherein, Training the one or more first neural network components comprises: training the one or more first neural network components configured to generate the generator in the generative adversarial network by using Euclidean similarity.

3. The method of claim 1, wherein, Training the one or more first neural network components comprises: calculating the loss function by using negative musical samples; and minimizing the loss function by using gradient descent, wherein the loss function is a softmax function.

4. The method of claim 1, wherein, Training the one or more second neural network components comprises: training the discriminator to distinguish between the generated embeddings and vectors or other embeddings sampled from a posterior distribution of all generated embeddings, and training the generator to fool the associated discriminator.

5. An electronic device, comprising: at least one memory; at least one speaker; and at least one processor, operatively coupled to the at least one memory and the at least one speaker, the at least one processor configured to: receive non-verbal inputs associated with musical passages; alternately: train one or more first neural network components configured to generate a generator in a generative adversarial network using the received non-verbal inputs associated with musical passages to generate a deep structured semantic model through the one or more first neural network components that generates a plurality of embeddings in a latent space that describe different musical passages and relationships between different musical passages, the training causing the generated embeddings to seek to minimize a loss function; adversarially train one or more second neural network components configured to generate a discriminator in the generative adversarial network in conjunction with the one or more first neural network components, the discriminator connected to a last layer of the one or more first neural network components, causing the one or more first neural network components to be trained to generate the deep structured semantic model through the one or more first neural network components such that the embeddings in the latent space have a predefined distribution; until the loss function converges; and generate the learned deep structured semantic model. Training the one or more first neural network components comprises: training the one or more first neural network components configured to generate the generator in the generative adversarial network by using Euclidean similarity. Training the one or more first neural network components comprises: calculating the loss function by using negative musical samples; and minimizing the loss function by using gradient descent, wherein the loss function is a softmax function. Training the one or more second neural network components comprises: training the discriminator to distinguish between the generated embeddings and vectors or other embeddings sampled from a posterior distribution of all generated embeddings, and training the generator to fool the associated discriminator.

6. The electronic device of claim 5, wherein, The at least one processor is further configured to train the one or more first neural network components configured to generate a generator in a generative adversarial network by using Euclidean similarity.

7. The electronic device of claim 5, wherein, The at least one processor is further configured to: compute a loss function by using negative music samples; and minimize the loss function by using gradient descent, wherein the loss function is a softmax function.

8. The electronic device of claim 5, wherein, The at least one processor is further configured to: train the discriminator to distinguish between the generated embeddings and vectors or other embeddings sampled from a posterior distribution of all generated embeddings, and train the generator to fool the associated discriminator.

9. A method of training a machine learning system, comprising: receiving non-verbal inputs associated with a musical passage, the non-verbal inputs including relevant inputs and irrelevant inputs; generating at least two embeddings by one or more first neural network components configured to generate a generator in a generative adversarial network; training one or more second neural network components configured to generate a classifier in a generative adversarial network using a combined embedding formed by concatenating the two generated embeddings, such that the classifier distinguishes between the relevant and irrelevant non-verbal inputs; adversarially training one or more third neural network components configured to generate a discriminator in a generative adversarial network in conjunction with the one or more first neural network components, the discriminator connected to a last layer of the one or more first neural network components, such that the one or more first neural network components are trained to generate a deep structured semantic model by the one or more first neural network components, such that the embeddings in a latent space have a predefined distribution; and generating a learned adversarial neighborhood model.

10. The method of claim 9, wherein, Training the one or more first neural network components comprises: training the one or more first neural network components configured to generate a classifier in a generative adversarial network by using cross-entropy with two classifications.

11. An electronic device, comprising: at least one memory; at least one speaker; and at least one processor, operatively coupled to the at least one memory and the at least one speaker, the at least one processor configured to: receive non-verbal inputs associated with a musical passage, the non-verbal inputs including relevant inputs and irrelevant inputs; generate at least two embeddings by one or more first neural network components configured to generate a generator in a generative adversarial network; train one or more second neural network components configured to generate a classifier in a generative adversarial network using a combined embedding formed by concatenating the two generated embeddings, such that the classifier distinguishes between the relevant and irrelevant non-verbal inputs; adversarially train one or more third neural network components configured to generate a discriminator in a generative adversarial network in conjunction with the one or more first neural network components, the discriminator connected to a last layer of the one or more first neural network components, such that the one or more first neural network components are trained to generate a deep structured semantic model by the one or more first neural network components, such that the embeddings in a latent space have a predefined distribution; generate a learned adversarial neighborhood model.

12. The electronic device of claim 11, wherein, The at least one processor is further configured to train one or more first neural network components configured to generate a classifier in a generative adversarial network by using cross-entropy with two classes.

13. A machine-readable medium containing instructions that, when executed, cause at least one processor of an electronic device to perform the method of any of claims 1-4 and 9-10.