Machine Learning (ML)-Based In-Audio Emotion and Voice Transformation Using Virtual Domain Mixing and Fake Pair Masking

The ML-based intra-audio emotion and voice transformation using virtual domain mixing and fake pair masking addresses the challenge of costly data collection by enabling efficient transformation of unseen speaker-emotion combinations, effectively converting emotional voice styles and speaker identities.

JP2025536954APending Publication Date: 2025-11-12SONY GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025522741
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-09-27
Filing Date
2023-10-09
Publication Date
2025-11-12

AI Technical Summary

Technical Problem

Conventional emotional voice conversion techniques require costly and time-consuming collection of emotional voice data for target speakers, limiting their applicability to seen speaker-emotion combinations.

Method used

Employing machine learning-based intra-audio emotion and voice transformation using virtual domain mixing and fake pair masking, which allows transformation of unseen speaker-emotion combinations by applying an ML model set to source, reference speaker, and reference emotion audio, and employing a fake-pair masking strategy to prevent overfitting.

Benefits of technology

Enables efficient transformation of emotional voice styles without requiring emotional data from target speakers, facilitating simultaneous transformation of speaker identity and speaking style, and preventing overfitting in the discriminator model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025536954000010
    Figure 2025536954000010
  • Figure 2025536954000011
    Figure 2025536954000011
  • Figure 2025536954000012
    Figure 2025536954000012
Patent Text Reader

Abstract

An electronic device and method for machine learning (ML)-based in-audio emotion and voice transformation using virtual domain mixing and fake pair masking are disclosed. The electronic device receives source audio associated with a first user, reference speaker audio associated with a second user, and reference emotion audio associated with a third user. The electronic device applies a set of ML models to generate transformed audio. The generated transformed audio is associated with the content of the source audio, the identity of the second user, and the emotion of the third user. The electronic device applies a source speaker classifier and a source emotion classifier to the transformed audio and retrains an adversarial model. Based on the retraining, the adversarial model can transform the input audio into output audio associated with the identity of the second user and the emotion of the third user.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-reference to related applications / incorporation by reference This application claims the benefit of priority to U.S. Provisional Patent Application Serial No. 63 / 380,425, filed October 22, 2022, which claims the benefit of priority to U.S. Patent Application No. 18 / 475,788, filed with the U.S. Patent and Trademark Office on September 27, 2023, the contents of which are incorporated herein by reference in their entireties.

[0002] Various embodiments of the present disclosure relate to machine learning-based media processing, and more particularly, to machine learning (ML)-based emotion and voice conversion in audio using virtual domain mixing and fake pair-masking. [Background technology]

[0003] Advances in the field of machine learning (ML)-based speech translation systems have led to the development of various ML models capable of performing emotional voice conversions. Emotional voice conversion can involve receiving a first audio associated with a first emotional style and a second audio associated with a second emotional style, and then converting the first audio into a third audio. This conversion can be such that the linguistic content of the first audio is preserved in the third audio, and the first emotional style is converted into the second emotional style. Conventional emotional voice conversion techniques can focus on speaker-dependent scenarios in which the emotional style of a voice associated with a speaker can be modified. An ML model trained for an emotional voice conversion task can generate an output voice signal upon receiving an input voice signal associated with a speaker. The output voice signal can be such that the speaker identity or emotional style associated with the input voice signal is converted to a target speaker identity or target emotion style. Such a transformation may require training or testing an ML model using emotional voice data associated with the target speaker, however, collecting emotional voice data associated with the target speaker may be costly and time-consuming and may not be feasible in some scenarios. Summary of the Invention

[0004] The limitations and disadvantages of conventional approaches will become apparent to those skilled in the art by comparing the described system with certain aspects of the present disclosure illustrated in the remainder of this application and with reference to the drawings.

[0005] Provided are electronic devices and methods for machine learning (ML)-based in-audio emotion and voice transformation using virtual domain mixing and fake pair masking substantially as shown and / or described in connection with at least one of the drawings and more fully set forth in the claims.

[0006] These and other features and advantages of the present disclosure will become apparent from a consideration of the following detailed description of the disclosure when taken in conjunction with the accompanying drawings, in which like reference characters refer to like elements throughout. [Brief explanation of the drawings]

[0007] [Figure 1] FIG. 1 is a block diagram illustrating an exemplary network environment for machine learning (ML)-based in-audio emotion and voice transformation using virtual domain mixing and fake pair masking, according to an embodiment of the present disclosure. [Figure 2] FIG. 2 is a block diagram illustrating the example electronic device of FIG. 1 in accordance with an embodiment of the present disclosure. [Figure 3A] FIG. 1 illustrates an exemplary processing pipeline for ML-based intra-audio emotion and voice transformation with virtual domain mixing and fake pair masking, according to an embodiment of the present disclosure. [Figure 3B] FIG. 1 illustrates an exemplary processing pipeline for ML-based intra-audio emotion and voice transformation with virtual domain mixing and fake pair masking, according to an embodiment of the present disclosure. [Figure 4] FIG. 1 illustrates an exemplary scenario of ML-based intra-audio emotion and voice transformation using virtual domain mixing and fake pair masking, according to an embodiment of the present disclosure. [Figure 5] FIG. 1 illustrates an exemplary scenario of application of ML models for in-audio emotion and voice transformation using virtual domain mixing and fake pair masking, according to an embodiment of the present disclosure. [Figure 6]1 is a flowchart illustrating the operation of an exemplary method for machine learning (ML)-based in-audio emotion and voice transformation using virtual domain mixing and fake pair masking, according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0008] Electronic devices and methods for machine learning (ML)-based in-audio emotion and voice transformation using virtual domain mixing and fake pair masking may include the following embodiments. Exemplary aspects of the present disclosure may provide an electronic device that can receive source audio (e.g., a voice having a neutral emotion) associated with a first user (i.e., a source speaker). The electronic device can receive reference-speaker audio (e.g., a voice indicating the identity of the target speaker) associated with a second user (i.e., a target speaker). The electronic device can receive reference-emotion audio (e.g., a voice having a non-neutral target emotion) associated with a third user (which may be the target speaker or another speaker). The electronic device can apply a machine learning (ML) model set to the received source audio, the received reference speaker audio, and the received reference emotion audio. The electronic device can generate transformed audio based on the application of the ML model set. The generated transformed audio can be related to the content of the source audio (i.e., the linguistic content of the voice of the first user or source speaker), the identity of the second user (i.e., the target speaker), and the emotion of the third user (i.e., the target emotion). The electronic device can apply each of a source speaker classifier and a source emotion classifier to the generated transformed audio. The electronic device can retrain an adversarial model based on application of each of the source speaker classifier and the source emotion classifier.Based on retraining the adversarial model, input audio (e.g., input voice with a neutral emotion) associated with a first user can be transformed into output audio (e.g., output voice) associated with the identity of a second user and the emotion of a third user.

[0009] It will be appreciated that an emotional voice conversion (EVC) system can convert the emotion associated with an input speech signal from one emotional style to another without modifying the linguistic content of the input speech signal. However, such emotion conversion may generally only be possible for seen speaker-emotion combinations. That is, the EVC system can convert the current emotional style of an audio signal that can be associated with a target speaker to a target emotional style based on the availability of emotional and neutral data associated with the target speaker during training of the EVC system. Collecting emotional data for the target speaker along with neutral data may be costly and time-consuming, and may not be possible in some scenarios.

[0010] To address the above-mentioned problems, the disclosed electronic device and method can employ ML-based intra-audio emotion and voice transformation using virtual domain mixing and fake pair masking. The disclosed electronic device can apply an ML model set to an audio signal to transform the emotion and / or voice associated with a speaker of the audio signal. This transformation can be achieved even if the training or test data available for training or testing the ML model set does not include emotion data associated with the speaker. That is, the disclosed electronic device can use an ML model set for emotion and voice transformation of unseen speaker-emotion combinations. Voice-emotion transformation can be achieved based on emotion data associated with supporting speakers. Furthermore, in some embodiments, the disclosed electronic device can simultaneously transform speaker identity and speaking style. In such simultaneous transformation, a first ML model can be used to determine a speaker style associated with the reference speaker audio, and a second ML model can be used to determine an emotional style associated with the reference emotional audio. Furthermore, the disclosed electronic device can use virtual domain mixing (VDM) for random generation of speaker-emotion pair combinations based on emotion data associated with the supporting speakers. The disclosed electronic device can employ a fake-pair masking strategy to prevent overfitting of the discriminator model due to the use of randomly generated speaker-emotion pairs for training the adversarial model.

[0011] FIG. 1 is a block diagram illustrating an exemplary network environment for machine learning (ML)-based in-audio emotion and voice transformation using virtual domain mixing and fake pair masking, according to an embodiment of the present disclosure. FIG. 1 illustrates a network environment 100. The network environment 100 may include an electronic device 102, a server 104, and a database 106. The electronic device 102 may communicate with the server 104 via one or more networks (e.g., a communications network 108). The electronic device 102 may include an ML model set 110, a source speaker classifier 112A, a source emotion classifier 112B, an adversarial model 112C, and an annealing model 112D. The ML model set 110 may include a first ML model 110A, a second ML model 110B, and a third ML model 110C. The database 106 may include audio data 114. The audio data 114 may include source audio 114 A, reference speaker audio 114 B, and reference emotion audio 114 C. Also shown in FIG.

[0012] The electronic device 102 may include suitable logic, circuitry, interfaces, and / or code that may be configured to receive source audio 114A associated with a first user. The electronic device 102 may receive reference speaker audio 114B associated with a second user. The electronic device 102 may receive reference emotion audio 114C associated with a third user. The electronic device 102 may apply the ML model set 110 to the received source audio 114A, the received reference speaker audio 114B, and the received reference emotion audio 114C. The electronic device 102 may generate transformed audio based on the application of the ML model set 110. The generated transformed audio may be related to the content of the source audio 114A, the identity of the second user, and the emotion of the third user. The electronic device 102 may apply each of the source speaker classifier 112A and the source emotion classifier 112B to the generated transformed audio. The electronic device 102 can retrain the adversarial model 112C based on application of each of the source speaker classifier 112A and the source emotion classifier 112B. Based on the retraining, input audio associated with a first user can be transformed into output audio associated with the identity of a second user and the emotion of a third user.

[0013] Examples of electronic devices 102 may include, but are not limited to, computing devices, smartphones, mobile phones, gaming devices, mainframe machines, servers, computer workstations, machine learning devices (e.g., enabled with or hosting computing, memory, and networking resources), and / or consumer electronics (CE) devices.

[0014] The server 104 may include suitable logic, circuitry, interfaces, and / or code that may be configured to receive the source audio 114A, the reference speaker audio 114B, and the reference emotional audio 114C from the electronic device 102. The server 104 may include an ML model set 110 and may generate transformed audio based on application of the ML model set 110 to each of the received source audio 114A, the received reference speaker audio 114B, and the received reference emotional audio 114C. The generated transformed audio may be related to the content of the source audio 114A, the identity of the second user, and the emotion of the third user. The server 104 may further include a source speaker classifier 112A and a source emotion classifier 112B that may be applied to the generated transformed audio for retraining the adversarial model 112C based on application of each of the source speaker classifier and source emotion classifier (which may be included in the server 104). The retrained adversarial model 112C can facilitate converting input audio received from the electronic device 102 into output audio associated with the identity of the second user and the emotion of the third user. The server 104 can transmit the output audio to the electronic device 102.

[0015] Server 104 may be implemented as a cloud server and may perform operations through web applications, cloud applications, HTTP requests, repository operations, file transfers, etc. Other implementations of server 104 may include, but are not limited to, a database server, a file server, a web server, a media server, an application server, a mainframe server, a machine learning server (e.g., enabled with or hosting computing, memory, and networking resources), or a cloud computing server.

[0016] In at least one embodiment, the server 104 can be implemented as multiple distributed cloud-based resources using multiple techniques known to those skilled in the art. Those skilled in the art will appreciate that the scope of the present disclosure may not be limited to the implementation of the server 104 and the electronic device 102 as two independent entities. In some embodiments, the functionality of the server 104 may be incorporated, in whole or at least in part, into the electronic device 102 without departing from the scope of the present disclosure. In some embodiments, the server 104 may host the database 106. Alternatively, the server 104 may be separate from the database 106 and communicatively coupled to the database 106.

[0017] The database 106 may include suitable logic, interfaces, and / or code that can be configured to store the audio data 114. The database 106 may be derived from data from a relational or non-relational database, or from a comma-separated values ​​(csv) file set in traditional storage or big data storage. The database 106 may be stored or cached on a device, such as a server (e.g., server 104) or the electronic device 102. The device that stores the database 106 may be configured to receive queries for the audio data 114 from the electronic device 102. In response, the device of the database 106 may be configured to retrieve and provide the audio data 114 to the electronic device 102 based on the received query.

[0018] In some embodiments, database 106 may be hosted on multiple servers stored in the same or different locations. The operations of database 106 may be performed using hardware, including a processor, a microprocessor (e.g., that performs or controls the execution of one or more operations), a field programmable gate array (FPGA), or an application specific integrated circuit (ASIC). In some other cases, database 106 may be implemented using software.

[0019] The communication network 108 may include a communication medium that enables the electronic device 102 and the server 104 to communicate with each other. The communication network 108 may be either a wired connection or a wireless connection. Examples of the communication network 108 may include, but are not limited to, the Internet, a cloud network, a cellular or wireless mobile network (such as Long Term Evolution and Fifth Generation (5G) New Radio (NR)), a satellite communication system (e.g., using a network of low-earth orbit satellites), a Wireless Fidelity (Wi-Fi) network, a personal area network (PAN), a local area network (LAN), or a metropolitan area network (MAN). The various devices in the network environment 100 may be configured to connect to the communication network 108 according to various wired and wireless communication protocols. Examples of such wired and wireless communication protocols include, but are not limited to, at least one of Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), ZigBee, EDGE, IEEE802.11, Light Fidelity (Li-Fi), 802.16, IEEE802.11s, IEEE802.11g, multi-hop communication, wireless access point (AP), device-to-device communication, cellular communication protocols, and Bluetooth (BT) communication protocols.

[0020] The first ML model 110A can be a classifier model that can be trained to identify relationships between inputs, such as features, in a training dataset and output labels. The first ML model 110A can be applied to the received reference speaker audio 114B and a first domain code associated with the received reference speaker audio 114B. Based on the application of the first ML model 110A, a speaker style code associated with the received reference speaker audio 114B can be obtained. The first ML model 110A can be defined by hyperparameters, such as the number of weights, a cost function, an input size, and the number of layers. The parameters of the first ML model 110A can be adjusted to update the weights toward a global minimum of the cost function of the first ML model 110A. The first ML model 110A can be trained to output prediction / classification results for an input set after multiple epochs of training on feature information in the training dataset.

[0021] The first ML model 110A may include electronic data that may be implemented, for example, as a software component of an application executable on the electronic device 102. The first ML model 110A may rely on libraries, external scripts, or other logic / instructions for execution by a processing device. The first ML model 110A may include code and routines configured to enable a computing device to perform one or more operations, such as determining a speaker style code. Additionally or alternatively, the first ML model 110A may be implemented using hardware, including a processor, a microprocessor (e.g., performing or controlling the execution of one or more operations), a field programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). Alternatively, in some embodiments, the first ML model 110A may be implemented using a combination of hardware and software.

[0022] In one embodiment, the first ML model 110A can be a neural network (NN) model. The NN model can be a computer network or a system of artificial neurons arranged as nodes in a series of NN layers. The series of NN layers of the NN model can include an input NN layer, one or more hidden NN layers, and an output NN layer. Each layer of the series of NN layers can include one or more nodes (or artificial neurons, represented, for example, by a circle). The output of every node in the input NN layer can be connected to at least one node in the hidden NN layer(s). Similarly, the input of each hidden NN layer can be connected to the output of at least one node in another layer of the NN model. The output of each hidden NN layer can be connected to the input of at least one node in another NN layer of the NN model. The node(s) in the final NN layer can receive inputs from at least one hidden NN layer and output results. The number of NN layers and the number of nodes in each NN layer can be determined from hyperparameters of the NN model. Such hyperparameters can be set before, during, or after training the NN model on a training dataset.

[0023] Each node in a neural network model can correspond to a mathematical function (e.g., a sigmoid function or a rectified linear unit) with a set of parameters that can be adjusted during training of the network. The parameter set can include, for example, weight parameters and regularization parameters. Each node can calculate an output using a mathematical function based on one or more inputs from nodes in other neural network layer(s) (e.g., previous layer(s)) of the neural network. All or some of the nodes in a neural network can correspond to the same or different mathematical functions. In training a neural network model, one or more parameters of each node in the neural network can be updated based on whether the output of the final neural network layer for a given input (from the training dataset) matches the correct result based on the neural network's loss function. The above process can be repeated for the same or different inputs until a minimum value of the loss function is achieved and the training error is minimized. Several training methods are known in the art, such as gradient descent, stochastic gradient descent, batch gradient descent, gradient boosting, and metaheuristic methods.

[0024] The second ML model 110B may be a classifier model that can be trained to identify relationships between inputs, such as features in a training dataset, and output labels. The second ML model 110B may be applied to the received reference emotional audio 114C and a second domain code associated with the received reference emotional audio 114C. Based on the application of the second ML model 110B, an emotion style code associated with the received reference emotional audio 114C may be determined. Details regarding the second ML model 110B may be similar to those of the first ML model 110A. Accordingly, such details are omitted for the sake of brevity of this disclosure.

[0025] The third ML model 110C can be a machine learning model that can be applied to the received source audio 114A, the determined speaker style code, and the determined emotional style code. The generation of the transformed audio can be based on the application of the third ML model 110C. Details regarding the third ML model 110C can be similar to the first ML model 110A. Accordingly, such details are omitted for the sake of brevity of this disclosure.

[0026] The source speaker classifier 112A may be a classifier model that can be applied to the generated transformed audio. Based on the application of the source speaker classifier 112A to the generated transformed audio, the domain of the source speaker of the generated transformed audio may be determined. Details regarding the source speaker classifier 112A may be similar to those of the first ML model 110A. Accordingly, such details are omitted for the sake of brevity of this disclosure.

[0027] The source emotion classifier 112B can be a classifier model that can be applied to the generated transformed audio. Based on application of the source emotion classifier 112B to the generated transformed audio, a domain of emotion associated with the generated transformed audio can be determined. Further, details regarding the source emotion classifier 112B can be similar to those of the first ML model 110A. Accordingly, such details are omitted for the sake of brevity of this disclosure.

[0028] The adversarial model 112C can be an ML model that can be retrained based on application of each of the source speaker classifier 112A and the source emotion classifier 112B to the transformed audio. The retrained adversarial model 112C can facilitate the third ML model 110C to transform input audio associated with the first user into output audio associated with the identity of the second user and the emotion of the third user. In one embodiment, the adversarial model 112C can include a classifier model. The classifier model can be a type of classifier model that can classify whether the output generated by the generator model (i.e., the third ML model 110C) is authentic or fake. The classifier model can classify the transformed audio as authentic or fake. Details regarding the adversarial model 112C and the classifier model can be similar to those of the first ML model 110A. Accordingly, such details are omitted for the sake of brevity of this disclosure.

[0029] The annealing model 112D may be an ML model that can be applied to a fundamental frequency loss and a norm consistency loss associated with the generated transformed audio. Based on the application of the annealing model 112D, a weight set associated with the fundamental frequency loss and the norm consistency loss may be determined. Furthermore, details regarding the annealing model 112D may be similar to those of the first ML model 110A. Accordingly, such details are omitted for the sake of brevity of this disclosure.

[0030] The audio data 114 may include source audio 114A, reference speaker audio 114B, and reference emotion audio 114C. The source audio 114A may be source audio data, such as voice data associated with a source speaker (i.e., a first user) having a neutral emotion. The identity and / or emotion (of the source speaker) associated with the voice data (i.e., voice content) may need to be converted to voice data associated with the identity of a target speaker (e.g., a second user) and the emotion of the target speaker or another speaker (e.g., a third user). The reference speaker audio 114B may be voice data associated with a target speaker (i.e., a second user) having a neutral or non-neutral emotion. The reference emotion audio 114C may be voice data associated with a speaker (i.e., a third user) having a target emotion. In one embodiment, the source audio 114A may correspond to a neutral-emotion spectrogram associated with a first user. The reference speaker audio 114B may correspond to a user-identity spectrogram associated with a second user. The reference emotion audio 114C may correspond to a non-neutral emotion spectrogram associated with a third user. In some embodiments, the first user may be the same as the second user. In other embodiments, the second user may be the same as the third user. In yet other embodiments, the first user, second user, and third user may be different users.

[0031] In operation, the electronic device 102 can be configured to receive source audio 114A associated with a first user. The source audio 114A can be audio that can be associated with the first user's identity and a neutral emotion. The identity and / or (neutral) emotion that the source audio 114A can be associated with may need to be translated into the identity and target emotion of a target speaker. In one example, the electronic device 102 can send a request to the database 106 to retrieve the source audio 114A. Upon receiving the request, the database 106 can verify the request, and based on the verification, the server 104 can send the source audio 114A to the electronic device 102. Details regarding receiving the source audio 114A are further shown, for example, in FIG. 3A (302).

[0032] The electronic device 102 can be further configured to receive reference speaker audio 114B associated with a second user. The reference speaker audio 114B can be voice content that can be associated with the identity of the second user. The second user can be a target speaker. The electronic device 102 can determine a speaking style code based on the reference speaker audio 114B. After converting the source audio 114A, the audio content of the source audio 114A can be associated with the identity of the target speaker. In an embodiment, the electronic device 102 can send a request to the database 106 to retrieve the reference speaker audio 114B. Upon receiving the request, the database 106 can verify the request, and based on the verification, the server 104 can send the reference speaker audio 114B to the electronic device 102. Details regarding receiving the reference speaker audio 114B are further shown, for example, in FIG. 3A (304).

[0033] The electronic device 102 may be further configured to receive reference emotional audio 114C associated with a third user. The reference emotional audio 114C may be voice content that may be associated with an emotion of the third user, which may be a target emotion. The electronic device 102 may determine an emotional style code based on the reference emotional audio 114C. After converting the source audio 114A, the audio content of the source audio 114A may be associated with the identity of the target speaker and the target emotion. In some embodiments, the electronic device 102 may send a request to the database 106 to retrieve the reference emotional audio 114C. Upon receiving the request, the database 106 may verify the request, and based on the verification, the server 104 may send the reference emotional audio 114C to the electronic device 102. Details regarding receiving the reference emotional audio 114C are further shown, for example, in FIG. 3A (306).

[0034] The electronic device 102 may be further configured to apply the ML model set 110 to the source audio 114A, the reference speaker audio 114B, and the reference emotional audio 114C. The source audio 114A, the reference speaker audio 114B, and the reference emotional audio 114C may be provided as inputs to the ML model set 110. Details regarding the application of the ML model set 110 are further shown, for example, in FIG. 3A (308).

[0035] The electronic device 102 may be further configured to generate transformed audio based on application of the ML model set 110. The generated transformed audio may be related to the content of the source audio 114A, the identity of the second user (i.e., the target speaker), and the emotion of the third user (i.e., the target emotion). The ML model set 110 may transform the identity (of the first user) and emotion (neutral emotion) to which the content of the source audio 114A may be associated. The transformation may be based on the identity (of the second user or the target speaker) to which the reference speaker audio 114B may be associated and the emotion (i.e., the target emotion) to which the reference emotional audio 114C may be associated. As a result of this transformation, transformed audio may be generated that may include a speaking style related to the reference speaker audio 114B and an emotional style related to the reference emotional audio 114C. Details regarding the generation of the transformed audio are further shown, for example, in FIG. 3A (310).

[0036] The electronic device 102 may be further configured to apply a source speaker classifier 112A and a source emotion classifier 112B, respectively, to the generated transformed audio. Based on the application of the source speaker classifier 112A and the source emotion classifier 112B, the electronic device 102 may determine a source speaker domain and an emotion domain associated with the generated transformed audio. The output of the speaker classifier 112A may indicate whether an identity associated with the transformed audio belongs to a target speaker (whose identity may be associated with the reference speaker audio 114B). Similarly, the output of the source emotion classifier 112B may indicate whether an emotion associated with the transformed audio is a target emotion (i.e., an emotion associated with the reference emotion audio 114C). Details regarding the application of the source speaker classifier 112A and the source emotion classifier 112B, respectively, are further illustrated, for example, in FIG. 3B (312).

[0037] The electronic device 102 may be further configured to retrain the adversarial model 112C based on the application of each of the source speaker classifier 112A and the source emotion classifier 112B. Retraining the adversarial model 112C may facilitate the third ML model 110C to convert input audio associated with the first user into output audio. The output audio may be associated with the identity of the second user and the emotion of the third user. Based on the retraining of the adversarial model 112C, the third ML model 110C may convert input audio associated with the first user into output audio associated with the identity of the second user and the emotion of the third user. Details regarding the retraining of the adversarial model 112C are further shown, for example, in FIG. 3B.

[0038] Figure 2 is a block diagram illustrating the example electronic device of Figure 1, in accordance with an embodiment of the present disclosure. The description of Figure 2 is provided with reference to elements of Figure 1. Figure 2 illustrates an example electronic device 102. The electronic device 102 may include circuitry 202, memory 204, input / output (I / O) devices 206, a network interface 208, an ML model set 110, a source speaker classifier 112A, a source emotion classifier 112B, an adversarial model 112C, and an annealing model 112D. The ML model set 110 may include a first ML model 110A, a second ML model 110B, and a third ML model 110C. The memory 204 may include audio data 114. The input / output (I / O) devices 206 may include a display device 210.

[0039] The circuit 202 may include suitable logic, circuits, and / or interfaces that can be configured to execute program instructions associated with different operations performed by the electronic device 102. These operations may include receiving source audio, receiving reference speaker audio, receiving reference emotion audio, applying a set of ML models, generating transformed audio, applying a classifier, and retraining an adversarial model. The circuit 202 may include one or more processing units that may be implemented as independent processors. In some embodiments, one or more processing units may be implemented as an integrated processor or a group of processors that collectively perform the functions of one or more specialized processing units. The circuit 202 may be implemented based on multiple processor technologies known in the art. Example implementations of the circuit 202 may be an X86-based processor, a graphics processing unit (GPU), a reduced instruction set computing (RISC) processor, an application-specific integrated circuit (ASIC) processor, a complex instruction set computing (CISC) processor, a microcontroller, a central processing unit (CPU), and / or other control circuitry.

[0040] The memory 204 may include suitable logic, circuitry, interfaces, and / or code that may be configured to store one or more instructions executed by the circuit 202. The one or more instructions stored in the memory 204 may be configured to perform different operations of the circuit 202 (and / or the electronic device 102). The memory 204 may further be configured to store audio data 114 (which may be retrieved from the database 106) and converted audio. Example implementations of the memory 204 may include, but are not limited to, random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), hard disk drive (HDD), solid-state drive (SSD), CPU cache, and / or a secure digital (SD) card.

[0041] The I / O device 206 may include suitable logic, circuitry, interfaces, and / or code that can be configured to receive input and provide output based on the received input. For example, the I / O device 206 may receive a first user input indicating a request to convert the source audio 114A. The first user input may include the source audio 114A. For example, the I / O device 206 may include a microphone through which the user 116 can record their own voice as the source audio 114A. Alternatively, the user 116 may select the source audio 114A from a set of audio files stored on the electronic device 102. The I / O device 206 may be further configured to render the converted audio. The I / O device 206 may include a display device 210. Examples of the I / O device 206 may include, but are not limited to, a display (e.g., a touch screen), a keyboard, a mouse, a joystick, a microphone, or a speaker. Examples of the I / O device 206 may further include a Braille I / O device, such as a Braille keyboard and a Braille reader.

[0042] The network interface 208 may include suitable logic, circuitry, interfaces, and / or code that may be configured to facilitate communications between the electronic device 102 and the server 104 over the communications network 108. The network interface 208 may be implemented using various known technologies to support wired or wireless communications between the electronic device 102 and the communications network 108. The network interface 208 may include, but is not limited to, an antenna, a radio frequency (RF) transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a coder-decoder (CODEC) chipset, a subscriber identity module (SIM) card, or a local buffer circuit.

[0043] The network interface 208 may be configured to communicate via wireless communication with a network such as the Internet, an intranet, a wireless network, a cellular telephone network, a wireless local area network (LAN), or a metropolitan area network (MAN). The wireless communication may be configured to use one or more of a number of communication standards, protocols, and technologies, such as Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (W-CDMA), Long Term Evolution (LTE), Fifth Generation (5G) New Radio (NR), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wireless Fidelity (WiFi) (such as IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, or IEEE 802.11n), Voice over Internet Protocol (VoIP), Light Fidelity (Li-Fi), Worldwide Interoperability for Microwave Access (Wi-MAX), protocols for email, instant messaging, and short message service (SMS).

[0044] The display device 210 may include suitable logic, circuits, and interfaces that may be configured to render information based on instructions or input received from the circuit 202 or the I / O device 206. The display device 210 may be a touchscreen that allows a user (e.g., the user 116) to provide user input via the display device 210. The touchscreen may be at least one of a resistive touchscreen, a capacitive touchscreen, or a thermal touchscreen. The display device 210 may be implemented through a number of known technologies, such as, but not limited to, at least one of a liquid crystal display (LCD) display, a light-emitting diode (LED) display, a plasma display, or an organic LED (OLED) display technology, or other display devices. According to some embodiments, the display device 210 may represent a display screen of a head-mounted device (HMD), a smart glasses device, a see-through display, a projection display, an electrochromic display, or a transparent display. Various operations of the circuit 202 for ML-based intra-audio emotion and voice conversion using virtual domain mixing and fake pair masking are further described, for example, in FIGS. 3A and 3B.

[0045] 3A and 3B collectively illustrate an exemplary processing pipeline for ML-based intra-audio emotion and voice conversion using virtual domain mixing and fake pair masking, according to an embodiment of the present disclosure. The description of FIGS. 3A and 3B is provided with reference to elements in FIGS. 1 and 2. FIGS. 3A and 3B illustrate an exemplary processing pipeline 300 illustrating exemplary operations 302-318 for ML-based emotional voice conversion based on virtual domain mixing and fake pair masking. The exemplary operations 302-318 can be performed by any computer system, such as the electronic device 102 of FIG. 1 or the circuit 202 of FIG. 2. FIGS. 3A and 3B further include source audio 114A, reference speaker audio 114B, reference emotion audio 114C, the ML model set 110, transformed audio 310A, a source speaker classifier 112A, a source emotion classifier 112B, an adversarial model 112C, input audio 316A, and output audio 318A.

[0046] Referring to FIG. 3A, a source audio receiving operation may be performed at 302. The circuit 202 may be configured to receive source audio 114A associated with a first user. The source audio 114A may be voice data associated with the identity of the source speaker and the emotion of the source speaker. The identity and / or emotion may need to be transformed. For example, the source audio 114A (i.e., voice data) may be the voice of a first user, such as user 116, which may be recorded or received as the source audio 114A.

[0047] In one embodiment, the source audio 114A may correspond to a neutral emotion spectrogram associated with a first user. The spectrogram may be extracted from the source audio 114A, which may be a sentence spoken with a neutral emotion by the first user (i.e., the source speaker). Note that a spectrogram (herein a neutral emotion spectrogram) may be understood to be usable to visually represent signal strength, such as intensity and loudness, of an audio signal (herein a source audio 114A) over time across a spectrum of frequencies.

[0048] At 304, a reference speaker audio receiving operation can be performed. The circuit 202 can be configured to receive reference speaker audio 114B associated with a second user. The reference speaker audio 114B can be voice data associated with the identity of the second user (i.e., the target speaker). The circuit 202 can determine a speaker style code using the reference speaker audio 114B. For example, the reference speaker audio 114B can be a voice recording of a sentence spoken by the target speaker. The sentence can be spoken with a neutral emotion. In one embodiment, the reference speaker audio 114B can correspond to a user identity spectrogram associated with the identity of the second user (i.e., the target speaker). The user identity spectrogram can be extracted from the reference speaker audio 114B (i.e., the sentence spoken by the second user).

[0049] At 306, a reference emotional audio receiving operation may be performed. The circuit 202 may be configured to receive reference emotional audio 114C associated with a third user. The reference emotional audio 114C may be voice data that may be associated with the emotion of the third user (i.e., the target emotion). The circuit 202 may determine an emotional style code using the reference emotional audio 114C. For example, the reference emotional audio 114C may be a voice recording of a sentence spoken by the third user with an angry emotion, which may be the target emotion. The emotion associated with the source audio 114A may be transformed from neutral to the target emotion (i.e., the angry emotion). In some embodiments, the reference emotional audio 114C may correspond to a non-neutral emotional spectrogram (e.g., an angry emotion) associated with the third user. The non-neutral emotional spectrogram may be extracted from the reference emotional audio 114C (i.e., the sentence spoken by the third user). In some embodiments, the first user may be the same as the second user. In other embodiments, the second user may be the same as the third user. In yet another embodiment, the first user, the second user, and the third user can be different users.

[0050] An ML model set application operation can be performed at 308. The circuit 202 can be configured to apply the ML model set 110 to each of the source audio 114A, the reference speaker audio 114B, and the reference emotional audio 114C. The ML model set 110 can include a first ML model 110A, a second ML model 110B, and a third ML model 110C.

[0051] In one embodiment, the circuit 202 can be configured to apply a first ML model 110A (e.g., a speaker style encoder model) of the ML model set 110 to the reference speaker audio 114B and a first domain code associated with the received reference speaker audio 114B, where the first domain code can be a speaker domain code. The received reference speaker audio 114B and the first domain code can be provided as inputs to the first ML model 110A (i.e., the speaker style encoder model).

[0052] The speaker style encoder can extract a speaker style code from the first domain code. In one embodiment, a speaker style encoder model can process the reference speaker audio 114B with respect to a set of domain codes to obtain a set of shared features. A first domain projection can then be used to map the obtained set of shared features to the first domain code. Thus, the circuit 202 can be configured to determine a speaker style code associated with the reference speaker audio 114B based on application of the first ML model 110A to the reference speaker audio 114B. The speaker style code can be obtained according to equation (1): TIFF2025536954000001.tif13150(1) Here, "S spk " can be a function related to the speaker style encoder model (i.e., the first ML model 110A), and "h spk " can be a speaker style code, and "R spk " can be a reference spectrogram (i.e., a user identity spectrogram) associated with (the identity of) a second user (i.e., a target speaker), and "y spk " may correspond to the first domain code.

[0053] The circuit 202 may be configured to apply a second ML model 110B (e.g., an emotion style encoder model) of the ML model set 110 to the reference emotional audio 114C and a second domain code associated with the reference emotional audio 114C. Here, the second domain code may be an emotional domain code. The reference emotional audio 114C and the second domain code may be provided as inputs to the second ML model 110B (i.e., the emotion style encoder model). The emotion style encoder model may extract the emotion style code from the second domain code. In an embodiment, the emotion style encoder model may process the reference emotional audio 114C for a set of domain codes to obtain a set of common features. Then, a second domain projection may be used to map the obtained set of common features to the second domain code. Thus, the circuit 202 may be configured to determine an emotion style code associated with the reference emotional audio 114C based on application of the second ML model 110B to the reference emotional audio 114C. In one example, the emotional style code can be obtained according to equation (2): TIFF2025536954000002.tif7150(2) Here, "S emo " can be a function related to the emotional style encoder model (i.e., the second ML model 110B), and "h emo " can be an emotional style code, and "R emo " can be a reference spectrogram (i.e., a non-neutral emotion (i.e., target emotion) spectrogram) associated with a third user, and "y emo " can correspond to a second domain code.

[0054] The circuit 202 can be configured to apply a third ML model 110C of the ML model set 110 to the source audio 114A, the determined speaker style code, and the determined emotional style code. The source audio 114A, the speaker style code, and the emotional style code can be provided as inputs to the third ML model 110C.

[0055] In one embodiment, the third ML model 110C can be a generator model that can convert the source audio 114A into transformed audio having a target style specified by one or more style embeddings (i.e., speaker identity embeddings and emotion embeddings), where the style embeddings can be speaker style codes (associated with the identity of the target speaker, i.e., the second user) and emotion style codes (associated with the target emotion (e.g., the emotion of anger)).

[0056] In one embodiment, the generator model may include an encoder model, an adder model, and a decoder model. The encoder model may determine an encoding vector associated with the source audio 114A. The adder model may determine a sum of two or more inputs. The decoder model may generate the transformed audio 310A.

[0057] In one embodiment, the circuit 202 can be configured to apply an encoder model to the source audio 114A. The source audio 114A can be provided as an input to the encoder model. Based on the application of the encoder model to the source audio 114A, an encoding vector can be determined as an output of the encoder model. In one example, the encoding vector can correspond to a latent feature vector "hx" associated with the received source audio 114A.

[0058] The circuit 202 may be further configured to apply a fundamental frequency network to the source audio 114A to determine a fundamental frequency. The fundamental frequency may correspond to the pitch of the audio waveform. The fundamental frequency network may be a joint detection and classification (JDC) network including a series of convolutional layers and one or more bidirectional long short-term memory (BLSTM) units. The JDC network may be pre-trained to extract the fundamental frequency from a source spectrogram (i.e., a neutral emotion spectrogram to which the source audio 114A may correspond, which may be associated with a first user).

[0059] The circuit 202 can be further configured to apply an adder model to the determined coding vector and the determined fundamental frequency. The circuit 202 can be further configured to determine a first vector based on application of the adder model. The adder model can accumulate the coding vector and the fundamental frequency to determine the first vector.

[0060] The circuit 202 may be further configured to apply a decoder model to the first vector, the speaker style code, and the emotional style code. The determined first vector, the determined speaker style code, and the determined emotional style code may be provided as inputs to the decoder model.

[0061] At 310, a transformed audio 310A determination operation can be performed. The circuit 202 can be configured to determine the transformed audio 310A based on application of the ML model set 110. The generated transformed audio 310A can be related to the content of the source audio 114A (i.e., a sentence spoken by a first user or source speaker, which can be received as the source audio 114A). The transformed audio 310A can further be related to the identity of a second user (i.e., a target speaker) (to which the reference speaker audio 114B can be related) and the emotion of a third user (i.e., a target emotion). The ML model set 110 can process the source audio 114A, the reference speaker audio 114B, and the reference emotion audio 114C to transform the source audio 114A into the transformed audio 310A. In one example, the source audio 114A can be related to the first user or source speaker, “John.” The source audio 114A can be a first sentence spoken with a neutral emotion and can correspond to a spectrogram of a neutral emotion. The reference speaker audio 114B can be associated with the identity of a second user or target speaker, "Mark." The reference speaker audio 114B can be a second sentence spoken with a neutral or non-neutral emotion and can correspond to a spectrogram of a user identity (Mark's identity). The reference emotion audio 114C can be associated with a target emotion (e.g., an emotion of "anger"). The reference emotion audio 114C can be a third sentence spoken by a third user, "Tom," with an emotion of "anger." The source audio 114A needs to be transformed so that the transformed audio (e.g., transformed audio 310A) relates to the linguistic content of the first sentence, the speaking style (i.e., identity) of the second user (i.e., target speaker "Mark"), and the emotional style of the third user "Tom," and thus the transformed audio 310A can be generated based on the source audio 114A, the reference speaker audio 114B, and the reference emotional audio 114C.It should be noted that the linguistic content of the received source audio 114A, the received reference speaker audio 114B, and the received reference emotion audio 114C may or may not be similar. On the other hand, the linguistic content of the source audio 114A and the linguistic content of the transformed audio 310A may be similar or identical. The transformed audio 310A may resemble audio including a first sentence, spoken by a second user (i.e., the target speaker, "Mark"), and having an emotion (i.e., target emotion) of a third user (i.e., "Tom") (of anger).

[0062] In some embodiments, generation of the transformed audio 310A can be further based on application of a third ML model 110C. As described above, the third ML model 110C of the ML model set 110 can be applied to each of the source audio 114A, the speaker style code, and the emotional style code (which are inputs to the third ML model 110C or generator model). The third ML model 110C can transform the received source audio 114A into the transformed audio 310A based on the speaker style code and the emotional style code. Because the determined speaker style code can be associated with the reference speaker audio 114B and the emotional style code can be associated with the reference emotional audio 114C, the transformed audio 310A can be associated with the identity of the second user and the emotion of the third user.

[0063] In some embodiments, generating the transformed audio 310A can be based on application of a decoder model, where the decoder model can be applied to the determined first vector, the determined speaker style code, and the determined emotional style code. The decoder model can process the determined first vector, the determined speaker style code, and the determined emotional style code to generate the transformed audio 310A.

[0064] Referring to FIG. 3B, a classifier application operation can be performed at 312. The circuit 202 can be configured to apply each of a source speaker classifier 112A and a source emotion classifier 112B to the generated transformed audio 310A. The source speaker classifier 112A can determine whether the generated transformed audio 310A is associated with the identity of a second user (i.e., a target speaker). The source emotion classifier 112B can determine whether the generated transformed audio 310A is associated with the emotion of a third user (i.e., a target emotion). For example, the source speaker classifier 112A can determine whether the speaker of the linguistic content of the transformed audio 310A is “Mark,” and the source emotion classifier 112B can determine whether the linguistic content of the transformed audio 310A is spoken with the emotion of “anger” (i.e., the emotion of “Tom”). These determinations can be based on a first domain code and a second domain code.

[0065] An adversarial model retraining operation can be performed at 314. The circuit 202 can be configured to retrain the adversarial model 112C based on the application of each of the source speaker classifier 112A and the source emotion classifier 112B.

[0066] The adversarial model 112C can be retrained to optimize an optimization function. The optimization function can be related to an adversarial loss, an adversarial source classifier loss, a style reconstruction loss, and a style diversification loss. The adversarial loss can be determined according to equation (3): TIFF2025536954000003.tif14161(3) Here, "L adv” can be the adversarial loss, “X” can be the source spectrogram (i.e., the neutral sentiment spectrogram associated with the first user), and “y src ” can be the source domain, “D” is a function related to the discriminator model, and “y drg " can be the target domain, and "h spk " can be a speaker style code, and "h emo ” can be an emotional style code. Note that “D(·,y)” can indicate whether the output is genuine or fake for domain “y∈Y”.

[0067] The source speaker classifier 112A and the source emotion classifier 112B can be trained to classify the speaker identity and speaker emotion associated with the transformed audio 310A based on cross entropy loss. The adversarial source classifier loss can be determined according to equation (4): TIFF2025536954000004.tif28163(4) where “X” can be the source spectrogram (i.e., the neutral emotion spectrogram associated with the first user), and “y spk " can correspond to the first domain code, and "h spk " may be a speaker style code, "CE(.)" may be a cross-entropy loss, "G(.)" may be a generator function associated with the third ML model 110C, and "h emo ” can be the emotion style code, “D” can be a function related to the discriminator model, and “y src " can be the source domain, and "y emo " can correspond to a second domain code.

[0068] The style reconstruction loss can be used to determine style codes, such as speaker style codes and emotion style codes, associated with the transformed audio 310A. The style reconstruction loss can be determined according to equation (5): TIFF2025536954000005.tif21161(5) Here, "L sty ” can be the style reconstruction loss, “X” can be the source spectrogram (i.e., the neutral emotion spectrogram associated with the first user), and “y spk " can correspond to the first domain code, and "h spk " can be a speaker style code, and "S spk " can be a function associated with the speaker style encoder model (i.e., the first ML model 110A), "G(.)" can be a generator function associated with the third ML model 110C, and "h emo " can be an emotional style code, and "y emo " can correspond to a second domain code, and "S emo " can be a function associated with the emotional style encoder model (i.e., the second ML model 110B).

[0069] The style diversification loss can be used to ensure diversity in the generated samples with different style embeddings. The style diversification loss can be determined according to equation (6), TIFF2025536954000006.tif41161(6) Here, "L ads ” is the style diversification loss, “X” can be the source spectrogram (i.e., the neutral emotion spectrogram associated with the first user), and “y spk " can correspond to the first domain code, and "h spk " can be a speaker style code (generated as an output of the first ML model 110A), "G(.)" can be a generator function associated with the third ML model 110C, and "h emo" can be the emotion style code (generated as the output of the second ML model 110B), and "h' emo " can be another emotional style code, and "h' spk " can be another speaker style code. The adversarial model 112C can be retrained based on each of the adversarial loss, the adversarial source classifier loss, the style reconstruction loss, and the style diversification loss.

[0070] In an embodiment, the circuit 202 may be further configured to determine a fundamental frequency loss and a norm consistency loss associated with the generated transformed audio 310A. The determination of the fundamental frequency loss may be based on the following equation (7): TIFF2025536954000007.tif14161(7) Here, "L fo " can be the fundamental frequency loss, "X" can be the source spectrogram, "F(X)" (i.e., a neutral emotion spectrogram associated with the first user), and the fundamental frequency in Hertz of "X" can be determined; TIFF2025536954000008.tif7150 can be the temporal mean. Since the mean fundamental frequency of male and female speakers may be different, the fundamental frequency loss can be determined taking the temporal mean into account.

[0071] The norm consistency loss can be determined based on equation (8), TIFF2025536954000009.tif14161(8) Here, "L norm " can be a norm consistency loss, "X" can be a source spectrogram (i.e., a neutral emotion spectrogram associated with the first user), and "t" can be a frame index. The norm consistency loss can be used to preserve silence intervals that may be present within "X".

[0072] The circuit 202 may then be configured to apply the annealing model 112D to the fundamental frequency loss and the norm consistency loss. The circuit 202 may further be configured to determine a weight set associated with the determined fundamental frequency loss and the determined norm consistency loss based on the application of the annealing model 112D. The weight set may be determined by gradually linearly decreasing the first weight set associated with the fundamental frequency loss and the norm consistency loss in each iteration. The third ML model 110C may be retrained further based on the determined weight set. Furthermore, in some cases, to determine the effectiveness of the annealing model 112D, an ML model may be trained for annealing, another ML model may be trained for the fundamental frequency loss, and yet another ML model may be trained for the norm consistency loss.

[0073] In one embodiment, the adversarial model 112C may include a discriminator model. Note that a discriminator model may be understood as a type of classifier model that can classify whether an output generated by a generator model (i.e., the third ML model 110C) is authentic or fake. If it is determined that the identity associated with the transformed audio 310A belongs to the second user or target speaker and the emotion associated with the transformed audio 310 is the target emotion, the output may be authentic. In one example, the discriminator model may provide a binary output, where a binary output of “1” may indicate that the output generated by the generator model is authentic, and a binary output of “0” may indicate that the output generated by the generator model is fake. The generator model (i.e., the third ML model 110C) may be retrained based on whether the output generated by the generator model is authentic or fake. The classifier model of the present disclosure can classify whether the transformed audio 310A is authentic or fake based on the application of the source speaker classifier 112A and the source emotion classifier 112B to the transformed audio 310A. During the training phase, the outputs of the classifiers (i.e., the source speaker classifier 112A and the source emotion classifier 112B) can be used to calculate an adversarial source classifier loss. The classifiers (i.e., the source speaker classifier 112A and the source emotion classifier 112B) enable the generator model to generate samples that include a target speaker style and a target emotion style. However, an adversarial network (e.g., including a classifier model) can also be omitted during the inference phase.

[0074] The third ML model 110C (i.e., the generator model) can be retrained based on the output of the discriminator model (i.e., the classifier). For example, the third ML model 110C can be retrained using the adversarial loss calculated using the classifier during the training phase. The reference speaker audio 114B and the reference emotion audio 114C can be associated with a speaker identity and emotion pair that is invisible (to the generator model and the discriminator model). Thus, the transformed audio 310A generated based on the source audio 114A, the reference speaker audio 114B, and the reference emotion audio 114C may be classified as fake based on application of the discriminator model to the transformed audio 310A. To mitigate such labeling (e.g., fake) by the discriminator model, a fake pair masking strategy can be employed. A fake pair masking strategy can involve masking the output of the third ML model 110C (i.e., the generator model) for retraining the adversarial model 112C (i.e., the classifier model) in scenarios where the transformed audio (produced as the output of the generator or third ML model 110C) is associated with unseen speaker-emotion pairs. Fake masking can be involved because the classifier model is blind to the speaker identity-emotion pairs associated with the transformed audio 310A. Thus, the classifier model can be applied to the fake-masked transformed audio 310A for retraining the classifier model.

[0075] The circuit 202 may be further configured to apply a classifier model to the generated transformed audio 310A based on determining that the speaker identity-emotion pair associated with the transformed audio 310A is a visible pair to the classifier model. The classifier model may then be applied to the generated transformed audio 310A to determine whether the generated transformed audio 310A is authentic or fake. Based on the output of the classifier model, the third ML model 110C may be retrained.

[0076] At 316, an input audio receiving operation can be performed. The circuit 202 can be configured to receive input audio 316A. The input audio 316A can be associated with a first user, such as user 116. Here, it may be necessary to transform the identity and / or emotion associated with the linguistic content of the input audio 316A after retraining the adversarial model 112C. In one example, the input audio 316A can be recorded via a voice recording device and transmitted to the electronic device 102. In another example, the input audio 316A can be pre-recorded and stored in the database 106 or memory 204. The input audio 316A can be retrieved from the database 106 or memory 204.

[0077] An output audio determination operation may be performed at 318. The third ML model 110C may be configured to convert the input audio 316A associated with the first user into output audio 318A associated with the identity of the second user and the emotion of the third user. Referring to FIG. 3B, the output audio 318A may be associated with the identity of the second user, such as "Mark," and further associated with the emotion of "anger" of a third user, such as "Tom." That is, the output audio 318A may include the audio content of the input audio 316A, but be the voice of the second user, "Mark," with the emotion of "anger" of "Tom."

[0078] In some embodiments, the output audio 318A may correspond to a non-human voice. In one example, the input audio 316A may be a human voice associated with a first user (e.g., "John"), and the output audio 318A may be audio associated with the identity of a "dog" (which may be a second user). Further, the output audio 318A may be associated with the emotion of "anger" of a third user, such as "Mark."

[0079] In one embodiment, the input audio 316A can be a doorbell sound, and the output audio 318A can be a human voice. In one example, the doorbell sound can be the phrase "please open the door." The input audio 316A can be converted to output audio 318A, which can be associated with the identity of a second user, "Mark." The output audio 318A can further be associated with the emotion of "anger" of a third user, "Tom." That is, the output audio 318A can be the phrase "please open the door," which sounds like the second user, "Mark," is speaking in an angry tone. In this manner, the doorbell sound can also be converted to output audio 318A, which can be a human voice.

[0080] In one embodiment, the input audio 316A can be a human voice and the output audio 318A can be a doorbell sound. In one example, the input audio 316A can include the phrase "open the door" in the voice of a first user, "Mary." This input audio 316A can be converted to output audio 318A, which can be associated with a doorbell sound.

[0081] The disclosed electronic device 102 can convert source audio into transformed audio and use the audio to train a generator model (i.e., the third ML model 110C) and a discriminator model (i.e., the adversarial model 112C). The audio can be associated with a speaker-emotion pair that is unseen to the discriminator model. Furthermore, the electronic device 102 can simultaneously transform the speaker identity and emotional style associated with the source audio through virtual domain mixing using the third ML model 110C. Based on a determination that the speaker identity-emotion pair associated with the transformed audio is an unseen pair, the electronic device 102 can employ fake pair masking to prevent the discriminator model from being applied to the transformed audio generated by the generator model. The electronic device 102 can apply an annealing model 112D to a fundamental frequency loss and a norm consistency loss associated with the source audio to determine a weight set associated with the determined fundamental frequency loss and the determined norm consistency loss. The third ML model 110C can be retrained further based on the determined weight set.

[0082] Figure 4 illustrates an exemplary scenario for ML-based intra-audio emotion and voice transformation using virtual domain mixing and fake pair masking, according to an embodiment of the present disclosure. Figure 4 is described in relation to elements in Figures 1, 2, 3A, and 3B. Figure 4 illustrates an exemplary scenario 400. The exemplary scenario 400 includes the first ML model 110A, the second ML model 110B, the third ML model 110C, the source speaker classifier 112A, the source emotion classifier 112B, the adversarial model 112C, the source audio 114A, the reference speaker audio 114B, and the reference emotion audio 114C of Figure 1. Also shown are a first domain code 402, a second domain code 404, transformed audio 406, a classifier model 408, a first domain code 410, a second domain code 412, an adder 414, and a fake pair masking element 416. The adversarial model 112C may include a classifier model 408, a first domain code 410, a second domain code 412, and a fake pair masking element 416. A series of operations related to the scenario 400 will now be described.

[0083] 4, the reference speaker audio 114B and the first domain code 402 may be provided as input to the first ML model 110A. Based on application of the first ML model 110A to the reference speaker audio 114B and the first domain code 402, a speaking style code "h" associated with the reference speaker audio 114B may be determined. spk Further, the reference emotional audio 114C and the second domain code 404 may be provided as inputs to the second ML model 110B. Based on application of the second ML model 110B to the reference emotional audio 114C and the second domain code 404, an emotional style code "h" associated with the reference emotional audio 114C may be determined. emo Then, adder 414 determines the speaker style code "h spk " and the emotional style code "h emo " can be determined.

[0084] Referring to FIG. 4, source audio 114A and speaker style code “h spk " and the emotional style code "h emo " may be provided as an input to a third ML model 110C. Based on application of the third ML model 110C, a transformed audio 406 may be determined. The transformed audio 406 and the first domain code 410 may be provided as inputs to a source speaker classifier 112A to classify whether an identity associated with the transformed audio 406 belongs to a second user. In an embodiment, the first domain code 410 may be similar to the first domain code 402. The transformed audio 406 and the second domain code 412 may be provided as inputs to a source emotion classifier 112B to classify whether an emotion associated with the transformed audio 406 is an emotion of a third user. In an embodiment, the second domain code 412 may be similar to the second domain code 404.

[0085] Further, the transformed audio 406 may be provided to a fake pairs masking element 416. The circuit 202 may determine whether the identity-emotion pair associated with the transformed audio 406 is a visible pair to the classifier model 408. If the identity-emotion pair is a visible pair, the fake pairs masking element 416 may provide the transformed audio 406 to the classifier model 408. The classifier model 408 may determine whether the transformed audio 406 is authentic or fake. Based on the output of the classifier model 408, the third ML model 110C and the classifier model 408 may be retrained.

[0086] On the other hand, if the identity-emotion pair associated with the transformed audio 406 is an unseen pair, the fake pair masking element 416 can mask the identity-emotion pair without providing the transformed audio 406 to the classifier model 408.

[0087] It should be noted that scenario 400 of FIG. 4 is for illustrative purposes and should not be construed as limiting the scope of the present disclosure.

[0088] FIG. 5 illustrates an exemplary scenario for applying a third ML model, according to an embodiment of the present disclosure. FIG. 5 is described with reference to elements in FIGS. 1, 2, 3A, 3B, and 4. FIG. 5 illustrates an exemplary scenario 500. The exemplary scenario 500 includes the source audio 114A, the first ML model 110A, and the third ML model 110C of FIG. 1. Also shown are reference audio 502, a fundamental frequency network 504, a fundamental frequency vector 506, a domain code 508, a style vector 510, and transformed audio 512. The third ML model 110C may include an encoder model 514, an encoding vector 516, a summer 518, a first vector 520, and a decoder model 522. A series of operations related to the scenario 500 will now be described.

[0089] 5, reference audio 502 and domain code 508 may be provided as input to first ML model 110A. In one embodiment, reference audio 502 may be voice content associated with the identity and emotion (neutral or non-neutral) of a second user. Reference audio 502 may correspond to a user identity and a non-neutral (or neutral) emotion spectrogram associated with the second user. A style vector 510 may be obtained based on application of first ML model 110A to reference audio 502 and domain code. Here, style vector 510 may be associated with the speaker identity and emotion of the second user.

[0090] Further, an encoder model 514 may be applied to the received source audio 114A. An encoding vector 516 may be determined based on the application of the encoder model 514. The source audio 114A may also be provided as an input to a fundamental frequency network 504. A fundamental frequency vector 506 may be determined based on the application of the fundamental frequency network 504 to the received source audio 114A. The determined encoding vector 516 and the fundamental frequency vector 506 may be provided as input to an adder 518 to determine a first vector 520. A decoder model 522 may be applied to the determined first vector 520 and the determined style vector 510. A transformed audio 512 may be obtained based on the application of the decoder model 522. Here, the transformed audio 512 may be associated with the linguistic content of the source audio 114A, the emotion of the second user, and the identity of the second user having the reference audio 502. In this manner, the scenario 500 may be applicable in a situation where the reference audio 502 is associated with a non-neutral emotion.

[0091] It should be noted that the scenario 500 of FIG. 5 is for illustrative purposes and should not be construed as limiting the scope of the present disclosure.

[0092] FIG. 6 is a flowchart illustrating the operation of an exemplary method for machine learning (ML)-based emotion and voice transformation in audio using virtual domain mixing and fake pair masking, according to an embodiment of the present disclosure. The description of FIG. 6 is provided with reference to elements in FIGS. 1, 2, 3A, 3B, 4, and 5. FIG. 6 illustrates a flowchart 600. The flowchart 600 may include operations 602-616 and may be performed by the electronic device 102 of FIG. 1 or the circuit 202 of FIG. 2. The flowchart 600 may start at 602 and proceed to 604.

[0093] At 604, source audio 114A associated with the first user can be received. The circuit 202 can be configured to receive the source audio 114A associated with the first user. Details regarding receiving the source audio 114A are further shown, for example, in FIG. 3A (302).

[0094] At 606, reference speaker audio 114B associated with the second user can be received. The circuit 202 can be configured to receive the reference speaker audio 114B associated with the second user. Details regarding receiving the reference speaker audio 114B are further shown, for example, in FIG. 3A (304).

[0095] At 608, reference emotional audio 114C associated with the third user may be received. The circuit 202 may be configured to receive the reference emotional audio 114C associated with the third user. Details regarding receiving the reference emotional audio 114C are further shown, for example, in FIG. 3A (306).

[0096] At 610, a machine learning (ML) model set 110 can be applied to the received source audio 114A, the received reference speaker audio 114B, and the received reference emotional audio 114C. The circuit 202 can be configured to apply the ML model set 110 to the received source audio 114A, the received reference speaker audio 114B, and the received reference emotional audio 114C. Details regarding the application of the ML model set 110 are further shown, for example, in FIG. 3A (308).

[0097] At 612, transformed audio 310A can be generated based on application of the ML model set 110, the transformed audio 310A can be associated with the content of the source audio, the identity of the second user, and the emotion of the third user. The circuit 202 can be configured to generate transformed audio 310A can be associated with the content of the source audio, the identity of the second user, and the emotion of the third user based on application of the ML model set 110. Details regarding generating the transformed audio are further shown, for example, in FIG. 3A (310).

[0098] At 614, each of the source speaker classifier 112A and the source emotion classifier 112B can be applied to the generated transformed audio 310A. The circuit 202 can be configured to apply each of the source speaker classifier 112A and the source emotion classifier 112B to the generated transformed audio 310A. Details regarding the application of the source speaker classifier 112A and the source emotion classifier 112B are further shown, for example, in FIG. 3B (312).

[0099] At 616, the adversarial model 112C may be retrained based on application of each of the source speaker classifier 112A and the source emotion classifier 112B, and based on the retraining, the input audio 316A associated with the first user may be converted into output audio 318A associated with the identity of the second user and the emotion of the third user. The circuit 202 may be configured to retrain the adversarial model 112C based on application of each of the source speaker classifier 112A and the source emotion classifier 112B, and based on the retraining, the input audio 316A associated with the first user may be converted into output audio 318A associated with the identity of the second user and the emotion of the third user. Details regarding the retraining of the adversarial model 112C are further shown, for example, in FIG. 3B (314). Control may proceed to end.

[0100] Although flowchart 600 is depicted as discrete operations such as 604, 606, 608, 610, 612, 614, and 616, the disclosure is not so limited. Thus, in some embodiments, such discrete operations may be further divided into additional operations, combined into fewer operations, or eliminated, depending on the implementation, without departing from the essence of the disclosed embodiments.

[0101] Various embodiments of the present disclosure may provide a non-transitory computer-readable medium and / or storage medium having stored thereon computer-executable instructions executable by a machine and / or computer to operate an electronic device (e.g., electronic device 102 of FIG. 1 ). Such instructions may cause electronic device 102 to perform operations that may include receiving source audio (e.g., source audio 114A) associated with a first user. The operations may further include receiving reference speaker audio (e.g., reference speaker audio 114B) associated with a second user. The operations may further include receiving reference emotional audio (e.g., reference emotional audio 114C) associated with a third user. The operations may further include applying a machine learning (ML) model set (e.g., ML model set 110) to the received source audio 114A, the received reference speaker audio 114B, and the received reference emotional audio 114C. The operations may further include generating transformed audio (e.g., transformed audio 310A) that can be associated with the content of the source audio, the identity of the second user, and the emotion of the third user based on application of the ML model set 110. The operations may further include applying each of a source speaker classifier (e.g., source speaker classifier 112A) and a source emotion classifier (e.g., source emotion classifier 112B) to the generated transformed audio 310A. The operations may further include retraining an adversarial model (e.g., adversarial model 112C) based on application of each of the source speaker classifier 112A and the source emotion classifier 112B. Based on the retraining, input audio associated with the first user (e.g., input audio 316A) can be transformed into output audio (e.g., output audio 318A) that is associated with the identity of the second user and the emotion of the third user.

[0102] An exemplary aspect of the present disclosure may provide an electronic device (such as the electronic device 102 of FIG. 1 ) including a circuit (such as a circuit 202). The circuit 202 may be configured to receive source audio 114A associated with a first user. The circuit 202 may be configured to receive reference speaker audio 114B associated with a second user. The circuit 202 may be configured to receive reference emotion audio 114C associated with a third user. The circuit 202 may be configured to apply an ML model set 110 to the received source audio 114A, the received reference speaker audio 114B, and the received reference emotion audio 114C. The circuit 202 may be configured to generate transformed audio 310A that can be associated with the content of the source audio 114A, the identity of the second user, and the emotion of the third user based on the application of the ML model set 110. The circuit 202 may be configured to apply each of a source speaker classifier 112A and a source emotion classifier 112B to the generated transformed audio 310A. The circuit 202 can be configured to retrain the adversarial model 112C based on application of each of the source speaker classifier 112A and the source emotion classifier 112B, and based on the retraining, can convert input audio 316A associated with a first user into output audio 318A associated with the identity of a second user and the emotion of a third user.

[0103] In an embodiment, source audio 114A may correspond to a neutral emotion spectrogram associated with a first user, reference speaker audio 114B may correspond to a user identity spectrogram associated with a second user, and reference emotion audio 114C may correspond to a non-neutral emotion spectrogram associated with a third user.

[0104] In one embodiment, the circuit 202 may be further configured to apply a first ML model 110A of the ML model set 110 to the received reference speaker audio 114B and a first domain code (e.g., first domain code 402) associated with the received reference speaker audio 114B. The circuit 202 may be further configured to determine a speaker style code associated with the received reference speaker audio 114B based on the application of the first ML model 110A. The circuit 202 may be further configured to apply a second ML model 110B of the ML model set 110 to the received reference emotional audio 114C and a second domain code (e.g., second domain code 404) associated with the received reference emotional audio 114C. The circuit 202 may be further configured to determine an emotional style code associated with the received reference emotional audio 114C based on the application of the second ML model 110B. The circuit 202 may be further configured to apply a third ML model 110C of the ML model set 110 to the received source audio 114A, the determined speaker style code, and the determined emotional style code, and generation of the transformed audio 310A may be further based on application of the third ML model 110C.

[0105] In one embodiment, the first ML model 110A can be a speaker style encoder model.

[0106] In one embodiment, the second ML model 110B can be an emotional style encoder model.

[0107] In one embodiment, the third ML model 110C may correspond to the generator model.

[0108] In one embodiment, the generator model may include an encoder model (eg, encoder model 514), an adder model, and a decoder model (eg, decoder model 522).

[0109] In one embodiment, the circuit 202 may be further configured to apply an encoder model 514 to the received source audio 114A. The circuit 202 may be further configured to determine an encoding vector (e.g., encoding vector 516) based on application of the encoder model 514. The circuit 202 may be further configured to apply a fundamental frequency network (e.g., fundamental frequency network 504) to the received source audio 114A. The circuit 202 may be further configured to determine a fundamental frequency (e.g., fundamental frequency vector 506) based on application of the fundamental frequency network 504. The circuit 202 may be further configured to apply an adder model (e.g., adder 518 of FIG. 5 ) to the determined encoding vector 516 and the determined fundamental frequency vector 506. The circuit 202 may be further configured to determine a first vector (e.g., first vector 520) based on application of the adder model. The circuit 202 can be further configured to apply a decoder model 522 to the determined first vector 520, the determined speaker style code, and the determined emotional style code, and generation of the transformed audio 310A can be further based on application of the decoder model 522.

[0110] In one embodiment, the adversarial model 112C may include a classifier model (eg, classifier model 408).

[0111] In an embodiment, the circuit 202 may be further configured to apply a classifier model 408 to the generated transformed audio 406 based on determining that the reference speaker audio 114B and the reference emotional audio 114C correspond to a visible pair.

[0112] In an embodiment, the circuit 202 can be further configured to determine a fundamental frequency loss and a norm consistency loss associated with the generated transformed audio 310. The circuit 202 can be further configured to apply an annealing model (e.g., the annealing model 112D) to the determined fundamental frequency loss and the determined norm consistency loss. The circuit 202 can be further configured to determine a weight set associated with the determined fundamental frequency loss and the determined norm consistency loss based on the application of the annealing model 112D, and the retraining of the third ML model 110C can be further based on the determined weight set.

[0113] In some embodiments, the output audio 318A may correspond to a non-human voice.

[0114] In one embodiment, the input audio 316A may relate to a doorbell sound and the output audio 318A may correspond to a human voice.

[0115] In one embodiment, the input audio 316A may relate to a human voice and the output audio 318A may correspond to a doorbell sound.

[0116] The present disclosure may also be located in a computer program product, which includes all features enabling the implementation of the methods described herein and which is capable of executing these methods when loaded into a computer system. A computer program in this context means any expression, in any language, code or notation, of a set of instructions intended to cause a system having information processing capabilities to perform a particular function, either directly, or after a) conversion into another language, code or notation, or b) reproduction in a different content form, or both.

[0117] While the present disclosure has been described with reference to several embodiments, those skilled in the art will recognize that various modifications may be made and equivalents may be substituted without departing from the scope of the disclosure. Additionally, many modifications may be made to adapt a particular situation or material to the teachings of the disclosure without departing from the scope of the disclosure. Therefore, it is not intended that the disclosure be limited to the particular embodiments disclosed, but rather, it is intended to include all embodiments falling within the scope of the appended claims.

Claims

1. 1. An electronic device comprising: receiving source audio associated with a first user; receiving a reference speaker audio associated with a second user; receiving reference emotional audio associated with a third user; applying a set of machine learning (ML) models to the received source audio, the received reference speaker audio, and the received reference emotion audio; generating transformed audio related to the content of the source audio and the identity of the second user and further related to an emotion corresponding to the third user based on application of the set of ML models; applying each of a source speaker classifier and a source emotion classifier to the generated transformed audio; retraining an adversarial model based on application of each of the source speaker classifier and the source emotion classifier; a circuit configured as follows: input audio associated with the first user is transformed into output audio associated with the identity of the second user and associated with the emotion of the third user based on the retraining. An electronic device characterized by:

2. the source audio corresponds to a neutral emotional source spectrogram associated with the first user; the reference speaker audio corresponds to a neutral emotional user identity spectrogram associated with the second user; the reference emotional audio corresponds to a non-neutral emotional spectrogram associated with the third user. The electronic device of claim 1 .

3. The circuit comprises: applying a first ML model of the set of ML models to the received reference speaker audio and a first domain code associated with the received reference speaker audio; determining a speaker style code associated with the received reference speaker audio based on application of the first ML model; applying a second ML model from the set of ML models to the received reference emotional audio and a second domain code associated with the received reference emotional audio; determining an emotional style code associated with the received reference emotional audio based on application of the second ML model; applying a third ML model of the set of ML models to the received source audio, the determined speaker style code, and the determined emotional style code; and wherein generating the transformed audio is further based on applying the third ML model. The electronic device of claim 1 .

4. the first ML model is a speaker style encoder model; The electronic device of claim 3 .

5. the second ML model is an emotional style encoder model; The electronic device of claim 3 .

6. the third ML model corresponds to a generator model; The electronic device of claim 3 .

7. the generator model includes an encoder model, an adder model, and a decoder model; 7. The electronic device of claim 6.

8. The circuit comprises: applying the encoder model to the received source audio; determining an encoding vector based on application of the encoder model; applying a fundamental frequency network to the received source audio; determining a fundamental frequency based on application of the fundamental frequency network; applying the adder model to the determined coding vector and the determined fundamental frequency; determining a first vector based on application of the adder model; applying the decoder model to the determined first vector, the determined speaker style code, and the determined emotional style code. and wherein generating the transformed audio is further based on applying the decoder model.

8. The electronic device of claim 7.

9. the adversarial model includes a classifier model; The electronic device of claim 1 .

10. the circuitry is further configured to apply the classifier model to the generated transformed audio based on a determination that the reference speaker audio and the reference emotion audio correspond to a visible pair.

10. The electronic device of claim 9.

11. The circuit comprises: determining a fundamental frequency loss and a norm consistency loss associated with the generated transformed audio; applying an annealing model to the determined fundamental frequency loss and the determined norm consistency loss; determining a set of weights associated with the determined fundamental frequency loss and the determined norm consistency loss based on application of the annealing model; wherein retraining the third ML model is further based on the determined weight set. The electronic device of claim 1 .

12. the output audio corresponds to a non-human voice. The electronic device of claim 1 .

13. the input audio relates to a doorbell sound and the output audio corresponds to a human voice; The electronic device of claim 1 .

14. the input audio content relates to a human voice and the output audio corresponds to a doorbell sound; The electronic device of claim 1 .

15. In an electronic device, receiving source audio associated with a first user; receiving a reference speaker audio associated with a second user; receiving reference emotional audio associated with a third user; applying a set of machine learning (ML) models to the received source audio, the received reference speaker audio, and the received reference emotion audio; generating transformed audio associated with the identity of the second user and further associated with an emotion corresponding to the third user based on application of the set of ML models; applying each of a source speaker classifier and a source emotion classifier to the generated transformed audio; retraining an adversarial model based on application of each of the source speaker classifier and the source emotion classifier; and wherein input audio associated with the first user is transformed into output audio associated with the identity of the second user and associated with the emotion of the third user based on the retraining. A method characterized by:

16. the source audio corresponds to a neutral emotional source spectrogram associated with the first user; the reference speaker audio corresponds to a neutral emotional user identity spectrogram associated with the second user; the reference emotional audio corresponds to a non-neutral emotional spectrogram associated with the third user.

16. The method of claim 15.

17. The circuit comprises: applying a first ML model of the set of ML models to the received reference speaker audio and a first domain code associated with the received reference speaker audio; determining a speaker style code associated with the received reference speaker audio based on application of the first ML model; applying a second ML model from the set of ML models to the received reference emotional audio and a second domain code associated with the received reference emotional audio; determining an emotional style code associated with the received reference emotional audio based on application of the second ML model; applying a third ML model of the set of ML models to the received source audio, the determined speaker style code, and the determined emotional style code; and wherein generating the transformed audio is further based on applying the third ML model.

16. The method of claim 15.

18. the first ML model is a speaker style encoder model; 18. The method of claim 17.

19. the second ML model is an emotional style encoder model; 18. The method of claim 17.

20. A non-transitory computer-readable medium having computer-executable instructions stored thereon, the instructions, when executed by an electronic device, receiving source audio associated with a first user; receiving a reference speaker audio associated with a second user; receiving reference emotional audio associated with a third user; applying a set of machine learning (ML) models to the received source audio, the received reference speaker audio, and the received reference emotion audio; generating transformed audio associated with the identity of the second user and further associated with an emotion corresponding to the third user based on application of the set of ML models; applying each of a source speaker classifier and a source emotion classifier to the generated transformed audio; retraining an adversarial model based on application of each of the source speaker classifier and the source emotion classifier; wherein input audio associated with the first user is converted to output audio associated with the identity of the second user and associated with the emotion of the third user based on the retraining.

1. A non-transitory computer-readable medium comprising:

Citation Information

Patent Citations

  • Data conversion leaning apparatus, data conversion apparatus, method and program

    JP2020140244A