Methods and systems for assigning unique voices to electronic devices

By processing user commands through NLP, speaker embeddings that are infused with emotion are generated and injected, solving the problem of the lack of emotion in the voice of virtual assistant devices and realizing emotional interaction and intimacy between the device and the user.

CN116391225BActive Publication Date: 2026-04-03SAMSUNG ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-07
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

The voices generated by existing virtual assistant devices lack emotional connection, resulting in a lack of emotional interaction between users and devices.

Method used

The user's commands are preprocessed using Natural Language Processing (NLP), categorizing utterance and non-utterance information, acquiring device-specific information, generating contextual parameters, selecting speaker embeddings stored in the database, and outputting the results to the user for playback, while injecting human emotions to generate unique speech.

Benefits of technology

It enables emotional interaction between users and devices, enhancing users' intimacy and emotional connection with their devices by assigning unique voices to each device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116391225B_ABST
    Figure CN116391225B_ABST
Patent Text Reader

Abstract

This invention discloses a method in an interactive computing system, comprising: preprocessing input natural language (NL) from a user command based on natural language processing (NLP) for classifying utterance and non-utterance information; obtaining NLP results from the user command; acquiring device-specific information from one or more IoT devices operating in the environment based on the NLP results; generating one or more context parameters based on the NLP results and the device-specific information; selecting at least one speaker embedding stored in a database for one or more IoT devices based on the one or more context parameters; and outputting the selected at least one speaker embedding for playback to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to generating voice messages for electronic devices. More specifically, this disclosure relates to assigning unique voice to one or more electronic devices, and thereby generating the assigned voice for playback to a user. Background Technology

[0002] With the latest developments in electronic devices, interactive computer systems, such as virtual assistant devices, have evolved to perform various activities to assist users. A virtual assistant device is a device that typically communicates with multiple IoT (Internet of Things) devices to control them. In examples, a virtual assistant device may operate as a standalone device or may be provided as an application or as an extended browser running on computing devices included in cloud networks or embedded systems. Without departing from the scope of this disclosure, the terms "interactive computer system" and "virtual assistant device" may be used interchangeably below.

[0003] Virtual assistant devices assist users in various ways, such as monitoring the home, managing home IoT devices, and providing suggestions for queries. For example, digital appliances (DAs) like smart refrigerators, smart TVs, and smart washing machines use voice assistants to interact with users and answer their commands. For assigned tasks, the voice assistant responds to the user using speech generated by a text-to-speech (TTS) generation system. Currently, the speech or utterance generated by each smart device sounds similar because the TTS system is not device-specific. These virtual assistant devices have similar robotic voice responses. For example, the current state of technology can at best offer users the option to choose a Hollywood character's voice as their voice assistant. However, there is no emotional connection between the user and the voice assistant device. Therefore, the speech generated by this computer interaction system is inherently robotic and lacks emotion. This results in a lack of emotional connection between the user and the voice assistant device. If all digital devices had their own unique voices, users would feel a closer connection to the devices, just as they would with other people / animals.

[0004] Therefore, a solution is needed to overcome the above-mentioned defects. Summary of the Invention

[0005] Technical issues

[0006] There is no emotional connection between the user and the voice assistant device. Therefore, the voice produced by this computer interaction system is inherently robotic and lacks emotion. This results in a lack of emotional connection between the user and the voice assistant device.

[0007] Technical solution

[0008] This overview is provided to present the selection of concepts in a simplified form, which will be further described in the detailed description of the inventive concepts. This summary is not intended to identify key or essential inventive concepts of this disclosure, nor is it intended to define the scope of the inventive concepts.

[0009] According to embodiments of this disclosure, a method in an interactive computing system includes: preprocessing input natural language (NL) from a user command based on natural language processing (NLP) for classifying utterance and non-utterance information; obtaining NLP results from the user command; acquiring device-specific information from one or more IoT devices operating in an environment based on the NLP results; generating one or more context parameters based on the NLP results and the device-specific information; selecting at least one speaker embedding stored in a database for one or more IoT devices based on the one or more context parameters; and outputting the selected at least one speaker embedding for playback to a user.

[0010] The method may further include: processing at least one selected speaker embedding based on a text-to-speech mechanism, and replaying it to the user in natural language.

[0011] Obtaining NLP results can include acquiring acoustic information that forms part of non-verbal information, in order to generate one or more contextual parameters based on the acoustic information.

[0012] Generating one or more context parameters may include selecting a speech embedding from at least one speaker embedding for playback by IoT devices in one or more IoT devices.

[0013] Preprocessing may include: the step of collecting multiple audio samples from multiple speakers to generate one or more speaker embeddings; and the step of storing the generated one or more speaker embeddings in a database for selecting one or more speaker embeddings for one or more IoT devices.

[0014] The method may further include artificially injecting human emotions into one or more generated speaker embeddings, wherein the generated one or more speaker embeddings include different types of tone and texture for multiple audio samples; and storing the generated one or more speaker embeddings with human emotions in a database.

[0015] Preprocessing may include: extracting device-specific information of one or more IoT devices operating in a network environment; associating the device-specific information with at least one speaker embedding to generate a mapping of speaker embedding sets for each of the one or more IoT devices based on the device-specific information; assigning different speech from the speaker embedding sets to each of the one or more IoT devices based on relevance; and storing the assigned speaker embedding sets in a database as a mapping for selecting encoded speaker embeddings for one or more IoT devices.

[0016] Preprocessing may include: extracting device-specific information of one or more IoT devices operating in a network environment; identifying similar IoT devices from the one or more IoT devices operating in the environment based on the device-specific information; associating the device-specific information with at least one speaker embedding to generate a mapping of speaker embedding sets for each of the identified similar IoT devices based on the device-specific information; assigning different voices to each of the identified similar IoT devices included in the speaker embedding sets; and storing the assigned speaker embedding sets in a database as a mapping for selecting speaker embeddings for one or more IoT devices.

[0017] Obtaining NLP results may include determining one of the following: success, failure, or follow-up of an IoT event on one or more IoT devices due to a user command or a device-specific event, wherein the determined result corresponds to an NLP result.

[0018] Obtaining acoustic information from non-verbal information can include identifying one or more audio events in the surrounding environment of one or more IoT devices, wherein the one or more audio events in the surrounding environment correspond to acoustic information.

[0019] According to another embodiment of this disclosure, an interactive computing system includes: one or more processors; and a memory configured to store instructions executable by the one or more processors; and wherein the one or more processors are configured to: preprocess input natural language (NL) from a user command based on natural language processing (NLP) for classifying utterance and non-utterance information; obtain NLP results from the user command; obtain device-specific information from one or more IoT devices operating in an environment based on the NLP results; generate one or more context parameters based on the NLP results and the device-specific information; select at least one speaker embedding stored in a database for one or more IoT devices based on the one or more context parameters; and output the selected at least one speaker embedding for playback to a user.

[0020] One or more processors may also be configured to process at least one selected speaker embedding based on a text-to-speech mechanism for playback to the user in natural language.

[0021] One or more processors may also be configured to: acquire acoustic information, including acquiring a portion of non-verbal information, to generate one or more contextual parameters based on the acoustic information.

[0022] One or more processors may also be configured to select a voice embedding from at least one speaker embedding for playback by an IoT device in one or more IoT devices.

[0023] One or more processors may also be configured to: collect multiple audio samples from multiple speakers to generate one or more speaker embeddings; store the generated one or more speaker embeddings with human emotions in a database for selecting one or more speaker embeddings for one or more IoT devices.

[0024] One or more processors may also be configured to: artificially inject human emotions into one or more generated speaker embeddings, wherein the generated one or more speaker embeddings include different types of pitch and texture for multiple audio samples; and store the generated one or more speaker embeddings with human emotions in a database.

[0025] One or more processors may also be configured to: extract device-specific information of one or more IoT devices operating in a network environment; associate the device-specific information with at least one speaker embedding to generate a mapping of speaker embedding sets for each of the one or more IoT devices based on the device-specific information; assign distinct speech from the speaker embedding sets to each of the one or more IoT devices based on relevance; and store the assigned speaker embedding sets in a database as a mapping for selecting encoded speaker embeddings for one or more IoT devices.

[0026] One or more processors may also be configured to: extract device-specific information of one or more IoT devices operating in a network environment; identify similar IoT devices from the one or more IoT devices operating in the environment based on the device-specific information; associate the device-specific information with at least one speaker embedding to generate a mapping of speaker embedding sets for each of the identified similar IoT devices based on the device-specific information; assign different voices to each of the identified similar IoT devices included in the speaker embedding sets; and store the assigned speaker embedding sets in a database as a mapping for selecting speaker embeddings for one or more IoT devices.

[0027] One or more processors may also be configured to determine one of the success, failure, or follow-up of an IoT event on one or more IoT devices due to a user command or a device-specific event, wherein the determined result corresponds to an NLP result.

[0028] One or more processors may also be configured to determine one or more audio events of the surrounding environment of one or more IoT devices, wherein the one or more audio events of the surrounding environment correspond to acoustic information.

[0029] Beneficial effects

[0030] According to embodiments of this disclosure, a method in an interactive computing system includes preprocessing input natural language (NL) from a user command based on natural language processing (NLP) for classifying utterance and non-utterance information, obtaining NLP results from the user command, acquiring device-specific information from one or more IoT devices operating in the environment based on the NLP results, generating one or more context parameters based on the NLP results and the device-specific information, selecting at least one embedded speaker stored in a database for one or more IoT devices based on the one or more context parameters, and outputting the selected at least one speaker embedding for playback to a user. Attached Figure Description

[0031] These and other features, aspects, and advantages of this disclosure will be better understood when the following detailed description is read with reference to the accompanying drawings, in which the same characters denote the same parts, wherein:

[0032] Figure 1 A network environment 100 is shown for an implementation of a system 101 for providing voice assistance to a user, according to an embodiment of the present disclosure.

[0033] Figure 2 A schematic block diagram of a system 101 for providing device operation notifications to a user according to an embodiment of the present disclosure is shown.

[0034] Figure 3 A detailed implementation of a component for generating a speaker embedding 300 according to an embodiment of the present disclosure is shown.

[0035] Figure 4 A detailed implementation of a component for artificially injecting human emotions into a speaker embedding, according to an embodiment of the present disclosure, is shown.

[0036] Figure 5 A detailed implementation of the component for generating a mapping of speaker embeddings according to this disclosure is shown.

[0037] Figure 6 Detailed implementations of components / modules for dynamically selecting speaker embeddings according to embodiments of this disclosure are shown.

[0038] Figure 7 An exemplary scenario is shown in an interactive computing system for presenting voice assistance, according to embodiments of the present disclosure.

[0039] Figure 8 A flowchart illustrating an interactive computing system for presenting voice assistance according to an embodiment of the present disclosure is shown.

[0040] Figure 9 A sequence flow for presenting a voice-assisted interactive computing system according to an embodiment of the present disclosure is shown.

[0041] Figure 10 A flowchart illustrating a system for presenting voice assistance according to an embodiment of the present disclosure is shown.

[0042] Figure 11 The sequence flow of a system for presenting voice assistance is shown according to an embodiment of the present disclosure.

[0043] Figures 12-18 Various use cases of system-based implementations according to embodiments of this disclosure are illustrated.

[0044] Furthermore, those skilled in the art will understand that the elements in the accompanying drawings are shown for simplicity and may not necessarily be drawn to scale. For example, flowcharts illustrate methods based on the most prominent steps involved to aid in understanding various aspects of this disclosure. Additionally, depending on the construction of the device, one or more components of the device may be represented in the drawings using conventional symbols, and the drawings may show only those specific details relevant to understanding embodiments of this disclosure, so as not to obscure details that would be readily apparent to those of ordinary skill in the art to benefit from the description herein. Detailed Implementation

[0045] To facilitate an understanding of the principles of the inventive concept, reference will now be made to the embodiments shown in the accompanying drawings, and these embodiments will be described using specific language. However, it should be understood that the scope of the inventive concept is not limited, and such changes and further modifications to the illustrated system, as well as such further applications of the principles of the inventive concept shown therein, are considered to be commonly thought of by those skilled in the art to which the inventive concept pertains.

[0046] Those skilled in the art will understand that the foregoing general description and the following detailed description are explanations of the inventive concept and are not intended to limit the inventive concept.

[0047] In this specification, references to "one aspect," "another aspect," or similar language mean that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment of this disclosure. Therefore, phrases such as "in one embodiment," "in another embodiment," and similar language appearing throughout the specification may, but not necessarily all, refer to the same embodiment.

[0048] The terms “comprising,” “including,” or any other variations thereof are intended to cover non-exclusive inclusion, such that a process or method that includes the list of steps may include not only those steps but also other steps not expressly listed or inherent to the process or method. Similarly, without further constraints, one or more devices or subsystems or elements, structures, or components described by “comprising…a” do not exclude the presence of other devices or other subsystems or other elements, other structures or other components, or additional devices or additional subsystems or additional elements, additional structures, or additional components.

[0049] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the concepts of this invention pertain. The systems, methods, and examples provided herein are illustrative only and not restrictive.

[0050] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0051] Figure 1This illustration shows a network environment 100 for providing voice assistance to a user using a system 101 implemented according to embodiments of the present disclosure. Specifically, the present disclosure provides or assigns a unique voice to each digital device, electronic device, or IoT device connected to system 101. For example, system 101 may include, but is not limited to, computer interaction systems, virtual assistant systems, or electronic systems. Providing a unique or different voice to each digital device, IoT device, or electronic device can be, for example, but not limited to, conversational style attributes in natural human language infused with emotion, dynamic responses of the system in response to user inquiries, or the occurrence of device events infused with emotion, human voice responses, etc. Thus, as an example, consider a scenario where a user gives a command to a smart washing machine to wash clothes by saying “Start washing in quick mode.” System 101 then processes the command and determines the characteristics of the IoT device, such as the type of IoT device. In this case, it will determine whether the IoT device is a washing machine, the capacity of the IoT device, the operating status of the IoT device, and whether the IoT device is overloaded, etc. Therefore, system 101 can assign and generate a unique response. When the washing machine is determined to be overloaded with clothes and has been running for 8 hours, a unique response, such as, but not limited to, uttering the phrase "Start washing in fast mode" in a tired voice, with the groaning of the motor in the background, provides a human-like connection to the user. In contrast, traditional existing voice-assisted devices will provide a robotic-type response without any emotional infusion.

[0052] Examples of environment 100 may include, but are not limited to, homes, offices, buildings, hospitals, schools, public places, etc. that support IoT. In other words, an environment in which IoT devices are implemented can be understood as an example of environment 100. Examples of devices that can implement system 101 include, but are not limited to, smartphones, tablets, distributed computing systems, servers and cloud servers, or dedicated embedded systems.

[0053] like Figure 1 As shown, in the example, system 101 can interact with multiple IoT devices 103-1, 103-2, 103-3, and 103-4. 3 -3 and 10 3-4. The devices 103-1, 103-2, 103-3, and 103-4 can be communicatively coupled. In the example, IoT devices 103-1, 103-2, 103-3, and 103-4 can include, but are not limited to, washing machines, televisions, mobile devices, speakers, refrigerators, air conditioners, heating equipment, monitoring systems, home appliances, alarm systems, and sensors. It is understood that each of the above examples is a smart device because it can connect to one or more remote servers or IoT cloud servers 105 and is part of a network environment such as environment 100. In the example, system 101 can be coupled to one or more remote servers or IoT cloud servers 105 for accessing data, processing information, etc. In one implementation, IoT devices 103-1, 103-2, 103-3, and 103-4 can also be used without connecting to any networked environment, i.e., in offline mode. In a further implementation, the system can be further coupled to any electronic device other than IoT devices. Here, the data can include, but is not limited to, device type, device capabilities, device hardware characteristics, device location in the home IoT setup, IoT device usage history, user history, etc. In the example, the generated data can be stored in a remote or IoT cloud server 105. Furthermore, system 101 can be configured to receive one or more commands as input from the user or other devices near the user. For example, other devices may also include IoT devices. Additionally, system 101 can be configured to receive one or more commands as input, in the form of voice or text, or any IoT event 107, from the user or other devices present near the user, or from any digital device. For example, other devices may also include IoT devices, or smartphones, laptops, tablets, etc. For example, the terms device, electronic device, digital device, and IoT device may be used interchangeably without departing from the scope of the inventive concept.

[0054] Figure 2 A schematic block diagram of a system 101 for providing device action notifications to a user according to an embodiment of the present disclosure is shown. In an example embodiment, system 101 may include a processor 203, a memory 205, a database 207, a module 209, an audio unit 211, a transceiver 213, and an AI module 215. In the example, the memory 205, database 207, module 209, audio unit 211, transceiver 213, and AI module 215 are coupled to the processor 203.

[0055] In the example, processor 203 can be a single processing unit or multiple units, all of which can include multiple computing units. Processor 203 can be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuits, and / or any device that manipulates signals based on operating instructions. Among other capabilities, processor 203 is configured to fetch and execute computer-readable instructions and data stored in memory 205.

[0056] The memory 205 may include any non-transitory computer-readable medium known in the art, including, for example, volatile memory such as static random access memory (SRAM) and dynamic random access memory (DRAM), and / or non-volatile memory such as read-only memory (ROM), erasable programmable ROM, flash memory, hard disk, optical disk and magnetic tape.

[0057] In the example, module 209 may include a program, subroutine, part of a program, software component, or hardware component capable of performing the tasks or functions described herein. As used herein, module 209 may be implemented independently of other modules on a hardware component such as a server, or the module may reside on the same server as other modules, or within the same program. Module 209 may be implemented on a hardware component such as a processor, microprocessor, microcomputer, microcontroller, digital signal processor, central processing unit, state machine, logic circuit, and / or any device that manipulates signals based on operating instructions. When executed by processor 203, module 209 may be configured to perform any of the described functions.

[0058] Database 207 can be implemented using integrated hardware and software. The hardware may include a hardware disk controller with programmable search capabilities or a software system running on general-purpose hardware. Examples of databases include, but are not limited to, in-memory databases, cloud databases, distributed databases, embedded databases, etc. Among other things, database 207 serves as a repository for storing data processed, received, and generated by one or more of processors 203 and module 209.

[0059] Audio unit 211 may include a speaker and / or a microphone to generate audio output. The audio output can be achieved using various technologies such as Automatic Speech Recognition (ASR), Natural Language Understanding (NLU), Natural Language Processing (NLP), and Natural Language Generation (NLG). Audio unit 211 can, based on the receipt of commands, generate text-to-speech operations for the user in natural language related to IoT devices 103-1, 103-2, 103-3, and 103-4 via speaker embedding 218. The generation of speaker embeddings will be explained in detail in the following paragraphs.

[0060] Transceiver 213 can be both a transmitter and a receiver. Transceiver 213 can communicate with users and / or other IoT devices via any wireless standard such as 3G, 4G, 5G, etc., or use other wireless technologies such as Wi-Fi, Bluetooth, etc. Transceiver 213 can be configured to communicate with users and / or other IoT devices via wireless standards such as 3G, 4G, 5G, etc. Figure 1 Other IoT devices shown receive one or more user commands / text commands / IoT events 221 or commands.

[0061] The AI ​​module 215 for speaker embedding 218 for playback to the user may include multiple neural network layers. Examples of neural networks include, but are not limited to, convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), and restricted Boltzmann machines (RBMs). A learning technique is a method for training a predetermined target device (e.g., a robot) using multiple learning data to prompt, allow, or control the target device to make determinations or predictions. Examples of learning techniques include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. At least one of multiple CNN models can be implemented, thereby enabling the execution of the mechanisms of this subject matter through an AI model. AI-related functions can be executed via non-volatile memory, volatile memory, and a processor. The processor may include one or more processors. Here, the one or more processors may be general-purpose processors such as a central processing unit (CPU), application processor (AP), etc., graphics processing units only such as a graphics processing unit (GPU), a vision processing unit (VPU), and / or AI-specific processors such as a neural processing unit (NPU). One or more processors control the processing of input data based on predefined operating rules or artificial intelligence (AI) models stored in non-volatile and volatile memory. These predefined operating rules or AI models are provided through training or learning.

[0062] In an embodiment, system 101 can be configured to preprocess input natural language (NL) from user commands based on natural language processing (NLP). Subsequently, system 101 categorizes utterance and non-utterance information into utterance information. For example, system 101 categorizes commands in the user's utterance voice between surrounding non-utterance information. For example, the command might be "start washing in fast mode," while non-utterance information might include, for example, a noisy environment, a laundry scene, a kitchen chimney, a silent home, television sounds, human voices, pet sounds, etc. Therefore, the categorization between commands and surrounding non-utterance information helps obtain acoustic information to dynamically determine the most suitable voice for the device. Specifically, system 101 can be configured to determine a suitable speaker embedding to generate the most appropriate voice for the device. For example, the speaker embedding is based on the speaker's / user's voice's voiceprint / spectral mapping and / or the emotions expressed during the speech, or the spectrogram of the speaker's / user's voice infused with human emotions (such as happiness, sadness, anxiety, fatigue, etc.). Speaker embeddings can be dynamically generated during runtime and / or can be pre-stored.

[0063] After classification, system 101 can be configured to obtain NLP results from the utterance information. The NLP results can, for example, determine at least one of the following: success, failure, or follow-up of an IoT event on one or more IoT devices due to a user command. For example, success can be understood as the completion of an IoT event, such as the successful completion of laundry. Follow-up can be understood as any IoT event that requires follow-up. After successful laundry completion, television playback will be queued according to a predefined user command. Furthermore, failure of an IoT event can be understood as a capability, connectivity, latency, or memory problem when executing the user command.

[0064] Subsequently, system 101 obtains device-specific information from one or more target devices. Alternatively, the device-specific information can be obtained from a remote server / IoT cloud server 105 or database 207. The device-specific information can be associated with one or more IoT devices or network devices in the environment. The device-specific information may include at least one of the following: device physical characteristics, software characteristics, device type, device capabilities, hardware characteristics, device location in a networked environment, usage history, device settings, and information such as device error states, location, and device events. Without departing from the scope of the inventive concept, the device-specific information may be referred to as characteristics of one or more devices.

[0065] For example, the physical characteristics of a device can include its model, shape, and size. For instance, the shape and size of a device can be categorized as heavy-duty, light-duty, medium-duty, large-sized, medium-sized, or small-sized. For example, heavy-duty equipment can include washing machines and home hubs. Light-duty equipment can include IoT sensors and wearable devices. Medium-duty equipment can include mobile devices and smart speakers.

[0066] Software characteristics can include the software version installed on the device, or whether the device adaptively updates software or has available software functions, supports domains for generating voice commands, etc. Therefore, software characteristics can be used to determine the device's operational status or potential errors. Device types can include, for example, cleaning devices, storage devices, audio / video devices, monitoring devices, cooling devices, static / non-removable devices, removable devices, etc.

[0067] Device capabilities can include loading capacity or user-command-based event processing capabilities. For example, loading capacity could refer to a refrigerator with a 180-liter capacity, a washing machine with a capacity to hold 15kg of clothes, etc. Hardware characteristics can include the size of the device, its relative mobility, its power consumption, and its computing speed. Device location in IoT can include the geographic location of the IoT device. Usage history can include user history data, such as the user's preferred operating mode, the frequency of use of any IoT device, etc. Device settings can include specific settings of the device, such as the washing machine being set to a heavy wash mode, or the air conditioner being set to a medium cooling mode, etc. Furthermore, in the example, device error states could include the washing machine running out of water during a wash or a power outage during the operation of the IoT device. In another example, user-command-based event processing capabilities can include event scenarios in which the user's commands can be processed by the device or not. Consider a scenario where a user issues a microwave-related command to the washing machine. Therefore, the washing machine cannot process microwave-related events. In yet another example, user-command-based event processing capabilities can be related to software characteristics. For example, if an X app for playing movies / TV series is not installed on the mobile device, any user commands related to that X app will be beyond the capabilities of the mobile device. Furthermore, the implementation details of device-specific information will be explained in detail in the following paragraphs.

[0068] After acquiring device-specific information, system 101 generates one or more context parameters. Alternatively, without departing from the scope of the present invention, context parameters may be referred to as hyperparameters. Context parameters may include a simplified set of parameters generated based on the analysis of NLP results and device-specific information. For example, Table 1 may represent the simplified set of parameters:

[0069] Table 1

[0070]

[0071] As an example scenario, the user provides the command "Start the washing machine in fast mode". Therefore, system 101 preprocesses the command to classify utterance and non-utterance information. Thus, in this case, the utterance information is "washing clothes in fast mode", and the non-utterance information is acoustic information obtained from non-utterance information determined based on audio events from the surrounding environment of the IoT device. Therefore, based on the acoustic information, system 101 identifies the audio scene. That is, it identifies the laundry room. Therefore, the context parameter corresponding to the audio scene is "laundry room". Furthermore, for example, the washing machine is relatively light, so the context parameter corresponding to its size is "light". Additionally, if the power is ON, the context parameter corresponding to the result can be considered "successful". Since no other device or user relies on the washing machine operation, the context parameter corresponding to complexity is considered "easy". Furthermore, if the clothes in the washing machine exceed its capacity, the context parameter corresponding to the operation state ID is "overload". Therefore, based on device-specific information and acoustic information, a simplified set of context parameters is obtained, as shown in Table 1.

[0072] Furthermore, after generating context parameters and based on these parameters, a suitable speech embedding is selected from the speaker embeddings, and the selected speaker embedding is output for playback to the user. Specifically, the selected speaker embedding and text are provided as input to the text-to-speech (TTS) engine for further processing to output audio / sound of the text with the characteristics of the selected speaker embedding for playback to the user. Now, in the same example of the washing machine, because the washing machine is not very heavy and it is overloaded, a female voice speaker embedding with an emotion of fatigue is selected. The generation of speaker embeddings will be explained in detail in the following paragraphs.

[0073] Figure 3Detailed implementations of components for generating a speaker embedding 300 according to embodiments of the present disclosure are shown. In one implementation, system 101 includes a plurality of user audio recording units 303-1 and 303-2, a plurality of noise suppression units 303-1 and 303-2, a plurality of feature extraction units 307-1 and 307-2, and a plurality of speaker encoders 309-1 and 309-2. For example, system 101 may be configured to collect a plurality of audio samples from a speaker 301-1 via user audio recording unit 303-1 in the plurality of user audio recording units 303-1 and 303-2. Subsequently, noise suppression unit 305-1 in the plurality of noise suppression units 305-1 and 305-2 may be configured to remove unwanted background noise from the collected speech. After removing unwanted noise, feature extraction unit 307-1 in multiple feature extraction units 307-1 and 307-2 can be configured to extract speech features by applying mathematical transformations known in the art, as a voiceprint / spectral map of the speech of speaker 301-1. The extracted voiceprint / spectral map represents a unique identifier of the utterance of speaker 301-1. The extracted features can then be configured to be provided as input to speaker encoder 309-2 in multiple speaker encoders 309-1 and 309-2 to generate / output speaker embedding 311-1. Finally, embedding 311-1 represents the speaker's voice. The above steps can be performed on multiple users 301-1, 301-2 to 301-n to generate one or more speaker embeddings 311-1, 311-2 to 311-n. The generated one or more speaker embeddings 311-1, 311-2 to 311-n are encoded voiceprints or speaker encoders.

[0074] After collecting the generation of one or more speaker embeddings, system 101 can be configured to artificially inject human emotions into the generated one or more speaker embeddings. Figure 4 This illustration shows a detailed implementation of components for artificially injecting human emotions into speaker embeddings according to embodiments of the present disclosure. In one implementation, block 401 relates to the generation of one or more speaker embeddings, and above... Figure 3The generation of this feature has already been explained. Therefore, for the sake of brevity, the description of box 401 is omitted here. Now, after generating one or more speaker embeddings from block 401, the generated one or more speaker embeddings can be configured to be input into a deep neural network (DNN) 403. The generated one or more speaker embeddings can be configured to inject human emotions by utilizing an emotion embedding library 405. The deep neural network (DNN) 403 can be further configured to generate one or more emotion-injected speaker embeddings 407, including different types of pitch and texture for multiple audio samples. The generated one or more emotion-injected speaker embeddings 407 are then configured to be stored or recorded in a speaker embedding library 409 for selecting speaker embeddings for one or more target IoT devices. For example, the generated speaker embeddings can include various types of human speech, such as male voices, female voices, children's voices, elderly voices, young people's voices, cartoon character voices, famous people's voices, robot voices, AI-generated voices, and so on. In an embodiment, a speaker embedding is one or more voiceprints / spectral maps of a speaker's speech. For example, without departing from the scope of the present invention, the generated speaker embedding may be referred to as a first speaker embedding. As a further example, the speaker embedding is defined in a subspace within the latent space, which is restricted to a defined number of coordinates.

[0075] In an alternative embodiment, system 101 may skip the mechanism of artificially injecting human emotions into the generated one or more speaker embeddings. Therefore, the generated one or more speaker embeddings may include natural user speech without any injected emotions.

[0076] In this embodiment, speaker embeddings are generated based on the emotions expressed during speech delivery. Therefore, speaker embeddings can be generated dynamically at runtime or preferentially during the training phase. When these speaker embeddings are provided to the text-to-speech (TTS) engine, audio is generated from the user's speech based on the characteristics of the speaker embeddings.

[0077] Figure 5 Detailed implementation of components for generating a mapping of speaker embeddings according to this disclosure is shown. System 101 includes a speaker embedding planner unit 501. The speaker embedding planner unit 501 generates a mapping of speaker embeddings based on device-specific information. The mapping of speaker embeddings is generated based on at least one fixed or transient attribute of one or more devices.

[0078] For example, a device's fixed attributes may include hardware characteristics and capabilities. A device's transient attributes may include device error states, location, and device events. These fixed and transient attributes are collectively referred to as device attribute 506. Device attribute 506 can be provided as input to a target audio module selector (not shown). The target audio module selector further provides device attribute 506 to a metadata extractor and formatter 511. The metadata extractor and formatter 511 can be configured to extract device-specific information from an IoT server / IoT cloud 503, such as device type, device capabilities, hardware characteristics, device location in a home IoT setup, and usage history. Subsequently, the metadata extractor and formatter 511 can be configured to manipulate device attribute 506 and device-specific information to output formatted metadata, which serves as input to a target embedding mapper 513. The target embedding mapper 513 can also be configured to receive a verbosity value 507 as input.

[0079] For example, the speech detail value 507 is a set of values ​​that define the output detail of the speaker embedding. Detail includes, but is not limited to, speech gender (male / female / child), accent, time-stretch, pitch, output length (short / clear or long / detailed), etc. These values ​​will be used by the target embedding mapper 513 to assign appropriate speech to the target IoT device.

[0080] Target embedding mapper 513 can be configured to operate on formatted metadata, utterance detail value 507, and speaker embeddings from speaker embedding repository 505. Subsequently, target embedding mapper 513 can be configured to associate or map device-specific information with one or more of the speaker embeddings to generate a mapping of speaker embedding sets for each IoT device based on the device-specific information. For example, boxes 515-1 and 515-2 are mappings for IoT devices (IoT device 1 and IoT device 2), respectively. Mappings 515-1 and 515-2 depict all possible assigned embeddings, covering all possible scenarios and emotions of speech that may occur in the IoT environment. For example, mapping 515-1 depicts a set of speaker embeddings assigned various speech generated from a first speaker embedding based on a mapping of extracted features. For example, speaker embedding 1 could represent a FamilyHub1_emb_sad.pkl file assigned sad speech, speaker embedding 2 could represent FamilyHub1_emb_happy.pkl, and so on. Table 2 shows the generative mapping of the speaker embedding set.

[0081] Table 2

[0082]

[0083] The generation of speaker embedding maps can be implemented in any newly added IoT device. After generating the mapping of the speaker embedding set, the target embedding mapper 513 can be configured to assign different voices to each of the speaker embedding sets of each IoT device based on relevance / mapping. The mapping of the speaker embedding set of each IoT device will be used to dynamically generate natural language output. The mapping of the speaker embedding set of each IoT device can be stored / recorded in an encoded format in an IoT server / IoT cloud 503 or an encoded embedding unit 509. Therefore, according to this disclosure, each IoT device is assigned a unique voice. Without departing from the scope of the inventive concept, the speaker embedding set may be referred to as a second speaker embedding.

[0084] According to embodiments of this disclosure, each speaker embedding in the set is associated with a corresponding tag. As an example, the tag may correspond to an identifier of the speaker embedding. The identifier may include a reference number or the user's name. For example, speaker-embeddings.ram or speaker-embeddings.old man, speaker-embeddings.1, etc.

[0085] In another embodiment of the invention, the speaker embedding planner unit 501 can be configured to identify similar IoT devices from one or more IoT devices operating in the environment based on device-specific information including IoT device attributes 506. For example, the speaker embedding planner unit 501 can identify IoT device 1 and IoT device 2 as similar devices or devices of the same category, such as refrigerators. Then, the target embedding mapper 513 can be configured to associate / map the device-specific information with one or more speaker embeddings to generate a mapping of speaker embedding sets for each of the identified similar IoT devices based on the device-specific information. Subsequently, the target embedding mapper 513 can be configured to assign a different voice to each of the speaker embedding sets for each IoT device based on the association / mapping. For example, IoT device 1 is identified as the latest model of refrigerator compared to IoT device 2. Then, in this case, the speaker embedding assigned to IoT device 2 will be an older speaker embedding, and the speaker embedding assigned to IoT device 2 will be a younger speaker embedding. Therefore, the appropriate speaker embedding will be selected at runtime. Subsequently, the mapping of speaker embedding sets for each IoT device can be stored in an encoded format in the IoT server / IoT cloud 503 or the encoded embedding unit 509. The detailed implementation of the mechanism of system 101, used to generate speakers embedded in the interactive computing system for presenting a voice assistant, will be explained in the following sections.

[0086] Figure 6Detailed implementations of components / modules for dynamically selecting speaker embeddings according to embodiments of the present disclosure are shown. In one implementation, a user can provide a voice command 601-1 or a text command 601-2, or provide the occurrence of a device event 601-3 as input. According to an embodiment, an automatic speech recognizer 603 is configured to receive the voice command 601-1. The automatic speech recognizer 603 is configured to recognize the user's speech information, which is further provided as input to an NLP unit 605. Furthermore, the NLP unit 605 can be configured to receive the text command 601-2 and the IoT event 601-3. Therefore, the NLP unit 605 can be configured to convert the speech information, text command, or IoT event to generate an NLP result. The NLP result can be further provided to a task analyzer 607. The NLP result can be used to determine at least one of the success, failure, or follow-up of an IoT event on one or more IoT devices caused by a user command. For example, success can be understood as the completion of an IoT event, such as the successful completion of laundry washing. Furthermore, the follow-up of IoT events on one or more IoT devices can be understood as any IoT event requiring action due to a user's previous actions. For example, a user provides the command "Call Vinay". The system then analyzes the task and might generate output such as, "Do you want to call a mobile phone or a home number?" The user might then say "Mobile phone".

[0087] For example, after successfully completing the laundry process, television playback will be scheduled according to predefined user commands. Furthermore, IoT event failures can be attributed to issues with capability, connectivity, latency, or memory when executing user commands.

[0088] Simultaneously, voice command 601-1 and device-specific information can be configured to be fed into non-verbal classifier 609. Non-verbal classifier 609 can be configured to classify non-verbal information corresponding to surrounding information of the IoT device or user. In one implementation, non-verbal classifier 609 can be configured to identify surrounding portable IoT devices 619 and surrounding audio scenes 616 based on voice commands and device-specific information, and output the non-verbal classifier result 615 to task analyzer 607 as input for predicting appropriate audio scenes.

[0089] According to an embodiment, the task analyzer 607 may include an acoustic prediction module 607-1, a complexity analysis module 607-3, a result analysis module 607-4, a target device retrieval feature module 607-7, and a target device status module 607-9. In one implementation, the task analyzer 607 may be configured to analyze NLP results and device-specific information. The detailed implementation of the task analyzer 607 and its components is explained below.

[0090] According to another embodiment, the acoustic prediction module 607-1 can be configured to predict an acoustic or audio scene of the user's surroundings based on the non-verbal classifier result 615 determined by the non-verbal classifier 609. For example, the acoustic prediction module 607-1 can predict an acoustic or audio scene. The acoustic or audio scene can be, for example, a noisy home, a silent home, an alarm clock event, the direction of a living voice, or speech. For example, a noisy home may include television playback, motor running, running water, etc. As a further example, a living voice may include human or pet sounds, etc. Based on the acoustic scene, the audio scene can be predicted as a happy, angry, or sad scene. In a further implementation, the TTS output volume can be increased based on the acoustic scene information. In a further implementation, background effects can be added based on the acoustic scene information.

[0091] According to another embodiment, the complexity analysis module 607-3 can be configured to analyze the complexity level of user commands. For example, the complexity analysis module 607-3 can be configured to analyze whether the execution of the command is on a single device or multiple devices, the computational overhead caused by the execution of the command, the time spent executing the command, and so on. Based on the analysis of the complexity analysis module 607-3, tasks can be classified as simple tasks, normal tasks, or difficult tasks. For example, turning on a light can be classified as a simple task, turning off the air conditioner at 10 pm can be classified as a normal task, and placing an order on Amazon from a shopping list can be classified as a difficult task because it involves multiple steps to complete the task.

[0092] According to another embodiment, the result analysis module 607-5 can be configured to analyze the response of the IoT device. For example, the result analysis module 607-5 can be configured to analyze whether the IoT device response has been successfully executed, whether the execution of the response has failed, or whether the execution is incomplete, based on user commands. If a fault occurs, the cause of the fault is analyzed. For example, regardless of whether the IoT device is overloaded, there may be connectivity problems, latency problems, memory problems, etc. Based on the analysis, the result analysis module 607-5 generates the analysis results. The result analysis module 607-5 can be configured to output a mapping between the analysis results and their causes.

[0093] According to another embodiment, the target device feature module 607-7 can be configured to analyze the hardware characteristics of the IoT device, such as the size, mobility, location, power consumption, computational head, and capabilities of the IoT device. The target device feature module 607-7 can also be configured to analyze the software characteristics of the IoT device, such as the software capabilities, operating status, and error scenarios of the IoT device. Subsequently, the target device feature module 607-7 can be further configured to classify the IoT device based on the analysis, such as large, small, medium, mobile, or immobile.

[0094] According to another embodiment, the target device status module 607-9 can be configured to obtain the current operating status of the IoT device regarding connectivity, error status, and overload, and the IoT device is classified accordingly.

[0095] Following the analysis of the various components of the task analyzer 607 as described above, the task analyzer 607 can be configured to generate one or more context parameters, including a simplified set of parameters, based on the analysis results from the NLP results and the analysis results from the various modules of the task analyzer 607, as described above. Examples of context parameters are shown in Table 1.

[0096] According to another embodiment, the embedding selector 611 can be configured to receive generated contextual parameters from the task analyzer 607 and / or user preference 613. For example, the user preference relates to utterance detail, which may include at least one of the following: the gender of the speech, the accent of the speech, the time stretching in the speech, the length of the speech, and the user's preferred speech. For example, the user pre-selects a male voice. Then, in this case, for the same washing machine example as explained in the preceding paragraphs, the embedding selector 611 can select a speaker with a male and tired voice embedded, rather than a female voice as explained therein.

[0097] Subsequently, the embedding selector 611 can be configured to select the most suitable speaker embedding for the target IoT device from the encoded embedding data repository 614 based on context parameters and / or user preferences 613. The selected speaker embedding is processed based on a text-to-speech mechanism to generate a text-to-speech (TTS) audio response to the user's command. Specifically, the speaker embedding is provided to the decoder and vocoder 617 to generate a spectrogram from the speaker embedding, which is then converted into the corresponding audio signal / audio utterance for playback to the user in natural language.

[0098] In another embodiment, the embedding selector 611 can be configured to retrieve user preferences related to utterance detail from device-specific information. Subsequently, the embedding selector 611 can be configured to reanalyze one or more generated contextual information. Then, the embedding selector 611 can be configured to associate the utterance detail-related user preferences with one or more speaker embeddings to generate a modified map of the speaker embedding set based on the reanalyzed contextual information. The modified map of the speaker embedding set is then stored in a database for selecting speaker embeddings for one or more target IoT devices for future use.

[0099] In an embodiment, the sound injector 619 can be configured to receive an audio signal generated by the decoder and vocoder 617. The sound injector 619 can be configured to select superimposed synthesized sounds based on context parameters, acoustic information, and device-specific information. Subsequently, the sound injector 619 can be configured to combine the superimposed synthesized sounds with the generated audio sounds and output the injected superimposed synthesized sounds in natural language to generate a speech stream for playback to the user. For example, a beeping sound signal can be superimposed to notify of an error status, and a siren signal can be superimposed to notify of a security alarm, etc. The injection of superimposed sounds is optional. Therefore, the audio signal output by the decoder and vocoder 617 can be played back as a TTS (Text-to-Speech) device by a voice assistance device.

[0100] Figure 7 An exemplary scenario is illustrated in an interactive computing system for presenting voice assistance according to embodiments of the present disclosure. The exemplary scenario implements... Figure 6 For simplicity, the same reference numerals are used for the details. In the scenario, the user provides the command "Start the washing machine in quick mode," as shown in Box 1. Therefore, System 101 preprocesses the command to classify utterance and non-utterance information. NLP unit 605 generates "Sorry, I can't do that right now," as shown in Box 2. Meanwhile, non-utterance classifier 609 identifies the audio scene as a laundry room and the user nearby, as shown in Box 3. Result analysis unit 607-5 generates its analysis result: "Failure," reason: "No water at the inlet," and complexity analysis module 607-3 generates its analysis conclusion: Complexity: Easy, as shown in Box 4. In addition, device-specific information 610 can determine the device-specific information as type: washing machine, size: lightweight, ID: unique, repeat: no, status: stopped / normal, as shown in Box 5. Subsequently, task analyzer 607 can generate context parameters, as shown in Box 6. It can be seen that the context parameters are a simplified set of parameters. Based on the context parameters and user preferences, the embedding selector 611 selects the most suitable speaker embedding, as shown in box 8.

[0101] Figure 8A flowchart illustrating an interactive computing system for presenting voice assistance according to an embodiment of the present disclosure is shown. As described above, method 800 can be implemented by system 101 using its components. In embodiments, method 800 can be executed by processor 203, memory 205, database 207, module 209, audio unit 210, transceiver 213, AI module 215, and audio unit 211. Furthermore, for the sake of brevity, ... Figures 1-6 The details of this disclosure are explained in detail in the description, and therefore are not disclosed here. Furthermore, references will be made to... Figure 9 To explain Figure 8 .

[0102] In box 801, method 800 includes generating a first speaker embedding for one or more speakers to define a subspace within the latent space. For example... Figure 9 As shown, when an audio recorder receives audio input, the audio signal typically includes noise, hence it is called a "noisy audio signal." A noise suppressor removes the noise and outputs a clean audio signal, called a "clean audio signal." A feature extractor extracts audio features and generates a speaker encoder, which is then infused with emotion to generate a first speaker embedding. This first speaker embedding is stored in a speaker embedding repository. The same process is performed on one or more speakers. Figure 3 The first speaker embedding is described in detail in [the document]. Therefore, for the sake of brevity, it is not disclosed in this paper.

[0103] In box 803, method 800 includes extracting characteristics of one or more network devices in a network environment that are related to the operation and / or nature of the devices. For example... Figure 9 As shown, the metadata extractor extracts metadata from the IoT server input and IoT device attributes, and provides this metadata to the target audio selector.

[0104] In box 805, method 800 includes logging a second speaker embedding based on features extracted relative to one or more network device mappings. Figure 9 As shown, the target audio selector correlates features of the extracted metadata to generate a second speaker embedding stored in the speech embedding data, which is used to select the appropriate speaker embedding to generate a TTS response for user playback.

[0105] In box 807, method 800 includes command-based reception, generating text into speech operations via second speaker embedding about one or more devices.

[0106] As another example, method 800 also includes selecting different voices for one or more network devices based on a mapping between device characteristics and tags associated with a second speaker embedding recorded during text-to-speech operations. The method then further assigns different voices relative to the one or more network devices during text-to-speech operations.

[0107] Figure 10 A flowchart illustrating an interactive computing system for presenting voice assistance according to an embodiment of the present disclosure is shown. As described above, method 1000 can be implemented by system 101 using its components. In embodiments, method 1000 can be executed by processor 203, memory 205, database 207, module 209, audio unit 210, transceiver 213, AI module 215, and audio unit 211. Furthermore, for the sake of brevity, ... Figures 1-6 The details of this disclosure are explained in detail in the description, and therefore are not disclosed here. Furthermore, references will be made to... Figure 11 To explain Figure 10 .

[0108] In box 1001, method 1000 includes preprocessing the input natural language (NL) text from the user command based on natural language processing (NLP) to classify utterance and non-utterance information. Figure 11 As shown, the automatic speech recognition classifies voice commands as follows: audio-1 as non-speech information and audio-2 as speech information, and executes the steps in box 1001.

[0109] Furthermore, the preprocessing step of method 1000 includes collecting multiple audio samples from multiple speakers to generate one or more speaker embeddings. Then, the preprocessing step includes storing the generated one or more speaker embeddings with human emotion in a database for selecting speaker embeddings for one or more target IoT devices. Additionally, the preprocessing step includes extracting device-specific information of the one or more IoT devices operating in a network environment. Subsequently, the preprocessing step includes associating the second device-specific information with the one or more speaker embeddings to generate a mapping of speaker embedding sets for each IoT device based on the device-specific information. Then, the preprocessing step includes assigning a different voice to each of the speaker embedding sets for each IoT device based on relevance, and storing the assigned speaker embedding sets in the database as a mapping for selecting speaker embeddings encoded for one or more target IoT devices.

[0110] In another example, method 1000 further includes artificially injecting human emotions into one or more generated speaker embeddings, wherein the generated one or more speaker embeddings include different types of pitch and texture for multiple audio samples, and the generated one or more speaker embeddings with human emotions are stored in a database.

[0111] In another example, the preprocessing step of method 1000 includes extracting device-specific information of one or more IoT devices operating in a network environment, and then identifying similar IoT devices from among the one or more IoT devices operating in the environment based on the device-specific information. Subsequently, the preprocessing step includes associating the device-specific information with one or more speaker embeddings to generate a mapping of speaker embedding sets for each of the identified similar IoT devices based on the device-specific information. Then, the preprocessing step includes assigning different voices to each of the identified similar IoT devices from the speaker embedding sets, and storing the assigned speaker embedding sets in a database as a mapping for selecting speaker embeddings for one or more target IoT devices.

[0112] In box 1003, method 1000 includes obtaining NLP results from user commands. Figure 11 As can be seen, the task analyzer acquires discourse information. For example, the acquisition step includes acquiring acoustic information that forms part of the non-discourse information to generate contextual parameters based on the acoustic information. Furthermore, acquiring NLP results from discourse information in the NL text includes determining one of the success, failure, or follow-up of an IoT event on one or more IoT devices due to a user command or device-specific event, where the determined result corresponds to the NLP result. Additionally, acquiring acoustic information from non-discourse information in the input NL text includes identifying one or more audio events in the surrounding environment of one or more target IoT devices, where the audio events in the surrounding environment correspond to the acoustic information.

[0113] In box 1003, method 1000 includes obtaining device-specific information from one or more target IoT devices operating in an environment based on NLP results. For example, the device-specific information is associated with one or more IoT devices and includes at least one of the following: device physical characteristics, software characteristics, device type, device capabilities, hardware characteristics, device location in the IoT environment, usage history, device settings, and information such as device error states, location, and device events. Figure 11 It can be seen that the acquisition of device-specific information includes the acquisition of data on the characteristics of the active device and the target device.

[0114] In box 1005, method 1000 includes generating one or more context parameters based on one or more of NLP results and device-specific information. Figure 11As can be seen, the task analyzer generates resulting features, which are referred to as context parameters in the method executed in box 1005. The generation of one or more context parameters involves selecting a speech embedding from the speaker embedding for playback by the IoT device. For example... Figure 11 As shown, the embedding selector selects the embedding used to generate the waveform, thereby performing the step at block 1005.

[0115] In box 1007, method 1000 includes selecting at least one speaker embedding stored in a database for one or more target IoT devices based on one or more context parameters. Figure 11 As shown, the speaker embedding selection is performed on a database known as speech embedding data.

[0116] In box 1011, method 1000 includes outputting the selected speaker embedding for playback to the user. Furthermore, the speaker embedding is dynamically generated as natural language output. Figure 11 As can be seen, the embedding selector selects the embedding for speaker selection from the speech embedding data and generates a response for execution by the waveform generator.

[0117] In another example, method 1000 also includes analyzing NLP results and device-specific information, wherein one or more of the context parameters include a simplified set of parameters generated based on the analysis of the NLP results and device-specific information. The analysis of the NLP results and device-specific information is performed by... Figure 11 The task analyzer shown is executed.

[0118] In another example, method 1000 further includes retrieving user preferences related to utterance detail from device-specific information, and then reanalyzing one or more generated contextual information. Subsequently, method 100 includes associating the user preferences related to utterance detail with one or more speaker embeddings to generate a modified mapping of a set of speaker embeddings based on the reanalyzed contextual information, wherein utterance detail includes at least one of the following: gender of the speech, accent of the speech, time stretching of the speech, length of the speech, and user-preferred speech; and the method includes storing the modified mapping of the set of speaker embeddings for selecting speaker embeddings for one or more target IoT devices.

[0119] In yet another example, method 1000 further includes generating audio sound by utilizing a text-to-speech mechanism based on processing of selected speaker embeddings, then selecting an overlay synthetic sound based on contextual parameters, acoustic information, and device-specific information, and then combining the overlay synthetic sound with the generated audio sound. Subsequently, the method also includes playing back the audio sound infused with the overlay synthetic sound as a natural language output to the user. Figure 11It can be seen that non-speech sound superposition can be injected by a waveform generator and can be used as the output of generated speech with sound superposition.

[0120] In yet another example, method 1000 also includes processing the selected speaker embedding based on a text-to-speech mechanism to replay the audio signal to the user in natural language. In an exemplary scenario, system 101 can generate audio signals for playback to the user based on various device physical characteristics. For example, heavy devices such as washing machines and home hubs can be configured to generate coarser speech. Lighter devices such as IoT sensors and wearable devices can be configured to generate lighter speech. Similarly, media devices such as mobile phones and smart speakers can be configured to generate media speech. In an alternative implementation, the generated audio signals can be stored preferentially, allowing the device to generate unique or different speech based on the different speech assigned to each device, even when the device is in offline mode. Figure 12 An exemplary scenario is illustrated where different voices are generated for different devices based on their physical attributes. For example, voice 1 can be assigned to device 1, which belongs to the light category; voice 2 can be assigned to device 2, which belongs to the heavy category; and voice 3 can be assigned to device 3, which belongs to the medium category.

[0121] In another exemplary scenario, system 101 can generate an audio signal for playback to a user based on the duration of device use and the device's age. For example, if the device is increasingly older, the assigned voice will be similar to an older version for the elderly, etc. In another example, if the system identifies two similar devices, it will assign two different voices to these devices based on when the devices were introduced, their location in the networked IoT environment, their physical characteristics, software characteristics, device type, device functions, and hardware characteristics.

[0122] In yet another exemplary scenario, Figure 13 The system 101 is shown to be able to generate audio signals for playback to the user based on device capabilities. For example, device 1 might get stuck while performing an operation, and when the device capability corresponds to the operational problem, device 1 might generate a response in an anxious voice, "I'm stuck, help me!!" Similarly, for example, device 2 might be overloaded while performing an operation, and when the device capability corresponds to capacity, device 2 might generate a response in a tired voice, "I'm overloaded, please get some clothes." In another example, when the device capability corresponds to an error event or event completion, device 1 might generate a response in an annoyed voice, "You forgot to close my door. Help me now" or "I've washed all the clothes."

[0123] According to another example, System 101 is able to generate audio signals for playback to the user based on task complexity. For example, if the device has completed a challenging task, its voice will convey happiness and a sense of accomplishment. The user can perceive the complexity of the task. Difficult / complex tasks can be performed on multiple devices. Third-party vendor communication tasks requiring urgent handling, such as online orders during Black Friday sales, can be responded to by the voice assistant with a similar tone. If the device is unable to complete the task due to its limitations, the voice will have a sense of sadness. If the device cannot complete the task but can suggest alternative options to the user, the voice texture will be different, such as... Figure 14 As shown.

[0124] According to another example, system 101 can generate multiple audio signals based on device and user profile characteristics for playback to users with similar devices, such as... Figure 15 As shown. For example, a husband's call will generate a notification using a male voice, or a wife's call will generate a notification using a female voice; or a father's call will generate a notification using an elderly person's voice.

[0125] According to another example, when multiple devices of the same type exist in a home, system 101 can generate audio signals for playback to the user. For example, a smart TV in the living room can generate a different voice signal than a smart TV in the kitchen. In another example, system 101 may be able to generate audio signals for playback to the user in response to different IoT-based events. For example, as... Figure 16 As shown, system 101 can generate speech in regular speech when receiving a notification, or in curious speech with a water flow effect when receiving a notification about a tap being turned on, or in curious speech with a dog barking effect when receiving a notification about a dog barking, or in hypersound with a siren effect when a safety alarm for receiving smoke is detected.

[0126] According to another example, see reference Figure 17 System 101 can generate speech signals based on acoustic scenarios and device settings, and modify the pitch of the output speech. As can be seen in Scenario 1, if the acoustic scenario is identified as a noisy room and the mobile device is in sound mode, the speech generated by the device will be modified with each iteration of the signal. In another scenario, Scenario 2, if the acoustic scenario is identified as a quiet room, the mobile device is in vibration mode, or an infant is asleep (IoT), the speech generated by the device will be a background beeping sound.

[0127] In another exemplary scenario, based on user proximity, the device can relay messages based on user distance and noise levels. For example, in the event of a washing machine malfunction, when the user is far away, the device can relay the washing machine's message to the user in the machine's voice. This ensures that the user receives the message and can identify which device the message actually originated from via voice recognition.

[0128] According to another exemplary scenario, system 101 can adjust the volume of the voice signal based on the type of event and the detection of the user's presence, such as... Figure 18 As shown. Therefore, as can be seen, when the presence of a user is detected in the living room, device 1 will generate a loud voice signal. Thus, as can be seen from the above, this disclosure improves the end-user experience and increases the user's connection with the AI-enabled virtual assistant system.

[0129] Furthermore, the actions in any flowchart do not need to be performed in the order shown; nor is it necessary to execute all actions. Moreover, actions that do not depend on other actions can be performed in parallel with other actions. The scope of the embodiments is by no means limited to these specific instances. Many variations are possible, such as differences in structure, size, and material use, whether explicitly stated in the specification. The scope of the embodiments is at least as broad as that given by the appended claims.

[0130] The advantages, other advantages, and solutions to problems have been described above with reference to specific embodiments. However, the benefits, advantages, solutions to problems, and any components that may cause any benefit, advantage, or solution to occur or become more apparent should not be construed as key, essential, or fundamental features or components of any or all claims.

[0131] While specific language has been used to describe the subject matter, it is not intended to impose any limitation. As will be apparent to those skilled in the art, various working modifications can be made to the method to achieve the inventive concepts taught herein. The accompanying drawings and the foregoing description provide examples of embodiments. Those skilled in the art will understand that one or more of the described elements can be well combined into a single functional element. Alternatively, certain elements may be separated into multiple functional elements. Elements from one embodiment may be added to another embodiment.

Claims

1. A method in an interactive computing system, the method comprising: Based on natural language processing (NLP), the input natural language (NL) from user commands is preprocessed for classifying discourse information and non-discourse information; Obtain NLP results from user commands; Obtain device-specific information from one or more IoT devices operating in an environment based on NLP results; Generate one or more context parameters based on NLP results and device-specific information; Select at least one speaker embedding stored in a database for one or more IoT devices based on one or more context parameters; as well as The output includes at least one speaker embedding for playback to the user. The one or more context parameters include: The size of the target IoT device The complexity is determined based on the dependence of other IoT devices in the one or more IoT devices on the operation of the target IoT device. Audio scenes indicating the operating location of target IoT devices, and The operating status of the target IoT device The selection of at least one speaker embedding includes selection based on the size of the target IoT device, the complexity, the audio scene, and the operating status of the target IoT device.

2. The method according to claim 1, further comprising: The text-to-speech mechanism processes at least one speaker embedding and plays it back to the user in natural language.

3. The method according to claim 1, wherein, Acquiring acoustic information, which forms part of non-discourse information, to generate one or more contextual parameters based on the acoustic information. Acoustic information obtained from non-verbal information includes: Identify one or more audio events in the surrounding environment of one or more IoT devices, wherein the one or more audio events in the surrounding environment correspond to acoustic information.

4. The method according to claim 1, wherein, Generating one or more context parameters includes selecting a speech embedding from at least one speaker embedding for playback by IoT devices in one or more IoT devices.

5. The method according to claim 1, wherein, Preprocessing includes: Collect multiple audio samples from multiple speakers to generate one or more speaker embeddings; One or more speaker embeddings generated are stored in a database and used to select one or more speaker embeddings for one or more IoT devices; Human emotions are artificially injected into one or more generated speaker embeddings, wherein the generated speaker embeddings include different types of pitch and texture for multiple audio samples; and One or more speakers with human emotions are generated and embedded in the database.

6. The method according to claim 1, wherein, Preprocessing includes: Extract device-specific information from one or more IoT devices operating in a network environment; Associating device-specific information with at least one speaker embedding to generate a mapping of speaker embedding sets for each of one or more IoT devices based on the device-specific information; Assigning different voices from a set of speaker embeddings to each of one or more IoT devices based on relevance; and The assigned speaker embedding set is stored in the database as a mapping for selecting coded speaker embeddings for one or more IoT devices.

7. The method according to claim 1, wherein, Preprocessing includes: Extract device-specific information from one or more IoT devices operating in a network environment; Identify similar IoT devices from one or more IoT devices operating in the environment based on device-specific information; Associating device-specific information with at least one speaker embedding to generate a mapping of speaker embedding sets for each of the similar IoT devices identified based on the device-specific information; Assigning different voices to each of similar IoT devices that include the recognition in the speaker embedding set; and The assigned set of speaker embeddings is stored in a database as a mapping for selecting speaker embeddings for one or more IoT devices.

8. The method according to claim 1, wherein, The following are included: Determine one of the following—success, failure, or follow-up—on one or more IoT devices due to a user command or a device-specific event, where the determined result corresponds to an NLP result.

9. An interactive computing system, comprising: One or more processors; as well as Memory is configured to store instructions that can be executed by one or more processors; as well as One or more processors are configured as follows: Based on natural language processing (NLP), the input natural language (NL) from user commands is preprocessed to classify utterance information and non-utterance information; Obtain NLP results from user commands; Obtain device-specific information from one or more IoT devices operating in an environment based on NLP results; Generate one or more context parameters based on NLP results and device-specific information; Select at least one speaker embedding stored in a database for one or more IoT devices based on one or more context parameters; and The output selects at least one speaker embedding for playback to the user. The one or more context parameters include: The size of the target IoT device The complexity is determined based on the dependence of other IoT devices in the one or more IoT devices on the operation of the target IoT device. Audio scenes indicating the operating location of target IoT devices, and The operating status of the target IoT device The selection of at least one speaker embedding includes selection based on the size of the target IoT device, the complexity, the audio scene, and the operating status of the target IoT device.

10. The system according to claim 9, wherein, One or more processors are also configured to process at least one selected speaker embedding based on a text-to-speech mechanism for playback to the user in natural language.

11. The system according to claim 9, wherein, One or more processors are also configured to: Acoustic information is obtained as part of non-discourse information in order to generate one or more contextual parameters based on the acoustic information. Identify one or more audio events in the surrounding environment of one or more IoT devices, wherein the one or more audio events in the surrounding environment correspond to acoustic information.

12. The system according to claim 9, wherein, One or more processors are also configured to select a voice embedding from at least one speaker embedding for playback by an IoT device in one or more IoT devices.

13. The system according to claim 9, wherein, One or more processors are also configured to: Collect multiple audio samples from multiple speakers to generate one or more speaker embeddings; One or more speaker embeddings with human emotions are generated and stored in a database for selecting one or more speaker embeddings for one or more IoT devices; Human emotions are artificially injected into one or more speaker embeddings, wherein the one or more speaker embeddings include different types of pitch and texture for multiple audio samples; as well as One or more speakers with human emotions are generated and embedded in the database.

14. The system according to claim 9, wherein, One or more processors are also configured to: Extract device-specific information from one or more IoT devices operating in a network environment; Associating device-specific information with at least one speaker embedding to generate a mapping of speaker embedding sets for each of one or more IoT devices based on the device-specific information; Assigning different voices from a set of speaker embeddings to each of one or more IoT devices based on relevance; The assigned speaker embedding set is stored in a database as a mapping for selecting coded speaker embeddings for one or more IoT devices; as well as Determine one of the following—success, failure, or follow-up—on one or more IoT devices due to a user command or a device-specific event, where the determined result corresponds to an NLP result.

15. The system according to claim 9, wherein, One or more processors are also configured to: Extract device-specific information from one or more IoT devices operating in a network environment; Identify similar IoT devices from one or more IoT devices operating in the environment based on device-specific information; Associating device-specific information with at least one speaker embedding to generate a mapping of speaker embedding sets for each of the similar IoT devices identified based on the device-specific information; Assign different voices to each of similar IoT devices that include the recognition in the speaker embedding set; as well as The assigned set of speaker embeddings is stored in a database as a mapping for selecting speaker embeddings for one or more IoT devices.

Citation Information

Patent Citations

  • Robot emotion generating and expressing system

    CN103218654A

  • System and method for providing smart objects virtual communication

    US20200143235A1