System and method for false wake-up suppression

By delaying wake-up decisions and combining multiple data sources and features, the problem of wrong wake-up in voice assistants is solved, significantly improving the accuracy and user experience of the system.

CN119948562APending Publication Date: 2025-05-06SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380067254.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-06-02
Filing Date
2023-06-27
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Error wakeup leads to poor user experience, waste of resources, and privacy issues in voice assistants, especially when audio training samples are lacking during seamless registration.

Method used

By using multiple data sources and output from automatic speech recognition (ASR) models, wake-up decisions are delayed until post-ASR processing, combining audio, text, and contextual features, predict the probability of false wake-up and suppress its occurrence.

Benefits of technology

It significantly improves the wake-up detection results, improves the accuracy of the system, reduces the occurrence of false wake-ups, improves user experience and protects user privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948562A_ABST
    Figure CN119948562A_ABST
Patent Text Reader

Abstract

A method includes obtaining a speech signal. The method further includes predicting a first likelihood of speaking a wake-up word or phrase in the speech signal using a first machine learning model trained to receive the speech signal as input. The method also includes, in response to the first likelihood exceeding a first threshold, performing automatic speech recognition on the speech signal to determine a textual representation of the speech signal. The method further includes predicting a second likelihood that the wake-up word or phrase is spoken in the speech signal using a second machine learning model trained to receive at least one of the textual representation, an audio feature associated with the speech signal, and a contextual feature associated with an electronic device. Further, the method includes, in response to the second likelihood exceeding a second threshold, generating an instruction for performing an action requested in the voice signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates generally to machine learning systems and more particularly to a system and method for false wakeup suppression. Background Art

[0002] When voice assistants wake up when the user did not intend to speak to them, it results in a poor user experience and a waste of device resources. False wakeups can also cause privacy issues between users. In addition, the false wakeup problem is more serious when a seamless registration process is used that does not require audio training samples from the user. Previously, during registration of the device (such as setting up the device after purchase), users were required to provide audio training samples of themselves saying the wake-up word to train the wake-up word detection system. However, users may not want to provide audio training samples, so seamless registration is increasingly implemented instead. Although seamless registration can improve the user experience by eliminating steps during registration, the lack of audio training samples can result in reduced accuracy of the wake-up word detection system. Summary of the invention

[0003] Technical Solution

[0004] The present disclosure relates to a system and method for false wakeup suppression.

[0005] In an embodiment, a method includes: obtaining a voice signal by at least one processing device of an electronic device. The method also includes predicting, by the at least one processing device, a first likelihood that a wake-up word or phrase is spoken in the voice signal using a first machine learning model trained to receive the voice signal as input. The method also includes, in response to the first likelihood exceeding a first threshold, performing automatic speech recognition on the voice signal by the at least one processing device to determine a text representation of the voice signal. The method also includes predicting, by the at least one processing device, a second likelihood that the wake-up word or phrase is spoken in the voice signal using a second machine learning model, the second machine learning model being trained to receive the text representation, at least one of audio features associated with the voice signal, and context features associated with the electronic device. In addition, the method includes generating, by the at least one processing device, an instruction to perform an action requested in the voice signal in response to the second likelihood exceeding a second threshold.

[0006] In an embodiment, an electronic device comprises: at least one processing device configured to obtain a speech signal. The at least one processing device is also configured to use a first machine learning model trained to receive the speech signal as input to predict a first likelihood that a wake-up word or phrase is spoken in the speech signal. The at least one processing device is also configured to perform automatic speech recognition on the speech signal to determine a text representation of the speech signal in response to the first likelihood exceeding a first threshold. The at least one processing device is also configured to use a second machine learning model to predict a second likelihood that the wake-up word or phrase is spoken in the speech signal, the second machine learning model being trained to receive the text representation, audio features associated with the speech signal, and at least one of contextual features associated with the electronic device. In addition, the at least one processing device is configured to generate instructions for performing an action requested in the speech signal in response to the second likelihood exceeding a second threshold.

[0007] In an embodiment, a machine-readable medium includes instructions that, when executed, cause at least one processor of an electronic device to obtain a voice signal. The machine-readable medium also includes instructions that, when executed, cause the at least one processor to use a first machine learning model trained to receive the voice signal as input to predict a first likelihood of a wake-up word or phrase being spoken in the voice signal. The machine-readable medium also includes instructions that, when executed, cause the at least one processor to perform automatic speech recognition on the voice signal in response to the first likelihood exceeding a first threshold to determine a text representation of the voice signal. The machine-readable medium also includes instructions that, when executed, cause the at least one processor to use a second machine learning model to predict a second likelihood of speaking the wake-up word or phrase in the voice signal, the second machine learning model being trained to receive the text representation, at least one of the audio features associated with the voice signal, and the context features associated with the electronic device. In addition, the machine-readable medium includes instructions that, when executed, cause the at least one processor to generate instructions for executing an action requested in the voice signal in response to the second likelihood exceeding a second threshold.

[0008] Other technical features may be apparent to those skilled in the art from the following drawings, descriptions and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] For a more complete understanding of the present disclosure and its advantages, reference is now made to the following description taken in conjunction with the accompanying drawings, wherein like reference numerals represent like parts:

[0010] Figure 1 An example network configuration including electronic devices according to the present disclosure is shown;

[0011] Figure 2A An example wake word detection and false wake word suppression system according to the present disclosure is shown;

[0012] Figure 2B An example system is shown in which a false wakeup suppression classifier model according to the present disclosure is stored on a server;

[0013] Figure 3 An example wake-up verification process according to the present disclosure is shown;

[0014] Figure 4A Example audio and contextual input features according to the present disclosure are shown;

[0015] Figure 4B An example audio and contextual feature input process according to the present disclosure is shown;

[0016] Figure 5 An example false wakeup suppression classifier model architecture according to the present disclosure is shown;

[0017] Figure 6 Another example false wakeup suppression classifier model architecture according to the present disclosure is shown;

[0018] Figure 7 An example process for training a false arousal suppression classifier model according to the present disclosure is shown; and

[0019] Fig. 8A and 8B An example method for false wakeup suppression according to the present disclosure is shown;

[0020] Fig. 9 A flow chart of a method for false wakeup suppression according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0021] Before proceeding to the following detailed description, it may be advantageous to set forth definitions of certain words and phrases used throughout this patent document. The terms "send," "receive," and "communicate," and their derivatives, encompass both direct and indirect communications. The terms "include," "comprise," and their derivatives, mean including, but not limited to. The term "or" is inclusive, meaning and / or. The phrase "associated with," and its derivatives, means including, included within, interconnected with, including, contained within, connected to or connected with, coupled to or coupled with, communicable with, collaborate with, interlaced, juxtaposed, close to, bound to or bound with, having, having the nature of, having a relationship to or with, etc.

[0022] In addition, the various functions described below can be implemented or supported by one or more computer programs, each of which is formed by a computer-readable program code and implemented in a computer-readable medium. The terms "application" and "program" refer to one or more computer programs, software components, instruction sets, processes, functions, objects, classes, instances, related data or parts thereof suitable for implementation with suitable computer-readable program codes. The phrase "computer-readable program code" includes any type of computer code, including source code, object code and executable code. The phrase "computer-readable medium" includes any type of medium that can be accessed by a computer, such as a read-only memory (ROM), a random access memory (RAM), a hard drive, a compact disc (CD), a digital video disc (DVD) or any other type of memory. "Non-transitory" computer-readable media excludes wired, wireless, optical or other communication links that transmit temporary electrical signals or other signals. Non-transitory computer-readable media include media that can permanently store data and media that can store data and rewrite it later, such as rewritable optical discs or erasable memory devices.

[0023] As used herein, terms and phrases such as "having", "may have", "include" or "may include" a feature (such as a number, function, operation or component such as a part) indicate the presence of the feature and do not exclude the presence of other features. In addition, as used herein, the phrases "A or B", "at least one of A and / or B" or "one or more of A and / or B" may include all possible combinations of A and B. For example, "A or B", "at least one of A and B" and "at least one of A or B" may indicate all of the following: (1) including at least one A, (2) including at least one B, or (3) including at least one A and at least one B. In addition, as used herein, the terms "first" and "second" may modify various components regardless of importance and do not limit the components. These terms are only used to distinguish one component from another. For example, a first user device and a second user device may indicate user devices that are different from each other, regardless of the order or importance of the devices. Without departing from the scope of the present disclosure, a first component may be represented as a second component, and vice versa.

[0024] It will be understood that when an element (such as a first element) is referred to as being (operably or communicatively) "coupled with" or "coupled to" another element (such as a second element) or "connected with" or "connected to" another element (such as a second element), it may be coupled or connected to the other element directly or via a third element. In contrast, it will be understood that when an element (such as a first element) is referred to as being "directly coupled with" or "directly coupled to" another element (such as a second element) or "directly connected with" or "directly connected to" another element (such as a second element), no other elements (such as a third element) are interposed between the element and the other element.

[0025] As used herein, the phrase "configured (or set) to" may be used interchangeably with the phrases "suitable for," "capable of," "designed to," "adapted to," "manufactured to," or "capable of," as the case may be. The phrase "configured (or set) to" does not necessarily mean "specifically designed in hardware to." On the contrary, the phrase "configured to" may indicate that a device may perform an operation together with another device or component. For example, the phrase "a processor configured (or set) to perform A, B, and C" may represent a general-purpose processor (such as a CPU or application processor) that can perform operations by executing one or more software programs stored in a memory device, or a dedicated processor (such as an embedded processor) for performing operations.

[0026] The terms and phrases used herein are provided only for describing some embodiments of the present disclosure, rather than limiting the scope of other embodiments of the present disclosure. It should be understood that, unless the context clearly indicates otherwise, the singular forms "one", "an" and "said" include plural references. All terms and phrases used herein (including technical terms and phrases and scientific terms and phrases) have the same meaning as the meaning generally understood by those of ordinary skill in the art to which the embodiments of the present disclosure belong. It will be further understood that terms and phrases (such as, those terms and phrases defined in a common dictionary) should be interpreted as having a meaning consistent with its meaning in the context of the relevant field, and will not be interpreted with an idealized or overly formal meaning, unless explicitly defined as such herein. In some cases, the terms and phrases defined herein may be interpreted as excluding embodiments of the present disclosure.

[0027] Examples of "electronic devices" according to embodiments of the present disclosure may include at least one of the following: a smart phone, a tablet personal computer (PC), a mobile phone, a video phone, an e-book reader, a desktop PC, a laptop computer, a netbook computer, a workstation, a personal digital assistant (PDA), a portable multimedia player (PMP), an MP3 player, a mobile medical device, a camera, or a wearable device (such as smart glasses, a head mounted device (HMD), electronic clothing, an electronic bracelet, an electronic necklace, an electronic accessory, an electronic tattoo, a smart mirror, or a smart watch). Other examples of electronic devices include smart home appliances. Examples of smart home appliances may include a television, a digital video disc (DVD) player, an audio player, a refrigerator, an air conditioner, a vacuum cleaner, an oven, a microwave, a washing machine, a dryer, an air purifier, a set-top box, a home automation control panel, a security control panel, a TV box (such as Samsung HOMESYNC, Apple TV, or Google TV), a smart speaker or a speaker with an integrated digital assistant (such as Samsung GALAXY HOME, Apple HOMEPOD, or Amazon ECHO), a game console (such as XBOX, PLAYSTATION, or NINTENDO), an electronic dictionary, an electronic key, a portable camera, or at least one of an electronic photo frame. Other examples of electronic devices include at least one of the following: various medical devices (such as various portable medical measuring devices (such as blood sugar measuring devices, heart rate measuring devices, or body temperature measuring devices), magnetic resource angiography (MRA) devices, magnetic resource imaging (MRI) devices, computed tomography (CT) devices, imaging devices, or ultrasound devices), navigation devices, global positioning system (GPS) receivers, event data recorders (EDR), flight data recorders (FDR), car infotainment devices, navigation electronic devices (such as navigation navigation devices or gyrocompasses), avionics equipment, security devices, vehicle head units, industrial or household robots, automated teller machines (ATMs), point of sale (POS) devices, or Internet of Things (IoT) devices (such as light bulbs, various sensors, electric or gas meters, sprinklers, fire alarms, thermostats, street lights, ovens, fitness equipment, hot water tanks, heaters, or boilers). Other examples of electronic devices include at least a portion of a piece of furniture or a building / structure, an electronic board, an electronic signature receiving device, a projector, or various measuring devices (such as devices for measuring water, electricity, gas, or electromagnetic waves). Note that according to various embodiments of the present disclosure, the electronic device may be one or a combination of the devices listed above. According to some embodiments of the present disclosure, the electronic device may be a flexible electronic device. The electronic devices disclosed herein are not limited to the devices listed above, and may include any other electronic devices now known or later developed.

[0028] In the following description, an electronic device is described with reference to the accompanying drawings according to various embodiments of the present disclosure. As used herein, the term "user" may refer to a person using an electronic device or another device (such as an artificial intelligence electronic device).

[0029] Definitions for certain other words and phrases may be provided throughout this patent document. Those of ordinary skill in the art should understand that in many, if not most instances, such definitions apply to prior, as well as future uses of such defined words and phrases.

[0030] None of the description in this application should be read as implying that any particular element, step, or function is an essential element that must be included in the claims scope. The scope of the patented subject matter is limited solely by the claims.

[0031] The following discussion is described with reference to the accompanying drawings. Figures 1 to 8B As well as various embodiments of the present disclosure. However, it should be understood that the present disclosure is not limited to these embodiments, and all changes and / or equivalent forms or replacement forms thereof also belong to the scope of the present disclosure. Throughout the specification and the drawings, the same or similar reference numerals may be used to refer to the same or similar elements.

[0032] As described above, when voice assistants (such as BIXBY, SIRI, and ALEXA) wake up when the user does not intend to speak to them, it will result in a poor user experience and a waste of device resources. False wake-ups can also cause privacy issues between users because users may think that the content of their speech is being transmitted to other electronic devices over the Internet. In addition, the false wake-up problem is more serious when a seamless registration process that does not require audio training samples from the user is used. Previously, during device registration (such as setting up the device after purchase), users were required to provide audio training samples of the wake-up words they said themselves for training the wake-up word detection system. However, users may not want to provide audio training samples, so seamless registration is increasingly implemented instead. Although seamless registration can improve the user experience by eliminating steps during registration, the lack of use of audio training samples may result in reduced accuracy of the wake-up word detection system.

[0033] The present disclosure provides various techniques for post-automatic speech recognition (ASR) false wakeup suppression that utilize multiple data sources and use one or more outputs from an ASR model in the first few seconds instead of making a wakeup decision. By delaying wakeup decision making to post-ASR processing, wakeup detection results are significantly improved, potentially up to 93% to 96% or better. In some embodiments, multiple data sources associated with audio input are collected, and wakeup decisions are performed using the multiple data sources and the ASR text output provided by the ASR model (rather than making a wakeup decision with audio data alone).

[0034] As described in the present disclosure, this is achieved by implementing a false wakeup suppression model, where the false wakeup suppression model takes the output from the audio wakeup classifier model (as well as any data derived from the output of the wakeup classifier model or from the input audio signal) and the output from the ASR model and contextual features that provide context for the environment in which the wakeup detection operation is being performed. For example, information derived from the audio signal, the output ASR text from the ASR model, and the client device context can be used to predict the probability of a false wakeup. In various embodiments, the threshold for determining whether an intended wakeup event is likely to exist can be adjusted based on the prior probabilities of false and correct wakeup events determined using previously processed audio signals to tune the system. As described in the present disclosure, in some cases, the false wakeup suppression model can be constructed using other configurations of a transformer-based language model and a multilayer perceptron model, a transformer-based language model and a random forest classifier model, or a machine learning model.

[0035] Figure 1 An example network configuration 100 including electronic devices according to the present disclosure is shown. Figure 1 The embodiment of the network configuration 100 shown in FIG. 1 is for illustration only. Other embodiments of the network configuration 100 may be used without departing from the scope of the present disclosure.

[0036] According to an embodiment of the present disclosure, an electronic device 101 is included in a network configuration 100. The electronic device 101 may include at least one of a bus 110, a processor 120, a memory 130, an input / output (I / O) interface 150, a display 160, a communication interface 170, or a sensor 180. In some embodiments, the electronic device 101 may exclude at least one of these components, or may add at least one other component. The bus 110 includes circuits for connecting components 120-180 to each other and for transmitting communications (such as control messages and / or data) between components.

[0037] The processor 120 includes one or more processing devices, such as one or more microprocessors, microcontrollers, digital signal processors (DSPs), application specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs). In some embodiments, the processor 120 includes one or more of a central processing unit (CPU), an application processor (AP), a communication processor (CP), or a graphics processor unit (GPU). The processor 120 is capable of controlling at least one of the other components of the electronic device 101 and / or performing operations or data processing related to communication or other functions.

[0038] The functions related to artificial intelligence according to the present disclosure are operated by at least one processor 120 and a memory 130. The processor 120 may include one or more processors. In this case, the one or more processors may be a general-purpose processor (such as a CPU, an AP, or a digital signal processor (DSP)), a graphics processor only (such as a GPU or a visual processing unit (VPU)), or an artificial intelligence processor only (such as an NPU). For example, when one or more processors are processors dedicated to artificial intelligence, the processor dedicated to artificial intelligence may be designed as a hardware structure dedicated to processing a specific artificial intelligence model.

[0039] The processor 120 controls the input data to be processed according to the predefined operation rules or artificial intelligence model stored in the memory 130. Optionally, when the processor 120 is a processor dedicated to artificial intelligence, the processor dedicated to artificial intelligence can be designed with a hardware structure dedicated to processing a specific artificial intelligence model.

[0040] The predefined action rule or artificial intelligence model is characterized in that it is created by learning. Here, creation by learning means learning a basic artificial intelligence model using a plurality of learning data through a learning algorithm, so that a predefined action rule or artificial intelligence model set to perform the desired characteristics (or purpose) is created. This learning can be performed in the device itself that executes the artificial intelligence according to the present disclosure, or by a separate server and / or system. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but are not limited to the above examples. The artificial intelligence model can be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values, and the neural network operation is performed by the operation between the operation result of the previous layer and the multiple weight values. The multiple weights possessed by the multiple neural network layers can be optimized by the learning results of the artificial intelligence model. For example, multiple weights can be updated so that the loss value or cost value obtained from the artificial intelligence model is reduced or minimized during the learning process. The artificial neural network may include a deep neural network (DNN), such as a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), or a deep Q network, but is not limited to the above examples.

[0041] As described below, the processor 120 can receive and process input (such as an audio input such as voice data or a signal received from an audio input device such as a microphone), and use the input to perform a wake-up word detection and wake-up determination process. This may include using at least a first machine learning model to predict a first likelihood that the wake-up word is spoken in the input, and a second machine learning model to predict a second likelihood that the wake-up word is spoken in the input, and if both probabilities are above a corresponding threshold, triggering further actions. The processor 120 can also instruct other devices to perform certain operations (such as outputting audio using an audio output device such as a speaker) or displaying content on one or more displays 160. The processor 120 can also receive input (such as data samples to be used in training a machine learning model), and manage such training by inputting samples into the machine learning model, receiving output from the machine learning model, and executing a learning function (such as a loss function) to improve the machine learning model.

[0042] The memory 130 may include volatile and / or non-volatile memory. For example, the memory 130 may store commands or data related to at least one other component of the electronic device 101. According to an embodiment of the present disclosure, the memory 130 may store software and / or programs 140. The programs 140 include, for example, a kernel 141, middleware 143, an application programming interface (API) 145, and / or an application program (or "application") 147. At least a portion of the kernel 141, the middleware 143, or the API 145 may be represented as an operating system (OS).

[0043] The kernel 141 may control or manage system resources (such as bus 110, processor 120, or memory 130) for executing operations or functions implemented in other programs (such as middleware 143, API 145, or application 147). The kernel 141 provides an interface that allows the middleware 143, API 145, or application 147 to access the various components of the electronic device 101 to control or manage system resources. Application 147 includes one or more applications that receive audio data, predict the wake-up word in the speech included in the audio data, perform speech recognition on the speech, perform false wake-up word suppression or verification operations using the audio data, and if the false wake-up word verification operation passes, perform tasks related to the speech content. These functions may be performed by a single application or by multiple applications, each of which performs one or more of these functions. For example, the middleware 143 may be used as a relay to allow the API 145 or application 147 to transmit data to the kernel 141. Multiple applications 147 may be provided. The middleware 143 can control work requests received from the applications 147, such as by assigning a priority for using system resources (such as the bus 110, the processor 120, or the memory 130) of the electronic device 101 to at least one of the plurality of applications 147. The API 145 is an interface that allows the application 147 to control functions provided from the kernel 141 or the middleware 143.

[0044] The I / O interface 150 serves as an interface that can, for example, transmit commands or data input from a user or other external devices to other components of the electronic device 101. The I / O interface 150 can also output commands or data received from other components of the electronic device 101 to the user or other external devices.

[0045] The display 160 includes, for example, a liquid crystal display (LCD), a light emitting diode (LED) display, an organic light emitting diode (OLED) display, a quantum dot light emitting diode (QLED) display, a microelectromechanical system (MEMS) display, or an electronic paper display. The display 160 may also be a depth perception display, such as a multi-focal display. The display 160 is capable of displaying, for example, various contents (such as text, images, videos, icons, or symbols) to the user. The display 160 may include a touch screen and may receive, for example, a touch, gesture, proximity, or hover input using an electronic pen or a body part of the user.

[0046] The communication interface 170, for example, can establish communication between the electronic device 101 and an external electronic device (such as the first electronic device 102, the second electronic device 104, or the server 106). For example, the communication interface 170 can be connected to the network 162 or 164 through wireless or wired communication to communicate with the external electronic device. The communication interface 170 can be a wired or wireless transceiver or any other component for sending and receiving signals.

[0047] The electronic device 101 also includes one or more sensors 180, which can measure physical quantities or detect the activation state of the electronic device 101, and convert the measured or detected information into electrical signals. The sensor 180 may also include one or more buttons for touch input, one or more microphones, gesture sensors, gyroscopes or gyro sensors, air pressure sensors, magnetic sensors or magnetometers, acceleration sensors or accelerometers, grip sensors, proximity sensors, color sensors (such as, RGB sensors), biophysical sensors, temperature sensors, humidity sensors, illumination sensors, ultraviolet (UV) sensors, electromyography (EMG) sensors, electroencephalogram (EEG) sensors, electrocardiogram (ECG) sensors, infrared (IR) sensors, ultrasonic sensors, iris sensors or fingerprint sensors. The sensor 180 may also include an inertial measurement unit, which may include one or more accelerometers, gyroscopes and other components. In addition, the sensor 180 may include a control circuit for controlling at least one of the sensors included here. Any of these sensors 180 may be located in the electronic device 101.

[0048] The first external electronic device 102 or the second external electronic device 104 may be a wearable device or a wearable device (such as an HMD) in which an electronic device may be installed. When the electronic device 101 is installed in the electronic device 102 (such as an HMD), the electronic device 101 may communicate with the electronic device 102 through the communication interface 170. The electronic device 101 may be directly connected to the electronic device 102 to communicate with the electronic device 102 without involving a separate network. The electronic device 101 may also be an augmented reality wearable device including one or more cameras, such as glasses.

[0049] Wireless communication can use, for example, WiFi, long term evolution (LTE), long term evolution-advanced (LTE-A), 5th generation wireless system (5G), millimeter wave or 60GHz wireless communication, wireless USB, code division multiple access (CDMA), wideband code division multiple access (WCDMA), universal mobile telecommunications system (UMTS), wireless broadband (WiBro) or global system for mobile communications (GSM) as at least one of the communication protocol. Wired connection can include, for example, universal serial bus (USB), high-definition multimedia interface (HDMI), recommended standard 232 (RS-232) or plain old telephone service (POTS). Network 162 includes at least one communication network, such as a computer network (e.g., local area network (LAN) or wide area network (WAN)), the Internet or a telephone network.

[0050] The first external electronic device 102 and the second external electronic device 104 and the server 106 may each be a device of the same or different type as the electronic device 101. According to some embodiments of the present disclosure, the server 106 includes a group of one or more servers. In addition, according to some embodiments of the present disclosure, all or some operations performed on the electronic device 101 may be performed on another one or more other electronic devices (such as the electronic devices 102 and 104 or the server 106). In addition, according to some embodiments of the present disclosure, when the electronic device 101 should automatically or upon request perform a certain function or service, the electronic device 101 may request another device (such as the electronic device 102 and the electronic device 104 or the server 106) to perform at least some functions associated with it, instead of performing the function or service alone, or additionally request another device (such as the electronic device 102 and the electronic device 104 or the server 106) to perform at least some functions associated with it. Another electronic device (such as the electronic device 102 and the electronic device 104 or the server 106) is capable of performing the requested function or additional function, and transmitting the result of the execution to the electronic device 101. The electronic device 101 can provide the requested function or service by processing the received result as it is or additionally. To this end, for example, cloud computing, distributed computing, or client-server computing technology can be used. Figure 1 The electronic device 101 is shown to include a communication interface 170 that communicates with an external electronic device 104 or a server 106 via a network 162 , but according to some embodiments of the present disclosure, the electronic device 101 may operate independently without a separate communication function.

[0051] The server 106 may include components that are the same or similar to the electronic device 101 (or a suitable subset thereof). The server 106 may support driving the electronic device 101 by performing at least one of the operations (or functions) implemented on the electronic device 101. For example, the server 106 may include a processing module or processor that can support the processor 120 implemented in the electronic device 101. As described below, the server 106 may receive and process inputs (such as audio inputs such as voice data or signals received from an audio input device such as a microphone), and use the inputs to perform wake-up word detection and wake-up determination processes. This may include using at least a first machine learning model to predict a first likelihood of saying a wake-up word in the input, and a second machine learning model to predict a second likelihood of saying a wake-up word in the input, and if both probabilities are above a corresponding threshold, further actions are triggered. The server 106 may also instruct other devices to perform certain operations (such as outputting audio using an audio output device such as a speaker) or display content on one or more displays 160. Server 106 may also receive inputs (such as data samples to be used in training a machine learning model) and manage such training by inputting the samples into the machine learning model, receiving outputs from the machine learning model, and executing learning functions (such as a loss function) to improve the machine learning model.

[0052] although Figure 1 An example of a network configuration 100 including an electronic device 101 is shown, but Figure 1 Various changes may be made. For example, network configuration 100 may include any suitable number of each component in any suitable arrangement. In general, computing and communication systems have a wide variety of configurations, and Figure 1 The scope of the present disclosure is not limited to any particular configuration. Figure 1 One operating environment is shown in which the various features disclosed in this patent document may be used, but these features may be used in any other suitable system.

[0053] Figure 2A An example wake-up word detection and false wake-up word suppression system 200 according to the present disclosure is shown. For ease of explanation, the system 200 is described as involving the use of Figure 1 The electronic device 101 in the network configuration 100. However, the system 200 may be used with any other suitable electronic device (such as, the server 106) and may be used in any other suitable system.

[0054] As in Figure 2AAs shown in , the system 200 includes an electronic device 101, which includes a processor 120. The processor 120 is operably coupled to or otherwise configured to use one or more machine learning models, such as an audio wakeup classifier model 201 (which may include a wakeup word detector model), an automatic speech recognition (ASR) model 202, and a false wakeup suppression classifier model 204. The audio wakeup classifier model 201 can be trained to recognize one or more wakeup words or phrases, and act as a first gatekeeper when determining whether the electronic device 101 should activate or further use a voice assistant to process received audio data including user requests or commands and act according to the received audio data. The processor 120 can also be operably coupled to or otherwise configured to use one or more other models 205 or other processes, such as one or more natural language understanding (NLU) models, one or more audio feature acquisition processes, one or more context feature acquisition processes, one or more command routers or routing processes, one or more action plan building processes, etc. It will be understood that the machine learning models 201-205 can be stored in a memory of the electronic device 101, such as memory 130, and accessed by the processor 120 to perform automatic speech recognition tasks or other tasks. However, the machine learning models 201-205 can be stored in any other suitable manner.

[0055] The system 200 also includes an audio input device 206 (such as a microphone), an audio output device 208 (such as a speaker or headphones), and a display 210 (such as a screen or monitor, such as the display 160). The processor 120 receives audio input from the audio input device 206 and provides the audio input to the trained audio wakeup classifier model 201. The trained audio wakeup classifier model 201 detects whether a wakeup word or phrase is included in an utterance within the audio data, and outputs a result to the processor 120, such as one or more predictions or probabilities that the utterance includes the wakeup word or phrase. If a wakeup word or phrase is detected, such as based on determining that the likelihood that the speech signal includes the wakeup word or phrase is above a specified threshold, the processor 120 provides the audio data to the ASR model 202, which processes the audio data and provides a predicted text output based on the audio data. The processor 120 provides the predicted text output, audio features associated with the input audio data, and contextual features associated with the specific context or environment in which the wake-up word detection is being performed to the false wake-up suppression classifier model 204 in order to predict whether the user intended to trigger a wake-up of the voice assistant, such as whether the utterance was merely random user utterance.

[0056] If the processor 120 determines using the models 201-204 that the utterance may include a wake word or phrase, such as if one or more probabilities are above one or more thresholds, the processor 120 may instruct at least one action of the electronic device 101 or another device or system. For example, in response to positive detection of a wake word or phrase, the processor 120 may instruct one or more further actions corresponding to one or more instructions or requests provided in the utterance.

[0057] As a specific example, assume that an utterance received from a user via the audio input device 206 includes a wake word or phrase (such as, "HEY BIXBY, CALL MOM"). Here, the trained audio wakeup classifier model 201 detects a first likelihood that the wakeup word "BIXBY" or the phrase "HEY, BIXBY" is present, and the processor 120 determines that the first likelihood is above a threshold, which triggers further processing of the utterance. The audio data is provided to the ASR model 202, and the text output from the ASR model and additional audio features and contextual features are provided to the false wakeup suppression classifier model 204. The false wakeup suppression classifier model 204 provides a second likelihood that the wakeup word "BIXBY" or the phrase "HEY, BIXBY" is present. Upon determining that the second likelihood is above the threshold, the processor 120 instructs the audio output device 208 to output "Call Mom". The processor 120 also causes a phone application or other communication application to start a communication session with a "Mom" contact stored on the electronic device 101 or otherwise associated with the user of the electronic device 101.

[0058] As another specific example, assume that the utterance "Hey BIXBY, start the timer" is received. The trained audio wakeup classifier model 201 detects a first possibility that the wakeup word "BIXBY" or the phrase "Hey, BIXBY" is present, and the processor 120 may determine that the first possibility is above a threshold, which triggers further processing of the utterance. The audio data is provided to the ASR model 202, and the text output from the ASR model as well as additional audio features and contextual features are provided to the false wakeup suppression classifier model 204. The false wakeup suppression classifier model 204 provides a second possibility that the wakeup word "BIXBY" or the phrase "Hey, BIXBY" is present. The processor 120 may determine that the second possibility is above a threshold, and the processor 120 may instruct execution of a timer application and display a timer on the display 210 of the electronic device 101.

[0059] In various embodiments, it will be understood that trained machine learning models such as the audio wakeup classifier model 201 and the false wakeup suppression classifier model 204 can operate to detect or predict whether a wakeup word or phrase is in an utterance. Based on this determination, the utterance may or may not be provided to another machine learning model (such as at least one NLU model) for further processing of the utterance in order to recognize the command given by the user. In addition, in various embodiments, based on the domain and / or command included in the utterance, the router can route the processed utterance to different models or sub-assistants associated with the voice assistant and associated with different domains or applications (such as one or more travel domains / applications, one or more music domains / applications, one or more phone domains / applications, etc.). In addition, in various embodiments, the audio wakeup classifier model 201 and the false wakeup suppression classifier model 204 act as gatekeepers to provide a lightweight solution for detecting whether a wakeup word or phrase is actually present in an utterance before submitting additional resources to be processed by the electronic device 101.

[0060] In addition, it will be understood that the system 200 delays final wake-up decision making until after the ASR model 202 processes at least a portion of the audio data, using the output from the ASR model 202 to perform wake-up verification using the false wake-up suppression classifier model 204. By delaying the wake-up decision making until post-ASR processing, the wake-up detection results are significantly improved, for example, by more than 90%. For example, as described in the present disclosure, it has been found that the use of deep learning methods improves wake-up detection by up to 96% or more, and it has been found that the use of random forest methods improves wake-up detection by up to 93% or more. In addition, in some embodiments, the system 200 or at least the models 201 to 204 can be deployed on a client electronic device so that the wake-up word or phrase detection and wake-up determination process can be performed on the device without transmitting any data over a public network such as the Internet, which can avoid sending user speech data to the server 106 or other external destinations. This can greatly alleviate user privacy issues because utterances that do not actually include a wake-up word or phrase and are therefore not intended to invoke a voice assistant provided by the user are not transmitted or stored outside the user's electronic device. However, it will be appreciated that one or more components of system 200 may be distributed to other devices depending on the particular implementation of system 200 or based on available computing resources.

[0061] although Figure 2A One example of a wake word detection and false wake word suppression system 200 is shown, but may be used for Figure 2AVarious changes may be made. For example, the audio input device 206, the audio output device 208, and the display 210 may be connected to the processor 120 within the electronic device 101, such as via a wired connection or circuit. In other embodiments, the audio input device 206, the audio output device 208, and the display 210 may be external to the electronic device 101 and connected via a wired or wireless connection. In addition, in some cases, one or more of the audio wake-up classifier model 201, the ASR model 202, the false wake-up suppression classifier model 204, and the other machine learning models 205 may be stored as separate models called by the processor 120 to perform certain tasks, or may be included in and form part of one or more larger machine learning models.

[0062] Furthermore, in some embodiments, one or more of the machine learning models (such as one or more of the ASR model 202, the false wakeup suppression classifier model 204, and the other machine learning models 205) may be stored remotely from the electronic device 101, such as on the server 106. As a specific example, Figure 2B An example system 207 is shown in which an ASR model 202, a false wakeup suppression classifier model 204, and other machine learning models 205 according to the present disclosure are stored on a server 106. Here, the electronic device 101 may send a request including input (such as captured audio data) to the server 106, such as by using the communication interface 170, and the server 106 includes its own processor 120 and communication interface 170 to process the input using the machine learning models 202, 204, 205. The results provided by the machine learning models 202, 204, 205 may be sent back to the electronic device 101. In some embodiments, the electronic device 101 may be replaced by a server 106 that receives audio input from a client device and sends instructions back to the client device to perform a function associated with the instructions included in the utterance.

[0063] Figure 3 An example wake-up verification process 300 according to the present disclosure is shown. For ease of explanation, the process 300 is described as involving the use of Figure 1 The process 300 may be used with any other suitable electronic device (such as, server 106) or combination of electronic devices (such as electronic device 101 and server 106), and may be used in any other suitable system.

[0064] like Figure 3As shown in , process 300 includes a first machine learning model (audio wake-up classifier model 201) receiving an audio signal 302, such as a signal received via an audio input device such as audio input device 206. The audio wake-up classifier model 201 is trained to determine a first likelihood or probability that the audio signal 302 is a user utterance that includes a particular wake-up word or phrase. If the processor 120 determines that the first likelihood is equal to or above a specified threshold, such as 0.5 or 50%, the audio signal 302 is provided to the ASR model 202. The audio wake-up classifier model 201 is thus used as a preliminary gatekeeper for the wake-up process of invoking a voice assistant. Therefore, if the first likelihood is below the threshold, the process 300 will end.

[0065] The ASR model 202 processes the audio signal 302 and outputs at least a text output for further processing. The text output is provided to a second machine learning model (false arousal suppression classifier model 204). Figure 3 As shown in , the false wakeup suppression classifier model 204 is included after the ASR model 202. The process 300 using the novel false wakeup suppression classifier model 204 is used to detect and suppress false wakeup events after ASR. The false wakeup suppression classifier model 204 is a classifier that detects whether a natural language (NL) string is from an unintentional wakeup or an intentional wakeup, and the natural language (NL) string may include or be a part of the ASR text output provided by the ASR model 202. Here, the process 300 uses the ASR text output to predict the probability of a false wakeup. The false wakeup suppression classifier model 204 can also receive audio features 304 and context features 306 as inputs to enhance the ability to detect false wakeups and suppress further actions. The false wakeup suppression classifier model 204 can use additional inputs as modifiers and / or constraints when performing classification to output a second possibility or probability that the audio signal 302 includes an utterance with a wakeup word or phrase.

[0066] As an example, Figure 4A 4 shows example audio and context input features 400 that may be used by the false wakeup suppression classifier model 204 to determine the likelihood that a wakeup word or phrase is included in the audio signal 302. Figure 4AAs shown in , features 400 include audio and context features that can be provided as input to the false wakeup suppression classifier model 204. Audio features (such as, audio features 304) can include natural language strings from the ASR model 202. Additional input features derived from the natural language strings can also be used as audio feature inputs to the false wakeup suppression classifier model 204, such as bag of words input, total word count input, total character count input, unique word count input, ratio of character count to audio time, number of stop words, and the like. The audio features 304 can also include features based on the audio signal 302, such as background noise level, audio time (such as in seconds or milliseconds), signal-to-noise ratio, pre-inverse text normalization (Pre-ITN) data, and the like. Features 400 can additionally or alternatively include context features (such as, context features 306), which are non-audio signal context information associated with or obtained from the client electronic device. As shown in FIG. Figure 4A As shown in , contextual features 306 may include the client launch method (how to activate the assistant, such as by button, voice, etc.), the current foreground application or package running on the electronic device, the user's gender, whether the device is operating in hands-free mode, etc.

[0067] As shown in the example audio and context feature input process 401 according to the present disclosure Figure 4B As shown in , the audio features 304 and the context features 306 can be provided as a set of inputs 402 to the false wakeup suppression classifier model 204. Including the false wakeup suppression classifier model 204 after the ASR model 202 allows the audio wakeup classifier model 201 to maintain its normal function of acting as an initial gatekeeper to detect whether the audio signal 302 may include utterances intended to initiate device wakeup. Using this set of inputs 402, the false wakeup suppression classifier model 204 can verify whether a wakeup event has actually occurred (the wakeup may be intentional) or whether the wakeup should be suppressed (the wakeup may be unintentional). The false wakeup suppression classifier model 204 can output a probability 404 indicating whether the wakeup event initially detected by the audio wakeup classifier model 201 is a false wakeup event, which can be compared to a threshold to suppress the wakeup event or proceed forward through further voice assistant processes. Using the false wakeup suppression classifier model 204 can enhance the user experience, avoid transmitting unintentional utterances to another electronic device, and / or reduce resource consumption, such as consumption of battery or computing resources.

[0068] As a specific example, in various embodiments, the false wakeup suppression classifier model 204 is trained using a labeled dataset of audio signals divided into a training set and a test set. The labeled dataset may be used to train a machine learning model (such as a random forest model, a multilayer perceptron model, and / or a deep learning transformer model (such as a bidirectional encoder representation from transformers (BERT))) of the false wakeup suppression classifier model 204 to predict false wakeups. During the inference process, the false wakeup suppression classifier model 204 may be deployed at an inference server to process input sent by a client device, or may be deployed at a client device (such as a phone, tablet, or computer) to improve user privacy and / or reduce response time. The false wakeup suppression classifier model 204 may output a probability score (such as a score between 0 and 1 (a higher score indicates a likelihood that the wakeup event is valid)) for comparison with a threshold to predict false wakeups.

[0069] The thresholds may also be adjusted for the a priori probabilities of false and correct wakeups to tune the system. For example, in some embodiments, the thresholds may be adjusted based on the performance of the audio wakeup classifier model 201. For example, if the audio wakeup classifier 201 has a low accuracy determined based on previous wakeup predictions, the thresholds of the false wakeup suppression classifier model 204 may be adjusted so that the false wakeup suppression classifier model 204 is able to catch false wakeups that the audio wakeup classifier 201 missed.

[0070] In some embodiments, the false wakeup suppression classifier model 204 may be trained for different types of audio background noise, and thresholds may be preset for specific environments. For example, the false wakeup suppression classifier model 204 may be trained to recognize that the user and the electronic device are currently in a vehicle based on road noise, wind noise, or other sounds present in the audio signal 302. In response, the threshold may be set lower (such as by reducing the 0.5 probability threshold to a 0.3 probability threshold) to make the output probability 404 more likely to exceed the threshold and trigger further voice assistant processing of the utterance, because the user in the vehicle may be more likely to want to activate the voice assistant to keep the user's hands available. Contextual features may also be used to reinforce this determination, such as if a global positioning system (GPS) application or other navigation application indicates that the device is moving at a rate or speed commensurate with the vehicle.

[0071] As another specific example, the false wakeup suppression classifier model 204 may be trained to recognize that the user and the electronic device are currently in a crowded area, such as an airport, a restaurant, etc., based on the amount of background chat or other sounds in the audio signal 302. Here, in response, the threshold may be set higher (such as by increasing the 0.5 probability threshold to the 0.7 probability threshold) so that the output probability 404 will be less likely to exceed the threshold and suppress further voice assistant processing of the utterance because the background noise may increase the likelihood of triggering a false wakeup detection based on the utterance in the background noise rather than the user's utterance. The overall background noise level may also play a similar role, such as when high background noise levels cause the threshold to increase and low background noise levels cause the threshold to decrease.

[0072] Refer again Figure 3 , based on the output threshold from the false wakeup suppression classifier model 204, the wakeup event can be suppressed and the process 300 ends at step 307. However, if the output probability exceeds the threshold, further voice assistant processing can be performed. For example, the router 308 can receive at least the output from the ASR model 202 and route the processed utterances to different models associated with the voice assistant and associated with different domains or applications (such as, one or more travel domains / applications, one or more music domains / applications, one or more phone domains / applications, etc.). Based on the routing, the utterance can be provided to at least one NLU model 310, which is trained to recognize and perform domain-specific actions. The NLU model 310 can further process the utterance to recognize the command given by the user and perform or trigger an action 312 based on the command, such as performing navigation to a specific place using a navigation application, playing a song selection using a music player application, starting a timer using a timer application, etc.

[0073] although Figure 3 An example of the wake-up verification process 300 is shown, but the Figure 3 For example, they may be combined, further subdivided, duplicated, rearranged, or omitted according to specific needs. Figure 3 In addition, if necessary or desired, one or more additional components and functions may be included. In addition, although shown as a series of steps, Figure 3The various steps in may overlap, occur in parallel, occur in a different order, or occur any number of times. In some embodiments, process 300 may be implemented by an electronic device such as a client device, or process 300 may also be performed using a distributed architecture. For example, the audio wakeup classifier model 201, the ASR model 202, and the false wakeup suppression classifier model 204 may be executed on a client electronic device (such as, electronic device 101) or at a server (such as, server 106). When deployed on a client electronic device, one or more of the models may be compressed using quantization, weight pruning, or other techniques. Although executing models 201 to 204 on a client electronic device may reduce privacy issues and may result in faster response times, in some embodiments, models 201 to 204 may be executed on a server or other remote electronic device based on available resources.

[0074] In various embodiments, the router 308 and the NLU model 310 may also be executed by the client electronic device or by the server. When executed by the server, the server may provide the client electronic device with a determined action 312 to be performed by the client electronic device. In addition, the action 312 may be performed by the client electronic device, or another electronic device (such as an external display, speaker, etc.) may be instructed to perform the action. In some embodiments, the wakeup classification performed by the audio wakeup classifier model 201 may be performed by the client electronic device, and the client electronic device may provide an audio signal 302 received via an audio input device of the client electronic device to the server. The ASR model 202, the false wakeup suppression classifier model 204, the router 308, and / or the NLU model 310 may be executed by the server based on audio data provided from the client electronic device.

[0075] Additionally, in some embodiments, the process 300 and its associated architecture can be implemented for any language. In some embodiments, the false wakeup suppression classifier model 204 can predict false wakeups before the ASR model 202 completes the full ASR. For example, the ASR model 202 can provide a real-time text data stream as it processes the audio signal 302, and the false wakeup suppression classifier model 204 can provide probabilities based on partial ASR output. Additionally, in some embodiments, continuous adaptation can be performed on the false wakeup suppression classifier model 204. For example, a validated heuristic can be used to continuously identify correct wakeup utterances and incorrect wakeup utterances after the system has been deployed. The training data can be augmented with samples from the new data and used to retrain the false wakeup suppression classifier model 204 using the augmented training data.

[0076] although Figure 4A and Figure 4BAn example audio and context input feature 400 and an example audio and context feature input process 401 are shown, respectively, but may be used for Figure 4A and Figure 4B For example, they may be combined, further subdivided, duplicated, rearranged, or omitted according to specific needs. Figure 4A and Figure 4B In addition, if necessary or desired, one or more additional components and functions may be included. As a specific example, the system may be configured to be configured to be a user-friendly interface, such as based on a predetermined setting or dynamically based on current conditions. Figure 4A and Figure 4B All of the audio and context features shown in can be provided to the false wakeup suppression classifier model 204, or a subset of the audio and context features can be provided. For example, continuous learning can show that certain audio and context features or combinations of certain features provide more accurate results, such as if it is found that features such as signal-to-noise ratio, total number of characters, audio time, client startup method, and number of characters and audio time features provide more optimized results or probability predictions.

[0077] Figure 5 An example false wakeup suppression classifier model architecture 500 according to the present disclosure is shown. For ease of explanation, Figure 5 The architecture 500 shown in FIG. 5 is described as Figure 1 The network configuration 100 is implemented on or supported by the electronic device 101. However, Figure 5 The architecture 500 shown in may be used with any other suitable apparatus and may be used in any other suitable system, such as when the architecture 500 is implemented on or supported by the server 106 .

[0078] like Figure 5 As shown in , the architecture 500 includes a deep transformer language model 502 such as a BERT model, which receives a text string provided by an ASR model and corresponding to an input audio signal as input. The deep transformer language model 502 is trained to help use the surrounding text to establish context to provide the meaning of the language in the text. The deep transformer language model can connect each output element to each input element, and can dynamically calculate the weights between them based on their connection. In an embodiment where the deep transformer language model 502 is a BERT model, the input text can be processed bidirectionally, meaning that it can be read in two directions at a time and used to perform natural language understanding using text input. In this example, the deep transformer language model 502 provides an output to a multilayer perceptron 508. The multilayer perceptron 508 is used to ultimately provide a probability of whether the speech should trigger the wake-up process or whether the wake-up process should be suppressed.

[0079] In addition to the output received from the deep transformer language model 502, the multilayer perceptron 508 can receive both audio features and contextual features. The audio, text, and derived feature acquisition operation 504 collects a plurality of audio features based on or derived from the audio signal and the ASR text, such as audio time, signal-to-noise ratio, background noise level, and contextual features. Figure 4A and Figure 4B In various embodiments, the audio, text, and derived feature acquisition operation 504 may encode these audio features and provide them as additional inputs to the multilayer perceptron 508. In addition, the acquired audio features (such as specific background noise detected in the audio signal) may be used to adjust the probability threshold for arousal suppression.

[0080] Additionally or alternatively, the context feature acquisition operation 506 may collect a plurality of context information, such as user gender, foreground application identification, client launch method, hands-free settings, or other information about the user's current state. Figure 4A and 4B Those other contextual information described. In various embodiments, the context feature acquisition operation 506 can encode the context features and provide them as additional inputs to the multilayer perceptron 508. It will be understood that the multilayer perceptron 508 includes multiple layers, including an input layer, an intermediate layer or a weighted layer, and an output layer that provides a final probability as to whether the wakeup is valid. Based on the output probability, the threshold determination operation 510 determines whether the probability output by the multilayer perceptron 508 exceeds a threshold. If not, the wakeup process is suppressed. If so, the wakeup process is allowed to continue so that the audio input can be further processed for domain / intent detection in order to trigger one or more actions of one or more electronic devices.

[0081] although Figure 5 An example of a false wakeup suppression classifier model architecture 500 is shown, but may be used for Figure 5 For example, they may be combined, further subdivided, duplicated, rearranged, or omitted according to specific needs. Figure 5 In addition, one or more additional components and functions may be included if needed or desired. In addition, it will be understood that various types of deep transformer models may be used and various other types of classification models may be used without departing from the scope of the present disclosure.

[0082] Figure 6 Another example of a false wakeup suppression classifier model architecture 600 according to the present disclosure is shown. For ease of explanation, Figure 6 The architecture 600 shown in FIG. 6 is described as Figure 1 The network configuration 100 is implemented on or supported by the electronic device 101. However, Figure 6 The architecture 600 shown in may be used with any other suitable apparatus and may be used in any other suitable system, such as when the architecture 600 is implemented on or supported by the server 106 .

[0083] like Figure 6 As shown in , the architecture 600 includes a deep transformer language model 602, such as a BERT model, which receives as input a text string provided by an ASR model and corresponding to an input audio signal. The deep transformer language model 602 is trained to help establish context using surrounding text to provide meaning for the language in the text. In an embodiment where the deep transformer language model 602 is a BERT model, the input text can be processed bidirectionally. In this example, the deep transformer language model 602 provides an output to a random forest classifier 608. The random forest classifier 608 is used to ultimately provide a probability of whether the utterance should trigger the wake-up process or whether the wake-up process should be suppressed.

[0084] In addition to the output received from the deep transformer language model 602, the random forest classifier 608 can receive both audio features and contextual features. The audio, text, and derived feature acquisition operation 604 collects a plurality of audio features based on or derived from the audio signal and the ASR text, such as audio time, signal-to-noise ratio, background noise level, and contextual features such as the audio signal. Figure 4A and Figure 4B In various embodiments, the audio, text, and derived feature acquisition operation 604 may encode these audio features and provide them as additional inputs to the random forest classifier 608. In addition, the acquired audio features (such as specific background noise detected in the audio signal) may also be used to adjust the probability threshold of arousal suppression.

[0085] Additionally or alternatively, the context feature acquisition operation 606 may collect a plurality of context information, such as user gender, foreground application identification, client launch method, hands-free settings, or other information about the user's current state. Figure 4A and 4BThose other contextual information described. In various embodiments, the context feature acquisition operation 606 may encode context features and provide them as additional inputs to the random forest classifier 608. It will be understood that the random forest classifier 608 may include multiple decision trees that are used to provide multiple probability outputs that are combined, such as via averaging, to provide a final probability as to whether the wake-up is valid. Based on the output probability, the threshold determination operation 610 determines whether the probability output by the random forest classifier 608 exceeds a threshold. If not, the wake-up process is suppressed. If yes, the wake-up process is allowed to continue so that the audio input can be further processed for domain / intent detection to trigger one or more actions of one or more electronic devices.

[0086] although Figure 6 An example of a false wakeup suppression classifier model architecture 600 is shown, but may be used for Figure 6 For example, they may be combined, further subdivided, duplicated, rearranged, or omitted according to specific needs. Figure 6 In addition, if needed or desired, one or more additional components and functions may be included. In addition, it will be understood that various types of deep transformer models may be used and various other types of classification models may be used without departing from the scope of the present disclosure.

[0087] Figure 7 An example process 700 for training a false wakeup suppression classifier model according to the present disclosure is shown. For ease of explanation, Figure 7 The process 700 shown in FIG. 1 is described as using Figure 1 The process 700 is performed by the server 106 in the network configuration 100. However, the process 700 can be used with any other suitable device (such as the electronic device 101) and can be used in any other suitable system.

[0088] like Figure 7As shown in , during process 700, a labeled dataset of audio signals (which in some cases may be manually labeled) is generated and divided into a set of training audio signals 701 and a set of test audio signals 703. The training audio signals 701 are used as input to provide ASR text output using a trained ASR model 702. The ASR text generated from the training audio signals 701 is provided to a false wakeup suppression classifier model 704 to be trained under process 700. In addition, features derived from the basic features of the training audio signals 701 are created and labeled, and these labeled audio and contextual features 705 are used as additional input to the false wakeup suppression classifier model 704. Some of the labeled audio and contextual features 705 derived from the basic features may include total unique words, total stop words, character to audio time ratio, or such as with respect to Figure 4A and 4B Using the input ASR text and the input labeled audio and context features 705, the false arousal suppression classifier model 704 outputs a probability score 706 for predicting and suppressing false arousals.

[0089] Using the output probability score 706 from the false awakening suppression classifier model 704, the loss function 708 determines an error or loss based on the reference true value 707, and modifies or updates the false awakening suppression classifier model 704 based on the error or loss. For example, when the output of the false awakening suppression classifier model 704 is different from the reference true value 707, the difference can be used to calculate the loss as defined by the loss function 708. The loss function 708 can use any suitable measure of the loss associated with the output generated by the false awakening suppression classifier model 704, such as a cross entropy loss or a mean squared error. Based on the calculated loss, the parameters of the false awakening suppression classifier model 704 can be adjusted.

[0090] At decision block 710, it is determined whether the initial training of the false awakening suppression classifier model 704 is complete, such as by determining whether the false awakening suppression classifier model 704 is predicting false awakenings with an acceptable level of accuracy using the input training data. If not, the process 700 loops back to provide the same or additional training audio signal 701 and the same or additional labeled audio and context features 705 to continue training the false awakening suppression classifier model 704. The process 700 may loop any number of times here to obtain additional outputs from the false awakening suppression classifier model 704 to be compared with the reference true value, so that the additional loss may be determined using the loss function 708. Over time, the false awakening suppression classifier model 704 generates more accurate outputs that more closely match the reference true value, and the loss of the metric becomes smaller. The amount of training data used may vary depending on the number of training cycles and may include a large amount of training data. At some point, the loss of the metric may drop below a specified threshold, and the desired accuracy may be determined at decision block 710 to be achieved, indicating that the initial training of the false awakening suppression classifier model 704 is complete. Therefore, a trained false awakening suppression classifier model 712 is obtained.

[0091] As described in the present disclosure, the trained false arousal suppression classifier model 712 can be of various architectures, such as architectures utilizing random forests, deep learning transformers, and / or multi-layer perceptrons. It has been found that various architectures can provide highly accurate arousal detection. For example, it has been found that an architecture using random forests can provide approximately 93% accuracy or better, and an architecture using a multi-layer perceptron can provide approximately 96% accuracy or better.

[0092] In some embodiments, the trained false arousal suppression classifier model 712 may be deployed at this point in process 700. In some embodiments, such as in Figure 7 In the embodiment shown in , a test audio signal 703 can be used to test the trained false awakening suppression classifier model 712. The test audio signal 703 is provided to the trained ASR model 702, and the ASR text and audio and context features 711 are input to the trained false awakening suppression classifier model 712. At decision block 714, it is determined whether the trained false awakening suppression classifier model 712 performs as expected based on the output of the trained false awakening suppression classifier model 712. It will be understood that multiple tests can be performed and the overall performance of the multiple tests can be evaluated at decision block 714.

[0093] Use as Figure 7The test data set shown in can help evaluate the machine learning model and confirm that the model is working as expected or whether problems have occurred during training, such as whether the model overfits the training data. If the performance of the model is not as expected at decision block 714, process 700 loops back to the beginning of process 700 to retrain the model using the test data (or continue to train the model) until the desired performance is achieved. Once it is determined at decision block 714 that the trained false wakeup suppression classifier model 712 is performing as expected, process 700 ends and the trained false wakeup suppression classifier model 712 can be deployed. As a specific example, process 700 can be performed at Figure 1 The method is executed on the server 106 in the network configuration 100, and the trained false wakeup suppression classifier model 712 can be deployed to the client electronic device 101 for use.

[0094] although Figure 7 One example of a process 700 for training a false arousal suppression classifier model is shown, but may be used for Figure 7 For example, although shown as a series of steps, Figure 7 The various steps in may overlap, occur in parallel, occur in a different order, or occur any number of times. In addition, in some embodiments, the training audio signal 701 and the test audio signal 703 may not be provided to the ASR model 702. For example, instead of relying on the ASR model 702, labeled audio features may be derived from the training audio signal, and ASR text may be manually created from the training audio samples and provided to the false wakeup suppression classifier model 704 for training.

[0095] Fig. 8A and 8B An example method 800 for false wakeup suppression according to the present disclosure is shown. For ease of explanation, the method 800 shown in FIG. 8 is described as being composed of Figure 1 The method 800 shown in FIG. 8 is performed by the electronic device 101 in the network configuration 100. However, the method 800 shown in FIG. 8 may be used with any other suitable device (such as the server 106) and may be used in any other suitable system.

[0096] As shown in FIG8 , in step 802, a voice signal is obtained from an audio input device. This may include, for example, the processor 120 of the electronic device 101 obtaining audio data from the audio input device 206 and storing the obtained audio data at least temporarily in a memory. In step 804, a machine learning model is used to predict a first possibility that the voice signal includes a wake-up indication (such as a wake-up word or other command that triggers the voice assistant wake-up process). This may include, for example, the processor 120 of the electronic device 101 using the audio wake-up classifier model 201 to obtain the probability that the voice signal includes a wake-up indication or command from the user. In step 806, it is determined whether the first possibility exceeds a first threshold. This may include, for example, the processor 120 of the electronic device 101 comparing the probability output by the audio wake-up classifier model 201 with a specified probability threshold and determining whether the probability exceeds the specified threshold. If the first possibility does not exceed the first threshold, the method 800 ends in step 807. This may include, for example, the processor 120 of the electronic device 101 determining that the first possibility does not exceed the first threshold and stopping the wake-up process because the processor 120 has determined that the desired wake-up command is unlikely to exist in the voice signal.

[0097] If the first likelihood exceeds the first threshold, automatic speech recognition is performed on the speech signal at step 808 to determine and output a text representation of the speech signal. This can include, for example, the processor 120 of the electronic device 101 inputting the speech signal into the ASR model 202 and receiving an ASR text output from the ASR model 202. At step 810, at least one of the text representation, audio features associated with the speech signal, and context features is provided to a second machine learning model. This can include, for example, the processor 120 of the electronic device 101 using the audio, text, and derived feature acquisition operation 604 to acquire the ASR text output and audio features, using the context feature acquisition operation 606 to acquire context features, and inputting the ASR text, audio features, and / or context features into the false wakeup suppression classifier model 204. In some embodiments, such as if the ASR model 202 is configured to output a text stream, the text representation can be a partial ASR output. Audio features can include word bags, total word count, total character count, unique word count, stop word count, audio time, signal-to-noise ratio, and information such as about Figure 4A and Figure 4B The contextual features may include at least one user feature, a launch method, foreground application information, and one or more other contextual features, such as information about the audio source. Figure 4A and 4B Those described.

[0098] In step 812, determine whether to adjust the second threshold. This may include, for example, the processor 120 of the electronic device 101 determining whether to adjust the second threshold based on one or more parameters such as the accuracy of the first machine learning model and / or based on one or more of the audio or context features. If the second threshold is not adjusted, the method 800 moves to step 814. If the second threshold is to be adjusted, the method 800 moves to step 813. In step 813, the second threshold is adjusted based on at least one parameter. For example, the processor 120 may determine the background noise environment associated with the speech signal and set the second threshold based on the determined background noise environment. As another example, the processor 120 may determine the accuracy of the first machine learning model used to predict the first possibility based on multiple previous speech signals, and adjust the second threshold lower or higher based on the determined accuracy. Then, the method 800 moves to step 814.

[0099] In step 814, a second possibility that the voice signal includes a wake-up instruction is predicted using a second machine learning model. This may include, for example, the processor 120 of the electronic device 101 using the false wake-up suppression classifier model 204 to obtain a second probability that the voice signal includes a wake-up instruction or command from the user. Here, the false wake-up suppression classifier model 204 may use a text representation, an audio feature associated with the voice signal, and a context feature associated with the electronic device. For example, in order to obtain a second possibility, a deep transformer model included in the second machine learning model may be used to perform natural language processing on the text representation, and one or more results may be output. One or more results of the natural language processing may be provided to a multilayer perceptron model of the second machine learning model together with the audio features and the context features, and an output related to the second possibility may be received from the multilayer perceptron model, the output indicating whether to suppress the wake-up operation of the voice assistant. As another example, in order to obtain a second possibility, a deep transformer model included in the second machine learning model may be used to perform natural language processing on the text representation, and one or more results may be output. One or more results of the natural language processing may be provided to a random forest classifier model of a second machine learning model along with the audio features and the context features, and an output related to a second possibility may be received from the random forest classifier model, indicating whether to suppress the wake-up operation of the voice assistant.

[0100] At step 816, it is determined whether the second likelihood exceeds a second threshold. This may include, for example, the processor 120 of the electronic device 101 comparing the probability output by the false wakeup suppression classifier model 204 with a specified threshold. If the probability is below the specified threshold, the wakeup process is suppressed at step 817, and the method 800 ends at step 807. If the probability exceeds the specified threshold, the wakeup process may continue by moving to step 818.

[0101] In step 818, the instruction is routed to a natural language understanding model associated with the determined domain based on the text representation. This may include, for example, the processor 120 of the electronic device 101 using the router 308 to determine which domains and associated applications will be used for one or more commands in the utterance. In step 820, the natural language understanding model to which the process is routed at step 818 is used to determine one or more actions to be performed. This may include, for example, the processor 120 of the electronic device 101 using the NLU model 310 associated with the domain, sub-assistant and / or application to identify one or more actions to be performed by the electronic device 101 or another electronic device. In step 822, an instruction for performing one or more determined actions requested in the voice signal is generated. This may include, for example, the processor 120 of the electronic device 101 creating a command for one or more applications to execute in order to satisfy a request in the user's utterance, such as starting a timer, playing a song, performing navigation to a location, etc.

[0102] At step 824, it is determined whether to perform continuous adjustment. This may include, for example, the processor 120 of the electronic device 101 determining whether the false wakeup suppression classifier model 204 has associated settings or parameters for performing continuous learning operations. If not, the method 800 ends at step 830. Otherwise, at step 826, the accuracy of the second machine learning model is determined. This may include, for example, the processor 120 of the electronic device 101 continuously or repeatedly identifying correct wakeup and false wakeup utterances using a validated heuristic after the false wakeup suppression classifier model 204 has been deployed to improve the false wakeup suppression classifier model 204. In order to continuously improve the second machine learning model, the training data is enhanced, and the second machine learning model is retrained using the enhanced training data. This may include, for example, the processor 120 of the electronic device 101 enhancing the training data, retraining the false wakeup suppression classifier model 204 using the enhanced training data, and providing the updated false wakeup suppression classifier model 204 for further use. The method 800 ends at step 830.

[0103] Although Figure 8 shows one example of a method 800 for false wakeup suppression, various changes may be made to Figure 8. For example, although shown as a series of steps, the various steps in Figure 8 may overlap, occur in parallel, occur in a different order, or occur any number of times.

[0104] Fig. 9 An example method 800 for false wakeup suppression according to the present disclosure is shown. For ease of explanation, Fig. 9 The method 900 shown in FIG. 1 is described as Figure 1 The network configuration 100 is executed in the electronic device 101. However, Fig. 9The method 900 shown in may be used with any other suitable apparatus, such as the server 106, and may be used in any other suitable system.

[0105] Reference Fig. 9 , in step 910 (similar to step 802), the electronic device 101 obtains a speech signal. In step 920 (operating similarly to step 804), the electronic device 101 uses a first machine learning model trained to receive a speech signal as input to predict a first likelihood that a wake-up word or phrase is spoken in the speech signal. In step 930 (similar to steps 806 and 808), in response to the first likelihood exceeding a first threshold, the electronic device 101 determines a text representation of the speech signal based on performing automatic speech recognition (ASR) on the speech signal. In step 940 (similar to steps 810 and 814), the electronic device 101 uses a second machine learning model to predict a second likelihood that a wake-up word or phrase is spoken in the speech signal, the second machine learning model being trained to receive at least one of the following: a text representation of the speech signal, an audio feature associated with the speech signal, and a context feature associated with the electronic device 101. In step 950 (similar to steps 816 , 822 ), in response to the second likelihood exceeding the second threshold, the electronic device 101 generates instructions for performing the action requested in the speech signal.

[0106] In an embodiment, similar to step 817, when the second possibility does not exceed the second threshold, the electronic device 101 suppresses the wake-up operation of the voice assistant.

[0107] although Fig. 9 An example of a method 900 for false wakeup suppression is shown, but may be used for Fig. 9 For example, although shown as a series of steps, Fig. 9 The various steps in may overlap, occur in parallel, occur in a different order, or occur any number of times.

[0108] Although the present disclosure has been described with exemplary embodiments, various changes and modifications may be suggested to one skilled in the art. The present disclosure is intended to encompass such changes and modifications as fall within the scope of the appended claims.

Claims

1. A method (900) performed by an electronic device, the method (900) comprising: Obtaining (910) a speech signal; predicting (920) a first likelihood that a wake word or phrase is spoken in the speech signal using a first machine learning model trained to receive the speech signal as input; In response to the first likelihood exceeding a first threshold, determining (930) a textual representation of the speech signal based on performing automatic speech recognition (ASR) on the speech signal; predicting (940) a second likelihood that the wake word or phrase was spoken in the speech signal using a second machine learning model trained to receive at least one of: the textual representation of the speech signal, audio features associated with the speech signal, and contextual features associated with the electronic device; as well as In response to the second likelihood exceeding the second threshold, instructions are generated (950) for performing an action requested in the speech signal.

2. The method according to claim 1, further comprising: If the second possibility does not exceed the second threshold, the wake-up operation of the voice assistant is suppressed.

3. The method according to any one of claims 1 to 2, wherein: The steps of predicting a second likelihood using a second machine learning model include: performing natural language processing on the text representation using a deep transformer model included in a second machine learning model and outputting one or more results; providing the one or more results of the natural language processing, the audio features, and the context features to a multilayer perceptron model or a random forest classifier model included in a second machine learning model; and A second possibility is determined using the multilayer perceptron model or the random forest classifier model, wherein the second possibility indicates whether to suppress a wake-up operation of the voice assistant.

4. The method according to any one of claims 1 to 3, further comprising: determining a background noise environment associated with the speech signal; as well as The second threshold is set based on the determined background noise environment.

5. The method according to any one of claims 1 to 4, wherein: The audio features include at least one of: bag of words, total number of words, total number of characters, number of unique words, number of stop words, audio time, background noise level, and signal-to-noise ratio.

6. The method according to any one of claims 1 to 5, wherein: The contextual characteristics include at least one of the following items: at least one user characteristic, a startup method, a hands-free setting, and foreground application information.

7. The method according to any one of claims 1 to 6, further comprising: determining an accuracy of a first machine learning model for predicting a first likelihood based on a plurality of prior speech signals; as well as The second threshold is adjusted based on the determined accuracy.

8. An electronic device (101), comprising: At least one processing device (120) is configured to: Obtaining (910) a speech signal; predicting (920) a first likelihood that a wake word or phrase is spoken in the speech signal using a first machine learning model trained to receive the speech signal as input; In response to the first likelihood exceeding a first threshold, determining (930) a textual representation of the speech signal based on performing automatic speech recognition (ARS) on the speech signal; predicting (940) a second likelihood that the wake word or phrase was spoken in the speech signal using a second machine learning model trained to receive at least one of: the textual representation of the speech signal, audio features associated with the speech signal, and contextual features associated with the electronic device; as well as In response to the second likelihood exceeding the second threshold, instructions are generated (950) for performing an action requested in the speech signal.

9. The electronic device according to claim 8, wherein the at least one processing device is further configured to: If the second possibility does not exceed the second threshold, the wake-up operation of the voice assistant is suppressed.

10. The electronic device according to any one of claims 8 to 9, wherein: To predict a second likelihood using the second machine learning model, the at least one processing device is further configured to: performing natural language processing on the text representation using a deep transformer model included in a second machine learning model, and outputting one or more results; providing the one or more results of the natural language processing, the audio features, and the context features to a multilayer perceptron model or a random forest classifier model included in a second machine learning model; and A second possibility is determined using the multilayer perceptron model or the random forest classifier model, wherein the second possibility indicates whether to suppress a wake-up operation of the voice assistant.

11. The electronic device according to any one of claims 8 to 10, wherein: The at least one processing device is further configured to: determining a background noise environment associated with the speech signal; and The second threshold is set based on the determined background noise environment.

12. The electronic device according to any one of claims 8 to 11, wherein: The audio features include at least one of: bag of words, total number of words, total number of characters, number of unique words, number of stop words, audio time, background noise level, and signal-to-noise ratio.

13. The electronic device according to any one of claims 8 to 12, wherein: The contextual characteristics include at least one of the following items: at least one user characteristic, a startup method, a hands-free setting, and foreground application information.

14. The electronic device according to any one of claims 8 to 13, wherein: The at least one processing device is further configured to: determining an accuracy of a first machine learning model for predicting a first likelihood based on a plurality of prior speech signals; and The second threshold is adjusted based on the determined accuracy.

15. A computer-readable medium comprising instructions, which, when executed, cause at least one processor of an electronic device to perform operations corresponding to the method of any one of claims 1-7.